[{"content":"I arrived at the Guangdong-Hong Kong-Macao Greater Bay Area Summer Camp in July 2026 as a student from @Guangdong University of Technology , curious about two things: where ideas go after they leave the classroom, and what Hong Kong feels like beyond the photographs. Over the next few days, those questions kept changing shape. They appeared in student ventures, working products, local farms, a robot we assembled, and factory lines moving at full scale. The strongest thread, though, was the people: students with very different backgrounds becoming teammates, then friends.\nPeople Before Programmes The first evening had the familiar awkwardness of any new group: careful introductions, half-remembered names, and conversations that needed a little time to warm up. Around the welcome-dinner table, that formality gradually disappeared. Before there was a team, there was simply a group of people learning how to talk to one another.\nMy teammates came from @Lingnan University , @Macau University of Science and Technology , @The Education University of Hong Kong , @Hong Kong Metropolitan University , @Zhejiang University , @Shanghai Jiao Tong University , and @The Hong Kong Polytechnic University . We came from different disciplines and were at different stages of study. The mix became most visible when one question produced several completely different lines of thought—and gave us more to talk about than any standard introduction could.\nWelcome dinner in Hong Kong At PolyU, I was drawn less to the language of entrepreneurship than to the support behind it. Students are given time, guidance, and room to keep developing an early idea. Project-building did not feel like a side experiment there. If I return, I hope it will be as a participant rather than a visitor.\nVisit to The Hong Kong Polytechnic University At @Cyberport , \u0026quot;incubation\u0026quot; stopped sounding like a word from a slide deck. With actual products in front of us, the questions became sharper: who would use this, what problem was it solving, and what would it take to work beyond a demonstration? The ideas were no longer floating on their own; they had users and constraints waiting for them outside the room.\nHong Kong in Motion Street scene in Central, Hong Kong Between scheduled visits, Hong Kong refused to become mere background. Central compressed towers, slopes, footbridges, older buildings, and narrow streets into the same field of view. A few minutes on foot could completely change the character of the street. The city was dense, but not unreadable; different periods and everyday rhythms remained visible beside one another.\nSome of my clearest memories came through a window rather than during a formal visit. At sunset, light appeared and disappeared between buildings as we moved through the city. After a day organised around destinations, it felt good to stop interpreting everything and simply watch.\nHong Kong at sunset Then came blue hour: distant lights switching on one by one, people continuing through their routines, and the surroundings growing quieter without ever becoming still. The photograph captured the view, but the transition was the part I wanted to keep.\nHong Kong during blue hour Victoria Harbour opened all that density into distance. I knew the skyline long before I stood in front of it, yet no single landmark mattered as much as the lights, water, and sea breeze together. It was the familiar image of Hong Kong, but finally attached to a real evening of my own.\nVictoria Harbour at night Switching Scales The programme kept changing lenses. At @Good Old Soil (好老土) , a Hong Kong community organisation supporting local farmers, the work centred on land, production, distribution, and livelihood. There was no flashy prototype in the fields. The work began with paying attention to a place, its farmers, and what local agriculture needed in order to keep moving.\nField visit with a local farming social organisation The embedded-AI workshop changed the pace completely. Our team assembled and debugged a robot, and the feedback was immediate: software, hardware, and physical behaviour either worked together or they did not. We divided the work, worked through each problem, and eventually watched a collection of parts become a functioning system. That was easily one of the most satisfying moments of the camp.\nThen the scale jumped again. At @Haid Group , technology ran through modern agriculture and animal husbandry—from feed and breeding to animal healthcare, intelligent farming, and food processing. The visit revealed the technical depth inside an industry that rarely presents itself with the visual language of a technology company.\nAt @Yikang Medical (广州一康) , our hands-on exposure to medical-equipment processing shifted the emphasis again. Here, engineering was about precision, standardisation, and the repeatable steps that make equipment dependable in real clinical use.\nThe production line at @GAC Aion moved to yet another rhythm. Separate processes connected continuously as components became complete vehicles. The showroom presented the polished result; the factory floor revealed the automation, timing, and coordination between people and machines behind it.\nBack at GDUT Group photograph at Guangdong University of Technology By the time we returned to Guangdong University of Technology, the university names and academic backgrounds that had filled our first introductions no longer felt like labels. We had spent days moving, building, listening, asking questions, and getting used to one another. The photograph outside the Engineering Training Center is informal, and that is exactly why I like it. It feels much closer to the people the programme had brought together.\nOur Turn on Stage After days of listening to universities and companies explain their work, the final roadshow put us on the other side of the room. Our team presented Huisheng Zhiyuan, a project focused on the intelligent production of organoids. It was selected as the sole Outstanding Project among all participating projects.\nI delivered the presentation on behalf of the team. The person speaking onstage is easy to see; the collective work behind that moment is not. The recognition belonged to the project and the team. My part was to communicate our work as clearly as I could.\nPresenting the Huisheng Zhiyuan project at the summer camp\u0026#39;s final roadshow An Honour—and a Beginning Being named an Outstanding Camper was a genuine honour, but I do not want that title to become the whole story. The photograph with the president of Guangdong University of Technology records the formal moment. What gave it meaning was everything around it: the guidance of teachers and organisers, the work of my teammates, and the curiosity and encouragement shared throughout the programme.\nPhotograph with the president of Guangdong University of Technology after receiving the Outstanding Camper honour Looking through the photographs now, words such as \u0026quot;innovation\u0026quot; and \u0026quot;entrepreneurship\u0026quot; no longer arrive as abstractions. They bring back a classroom at PolyU, products at Cyberport, fields in Hong Kong, a robot taking shape, factory floors, and the few concentrated minutes of a final presentation.\nMore than any single visit, I am grateful for the people who were there with me. To every friend I met from different cities, universities, disciplines, and stages of study: thank you for the conversations, laughter, encouragement, and memories we now share. The summer camp reached its final day, but our story did not end there. In many ways, it is only just beginning.\nFriends and memories from the Greater Bay Area Summer Camp ","permalink":"/blog/greater-bay-area-summer-camp/","summary":"\u003cp\u003eI arrived at the Guangdong-Hong Kong-Macao Greater Bay Area Summer Camp in July 2026 as a student from \u003ca href=\"https://english.gdut.edu.cn/\" class=\"mention-link\"\u003e@Guangdong University of Technology\u003c/a\u003e\n, curious about two things: where ideas go after they leave the classroom, and what Hong Kong feels like beyond the photographs. Over the next few days, those questions kept changing shape. They appeared in student ventures, working products, local farms, a robot we assembled, and factory lines moving at full scale. The strongest thread, though, was the people: students with very different backgrounds becoming teammates, then friends.\u003c/p\u003e","title":"Greater Bay Area Summer Camp: Across Cities and Industries—and Honoured as an Outstanding Camper"},{"content":"On July 9, 2026, our three-person team won the Gold Award at the Guangzhou TeamOne AI Technology Challenge, receiving a RMB 50,000 prize for Marketing Hub.\nMarketing Hub is an AI-powered marketing workspace designed to turn an initial idea into a structured content production process. The platform covers idea development, brand context management, content generation, visual workflow orchestration, project assets, and AI usage governance.\nThe product starts from a lightweight idea-entry experience: a user can write down a rough product, audience, or channel direction, then let the system expand it into a more structured marketing workflow.\nMarketing Hub idea entry screen The workflow canvas connects brand context, copywriting, visual prompt generation, compliance review, and image asset generation into a visible production chain, with task status and generated assets kept in the same workspace.\nMarketing Hub workflow canvas In this project, I was mainly responsible for core code development, overall architecture design, and top-level system design. My work focused on connecting product requirements with an implementable engineering structure: defining the main modules, organizing the workflow logic, and making the system coherent enough to support a complete demonstration and evaluation process.\nThis award is an important milestone for the project, and it also gave me a concrete reminder that engineering work becomes more valuable when it can be presented, tested, and recognized in a real setting.\n","permalink":"/blog/marketing-hub-teamone-gold/","summary":"\u003cp\u003eOn July 9, 2026, our three-person team won the \u003cstrong\u003eGold Award\u003c/strong\u003e at the \u003cstrong\u003eGuangzhou TeamOne AI Technology Challenge\u003c/strong\u003e, receiving a \u003cstrong\u003eRMB 50,000 prize\u003c/strong\u003e for \u003cstrong\u003eMarketing Hub\u003c/strong\u003e.\u003c/p\u003e\n\u003cp\u003eMarketing Hub is an AI-powered marketing workspace designed to turn an initial idea into a structured content production process. The platform covers idea development, brand context management, content generation, visual workflow orchestration, project assets, and AI usage governance.\u003c/p\u003e\n\u003cp\u003eThe product starts from a lightweight idea-entry experience: a user can write down a rough product, audience, or channel direction, then let the system expand it into a more structured marketing workflow.\u003c/p\u003e","title":"Marketing Hub Won Gold at the Guangzhou TeamOne AI Technology Challenge"},{"content":"On July 8, 2026, I completed the on-site delivery of the Rocket 3D Reconstruction System at the Xi'an Institute of Optics and Precision Mechanics, Chinese Academy of Sciences (XIOPM).\nThis project marked a complete engineering cycle from system design to on-site handover. My responsibilities included core code development, overall architecture design, top-level design, system integration, and final delivery.\nThe delivery work required the system to move beyond a development environment and operate as a finished project in a real institutional setting. It involved not only implementation, but also coordination, deployment preparation, issue handling, and final acceptance-oriented communication.\nFor me, this was my first truly landed project, and it made the meaning of “delivery” much more concrete.\n","permalink":"/blog/xioptics-rocket-3d-reconstruction-delivery/","summary":"\u003cp\u003eOn July 8, 2026, I completed the on-site delivery of the \u003cstrong\u003eRocket 3D Reconstruction System\u003c/strong\u003e at the \u003cstrong\u003eXi'an Institute of Optics and Precision Mechanics, Chinese Academy of Sciences (XIOPM)\u003c/strong\u003e.\u003c/p\u003e\n\u003cp\u003eThis project marked a complete engineering cycle from system design to on-site handover. My responsibilities included \u003cstrong\u003ecore code development\u003c/strong\u003e, \u003cstrong\u003eoverall architecture design\u003c/strong\u003e, \u003cstrong\u003etop-level design\u003c/strong\u003e, \u003cstrong\u003esystem integration\u003c/strong\u003e, and \u003cstrong\u003efinal delivery\u003c/strong\u003e.\u003c/p\u003e\n\u003cp\u003eThe delivery work required the system to move beyond a development environment and operate as a finished project in a real institutional setting. It involved not only implementation, but also coordination, deployment preparation, issue handling, and final acceptance-oriented communication.\u003c/p\u003e","title":"My First Delivered Project: Rocket 3D Reconstruction System at XIOPM"},{"content":"Our Multi-Agent System helps users navigate digital government tasks. It leverages collaborative agents to automate workflows, ensuring accessible and secure service for everyone.\nKey Features:\nAutomated workflow processing through Agent collaboration. Built to enhance accessibility for digital government services. Secure environment for public data handling. View the source code and documentation on GitHub using the Code button above.\n","permalink":"/projects/intelligent-government-assistant/","summary":"\u003cp\u003eOur Multi-Agent System helps users navigate digital government tasks. It leverages collaborative agents to automate workflows, ensuring accessible and secure service for everyone.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eKey Features:\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003eAutomated workflow processing through Agent collaboration.\u003c/li\u003e\n\u003cli\u003eBuilt to enhance accessibility for digital government services.\u003c/li\u003e\n\u003cli\u003eSecure environment for public data handling.\u003c/li\u003e\n\u003c/ul\u003e\n\u003cp\u003e\u003cem\u003eView the source code and documentation on GitHub using the Code button above.\u003c/em\u003e\u003c/p\u003e","title":"Intelligent Government Service Assistant"},{"content":"Vibe is an AI-powered fashion community application built exclusively for the OpenHarmony ecosystem. It provides users with intelligent fashion recommendations and serves as a vibrant platform for fashion enthusiasts to connect, share, and discover new trends.\nKey Features:\nAI-driven outfit recommendations and styling suggestions. Native integration with the OpenHarmony operating system. Social community features for user interaction and sharing. You can check out the repository via the Code button above.\n","permalink":"/projects/vibe/","summary":"\u003cp\u003eVibe is an AI-powered fashion community application built exclusively for the OpenHarmony ecosystem. It provides users with intelligent fashion recommendations and serves as a vibrant platform for fashion enthusiasts to connect, share, and discover new trends.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eKey Features:\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003eAI-driven outfit recommendations and styling suggestions.\u003c/li\u003e\n\u003cli\u003eNative integration with the OpenHarmony operating system.\u003c/li\u003e\n\u003cli\u003eSocial community features for user interaction and sharing.\u003c/li\u003e\n\u003c/ul\u003e\n\u003cp\u003e\u003cem\u003eYou can check out the repository via the Code button above.\u003c/em\u003e\u003c/p\u003e","title":"Vibe - AI Fashion Community"},{"content":"Congratulations! The manuscript has been submitted to the IEEE Journal of Biomedical and Health Informatics, Currently, the manuscript is under peer review. I will update this post with the abstract and key methodology once furtherprogress is made.\n","permalink":"/research/ieee_jbhi/","summary":"\u003cp\u003eCongratulations! The manuscript has been submitted to the IEEE Journal of Biomedical and Health Informatics, Currently, the manuscript is under peer review. I will update this post with the abstract and key methodology once furtherprogress is made.\u003c/p\u003e","title":"Manuscript Submitted to IEEE JBHI"},{"content":"写在前面 机器学习系统越来越依赖真实用户数据：医疗图像、位置轨迹、搜索日志、推荐行为、设备传感器和文本交互都可能进入训练流程。数据越细，模型越容易学到有用模式，也越可能记住个体信息。攻击者不一定需要拿到原始数据，只要能访问模型输出、梯度更新或模型参数，就可能推断某个样本是否参与训练，甚至重构训练样本的部分特征。\n差分隐私（Differential Privacy, DP）提供了一种可证明的隐私保护框架。它的核心目标不是让数据“绝对不可见”，而是限制单个个体对算法输出的影响：无论某个人的数据是否出现在数据集中，外部观察者看到输出后都不应显著改变对这个人的判断。\n本文按“定义、机制、组合、机器学习实践”的顺序展开。重点不是堆概念，而是回答三个实践问题：\n$\\epsilon,\\delta$ 到底控制什么？ 噪声为什么要按敏感度来加？ 在深度学习中如何使用 DP-SGD，并避免常见误区？ 1. 隐私风险从哪里来 1.1 成员推断攻击 成员推断攻击（membership inference attack）试图判断某条记录是否在训练集中出现过。若模型对训练样本的置信度系统性高于未见样本，攻击者就可以利用输出概率、损失值或预测排名进行判断。\n例如，一个疾病风险预测模型若对某个病人的记录表现出异常高的确定性，攻击者可能推断该病人曾参与某个敏感医疗数据集。即便模型没有直接输出训练数据，这种参与关系本身也可能是隐私信息。\n1.2 模型反演与梯度泄露 模型反演攻击（model inversion）尝试根据模型输出恢复输入特征；梯度泄露攻击则利用联邦学习或分布式训练中的梯度更新重构训练样本。研究表明，在小批量训练、过参数化模型和高维输入场景中，梯度可能携带大量样本信息。\n联邦学习只能减少原始数据离开本地设备的风险，并不自动保证隐私。如果上传的梯度或模型更新没有额外保护，攻击者仍然可能从更新中推断本地数据。因此，联邦学习常与安全聚合、差分隐私和可信执行环境组合使用。\n2. 差分隐私的形式化定义 2.1 相邻数据集 差分隐私比较的是两个相邻数据集 $D$ 和 $D^{\\prime}$。常见定义有两种：\nadd/remove 邻接：$D^{\\prime}$ 比 $D$ 多或少一条记录。 replace-one 邻接：$D$ 和 $D^{\\prime}$ 记录数相同，但其中一条记录不同。 不同邻接定义会影响敏感度和隐私预算解释。阅读论文或工具文档时，需要确认使用的是哪一种。\n2.2 $(\\epsilon,\\delta)$-差分隐私 随机算法 $M$ 满足 $(\\epsilon,\\delta)$-差分隐私，若对任意相邻数据集 $D,D^{\\prime}$ 和任意输出集合 $S$，都有\n$$ \\Pr[M(D)\\in S]\\le e^\\epsilon \\Pr[M(D^{\\prime})\\in S]+\\delta. $$其中：\n$\\epsilon$ 控制输出分布的最大可区分程度。越小，隐私越强，噪声通常越大。 $\\delta$ 是允许极小概率失败的松弛项。实际中通常要求 $\\delta$ 小于数据集规模倒数，例如 $\\delta \u003c 1/n$ 或更严格。 当 $\\delta=0$ 时，称为纯差分隐私（pure DP）；当 $\\delta\u003e0$ 时，称为近似差分隐私（approximate DP）。深度学习中常用的是近似 DP，因为高斯机制和 DP-SGD 更自然地落在 $(\\epsilon,\\delta)$ 框架下。\n2.3 隐私损失随机变量 为了理解 $\\epsilon$，可以定义隐私损失：\n$$ L(o)=\\log\\frac{\\Pr[M(D)=o]}{\\Pr[M(D^{\\prime})=o]}. $$差分隐私要求这个比值不要太大。也就是说，观察到输出 $o$ 后，攻击者不能非常确定地区分数据来自 $D$ 还是 $D^{\\prime}$。这就是“保护个体参与”的数学形式。\nDP Concept Graph 3. 差分隐私的三个基本性质 3.1 后处理不损害隐私 如果 $M(D)$ 已经满足差分隐私，那么对输出再做任意不访问原始数据的函数处理 $f(M(D))$，仍然满足同样的差分隐私。这叫后处理性质（post-processing）。\n这个性质很实用。例如模型训练时加入 DP 噪声后，后续把模型转换格式、压缩、部署、做推理接口封装，都不会额外消耗隐私预算。前提是这些步骤不再次访问未保护的训练数据。\n3.2 组合会消耗隐私预算 如果对同一数据集重复运行多个 DP 算法，总隐私损失会累积。最简单的顺序组合为：\n$$ \\epsilon_{\\text{total}}=\\sum_i \\epsilon_i. $$实际机器学习中，训练会迭代很多步，因此不能简单地每一步都用很大的 $\\epsilon$。DP-SGD 需要使用隐私会计（privacy accountant）来累计多轮采样、裁剪和加噪后的总隐私预算。常见会计方法包括 moments accountant、Rényi Differential Privacy (RDP) 和 Gaussian DP 等。\n3.3 并行组合可以节省预算 如果不同 DP 查询作用在互不重叠的数据子集上，总隐私预算可以取最大值而不是求和。这叫并行组合。它在分组统计、分区查询和某些联邦学习场景中有用。\n4. 敏感度：噪声尺度的来源 4.1 全局敏感度 对查询函数 $f:D\\to \\mathbb{R}^k$，其 $L_1$ 全局敏感度定义为\n$$ \\Delta_1 f=\\max_{D,D^{\\prime}}\\lVert f(D)-f(D^{\\prime})\\rVert_1. $$它表示改变一条记录时，查询结果最多会变化多少。差分隐私的噪声尺度不是随便选的，而是由敏感度和隐私预算共同决定。\n例如，统计数据库中满足某条件的人数。改变一条记录最多让计数变化 1，因此敏感度为 1。若统计平均值，则必须先限制每个样本的取值范围，否则单个异常值可能让平均值变化非常大，敏感度无法控制。\n4.2 裁剪是机器学习中的敏感度控制 深度学习中，单个样本的梯度范数可能很大。如果直接对梯度加噪，敏感度没有上界，无法给出有效 DP 保证。DP-SGD 的第一步就是对每个样本的梯度进行裁剪：\n$$ \\bar{g}_i = g_i \\cdot \\min\\left(1,\\frac{C}{\\|g_i\\|_2}\\right), $$其中 $C$ 是裁剪阈值。裁剪后每个样本对 batch 平均梯度的贡献被限制住，后续加高斯噪声才有意义。\n5. 基本机制 5.1 拉普拉斯机制 对数值查询 $f(D)$，拉普拉斯机制输出\n$$ M(D)=f(D)+\\eta,\\qquad \\eta\\sim \\operatorname{Laplace}\\left(0,\\frac{\\Delta_1 f}{\\epsilon}\\right). $$噪声尺度为 $\\Delta_1 f/\\epsilon$。敏感度越大，需要的噪声越大；$\\epsilon$ 越小，隐私越强，噪声也越大。\n下面的代码可视化不同尺度的拉普拉斯分布：\n1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 import numpy as np import matplotlib.pyplot as plt def laplace_pdf(x, mu, b): return 0.5 / b * np.exp(-np.abs(x - mu) / b) x = np.linspace(-10, 10, 1000) for b in [0.5, 1.0, 2.0]: plt.plot(x, laplace_pdf(x, 0, b), label=f\u0026#34;b={b}\u0026#34;) plt.title(\u0026#34;Laplace distribution with different scales\u0026#34;) plt.xlabel(\u0026#34;noise\u0026#34;) plt.ylabel(\u0026#34;density\u0026#34;) plt.legend() plt.show() 实践建议：拉普拉斯机制适合简单统计查询，如计数、求和、直方图。用于均值时必须先裁剪或限制取值范围。\n5.2 高斯机制 高斯机制输出\n$$ M(D)=f(D)+\\mathcal{N}(0,\\sigma^2 I). $$它通常用于 $(\\epsilon,\\delta)$-DP。与拉普拉斯机制相比，高斯噪声更适合 $L_2$ 敏感度和高维向量场景，因此在 DP-SGD 中更常见。\n在机器学习中，加噪对象不是最终模型输出，而是每一步的裁剪后平均梯度：\n$$ \\tilde{g}=\\frac1B\\left(\\sum_{i=1}^{B}\\bar{g}_i+\\mathcal{N}(0,\\sigma^2 C^2 I)\\right), $$其中 $B$ 是 batch size，$C$ 是裁剪阈值，$\\sigma$ 是 noise multiplier。\n5.3 指数机制 当输出不是数值，而是类别、候选项或离散选择时，可以使用指数机制。设效用函数为 $q(D,r)$，衡量输出候选 $r$ 的好坏；其敏感度为 $\\Delta q$。指数机制按如下概率采样：\n$$ \\Pr[M(D)=r]\\propto \\exp\\left(\\frac{\\epsilon q(D,r)}{2\\Delta q}\\right). $$效用越高的候选越容易被选中，但采样仍有随机性，从而限制单条记录对最终选择的影响。\n典型应用包括：隐私保护的投票、选择最优模型、选择特征、选择聚类中心等。它的关键是设计效用函数，并控制该函数对单条记录的敏感度。\n6. DP-SGD：深度学习中的差分隐私 6.1 标准 SGD 为什么不够 标准 SGD 每次用一个 batch 的梯度更新参数：\n$$ \\theta_{t+1}=\\theta_t-\\eta \\frac1B\\sum_{i=1}^{B}g_i. $$如果某个样本梯度非常特殊，它可能显著影响更新方向。模型经过多轮训练后，可能记住该样本的模式。过拟合越严重，成员推断攻击通常越容易成功。\nModel Inversion Attack 6.2 DP-SGD 的步骤 DP-SGD 每轮训练通常包含四步：\n对 batch 中每个样本计算 per-sample gradient。 将每个样本梯度裁剪到范数不超过 $C$。 对裁剪后的梯度和加入高斯噪声。 用加噪后的平均梯度更新模型参数。 伪代码如下：\n1 2 3 4 5 6 7 8 for each training step: sample a mini-batch for each example i: compute gradient g_i clip g_i to norm C add Gaussian noise to summed gradients average noisy gradients update model parameters DP-SGD 的核心代价来自两部分：裁剪会改变真实梯度方向，噪声会增加优化方差。因此它通常会降低模型精度，尤其在数据量小、模型大、任务难或隐私预算很严格时更明显。\n7. 用 Opacus 训练一个 DP 模型 7.1 最小代码结构 Opacus 是 PyTorch 生态中常用的 DP-SGD 工具。典型流程是先定义模型、优化器和 dataloader，然后用 PrivacyEngine 包装训练过程。\n1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 import torch import torch.nn as nn import torch.optim as optim from opacus import PrivacyEngine model = nn.Sequential( nn.Flatten(), nn.Linear(28 * 28, 128), nn.ReLU(), nn.Linear(128, 10), ) optimizer = optim.SGD(model.parameters(), lr=0.05, momentum=0.9) criterion = nn.CrossEntropyLoss() privacy_engine = PrivacyEngine() model, optimizer, train_loader = privacy_engine.make_private( module=model, optimizer=optimizer, data_loader=train_loader, noise_multiplier=1.1, max_grad_norm=1.0, ) for epoch in range(epochs): model.train() for x, y in train_loader: optimizer.zero_grad() logits = model(x) loss = criterion(logits, y) loss.backward() optimizer.step() epsilon = privacy_engine.get_epsilon(delta=1e-5) print(f\u0026#34;epoch={epoch}, epsilon={epsilon:.2f}\u0026#34;) 7.2 关键参数如何理解 参数 作用 调大后的影响 调小后的影响 noise_multiplier 高斯噪声强度 隐私更强，精度可能下降 精度可能更好，隐私变弱 max_grad_norm 梯度裁剪阈值 裁剪更少，但需要更大噪声 裁剪更强，可能欠拟合 batch size 每步样本数 梯度更稳定，但采样率变化会影响会计 噪声影响更明显 epochs 训练轮数 精度可能提升，但隐私预算消耗增加 隐私消耗少，但可能未收敛 learning rate 更新步长 收敛快但不稳定 稳定但可能慢 实践中不要只调 epsilon。更常见的流程是先设定隐私目标，再搜索 noise_multiplier、max_grad_norm、batch size 和学习率。\n7.3 一个可执行的调参顺序 先训练非 DP baseline，确认模型和数据管道正常。 固定模型结构和 batch size，加入 DP-SGD。 从 max_grad_norm=1.0、noise_multiplier=1.0 附近开始。 若训练 loss 不下降，先调学习率和裁剪阈值，不要马上降低噪声。 若 epsilon 太大，提高 noise multiplier 或减少 epochs。 若精度下降过多，尝试更小模型、更强正则、更好预训练或更大数据量。 8. 结果解释：精度下降不是失败 假设在 MNIST 上比较标准 SGD 和 DP-SGD，可能得到类似结果：\n设置 测试准确率 隐私说明 标准 SGD 96% 左右 没有显式 DP 保证 DP-SGD 88% 到 92% 取决于 $\\epsilon,\\delta$ 和噪声强度 这种差距并不意外。DP 的目标不是免费提升泛化，而是在可接受性能下给出可审计的个体隐私保护。对于医疗、金融和用户行为数据，牺牲少量精度换取明确隐私边界通常是合理的工程选择。\n需要避免一种错误表达：不能只说“我们加了噪声，所以保护隐私”。严谨表述应包含邻接定义、机制、$\\epsilon,\\delta$、隐私会计方法、训练轮数、采样率和裁剪阈值。\n9. 常见误区 9.1 只对最终模型参数加噪 训练完成后再给模型参数加一点噪声，通常不能提供可靠 DP 保证。因为模型在训练过程中已经充分接触原始数据，可能已经记住敏感样本。DP-SGD 的关键是在每一步优化时控制单样本贡献并加噪。\n9.2 不裁剪梯度就加噪 没有裁剪就没有稳定敏感度上界。噪声尺度无法校准，隐私保证也不成立。\n9.3 只报告 epsilon，不报告 delta $(\\epsilon,\\delta)$ 是一组参数。只报告 $\\epsilon$ 会隐藏失败概率假设。还应报告数据集规模和 $\\delta$ 的选取理由。\n9.4 把 DP 当成防御所有攻击的方法 差分隐私主要限制个体参与信息泄露，不等于防止模型被盗、不等于防止提示注入、不等于防止训练数据本身被系统管理员看到。完整隐私系统通常还需要访问控制、加密、安全审计、数据最小化和合规流程。\n10. 实践清单 如果要在项目中使用差分隐私，可以按下面清单推进：\n明确保护对象：保护用户是否参与训练，还是保护具体属性值？ 明确数据边界：一条 record 是一个样本、一个用户，还是一个用户的所有样本？ 选择邻接定义：add/remove 还是 replace-one。 对输入做裁剪或范围限制，确保敏感度可控。 选择机制：简单统计用拉普拉斯或高斯机制；模型训练用 DP-SGD；离散选择用指数机制。 记录隐私会计方法和最终 $\\epsilon,\\delta$。 与非 DP baseline 对比，报告精度、训练稳定性和隐私预算。 做成员推断攻击评估，验证 DP 是否降低攻击成功率。 保留实验配置，确保隐私声明可复现。 11. 小结 差分隐私的核心思想是限制单条记录对输出分布的影响。敏感度决定噪声尺度，隐私预算控制可区分程度，组合定理解释多次查询或多轮训练时隐私如何累积。对于机器学习，DP-SGD 通过 per-sample gradient clipping 和 Gaussian noise 把这个思想嵌入优化过程。\n实际使用 DP 时，最重要的是完整报告：数据邻接定义、裁剪阈值、噪声强度、采样率、训练轮数、会计方法和最终 $\\epsilon,\\delta$。只有这些信息齐全，隐私保证才是可解释、可复现、可审计的。\n参考文献 [1] Cynthia Dwork. Differential Privacy. ICALP, 2006.\n[2] Cynthia Dwork, Aaron Roth. The Algorithmic Foundations of Differential Privacy. Now Publishers, 2014.\n[3] Frank McSherry, Kunal Talwar. Mechanism Design via Differential Privacy. FOCS, 2007.\n[4] Martin Abadi, Andy Chu, Ian Goodfellow, H. Brendan McMahan, Ilya Mironov, Kunal Talwar, Li Zhang. Deep Learning with Differential Privacy. ACM CCS, 2016.\n[5] Ilya Mironov. Rényi Differential Privacy. IEEE CSF, 2017.\n[6] H. Brendan McMahan, Daniel Ramage, Kunal Talwar, Li Zhang. Learning Differentially Private Recurrent Language Models. ICLR, 2018.\n[7] Nicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos, Dawn Song. The Secret Sharer: Evaluating and Testing Unintended Memorization in Neural Networks. USENIX Security, 2019.\n[8] Opacus Team. Opacus: User-Friendly Differential Privacy Library in PyTorch. Meta AI.\n[9] Bangzhou Xin. Research on Differential Privacy Techniques in Machine Learning. University of Science and Technology of China, 2022.\n[10] Xiaoguang Li, Hui Li, et al. A Survey on Differential Privacy. Journal of Information Security, 2018.\n","permalink":"/blog/dp_learn/","summary":"\u003ch2 id=\"写在前面\"\u003e写在前面\u003c/h2\u003e\n\u003cp\u003e机器学习系统越来越依赖真实用户数据：医疗图像、位置轨迹、搜索日志、推荐行为、设备传感器和文本交互都可能进入训练流程。数据越细，模型越容易学到有用模式，也越可能记住个体信息。攻击者不一定需要拿到原始数据，只要能访问模型输出、梯度更新或模型参数，就可能推断某个样本是否参与训练，甚至重构训练样本的部分特征。\u003c/p\u003e\n\u003cp\u003e差分隐私（Differential Privacy, DP）提供了一种可证明的隐私保护框架。它的核心目标不是让数据“绝对不可见”，而是限制单个个体对算法输出的影响：无论某个人的数据是否出现在数据集中，外部观察者看到输出后都不应显著改变对这个人的判断。\u003c/p\u003e\n\u003cp\u003e本文按“定义、机制、组合、机器学习实践”的顺序展开。重点不是堆概念，而是回答三个实践问题：\u003c/p\u003e\n\u003col\u003e\n\u003cli\u003e$\\epsilon,\\delta$ 到底控制什么？\u003c/li\u003e\n\u003cli\u003e噪声为什么要按敏感度来加？\u003c/li\u003e\n\u003cli\u003e在深度学习中如何使用 DP-SGD，并避免常见误区？\u003c/li\u003e\n\u003c/ol\u003e\n\u003ch2 id=\"1-隐私风险从哪里来\"\u003e1. 隐私风险从哪里来\u003c/h2\u003e\n\u003ch3 id=\"11-成员推断攻击\"\u003e1.1 成员推断攻击\u003c/h3\u003e\n\u003cp\u003e成员推断攻击（membership inference attack）试图判断某条记录是否在训练集中出现过。若模型对训练样本的置信度系统性高于未见样本，攻击者就可以利用输出概率、损失值或预测排名进行判断。\u003c/p\u003e\n\u003cp\u003e例如，一个疾病风险预测模型若对某个病人的记录表现出异常高的确定性，攻击者可能推断该病人曾参与某个敏感医疗数据集。即便模型没有直接输出训练数据，这种参与关系本身也可能是隐私信息。\u003c/p\u003e\n\u003ch3 id=\"12-模型反演与梯度泄露\"\u003e1.2 模型反演与梯度泄露\u003c/h3\u003e\n\u003cp\u003e模型反演攻击（model inversion）尝试根据模型输出恢复输入特征；梯度泄露攻击则利用联邦学习或分布式训练中的梯度更新重构训练样本。研究表明，在小批量训练、过参数化模型和高维输入场景中，梯度可能携带大量样本信息。\u003c/p\u003e\n\u003cp\u003e联邦学习只能减少原始数据离开本地设备的风险，并不自动保证隐私。如果上传的梯度或模型更新没有额外保护，攻击者仍然可能从更新中推断本地数据。因此，联邦学习常与安全聚合、差分隐私和可信执行环境组合使用。\u003c/p\u003e\n\u003ch2 id=\"2-差分隐私的形式化定义\"\u003e2. 差分隐私的形式化定义\u003c/h2\u003e\n\u003ch3 id=\"21-相邻数据集\"\u003e2.1 相邻数据集\u003c/h3\u003e\n\u003cp\u003e差分隐私比较的是两个相邻数据集 $D$ 和 $D^{\\prime}$。常见定义有两种：\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003eadd/remove 邻接：$D^{\\prime}$ 比 $D$ 多或少一条记录。\u003c/li\u003e\n\u003cli\u003ereplace-one 邻接：$D$ 和 $D^{\\prime}$ 记录数相同，但其中一条记录不同。\u003c/li\u003e\n\u003c/ul\u003e\n\u003cp\u003e不同邻接定义会影响敏感度和隐私预算解释。阅读论文或工具文档时，需要确认使用的是哪一种。\u003c/p\u003e\n\u003ch3 id=\"22--差分隐私\"\u003e2.2 $(\\epsilon,\\delta)$-差分隐私\u003c/h3\u003e\n\u003cp\u003e随机算法 $M$ 满足 $(\\epsilon,\\delta)$-差分隐私，若对任意相邻数据集 $D,D^{\\prime}$ 和任意输出集合 $S$，都有\u003c/p\u003e\n$$\n\\Pr[M(D)\\in S]\\le e^\\epsilon \\Pr[M(D^{\\prime})\\in S]+\\delta.\n$$\u003cp\u003e其中：\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e$\\epsilon$ 控制输出分布的最大可区分程度。越小，隐私越强，噪声通常越大。\u003c/li\u003e\n\u003cli\u003e$\\delta$ 是允许极小概率失败的松弛项。实际中通常要求 $\\delta$ 小于数据集规模倒数，例如 $\\delta \u003c 1/n$ 或更严格。\u003c/li\u003e\n\u003c/ul\u003e\n\u003cp\u003e当 $\\delta=0$ 时，称为纯差分隐私（pure DP）；当 $\\delta\u003e0$ 时，称为近似差分隐私（approximate DP）。深度学习中常用的是近似 DP，因为高斯机制和 DP-SGD 更自然地落在 $(\\epsilon,\\delta)$ 框架下。\u003c/p\u003e","title":"差分隐私基础：从数学定义到机器学习实践"},{"content":"写在前面 仿射变换（affine transformation）是线性代数和几何之间非常实用的一座桥。它既可以解释“换坐标系后为什么同一个点会有不同坐标”，也可以解释图像处理中常见的平移、旋转、缩放、错切、投影近似和椭圆生成。理解仿射变换的关键不是记公式，而是把三件事分清：\n点本身是几何对象，坐标只是它在某个坐标系下的表示。 矩阵描述的是“基向量如何被送到新的方向和尺度”。 仿射变换保持直线、平行性和线段比例，但通常不保持长度、角度和圆。 本文从最基础的坐标和矩阵讲起，然后推到直线、面积、圆锥曲线和实际编程。目标是：读完之后能把仿射变换用于推导，也能把它写进代码。\n1. 坐标：点不变，表示会变 1.1 坐标依赖于基底 在平面直角坐标系 $xOy$ 中，点 $P=(4,3)$ 通常被理解为：从原点出发，沿 $x$ 轴正方向走 4 个单位，再沿 $y$ 轴正方向走 3 个单位。\n用向量语言写，若 $\\vec{i},\\vec{j}$ 是标准正交基，则\n$$ \\overrightarrow{OP}=4\\vec{i}+3\\vec{j}. $$如果换一组基底，例如\n$$ \\vec{i}^{\\prime}=2\\vec{i},\\qquad \\vec{j}^{\\prime}=\\frac12\\vec{j}, $$则同一个向量也可以写成\n$$ \\overrightarrow{OP}=2\\vec{i}^{\\prime}+6\\vec{j}^{\\prime}. $$所以同一个点在新基底下的坐标变成 $(2,6)$。这不是点移动了，而是表示方式变了。很多几何问题容易混乱，就是因为把“点的运动”和“坐标表示的改变”混在一起。\n1.2 主动变换与被动变换 理解仿射变换时，建议区分两种视角：\n主动变换：点真的被送到另一个位置。例如把图形旋转 30 度。 被动变换：点不动，只是观察它的坐标系变了。例如把坐标轴旋转 30 度。 两者公式很像，但矩阵方向往往相反。工程实践中，大多数图形库采用主动变换：给定点坐标 $x$，计算变换后的点 $x^{\\prime}$。数学推导中，换基和坐标变换经常采用被动视角。写代码时必须确认自己使用的是列向量约定还是行向量约定，否则旋转方向、矩阵乘法顺序都会出错。\n2. 线性方程组与行列式 2.1 二阶行列式的几何意义 二阶矩阵\n$$ A= \\begin{pmatrix} a \u0026 b \\\\ c \u0026 d \\end{pmatrix} $$的行列式为\n$$ \\det(A)=ad-bc. $$它的几何意义是面积缩放因子。若两个基向量张成的平行四边形原面积为 1，经过矩阵 $A$ 变换后面积变成 $|\\det(A)|$。符号则表示方向是否翻转：$\\det(A)\u003e0$ 保持方向，$\\det(A)\u003c0$ 翻转方向，$\\det(A)=0$ 会把平面压扁到一条线或一个点。\n这一点非常重要：仿射变换下长度和角度可能变化，但面积会统一乘以同一个因子 $|\\det(A)|$。\n2.2 线性方程组与可逆性 考虑线性方程组\n$$ \\begin{cases} a_{11}x+a_{12}y=b_1,\\\\ a_{21}x+a_{22}y=b_2. \\end{cases} $$写成矩阵形式为\n$$ A\\begin{bmatrix}x \\\\ y\\end{bmatrix} = \\begin{bmatrix}b_1 \\\\ b_2\\end{bmatrix}, \\qquad A= \\begin{bmatrix} a_{11} \u0026 a_{12}\\\\ a_{21} \u0026 a_{22} \\end{bmatrix}. $$当\n$$ \\det(A)=a_{11}a_{22}-a_{12}a_{21}\\neq 0 $$时，方程组有唯一解。几何上，这说明两条直线有唯一交点，也说明矩阵 $A$ 没有把平面压扁，因此可以反向恢复坐标。\n例如，用非标准基底\n$$ \\vec{i}=(1,2),\\qquad \\vec{j}=(3,4) $$表示点 $P=(5,10)$，需要解\n$$ x\\vec{i}+y\\vec{j}=\\overrightarrow{OP}. $$展开得到\n$$ \\begin{cases} x+3y=5,\\\\ 2x+4y=10. \\end{cases} $$解得 $(x,y)=(5,0)$。这说明 $P$ 正好在 $\\vec{i}$ 方向上走 5 个单位即可到达。\n3. 什么是仿射变换 3.1 定义 二维仿射变换可以写成\n$$ T(x)=Ax+b, $$其中 $A$ 是 $2\\times2$ 矩阵，$b$ 是二维平移向量。若写成坐标形式：\n$$ \\begin{bmatrix} x^{\\prime} \\\\ y^{\\prime} \\end{bmatrix} = \\begin{bmatrix} a_{11} \u0026 a_{12} \\\\ a_{21} \u0026 a_{22} \\end{bmatrix} \\begin{bmatrix} x \\\\ y \\end{bmatrix} + \\begin{bmatrix} t_x \\\\ t_y \\end{bmatrix}. $$线性变换只能描述旋转、缩放、错切、反射等“固定原点”的操作；仿射变换在线性变换后增加平移，因此可以描述更完整的平面图形变换。\n3.2 齐次坐标 为了把平移也写成矩阵乘法，可以引入齐次坐标：\n$$ \\begin{bmatrix} x^{\\prime} \\\\ y^{\\prime} \\\\ 1 \\end{bmatrix} = \\begin{bmatrix} a_{11} \u0026 a_{12} \u0026 t_x \\\\ a_{21} \u0026 a_{22} \u0026 t_y \\\\ 0 \u0026 0 \u0026 1 \\end{bmatrix} \\begin{bmatrix} x \\\\ y \\\\ 1 \\end{bmatrix}. $$这在计算机图形学中非常常见，因为多个变换可以直接通过矩阵相乘组合起来。例如先缩放、再旋转、再平移，可以写成一个矩阵：\n$$ T = T_{\\text{translate}}T_{\\text{rotate}}T_{\\text{scale}}. $$注意矩阵乘法一般不交换，先旋转后平移和先平移后旋转不是同一个操作。\n4. 常见仿射变换 4.1 缩放 缩放矩阵为\n$$ A= \\begin{pmatrix} s_x \u0026 0\\\\ 0 \u0026 s_y \\end{pmatrix}. $$若 $s_x=s_y$，图形等比例缩放；若 $s_x\\neq s_y$，圆会变成椭圆。面积缩放因子为\n$$ |\\det(A)|=|s_xs_y|. $$4.2 旋转 采用列向量约定时，逆时针旋转 $\\theta$ 的矩阵为\n$$ R(\\theta)= \\begin{pmatrix} \\cos\\theta \u0026 -\\sin\\theta\\\\ \\sin\\theta \u0026 \\cos\\theta \\end{pmatrix}. $$旋转矩阵满足\n$$ R^\\top R=I,\\qquad \\det(R)=1. $$因此旋转保持长度、角度和面积，是一种特殊的仿射变换。\n4.3 错切 错切矩阵可以写成\n$$ A= \\begin{pmatrix} 1 \u0026 k\\\\ 0 \u0026 1 \\end{pmatrix}. $$它会把竖直方向保持不变，把 $x$ 坐标按 $y$ 的大小平移。错切不保持角度，但 $\\det(A)=1$，所以保持面积。\n4.4 平移 平移为\n$$ T(x)=x+b. $$平移不改变图形形状、面积、长度和角度，但它不是线性变换，因为 $T(0)\\neq0$。这就是为什么仿射变换比线性变换多一个平移项。\n5. 仿射变换保持什么 5.1 直线仍然是直线 设直线参数形式为\n$$ x(t)=p+t v. $$经过仿射变换 $T(x)=Ax+b$ 后，\n$$ T(x(t))=A(p+tv)+b=(Ap+b)+t(Av). $$这仍然是参数 $t$ 的一次表达式，所以直线会变成直线。若 $A$ 可逆，非退化直线不会被压成点。\n5.2 平行性保持 两条平行直线方向向量相同或成比例。若方向向量分别为 $v$ 和 $\\lambda v$，变换后为 $Av$ 和 $A(\\lambda v)=\\lambda Av$，仍然成比例。因此仿射变换保持平行性。\n5.3 线段比例保持 若点 $C$ 在线段 $AB$ 上，且\n$$ C=(1-t)A+tB, $$经过仿射变换后：\n$$ T(C)=(1-t)T(A)+tT(B). $$所以中点仍然是中点，三等分点仍然是三等分点。这个性质在计算机视觉和图形插值中很有用。\n5.4 面积按行列式统一缩放 若一个区域 $S$ 经过线性部分 $A$ 变换，则面积满足\n$$ \\operatorname{Area}(T(S))=|\\det(A)|\\operatorname{Area}(S). $$因此两个区域的面积比会保持不变：\n$$ \\frac{\\operatorname{Area}(T(S_1))}{\\operatorname{Area}(T(S_2))}= \\frac{\\operatorname{Area}(S_1)}{\\operatorname{Area}(S_2)}. $$仿射几何关注的正是这些不会随仿射变换改变的性质。\n6. 直线方程如何变换 6.1 用法向量表示直线 直线可以写成\n$$ n^\\top x+c=0, $$其中 $n=(A,B)^\\top$ 是法向量，$c=C$，对应常见形式\n$$ Ax+By+C=0. $$若点经过仿射变换\n$$ x^{\\prime}=Mx+b, $$且 $M$ 可逆，则\n$$ x=M^{-1}(x^{\\prime}-b). $$代回直线方程：\n$$ n^\\top M^{-1}(x^{\\prime}-b)+c=0. $$整理可得新直线：\n$$ (M^{-\\top}n)^\\top x^{\\prime} + c - n^\\top M^{-1}b=0. $$所以法向量按 $M^{-\\top}$ 变换，而不是按 $M$ 直接变换。这一点在图形学中也会出现：法向量的变换通常使用 inverse-transpose matrix。\n6.2 实用检查 如果只变换点，用 $x^{\\prime}=Mx+b$；如果变换直线或平面法向量，要用 $M^{-\\top}$。这两者不能混用。一个简单测试是：变换前点在直线上，变换后点也必须满足新直线方程。\n7. 圆、椭圆与切线 7.1 椭圆可以看成圆的仿射像 单位圆为\n$$ u^2+v^2=1. $$经过缩放\n$$ x=au,\\qquad y=bv $$后得到\n$$ \\frac{x^2}{a^2}+\\frac{y^2}{b^2}=1. $$因此椭圆可以看作单位圆在两个方向上不同尺度缩放后的结果。面积也立即得到：\n$$ S_{\\text{ellipse}}=|\\det(A)|S_{\\text{circle}}=ab\\pi. $$7.2 椭圆切线 单位圆上点 $(u_0,v_0)$ 的切线为\n$$ u_0u+v_0v=1. $$令\n$$ u=\\frac{x}{a},\\qquad v=\\frac{y}{b}, $$且椭圆上对应点为 $(x_0,y_0)=(au_0,bv_0)$，则切线变成\n$$ \\frac{x_0x}{a^2}+\\frac{y_0y}{b^2}=1. $$这比直接对椭圆方程求导更能说明几何来源：椭圆的切线关系来自圆的切线关系，再经过仿射变换得到。\n7.3 圆锥曲线的统一视角 圆、椭圆、抛物线和双曲线都可以用二次型描述：\n$$ x^\\top Qx+q^\\top x+c=0. $$仿射变换会改变 $Q,q,c$，但曲线的退化性、相交关系、切线关系等可以通过矩阵形式统一处理。计算机视觉中的图像配准、平面单应、二次曲线拟合都大量使用这种写法。\n8. 编程中的仿射变换 8.1 用 NumPy 变换点集 下面采用列向量约定，但为了方便批量计算，代码中把点集存成 $N\\times2$，最后使用右乘转置：\n1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 import numpy as np points = np.array([ [0.0, 0.0], [1.0, 0.0], [1.0, 1.0], [0.0, 1.0], ]) theta = np.deg2rad(30) R = np.array([ [np.cos(theta), -np.sin(theta)], [np.sin(theta), np.cos(theta)], ]) S = np.array([ [2.0, 0.0], [0.0, 0.8], ]) t = np.array([3.0, 1.0]) A = R @ S transformed = points @ A.T + t print(transformed) 需要注意两点：\n若点按行存储，使用 points @ A.T + t。 若点按列存储，使用 A @ points + t。 两种写法都可以，但不能在同一个项目里混用。\n8.2 变换图像时的反向采样 图像处理中，若直接把原图每个像素 $x$ 映射到目标位置 $x^{\\prime}=Ax+b$，目标图可能出现空洞。因此实际图像 warp 通常采用反向采样：\n$$ x=A^{-1}(x^{\\prime}-b). $$也就是遍历目标图每个像素，反查它在原图中的位置，再用最近邻、双线性或双三次插值取值。OpenCV、Pillow、skimage 等库基本都采用这个思路。\n8.3 调试仿射变换的清单 实际写代码时，可以按下面顺序排查：\n明确使用行向量还是列向量。 明确旋转角度单位是度还是弧度。 明确坐标系 $y$ 轴向上还是向下。图像坐标通常 $y$ 轴向下。 检查矩阵乘法顺序。先做的变换通常离点更近。 检查 $\\det(A)$ 是否接近 0，避免不可逆或数值不稳定。 若变换法向量，使用 $A^{-\\top}$。 9. 小结 仿射变换可以概括为\n$$ T(x)=Ax+b. $$它保持直线、平行性、线段比例和面积比，但不一定保持长度、角度和圆。行列式给出面积缩放因子，逆矩阵负责把目标坐标反查回原坐标，齐次坐标则让平移和线性部分统一成一个矩阵。\n从学习路线看，仿射变换最值得掌握的是三种能力：看懂公式里的几何含义，能从几何问题写出矩阵，能在代码中稳定地实现和调试。掌握这三点之后，椭圆、图像变换、相机几何和图形渲染中的很多公式都会变得更自然。\n参考文献 [1] Gilbert Strang. Introduction to Linear Algebra. Wellesley-Cambridge Press, 2016.\n[2] Marcel Berger. Geometry I. Springer, 1987.\n[3] H. S. M. Coxeter. Introduction to Geometry. Wiley, 1969.\n[4] Richard Hartley, Andrew Zisserman. Multiple View Geometry in Computer Vision. Cambridge University Press, 2004.\n[5] Richard Szeliski. Computer Vision: Algorithms and Applications. Springer, 2022.\n[6] John D. Foley, Andries van Dam, Steven K. Feiner, John F. Hughes. Computer Graphics: Principles and Practice. Addison-Wesley, 1996.\n","permalink":"/blog/affine_geometry/","summary":"\u003ch2 id=\"写在前面\"\u003e写在前面\u003c/h2\u003e\n\u003cp\u003e仿射变换（affine transformation）是线性代数和几何之间非常实用的一座桥。它既可以解释“换坐标系后为什么同一个点会有不同坐标”，也可以解释图像处理中常见的平移、旋转、缩放、错切、投影近似和椭圆生成。理解仿射变换的关键不是记公式，而是把三件事分清：\u003c/p\u003e\n\u003col\u003e\n\u003cli\u003e点本身是几何对象，坐标只是它在某个坐标系下的表示。\u003c/li\u003e\n\u003cli\u003e矩阵描述的是“基向量如何被送到新的方向和尺度”。\u003c/li\u003e\n\u003cli\u003e仿射变换保持直线、平行性和线段比例，但通常不保持长度、角度和圆。\u003c/li\u003e\n\u003c/ol\u003e\n\u003cp\u003e本文从最基础的坐标和矩阵讲起，然后推到直线、面积、圆锥曲线和实际编程。目标是：读完之后能把仿射变换用于推导，也能把它写进代码。\u003c/p\u003e\n\u003ch2 id=\"1-坐标点不变表示会变\"\u003e1. 坐标：点不变，表示会变\u003c/h2\u003e\n\u003ch3 id=\"11-坐标依赖于基底\"\u003e1.1 坐标依赖于基底\u003c/h3\u003e\n\u003cp\u003e在平面直角坐标系 $xOy$ 中，点 $P=(4,3)$ 通常被理解为：从原点出发，沿 $x$ 轴正方向走 4 个单位，再沿 $y$ 轴正方向走 3 个单位。\u003c/p\u003e\n\u003cimg src=\"./images/image3.png\" /\u003e\n\u003cp\u003e用向量语言写，若 $\\vec{i},\\vec{j}$ 是标准正交基，则\u003c/p\u003e\n$$\n\\overrightarrow{OP}=4\\vec{i}+3\\vec{j}.\n$$\u003cp\u003e如果换一组基底，例如\u003c/p\u003e\n$$\n\\vec{i}^{\\prime}=2\\vec{i},\\qquad \\vec{j}^{\\prime}=\\frac12\\vec{j},\n$$\u003cp\u003e则同一个向量也可以写成\u003c/p\u003e\n$$\n\\overrightarrow{OP}=2\\vec{i}^{\\prime}+6\\vec{j}^{\\prime}.\n$$\u003cp\u003e所以同一个点在新基底下的坐标变成 $(2,6)$。这不是点移动了，而是表示方式变了。很多几何问题容易混乱，就是因为把“点的运动”和“坐标表示的改变”混在一起。\u003c/p\u003e\n\u003ch3 id=\"12-主动变换与被动变换\"\u003e1.2 主动变换与被动变换\u003c/h3\u003e\n\u003cp\u003e理解仿射变换时，建议区分两种视角：\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e主动变换：点真的被送到另一个位置。例如把图形旋转 30 度。\u003c/li\u003e\n\u003cli\u003e被动变换：点不动，只是观察它的坐标系变了。例如把坐标轴旋转 30 度。\u003c/li\u003e\n\u003c/ul\u003e\n\u003cp\u003e两者公式很像，但矩阵方向往往相反。工程实践中，大多数图形库采用主动变换：给定点坐标 $x$，计算变换后的点 $x^{\\prime}$。数学推导中，换基和坐标变换经常采用被动视角。写代码时必须确认自己使用的是列向量约定还是行向量约定，否则旋转方向、矩阵乘法顺序都会出错。\u003c/p\u003e\n\u003ch2 id=\"2-线性方程组与行列式\"\u003e2. 线性方程组与行列式\u003c/h2\u003e\n\u003ch3 id=\"21-二阶行列式的几何意义\"\u003e2.1 二阶行列式的几何意义\u003c/h3\u003e\n\u003cp\u003e二阶矩阵\u003c/p\u003e\n$$\nA=\n\\begin{pmatrix}\na \u0026 b \\\\\nc \u0026 d\n\\end{pmatrix}\n$$\u003cp\u003e的行列式为\u003c/p\u003e\n$$\n\\det(A)=ad-bc.\n$$\u003cp\u003e它的几何意义是面积缩放因子。若两个基向量张成的平行四边形原面积为 1，经过矩阵 $A$ 变换后面积变成 $|\\det(A)|$。符号则表示方向是否翻转：$\\det(A)\u003e0$ 保持方向，$\\det(A)\u003c0$ 翻转方向，$\\det(A)=0$ 会把平面压扁到一条线或一个点。\u003c/p\u003e","title":"仿射变换入门：从坐标、矩阵到椭圆与不变量"}]