human_decomposition_military_org_20260416_deep_research
方法论库 · 引用级 · none
本页是 <code>contexts/methodology/human_decomposition_military_org_20260416_deep_research.md</code> 的逐字投影(仅隐私清洗,零改写)。
时点提示:本页是仓内文件 contexts/methodology/human_decomposition_military_org_20260416_deep_research.md 的逐字投影(仅做隐私清洗:仓库根绝对路径→相对路径、家目录→~/;除此零改写)。若源文件后续有修订,以仓内真源为准。
报告元数据(frontmatter)
name: human_decomposition_military_org_20260416_deep_research
description: 军事/组织拆解原则迁移
domain: cross
consumption:
surface: none
trigger: ""
consumer: orchestrator
status: library
promoted_to: null
superseded_by: contexts/methodology/task_decomposition_history_methodologies_survey_20260416.md人类历史上的任务分解方法论与反例:军事、组织、认知科学的深度调研
调研日期:2026-04-16
调研者:Claude (Opus 4.7, 1M context, sub-agent)
调研范围:军事指挥链与任务分权、组织理论、认知科学中的分解、分解失败案例、跨域同构
摘要
人类在大规模协作中面对的核心问题,是如何把一个复杂目标切分给很多双手去做,又不让切分的裂缝变成失败的断层。本调研梳理了从拿破仑军团制、普鲁士 Auftragstaktik,到 Mintzberg 组织学配置、Amazon 两张披萨团队、Spotify 模型的完整谱系,并回访 Healthcare.gov、NHS NPfIT、瀑布模型被误读的历史等反例。一个贯穿所有成功实践的底层原则是 Parnas 1972 年提出的信息隐藏:每个模块隐藏自己的局部决策,对外只暴露稳定的接口与意图。失败模式也高度同构:把过程分解混淆为结构分解、把自治理解为放任、忽视 span of control 的认知容量上限。末尾专题从这些共性中提炼出对 agent 编排的四个核心启示。
一、军事指挥链与任务分权
1.1 拿破仑军团制(Corps d'Armée):决策权的前移
拿破仑不是军团制的发明者,Bourcet 和 Guibert 在 18 世纪末已经提出过类似构想,但他是第一个把军团变成常设编制并贯通指挥、情报、后勤的人。Saber and Scroll Journal 的分析指出:「Each Corps was essentially a miniature army; each possessed cavalry, artillery and infantry and each was large enough that it could fight independently until another Corps could come to its support.」[https://saberandscroll.scholasticahq.com/article/28520-napoleon-apex-of-the-military-revolution.pdf]
核心设计原则:
第一,每个军团都是自足的小型军队,包含步兵、骑兵、炮兵、参谋。一个典型军团约 28,000 人。这意味着任何一个军团在等待援军的一天之内不会被完整消灭(「no single corps of roughly 28,000 men could be overwhelmed in one day」[Calhoun NPS])。
第二,军团之间保持不超过一天行军距离,形成可相互支援的网络。这是一种冗余而不浪费的拓扑:既不是集中兵力也不是散落兵力,而是一种弹性耦合。
第三,Napoleon 是「centralized control with decentralized operations」的早期体现者。他在整体上保持指挥权(选取目标、决定主力方向),但把战术执行下放给军团指挥官,让他们根据本地情报自己决定走哪条路、何时投入战斗。
第四,信息系统层面。Napoleon 建立了以 Berthier 为核心的参谋部,标准化了情报传递与命令格式。这是把 Auftragstaktik 得以运作的物理基础设施:没有标准化的消息通道,再好的放权理念也不能落地。[https://calhoun.nps.edu/server/api/core/bitstreams/60f29c3d-a116-4d83-863b-6149bee61da7/content]
1.2 Auftragstaktik:普鲁士-德意志传统
Auftragstaktik 在字面上是「任务战术」(Auftrag = 任务,Taktik = 战术)。关键事实,而这也是很多现代文献搞混的:Moltke 本人从来没用过 Auftragstaktik 这个词。这个术语首次出现在德意志军队的正式文件中是在二战前后。Moltke 只区分两种命令:Befehl(详细指令)和 Weisung(指示/意图)。[DTIC ADA569668]
Moltke 1869 年《高级指挥官指令》(Instructions for Large Unit Commanders) 的核心段落(原译):
Orders should only go as far ahead as conditions can reasonably be predicted. These change very rapidly in war. Seldom will orders that anticipate far in advance and in detail succeed completely to execution. . . . The higher the authority, the shorter and more general will the orders be. The next lower command adds what further precision appears necessary. The detail of execution is left to the verbal order, to the command. Each thereby retains freedom of action and decision within his authority.
[https://apps.dtic.mil/sti/pdfs/ADA599111.pdf]
翻译:命令的前瞻性应只覆盖可预测的范围。战争中情况变化极快。试图提前详细规划的命令鲜少能完整执行到底。职位越高,命令应当越短越宽泛。下一级指挥官补充必要的精确度。执行细节留给口头命令和现场指令。每一级都在自己的权限内保有行动和决断的自由。
这段话的深刻之处:它不是在讲 management style,它是在讲一种认识论——战场的不确定性随时间和层级指数增长,所以命令的精确度应该随层级下降而增加,而不是相反。顶层设定 what 和 why,越往下越补充 how。
Wikipedia 对 Mission-type tactics 的描述捕获了另一个关键约束:「Building a high level of trust, competency and understanding is crucial for the success of such a doctrine.」[https://en.wikipedia.org/wiki/Mission-type_tactics] 也就是说,Auftragstaktik 不是一个可以即插即用的方法,它依赖漫长的军官教育投资。Moltke 同时代做的另一件重要事情是建立了普鲁士军事学院体系,让军官之间形成共同的战术语言。失去这层公共认知,放权会直接变成混乱。
1.3 现代美军 ADP 6-0 的 Mission Command
美军在 1986 年 FM 100-5 中首次正式纳入 mission orders,2012 年发布 ADP 6-0 正式命名为 Mission Command。2019 年修订版列出的七条原则:
- Competence(能力)
- Mutual trust(相互信任)
- Shared understanding(共享理解)
- Commander's intent(指挥官意图)
- Mission orders(任务式命令)
- Disciplined initiative(守纪的主动性)
- Risk acceptance(风险接纳)
[https://armypubs.army.mil/epubs/DR_pubs/DR_a/ARN34403-ADP_6-0-000-WEB-3.pdf]
ADP 6-0 对 mission orders 的定义值得特别摘录:
Mission orders are directives that convey the commander's desired end state, but not the method by which to achieve them. In other words, these orders are not prescriptive in nature.
[DTIC AD1001514]
翻译:任务式命令传达指挥官期望的最终状态,而不传达达到它的方法。换句话说,这类命令不是指示性的。
这里值得警惕的是 ADP 6-0 自己的一个问题。National Defense University Press 发表的 "Beyond Auftragstaktik: The Case Against Hyper-Decentralized Command" 指出,把 mission command 等同于极度去中心化是危险的误读:「centralized control required to do so, however, risks sacrificing adaptability. Moltke's dictum underscores the idea that the chaos of battle will eventually render preplanned controls obsolete, at which point they merely limit freedom of action.」但作者同时强调,Jominian(严格集中)和 extreme mission command 都不是答案,真正的 Moltke 传统是一种有节制的去中心化。[https://ndupress.ndu.edu/Media/News/News-Article-View/Article/2076032/beyond-auftragstaktik-the-case-against-hyper-decentralized-command/]
1.4 Boyd 的 OODA Loop 与 Organic Design
John Boyd 在 1970 年代通过研究朝鲜战争中 F-86 vs MiG-15 的交战记录(F-86 胜率 10:1),发现优势不在飞机性能,而在飞行员能更快地 Observe-Orient-Decide-Act 的能力。Boyd 把这个循环推广为一种普遍的冲突理论:能够进入对手的 OODA 循环内部、比对手更快地完成一个周期的一方,就能制造混乱并取得优势。[https://en.wikipedia.org/wiki/OODA_loop]
Boyd 的 Organic Design for Command and Control(1987)比 OODA 更深刻的地方是 Orientation = Schwerpunkt 这个论断:「Orientation is the Schwerpunkt. It shapes the way we interact with the environment—hence orientation shapes the way we observe, the way we decide, the way we act.」[Team One Network]
Orientation(方向感)不是一个简单的情境感知,而是基因遗传、文化传统、过往经验、当前情境的交互投射。Boyd 认为它是决定性环节。Observe 只是采集原始数据,Decide 和 Act 只是下游结果,真正的胜负在于你如何 orient。这对 agent 编排的启发是:sub-agent 的 orientation 质量决定了后续所有决策的质量,盲目给 sub-agent 抛任务而不灌输上下文,等于让它从零 orientation 开始。
1.5 Span of Control 的军事来源
Span of control 的概念由英国将军 Sir Ian Hamilton 在 1921 年的《The Soul and Body of an Army》中首次提出。Hamilton 基于对英军指挥官的观察总结:「leaders could not effectively control more than three to six people.」[https://www.thefivecoatconsultinggroup.com/the-coronavirus-crisis/span-of-control]
Graicunas 在 1930 年代从数学上深化:relationships 的复杂度不是线性增长,而是指数增长。一个主管管 4 个下属有 44 种关系,管 6 个下属则有 222 种关系。[同上]
美军现代实际编制的数据(Fivecoat):
- Infantry Squad Leader(9 人)→ 2 个直接下属
- Infantry Platoon Leader(38 人)→ 6 个直接下属
- Infantry Company Commander(130 人)→ 7 个直接下属
- Infantry Battalion Commander(700 人)→ 9 个直接下属
- Brigade Commander(3,000 人)→ 10 个直接下属
注意这里的模式:一线战斗单位(班、排)的 span 最窄(2-6),因为需要高度的战场协调;越往上走 span 可以扩大到 9-10,因为协调性质从实时战术转为策略规划。
美国海军 OPNAVINST 3120.32D 直接写入了规章:「Ordinarily, a supervisor should be immediately responsible for not less than three or more than seven subordinates.」[https://www.secnav.navy.mil/doni/Directives/03000%20Naval%20Operations%20and%20Readiness/03-100%20Naval%20Operations%20Support/3120.32D%20W%20CH-1.pdf]
二、组织理论
2.1 Mintzberg 的五种组织配置
Mintzberg(1979-1983)提出五种结构原型,每种围绕一个占主导的组织部分和一种核心协调机制:
| 配置 | 占主导部分 | 协调机制 | 适合环境 |
|---|---|---|---|
| Simple Structure | Strategic Apex(战略顶层) | Direct Supervision | 创业公司、小公司、危机情境 |
| Machine Bureaucracy | Technostructure(技术结构) | Standardization of Work Processes | 大规模、稳定、高产量的制造与后台 |
| Professional Bureaucracy | Operating Core(操作核心) | Standardization of Skills | 医院、学校、律师事务所 |
| Divisionalized Form | Middle Line(中层管理) | Standardization of Outputs | 多业务/多地域的大企业 |
| Adhocracy | Support Staff | Mutual Adjustment | 研发、咨询、创新密集 |
[https://umbrex.com/resources/frameworks/organization-frameworks/mintzberg-organizational-configurations/]
Mintzberg 的核心洞察:协调机制决定了结构。Simple Structure 通过人直接看着人来协调;Machine Bureaucracy 通过标准化流程;Professional Bureaucracy 通过标准化技能(医生不需要老板盯着,他们的医学院训练已经把流程内化了);Adhocracy 通过面对面的相互调适。
这个框架对 agent 编排的意义在于:它揭示了协调机制和组织结构是耦合的。你不能选了一个松散的 agent 网络结构,却期望它像 Machine Bureaucracy 一样通过标准化流程自动协调。每种结构带来自己的失败模式:Simple Structure 扩展性差;Machine Bureaucracy 僵化;Adhocracy 容易碎片化。
2.2 Matrix Organization 的失败案例
矩阵组织(同时按功能和项目分工)在 1970-1990 年代被 Philips、ABB、Digital Equipment 等跨国公司广泛采用。Eric Viardot 的评估:「Matrix structure usually ends in failure because it accrues more disadvantages than advantages when compared to divisional and functional models.」[https://www.linkedin.com/pulse/infamous-matrix-structure-other-types-organizations-eric-viardot]
具体失败模式(Leadership Edge 研究,PDF):
第一,双重汇报链(dual reporting)。一个员工同时向项目经理和职能经理汇报,当两人指示冲突时员工陷入瘫痪。研究显示 87% 的中层经理和 23% 的高层经理把「角色与责任不清」列为矩阵组织的核心问题。[theleadershipedge.com/wp-content/uploads/2025/06/Challenges_Strategies_of_Matrix_Orgs.pdf]
第二,责任与权力错位。HR 可能有全球政策的责任,但没有在各区域实施的权力。这种结构天然制造政治斗争。
第三,绩效评估困难。员工参与多个项目、向多个经理汇报,没有哪个经理能独立评估其表现。
矩阵组织失败的根因是它违反了 Fayol 从军事中继承的 unity of command 原则:一个下属不应同时向多个上级汇报。现代解法通常是让一条链「硬」、另一条链「软」(dotted line),或者引入 matrix guardian 这样的制度化协调角色。[epicflow.com]
2.3 Spotify Model:从神话到自我批判
2012 年 Spotify 的 Henrik Kniberg 和 Anders Ivarsson 发表的白皮书介绍了 Squad / Tribe / Chapter / Guild 四层结构:
- Squad:8-10 人的跨职能自组织团队,负责一个产品切片
- Tribe:相关 squad 的集合,约 100-150 人(Dunbar 数字上限)
- Chapter:同一职能域(如前端、数据)的横向社区
- Guild:跨部落的兴趣社群
2020 年,前 Spotify 产品经理 Jeremiah Lee 发表了著名的 "Spotify's Failed #SquadGoals",第一次系统揭示这个模型根本没按白皮书描述的方式运作过。核心摘录:
The Spotify model is revealed as a collection of cross-functional teams with too much autonomy and a poor management structure. Don't fall for it. Had Spotify referred to these ideas by their original names, perhaps it could have evaluated them more fairly when they failed instead of having to confront changing its cultural identity simply to find internal processes that worked well.
[https://www.jeremiahlee.com/posts/failed-squad-goals/]
Lee 和后续分析者(Chameleon.io, DecoCMS)总结的失败原因:
第一,autonomy 没有对应的 alignment。白皮书说的是 "Aligned Autonomy",但 Spotify 实际只交付了前半部分。Anders Ivarsson 后来承认,原计划是多篇系列文章涵盖 autonomy / alignment / accountability,但后两篇从没写出来。
第二,Chapter Lead 同时要做技术导师和人员经理。这两个角色本质是矛盾的:技术导师关心代码,人员经理关心人的成长。一个 Chapter Lead 要管理分布在多个 squad 的工程师,不共事的情况下根本没有依据评估绩效。
第三,协调成本随 squad 数量指数增长。10 个 squad 时「最小化依赖」是可能的,到 100 个 squad 时一切都相互依赖,没有明确的上级协调机制。
第四,cool-sounding names 的诅咒。Squad / Tribe / Chapter / Guild 这些名字看起来像《权力的游戏》,但对试图复制的公司来说,它们把本来可以用 team / department / functional area / interest group 这些标准词汇理解的东西神秘化了。结果是公司复制了结构没复制文化。
第五,Market dynamics 变化。2012 年 Spotify 处于音乐流媒体的爆发期,实验性策略正确;到 2020 年市场成熟,执行效率比探索更重要,旧组织结构不再匹配。
Jeremiah Lee 在 podcast 中的一个总结很尖锐:「It was not really agile. It was just not waterfall.」[Chameleon.io] 这句话指向一个深层问题:放弃瀑布不等于实现了 agile。去掉旧的控制结构如果没有新的协调机制填入,组织会陷入结构性混乱。
2.4 Coase 的公司边界理论
Ronald Coase 在 1937 年《The Nature of the Firm》中提出一个根本问题:如果市场通过价格机制就能协调经济活动,为什么还需要公司这种内部协调的组织存在?他的答案是 transaction cost(交易成本):市场有发现价格、谈判合同、执行合同的成本,公司内部则有信息流、激励、监督、绩效评估的成本。公司的边界由这两类成本的边际相等决定。[https://strategicmanagementreview.net/assets/articles/Bylund.pdf]
Williamson 1975 年的扩展:当资产专用性(asset specificity)高、契约不完备时,市场交易容易遭遇 hold-up 问题,这时纳入公司内部更优。
对 agent 编排的启示:what to delegate 是一个经济学问题而不仅是技术问题。如果协调一个 sub-agent 需要的上下文传递成本、验证成本超过自己做,就不应该 delegate。这是为什么 Opus 模式下把设计和质量把关留给主模型、把执行和调研下放给 sub-agent 是一个经济最优解——设计和把关的 trust cost 太高,不适合外包。
2.5 Amazon 的 Two-Pizza Teams
Jeff Bezos 的规则:任何团队应小到两张披萨能喂饱,也就是 5-10 人。背后的认知科学依据:
第一,Ringelmann 效应(社交懒散):个体贡献随团队规模增加而下降。Bibb Latané 的实验(蒙眼戴耳机喊叫)显示 6 人组只发出相当于个体 36% 能力的音量。即使控制了协调损失(让人相信自己在小组里但其实独自喊),仍然只有 74%。[https://buffer.com/resources/small-teams-why-startups-often-win-against-google-and-facebook]
第二,Hackman 的研究:团队沟通通道数随 n(n-1)/2 增长。Staats-Milkman-Fox 实验让 2 人组和 4 人组用乐高拼同一个图形,小组平均 36 分钟,大组 52 分钟(多花 44%),但大组对自己需要的时间预期比实际少一倍。[https://blog.nuclino.com/two-pizza-teams-the-science-behind-jeff-bezos-rule]
第三,Bezos 的反直觉观点:「Communication is a sign of dysfunction. It means people aren't working together in a close, organic way. We should be trying to figure out a way for teams to communicate less with each other, not more.」[同上]
Two-pizza teams 的完整设计不只是规模,还包括 single-threaded ownership:一个团队只管一个服务或产品线,不跨多个服务。配合 Bezos 的 API 强制令(任何内部团队之间只能通过明确的 API 交互,不允许绕过),Amazon 把公司变成了「一台制造机器的机器」(a machine that makes the machine)。[https://www.theguardian.com/technology/2018/apr/24/the-two-pizza-rule-and-the-secret-of-amazons-success]
三、认知科学中的分解
3.1 Sweller 的认知负荷理论(Cognitive Load Theory)
John Sweller 在 1988 年提出的 CLT 把工作记忆的负荷分为三类:
- Intrinsic load:任务本身的内在复杂度。由元素交互度(element interactivity)决定,只能通过学习者专业化程度来缓解。
- Extraneous load:由不良教学或呈现方式引入的额外负荷,纯粹有害,应被削减。
- Germane load:用于把新信息整合进长期记忆中的 schema 的有效认知努力。
2010 年 Sweller 的修订(Educational Psychology Review, Vol 22, pp. 123-138)把 germane load 重新定义为 intrinsic load 的一部分,不是独立来源。[https://link.springer.com/article/10.1007/s10648-010-9128-5]
CLT 对任务分解的直接启示:工作记忆容量是硬约束。Miller 的 7±2 或 Cowan 的 4±1 都说明人类同时持有的元素数量极有限。任务分解的目的不仅是把工作切小,更是降低每一步的 element interactivity,让专家可以通过 schema 识别快速处理,不专家可以在较低 intrinsic load 下学习。
3.2 Chase 和 Simon 的 Chunking 实验
1973 年 Chase 和 Simon 的经典棋盘实验:让象棋大师、熟练棋手和新手分别观察 5 秒棋局,然后复盘。
- 合法残局:大师可复盘约 20-25 颗棋,熟练者 10-15 颗,新手 4-6 颗
- 随机摆放(无意义棋局):三组表现几乎一样,都在 4-6 颗
解释:大师并没有更好的短期记忆容量,他们记忆的是 chunk(有意义的模式团)。Simon 和 Gilmartin 估计一个象棋大师在长期记忆中持有 10,000 到 100,000 个 chunks(文献常引 50,000)。识别出一个 chunk 后,大师会用一个「指针」在工作记忆中代替整块内容,从而突破了 7±2 的限制。[https://pubmed.ncbi.nlm.nih.gov/9709441/]
这个实验对任务分解的启示非常深:专家和新手在面对同一个任务时,适合的分解粒度是不同的。对专家合适的「一步」可能包含几十个新手需要独立处理的细节。把一个任务切到「原子」级别其实可能会破坏专家的 chunk 识别能力,反而让他更慢。这是为什么好的工程组织会区分 junior 和 senior 的 ticket 粒度。
3.3 Hutchins 的分布式认知
Edwin Hutchins 1995 年《Cognition in the Wild》记录了他在美国海军 USS Palau 上观察导航团队的民族志研究。Hutchins 的核心论点:
[T]he navigation team can be seen as a cognitive and computational system... cultural activity systems have cognitive properties of their own that are different from the cognitive properties of the individuals who participate in them.
[https://direct.mit.edu/books/monograph/4892/Cognition-in-the-Wild]
翻译:导航团队可以被看作一个认知与计算系统……文化活动系统自身具有与其参与者的认知特性不同的认知特性。
Hutchins 提出认知过程的三种分布方式:
- 分布于社会群体成员之间(socially distributed)
- 分布于内部心智与外部结构之间(人与工具/环境的耦合)
- 分布于时间中(早期事件的产物改变后期事件的性质)
在 Palau 的实际案例中,没有任何一个船员掌握导航的全部知识。海图、罗盘、三角板、通话线路、笔记本、记录员、观察员、舵手、船长,是一个整体的认知系统。即使船长临时退出,这个系统仍能完成导航——因为知识和计算已经被嵌入到工具和程序中,不依赖任何单一头脑。
对 agent 编排的启示:不要把 agent 当作需要在单个调用中完成所有推理的孤立实体。设计好的工具(memory、checklist、structured output)能把一部分认知负担从 agent 转移到系统本身。
3.4 Hierarchical Task Analysis (HTA)
HTA 是人因工程和 HCI 领域的经典任务分析方法(Annett & Duncan, 1967 起源)。核心步骤:
- 识别主任务目标
- 把主目标分解为 sub-operations,附带 plan 说明执行的条件与顺序
- 对每个 sub-operation 决定是否进一步分解
- 迭代
- 分析分解以发现任务操作的低效
- 建议改进
HTA 的停止规则是 P × C:在某一粒度上,继续分解的代价 C 超过从中获益的概率 P 乘以获益大小时,应停止。[https://www.humanreliability.com/human-factors/hierarchical-task-analysis-hta/]
HTA 的示例(textual form):
0. 清洁房子
1. 取出吸尘器
2. 装上合适的吸头
3. 清洁房间
3.1 清洁走廊
3.2 清洁起居室
3.3 清洁卧室
4. 清空集尘袋
5. 收起吸尘器
Plan 0: 依次做 1,2,3,5; 集尘袋满时做 4
Plan 3: 按需做 3.1, 3.2, 3.3 任意顺序这个结构对 agent 编排而言简直是范本:顶层目标 → 分步子目标 → 元计划(plan)说明依赖和触发条件。Plan 0 的「集尘袋满时做 4」是一个条件触发,这是把状态监测耦合到分解中的经典做法。
四、分解失败案例的 Post-Mortem
4.1 瀑布模型被误读的历史
Winston Royce 1970 年在 IEEE WESCON 发表的 "Managing the Development of Large Software Systems" 是「瀑布模型」的源头,但这个模型被后世严重误读。
Royce 论文第一段原文:
I am going to describe my personal views about managing large software developments. I have had various assignments during the past nine years, mostly concerned with the development of software packages for spacecraft mission planning, commanding and post-flight analysis... I have become prejudiced by my experiences and I am going to relate some of these prejudices in this presentation.
[https://www.praxisframework.org/files/royce1970.pdf]
紧接着 Figure 2 呈现了线性的瀑布式流程,然后 Royce 写下了决定性的一句:
I believe in this concept, but the implementation described above is risky and invites failure.
[同上]
翻译:我相信这个概念,但上述的实现方式是有风险的,会招致失败。
Royce 随后提出了五项修改建议,其核心是让相邻阶段之间可迭代,并强调 "continuous customer involvement"。这实际上是一个准 agile 的愿景。但他的论文被后续的 DoD-STD-2167 (1985) 引用时,只用了 Figure 2 的严格线性版本,剥离了他自己的警告和修改建议。[https://notafactoryanymore.com/2015/02/18/waterfall-or-agile-reflections-on-winston-royces-original-paper/]
PMWorld Journal 2018 年 Johnny Morgan 的重读文章引用 Girvan & Paul 的评价:
Unfortunately, Royce's paper was widely misunderstood. He presented the above model as 'risky and invites failure' and was proposing modifications to make it much more iterative and incremental. However, that element of his work is largely forgotten, and his waterfall picture remains in common use.
[https://pmworldlibrary.net/wp-content/uploads/2018/07/pmwj72-Jul2018-Morgan-applying-1970-waterfall-lessons-umd-paper.pdf]
这个历史本身就是一个文档认知失败案例:原作者明确标注了「有风险,会失败」,但后继者从插图中抽取了结构而丢弃了警告。对 agent 编排的启示是惊人的——任何设计文档都会被后续读者简化,所以警告必须嵌入结构本身而不是注释。
4.2 Healthcare.gov 的崩溃(2013)
2013 年 10 月 1 日 Healthcare.gov 启动当天,只有 6 人成功注册。到问题修复,美国政府花费超过 6.3 亿美元。GAO 和多个独立审计报告揭示的失败模式:
第一,契约碎片化。总承包商 CGI 下有 16 个官方子承包商,加起来 55 个以上的契约方,「no focal point for project responsibility and accountability」。[https://medium.com/@ketan.keshav7/the-630-million-lesson]
第二,压缩测试窗口以保政治截止日期。内部 CMS 报告和 McKinsey 外部咨询都警告需求仍在变化、测试时间不足。管理层把必要的 7 个月端到端测试压缩到 1 个月。[同上]
第三,权力分裂。Forbes 报道:「three different parts of the bureaucracy contending for control—the IT shop, the policy shop, and the communications shop—key decisions were often delayed, guidance to contractors was inconsistent, and nobody was truly in charge.」[https://www.forbes.com/sites/lorenthompson/2013/12/03/healthcare-gov-diagnosis]
第四,不现实的需求工程。75 屏幕的流程、1000+ 屏幕的完整系统、5 个联邦机构 + 36 个州 + 300 家保险商 + 4000+ 保险计划。没有分层的信息隐藏,一个用户请求要穿透整个协作矩阵。[同上]
Anthopoulos 等人在 Government Information Quarterly 的学术分析(表 3-7)把失败归因到 PMBOK 九大知识域的每一个都出了问题,其中 scope management("complex architecture design; evolving requirements")和 communications management("complex due to the size of stakeholders")是重灾区。[http://greatquestion.com.au/wp-content/uploads/2023/09/Anthopoulos-et-al-Why-e-government-projects-fail]
根因诊断:Healthcare.gov 把 55 个团队组织在一起,但缺乏军事 Auftragstaktik 式的「共享意图 + 明确接口」。CMS 既不扮演 commander,也没有确立 commander's intent,于是每个承包商按自己的理解工作,最后集成阶段发现互不兼容。
4.3 UK NHS NPfIT/Connecting for Health(2002-2011)
英国 National Programme for IT 本来预算 60 亿英镑,最终花费超过 100 亿英镑,是公共部门 IT 项目史上最昂贵的失败之一。英国下议院 PAC 委员会在 2013 年定性为「one of the worst and most expensive contracting fiascos in the history of the public sector」。[https://www.bbc.com/news/uk-politics-24130684]
剑桥大学 Computer Laboratory 的 Ross Anderson 等人的 post-mortem 论文(2014)记录了多层失败:
第一,中心化集采破坏了既有系统。NPfIT 只资助两个主要软件供应商 (Cerner Millennium 和旧版 Lorenzo) 的新软件,不资助各医院 Trust 自己既有软件的维护。结果很多医院被迫停止自己能正常运作的系统去等待 NPfIT 交付,而 NPfIT 迟迟不来。[https://www.cl.cam.ac.uk/archive/rja14/Papers/npfit-mpp-2014-case-history.pdf]
第二,强势项目总监模式的反噬。NPfIT 项目总监 Richard Granger 有强烈个性,可以说「既是优势也是诅咒」。强势领导在项目初期推动决策,但一旦项目出问题后,没有足够的 check and balance 能力挑战他的判断。[https://www.computerweekly.com/opinion/Six-reasons-why-the-NHS-National-Programme-For-IT-failed]
第三,忽视一线用户。医生和 GP 从项目启动就表达担忧,认为系统不符合临床工作流,但这些反馈在集中化决策中被屏蔽。最终上线的系统「slow, cumbersome, insufficiently explained and poorly implemented」,医生要花 20 分钟教 GP 怎么用新系统。[同上]
第四,供应商之间的互联失败。NPfIT 架构假设供应商会顺利集成,但 Fujitsu 最终退出、CSC 陷入合同纠纷、Accenture 持续遭遇性能问题。
PAC 委员会的总结:「it was a failure of management for the NHS to have got itself in that mess with its supplier.」[https://publications.parliament.uk/pa/cm201314/cmselect/cmpubacc/294/294.pdf]
根因诊断:NPfIT 和 Healthcare.gov 高度相似。核心都是 over-centralization without commander's intent。架构上试图统一,但没有建立统一所需的共享认知、信任和接口标准。
4.4 SAFe 的结构性批评
Scaled Agile Framework (SAFe) 是目前全球最流行的「规模化敏捷」框架(State of Agile 2020 报告中 37% 采用率),但在敏捷专家社区中广受批评。Steve Denning 在 Forbes 2019 年的文章直接定性:
A particularly worrying variant is the Scaled Agile Framework or SAFe. Essentially this is codified bureaucracy, in which the customer is almost totally absent. It is now pervasive in large firms because it gives the management a mandate to call themselves agile and keep doing what they have always done.
[https://www.forbes.com/sites/stevedenning/2019/05/23/understanding-fake-agile/]
Willem-Jan Ageling 的系统性批评(altexsoft.com, liberators.io)提出 SAFe 的结构性问题:
- 类瀑布式的 PI(Program Increment):8-10 周的计划周期,远长于 Scrum 的 2 周 sprint,已经失去了 agile 的短反馈特性
- 角色过度规定:Release Train Engineer、Product Manager、Business Owner、System Architect、Release Management、Portfolio Manager 等多层角色,本质是 Machine Bureaucracy 的敏捷包装
- 客户缺席:SAFe 的经典框架图中客户处于次要位置,违反了敏捷宣言的 "customer collaboration over contract negotiation"
Maarten Dalmijn 在 mdalmijn.com 的尖锐批评:"did you ever hear the story of the self-organizing Scrum Team who decided to adopt SAFe? Me neither. No self-organizing team in their right mind would ever cripple themselves by adopting SAFe."[mdalmijn.com/p/safe-is-a-marketing-framework-not-an-agile-scaling-framework]
SAFe 流行的真正原因:它让企业管理层能同时保留现有的集中化控制结构又能声称自己敏捷。它满足了企业的政治需求而不是工程需求。这是一个现代的 Mission Command 失败案例——用程序语言取代了意图语言。
4.5 Brooks 定律的现代验证
Fred Brooks 1975 年《The Mythical Man-Month》的中心定律:Adding manpower to a late software project makes it later.
Brooks 的推理逻辑:
- 软件任务有不可分割的顺序约束
- 新人需要训练时间,训练消耗老人的生产力
- 沟通通道数量随 n(n-1)/2 增长
- 所以在某个点之后,增加人数的净贡献是负的
Brooks 自己的公式:「The number of months of a project depends upon its sequential constraints. The maximum number of men depends upon the number of independent subtasks.」[https://en.wikipedia.org/wiki/The_Mythical_Man-Month]
Eric Raymond 在《The Cathedral and the Bazaar》中对 Brooks 定律提出了部分反驳:开源项目如 Linux 能有效地接受大量贡献者,这在 Brooks 定律的前提下是不可能的。Raymond 的解释是:开源项目的贡献通常是高度解耦的(bug 修复、小特性),不会触发 Brooks 定律中的沟通爆炸。[https://effectiviology.com/brooks-law/]
这个反例本身强化了信息隐藏的价值:Linux 能扩展是因为 Linus 的仁慈独裁者架构加上模块化代码,让贡献者之间不需要相互协调。Agent 编排的启示直接:只有在模块化良好的前提下,加更多 sub-agent 才能线性扩展产出,否则越加越慢。
五、跨域同构分析(专题)
从以上材料中,可以提炼出五类贯穿军事、组织、认知、软件的共同模式。这些模式既解释了为什么成功实践彼此相似,也预测了 agent 编排的失败模式。
5.1 Parnas Information Hiding 作为万能原则
David Parnas 1972 年《On the Criteria to Be Used in Decomposing Systems into Modules》是软件工程史上最被引用的分解论文之一。Parnas 的核心主张:
We propose instead that one begins with a list of difficult design decisions or design decisions which are likely to change. Each module is then designed to hide such a decision from the others.
[https://www.riverandsoftware.com/p/criteria-to-be-used-in-modularisation-paper]
翻译:我们建议从一组难以决定或可能变化的设计决策开始。每个模块被设计为把这类决策隐藏于其他模块之外。
Parnas 提出两种分解 KWIC(Key Word In Context)索引系统的方法:
- Decomposition 1:按流程步骤分(read, shift, sort, output)
- Decomposition 2:按信息隐藏分(每个模块藏一个 likely-to-change 的设计决策)
两种分解产出相同功能,但 Decomposition 2 对需求变化更稳定,因为变化通常发生在一个模块内部,不会级联到其他模块。
这个原则的跨域同构:
| 领域 | 信息隐藏的体现 |
|---|---|
| 拿破仑军团 | 军团指挥官决定走哪条路、何时交战,上级不干预 |
| Auftragstaktik | 下级对「怎么做」的决策对上级透明但不被上级接管 |
| Mintzberg Professional Bureaucracy | 专业技能的标准化让操作核心不需要被管理层了解内部 |
| Two-pizza team | Single-threaded ownership:团队隐藏自己的内部架构,对外只暴露 API |
| HTA | Plan 层说明触发条件,不暴露 sub-operation 的实现细节 |
| Distributed cognition | 每个岗位只知道自己那部分,通过工具和协议协调 |
失败案例的共同反模式:
- Healthcare.gov:55 个承包商之间没有明确的信息边界,一个需求变化波及整个集成
- NHS NPfIT:全国统一软件强迫所有 Trust 暴露内部流程
- Matrix org:双重汇报链暴露了员工的任务到两套政治结构
- SAFe:规定的程序语言强迫团队暴露执行细节到程序层面
Claim for agent orchestration:每个 sub-agent 应该隐藏自己的工具调用、推理链、中间状态,只对主 agent 暴露任务成果和不确定性。任何把 sub-agent 中间状态显式暴露给其他 sub-agent 的设计,都在制造沟通通道爆炸。
5.2 Span of Control 的 3-7 数字在跨域一致
| 领域 | Span of Control 典型范围 | 来源 |
|---|---|---|
| 美军(步兵分队) | 2-6 | Hamilton 1921, Fivecoat |
| 美军(营/旅/高层) | 7-10 | Fivecoat |
| 美国海军指导文件 | 3-7 | OPNAVINST 3120.32D |
| Amazon Two-pizza | 5-10 | Bezos rule |
| Spotify Squad | 8-10 | Kniberg 2012 |
| 认知容量(Miller) | 7±2 | Miller 1956 |
| 认知容量(Cowan) | 4±1 | Cowan 2001 |
| Graicunas 管理公式极限 | 4-6(关系数 < 100) | Graicunas 1933 |
| 社会学(Dunbar inner circles) | 5 (support clique), 15 (sympathy group) | Dunbar 1993 |
这个数字的跨域一致性不是巧合。背后的三重锚点:
- 工作记忆容量:Miller 7±2 或 Cowan 4 chunks
- 关系复杂度:n(n-1)/2 的组合爆炸
- 注意力经济:人类总时间有限,每个关系都需要维护
对 agent 编排的启示:一个 agent(不管是 Claude Opus 还是用户本人)同时管理的 sub-agent 数应该保持在 3-7,接近上限时应引入中层协调者(middle management),而不是继续线性增加。Dunbar 数字 150 作为部落规模上限,对应的是单个 agent 系统能有效协调的 sub-agent 总量不超过约 150,超过则需要分部(tribes/divisions)。
5.3 Mission Command / Auftragstaktik 与 Agent 编排的结构同构
Moltke 1869 年原则 → Agent 编排映射:
| Moltke 原则 | Agent 编排对应 |
|---|---|
| 命令只覆盖可预测范围 | Task 指派不应过度规定执行路径 |
| 职位越高命令越短越宽泛 | Main agent 的 prompt 不应具体到每步工具调用 |
| 下一级补充必要精确度 | Sub-agent 在自己的 context 里展开具体策略 |
| 每一级保有行动自由 | Sub-agent 不应被实时干预 |
| 基于共同训练建立信任 | Sub-agent 必须共享 rules/SOUL.md 等根基 |
Spotify 失败的根因:autonomy without alignment。这个模式对 agent 编排的警示是:给 sub-agent 自由度前,必须先建立 alignment 机制。如果不读取 rules/ 文件就派出 sub-agent,等于在给一个陌生士兵发战场自主权。
ADP 6-0 的七条原则中,有四条直接对应 agent 编排的成功条件:
- Competence → sub-agent 的模型能力必须匹配任务复杂度(Haiku 不适合做架构决策)
- Shared understanding → sub-agent 必须读过 workspace rules
- Commander's intent → 任务 prompt 应明确 why 和 desired end state
- Mission orders → 任务 prompt 应避免过度规定 how
5.4 人类分解失败的共同模式
从调研的五大失败案例(Spotify、Healthcare.gov、NHS NPfIT、SAFe、瀑布模型误读)中提炼的共同反模式:
模式 F1:过程分解冒充结构分解
- 瀑布模型:按时间阶段分解被误读为按模块分解
- Royce 自己在原论文 Figure 1 就是按流程(分析→编码→测试)分解的 Decomposition 1
- Parnas 的一生工作都在反对这种做法
模式 F2:缺失 Commander's Intent
- Healthcare.gov 的 55 个承包商没有共同的 desired end state
- NHS NPfIT 的中心化供应商之间缺乏意图层面的对齐
- Spotify 的 squad 各自优化本地目标,丢失全局目标
模式 F3:Autonomy 与 Alignment 失衡
- Spotify 把白皮书的「Aligned Autonomy」砍半
- Matrix org 给员工多重自由度但没对应的协调结构
- SAFe 相反,alignment 过度而 autonomy 被稀释
模式 F4:Span of Control 被突破
- Healthcare.gov 的 CMS 直接协调 55 个承包商
- NPfIT 项目总监一人承担了本应分层的信任结构
- Spotify 的 Chapter Lead 跨 squad 管理让信任链断裂
模式 F5:信息隐藏边界被暴露
- Matrix dual reporting 让员工的任务对多方透明
- NPfIT 的全国统一软件让每个 Trust 暴露内部流程
- Healthcare.gov 的屏幕流把五个联邦机构的协调暴露给终端用户
对 agent 编排的预测:这些失败模式在 agent 系统中完全同构,而且更快显现。F2(缺 commander's intent)直接对应「未给 sub-agent 说明为什么做」;F3(失衡)对应给 sub-agent 大 prompt 但没给 rules 文件;F4 对应主 agent 直接管太多 sub-agent 不引入中层;F5 对应让多个 sub-agent 读取彼此的中间工作文件。
5.5 成功分解的共同结构
从 Napoleonic corps、Auftragstaktik、Two-pizza、HTA、Parnas 中可以抽象出四个共同的成功条件:
- 意图透明,细节隐藏:上层明确 what + why,下层自主决定 how
- span of control 受认知容量约束:3-7 范围内保持决策者有效
- 标准化的接口而非过程:协调机制是 API/protocol 而非工作流程
- 冗余但不浪费的拓扑:军团之间互相支援,但不重复执行
这四条构成了任何跨域任务分解系统的必要条件,缺一即触发某类失败模式。
六、Claim 验证状态表
| Claim | 证据强度 | 核心来源 |
|---|---|---|
| Moltke 本人从未使用 Auftragstaktik 一词 | 强 | DTIC ADA569668;维基百科 Mission-type tactics |
| ADP 6-0 七条原则(2019 版) | 强 | armypubs.army.mil ARN34403 原文 |
| Napoleon corps 制是「centralized control + decentralized ops」 | 强 | Calhoun NPS;eARMOR 双重确认 |
| Span of control 3-7 跨域一致 | 强 | Hamilton 1921;OPNAVINST 3120.32D;Fivecoat;Graicunas |
| Bezos two-pizza 规则具体在 5-10 人范围 | 强 | AWS 白皮书;Nuclino;Buffer 多源 |
| Spotify 2020 官方从未正式用过完整模型 | 强 | Jeremiah Lee 原文;Anders Ivarsson 证言 |
| Royce 1970 论文原文有「risky and invites failure」警告 | 强 | 原论文 PDF (Praxis Framework) 第 2 页 |
| Healthcare.gov 有超过 55 家承包商 | 强 | GAO 报告;Forbes 2013;PM360 |
| NHS NPfIT 总成本超 100 亿英镑 | 强 | BBC 2013;PAC 委员会报告 |
| Parnas 1972 提出信息隐藏作为分解标准 | 强 | CACM Vol 15 No 12 原刊;Princeton 历史课程 |
| Brooks 定律(增加人手延长进度)得到广泛经验支持 | 强 | 原书 1975;现代多篇验证 |
| Hutchins distributed cognition 在 Palau 军舰的实地研究 | 强 | MIT Press 出版的《Cognition in the Wild》 |
| Chase-Simon chunking 估计大师有 50,000 chunks | 中 | Simon-Gilmartin 1973 原始估计;后续修订争议 |
| Sweller germane load 2010 年被重新定义为非独立来源 | 强 | Educational Psychology Review Vol 22 原文 |
| Dunbar 150 数字在不同文化中稳定 | 中 | Dunbar 原研究;Bayesian 重新分析有 69-109 范围的争议 |
七、URL 索引(完整 25 个)
军事指挥与 Auftragstaktik
- https://apps.dtic.mil/sti/tr/pdf/ADA569668.pdf — DTIC, Auftragstaktik: The Basis for Modern Military Command
- https://apps.dtic.mil/sti/pdfs/ADA599111.pdf — DTIC, Initiative Within the Philosophy of Auftragstaktik
- https://apps.dtic.mil/sti/tr/pdf/ADA563054.pdf — DTIC, Moltke's Mission Command Philosophy in the Twenty-First Century
- https://en.wikipedia.org/wiki/Mission-type_tactics — Wikipedia, Mission-type tactics
- https://ndupress.ndu.edu/Media/News/News-Article-View/Article/2076032/beyond-auftragstaktik-the-case-against-hyper-decentralized-command/ — NDU Press, Beyond Auftragstaktik
- https://armypubs.army.mil/epubs/DR_pubs/DR_a/ARN34403-ADP_6-0-000-WEB-3.pdf — U.S. Army, ADP 6-0 Mission Command (2019)
- https://apps.dtic.mil/sti/tr/pdf/AD1001514.pdf — DTIC, Historical Roots of Mission Command
拿破仑与战略思想
- https://calhoun.nps.edu/server/api/core/bitstreams/60f29c3d-a116-4d83-863b-6149bee61da7/content — Naval Postgraduate School, Command and Control of the Grand Armée
- https://saberandscroll.scholasticahq.com/article/28520-napoleon-apex-of-the-military-revolution.pdf — Saber and Scroll Journal, Napoleon Apex of Military Revolution
- https://www.benning.army.mil/Armor/eARMOR/content/issues/2014/MAR_JUN/Chavous.html — eARMOR, Napoleon Bonaparte Contributions to Modern Warfare
- https://www.thecollector.com/how-napoleon-bonaparte-build-greatest-army/ — The Collector, How Napoleon Built the Greatest Army
OODA Loop 与 Boyd
- https://en.wikipedia.org/wiki/OODA_loop — Wikipedia, OODA loop
- https://www.colonelboyd.com/boydswork — John Boyd Homepage, Organic Design for Command and Control
- https://teamonenetwork.com/wp-content/uploads/2019/03/COMING-FULL-CIRCLE-WITH-BOYD%E2%80%99S-OODA-LOOP-IDEAS.pdf — Team One Network, Coming Full Circle with Boyd
Span of Control
- https://www.thefivecoatconsultinggroup.com/the-coronavirus-crisis/span-of-control — Fivecoat Consulting, Span of Control with Army examples
- https://www.secnav.navy.mil/doni/Directives/03000%20Naval%20Operations%20and%20Readiness/03-100%20Naval%20Operations%20Support/3120.32D%20W%20CH-1.pdf — U.S. Navy, OPNAVINST 3120.32D
- https://cgsc.contentdm.oclc.org/digital/api/collection/p4013coll3/id/1747/download — CGSC, Span of Control and the Operational Commander
组织理论
- https://umbrex.com/resources/frameworks/organization-frameworks/mintzberg-organizational-configurations/ — Umbrex, Mintzberg Organizational Configurations
- https://blog.nuclino.com/two-pizza-teams-the-science-behind-jeff-bezos-rule — Nuclino, Science behind Two-Pizza teams
- https://d1.awsstatic.com/executive-insights/en_US/two_pizza_teams_eBook.pdf — AWS, Powering Innovation with Two-Pizza Teams
- https://www.jeremiahlee.com/posts/failed-squad-goals/ — Jeremiah Lee, Spotify's Failed SquadGoals
- https://theleadershipedge.com/wp-content/uploads/2025/06/Challenges_Strategies_of_Matrix_Orgs.pdf — Leadership Edge, Challenges of Matrix Organizations
认知科学
- https://link.springer.com/article/10.1007/s10648-010-9128-5 — Sweller 2010, Element Interactivity and Cognitive Load
- https://pubmed.ncbi.nlm.nih.gov/9709441/ — PubMed, Expert Chess Memory Revisiting Chunking Hypothesis
- https://direct.mit.edu/books/monograph/4892/Cognition-in-the-Wild — MIT Press, Hutchins, Cognition in the Wild
- https://www.humanreliability.com/human-factors/hierarchical-task-analysis-hta/ — Human Reliability, What is HTA
失败案例
- https://www.praxisframework.org/files/royce1970.pdf — Royce 1970 原论文
- https://pmworldlibrary.net/wp-content/uploads/2018/07/pmwj72-Jul2018-Morgan-applying-1970-waterfall-lessons-umd-paper.pdf — Morgan 2018 重读
- http://greatquestion.com.au/wp-content/uploads/2023/09/Anthopoulos-et-al-Why-e-government-projects-fail-An-analysis-of-the-healthcare.gov-website-2016-1.pdf — Anthopoulos et al, Healthcare.gov 分析
- https://www.forbes.com/sites/lorenthompson/2013/12/03/healthcare-gov-diagnosis-the-government-broke-every-rule-of-project-management/ — Forbes, Healthcare.gov 诊断
- https://www.henricodolfing.ch/case-study-1-the-10-billion-it-disaster-at-the-nhs/ — Dolfing, NHS NPfIT 10B Disaster
- https://www.cl.cam.ac.uk/archive/rja14/Papers/npfit-mpp-2014-case-history.pdf — Anderson et al 2014, NPfIT Case History
- https://www.forbes.com/sites/stevedenning/2019/05/23/understanding-fake-agile/ — Steve Denning, Understanding Fake Agile
分解原理
- https://www.riverandsoftware.com/p/criteria-to-be-used-in-modularisation-paper — Parnas 1972 原文解读
- https://en.wikipedia.org/wiki/The_Mythical_Man-Month — Wikipedia, The Mythical Man-Month
- https://effectiviology.com/brooks-law/ — Effectiviology, Brooks' Law 深度解析
八、交付清单
8.1 关键词列表(便于后续检索)
- Auftragstaktik / Mission Command / Befehlstaktik
- Commander's intent / Mission orders
- Moltke the Elder / Moltke 1869 Instructions
- Boyd / OODA loop / Organic Design / Schwerpunkt
- Napoleonic corps d'armée / Berthier staff
- Span of control / Hamilton 1921 / Graicunas
- Mintzberg configurations / Machine Bureaucracy / Adhocracy
- Spotify model / Squad / Tribe / Chapter / Guild / Jeremiah Lee
- Two-pizza team / Ringelmann effect / single-threaded ownership
- Matrix organization / dual reporting / matrix guardian
- Coase transaction cost / Williamson / asset specificity
- Parnas information hiding / Decomposition 1 vs 2
- Cognitive Load Theory / intrinsic / extraneous / germane
- Chase-Simon chunking / chess expertise / 50000 chunks
- Distributed cognition / Hutchins / Palau / Cognition in the Wild
- Hierarchical Task Analysis / HTA / plan / P×C stopping rule
- Waterfall / Royce 1970 / risky and invites failure
- Healthcare.gov / 55 contractors / CMS / GAO report
- NHS NPfIT / Connecting for Health / Richard Granger / £10B
- SAFe / Scaled Agile Framework / fake agile / Denning
- Brooks law / Mythical Man-Month / n(n-1)/2 communication
- Dunbar number / 150 / inner circles / neocortex ratio
8.2 对 Agent 编排最有启发的 3 个原则
原则一:Commander's Intent 是唯一可跨层级传递的东西
从 Moltke 到 ADP 6-0 到 Amazon two-pizza,所有成功分解的共同特征是高层只传达 what + why + desired end state,不传达 how。执行细节必然在下层动态确定,因为只有下层掌握本地信息。对 agent 编排的直接映射:给 sub-agent 的 prompt 应该用 2-3 句话说清楚任务目标、成功标准、约束边界,剩下的交给 sub-agent 的自主工具调用决定。把 prompt 写成一个 50 步 checklist 等于把 Mission Command 退化回 Jominian centralization,同时损失速度和适应性。
原则二:Span of Control 的 3-7 数字是认知硬约束,不可违反
军事、管理、认知科学、工作记忆研究四条独立证据链都指向 3-7 这个数字。超出这个范围时必然需要引入中层协调(tribe、brigade、division)。对 agent 编排的直接映射:一个 orchestrator agent 应该同时管理不超过 7 个 worker sub-agent,超过这个数字的任务必须先做二级分解——不是线性增加 sub-agent,而是引入 middle-manager agent 做二层调度。Healthcare.gov 的 55 承包商和 NPfIT 的集中供应商网络都是这一原则的失败案例。
原则三:Information Hiding 是所有成功分解的底层结构
Parnas 1972 年的观察:按流程分解(decomposition 1)脆弱,按「隐藏可能变化的决策」分解(decomposition 2)稳健。军团制、Auftragstaktik、Two-pizza team、HTA 都在不同层面实现了同样原则:每个单元隐藏自己的内部实现,对外暴露稳定的意图/接口/成果。反例是 Matrix org 的 dual reporting(员工内部工作暴露于两条管理线)、Healthcare.gov 的跨承包商集成(每个模块的内部假设都暴露给其他模块)。对 agent 编排的直接映射:sub-agent 的中间思考、工具调用顺序、内部尝试不应被主 agent 或其他 sub-agent 窥视,只应传递最终结论和未决不确定性。否则沟通通道爆炸会触发 Brooks 定律。
8.3 Span of Control 在不同领域的数字对比
| 来源 | 数字范围 | 备注 |
|---|---|---|
| Hamilton 1921(英军指挥官) | 3-6 | span of control 概念源头 |
| Graicunas 1933 数学分析 | 4-6 | 基于 n(n-1)/2 公式,关系数 < 100 |
| Miller 1956(工作记忆) | 7±2 | chunk 数量上限 |
| Cowan 2001(重新估计工作记忆) | 4±1 | 在无排练条件下的真实值 |
| U.S. Navy OPNAVINST 3120.32D | 3-7 | 正式军规定义 |
| 美军步兵分队(squad) | 2 | 班长直接管 2 个副手 |
| 美军步兵排(platoon) | 6 | 排长管 4 班长 + 排军士 + 火力支援士官 |
| 美军步兵连(company) | 7 | 连长管 3 排长 + 首席士官 + 副连长 + 火力支援官 |
| 美军步兵营(battalion) | 9 | 营长管 6 连长 + 营军士 + 作训官 + 副营长 |
| 美军步兵旅(brigade) | 10 | 旅长管 7 营长 + 旅军士 + 作训官 + 副旅长 |
| Amazon two-pizza team | 5-10 | 披萨能喂饱的上限 |
| Spotify squad | 8-10 | 2012 白皮书规定 |
| Dunbar inner circle | 5 | support clique |
| Dunbar sympathy group | 15 | 可维持亲密关系 |
| Dunbar active network | 50 | 偶尔互动 |
| Dunbar 总数 | 150 | 认识且能维持稳定关系的上限 |
| 现代企业组织建议(ERC) | 4-5 直接汇报(高层);6-7(中层);8-15(基层) | 按职位性质分层 |
| 典型 Agile team | 5-9 | Scrum 指南的 dev team size |
模式观察:3-7 出现在基础作战单位(需高协调)、工作记忆(同时追踪)、海军规定(制度化)。7-10 出现在更高层级(协调由中层做大量了)。Dunbar 15 是亲密关系上限,150 是认知网络上限。跨领域一致的范围 3-7 正好对应单个指挥者同时实时跟踪的硬上限。
8.4 推荐 Round 3 扩张方向
方向 A:失败模式的量化研究
本次调研主要是定性案例。Round 3 可深入的是:Standish Group 的 CHAOS Report 或 McKinsey-Oxford 2012 年大型 IT 项目研究,提取失败率随项目规模的量化关系,验证 Brooks 定律在组织层的形式。具体数据点:200M+ 预算的政府 IT 项目失败率 vs 10M 预算的失败率。
方向 B:Auftragstaktik 在非军事场景的具体应用
Stanley McChrystal 在 JSOC 如何把 Mission Command 升级到反恐作战(Team of Teams 2015)。Ray Dalio 的 Bridgewater 的 "radical transparency" 是否是 Mission Command 的变体。Toyota Production System 中的 andon cord 是现场决策权前移的东方版 Auftragstaktik。这些对 agent 编排的直接映射价值高。
方向 C:Span of Control 在 AI / Agent 系统中的实证
本次调研建立了跨域 3-7 的普适性,但 Round 3 可以进入 AI agent 层面:GPT-4 / Claude / Gemini 在多 agent 协调场景下的有效 span 是多少?当一个 orchestrator agent 要管理 5、7、10、15 个 worker agent 时,成功率/幻觉率/输出质量的变化曲线。这个研究会首次把人类 span of control 数字迁移到 AI 系统并得到独立验证。
方向 D:Parnas Information Hiding 在 LLM 时代的变奏
传统 Information Hiding 假设模块边界稳定。但 LLM 的能力决定了一个 agent 可以动态「学习」另一个 agent 的内部实现(通过 few-shot examples)。这对 information hiding 构成根本挑战:是否应该在 agent 系统中引入比 Parnas 更强的隔离机制?是否需要「encrypted intent」这种新概念?
方向 E:东亚军事与组织学传统的对照
本次调研主要来自西方。中国明清八旗制度、日本战国家臣团、毛泽东「十大军事原则」中的「集中优势兵力各个歼灭」,这些传统在任务分解上的独特贡献值得发掘。特别是日本 Toyota 的 hoshin kanri(方针管理)是 Auftragstaktik 与东方层级文化的融合,对理解普适 vs 文化特定很有价值。
结语
人类用了两个世纪打磨从拿破仑军团到 two-pizza team 的分解方法论,核心共识相当少:意图要明确但细节要留白;span 要受认知约束;边界要隐藏实现;冗余但不浪费的拓扑。失败也惊人一致:指令过度规定、span 突破上限、信息隐藏边界被暴露、autonomy 和 alignment 失衡。这些原则对 agent 编排不是类比,它们是结构同构——同样的信息论约束和同样的认知容量限制,驱动着人类和 AI 在多主体协调中的边界。
对「等天黑」的用户画像而言,本调研最该带走的一件东西是:Auftragstaktik 的成功依赖军官共享教育背景,没有共享教育的 Mission Command 会退化为放任。这直接对应你已经发现的 sub-agent 在缺少 rules 文件注入时的 40% 合规率问题。解法不是写更详细的任务 prompt,而是确保每个 sub-agent 先读 SOUL.md / USER.md / COMMUNICATION.md 建立 shared understanding,就像普鲁士军官在战争学院里建立的共同战术语言。这是普鲁士传统最值得移植的一面。