⚠️ run_id 身份歧义(5 个)
run_id 是 run 的身份。同一个 run_id 在磁盘上有不止一份时,
"引用该 run 的评价"就不再指向单一对象。本站按确定性规则取了其中一份,但这不解决歧义本身——
需要有人决定保留哪一条历史(或给其中一条加后缀重命名)。
RUN-20260912-9105 存在 2 份:
采用 runs/_baseline-main/RUN-20260912-9105(34 条事件),另有 runs/RUN-20260912-9105(34 条事件)。事件类型与序号一致但事件内容不同:这是同一个 run_id 下的两条历史,引用该 run 的评价存在歧义。RUN-20260912-9104 存在 2 份:
采用 runs/_baseline-main/RUN-20260912-9104(37 条事件),另有 runs/RUN-20260912-9104(37 条事件)。事件类型与序号一致但事件内容不同:这是同一个 run_id 下的两条历史,引用该 run 的评价存在歧义。RUN-20260912-9103 存在 2 份:
采用 runs/_baseline-main/RUN-20260912-9103(14 条事件),另有 runs/RUN-20260912-9103(14 条事件)。事件类型与序号一致但事件内容不同:这是同一个 run_id 下的两条历史,引用该 run 的评价存在歧义。RUN-20260912-9102 存在 2 份:
采用 runs/_baseline-main/RUN-20260912-9102(13 条事件),另有 runs/RUN-20260912-9102(13 条事件)。事件类型与序号一致但事件内容不同:这是同一个 run_id 下的两条历史,引用该 run 的评价存在歧义。RUN-20260912-9101 存在 2 份:
采用 runs/_baseline-main/RUN-20260912-9101(20 条事件),另有 runs/RUN-20260912-9101(20 条事件)。事件类型与序号一致但事件内容不同:这是同一个 run_id 下的两条历史,引用该 run 的评价存在歧义。
Stage 1 Baseline Report
这是时点快照:报告绑定生成时刻的代码版本与一批固定 run。之后新增的 run 不会自动改写它——
需要重跑 npm run baseline 才会更新。因此"限制条目里的样本量"指的是生成当时的样本,不是当前全部。
- 总体结论
- NO_GO
- 报告指纹
f4d413690aea5990
- 绑定代码版本
59365184fbe28d216e1f5a95382225b9ff90bb5a
- 生成时间
- 2026-09-17T13:36:17.598Z
- EvalRecord
- 9 条|P0/P1 证据完整:是|裁决:0 条
Gate 判定
| Gate | 项 | 结论 | 理由 |
|---|
| G1-1 | Event Integrity | PASS | 无序列重复/倒退/断点;核心事件字段完整 |
| G1-2 | Replay | PASS | 任一 fixture 均可重建 Director→Role→PublicTurn→StateUpdate 链,并解释版本 |
| G1-3 | Eval Traceability | PASS | P0/P1 结论 100% 指向具体证据;无结论性记录缺少 target_refs |
| G1-4 | Gold / Adjudication | BLOCKED | gold_failures 尚无 confirmed 锚点(draft 3,rejected 0);gold_good_snippets 尚无 confirmed 锚点(draft 2,rejected 0);裁决流程尚未被真实使用(无 AdjudicationRecord) |
| G1-5 | Independence | PASS | 无外部项目路径 / 数据库 / 服务依赖 |
| G1-6 | Baseline Report | PASS | 报告明确区分"已工程验证"与"尚待 Stage 2 验证",未把未验证能力写成已通过 |
Failure Coverage(F01–F17)
| Code | 名称 | 自动检查 | 检测层级 |
|---|
F01 | ROLE_DRIFT | — | human_only |
F02 | ECHOING | — | auto_candidate |
F03 | FORCED_DISAGREEMENT | — | human_only |
F04 | TOPIC_DRIFT | AC-07 | auto_candidate |
F05 | UNSUPPORTED_FACT | AC-02 | auto_candidate |
F06 | THEORY_DUMP | — | human_only |
F07 | MECHANICAL_TURN_TAKING | AC-03 | auto_candidate |
F08 | ENDURANCE_COLLAPSE | — | stage2 |
F09 | ROLE_LEAKAGE | — | human_only |
F10 | ACTION_VAGUENESS | AC-06 | auto_candidate |
F11 | CONDITION_INSENSITIVITY | — | stage2 |
F12 | NO_SELF_REVISION | — | human_only |
F13 | CASE_ANSWER_LEAKAGE | AC-08 | auto_candidate |
F14 | PROFESSIONAL_SAFETY | AC-09 | auto_candidate |
F15 | DIRECTOR_REASON_MISMATCH | AC-04 | auto_candidate |
F16 | RAG_HOMOGENIZATION | AC-05 | auto_candidate |
F17 | REPLAY_INTEGRITY | AC-01 | auto_final |
跨 run 分析
按数据来源分组:mock 与真实模型混在一起会被判成 mixed(跨 run 比较无效),分组后每组才可解读。
真实模型(5 期 / 34 轮)
全部 run 使用真实模型(openai-compatible:openai/gpt-5.5#strong, openai-compatible:openai/gpt-5.5#fast)
结构性问题:0
| Code | 复现 run 数 | 出现在 |
|---|
F15 | 2/5 | RUN-20260912-9106-8t1, RUN-20260912-9106-8t3 |
F10 | 1/5 | RUN-20260912-9106-1789403245_run2 |
角色发言占比
| 角色 | 出现 run 数 | 平均 | 最小 | 最大 |
|---|
| GAME-R01 | 5/5 | 59.0% | 50.0% | 62.5% |
| GAME-R04 | 5/5 | 33.5% | 25.0% | 40.0% |
| GAME-R05 | 3/5 | 7.5% | 0.0% | 12.5% |
后半段增益(Endurance 信号)
| run_id | 后半段增量类型 | 仍有新增价值 |
|---|
| RUN-20260912-9106-8t3 | evidence, evidence, condition | 是 |
| RUN-20260912-9106-8t2 | counterexample, condition, condition | 是 |
| RUN-20260912-9106-8t1 | evidence, condition, condition | 是 |
| RUN-20260912-9106-1789403245_run2 | action, condition | 是 |
| RUN-20260912-9106-1789403017007 | action, condition | 是 |
观察(候选模式,需人工确认)
mock / fixture(5 期 / 16 轮)
全部 run 使用 mock / fixture 网关(fixture-gateway-problem_text#strong, fixture-gateway-problem_text#fast, fixture-gateway-retrieval#strong, fixture-gateway-retrieval#fast, fixture-gateway-chain#strong, fixture-gateway-chain#fast, fixture-gateway-fail_on_second_role#strong, fixture-gateway-fail_on_second_role#fast):结论仅能验证 harness,不得写入 Stage 2 判定
结构性问题:0
| Code | 复现 run 数 | 出现在 |
|---|
F10 | 1/5 | RUN-20260912-9105 |
F15 | 1/5 | RUN-20260912-9105 |
F16 | 1/5 | RUN-20260912-9104 |
角色发言占比
| 角色 | 出现 run 数 | 平均 | 最小 | 最大 |
|---|
| GAME-R01 | 5/5 | 84.0% | 60.0% | 100.0% |
| GAME-R04 | 2/5 | 8.0% | 0.0% | 20.0% |
| GAME-R05 | 2/5 | 8.0% | 0.0% | 20.0% |
后半段增益(Endurance 信号)
| run_id | 后半段增量类型 | 仍有新增价值 |
|---|
| RUN-20260912-9105 | none, none | 否 |
| RUN-20260912-9104 | condition, condition | 是 |
| RUN-20260912-9103 | interpretation | 是 |
| RUN-20260912-9102 | evidence | 是 |
| RUN-20260912-9101 | condition | 是 |
观察(候选模式,需人工确认)
- 数据来源警告:全部 run 使用 mock / fixture 网关(fixture-gateway-problem_text#strong, fixture-gateway-problem_text#fast, fixture-gateway-retrieval#strong, fixture-gateway-retrieval#fast, fixture-gateway-chain#strong, fixture-gateway-chain#fast, fixture-gateway-fail_on_second_role#strong, fixture-gateway-fail_on_second_role#fast):结论仅能验证 harness,不得写入 Stage 2 判定。下面的"模式"里凡涉及生成行为的,都可能来自桩而非系统。
- RUN-20260912-9103:未写入 RUN_CLOSED(运行在结束前中断)
- RUN-20260912-9105:后半段只出现 none/none,未见新增价值(Endurance 风险)
- 1/5 条 run 有过半数轮次增量为 none(RUN-20260912-9105)——同义复述风险
全部(混合时不可比较)(10 期 / 50 轮)
本批混用了 mock 与真实模型(mock: fixture-gateway-problem_text#strong, fixture-gateway-problem_text#fast, fixture-gateway-retrieval#strong, fixture-gateway-retrieval#fast, fixture-gateway-chain#strong, fixture-gateway-chain#fast, fixture-gateway-fail_on_second_role#strong, fixture-gateway-fail_on_second_role#fast)——跨 run 比较无效,请拆分后重跑分析
结构性问题:0
| Code | 复现 run 数 | 出现在 |
|---|
F15 | 3/10 | RUN-20260912-9105, RUN-20260912-9106-8t1, RUN-20260912-9106-8t3 |
F10 | 2/10 | RUN-20260912-9105, RUN-20260912-9106-1789403245_run2 |
F16 | 1/10 | RUN-20260912-9104 |
角色发言占比
| 角色 | 出现 run 数 | 平均 | 最小 | 最大 |
|---|
| GAME-R01 | 10/10 | 71.5% | 50.0% | 100.0% |
| GAME-R04 | 7/10 | 20.7% | 0.0% | 40.0% |
| GAME-R05 | 5/10 | 7.7% | 0.0% | 20.0% |
后半段增益(Endurance 信号)
| run_id | 后半段增量类型 | 仍有新增价值 |
|---|
| RUN-20260912-9106-8t3 | evidence, evidence, condition | 是 |
| RUN-20260912-9106-8t2 | counterexample, condition, condition | 是 |
| RUN-20260912-9106-8t1 | evidence, condition, condition | 是 |
| RUN-20260912-9106-1789403245_run2 | action, condition | 是 |
| RUN-20260912-9106-1789403017007 | action, condition | 是 |
| RUN-20260912-9105 | none, none | 否 |
| RUN-20260912-9104 | condition, condition | 是 |
| RUN-20260912-9103 | interpretation | 是 |
| RUN-20260912-9102 | evidence | 是 |
| RUN-20260912-9101 | condition | 是 |
观察(候选模式,需人工确认)
- 数据来源警告:本批混用了 mock 与真实模型(mock: fixture-gateway-problem_text#strong, fixture-gateway-problem_text#fast, fixture-gateway-retrieval#strong, fixture-gateway-retrieval#fast, fixture-gateway-chain#strong, fixture-gateway-chain#fast, fixture-gateway-fail_on_second_role#strong, fixture-gateway-fail_on_second_role#fast)——跨 run 比较无效,请拆分后重跑分析。下面的"模式"里凡涉及生成行为的,都可能来自桩而非系统。
- RUN-20260912-9103:未写入 RUN_CLOSED(运行在结束前中断)
- RUN-20260912-9105:后半段只出现 none/none,未见新增价值(Endurance 风险)
- 1/10 条 run 有过半数轮次增量为 none(RUN-20260912-9105)——同义复述风险
Gold Eval Set 状态
| 集合 | 总量 | confirmed | draft | rejected |
|---|
| gold_failures | 3 | 0 | 3 | 0 |
| gold_good_snippets | 2 | 0 | 2 | 0 |
| gold_borderline | 1 | 0 | 1 | 0 |
| gold_replay_cases | 2 | 0 | 2 | 0 |
G1-4 要求 gold_failures 与 gold_good_snippets 各至少 1 条 confirmed,且裁决流程被真实使用。当前 confirmed 合计 0。