三角色 AI 专业谈话系统

⚠️ run_id 身份歧义(5 个)

run_id 是 run 的身份。同一个 run_id 在磁盘上有不止一份时, "引用该 run 的评价"就不再指向单一对象。本站按确定性规则取了其中一份,但这不解决歧义本身—— 需要有人决定保留哪一条历史(或给其中一条加后缀重命名)。

Stage 1 Baseline Report

这是时点快照:报告绑定生成时刻的代码版本与一批固定 run。之后新增的 run 不会自动改写它—— 需要重跑 npm run baseline 才会更新。因此"限制条目里的样本量"指的是生成当时的样本,不是当前全部。

总体结论
NO_GO
报告指纹
f4d413690aea5990
绑定代码版本
59365184fbe28d216e1f5a95382225b9ff90bb5a
生成时间
2026-09-17T13:36:17.598Z
EvalRecord
9 条|P0/P1 证据完整:是|裁决:0 条

Gate 判定

Gate结论理由
G1-1Event IntegrityPASS无序列重复/倒退/断点;核心事件字段完整
G1-2ReplayPASS任一 fixture 均可重建 Director→Role→PublicTurn→StateUpdate 链,并解释版本
G1-3Eval TraceabilityPASSP0/P1 结论 100% 指向具体证据;无结论性记录缺少 target_refs
G1-4Gold / AdjudicationBLOCKEDgold_failures 尚无 confirmed 锚点(draft 3,rejected 0);gold_good_snippets 尚无 confirmed 锚点(draft 2,rejected 0);裁决流程尚未被真实使用(无 AdjudicationRecord)
G1-5IndependencePASS无外部项目路径 / 数据库 / 服务依赖
G1-6Baseline ReportPASS报告明确区分"已工程验证"与"尚待 Stage 2 验证",未把未验证能力写成已通过

Failure Coverage(F01–F17)

Code名称自动检查检测层级
F01ROLE_DRIFThuman_only
F02ECHOINGauto_candidate
F03FORCED_DISAGREEMENThuman_only
F04TOPIC_DRIFTAC-07auto_candidate
F05UNSUPPORTED_FACTAC-02auto_candidate
F06THEORY_DUMPhuman_only
F07MECHANICAL_TURN_TAKINGAC-03auto_candidate
F08ENDURANCE_COLLAPSEstage2
F09ROLE_LEAKAGEhuman_only
F10ACTION_VAGUENESSAC-06auto_candidate
F11CONDITION_INSENSITIVITYstage2
F12NO_SELF_REVISIONhuman_only
F13CASE_ANSWER_LEAKAGEAC-08auto_candidate
F14PROFESSIONAL_SAFETYAC-09auto_candidate
F15DIRECTOR_REASON_MISMATCHAC-04auto_candidate
F16RAG_HOMOGENIZATIONAC-05auto_candidate
F17REPLAY_INTEGRITYAC-01auto_final

跨 run 分析

按数据来源分组:mock 与真实模型混在一起会被判成 mixed(跨 run 比较无效),分组后每组才可解读。

真实模型(5 期 / 34 轮)

全部 run 使用真实模型(openai-compatible:openai/gpt-5.5#strong, openai-compatible:openai/gpt-5.5#fast)

结构性问题:0

Code复现 run 数出现在
F152/5RUN-20260912-9106-8t1, RUN-20260912-9106-8t3
F101/5RUN-20260912-9106-1789403245_run2
角色发言占比
角色出现 run 数平均最小最大
GAME-R015/559.0%50.0%62.5%
GAME-R045/533.5%25.0%40.0%
GAME-R053/57.5%0.0%12.5%
后半段增益(Endurance 信号)
run_id后半段增量类型仍有新增价值
RUN-20260912-9106-8t3evidence, evidence, condition
RUN-20260912-9106-8t2counterexample, condition, condition
RUN-20260912-9106-8t1evidence, condition, condition
RUN-20260912-9106-1789403245_run2action, condition
RUN-20260912-9106-1789403017007action, condition
观察(候选模式,需人工确认)

mock / fixture(5 期 / 16 轮)

全部 run 使用 mock / fixture 网关(fixture-gateway-problem_text#strong, fixture-gateway-problem_text#fast, fixture-gateway-retrieval#strong, fixture-gateway-retrieval#fast, fixture-gateway-chain#strong, fixture-gateway-chain#fast, fixture-gateway-fail_on_second_role#strong, fixture-gateway-fail_on_second_role#fast):结论仅能验证 harness,不得写入 Stage 2 判定

结构性问题:0

Code复现 run 数出现在
F101/5RUN-20260912-9105
F151/5RUN-20260912-9105
F161/5RUN-20260912-9104
角色发言占比
角色出现 run 数平均最小最大
GAME-R015/584.0%60.0%100.0%
GAME-R042/58.0%0.0%20.0%
GAME-R052/58.0%0.0%20.0%
后半段增益(Endurance 信号)
run_id后半段增量类型仍有新增价值
RUN-20260912-9105none, none
RUN-20260912-9104condition, condition
RUN-20260912-9103interpretation
RUN-20260912-9102evidence
RUN-20260912-9101condition
观察(候选模式,需人工确认)

全部(混合时不可比较)(10 期 / 50 轮)

本批混用了 mock 与真实模型(mock: fixture-gateway-problem_text#strong, fixture-gateway-problem_text#fast, fixture-gateway-retrieval#strong, fixture-gateway-retrieval#fast, fixture-gateway-chain#strong, fixture-gateway-chain#fast, fixture-gateway-fail_on_second_role#strong, fixture-gateway-fail_on_second_role#fast)——跨 run 比较无效,请拆分后重跑分析

结构性问题:0

Code复现 run 数出现在
F153/10RUN-20260912-9105, RUN-20260912-9106-8t1, RUN-20260912-9106-8t3
F102/10RUN-20260912-9105, RUN-20260912-9106-1789403245_run2
F161/10RUN-20260912-9104
角色发言占比
角色出现 run 数平均最小最大
GAME-R0110/1071.5%50.0%100.0%
GAME-R047/1020.7%0.0%40.0%
GAME-R055/107.7%0.0%20.0%
后半段增益(Endurance 信号)
run_id后半段增量类型仍有新增价值
RUN-20260912-9106-8t3evidence, evidence, condition
RUN-20260912-9106-8t2counterexample, condition, condition
RUN-20260912-9106-8t1evidence, condition, condition
RUN-20260912-9106-1789403245_run2action, condition
RUN-20260912-9106-1789403017007action, condition
RUN-20260912-9105none, none
RUN-20260912-9104condition, condition
RUN-20260912-9103interpretation
RUN-20260912-9102evidence
RUN-20260912-9101condition
观察(候选模式,需人工确认)

Gold Eval Set 状态

集合总量confirmeddraftrejected
gold_failures3030
gold_good_snippets2020
gold_borderline1010
gold_replay_cases2020

G1-4 要求 gold_failuresgold_good_snippets 各至少 1 条 confirmed,且裁决流程被真实使用。当前 confirmed 合计 0。