test(benchmark): pin the runtime-observation path into the study campaign block - #5255
Conversation
cocolord
left a comment
There was a problem hiding this comment.
APPROVE — 未发现阻塞项。exact head 2a9b5877a6c9aec856b276e5a186d0dafa6445c5 为此前没有覆盖的 runtime_observation → study dashboard campaign 投影补上了真实 reducer 路径测试;三项故障注入均能被新增 case 捕获,focused/related suite、架构 census、risk-based premerge 和远端 required checks 均通过。
动机
本 PR 补的是 benchmark study dashboard 的一个明确覆盖缺口:产品 reducer 已接受 runtime_observation upload envelope,并把 observation 数量与 typed classification 聚合到 campaign,但既有 suite 没有任何用例把这种 record kind 送进该路径。这样一来,即使 reducer 漏收记录、丢掉分类聚合,或者 dashboard payload 写出与 factorial list 不一致的摘要,原有 20 个 projection tests 仍可能保持绿色。
新增 case 把 RFC #5203 中 campaign projection 的 coverage slice 变成可执行约束。它不改变用户可见 payload、CLI、schema 或默认行为;对长期维护的改善是让 dashboard campaign 摘要与其 typed runtime/factorial authority 不再只有实现自证,而是具备可杀死实际回归的测试。这个边界本身完整、可独立回滚,后续示例 payload 是否修正仍可由其 owner 单独决策。
改动思路
测试复用现有生产入口:先由 build_benchmark_runtime_observation 从 admission、receipt 与 runner-owner facts 生成三种合法 typed observation,再用 build_benchmark_upload_envelope(..., record_kind="runtime_observation") 包装,与既有 board rows 和 insight 一起交给 build_benchmark_study_dashboard。权威规则仍全部留在原有 runtime enum、envelope normalizer 与 study reducer 中,测试没有复制第二套分类算法,也没有引入新的 state owner。
正向路径验证三条 observation 最终形成 runtime_observation_count == 3 以及三个 classification bucket;同时验证 bucket key 来自 BenchmarkRuntimeClassification 且总和等于 observation count。负向证据通过三次独立 mutation 覆盖 record-kind admission、runtime aggregation 和 factorial summary/list consistency:在 immutable base 上原有 suite 对这些变异仍全绿,在 exact head 上新增 case 分别失败,说明新增断言确实提供了此前不存在的回归敏感度。
具体改动
关键代码讲解
tests/capabilities/test_benchmark_study_projection.py::test_dashboard_projects_runtime_observations_into_the_campaign_block 是本 PR 唯一手写行为改动。它构造 baseline/treatment board rows、一个 insight 与三种 runtime observations,然后从完整 dashboard 返回值读取 campaign 和 authority,覆盖真实公开 builder,而不是直接调用 reducer 内部 helper。
loopx/capabilities/benchmark_toolkit/runtime_observation.py::build_benchmark_runtime_observation 是未修改的 typed state owner。新增用例通过它产生 not_admitted、running_qualified、terminal_pending_reconcile,再用 BenchmarkRuntimeClassification 枚举校验投影 key,避免把 prose 或任意字符串当作合法状态分类。
loopx/capabilities/benchmark_toolkit/study_projection.py::build_benchmark_study_dashboard 是实际消费边界:它从 active envelopes 过滤 runtime_observation,按 classification 聚合 runtime_counts,并把 count、typed buckets、factorial counts 与 comparison source 投影到 dashboard。测试对这些既有输出建立一致性关系,没有新建平行实现。
loopx/semantics/project_registry_io_manifest_v1.json 只更新一处生成 census 的源码行号。相同校验在 immutable base 上因 contract.py::check_contract 的旧 anchor 失败,在 head 上显示 250 sites current;这是让该 exact head 满足现有架构门禁的机械刷新,不改变 runtime contract。整个 PR 为两个文件、+80/-1,其中 79 行是一个同域 regression case,没有生产逻辑或公共协议增量。
对主干的风险
未发现阻塞 finding。最强回归场景是 reducer 停止接纳 runtime_observation、不再累加分类,或让 campaign 的 factorial count 与实际 list 分离;新增 case 对三个场景的独立 mutation 都会失败,而 base 的原有 20 项测试仍通过。head 的 projection file 为 21/21,通过五个相关 benchmark toolkit 文件的 103/103;架构 census 两文件 9/9,manifest generator 报 250 sites current;Ruff lint、diff checks 与 Python compile 通过。Ruff format 在 base/head 都报告相同四处既存 hunk,新增代码没有扩大该差异。
完整 risk-based premerge 在补齐仓库声明的 npm dev dependencies 后通过 4 项 direct checks 与 6 项 catalog checks,包括 semantic vocabulary、benchmark permission、candidate-source 和 artifact-boundary smokes。最初 premerge 的 semantic smoke 失败仅因为 checkout 缺少 TypeScript package;依赖安装后同一 exact head 独立重跑及完整 premerge 均通过。GitHub 的 Sign-off、dependency review、Python/TypeScript shards、dashboard acceptance、pytest、Sonar 和 merge-gate 在该 head 全绿。
该 PR 是 coverage-only 边界:不会修改默认 payload、渲染、调度、权限或外部 effect,因此 default-off、authority lifecycle 与行为披露没有新增兼容性负担。79 行测试同时包含少量 factorial/authority consistency 断言,范围略宽于标题中的 runtime observation,但它们来自同一 campaign payload、使用同一个真实 reducer,且 mutation 证明能捕获当前默认示例中已出现过的矛盾形状;作为同域防回归补强是相称的。
语义与 CI 对齐
本 PR 复用现有 BenchmarkRuntimeClassification、upload-envelope record kind、study dashboard schema 与 factorial projection authority,没有创建或扩展共享词汇。新增 census anchor 与生成器输出一致,semantic-vocabulary-drift smoke 和 required CI 均通过;typed classification 与聚合关系由真实入口 mutation 验证,而不是从当前输出反推预期。
我的整体评价
这是一个窄而完整的测试增量:它命中已发布的 reducer 与 dashboard consumer seam,补足此前可复现的盲区,同时不把 issue #5203 的后续 payload/UI 决策偷偷带入本 PR。对用户体验没有默认变化,对长期维护则明确提升了 runtime health campaign summary 的可信度。
代码体积与问题相称,现有 owner 得到复用,typed state、authority 边界和 public/private boundary 均保持不变。唯一残余风险是该 case 没有启动 CLI subprocess 或浏览器渲染,但既有 sibling cases 已覆盖其它 record kinds 的 CLI 流程,本 PR 又未改 CLI/UI;结合真实 public builders、三项 mutation、103 项相关 suite、完整 premerge 与全绿远端检查,现有覆盖足以批准该 exact head。
English verdict: APPROVE — exact head 2a9b5877a6c9aec856b276e5a186d0dafa6445c5 adds durable real-path coverage for runtime-observation aggregation in the study dashboard without changing production behavior. Three independent mutations are killed only by the new case; the 103-test related suite, architecture census and manifest check, risk-based premerge, and all required remote checks pass. No blocking finding remains.
|
This pull request has merge conflicts with Choose the remote for the base repository, not an out-of-date fork. git fetch upstream
git rebase upstream/main
# Resolve each conflict, git add the resolved files, then git rebase --continue.
git push --force-with-lease origin HEADFor a same-repository clone whose Keep the DCO |
…aign block Signed-off-by: karenchuu <25980598+karenchuu@users.noreply.github.com>
2a9b587 to
8ff0b3e
Compare
cocolord
left a comment
There was a problem hiding this comment.
动机
这个 PR 补的是 benchmark study reducer 的真实覆盖缺口:生产代码已接收 runtime_observation envelope,并将 observation 数量和 typed classification 聚合到 dashboard 的 campaign,但原有测试没有任何一条把这种 record kind 送入该路径。因此 reducer 即使漏收记录、漏算分类,既有 projection suite 仍可能全绿。当前修改只增加回归测试,不改 payload、CLI、UI、调度、评分或 benchmark 执行。
改动思路
测试复用现有公开 builder:用 build_benchmark_runtime_observation 从 admission、receipt 和 runner-owner facts 生成三种合法 observation,再经现有 upload envelope 进入 build_benchmark_study_dashboard。断言从相同 reducer 输出建立关系:observation 总数等于分类 bucket 总和,分类 key 属于 BenchmarkRuntimeClassification,factorial count 与实际 list 一致,authority source 与 list 中 schema 一致。测试不复制分类算法,也不新增 state owner。
具体改动
完整 diff 只有 tests/capabilities/test_benchmark_study_projection.py 增加 79 行。新增 test_dashboard_projects_runtime_observations_into_the_campaign_block 构造两条 board row、一个 insight 和三条 runtime observation,覆盖 not_admitted、running_qualified、terminal_pending_reconcile 三种分类,然后通过完整 dashboard 结果检查 campaign 与 authority。
关键内容讲解
- 输入通过
build_benchmark_upload_envelope(..., record_kind="runtime_observation")进入正式 normalization/filter 路径,不直接调用 reducer 私有 helper。 - 分类合法性由现有
BenchmarkRuntimeClassification枚举约束;新增断言验证 bucket key 和总量关系,而不是把任意 prose 当状态。 - factorial 的 count/list/source 断言位于同一个 campaign payload 上,能捕获摘要与明细分叉,但不改变其生产定义。
- 仓库检索确认此前没有其他测试或 example 覆盖 runtime-observation 到 study campaign 的这条路径;同作者同期只有这一条 benchmark test PR,不构成重复的同形批量提交。
对主干的风险
这是 coverage-only 变更,主要风险不是运行时回归,而是新增测试是否只固定偶然输出、增加脆弱维护成本。这里的断言来自独立的公开 contract:typed classification、总数守恒、list/count 一致性和 authority lineage;此前三种生产 mutation 在 base 的 20 项 suite 中均能漏过,而加入该测试后都会失败。当前 exact head 上目标文件 21/21、相关 benchmark suite 156/156、Ruff、Python compile 与 4 项 risk-based premerge canary 均通过。
远端 required CI 仍红,但固定 base 与当前 head 的六个失败测试 ID/断言签名一致,涉及 quota canonical projection、global gate 与 scheduler ACK,不在本 PR 唯一测试文件及 benchmark reducer 路径内。因此这些失败不归因于本 PR,但 merge readiness 仍应 hold,直到 required checks 变绿或主干基线被修复。未运行浏览器验证,因为本 PR 不改 UI 或用户可见 payload。
我的整体评价
APPROVE。exact head 8ff0b3ef5e76abe3463de562ae6786d917cb2b7b 是一个窄而耐久的测试增量:它覆盖已发布 reducer 和 dashboard consumer seam,能杀死此前漏检的实际回归,又不引入生产代码、私有 fixture、原始 benchmark 证据或一次性实验逻辑。与此前已审 head 相比,测试文件和被测 production reducer 的哈希完全一致,而新 head 还移除了无关的 manifest 行号刷新;我已重新执行当前 head 的 focused/related suite。未发现阻塞项,approval 不代表 merge 授权,红色 required CI 仍是单独的合并门禁。
English verdict: APPROVE 8ff0b3ef5e76abe3463de562ae6786d917cb2b7b. This test-only exact head adds durable, non-duplicative real-reducer coverage for runtime-observation campaign aggregation. The focused 21-test and related 156-test suites, lint/compile, and risk-based premerge checks pass; prior mutation evidence remains valid because the test and owning production modules are byte-identical. The six required-CI failures match the immutable base and remain a separate merge-readiness hold.
Goal And Delivered Outcome
Outcome basis / optional anchor: task issue [Task][RFC]: Qualify the integrated benchmark study upload and dashboard contract #5203 (
[Task][RFC]forbenchmark-study-upload-dashboard-v0). Claimed on that issue with the milestone andthe exact files named, as its contribution route asks. Base
6643f3670.Goal/source and gap: RFC §7 requires that the derived campaign/arm/case/run views come
from the shipped reducer. Five of the six acceptance items are already evidenced by
existing tests; the coverage half of the fourth was not. §3.2's fifth record kind,
runtime_observation, reached no test at all:tests/andexamples/finds no case uploadingrecord_kind="runtime_observation", althoughstudy_projection.py:47-53accepts itand
:1304-1311derivesruntime_classification_countsstraight from those records;benchmark-study-page.tsx:154-156renders had no case ableto tell a derived runtime-health value from an invented one.
Observable before → after, with the validation row that proves it: one new case,
test_dashboard_projects_runtime_observations_into_the_campaign_block, uploads two boardrows, one insight and three runtime observations through the existing envelope
builder, then pins the campaign block to the reducer's own definitions: the observation
count equals what was uploaded, the classification keys are members of
BenchmarkRuntimeClassificationand sum to that count, andfactorial_contrast_countequals
len(factorial_contrasts)with a matchingfactorial_comparison_source. Beforethis case the product could return anything there and the suite stayed green (mutation
rows below).
Issue/task and intended base: Related to [Task][RFC]: Qualify the integrated benchmark study upload and dashboard contract #5203, not
Closes— the acceptancereconciliation continues in the remaining-work note.
Scope And Continuation
Completed scope and remaining work: this closes the coverage half of criterion 4 for the
projection. It deliberately does not touch the shipped default payload, and while writing
this case I found two aggregates in it that this reducer cannot emit:
apps/presentation/dashboard/public/benchmark-study.example.json:32lists"terminal_qualified": 4, which is not a member ofBenchmarkRuntimeClassification(
runtime_observation.py:29-36);:29-30claimfactorial_contrast_count: 1while:88is"factorial_contrasts": [],although the reducer defines that count as
len(board["factorial_contrasts"])(
study_projection.py:1341).I am not repairing that file here on purpose: it is the dashboard's default user-visible
payload, so replacing it changes what the page shows and needs a re-hashed
loopx/web/chat/bundle-manifest.json, and the contributor board keeps the dashboardsurface maintainer-owned (GH-C101). The new case is what turns that discrepancy into a
checkable statement instead of an opinion — it is exactly mutation M2 below. Reported on
[Task][RFC]: Qualify the integrated benchmark study upload and dashboard contract #5203 rather than decided unilaterally.
Slice boundary / successor: complete within this scope for test coverage. The payload
repair is the named successor, once the dashboard owner says whether the current numbers
are intentional illustrative data.
Validation
2a9b5877a(two commits: the change itself, plus a one-field refresh ofloopx/semantics/project_registry_io_manifest_v1.json. That census test fails onmainat6643f3670by itself -loopx/contract.py::check_contractis recorded at line 1014 and reads at 1027 - and the refresh is what feat(goal-channel): compose scoped claims and private reconnect context (#5198) #5248/fix: settle declared completions before terminal closeout #5249/feat(explore): add typed research evidence and composition shadow #5250 each carry to get green; without it this PR'spytestlane inherits main's red. Regenerated withscripts/generate_project_registry_io_manifest.py(250 sites, 0 unclassified); only that one field moved.staticpassedruff checkclean on the changed file. Limitation, stated rather than hidden: this file already failsruff format --checkat the base commit (four hunks at lines 593, 616, 730 and 872), so I did not reformat it; every line I added is format-clean, checked by comparing the--diffhunk list before and after this change.unitpassedpytest tests/capabilities/test_benchmark_study_projection.py→ 21 passed (20 existing plus this case).integrationpassedpytestover the five benchmark toolkit capability test files (study_projection,experiment_board,four_arm_contract,runtime_observation,behavior_finding) → 103 passed / 0 failed.regression_paritypassedstudy_projection.py, reverted after each, all caught by the new case: (M1) removingruntime_observationfrom the accepted record kinds; (M2) hardcodingfactorial_contrast_count: 1beside an emptyfactorial_contrastslist — the exact shape of the shipped payload discrepancy; (M3) dropping the classification aggregation. On the base commit, all three leave the existing suite green.real_entrypointpassedbuild_benchmark_upload_envelopeandbuild_benchmark_study_dashboard, the same functions the CLI and local readback use; no private helper is called directly. Limitation: it does not runloopx benchmark-studyas a subprocess — sibling cases already cover the CLI path for the other record kinds.manualnot_applicableseam, which had none. It cannot cover the shipped example payload named above, nor the
dashboard rendering;
examples/dashboard-benchmark-study-browser-smoke.mjswas not runbecause nothing in the render path changed.
See validation disclosure guidance.
Frontend / Visual Evidence
N/A — no user-visible surface changes. The payload repair this PR documents is left as the
successor precisely because it would be visible.
Repro And Regression Proof
regression_parity) toloopx/capabilities/benchmark_toolkit/study_projection.pyon the base commit and runpytest tests/capabilities/test_benchmark_study_projection.py: the whole file staysgreen. With this case present, each one fails it.