Skip to content

test(benchmark): pin the runtime-observation path into the study campaign block - #5255

Merged
huangruiteng merged 1 commit into
loopx-project:mainfrom
karenchuu:karenchuu/benchmark-study-runtime-observation-coverage
Sep 29, 2026
Merged

huangruiteng merged 1 commit into
loopx-project:mainfrom
karenchuu:karenchuu/benchmark-study-runtime-observation-coverage

Conversation

@karenchuu

@karenchuu karenchuu commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Goal And Delivered Outcome

  • Outcome basis / optional anchor: task issue [Task][RFC]: Qualify the integrated benchmark study upload and dashboard contract #5203 ([Task][RFC] for
    benchmark-study-upload-dashboard-v0). Claimed on that issue with the milestone and
    the exact files named, as its contribution route asks. Base 6643f3670.

  • Goal/source and gap: RFC §7 requires that the derived campaign/arm/case/run views come
    from the shipped reducer. Five of the six acceptance items are already evidenced by
    existing tests; the coverage half of the fourth was not. §3.2's fifth record kind,
    runtime_observation, reached no test at all:

    • a grep over tests/ and examples/ finds no case uploading
      record_kind="runtime_observation", although study_projection.py:47-53 accepts it
      and :1304-1311 derives runtime_classification_counts straight from those records;
    • so the campaign block that benchmark-study-page.tsx:154-156 renders had no case able
      to tell a derived runtime-health value from an invented one.
  • Observable before → after, with the validation row that proves it: one new case,
    test_dashboard_projects_runtime_observations_into_the_campaign_block, uploads two board
    rows, one insight and three runtime observations through the existing envelope
    builder, then pins the campaign block to the reducer's own definitions: the observation
    count equals what was uploaded, the classification keys are members of
    BenchmarkRuntimeClassification and sum to that count, and factorial_contrast_count
    equals len(factorial_contrasts) with a matching factorial_comparison_source. Before
    this case the product could return anything there and the suite stayed green (mutation
    rows below).

  • Issue/task and intended base: Related to [Task][RFC]: Qualify the integrated benchmark study upload and dashboard contract #5203, not Closes — the acceptance
    reconciliation continues in the remaining-work note.

Scope And Continuation

  • Completed scope and remaining work: this closes the coverage half of criterion 4 for the
    projection. It deliberately does not touch the shipped default payload, and while writing
    this case I found two aggregates in it that this reducer cannot emit:

    • apps/presentation/dashboard/public/benchmark-study.example.json:32 lists
      "terminal_qualified": 4, which is not a member of BenchmarkRuntimeClassification
      (runtime_observation.py:29-36);
    • :29-30 claim factorial_contrast_count: 1 while :88 is "factorial_contrasts": [],
      although the reducer defines that count as len(board["factorial_contrasts"])
      (study_projection.py:1341).

    I am not repairing that file here on purpose: it is the dashboard's default user-visible
    payload, so replacing it changes what the page shows and needs a re-hashed
    loopx/web/chat/bundle-manifest.json, and the contributor board keeps the dashboard
    surface maintainer-owned (GH-C101). The new case is what turns that discrepancy into a
    checkable statement instead of an opinion — it is exactly mutation M2 below. Reported on
    [Task][RFC]: Qualify the integrated benchmark study upload and dashboard contract #5203 rather than decided unilaterally.

  • Slice boundary / successor: complete within this scope for test coverage. The payload
    repair is the named successor, once the dashboard owner says whether the current numbers
    are intentional illustrative data.

Validation

Check kind Result Public-safe evidence / limitation
static passed ruff check clean on the changed file. Limitation, stated rather than hidden: this file already fails ruff format --check at the base commit (four hunks at lines 593, 616, 730 and 872), so I did not reformat it; every line I added is format-clean, checked by comparing the --diff hunk list before and after this change.
unit passed pytest tests/capabilities/test_benchmark_study_projection.py → 21 passed (20 existing plus this case).
integration passed pytest over the five benchmark toolkit capability test files (study_projection, experiment_board, four_arm_contract, runtime_observation, behavior_finding) → 103 passed / 0 failed.
regression_parity passed Three product mutations applied one at a time to study_projection.py, reverted after each, all caught by the new case: (M1) removing runtime_observation from the accepted record kinds; (M2) hardcoding factorial_contrast_count: 1 beside an empty factorial_contrasts list — the exact shape of the shipped payload discrepancy; (M3) dropping the classification aggregation. On the base commit, all three leave the existing suite green.
real_entrypoint passed The case enters through build_benchmark_upload_envelope and build_benchmark_study_dashboard, the same functions the CLI and local readback use; no private helper is called directly. Limitation: it does not run loopx benchmark-study as a subprocess — sibling cases already cover the CLI path for the other record kinds.
manual not_applicable No UI or payload changes in this PR.
  • Coverage and gaps: the added case covers the runtime-observation → campaign projection
    seam, which had none. It cannot cover the shipped example payload named above, nor the
    dashboard rendering; examples/dashboard-benchmark-study-browser-smoke.mjs was not run
    because nothing in the render path changed.

See validation disclosure guidance.

Frontend / Visual Evidence

N/A — no user-visible surface changes. The payload repair this PR documents is left as the
successor precisely because it would be visible.

Repro And Regression Proof

  • Repro before the fix: apply any of M1-M3 (see regression_parity) to
    loopx/capabilities/benchmark_toolkit/study_projection.py on the base commit and run
    pytest tests/capabilities/test_benchmark_study_projection.py: the whole file stays
    green. With this case present, each one fails it.

cocolord
cocolord previously approved these changes Sep 28, 2026

@cocolord cocolord left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

APPROVE — 未发现阻塞项。exact head 2a9b5877a6c9aec856b276e5a186d0dafa6445c5 为此前没有覆盖的 runtime_observation → study dashboard campaign 投影补上了真实 reducer 路径测试;三项故障注入均能被新增 case 捕获,focused/related suite、架构 census、risk-based premerge 和远端 required checks 均通过。

动机

本 PR 补的是 benchmark study dashboard 的一个明确覆盖缺口:产品 reducer 已接受 runtime_observation upload envelope,并把 observation 数量与 typed classification 聚合到 campaign,但既有 suite 没有任何用例把这种 record kind 送进该路径。这样一来,即使 reducer 漏收记录、丢掉分类聚合,或者 dashboard payload 写出与 factorial list 不一致的摘要,原有 20 个 projection tests 仍可能保持绿色。

新增 case 把 RFC #5203 中 campaign projection 的 coverage slice 变成可执行约束。它不改变用户可见 payload、CLI、schema 或默认行为;对长期维护的改善是让 dashboard campaign 摘要与其 typed runtime/factorial authority 不再只有实现自证,而是具备可杀死实际回归的测试。这个边界本身完整、可独立回滚,后续示例 payload 是否修正仍可由其 owner 单独决策。

改动思路

测试复用现有生产入口:先由 build_benchmark_runtime_observation 从 admission、receipt 与 runner-owner facts 生成三种合法 typed observation,再用 build_benchmark_upload_envelope(..., record_kind="runtime_observation") 包装,与既有 board rows 和 insight 一起交给 build_benchmark_study_dashboard。权威规则仍全部留在原有 runtime enum、envelope normalizer 与 study reducer 中,测试没有复制第二套分类算法,也没有引入新的 state owner。

正向路径验证三条 observation 最终形成 runtime_observation_count == 3 以及三个 classification bucket;同时验证 bucket key 来自 BenchmarkRuntimeClassification 且总和等于 observation count。负向证据通过三次独立 mutation 覆盖 record-kind admission、runtime aggregation 和 factorial summary/list consistency:在 immutable base 上原有 suite 对这些变异仍全绿,在 exact head 上新增 case 分别失败,说明新增断言确实提供了此前不存在的回归敏感度。

具体改动

关键代码讲解

tests/capabilities/test_benchmark_study_projection.py::test_dashboard_projects_runtime_observations_into_the_campaign_block 是本 PR 唯一手写行为改动。它构造 baseline/treatment board rows、一个 insight 与三种 runtime observations,然后从完整 dashboard 返回值读取 campaign 和 authority,覆盖真实公开 builder,而不是直接调用 reducer 内部 helper。

loopx/capabilities/benchmark_toolkit/runtime_observation.py::build_benchmark_runtime_observation 是未修改的 typed state owner。新增用例通过它产生 not_admitted、running_qualified、terminal_pending_reconcile,再用 BenchmarkRuntimeClassification 枚举校验投影 key,避免把 prose 或任意字符串当作合法状态分类。

loopx/capabilities/benchmark_toolkit/study_projection.py::build_benchmark_study_dashboard 是实际消费边界:它从 active envelopes 过滤 runtime_observation,按 classification 聚合 runtime_counts,并把 count、typed buckets、factorial counts 与 comparison source 投影到 dashboard。测试对这些既有输出建立一致性关系,没有新建平行实现。

loopx/semantics/project_registry_io_manifest_v1.json 只更新一处生成 census 的源码行号。相同校验在 immutable base 上因 contract.py::check_contract 的旧 anchor 失败,在 head 上显示 250 sites current;这是让该 exact head 满足现有架构门禁的机械刷新,不改变 runtime contract。整个 PR 为两个文件、+80/-1,其中 79 行是一个同域 regression case,没有生产逻辑或公共协议增量。

对主干的风险

未发现阻塞 finding。最强回归场景是 reducer 停止接纳 runtime_observation、不再累加分类,或让 campaign 的 factorial count 与实际 list 分离;新增 case 对三个场景的独立 mutation 都会失败,而 base 的原有 20 项测试仍通过。head 的 projection file 为 21/21,通过五个相关 benchmark toolkit 文件的 103/103;架构 census 两文件 9/9,manifest generator 报 250 sites current;Ruff lint、diff checks 与 Python compile 通过。Ruff format 在 base/head 都报告相同四处既存 hunk,新增代码没有扩大该差异。

完整 risk-based premerge 在补齐仓库声明的 npm dev dependencies 后通过 4 项 direct checks 与 6 项 catalog checks,包括 semantic vocabulary、benchmark permission、candidate-source 和 artifact-boundary smokes。最初 premerge 的 semantic smoke 失败仅因为 checkout 缺少 TypeScript package;依赖安装后同一 exact head 独立重跑及完整 premerge 均通过。GitHub 的 Sign-off、dependency review、Python/TypeScript shards、dashboard acceptance、pytest、Sonar 和 merge-gate 在该 head 全绿。

该 PR 是 coverage-only 边界:不会修改默认 payload、渲染、调度、权限或外部 effect,因此 default-off、authority lifecycle 与行为披露没有新增兼容性负担。79 行测试同时包含少量 factorial/authority consistency 断言,范围略宽于标题中的 runtime observation,但它们来自同一 campaign payload、使用同一个真实 reducer,且 mutation 证明能捕获当前默认示例中已出现过的矛盾形状;作为同域防回归补强是相称的。

语义与 CI 对齐

本 PR 复用现有 BenchmarkRuntimeClassification、upload-envelope record kind、study dashboard schema 与 factorial projection authority,没有创建或扩展共享词汇。新增 census anchor 与生成器输出一致,semantic-vocabulary-drift smoke 和 required CI 均通过;typed classification 与聚合关系由真实入口 mutation 验证,而不是从当前输出反推预期。

我的整体评价

这是一个窄而完整的测试增量:它命中已发布的 reducer 与 dashboard consumer seam,补足此前可复现的盲区,同时不把 issue #5203 的后续 payload/UI 决策偷偷带入本 PR。对用户体验没有默认变化,对长期维护则明确提升了 runtime health campaign summary 的可信度。

代码体积与问题相称,现有 owner 得到复用,typed state、authority 边界和 public/private boundary 均保持不变。唯一残余风险是该 case 没有启动 CLI subprocess 或浏览器渲染,但既有 sibling cases 已覆盖其它 record kinds 的 CLI 流程,本 PR 又未改 CLI/UI;结合真实 public builders、三项 mutation、103 项相关 suite、完整 premerge 与全绿远端检查,现有覆盖足以批准该 exact head。

English verdict: APPROVE — exact head 2a9b5877a6c9aec856b276e5a186d0dafa6445c5 adds durable real-path coverage for runtime-observation aggregation in the study dashboard without changing production behavior. Three independent mutations are killed only by the new case; the 103-test related suite, architecture census and manifest check, risk-based premerge, and all required remote checks pass. No blocking finding remains.

@mergify

mergify Bot commented Sep 28, 2026

Copy link
Copy Markdown

This pull request has merge conflicts with main and cannot be merged
until they are resolved. Please rebase or merge the base branch, @karenchuu.

Choose the remote for the base repository, not an out-of-date fork.
For a fork clone, first inspect git remote -v; upstream must point
to https://github.com/loopx-project/loopx.git. If it is absent, add it
with git remote add upstream https://github.com/loopx-project/loopx.git.
Then run:

git fetch upstream
git rebase upstream/main
# Resolve each conflict, git add the resolved files, then git rebase --continue.
git push --force-with-lease origin HEAD

For a same-repository clone whose origin points to
https://github.com/loopx-project/loopx.git, use origin instead of
upstream for fetch/rebase. If you prefer merging the base, use
git merge <base-remote>/main and push normally.

Keep the DCO Signed-off-by trailer on every commit when you rebase.
https://docs.github.com/en/pull-requests/collaborating-with-pull-requests/working-with-forks/syncing-a-fork

@mergify mergify Bot added the needs-rebase Mergify: the pull request has merge conflicts with its base branch label Sep 28, 2026
…aign block

Signed-off-by: karenchuu <25980598+karenchuu@users.noreply.github.com>
@karenchuu
karenchuu force-pushed the karenchuu/benchmark-study-runtime-observation-coverage branch from 2a9b587 to 8ff0b3e Compare September 29, 2026 04:18
@mergify mergify Bot removed the needs-rebase Mergify: the pull request has merge conflicts with its base branch label Sep 29, 2026

@cocolord cocolord left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

动机

这个 PR 补的是 benchmark study reducer 的真实覆盖缺口:生产代码已接收 runtime_observation envelope,并将 observation 数量和 typed classification 聚合到 dashboard 的 campaign,但原有测试没有任何一条把这种 record kind 送入该路径。因此 reducer 即使漏收记录、漏算分类,既有 projection suite 仍可能全绿。当前修改只增加回归测试,不改 payload、CLI、UI、调度、评分或 benchmark 执行。

改动思路

测试复用现有公开 builder:用 build_benchmark_runtime_observation 从 admission、receipt 和 runner-owner facts 生成三种合法 observation,再经现有 upload envelope 进入 build_benchmark_study_dashboard。断言从相同 reducer 输出建立关系:observation 总数等于分类 bucket 总和,分类 key 属于 BenchmarkRuntimeClassification,factorial count 与实际 list 一致,authority source 与 list 中 schema 一致。测试不复制分类算法,也不新增 state owner。

具体改动

完整 diff 只有 tests/capabilities/test_benchmark_study_projection.py 增加 79 行。新增 test_dashboard_projects_runtime_observations_into_the_campaign_block 构造两条 board row、一个 insight 和三条 runtime observation,覆盖 not_admitted、running_qualified、terminal_pending_reconcile 三种分类,然后通过完整 dashboard 结果检查 campaign 与 authority。

关键内容讲解

  • 输入通过 build_benchmark_upload_envelope(..., record_kind="runtime_observation") 进入正式 normalization/filter 路径,不直接调用 reducer 私有 helper。
  • 分类合法性由现有 BenchmarkRuntimeClassification 枚举约束;新增断言验证 bucket key 和总量关系,而不是把任意 prose 当状态。
  • factorial 的 count/list/source 断言位于同一个 campaign payload 上,能捕获摘要与明细分叉,但不改变其生产定义。
  • 仓库检索确认此前没有其他测试或 example 覆盖 runtime-observation 到 study campaign 的这条路径;同作者同期只有这一条 benchmark test PR,不构成重复的同形批量提交。

对主干的风险

这是 coverage-only 变更,主要风险不是运行时回归,而是新增测试是否只固定偶然输出、增加脆弱维护成本。这里的断言来自独立的公开 contract:typed classification、总数守恒、list/count 一致性和 authority lineage;此前三种生产 mutation 在 base 的 20 项 suite 中均能漏过,而加入该测试后都会失败。当前 exact head 上目标文件 21/21、相关 benchmark suite 156/156、Ruff、Python compile 与 4 项 risk-based premerge canary 均通过。

远端 required CI 仍红,但固定 base 与当前 head 的六个失败测试 ID/断言签名一致,涉及 quota canonical projection、global gate 与 scheduler ACK,不在本 PR 唯一测试文件及 benchmark reducer 路径内。因此这些失败不归因于本 PR,但 merge readiness 仍应 hold,直到 required checks 变绿或主干基线被修复。未运行浏览器验证,因为本 PR 不改 UI 或用户可见 payload。

我的整体评价

APPROVE。exact head 8ff0b3ef5e76abe3463de562ae6786d917cb2b7b 是一个窄而耐久的测试增量:它覆盖已发布 reducer 和 dashboard consumer seam,能杀死此前漏检的实际回归,又不引入生产代码、私有 fixture、原始 benchmark 证据或一次性实验逻辑。与此前已审 head 相比,测试文件和被测 production reducer 的哈希完全一致,而新 head 还移除了无关的 manifest 行号刷新;我已重新执行当前 head 的 focused/related suite。未发现阻塞项,approval 不代表 merge 授权,红色 required CI 仍是单独的合并门禁。

English verdict: APPROVE 8ff0b3ef5e76abe3463de562ae6786d917cb2b7b. This test-only exact head adds durable, non-duplicative real-reducer coverage for runtime-observation campaign aggregation. The focused 21-test and related 156-test suites, lint/compile, and risk-based premerge checks pass; prior mutation evidence remains valid because the test and owning production modules are byte-identical. The six required-CI failures match the immutable base and remain a separate merge-readiness hold.

@huangruiteng
huangruiteng merged commit 8007719 into loopx-project:main Sep 29, 2026
41 of 53 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants