Skip to content

feat(capabilities): bounded Goal improvement intent M1 backend preview - #5493

Merged
huangruiteng merged 2 commits into
mainfrom
codex/autonomous-capability-policy-m1-20261003
Oct 3, 2026
Merged

huangruiteng merged 2 commits into
mainfrom
codex/autonomous-capability-policy-m1-20261003

Conversation

@huangruiteng

@huangruiteng huangruiteng commented Oct 2, 2026 •

Copy link
Copy Markdown
Collaborator

Summary / 摘要

Add the bounded improvement-intent M1 backend preview for Goal-scoped capability organization. This consumes the direct-first design in the portfolio RFC; it is not the RFC's separately numbered external-evidence connector milestone.

实现 Goal 能力组织的“主动改进意图”M1 后端预览:原配置 owner、既有 before_plan/replan 入口、有界发现/试用建议和效果/回滚引用。不是第二个能力启用开关,也不是外部证据 connector 里程碑。

  • The original Goal registry and configure-goal / settings preview-apply-CAS writer own control_plane.capability_improvement. Default off, bounded minutes 1–30 and proposed trials 0–2. Invalid authoring cannot persist; corrupt stored intent remains visible and explicitly clearable.
  • TS owns normalization and advice inside the existing packaged runtime fingerprint boundary. Python only adapts IO. Reuse the existing bounded coordinator projection and its unchanged size budget; extract the existing subagent enablement calculation instead of duplicating it.
  • No explicit gap means no discovery. Applicable enabled capabilities take precedence over trial proposals. A trial requires explicit applicability, disabled state, original-owner configuration, effect-baseline and rollback references. Caller observations remain unauthenticated advice, never execution authority.
  • No new memory/adoption store, scheduler, automatic catalog scan, provider execution or autonomous installation. Existing Todo/lease/quota, provider readiness and protected-effect admission remain authoritative. Off preserves existing coordinator output.

原 Goal 配置、事务、CAS 与读回复用同一个 owner;默认关闭、失败建议不阻塞工作。直接能力无需加入 portfolio,试用仍走原 owner。没有增加记忆/采用账本、调度器、自动扫描或执行权限;未知效果仍是未知,不以调用、PR 或测试通过代替真实效用。

User entry points and remaining work / 用户入口与剩余工作

Implemented: CLI configuration + readback, explicit cold-path agent-context advice, live quota before-plan/replan injection, and the original Goal settings API/editor field descriptors.

Partial product delivery: the existing schema-driven App editor receives the three fields, but packaged App interaction, localization and Lark guidance/control/receipt readback still require companion qualification through that same projection. No frontend interaction or Lark end-to-end acceptance is claimed. Companion scope follows the RFC delivery path, without a second configuration or chat source of truth. The packaged frontend build below proves packaging, not those interactions.

已完成 CLI/共享设置 API/上下文注入;打包 App 交互与本地化、Lark 控制和回执读回仍未验收,故产品为 partial。未修改 App 源文件,不声称已完成端到端、真实领域效用或默认采用。此 Core 行为改动留给维护者审查与合并,不自合并或升级默认安装。

Validation / 验证

  • Stable acceptance command: LOOPX_USAGE_PING=0 uv run --extra test pytest -q tests/capabilities/test_capability_improvement.py — 10 passed, 13.46 s; real CLI preview-no-write / apply-readback / clear, same-owner settings CAS, live quota before-plan and typed replan, invalid-input negatives.
  • TS policy + existing agent-context tests — 22 passed, no skips; project control-plane typecheck and explicit new-test typecheck passed.
  • Configuration UI/transaction/machine-contract family — 49 passed.
  • Independently installed wheel from outside the source checkout — 7 CLI/resource checks passed: default off, preview-no-write, apply-readback, direct-first, replan trial, clear/off output parity, installed TS + bilingual-doc hash parity. Disposable synthetic registry only; no default adoption or live Goal mutation.
  • Packaged Chat build/manifest verification, wheel build, CLI output-budget regression, focused Ruff, project Mypy and semantic drift smoke passed. Registry I/O manifest update only adjusts two source line offsets; 273 sites and zero unclassified direct sites. The diff-scoped vocabulary advisory detects no supported carriers; it does not analyze unnamed literals or prove absence of semantics.
  • Broader related family on the final head had 85 passes and 1 failure, 34.78 s. The unchanged test_goal_configuration_service_rechecks_revision_before_write is stopped before CAS by this machine's conflicting implicit runtime roots. The exact unchanged baseline reproduces the same failure. Neither that test nor routing behavior was weakened; this is not a full-green repository/CI claim.

新增负向验收、关闭兼容性与真实安装包路径均通过;既有隐式默认目录冲突在原基线也复现,原断言保留并披露。未跑通全部仓库/CI,不以局部验证或打包成功冒充完整产品验收。私有运行材料、日志和实验产物仅保留本地,不提交。

Reuse and simplification / 复用与精简

Configuration reuses the original registry transaction/catalog/settings API; guidance reuses projectAgentContext and the existing quota entry. One TS owner handles all policy rules. Subagent configuration remains owned by its original helper. No generic execution engine or composition DSL was added.

配置及读回复用原事务/目录/设置 API,上下文复用既有有界调度器,子 Agent 配置判断由原 helper 唯一持有;未创建并行 Python 决策源或新增通用执行器。

Signed-off-by: huangruiteng <huangrt01@163.com>

@huangruiteng huangruiteng left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approval conclusion (author-owned PR; GitHub blocks formal self-approval)

Reviewer: model_agent | GPT-5 | OpenAI

English verdict: APPROVE

Exact head: 5493@1685de24ed13682055bde4555b9f50b60d9fff7d; immutable common baseline: 649a54eeeb00953e0f501af44d1b5e59f903f349. 完整 PR 评审;当前目标不等待 CI,本轮未查询或轮询 CI。作者所属账号只能发布 COMMENTED 结论,不是独立维护者批准。

动机

维护者和执行任务的 Agent,需要在工作中判断是否值得寻找或试用另一项能力,但不应让例行任务反复进入探索。

例如,任务缺少一种能力时,原来没有可设置的有界改进建议入口;现在可以对一个 Goal 开启分钟数和试用数量受限的建议,已有可用能力优先,试用建议必须附有效果与回滚依据。没有明确缺口时仍继续原任务。

本轮真实 CLI 与 HTTP 验证证明:默认关闭时三个生命周期阶段的完整输出不变;开启后可以预览、写入和读回同一份配置,过期预览拒绝写入,重新预览后恢复。建议仍不授予执行权限。

本次不自动扫描、安装或执行能力,不建立新的采用/记忆账本,也不宣称探索产生了真实领域效用或完成整个能力组织产品。

打包 App 交互、本地化、Lark 控制与回执、真实使用者采用和持续效用仍由既有能力组织交付路径验收;当前是可独立使用的后端预览增量。

配置里的“有界”是维护者明确选择的改进投入,不是能力授权。这个增量的价值在于让例行工作不被探索打断,同时把值得试用的情况交回原能力 owner;测试数量、建议出现或试用次数都不证明真实效用。接受后仍不能把产品交付、安装或采用宣告完成。

改动思路

复用现有 Goal 配置、事务与预览/应用 CAS(按预览版本比较后写入),不另建开关或状态库。新增决策归现有 TypeScript 控制面:Python 只传送原始配置与观察。原能力 owner 仍独占启用、准备就绪、权限与执行,建议携带引用也不能代替这些检查。

上下文继续使用原有有界 provider 组合。只有选定 Goal 的 bounded 模式在 before_plan/replan 加入建议;已有启用能力优先,没有明确缺口不探索,试用还需要配置、效果基线与回滚引用。抽取原子 Agent 配置 helper 是同一变更原因的行为保持重构,不引入第二个决策源或泛化执行框架。

具体改动

全量 22 文件、+817/-8,包含生产、双语文档、真实入口回归和两处 registry IO 行号坐标;273 个 IO site 的数量未变。原 settings editor、共享 API schema、CLI 与 Lark 消费者也检查了:现有 schema 驱动配置字段可承载这三个值,但这不等于打包 UI 或 Lark 旅程已验收。

书面基准:docs/architecture/rfcs/goal-scoped-capability-portfolio-v0.md,spec_revision 649a54eeeb00953e0f501af44d1b5e59f903f349。该共同基线先于本 PR;后来的 main 中同一 RFC 内容相同。此次 README 和作者宣告仅用于行为披露,不能作为独立验收标准。

  • §1 Decision summary:Direct use first; portfolio/improvement advice not prerequisite to existing capability use. implemented:对应配置/上下文 owner 及下面真实正反路径。
  • §5 Ownership and authority:Original capability/configuration/effect owners retain enablement and admission. implemented:对应配置/上下文 owner 及下面真实正反路径。
  • §5 Activation and proportional execution:Goal-scoped default-off, bounded intent, no mandatory exploration for routine work. implemented:对应配置/上下文 owner 及下面真实正反路径。
  • §5 Runtime injection:Reuse existing bounded context provider, preserve off and unrelated phases. implemented:对应配置/上下文 owner 及下面真实正反路径。
  • §7 Safety, privacy and compatibility:Failure must not block useful work; references are advice, not private raw evidence or authority. implemented:对应配置/上下文 owner 及下面真实正反路径。
  • §9 Validation acceptance:Real preview/apply/CAS/clear/readback, off parity and scope/failure recovery. implemented:对应配置/上下文 owner 及下面真实正反路径。
  • §11 Normative delivery plan:Backend preview remains partial; packaged frontend/Lark/utility acceptance deferred to same existing owner. deferred:同一既有交付路径继续持有前端/Lark/真实效用;本次不关闭它。

关键代码讲解

  • loopx/control_plane/capabilities/goal_capability_organization.ts:11 — normalizeImprovementPolicy:Reject unknown fields and bool/noninteger bounds; mode off/bounded, minutes1–30, trials0–2.
  • loopx/control_plane/capabilities/goal_capability_organization.ts:45 — planCapabilityImprovement:No gap=>continue; enabled applicable direct first; disabled trial requires original-config/effect/rollback refs; no execution authority.
  • loopx/control_plane/capabilities/goal_capability_organization.ts:82 — improvementContextProvider:Only before_plan contribution; original3072 budget and guidance_only authority; raw observation not emitted.
  • loopx/control_plane/goal_agent_context.ts:8 — evaluateGoalAgentContext:Compose original subagent provider; improvement provider loaded only bounded Goal gate.
  • loopx/capabilities/goal_capability_organization/goal_configuration.py:21 — apply_change:IO adapter invokes typed plan; original Goal transaction and shared HTTP CAS retain persistence.

独立验收不是把输出重新包装成 oracle:在共同基线与 head,默认关闭时三个阶段的完整公共输出相等,没有删除诊断、权限、指引或失败详情;只读调用的 registry 字节未变。当前 head 的 23 个 CLI/HTTP 观察、25 项独立断言通过,包括他 Goal、开启后新 Goal、不正确阶段、超预算拒绝、无缺口、direct-first、缺少回滚依据,以及损坏配置不阻塞原子 Agent 上下文。清除后三个阶段恢复原输出。

实际 ChatHTTPServer 与原 file registry writer 验证:预览 201 不写入,另一真实 CLI 更改同一配置后旧预览应用返回 409 且无写入;重新预览 201、应用 200、单独读回配置成功。这不是 mixin/mock 的完成状态。试用敏感性另用临时源码副本去掉 rollback 判断,在相同真实 agent-context 命令上错误地产生 propose_reversible_trial,使预先独立的 continue_current_work 断言失败;恢复精确源码后通过,未修改 PR head。

当前源码 Python 63 passed、1 failed;TS 22/22 passed、无 skips。差异词汇 advisory 先于全树 semantic drift,随后完整控制面 tsc、Ruff、配置的 19-source mypy、diff check 全部通过。这些是有界证据,不证明全部仓库、真实领域效用或安装后的产品采用。

对主干的风险

保留失败原貌:未改的 test_goal_configuration_service_rechecks_revision_before_write 在共同基线与 head 同样被临时 TS 锁文件误判为第二个默认运行目录,先于 stale-CAS 检查拒绝。隔离当前/legacy 根的锁内追踪同样复现;相关 paths.py、锁实现和 configure service 在这两版字节一致。较新 main 的独立 f955e8305 修复了该判断,9ac4efa 上测试通过。故这是可独立归因的旧基线缺陷,不是本次加剧的失败,也不只是“环境噪声”。没有删除测试或降低断言;本轮不声称当前 branch 全绿。

当前 remote base 已前进,批准源码增量与组合后的集成/合并就绪分开:不能用较新 main 单独通过来声称 PR 分支已集成,也不要求本 PR 再实现同一锁发现修复。后续维护者集成要保留这项读回。建议的候选上限只是观察边界,空候选不是 catalog 不存在的证据;mutable 引用或 enabled 标记不能产生原能力执行权。关闭与损坏模式不注入新指引,默认行为未扩大到其他 Goal。

相关 future-facing pass 已采用:共享原子 Agent 配置计算和原 3072 字节组合预算;无新采用账本、重复 planner、调度器或执行器。CLI/共享设置 API/上下文变了;打包 App、本地化和 Lark companion 的剩余验收在原能力 owner,不能以“后端-only”规避,也不建仪式性新 Todo。回滚仍是原 Goal 配置 off/clear 和源码 revert,未自动安装或切换运行中的 Goal。

我的整体评价

APPROVE 这个 justified_increment:长程上减少重复权威与无界探索,用户体验上提供同一配置入口、可读回的有界建议和真实拒绝/恢复;可独立使用而不扩张执行权限。默认关闭、作用域、回滚依据、真实 CAS 与敏感性证据成立,没有本次 scoped 阻塞项。

通过只覆盖本次后端预览源码,不覆盖整个 M1/M2、打包前端/Lark、安装采用或真实效用。旧基线红项已归因但单独保留 readiness/组合验收;发布后还须执行 approval-closeout,过期阻塞评审只有逐项证明确已修复/被反证且有权限才撤销。不会自动合并、安装、启用能力或宣告 Goal 完成。

@huangruiteng
huangruiteng merged commit 7d829aa into main Oct 3, 2026
7 checks passed
@huangruiteng
huangruiteng deleted the codex/autonomous-capability-policy-m1-20261003 branch October 3, 2026 06:22
@Duang777

Duang777 commented Oct 3, 2026

Copy link
Copy Markdown
Collaborator

Post-merge CI found a deterministic baseline failure in test-shard (1): tests/capabilities/test_capability_extension_registry.py::test_builtin_catalog_preserves_order_and_marks_provider expects BUILTIN_IDS, while the catalog now appends goal-capability-organization.

I reproduced the same failure on immutable main@60f0e64e with the exact focused test (1 failed), and neither the test nor capability catalog is changed by dependent PR #5338. The post-merge shard for #5493 shows the same failure. This needs a baseline follow-up; I am keeping #5338 scoped and unchanged.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants