Skip to content

fix: drift the prompt audit found in the instruction files - #35

Merged
mauricios merged 2 commits into
mainfrom
fix/prompt-audit-drift
Oct 6, 2026
Merged

mauricios merged 2 commits into
mainfrom
fix/prompt-audit-drift

Conversation

@mauricios

@mauricios mauricios commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

A prompt audit (/claude-api prompt-audit) of the framework's 55 instruction files found no dated prompting patterns. There's no pressure language beyond one line, and no thinking scaffolds, word caps, narration bans, formatting bans, or retired model names. Every command, agent, path, and named section the files cite resolves. What it did find is text that later decisions left behind, and two habits that cost tokens. This PR fixes the 10 findings the audit rated high or medium confidence, and the one flag that needed a decision (L1).

Audit assumptions:

  • Target models: files with no pinned model were checked against Claude Opus 5.5. Each agent was checked against its pin: Sonnet 5.5 or Opus 5.5.
  • Out of scope: the plugin's generated copies (regenerated here), eval fixtures, settings files, and GEMINI.md and .cursor/rules/, which non-Claude tools read.

High confidence: contradicted by the repository

# Where What was wrong Fix
H1 skeleton/.claude/skills/orchestrate/SKILL.md:52-55 Said the reviewers and @architect run on haiku, and @spec-analyzer on sonnet. The frontmatter it cites, CLAUDE.md:62, COST-MODEL.md's table, and decision 0012 (reviewers "from haiku to sonnet") disagree States the real tiering once
H2 same file, :62-67 The example plan, the template Claude copies, showed (haiku) agents, "~3 Haiku-tier invocations (~minimal)", and "<30s" Sonnet agents, the cost model's start-up cost, and no unbacked time
H3 same file, :110-115 "3 parallel agents, Haiku-tier: ~15-25k input". COST-MODEL.md:77 (newer) says each subagent starts from ~50–60k tokens Takes the figure from the cost model
H4 skeleton/docs/COST-MODEL.md:177 "Use Haiku for … most reviews", against its own per-agent table and decision 0012 Drops "and most reviews"
H5 CLAUDE.md:20 The plugin's generated files listed without the reference docs, which .generated now includes (decision 0019). An agent could hand-edit them Lists them
H6 CLAUDE.md:35 The rebuild rule covered only skeleton/.claude/, but the build also reads four docs in skeleton/docs/ Names them
H7 skeleton/.claude/rules/code-quality.md:91 "Explain why, not what". It loads with git-workflow.md, and both that file and /commit (both newer) say the subject says what changed Matches them

Medium confidence

# Where What Fix
M1 all 8 skeleton/.claude/agents/*/agent.md Each began by reading AGENTS.md and CLAUDE.md. Claude Code already loads both into a custom subagent's starting context, so every invocation paid for them twice Drops the re-read, keeping a pointer where the line said what to look for (conventions, run commands, test ports)
M2 deep-spec-analysis.js:87, deep-review.js:97 Each agent's own lens came before the context all of them share, so one run's agents couldn't reuse a cached prefix. The plugin's deep-spec-analysis context now carries ~10 KB of spec model ${context} first
M3 skeleton/.claude/rules/deployment.md:27 The only caps "NEVER" in the rules, with no reason A plain rule with its reason

Decided: L1, do failing tests block a merge?

deployment.md:41 said "Tests must pass before merge". git-workflow.md:69-70 said CI "is a signal, not the gate … red ones are information". Decided: tests pass before merge; a failure the reviewer accepts is explained in the pull request. Both files now say this. The human review stays the gate, as AGENTS.md, CONTRIBUTING.md, and the constitution example already say.

Not changed: low-confidence flags

  • L2, low confidence. Some web- and TypeScript-specific checks sit in stack-agnostic agents and rules.
  • L3, low confidence. The skill template restates one rule in several sections, for example /orchestrate's no-auto-progression rule appears six times. This is the deliberate template from decision 0013.

Verification

  • Plugin copies: regenerated with scripts/build-aplyca-adf.sh.
  • Static evals: all pass (check-skills 158, test-hooks 82, test-modules 45, test-plugin 18).
  • Out-of-band dependencies: no eval or script matched the changed text.
  • Each finding: rechecked against the repository, using the frontmatter, .generated, the build script's REFERENCE_DOCS, and git blame.
  • Subagent context: the claim that subagents load CLAUDE.md and AGENTS.md comes from Claude Code's subagent docs, § What loads at startup.
  • Worth running after merge:
    • one agent session, such as the plugin-docs suite's agent-spec-model, to see an agent work without the re-read;
    • a /deep-spec-analysis run, comparing its agents' cache_read_input_tokens.

Upgrade impact: in the CHANGELOG under Unreleased. In short: overwrite the agents, orchestrate, the two workflows, code-quality.md, git-workflow.md, and COST-MODEL.md, and merge two lines in deployment.md. The root CLAUDE.md is framework-internal.

🤖 Generated with Claude Code

mauricios and others added 2 commits October 5, 2026 23:41
A prompt audit of the framework's 55 instruction files found no dated prompting patterns, but text
later decisions had left behind, and two habits that cost tokens:

- /orchestrate named Haiku for the reviewers and @architect, and priced a review from Haiku-sized
  contexts; the agents run on Sonnet and Opus since decision 0012. Its plan, example, and cost note
  now match the frontmatter and the cost model.
- COST-MODEL.md recommended Haiku for "most reviews", against its own per-agent table.
- All eight agents began by reading AGENTS.md and CLAUDE.md, which Claude Code already loads into a
  subagent's context.
- deep-spec-analysis and deep-review put the shared context first, so one run's agents can reuse a
  cached prefix.
- code-quality.md's "explain why, not what" contradicted git-workflow.md and /commit;
  deployment.md's caps NEVER is a plain rule with its reason.
- The root CLAUDE.md now lists the reference docs among the plugin's generated files and in the
  rebuild rule (decision 0019).

Plugin copies regenerated with scripts/build-aplyca-adf.sh; static evals pass.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
deployment.md said "Tests must pass before merge"; git-workflow.md called red checks "information".
Both now say: tests pass before merge; a failure the reviewer accepts is explained in the pull
request. The human review stays the gate. Resolves the prompt audit's L1 flag.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@mauricios
mauricios marked this pull request as ready for review October 6, 2026 04:48
@mauricios
mauricios merged commit ff0fdd5 into main Oct 6, 2026
1 check passed
@mauricios
mauricios deleted the fix/prompt-audit-drift branch October 6, 2026 04:48
mauricios added a commit that referenced this pull request Oct 6, 2026
A patch release: fixes to the framework's instruction files (#35), with nothing new to adopt.
Unreleased becomes v1.2.1, with the upgrade from v1.2.0: /aplyca-adf:upgrade moves the pin; a
packaged project updates the three rule files it commits.

plugin.json 1.2.1; both READMEs name the release.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant