Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 3 additions & 1 deletion AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -60,7 +60,9 @@ each tier is a superset of the confidence of the one above it, so a deep change
runs all three.

**1. cheap — always, before every commit that touches `plugins/**` or `evals/**`.**
Deterministic, offline, free, under a second:
Deterministic, offline, free. ~14s for 1296 checks across 25 plugins, measured
2026-09-18 — it scales with the plugin count, so re-measure rather than trusting
this number:

```sh
evals/cheap/run.sh # exit 0 required to commit
Expand Down
2 changes: 1 addition & 1 deletion docs/research/agentic-patterns-corpus.json

Large diffs are not rendered by default.

6 changes: 3 additions & 3 deletions docs/research/agentic-patterns-corpus.md
Original file line number Diff line number Diff line change
Expand Up @@ -94,7 +94,7 @@ Adoption alone is not fit. These were rejected with reasons:
- Per-turn predictive model routing (GPT-5-style routers) — universal in consumer products, but a mis-route has no recovery path and per-turn switching destroys cache affinity (up to 12.5x prefix swing). Pin model and effort at BEGIN, hold to END
- Semantic/vector response caching — widely marketed, but the coding-agent adoption claim has no primary source and agent steps are not repeated queries; the real practice is prefix/KV stability, already absorbed by criteria-pin
- TDD-Agent dual-track refinement — test-first is confirmed and already Red Gate's BEGIN, but letting the agent refine the tests alongside the code is the exact reward-hacking mutation control exists to block. Name the rejection in the protocol
- Graph/knowledge-graph memory at round level — 3-5x the cost of flat RAG and needs a hand-built ontology; a repo already has git, grep and a type checker as a better graph. Keep it out of rounds entirely
- Graph/knowledge-graph memory at round level — 3-5x the cost of flat RAG and needs a hand-built ontology; a repo already has git, grep and a type checker as a better graph. Keep it out of rounds entirely [Revisited 2026-09-10 in [`harness-knowledge-graph.md`](harness-knowledge-graph.md): upheld, and now the load-bearing reason — but scoped to git-native repo data, not universal.]
- Self-modifying agent archives (Darwin Gödel Machine) — adopt the two invariants it proves by violating them, never the mechanism; a harness that can rewrite its own checker has no gate
- Concurrent fast-path/CoT racing — an inference-layer technique with no agent-orchestration deployments; racing two MIDDLE writers breaks single-writer and doubles cost to shave seconds off a human-gated round

Expand Down Expand Up @@ -677,12 +677,12 @@ Scout claims that did not survive verification are called out in the implication
**Knowledge-graph + agentic RAG (GRAG-ProSafe class)**
*Mechanism:* Four-stage LLM extraction turns unstructured reports into a dynamic knowledge graph, then multi-hop retrieval plus chain-of-thought reasoning answers causal questions over it. GRAG-ProSafe built 1637 nodes / 2285 edges from 198 iron-and-steel accident reports, scoring 0.868 faithfulness, 0.824 answer relevancy, 0.805 factual correctness.
*Why leaders use it:* Multi-hop causal questions that flat vector RAG cannot answer in knowledge-dense, audit-bound domains such as industrial safety and root-cause analysis.
*Failure mode:* Adoption evidence does not hold up: this is a single Expert Systems with Applications paper on one 198-document corpus, not deployed production practice across leaders.
*Failure mode:* Adoption evidence does not hold up: this is a single Expert Systems with Applications paper on one 198-document corpus, not deployed production practice across leaders. [Revisited 2026-09-10 in [`harness-knowledge-graph.md`](harness-knowledge-graph.md): this evidence-quality objection no longer holds — Harness ships a production software-delivery knowledge graph at enterprise scale. The domain-fit objection in the rejected list stands.]
*Red Gate fit:* Should NOT enter Red Gate's loop — graph construction cost dwarfs the payoff at 24-skill scale. The only plausible use is DETECT recurrence over accumulated exhaust, and a flat index over dev-diary entries reaches that far more cheaply.
*Sources:* https://www.sciencedirect.com/science/article/abs/pii/S0957417425035626

**Implications:**
- Nothing here is outright vapor, but two are demoted. Knowledge-graph agentic RAG rests on one 198-document academic system, not leader adoption — treat as research, not roadmap. Circuit breakers are a blog-sourced restatement of what stop-rule and the budget pool already do.
- Nothing here is outright vapor, but two are demoted. Knowledge-graph agentic RAG rests on one 198-document academic system, not leader adoption — treat as research, not roadmap. [Revisited 2026-09-10 in [`harness-knowledge-graph.md`](harness-knowledge-graph.md): this evidence-quality objection no longer holds — Harness ships a production software-delivery knowledge graph at enterprise scale. The domain-fit objection in the rejected list stands.] Circuit breakers are a blog-sourced restatement of what stop-rule and the budget pool already do.
- Three scout citations do not hold up and were corrected: AgentTrace is arXiv 2602.10133 (not 2604.26152); STORM is a May 2026 research system, not a shipping OpenAI Agents SDK feature — the scout conflated it with SDK handoffs; and the NVIDIA/Gretel price was reported as nine figures above a $320M valuation, terms undisclosed.
- The two highest-value absorptions are structural, not additive. (1) Make the between-round human gate a durable checkpoint and adopt LangGraph's replay discipline — read-only before the gate, writes after — so a resumed round cannot double-apply side effects. (2) Turn EMIT exhaust into OTel-shaped spans keyed to the pinned verifier id, so CONSOLIDATE and DETECT operate on structured traces instead of diary prose.
- Cost is a harness property, not a model choice: 41% blended cost and 38% token reduction came from swapping orchestration alone. Red Gate's pointer envelopes should be governed by an explicit cache contract — stable prefix for criteria and skill text, 4 breakpoints budgeted, and an acknowledgement that human gates exceed the 5-minute TTL.
Expand Down
Loading
Loading