Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
32 commits
Select commit Hold shift + click to select a range
06772d8
feat: add jori coordination plugin
JRichlen Sep 5, 2026
80b9096
docs(jori): add sourced model routing guidance
JRichlen Sep 5, 2026
5902af6
Merge branch 'main' into feat/jori-marketplace-main
JRichlen Sep 6, 2026
53dbc1d
Apply batched suggestions from code review
JRichlen Sep 6, 2026
8e9a2f2
Merge branch 'main' into feat/jori-marketplace-main
JRichlen Sep 6, 2026
cc5a906
fix(jori): clarify GitHub routing controls
JRichlen Sep 5, 2026
1c93242
docs(examples): correct marketplace coverage counts
JRichlen Sep 6, 2026
2567cf0
feat(evals): stage GLM subject and price monitor
JRichlen Sep 6, 2026
b12d894
Merge branch 'main' into feat/jori-marketplace-main
JRichlen Sep 6, 2026
f8256c9
Merge branch 'main' into feat/jori-marketplace-main
JRichlen Sep 6, 2026
7ca489e
Merge remote-tracking branch 'origin/feat/jori-marketplace-main' into…
JRichlen Sep 6, 2026
6d5342c
redgate: make criteria-index run ordering locale-stable (LC_ALL=C sort)
JRichlen Sep 6, 2026
cb64d58
redgate: point hooks.json at hooks/hooks-handlers/ where the handlers…
JRichlen Sep 6, 2026
8d18b50
evals/agentic: core, measurement and protocol lanes (wave 1, T01-T10,…
JRichlen Sep 6, 2026
456df7d
evals/agentic + evals/redteam: registry/corpus, native adapters, red-…
JRichlen Sep 6, 2026
0f2fb0a
evals/agentic: integration lane — public CLI, catalog index, lifecycl…
JRichlen Sep 7, 2026
3a9b3a3
evals/agentic + evals/redteam: independent adversarial review and rep…
JRichlen Sep 7, 2026
18a9231
ci: install the pinned tooling the agentic and red-team gates verify …
JRichlen Sep 7, 2026
bc679cb
evals/agentic: deepen graveyard, redgate and egress-gate to 8 cards e…
JRichlen Sep 7, 2026
c57199e
ci/adapters: surface the CLI help text when flag conformance cannot r…
JRichlen Sep 7, 2026
3d7075d
ci: expose node on the driver's allowlisted PATH so the npm codex wra…
JRichlen Sep 7, 2026
354dd01
adapters: resolve a driver's binary by its declared basename, not the…
JRichlen Sep 7, 2026
d1b2f4f
evals: first native evidence (T26-T29 live on Claude Code 2.1.263) an…
JRichlen Sep 8, 2026
f6905cb
fix(evals): verify task artifacts and preserve honest test outcomes
JRichlen Sep 9, 2026
ec8d22d
fix(evals): repair scoring, roster coverage, and logger lifecycle
JRichlen Sep 9, 2026
0e7277a
Merge current main with verified test repairs
JRichlen Sep 9, 2026
5b306e9
test(evals): separate plan structure from live roster coverage
JRichlen Sep 9, 2026
ba3e913
fix(evals): bound paid CI concurrency and clarify grading contracts
JRichlen Sep 9, 2026
1d32202
Merge main into feat/agentic-test-framework
claude Sep 25, 2026
192f2a7
evals: repin fleet-playbook-curator task input after main's #138
claude Sep 25, 2026
814f4ff
jori evals: grade calibration controls as quoted element checklists
claude Sep 25, 2026
656f7f0
Merge remote-tracking branch 'origin/main' into feat/agentic-test-fra…
claude Sep 25, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
The table of contents is too big for display.
Diff view
Diff view
  •  
  •  
  •  
9 changes: 9 additions & 0 deletions .claude-plugin/marketplace.json
Original file line number Diff line number Diff line change
Expand Up @@ -435,6 +435,15 @@
"meta"
]
},
{
"name": "jori",
"description": "Coordinate complex work with bounded agents, evidence, and proportional dashboards.",
"version": "0.0.1",
"source": "./plugins/jori",
"author": { "name": "Jordan Richlen" },
"license": "MIT",
"keywords": ["orchestration", "multi-agent", "coordination", "evidence", "dashboards"]
},
{
"name": "agent-compiler",
"description": "Compile deterministic, content-hashed agents from small behavior modules: a skill turns fuzzy intent into a typed AgentQuery, a stdlib-only kernel resolves modules, expands dependencies, fails closed on conflicts and over-ceiling effects, and emits an immutable AgentImage plus a rendered harness-native subagent — never inventing behavior text without provenance.",
Expand Down
110 changes: 106 additions & 4 deletions .github/workflows/evals.yml
Original file line number Diff line number Diff line change
Expand Up @@ -39,8 +39,64 @@ jobs:
# Full history: the timeline guard verifies /commit/<sha> receipts
# against the object store, which a depth-1 checkout cannot see.
fetch-depth: 0
- uses: actions/setup-node@v4
with:
node-version: 22
- name: install the pinned offline tooling the agentic and red-team gates verify against
# The always-on gate is host-portable, not host-free: evals/redteam checks
# the promptfoo 0.122.0 pin by package.json + dist-manifest digest at
# PROMPTFOO_HOME, and the agentic adapter lane reads the INSTALLED
# claude/codex --help to refuse any driver flag the binaries do not
# support (T25). Install exactly the pinned versions into the runner's
# temp dir (never npx, never @latest) and export the two locations. No
# login, no model call: --help and --version are all the gate touches.
run: |
set -euo pipefail
sudo apt-get update -qq
sudo apt-get install -y bubblewrap patch
mkdir -p "$RUNNER_TEMP/pinned-tools" && cd "$RUNNER_TEMP/pinned-tools"
npm init -y >/dev/null
npm install --no-audit --no-fund --no-save promptfoo@0.122.0 @anthropic-ai/claude-code@2.1.263 @openai/codex@0.153.4 winston@3.19.0 winston-transport@4.9.0
echo "PROMPTFOO_HOME=$RUNNER_TEMP/pinned-tools/node_modules/promptfoo" >> "$GITHUB_ENV"
echo "$RUNNER_TEMP/pinned-tools/node_modules/.bin" >> "$GITHUB_PATH"
# The adapter driver runs each CLI under an env ALLOWLIST whose PATH is
# /usr/bin:/bin (contract 10.1) -- it never inherits the runner's PATH.
# The npm codex wrapper is a `#!/usr/bin/env node` script, so node must
# be reachable on that allowlisted PATH; expose the setup-node binary
# there (a provisioning step, not a framework change).
sudo ln -sf "$(command -v node)" /usr/bin/node
# Diagnostics only: prove the three binaries answer --version/--help here.
export PATH="$RUNNER_TEMP/pinned-tools/node_modules/.bin:$PATH"
node "$RUNNER_TEMP/pinned-tools/node_modules/promptfoo/dist/src/entrypoint.js" --version
claude --version
codex --version || true
codex exec --help 2>&1 | head -5 || true
- name: apply the pinned logger lifecycle correction
run: python3 evals/redteam/bin/provision-logger.py --promptfoo-home "$PROMPTFOO_HOME" --apply
- name: provision scoped bubblewrap namespaces and prove the positive baseline
run: bash ci/provision-bubblewrap.sh
- name: run cheap eval gate
run: evals/cheap/run.sh
- name: verify task grading, uncertainty, and computation isolation
run: |
python3 -m unittest \
evals.agentic.tests.test_corpus \
evals.agentic.tests.test_controls \
evals.agentic.tests.test_corpus_integrity \
evals.agentic.tests.test_grader_forgery \
evals.agentic.tests.test_task_contracts \
evals.agentic.tests.test_eval_ladder \
evals.agentic.tests.test_effect_observer \
evals.agentic.tests.test_redteam_inference \
evals.agentic.tests.test_redteam_effect_validity \
evals.agentic.tests.test_redteam_task_validity \
evals.agentic.tests.test_redteam_task_exposure \
evals.agentic.tests.test_redteam_task_tools \
evals.agentic.tests.test_redteam_tranche_tools \
evals.agentic.tests.test_redteam_grader_faults \
evals.agentic.tests.test_redteam_logger \
evals.agentic.tests.test_redteam_design.TwoByTwoCompleteness \
evals.agentic.tests.test_redteam_controls.SafeVulnerableControls

# ── counterfeit tier ──────────────────────────────────────────────────────
# A green cheap tier only shows that GOOD plugins pass. The counterfeit corpus
Expand All @@ -67,6 +123,10 @@ jobs:
corpus:
- 'evals/cheap/**'
- 'evals/counterfeits/**'
- 'evals/agentic/**'
- 'evals/redteam/**'
- 'ci/provision-bubblewrap.sh'
- '.github/workflows/evals.yml'
- 'plugins/**'

counterfeit-run:
Expand All @@ -76,6 +136,42 @@ jobs:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with:
node-version: 22
- name: install the pinned offline tooling the agentic and red-team gates verify against
# The always-on gate is host-portable, not host-free: evals/redteam checks
# the promptfoo 0.122.0 pin by package.json + dist-manifest digest at
# PROMPTFOO_HOME, and the agentic adapter lane reads the INSTALLED
# claude/codex --help to refuse any driver flag the binaries do not
# support (T25). Install exactly the pinned versions into the runner's
# temp dir (never npx, never @latest) and export the two locations. No
# login, no model call: --help and --version are all the gate touches.
run: |
set -euo pipefail
sudo apt-get update -qq
sudo apt-get install -y bubblewrap patch
mkdir -p "$RUNNER_TEMP/pinned-tools" && cd "$RUNNER_TEMP/pinned-tools"
npm init -y >/dev/null
npm install --no-audit --no-fund --no-save promptfoo@0.122.0 @anthropic-ai/claude-code@2.1.263 @openai/codex@0.153.4 winston@3.19.0 winston-transport@4.9.0
echo "PROMPTFOO_HOME=$RUNNER_TEMP/pinned-tools/node_modules/promptfoo" >> "$GITHUB_ENV"
echo "$RUNNER_TEMP/pinned-tools/node_modules/.bin" >> "$GITHUB_PATH"
# The adapter driver runs each CLI under an env ALLOWLIST whose PATH is
# /usr/bin:/bin (contract 10.1) -- it never inherits the runner's PATH.
# The npm codex wrapper is a `#!/usr/bin/env node` script, so node must
# be reachable on that allowlisted PATH; expose the setup-node binary
# there (a provisioning step, not a framework change).
sudo ln -sf "$(command -v node)" /usr/bin/node
# Diagnostics only: prove the three binaries answer --version/--help here.
export PATH="$RUNNER_TEMP/pinned-tools/node_modules/.bin:$PATH"
node "$RUNNER_TEMP/pinned-tools/node_modules/promptfoo/dist/src/entrypoint.js" --version
claude --version
codex --version || true
codex exec --help 2>&1 | head -5 || true
- name: apply the pinned logger lifecycle correction
run: python3 evals/redteam/bin/provision-logger.py --promptfoo-home "$PROMPTFOO_HOME" --apply
- name: provision scoped bubblewrap namespaces and prove the positive baseline
run: bash ci/provision-bubblewrap.sh
- name: prove the cheap gate discriminates (counterfeit corpus)
run: evals/counterfeits/run.sh

Expand Down Expand Up @@ -306,9 +402,15 @@ jobs:

behavioral-run:
name: behavioral tier — promptfoo (${{ matrix.plugin }})
needs: [behavioral-detect, grader-model]
needs: [behavioral-detect, grader-model, routing-eval]
# Share the subject-provider budget with routing without overlapping calls.
# !cancelled() overrides the implicit success() dependency gate: a routing
# FAILURE must not skip these independent behavioral measurements.
# No packs → nothing to run. Fork PRs have no secrets, same as the other tiers.
if: >-
!cancelled() &&
needs.behavioral-detect.result == 'success' &&
needs.grader-model.result == 'success' &&
needs.behavioral-detect.outputs.plugins != '[]' &&
(github.event_name != 'pull_request' ||
(github.event.pull_request.head.repo.full_name == github.repository &&
Expand Down Expand Up @@ -402,7 +504,7 @@ jobs:
# real arbiter: it excludes FAULTs and scores a per-scenario floor over the
# valid samples. A total crash that writes no results.json still fails,
# because pass-rate.sh exits non-zero on a missing/unreadable file.
run: npx --yes promptfoo@0.122.0 eval --output results.json || echo "::warning::promptfoo exited non-zero — the statistical gate scores results.json and decides this leg"
run: npx --yes promptfoo@0.122.0 eval --max-concurrency 1 --output results.json || echo "::warning::promptfoo exited non-zero — the statistical gate scores results.json and decides this leg"
- name: statistical gate (per-scenario pass rate, FAULT-excluding)
if: steps.touched.outputs.paid == 'true'
# repeat:3 in each pack expands every scenario into 3 rows. This enforces a
Expand Down Expand Up @@ -597,7 +699,7 @@ jobs:
# distinction unreachable. Only a missing/empty results.json (the eval
# itself never ran) fails this step.
run: |
npx --yes promptfoo@0.122.0 eval -c promptfooconfig.yaml --output results.json \
npx --yes promptfoo@0.122.0 eval --max-concurrency 1 -c promptfooconfig.yaml --output results.json \
|| echo "promptfoo exited non-zero — verdict delegated to the statistical gate"
test -s results.json
- name: promptfoo eval (redgate trajectory)
Expand All @@ -611,7 +713,7 @@ jobs:
OPENROUTER_API_KEY: ${{ secrets.OPENROUTER_API_KEY }}
PROMPTFOO_RETRY_5XX: "true"
run: |
npx --yes promptfoo@0.122.0 eval -c promptfooconfig.yaml --output trajectory-results.json \
npx --yes promptfoo@0.122.0 eval --max-concurrency 1 -c promptfooconfig.yaml --output trajectory-results.json \
|| echo "promptfoo exited non-zero — verdict delegated to the statistical gate"
test -s trajectory-results.json
- name: statistical gate (per-scenario k-of-N pass rate)
Expand Down
79 changes: 79 additions & 0 deletions .github/workflows/model-pricing.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,79 @@
name: model-pricing

on:
schedule:
- cron: '25 13 * * *'
workflow_dispatch:
pull_request:
paths:
- '.github/workflows/model-pricing.yml'
- 'ci/model-pricing/**'

permissions:
contents: read

concurrency:
group: model-pricing-state
cancel-in-progress: false

jobs:
test:
name: model pricing — offline controls
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
with:
persist-credentials: false
- name: verify detector and budget invariants without network
run: python3 ci/model-pricing/test_monitor.py

monitor:
name: model pricing — detect and stage
needs: test
# No PR credentials/inference. Dispatch from a feature branch is refused.
if: >-
github.event_name != 'pull_request' &&
github.ref == format('refs/heads/{0}', github.event.repository.default_branch) &&
vars.OPENROUTER_PRICE_MONITOR_ENABLED == 'true'
runs-on: ubuntu-latest
timeout-minutes: 10
permissions:
contents: write # only the dedicated state branch; never main or a PR
steps:
- uses: actions/checkout@v4
- name: read public prices and persist evidence
run: python3 ci/model-pricing/monitor.py scan
- name: run authorized bounded strategy on pending material changes
# policy.json is an additional default-off gate, with zero spend shipped.
env:
PRICE_STRATEGY_KEY: ${{ secrets.PRICE_STRATEGY_KEY }}
run: python3 ci/model-pricing/monitor.py review
- name: summarize actionable change
if: always()
env:
JOB_STATUS: ${{ job.status }}
run: |
python3 - <<'PY'
import json, os
from pathlib import Path
root = Path('work/model-pricing')
status = json.loads((root / 'status.json').read_text()) if (root / 'status.json').exists() else {}
if os.environ['JOB_STATUS'] != 'success':
message = 'Price monitor stopped. Inspect sanitized logs and restore missing state or authorization; no automatic retry or promotion.'
elif (root / 'proposal.json').exists():
message = 'A bounded agent strategy and config-change plan are staged in the model-pricing artifact. Quality is unvalidated; human review and existing gates are required. No settings were applied.'
elif status.get('status') == 'material_change':
message = 'Material price/control changes are staged in the model-pricing artifact. Agent analysis remains subject to its separate budget and provider authorization.'
else:
message = ''
if message:
with open(os.environ['GITHUB_STEP_SUMMARY'], 'a') as stream:
stream.write(message + '\n')
PY
- uses: actions/upload-artifact@v4
if: always()
with:
name: model-pricing-${{ github.run_id }}
path: work/model-pricing/*.json
retention-days: 90
if-no-files-found: ignore
11 changes: 11 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,10 @@
/delete-originals.sh
/summary.tsv
*.bundle
# Fixed, synthetic full-history archives are required inputs for the archive
# restoration test. Keep runtime bundles ignored everywhere else.
!/evals/agentic/tasks/graveyard/graveyard-pos-02/task/legacy-service-source.bundle
!/evals/agentic/tasks/graveyard/graveyard-pos-02/fixtures/pass/archive/legacy-service.bundle
.DS_Store

# Eval artifacts. Secrets and generated output must never be committed.
Expand Down Expand Up @@ -39,3 +43,10 @@ plugins/*/evals/promptfoo/results.matrix.json
plugins/*/evals/promptfoo/promptfooconfig.regrade.*.yaml
plugins/*/evals/promptfoo/replay.*.json
plugins/*/evals/promptfoo/results.regrade.*.json

# Agentic test framework run manifests (evals/agentic/run.py, contract §9) — emitted per run, never committed.
evals/agentic/manifests/runs/
*.out.json

# red-team runner scratch (promptfoo home, logs, results) — generated, never committed
evals/redteam/.artifacts/
1 change: 1 addition & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -46,6 +46,7 @@ multiple machines and want it to resolve identically every time.
| [**stop-rule**](plugins/stop-rule/) | A halting discipline for iterative fix loops: declare an attempt bound up front, count honestly, and at the bound stop and report state with ranked hypotheses — never attempt N+1 on momentum. |
| [**redgate**](plugins/redgate/) | Run any idea through Red Gate: rounds of ARM/TRACE/JUDGE with graduated autonomy — round gates classified PATCH/MINOR/MAJOR via semver-gate, so derived work auto-passes inside a human-approved mandate while scout decisions, plan approval, and irreversible actions always block on the human. Cross-harness (Claude Code, Codex, Copilot via APM); the red gate is executed, not asked. |
| [**recurrence-detector**](plugins/recurrence-detector/) | Close the growth loop's DETECT step: cluster the exhaust every run sheds (stop-reports, findings, unmet criteria, diary entries) by failure shape, and surface any shape seen at least 3 times as a named candidate invariant with its sightings cited. Proposes; never scaffolds. |
| [**jori**](plugins/jori/) | Coordinate complex work with bounded agents, evidence, and proportional dashboards. |
| [**agent-compiler**](plugins/agent-compiler/) | Compile deterministic, content-hashed agents from small behavior modules: fuzzy intent becomes a typed AgentQuery, then a stdlib-only kernel resolves modules, fails closed on conflicts and over-ceiling effects, and emits an immutable AgentImage with per-line provenance. |
| [**eval-ladder**](plugins/eval-ladder/) | Design and audit an agent system's eval ladder — the cheapest rung that catches each regression, the blind spot beside every green, judges validated by TPR/TNR, pass^k for irreversible actions. Use when designing, auditing, or defending a test/eval strategy for an agent, skill, or prompt; when adding an eval tier or LLM judge; or when a suite is all-green and you cannot say what it would catch. |

Expand Down
Loading
Loading