When a supplier delivery fails, the cakes made from it are already promised to customers. PromisePatch works out which promises the failure breaks. It fixes only the ones it has permission to fix, hands the rest to the owner, and leaves every other order alone.
A made-to-order bakery takes orders against supplies it expects to receive. One morning the raspberries don't arrive.
Reordering raspberries is the easy part. The hard part is the promises already made to customers. Some orders can be changed because the customer already agreed to a swap. Some can be changed only if the customer says yes now. Some cannot be changed at all and need the owner. Most are not affected and must not be touched. Get one wrong and a customer receives a cake they refused, or a message about an order that was fine.
The Hollow Oak bakery has six orders for today. A worker says "today's raspberry delivery didn't arrive" and clarifies "just raspberries — the strawberries came."
| order | what PromisePatch finds | permission | outcome |
|---|---|---|---|
| A | a swap the customer's standing preference already allows | already given | recovered: amended automatically |
| B | a swap the customer has not agreed to yet | must ask | recovered: one Telegram message, YES on a signed link, re-checked, then amended |
| C | the customer forbids substitutions | none | owner: escalated, scheduled work held |
| D | no approved alternative recipe exists | none | owner: escalated, scheduled work held |
| E | no raspberries in it | not needed | untouched |
| F | no raspberries in it | not needed | untouched |
Untouched means zero effects: no message, no write, no reservation change, no task hold and no audit event for E or F. That held in each of five restart rehearsals on the deployed app (bridge-release.md, funnel). The bakery, its customers and the order system are a labelled fixture and a simulator. The Telegram message and the web approval are real.
Finding a possible substitution is not the same as having permission to apply it.
Any capable agent can find a recipe that works without raspberries. What this repository demonstrates is what PromisePatch does between that candidate and a write to an order:
| step | what PromisePatch does |
|---|---|
| candidate action | the model reads the report; a pre-authored SubstitutionPolicy names the only versions a recovery may use |
| permission | each affected order is checked for what covers it: a recorded preference, a customer's literal YES, or nothing, in which case it goes to the owner |
| human authority | the plan is applied only after a person approves that exact plan; a service credential cannot create that approval |
| fresh-state revalidation | before writing, the worker re-reads current state and runs ten checks |
| write or refusal | it amends the order, or refuses with a named reason such as STALE and writes nothing |
The model only understands. Deterministic rules and fresh state authorize. This is a claim about what PromisePatch does, shown in the evidence below, not about what other agents cannot do.
A recovery being valid once does not mean it stays valid. Before it writes, PromisePatch re-reads fresh state and runs ten checks. If the world has changed, the old permission is rejected.
Taken live on the deployed release on 2026-10-02 (revalidation-proof.md):
| order B | |
|---|---|
1. the customer says YES |
to swapping one cake, order version v1 |
| 2. the order changes | the order system now says two cakes, version v2 |
| 3. PromisePatch re-reads fresh state | check 2, order state and version unchanged, fails: expected v1, found v2 |
| 4. the stale yes is rejected | STALE. No write: the order system keeps its own v2, unamended |
| 5. what follows | the customer is told the request no longer applies; B is re-planned and, with no new plan confirmation, went to the owner |
A control run with the same steps and no order change passed all ten checks and amended B from v1 to v2. The change was a deliberate edit in the simulated order system, and the worker was paused so the yes, the change and the re-check happened in that order. The live run exercised check 2. The other checks refusing, and the later commit-time gate, are proved by tests.
A worker can drive the case by conversation. Each layer has one job, and only the last one can authorize anything.
| layer | its job | can it authorize? |
|---|---|---|
| Amazon Bedrock (Nova 2 Lite) | understands the worker's sentence and picks one of five tools: report, clarify, confirm, withdraw, status | no |
| Simulated Alexa+ via MCP | carries that action to the real, authenticated MCP endpoint | no: it can spend a human approval that already exists, and cannot create one |
| Deterministic rules and fresh state | decide whether the action is actually allowed right now | yes |
Taken live on the deployed release on 2026-10-02 (alexa-mcp-confirm-proof.md):
- The worker approves the plan in the browser. This writes the one human approval.
- The worker types "Yes, go ahead." into the Simulated Alexa+ panel.
- Bedrock selects
CONFIRM. - The real MCP
confirmspends that browser approval. Afterwards there is still exactly one approval, attributed to the worker, not to the MCP credential. - The workflow proceeds: A is amended, B's customer is asked, C and D go to the owner, and E and F are untouched.
Bedrock choosing CONFIRM is not enough on its own. The server also checks that the worker's words
are a plain yes. In an earlier live turn, "Yes, confirm the plan" was refused for that reason, and
nothing was confirmed (bridge-release.md). That MCP cannot create an
approval is enforced by design (ADR-0018)
and proved in CI. The live run shows the success path.
This is not a native Alexa+ integration, and none is claimed. The Alexa+ experience is simulated: a panel in the app drives a case-scoped MCP client on the server (ADR-0028). That client calls the same MCP endpoint any Alexa+ agent or other MCP client would call.
| real | simulated or constructed |
|---|---|
| The AWS deployment: EC2, private encrypted RDS PostgreSQL, a Let's Encrypt certificate | The external order system: the project's own simulator, a separate application with its own store, not a real point of sale |
| Amazon Bedrock calls, made with the instance role | The Alexa+ experience: simulated by an MCP client on the server, not a native Alexa+ skill |
The MCP endpoint: Streamable HTTP, protocol 2025-11-25, bearer-authenticated |
The bakery: its customers, recipes and orders are a labelled fixture |
| Telegram messages delivered to a real phone, and signed customer approval links | The order change in the stale-yes proof, made on purpose |
| Persistence, the durable worker, revalidation and authorization |
The model understands; the deterministic protocol authorizes. Every rule below is enforced in code and tests, not by convention. Import-linter contracts stop the model boundary, the MCP server and the conversational client from reaching the domain or the database.
- Model output is never authority. The model only proposes a reading of a sentence, which must match the bakery's own vocabulary. It cannot write a row, record consent, approve a plan or choose a recovery. The canonical raspberry report costs zero model calls, and a test asserts it.
- Worker plan approval and customer consent are different things, with different parsers, records and words. A customer's yes never spends a worker approval, and a worker's approval never counts as a customer's yes.
- Customer consent is a literal
YESorNO, trimmed and case-insensitive. Any other reply decides nothing: it is stored word for word and gets one confirmation prompt. - A service credential is not a person. Holding the MCP bearer token proves a process. A report
carried over MCP is attested as the server's configured worker (
PP_SURFACE_WORKER_ID), and an MCPconfirmcan only spend an approval a human already wrote, in a signed-in browser session or on the operator console. - A confirmation binds to the plan that was read out. Stale, wrong-case, replayed and repeated confirmations fail closed.
- A yes is perishable. When a customer's answer arrives, ten checks run against a fresh
snapshot. A change that is no longer true is refused as
STALEand nothing is sent; an expired answer goes to the owner; an unauthorized one is refused. The commit and the amendment's first dispatch each judge freshness again (ADR-0024, ADR-0026), and a stale finding there sends nothing and escalates the track to the owner. - Recovery only chooses from recipe versions written in advance. Nothing invents a substitute at runtime.
- Unknown or conflicting state fails closed to
BLOCKED, never toUNAFFECTED. - Work that has started is never reported as stopped. Scheduled work on a blocked promise is held, and work that has already started is escalated to its owner instead.
- A withdrawal is never an undo. It stops future work and reverses no physical fact.
More detail: semantic-boundary.md,
mcp-human-confirmation-boundary.md,
started-work-contract.md,
bounded-withdrawal.md and the ADRs in docs/adr/.
flowchart LR
worker["Bakery worker<br/>browser · voice or text"]
agent["Simulated Alexa+<br/>case-scoped MCP client<br/>(or any MCP client)"]
mcp["MCP server<br/>Streamable HTTP · 2025-11-25<br/>no database access"]
api["Intent API<br/>actor and clock are server-derived"]
model["Semantic boundary<br/>Amazon Bedrock<br/>proposes a reading or a tool"]
engine["promise_graph<br/>reach · partition · revalidate<br/>pure, deterministic"]
wf["Durable workflow<br/>case state machine + step ledger<br/>PostgreSQL"]
oms["External order system<br/>simulated · system of record"]
tg["Telegram<br/>outbound message"]
link["Signed web link<br/>literal YES or NO"]
customer(("Customer"))
worker -->|"panel turn"| agent
agent -. "which tool?" .-> model
agent -->|bearer token| mcp -->|service token| api
worker -->|"session: one of two<br/>plan-approval channels"| api
api --> wf
wf -. "the words" .-> model
model -. "a candidate reading,<br/>never authority" .-> wf
wf <--> engine
wf -->|"governed amendment,<br/>after revalidation"| oms
oms -->|signed events| wf
wf --> tg --> customer --> link --> wf
classDef authority fill:#1B2644,stroke:#8390F2,stroke-width:2px,color:#F7F8FC
classDef understanding fill:#1B2644,stroke:#AEB6C8,stroke-dasharray:5 4,color:#F7F8FC
classDef outside fill:#111A2B,stroke:#76819A,color:#F7F8FC
class engine,wf,api authority
class model understanding
class worker,agent,mcp,oms,tg,link,customer outside
Solid indigo borders mark where authority lives. The dashed node is understanding only.
promise_graph(packages/promise-graph) is a pure package. It handles reachability, temporal availability, allocation, impact classification, recovery validation, snapshot fingerprints and the revalidation checklist. It does no I/O, reads no environment and never calls the clock, so every customer-affecting decision can be tested without the cloud.- The backend (
apps/backend) runs asapi,workerandmcp: a persisted case state machine with a step ledger and an audited PostgreSQL write boundary. The worker is stateless: restarting it is the recovery mechanism, and outstanding work resumes from its rows. - The external order system (
apps/order-simulator) is a separate application with its own store. PromisePatch mirrors it and pushes governed amendments; neither reads the other's storage (order-system.md). - The case workspace (
apps/frontend) renders the case and never decides. - MCP exposes the five intent tools over an authenticated endpoint, pinned to protocol revision
2025-11-25by a test. Unauthenticated callers are refused before the protocol layer, and unlistedOriginandHostvalues are rejected (p5.1-mcp-transport-spine.md, p5.2-mcp-clarification-and-confirmation.md). In the browser, voice uses the browser's own speech recognition and reaches the same services as a typed turn. No spoken phrase carries authority a typed one could not.
The live app at https://184.194.40.87.sslip.io runs one EC2 t4g.small in us-east-1,
behind Caddy with a Let's Encrypt certificate, against a private, encrypted RDS PostgreSQL. IMDSv2
is required, and the instance role calls Amazon Bedrock, so no AWS key is held anywhere.
| release | 740a062838e0ea2620499abed27d653c42fc05f7, image 740a062838e0, as GET /healthz reported it on 2026-10-02. Repository HEAD may contain documentation-only commits on top of it |
| product gate | pr run 36925136266, 13 of 13 jobs on that exact SHA, the whole-stack browser suite included |
| revalidated on it | the Alexa+ bridge verified live with a real Bedrock turn over the real MCP endpoint; the v2 effect set 16/16; five deployed restart rehearsals, each PASS; the local demo contract, 47 assertions (bridge-release.md) |
| proved on it afterwards | a live stale-yes refusal (revalidation-proof.md) and a live MCP confirm spending a browser approval (alexa-mcp-confirm-proof.md) |
| customer channel | Telegram outbound is live, one message per rehearsal; customers answer on the signed web link (deployed-customer-channel.md) |
Each number has a caveat, and the caveat is part of the result.
| result | what it is | record |
|---|---|---|
| 0/2 | Untouched orders that received any effect, in each of five deployed rehearsals. Each restarted the worker at a different point: while waiting for the customer, across the plan confirmation, across the answer, after resolution, and the first again. Per rehearsal: 2 amendments, 1 customer message, 2 task holds, 3 outbox rows, all delivered on attempt 1. | g8-demo-funnel.md, repeated in bridge-release.md |
| 16/16 | The v2 release condition, taken once on the current release. v2 is a separately versioned label correction of the effect-set manifest in which one label moved: S12's hold, because started kitchen work is never held (ADR-0017). It is not a re-score of v1. | g8-effect-set-release-condition.md, bridge-release.md |
| 11/16 | The permanent headline: the first scored run of the sixteen frozen scenarios against the v1 manifest, labelled by hand before the runner existed. Five failed. Four were implementation defects, since fixed under published SHAs; the fifth is the S12 label above. This result is never replaced. | effect-set-first-scored-run.md |
| 9/10 | Voice turns in which a truthful spoken response began within four seconds of speech ending. It met the predeclared threshold of K ≥ 9 exactly. | g7-ten-turn-voice-measurement.md |
What those numbers do not say:
- The effect sets are developer-authored, finite and public. They are not an independent or held-out benchmark. The effect-set CI workflow stays red on purpose, because it judges v1.
- The funnel is one fixture measured five times per release. It shows the demo repeats, not a reliability rate.
- The voice result comes from a second run. The first run (
K = 1/10) was voided after its intervals had been computed, which the predeclared protocol forbids, and the claims audit says a strict reader may treat run 2 as a best-of-two. Both runs are published in full. They ran on the local stack, with no public-internet round trip, using the browser's own speech APIs.
Earlier releases (historical, not current)
| release | commit / image | product gate (pr) |
record |
|---|---|---|---|
| post-intake release | 283f63f2845f8c5e93b2a791eebc15bc4de3f4d7 / 283f63f2845f |
run 36759222324, 13 of 13 |
post-intake-release.md: migration 0010, v2 16/16, five deployed rehearsals |
| G8 freeze | 56c302366b3ddc0d824c1588a4a9ddbd193ed891 / 4529a802e34e |
run 36310794944, 13 of 13 |
g8-closeout.md: 22 of 22 rows closed; the first five rehearsals (R1 · R2 · R3 · R4 · R5) |
The G8 freeze of 2026-09-27 was reopened by
ADR-0027
and re-established at 283f63f, then reopened for the ADR-0028 bridge and re-established at
740a062. Earlier deployment history: p6.2-first-deployment.md,
phase7-rc-deployment.md,
phase7-approval-log-privacy-repair.md and
customer-disclosure-hardening.md.
On the live app. Press Look around a real case. It opens a read-only observer session with no account. You can read everything, including the evidence drawer, and change nothing, because the domain refuses every write from that principal (ADR-0016).
From a clone, with nothing but Python and uv. These need no database, no container and no credential:
uv run python scripts/verify_effect_set_manifest.py # recompute the frozen v1 manifest hash
uv run python scripts/run_effect_sets.py --check # prove the clone is complete and intact
uv run pytest packages/promise-graph # the deterministic engine's own suiteThe engine also runs standalone, outside the workspace. See
packages/promise-graph/README.md and its
fresh-clone proof.
The whole storyboard, locally. Bring up the local stack, then run the demo-contract runner. It executes the canonical storyboard as 49 assertions through the browser path (47 when the plan is confirmed on the operator console, as in the current release's local run), through the product's own transports, and reads its evidence in read-only transactions. The fixture and the worker restart stay the operator's actions (g8-demo-contract-runner.md):
PP_INTERNAL_SERVICE_TOKEN="$(grep '^PP_INTERNAL_SERVICE_TOKEN=' docker/env/api.env | cut -d= -f2-)" \
uv run python scripts/with_local_env.py -- \
uv run python scripts/demo_contract.py --api http://127.0.0.1:58000 --order-system http://127.0.0.1:58100| claim | record | what it proves |
|---|---|---|
| a stale yes is refused live | revalidation-proof.md | on 740a062838e0: a customer's v1 yes refused as STALE after the order moved to v2, with no amendment for that order; plus a control run that applied |
| MCP spends, never creates, a human approval | alexa-mcp-confirm-proof.md | on 740a062838e0: browser approval, then "Yes, go ahead.", Bedrock CONFIRM, real MCP confirm; still one approval; no customer consent created |
| the current release, revalidated | bridge-release.md | 740a062, image 740a062838e0, pr run 36925136266 13 of 13; bridge verified live; v2 16/16; R1–R5 PASS; demo contract 47 |
| untouched means untouched | g8-demo-funnel.md | the funnel 6 → 1/1/2 + 2, and 0/2 untouched orders affected, in all five rehearsals |
| the storyboard is executable | g8-demo-contract-runner.md | 49 assertions through the intent API and the signed link, with no direct consent insert; 47 when confirmed on the console |
| the immutable headline | effect-set-first-scored-run.md | 11/16 against frozen v1, with every diff published |
| the separate release condition | g8-effect-set-release-condition.md | 16/16 against the v2 label correction, and the fix SHA for each v1 failure |
| adversarial faults | g8-adversarial-proof-map.md | all eleven named faults, from a lost MCP response and model self-confirmation to crashes on either side of external acceptance, each proved |
| the voice number | g7-ten-turn-voice-measurement.md | 9/10 in run 2, with void run 1 and every timing published |
| MCP transport | p5.1-mcp-transport-spine.md | Streamable HTTP, 2025-11-25, bearer and Origin/Host refusals, tested with the official SDK |
| real customer loop | deployed-customer-channel.md · customer-approval-link.md | one Telegram delivery and a web YES, revalidated, then EXT-B amended once |
| earlier releases, historical | post-intake-release.md · g8-closeout.md | 283f63f2845f and 4529a802e34e, each with its own CI run and five rehearsals |
| engine from a clean clone | g8-standalone-fresh-clone-proof.md | 335 tests passed from a fresh public clone |
| effect-set clone check | g8-effect-set-fresh-clone-proof.md | uv sync --frozen and both manifests' checks exit 0 |
| development evidence | g8-development-evidence.md | curated, redacted evaluation results, failures kept |
| provenance | g8-contribution-provenance.md | every commit is dated inside the submission window |
| claims against evidence | claims-audit.md | an audit of this repository's own claims, overclaims included |
| integration cost | prerequisites-integration-cost-and-limitations.md | what adopting this would require, and what is not established |
The initial target is small custom-order food businesses, where an ingredient disruption touches customer promises and production work at the same time. This is a hypothesis, not a finding: no pilot users, demand, savings or time reductions are claimed. Production adoption would require a real order and catalog integration, substitution and permission policies authored for that business, a configured customer channel, and onboarding and validation with actual operators. See prerequisites-integration-cost-and-limitations.md.
- One bakery, fixture data, a simulated order system. It is not Square or a production point of sale. The Telegram message and the web approval are the real parts.
- The Alexa+ experience is simulated by an MCP client on the server. No native Alexa+ skill
exists. The live
confirmproof is one run of the success path; the refusal when no approval exists is proved in CI. - One refusal kind has been exercised live, once.
STALE, through check 2, on the deployed release (revalidation-proof.md). The other checks refusing,EXPIRED,UNAUTHORIZED,NOOPand the commit-time freshness gate are proved by tests only. - Telegram inbound is deliberately not built. A second route for the word
YESwould be a second consent parser. The signed link proves possession of the message, not identity. - Telegram's Bot API has no idempotency key. A retry after an uncertain send can deliver a duplicate message. Order amendments carry a stable idempotency key, so a duplicate is a second message, never a second amendment (customer-message-transport.md).
- MCP intake is a trusted reporting channel. A report over MCP is attested as the server's configured worker, not by a person the server authenticated (p5.1-mcp-transport-spine.md).
- The effect sets are developer-authored, and both evaluation holdouts remain sealed.
- The
SUR-1comparative benchmark says nothing comparative about models. Two of its arms called the model zero times (sur1-fifth-scored-run.md). - The voice result carries the caveats above, and it is not a production latency SLA.
- Operations debt is recorded, not fixed.
deploy.sh stackcannot release against the drifted stack template (non-destructive-release.md §10.1). The demo restore needs settings no single deployed container holds (demo-world-restore.md). CloudWatch keeps pre-redaction lines until its 14-day retention expires them. - Deployment smoke shows 9/12 from the operator's machine. A local TLS-intercepting proxy
times out three refusal probes; on the host they answer
401,403and421. - G7's demo-narrative comprehension check was not performed, by the project owner's decision.
- Python 3.12 and uv
- Docker with Compose v2, for the local stack
- Node.js 24, for the frontend
The stack is a disposable PostgreSQL 16 in a Docker volume, the repository's own migrations, the Hollow Oak fixture, the API, the durable workflow worker, the MCP endpoint, the frontend and the External Order System simulator. It needs no hosted database and no cloud account.
uv run python scripts/bootstrap_local_env.pyThat generates docker/env/*.env with fresh credentials. The files are never committed, and
there are no default passwords: read the seeded demo logins out of docker/env/migrate.env.
docker compose up --detach --wait--wait returns only once PostgreSQL is healthy, the migrations have exited zero, the fixture
has loaded, /readyz reports ready and the frontend is serving. Then open
http://localhost:55173 and sign in as maya with PP_DEMO_WORKER_PASSWORD.
docker compose run --rm seed # reload the fixture
docker compose restart worker # restart the worker; outstanding work resumes
docker compose down # stop, keeping the database
docker compose down --volumes # stop and discard the database| Service | On the host | Inside the network |
|---|---|---|
| frontend | http://localhost:55173 | frontend:5173 |
| api | http://localhost:58000 | api:8000 |
| worker | no port; docker compose logs worker |
-- |
| mcp | http://localhost:58001/mcp | mcp:8001 |
| order-simulator | http://localhost:58100 | order-simulator:8100 |
| postgres | 127.0.0.1:55432 |
postgres:5432 |
The host ports are deliberately not 5173, 8000 and 5432: those are usually already taken on a
machine that develops this project. If a port is reserved on your machine, override it with the
PROMISEPATCH_*_PUBLISHED_PORT variables.
A few properties are worth knowing before you use it:
- Neither the API nor the worker holds an administrative credential. Both connect as
promisepatch_app, which owns nothing, migrates nothing and cannot truncate a table. Migrations andpp reset-demo-staterun in separate containers with the administrative connection. That is why there is no HTTP reset endpoint. - The worker is stateless, so restarting it is the recovery mechanism. Every piece of
outstanding work is a row and everything the process holds is a lease;
docker compose restart worker-- or killing it outright -- loses nothing, and the work resumes as the leases expire. Several workers can run at once without coordinating. pp reset-demo-staterecreates the fixture workers, so it signs everyone out. It is an operator command that replaces every domain row PromisePatch owns, and the sessions go with them.pp restore-demo-worldis the whole demo repair, and it is destructive. It reseeds, resets the External Order System's own order book, opens the canonical case, and puts back a demo customer binding the reset would otherwise erase. It requires--confirm destroy-and-restore, refuses anything that is not a canonical demo world before it destroys anything, and never confirms a plan or sends a message.--dry-runwrites nothing. See docs/demo-world-restore.md.- The MCP endpoint is a separate process, and cannot reach the database. It reaches a case
the way any other client would: an authenticated HTTP call to the API's
/internal/intents. Point a client at it with the bearer token fromdocker/env/mcp.env. - The customer channel defaults to a fake provider everywhere except the deployment, so CI,
the tests and the local stack send nothing. Telegram is selected by
PP_CUSTOMER_CHANNEL_PROVIDER=telegram(customer-message-transport.md). - The model defaults to a deterministic fake, so the suites, the local stack and CI run with
no AWS credentials of any kind.
PP_LLM_PROVIDER=bedrockswitches to Amazon Bedrock using whatever the AWS SDK already authenticates with.
Behind an antivirus or corporate proxy that terminates TLS, put that root certificate in
docker/env/extra-ca.crt before building; the file is created empty and is otherwise ignored.
.env at the repository root points Alembic, the CLI and the integration suite at whichever
database you configured. docker/env/host.env points them at the disposable local one instead,
and scripts/with_local_env.py runs a single command with it:
uv run python scripts/with_local_env.py -- uv run pytest apps/backendpr is the product gate: ruff, mypy in three groups, pytest with a coverage floor, the
Hypothesis CI profile, import-linter, the MCP protocol suite, the semantic boundary and
evaluation, the order system, the frontend, gitleaks, the backend against PostgreSQL and the
whole-stack browser suite.
The engine suite needs nothing at all:
uv run pytest packages/promise-graphThe backend suite needs a database. With the local stack running, stop the worker first: it shares the local database and will claim the steps a workflow test just enqueued.
docker compose stop workeruv run python scripts/with_local_env.py -- uv run pytest apps/backendThe order-system boundary, end to end across both applications:
uv run python scripts/with_local_env.py -- uv run pytest apps/backend/tests/test_order_system_boundary.pyThe semantic boundary needs neither a database nor an AWS account:
uv run pytest apps/backend/tests/test_semantic_contracts.py \
apps/backend/tests/test_semantic_provider.py \
apps/backend/tests/test_semantic_grounding.py \
apps/backend/tests/test_explanations.py \
apps/backend/tests/test_bedrock_semantic.pyThe MCP protocol suite needs no database, no credential and no model. It starts the real server on a loopback socket and drives it with the official SDK's client:
uv run pytest apps/backend/tests/test_mcp_protocol.py apps/backend/tests/test_status_view.py \
apps/backend/tests/test_plan_identity.pyWhat a tool call causes needs the database. This suite drives SDK client, Streamable HTTP, MCP server, the service-token hop, the intent API, the domain and PostgreSQL, then asserts the rows:
uv run python scripts/with_local_env.py -- uv run pytest apps/backend/tests/test_intent_api.pyTests that call Amazon Bedrock for real are marked bedrock_live and deselected by default:
PP_LLM_PROVIDER=bedrock uv run pytest -m bedrock_liveThe order system's own suite:
uv run pytest apps/order-simulator packages/order-contractThe effect-set harness runs the frozen scenarios against the real system. With the local stack
up and the worker stopped, this is a development run that computes no score; --scored
reproduces the measurement, and --manifest selects v2
(effect-set-run-protocol.md fixes what counts as scored):
uv run python scripts/with_local_env.py -- uv run python scripts/run_effect_sets.pyuv run pytest apps/backend/tests/test_effect_set_judge.py scripts/tests/test_run_effect_sets.py \
scripts/tests/test_effect_set_manifest.pyThe frozen v1 manifest is docs/effect-sets/scenarios.v1.json, promisepatch-effect-sets
v1.0.0, content hash d41f5afcd01eda8e6fa4c28784f1fb0c238bbc27711019aac670914db62b2cdc. Its
labels are not derived from PromisePatch's output (effect-set-manifest.md).
The semantic and explanation evaluations run offline with zero provider calls:
python -m evals replay and python -m evals explanation-replay
(evals/README.md, explanation-quality-gate.md).
The frontend gates, and the browser suite against the local stack (it reloads the fixture):
cd apps/frontend && npm ci && npm run typecheck && npm run lint && npm test && npm run buildcd apps/frontend && npx playwright install chromium && npm run e2eApache-2.0. See LICENSE.
The G8 freeze at 56c3023 was reopened by ADR-0027 and re-established at 283f63f, then
reopened for the ADR-0028 bridge and re-established at the deployed release 740a062, image
740a062838e0. Repository HEAD may contain documentation-only commits on top of it.