Skip to content

v1.4.2: fail closed on optimizer evidence quality - #13

Merged
ipeterpetrus merged 26 commits into
mainfrom
fix/v1.4.2-optimizer-evidence-integrity
Sep 22, 2026
Merged

ipeterpetrus merged 26 commits into
mainfrom
fix/v1.4.2-optimizer-evidence-integrity

Conversation

@ipeterpetrus

@ipeterpetrus ipeterpetrus commented Sep 22, 2026 •

Copy link
Copy Markdown
Owner

Patch-release candidate for the legacy optimizer's evidence integrity. Not 1.5, not a
schema-4 port, and nothing new is activated. v1.4.1 is not moved and its released runtime is
unaffected.

This PR has four heads, and every blocked one is still described below. Each repair closed
the defect it was aimed at and exposed the next one, so the history is the evidence: nothing here
has been rewritten to make the branch look like it arrived clean.

0. The head history, in order

1. fe78cbc   the original repair: six defects in the legacy optimizer's evidence integrity
             -> BLOCKED by an independent acceptance review
2.            blocker B1: container damage was PERMANENT. One torn line disabled promotion for
             every scope sharing the history file, forever, with no path back.
3. e6c1f4f   the epoch repair: a loss is a BOUNDARY at a physical position, not a verdict
             -> BLOCKED by an adversarial review of that head
4.            blocker B-UTF8: attribution was read off the line that had just FAILED validation.
             Invalid UTF-8 in `scope_id` survived `errors="replace"` as U+FFFD, passed for a
             readable label, and cut a scope that does not exist — while the real population kept
             crossing the loss and a candidate was written from both sides of it.
5. 71016ee   the trust repair: a rejected line is never an authority for its own scope, the
             history is decoded strictly, and a scope-local boundary has exactly one trusted
             source — a record that PASSED validation and whose own counters prove the loss.
             -> author-adversarial review found one MEDIUM
6.            two MATERIALLY DIFFERENT observations sharing one run_id were collapsed into a
             single retry, so the second loss never opened a boundary and a candidate rested on
             evidence from both sides of it.
7. this head  the identity repair: run_id is a claim, not proof. Same identity plus a different
             persisted observation is a RUN_ID_CONFLICT, not a retry, and it cuts where it
             appears.

0.0 What the adversarial review blocked (head 2 → head 3)

Reproduced before anything was designed for it: six clean rev records, one SameWrite-shaped line
holding invalid UTF-8 inside scope_id (and independently failing validation), six more clean rev
records. On e6c1f4f: CANDIDATE, exit 10, one candidate file, records_in_epoch=12,
damage.scope_local={"\ufffd\ufffd\ufffd": 1}, and the specification's evidence_run_ids naming
records from both sides of the loss.

The defect was never about Unicode. It was provenance, and the repair says so as a rule:

A record that failed validation is not a trustworthy authority for its own scope attribution.

  • The history is read as BYTES, decoded strictly, one physical line at a time. errors="replace"
    destroyed the evidence that a decode had failed. A line that cannot decode is now a named
    rejection (line is not valid UTF-8) and a loss, and the reader continues at the next line rather
    than abandoning the file. The size cap measures the bytes it rejects.
  • Every loss a rejected line represents is FILE-GLOBAL. No rejected line may name a scope,
    whatever its scope_id looks like — rev, agent-b, a path, a 64-character label, or bytes that
    never decoded. This over-blocks on purpose: one corrupt line cuts scopes that were never damaged.
    That is the chosen half of the trade, because epochs recover and a fail-open crossing of a real
    loss does not. It also closes the MEDIUM (history.quality semantics) and both LOWs (a rejected
    label echoed into a public map, and that map's unbounded cardinality) in one move instead of three
    separate patches: 2000 rejected lines carrying 2000 distinct labels now add zero public keys.
  • A scope-local boundary keeps exactly one source, and it was arrived at by testing a hypothesis
    rather than assuming one. A record that PASSES valid_record() and whose own canonical loss
    counters report a loss was measured on the old head to keep history.quality DEGRADED at +2,
    +6 and +60 healthy records
    , while the identical populations without it promoted every time —
    the same fail-stuck shape the epoch model exists to remove, reached through the one door the epoch
    model did not watch. Such a record now opens a scope-local recovery boundary after itself and
    belongs to the epoch it closes, never to the one it opens. It is trusted where a rejected line is
    not: it passed structural validation, its scope_id is the field every accepted record already
    publishes through scope.known, and the damage fact comes from RECORD_LOSS_COUNTERS.
  • Deduplication identity drops the scope-local half of the epoch stamp. Across a file-global
    loss the reader cannot tell whether a repeated run_id is the same run, so both copies stand.
    Across a trusted boundary the file is intact and the repeat is the same run retrying, so its
    cleaner copy cannot enter the recovered epoch and launder the loss its twin reported. Both
    physical orders are frozen cases (VD_10, VD_11).

PARTIAL (a bound the caller asked for), UNKNOWN (a schema that cannot attest completeness) and
INVALID/EMPTY (readable but not comparable) deliberately create no boundary: the repair
target is loss continuity, not every ineligible record.

Found while verifying privacy, not reported by anyone: valid_record() echoed a rejected line's
own schema_version VALUE into history.rejected, which reaches --json, the human report and any
log that keeps them. Truncating it is not a bound — thirty characters of a credential is still the
credential — so only a number is echoed and any other value is named by type.

The limit this repair does not close, stated rather than implied: a legacy flat record carries no
integrity tag, so a corruption that turns one valid record into a different valid record
(scope_id flipped from agent-a to agent-b, a share vector rewritten to another that still sums
to ~100) is indistinguishable from a record the producer meant to write. Nothing here detects that,
and nothing can — the information needed is absent from the format.
LEGACY_STRUCTURALLY_VALID_CORRUPTION_LIMITATION=YES.

Two frozen B1 expectations are superseded rather than quietly re-run, each with its reason and
its replacement recorded in docs/V142_COUNTEREXAMPLES.md §6.6: B1_13e (its middle copy carried a
loss counter, so under the trust model it now opens an epoch instead of sitting inside one) and
B1_15 (its two boundaries came from rejected lines, so the same shape is rebuilt from trusted
ones). The property each protected keeps a case.

New frozen cases: TUTF_01–TUTF_09, TSCOPE_01–TSCOPE_07 with seven label canaries and a
2000-label cardinality attack, VD_01–VD_11, MSG_01/MSG_02, and TPRIV. New mutants:
M_UTF8_REPLACEMENT_ATTRIBUTED, M_REJECTED_SCOPE_TRUSTED, M_REJECTED_SCOPE_GHOST_KEY,
M_REJECTED_SCOPE_CARDINALITY, M_REJECTED_VALUE_ECHOED, M_GLOBAL_DAMAGE_NOT_CUT,
M_GLOBAL_DAMAGE_POISONS_FOREVER, M_VALID_DEGRADED_POISONS_SCOPE_FOREVER,
M_VALID_DEGRADED_CUTS_ALL_SCOPES, M_DEGRADED_RECORD_INCLUDED_POST_BOUNDARY,
M_DEGRADED_RETRY_NO_RECOVERY.

0.0.1 And what an author review of the trust repair itself found

Two rounds of cross-family author review ran on this head. The first returned NEEDS-FIX with a
HIGH, reproduced before it was believed and fixed in 189db57:

The scope-local boundary was opened at READ time, before deduplication, so a copy that was
then discarded as duplicate run_id (retry) still moved the counter — while this reader's own
rule calls a deduplicated retry bookkeeping rather than a loss.

DEGRADED X  ·  6 healthy records  ·  DEGRADED X again
  before: INSUFFICIENT_DATA, records_in_epoch=0, scope_local={"a": 2}
   after: CANDIDATE, records_in_epoch=6, scope_local={"a": 1}, one counted retry

Six healthy records sat between the two copies and none survived, and repeating
6 healthy + one more copy of X held the scope down indefinitely — the fail-stuck shape this
repair exists to remove, rebuilt out of its own recovery mechanism.
A boundary is now opened at
most once per (scope, file-global epoch, run_id). A record with no run_id cannot be shown to be
a retry and stays its own observation; a genuinely different second loss still cuts; across a
file-global loss the identity differs, because there the reader cannot tell whether a repeated id is
the same run at all. Frozen as VD_12, VD_12b, VD_13, VD_13b, VD_13c, with the mutant
M_DEGRADED_RETRY_CUTS_TWICE carrying its own positive control.

The same commit carries a second finding, made while verifying that a migration must not read as
file damage rather than reported by anyone: a line holding another tool's record_type in a
shared history was classified unknown record_type and counted as a loss, so a foreign entry cut
the file for every scope. The distinction no shares already draws one check further down now
applies here too — a line that also carries our fields is a corrupted record of ours and stays a
loss; a line that carries none of them is another tool's entry and creates no boundary
(M_FOREIGN_RECORD_TYPE_IS_DAMAGE).

Round 2 returned NEEDS-FIX too, with a HIGH that goes to the root of the trust model's own
rationale. §0.0 justified the one trusted scope-local boundary by saying the record "passed
structural validation and its scope_id is the field every accepted record already publishes". The
first half was true; the second was an assumption. valid_record() never checked scope_id, and
scope_of() is str(r.get("scope_id") or "default") — so a record that passes validation carrying
scope_id = ["rev"] becomes the scope "['rev']":

6 clean rev · one valid DEGRADED record with scope_id = ["rev"] · 6 clean rev
  before: CANDIDATE, records_in_epoch=12, damage {"file_global": 0, "scope_local": {"['rev']": 1}}
          candidate names rid-rev-0..5 AND rid-rev-20..25 — both sides of the loss
   after: CANDIDATE, records_in_epoch=6,  damage {"file_global": 1, "scope_local": {}}
          scope.known == ["rev"], candidate names only rid-rev-20..25

The blocked B-UTF8 defect, rebuilt through the one door this repair opened. The fix is the
producer's own contract rather than a new invention: tools/carry.py writes exactly one shape,
str(scope_id or "default")[:64], so a value of another type or another length was not written by
it. Such a record is refused with a STATIC reason — its content must never be echoed — and is a
file-global loss like any other unattributable one. A control byte inside a label stays accepted
deliberately, because the producer's cap truncates length without stripping control characters, that
label already reaches scope.known on every head of this branch, and calling it corruption would
invent damage where the file is intact. Frozen as TSCOPE_08 with four forged shapes and five
producer-writable controls, plus M_SCOPE_LABEL_UNCHECKED; §6.11 of the counterexample document
states the residual.

This is author review, not acceptance. The verdict on this head belongs to a new independent
reviewer.

0.0.2 And what the author-adversarial review of 71016ee found (head 3 → head 4)

One MEDIUM, reproduced before it was believed and fixed in 6ef9988. It is the last collapse this
reader still made, and it is the same shape as the two blocked heads above, through the last door
left open: identity.

Two materially different valid observations carrying one run_id were treated as a single retry,
so the second loss never opened a boundary.

DEGRADED X1(ts 0,  share 30, unreadable 2, sessions 40, turns 1000, scanned 80)
6 healthy
DEGRADED X2(ts 40, share 66, unreadable 9, sessions 91, turns 2400, scanned 150)
6 healthy                                        [one scope, one file-global epoch]

  before: CANDIDATE, records_in_epoch 12, one scope-local boundary for TWO reported losses,
          "duplicate run_id (retry)" = 1, candidate evidence = 6 records from each side of X2
   after: CANDIDATE, records_in_epoch 6, two boundaries, run_id_conflicts = 1, zero retries,
          candidate evidence = the six records after X2 and nothing else

Nine persisted fields differ between the two records; both pass validation and are independently
DEGRADED. Controls on the old head: a different run_id already gave two boundaries and an epoch
of six, an identical duplicate gave one and twelve, and a file-global loss between the copies gave
two plus one.

run_id is an identity CLAIM, not proof of semantic equality. Two records are the same run only
when their identity and their persisted observation match. Equivalence is a digest of the whole
parsed record with keys sorted — a hand-picked subset would be the same mistake in a new spelling —
so key order and whitespace cannot make identical observations differ, a persisted field this reader
does not know still can, and the reader's own annotations (_history_epoch,
_evidence_quality_floor) are excluded, because reading a file must not change what a record is.

A same-identity pair whose observations differ is a RUN_ID_CONFLICT: the later record is kept,
it is never reported as a retry, and it opens a recoverable scope-local boundary at its own physical
position, belonging to the epoch it closes. Complete-versus-complete counts too, and §7.3 of the
counterexample document argues that row rather than inferring it. Across a file-global loss nothing
changes: the reader cannot establish continuity there, so the later copy is a fresh identity.

One classification now drives the boundary, the deduplication, the quality floor and the
diagnostics.
The separate second pass is gone — a boundary layer and a dedup layer holding two
notions of identity is exactly how they came to disagree about one pair.

history.run_id_conflicts is a bounded integer, deliberately outside history.rejected (a
conflicting record is accepted, not a line that failed to become one) and deliberately not a map:
six hundred distinct conflicting ids produce the integer 600 and zero new keys.
output_schema_version stays 2.

Five earlier cases used one run_id for two different observations and asserted a retry. Each is
replaced rather than quietly re-run, with the reason recorded in §7.6 before the edit, and each
keeps the property it protected — now enforced by a boundary instead of by a quality floor, which is
strictly stronger. A consequence stated rather than hidden: a true retry now has the same quality by
construction, so QUALITY_FLOOR is a guard rather than a live path.

The limit this cannot close (§7.4): two physically distinct losses under one run_id with
identical persisted content stay indistinguishable from one identical retry. Line position is
deliberately not used as identity — that would make every true retry open a fresh boundary and
rebuild the fail-stuck behaviour the previous round removed.
IDENTICAL_REUSED_RUNID_LIMITATION=YES.

New frozen rows RID_01–RID_15; new mutants M_RUNID_CONFLICT_TREATED_AS_RETRY,
M_RUNID_CONFLICT_SECOND_LOSS_SUPPRESSED, M_EXACT_RETRY_OPENS_SECOND_BOUNDARY,
M_RUNID_CONFLICT_CROSSES_EPOCH, M_RUNID_CONFLICT_CROSS_SCOPE_COLLIDES,
M_RUNID_CONFLICT_GLOBAL_EPOCH_COLLIDES, M_CLEAN_RUNID_CONFLICT_IGNORED,
M_PRIVATE_ANNOTATION_IN_FINGERPRINT.

This is author work, not acceptance. The verdict on this head belongs to a fresh independent lane.

0.1 What the acceptance review blocked (head 1 → head 2)

The first head of this branch (fe78cbc) passed its own four adversarial rounds and CI. An
independent acceptance review then reproduced the six original defects on main, confirmed all six
were closed here, and blocked the PR on a defect this branch introduced:

container damage was permanent. One torn line — the crash fragment docs/MULTI_AGENT.md
§Concurrency calls an expected event, the one the reader is designed to "reject exactly that line
and count it" — set the whole file DEGRADED forever. Nothing in this product expires, rotates or
repairs a history ("The optimizer does not schedule, expire or rotate"), and --accept-partial
cannot adopt a loss by design, so one crash disabled promotion for every scope sharing
~/logs/carry_history.jsonl, permanently
.

Measured on that head, every one of these stayed PARTIAL_EVIDENCE with zero candidates: 6 good + torn; + 2 more; + 6 more; + 60 more; a sibling scope's records; a loss before any record; and
a legitimate zero-carry record. The writer was confirmed to leave the fragment in place: after a
crash it appends a newline and its own record, and the fragment stays line 1 forever.

The replacement is a boundary, not a verdict. An unattributable loss cuts the promotion history
at that line's physical position in the file. Records before the cut never join records after
it, the newest epoch is damage-free by construction, and the loss stays reported.

  • damage_boundary() is the one classifier every rejection passes through. A line that parses and
    carries a readable scope_id cuts that scope; a line that cannot say whose record it was —
    unparseable, oversized, a file that would not open — cuts every scope, because guessing would
    be the fail-open half. A foreign line, a schema-4 refusal by design and a deduplicated retry are
    not losses and cut nothing.
  • Records are stamped at read time with the epoch they were written in. Physical order, never the
    clock: the clock is exactly what a damaged history cannot be trusted about.
  • active_records() is chosen before anything is anchored, so a pre-loss record cannot pick the
    scope, the workload class, the corpus size or the trend for the population after it.
  • Deduplication is per epoch: the same run_id after a loss is that population's own observation;
    inside one epoch the retry is still dropped, counted, and still cannot launder the survivor.
  • The gap stays visible — history.rejected counts it, history.damage says how many boundaries
    the file holds, and the human report names them — while history.quality describes only the
    evidence eligible for the current analysis.

No time window, no expiry, no ratio, no override flag, no new status, and --accept-partial is
unchanged: it adopts a caller's chosen bound and never a loss.

A cross-family lane then attacked the epoch model itself and found two more places where the repair
still consulted the population it had just cut — both reproduced before being believed, both fixed
here, both pinned by a case and a mutant: a ledger-only finding was still filed under the scope of
the pre-loss records
(and the scope travels into the candidate id), and the epoch stamp being a
pair of counts meant two scopes could share one stamp, so deduplication treated one scope's record
as another's retry and pushed its quality onto it. The long-run simulator was also discarding the
epoch it was handed. §5.3 of the counterexample document has both cases.

This is author review, not acceptance: the verdict on this head belongs to a new independent
reviewer.

The same review's two MEDIUMs are repaired with it: history.quality is EMPTY (never COMPLETE)
when nothing is comparable, and a zero-carry sweep is readable evidence, not corruption —
tools/carry.py says so in its own report ("A session whose every item lands on its final turn
carries nothing"), the 1.3 writer emits shares: {} for it, the current writer guards if C else {}, and running accumulate() on such a transcript here returns sessions=1 turns=6 carry=0. The
contradictions stay INVALID: carry with no shares, shares with no carry, sessions without turns,
sessions beyond what was scanned.

docs/V142_COUNTEREXAMPLES.md §5 freezes thirteen B1 rows (status, exit, candidate files,
comparable count, active-epoch size, quality, boundary count) written before the repair existed, §5.1
records zero deviations from them, and §5.2 lists the three earlier expectations the epoch model
supersedes — each replaced together with a companion case proving the protection it carried is still
there.

1. Why PR #7 is not merged directly

PR #7 identified real defects, and an
independent audit against 77e3677 confirmed that most of them are still live. It is not
merged here for two reasons, neither of them about the quality of its analysis:

  • it was written against the pre-v1.4 architecture and now conflicts with main (17 files,
    CONFLICTING), and
  • main has since grown a typed evidence kernel (tools/evidence/*) that did not exist then, so
    some of PR SameWrite 1.3.1 — fail closed on partial history evidence #7's own additions are either unnecessary here or would duplicate a stronger design.

Its invariants, counterexamples and test matrices are treated as evidence; its implementation
is not copied. PR #7 stays open until this repair is reviewed and released, and the mapping in §7
is what would justify closing it as superseded.

2. The defects, reproduced on current main first

docs/V142_COUNTEREXAMPLES.md freezes thirteen cases — expected status, strict exit code,
candidate-file count, comparable count and evidence quality — written before the repair existed
and then run against 77e3677:

R142_01    CANDIDATE            exit 10  files 1  comparable 6   bounded history promotes
R142_01P   CANDIDATE            exit 10  files 1  comparable 6   --accept-partial changed nothing
R142_02    INSUFFICIENT_DATA    exit 20  files 0  comparable 0   INVALID record anchored, 6 stranded
R142_03    HOST_BEHAVIOR_SHIFT  exit 30  files 1  comparable 6   refusing status still wrote the file
R142_04    CANDIDATE            exit 10  files 1  comparable 6   sessions=turns=carry_bytes=0 promoted
R142_05    carry quality=COMPLETE, malformed=1 / run reports NO_ACTION, exit 0
R142_06    the listing scan read the OLDEST transcript under --max-files 1
U1         CANDIDATE            exit 10  files 1  comparable 6   schema-1 record read as COMPLETE

The four "already correct" controls (P1, P4, S_NOACTION, S_LOCK) behaved correctly before and
still do.

Scope of the exposure, stated plainly: the current writer produces schema-4 envelope records, which
this optimizer refuses by name, so these defects bite a history file written by 1.2/1.3 (or one
authored by hand), not a fresh one. The gate was still absent.

3. The semantic repair

One law, asked of the evidence each finding actually rests on, and one gate on the side effect.

  • record_quality() / sweep_quality() — the worst of what evidence claims and what its own
    numbers can show, over the counters the producer has written since schema 2
    (unreadable, oversize, skipped_by_limit). Nothing is invented for an older schema: schema
    0/1 predates evidence_quality entirely, so a schema-1 record carrying COMPLETE is UNKNOWN,
    and a schema-2 record without the counters cannot attest completeness either.
  • PARTIAL and DEGRADED are different things. A bound the caller asked for is PARTIAL and
    is exactly what --accept-partial adopts. Evidence that was selected and then lost — unreadable,
    oversize, malformed, identity-changed, conflicting — is DEGRADED, and no flag accepts it.
    carry.LOSS_FIELDS / carry.BOUND_FIELDS is the single vocabulary both sides read.
  • Malformed records are a loss. carry.sweep_label() no longer reports COMPLETE for a sweep
    whose transcript lost records to torn JSON.
  • eligible_anchor() chooses the comparison anchor — and the analysed scope — from records
    that survive the eligibility filter, so the newest INVALID record can no longer decide the
    corpus size for everyone else and then remove itself.
  • The history trend is gated. It was if True:. It is now a CANDIDATE on eligible history
    and an OBSERVED finding carrying the reason otherwise.
  • Per-finding evidence. Live findings gate on the sweep, the trend on the eligible history, the
    guard on its own ledger sample. A partial history no longer blocks a live finding it never
    supported, and a clean live sweep no longer waves a bounded history through.
  • emit_candidates(findings, outdir, status) writes nothing unless the status is CANDIDATE.
    The gate is inside the emitter, not at its one caller, so the CLI, the long-run simulator and any
    future scheduler all route through it.
  • carry.bounded_paths() is the one definition of "the newest N", used by both the carry sweep
    and the skill-listing scan.

CANDIDATE now precedes PARTIAL_EVIDENCE in overall_status(). That is not a loosening: a
finding only reaches CANDIDATE after its own evidence passed the gate, and the old order refused
a valid history-only finding because today's live sweep happened to be bounded.

4. Counterexamples and positive controls

tests/test_evidence_integrity.py — 232 assertions: one case per frozen matrix row, the law's own
unit cases, and the cases the review rounds added (R142_07 … R142_13). A fail-closed patch can look
excellent by refusing everything, so the controls are part of the gate:

  • P1 a clean COMPLETE population still promotes (CANDIDATE, exit 10, one file).
  • P2 an explicitly accepted bounded population still promotes (--accept-partial).
  • P3 an INVALID record in another scope does not poison this one.
  • P4 schema-4 evidence stays refused by name; history_schema_supported is still [0, 1, 2].
  • S_NOACTION / S_LOCK the strict-exit contract, with the filesystem checked, not the wording.

5. Mutation coverage

Twenty-three new mutants, each restoring exactly one repaired defect, each RED on the mutant and
GREEN on the real source (tests/test_mutation.py, 55/55) — the last eight are the epoch model's:

M_PARTIAL_PROMOTES · M_INVALID_ANCHOR · M_HOST_SHIFT_WRITES · M_MALFORMED_COMPLETE
M_BOUND_SAMPLE_DIVERGES · M_TREND_QUALITY_BYPASS · M_DEDUP_LAUNDERS
M_IMPOSSIBLE_COUNTERS · M_CONTAINER_DAMAGE_IGNORED · M_LEDGER_TORN_PROMOTES
M_EMIT_FAILURE_SILENT · M_BARE_PARTIAL_CLAIM · M_RECORD_LOSS_IGNORED
M_STRUCTURAL_REJECT_NOT_DAMAGE · M_LEDGER_UNKNOWN_EVENT
M_DAMAGE_POISONS_FOREVER · M_DAMAGE_IGNORED_COMPLETELY · M_EPOCH_MERGES_ACROSS_GAP
M_SCOPE_DAMAGE_GLOBALIZED · M_UNATTRIBUTABLE_DAMAGE_SCOPED · M_EMPTY_COMPARABLE_REPORTS_COMPLETE
M_ZERO_CARRY_BECOMES_DAMAGE · M_PRE_DAMAGE_RECORD_ANCHORS

None is a string mutant: every oracle observes a status, a comparable count, a file on disk, a
sweep's quality or a JSON field through a real run.

Four fixtures claimed COMPLETE without the acquisition counters a real sweep always writes
(multi-agent, mutation, long-run, readiness). Each now states them, and each edit is locked by an
added assertion
that the same record without them refuses to promote, so the fixture change
cannot hide the gate it was making room for.

6. What did NOT change

  • SCHEMA4_OPTIMIZER_SUPPORT=NO — SCHEMA_SUPPORTED is still (0, 1, 2) and a v1.4 envelope is
    still refused by name, with the refusal counted.
  • Typed promotion stays shadow-only: tests/test_release_shadow_only.py passes, no promotion
    or candidate-persistence entry point is added, and nothing branches on the product version.
  • The schema-4 contract is untouched — proved, not asserted: the encoded v1.4 record for a clean,
    a torn and a bounded sweep has a byte-identical sha256 before and after this branch, and the
    typed reader's loss_observed verdict is unchanged. Only the legacy quality label moved, for
    the torn sweep, which is the defect being repaired.
  • SKILL_BODY_DELTA=0, SKILL_DESCRIPTION_DELTA=0, HOOK_BEHAVIOR_DELTA=0 — zero files under
    skills/ or hooks/; CANONICAL_BODY_SHA256=7edec9f21e0bd50583e388e0bdc177f92ba8f0db6ce2bb61acd52b39d81e7767.
    No model-facing byte changes, so no benchmark is required.
  • Product version stays 1.4.1 in all three manifests. Packaging, tagging and release are a
    separate gate after review, as they were for 1.4.1.

7. PR #7 property → v1.4.2 replacement

PR #7 status on main replaced by
H1 impossible/unattested COMPLETE promoted APPLIES record_quality() + _derived_record_quality(); cases R142_04, U1, law units
H2 torn JSON leaves the sweep COMPLETE APPLIES carry.LOSS_FIELDS + sweep_label(); case R142_05, mutant M_MALFORMED_COMPLETE
H3 candidate stamped effective_evidence_quality: COMPLETE DOES NOT APPLY (introduced inside PR #7; main never writes that field) not reintroduced — the quality travels in the finding's evidence line instead
H4 refusing status still writes a candidate APPLIES status gate inside emit_candidates(); case R142_03, mutant M_HOST_SHIFT_WRITES
H5 substring artifact check, unreadable = current, invalid UTF-8 crashes DOES NOT APPLY (introduced inside PR #7; main only tests existence and derives no trust from contents — nothing reads a candidate file back) ARTIFACT_VALIDATION=DEFERRED_NOT_REQUIRED_FOR_THIS_REPAIR; the typed kernel's classifier already supersedes it for v1.4
M1 candidates_existing turned into prose DOES NOT APPLY (main reports bare ids) unchanged, now covered by the JSON assertions
M2 newest INVALID becomes the anchor APPLIES eligible_anchor(); case R142_02, mutant M_INVALID_ANCHOR
M3 two definitions of a bounded sample APPLIES carry.bounded_paths(); case R142_06, mutant M_BOUND_SAMPLE_DIVERGES

8. Regression

ALL_TESTS=PASS                     seventeen Python suites, 1460 assertions (CI enforces the number)
MUTATION_TESTS=PASS                55/55 mutants RED on the mutant, GREEN on the real source
SHELL_OFFLINE=PASS                 tests/test_oneliner_cleanup.sh; acceptance_public and
                                   acceptance_upgrade pass; acceptance_openclaw needs OPENCLAW_BIN
                                   and fails identically on main (verified in a worktree of main,
                                   byte-identical output) — environmental, not this branch
TYPED_EVIDENCE_REGRESSION=PASS     test_evidence_boundary (472) + test_evidence_phase2 (87)
MESSAGE_ACCOUNTING_REGRESSION=PASS test_multiblock (164) + test_mutation, issue #1 accounting intact
SHADOW_ONLY_REGRESSION=PASS        test_release_shadow_only
AI-VOS                             longrun LADDER=PASS, POLICY_MUTATION=NONE, CANDIDATE_SPAM=NO
                                   readiness 28 PASS / 1 FAIL — the same single pre-existing failure
                                   main has (26/1); this branch adds two passes and fixes none of it

TOTAL_SAVINGS=NOT_PROVEN and WORLD_BEST_CLAIM=NOT_TESTED are unchanged.

9. Adversarial review

Two reviewer families, asked for counterexamples rather than an opinion, on the diff itself.

Round 1 — eight defects, all in the repair, all closed here:

  • a retry reusing a run_id could launder its own sweep's quality (found by re-attacking the law)
  • schema_version as a string or a float was read as that schema
  • zero counters with no sessions / turns / carry_bytes at all read COMPLETE — a share
    vector with no population behind it (both families found this one)
  • one unreadable mtime sent the whole bounded selection back to a discovery-order slice
  • sessions=40 out of scanned=1 — counters that cannot describe one sweep — read COMPLETE
  • damage to the history container (a torn line in the file) never reached the evidence decision
  • a ledger that lost a line still retired the guard
  • a run could report CANDIDATE and exit 10 while every candidate write failed

Each has a case in the suite and, where it is a behaviour rather than a value, a mutant.

Round 3 — four more, every one of them a word taken for evidence: a PARTIAL claim no counter
could explain was adopted by --accept-partial; the record-level loss vocabulary was two counters
wide while the contract promises more (and the dedup floor inherited the same short list); container
damage was an allowlist, so a rejection reason added later would default to "not damage"; and a
ledger line that is neither checked nor denied counted as nothing at all.

Round 4 — no surviving HIGH. One was claimed (a missing isinstance guard in load_ledger)
and does not exist: the guard is there and a list-valued ledger line is counted as rejected,
verified by running it. Four MEDIUMs were real and are fixed: the live sweep now refuses the same
impossibility the record law refuses; the bounded selection is made ONCE per run and handed to both
consumers, so a transient getmtime failure between two calls cannot produce two samples; a line
that claims to be one of our records and carries no shares is damage rather than a foreign line;
and a candidate write that fails mid-way no longer leaves its .tmp file behind.

Two limits are stated rather than papered over: container damage is a property of the FILE, so a
torn line degrades every scope in it — an unparseable line has no scope to attribute it to — and
HOST_BEHAVIOR_SHIFT still closes the emitter for findings that do not rest on the shifted
population, which is the release contract's ordering, not this patch's.

On the anchor fallback, one reviewer read the eligible_anchor(recs) or newest fallback in
main() as a way for an ineligible record to choose the analysed scope. The observation reproduces — and stops there. The fallback is reached only when nothing
in the file is eligible, so there is no population to strand and nothing to promote: the reviewer's
exact input gives PARTIAL_EVIDENCE, exit 40, zero files. Both halves are now tests (R142_14),
including the same record added to a real population in another scope, where the eligible records
still choose the scope and still promote.

One deliberate narrowing came out of this: the first cut of container damage degraded the history
for any rejected line, so one stray hand-written line in a shared file would have blocked a real
population forever. Damage is now a torn line, an oversized line or an unreadable file; a foreign
line that parses, and a refusal by design, are counted and reported without degrading. A corrupted
carry record that still parses as JSON is reported and does not degrade — stated as a limit
rather than implied as coverage.

10. What a reviewer should attack

The law's calibration is the judgement call, not the plumbing: a schema-1 history can no longer
promote anything
, because the generation that wrote it had no field to say how the sweep went. If
you think a record with no acquisition facts should still carry a promotion, this is the line to
argue with — it is one function, record_quality(), and every case that depends on it is in
docs/V142_COUNTEREXAMPLES.md.

ipeterpetrus and others added 26 commits September 21, 2026 21:38
Thirteen cases, every expected status, exit code, candidate-file count
and comparable count written down BEFORE any repair exists, and then run
against 77e3677 to record what it actually does.

Seven of them are defects the current legacy optimizer really has, on
these exact fixtures:

- a history built entirely from bounded sweeps promotes a trend
  candidate, and --accept-partial changes nothing because history
  quality is never read at all
- the newest INVALID record picks the comparison anchor and then removes
  itself, stranding six eligible records (comparable 0)
- HOST_BEHAVIOR_SHIFT reports exit 30 while writing the candidate file
- six records claiming COMPLETE with sessions=turns=carry_bytes=0
  promote
- a transcript that lost a record to torn JSON still reports the sweep
  COMPLETE, and the run reports NO_ACTION
- --max-files 1 gives the carry sweep and the skill-listing scan two
  different single-source samples
- a schema-1 record, from a generation that had no evidence_quality
  field at all, is read as COMPLETE

Four cases are positive controls (a clean COMPLETE population must still
promote; an invalid record in another scope must not poison this one;
schema-4 evidence must stay refused by name; a flat population must stay
NO_ACTION) and two pin the strict-exit contract.

No source file is touched by this commit: the suite is expected to fail
here, and the matrix in docs/V142_COUNTEREXAMPLES.md records both the
frozen expectations and what main did.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…e emission gate

The legacy optimizer decided five different things about evidence in five
different places, and one of them decided nothing at all. This makes the
question one function, asks it of the evidence each finding actually
rests on, and puts the candidate-writing gate where every caller has to
pass through it.

What changed, in the order the counterexamples found it:

- `record_quality()` / `sweep_quality()` are the law. The worst of what a
  record claims and what its own numbers can show, over the counters the
  producer has written since schema 2. A schema-1 record has no
  evidence_quality field at all, so a schema-1 record carrying COMPLETE
  is a claim its writer could not have made: UNKNOWN. A schema-2 record
  without the acquisition counters cannot attest completeness either.
- PARTIAL and DEGRADED are now different things. A bound the caller asked
  for (`--max-files`) is PARTIAL and is exactly what `--accept-partial`
  adopts; evidence that was selected and then lost is DEGRADED and no
  flag accepts it. carry.LOSS_FIELDS / BOUND_FIELDS is the single
  vocabulary both sides read.
- A sweep that lost records to torn JSON is no longer COMPLETE.
  `malformed` joins the loss counters, so carry.sweep_label() degrades
  the sweep and the optimizer reads it as DEGRADED.
- `eligible_anchor()` picks the comparison anchor from records that
  survive their own filter. The newest INVALID record used to choose the
  scope, the workload class and the corpus size for everyone else and
  then remove itself, stranding a whole eligible population.
- The history trend is gated. It was `if True:` — a trend over bounded
  sweeps was indistinguishable from one swept in full. Now it is a
  CANDIDATE on eligible history and an OBSERVED finding, with the reason
  in its evidence line, on anything else.
- Each finding is gated on ITS OWN evidence: live findings on the sweep,
  the trend on the eligible history, the guard on its own ledger sample.
  A partial history no longer blocks a live finding it never supported.
- `emit_candidates()` takes the status and writes nothing unless it is
  CANDIDATE. HOST_BEHAVIOR_SHIFT reported exit 30 while leaving a
  specification on disk; the gate now lives in the emitter, so the CLI,
  the long-run simulator and any future caller route through it.
- `carry.bounded_paths()` is the one definition of "the newest N", used
  by both the carry sweep and the skill-listing scan. One `--max-files 1`
  run was analysing two different single-source populations.

Schema-4 evidence stays unsupported by this optimizer and the typed
promotion stays shadow-only. Proved rather than asserted: the encoded
v1.4 record for a clean, a torn and a bounded sweep has a byte-identical
sha256 before and after this commit, and the typed reader's loss_observed
verdict is unchanged. Only the legacy `quality` label moved, for the torn
sweep, which is the defect being repaired.

Four fixtures claimed COMPLETE without the acquisition counters a real
sweep always writes (multi-agent, mutation, long-run, readiness). Each
now states them, and each edit is locked by an added assertion that the
same record WITHOUT them refuses to promote — so the fixture change
cannot hide the gate it was making room for.

skills/ and hooks/ are untouched: CANONICAL_BODY_SHA256 is unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Six new mutants, each restoring exactly one of the defects the frozen
matrix found, each RED on the mutant and GREEN on the real source:

  M_PARTIAL_PROMOTES       may_promote() returns True
  M_INVALID_ANCHOR         the anchor is chosen before the filter again
  M_HOST_SHIFT_WRITES      the emitter's status gate is deleted
  M_MALFORMED_COMPLETE     `malformed` leaves the loss vocabulary
  M_BOUND_SAMPLE_DIVERGES  the listing scan takes its own slice again
  M_TREND_QUALITY_BYPASS   the history trend stops asking

None of them is a string mutant: every one changes behaviour that the
oracle observes through a real run — a status, a comparable count, a
file on disk, a sweep's quality, a JSON listing field.

Two existing mutation patterns moved with the code and were re-pointed
(the PARTIAL gate and the trend's direction), and the shared record
fixture states its acquisition counters for the same reason the other
fixtures do.

README: 1311 assertions in seventeen suites (CI enforces this number).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Three from an adversarial pass over the new law, one from a cross-family
review of the same diff. Each has a case in the suite; the first also has
a mutant.

- A retry that reuses a run_id could launder its own sweep. The duplicate
  is dropped as an OBSERVATION, but the quality it reported is part of
  what that id actually saw, so the worse of the two now travels with the
  record that survives (QUALITY_FLOOR, in memory only — nothing is
  written back to the history file).
- `schema_version` of "2" (a string) or 2.0 (a float) was accepted as
  schema 2 by int(). A version this reader cannot name is not a newer
  generation to trust; it is an older one to doubt.
- Zero counters say "nothing went wrong"; they do not say a sweep
  happened. A schema-2 record carrying zeroed counters and no `sessions`,
  `turns` or `carry_bytes` at all read COMPLETE — a share vector with no
  population behind it. It now reads UNKNOWN. (cross-family review)
- One source whose mtime could not be read sent the ENTIRE bounded
  selection back to a discovery-order slice — the exact sample
  bounded_paths() exists to avoid. The fallback is now per source: a file
  that cannot be dated cannot claim to be the newest, and the rest still
  order by mtime. (cross-family review)

README: 1330 assertions.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Four rows, each from an attack on the repair rather than on the original
defect, kept in their own section so the frozen matrix stays readable as
what it was when it was frozen. No expectation in section 1 changed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
… nobody checked"

A cross-family adversarial pass over the repair itself. Every one of
these produced a CANDIDATE and a file on disk before this commit.

- Counters that cannot describe one sweep. `sessions=40` out of
  `scanned=1` — or out of `scanned=0` — read COMPLETE, because the law
  checked each counter's type and sign but never their relation. A
  session is a transcript that was scanned and cleared the turn floor, so
  the producer can never report more sessions than it scanned. That
  combination is now INVALID, and `scanned` joins the fields a schema-2
  record must carry before its zeroes mean anything.
- Damage to the history CONTAINER never reached the evidence decision. A
  torn line in the history file was counted in `history.rejected` and
  then ignored: the trend built from the surviving records was promoted
  as if nothing had been lost. `container_quality()` reads it as a loss
  (DEGRADED, so no flag adopts it) and both `analyse()` and
  `overall_status()` take the worst of it and the records. A refusal by
  design — current-generation evidence this optimizer does not read — and
  a deduplicated retry are NOT damage; both have controls.
- A ledger that lost a line still retired the guard. 100 readable writes
  plus one torn line promoted `noop-guard-retire`. `ledger_usable` now
  requires `rejected == 0`, and a ledger that exists but cannot be opened
  is no longer indistinguishable from no ledger at all.
- A run could report CANDIDATE and exit 10 while every candidate file
  failed to write. The status has to say what happened: nothing landed,
  so the run is INTERNAL_ERROR (exit 50), with the failure named in the
  JSON.

Four more mutants, each RED on the mutant and GREEN on the real source
(41/41 now): M_IMPOSSIBLE_COUNTERS, M_CONTAINER_DAMAGE_IGNORED,
M_LEDGER_TORN_PROMOTES, M_EMIT_FAILURE_SILENT.

README: 1360 assertions.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…were foreign

The first cut degraded the history for any rejected line, including a
well-formed JSON line that simply is not a carry record. In a shared file
one stray append would then block a real population forever — a
fail-closed patch that refuses evidence it should still read is also a
defect.

Damage is now what it says: a torn line, a line past the size cap, a file
that could not be opened. A foreign line and a refusal by design are
counted and reported, and do not degrade; both have controls in the
suite. The limit is written down rather than implied: a corrupted carry
record that still parses as JSON is reported and does not degrade.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A confirmation-round reviewer read the `eligible_anchor(recs) or newest`
fallback in main() as a way back in for an ineligible record. Their exact
input reproduces the observation — the run names the scope of the one
INVALID record it found — and it stops there: status PARTIAL_EVIDENCE,
exit 40, nothing on disk, because the fallback is only reached when
NOTHING in the file is eligible, and then there is no population to
strand and nothing to promote.

Both halves are now tests rather than an argument: the reviewer's input,
and the same record added to a real population in another scope, where
the eligible records still choose the scope and still promote.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…locks

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The confirmation lane attacked the repair rather than the original defect
and found four more places where a word was taken for evidence.

- A record could say PARTIAL while every counter it carried said nothing
  happened, and --accept-partial adopted it. A bound shows up in
  skipped_by_limit and a loss shows up in a loss counter; a PARTIAL claim
  that no counter can explain cannot say WHY it was partial, so it is
  UNKNOWN and no flag adopts it.
- The record-level loss vocabulary was two counters wide (unreadable,
  oversize) while the contract promises more. A record carrying
  malformed, malformed_lines, identity_changed, conflicted_sources or
  records_rejected read COMPLETE — and a retry carrying one could launder
  itself through dedup, because the floor was computed with the same
  short list. REQUIRED (what the schema-2 writer always wrote, so its
  absence means UNKNOWN) and LOSS (what must be honoured when present)
  are now separate lists.
- Container damage was an allowlist of three reasons, so a rejection
  valid_record() learns to make later would default to "not damage". It
  is now the other way round: a rejected line is damage unless it is one
  of three named exceptions — a refusal by design, a deduplicated retry,
  or a well-formed line that was never a carry record. An unsupported
  schema-3 record and a record whose shares do not sum to a population
  now close the emitter instead of being footnotes.
- A ledger line that is neither `checked` nor `denied` was counted as
  nothing at all, so a ledger full of unknown events looked like a clean
  sample. It counts as rejected, which the guard's own gate reads.

Four more mutants (45/45). The schema-4 contract is still untouched: the
encoded v1.4 record for a clean, a torn and a bounded sweep has the same
sha256 as on main.

README: 1399 assertions.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…, no half files

A fourth adversarial pass returned no HIGH that survived checking (the
one HIGH claimed a missing isinstance guard in load_ledger that is
present, and a list-valued ledger line is counted as rejected, verified
by execution). Four MEDIUMs were real:

- `sweep_quality()` did not apply the impossibility rule its record-level
  twin applies: a live sweep reporting sessions with zero turns read
  COMPLETE. It is INVALID, like the record.
- `bounded_paths()` was called twice per run — once by the sweep, once
  for the listing — so an mtime that became unreadable between the two
  calls produced two different samples again. The selection is made ONCE
  in main() and handed to both; `accumulate(selected=...)` takes it, and
  still reports the bound against the whole discovered population.
- A line that CLAIMS to be one of our records and carries no shares is a
  corrupted record, not another tool's entry. It now reads as container
  damage; a line that claims nothing still does not.
- A candidate write that failed mid-way left its `.tmp-<pid>` file
  behind. It is removed on the failure path.

Two limits stay, stated rather than papered over: container damage is a
property of the FILE, so a torn line degrades every scope in it (an
unparseable line has no scope to attribute it to), and
HOST_BEHAVIOR_SHIFT still closes the emitter for findings that do not
rest on the shifted population — that ordering is the release contract's,
not this patch's.

45/45 mutants. README: 1408 assertions. The schema-4 encoded record for a
clean, a torn and a bounded sweep still has the same sha256 as main.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
An independent acceptance review BLOCKED the first head of this branch.
The six original defects were closed, but the container-damage repair
introduced a new one: damage was permanent. One torn line — the crash
fragment docs/MULTI_AGENT.md calls an expected event, the one the reader
is designed to "reject exactly that line and count it" — set the whole
file DEGRADED forever, for every scope sharing it, with no flag able to
adopt a loss and nothing in the product that expires or rotates a
history. Measured on the current head, all of these stay PARTIAL_EVIDENCE
with zero candidates: 6 good + torn; + 2 more; + 6 more; + 60 more; a
sibling scope's records; a loss before any record; and a legitimate
zero-carry record.

This commit freezes what the repair must do, before it exists: thirteen
rows in docs/V142_COUNTEREXAMPLES.md §5 with status, exit code, candidate
files, comparable count, active-epoch size, history quality and boundary
count, plus the two MEDIUMs the same review found (history.quality read
COMPLETE with nothing comparable; a producer-valid zero-carry record read
as corruption).

The model being frozen: an unattributable loss cuts the promotion history
at that line's PHYSICAL position. Evidence before the cut never joins
evidence after it, the newest epoch is by construction damage-free, and
the loss stays reported. No time window, no expiry, no ratio, no new
override flag, and --accept-partial still adopts a chosen bound and never
a loss.

21 assertions fail here, each naming its frozen expectation.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
… file

B1, the blocker an independent acceptance review found in this branch's
first head. Container damage was permanent: one torn line — the crash
fragment docs/MULTI_AGENT.md calls an expected event — set the whole file
DEGRADED, and since nothing here expires, rotates or repairs a history
and no flag adopts a loss, a single crash disabled promotion for every
scope sharing ~/logs/carry_history.jsonl, forever. Measured before this
commit: 6 good + torn + 60 good still returned PARTIAL_EVIDENCE with zero
candidates.

A loss is now a BOUNDARY at the rejected line's physical position in the
file, not a verdict on the file:

- `damage_boundary()` is the one classifier every rejection goes through,
  including the ones that never became an object. A line that parses and
  carries a readable `scope_id` cuts that scope; a line that cannot say
  whose record it was — unparseable, oversized, a file that would not
  open — cuts every scope, because guessing would be the fail-open half.
  A foreign line, a refusal by design and a deduplicated retry are not
  losses and cut nothing.
- Records are stamped at READ time with the epoch they were written in:
  (file-global losses before them, losses attributed to their own scope
  before them). Physical order, never the clock, because the clock is
  what a damaged history cannot be trusted about.
- `active_records()` is the analysed population: what came after the
  newest loss that applies to its scope. It is selected BEFORE anything
  is anchored, so a pre-loss record cannot choose the scope, the workload
  class, the corpus size or the trend for the population after it.
- The active epoch is therefore damage-free by construction, which is why
  the container no longer gates: `analyse()` and `overall_status()` ask
  only about the evidence eligible right now.
- The loss stays visible: `history.rejected` counts it and the new
  `history.damage` says how many boundaries the file holds, in the JSON
  and in the human report.

Deduplication is now per epoch. The same run_id after a loss is that
population's own observation; inside one epoch the retry is still
dropped, counted, and still cannot launder the survivor's quality.

Two MEDIUMs from the same review, repaired here because the epoch model
makes both sharper:

- `history.quality` is the quality of the evidence eligible for THIS
  analysis, and is EMPTY when nothing is comparable. It read COMPLETE
  with zero comparable records. `worst_quality([])` still means COMPLETE:
  a finding resting on no sampled evidence is not degraded by sampling it
  never used.
- A zero-carry sweep is readable evidence, not corruption. tools/carry.py
  says so in its own report — "A session whose every item lands on its
  final turn carries nothing" — the 1.3 writer emits `shares: {}` for it
  and the current one guards `if C else {}`. Proved by running it:
  accumulate() on such a transcript returns sessions=1 turns=6 carry=0.
  The record reads EMPTY, is left out of comparable, and cuts nothing.
  The contradictions stay INVALID: carry with no shares, shares with no
  carry, sessions without turns, sessions beyond what was scanned.

--accept-partial is unchanged and still adopts only a caller's chosen
bound, never a loss. No time window, no expiry, no ratio, no override
flag, no new status. Schema-4 stays unsupported and its encoded bytes are
unchanged (clean, malformed and bounded sweeps all hash identical to
main). 1460 assertions, 53 mutants.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A run over a damaged file now says how many records the scope holds, how
many of them came after the newest loss, and what the loss was — the same
distinction the JSON makes between what the file holds and what the
current analysis rests on.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…es inside one epoch

The chosen semantics, pinned rather than inherited: the copy that lands
after a loss is judged on its own evidence (clean before, bounded after
-> PARTIAL), the copy excluded on the far side of the loss does not
poison the epoch that follows it (bounded before, clean after ->
COMPLETE, and no refusal), and three copies inside ONE epoch still leave
the worst of them on the survivor (DEGRADED, two retries counted).

Which copy comes first is exactly what a crash decides, so both orders
are tests now.

README: 1463 assertions.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Found by a cross-family lane attacking the epoch model itself, and both
reproduced before being believed.

- A finding that rests on the ledger alone may promote after a loss —
  that is per-finding evidence working — but the SCOPE it was filed under
  still came from the records before the loss, and the scope travels into
  the candidate id. A history of six `stale` records followed by a torn
  line produced `noop-guard-retire-stale-…` and wrote it. The scope now
  comes from the current epoch or from nothing: `default` when the epoch
  is empty, never a population that no longer exists.
- The epoch stamp is a pair of COUNTS, so a record of scope "a" and one
  of scope "b" can carry the same numbers while belonging to different
  epochs. Deduplication keyed on the stamp alone therefore treated
  scope "b"'s record as scope "a"'s retry and pushed its PARTIAL quality
  onto scope "a" through QUALITY_FLOOR, turning a clean population into
  PARTIAL_EVIDENCE. Identity now carries the scope.
- The long-run simulator discarded the epoch it was handed, so its
  analysis could combine records across a loss. It uses active_records()
  and the same damage summary the CLI does.

Two mutants (M_STALE_SCOPE_ANCHOR, M_DEDUP_IGNORES_SCOPE) and four
assertions (B1_14, B1_15) pin all of it; the older anchor mutant moved
with the code. 1469 assertions, 55 mutants, schema-4 bytes still equal to
main, skill body untouched.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
B1_14 (a stale population naming a post-loss candidate) and B1_15 (two
scopes sharing one epoch stamp), in the same post-freeze section as the
rest, with the long-run simulator's missing epoch named too.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…enting it

The adversarial review of the epoch model found attribution taken from the line
that had just failed validation: invalid UTF-8 in scope_id survived
errors="replace" as U+FFFD, read as a readable label, and cut a scope that does
not exist while the real population kept crossing the loss.

Frozen here, before any code: rejected lines are never an authority for their own
scope (every loss is file-global), the history is decoded strictly per physical
line, and the one trusted source of a scope-local boundary is a record that
passed validation and whose own canonical loss counters prove the loss.

Section 6.3 records the hypothesis reproduced first, with a matched control:
VALID_DEGRADED_FAIL_STUCK=YES. Section 6.6 records the two B1 expectations this
trust model supersedes, and what replaces each.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…oundary

38 failures on e6c1f4f, one per row of docs/V142_COUNTEREXAMPLES.md §6:

  TUTF_01..08   invalid UTF-8 in a physical line is an unattributable loss.
                Today errors="replace" turns it into U+FFFD, the reader calls
                that a readable scope_id, and a scope that does not exist is
                cut while the real population keeps crossing the gap — the
                candidate names run ids from both sides of the loss.
  TSCOPE_01..07 a record that failed validation is not an authority for its
                own scope_id, whatever the string looks like. Canaries and a
                2000-label cardinality attack pin that no rejected label can
                become a public map key.
  VD_01..11     a record that PASSED validation and whose own canonical loss
                counters prove a loss is the one trusted source of a
                scope-local boundary. Reproduced first, with a matched
                control: today that record keeps history.quality DEGRADED
                forever, at +2, +6 and +60 healthy records.
  MSG_01/02     file-global recovery, already correct, pinned against regress.

The write() helper now takes bytes: a line that is not valid UTF-8 cannot be
written through a text handle, and that line is the input under test.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Closes the HIGH blocker the adversarial review of the epoch model found, and
the MEDIUM and two LOWs that shared its root cause. The defect was not Unicode.
It was provenance: attribution was read off the very line that had just failed
validation.

Reproduced first, on e6c1f4f: six clean records, one SameWrite-shaped line whose
`scope_id` holds invalid UTF-8 and which also fails validation, six more clean
records. errors="replace" turned the bytes into U+FFFD, damage_boundary() called
that a readable label and cut a scope nobody has, the real population was never
cut, and the run wrote a candidate whose evidence_run_ids span both sides of the
loss (CANDIDATE / exit 10 / epoch 12 / comparable 12).

Three changes, one rule:

  A record that failed validation is not a trustworthy authority for its own
  scope attribution.

* The history is read as BYTES and decoded strictly, one physical line at a
  time. A line that cannot decode is a named rejection ("line is not valid
  UTF-8") and a loss; the reader continues at the next line rather than
  abandoning the file. The size cap now measures the bytes it rejects.
* Every loss a rejected line represents is FILE-GLOBAL. No rejected line may
  name a scope, whatever its scope_id looks like. This over-blocks — one corrupt
  line cuts scopes that were never damaged — and that is the chosen half of the
  trade: epochs recover, a fail-open crossing of a real loss does not. It also
  removes the ghost keys and the unbounded cardinality in one move rather than
  three patches (2000 rejected labels now add zero public map keys).
* `history.damage.scope_local` keeps exactly one source, found by testing the
  hypothesis rather than assuming it: a record that PASSED validation and whose
  own canonical loss counters prove evidence was lost. Measured on the old head,
  such a record kept history.quality DEGRADED at +2, +6 and +60 healthy records,
  while the identical populations without it promoted every time — the same
  fail-stuck shape the epoch model exists to remove, through the one door it did
  not watch. It now opens a scope-local recovery boundary after itself, and
  belongs to the epoch it closes rather than the one it opens.

Deduplication identity drops the scope-local half of the epoch stamp: across a
file-global loss the reader cannot tell whether a repeated run_id is the same
run, so both copies stand; across a trusted boundary the file is intact and the
repeat is the same run retrying, so its cleaner copy cannot enter the recovered
epoch and launder the loss its twin reported.

Also found while verifying privacy: valid_record() echoed a rejected line's own
`schema_version` VALUE into `history.rejected`, which reaches --json, the human
report and any log that keeps them. Truncating it is not a bound — thirty
characters of a credential is still the credential — so only a number is echoed
and anything else is named by type.

PARTIAL (a bound the caller asked for), UNKNOWN (a schema that cannot attest)
and INVALID/EMPTY (readable but not comparable) deliberately create no boundary:
the repair target is loss continuity, not every ineligible record.

Two frozen B1 expectations are superseded rather than quietly re-run, with the
reason and the replacement recorded in docs/V142_COUNTEREXAMPLES.md §6.6.

17 suites, 1545 assertions, 0 failures. 66 mutants, all RED on the mutation and
GREEN on the real source. Six planted bad implementations — lenient decode,
trusted rejected scope, ignored loss, permanent global cut, no candidate ever,
no recovery boundary — are each caught by the suites. Schema-4 encoded bytes are
identical to base main for clean, malformed and bounded acquisitions, and the
legacy optimizer still refuses schema-4 by name.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…pe is not our loss

Two findings, both reproduced before they were believed.

A cross-family review of the trust repair found a HIGH in the recovery mechanism
itself. The scope-local boundary was opened at READ time, before deduplication,
so a copy that was then discarded as `duplicate run_id (retry)` still moved the
counter — and this reader's own rule calls a deduplicated retry bookkeeping
rather than a loss. Its fixture, measured on the previous head:

    DEGRADED X · 6 healthy records · DEGRADED X again
    -> INSUFFICIENT_DATA, records_in_epoch=0, scope_local={"a": 2}

Six healthy records sat between the two copies and none of them survived, and
repeating `6 healthy + one more copy of X` held the scope down indefinitely: the
fail-stuck shape this repair exists to remove, rebuilt out of its own recovery
mechanism. A boundary is now opened at most once per `(scope, file-global epoch,
run_id)`. A record with no `run_id` cannot be shown to be a retry and stays its
own observation; across a file-global loss the identity differs, because there
the reader cannot tell whether a repeated id is the same run at all. A genuinely
different second loss still cuts (`VD_13`).

The second was found while verifying §6 of the task — a migration must not read
as file damage — rather than reported by anyone. A line carrying another tool's
`record_type` in a shared history was classified `unknown record_type` and
counted as a LOSS, so a foreign entry cut the file for every scope. The
distinction `no shares` already draws one check further down now applies here
too: a line that ALSO carries our fields (schema_version / run_id / carry_bytes)
with an unknown record_type is a corrupted record of ours and stays a loss; a
line that carries none of them is another tool's entry and is counted without a
boundary.

New frozen cases VD_12, VD_12b, VD_13, VD_13b, VD_13c and a foreign-record_type
control; new mutants M_DEGRADED_RETRY_CUTS_TWICE and
M_FOREIGN_RECORD_TYPE_IS_DAMAGE, both with their positive control in the same
body. docs/V142_COUNTEREXAMPLES.md §6.10 records the finding and its fixture.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…d before acceptance

Round 2 of author review returned NEEDS-FIX with a HIGH that goes to the root of
the trust model's own rationale. §6.2 justified the one trusted scope-local
boundary by saying "the record passed structural validation and its scope_id is
the same field every accepted record publishes". The first half was true. The
second was an assumption: valid_record() never checked scope_id at all, and
scope_of() is str(r.get("scope_id") or "default").

Reproduced before it was believed:

    6 clean rev · one valid DEGRADED record with scope_id = ["rev"] · 6 clean rev
      CANDIDATE / exit 10 / one candidate file
      records_in_epoch=12, comparable=12
      damage {"file_global": 0, "scope_local": {"['rev']": 1}}
      candidate names rid-rev-0..5 AND rid-rev-20..25 — both sides of the loss

That is the blocked B-UTF8 defect rebuilt through the one door this repair
opened: attribution taken from a field nobody had checked, a cut landing on a
population that does not exist, the real one crossing the loss.

The rule is the producer's own contract, not a new invention. tools/carry.py
writes exactly one shape — `str(scope_id or "default")[:64]` — so a value of
another type, or a string longer than that cap, was not written by it. Such a
record is refused with a STATIC reason (its content must never be echoed) and is
a file-global loss like any other unattributable one. Measured after the fix, the
reviewer's fixture gives file_global=1, an active epoch of six, scope.known ==
["rev"], and a candidate naming only rid-rev-20..25.

A control byte inside a label stays ACCEPTED, deliberately: the producer's cap
truncates length but does not strip control characters, such a label is already
published through scope.known on every head of this branch, and calling it
corruption would invent damage where the file is intact. Stated in
docs/V142_COUNTEREXAMPLES.md §6.11 rather than hidden.

Frozen as TSCOPE_08 with four forged shapes and five producer-writable controls;
mutant M_SCOPE_LABEL_UNCHECKED carries its own positive control.

17 suites, 1563 assertions, 0 failures. 69 mutants, 0 survivors.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The author-adversarial review of 71016ee found that two materially different
valid observations sharing one run_id were collapsed into a single retry, so the
second loss never opened a boundary and a candidate was written whose evidence
spanned it. Reproduced on that head first, with four controls.

Frozen here, before any code: run_id is an identity CLAIM, not proof of semantic
equality. Two records are one retry only when their identity AND their canonical
persisted observation match; a same-identity pair whose observations differ is a
RUN_ID_CONFLICT — an integrity event that opens a recoverable scope-local
boundary at the later record's own position and is never reported as a retry.

Section 7.3 argues the complete-versus-complete row rather than inferring it, and
7.4 records the limit no reading of this format can close: identical content under
a reused run_id stays indistinguishable from one identical retry.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Fourteen failures on 71016ee, one per row of docs/V142_COUNTEREXAMPLES.md §7.

The rows that already pass are the ones the repair must NOT change: an exact
duplicate (RID_01), a key-order-only difference (RID_10), a whitespace-only
difference (RID_11), the same id in another scope (RID_08) and the same id across
a file-global loss (RID_09). They are the positive controls against reintroducing
the fail-stuck behaviour the previous round removed.

The rows that fail are the defect: a materially different observation under one
run_id is reported as "duplicate run_id (retry)", opens no boundary, and lets a
candidate rest on evidence from both sides of the second loss.

The last three checks pin the identity itself: a reader annotation
(_history_epoch, _evidence_quality_floor) may never change what a record is,
while any persisted field — including one this reader does not know — must.
optimize.observation_digest does not exist yet, which is the point.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…equality

The author-adversarial review of 71016ee found the last collapse this reader
still made: two materially different valid observations carrying one run_id were
treated as a single retry, so the second loss never opened a boundary.

Reproduced first, with four controls:

    DEGRADED X1(ts 0,  share 30, unreadable 2, sessions 40, turns 1000, scanned 80)
    6 healthy
    DEGRADED X2(ts 40, share 66, unreadable 9, sessions 91, turns 2400, scanned 150)
    6 healthy                                    [one scope, one file-global epoch]

      CANDIDATE / exit 10 / 1 candidate file
      records_in_epoch 12, comparable 12, scope_local boundaries 1 for TWO losses
      "duplicate run_id (retry)" = 1
      candidate evidence = 6 records from before X2 and 6 from after it

Nine persisted fields differ; both records pass validation and are independently
DEGRADED. The controls: a different run_id already gave two boundaries and an
epoch of six, an identical duplicate gave one and twelve, and a file-global loss
between the copies gave two plus one.

Two records are the same run only when their identity AND their persisted
observation match. Equivalence is a digest of the WHOLE parsed record with keys
sorted — a hand-picked subset would be the same mistake in a new spelling — so
key order and whitespace cannot make identical observations differ, a persisted
field this reader does not know still can, and the reader's own annotations
(_history_epoch, _evidence_quality_floor) are excluded, because reading a file
must not change what a record is.

A same-identity pair whose observations differ is a RUN_ID_CONFLICT: the later
record is kept, it is never reported as a retry, and it opens a recoverable
scope-local boundary at its own physical position, belonging to the epoch it
closes. Complete-versus-complete counts too, and §7.3 of the counterexample
document argues that row rather than inferring it: nothing was lost, but the file
states two things under one identity and a trend drawn across that point is drawn
over a file whose identity discipline has already failed.

One classification drives the boundary, the deduplication, the quality floor and
the diagnostics. The separate second pass is gone: a boundary layer and a dedup
layer holding two notions of identity is exactly how they came to disagree about
one pair. Across a file-global loss nothing changes — the reader cannot establish
continuity there, so the later copy is a fresh identity.

history.run_id_conflicts is a bounded integer, deliberately outside
history.rejected (a conflicting record is accepted, not a line that failed to
become one) and deliberately not a map: 600 distinct conflicting ids produce the
integer 600 and zero new keys. output_schema_version stays 2.

Five earlier cases used one run_id for two different observations and asserted a
retry. Each is replaced rather than quietly re-run, with the reason recorded in
§7.6 before the edit, and each keeps the property it protected — now enforced by
a boundary instead of by a quality floor, which is strictly stronger. A
consequence stated rather than hidden: a true retry now has the same quality by
construction, so QUALITY_FLOOR is a guard rather than a live path.

§7.4 records the limit no reading of this format can close: identical persisted
content under a reused run_id stays indistinguishable from one identical retry.
Line position is deliberately not used as identity — that would make every true
retry open a fresh boundary and rebuild the fail-stuck behaviour just removed.

17 suites, 1601 assertions, 0 failures. 77 mutants, 0 survivors.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant