v1.4.2: fail closed on optimizer evidence quality - #13
Merged
Merged
Conversation
Thirteen cases, every expected status, exit code, candidate-file count and comparable count written down BEFORE any repair exists, and then run against 77e3677 to record what it actually does. Seven of them are defects the current legacy optimizer really has, on these exact fixtures: - a history built entirely from bounded sweeps promotes a trend candidate, and --accept-partial changes nothing because history quality is never read at all - the newest INVALID record picks the comparison anchor and then removes itself, stranding six eligible records (comparable 0) - HOST_BEHAVIOR_SHIFT reports exit 30 while writing the candidate file - six records claiming COMPLETE with sessions=turns=carry_bytes=0 promote - a transcript that lost a record to torn JSON still reports the sweep COMPLETE, and the run reports NO_ACTION - --max-files 1 gives the carry sweep and the skill-listing scan two different single-source samples - a schema-1 record, from a generation that had no evidence_quality field at all, is read as COMPLETE Four cases are positive controls (a clean COMPLETE population must still promote; an invalid record in another scope must not poison this one; schema-4 evidence must stay refused by name; a flat population must stay NO_ACTION) and two pin the strict-exit contract. No source file is touched by this commit: the suite is expected to fail here, and the matrix in docs/V142_COUNTEREXAMPLES.md records both the frozen expectations and what main did. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…e emission gate The legacy optimizer decided five different things about evidence in five different places, and one of them decided nothing at all. This makes the question one function, asks it of the evidence each finding actually rests on, and puts the candidate-writing gate where every caller has to pass through it. What changed, in the order the counterexamples found it: - `record_quality()` / `sweep_quality()` are the law. The worst of what a record claims and what its own numbers can show, over the counters the producer has written since schema 2. A schema-1 record has no evidence_quality field at all, so a schema-1 record carrying COMPLETE is a claim its writer could not have made: UNKNOWN. A schema-2 record without the acquisition counters cannot attest completeness either. - PARTIAL and DEGRADED are now different things. A bound the caller asked for (`--max-files`) is PARTIAL and is exactly what `--accept-partial` adopts; evidence that was selected and then lost is DEGRADED and no flag accepts it. carry.LOSS_FIELDS / BOUND_FIELDS is the single vocabulary both sides read. - A sweep that lost records to torn JSON is no longer COMPLETE. `malformed` joins the loss counters, so carry.sweep_label() degrades the sweep and the optimizer reads it as DEGRADED. - `eligible_anchor()` picks the comparison anchor from records that survive their own filter. The newest INVALID record used to choose the scope, the workload class and the corpus size for everyone else and then remove itself, stranding a whole eligible population. - The history trend is gated. It was `if True:` — a trend over bounded sweeps was indistinguishable from one swept in full. Now it is a CANDIDATE on eligible history and an OBSERVED finding, with the reason in its evidence line, on anything else. - Each finding is gated on ITS OWN evidence: live findings on the sweep, the trend on the eligible history, the guard on its own ledger sample. A partial history no longer blocks a live finding it never supported. - `emit_candidates()` takes the status and writes nothing unless it is CANDIDATE. HOST_BEHAVIOR_SHIFT reported exit 30 while leaving a specification on disk; the gate now lives in the emitter, so the CLI, the long-run simulator and any future caller route through it. - `carry.bounded_paths()` is the one definition of "the newest N", used by both the carry sweep and the skill-listing scan. One `--max-files 1` run was analysing two different single-source populations. Schema-4 evidence stays unsupported by this optimizer and the typed promotion stays shadow-only. Proved rather than asserted: the encoded v1.4 record for a clean, a torn and a bounded sweep has a byte-identical sha256 before and after this commit, and the typed reader's loss_observed verdict is unchanged. Only the legacy `quality` label moved, for the torn sweep, which is the defect being repaired. Four fixtures claimed COMPLETE without the acquisition counters a real sweep always writes (multi-agent, mutation, long-run, readiness). Each now states them, and each edit is locked by an added assertion that the same record WITHOUT them refuses to promote — so the fixture change cannot hide the gate it was making room for. skills/ and hooks/ are untouched: CANONICAL_BODY_SHA256 is unchanged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Six new mutants, each restoring exactly one of the defects the frozen matrix found, each RED on the mutant and GREEN on the real source: M_PARTIAL_PROMOTES may_promote() returns True M_INVALID_ANCHOR the anchor is chosen before the filter again M_HOST_SHIFT_WRITES the emitter's status gate is deleted M_MALFORMED_COMPLETE `malformed` leaves the loss vocabulary M_BOUND_SAMPLE_DIVERGES the listing scan takes its own slice again M_TREND_QUALITY_BYPASS the history trend stops asking None of them is a string mutant: every one changes behaviour that the oracle observes through a real run — a status, a comparable count, a file on disk, a sweep's quality, a JSON listing field. Two existing mutation patterns moved with the code and were re-pointed (the PARTIAL gate and the trend's direction), and the shared record fixture states its acquisition counters for the same reason the other fixtures do. README: 1311 assertions in seventeen suites (CI enforces this number). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Three from an adversarial pass over the new law, one from a cross-family review of the same diff. Each has a case in the suite; the first also has a mutant. - A retry that reuses a run_id could launder its own sweep. The duplicate is dropped as an OBSERVATION, but the quality it reported is part of what that id actually saw, so the worse of the two now travels with the record that survives (QUALITY_FLOOR, in memory only — nothing is written back to the history file). - `schema_version` of "2" (a string) or 2.0 (a float) was accepted as schema 2 by int(). A version this reader cannot name is not a newer generation to trust; it is an older one to doubt. - Zero counters say "nothing went wrong"; they do not say a sweep happened. A schema-2 record carrying zeroed counters and no `sessions`, `turns` or `carry_bytes` at all read COMPLETE — a share vector with no population behind it. It now reads UNKNOWN. (cross-family review) - One source whose mtime could not be read sent the ENTIRE bounded selection back to a discovery-order slice — the exact sample bounded_paths() exists to avoid. The fallback is now per source: a file that cannot be dated cannot claim to be the newest, and the rest still order by mtime. (cross-family review) README: 1330 assertions. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Four rows, each from an attack on the repair rather than on the original defect, kept in their own section so the frozen matrix stays readable as what it was when it was frozen. No expectation in section 1 changed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
… nobody checked" A cross-family adversarial pass over the repair itself. Every one of these produced a CANDIDATE and a file on disk before this commit. - Counters that cannot describe one sweep. `sessions=40` out of `scanned=1` — or out of `scanned=0` — read COMPLETE, because the law checked each counter's type and sign but never their relation. A session is a transcript that was scanned and cleared the turn floor, so the producer can never report more sessions than it scanned. That combination is now INVALID, and `scanned` joins the fields a schema-2 record must carry before its zeroes mean anything. - Damage to the history CONTAINER never reached the evidence decision. A torn line in the history file was counted in `history.rejected` and then ignored: the trend built from the surviving records was promoted as if nothing had been lost. `container_quality()` reads it as a loss (DEGRADED, so no flag adopts it) and both `analyse()` and `overall_status()` take the worst of it and the records. A refusal by design — current-generation evidence this optimizer does not read — and a deduplicated retry are NOT damage; both have controls. - A ledger that lost a line still retired the guard. 100 readable writes plus one torn line promoted `noop-guard-retire`. `ledger_usable` now requires `rejected == 0`, and a ledger that exists but cannot be opened is no longer indistinguishable from no ledger at all. - A run could report CANDIDATE and exit 10 while every candidate file failed to write. The status has to say what happened: nothing landed, so the run is INTERNAL_ERROR (exit 50), with the failure named in the JSON. Four more mutants, each RED on the mutant and GREEN on the real source (41/41 now): M_IMPOSSIBLE_COUNTERS, M_CONTAINER_DAMAGE_IGNORED, M_LEDGER_TORN_PROMOTES, M_EMIT_FAILURE_SILENT. README: 1360 assertions. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…were foreign The first cut degraded the history for any rejected line, including a well-formed JSON line that simply is not a carry record. In a shared file one stray append would then block a real population forever — a fail-closed patch that refuses evidence it should still read is also a defect. Damage is now what it says: a torn line, a line past the size cap, a file that could not be opened. A foreign line and a refusal by design are counted and reported, and do not degrade; both have controls in the suite. The limit is written down rather than implied: a corrupted carry record that still parses as JSON is reported and does not degrade. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A confirmation-round reviewer read the `eligible_anchor(recs) or newest` fallback in main() as a way back in for an ineligible record. Their exact input reproduces the observation — the run names the scope of the one INVALID record it found — and it stops there: status PARTIAL_EVIDENCE, exit 40, nothing on disk, because the fallback is only reached when NOTHING in the file is eligible, and then there is no population to strand and nothing to promote. Both halves are now tests rather than an argument: the reviewer's input, and the same record added to a real population in another scope, where the eligible records still choose the scope and still promote. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…locks Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The confirmation lane attacked the repair rather than the original defect and found four more places where a word was taken for evidence. - A record could say PARTIAL while every counter it carried said nothing happened, and --accept-partial adopted it. A bound shows up in skipped_by_limit and a loss shows up in a loss counter; a PARTIAL claim that no counter can explain cannot say WHY it was partial, so it is UNKNOWN and no flag adopts it. - The record-level loss vocabulary was two counters wide (unreadable, oversize) while the contract promises more. A record carrying malformed, malformed_lines, identity_changed, conflicted_sources or records_rejected read COMPLETE — and a retry carrying one could launder itself through dedup, because the floor was computed with the same short list. REQUIRED (what the schema-2 writer always wrote, so its absence means UNKNOWN) and LOSS (what must be honoured when present) are now separate lists. - Container damage was an allowlist of three reasons, so a rejection valid_record() learns to make later would default to "not damage". It is now the other way round: a rejected line is damage unless it is one of three named exceptions — a refusal by design, a deduplicated retry, or a well-formed line that was never a carry record. An unsupported schema-3 record and a record whose shares do not sum to a population now close the emitter instead of being footnotes. - A ledger line that is neither `checked` nor `denied` was counted as nothing at all, so a ledger full of unknown events looked like a clean sample. It counts as rejected, which the guard's own gate reads. Four more mutants (45/45). The schema-4 contract is still untouched: the encoded v1.4 record for a clean, a torn and a bounded sweep has the same sha256 as on main. README: 1399 assertions. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…, no half files A fourth adversarial pass returned no HIGH that survived checking (the one HIGH claimed a missing isinstance guard in load_ledger that is present, and a list-valued ledger line is counted as rejected, verified by execution). Four MEDIUMs were real: - `sweep_quality()` did not apply the impossibility rule its record-level twin applies: a live sweep reporting sessions with zero turns read COMPLETE. It is INVALID, like the record. - `bounded_paths()` was called twice per run — once by the sweep, once for the listing — so an mtime that became unreadable between the two calls produced two different samples again. The selection is made ONCE in main() and handed to both; `accumulate(selected=...)` takes it, and still reports the bound against the whole discovered population. - A line that CLAIMS to be one of our records and carries no shares is a corrupted record, not another tool's entry. It now reads as container damage; a line that claims nothing still does not. - A candidate write that failed mid-way left its `.tmp-<pid>` file behind. It is removed on the failure path. Two limits stay, stated rather than papered over: container damage is a property of the FILE, so a torn line degrades every scope in it (an unparseable line has no scope to attribute it to), and HOST_BEHAVIOR_SHIFT still closes the emitter for findings that do not rest on the shifted population — that ordering is the release contract's, not this patch's. 45/45 mutants. README: 1408 assertions. The schema-4 encoded record for a clean, a torn and a bounded sweep still has the same sha256 as main. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
An independent acceptance review BLOCKED the first head of this branch. The six original defects were closed, but the container-damage repair introduced a new one: damage was permanent. One torn line — the crash fragment docs/MULTI_AGENT.md calls an expected event, the one the reader is designed to "reject exactly that line and count it" — set the whole file DEGRADED forever, for every scope sharing it, with no flag able to adopt a loss and nothing in the product that expires or rotates a history. Measured on the current head, all of these stay PARTIAL_EVIDENCE with zero candidates: 6 good + torn; + 2 more; + 6 more; + 60 more; a sibling scope's records; a loss before any record; and a legitimate zero-carry record. This commit freezes what the repair must do, before it exists: thirteen rows in docs/V142_COUNTEREXAMPLES.md §5 with status, exit code, candidate files, comparable count, active-epoch size, history quality and boundary count, plus the two MEDIUMs the same review found (history.quality read COMPLETE with nothing comparable; a producer-valid zero-carry record read as corruption). The model being frozen: an unattributable loss cuts the promotion history at that line's PHYSICAL position. Evidence before the cut never joins evidence after it, the newest epoch is by construction damage-free, and the loss stays reported. No time window, no expiry, no ratio, no new override flag, and --accept-partial still adopts a chosen bound and never a loss. 21 assertions fail here, each naming its frozen expectation. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
… file
B1, the blocker an independent acceptance review found in this branch's
first head. Container damage was permanent: one torn line — the crash
fragment docs/MULTI_AGENT.md calls an expected event — set the whole file
DEGRADED, and since nothing here expires, rotates or repairs a history
and no flag adopts a loss, a single crash disabled promotion for every
scope sharing ~/logs/carry_history.jsonl, forever. Measured before this
commit: 6 good + torn + 60 good still returned PARTIAL_EVIDENCE with zero
candidates.
A loss is now a BOUNDARY at the rejected line's physical position in the
file, not a verdict on the file:
- `damage_boundary()` is the one classifier every rejection goes through,
including the ones that never became an object. A line that parses and
carries a readable `scope_id` cuts that scope; a line that cannot say
whose record it was — unparseable, oversized, a file that would not
open — cuts every scope, because guessing would be the fail-open half.
A foreign line, a refusal by design and a deduplicated retry are not
losses and cut nothing.
- Records are stamped at READ time with the epoch they were written in:
(file-global losses before them, losses attributed to their own scope
before them). Physical order, never the clock, because the clock is
what a damaged history cannot be trusted about.
- `active_records()` is the analysed population: what came after the
newest loss that applies to its scope. It is selected BEFORE anything
is anchored, so a pre-loss record cannot choose the scope, the workload
class, the corpus size or the trend for the population after it.
- The active epoch is therefore damage-free by construction, which is why
the container no longer gates: `analyse()` and `overall_status()` ask
only about the evidence eligible right now.
- The loss stays visible: `history.rejected` counts it and the new
`history.damage` says how many boundaries the file holds, in the JSON
and in the human report.
Deduplication is now per epoch. The same run_id after a loss is that
population's own observation; inside one epoch the retry is still
dropped, counted, and still cannot launder the survivor's quality.
Two MEDIUMs from the same review, repaired here because the epoch model
makes both sharper:
- `history.quality` is the quality of the evidence eligible for THIS
analysis, and is EMPTY when nothing is comparable. It read COMPLETE
with zero comparable records. `worst_quality([])` still means COMPLETE:
a finding resting on no sampled evidence is not degraded by sampling it
never used.
- A zero-carry sweep is readable evidence, not corruption. tools/carry.py
says so in its own report — "A session whose every item lands on its
final turn carries nothing" — the 1.3 writer emits `shares: {}` for it
and the current one guards `if C else {}`. Proved by running it:
accumulate() on such a transcript returns sessions=1 turns=6 carry=0.
The record reads EMPTY, is left out of comparable, and cuts nothing.
The contradictions stay INVALID: carry with no shares, shares with no
carry, sessions without turns, sessions beyond what was scanned.
--accept-partial is unchanged and still adopts only a caller's chosen
bound, never a loss. No time window, no expiry, no ratio, no override
flag, no new status. Schema-4 stays unsupported and its encoded bytes are
unchanged (clean, malformed and bounded sweeps all hash identical to
main). 1460 assertions, 53 mutants.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A run over a damaged file now says how many records the scope holds, how many of them came after the newest loss, and what the loss was — the same distinction the JSON makes between what the file holds and what the current analysis rests on. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…es inside one epoch The chosen semantics, pinned rather than inherited: the copy that lands after a loss is judged on its own evidence (clean before, bounded after -> PARTIAL), the copy excluded on the far side of the loss does not poison the epoch that follows it (bounded before, clean after -> COMPLETE, and no refusal), and three copies inside ONE epoch still leave the worst of them on the survivor (DEGRADED, two retries counted). Which copy comes first is exactly what a crash decides, so both orders are tests now. README: 1463 assertions. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Found by a cross-family lane attacking the epoch model itself, and both reproduced before being believed. - A finding that rests on the ledger alone may promote after a loss — that is per-finding evidence working — but the SCOPE it was filed under still came from the records before the loss, and the scope travels into the candidate id. A history of six `stale` records followed by a torn line produced `noop-guard-retire-stale-…` and wrote it. The scope now comes from the current epoch or from nothing: `default` when the epoch is empty, never a population that no longer exists. - The epoch stamp is a pair of COUNTS, so a record of scope "a" and one of scope "b" can carry the same numbers while belonging to different epochs. Deduplication keyed on the stamp alone therefore treated scope "b"'s record as scope "a"'s retry and pushed its PARTIAL quality onto scope "a" through QUALITY_FLOOR, turning a clean population into PARTIAL_EVIDENCE. Identity now carries the scope. - The long-run simulator discarded the epoch it was handed, so its analysis could combine records across a loss. It uses active_records() and the same damage summary the CLI does. Two mutants (M_STALE_SCOPE_ANCHOR, M_DEDUP_IGNORES_SCOPE) and four assertions (B1_14, B1_15) pin all of it; the older anchor mutant moved with the code. 1469 assertions, 55 mutants, schema-4 bytes still equal to main, skill body untouched. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
B1_14 (a stale population naming a post-loss candidate) and B1_15 (two scopes sharing one epoch stamp), in the same post-freeze section as the rest, with the long-run simulator's missing epoch named too. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…enting it The adversarial review of the epoch model found attribution taken from the line that had just failed validation: invalid UTF-8 in scope_id survived errors="replace" as U+FFFD, read as a readable label, and cut a scope that does not exist while the real population kept crossing the loss. Frozen here, before any code: rejected lines are never an authority for their own scope (every loss is file-global), the history is decoded strictly per physical line, and the one trusted source of a scope-local boundary is a record that passed validation and whose own canonical loss counters prove the loss. Section 6.3 records the hypothesis reproduced first, with a matched control: VALID_DEGRADED_FAIL_STUCK=YES. Section 6.6 records the two B1 expectations this trust model supersedes, and what replaces each. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…oundary 38 failures on e6c1f4f, one per row of docs/V142_COUNTEREXAMPLES.md §6: TUTF_01..08 invalid UTF-8 in a physical line is an unattributable loss. Today errors="replace" turns it into U+FFFD, the reader calls that a readable scope_id, and a scope that does not exist is cut while the real population keeps crossing the gap — the candidate names run ids from both sides of the loss. TSCOPE_01..07 a record that failed validation is not an authority for its own scope_id, whatever the string looks like. Canaries and a 2000-label cardinality attack pin that no rejected label can become a public map key. VD_01..11 a record that PASSED validation and whose own canonical loss counters prove a loss is the one trusted source of a scope-local boundary. Reproduced first, with a matched control: today that record keeps history.quality DEGRADED forever, at +2, +6 and +60 healthy records. MSG_01/02 file-global recovery, already correct, pinned against regress. The write() helper now takes bytes: a line that is not valid UTF-8 cannot be written through a text handle, and that line is the input under test. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Closes the HIGH blocker the adversarial review of the epoch model found, and the MEDIUM and two LOWs that shared its root cause. The defect was not Unicode. It was provenance: attribution was read off the very line that had just failed validation. Reproduced first, on e6c1f4f: six clean records, one SameWrite-shaped line whose `scope_id` holds invalid UTF-8 and which also fails validation, six more clean records. errors="replace" turned the bytes into U+FFFD, damage_boundary() called that a readable label and cut a scope nobody has, the real population was never cut, and the run wrote a candidate whose evidence_run_ids span both sides of the loss (CANDIDATE / exit 10 / epoch 12 / comparable 12). Three changes, one rule: A record that failed validation is not a trustworthy authority for its own scope attribution. * The history is read as BYTES and decoded strictly, one physical line at a time. A line that cannot decode is a named rejection ("line is not valid UTF-8") and a loss; the reader continues at the next line rather than abandoning the file. The size cap now measures the bytes it rejects. * Every loss a rejected line represents is FILE-GLOBAL. No rejected line may name a scope, whatever its scope_id looks like. This over-blocks — one corrupt line cuts scopes that were never damaged — and that is the chosen half of the trade: epochs recover, a fail-open crossing of a real loss does not. It also removes the ghost keys and the unbounded cardinality in one move rather than three patches (2000 rejected labels now add zero public map keys). * `history.damage.scope_local` keeps exactly one source, found by testing the hypothesis rather than assuming it: a record that PASSED validation and whose own canonical loss counters prove evidence was lost. Measured on the old head, such a record kept history.quality DEGRADED at +2, +6 and +60 healthy records, while the identical populations without it promoted every time — the same fail-stuck shape the epoch model exists to remove, through the one door it did not watch. It now opens a scope-local recovery boundary after itself, and belongs to the epoch it closes rather than the one it opens. Deduplication identity drops the scope-local half of the epoch stamp: across a file-global loss the reader cannot tell whether a repeated run_id is the same run, so both copies stand; across a trusted boundary the file is intact and the repeat is the same run retrying, so its cleaner copy cannot enter the recovered epoch and launder the loss its twin reported. Also found while verifying privacy: valid_record() echoed a rejected line's own `schema_version` VALUE into `history.rejected`, which reaches --json, the human report and any log that keeps them. Truncating it is not a bound — thirty characters of a credential is still the credential — so only a number is echoed and anything else is named by type. PARTIAL (a bound the caller asked for), UNKNOWN (a schema that cannot attest) and INVALID/EMPTY (readable but not comparable) deliberately create no boundary: the repair target is loss continuity, not every ineligible record. Two frozen B1 expectations are superseded rather than quietly re-run, with the reason and the replacement recorded in docs/V142_COUNTEREXAMPLES.md §6.6. 17 suites, 1545 assertions, 0 failures. 66 mutants, all RED on the mutation and GREEN on the real source. Six planted bad implementations — lenient decode, trusted rejected scope, ignored loss, permanent global cut, no candidate ever, no recovery boundary — are each caught by the suites. Schema-4 encoded bytes are identical to base main for clean, malformed and bounded acquisitions, and the legacy optimizer still refuses schema-4 by name. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…pe is not our loss
Two findings, both reproduced before they were believed.
A cross-family review of the trust repair found a HIGH in the recovery mechanism
itself. The scope-local boundary was opened at READ time, before deduplication,
so a copy that was then discarded as `duplicate run_id (retry)` still moved the
counter — and this reader's own rule calls a deduplicated retry bookkeeping
rather than a loss. Its fixture, measured on the previous head:
DEGRADED X · 6 healthy records · DEGRADED X again
-> INSUFFICIENT_DATA, records_in_epoch=0, scope_local={"a": 2}
Six healthy records sat between the two copies and none of them survived, and
repeating `6 healthy + one more copy of X` held the scope down indefinitely: the
fail-stuck shape this repair exists to remove, rebuilt out of its own recovery
mechanism. A boundary is now opened at most once per `(scope, file-global epoch,
run_id)`. A record with no `run_id` cannot be shown to be a retry and stays its
own observation; across a file-global loss the identity differs, because there
the reader cannot tell whether a repeated id is the same run at all. A genuinely
different second loss still cuts (`VD_13`).
The second was found while verifying §6 of the task — a migration must not read
as file damage — rather than reported by anyone. A line carrying another tool's
`record_type` in a shared history was classified `unknown record_type` and
counted as a LOSS, so a foreign entry cut the file for every scope. The
distinction `no shares` already draws one check further down now applies here
too: a line that ALSO carries our fields (schema_version / run_id / carry_bytes)
with an unknown record_type is a corrupted record of ours and stays a loss; a
line that carries none of them is another tool's entry and is counted without a
boundary.
New frozen cases VD_12, VD_12b, VD_13, VD_13b, VD_13c and a foreign-record_type
control; new mutants M_DEGRADED_RETRY_CUTS_TWICE and
M_FOREIGN_RECORD_TYPE_IS_DAMAGE, both with their positive control in the same
body. docs/V142_COUNTEREXAMPLES.md §6.10 records the finding and its fixture.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…d before acceptance
Round 2 of author review returned NEEDS-FIX with a HIGH that goes to the root of
the trust model's own rationale. §6.2 justified the one trusted scope-local
boundary by saying "the record passed structural validation and its scope_id is
the same field every accepted record publishes". The first half was true. The
second was an assumption: valid_record() never checked scope_id at all, and
scope_of() is str(r.get("scope_id") or "default").
Reproduced before it was believed:
6 clean rev · one valid DEGRADED record with scope_id = ["rev"] · 6 clean rev
CANDIDATE / exit 10 / one candidate file
records_in_epoch=12, comparable=12
damage {"file_global": 0, "scope_local": {"['rev']": 1}}
candidate names rid-rev-0..5 AND rid-rev-20..25 — both sides of the loss
That is the blocked B-UTF8 defect rebuilt through the one door this repair
opened: attribution taken from a field nobody had checked, a cut landing on a
population that does not exist, the real one crossing the loss.
The rule is the producer's own contract, not a new invention. tools/carry.py
writes exactly one shape — `str(scope_id or "default")[:64]` — so a value of
another type, or a string longer than that cap, was not written by it. Such a
record is refused with a STATIC reason (its content must never be echoed) and is
a file-global loss like any other unattributable one. Measured after the fix, the
reviewer's fixture gives file_global=1, an active epoch of six, scope.known ==
["rev"], and a candidate naming only rid-rev-20..25.
A control byte inside a label stays ACCEPTED, deliberately: the producer's cap
truncates length but does not strip control characters, such a label is already
published through scope.known on every head of this branch, and calling it
corruption would invent damage where the file is intact. Stated in
docs/V142_COUNTEREXAMPLES.md §6.11 rather than hidden.
Frozen as TSCOPE_08 with four forged shapes and five producer-writable controls;
mutant M_SCOPE_LABEL_UNCHECKED carries its own positive control.
17 suites, 1563 assertions, 0 failures. 69 mutants, 0 survivors.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The author-adversarial review of 71016ee found that two materially different valid observations sharing one run_id were collapsed into a single retry, so the second loss never opened a boundary and a candidate was written whose evidence spanned it. Reproduced on that head first, with four controls. Frozen here, before any code: run_id is an identity CLAIM, not proof of semantic equality. Two records are one retry only when their identity AND their canonical persisted observation match; a same-identity pair whose observations differ is a RUN_ID_CONFLICT — an integrity event that opens a recoverable scope-local boundary at the later record's own position and is never reported as a retry. Section 7.3 argues the complete-versus-complete row rather than inferring it, and 7.4 records the limit no reading of this format can close: identical content under a reused run_id stays indistinguishable from one identical retry. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Fourteen failures on 71016ee, one per row of docs/V142_COUNTEREXAMPLES.md §7. The rows that already pass are the ones the repair must NOT change: an exact duplicate (RID_01), a key-order-only difference (RID_10), a whitespace-only difference (RID_11), the same id in another scope (RID_08) and the same id across a file-global loss (RID_09). They are the positive controls against reintroducing the fail-stuck behaviour the previous round removed. The rows that fail are the defect: a materially different observation under one run_id is reported as "duplicate run_id (retry)", opens no boundary, and lets a candidate rest on evidence from both sides of the second loss. The last three checks pin the identity itself: a reader annotation (_history_epoch, _evidence_quality_floor) may never change what a record is, while any persisted field — including one this reader does not know — must. optimize.observation_digest does not exist yet, which is the point. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…equality The author-adversarial review of 71016ee found the last collapse this reader still made: two materially different valid observations carrying one run_id were treated as a single retry, so the second loss never opened a boundary. Reproduced first, with four controls: DEGRADED X1(ts 0, share 30, unreadable 2, sessions 40, turns 1000, scanned 80) 6 healthy DEGRADED X2(ts 40, share 66, unreadable 9, sessions 91, turns 2400, scanned 150) 6 healthy [one scope, one file-global epoch] CANDIDATE / exit 10 / 1 candidate file records_in_epoch 12, comparable 12, scope_local boundaries 1 for TWO losses "duplicate run_id (retry)" = 1 candidate evidence = 6 records from before X2 and 6 from after it Nine persisted fields differ; both records pass validation and are independently DEGRADED. The controls: a different run_id already gave two boundaries and an epoch of six, an identical duplicate gave one and twelve, and a file-global loss between the copies gave two plus one. Two records are the same run only when their identity AND their persisted observation match. Equivalence is a digest of the WHOLE parsed record with keys sorted — a hand-picked subset would be the same mistake in a new spelling — so key order and whitespace cannot make identical observations differ, a persisted field this reader does not know still can, and the reader's own annotations (_history_epoch, _evidence_quality_floor) are excluded, because reading a file must not change what a record is. A same-identity pair whose observations differ is a RUN_ID_CONFLICT: the later record is kept, it is never reported as a retry, and it opens a recoverable scope-local boundary at its own physical position, belonging to the epoch it closes. Complete-versus-complete counts too, and §7.3 of the counterexample document argues that row rather than inferring it: nothing was lost, but the file states two things under one identity and a trend drawn across that point is drawn over a file whose identity discipline has already failed. One classification drives the boundary, the deduplication, the quality floor and the diagnostics. The separate second pass is gone: a boundary layer and a dedup layer holding two notions of identity is exactly how they came to disagree about one pair. Across a file-global loss nothing changes — the reader cannot establish continuity there, so the later copy is a fresh identity. history.run_id_conflicts is a bounded integer, deliberately outside history.rejected (a conflicting record is accepted, not a line that failed to become one) and deliberately not a map: 600 distinct conflicting ids produce the integer 600 and zero new keys. output_schema_version stays 2. Five earlier cases used one run_id for two different observations and asserted a retry. Each is replaced rather than quietly re-run, with the reason recorded in §7.6 before the edit, and each keeps the property it protected — now enforced by a boundary instead of by a quality floor, which is strictly stronger. A consequence stated rather than hidden: a true retry now has the same quality by construction, so QUALITY_FLOOR is a guard rather than a live path. §7.4 records the limit no reading of this format can close: identical persisted content under a reused run_id stays indistinguishable from one identical retry. Line position is deliberately not used as identity — that would make every true retry open a fresh boundary and rebuild the fail-stuck behaviour just removed. 17 suites, 1601 assertions, 0 failures. 77 mutants, 0 survivors. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This was referenced Sep 22, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Patch-release candidate for the legacy optimizer's evidence integrity. Not 1.5, not a
schema-4 port, and nothing new is activated.
v1.4.1is not moved and its released runtime isunaffected.
0. The head history, in order
0.0 What the adversarial review blocked (head 2 → head 3)
Reproduced before anything was designed for it: six clean
revrecords, one SameWrite-shaped lineholding invalid UTF-8 inside
scope_id(and independently failing validation), six more cleanrevrecords. On
e6c1f4f:CANDIDATE, exit 10, one candidate file,records_in_epoch=12,damage.scope_local={"\ufffd\ufffd\ufffd": 1}, and the specification'sevidence_run_idsnamingrecords from both sides of the loss.
The defect was never about Unicode. It was provenance, and the repair says so as a rule:
errors="replace"destroyed the evidence that a decode had failed. A line that cannot decode is now a named
rejection (
line is not valid UTF-8) and a loss, and the reader continues at the next line ratherthan abandoning the file. The size cap measures the bytes it rejects.
whatever its
scope_idlooks like —rev,agent-b, a path, a 64-character label, or bytes thatnever decoded. This over-blocks on purpose: one corrupt line cuts scopes that were never damaged.
That is the chosen half of the trade, because epochs recover and a fail-open crossing of a real
loss does not. It also closes the MEDIUM (
history.qualitysemantics) and both LOWs (a rejectedlabel echoed into a public map, and that map's unbounded cardinality) in one move instead of three
separate patches: 2000 rejected lines carrying 2000 distinct labels now add zero public keys.
rather than assuming one. A record that PASSES
valid_record()and whose own canonical losscounters report a loss was measured on the old head to keep
history.qualityDEGRADEDat +2,+6 and +60 healthy records, while the identical populations without it promoted every time —
the same fail-stuck shape the epoch model exists to remove, reached through the one door the epoch
model did not watch. Such a record now opens a scope-local recovery boundary after itself and
belongs to the epoch it closes, never to the one it opens. It is trusted where a rejected line is
not: it passed structural validation, its
scope_idis the field every accepted record alreadypublishes through
scope.known, and the damage fact comes fromRECORD_LOSS_COUNTERS.loss the reader cannot tell whether a repeated
run_idis the same run, so both copies stand.Across a trusted boundary the file is intact and the repeat is the same run retrying, so its
cleaner copy cannot enter the recovered epoch and launder the loss its twin reported. Both
physical orders are frozen cases (
VD_10,VD_11).PARTIAL(a bound the caller asked for),UNKNOWN(a schema that cannot attest completeness) andINVALID/EMPTY(readable but not comparable) deliberately create no boundary: the repairtarget is loss continuity, not every ineligible record.
Found while verifying privacy, not reported by anyone:
valid_record()echoed a rejected line'sown
schema_versionVALUE intohistory.rejected, which reaches--json, the human report and anylog that keeps them. Truncating it is not a bound — thirty characters of a credential is still the
credential — so only a number is echoed and any other value is named by type.
The limit this repair does not close, stated rather than implied: a legacy flat record carries no
integrity tag, so a corruption that turns one valid record into a different valid record
(
scope_idflipped fromagent-atoagent-b, a share vector rewritten to another that still sumsto ~100) is indistinguishable from a record the producer meant to write. Nothing here detects that,
and nothing can — the information needed is absent from the format.
LEGACY_STRUCTURALLY_VALID_CORRUPTION_LIMITATION=YES.Two frozen B1 expectations are superseded rather than quietly re-run, each with its reason and
its replacement recorded in
docs/V142_COUNTEREXAMPLES.md§6.6:B1_13e(its middle copy carried aloss counter, so under the trust model it now opens an epoch instead of sitting inside one) and
B1_15(its two boundaries came from rejected lines, so the same shape is rebuilt from trustedones). The property each protected keeps a case.
New frozen cases:
TUTF_01–TUTF_09,TSCOPE_01–TSCOPE_07with seven label canaries and a2000-label cardinality attack,
VD_01–VD_11,MSG_01/MSG_02, andTPRIV. New mutants:M_UTF8_REPLACEMENT_ATTRIBUTED,M_REJECTED_SCOPE_TRUSTED,M_REJECTED_SCOPE_GHOST_KEY,M_REJECTED_SCOPE_CARDINALITY,M_REJECTED_VALUE_ECHOED,M_GLOBAL_DAMAGE_NOT_CUT,M_GLOBAL_DAMAGE_POISONS_FOREVER,M_VALID_DEGRADED_POISONS_SCOPE_FOREVER,M_VALID_DEGRADED_CUTS_ALL_SCOPES,M_DEGRADED_RECORD_INCLUDED_POST_BOUNDARY,M_DEGRADED_RETRY_NO_RECOVERY.0.0.1 And what an author review of the trust repair itself found
Two rounds of cross-family author review ran on this head. The first returned
NEEDS-FIXwith aHIGH, reproduced before it was believed and fixed in
189db57:Six healthy records sat between the two copies and none survived, and repeating
6 healthy + one more copy of Xheld the scope down indefinitely — the fail-stuck shape thisrepair exists to remove, rebuilt out of its own recovery mechanism. A boundary is now opened at
most once per
(scope, file-global epoch, run_id). A record with norun_idcannot be shown to bea retry and stays its own observation; a genuinely different second loss still cuts; across a
file-global loss the identity differs, because there the reader cannot tell whether a repeated id is
the same run at all. Frozen as
VD_12,VD_12b,VD_13,VD_13b,VD_13c, with the mutantM_DEGRADED_RETRY_CUTS_TWICEcarrying its own positive control.The same commit carries a second finding, made while verifying that a migration must not read as
file damage rather than reported by anyone: a line holding another tool's
record_typein ashared history was classified
unknown record_typeand counted as a loss, so a foreign entry cutthe file for every scope. The distinction
no sharesalready draws one check further down nowapplies here too — a line that also carries our fields is a corrupted record of ours and stays a
loss; a line that carries none of them is another tool's entry and creates no boundary
(
M_FOREIGN_RECORD_TYPE_IS_DAMAGE).Round 2 returned
NEEDS-FIXtoo, with a HIGH that goes to the root of the trust model's ownrationale. §0.0 justified the one trusted scope-local boundary by saying the record "passed
structural validation and its
scope_idis the field every accepted record already publishes". Thefirst half was true; the second was an assumption.
valid_record()never checkedscope_id, andscope_of()isstr(r.get("scope_id") or "default")— so a record that passes validation carryingscope_id = ["rev"]becomes the scope"['rev']":The blocked B-UTF8 defect, rebuilt through the one door this repair opened. The fix is the
producer's own contract rather than a new invention:
tools/carry.pywrites exactly one shape,str(scope_id or "default")[:64], so a value of another type or another length was not written byit. Such a record is refused with a STATIC reason — its content must never be echoed — and is a
file-global loss like any other unattributable one. A control byte inside a label stays accepted
deliberately, because the producer's cap truncates length without stripping control characters, that
label already reaches
scope.knownon every head of this branch, and calling it corruption wouldinvent damage where the file is intact. Frozen as
TSCOPE_08with four forged shapes and fiveproducer-writable controls, plus
M_SCOPE_LABEL_UNCHECKED; §6.11 of the counterexample documentstates the residual.
This is author review, not acceptance. The verdict on this head belongs to a new independent
reviewer.
0.0.2 And what the author-adversarial review of
71016eefound (head 3 → head 4)One MEDIUM, reproduced before it was believed and fixed in
6ef9988. It is the last collapse thisreader still made, and it is the same shape as the two blocked heads above, through the last door
left open: identity.
Nine persisted fields differ between the two records; both pass validation and are independently
DEGRADED. Controls on the old head: a differentrun_idalready gave two boundaries and an epochof six, an identical duplicate gave one and twelve, and a file-global loss between the copies gave
two plus one.
run_idis an identity CLAIM, not proof of semantic equality. Two records are the same run onlywhen their identity and their persisted observation match. Equivalence is a digest of the whole
parsed record with keys sorted — a hand-picked subset would be the same mistake in a new spelling —
so key order and whitespace cannot make identical observations differ, a persisted field this reader
does not know still can, and the reader's own annotations (
_history_epoch,_evidence_quality_floor) are excluded, because reading a file must not change what a record is.A same-identity pair whose observations differ is a
RUN_ID_CONFLICT: the later record is kept,it is never reported as a retry, and it opens a recoverable scope-local boundary at its own physical
position, belonging to the epoch it closes. Complete-versus-complete counts too, and §7.3 of the
counterexample document argues that row rather than inferring it. Across a file-global loss nothing
changes: the reader cannot establish continuity there, so the later copy is a fresh identity.
One classification now drives the boundary, the deduplication, the quality floor and the
diagnostics. The separate second pass is gone — a boundary layer and a dedup layer holding two
notions of identity is exactly how they came to disagree about one pair.
history.run_id_conflictsis a bounded integer, deliberately outsidehistory.rejected(aconflicting record is accepted, not a line that failed to become one) and deliberately not a map:
six hundred distinct conflicting ids produce the integer
600and zero new keys.output_schema_versionstays 2.Five earlier cases used one
run_idfor two different observations and asserted a retry. Each isreplaced rather than quietly re-run, with the reason recorded in §7.6 before the edit, and each
keeps the property it protected — now enforced by a boundary instead of by a quality floor, which is
strictly stronger. A consequence stated rather than hidden: a true retry now has the same quality by
construction, so
QUALITY_FLOORis a guard rather than a live path.The limit this cannot close (§7.4): two physically distinct losses under one
run_idwithidentical persisted content stay indistinguishable from one identical retry. Line position is
deliberately not used as identity — that would make every true retry open a fresh boundary and
rebuild the fail-stuck behaviour the previous round removed.
IDENTICAL_REUSED_RUNID_LIMITATION=YES.New frozen rows
RID_01–RID_15; new mutantsM_RUNID_CONFLICT_TREATED_AS_RETRY,M_RUNID_CONFLICT_SECOND_LOSS_SUPPRESSED,M_EXACT_RETRY_OPENS_SECOND_BOUNDARY,M_RUNID_CONFLICT_CROSSES_EPOCH,M_RUNID_CONFLICT_CROSS_SCOPE_COLLIDES,M_RUNID_CONFLICT_GLOBAL_EPOCH_COLLIDES,M_CLEAN_RUNID_CONFLICT_IGNORED,M_PRIVATE_ANNOTATION_IN_FINGERPRINT.This is author work, not acceptance. The verdict on this head belongs to a fresh independent lane.
0.1 What the acceptance review blocked (head 1 → head 2)
The first head of this branch (
fe78cbc) passed its own four adversarial rounds and CI. Anindependent acceptance review then reproduced the six original defects on
main, confirmed all sixwere closed here, and blocked the PR on a defect this branch introduced:
Measured on that head, every one of these stayed
PARTIAL_EVIDENCEwith zero candidates:6 good + torn;+ 2 more;+ 6 more;+ 60 more; a sibling scope's records; a loss before any record; anda legitimate zero-carry record. The writer was confirmed to leave the fragment in place: after a
crash it appends a newline and its own record, and the fragment stays line 1 forever.
The replacement is a boundary, not a verdict. An unattributable loss cuts the promotion history
at that line's physical position in the file. Records before the cut never join records after
it, the newest epoch is damage-free by construction, and the loss stays reported.
damage_boundary()is the one classifier every rejection passes through. A line that parses andcarries a readable
scope_idcuts that scope; a line that cannot say whose record it was —unparseable, oversized, a file that would not open — cuts every scope, because guessing would
be the fail-open half. A foreign line, a schema-4 refusal by design and a deduplicated retry are
not losses and cut nothing.
clock: the clock is exactly what a damaged history cannot be trusted about.
active_records()is chosen before anything is anchored, so a pre-loss record cannot pick thescope, the workload class, the corpus size or the trend for the population after it.
run_idafter a loss is that population's own observation;inside one epoch the retry is still dropped, counted, and still cannot launder the survivor.
history.rejectedcounts it,history.damagesays how many boundariesthe file holds, and the human report names them — while
history.qualitydescribes only theevidence eligible for the current analysis.
No time window, no expiry, no ratio, no override flag, no new status, and
--accept-partialisunchanged: it adopts a caller's chosen bound and never a loss.
A cross-family lane then attacked the epoch model itself and found two more places where the repair
still consulted the population it had just cut — both reproduced before being believed, both fixed
here, both pinned by a case and a mutant: a ledger-only finding was still filed under the scope of
the pre-loss records (and the scope travels into the candidate id), and the epoch stamp being a
pair of counts meant two scopes could share one stamp, so deduplication treated one scope's record
as another's retry and pushed its quality onto it. The long-run simulator was also discarding the
epoch it was handed. §5.3 of the counterexample document has both cases.
This is author review, not acceptance: the verdict on this head belongs to a new independent
reviewer.
The same review's two MEDIUMs are repaired with it:
history.qualityisEMPTY(neverCOMPLETE)when nothing is comparable, and a zero-carry sweep is readable evidence, not corruption —
tools/carry.pysays so in its own report ("A session whose every item lands on its final turncarries nothing"), the 1.3 writer emits
shares: {}for it, the current writer guardsif C else {}, and runningaccumulate()on such a transcript here returnssessions=1 turns=6 carry=0. Thecontradictions stay
INVALID: carry with no shares, shares with no carry, sessions without turns,sessions beyond what was scanned.
docs/V142_COUNTEREXAMPLES.md§5 freezes thirteen B1 rows (status, exit, candidate files,comparable count, active-epoch size, quality, boundary count) written before the repair existed, §5.1
records zero deviations from them, and §5.2 lists the three earlier expectations the epoch model
supersedes — each replaced together with a companion case proving the protection it carried is still
there.
1. Why PR #7 is not merged directly
PR #7 identified real defects, and an
independent audit against
77e3677confirmed that most of them are still live. It is notmerged here for two reasons, neither of them about the quality of its analysis:
main(17 files,CONFLICTING), andmainhas since grown a typed evidence kernel (tools/evidence/*) that did not exist then, sosome of PR SameWrite 1.3.1 — fail closed on partial history evidence #7's own additions are either unnecessary here or would duplicate a stronger design.
Its invariants, counterexamples and test matrices are treated as evidence; its implementation
is not copied. PR #7 stays open until this repair is reviewed and released, and the mapping in §7
is what would justify closing it as superseded.
2. The defects, reproduced on current main first
docs/V142_COUNTEREXAMPLES.mdfreezes thirteen cases — expected status, strict exit code,candidate-file count, comparable count and evidence quality — written before the repair existed
and then run against
77e3677:The four "already correct" controls (P1, P4, S_NOACTION, S_LOCK) behaved correctly before and
still do.
Scope of the exposure, stated plainly: the current writer produces schema-4 envelope records, which
this optimizer refuses by name, so these defects bite a history file written by 1.2/1.3 (or one
authored by hand), not a fresh one. The gate was still absent.
3. The semantic repair
One law, asked of the evidence each finding actually rests on, and one gate on the side effect.
record_quality()/sweep_quality()— the worst of what evidence claims and what its ownnumbers can show, over the counters the producer has written since schema 2
(
unreadable,oversize,skipped_by_limit). Nothing is invented for an older schema: schema0/1 predates
evidence_qualityentirely, so a schema-1 record carryingCOMPLETEisUNKNOWN,and a schema-2 record without the counters cannot attest completeness either.
PARTIALandDEGRADEDare different things. A bound the caller asked for isPARTIALandis exactly what
--accept-partialadopts. Evidence that was selected and then lost — unreadable,oversize, malformed, identity-changed, conflicting — is
DEGRADED, and no flag accepts it.carry.LOSS_FIELDS/carry.BOUND_FIELDSis the single vocabulary both sides read.carry.sweep_label()no longer reportsCOMPLETEfor a sweepwhose transcript lost records to torn JSON.
eligible_anchor()chooses the comparison anchor — and the analysed scope — from recordsthat survive the eligibility filter, so the newest
INVALIDrecord can no longer decide thecorpus size for everyone else and then remove itself.
if True:. It is now aCANDIDATEon eligible historyand an
OBSERVEDfinding carrying the reason otherwise.guard on its own ledger sample. A partial history no longer blocks a live finding it never
supported, and a clean live sweep no longer waves a bounded history through.
emit_candidates(findings, outdir, status)writes nothing unless the status isCANDIDATE.The gate is inside the emitter, not at its one caller, so the CLI, the long-run simulator and any
future scheduler all route through it.
carry.bounded_paths()is the one definition of "the newest N", used by both the carry sweepand the skill-listing scan.
CANDIDATEnow precedesPARTIAL_EVIDENCEinoverall_status(). That is not a loosening: afinding only reaches
CANDIDATEafter its own evidence passed the gate, and the old order refuseda valid history-only finding because today's live sweep happened to be bounded.
4. Counterexamples and positive controls
tests/test_evidence_integrity.py— 232 assertions: one case per frozen matrix row, the law's ownunit cases, and the cases the review rounds added (R142_07 … R142_13). A fail-closed patch can look
excellent by refusing everything, so the controls are part of the gate:
COMPLETEpopulation still promotes (CANDIDATE, exit 10, one file).--accept-partial).INVALIDrecord in another scope does not poison this one.history_schema_supportedis still[0, 1, 2].5. Mutation coverage
Twenty-three new mutants, each restoring exactly one repaired defect, each RED on the mutant and
GREEN on the real source (
tests/test_mutation.py, 55/55) — the last eight are the epoch model's:None is a string mutant: every oracle observes a status, a comparable count, a file on disk, a
sweep's quality or a JSON field through a real run.
Four fixtures claimed
COMPLETEwithout the acquisition counters a real sweep always writes(multi-agent, mutation, long-run, readiness). Each now states them, and each edit is locked by an
added assertion that the same record without them refuses to promote, so the fixture change
cannot hide the gate it was making room for.
6. What did NOT change
SCHEMA4_OPTIMIZER_SUPPORT=NO—SCHEMA_SUPPORTEDis still(0, 1, 2)and a v1.4 envelope isstill refused by name, with the refusal counted.
tests/test_release_shadow_only.pypasses, no promotionor candidate-persistence entry point is added, and nothing branches on the product version.
a torn and a bounded sweep has a byte-identical sha256 before and after this branch, and the
typed reader's
loss_observedverdict is unchanged. Only the legacyqualitylabel moved, forthe torn sweep, which is the defect being repaired.
SKILL_BODY_DELTA=0,SKILL_DESCRIPTION_DELTA=0,HOOK_BEHAVIOR_DELTA=0— zero files underskills/orhooks/;CANONICAL_BODY_SHA256=7edec9f21e0bd50583e388e0bdc177f92ba8f0db6ce2bb61acd52b39d81e7767.No model-facing byte changes, so no benchmark is required.
separate gate after review, as they were for 1.4.1.
7. PR #7 property → v1.4.2 replacement
COMPLETEpromotedrecord_quality()+_derived_record_quality(); cases R142_04, U1, law unitsCOMPLETEcarry.LOSS_FIELDS+sweep_label(); case R142_05, mutantM_MALFORMED_COMPLETEeffective_evidence_quality: COMPLETEmainnever writes that field)emit_candidates(); case R142_03, mutantM_HOST_SHIFT_WRITESmainonly tests existence and derives no trust from contents — nothing reads a candidate file back)ARTIFACT_VALIDATION=DEFERRED_NOT_REQUIRED_FOR_THIS_REPAIR; the typed kernel's classifier already supersedes it for v1.4candidates_existingturned into prosemainreports bare ids)INVALIDbecomes the anchoreligible_anchor(); case R142_02, mutantM_INVALID_ANCHORcarry.bounded_paths(); case R142_06, mutantM_BOUND_SAMPLE_DIVERGES8. Regression
TOTAL_SAVINGS=NOT_PROVENandWORLD_BEST_CLAIM=NOT_TESTEDare unchanged.9. Adversarial review
Two reviewer families, asked for counterexamples rather than an opinion, on the diff itself.
Round 1 — eight defects, all in the repair, all closed here:
run_idcould launder its own sweep's quality (found by re-attacking the law)schema_versionas a string or a float was read as that schemasessions/turns/carry_bytesat all readCOMPLETE— a sharevector with no population behind it (both families found this one)
mtimesent the whole bounded selection back to a discovery-order slicesessions=40out ofscanned=1— counters that cannot describe one sweep — readCOMPLETECANDIDATEand exit 10 while every candidate write failedEach has a case in the suite and, where it is a behaviour rather than a value, a mutant.
Round 3 — four more, every one of them a word taken for evidence: a
PARTIALclaim no countercould explain was adopted by
--accept-partial; the record-level loss vocabulary was two counterswide while the contract promises more (and the dedup floor inherited the same short list); container
damage was an allowlist, so a rejection reason added later would default to "not damage"; and a
ledger line that is neither
checkednordeniedcounted as nothing at all.Round 4 — no surviving HIGH. One was claimed (a missing
isinstanceguard inload_ledger)and does not exist: the guard is there and a list-valued ledger line is counted as rejected,
verified by running it. Four MEDIUMs were real and are fixed: the live sweep now refuses the same
impossibility the record law refuses; the bounded selection is made ONCE per run and handed to both
consumers, so a transient
getmtimefailure between two calls cannot produce two samples; a linethat claims to be one of our records and carries no shares is damage rather than a foreign line;
and a candidate write that fails mid-way no longer leaves its
.tmpfile behind.Two limits are stated rather than papered over: container damage is a property of the FILE, so a
torn line degrades every scope in it — an unparseable line has no scope to attribute it to — and
HOST_BEHAVIOR_SHIFTstill closes the emitter for findings that do not rest on the shiftedpopulation, which is the release contract's ordering, not this patch's.
On the anchor fallback, one reviewer read the
eligible_anchor(recs) or newestfallback inmain()as a way for an ineligible record to choose the analysed scope. The observation reproduces — and stops there. The fallback is reached only when nothingin the file is eligible, so there is no population to strand and nothing to promote: the reviewer's
exact input gives
PARTIAL_EVIDENCE, exit 40, zero files. Both halves are now tests (R142_14),including the same record added to a real population in another scope, where the eligible records
still choose the scope and still promote.
One deliberate narrowing came out of this: the first cut of container damage degraded the history
for any rejected line, so one stray hand-written line in a shared file would have blocked a real
population forever. Damage is now a torn line, an oversized line or an unreadable file; a foreign
line that parses, and a refusal by design, are counted and reported without degrading. A corrupted
carry record that still parses as JSON is reported and does not degrade — stated as a limit
rather than implied as coverage.
10. What a reviewer should attack
The law's calibration is the judgement call, not the plumbing: a schema-1 history can no longer
promote anything, because the generation that wrote it had no field to say how the sweep went. If
you think a record with no acquisition facts should still carry a promotion, this is the line to
argue with — it is one function,
record_quality(), and every case that depends on it is indocs/V142_COUNTEREXAMPLES.md.