Skip to content

fix(models): Claude Fable 5.1 is selectable, and the alias help names the current generation (#2726) - #2739

Merged
vybe merged 5 commits into
devfrom
vybe/issue-2726
Sep 13, 2026
Merged

vybe merged 5 commits into
devfrom
vybe/issue-2726

Conversation

@trinity-ability

Copy link
Copy Markdown
Contributor

Anthropic shipped claude-fable-5-1 and Trinity's catalog never moved, so the model was unselectable everywhere the catalog feeds — absent from the operator picker and the admin fleet-default dropdown, and 422'd by the server-validated public-channel override. This is not drift (#2086 made drift impossible by deriving every list from one source); it is source-vs-reality staleness, the one control #2086 deliberately left human. The fix is one ModelEntry at index 1 plus one codegen run — no consumer was edited — with the (latest) marker moved off claude-fable-5, the agent-server alias help text restated so opus stops being documented as Opus 4.8, the fable alias accepted by the same router that advertises it, a one-line ops-dashboard label mapping, and four docs/user-docs/ sentences that were telling customers Fable 5 is the most capable model.

Fixes #2726

What changed

  • src/backend/services/model_catalog.py — one ModelEntry for claude-fable-5-1 / Claude Fable 5.1 / note Most capable — longest tasks (latest) at index 1, with its three booleans passed positionally (True, True, False) per the dataclass's own ⚠️, and the em-dash written as the \u2014 escape its neighbours use. claude-fable-5 keeps its id, its slot and both policy flags, and loses only the (latest) marker — no Legacy relabel, no reposition. Adds a Last synced: 2026-09-12 bump anchor (the model_context.py:40-44 idiom) so catalog staleness becomes provable by inspection.
  • src/frontend/src/constants/modelCatalog.js — regenerated via python3 scripts/gen_model_catalog.py. Never hand-edited; re-running the script on the pushed tree leaves git status empty.
  • docker/base-image/agent_server/routers/chat.py — GET /api/model's help text now states the durable rule before the perishable list: "each resolves to the current generation of its family — today sonnet (Sonnet 5), opus (Opus 5), haiku (Haiku 4.5), fable (Fable 5.1)". PUT /api/model accepts fable / fable[1m], and its 400 detail names the same four aliases the GET advertises.
  • src/backend/routers/ops.py — one "claude-fable-5-1": "Claude Fable 5.1" mapping entry. The prettifier's re.sub(r'-\d{8}$','') strips only an 8-digit suffix, so the point-release id kept its -1 and fell through to the title-case fallback as the literal "Claude Fable 5 1".
  • Tests — tests/unit/test_2086_model_catalog_parity.py 12→14 tests; new tests/unit/test_2726_ops_model_label.py (9 tests); tests/test_agent_chat.py exercises the fable alias.
  • Docs — feature-flows/model-selection.md (both preset tables, the alias sentence, the 400-detail row, changelog), feature-flows.md index row, one unconditional correction in workspace-model-choice.md, and four docs/user-docs/ files.

No consumer was edited — ModelSelector.vue, Settings.vue, settings_service.py, client_portal/service.py and the MCP tools all derive from the catalog. Editing one would be the #2086 regression.

Acceptance criteria

  • AC 1 — claude-fable-5-1 selectable end-to-end: operator picker (PRESET_MODELS projects MODEL_CATALOG), admin fleet-default dropdown (adminDefaultSelectable filter), public-channel override (is_valid_public_channel_model("claude-fable-5-1") is now true, so PUT /api/agents/{name}/public-channel-model answers 200 instead of 422 — pinned by test_fable_5_1_is_present_and_selectable_end_to_end).

    AC 1 lists four broken surfaces; three are ticked. The Workspace composer is deliberately NOT included. ent#403 excluded the whole Fable tier from that client-facing surface days before this issue was filed, on a written rationale recorded in docs/memory/feature-flows/workspace-model-choice.md, requirements/core-agent.md AC-1 and requirements/public-access.md §47.2 FR-2 — Opus-vs-Fable is an operator distinction, not a client one. That omission was never catalog staleness (workspace=False is set on purpose), so a refresh would never have changed it; reversing it is a product decision — a type-feature for the private tracker, not this P1 bug. The issue owner chose to keep the Workspace excluded. Both workspace kwargs are omitted, the entry is flag-identical to claude-fable-5, and WORKSPACE_MODELS is unchanged.

  • AC 2 — Fable 5.1 carries (latest); Fable 5 no longer does. Pinned by test_at_most_one_latest_marker_per_family_and_fable_is_5_1, which asserts the rule (at most one (latest) per family, any family) plus the two AC-specific ids.

  • AC 3 — the alias help text names the generation each alias resolves to; proven against the built image (below), not just the source.

  • AC 4 — the id is the exact undated claude-fable-5-1. The dated legacy ids are deliberately left alone (see Notes).

  • AC 5 — generated mirror in sync: test_generated_js_is_byte_fresh passes, and re-running the codegen on the pushed tree produces no diff.

  • AC 6 — docs/memory/feature-flows/model-selection.md preset tables + changelog updated, plus the index row.

  • AC 7 — no default changes. PLATFORM_DEFAULT_MODEL_VALUE and the recommended marker are untouched; the new entry is recommended=False, asserted both at import time and by test_recommended_is_exactly_one_and_is_the_platform_default.

Folded-in adjacent fix

PUT /api/model rejected fable with a 400 while GET /api/model on the same router advertised it in available_models — a pre-existing advertise-vs-accept asymmetry outside every AC. It is folded in here rather than deferred: same file, same handler pair, riding the same base-image rebuild; AC 3 makes the contradiction worse by naming a generation for an alias the sibling endpoint 400s; and accepting it restores an advertised contract rather than granting a new capability. It is still a behaviour change outside the ACs — a maintainer may reasonably ask for it as a separate PR, in which case dropping the valid_aliases and 400-detail hunks from chat.py leaves the rest of this PR intact.

Deployment note

The agent-server changes (the alias help text and fable acceptance) reach an agent only after ./scripts/deploy/build-base-image.sh and a container recreate. Existing agents keep serving the old note and keep 400ing fable until then. Zero functional impact in the meantime — model selection itself is unaffected, and nothing else in this PR depends on the image. The backend and frontend halves ship with the normal image build: the public-channel PUT accepts the new id the moment the backend restarts. No migration, no data backfill, no restart ordering; a stored public_channel_model of claude-fable-5 is unaffected because that id does not move and keeps both flags.

Verification

/verify-local --skip-agent on this tree — status pass: unit 15,349 passed / 31 skipped / 0 failed; build + import-smoke OK; boot + health OK; integration 70 passed / 13 skipped / 2 registry-deselected / 0 failed (isolated project trinity-verify-d410c1b4).

Base-image proof. The agent-server change can only be proven in the image, and the global-mode agent stage refuses while the operator's dev stack is live, so it was proven directly. Built trinity-agent-base:issue-2726-verify from this HEAD with --build-arg VERSION="$(cat VERSION)" (14 s, cache hit — the arg is load-bearing: Dockerfile:8 feeds ENV TRINITY_BASE_VERSION at line 18, above apt/Go/Node/Claude Code, so omitting it invalidates every layer below). trinity-agent-base:latest id was captured before and after and is unchanged; the throwaway image was removed; no container joined trinity-agent-network and agent-trinity-system was never touched.

A defect in the plan's own recipe, stated honestly. As first written, the §7 check ran docker run --rm --network none "$TAG" python - <<'PYCHK' — without -i. docker run does not forward stdin without it, so the heredoc program was never read: python got an empty program and exited 0 on any image. Measured: the verbatim form produced no output at all. The evidence below is from the corrected form, and the negative control proves it bites.

docker run --rm -i --network none "$TAG" python - <<'PYCHK'
# (a) the alias-help literal, (b) the valid_aliases list, (c) the 400 detail —
# extracted from the file's AST inside the container, hard-asserted, no `||`.
PYCHK

Three extracted literals, verbatim from the run:

  • (a) Claude model aliases (Anthropic API): each resolves to the current generation of its family — today sonnet (Sonnet 5), opus (Opus 5), haiku (Haiku 4.5), fable (Fable 5.1). Add [1m] suffix for the 1M extended-context beta (e.g. sonnet[1m]). — and no Opus 4.8.
  • (b) ['sonnet', 'opus', 'haiku', 'fable', 'sonnet[1m]', 'opus[1m]', 'haiku[1m]', 'fable[1m]']
  • (c) the 400 detail lists (sonnet, opus, haiku, fable).

id -u inside the image = 1000 (Invariant #17).

Negative controls — the proof is non-vacuous. The same script against origin/dev's chat.py exits 1 at arm (a). Two host mutants exit 1 at arms (b) and (c) respectively (new note + old valid_aliases; new valid_aliases + old 400 detail). All three arms bite independently — without (b) and (c) the folded-in fable fix would have ridden on a grep that matches the old file too.

Unit-level, re-run after rebase:

Command Result
test_2086_model_catalog_parity.py + test_ent243_prompt_tier.py 52 passed (was 49)
test_ent403_workspace_model.py + test_894_public_channel_model.py 56 passed (unchanged — the Workspace is untouched)
test_2726_ops_model_label.py (new) 9 passed
npx vitest run tests/unit/portalModelChoice.spec.js 54 passed
full frontend npm run test:unit 125 files / 2,732 passed
python3 scripts/gen_model_catalog.py && git status --porcelain empty — mirror byte-fresh

Meta-tests — each new guard was watched to fail before being trusted. The parity guard is derived (both halves come from the catalog), so it ratifies whatever the catalog says; a green first run is not evidence. Planting each violation turned the matching guard red: flipping the new entry's public_channel to False; restoring (latest) on Fable 5; dropping the id from the flag groups (the coverage assertion); and removing the ops mapping (3 failed).

Notes

  • Deviation from the dossier's D4. Fable 5 keeps its slot and its id — no Legacy (prior Fable) relabel and no reposition; only the marker drops. AC 2 asks that Fable 5 stop being labelled the latest, and dropping the marker is exactly that; Legacy would imply a deprecation Trinity has no policy behind (the reference says Fable 5 is still served) and would cost a visible picker reorder for a claim no AC makes.
  • The ops dashboard now reads "Claude Fable 5.1", but the JS twin stores/observability.js:281-291 deliberately still returns raw ids for the Fable family: it has no fable branch at all, and adding one would be a second, hand-synced copy of the same mapping — the exact shape refactor: centralize the selectable model catalog — one source of truth for three hand-synced lists #2086 exists to prevent. The Python mapping is the one the cost/token dashboards read.
  • Two cosmetics accepted, not fixed. (1) Typing the exact string claude-fable-5 in the operator picker now matches both rows (ModelSelector.vue:120 uses .includes), new one first — harmless, there is no auto-select (:163 guards highlightedIndex >= 0). (2) Settings.vue:360 renders {{ m.label }} — {{ m.note }}, so the new row shows a nested em-dash — identical in shape to what Fable 5 renders today, so not a regression; re-wording an operator string no AC touches would be scope creep.
  • The dated legacy ids (claude-haiku-4-5-20251001, claude-opus-4-5-20251101, claude-sonnet-4-5-20250929) were deliberately not renormalised to the reference's undated form — doing so would invalidate every stored agent setting pinned to them, for zero AC gain. Considered, not overlooked.
  • Owed .claude/agents/test-runner.md catalog rows — the submodule is read-only in this run, so the two rows are listed here instead: tests/unit/test_2086_model_catalog_parity.py 12→14 tests, and a new entry for tests/unit/test_2726_ops_model_label.py (9 tests).

Follow-ups (not filed — for the maintainer)

  1. MCP .describe() examples still cite claude-opus-4-8 — chat.ts:438,655, schedules.ts:176, loops.ts:166. Free text, no enum, nothing breaks; Update and actualize the available LLM model list across the app (add Opus 4.8, retire EOL models) #1080 refreshed these, refactor: centralize the selectable model catalog — one source of truth for three hand-synced lists #2086 then redefined the refresh contract as "one file, one script".
  2. Dated legacy ids vs the reference's undated ids — deliberately not renormalised here (stored settings), but worth a decision.
  3. The stale-prose class — this is its third instance (Update and actualize the available LLM model list across the app (add Opus 4.8, retire EOL models) #1080, bug: Fable 5 / Sonnet 5 missing from platform default-model dropdown and PUBLIC_CHANNEL_MODELS whitelist (drifted from ModelSelector presets) #1660, bug: selectable model catalog is stale — Claude Fable 5.1 (claude-fable-5-1) not selectable anywhere; tier labels + alias help text behind the current lineup #2726). refactor: centralize the selectable model catalog — one source of truth for three hand-synced lists #2086's contract is "one file, one script", but prose that enumerates or counts the catalog sits outside it: two tables in model-selection.md, a table and an entry count in workspace-model-choice.md, tier strings in requirements/public-access.md FR-2, counts in requirements/core-agent.md and runtimes.md, and four docs/user-docs/ files. Every refresh must hand-chase all of them, and it has already failed once (item 5). Wants a catalog canary diffing MODEL_CATALOG against the published lineup, or — far cheaper — a scheduled monthly "re-verify the catalog" issue. The Last synced: anchor added here makes staleness provable; it does not prevent it.
  4. model_catalog.py's "That is the whole maintenance surface — one file, one script" is true of the catalog and false of a model refresh — this PR is 14 files. Worth amending so the next person's estimate is right.
  5. docs/memory/requirements/runtimes.md:314 says "7 of 9 PRESET_MODELS entries" — already stale before this PR (10 since refactor: centralize the selectable model catalog — one source of truth for three hand-synced lists #2086), 11 now. Out of this PR's blast radius.
  6. docker/base-image/agent_server/services/headless_executor.py:1616 carries a three-alias docstring that omits fable.
  7. The backend PUT proxy collapses the agent-server 400 into a generic 503 (routers/chat.py:677-686), so an operator who sends an invalid model gets "Failed to set model" rather than the 400 detail this PR just improved.
  8. An offline AST guard for the agent-server help text — it has no parity-tested twin (it is not an Invariant Setup improvements #5 vendored mirror), so today only a built image can catch it drifting.

🤖 Generated with Claude Code

https://claude.ai/code/session_017YoUiMQgPj3BeFMpyYpfkB

trinity-ability and others added 5 commits September 12, 2026 13:16
…marker to it

Anthropic shipped claude-fable-5-1, the successor to Fable 5 in the same tier at
the same price, and the catalog had not moved — so the model was unselectable in
the operator picker, the admin fleet-default dropdown, and the per-agent
public-channel override (server-validated, so a 422 rather than a quiet absence).
#2086 made drift between those lists impossible; the control it left human is
exactly this: one ModelEntry, one codegen run.

Fable 5 keeps its slot, its id and both flags — only the (latest) marker moves,
since two adjacent rows reading "Most capable — longest tasks (latest)" is the
confusion the marker exists to prevent. It is NOT relabelled Legacy and NOT
repositioned: AC 2 asks that it stop being labelled the latest, and dropping the
marker is exactly that. No default moves — PLATFORM_DEFAULT_MODEL_VALUE and the
`recommended` marker are untouched.

The Workspace composer is deliberately NOT included: ent#403 excluded the whole
Fable tier from that client-facing surface three days before this issue was
filed, and reversing a written product decision is a type-feature, not a P1 bug.
Both Workspace kwargs are omitted, so the entry matches claude-fable-5 exactly.

The parity guard is derived — it goes green on whatever the catalog says — so the
AC is pinned by explicit assertions instead: a per-id end-to-end test including
the 422->200 gate, and a family-generic "at most one (latest) per tier" rule with
the Fable id pinned on top (deliberately brittle; the next Fable must edit it).
test_per_model_flags named 9 of the catalog's 10 entries, so an 11th could have
been added and gone entirely unchecked while green — it now asserts that every
catalog id falls in some flag group.

All three assertions were meta-tested by planting violations with the mirror
regenerated, so the derived halves stayed green and only the new pins could bite.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017YoUiMQgPj3BeFMpyYpfkB
… and accept the `fable` the same router advertises

GET /api/model's help text still said `opus` means Opus 4.8 and `fable` means
Fable 5 — two generations behind. It now states the durable rule ("each resolves
to the current generation of its family") before the perishable list, so the next
refresh has less to chase and a reader can see which half rots. The resolution is
the Claude Code CLI's, not Trinity's; the wording is sourced from the same
claude-api reference the catalog verifies ids against.

The same file advertised `fable` in `available_models` while the sibling PUT
/api/model 400'd it — `valid_aliases` omitted `fable` and `fable[1m]`, and the
error text listed three aliases where the GET lists four. Naming the generation
`fable` resolves to while still refusing the value would have made that
contradiction more confident, not less, so it rides the same base-image rebuild:
two array elements and one word, restoring an advertised contract rather than
granting a capability.

The integration case extends the existing valid-model list. It exercises the fix
only on an agent whose base image carries it; a stale image 400s to a backend 503,
which that test already skips on, so it cannot go falsely red.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017YoUiMQgPj3BeFMpyYpfkB
…Workspace deferral

Both preset tables gain the new id and drop Fable 5's `(latest)` marker; the
tier paragraph says plainly that Fable 5 is still current-generation and still
served, so nobody later reads the missing marker as a retirement. The alias line
and the Error Handling row can now honestly say `fable`, because the PUT accepts
it as of this change.

The Workspace deferral is written down rather than left silent: ent#403 excluded
the Fable tier from the client-facing composer on a stated rationale, and a
reader of this flow should find that decision recorded as revisited, not absent.

workspace-model-choice.md's positional-dataclass warning said "all ten entries"
— false at eleven, and it is the warning protecting the single most dangerous
edit in this file. Its twin in model_catalog.py's docstring was fixed with the
catalog change.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017YoUiMQgPj3BeFMpyYpfkB
…odel

Four customer-facing files still led with Fable 5 as the flagship. Fixing the
internal memory docs while the pages a customer actually reads assert the
opposite would be a half-fix of a ticket whose whole thesis is "labels behind
the current lineup".

Each edit also says Fable 5 is still offered and still supported, so a reader who
has it pinned on an agent or a schedule does not read the change as a retirement.

The FAQ heading is deliberately unchanged — faq/README.md links its slug, and the
question is still valid — so only the body moved. whats-new/v0.8.5.md is left
alone: a release note was true when written, the same rule that keeps
docs/archive/ out of scope.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017YoUiMQgPj3BeFMpyYpfkB
…able 5 1"

`_format_model_name` strips an 8-DIGIT date suffix, then prefix-matches a small
mapping, then falls back to title-case-with-hyphens-as-spaces. Every catalog
addition before this one degraded *cleanly* through that fallback
(`claude-opus-5` -> "Claude Opus 5"). `claude-fable-5-1` is the first that
degrades *wrongly*: `-1` is not a date, no mapping prefix hits, and the per-model
cost and token breakdowns rendered the literal "Claude Fable 5 1".

#2086 FR-7 already declared this prettifier a known-deferred follow-up, and the
id was reachable as free text before this change. What moves it is that making
the model selectable in the picker turns a theoretical mangling into a likely
one — so the exact entry ships with the thing that causes it.

One mapping line. `claude-fable-5` is deliberately left unmapped: its fallback is
already correct, and since the mapping is scanned with `startswith` and first
match wins, an entry for it would match `claude-fable-5-1` first and silently
relabel the newer model. Nothing in the function is restructured.

The JS twin `stores/observability.js::formatModelName` is NOT changed, and that
is a decision rather than an omission. It returns a bare family name with no
version ("Claude Opus"), and `costBreakdown()` keeps no `model_id` beside it — so
a one-line `includes('fable')` branch would collapse `claude-fable-5` and
`claude-fable-5-1` into two indistinguishable "Claude Fable" rows on a per-model
breakdown, where today both show their distinct raw ids. It would also not make
the two prettifiers agree ("Claude Fable" vs "Claude Fable 5.1"); agreement needs
version-aware structure on the JS side, which stays with FR-7.

The test executes the function rather than restating the mapping, and its second
half is derived from MODEL_CATALOG: no selectable model may render a version
split across a space — the observable signature of this bug class — so the next
point-release id is covered without editing the test. Both halves were watched to
fail with the mapping removed, then restored.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017YoUiMQgPj3BeFMpyYpfkB
@vybe
vybe marked this pull request as ready for review September 12, 2026 20:00

@vybe vybe left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

merge-train: batch validated on train/20260913-0705 (#2745) — lane C, /validate-pr + /review + /cso --diff clean; catalog parity 23/23; the chat.py alias fold-in is accepted as-is. Post-merge: base image rebuild + agent recreate for the alias to land.

@vybe
vybe merged commit 8d15edb into dev Sep 13, 2026
28 checks passed
dolho added a commit that referenced this pull request Sep 15, 2026
… 903 dev lines into the split packages

The modify/delete conflicts on `routers/settings.py` and
`services/git_service.py` are resolved by DELETING dev's monolith copies and
re-porting every hunk dev added to them since the fork into the file that
now owns it, symbol by symbol, with each function's body checked equal to
dev's modulo package qualification:

git_service (one dev commit, ent#615 / #2757 — the fleet-PAT fix):
  - `_AUTH_PATTERNS` marker            -> conflicts.py
  - `_git_remote_url` removed, `_remote_seturl_subcommand` docstring,
    `_credentialless_remote_url`, `rebind_origin_and_push` (root push +
    credential in the exec env), `update_remote_pat` (env write, not URL)
                                        -> remotes.py
  - the credential-helper install + embedded-token sweep block
    (`write_container_github_pat`, both alarms, `scrub_git_remote_tokens`,
    the fleet sweep, `spawn_git_remote_token_scrub`, all `_SCRUB_*`)
                                        -> NEW token_scrub.py (remotes.py
    would otherwise sit at 821 lines, over the threshold the split exists for)
  - `_agent_can_push`, `_agent_has_write_credentials` docstring,
    `sync_to_github`, `reset_to_main_preserve_state`   -> sync.py
  - `initialize_git_in_container` (seeds before writing a remote)
                                        -> provisioning.py
  Package `__init__` re-exports every new name; the duplicate
  `REBIND_PUSH_TIMEOUT_S` the hunk would have introduced is dropped.

settings (five dev commits — #2715, #2619, #2707, #2741, #2739):
  - 11 changed routes replaced in place across flags/credentials/
    integrations/generic
  - 11 new symbols placed beside their dev-order predecessors; the #2715
    Resend/Gemini routes + their two helpers go to NEW provider_keys.py
    (credentials.py would otherwise reach 1,045 lines), included on the
    package router right after `credentials` and before `generic`
  - `_ANTHROPIC_KEY_ALIASES` / `_adopt_after_instance_key_removed` reached
    from generic.py through the sibling module object, per the package rule

ops: `_format_model_name`'s #2739 `claude-fable-5-1` entry lands in
`ops_costs_service.py`, where the split moved the function; the #2726
test imports from there.

Dev's tests that patch monolith attributes are re-pointed the way the
split re-pointed every earlier one: the ent#615 exec recorder is installed
on each execing sibling and `_detect_git_dir` on `gitignore`; #2572's `db`
fake on `credentials` and `generic`; ent#553's source read on `flags`;
#1677's emitter allowlist and the ent#615 source reads on `token_scrub`.
`_PRE_SPLIT_ROUTES`' post-split allowlist records the six #2715 routes; the
git_service import-surface pin drops `_git_remote_url` (gone by design) for
its ent#615 replacements.

Content conflicts: `backend.md` (dev's facts under the package names),
`test_ent123_tokenless_clone.py` (dev's helper patch, on `gs.sync`).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VpvcfgWkmQPD7DrDLmATTf
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants