fix(models): Claude Fable 5.1 is selectable, and the alias help names the current generation (#2726) - #2739
Merged
Conversation
…marker to it Anthropic shipped claude-fable-5-1, the successor to Fable 5 in the same tier at the same price, and the catalog had not moved — so the model was unselectable in the operator picker, the admin fleet-default dropdown, and the per-agent public-channel override (server-validated, so a 422 rather than a quiet absence). #2086 made drift between those lists impossible; the control it left human is exactly this: one ModelEntry, one codegen run. Fable 5 keeps its slot, its id and both flags — only the (latest) marker moves, since two adjacent rows reading "Most capable — longest tasks (latest)" is the confusion the marker exists to prevent. It is NOT relabelled Legacy and NOT repositioned: AC 2 asks that it stop being labelled the latest, and dropping the marker is exactly that. No default moves — PLATFORM_DEFAULT_MODEL_VALUE and the `recommended` marker are untouched. The Workspace composer is deliberately NOT included: ent#403 excluded the whole Fable tier from that client-facing surface three days before this issue was filed, and reversing a written product decision is a type-feature, not a P1 bug. Both Workspace kwargs are omitted, so the entry matches claude-fable-5 exactly. The parity guard is derived — it goes green on whatever the catalog says — so the AC is pinned by explicit assertions instead: a per-id end-to-end test including the 422->200 gate, and a family-generic "at most one (latest) per tier" rule with the Fable id pinned on top (deliberately brittle; the next Fable must edit it). test_per_model_flags named 9 of the catalog's 10 entries, so an 11th could have been added and gone entirely unchecked while green — it now asserts that every catalog id falls in some flag group. All three assertions were meta-tested by planting violations with the mirror regenerated, so the derived halves stayed green and only the new pins could bite. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017YoUiMQgPj3BeFMpyYpfkB
… and accept the `fable` the same router advertises
GET /api/model's help text still said `opus` means Opus 4.8 and `fable` means
Fable 5 — two generations behind. It now states the durable rule ("each resolves
to the current generation of its family") before the perishable list, so the next
refresh has less to chase and a reader can see which half rots. The resolution is
the Claude Code CLI's, not Trinity's; the wording is sourced from the same
claude-api reference the catalog verifies ids against.
The same file advertised `fable` in `available_models` while the sibling PUT
/api/model 400'd it — `valid_aliases` omitted `fable` and `fable[1m]`, and the
error text listed three aliases where the GET lists four. Naming the generation
`fable` resolves to while still refusing the value would have made that
contradiction more confident, not less, so it rides the same base-image rebuild:
two array elements and one word, restoring an advertised contract rather than
granting a capability.
The integration case extends the existing valid-model list. It exercises the fix
only on an agent whose base image carries it; a stale image 400s to a backend 503,
which that test already skips on, so it cannot go falsely red.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017YoUiMQgPj3BeFMpyYpfkB
…Workspace deferral Both preset tables gain the new id and drop Fable 5's `(latest)` marker; the tier paragraph says plainly that Fable 5 is still current-generation and still served, so nobody later reads the missing marker as a retirement. The alias line and the Error Handling row can now honestly say `fable`, because the PUT accepts it as of this change. The Workspace deferral is written down rather than left silent: ent#403 excluded the Fable tier from the client-facing composer on a stated rationale, and a reader of this flow should find that decision recorded as revisited, not absent. workspace-model-choice.md's positional-dataclass warning said "all ten entries" — false at eleven, and it is the warning protecting the single most dangerous edit in this file. Its twin in model_catalog.py's docstring was fixed with the catalog change. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017YoUiMQgPj3BeFMpyYpfkB
…odel Four customer-facing files still led with Fable 5 as the flagship. Fixing the internal memory docs while the pages a customer actually reads assert the opposite would be a half-fix of a ticket whose whole thesis is "labels behind the current lineup". Each edit also says Fable 5 is still offered and still supported, so a reader who has it pinned on an agent or a schedule does not read the change as a retirement. The FAQ heading is deliberately unchanged — faq/README.md links its slug, and the question is still valid — so only the body moved. whats-new/v0.8.5.md is left alone: a release note was true when written, the same rule that keeps docs/archive/ out of scope. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017YoUiMQgPj3BeFMpyYpfkB
…able 5 1" `_format_model_name` strips an 8-DIGIT date suffix, then prefix-matches a small mapping, then falls back to title-case-with-hyphens-as-spaces. Every catalog addition before this one degraded *cleanly* through that fallback (`claude-opus-5` -> "Claude Opus 5"). `claude-fable-5-1` is the first that degrades *wrongly*: `-1` is not a date, no mapping prefix hits, and the per-model cost and token breakdowns rendered the literal "Claude Fable 5 1". #2086 FR-7 already declared this prettifier a known-deferred follow-up, and the id was reachable as free text before this change. What moves it is that making the model selectable in the picker turns a theoretical mangling into a likely one — so the exact entry ships with the thing that causes it. One mapping line. `claude-fable-5` is deliberately left unmapped: its fallback is already correct, and since the mapping is scanned with `startswith` and first match wins, an entry for it would match `claude-fable-5-1` first and silently relabel the newer model. Nothing in the function is restructured. The JS twin `stores/observability.js::formatModelName` is NOT changed, and that is a decision rather than an omission. It returns a bare family name with no version ("Claude Opus"), and `costBreakdown()` keeps no `model_id` beside it — so a one-line `includes('fable')` branch would collapse `claude-fable-5` and `claude-fable-5-1` into two indistinguishable "Claude Fable" rows on a per-model breakdown, where today both show their distinct raw ids. It would also not make the two prettifiers agree ("Claude Fable" vs "Claude Fable 5.1"); agreement needs version-aware structure on the JS side, which stays with FR-7. The test executes the function rather than restating the mapping, and its second half is derived from MODEL_CATALOG: no selectable model may render a version split across a space — the observable signature of this bug class — so the next point-release id is covered without editing the test. Both halves were watched to fail with the mapping removed, then restored. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017YoUiMQgPj3BeFMpyYpfkB
vybe
marked this pull request as ready for review
September 12, 2026 20:00
vybe
approved these changes
Sep 13, 2026
vybe
left a comment
Contributor
There was a problem hiding this comment.
merge-train: batch validated on train/20260913-0705 (#2745) — lane C, /validate-pr + /review + /cso --diff clean; catalog parity 23/23; the chat.py alias fold-in is accepted as-is. Post-merge: base image rebuild + agent recreate for the alias to land.
5 tasks done
dolho
added a commit
that referenced
this pull request
Sep 15, 2026
… 903 dev lines into the split packages The modify/delete conflicts on `routers/settings.py` and `services/git_service.py` are resolved by DELETING dev's monolith copies and re-porting every hunk dev added to them since the fork into the file that now owns it, symbol by symbol, with each function's body checked equal to dev's modulo package qualification: git_service (one dev commit, ent#615 / #2757 — the fleet-PAT fix): - `_AUTH_PATTERNS` marker -> conflicts.py - `_git_remote_url` removed, `_remote_seturl_subcommand` docstring, `_credentialless_remote_url`, `rebind_origin_and_push` (root push + credential in the exec env), `update_remote_pat` (env write, not URL) -> remotes.py - the credential-helper install + embedded-token sweep block (`write_container_github_pat`, both alarms, `scrub_git_remote_tokens`, the fleet sweep, `spawn_git_remote_token_scrub`, all `_SCRUB_*`) -> NEW token_scrub.py (remotes.py would otherwise sit at 821 lines, over the threshold the split exists for) - `_agent_can_push`, `_agent_has_write_credentials` docstring, `sync_to_github`, `reset_to_main_preserve_state` -> sync.py - `initialize_git_in_container` (seeds before writing a remote) -> provisioning.py Package `__init__` re-exports every new name; the duplicate `REBIND_PUSH_TIMEOUT_S` the hunk would have introduced is dropped. settings (five dev commits — #2715, #2619, #2707, #2741, #2739): - 11 changed routes replaced in place across flags/credentials/ integrations/generic - 11 new symbols placed beside their dev-order predecessors; the #2715 Resend/Gemini routes + their two helpers go to NEW provider_keys.py (credentials.py would otherwise reach 1,045 lines), included on the package router right after `credentials` and before `generic` - `_ANTHROPIC_KEY_ALIASES` / `_adopt_after_instance_key_removed` reached from generic.py through the sibling module object, per the package rule ops: `_format_model_name`'s #2739 `claude-fable-5-1` entry lands in `ops_costs_service.py`, where the split moved the function; the #2726 test imports from there. Dev's tests that patch monolith attributes are re-pointed the way the split re-pointed every earlier one: the ent#615 exec recorder is installed on each execing sibling and `_detect_git_dir` on `gitignore`; #2572's `db` fake on `credentials` and `generic`; ent#553's source read on `flags`; #1677's emitter allowlist and the ent#615 source reads on `token_scrub`. `_PRE_SPLIT_ROUTES`' post-split allowlist records the six #2715 routes; the git_service import-surface pin drops `_git_remote_url` (gone by design) for its ent#615 replacements. Content conflicts: `backend.md` (dev's facts under the package names), `test_ent123_tokenless_clone.py` (dev's helper patch, on `gs.sync`). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VpvcfgWkmQPD7DrDLmATTf
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Anthropic shipped
claude-fable-5-1and Trinity's catalog never moved, so the model was unselectable everywhere the catalog feeds — absent from the operator picker and the admin fleet-default dropdown, and 422'd by the server-validated public-channel override. This is not drift (#2086 made drift impossible by deriving every list from one source); it is source-vs-reality staleness, the one control #2086 deliberately left human. The fix is oneModelEntryat index 1 plus one codegen run — no consumer was edited — with the(latest)marker moved offclaude-fable-5, the agent-server alias help text restated soopusstops being documented as Opus 4.8, thefablealias accepted by the same router that advertises it, a one-line ops-dashboard label mapping, and fourdocs/user-docs/sentences that were telling customers Fable 5 is the most capable model.Fixes #2726
What changed
src/backend/services/model_catalog.py— oneModelEntryforclaude-fable-5-1/Claude Fable 5.1/ noteMost capable — longest tasks (latest)at index 1, with its three booleans passed positionally (True, True, False) per the dataclass's own\u2014escape its neighbours use.claude-fable-5keeps its id, its slot and both policy flags, and loses only the(latest)marker — noLegacyrelabel, no reposition. Adds aLast synced: 2026-09-12bump anchor (themodel_context.py:40-44idiom) so catalog staleness becomes provable by inspection.src/frontend/src/constants/modelCatalog.js— regenerated viapython3 scripts/gen_model_catalog.py. Never hand-edited; re-running the script on the pushed tree leavesgit statusempty.docker/base-image/agent_server/routers/chat.py—GET /api/model's help text now states the durable rule before the perishable list: "each resolves to the current generation of its family — today sonnet (Sonnet 5), opus (Opus 5), haiku (Haiku 4.5), fable (Fable 5.1)".PUT /api/modelacceptsfable/fable[1m], and its 400 detail names the same four aliases theGETadvertises.src/backend/routers/ops.py— one"claude-fable-5-1": "Claude Fable 5.1"mapping entry. The prettifier'sre.sub(r'-\d{8}$','')strips only an 8-digit suffix, so the point-release id kept its-1and fell through to the title-case fallback as the literal "Claude Fable 5 1".tests/unit/test_2086_model_catalog_parity.py12→14 tests; newtests/unit/test_2726_ops_model_label.py(9 tests);tests/test_agent_chat.pyexercises thefablealias.feature-flows/model-selection.md(both preset tables, the alias sentence, the 400-detail row, changelog),feature-flows.mdindex row, one unconditional correction inworkspace-model-choice.md, and fourdocs/user-docs/files.No consumer was edited —
ModelSelector.vue,Settings.vue,settings_service.py,client_portal/service.pyand the MCP tools all derive from the catalog. Editing one would be the #2086 regression.Acceptance criteria
AC 1 —
claude-fable-5-1selectable end-to-end: operator picker (PRESET_MODELSprojectsMODEL_CATALOG), admin fleet-default dropdown (adminDefaultSelectablefilter), public-channel override (is_valid_public_channel_model("claude-fable-5-1")is now true, soPUT /api/agents/{name}/public-channel-modelanswers 200 instead of 422 — pinned bytest_fable_5_1_is_present_and_selectable_end_to_end).AC 1 lists four broken surfaces; three are ticked. The Workspace composer is deliberately NOT included. ent#403 excluded the whole Fable tier from that client-facing surface days before this issue was filed, on a written rationale recorded in
docs/memory/feature-flows/workspace-model-choice.md,requirements/core-agent.mdAC-1 andrequirements/public-access.md§47.2 FR-2 — Opus-vs-Fable is an operator distinction, not a client one. That omission was never catalog staleness (workspace=Falseis set on purpose), so a refresh would never have changed it; reversing it is a product decision — atype-featurefor the private tracker, not this P1 bug. The issue owner chose to keep the Workspace excluded. Bothworkspacekwargs are omitted, the entry is flag-identical toclaude-fable-5, andWORKSPACE_MODELSis unchanged.AC 2 — Fable 5.1 carries
(latest); Fable 5 no longer does. Pinned bytest_at_most_one_latest_marker_per_family_and_fable_is_5_1, which asserts the rule (at most one(latest)per family, any family) plus the two AC-specific ids.AC 3 — the alias help text names the generation each alias resolves to; proven against the built image (below), not just the source.
AC 4 — the id is the exact undated
claude-fable-5-1. The dated legacy ids are deliberately left alone (see Notes).AC 5 — generated mirror in sync:
test_generated_js_is_byte_freshpasses, and re-running the codegen on the pushed tree produces no diff.AC 6 —
docs/memory/feature-flows/model-selection.mdpreset tables + changelog updated, plus the index row.AC 7 — no default changes.
PLATFORM_DEFAULT_MODEL_VALUEand therecommendedmarker are untouched; the new entry isrecommended=False, asserted both at import time and bytest_recommended_is_exactly_one_and_is_the_platform_default.Folded-in adjacent fix
PUT /api/modelrejectedfablewith a 400 whileGET /api/modelon the same router advertised it inavailable_models— a pre-existing advertise-vs-accept asymmetry outside every AC. It is folded in here rather than deferred: same file, same handler pair, riding the same base-image rebuild; AC 3 makes the contradiction worse by naming a generation for an alias the sibling endpoint 400s; and accepting it restores an advertised contract rather than granting a new capability. It is still a behaviour change outside the ACs — a maintainer may reasonably ask for it as a separate PR, in which case dropping thevalid_aliasesand 400-detail hunks fromchat.pyleaves the rest of this PR intact.Deployment note
The agent-server changes (the alias help text and
fableacceptance) reach an agent only after./scripts/deploy/build-base-image.shand a container recreate. Existing agents keep serving the old note and keep 400ingfableuntil then. Zero functional impact in the meantime — model selection itself is unaffected, and nothing else in this PR depends on the image. The backend and frontend halves ship with the normal image build: the public-channel PUT accepts the new id the moment the backend restarts. No migration, no data backfill, no restart ordering; a storedpublic_channel_modelofclaude-fable-5is unaffected because that id does not move and keeps both flags.Verification
/verify-local --skip-agenton this tree — status pass: unit 15,349 passed / 31 skipped / 0 failed; build + import-smoke OK; boot + health OK; integration 70 passed / 13 skipped / 2 registry-deselected / 0 failed (isolated projecttrinity-verify-d410c1b4).Base-image proof. The agent-server change can only be proven in the image, and the global-mode agent stage refuses while the operator's dev stack is live, so it was proven directly. Built
trinity-agent-base:issue-2726-verifyfrom this HEAD with--build-arg VERSION="$(cat VERSION)"(14 s, cache hit — the arg is load-bearing:Dockerfile:8feedsENV TRINITY_BASE_VERSIONat line 18, above apt/Go/Node/Claude Code, so omitting it invalidates every layer below).trinity-agent-base:latestid was captured before and after and is unchanged; the throwaway image was removed; no container joinedtrinity-agent-networkandagent-trinity-systemwas never touched.Three extracted literals, verbatim from the run:
Claude model aliases (Anthropic API): each resolves to the current generation of its family — today sonnet (Sonnet 5), opus (Opus 5), haiku (Haiku 4.5), fable (Fable 5.1). Add [1m] suffix for the 1M extended-context beta (e.g. sonnet[1m]).— and noOpus 4.8.['sonnet', 'opus', 'haiku', 'fable', 'sonnet[1m]', 'opus[1m]', 'haiku[1m]', 'fable[1m]'](sonnet, opus, haiku, fable).id -uinside the image = 1000 (Invariant #17).Negative controls — the proof is non-vacuous. The same script against
origin/dev'schat.pyexits 1 at arm (a). Two host mutants exit 1 at arms (b) and (c) respectively (new note + oldvalid_aliases; newvalid_aliases+ old 400 detail). All three arms bite independently — without (b) and (c) the folded-infablefix would have ridden on agrepthat matches the old file too.Unit-level, re-run after rebase:
test_2086_model_catalog_parity.py+test_ent243_prompt_tier.pytest_ent403_workspace_model.py+test_894_public_channel_model.pytest_2726_ops_model_label.py(new)npx vitest run tests/unit/portalModelChoice.spec.jsnpm run test:unitpython3 scripts/gen_model_catalog.py && git status --porcelainMeta-tests — each new guard was watched to fail before being trusted. The parity guard is derived (both halves come from the catalog), so it ratifies whatever the catalog says; a green first run is not evidence. Planting each violation turned the matching guard red: flipping the new entry's
public_channeltoFalse; restoring(latest)on Fable 5; dropping the id from the flag groups (the coverage assertion); and removing the ops mapping (3 failed).Notes
Legacy (prior Fable)relabel and no reposition; only the marker drops. AC 2 asks that Fable 5 stop being labelled the latest, and dropping the marker is exactly that;Legacywould imply a deprecation Trinity has no policy behind (the reference says Fable 5 is still served) and would cost a visible picker reorder for a claim no AC makes.stores/observability.js:281-291deliberately still returns raw ids for the Fable family: it has nofablebranch at all, and adding one would be a second, hand-synced copy of the same mapping — the exact shape refactor: centralize the selectable model catalog — one source of truth for three hand-synced lists #2086 exists to prevent. The Python mapping is the one the cost/token dashboards read.claude-fable-5in the operator picker now matches both rows (ModelSelector.vue:120uses.includes), new one first — harmless, there is no auto-select (:163guardshighlightedIndex >= 0). (2)Settings.vue:360renders{{ m.label }} — {{ m.note }}, so the new row shows a nested em-dash — identical in shape to what Fable 5 renders today, so not a regression; re-wording an operator string no AC touches would be scope creep.claude-haiku-4-5-20251001,claude-opus-4-5-20251101,claude-sonnet-4-5-20250929) were deliberately not renormalised to the reference's undated form — doing so would invalidate every stored agent setting pinned to them, for zero AC gain. Considered, not overlooked..claude/agents/test-runner.mdcatalog rows — the submodule is read-only in this run, so the two rows are listed here instead:tests/unit/test_2086_model_catalog_parity.py12→14 tests, and a new entry fortests/unit/test_2726_ops_model_label.py(9 tests).Follow-ups (not filed — for the maintainer)
.describe()examples still citeclaude-opus-4-8—chat.ts:438,655,schedules.ts:176,loops.ts:166. Free text, no enum, nothing breaks; Update and actualize the available LLM model list across the app (add Opus 4.8, retire EOL models) #1080 refreshed these, refactor: centralize the selectable model catalog — one source of truth for three hand-synced lists #2086 then redefined the refresh contract as "one file, one script".model-selection.md, a table and an entry count inworkspace-model-choice.md, tier strings inrequirements/public-access.mdFR-2, counts inrequirements/core-agent.mdandruntimes.md, and fourdocs/user-docs/files. Every refresh must hand-chase all of them, and it has already failed once (item 5). Wants a catalog canary diffingMODEL_CATALOGagainst the published lineup, or — far cheaper — a scheduled monthly "re-verify the catalog" issue. TheLast synced:anchor added here makes staleness provable; it does not prevent it.model_catalog.py's "That is the whole maintenance surface — one file, one script" is true of the catalog and false of a model refresh — this PR is 14 files. Worth amending so the next person's estimate is right.docs/memory/requirements/runtimes.md:314says "7 of 9PRESET_MODELSentries" — already stale before this PR (10 since refactor: centralize the selectable model catalog — one source of truth for three hand-synced lists #2086), 11 now. Out of this PR's blast radius.docker/base-image/agent_server/services/headless_executor.py:1616carries a three-alias docstring that omitsfable.routers/chat.py:677-686), so an operator who sends an invalid model gets "Failed to set model" rather than the 400 detail this PR just improved.🤖 Generated with Claude Code
https://claude.ai/code/session_017YoUiMQgPj3BeFMpyYpfkB