Skip to content

Model aliases and pinned models in routing rules, v0.62.0 - #290

Merged
fylorn merged 31 commits into
mainfrom
feat/model-aliases
Oct 5, 2026
Merged

fylorn merged 31 commits into
mainfrom
feat/model-aliases

Conversation

@fylorn

@fylorn fylorn commented Oct 5, 2026

Copy link
Copy Markdown
Contributor

Closes #283. Includes the v0.62.0 version bump and its release notes, so this PR is the release commit.

Model aliases

  • A top-level aliases table: an alias is another name for the same model, with the names upstreams use for it, in order. An upstream serves the alias when it offers one of the names within models_only, and receives the first it offers; every such upstream takes part in failover with its own name.
  • GET /v1/models lists aliases; an alias takes precedence over an unlisted real model of the same name; listing and admission stay one function.
  • Inheritance runs one way: a key's allow and a rule's when.model written for an upstream model name also cover the aliases that list it; written for an alias, they cover only the alias.
  • set.model may name an alias; phase-two renames are sent as written. Replay, the L3 speed test and plugin renames resolve per upstream.
  • Control plane: /aliases CRUD, POST /alias-preview, GET /aliases/{name}/usage, rename that updates key and rule references; Claude-only suggestions for the same model under different names across upstreams.

Pinned models

to in a routing rule may be a list of {provider, model}, tried in order and sent as written (no aliases, no set.model). The key's allow is checked per candidate, so failover never reaches a model the key may not use.

Answers carry the client's model name

When the sent name differs from the client's and the upstream answered with the model it was sent, the answer (all four formats, streamed or not, and openai-model/x-openai-model) carries the client's name. A different model in the answer is passed through. Recording and the check-up see the upstream's own answer.

Also

  • /v1/models carries context windows (context_window, context_length, max_input_tokens, supports_1m, Gemini inputTokenLimit/outputTokenLimit); Claude tiers include fable and mythos.
  • WebSocket parity: Responses frames resolve names, check allow and get the rules' set/deny per frame; Realtime routes by the query's model.
  • allow is now also enforced when no upstream reports a model list.
  • Control-plane paths where a fixed segment shadowed {name} are moved (/provider-test, /provider-preview, /proxy-test, /plugin-*, /alias-preview), with a test covering every resource; plugin ids are no longer reserved.
  • Numeric-looking names in to and aliases load as strings.
  • Protocol 37 → 38; request store schema unchanged; configuration format only gains fields.

Built in parallel lanes (config/engine skeleton, catalog/admission, routing, answer names, listing metadata, control plane, replay), then reviewed as a whole; the review's findings (pinned failover vs allow, phase-two admission, alias rename refs, preview path, WebSocket) are fixed here. Locally: cargo fmt --all --check, cargo clippy --workspace --all-targets -D warnings, cargo test --workspace, scripts/release_notes_test.py.

🤖 Generated with Claude Code

fylorn and others added 30 commits October 5, 2026 09:43
Each model in GET /v1/models and GET /v1/models/{id} now carries what
clients read about it, taken from the price table (the same number the
control plane shows as ModelRow.context_window):

- OpenAI shape: context_window, context_length and max_input_tokens
- Anthropic shape: the same, plus supports_1m (true at 1,000,000 or more)
- Gemini shape: inputTokenLimit and outputTokenLimit

When the price table has no entry for the model, the fields are left out.
model_meta(book, provider, model) returns this metadata, so an alias can
take it from the model that serves it.

anthropic_family_tier also recognises fable and mythos.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…nd resolve

- tw-config: top-level `aliases` (`Aliases` / `Alias { name, models }`), a map
  whose values are a model name or a list of them, kept in the order written.
  Validation with `config.alias_*` codes: empty name, `*`/`?`, leading `__`,
  duplicate, empty list, empty model name, listing another alias, listing only
  itself. `check_aliases` is public for the alias preview.
- tw-engine: `Rule.to` is now `Option<Target>`, `Target::Name(String)` or
  `Target::Models(Vec<Pinned>)` (a string or a list of `{provider, model}`).
  A pinned list makes the listed upstreams the candidates in list order, with
  no group ordering; `Decision.pinned` carries each candidate's model and
  `models_asked` uses it. Validation with `engine.pinned_*` codes.
- tw-gateway: `models::resolve(cfg, catalog, provider, name)`, the name sent
  to an upstream for what the client asked (first listed alias model the
  upstream can serve; non-aliases unchanged). A hop sends the pinned model
  as written.
- Provider renames follow pinned entries; references count them.
- Control API types are unchanged for now: a pinned target shows as no
  target in `RuleView.to`.
- Manual tables regenerated; new section `routes[].rules[].to[]`.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
When the model name sent upstream (S) differs from the one the client asked
for (N) -- a rule's set.model, a plugin renaming it, and soon aliases and
pinned models -- the model the upstream writes in its answer (A) is now
written back as N, provided A is the same model as S
(model_name::same, dated snapshots included). A different model stays
visible, so a real substitution still shows (ChatGPT rerouting, Codex's
openai-model warning). A missing or empty model is not filled in, and
upstream errors are passed through untouched.

- tw-gateway answer_model: a streaming rewriter that counts JSON levels like
  the usage sniffer and holds back only the model string itself, so streams
  are not buffered. It rewrites exactly the fields the sniffer reads: root
  model / modelVersion / model_version, message.model, response.model --
  SSE frames, Gemini JSON arrays, whole bodies and WebSocket messages.
- relay: runs after format conversion and before reply hooks (plugin ctx
  keeps model = S, requested_model = N; frames plugins synthesize copy the
  renamed envelope); the recorder and the upstream check-up still see A
  because ending.feed runs on the raw bytes first. openai-model and
  x-openai-model response headers follow the same rule. Failover uses the
  name the answering hop sent.
- WebSocket: a model a plugin renamed in response.create is written back.
- model_name moves from tw-store to tw-pricing (re-exported as
  tw_store::model_name) so the gateway can use it without depending on the
  store.
- tests/models.rs: the fake upstream echoes the received model under `sent`,
  since `model` in an answer is now rewritten.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…alias's served name

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… aliases

- RuleView.to / RuleInput.to are a RuleTarget: a string (upstream, group or
  __all__) or a list of PinnedModel, written exactly like `to` in the config
  (TypeScript: string | Array<PinnedModel>). Both directions map in
  tw-control, so the overview gives pinned targets back and saving a route,
  or a dry-run draft, keeps them; an empty list is left to the engine, which
  refuses it with engine.pinned_empty.
- GET /models lists every alias (KnownModel.alias: its models; providers:
  listed, enabled upstreams where resolve finds a name), once even when a
  real model shares its name; real models carry the aliases that list them.
- ModelRow.aliases: the aliases that list a model on an upstream's list.
- CONTROL_API_VERSION 37 -> 38, with the notes for the whole alias feature.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… them

- tw-engine `Catalog`: `with_aliases` registers each alias under the listed
  upstreams that have any of its models (the first in list order is what goes
  out). Names come in two spaces: client names (`all`, `providers_for`,
  listing, admission) include aliases, and an alias takes its name from a real
  model it does not list; sent names (`offers`, `count_for`) stay each
  upstream's own list, so a pinned model or a phase-two rename is judged as
  written. Upstreams without a list register nothing, as before. New
  `alias`, `served`, `first_served`.
- `resolve_allowed` / `admits` stay one function and inherit from real names
  to aliases: an `allow` pattern that matches any model of an alias lets the
  alias through; naming the alias never exposes its models.
- tw-gateway: the published catalog carries the config's alias table.
  `models::sent_to` gives the name a candidate is sent (pinned as written,
  otherwise `resolve`) or why it cannot serve; `serving` takes the decision
  and uses it, so an alias skips upstreams that serve none of its models and
  a pinned candidate is judged on its pinned model. `fit` is the as-written
  check it was.
- `/v1/models` and `/v1/models/{id}` list aliases, with the metadata of the
  first model of the alias some upstream serves, looked up for that upstream.
- Admission: pinned candidates are judged at their own upstream; errors say
  which case it is, with new codes `gw.model.alias_unserved`,
  `gw.model.alias_unserved_rewritten`, `gw.model.pinned_not_offered`,
  `gw.model.pinned_not_allowed`.
- The pinned-model test reads `sent`, which the fake upstream answers since
  the response-name change.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…stream its own name

Replay used to resend the stored body unchanged, so a request for an alias,
or one a rule had rewritten, went out under the client's name. It now sends
the chosen upstream the name it took the first time when the request went
there under a different name (the attempt chain's recorded model), and
otherwise the alias resolved for that upstream. The model in the body, or
in the Gemini path, is the only thing changed, and the quote prices that
name. An upstream that offers none of the alias's models is refused with
control.replay_alias_unserved instead of being sent a name it cannot serve.

The inference speed test resolves the model per upstream before fitting,
pricing and sending; an alias an upstream cannot serve is skipped as
out_of_scope or not_offered, as a real name would be.

A plugin's params.model is a client-side name and may be an alias: it is
resolved per hop. The key's allow check stays on the name the plugin wrote.
An upstream that cannot serve the alias is skipped (gw.plugin.alias_unserved
on the attempt chain) and the next one is tried, so plug::attempt now tells
a whole-request refusal (Stop::Request) from a skipped hop (Stop::Hop).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…nned failover, dry-run models

- tw-engine: the engine carries the alias table (`Engine::with_aliases`,
  `alias_models`; `Config::engine()` passes it). `When::matches` /
  `matches_with_provider` take the request alias's models, so a `when.model`
  glob on an upstream name also matches the aliases that list it; a condition
  written on an alias matches only that alias. Both phases inherit.
- `Engine::asked` / `asked_of` give each candidate's asked name with its origin
  (`Origin::Client | Rule | PhaseTwo | Pinned`); `models_asked` is the same list
  without origins. Client and phase-one names are client-side (may be an alias);
  pinned and phase-two names go out as written. `Outcome2::Proceed` now says
  which model phase two set.
- tw-gateway `sent`: the name each upstream receives (`resolve` for client-side
  names, as written otherwise) and why it differs from the client's name.
  Routing skips candidates by the sent name (an upstream that serves none of an
  alias's models is skipped), `cheapest` prices and the context-window check use
  the sent name, and each hop re-resolves after failover; `Served.model` and the
  attempt record are the sent name. A hop whose upstream no longer serves the
  alias fails with `gw.model.alias_not_served` instead of sending the alias.
- Pinned targets fail over in list order with no group ordering; session
  affinity applies as for any candidate list (group `None`).
- Dry run: `candidate_models` (one per candidate: `sent_model`, `model_via` =
  alias | rule | pinned), rule trace and mismatch with inheritance, cheapest by
  sent name.
- Tests: engine inheritance/origins/pinned/phase-two; gateway end-to-end alias
  failover, set.model to an alias, pinned failover in list order, phase-two
  rename as written, inheritance; dry-run fields and cheapest by sent name.
  C0's pinned e2e test reads the echoed `sent` (C3 changed the fake upstream).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…eview, usage, rename

- tw-api: AliasInput/AliasSave/AliasModel/AliasView/AliasSuggestion/AliasesView/
  AliasPreviewRequest/AliasPreview/AliasRuleRef/AliasUsage/AliasWritten, and six
  endpoints: GET/POST /aliases, POST /aliases/preview, PUT/DELETE /aliases/{name},
  GET /aliases/{name}/usage.
- tw-config: upsert_alias/remove_alias edit the aliases table in place (a new
  alias goes to the end, a rename keeps its place, one model is written as a
  string, several as a list; an inline or empty table is rewritten as a block);
  refs::alias_refs/rename_alias find and rewrite the whole-value references in
  key allow lists and in when.model / set.model (not phase-two set.model, which
  is sent as written).
- tw-yaml: rename_key swaps a mapping key's bytes in place, with a self-check;
  keys added by put are rendered as scalars, so a user-chosen name that needs
  quotes ("a: b", "yes", "1.5") reads back as itself.
- tw-control: served_by via resolve per enabled upstream, shadows and per-model
  providers from each upstream's own listing, context window from the default
  price sheet, 24-hour requests and cost from the per-model summary; Claude-only
  same-model suggestions on an exact key (path prefixes, Bedrock geo/anthropic./
  -vN[:M], Vertex @Date, dots in versions; dates kept); a preview reports the
  same problems as config validation, with `original` so an edit is not a
  duplicate of itself.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sent names now come from models::sent_to (pinned as written, resolve, fit);
sent::name adds only the phase-two case, judged as written. Route skipping
goes through models::serving with phase-two names passed as pinned for that
upstream. A hop that finds its upstream can no longer serve the name reports
the skip with the existing explanation, so gw.model.alias_not_served is gone.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…odel rows of the manual

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…names

- Move the alias preview to `POST /alias-preview`: `POST /aliases/preview`
  shadowed `PUT`/`DELETE /aliases/{name}`, so an alias named `preview` could
  be neither updated nor deleted (405). A test keeps any fixed path out from
  under `/aliases/`.
- Renaming an alias whose old name stays in its own model list leaves the
  keys' `allow` entries and the rules' `when.model` alone: the old name now
  means that upstream model, and inheritance already carries it to the
  renamed alias. `set.model` is still rewritten, since it is a name to send
  and does not inherit. `rename_alias` takes the renamed alias and returns
  what it rewrote, which is what `renamed_in` reports.
- `served_by` in `GET /aliases` and the preview follow admission: when some
  upstream has a model list and no listed upstream offers any of the alias's
  models, the alias is refused before any candidate is tried, so it is served
  by nobody (and every model is unserved in the preview). With no lists at
  all the answer stays resolve-based.
- A rule's `to` and an alias's single model accept unquoted numbers and
  booleans as names (`to: 2024`, `x: 1.5`), as 0.61.0 did for `to`, instead
  of failing to load and sending the gateway into safe mode.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…wn model name

The Responses WebSocket path sent every frame's model as written: a client
alias, or a plugin renaming to an alias, went out as the alias itself, and a
rule that pinned models for the client routed the upgrade to the pinned
upstream but still sent the client's model.

Each response.create now follows the same naming as an HTTP hop: the pinned
model when the deciding rule pinned one for the connected upstream, a
phase-two rename as written, otherwise the client-side name (the frame's,
after a phase-one set.model, or after a plugin rename) resolved through the
alias table for that upstream. Phase two runs per frame with the frame's
facts; a phase-two deny refuses that frame.

When the upstream serves none of the requested alias's models, that frame
is not sent: the client gets a response.failed (gw.ws.alias_unserved,
gw.ws.alias_unserved_rewritten, or gw.plugin.alias_unserved for a plugin
rename) and the connection stays open. The answer's model name is written
back to the client's name for every case, and the attempt records the
model fixed at the upgrade (a pinned model or a phase-one rewrite).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…es as written

- Failover no longer reaches a candidate whose model the key may not use.
  Admission lets a request through when any candidate's model is allowed;
  each candidate is now checked again when routing and on every hop, and
  one the key's allow does not cover is skipped as `not_allowed`. Pinned
  models and phase-two renames are matched against `allow` as written;
  names on the client side (including aliases) keep the inheritance from
  upstream model names.
- Admission takes the origin-tagged candidates (`Engine::asked`), so a
  phase-two rename is judged like a pinned model: by the upstream's own
  list and the name itself, not through the alias table. A rename that
  happens to share an alias's name is admitted when that upstream lists it.
- The key's allow is enforced at admission even while no upstream has
  listed its models, with the usual "may not use" message.
- A plugin that renames the model to an alias gets the same allow
  inheritance as a request for that alias (`Catalog::allows` is public).
- `ServeSkip::NotAllowed` (`not_allowed`): dry runs given a key list the
  candidates skipped for it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
axum prefers a fixed path segment to a {param} one, so a fixed endpoint
under /providers/, /proxies/ or /plugins/ shadowed every resource of that
name: an upstream named `test` or `preview`, or a proxy named `test`, got
405 on PUT and DELETE, and the plugin ids `order`, `inspect`, `rewrite`
and `confirmed` had to be reserved (PUT /plugins/order reached the
reorder handler).

Moved, endpoint names unchanged:
  POST /providers/test      -> POST /provider-test
  POST /providers/preview   -> POST /provider-preview
  POST /proxies/test        -> POST /proxy-test
  POST /plugins/inspect     -> POST /plugin-inspect
  POST /plugins/rewrite     -> POST /plugin-rewrite
  POST /plugins/confirmed   -> POST /plugin-confirmed
  PUT  /plugins/order       -> PUT  /plugin-order
  POST /aliases/preview     -> POST /alias-preview (same edit as fix/ma-control)

- ep.rs: a test that no endpoint has a fixed segment where another
  endpoint with the same prefix takes a parameter, whatever the method.
  {guard} is a closed set; a fixed word there must not be a guard.
- Plugin ids are no longer reserved: config.plugin.reserved_id and
  control.plugin.reserved_id are removed, and a plugin named "Order"
  gets the id `order`.
- Handler tests: upstreams named `test` and `preview`, a proxy named
  `test`, and plugins with the ids `order`, `inspect`, `rewrite` and
  `confirmed` can be edited and deleted.
- Protocol 38 notes, 0.62.0 release notes and the configuration
  reference follow.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…d resolve Realtime's model

Brings the WebSocket paths to parity with HTTP for model access and rule
parameters.

- Each response.create on a Responses connection is checked against the
  key's allow the way HTTP checks a request: names on the client side (the
  frame's, a phase-one rewrite, a plugin rename) inherit from upstream model
  names listed by an alias (Catalog::allows); a pinned model or a phase-two
  rename is matched as written. A refused frame gets a response.failed with
  the HTTP messages (gw.model.not_allowed, not_allowed_rewritten,
  pinned_not_allowed, gw.plugin.model_not_allowed) and the connection stays
  open.
- The routing target stays the one decided at the upgrade, but the rest of
  the rules is evaluated per frame with that frame's facts, as for an HTTP
  request: phase one's accumulated set (and a deny that matches the frame's
  model), then phase two. set.max_tokens and set.thinking are written into
  the frame with the same code HTTP uses for a passthrough Responses body
  (forward::apply_set); max_tokens is left out for a Codex backend upstream,
  which HTTP strips too.
- A Realtime connection (/v1/realtime) takes its model from the query
  string: the upgrade is routed with it, refused when the key may not use it
  (no upstream that serves it and is allowed: 400 with the HTTP message),
  and the query sent upstream carries the connected upstream's own name
  (pinned as written, alias resolved). The attempt records the sent name and
  the request row the client's.
- Catalog::allows is public (same change as fix/ma-admission).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

# Conflicts:
#	release-notes/0.62.0.md
…eason

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@fylorn
fylorn merged commit 51ef803 into main Oct 5, 2026
5 checks passed
@fylorn
fylorn deleted the feat/model-aliases branch October 5, 2026 05:02
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Model aliases: a client-visible name for an upstream model

1 participant