Skip to content

Load-balance weights and distribution by speed and reliability, slow-start switch, upstream concurrency, model specs, key usage limits, per-turn WebSocket requests - #292

Merged
fylorn merged 22 commits into
mainfrom
integ/routing-2026-10
Oct 5, 2026
Merged

fylorn merged 22 commits into
mainfrom
integ/routing-2026-10

Conversation

@fylorn

@fylorn fylorn commented Oct 5, 2026

Copy link
Copy Markdown
Contributor

A routing and limits round: weights on load-balance groups and distribution by speed and reliability, moving a slow-starting stream to the next upstream, a concurrency cap per upstream, model specs set by hand, usage limits per gateway key, and Responses WebSocket turns counted as requests. Control-plane protocol 38 → 39. Answers #289.

Load-balance groups

  • A member may carry a weight: providers: [{ name: relay, weight: 7 }, official] (1–100; a plain name is 1 and is written back as a plain name, so existing configurations round-trip unchanged). Weights only on load-balance (engine.group_weight_not_load_balance, engine.group_weight_out_of_range).
  • Distribution is smooth weighted round-robin (7 : 3 interleaves). Its state lives in the gateway and reaches the pure engine through Facts, so the dry-run still predicts the data plane exactly. A request that stays on an upstream for its conversation is charged to it, so the weight is the long-run share of requests. Paused and full members sit out a round; the member that leads after stickiness is always charged.
  • groups[].balance_by: weights | latency | health | latency-health multiplies the weights by a factor: latency = (median of the members' time to first content / this member's)², clamped to [0.1, 10]; health = success rate², floored at 0.05. A member without enough samples counts as average (1.0).
  • WebSocket upgrades now use the same ordering as HTTP (pipeline::arrange).

Speed and success samples

  • A latency sample is the time from sending a hop to the first content of its answer (streamed answers only, any position in the candidate list). Before, it ran from the request's arrival and ended at a point that depended on the hop's position. url-test keeps using the startup link test until an upstream has 3 real samples; load-balance factors use real samples only.
  • A success tracker per upstream (last 50 outcomes within 30 minutes, at least 5) uses the breaker's fault classification. Busy skips, slow-start switches and client cancellations record nothing.

Slow stream start — failover.next_on_slow_start (default off): a streamed answer with no content after stream_start_wait_secs is cancelled and the next upstream is asked, if one could take the request at that moment; otherwise the current one keeps streaming. The last upstream always waits. The slow upstream is not paused. The abandoned attempt is recorded with outcome slow_start and its usage (AttemptUsage), marked as possibly billed. config.slow_start_too_short: with switching on, the wait must be at least 5 s. No keepalive while holding: response headers are not sent yet, and Google's Python SDK rejects SSE comment lines.

Concurrency — providers[].max_concurrent (1–1000) and failover.slot_wait_secs (default 30, one wait budget per request, shared with the key's rolling limits). A conversation that stays on a full upstream waits for a slot, then moves on; other candidates are skipped at once (ServeSkip::Busy). When every candidate is full the request waits for the first free slot; if none frees and no hop reached an upstream, 429 gw.busy_all (Retry-After); otherwise the last real failure. The per-key clients[].max_concurrent permit is now held until the answer has been relayed (it was released when the response started, so streamed bodies were not counted).

Model specs — providers[].model_specs: { <model>: { context_window?, max_output_tokens? } } win over the price table field by field, wherever core reports or uses a model's limits (the listing in every shape, ModelRow, aliases, the re-route when input outgrows the context window, the default max_tokens on conversion to Anthropic). PUT /provider-model-spec (ModelSpecSave → ConfigWritten) sets or clears one entry. A real model in /v1/models is now described by the first upstream offering it.

Usage limits per gateway key — clients[].limits: [{ per: minute|hour|day|week|month, requests|tokens|cost: N, cache_reads?: bool }], all must pass. Minute and hour roll; day, week (from Monday) and month follow the machine's local time. Tokens are uncached input + cache writes + output, plus cache reads when cache_reads: true. Cost is in USD (at least 0.01); unpriced models and billing: free upstreams count 0. In-flight requests reserve their estimate and settle from the recorded usage; a request that never reached an upstream counts for nothing. Totals are rebuilt from the request store at startup and when a period's start moves (time-zone change). Calendar limits answer 429 with Retry-After to the reset and x-should-retry: false (insufficient_quota in OpenAI bodies); rolling limits wait within the request's wait budget, then 429 with the exact Retry-After. Event::KeyLimitAlert at 80% and on reaching a limit. Validation: config.key_limit_*, including config.key_limit_retention (retention must cover the period after a restart).

Responses WebSocket — each response.create is its own request: started/routed/first token/ending events, a request-store row with the turn's usage and cost, per-turn key limits, key permit and upstream slot, latency and success samples; a refused or busy turn gets response.failed and the connection stays usable. An upstream that closes mid-turn fails the turn (gw.ws.upstream_closed). Realtime stays one row per connection, now with the usage of its response.done events, counted against key limits when it closes.

Upgrade notes for the release

  • Protocol 38 → 39 (the CONTROL_API_VERSION doc comment lists the changes). A UI written for 38 drops weights and limits when it saves groups and keys.
  • Request store schema unchanged.
  • The configuration only gains fields; every 0.62.x configuration loads, except one that lists an upstream twice in a group or a group member that is not an upstream (Configuration validation refuses a group member listed twice or that is not an upstream #291).
  • url-test ordering changes, since samples now measure time to first content per hop.
  • A key's max_concurrent now also covers streamed bodies, so a key at its limit queues more than before.
  • Responses WebSocket connections show one row per turn instead of one per connection.

Message codes — new: config.key_limit_cache_reads, config.key_limit_cost_too_small, config.key_limit_duplicate, config.key_limit_empty, config.key_limit_not_positive, config.key_limit_retention, config.key_limit_two_measures, config.model_spec_blank_model, config.model_spec_empty, config.model_spec_wildcard, config.model_spec_zero, config.provider_concurrency_range, config.slow_start_too_short, control.group.balance_not_load_balance, control.group.weight_not_load_balance, control.group.weight_not_member, control.group.weight_out_of_range, engine.group_balance_not_load_balance, engine.group_weight_not_load_balance, engine.group_weight_out_of_range, gw.busy_all, gw.busy_upstream, gw.key_limit.{requests,tokens,cost}_{per_period,rolling}, gw.slow_start, gw.ws.upstream_closed. None removed.

Built in parallel lanes and integrated locally; an adversarial review before this PR found 10 defects (slow-start switch dropping a working upstream for a full one, a busy 429 hiding a real error, link-test seeds used as speeds, refusals consuming request limits, Realtime outside token/cost limits, two validation gaps, a time-zone change resetting totals, an upstream close counted as a client cancel, an uncharged sticky leader, WebSocket upgrades ignoring load-balance), each fixed with a test that failed before the fix.

Checked locally: cargo fmt, cargo clippy --workspace --all-targets -D warnings, 3013 workspace tests.

🤖 Generated with Claude Code

fylorn and others added 22 commits October 5, 2026 16:05
A load-balance group could only take turns one by one, so an upstream
with twice the capacity (or half the price) got the same share as the
rest. A member can now be written as `{ name, weight }` (1-100); a plain
name is weight 1 and is written back as a plain name, so existing
configurations round-trip byte for byte, through the control plane's
writer too. Weights on other group types, or out of range, are refused
at load time (engine.group_weight_*) and when saving a group
(control.group.weight_*).

Distribution is nginx's smooth weighted round-robin, which interleaves
7:3 instead of sending seven in a row. It replaces the `seq`-based
rotation (and with it EventBus::peek_id, its only user). The per-group
state lives in the gateway (tw_gateway::balance) and reaches the engine
through Facts, so ordering stays a pure function and the dry-run reads
the same state without advancing it.

The state is charged at decision time, under one lock from ordering to
charging, so a burst of simultaneous requests spreads out. It charges
the member that leads after conversation stickiness: a conversation
that stays on A counts toward A, and new conversations make up the
difference, so the long-run ratio holds. Members that cannot serve the
request or are cooling down sit the round out and the rest share by
weight; otherwise their turn would fall to whoever follows them in the
group. A group whose members or weights change starts over.

The overview gives every member's weight for load-balance groups
(GroupView.weights), saving takes them (GroupInput.weights), and the
dry-run shows each candidate's weight (DryRunCandidate.weight).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A load-balance group spreads new conversations by its members' weights only,
so a member that is slow or keeps failing gets as many new conversations as
a fast, healthy one until the pauses catch it. groups[].balance_by (weights,
latency, health, latency-health; default weights, omitted when default)
multiplies each member's weight by a factor.

The engine stays pure: balance_factors(by, members, facts) computes the
factor per member from Facts. Latency uses the median TTFB that url-test
already uses ((median / own)^2, clamped to 0.1..10); health uses a new
Facts.success ((rate)^2, floored at 0.05 so a flaky upstream still gets an
occasional request and its recovery is seen). A member without enough
samples counts as 1.0, so a new upstream is tried but not flooded.

The smooth weighted round-robin turns on effective weights: weight x factor
x 1000, rounded, at least 1 (weighted::effective), computed in the one place
the schedule reads weights (weighted::round), for picking the leader and for
charging it alike. The x1000 keeps fractional factors meaningful with integer
current weights; equal factors scale every member alike, so the interleaving
is exactly what the weights alone give. Factors are taken over all candidates
before the paused ones sit the round out, which is also what the dry-run
shows.

The success rate comes from Health itself: every record_* call also lands
in a per-upstream window (last 50 outcomes within 30 minutes, at least 5 to
report), so it counts exactly what the breaker counts and needs no new
classification at the call sites. Client errors count as answered, a
missing model and client cancellations are not counted.

The data plane and the dry-run build their facts in one function
(AppState::group_facts, now taking the group), which fills TTFB and success
rates only for groups that use them; the dry-run reads the same snapshot and
the same round-robin state without advancing it, and shows each candidate's
weight, TTFB, success rate and factor. An end-to-end test checks that the
dry-run names the upstream the next new conversation takes, request after
request, while the factors move.

balance_by on any other group type is refused at load time
(engine.group_balance_not_load_balance) and by the control plane
(control.group.balance_not_load_balance).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Some upstreams accept a request and then send nothing for a long time: a
relay queueing it, an overloaded provider that does not say so, response
headers that do not come. The client sees a hung request while another
candidate could already be answering.

failover.next_on_slow_start (off by default) gives up on such an upstream
when a streamed answer has no content stream_start_wait_secs after the
request was sent, covering both the wait for response headers and the
held start of the stream. What counts as content is what the stream-start
hold already judges (text, thinking/reasoning, tool calls; not role or
start frames). Giving up drops the pending request or the response, so
the connection closes and the upstream stops generating.

- The last candidate never switches. "Last" accounts for the candidates
  that could not take the request anyway (paused by then, or the hop could
  not be sent: phase-two denial, no servable name, conversion failure,
  forced server tool), so the effectively last one waits instead of
  trading a slow answer for a certain failure.
- The slow upstream is neither paused nor counted as a failure, nor as an
  outcome in the success rate load-balance groups weigh by.
- The abandoned try is an attempt with outcome slow_start, with the usage
  the upstream reported at the start (Anthropic message_start) or else the
  gateway's input estimate marked as estimated (AttemptView.usage).
- Only requests the client asked to stream; a non-streamed request is
  unaffected, and with the setting off the hold behaves as before.
- Validation refuses switching with a wait under 5 seconds
  (config.slow_start_too_short); FailoverView.next_on_slow_start shows it.

No keepalive is sent while holding: response headers are only sent once an
upstream is chosen, and committing a 200 early would turn every later 429,
5xx and the last upstream's own 4xx into in-stream errors that clients do
not retry on; Google's Python SDK also fails on SSE comment lines.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…t win over the price table

A relay's own models are often missing from the price table, and the table is
sometimes wrong, so clients were told no context window (or a wrong one) and
the gateway could not tell whether a conversation still fits its model.

`providers[].model_specs` maps an exact model id to `context_window` and/or
`max_output_tokens`. One function in tw-config (`model_specs::resolve`, reached
through `Provider::model_limits` / `Config::model_limits`) decides the
precedence, field by field, and every reader goes through it: the `/v1/models`
listing in all three shapes (a real model is now described by the first
upstream offering it, so its spec applies; aliases as before), the upstream
model rows, alias context windows, the pipeline's "input outgrew the held
decision's context window" check, and the output limit filled in when a
request is converted to Anthropic.

`ModelRow` gains `context_window_source`, `max_output_tokens` and
`max_output_tokens_source` so the UI can show which number is manual. `PUT
/provider-model-spec` sets or (both null) removes one entry, editing only that
upstream's `model_specs` so the rest of the entry stays byte for byte; the
provider dialog keeps existing specs, and they follow renames and deletes
because they live inside the entry. Validation: `config.model_spec_*`.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Some relays and accounts accept only a few requests at a time and refuse
the rest, so a busy upstream turned ordinary load into failed requests.
`providers[].max_concurrent` (1-1000) makes the gateway respect that limit
before a request is sent, instead of learning it from a refusal.

A request takes a slot when its hop is about to be sent (before plugins and
conversion run) and gives it back when the answer has been relayed to its
end or the client is gone. The engine's order is unchanged:

- the upstream that conversation stickiness kept (its cache, or the same
  turn) is waited for, since moving loses the cache;
- any other full upstream is skipped at once (attempt skipped: busy);
- when every remaining candidate is full, the request waits for the first
  to free, in candidate order, then answers 429 with Retry-After in the
  client's API shape (gw.busy_all, recorded as rate_limited).

All waiting shares one budget per request, `failover.slot_wait_secs`
(default 30, 0 = never wait), because the client receives nothing while
it waits. Waiting is not a failure: no pause, no breaker count, and no
outcome in the success rate load-balance groups weigh by. Token
counting and WebSocket connections take no slot. Limits live in gateway
state across reloads and are resized in place; a removed limit releases
its waiters.

With failover.next_on_slow_start, a full upstream still counts as one that
can take the request when deciding whether the slow one is the last: it may
free up within the slot wait, so the slow upstream is given up on and the
request may then wait for the full one (or get the busy 429). Giving up on a
slow upstream returns its slot at once.

Traffic shows `AttemptView.queued_ms` and `skipped: busy`; the control
plane exposes `ProviderView/ProviderInput.max_concurrent` and
`FailoverView.slot_wait_secs`.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The permit for `clients[].max_concurrent` was a local in pipeline(), so it
was returned the moment the response headers went out. A key limited to
one request could run any number of streams at once; the limit covered
only the wait for headers. The permit now moves into the response body
stream next to the in-flight counter and the upstream slot, and is
returned when the answer ends or the client goes away.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…our, day, week or month

A key on the LAN, or one handed to a script, can now be capped by what it
spends rather than only by how many requests run at once. Limits live on the
key (`clients[].limits`), each entry counts one measure, and every entry has
to pass.

- Day, week and month follow the calendar in the core machine's local time
  zone. Used up means refused until the next period: 429 with Retry-After to
  the reset and x-should-retry: false, and code insufficient_quota in the
  OpenAI shapes, so the official SDKs and Codex stop instead of retrying.
- Minute and hour are rolling. Used up means waiting for the next free slot
  when it frees within failover.slot_wait_secs (the same setting as the wait
  for an upstream's free slot, counted separately), otherwise 429 with exact
  Retry-After and retry-after-ms. The 429 for upstreams that are all at their
  max_concurrent now sets its Retry-After through the same field, so it also
  carries retry-after-ms and x-should-retry: true.
- Admission is one step after routing: calendar limits, then the existing
  max_concurrent gate (its permit still travels into the response body, so
  it is held until the answer ends), then rolling limits. A refused request is recorded like
  a routing refusal (start, empty attempt chain, failure with its own code).
- Running requests hold their input estimate (and its input cost on the first
  candidate); the recorder settles each row to the usage and cost it writes,
  through a direct hook rather than the bus, so the in-memory totals equal
  what a restart adds back from the store. Token counts and local answers do
  not count; unpriced models and free upstreams count $0.
- A restart rebuilds this day, week and month from the request records, so a
  monthly limit requires retention.row_days >= 31.
- The key view carries each limit with used/max/resets_at/reached, and the
  models without a price for keys that have a cost limit; the key input takes
  the limits; renaming a key through the control plane carries its usage.
- KeyLimitAlert on the event bus once per period at 80% and at the limit, for
  Lite's system notifications.
- WebSocket connections are admitted once when they open: the store records
  one row per connection without usage.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…budget

Latency samples (url-test ordering and the load-balance latency factor)
were taken when relaying began and counted from the request's arrival, so
they included key-limit and slot waits, plugins and every earlier failed
or abandoned hop, and ended at a point that depended on the hop's position
(first content behind the opening watch, response headers for the last
candidate, nearly the whole answer when not streamed). An upstream dropped
by the slow-start switch recorded nothing and kept its old fast samples,
while the next upstream was charged the wait.

A sample is now the time from sending that hop to the first content of a
streamed answer, detected by the same first-token reader that reports
RequestFirstToken, wherever the hop sits in the candidate list. Answers
that are not streamed give no sample. An upstream given up on for a slow
start records the time it was given, a lower bound that ranks it slow.
L1 handshake seeds stand in only until an upstream has enough real
samples and are then dropped, never mixed into the median.

Load-balance: a member at its max_concurrent at decision time sits the
round out like a paused one. Charging it for a request the slot cap then
skipped made a busy member drift below its share.

failover.slot_wait_secs is now one deadline per request, set when the key
gate lets it in and shared by the rolling key-limit wait and the upstream
slot wait; before, each could use the full time, so one request could
wait twice as long.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A Responses WebSocket connection used to be one request row for its whole
life, with no usage and no cost: ws.rs never read `usage`, the key's usage
limits were checked once when the connection opened, and neither the key's
`max_concurrent` nor the upstream's slot applied at all. A client configured
for WebSocket could therefore spend without limit and without a trace of
what it cost.

Now each `response.create` is a request, from the frame to the answer's
`response.completed` / `response.failed` / `response.incomplete` (or an
`error`): RequestStarted, RequestHeaders, RequestRouted (one hop, the
connection's upstream, with queued_ms), first token and an ending that
carries the turn's usage, so the recorder prices it like any HTTP request
and key usage, checkup and traffic all see it. The session comes from the
frame (Codex sends `prompt_cache_key`), so a conversation's turns group
like HTTP ones. A connection closed mid-turn cancels the turn; an upstream
that breaks fails it.

Each turn goes through the same admission as HTTP, in the same order:
calendar limits refuse, the key's `max_concurrent` waits, rolling limits
wait within `slot_wait_secs` or refuse, then the upstream's slot waits
within `slot_wait_secs` (the connection is bound to one upstream, so there
is no next candidate) and fails as busy. A refused turn leaves a row and is
answered with `response.failed` (code `insufficient_quota` when a calendar
period is used up); the connection stays open. The key's permit and the
upstream slot are held from sending a turn until its answer ends, so an
idle connection holds nothing. While a turn waits, the upstream side keeps
relaying and later client frames queue behind it, so a turn waiting for a
slot held by the previous turn on the same connection cannot deadlock.

The connection itself leaves no row, except when there is no turn to carry
the outcome: an upgrade a rule denies and an upstream that cannot be
connected keep their single row. Realtime and other WebSocket paths keep
one row per connection and the admission at open.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ccess samples

Responses WebSocket turns became requests of their own in the previous
commit, but they still lived outside the routing signals the HTTP path now
uses:

- turn::admit gave the rolling key-limit wait and the upstream-slot wait the
  full slot_wait_secs each, so a turn could wait twice as long as an HTTP
  request. It now sets one deadline after the key's max_concurrent gate and
  shares it between the two waits, as HTTP admission does.
- No turn ever fed the latency tracker, the breaker or the success rate:
  WebSocket did not touch Health at all, so an upstream failing every turn
  was never paused and kept its full share in a health-balanced group. A
  turn now records a speed sample from the moment the upstream starts on it
  (sent, or the previous turn on the connection ended) to its first content,
  read by the same first-token reader as HTTP. It records success or failure
  once, like an HTTP hop: first content is success; a turn that ends before
  any content is judged by its last frame with the same reader and
  classifier as an HTTP stream that errors before content (upstream fault =
  failure, the request's own fault = success); a broken upstream connection
  is a failure; a busy turn, a refused one, a client that leaves and a
  guard cut record nothing. The connection itself counts like a hop too: a
  handshake the upstream refuses or that cannot be made, and credentials
  that cannot be read at upgrade, record a failure; a whole-connection row
  (Realtime) that connects records a success.
- A turn that ends on an upstream `error` frame may be followed by a
  `response.failed` for the same response. It was attributed to the next
  queued turn, which then ended as that failure while its own answer fell
  to the turn after it. Such a frame is now recognised by its response id
  and passed to the client without being counted for any turn. Only turns
  that ended on `error` are remembered, each for one closing frame, and a
  turn's own `response.created` or id always wins, so an upstream that
  reuses one id for every response cannot leave a turn waiting forever.

The upstream slot tests named their route "默认" without default_route, so
every request went through the built-in all-upstreams group instead of the
group the tests are about. The config now sets it, and the helper that reads
each request's route asserts the request went through that group.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The routing round adds view, input and event fields across the control
plane: upstream max_concurrent, failover slot_wait_secs and
next_on_slow_start, attempt queued_ms / skipped / usage and the slow_start
outcome, load-balance weights and balance_by with the dry-run factors,
manual model specs and PUT /provider-model-spec, key usage limits with
KeyLimitAlert, and one row per Responses WebSocket turn. A UI written for 38
would drop group weights and key limits when it saves, so the version moves
to 39, with the notes above the constant describing it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Two ways a full upstream (max_concurrent) made failover worse than no
failover at all:

- The slow-start switch counted a full candidate as "a later one that can
  take the request". The request gave up on an upstream that was about to
  answer, then waited for the full one; with the recommended settings the
  shared wait deadline had already passed, so the client got an immediate
  429. At the deadline only candidates with a free slot now count; otherwise
  the slow upstream is treated as the last one and keeps streaming.

- When the wait for full candidates ended without a slot, the busy 429 won
  over a real failure: an upstream that answered 401 was hidden behind
  "every upstream is busy", with x-should-retry: true. The busy 429 is now
  only for requests that never reached an upstream; otherwise the last real
  failure is answered, as it is when no candidate is full. Errors that
  failover cannot fix (request errors) were already passed through at once.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
balance_by: latency read the same latency snapshot as url-test, and that
snapshot hands out the startup link-test seed until an upstream has three
real samples. The seed measures a TCP/TLS handshake (tens of ms), the real
samples measure time to first content (seconds), so a member with only a
seed looked tens of times faster and took the 10x factor cap. The latency
factor now uses real samples only; a member with just a seed is unmeasured
and neutral (1.0), as documented. url-test keeps using the seed, which is
what keeps it from acting like fallback right after startup.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A load-balance member that is full (max_concurrent) or paused sits out of
the pick for new conversations. advance() charged only members of that
round, so when stickiness kept a conversation on a full member, nothing was
charged: the member served after waiting for its slot, yet its share went
unrecorded, and the busier it was the more it got beyond its weight. The
member that actually leads after stickiness now always joins the round it
is charged in, so new conversations make up the difference as they do for
any other sticky request.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Day and week totals are rebuilt from the request records after a restart,
  just like the month's, but only month limits checked retention.row_days.
  A weekly limit with row_days 3 silently forgot the start of the week.
  Every calendar limit now needs records covering its period (day 1,
  week 7, month 31). The sentence names the period, so it gets a new code,
  config.key_limit_retention, replacing config.key_limit_month_retention.
- A cost limit below $0.01 rounds to zero micro-dollars or is used up by a
  single request; it is refused with config.key_limit_cost_too_small.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The schedule charges every request to the member that leads it, sticky
conversations included, so a weight is the long-run share of requests;
conversations in progress stay where they are and count toward that
member's share. The groups[].type cell, the prose around it and the API
comments on GroupView.weights and BalanceBy still said the weights share
out new conversations.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The test saved a cost limit of one micro-dollar with cache_reads, which the
new minimum of $0.01 now refuses first; use $1 so it still checks that
cache_reads is only for token limits.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Key limits counted every admitted request. A 429 because every upstream was
at its max_concurrent (sent with x-should-retry: true), a content-filter
denial or a phase-two rule refusal all counted, so a client retrying as
told burned its requests-per-minute limit although nothing was sent
anywhere and nothing was spent.

A request now counts only if it may have reached an upstream: some attempt
was sent (the upstream answered, or it failed after sending: timeout,
broken on the way, error before content). Skipped, refused, unconvertible,
credential-less and unreachable hops were never sent. A request with an
empty chain that failed was refused by the gateway; one that ended without
failing (the client left before the chain was reported) may still have been
in flight and counts as before. The rule lives once in tw-api
(RoutingView::reached_upstream); the recorder hands its verdict to the
settle hook, and the restart rebuild applies the same function in SQL
(tw_reached), so both totals agree. A request that does not count gives back
its reservation and the request it took in the rolling window, at
settlement and when its hold is dropped before it starts.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Books::period started a period's totals at zero whenever its computed start
differed from the one kept. A rollover is fine that way, but a time zone
change moves the start of today, this week and this month to a moment that
already has requests in it, and those totals dropped to zero: changing the
machine's time zone wiped out what a key had spent today.

When a key's period start changes, its totals are now read back from the
request store with the same query the startup rebuild uses. The gateway asks
once per key and period start (KeyLimits::reread_with), outside its lock;
tw_control::key_limits::follow does the read on a task, under the recorder's
lock, where every settlement also happens, so the totals it hands back are
exactly what has been settled. An answer for a period that has moved on
again is dropped. twcore wires it up after the startup rebuild. Until the
answer arrives the period counts from zero, as it did before.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
On a Responses WebSocket, a close frame or end of stream from the upstream
ended the pump the same way as the client leaving, so turns still in flight
were recorded as cancelled by the client, and one broken off before any
content counted nothing against the upstream's health. An upstream close is
now its own ending: the turns in flight fail with the new code
gw.ws.upstream_closed, and one without content yet counts as an upstream
failure, as an HTTP stream broken before content does. A client close is
still a cancellation with no health record. Whole-connection rows (Realtime
and other paths) still finish normally when either side closes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
An upgrade connected to the first available candidate, so a load-balance
group sent every new WebSocket connection to its first member, and
url-test/cheapest orders were ignored too; the dry-run kept predicting a
rotation that WebSocket traffic never followed.

The ordering step of the HTTP route (group order with weights and
balance_by, busy and paused members sitting out, stickiness, then charging
the leader) is now one function, pipeline::arrange, used by both paths. The
upgrade first drops candidates that cannot serve it (disabled ones; for a
Realtime connection also the ones whose alias or key scope does not fit),
orders the rest the same way and charges the member it actually connects
to. An upgrade has no body to recognise a conversation by, so it has no
stickiness.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A Realtime connection is one row, admitted with an empty estimate, and the
row carried no usage, so token and cost limits never saw anything spent on
it. The upstream reports each answer's usage in response.done
(response.usage: input_tokens including cached_tokens under
input_token_details, output_tokens). The connection now adds those up into
its row; the recorder prices it like any other row and settles it into the
key's limits when the connection closes. Limits are still checked when the
connection opens. Audio and image tokens are not told apart: the row is
priced at the model's per-token price, as other rows are. Other
whole-connection paths have no known usage format and still count only as a
request.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@fylorn
fylorn merged commit c2fbbc2 into main Oct 5, 2026
5 checks passed
@fylorn
fylorn deleted the integ/routing-2026-10 branch October 5, 2026 12:36
@fylorn fylorn mentioned this pull request Oct 5, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant