Repository navigation
Load-balance weights and distribution by speed and reliability, slow-start switch, upstream concurrency, model specs, key usage limits, per-turn WebSocket requests - #292
Merged
Conversation
A load-balance group could only take turns one by one, so an upstream
with twice the capacity (or half the price) got the same share as the
rest. A member can now be written as `{ name, weight }` (1-100); a plain
name is weight 1 and is written back as a plain name, so existing
configurations round-trip byte for byte, through the control plane's
writer too. Weights on other group types, or out of range, are refused
at load time (engine.group_weight_*) and when saving a group
(control.group.weight_*).
Distribution is nginx's smooth weighted round-robin, which interleaves
7:3 instead of sending seven in a row. It replaces the `seq`-based
rotation (and with it EventBus::peek_id, its only user). The per-group
state lives in the gateway (tw_gateway::balance) and reaches the engine
through Facts, so ordering stays a pure function and the dry-run reads
the same state without advancing it.
The state is charged at decision time, under one lock from ordering to
charging, so a burst of simultaneous requests spreads out. It charges
the member that leads after conversation stickiness: a conversation
that stays on A counts toward A, and new conversations make up the
difference, so the long-run ratio holds. Members that cannot serve the
request or are cooling down sit the round out and the rest share by
weight; otherwise their turn would fall to whoever follows them in the
group. A group whose members or weights change starts over.
The overview gives every member's weight for load-balance groups
(GroupView.weights), saving takes them (GroupInput.weights), and the
dry-run shows each candidate's weight (DryRunCandidate.weight).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A load-balance group spreads new conversations by its members' weights only, so a member that is slow or keeps failing gets as many new conversations as a fast, healthy one until the pauses catch it. groups[].balance_by (weights, latency, health, latency-health; default weights, omitted when default) multiplies each member's weight by a factor. The engine stays pure: balance_factors(by, members, facts) computes the factor per member from Facts. Latency uses the median TTFB that url-test already uses ((median / own)^2, clamped to 0.1..10); health uses a new Facts.success ((rate)^2, floored at 0.05 so a flaky upstream still gets an occasional request and its recovery is seen). A member without enough samples counts as 1.0, so a new upstream is tried but not flooded. The smooth weighted round-robin turns on effective weights: weight x factor x 1000, rounded, at least 1 (weighted::effective), computed in the one place the schedule reads weights (weighted::round), for picking the leader and for charging it alike. The x1000 keeps fractional factors meaningful with integer current weights; equal factors scale every member alike, so the interleaving is exactly what the weights alone give. Factors are taken over all candidates before the paused ones sit the round out, which is also what the dry-run shows. The success rate comes from Health itself: every record_* call also lands in a per-upstream window (last 50 outcomes within 30 minutes, at least 5 to report), so it counts exactly what the breaker counts and needs no new classification at the call sites. Client errors count as answered, a missing model and client cancellations are not counted. The data plane and the dry-run build their facts in one function (AppState::group_facts, now taking the group), which fills TTFB and success rates only for groups that use them; the dry-run reads the same snapshot and the same round-robin state without advancing it, and shows each candidate's weight, TTFB, success rate and factor. An end-to-end test checks that the dry-run names the upstream the next new conversation takes, request after request, while the factors move. balance_by on any other group type is refused at load time (engine.group_balance_not_load_balance) and by the control plane (control.group.balance_not_load_balance). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Some upstreams accept a request and then send nothing for a long time: a relay queueing it, an overloaded provider that does not say so, response headers that do not come. The client sees a hung request while another candidate could already be answering. failover.next_on_slow_start (off by default) gives up on such an upstream when a streamed answer has no content stream_start_wait_secs after the request was sent, covering both the wait for response headers and the held start of the stream. What counts as content is what the stream-start hold already judges (text, thinking/reasoning, tool calls; not role or start frames). Giving up drops the pending request or the response, so the connection closes and the upstream stops generating. - The last candidate never switches. "Last" accounts for the candidates that could not take the request anyway (paused by then, or the hop could not be sent: phase-two denial, no servable name, conversion failure, forced server tool), so the effectively last one waits instead of trading a slow answer for a certain failure. - The slow upstream is neither paused nor counted as a failure, nor as an outcome in the success rate load-balance groups weigh by. - The abandoned try is an attempt with outcome slow_start, with the usage the upstream reported at the start (Anthropic message_start) or else the gateway's input estimate marked as estimated (AttemptView.usage). - Only requests the client asked to stream; a non-streamed request is unaffected, and with the setting off the hold behaves as before. - Validation refuses switching with a wait under 5 seconds (config.slow_start_too_short); FailoverView.next_on_slow_start shows it. No keepalive is sent while holding: response headers are only sent once an upstream is chosen, and committing a 200 early would turn every later 429, 5xx and the last upstream's own 4xx into in-stream errors that clients do not retry on; Google's Python SDK also fails on SSE comment lines. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…t win over the price table A relay's own models are often missing from the price table, and the table is sometimes wrong, so clients were told no context window (or a wrong one) and the gateway could not tell whether a conversation still fits its model. `providers[].model_specs` maps an exact model id to `context_window` and/or `max_output_tokens`. One function in tw-config (`model_specs::resolve`, reached through `Provider::model_limits` / `Config::model_limits`) decides the precedence, field by field, and every reader goes through it: the `/v1/models` listing in all three shapes (a real model is now described by the first upstream offering it, so its spec applies; aliases as before), the upstream model rows, alias context windows, the pipeline's "input outgrew the held decision's context window" check, and the output limit filled in when a request is converted to Anthropic. `ModelRow` gains `context_window_source`, `max_output_tokens` and `max_output_tokens_source` so the UI can show which number is manual. `PUT /provider-model-spec` sets or (both null) removes one entry, editing only that upstream's `model_specs` so the rest of the entry stays byte for byte; the provider dialog keeps existing specs, and they follow renames and deletes because they live inside the entry. Validation: `config.model_spec_*`. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Some relays and accounts accept only a few requests at a time and refuse the rest, so a busy upstream turned ordinary load into failed requests. `providers[].max_concurrent` (1-1000) makes the gateway respect that limit before a request is sent, instead of learning it from a refusal. A request takes a slot when its hop is about to be sent (before plugins and conversion run) and gives it back when the answer has been relayed to its end or the client is gone. The engine's order is unchanged: - the upstream that conversation stickiness kept (its cache, or the same turn) is waited for, since moving loses the cache; - any other full upstream is skipped at once (attempt skipped: busy); - when every remaining candidate is full, the request waits for the first to free, in candidate order, then answers 429 with Retry-After in the client's API shape (gw.busy_all, recorded as rate_limited). All waiting shares one budget per request, `failover.slot_wait_secs` (default 30, 0 = never wait), because the client receives nothing while it waits. Waiting is not a failure: no pause, no breaker count, and no outcome in the success rate load-balance groups weigh by. Token counting and WebSocket connections take no slot. Limits live in gateway state across reloads and are resized in place; a removed limit releases its waiters. With failover.next_on_slow_start, a full upstream still counts as one that can take the request when deciding whether the slow one is the last: it may free up within the slot wait, so the slow upstream is given up on and the request may then wait for the full one (or get the busy 429). Giving up on a slow upstream returns its slot at once. Traffic shows `AttemptView.queued_ms` and `skipped: busy`; the control plane exposes `ProviderView/ProviderInput.max_concurrent` and `FailoverView.slot_wait_secs`. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The permit for `clients[].max_concurrent` was a local in pipeline(), so it was returned the moment the response headers went out. A key limited to one request could run any number of streams at once; the limit covered only the wait for headers. The permit now moves into the response body stream next to the in-flight counter and the upstream slot, and is returned when the answer ends or the client goes away. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…our, day, week or month A key on the LAN, or one handed to a script, can now be capped by what it spends rather than only by how many requests run at once. Limits live on the key (`clients[].limits`), each entry counts one measure, and every entry has to pass. - Day, week and month follow the calendar in the core machine's local time zone. Used up means refused until the next period: 429 with Retry-After to the reset and x-should-retry: false, and code insufficient_quota in the OpenAI shapes, so the official SDKs and Codex stop instead of retrying. - Minute and hour are rolling. Used up means waiting for the next free slot when it frees within failover.slot_wait_secs (the same setting as the wait for an upstream's free slot, counted separately), otherwise 429 with exact Retry-After and retry-after-ms. The 429 for upstreams that are all at their max_concurrent now sets its Retry-After through the same field, so it also carries retry-after-ms and x-should-retry: true. - Admission is one step after routing: calendar limits, then the existing max_concurrent gate (its permit still travels into the response body, so it is held until the answer ends), then rolling limits. A refused request is recorded like a routing refusal (start, empty attempt chain, failure with its own code). - Running requests hold their input estimate (and its input cost on the first candidate); the recorder settles each row to the usage and cost it writes, through a direct hook rather than the bus, so the in-memory totals equal what a restart adds back from the store. Token counts and local answers do not count; unpriced models and free upstreams count $0. - A restart rebuilds this day, week and month from the request records, so a monthly limit requires retention.row_days >= 31. - The key view carries each limit with used/max/resets_at/reached, and the models without a price for keys that have a cost limit; the key input takes the limits; renaming a key through the control plane carries its usage. - KeyLimitAlert on the event bus once per period at 80% and at the limit, for Lite's system notifications. - WebSocket connections are admitted once when they open: the store records one row per connection without usage. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…budget Latency samples (url-test ordering and the load-balance latency factor) were taken when relaying began and counted from the request's arrival, so they included key-limit and slot waits, plugins and every earlier failed or abandoned hop, and ended at a point that depended on the hop's position (first content behind the opening watch, response headers for the last candidate, nearly the whole answer when not streamed). An upstream dropped by the slow-start switch recorded nothing and kept its old fast samples, while the next upstream was charged the wait. A sample is now the time from sending that hop to the first content of a streamed answer, detected by the same first-token reader that reports RequestFirstToken, wherever the hop sits in the candidate list. Answers that are not streamed give no sample. An upstream given up on for a slow start records the time it was given, a lower bound that ranks it slow. L1 handshake seeds stand in only until an upstream has enough real samples and are then dropped, never mixed into the median. Load-balance: a member at its max_concurrent at decision time sits the round out like a paused one. Charging it for a request the slot cap then skipped made a busy member drift below its share. failover.slot_wait_secs is now one deadline per request, set when the key gate lets it in and shared by the rolling key-limit wait and the upstream slot wait; before, each could use the full time, so one request could wait twice as long. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A Responses WebSocket connection used to be one request row for its whole life, with no usage and no cost: ws.rs never read `usage`, the key's usage limits were checked once when the connection opened, and neither the key's `max_concurrent` nor the upstream's slot applied at all. A client configured for WebSocket could therefore spend without limit and without a trace of what it cost. Now each `response.create` is a request, from the frame to the answer's `response.completed` / `response.failed` / `response.incomplete` (or an `error`): RequestStarted, RequestHeaders, RequestRouted (one hop, the connection's upstream, with queued_ms), first token and an ending that carries the turn's usage, so the recorder prices it like any HTTP request and key usage, checkup and traffic all see it. The session comes from the frame (Codex sends `prompt_cache_key`), so a conversation's turns group like HTTP ones. A connection closed mid-turn cancels the turn; an upstream that breaks fails it. Each turn goes through the same admission as HTTP, in the same order: calendar limits refuse, the key's `max_concurrent` waits, rolling limits wait within `slot_wait_secs` or refuse, then the upstream's slot waits within `slot_wait_secs` (the connection is bound to one upstream, so there is no next candidate) and fails as busy. A refused turn leaves a row and is answered with `response.failed` (code `insufficient_quota` when a calendar period is used up); the connection stays open. The key's permit and the upstream slot are held from sending a turn until its answer ends, so an idle connection holds nothing. While a turn waits, the upstream side keeps relaying and later client frames queue behind it, so a turn waiting for a slot held by the previous turn on the same connection cannot deadlock. The connection itself leaves no row, except when there is no turn to carry the outcome: an upgrade a rule denies and an upstream that cannot be connected keep their single row. Realtime and other WebSocket paths keep one row per connection and the admission at open. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ccess samples Responses WebSocket turns became requests of their own in the previous commit, but they still lived outside the routing signals the HTTP path now uses: - turn::admit gave the rolling key-limit wait and the upstream-slot wait the full slot_wait_secs each, so a turn could wait twice as long as an HTTP request. It now sets one deadline after the key's max_concurrent gate and shares it between the two waits, as HTTP admission does. - No turn ever fed the latency tracker, the breaker or the success rate: WebSocket did not touch Health at all, so an upstream failing every turn was never paused and kept its full share in a health-balanced group. A turn now records a speed sample from the moment the upstream starts on it (sent, or the previous turn on the connection ended) to its first content, read by the same first-token reader as HTTP. It records success or failure once, like an HTTP hop: first content is success; a turn that ends before any content is judged by its last frame with the same reader and classifier as an HTTP stream that errors before content (upstream fault = failure, the request's own fault = success); a broken upstream connection is a failure; a busy turn, a refused one, a client that leaves and a guard cut record nothing. The connection itself counts like a hop too: a handshake the upstream refuses or that cannot be made, and credentials that cannot be read at upgrade, record a failure; a whole-connection row (Realtime) that connects records a success. - A turn that ends on an upstream `error` frame may be followed by a `response.failed` for the same response. It was attributed to the next queued turn, which then ended as that failure while its own answer fell to the turn after it. Such a frame is now recognised by its response id and passed to the client without being counted for any turn. Only turns that ended on `error` are remembered, each for one closing frame, and a turn's own `response.created` or id always wins, so an upstream that reuses one id for every response cannot leave a turn waiting forever. The upstream slot tests named their route "默认" without default_route, so every request went through the built-in all-upstreams group instead of the group the tests are about. The config now sets it, and the helper that reads each request's route asserts the request went through that group. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The routing round adds view, input and event fields across the control plane: upstream max_concurrent, failover slot_wait_secs and next_on_slow_start, attempt queued_ms / skipped / usage and the slow_start outcome, load-balance weights and balance_by with the dry-run factors, manual model specs and PUT /provider-model-spec, key usage limits with KeyLimitAlert, and one row per Responses WebSocket turn. A UI written for 38 would drop group weights and key limits when it saves, so the version moves to 39, with the notes above the constant describing it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Two ways a full upstream (max_concurrent) made failover worse than no failover at all: - The slow-start switch counted a full candidate as "a later one that can take the request". The request gave up on an upstream that was about to answer, then waited for the full one; with the recommended settings the shared wait deadline had already passed, so the client got an immediate 429. At the deadline only candidates with a free slot now count; otherwise the slow upstream is treated as the last one and keeps streaming. - When the wait for full candidates ended without a slot, the busy 429 won over a real failure: an upstream that answered 401 was hidden behind "every upstream is busy", with x-should-retry: true. The busy 429 is now only for requests that never reached an upstream; otherwise the last real failure is answered, as it is when no candidate is full. Errors that failover cannot fix (request errors) were already passed through at once. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
balance_by: latency read the same latency snapshot as url-test, and that snapshot hands out the startup link-test seed until an upstream has three real samples. The seed measures a TCP/TLS handshake (tens of ms), the real samples measure time to first content (seconds), so a member with only a seed looked tens of times faster and took the 10x factor cap. The latency factor now uses real samples only; a member with just a seed is unmeasured and neutral (1.0), as documented. url-test keeps using the seed, which is what keeps it from acting like fallback right after startup. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A load-balance member that is full (max_concurrent) or paused sits out of the pick for new conversations. advance() charged only members of that round, so when stickiness kept a conversation on a full member, nothing was charged: the member served after waiting for its slot, yet its share went unrecorded, and the busier it was the more it got beyond its weight. The member that actually leads after stickiness now always joins the round it is charged in, so new conversations make up the difference as they do for any other sticky request. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Day and week totals are rebuilt from the request records after a restart, just like the month's, but only month limits checked retention.row_days. A weekly limit with row_days 3 silently forgot the start of the week. Every calendar limit now needs records covering its period (day 1, week 7, month 31). The sentence names the period, so it gets a new code, config.key_limit_retention, replacing config.key_limit_month_retention. - A cost limit below $0.01 rounds to zero micro-dollars or is used up by a single request; it is refused with config.key_limit_cost_too_small. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The schedule charges every request to the member that leads it, sticky conversations included, so a weight is the long-run share of requests; conversations in progress stay where they are and count toward that member's share. The groups[].type cell, the prose around it and the API comments on GroupView.weights and BalanceBy still said the weights share out new conversations. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The test saved a cost limit of one micro-dollar with cache_reads, which the new minimum of $0.01 now refuses first; use $1 so it still checks that cache_reads is only for token limits. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Key limits counted every admitted request. A 429 because every upstream was at its max_concurrent (sent with x-should-retry: true), a content-filter denial or a phase-two rule refusal all counted, so a client retrying as told burned its requests-per-minute limit although nothing was sent anywhere and nothing was spent. A request now counts only if it may have reached an upstream: some attempt was sent (the upstream answered, or it failed after sending: timeout, broken on the way, error before content). Skipped, refused, unconvertible, credential-less and unreachable hops were never sent. A request with an empty chain that failed was refused by the gateway; one that ended without failing (the client left before the chain was reported) may still have been in flight and counts as before. The rule lives once in tw-api (RoutingView::reached_upstream); the recorder hands its verdict to the settle hook, and the restart rebuild applies the same function in SQL (tw_reached), so both totals agree. A request that does not count gives back its reservation and the request it took in the rolling window, at settlement and when its hold is dropped before it starts. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Books::period started a period's totals at zero whenever its computed start differed from the one kept. A rollover is fine that way, but a time zone change moves the start of today, this week and this month to a moment that already has requests in it, and those totals dropped to zero: changing the machine's time zone wiped out what a key had spent today. When a key's period start changes, its totals are now read back from the request store with the same query the startup rebuild uses. The gateway asks once per key and period start (KeyLimits::reread_with), outside its lock; tw_control::key_limits::follow does the read on a task, under the recorder's lock, where every settlement also happens, so the totals it hands back are exactly what has been settled. An answer for a period that has moved on again is dropped. twcore wires it up after the startup rebuild. Until the answer arrives the period counts from zero, as it did before. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
On a Responses WebSocket, a close frame or end of stream from the upstream ended the pump the same way as the client leaving, so turns still in flight were recorded as cancelled by the client, and one broken off before any content counted nothing against the upstream's health. An upstream close is now its own ending: the turns in flight fail with the new code gw.ws.upstream_closed, and one without content yet counts as an upstream failure, as an HTTP stream broken before content does. A client close is still a cancellation with no health record. Whole-connection rows (Realtime and other paths) still finish normally when either side closes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
An upgrade connected to the first available candidate, so a load-balance group sent every new WebSocket connection to its first member, and url-test/cheapest orders were ignored too; the dry-run kept predicting a rotation that WebSocket traffic never followed. The ordering step of the HTTP route (group order with weights and balance_by, busy and paused members sitting out, stickiness, then charging the leader) is now one function, pipeline::arrange, used by both paths. The upgrade first drops candidates that cannot serve it (disabled ones; for a Realtime connection also the ones whose alias or key scope does not fit), orders the rest the same way and charges the member it actually connects to. An upgrade has no body to recognise a conversation by, so it has no stickiness. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A Realtime connection is one row, admitted with an empty estimate, and the row carried no usage, so token and cost limits never saw anything spent on it. The upstream reports each answer's usage in response.done (response.usage: input_tokens including cached_tokens under input_token_details, output_tokens). The connection now adds those up into its row; the recorder prices it like any other row and settles it into the key's limits when the connection closes. Limits are still checked when the connection opens. Audio and image tokens are not told apart: the row is priced at the model's per-token price, as other rows are. Other whole-connection paths have no known usage format and still count only as a request. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Merged
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
A routing and limits round: weights on load-balance groups and distribution by speed and reliability, moving a slow-starting stream to the next upstream, a concurrency cap per upstream, model specs set by hand, usage limits per gateway key, and Responses WebSocket turns counted as requests. Control-plane protocol 38 → 39. Answers #289.
Load-balance groups
providers: [{ name: relay, weight: 7 }, official](1–100; a plain name is 1 and is written back as a plain name, so existing configurations round-trip unchanged). Weights only onload-balance(engine.group_weight_not_load_balance,engine.group_weight_out_of_range).Facts, so the dry-run still predicts the data plane exactly. A request that stays on an upstream for its conversation is charged to it, so the weight is the long-run share of requests. Paused and full members sit out a round; the member that leads after stickiness is always charged.groups[].balance_by: weights | latency | health | latency-healthmultiplies the weights by a factor: latency = (median of the members' time to first content / this member's)², clamped to [0.1, 10]; health = success rate², floored at 0.05. A member without enough samples counts as average (1.0).pipeline::arrange).Speed and success samples
url-testkeeps using the startup link test until an upstream has 3 real samples; load-balance factors use real samples only.Slow stream start —
failover.next_on_slow_start(default off): a streamed answer with no content afterstream_start_wait_secsis cancelled and the next upstream is asked, if one could take the request at that moment; otherwise the current one keeps streaming. The last upstream always waits. The slow upstream is not paused. The abandoned attempt is recorded with outcomeslow_startand its usage (AttemptUsage), marked as possibly billed.config.slow_start_too_short: with switching on, the wait must be at least 5 s. No keepalive while holding: response headers are not sent yet, and Google's Python SDK rejects SSE comment lines.Concurrency —
providers[].max_concurrent(1–1000) andfailover.slot_wait_secs(default 30, one wait budget per request, shared with the key's rolling limits). A conversation that stays on a full upstream waits for a slot, then moves on; other candidates are skipped at once (ServeSkip::Busy). When every candidate is full the request waits for the first free slot; if none frees and no hop reached an upstream, 429gw.busy_all(Retry-After); otherwise the last real failure. The per-keyclients[].max_concurrentpermit is now held until the answer has been relayed (it was released when the response started, so streamed bodies were not counted).Model specs —
providers[].model_specs: { <model>: { context_window?, max_output_tokens? } }win over the price table field by field, wherever core reports or uses a model's limits (the listing in every shape,ModelRow, aliases, the re-route when input outgrows the context window, the default max_tokens on conversion to Anthropic).PUT /provider-model-spec(ModelSpecSave→ConfigWritten) sets or clears one entry. A real model in/v1/modelsis now described by the first upstream offering it.Usage limits per gateway key —
clients[].limits: [{ per: minute|hour|day|week|month, requests|tokens|cost: N, cache_reads?: bool }], all must pass. Minute and hour roll; day, week (from Monday) and month follow the machine's local time. Tokens are uncached input + cache writes + output, plus cache reads whencache_reads: true. Cost is in USD (at least 0.01); unpriced models andbilling: freeupstreams count 0. In-flight requests reserve their estimate and settle from the recorded usage; a request that never reached an upstream counts for nothing. Totals are rebuilt from the request store at startup and when a period's start moves (time-zone change). Calendar limits answer 429 withRetry-Afterto the reset andx-should-retry: false(insufficient_quotain OpenAI bodies); rolling limits wait within the request's wait budget, then 429 with the exactRetry-After.Event::KeyLimitAlertat 80% and on reaching a limit. Validation:config.key_limit_*, includingconfig.key_limit_retention(retention must cover the period after a restart).Responses WebSocket — each
response.createis its own request: started/routed/first token/ending events, a request-store row with the turn's usage and cost, per-turn key limits, key permit and upstream slot, latency and success samples; a refused or busy turn getsresponse.failedand the connection stays usable. An upstream that closes mid-turn fails the turn (gw.ws.upstream_closed). Realtime stays one row per connection, now with the usage of itsresponse.doneevents, counted against key limits when it closes.Upgrade notes for the release
CONTROL_API_VERSIONdoc comment lists the changes). A UI written for 38 drops weights and limits when it saves groups and keys.url-testordering changes, since samples now measure time to first content per hop.max_concurrentnow also covers streamed bodies, so a key at its limit queues more than before.Message codes — new:
config.key_limit_cache_reads,config.key_limit_cost_too_small,config.key_limit_duplicate,config.key_limit_empty,config.key_limit_not_positive,config.key_limit_retention,config.key_limit_two_measures,config.model_spec_blank_model,config.model_spec_empty,config.model_spec_wildcard,config.model_spec_zero,config.provider_concurrency_range,config.slow_start_too_short,control.group.balance_not_load_balance,control.group.weight_not_load_balance,control.group.weight_not_member,control.group.weight_out_of_range,engine.group_balance_not_load_balance,engine.group_weight_not_load_balance,engine.group_weight_out_of_range,gw.busy_all,gw.busy_upstream,gw.key_limit.{requests,tokens,cost}_{per_period,rolling},gw.slow_start,gw.ws.upstream_closed. None removed.Built in parallel lanes and integrated locally; an adversarial review before this PR found 10 defects (slow-start switch dropping a working upstream for a full one, a busy 429 hiding a real error, link-test seeds used as speeds, refusals consuming request limits, Realtime outside token/cost limits, two validation gaps, a time-zone change resetting totals, an upstream close counted as a client cancel, an uncharged sticky leader, WebSocket upgrades ignoring load-balance), each fixed with a test that failed before the fix.
Checked locally:
cargo fmt,cargo clippy --workspace --all-targets -D warnings, 3013 workspace tests.🤖 Generated with Claude Code