A routing gateway for Claude Code, built on LiteLLM. It sends each request to a cheaper or stronger model tier based on what the request is, and measures what that saves.
Skills are recognized by the SHA-256 hash of their SKILL.md, not by guessing from the prompt. A live dashboard shows each request's cost next to what it would have cost on the model Claude Code asked for.
▫️ Same idea as an 802.1Q trunk port:
- Tagged frame → the VLAN ID picks the VLAN. Here: a registered skill's hash picks its tier
- Untagged frame → goes to the native VLAN. Here: plain chat goes to the lowest tier
- Allowed-VLAN list → only listed VLANs get their own path. Here: a skill missing from
catalog.yamlgoes to the lowest tier
registered skills plain chat unregistered subagents compaction, permission
(catalog.yaml) skills, titles suggestions, checks
recaps
__________▼_________________▼________________▼________________▼_______________▼_________________▼__________
\ hash verified │ sticky tier, │ lowest tier │ one tier below │ the session's │ tier running /
\ → catalog tier │ else lowest │ │ the session │ tier │ the model /
\ │ │ │ (then kept) │ │ asked for /
\________________│_______________│________________│________________│_______________│________________/
│ │ │ │ │ │
└─────────────────┴────────────────┴───────┬────────┴───────────────┴─────────────────┘
┌────────────────────┼────────────────────┐
▼ ▼ ▼
complex moderate light
Opus 5.5 · Sonnet 5 · Haiku 4.5 ·
high medium low
└────────────────────┼────────────────────┘
▼
Anthropic
- 🔀 llm-trunk
llm-trunk is a local proxy on 127.0.0.1:4000, between Claude Code and Anthropic. LiteLLM does the proxying; this repo adds the routing policy and the tools that measure its cost.
▫️ Tiers (set in catalog.yaml; models in litellm/config.yaml):
| Tier | Model | Effort | Max input | Max output |
|---|---|---|---|---|
light |
Haiku 4.5 | low | 64k tokens | 4,000 tokens |
moderate |
Sonnet 5 | medium | 128k tokens | 4,000 tokens |
complex |
Opus 5.5 | high | 180k tokens | 8,000 tokens |
▫️ Routing rules:
| Request | Tier |
|---|---|
| Your own messages | The session's sticky tier, else light |
| A registered skill | Its catalog tier, which becomes the session's sticky tier |
| An unregistered skill | light, and the sticky tier ends |
| Subagents | One tier below the session (never below light); each keeps its first tier |
Compaction (/compact) |
The session's tier, with no input cap |
| Suggestions and away summaries | The session's tier, so the prompt cache isn't lost |
| Session titles | light |
| Auto mode's permission checks | The tier running the model Claude Code asked for, so the safety check isn't weakened |
| Any move to a cheaper tier while the conversation is still cached on its current one | Waits on the current tier (⚓ held) until moving pays for re-caching the conversation |
▫️ Key characteristics:
- Bounded stickiness: after a registered skill runs, your next messages stay on its tier. This ends after 5 minutes idle (when the prompt cache expires), or 30 minutes after the skill was last run
- Cache-aware: a move to a cheaper tier waits while the conversation is still cached where it is, until moving costs no more than staying has. Moves up always happen at once
- Deny rules: an edited skill, or input over the tier's cap, is rejected with the reason before it reaches Anthropic
- Measured: every request is logged with its type, tier, tokens, cost, and the model that answered
- A local lab: one machine (macOS or Ubuntu) with Docker Compose, not a production service
▫️ Setup:
- Each skill's
SKILL.mdis hashed and listed incatalog.yamlwith a tier - Docker Compose starts LiteLLM and Postgres, with llm-trunk's routing callback loaded
- Claude Code gets a budget-capped LiteLLM key, never the real Anthropic key
▫️ Every request:
- The callback works out what the request is: a permission check, a subagent, compaction, a background call, a skill, or a normal message
- A skill counts as registered only if it's in the catalog, its hash matches, and the key is allowed to use it
- The request is allowed or denied. If allowed, its model, effort and output limit are set to the tier's, whatever Claude Code asked for
- After Anthropic answers, one line is written to the gateway's log
⚠️ NOTE: Background calls, compaction and permission checks are recognized by the exact text Claude Code sends. If a Claude Code update rewords that text, they'll be routed as normal messages. After updating Claude Code, run the manual sanity suite.
▫️ Access per key:
- A skill's hash isn't a secret: anyone who can read the skill files can send a registered skill and get its tier
- To limit that, set
allowed_skillsin a key's metadata. Other skills then count as unregistered for that key. A key without it can reach every tier - Permission checks never go above the highest tier the key can reach
| # | What you do | Gateway decision |
|---|---|---|
| 1 | Ask a plain question in a new session | light (Haiku 4.5, low) |
| 2 | Run /design-review |
complex (Opus 5.5, high), sticky |
| 3 | Ask a follow-up | stays on complex |
| 4 | Run /change-review |
moderate by the rules, but ⚓ held on complex: Opus still has the conversation cached, and moving would re-cache it on Sonnet 5 |
| 5 | Ask Claude to use a subagent | light, one tier below the rules' moderate |
| 6 | Run /personal-notes (not in the catalog) |
light, and the next message stays there |
| 7 | Run /compact in a complex session over the cap |
stays on complex, no cap |
| 8 | Edit a skill file, then run the skill | denied: the skill changed since it was hashed |
| 9 | Paste a ~200 KB log into a light session |
denied: input over the 64k cap |
| 10 | Start a new session | back to light |
A denial shows up in Claude Code as, for example: API Error: 400 llm-trunk: ~89409 input tokens exceeds the light tier cap (64000) — run /compact or start a new session.
⚠️ NOTE: Changing to a tier with a different model or effort makes the next message uncached, which is why moves down wait (see Concepts 101). Subagents are the cheapest place to use a lower tier, since they start with a fresh context.
▫️ Prerequisites:
| macOS | Ubuntu | |
|---|---|---|
| Docker with Compose v2 | Docker Desktop | Docker Engine + Compose plugin, then sudo usermod -aG docker $USER and log in again |
| Python 3.11+ | brew install python |
sudo apt install python3 python3-venv |
curl, openssl |
preinstalled | sudo apt install curl openssl |
| Claude Code | ✓ | ✓ |
| An Anthropic API key | ✓ | ✓ |
▫️ Step 1 - Clone and configure:
git clone https://github.com/pdudotdev/llm-trunk
cd llm-trunk
cp .env.example .env
Fill in .env:
LITELLM_MASTER_KEY: must start withsk-, e.g.echo "sk-$(openssl rand -hex 24)"ANTHROPIC_API_KEY: the key the gateway uses to call AnthropicPOSTGRES_PASSWORD:openssl rand -hex 24
▫️ Step 2 - Install the Python tools:
python3 -m venv .venv
source .venv/bin/activate
pip install -r requirements-dev.txt
Run source .venv/bin/activate again in each new terminal.
▫️ Step 3 - Start the gateway:
docker compose up -d
curl http://127.0.0.1:4000/health/liveliness # "I'm alive!"
▫️ Step 4 - Register your skills:
python3 scripts/hash_skill.py path/to/.claude/skills/my-skill/SKILL.md
Add the printed values to catalog.yaml under skills, with a tier:
my-skill:
tier: moderate
sha256: "…"
bytes: 424Re-hash a skill whenever you edit its SKILL.md, or it will be denied. Skills you don't register still work, on the lowest tier.
▫️ Step 5 - Point Claude Code at the gateway:
./scripts/create_usage_key.sh
cp client-settings.json.example path/to/your-client-repo/.claude/settings.json
The script prints a key (sk-…, $50 budget per 30 days). Put it in the copied settings.json as ANTHROPIC_AUTH_TOKEN.
▫️ Step 6 - Use it:
cd path/to/your-client-repo && claude # terminal 1: work as usual
python3 scripts/dashboard.py # terminal 2, in llm-trunk: live costs
The dashboard starts empty; add --since 1h to load recent history. Scroll the feed with ↑/↓ or the mouse wheel, g jumps back to the newest.
⚠️ NOTE: Changes tocatalog.yamlandpricing.yamlapply immediately. Changes underpolicy/or tolitellm/config.yamlneeddocker compose restart litellm. Changes todocker-compose.ymlneeddocker compose up -d, which recreates the container and so erases the routing history.
⚠️ NOTE: The routing history is kept in the gateway container's log.docker compose downerases it.
▫️ Desktop Code tab:
The Claude Desktop app's Code tab ignores a project's settings.json, so it cannot be routed per-repo. Desktop Code sessions always use your default settings (subscription or machine-wide). To use the gateway with a project, use the CLI (claude command) or the VS Code extension instead.
▫️ Machine-wide routing (for API-only teams):
If your team has no Claude Code subscriptions and routes all work through the gateway, set these in ~/.claude/settings.json:
{
"env": {
"ANTHROPIC_BASE_URL": "http://127.0.0.1:4000",
"ANTHROPIC_AUTH_TOKEN": "sk-…"
}
}▫️ Prompt caching (the (c) tag)
Anthropic caches the start of each request. A conversation only grows at the end, so each new message reuses everything before it from the cache and pays full price only for the new part:
| Request | Contents | From cache | New |
|---|---|---|---|
| Turn 1 | tools + system + msg 1 | nothing | everything |
| Turn 2 | … + msg 1 + reply 1 + msg 2 | up to msg 1 | reply 1 + msg 2 |
| Turn 3 | … + msg 2 + reply 2 + msg 3 | up to msg 2 | reply 2 + msg 3 |
- Price: a cache read costs 0.1× the normal input price (0.05× on Opus 5.5); a cache write costs 1.25×
- Expiry: about 5 minutes without use
- Reset by: changing the model or the effort settings, or rewriting earlier history (e.g.
/compact). That's why changing tiers costs one uncached message - Per model: each model caches separately. Opus 5.5 and Sonnet 5 even read the cache at the same price ($0.20 per million tokens), so moving a cached conversation from Opus to Sonnet costs a full re-cache and saves almost nothing afterwards
▫️ Cache-aware switching (the ⚓ held route)
Moving a conversation up (a skill that needs a stronger model) always happens at once. Moving it down only saves money, so it waits while the conversation is still cached on its current tier, and compares:
| Cost of this message | |
|---|---|
| Stay | Read the conversation from the current model's cache, write the new part |
| Move | Write the whole conversation on the cheaper model, minus what it already has cached (often Claude Code's tool list and system prompt, from other recent sessions) |
It stays while move − stay is more than what staying has cost so far compared with the cheaper tier, and moves once staying has cost that much, or at once when the cache has expired (then moving is free). It's the rent-or-buy rule: it never pays more than about one extra re-cache, and a pause usually ends a hold for free. Set cache_aware: false in catalog.yaml to always move at once.
Example (from a real run): after /design-review, /change-review wanted Sonnet 5. Moving re-cached 56k tokens for $0.143; staying on Opus 5.5 cost about $0.021.
The dashboard's "without llm-trunk" figure assumes Claude Code, talking to Anthropic directly, would have kept its cache: a re-cache the gateway causes counts against it.
▫️ Claude Code's own requests (the (i) tag)
| Request | When it's sent | Tier |
|---|---|---|
| Session title | Once, after your first messages | light |
| Next-prompt suggestion | A few seconds after a reply | the session's |
| Away summary | About 3 minutes after you leave the terminal | the session's |
| Permission check | Before actions, in auto mode | the one running the model asked for |
None of these extend the sticky timer, so they can't keep an expensive tier alive while nobody is working.
| File | Role |
|---|---|
catalog.yaml |
Tiers and registered skills |
litellm/config.yaml |
The model and prices for each tier |
pricing.yaml |
Anthropic's list prices, for the "without llm-trunk" comparison and cache-aware switching |
docker-compose.yml |
LiteLLM and Postgres, reachable only from this machine |
.env.example · client-settings.json.example |
Templates for .env and your client repo's .claude/settings.json |
policy/ |
The routing callback and rules |
scripts/ |
Key creation, skill hashing, dashboard, report, rule checker |
scenarios/ |
The scripted Claude Code session |
tests/ |
Automated tests, plus the manual sanity suite in tests/sanity/ |
- Cache-aware switching: only move to a cheaper tier when the savings outweigh re-caching
- Keep a held conversation on its model but lower its effort (Opus 5.5's per-message effort keeps the cache)
- Per-department virtual keys, each limited with
allowed_skills, with spend shown per key
You're responsible for creating your own API keys, paying for your usage, and checking catalog.yaml against your own skill files before routing real work through llm-trunk.
Licensed under the GNU General Public License v3.0.
Wanna say hello? DM me on LinkedIn.