Downloading a 40 GB model to find out it crawls at 0.3 tokens a second is a bad afternoon. llmspec reads your hardware, works out what every model in its catalog would cost to run on it, and tells you which ones are worth your bandwidth — before you spend it.
$ llmspec fit -n 5
# Model Provider Params Quant Mode Fit Size Mem% Ctx tok/s Score
1 Qwen3 30B A3B Instruct 2507 Alibaba Qwen 30B/3B Q3_K_M MoE Good 14 GB 62% 32K 79.7 77.3
2 GPT-OSS 20B OpenAI 21B/4B Q6_K MoE Good 16 GB 79% 64K 43.5 77.2
3 Qwen2.5 7B Instruct Alibaba Qwen 7.6B Q4_K_M GPU Perfect 4.3 GB 83% 32K 34.9 76.8
4 Qwen3 30B A3B Alibaba Qwen 30B/3B Q3_K_M MoE Good 14 GB 62% 32K 79.7 75.3
5 GLM-4 9B Chat Zhipu AI 9.4B Q3_K_M GPU Marginal 4.3 GB 92% 32K 35.0 75.2
An 8 GB laptop GPU. Nothing here needs to be downloaded twice to find out it does not fit.
- Install · Quick start · The interactive interface
- Commands · Filtering · Benchmarking
- HTTP API · MCP · Verifying a download
- Configuration · Adding your own models
- How it decides · Supported hardware · Runtimes
cargo install --path .Or build it directly:
cargo build --release # target/release/llmspecNeeds Rust 1.85 or newer. The model catalog is compiled into the binary, so there is nothing else to install and nothing to fetch at runtime.
Pull requests run formatting, locked dependency tests, Clippy with warnings denied, and a generated-catalog consistency check. The Rust validation jobs use the repository’s declared Rust 1.85 toolchain and enable backtraces for actionable CI failures.
The catalog check also validates the generator’s Python syntax and rejects records with empty ids, non-positive parameter counts, or invalid contexts. Small regression tests exercise those invariants directly, including rejection of non-finite parameter counts.
CI cancels superseded runs for the same ref and applies explicit timeouts to each validation job, preventing stale or wedged checks from consuming runners. Checkout steps do not persist repository credentials into the runner workspace.
Release builds and crates.io publishing also use the committed lockfile, so published artifacts are built from the reviewed dependency graph. Release jobs use the same Rust toolchain as CI and have explicit timeouts for cross-platform builds and publishing.
llmspec # interactive interface
llmspec fit -n 10 # the ten best models for this machine
llmspec doctor # what was detected, and what was guessed
llmspec bench # measure real tokens/secNumeric hardware overrides are validated before probing the machine. Values
such as --cpu-cores 0 or --gpu-count 0 fail with an explanatory error.
Size overrides also reject non-finite values instead of allowing them into
memory calculations.
Four questions llmspec exists to answer:
"What should I download?"
llmspec fit --perfect -n 5"What fits in 8 GB and still runs at a readable speed?"
llmspec fit --memory 8G --min-tps 15 --max-size 6G"Can my machine run this specific model?"
llmspec info "Qwen2.5-Coder-32B""What would I need to buy to run it?"
llmspec plan "Llama-3.3-70B" --quant q4_k_m --context 32768plan answers for any machine, not this one. Add a throughput target and it
also names the memory bandwidth that speed needs, and the cards that have it:
llmspec plan "Llama-3.1-8B" --target-tps 60Running llmspec with no arguments opens the full interface: every model,
ranked, filterable, with a detail panel that tells you the exact command to
run the one you picked. Built with ratatui.
┌ llmspec ─────────────────────────────────────────────────────────────────┐
│ CPU AMD Ryzen 7 7840HS (8 cores / 16 threads) │
│ RAM 15.3 GB total, 6.3 GB available │
│ GPU NVIDIA GeForce RTX 4060 Laptop GPU — 8.0 GB VRAM, 272 GB/s │
│ Backend CUDA use case general runtimes Ollama │
└──────────────────────────────────────────────────────────────────────────┘
| Key | Action |
|---|---|
j k, ↑ ↓ |
Move between models |
g G, Home End |
Jump to first / last |
PgUp PgDn, Ctrl-U Ctrl-D |
Scroll by ten |
/ |
Search by name, provider, size or capability |
f |
Fit filter: all → runnable → perfect → good → marginal |
a |
Show: all → GGUF builds → already installed |
s |
Sort column |
u |
Target use case, and re-rank for it |
Enter |
Detail panel: memory, context, and the command to run it |
p |
Hardware plan: what this model would need from any machine |
m then c |
Mark a model, then compare it with the selected one |
d |
Download the selected model through Ollama |
r |
Re-probe local runtimes and installed models |
S |
Simulate different VRAM, RAM or core count |
A |
Edit the speed model's tunables |
t |
Cycle the colour theme |
h ? |
Help |
Esc |
Close the open panel or popup |
q |
Quit |
Esc only ever backs out one level; from the top it does nothing, so the
session cannot be ended by reflex. q and Ctrl-C quit.
The status bar along the bottom is clickable — each hint is a button that
sends the key it names — and the wheel moves between models. Capturing the
mouse means the terminal's own selection needs Shift held while dragging.
Downloads and runtime probes run in the background — the interface never blocks on the network. Twenty-one themes are included; the choice is remembered.
| Command | What it does |
|---|---|
llmspec |
Interactive interface |
llmspec fit |
Every model ranked for this machine |
llmspec recommend |
Top picks as JSON, for scripts and agents |
llmspec info <model> |
One model in full, with the command to run it |
llmspec plan <model> |
What hardware this model would need |
llmspec search <query> |
Search the catalog and rank the matches |
llmspec list |
The catalog, with no hardware analysis |
llmspec system |
Detected hardware |
llmspec gpus [filter] |
GPUs with known memory bandwidth, for use with --gpu |
llmspec doctor |
Diagnostic report; exits non-zero on warnings |
llmspec runtimes |
Local inference servers that are running |
llmspec bench |
Measure real tokens/sec against a running runtime |
llmspec serve |
Read-only HTTP API |
llmspec mcp |
Serve the analysis to an assistant over MCP |
llmspec verify <file> |
Check a downloaded model file for damage or truncation |
| Flag | Description |
|---|---|
--json |
Machine-readable output (the default for recommend) |
-u, --use-case |
general, coding, reasoning, chat, multimodal, embedding |
--force-runtime |
Score for one runtime: ollama, llamacpp, lmstudio, vllm, docker, mlx |
--memory SIZE |
Override VRAM, e.g. 24G — creates a synthetic GPU if none is found |
--ram SIZE |
Override system RAM, e.g. 128G |
--cpu-cores N |
Override the core count |
--gpu NAME |
Simulate a known GPU, e.g. "RTX 3090" — brings its VRAM and bandwidth (see llmspec gpus) |
--gpu-count N |
Number of simulated GPUs, with --gpu |
--kv-quant TYPE |
How the runtime stores the KV cache: f16, q8_0, q4_0 |
--max-context N |
Cap the context used for memory estimation |
--cli |
Force table output instead of the interface |
Sizes accept G/GB/GiB, M/MB/MiB, T/TB/TiB, case-insensitive.
The overrides make llmspec a shopping tool as much as a diagnostic one:
llmspec fit --memory 24G --ram 64G -u coding -n 10 # if I bought a 3090
llmspec fit --gpu "RTX 3090" --ram 64G -u coding # same, with the card's real bandwidth
llmspec fit --memory 0 --ram 64G # CPU-only server--force-runtime does two things: it shifts the throughput estimate to that
runtime's characteristics, and it hides models the runtime cannot load — a
GGUF loader never sees a model with no GGUF build.
fit narrows on the things people actually decide by:
llmspec fit --min-tps 20 # must be fast enough to read along
llmspec fit --max-size 8G # must fit the disk budget
llmspec fit --min-context 32768 # must hold a real working set
llmspec fit --perfect # must fit VRAM with headroom
llmspec fit --mode gpu # no CPU offload
llmspec fit --quant q4_k_m # placed at a specific quantization
llmspec fit --provider mistral # one publisherThey compose:
llmspec fit -u coding --min-tps 25 --max-size 10G --min-context 32768 -n 5"A coding model that runs at 25+ tokens a second, downloads in under 10 GB, and holds 32k of context."
Every other number llmspec prints is an estimate. bench is the ground truth:
it asks a running runtime to generate tokens and reports what came back.
llmspec bench # the first installed model on the first live runtime
llmspec bench qwen3:8b # a specific model
llmspec bench --all --runs 5 # everything installed, five timed runs each
llmspec bench --json # machine-readable$ llmspec bench qwen3:8b
Measured throughput
AMD Ryzen 7 7840HS · NVIDIA GeForce RTX 4060 Laptop GPU (8 GB VRAM) · 15 GB RAM · CUDA
Model Runtime tok/s range TTFT vs est.
------------------------------------------------------------------------------------
qwen3:8b Ollama 57.5 57.4-57.6 18ms 1.77x
What the estimate assumed
Qwen/Qwen3-8B at Q4_K_M, 16K context, 4.6 GB of weights read per token
If the ratio is far from 1.00x, check these against what the runtime actually
loaded; the remaining gap is the speed model itself, which is deliberately
conservative.
To match these measurements, set the efficiency factor to 0.97
(press A in the TUI; the value is saved for next time)
The first run is untimed — it pays for loading the model, and timing it would report disk speed rather than inference speed. Results are the median of the timed runs.
vs est. is measured over estimated. The estimate is conservative by design,
so a ratio above 1.0 is normal. What matters is that llmspec tells you exactly
what it assumed and hands you the one number that reconciles the two, instead
of leaving you to guess which knob to turn.
llmspec bench --calibrate applies that number for you: it fits the
efficiency factor to the runs it just measured and saves it with the GPU and
backend it was measured on. The calibration is only used on that machine, and
doctor reports whether one is in effect, where it came from and how old it
is.
llmspec serve exposes the analysis as read-only JSON. It binds loopback by
default, because it reports what hardware the machine has.
llmspec serve --host 127.0.0.1 --port 8228| Route | Returns |
|---|---|
GET /health |
Version, catalog size, route list |
GET /system |
Detected hardware |
GET /runtimes |
Local runtimes that are running |
GET /catalog |
The model database, unanalysed |
GET /models |
Ranked fit analysis |
GET /models/top |
The five best runnable models |
GET /models/{id} |
One model's analysis |
/models accepts limit, use_case, provider, search, quant, mode,
min_fit, perfect, include_too_tight, max_context, min_tps,
max_size_gb and min_context.
Numeric filters must be positive where applicable. Invalid, negative, zero,
non-finite or malformed values return HTTP 400 with a JSON error. An
ambiguous model name also returns HTTP 400 instead of silently selecting the
first catalog entry; use the model id or a more specific query.
curl "http://127.0.0.1:8228/models?use_case=coding&min_tps=20&limit=3"
curl "http://127.0.0.1:8228/models/Qwen%2FQwen2.5-7B-Instruct"The CLI applies the same fail-closed rule to --max-context and the
OLLAMA_CONTEXT_LENGTH environment variable: invalid values stop startup
with an explanatory error instead of being silently ignored.
The HTTP server also bounds request parsing: it accepts at most 64 query parameters, rejects duplicate keys, and rejects an individual query component larger than 4 KiB.
Built on std::net — serving adds no dependency.
llmspec mcp speaks the Model Context Protocol on stdin/stdout, so an
assistant can ask what this machine runs instead of being told.
claude mcp add llmspec -- llmspec mcpThere is nothing else to install: the server is the same binary, and the catalog is compiled into it.
| Tool | Answers |
|---|---|
system |
What hardware is here, and the bandwidth the estimates come from |
runtimes |
Which runtimes are running locally, and what each has installed |
fit |
What this machine can run, ranked — the fit command as a tool |
model |
One model's full analysis on this machine |
search |
Catalog lookup by name, provider, size or use case |
plan |
What hardware a model needs, and which GPUs reach a target tokens/sec |
verify |
Whether a model file on disk is intact |
fit takes use_case, limit, provider, min_fit, min_tokens_per_sec,
max_size_gb and include_unrunnable; plan takes context, quant and
target_tokens_per_sec.
The protocol has two eras — the newer revisions carry their version in each
request's _meta and answer server/discover, the older ones open with an
initialize handshake — and clients are still spread across both. llmspec
answers either, so it does not matter which one your client speaks.
Because stdout carries the protocol, every diagnostic goes to stderr. Clients show that as the server's log.
MCP messages are capped at 1 MiB. Oversized input receives a JSON-RPC invalid request response and is not passed to the JSON parser.
Numeric MCP arguments must also be finite; string values such as "NaN" and
"inf" are rejected as tool errors.
Runtime-reported parameter sizes follow the same rule: malformed, negative, zero and non-finite sizes are ignored rather than used to select a model. Provider benchmark samples also require a finite, positive elapsed duration before throughput is calculated. Hand-edited persisted speed factors and cached bandwidth values are sanitized on load; invalid values fall back to safe defaults rather than influencing model ranking.
A 40 GB download that stopped at 38 GB looks fine until the runtime chokes on
it. llmspec verify reads the header and says so in a fraction of a second,
without touching the weights.
Verification caps untrusted header and metadata lengths at 64 MiB and rejects claims that cannot fit in the remaining file before allocating or iterating. Duplicate GGUF metadata keys are treated as structural corruption rather than silently allowing the last value to win.
llmspec verify ~/.ollama/models/blobs/sha256-2bada8a745... GGUF, 4.4 GB
Contents
Version 3
Architecture qwen2
Name Qwen2.5 7B Instruct
Tensors 339
Parameters 7.6B
Types Q4_K (169), F32 (141), Q6_K (29)
Context 32K
Tensor data 4.4 GB
The file is structurally intact.
It reads GGUF and safetensors, and exits non-zero when the file is damaged, so it drops into a script between the download and the job:
error the file is 4.3 GB short: the tensor table describes 4.4 GB of data
but the file ends at 47.7 MB. The download did not finish.
The parameter count and quantization mix are summed from the tensor table, so they describe what the file contains rather than what its name claims. For safetensors it also checks that no two tensors claim the same bytes.
Every length in these formats comes from the file itself, so nothing here allocates on a number before proving the bytes exist to back it — a header claiming two billion tensors is refused immediately rather than honoured.
This is a structural check. It does not execute anything and cannot tell a well-formed malicious model from a well-formed honest one; it tells you the container is intact.
llmspec remembers the theme, the target use case and the speed tunables, so a session starts where the last one ended. Files live in:
| Platform | Location |
|---|---|
| Windows | %APPDATA%\llmspec\ |
| Linux, macOS | $XDG_CONFIG_HOME/llmspec/ or ~/.config/llmspec/ |
| Anywhere | $LLMSPEC_CONFIG_DIR overrides both |
config.json holds the settings. A malformed one falls back to defaults
rather than stopping llmspec from starting.
{
"theme": "dracula",
"use_case": "coding",
"speed": { "efficiency": 0.72, "gpu_factor": 1.0 }
}t cycles through the themes in the TUI, but the name can also be set by
hand. default follows the terminal's own palette; the rest are fixed RGB:
| Editor palettes | dracula nord solarized gruvbox monokai tokyo-night catppuccin-mocha rose-pine everforest kanagawa one-dark |
| llmspec's own | ocean forest sunset slate aurora matrix sakura |
| Legibility first | colourblind-safe high-contrast |
colourblind-safe is worth knowing about even if you never use it. Every
other theme separates the fit verdicts along the red-to-green axis, which is
the one axis a deuteranope cannot read; this one uses the Okabe-Ito palette
and runs them from blue to orange instead. high-contrast keeps every colour
above the WCAG AA contrast floor for projectors and glare.
A name this build does not have — a theme from a newer version, or a typo — falls back to the default rather than refusing to start. Configs from before themes were named stored a number instead; those are still read.
Anything newer than the build, or private, goes in models.json in the same
directory. It is merged into the catalog at startup.
{
"models": [
{
"id": "internal/support-bot-7b",
"name": "Support Bot 7B",
"provider": "Internal",
"params_b": 7.6,
"context_length": 32768,
"use_case": "chat",
"gguf": true,
"layers": 28, "hidden_size": 3584, "kv_heads": 4, "head_dim": 128
}
]
}| Field | Required | Notes |
|---|---|---|
id, name, provider |
yes | id is the upstream repository path |
params_b |
yes | Total parameters in billions |
context_length |
yes | Native maximum |
use_case |
yes | One of the six |
active_params_b |
no | Set it to declare a MoE model |
ollama |
no | Runtime tag, if one exists |
quality_tier |
no | 1–5 family reputation, default 3 |
layers, hidden_size, kv_heads, head_dim |
no | All four or none |
An entry whose id matches a shipped one replaces it — that is how a stale
record gets corrected locally.
With the geometry, the KV cache is sized exactly. Without it, llmspec falls back to a parameter-count heuristic, which is less precise on models that use multi-head rather than grouped-query attention.
Each model is scored 0–100 on four dimensions, weighted by use case:
| Dimension | What it measures |
|---|---|
| Quality | Parameter count, family reputation, quantization loss, use-case fit |
| Speed | Estimated tokens/sec from memory bandwidth and weight size |
| Fit | Memory efficiency — the plateau runs from 50% to 95% of the pool |
| Context | Context held against what the use case needs |
Reasoning leans on quality (0.55); chat leans on speed (0.35).
Quantization is chosen, not assumed. llmspec walks Q8_0 down to Q2_K and picks the best-scoring level that fits, retrying at shorter contexts when the full window will not.
Mixture-of-experts is modelled properly. Only the active experts need to be resident, so Mixtral 8x7B needs about the VRAM of a 13B model, not a 47B one.
A model that fits but crawls is not a recommendation. Scores scale towards zero below 3 tokens a second, and a model that does not fit anywhere scores 0 and sorts last.
| Fit level | Meaning |
|---|---|
| Perfect | Fits VRAM with 15% headroom |
| Good | Fits, or offloading cleanly |
| Marginal | Over 90% of VRAM, or CPU-only |
| Too Tight | Does not fit anywhere |
The full derivation — memory arithmetic, the bandwidth model, every weight and threshold, and the limitations — is in docs/ARCHITECTURE.md.
| Platform | Detection |
|---|---|
| Windows | CPU/RAM native; NVIDIA via nvidia-smi; AMD and Intel from the display driver's registry key |
| Linux | CPU/RAM native; NVIDIA via nvidia-smi; AMD via rocm-smi; Intel Arc via lspci |
| macOS | Apple Silicon detected from the chip; VRAM sized from unified memory |
Around 90 GPUs have a known memory bandwidth — NVIDIA consumer and datacenter,
AMD Radeon and Instinct, Intel Arc, and Apple M1 through M4. Unrecognised
cards fall back to a per-backend constant, and llmspec doctor says so rather
than quietly guessing.
Multi-GPU VRAM is summed, with a tensor-parallelism factor applied.
Two details worth knowing:
- On Windows, VRAM comes from the driver's 64-bit
qwMemorySize, notWin32_VideoController.AdapterRAM— that field is 32-bit and reports 4 GB for a 24 GB card. - On Apple Silicon, 75% of unified memory is treated as usable VRAM, matching
the default
iogpu.wired_limit_mbcap.
If detection is wrong, --memory, --ram and --cpu-cores override it.
llmspec talks to whatever inference server is already running.
| Runtime | Default endpoint | Environment override |
|---|---|---|
| Ollama | http://127.0.0.1:11434 |
OLLAMA_HOST |
llama.cpp (llama-server) |
http://127.0.0.1:8080 |
LLAMA_CPP_HOST |
| LM Studio | http://127.0.0.1:1234 |
LMSTUDIO_HOST |
| vLLM | http://127.0.0.1:8000 |
VLLM_HOST |
| Docker Model Runner | http://127.0.0.1:12434 |
DOCKER_MODEL_HOST |
| MLX | http://127.0.0.1:8080 |
MLX_HOST |
Ollama uses its own API; the rest speak the OpenAI-compatible /v1 surface.
Discovery does a TCP connect check first, so probing five absent runtimes
costs microseconds rather than five timeouts.
llmspec recognises a model across runtimes even when they disagree about its
name — Qwen/Qwen2.5-7B-Instruct, qwen2.5:7b, qwen2.5-7b-instruct and
Qwen2.5-7B-Instruct-Q4_K_M.gguf all resolve to the same catalog entry. The
matching is deliberately conservative: phi-4 and phi-4-mini stay distinct,
because a wrong "already installed" tick is worse than a missed one.
Ollama is the only runtime with a download API, so d in the interface works
there. For the others, the detail panel prints the command to run yourself.
Nothing leaves your machine. Every endpoint is loopback unless you point an environment variable elsewhere, and llmspec makes no other network calls.
240 models from 56 providers, embedded at build time: 39 mixture-of-experts architectures, 34 vision and multimodal models, 20 embedding and reranking models, from 23M to 1T parameters.
Meta · Alibaba Qwen · OpenAI · Google · Microsoft · Mistral AI · DeepSeek · Zhipu AI · IBM · Cohere · NVIDIA · Moonshot · MiniMax · xAI · AI21 · Databricks · Tencent · Baidu · ByteDance · LG AI · Allen AI · TII · Hugging Face · OpenBMB · Liquid AI · Arcee · Nous Research · Perplexity · Shanghai AI Lab · BAAI · Nomic · Jina · and more.
The catalog is generated by scripts/add_models.py, the only place a record is
written. CI runs it with --check, so a hand edit to data/models.json fails
the build instead of being lost on the script's next run.
cargo test # 287 tests
cargo clippy --all-targets
cargo fmtEnable the pre-commit hook once per clone:
git config core.hooksPath .githooksIt runs the same three checks as CI — cargo fmt --check, cargo clippy -- -D warnings, cargo test — and refuses the commit if any of them fails,
which takes about a second on a warm build. Every red CI run in this
repository so far has been a commit that cargo fmt alone would have caught.
Set LLMSPEC_SKIP_HOOKS=1 for the rare commit that has to go through anyway.
The test suite covers the memory arithmetic against hand calculations, GPU name matching including the collisions it is designed to avoid, placement and ranking under every sort order, runtime response parsing for both API shapes, HTTP routing and error statuses, config round-tripping, key handling, and TUI rendering — including cramped terminals, empty result sets and every theme.
Dependencies: clap, colored, ratatui, serde, serde_json, sysinfo,
ureq. crossterm is reached through ratatui's re-export rather than
declared directly, so the two cannot drift onto incompatible versions. No
unsafe code.
MIT
