English · Русский · 简体中文 · Español
Back to the project · Install · Launch profiles · Measurements
Start with AGENTS.md and the shared LocalForgeLLM skill. The framework covers three LLM work levels: install an existing model and stack; tune or modify them; build a derived artifact from full weights. Updates and interface integration apply across the levels. Image, video and audio generation have a separate planned branch.
| Read when | Guide |
|---|---|
| Understand the framework and component ownership | Architecture |
| Interpret a task, discover an implementation or resume a stack | Agent workflow and stack record |
| Install a published model and its dependencies | Level 1 — install |
| Change context, placement, performance or MoE behavior | Level 2 — tune and adapt |
| Select, build, switch or compare CUDA and Vulkan | Backend workflow and measured adaptation |
| Download a matching head; enable, disable or diagnose speculative decoding | MTP operations, memory and quality |
| Convert/calibrate/quantize full weights | Level 3 — build |
| Prune and rebuild full weights; place individual experts | REAP → APEX and expert caching |
| Fit or accelerate Ternary Bonsai 2 on 8–12 GB GPUs | Packing, calibration and GPU cache · Measurements |
| Add typed decisions and browser-use to a local agent | Jev API · Hermes MCP/skills in llama UI |
| Choose an engine or an optional weight-processing method | Engines · APEX |
| Connect or customize an agent/shell | Interface contracts · Pi · OpenShell · llama UI · Hermes |
| Update, repair or roll back a working stack | Operations |
| Establish capability, quality and resource use | Validation |
| Work on the future generative branch | Planned scope |
| Check versions and the evidence behind a guide | Sources |
The technical guides are maintained in English. The four language entry pages preserve the translated deployment examples below. Source-inspected configuration is distinguished from measured deployment results in verification scope.
- Give your coding agent the task, target context and memory budget. Specify a dense or MoE model or let it choose one. It inspects the CPU, GPU, RAM, storage and installed software.
- Using the framework, model and runtime documentation, it selects a compatible engine and weight format.
- It tunes CPU/GPU placement, threads, context, cache precision and batch sizes for your workload.
- It runs the task, checks the output, measures speed and RAM/VRAM use, and refines the settings. The result is a reusable launch profile with its runtime version and parameters.
The installation and commands below are two historical examples for one PC, using Qwen and Gemma with APEX GGUF. The framework applies the same selection and tuning workflow to dense and MoE models and other quantization methods.
| Evidence | Best recorded result / outcome | Procedure and data |
|---|---|---|
| Bonsai PTQ1_0, RTX 4060 8 GB / Ryzen 5 5600 | 27.14 tok/s short response; 20.22 with 31,018 input tokens; 6.24 GiB process VRAM snapshot | Case study · 32K profile |
| Bonsai PTQ1_0, later llama UI/MCP session on the same 4060 | 19.58 tok/s weighted; 7,355 output tokens / 9 completed requests; 1 cancelled request excluded | Session scope and CSV |
| Bonsai PQ2_0, community RTX 5070 12 GB / i7-10700K | 59.65 tok/s over 10K thinking tokens; 5.60× the report's forced-10K Opus row | ap3x0s report and CSV |
| Gemma APEX, RTX 4060 8 GB / Ryzen 5 5600 | 18.80 weighted phase mean; 19.38 best completed-response average | CUDA, RAM and MTP |
| Gemma REAP124 rebuild, RTX 4500 Ada 24 GB / 5995WX | GGUF −2.81%; no pruning decode gain; expert-cache experiment +8.07% decode | Tools, commands and quality limits |
| Jev with a local model and browser-use | Typed API routing and host tool integration; pinned browser-use/Cua review; no local Jev latency benchmark | API and harness contract |
These rows use different workloads. See each report for timing definitions and failures; no universal framework or model-quality ranking is implied. English graphics and regeneration.
Reference hardware: Linux x86_64, a Vulkan driver, RTX 4060 8 GB, Ryzen 5 5600 and 32 GB RAM. Reserve approximately 22 GB of SSD space for Gemma with vision, or 18 GB for Qwen. These are the tested configurations, not universal minimum requirements.
- Download the llama.cpp b10883 Vulkan archive and extract it into
runtime/. - Create
models/and download one profile below. Keep the original filenames; the vision projector ismmproj.gguf. - Start the selected profile, then open localhost:8080. Agent clients use
http://127.0.0.1:8080/v1, with model IDgemma-apexorqwen-apex. Run one model at a time.
| Profile | Downloads |
|---|---|
| Gemma 4 26B-A4B Heretic · I-Balanced | Model · 19.51 GB + vision · 1.19 GB |
| Qwen3.6 35B-A3B · I-Compact | Model · 17.29 GB · text profile |
Run from the directory containing runtime/ and models/. The archive creates runtime/llama-b10883/. Vulkan0 is the RTX 4060 in our setup; adapt it if your device order differs.
Requires a systemd user session. The 11 GiB service limit includes file cache; mmap allows model pages to be reclaimed and read from SSD again. This limits resident memory at the cost of possible I/O stalls. It does not reserve memory for other applications or guarantee against OOM.
mkdir -p logs
systemd-run --user --unit=gemma-apex --collect \
--property=MemoryMax=11G --property=MemorySwapMax=0 \
--property=OOMPolicy=kill --property=OOMScoreAdjust=1000 \
--property=LimitCORE=0 --property=CPUQuota=600% \
--property=Nice=5 --property=TimeoutStopSec=15 \
--property="StandardOutput=append:$PWD/logs/gemma-server.log" \
--property=StandardError=inherit \
"$PWD/runtime/llama-b10883/llama-server" \
--model "$PWD/models/gemma-4-26B-A4B-heretic-APEX-I-Balanced.gguf" \
--mmproj "$PWD/models/mmproj.gguf" \
--no-mmproj-offload --image-max-tokens 1120 \
--alias gemma-apex --host 127.0.0.1 --port 8080 --cors-origins localhost \
--device Vulkan0 --gpu-layers all --n-cpu-moe 27 --fit off \
--load-mode mmap --threads 6 --threads-batch 6 \
--ctx-size 32768 --parallel 1 --batch-size 1152 --ubatch-size 1152 \
--flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 \
--cache-ram 0 --ctx-checkpoints 1 \
--temp 1.0 --top-p 0.95 --top-k 64 --min-p 0 \
--log-colors off --log-timestampsStop with systemctl --user stop gemma-apex. The projector runs on the CPU. Keep the 1152 batch sizes with the 1120 image-token limit: smaller microbatches caused an image-processing assertion in this build.
For the optional MTP trial, follow the verified 0.46 GB head download and Gemma MTP settings. Add those options to a candidate copy of this command, retaining the main placement and context. Disable by restoring this original command through the same service manager. The target remains APEX I-Balanced; the separate head is Q8_0.
Runs in the foreground; stop with Ctrl+C. This profile has no service memory limit or vision projector.
./runtime/llama-b10883/llama-server \
--model ./models/Qwen3.6-35B-A3B-APEX-I-Compact.gguf \
--alias qwen-apex --host 127.0.0.1 --port 8080 --cors-origins localhost \
--device Vulkan0 --gpu-layers all --n-cpu-moe 36 --fit off \
--load-mode none --threads 6 --threads-batch 6 \
--ctx-size 32768 --parallel 1 --batch-size 512 --ubatch-size 128 \
--flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 \
--cache-ram 0 --ctx-checkpoints 4 \
--temp 1.0 --top-p 0.95 --top-k 20 --min-p 0 \
--log-colors off --log-timestampsThese profiles use APEX mixed-precision GGUF weights. It assigns precision by tensor role and layer; I- profiles use importance-matrix calibration. MoE activates a subset of experts per token, while llama.cpp splits execution between CPU and GPU. The complete model spans SSD, RAM and VRAM.
Ryzen 5 5600 · 32 GB RAM · Manjaro Linux · llama.cpp b10883 / Vulkan · September 10–11, 2026.
| Model / profile | Decode | Context window | RAM / model VRAM |
|---|---|---|---|
| Qwen3.6 35B-A3B · I-Compact | 21.11 tok/s | 32,768 | 13.43 GiB peak RSS / 4.07 GiB |
| Gemma 4 26B-A4B Heretic · I-Balanced + vision | 9.78 tok/s | 32,768 | 11 GiB service limit / 4.88 GiB snapshot |
| Gemma 4, same target + Q8 MTP head (September 11) | 14.83 tok/s | 32,768 | 11 GiB service limit / startup allocations |
September 10: Qwen had 28 completed requests, 19.37–22.09 tok/s, up to 29,087 reported context tokens; Gemma had one completed request with vision enabled, 599 input / 632 output tokens. September 11: Gemma MTP had two requests with final timings, 1835 logged output tokens and at most 2265 live context tokens; MTP vision quality was not tested. A configured 32K window is not a full-window stress test; decode rates exclude prompt processing.
Qwen resource monitoring covers 314 samples over the first 10 min 44 s of the session, including pauses:
| Resource | Average / maximum |
|---|---|
| Model RAM, RSS | 13.23 / 13.43 GiB |
| Model VRAM | 4.07 / 4.07 GiB |
| Model CPU, normalized to all 12 logical threads | 10.1% / 46.9% |
| Whole-GPU utilization, including desktop applications | 89.8% / 100% |
| GPU temperature | 50.4 / 54 °C |
Available system RAM fell to 2.53 GiB; free VRAM to 811 MiB. Gemma reached its 11 GiB cgroup ceiling, which includes file cache and is not the same measure as RSS. Its 4.88 GiB VRAM value is a snapshot; CPU/GPU utilization was not recorded for that profile.
Qwen generated 7,820 tokens across 28 completed requests. Weighted decode is (7820 − 28) / 369.07315 = 21.11 tok/s, following llama.cpp's first-token accounting. New-prompt processing averaged 67.36 tok/s. Gemma spent 97.59 s processing the prompt and 64.55 s decoding: 162.13 s total. These measurements describe serving speed, not a model-quality evaluation.
- Colibri: authors report 9.2 / 9.9 tok/s cold / warm, 40 GB peak RSS, Threadripper 3945WX + RTX 3070 8 GB, int4, 200-token decode.
- FreeToken: authors report 39.3 tok/s on RTX 4060 Laptop 8 GB + i9-13900H, 32 GiB LPDDR5, NVFP4, OpenCode coding workload. Installed RAM is not measured process consumption.
Our Qwen rate is 2.13× Colibri's cited warm rate and 46.3% below FreeToken's cited rate. The rows describe different CPUs, formats, workloads and averaging methods. For memory, our figure and Colibri's are process RSS; FreeToken's 32 GiB is installed capacity. Gemma's 11 GiB is a service limit including file cache. Keep these definitions when comparing configurations.
The September 11 no-MTP request measured 9.81 tok/s; two different MTP requests averaged 14.83 tok/s. The longer response sustained 15.03 tok/s during decode, but also took 47.29 s to process its new prompt. Main offload, CPU expert placement and the 32K allocation stayed fixed. Different requests and cache state prevent treating the +51.1% gap as causal MTP uplift. Formal quality checks remain pending.
The case study includes full-Q8 and QAT alternatives, the external measurements' methods, Colibri's missing Gemma result, FreeToken's modality limits and the sanitized timing data. Follow MTP operations for artifact discovery, verified downloads, switching and rollback.
The later CUDA, MTP and RAM adaptation adds an English infographic and 19 sanitized timing records across eight phases. CUDA + MTP with a 20 GiB cap averaged 18.80 tok/s across four responses; the initial backend switch alone did not improve the observed rate. The report separates runtime compilation time from total setup time. Use the backend workflow to apply and test these choices on another stack.