Skip to content

Latest commit

 

History

History
65 lines (47 loc) · 12 KB

File metadata and controls

65 lines (47 loc) · 12 KB

Source map and verification scope

Documentation

The initial implementation review was on September 10, 2026, with dated follow-ups below through September 18. The pinned snapshots make these guides reviewable; they are not a mandate to downgrade or upgrade a working installation. At execution time, identify the installed release and check changes before using a command or code map.

Project Inspected source Main entry points
llama.cpp / llama UI 41fc7584f0c1 docs/build.md, tools/server/, tools/quantize/, tools/imatrix/, tools/ui/, src/models/
llama.cpp b10883 placement follow-up, September 11, 2026 91f6a6cf3 Installed server/bench --help, common/arg.cpp, common/common.h, src/llama-model.cpp, src/models/qwen35.cpp; dense-FFN placement and graph, not a new model throughput result
llama.cpp CUDA/Vulkan follow-up, September 11, 2026 b10883 build guide, kernel cgroup v2 memory interface Local pinned checkout: docs/build.md, ggml/CMakeLists.txt, ggml/src/ggml-cuda/CMakeLists.txt, tools/ui/CMakeLists.txt; successful local CUDA build and device discovery, compile-resource record and sanitized session timings. Generic build examples are source-checked, not fresh builds performed for this documentation update
llama.cpp MTP follow-up, September 11, 2026 b10883 / 91f6a6cf36138 common/speculative.cpp, common/common.cpp, common/arg.cpp, common/common.h, src/models/qwen35.cpp, src/models/delta-net-base.cpp, src/llama-sampler.cpp, server and Vulkan dispatch; MTP guide includes upstream issue/patch status, not a claim of a verified local MTP fix
Gemma assistant and artifact follow-up, September 11, 2026 b10883 assistant source, Unsloth head revision Assistant dimensions, target KV sharing, exact artifact bytes/SHA-256, native enable/disable flags; download recipe
APEX 636cec7e3f8d scripts/generate_config.sh, scripts/quantize.sh, sensitivity/allocation scripts
Ornith custom Q8 / September 12 BF16 GGUF revision 12393612 Saved APEX custom recipe, 753 tensor rules, integrity/download verification and local resource report; Q8 build and failed residency tradeoff. No REAP, no imatrix, no independently reproduced parent-to-BF16 conversion
Cerebras REAP / September 12 experiment 1970473c51ca src/reap/pruning_metrics.py; local Gemma 4 observer/export adapter and self-check; REAP → APEX build and quality limits
Individual expert cache / September 12 adrianhoehne fork, 018789c Corrected strided top-k profiling; static hot cache, prefill bypass, 40 measured requests; sanitized aggregate CSV. Requires the recorded local reader correction for profiling
Bonsai 2 / September 18 model revision 6ed5e12, PrismML runtime d8f26eec Locally downloaded model card, file manifests, runtime help/logs, native KV mean-centering and saved API checks; packing and launch recipe
Jev / September 18 TypeSafe API, models, launch article, September 15 HTTP schema, jev-1.13.0 and moving aliases, typed outputs, input limits and vendor pricing/latency; integration guide. No new paid call or vendor benchmark rerun
Jev browser-use / September 18 browser-use/jev-ultrafast, 452c1ad Pinned upstream agent.py, browser.py, snapshot.js, model.py, questions.py, author performance report and local host adapter; source walkthrough and adaptation boundaries
Cua / September 18 trycua/cua, 05f29785 Driver README, Linux MCP contract, cua-driver-core/src/browser/tools.rs and engine.rs: exact bindings, current refs and mutation checks; proposed Jev mapping. Source review only; no Cua installation or browser execution
Bonsai llama UI MCP / September 18 PrismML server source Installed help, server MCP schema and text-result handling; recorded 47-tool listing and actual skill lookup/read cycle; interface boundary
FreeToken fb7f732de08a docs/install.md, docs/cli.md, docs/models.md, python/freetoken/
FreeToken Gemma comparison follow-up, September 11, 2026 0ffd5c8b2941 model guide, issue #188 Text-only serving of multimodal checkpoints; separate author-reported patched 16 GB GPU result, not a local or official 8 GB benchmark
Colibri fd93c41aa6ae Family guides, c/coli, family engine/converter
Pi 4bd3f48df0b1 packages/coding-agent/, packages/agent/, packages/ai/, packages/tui/
OpenShell a0814443f19c docs/sandboxes/, docs/providers/, architecture/, gateway/supervisor crates
Hermes 67764dc08633 website/docs/, agent/, hermes_cli/, tools/, web/, apps/desktop/

The Hugging Face CLI and MLX-LM are additional discovery references. Inspect their selected versions before installing or integrating them.

Existing measured stack

The Qwen/Gemma examples preserve the working llama.cpp b10883 configurations and their measurements. The legacy OpenShell/Pi connection uses OpenShell 0.0.110 and Pi 0.84.2. Current source maps are separate from those previously measured versions.

The September 11 Gemma MTP case study adds three sanitized timing rows: one no-MTP request and two MTP requests. It verifies 14.83 tok/s weighted MTP decode and 15.03 tok/s over a 102.05-second decode interval. It does not establish a causal +51.1% MTP gain, quality parity, MTP vision/tool support or performance on a filled 32K context. The comparison links primary benchmark reports and explicitly separates their hardware, quantization and methodology. Raw sessions and machine paths remain local.

The later CUDA/MTP/RAM report covers 19 completed responses across eight phases; its first three timing rows overlap that MTP case study. It records the backend transition, RAM-cap and MTP changes, 18.80 tok/s weighted decode in the four-response CUDA/MTP/20-GiB phase and 10 min 30.161 s for the initial CUDA compilation. Workload/cache differences prevent isolating a CUDA-only gain; compilation excludes other setup work. The reusable backend workflow keeps those distinctions explicit.

The Bonsai report adds six unique community runs from ap3x0s and three selected local timing rows. The subscriber supplied aggregate HTML, not raw logs; the source hash and reported-unit limitations are recorded. RTX 5070 / PQ2_0 results remain separate from RTX 4060 / PTQ1_0. The latter includes a 31,018-token lookup and memory sampling, plus the original reasoning-assertion failure and corrected rerun.

Its separate interactive-session CSV records nine completed generations and one cancelled stream: 7,355 output tokens and 19.58 tok/s weighted decode. Reasoning-token and tool-call counts from that user conversation are unavailable. The report separately labels the saved reasoning, two-call MCP skill loop and local field-text checks; no conversation content is published. Model shutdown was verified without starting new inference.

The REAP/APEX record documents a completed rented-server build, rather than an unrun generic recipe. Its seven aggregate CSV rows separate the 32-request ABBA chat benchmark from 40 expert-cache timing trials. The quality target was not established; the artifact remained experimental. No rented-server speed is relabeled as an RTX 4060 result.

The September 18 graphics refresh recalculates phase means and ratios from the tracked data with the offline generator/check. Existing assets retain English labels and the project's monochrome design. Local documentation links, embedded command syntax and SVG structure can be checked without running inference. PNG export is artifact generation; no visual/UI review or live Jev browser trial is claimed for this documentation task.

The framework documents three kinds of evidence:

  • Measured: the named existing deployment runs and their stated workloads/resource definitions.
  • Source-checked: configuration fields, command forms, architecture maps and implementation-specific behavior at the linked revision.
  • To validate on the target: performance, quality, backend compatibility, a complete new installation, a custom weight build or an interface extension.

Writing these guides does not establish that every documented engine/interface combination was installed and run. For a real stack task, use the validation procedure and replace estimates with measurements in its local record.

Documentation checks covered local links/anchors, shell syntax, JSON/YAML examples and paths in the pinned source trees. The documented custom APEX generator command was also executed without model weights: it emitted 680 tensor rules, with the requested expert precision verified across all 40 example layers. This checks recipe generation, not quantization quality or model performance.

Refresh the knowledge

When a relevant component changes, inspect its release/configuration diff and the code paths linked by its guide. Update the command, code map and source revision together, then verify the affected task. Keep migration notes where users may have the older working setup, especially OpenShell's managed-route versus native-provider distinction. Do not update only a version label while retaining obsolete behavior.