Local-first LLM context optimization under strict token budgets.
ContextForge selects and optionally compresses high-value context before a downstream LLM call. It combines exact tokenizer-based limits, five comparable selection strategies, local semantic embeddings, extractive compression with provenance, and separate quality and performance benchmarks.
Python 3.11+ · Version 1.0.0 · MIT licensed
Retrieval or application data
↓
ContextForge
↓
optimized context
↓
downstream LLM
ContextForge prepares context. It does not retrieve documents or generate the downstream answer.
Quickstart · Dashboard · Architecture · Strategies · Benchmark · API · Limitations
RAG systems, coding assistants, support agents, document-QA tools, and long-running agents routinely accumulate more source material than fits economically in a prompt. Naive truncation is predictable but can discard the relevant evidence. ContextForge makes the selection step explicit, budget-safe, inspectable, and measurable.
- Five replaceable selection strategies, from a sequential baseline to hybrid ranking
- Hard token limits measured with the configured tokenizer—not character estimates
- Local sentence-transformer relevance scoring with batched embeddings
- Deterministic extractive compression with source-unit provenance
- 12-case, 120-evaluation evidence-retention benchmark
- Controlled runtime profiling with bounded in-process model reuse
- One typed engine shared by CLI, Python, HTTP API, and dashboard integrations
- 372 deterministic automated tests; real models are replaced with fakes in tests
Install the core project from a clone:
python -m venv .venv
.venv\Scripts\Activate.ps1
python -m pip install -e .Run the bundled semantic-selection demo:
contextforge optimize --input examples/semantic_context.txt --query "Why did the API become slow?" --budget 64 --chunk-size 64 --strategy semanticThe default sentence-transformers/all-MiniLM-L6-v2 model may download on first use;
inference runs locally and requires no LLM API key.
For the API and dashboard:
python -m pip install -e ".[api]"
contextforge serveOpen http://127.0.0.1:8000/ for the dashboard or http://127.0.0.1:8000/docs for generated API documentation. The server binds only to localhost by default.
The packaged dashboard turns the project's reviewed data into a one-page engineering summary:
- grouped recall, precision, and F1 comparisons for all five strategies;
- the controlled 132-to-37-token compression result;
- before/after algorithm measurements from Milestone 8; and
- the role of ContextForge between application data and a downstream LLM.
It uses server-rendered static HTML, CSS, and vanilla JavaScript—no frontend framework,
CDN, telemetry, or remote analytics. GET /dashboard-data.json returns the validated
snapshot used by the charts. The page never launches an expensive benchmark or model
download on load.
The benchmark and compression portions of the snapshot are generated through the real production runners. Regenerate them with:
python scripts/generate_dashboard_snapshot.pyThe generator refuses to overwrite the snapshot if reviewed benchmark or compression results drift. Performance comparisons are explicitly marked as recorded controlled development-environment measurements because the pre-optimization implementation is no longer the current runtime.
flowchart TD
CLI[CLI] --> ENGINE[ContextForgeEngine]
PY[Python API] --> ENGINE
HTTP[HTTP API] --> ENGINE
DASH[Dashboard] --> DATA[Validated benchmark snapshot]
DATA --> DASH
ENGINE --> TOK[Tokenizer abstraction]
TOK --> CHUNK[Boundary-aware chunker]
CHUNK --> SEM[Optional semantic analysis]
SEM --> STRAT[Selection strategy]
STRAT --> OPT[ContextOptimizer]
OPT --> LIMIT{Compression enabled?}
LIMIT -->|no| RESULT[Optimized context + metrics + provenance]
LIMIT -->|yes| COMP[Extractive compressor]
COMP --> RESULT
OPT --> EVAL[Evidence evaluator]
EVAL --> BENCH[Benchmark report]
BENCH --> DATA
The core boundaries are deliberate:
Tokenizerisolates tokenization; current production counting usestiktoken.TextChunkerpreserves source spans and chooses natural boundaries before hard splits.EmbeddingModelkeeps local transformer inference replaceable and testable.- Strategies implement one selection contract and never mutate immutable chunks.
ContextOptimizerenforces selected-token budgets independently of ranking logic.ExtractiveCompressoris an optional post-selection stage with its own exact budget.- Metrics and evidence evaluation are separate from CLI, API, and dashboard formatting.
ContextForgeEngineowns orchestration once; every application adapter reuses it.
| Strategy | Main idea | Useful when | Important tradeoff |
|---|---|---|---|
| Sequential | Keep the earliest chunks that fit | A deterministic truncation baseline | Ignores query relevance |
| Semantic | Prioritize raw query similarity | Relevant evidence may appear anywhere | Can favor long or repetitive chunks |
| Value Per Token | Rank relevance relative to token cost | Budgets are tight and chunks vary in size | Greedy heuristic, not global knapsack optimization |
| Redundancy Aware | Penalize similarity to already-selected chunks | Sources repeat or paraphrase facts | Requires quadratic novelty comparisons |
| Hybrid | Balance relevance, density, novelty, position, and recency | Mixed sources with repetition and metadata | Fixed heuristic weights are not universally optimal |
Semantic currently has the strongest aggregate evidence recall on the included curated benchmark. Hybrid is not claimed to be universally best.
Query-aware strategies embed the query and chunks in one batch. Semantic sorts by cosine
similarity. Value Per Token uses relevance / token_count. Redundancy-aware selection
iteratively discounts candidates similar to selected content. Hybrid combines fixed
normalized components:
0.40 relevance + 0.20 density + 0.20 novelty + 0.10 position + 0.10 recency
Candidates that do not fit are skipped, and selected chunks are restored to original source order. All strategy weights and tie rules are deterministic engineering heuristics; they were not learned from the benchmark.
The committed dataset contains 12 curated cases with two token budgets per case. All five strategies receive the same query, chunks, timestamps, tokenizer, embeddings, and budget:
12 cases × 2 budgets × 5 strategies = 120 strategy evaluations
Ground-truth evidence annotations are used only after selection. They never influence ranking. Aggregates are macro averages over 24 case-budget configurations.
| Strategy | Recall | Precision | F1 |
|---|---|---|---|
| Sequential | 70.83% | 60.42% | 60.83% |
| Semantic | 87.50% | 75.00% | 76.81% |
| Value Per Token | 72.92% | 62.50% | 63.61% |
| Redundancy Aware | 72.92% | 64.58% | 64.31% |
| Hybrid | 79.17% | 72.92% | 71.25% |
- Evidence recall: required evidence items fully covered by selected source spans.
- Context precision: selected chunks contributing to retained required evidence.
- Evidence F1: harmonic mean of those two measurements.
Run the benchmark or emit its complete stable JSON schema:
contextforge benchmark
contextforge benchmark --details
contextforge benchmark --format jsonThis small benchmark measures exact evidence retention, not downstream LLM answer accuracy. Results can change with the tokenizer or embedding model and should not be generalized to every workload.
Compression operates after chunk selection and has a separate hard budget. It splits selected chunks into sentence-like units, removes exact duplicates, and greedily retains relevant, token-efficient, non-redundant units. Retained text is copied from the source; ContextForge does not generate an abstractive rewrite.
Verified controlled example:
| Measurement | Result |
|---|---|
| Materialized selected context | 132 tokens |
| Compression budget | 40 tokens |
| Compressed context | 37 tokens |
| Additional tokens saved | 95 tokens |
| Compression reduction | 71.97% |
| Retained source units | 0:1, 0:5 |
Both annotated connection-pool cause and remediation facts remain in this run. This is an inspectable example, not a general quality guarantee.
contextforge optimize --input examples/compression_context.txt --query "Why did the API become slow, and what fixed it?" --budget 160 --chunk-size 160 --strategy sequential --compress --compression-budget 40The performance harness is separate from the evidence benchmark. It supports deterministic engine measurements, warm real-model measurements, cold-process runs, and optional Python allocation observations.
Selected controlled Milestone 8 comparisons:
| Workload | Before | After | Observed speedup |
|---|---|---|---|
| Redundancy selection · 500 units | 437.525 ms | 144.470 ms | 3.03× |
| Hybrid selection · 500 units | 617.631 ms | 264.119 ms | 2.34× |
| Compression pipeline · 200 units | 301.321 ms | 43.061 ms | 7.00× |
On the 200-unit compression workload, exact tokenizer calls fell from 3,568 to 313 (91.23% fewer). Improvements came from reused vector norms, priority-first exact compression fit checks, and bounded process-local transformer-model reuse.
These are controlled observations from the documented development environment, not universal production speedups. Runtime depends on hardware, versions, and workload. Cold transformer startup can dominate end-to-end latency, and the process-local cache does not improve a fresh CLI process.
Measure the current machine:
contextforge perf
contextforge perf --scale all --format json
contextforge perf --mode warm-model --scale small
contextforge perf --mode cold-cli --scale smallConsider an incident assistant receiving product notes, deployment history, gardening discussion, duplicate alerts, and a database postmortem. Given “Why did the API become slow, and what fixed it?”, ContextForge can select the connection-pool exhaustion and remediation passages, then provide that smaller source-grounded context to the assistant.
The same boundary can support a RAG pipeline, coding assistant, support agent, document-QA tool, or long-running agent. ContextForge supplies context selection and compression; it does not supply retrieval, orchestration, or answer generation for those systems.
| Command | Purpose |
|---|---|
contextforge optimize |
Select and optionally compress context under hard budgets |
contextforge benchmark |
Compare all strategies on exact evidence annotations |
contextforge perf |
Measure controlled engine, warm-model, or cold-process workloads |
contextforge serve |
Run the local development API and dashboard |
Use contextforge <command> --help for all supported options.
from contextforge import ContextForgeEngine, ContextForgeRequest
raw_context = """Release notes were approved.
Database connection-pool exhaustion made API requests wait.
Increasing the pool restored normal latency."""
engine = ContextForgeEngine()
result = engine.optimize(
ContextForgeRequest(
context=raw_context,
query="Why did the API become slow?",
budget=40,
strategy="semantic",
chunk_size=30,
)
)
print(result.optimized_context)
print(result.selected_chunk_ids)ContextForgeResult includes final text, exact token metrics, selected chunk spans, and
optional compression-unit provenance. Tokenizer and embedding factories can be injected
without changing selection algorithms.
Install the optional API dependencies and run the development server:
python -m pip install -e ".[api]"
contextforge serve --host 127.0.0.1 --port 8000Endpoints:
GET /— engineering dashboardGET /dashboard-data.json— validated local dashboard snapshotGET /health— cheap health response; does not load the transformerPOST /v1/optimize— structured optimization request/resultGET /docsandGET /openapi.json— generated FastAPI documentation
PowerShell example:
$body = @{
context = "Gardening notes.`n`nDatabase connection-pool exhaustion made API requests wait."
query = "Why did the API become slow?"
budget = 24
strategy = "semantic"
chunk_size = 20
} | ConvertTo-Json
Invoke-RestMethod `
-Method Post `
-Uri http://127.0.0.1:8000/v1/optimize `
-ContentType "application/json" `
-Body $bodyThe synchronous route delegates to ContextForgeEngine. Request chunks, analysis,
strategy, and compression state remain request-local. Loaded transformer weights can be
reused through the existing locked two-entry cache; user contexts, embeddings, and
optimization results are not globally cached.
The API processes request content in memory and does not intentionally persist it, log request bodies, or add telemetry. Uvicorn access logs still contain request metadata. This development service is not internet-hardened; deployers must add appropriate access, request-size, timeout, worker, and concurrency controls.
TextChunker prefers paragraph, sentence, then whitespace boundaries near the configured
maximum and falls back to a source-aligned hard split. Every emitted chunk is measured by
the tokenizer. Optional overlap reuses the largest source suffix that does not exceed the
overlap allowance; its full token cost is charged again when selected.
The optimizer validates selected_tokens <= selection_budget. Compression measures every
materialized proposal and validates both compressed_tokens <= compression_budget and
compressed_tokens <= pre_compression_tokens. Unicode text and exclusive source spans
remain lossless.
- Local embeddings: avoid a remote API dependency, request transmission, and API keys.
- Exact token budgets: use tokenizer counts because characters and words are unreliable proxies for model tokens.
- Extractive compression: retain source wording and provenance instead of introducing generative hallucination risk.
- Separate benchmarks: evidence retention and runtime answer different questions and should never be collapsed into one “quality” score.
- Model reuse, not result caching: reuse expensive loaded weights without retaining user requests or embeddings globally.
- One shared engine: prevent CLI, Python, HTTP, and dashboard integrations from developing inconsistent optimization pipelines.
- Small dashboard stack: packaged HTML/CSS/JavaScript keeps the demonstration portable, reviewable, and free of frontend build dependencies.
Install all development and optional API dependencies:
python -m pip install -e ".[dev]"
python -m pytest
python -m pip check
python -m buildTests cover tokenization, natural boundaries, overlap, Unicode, immutable models, all five strategies, hard budgets, compression and provenance, benchmark calculations, performance infrastructure, model-cache behavior, application orchestration, HTTP validation, OpenAPI, and dashboard data/routes. Unit tests use deterministic injected embeddings and do not download a real model.
See CONTRIBUTING.md for the complete workflow and the release checklist for final validation.
- The curated evidence dataset is small; evidence retention is not downstream answer quality.
- Semantic similarity and fixed hybrid/redundancy weights are heuristic and not learned.
- Redundancy and hybrid ranking remain O(n²d) in embedding count and dimension.
- Compression novelty work is O(u²d), with expensive exact-tokenization worst cases.
- Chunking can approach quadratic repeated-suffix tokenization in adversarial inputs.
- Local embedding model downloads and cold startup can dominate short operations.
- Performance varies with hardware, Python/model versions, and workload shape.
- The API/dashboard server is development-oriented, with no authentication, rate limiting, persistent job system, or deployment hardening.
- ContextForge includes no retrieval layer, vector database, remote LLM, or answer generator.
- Deployers must establish suitable input-size, timeout, concurrency, and resource limits.
ContextForge v1.0.0 is feature-complete for its intended portfolio scope.
- Milestone 1 — complete: Deterministic tokenization, chunking, metrics, CLI, and baseline optimizer.
- Milestone 2 — complete: Local semantic relevance scoring.
- Milestone 3 — complete: Greedy value-per-token allocation.
- Milestone 4 — complete: Redundancy-aware context selection.
- Milestone 5 — complete: Hybrid ranking with structural and recency signals.
- Milestone 6 — complete: Reproducible evidence-retention evaluation and benchmarking.
- Milestone 7 — complete: Deterministic extractive compression and provenance.
- Milestone 8 — complete: Profiling, model reuse, and algorithm hot-path optimization.
- Milestone 9 — complete: Shared Python application service and optional HTTP API.
- Milestone 10 — complete: Web-based analytics/benchmark dashboard and v1.0 release polish.
Potential work beyond v1.0 includes additional tokenizer/provider adapters, learned ranking, and larger downstream-answer evaluations. These are research ideas, not committed roadmap milestones or current capabilities.
ContextForge is available under the MIT License.