Skip to content

Repository files navigation

ContextForge

Local-first LLM context optimization under strict token budgets.

ContextForge selects and optionally compresses high-value context before a downstream LLM call. It combines exact tokenizer-based limits, five comparable selection strategies, local semantic embeddings, extractive compression with provenance, and separate quality and performance benchmarks.

Python 3.11+ · Version 1.0.0 · MIT licensed

Retrieval or application data
             ↓
       ContextForge
             ↓
       optimized context
             ↓
       downstream LLM

ContextForge prepares context. It does not retrieve documents or generate the downstream answer.

Quickstart · Dashboard · Architecture · Strategies · Benchmark · API · Limitations

Why it exists

RAG systems, coding assistants, support agents, document-QA tools, and long-running agents routinely accumulate more source material than fits economically in a prompt. Naive truncation is predictable but can discard the relevant evidence. ContextForge makes the selection step explicit, budget-safe, inspectable, and measurable.

Project highlights

  • Five replaceable selection strategies, from a sequential baseline to hybrid ranking
  • Hard token limits measured with the configured tokenizer—not character estimates
  • Local sentence-transformer relevance scoring with batched embeddings
  • Deterministic extractive compression with source-unit provenance
  • 12-case, 120-evaluation evidence-retention benchmark
  • Controlled runtime profiling with bounded in-process model reuse
  • One typed engine shared by CLI, Python, HTTP API, and dashboard integrations
  • 372 deterministic automated tests; real models are replaced with fakes in tests

Quickstart

Install the core project from a clone:

python -m venv .venv
.venv\Scripts\Activate.ps1
python -m pip install -e .

Run the bundled semantic-selection demo:

contextforge optimize --input examples/semantic_context.txt --query "Why did the API become slow?" --budget 64 --chunk-size 64 --strategy semantic

The default sentence-transformers/all-MiniLM-L6-v2 model may download on first use; inference runs locally and requires no LLM API key.

For the API and dashboard:

python -m pip install -e ".[api]"
contextforge serve

Open http://127.0.0.1:8000/ for the dashboard or http://127.0.0.1:8000/docs for generated API documentation. The server binds only to localhost by default.

Engineering dashboard

The packaged dashboard turns the project's reviewed data into a one-page engineering summary:

  • grouped recall, precision, and F1 comparisons for all five strategies;
  • the controlled 132-to-37-token compression result;
  • before/after algorithm measurements from Milestone 8; and
  • the role of ContextForge between application data and a downstream LLM.

It uses server-rendered static HTML, CSS, and vanilla JavaScript—no frontend framework, CDN, telemetry, or remote analytics. GET /dashboard-data.json returns the validated snapshot used by the charts. The page never launches an expensive benchmark or model download on load.

The benchmark and compression portions of the snapshot are generated through the real production runners. Regenerate them with:

python scripts/generate_dashboard_snapshot.py

The generator refuses to overwrite the snapshot if reviewed benchmark or compression results drift. Performance comparisons are explicitly marked as recorded controlled development-environment measurements because the pre-optimization implementation is no longer the current runtime.

Architecture

flowchart TD
    CLI[CLI] --> ENGINE[ContextForgeEngine]
    PY[Python API] --> ENGINE
    HTTP[HTTP API] --> ENGINE
    DASH[Dashboard] --> DATA[Validated benchmark snapshot]
    DATA --> DASH

    ENGINE --> TOK[Tokenizer abstraction]
    TOK --> CHUNK[Boundary-aware chunker]
    CHUNK --> SEM[Optional semantic analysis]
    SEM --> STRAT[Selection strategy]
    STRAT --> OPT[ContextOptimizer]
    OPT --> LIMIT{Compression enabled?}
    LIMIT -->|no| RESULT[Optimized context + metrics + provenance]
    LIMIT -->|yes| COMP[Extractive compressor]
    COMP --> RESULT

    OPT --> EVAL[Evidence evaluator]
    EVAL --> BENCH[Benchmark report]
    BENCH --> DATA
Loading

The core boundaries are deliberate:

  • Tokenizer isolates tokenization; current production counting uses tiktoken.
  • TextChunker preserves source spans and chooses natural boundaries before hard splits.
  • EmbeddingModel keeps local transformer inference replaceable and testable.
  • Strategies implement one selection contract and never mutate immutable chunks.
  • ContextOptimizer enforces selected-token budgets independently of ranking logic.
  • ExtractiveCompressor is an optional post-selection stage with its own exact budget.
  • Metrics and evidence evaluation are separate from CLI, API, and dashboard formatting.
  • ContextForgeEngine owns orchestration once; every application adapter reuses it.

Selection strategies

Strategy Main idea Useful when Important tradeoff
Sequential Keep the earliest chunks that fit A deterministic truncation baseline Ignores query relevance
Semantic Prioritize raw query similarity Relevant evidence may appear anywhere Can favor long or repetitive chunks
Value Per Token Rank relevance relative to token cost Budgets are tight and chunks vary in size Greedy heuristic, not global knapsack optimization
Redundancy Aware Penalize similarity to already-selected chunks Sources repeat or paraphrase facts Requires quadratic novelty comparisons
Hybrid Balance relevance, density, novelty, position, and recency Mixed sources with repetition and metadata Fixed heuristic weights are not universally optimal

Semantic currently has the strongest aggregate evidence recall on the included curated benchmark. Hybrid is not claimed to be universally best.

Query-aware strategies embed the query and chunks in one batch. Semantic sorts by cosine similarity. Value Per Token uses relevance / token_count. Redundancy-aware selection iteratively discounts candidates similar to selected content. Hybrid combines fixed normalized components:

0.40 relevance + 0.20 density + 0.20 novelty + 0.10 position + 0.10 recency

Candidates that do not fit are skipped, and selected chunks are restored to original source order. All strategy weights and tie rules are deterministic engineering heuristics; they were not learned from the benchmark.

Evidence-retention benchmark

The committed dataset contains 12 curated cases with two token budgets per case. All five strategies receive the same query, chunks, timestamps, tokenizer, embeddings, and budget:

12 cases × 2 budgets × 5 strategies = 120 strategy evaluations

Ground-truth evidence annotations are used only after selection. They never influence ranking. Aggregates are macro averages over 24 case-budget configurations.

Strategy Recall Precision F1
Sequential 70.83% 60.42% 60.83%
Semantic 87.50% 75.00% 76.81%
Value Per Token 72.92% 62.50% 63.61%
Redundancy Aware 72.92% 64.58% 64.31%
Hybrid 79.17% 72.92% 71.25%
  • Evidence recall: required evidence items fully covered by selected source spans.
  • Context precision: selected chunks contributing to retained required evidence.
  • Evidence F1: harmonic mean of those two measurements.

Run the benchmark or emit its complete stable JSON schema:

contextforge benchmark
contextforge benchmark --details
contextforge benchmark --format json

This small benchmark measures exact evidence retention, not downstream LLM answer accuracy. Results can change with the tokenizer or embedding model and should not be generalized to every workload.

Extractive compression

Compression operates after chunk selection and has a separate hard budget. It splits selected chunks into sentence-like units, removes exact duplicates, and greedily retains relevant, token-efficient, non-redundant units. Retained text is copied from the source; ContextForge does not generate an abstractive rewrite.

Verified controlled example:

Measurement Result
Materialized selected context 132 tokens
Compression budget 40 tokens
Compressed context 37 tokens
Additional tokens saved 95 tokens
Compression reduction 71.97%
Retained source units 0:1, 0:5

Both annotated connection-pool cause and remediation facts remain in this run. This is an inspectable example, not a general quality guarantee.

contextforge optimize --input examples/compression_context.txt --query "Why did the API become slow, and what fixed it?" --budget 160 --chunk-size 160 --strategy sequential --compress --compression-budget 40

Performance engineering

The performance harness is separate from the evidence benchmark. It supports deterministic engine measurements, warm real-model measurements, cold-process runs, and optional Python allocation observations.

Selected controlled Milestone 8 comparisons:

Workload Before After Observed speedup
Redundancy selection · 500 units 437.525 ms 144.470 ms 3.03×
Hybrid selection · 500 units 617.631 ms 264.119 ms 2.34×
Compression pipeline · 200 units 301.321 ms 43.061 ms 7.00×

On the 200-unit compression workload, exact tokenizer calls fell from 3,568 to 313 (91.23% fewer). Improvements came from reused vector norms, priority-first exact compression fit checks, and bounded process-local transformer-model reuse.

These are controlled observations from the documented development environment, not universal production speedups. Runtime depends on hardware, versions, and workload. Cold transformer startup can dominate end-to-end latency, and the process-local cache does not improve a fresh CLI process.

Measure the current machine:

contextforge perf
contextforge perf --scale all --format json
contextforge perf --mode warm-model --scale small
contextforge perf --mode cold-cli --scale small

Real-world placement

Consider an incident assistant receiving product notes, deployment history, gardening discussion, duplicate alerts, and a database postmortem. Given “Why did the API become slow, and what fixed it?”, ContextForge can select the connection-pool exhaustion and remediation passages, then provide that smaller source-grounded context to the assistant.

The same boundary can support a RAG pipeline, coding assistant, support agent, document-QA tool, or long-running agent. ContextForge supplies context selection and compression; it does not supply retrieval, orchestration, or answer generation for those systems.

CLI

Command Purpose
contextforge optimize Select and optionally compress context under hard budgets
contextforge benchmark Compare all strategies on exact evidence annotations
contextforge perf Measure controlled engine, warm-model, or cold-process workloads
contextforge serve Run the local development API and dashboard

Use contextforge <command> --help for all supported options.

Python API

from contextforge import ContextForgeEngine, ContextForgeRequest

raw_context = """Release notes were approved.

Database connection-pool exhaustion made API requests wait.

Increasing the pool restored normal latency."""

engine = ContextForgeEngine()
result = engine.optimize(
    ContextForgeRequest(
        context=raw_context,
        query="Why did the API become slow?",
        budget=40,
        strategy="semantic",
        chunk_size=30,
    )
)

print(result.optimized_context)
print(result.selected_chunk_ids)

ContextForgeResult includes final text, exact token metrics, selected chunk spans, and optional compression-unit provenance. Tokenizer and embedding factories can be injected without changing selection algorithms.

HTTP API

Install the optional API dependencies and run the development server:

python -m pip install -e ".[api]"
contextforge serve --host 127.0.0.1 --port 8000

Endpoints:

  • GET / — engineering dashboard
  • GET /dashboard-data.json — validated local dashboard snapshot
  • GET /health — cheap health response; does not load the transformer
  • POST /v1/optimize — structured optimization request/result
  • GET /docs and GET /openapi.json — generated FastAPI documentation

PowerShell example:

$body = @{
  context = "Gardening notes.`n`nDatabase connection-pool exhaustion made API requests wait."
  query = "Why did the API become slow?"
  budget = 24
  strategy = "semantic"
  chunk_size = 20
} | ConvertTo-Json

Invoke-RestMethod `
  -Method Post `
  -Uri http://127.0.0.1:8000/v1/optimize `
  -ContentType "application/json" `
  -Body $body

The synchronous route delegates to ContextForgeEngine. Request chunks, analysis, strategy, and compression state remain request-local. Loaded transformer weights can be reused through the existing locked two-entry cache; user contexts, embeddings, and optimization results are not globally cached.

The API processes request content in memory and does not intentionally persist it, log request bodies, or add telemetry. Uvicorn access logs still contain request metadata. This development service is not internet-hardened; deployers must add appropriate access, request-size, timeout, worker, and concurrency controls.

Token and chunk guarantees

TextChunker prefers paragraph, sentence, then whitespace boundaries near the configured maximum and falls back to a source-aligned hard split. Every emitted chunk is measured by the tokenizer. Optional overlap reuses the largest source suffix that does not exceed the overlap allowance; its full token cost is charged again when selected.

The optimizer validates selected_tokens <= selection_budget. Compression measures every materialized proposal and validates both compressed_tokens <= compression_budget and compressed_tokens <= pre_compression_tokens. Unicode text and exclusive source spans remain lossless.

Design decisions

  • Local embeddings: avoid a remote API dependency, request transmission, and API keys.
  • Exact token budgets: use tokenizer counts because characters and words are unreliable proxies for model tokens.
  • Extractive compression: retain source wording and provenance instead of introducing generative hallucination risk.
  • Separate benchmarks: evidence retention and runtime answer different questions and should never be collapsed into one “quality” score.
  • Model reuse, not result caching: reuse expensive loaded weights without retaining user requests or embeddings globally.
  • One shared engine: prevent CLI, Python, HTTP, and dashboard integrations from developing inconsistent optimization pipelines.
  • Small dashboard stack: packaged HTML/CSS/JavaScript keeps the demonstration portable, reviewable, and free of frontend build dependencies.

Development and testing

Install all development and optional API dependencies:

python -m pip install -e ".[dev]"
python -m pytest
python -m pip check
python -m build

Tests cover tokenization, natural boundaries, overlap, Unicode, immutable models, all five strategies, hard budgets, compression and provenance, benchmark calculations, performance infrastructure, model-cache behavior, application orchestration, HTTP validation, OpenAPI, and dashboard data/routes. Unit tests use deterministic injected embeddings and do not download a real model.

See CONTRIBUTING.md for the complete workflow and the release checklist for final validation.

Limitations

  • The curated evidence dataset is small; evidence retention is not downstream answer quality.
  • Semantic similarity and fixed hybrid/redundancy weights are heuristic and not learned.
  • Redundancy and hybrid ranking remain O(n²d) in embedding count and dimension.
  • Compression novelty work is O(u²d), with expensive exact-tokenization worst cases.
  • Chunking can approach quadratic repeated-suffix tokenization in adversarial inputs.
  • Local embedding model downloads and cold startup can dominate short operations.
  • Performance varies with hardware, Python/model versions, and workload shape.
  • The API/dashboard server is development-oriented, with no authentication, rate limiting, persistent job system, or deployment hardening.
  • ContextForge includes no retrieval layer, vector database, remote LLM, or answer generator.
  • Deployers must establish suitable input-size, timeout, concurrency, and resource limits.

Project roadmap

ContextForge v1.0.0 is feature-complete for its intended portfolio scope.

  1. Milestone 1 — complete: Deterministic tokenization, chunking, metrics, CLI, and baseline optimizer.
  2. Milestone 2 — complete: Local semantic relevance scoring.
  3. Milestone 3 — complete: Greedy value-per-token allocation.
  4. Milestone 4 — complete: Redundancy-aware context selection.
  5. Milestone 5 — complete: Hybrid ranking with structural and recency signals.
  6. Milestone 6 — complete: Reproducible evidence-retention evaluation and benchmarking.
  7. Milestone 7 — complete: Deterministic extractive compression and provenance.
  8. Milestone 8 — complete: Profiling, model reuse, and algorithm hot-path optimization.
  9. Milestone 9 — complete: Shared Python application service and optional HTTP API.
  10. Milestone 10 — complete: Web-based analytics/benchmark dashboard and v1.0 release polish.

Potential work beyond v1.0 includes additional tokenizer/provider adapters, learned ranking, and larger downstream-answer evaluations. These are research ideas, not committed roadmap milestones or current capabilities.

License

ContextForge is available under the MIT License.

About

Local-first LLM context optimization engine with semantic ranking, extractive compression, benchmarking,

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages