Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
67 commits
Select commit Hold shift + click to select a range
11abc34
Expand benchmark evidence and add guarded Codex OAuth campaigns
Coding-Dev-Tools Sep 16, 2026
7370e14
Fix campaign ledger IDs and validate public result envelopes
Coding-Dev-Tools Sep 16, 2026
cc25ac5
Preserve interrupted coding runs and repair benchmark fixture validity
Coding-Dev-Tools Sep 16, 2026
fefee93
Record OAuth pilot outcomes and launch frozen local follow-up measure…
Coding-Dev-Tools Sep 16, 2026
d244b8b
Finalize benchmark expansion and receipt integrity fixes
Coding-Dev-Tools Sep 17, 2026
62a7a47
fix: address PR review findings
Coding-Dev-Tools Sep 17, 2026
692efe5
fix: mark campaign digests as non-security hashes
Coding-Dev-Tools Sep 17, 2026
22ff021
fix: exclude secret fields from campaign bindings
Coding-Dev-Tools Sep 17, 2026
ceee28c
chore: document integrity digest CodeQL boundary
Coding-Dev-Tools Sep 17, 2026
192cc7e
fix: place CodeQL suppression at digest sink
Coding-Dev-Tools Sep 17, 2026
b65f8dd
fix: allowlist non-security campaign integrity digest
Coding-Dev-Tools Sep 17, 2026
180828b
fix: suppress audited campaign digest alert
Coding-Dev-Tools Sep 17, 2026
39b164a
fix: keep campaign digest compatible with CodeQL
Coding-Dev-Tools Sep 17, 2026
0989033
fix: enable exact CodeQL alert suppressions
Coding-Dev-Tools Sep 17, 2026
c8fc9aa
chore: refresh benchmark evidence and CodeQL suppression
Coding-Dev-Tools Sep 17, 2026
df41f61
fix: preserve CodeQL analysis configuration identity
Coding-Dev-Tools Sep 17, 2026
7dca246
fix: stabilize CodeQL analysis category
Coding-Dev-Tools Sep 17, 2026
db6f81f
fix: match CodeQL baseline categories
Coding-Dev-Tools Sep 17, 2026
8e9eb5e
fix: keep approved CodeQL digests out of uploaded SARIF
Coding-Dev-Tools Sep 17, 2026
23ff1c3
fix: preserve Python 3.9 write compatibility
Coding-Dev-Tools Sep 17, 2026
e456209
test: refresh offline evidence after portability fix
Coding-Dev-Tools Sep 17, 2026
9dba84f
fix: close remaining benchmark review gaps
Coding-Dev-Tools Sep 17, 2026
2b1d8fd
fix: close remaining continuation review gaps
Coding-Dev-Tools Sep 17, 2026
ce2f225
fix: close remaining review findings
Coding-Dev-Tools Sep 17, 2026
091e119
chore: refresh public evidence after review fixes
Coding-Dev-Tools Sep 17, 2026
872adec
fix: close remaining review and floor-gate gaps
Coding-Dev-Tools Sep 17, 2026
1099057
fix: close remaining peer benchmark review gaps
Coding-Dev-Tools Sep 17, 2026
658c499
fix: validate evidence actions and complete benchmark review fixes
Coding-Dev-Tools Sep 19, 2026
872ec22
chore: preserve existing line endings in evidence references
Coding-Dev-Tools Sep 19, 2026
35c17ea
fix: reject incomplete campaigns and mismatched benchmark corpora
Coding-Dev-Tools Sep 19, 2026
5837009
fix: preserve ownership of external container state directories
Coding-Dev-Tools Sep 19, 2026
b59d5c6
test: produce fresh source-bound queue analysis fixtures
Coding-Dev-Tools Sep 19, 2026
538cc71
fix: bind startup ownership marker to the managed data volume
Coding-Dev-Tools Sep 19, 2026
9fb3505
fix: initialize private state for rootless container startup
Coding-Dev-Tools Sep 19, 2026
51a613a
fix: score invalid JSON candidate results deterministically
Coding-Dev-Tools Sep 19, 2026
61e1348
fix(container): own newly created private directory ancestors
Coding-Dev-Tools Sep 19, 2026
e636f01
fix(benchmarks): bind final artifacts and repair invalid exact metadata
Coding-Dev-Tools Sep 19, 2026
c486636
fix(container): repair restored descendants despite ownership marker
Coding-Dev-Tools Sep 19, 2026
61e4807
fix(container): reject hard-linked privileged startup inputs
Coding-Dev-Tools Sep 19, 2026
5195850
fix(benchmarks): retain bound literals and validate resumed evidence
Coding-Dev-Tools Sep 19, 2026
480a4b2
fix(benchmarks): bind published results to evaluated snapshots
Coding-Dev-Tools Sep 19, 2026
0c11ae5
fix(context): preserve restrictions around multiline exact values
Coding-Dev-Tools Sep 19, 2026
2a1915c
fix(benchmarks): preserve session and audit evidence across resumes
Coding-Dev-Tools Sep 21, 2026
0244c45
fix(evidence): bind source revisions and preserve benchmark populations
Coding-Dev-Tools Sep 21, 2026
ab4b5a1
fix(memory): preserve relevant evidence through packing and correction
Coding-Dev-Tools Sep 21, 2026
65f4bec
fix(benchmarks): bind evaluation to verified input snapshots
Coding-Dev-Tools Sep 21, 2026
db30b46
fix(memory): withhold incomplete exact evidence groups
Coding-Dev-Tools Sep 21, 2026
aa88244
fix(benchmarks): validate labels and bind adapter revision
Coding-Dev-Tools Sep 21, 2026
eb75062
fix(evidence): validate writes and campaign completion
Coding-Dev-Tools Sep 21, 2026
8bdc30f
fix(provenance): preserve campaign trust and effective recall depth
Coding-Dev-Tools Sep 21, 2026
e401b50
fix: preserve correction intent and peer evidence boundaries
Coding-Dev-Tools Sep 21, 2026
c56da42
Retain observed benchmark usage through failed attempts
Coding-Dev-Tools Sep 21, 2026
9e8ffde
Preserve campaign interruptions and synchronize integrity gate location
Coding-Dev-Tools Sep 21, 2026
5490120
fix: release external checkpoint ownership after process death
Coding-Dev-Tools Sep 21, 2026
98e8f5c
fix: keep compact exact bindings with admitted evidence
Coding-Dev-Tools Sep 21, 2026
a19a807
Preserve complete restriction units and conservative budget ceilings
Coding-Dev-Tools Sep 21, 2026
f17de56
fix: bind benchmark digests to unique public source identities
Coding-Dev-Tools Sep 21, 2026
833e3e9
fix: narrow captured repair digest before artifact verification
Coding-Dev-Tools Sep 21, 2026
c3a8629
Preserve recall packing provenance and legacy binding safety
Coding-Dev-Tools Sep 21, 2026
c35d840
Require complete source for exact recall bindings
Coding-Dev-Tools Sep 21, 2026
dd74eb4
Recover local benchmark locks without replaying interrupted work
Coding-Dev-Tools Sep 21, 2026
3d7f5fb
Bind producer lock probes to the inspected file identity
Coding-Dev-Tools Sep 21, 2026
dabc9a1
Validate benchmark eligibility and repair provenance
Coding-Dev-Tools Sep 21, 2026
950c9c9
Validate benchmark rates, cell identities, and campaign locks
Coding-Dev-Tools Sep 21, 2026
95bff91
Validate public campaign metadata and observed usage
Coding-Dev-Tools Sep 21, 2026
2be5953
Preserve scope and atomic correction history in campaign replay
Coding-Dev-Tools Sep 21, 2026
655fb8e
Prevent candidate output from forging coding oracle results
Coding-Dev-Tools Sep 21, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
The table of contents is too big for display.
Diff view
Diff view
  •  
  •  
  •  
2 changes: 1 addition & 1 deletion .claude-plugin/skill-assets.sha256
Original file line number Diff line number Diff line change
Expand Up @@ -3,4 +3,4 @@ e2e08499a70d8d62ef22797cda0bb07d46cd1c6a79e5e18c1103b6c3c755b6d5 .claude-plugin
4bc8979b9ffeb97190960e551dbf4ddc6f7aeeb7b86894fd2298a59ff0001efa skills/engraphis-memory/SKILL.md
055655db84af07561d002f0c69744313d8413c39f3e873f941f0fa0b1e76dc66 skills/engraphis-memory/references/CONVENTIONS.md
62019760766ff472a76a0f81437898f39e3c1fe2631732b7b7733e50c1ad837f skills/engraphis-memory/references/SCOPING.md
33874c7c7a1c0911b0e73c7d22addc9828963d5436cb315fe7c6c5587c6b911d skills/engraphis-memory/references/TOOLS.md
9e5f1c8e91ca5697e828ab9e468504b8e1aa28e34f61307c9fcd850969126be9 skills/engraphis-memory/references/TOOLS.md
2 changes: 1 addition & 1 deletion .github/codeql/codeql-config.yml
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,6 @@
#
# Changing to SHA-256 would invalidate all existing local vectors and break
# the documented compatibility invariant in regression tests. The release SARIF
# gate waives only the two exact call sites; the CodeQL query remains enabled.
# gate waives only three exact call sites; the CodeQL query remains enabled.

name: "Engraphis CodeQL config"
4 changes: 4 additions & 0 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -52,6 +52,8 @@ jobs:
run: python -m eval.grounded
- name: Code-agent arm gate
run: python -m eval.code_arm
- name: Evidence contract boundary gate
run: python -m eval.evidence_contracts

typecheck:
name: core + backends typecheck (Python 3.11)
Expand Down Expand Up @@ -125,6 +127,8 @@ jobs:
run: python -m eval.reinforcement
- name: Adversarial memory prompt-boundary gate
run: python -m eval.adversarial_memory_security
- name: Evidence contract boundary gate
run: python -m eval.evidence_contracts
- name: Build and smoke installed core artifacts
shell: bash
run: |
Expand Down
14 changes: 14 additions & 0 deletions .github/workflows/codeql.yml
Original file line number Diff line number Diff line change
Expand Up @@ -10,6 +10,7 @@ on:

permissions:
contents: read
packages: read
security-events: write

jobs:
Expand All @@ -33,10 +34,23 @@ jobs:
languages: ${{ matrix.language }}
build-mode: none
config-file: ./.github/codeql/codeql-config.yml
packs: ${{ matrix.language == 'python' && 'codeql/python-queries:AlertSuppression.ql' || 'codeql/javascript-queries:AlertSuppression.ql' }}
- name: Analyze
id: analyze
uses: github/codeql-action/analyze@b96794f015dfd88f77b49b1c93e0fa7110f94c63 # v4
with:
output: codeql-results
upload: never
category: ".github/workflows/codeql.yml:analyze/language:${{ matrix.language }}"
- name: Require clean CodeQL results
run: python scripts/check_codeql_sarif.py "${{ steps.analyze.outputs.sarif-output }}"
- name: Filter approved non-security digest results
run: >-
python scripts/check_codeql_sarif.py --filter-approved
"${{ steps.analyze.outputs.sarif-output }}" codeql-results-filtered
- name: Upload CodeQL results
uses: github/codeql-action/upload-sarif@b96794f015dfd88f77b49b1c93e0fa7110f94c63 # v4
with:
sarif_file: codeql-results-filtered
category: ".github/workflows/codeql.yml:analyze/language:${{ matrix.language }}"
wait-for-processing: true
202 changes: 160 additions & 42 deletions BENCHMARKS.md

Large diffs are not rendered by default.

4 changes: 4 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,6 +5,10 @@ All notable changes to Engraphis are documented here. Format loosely follows

## [Unreleased]

- Receipt-chain structural corruption remains fail-closed at the Store boundary without
bricking a completed service operation: affected responses now carry a content-free
`receipt_warning`, and graph/import workers preserve their completed state.

## [1.7.4] - 2026-09-13

- Writable SQLite files now default to WAL plus FULL synchronization, with an explicit
Expand Down
6 changes: 5 additions & 1 deletion Dockerfile
Original file line number Diff line number Diff line change
Expand Up @@ -18,7 +18,9 @@ ENV PYTHONUNBUFFERED=1 \
# Customer-side cloud session and entitlement display cache. Keep it on /data rather
# than the container's ephemeral home so reconnects do not lose rotated credentials.
# License issuance, trial state, leases, and revocations remain private services.
ENGRAPHIS_STATE_DIR=/data/.engraphis
ENGRAPHIS_STATE_DIR=/data/.engraphis \
# Dashboard-managed non-secret settings must survive a Railway redeploy with the volume.
ENGRAPHIS_ENV_FILE=/data/.engraphis/config.env

WORKDIR /app

Expand All @@ -33,6 +35,8 @@ RUN apt-get update \
COPY pyproject.toml README.md LICENSE NOTICE ./
COPY engraphis ./engraphis
COPY scripts ./scripts
# The declared distribution license assets are part of the package build metadata.
COPY deploy ./deploy

# Railway runs CPU workloads. Install the CPU-only PyTorch wheel before the embedding
# stack so pip cannot select PyPI's multi-gigabyte CUDA dependency chain. The public
Expand Down
56 changes: 41 additions & 15 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -44,7 +44,7 @@ by default, or accept an explicit workspace plus optional `from_ts`, `to_ts`, an
`release_version` filters.

<p align="center">
<img src="https://raw.githubusercontent.com/Coding-Dev-Tools/engraphis/main/docs/images/context-efficiency.svg" alt="Dark chart of local measurements and deterministic fixtures, including a local LoCoMo diagnostic marked with an asterisk. Cross-session handoff satisfaction rises from 3 of 15 queries with the last memories to 15 of 15 with proactive ranking or a consolidated summary. Intent-layered graph routing rises from 0 of 3 to 3 of 3 correct top-1 targets, and two-hop graph recall rises from 0 of 3 with one-hop expansion to 3 of 3 with Personalized PageRank. Consolidation-aware ranking selects the expected digest in 2 of 2 summary cases instead of 0 of 2 for the baseline. Structure-aware chunks reduce context from 740.3 to 214.3 tokens and the smallest evidence-holding memory from 162.2 to 42.4 tokens. A compact JSON-shape proxy uses 10,982 rather than 23,810 tokens. Grounded recall makes 11 of 11 correct decisions and packed context averages 85.38 tokens under a 1,500-token cap." width="100%">
<img src="https://raw.githubusercontent.com/Coding-Dev-Tools/engraphis/main/docs/images/context-efficiency.svg" alt="Dark chart of registered deterministic fixtures. Structure-aware chunks reduce retrieved context from 740.3 to 214.3 tokens and the smallest evidence-holding memory from 162.2 to 42.4 tokens. A compact JSON-shape proxy uses 11,138 rather than 24,590 tokens. Retrieved-candidate quality is labeled separately from packed-context quality, both measured in the selected report with packed-quality fields. Actual MCP transport and provider billing are not measured." width="100%">
<br>
<sup>Less repeated history means more room for the task, tools, and useful evidence.</sup>
</p>
Expand Down Expand Up @@ -73,17 +73,27 @@ its counting boundary explicit.
|---|---|---|---|
| Retrieved top-5 memory content, averaged per question | Whole documents: **740.3** tokens → structure-aware chunks: **214.3** tokens | **526.0 fewer tokens per question** (**71.1% lower**, about **3.5× smaller**) | Recall@5 **1.000** in both modes across 6 documents and 18 questions |
| Smallest returned memory that contains the reference evidence | Whole documents: **162.2** tokens → chunks: **42.4** tokens | **119.8 fewer tokens to evidence** (**73.9% lower**, about **3.8× smaller**) | The same 18 questions had a returned evidence-holding memory in both modes |
| Full versus compact recall payload proxy across one 26-question pass within a 260-timed-recall CodeMem run | Full proxy: **23,810** `engraphis.regex.v1` tokens → compact proxy: **10,982** tokens | **12,828 proxy tokens avoided** (**53.88% lower**) | 26 payload samples; 260 timed recalls; Recall@5, hit@5, and answer-token recall all **1.000** |
| Full versus compact recall payload proxy across one 26-question pass within a 260-timed-recall CodeMem run | Full proxy: **24,590** `engraphis.regex.v1` tokens → compact proxy: **11,138** tokens | **13,452 proxy tokens avoided** (**54.71% lower**) | 26 payload samples; 260 timed recalls; Recall@5, hit@5, and answer-token recall all **1.000** |
| Packed prompt-context usage in the same 26-question CodeMem sample pass | Hard budget: **1,500** tokens; observed mean: **85.38**; observed maximum: **108** | A hard cap prevents a recall from exceeding its configured context budget | This is usage accounting, not a before/after savings comparison |

These values are evidence IDs `offline-chunking` and `offline-performance` in
[`offline-fixtures-v9.json`](https://github.com/Coding-Dev-Tools/engraphis/blob/main/docs/benchmark-evidence/offline-fixtures-v9.json),
SHA-256
`455fc9d32a236e582a49aaaf9b84f30cae2573dc6ed982f4dd7dd845afcaf24c`.
The performance report keeps its legacy `quality` fields for all candidate chunks returned before
context packing and adds `packed_quality` for evidence admitted to the reader context. The checked-in
v19 artifact includes both quality views, with Recall@5, hit@5 and answer-token evidence coverage
of 1.000 for the 26-question fixture in each view. Both views measure retrieved evidence;
neither is an end-to-end question-answer score. Coding outcomes, external datasets, and staged
operational capacity remain separate pending evaluation tracks until their artifacts are selected.

These values are evidence IDs `offline-chunking` and `offline-performance` in
[`offline-fixtures-v70.json`](https://github.com/Coding-Dev-Tools/engraphis/blob/main/docs/benchmark-evidence/offline-fixtures-v70.json),
SHA-256
`a8a96cb4d096ee72131d60307cc95b95f5b09716e813d9ee07090578dc1439c3`.
[`BENCHMARKS.md`](https://github.com/Coding-Dev-Tools/engraphis/blob/main/BENCHMARKS.md#public-numeric-evidence-registry)
records the matching suite digest, exact commands, and per-command config digests. External,
model-dependent, consolidation, productivity, and latency results remain unpublished until the
same evidence exists for them.
records the matching suite digest, exact commands, and per-command config digests. The offline
fixture registry intentionally excludes external, model-dependent, consolidation, productivity,
and latency results. Completed retrieval-only diagnostics are published separately in the
[benchmark expansion results](https://github.com/Coding-Dev-Tools/engraphis/blob/main/docs/BENCHMARK_EXPANSION_RESULTS.md) with redacted immutable
artifacts; no generated-answer, official leaderboard, hosted-latency, or paid result is claimed
here.

The compact payload shape avoids duplicating full memory bodies when the packed context and source
list are enough. The evaluator tokenizes JSON-shaped full and compact payload proxies built from
Expand Down Expand Up @@ -546,11 +556,27 @@ For an agent prompt, prefer `engraphis_recall_context`: it returns one hard-budg
`source_tokens`, `saved_tokens`, `savings_ratio`, `packed_count`, `omitted_count`, and
`token_counter`), and optional diagnostics. Accounting is exact for the named counter; inject the
reader's tokenizer when reader-model token parity is required. `engraphis_recall` remains the compatible full-recall
surface; use `response_mode="compact"` when the packed context is enough and full memory bodies
would duplicate it. For advanced query-planning configuration, see the
[architecture guide](https://github.com/Coding-Dev-Tools/engraphis/blob/main/docs/ARCHITECTURE_V3.md#query-planning).

For bi-temporal reads, `valid_at` selects what was true at a Unix timestamp and `known_at` selects
surface; use `response_mode="compact"` when the packed context is enough and full memory bodies
would duplicate it. For advanced query-planning configuration, see the
[architecture guide](https://github.com/Coding-Dev-Tools/engraphis/blob/main/docs/ARCHITECTURE_V3.md#query-planning).

Benchmark-driven alternatives are opt-in: `packing_mode="coverage"` keeps complete evidence
units from more source memories, while `retrieval_recipe="conversation"` and
`retrieval_recipe="long_session"` select the measured depth/budget starting points. The
historical `legacy`/`default` settings remain unchanged. For a value that must survive a file
edit or tool call exactly, Smart and Classic `engraphis_remember` and the Python/service write
APIs accept source-bound `exact_value` plus its `exact_value_type`. MCP remember requires a
unique occurrence; Python/service writes can select a repeated occurrence with `exact_value_span`.
Packed binding metadata requires the complete memory source, preserving conditions in any
language. Boundary whitespace outside the bound value may be trimmed. Coverage withholds a
bound group that cannot fit; legacy keeps its selected text but omits the incomplete binding.
Corrections and content revisions clear the old binding when content changes;
pass `exact_value` to explicitly bind the replacement, with `exact_value_span=[start,end]`
for a repeated occurrence, or `clear_exact_value=true` to remove a binding. Unchanged content
and title-only revisions preserve valid bindings. History preserves the original record.
MCP response trimming removes binding metadata whenever its supporting context is omitted.

For bi-temporal reads, `valid_at` selects what was true at a Unix timestamp and `known_at` selects
what Engraphis had learned then. `as_of` remains a compatibility alias for `valid_at`; supplying
both is allowed only when they match.

Expand Down Expand Up @@ -762,7 +788,7 @@ file. It never searches the working directory for `.env`, and explicit process v
|---------|---------|-------------|
| `ENGRAPHIS_ENV_FILE` | `~/.engraphis/config.env` | Optional trusted config leaf selected before trusted values load. Its bounded dependency-free parser performs no interpolation. An explicit value must be an absolute path to an owner-private regular file; arbitrary working-directory `.env` files are ignored. |
| `ENGRAPHIS_DB_PATH` | Source: `<repo>/engraphis.db`; installed: platform user-data directory | SQLite database file. Installed defaults are `%LOCALAPPDATA%\engraphis\engraphis.db` (Windows), `~/Library/Application Support/engraphis/engraphis.db` (macOS), and `$XDG_DATA_HOME/engraphis/engraphis.db` or `~/.local/share/engraphis/engraphis.db` (Linux). The environment variable overrides every default; a relative value is resolved from the trusted `~/.engraphis/config.env` directory so launch CWD cannot select a different workspace database. |
| `ENGRAPHIS_SQLITE_DURABILITY` | `durable` | Writable file databases use WAL and FULL commit synchronization. Explicit `balanced` selects NORMAL, which can lose recent acknowledged writes after OS/power failure. Effective settings appear in diagnostics; see [SQLite durability](https://github.com/Coding-Dev-Tools/engraphis/blob/main/docs/SQLITE_DURABILITY.md). |
| `ENGRAPHIS_SQLITE_DURABILITY` | `durable` | Writable file databases use WAL and FULL commit synchronization. Explicit `balanced` selects NORMAL, which can lose recent acknowledged writes after OS/power failure. Effective settings appear in diagnostics; see [SQLite durability](https://github.com/Coding-Dev-Tools/engraphis/blob/main/docs/SQLITE_DURABILITY.md). |
| `ENGRAPHIS_HOST` | `127.0.0.1` | Server bind address |
| `ENGRAPHIS_PORT` | `8700` | Dashboard port. A platform-injected `$PORT` (Railway/Fly/Heroku) takes precedence over this value for the dashboard bind; Compose pins both to `ENGRAPHIS_COMPOSE_PORT` so the mapping stays in sync |
| `ENGRAPHIS_SERVICE_MODE` | `customer` | The public package supports only `customer`; hosted vendor, relay, compute, and worker roles are not distributed here |
Expand Down
9 changes: 9 additions & 0 deletions deploy/railway-template.json
Original file line number Diff line number Diff line change
Expand Up @@ -15,6 +15,11 @@
}
},
"variables": {
"ENGRAPHIS_HOST": {
"value": "0.0.0.0",
"prompt": "Bind the public Railway service on all IPv4 interfaces so the platform PORT and health probe are reachable.",
"required": true
},
"ENGRAPHIS_SERVICE_MODE": {
"value": "customer",
"required": true
Expand All @@ -27,6 +32,10 @@
"value": "/data/.engraphis",
"required": true
},
"ENGRAPHIS_ENV_FILE": {
"value": "/data/.engraphis/config.env",
"required": true
},
"ENGRAPHIS_API_TOKEN": {
"value": "${{ secret(48) }}",
"secret": true,
Expand Down
Loading
Loading