diff --git a/docs/decisions/implemented/2026-09-30-lexical-tie-ranking.md b/docs/decisions/implemented/2026-09-30-lexical-tie-ranking.md new file mode 100644 index 00000000..68725834 --- /dev/null +++ b/docs/decisions/implemented/2026-09-30-lexical-tie-ranking.md @@ -0,0 +1,39 @@ +# Resolve lexical score ties after evidence selection + +[中文](2026-09-30-lexical-tie-ranking.zh-CN.md) + +**Status:** implemented +**Approved:** explicit +**Relates to:** [independent relevance training](2026-09-15-independent-relevance-training.md) + +## Problem + +Equal first-stage lexical scores leave some selected evidence in a weak order. +The tested small neural ranker did not demonstrate a material, reliable gain over +product ranking on real recall traces. The measured outcomes are in +[the retrieval experiment](../../experiments/retrieval-quality/lexical-tie-ranking-2026-09-30.md). + +## Decision + +The FTS5-only path reorders adjacent, exactly equal-`combinedScore` records after +the Active Graph budget has selected them. The tie-break uses IDF-weighted coverage +of the query's word terms over a selected statement and its first 500 evidence +characters. It preserves the selected set, non-tied order, QPP decisions, and +hybrid retrieval. The small neural ranker is a failed experiment for this task; +no neural ranker or model artifact is added by this change. + +## Alternatives considered + +- Add the tested MLP to product ranking: rejected because its same-candidate + result was weaker than the existing ranker, and its top-tie gain was not + statistically supported. +- Reorder all candidates or alter candidate generation: outside the measured + rule and would change the evidence budget or recall set. +- Keep existing lexical tie order: loses the measured LoCoMo ranking gain. + +## Consequences + +The rule uses transient CPU work and no model weights. It improves LoCoMo ranking +on the pinned sample while LongMemEval and BEAM have small regressions; R@20 and +candidate membership do not improve. It is not evidence of better semantic +candidate recovery or answer quality. diff --git a/docs/decisions/implemented/2026-09-30-lexical-tie-ranking.zh-CN.md b/docs/decisions/implemented/2026-09-30-lexical-tie-ranking.zh-CN.md new file mode 100644 index 00000000..35c1cb1f --- /dev/null +++ b/docs/decisions/implemented/2026-09-30-lexical-tie-ranking.zh-CN.md @@ -0,0 +1,34 @@ +# 在选出证据后处理词法分数并列 + +[English](2026-09-30-lexical-tie-ranking.md) + +**Status:** implemented +**Approved:** explicit +**Relates to:** [独立相关性训练](2026-09-15-independent-relevance-training.zh-CN.md) + +## Problem + +首阶段词法分数相同的证据可能以较弱的顺序进入上下文。试过的小型神经网络排序器, +在真实召回轨迹上没有证明比产品现有排序有实质且可靠的收益。具体结果见 +[检索实验](../../experiments/retrieval-quality/lexical-tie-ranking-2026-09-30.md)。 + +## Decision + +仅在 FTS5 词法路径中,于 Active Graph 预算选出证据后,对相邻且 +`combinedScore` 完全相同的记录排序。并列判定使用查询词在记忆陈述及前 500 +个证据字符中的 IDF 加权覆盖度。候选集合、非并列顺序、QPP 决策和混合检索 +保持原样。本任务的小型神经网络排序实验失败;此变更不加入神经网络排序器 +或模型产物。 + +## Alternatives considered + +- 把试过的 MLP 接入产品:同一候选集的效果弱于现有排序;限制在顶部并列组 + 后的小幅收益也没有统计证据支持,因此拒绝。 +- 重排所有候选或修改候选生成:超出了已测规则,也会改变证据预算或召回集合。 +- 保留原词法并列顺序:会失去 LoCoMo 样本上测得的排序收益。 + +## Consequences + +规则消耗少量临时 CPU 时间,不存储模型权重。固定样本上的 LoCoMo 排序改善, +LongMemEval 和 BEAM 小幅回退;R@20 和候选集合没有改善。这不能证明语义 +候选召回或最终回答质量提高。 diff --git a/docs/design/design.md b/docs/design/design.md index b9a613bf..dab61c1d 100644 --- a/docs/design/design.md +++ b/docs/design/design.md @@ -1855,6 +1855,13 @@ Search signals are separated by purpose: - regular expression only as an advanced/debug fallback over a bounded candidate set or raw session subset. +On the FTS5-only path, after the Active Graph budget selects the ranked evidence, +adjacent candidates with exactly equal `combinedScore` are ordered by IDF-weighted +coverage of the query's word terms in their statements and bounded evidence +excerpts. The rule does not change candidate membership, QPP expansion, or hybrid +retrieval. Its evaluation and the rejected neural ranking arm are recorded in +[lexical tie ranking](../decisions/implemented/2026-09-30-lexical-tie-ranking.md). + Arbitrary model-generated regex is not a relevance ranker and must not scan the entire store by default. Surface anchors are admitted beside semantic candidates before the shared AG budget and projection. Plain prose produces no surface diff --git a/docs/experiments/retrieval-quality/lexical-tie-ranking-2026-09-30.md b/docs/experiments/retrieval-quality/lexical-tie-ranking-2026-09-30.md new file mode 100644 index 00000000..d3151b89 --- /dev/null +++ b/docs/experiments/retrieval-quality/lexical-tie-ranking-2026-09-30.md @@ -0,0 +1,57 @@ +# Lexical tie ranking and neural ranker outcome + +## Question and protocol + +Can a CPU-only, model-free rule improve the order of evidence already selected by +NMG, under the 100 MB additional resident-memory budget? The comparison calls +the product search once per question and scores the same returned candidates +before and after reordering adjacent records with exactly equal `combinedScore`. +The rule uses `Intl.Segmenter` word terms (NFKC, lowercase, at least two Unicode +code points) and IDF-weighted query-term coverage over each statement plus the +first 500 evidence characters. Query-term document frequency is computed over +the selected ranked prefix. Unranked chain and block supplements are untouched. + +The pinned samples contain LoCoMo 1,540 questions from 10 users, LongMemEval +100, BEAM 400, PersonaMem 500, and HaluMem 352. The same Active Graph and appended +budgets are used in both arms. Official gold evidence labels are used only by +the retrieval scorer, never for fitting or selecting this rule. The rule was fixed +before the held-out comparison. The paired unit for the exact sign-flip check is +a user, not a question; a p-value from only two users has little resolution. + +## Held-out lexical results + +| Dataset | R@1 before → after | R@5 before → after | MRR(Q) before → after | Paired MRR direction | +| --- | ---: | ---: | ---: | --- | +| LoCoMo | 7.71% → 17.42% | 14.86% → 29.09% | 0.1830 → 0.3364 | 441 queries improved, 70 worsened; 10/10 users improved; p=0.00195 | +| LongMemEval | 50.59% → 50.00% | 66.47% → 66.47% | 0.8823 → 0.8800 | 3 improved, 4 worsened | +| BEAM | 2.76% → 2.76% | 6.60% → 6.48% | 0.1352 → 0.1338 | 7 improved, 5 worsened | +| PersonaMem | 1.84% → 2.24% | 8.76% → 8.63% | 0.1246 → 0.1326 | 25 improved, 14 worsened; unadjusted p=0.017 | +| HaluMem | 0.00% → 4.23% | 5.04% → 9.27% | 0.0285 → 0.1012 | 38 improved, 2 worsened; only 2 users, p=0.5 | + +R@20 is unchanged on every sample because the rule only permutes selected +evidence. The offline rerank averaged roughly 0.13–2.0 ms per query across the +five samples and requires no model weights. The pre-integration paired report and +the post-integration five-dataset report are in the ignored local evaluation +workspace at `evals/results/retrieval/program-ranking-probe/`. After integration, +all scoring fields in the product lexical arm match the pre-integration candidate +arm; a cached-embedding LongMemEval control retained its original hybrid metrics. + +These are ranking metrics, not an answer-quality or candidate-recovery result. +Absolute lexical recall remains low on BEAM, PersonaMem, and HaluMem. The small +LongMemEval and BEAM regressions are real paired outcomes of this fixed rule. + +## Neural ranking outcome + +A small MLP trained on multilingual SemRel with 17 lexical-overlap features did +not meet the adoption bar on NMG's independent real-recall traces: 19 queries, +43 explicitly useful memories, and no verified negative labels. On the same +candidate sets, the existing ranker had Hit@1 16/19 and MRR(Q) 0.921; the MLP +had MRR(Q) 0.861. Restricting the MLP to the highest equal-usefulness group +gave Hit@1 17/19 and MRR(Q) 0.947, but the paired result was not significant +(exact p=1.0). The available labels also cannot establish precision on irrelevant +memories. Changing a seed is not evidence that the route succeeds. + +The neural ranker therefore failed this task's requirement of a demonstrated, +material gain over product ranking. No neural model is included in the lexical +tie-ranking change. This conclusion applies to the tested ranking task and data; +it does not turn untested semantic candidate recovery into a measured result. diff --git a/evals/retrieval/README.md b/evals/retrieval/README.md index f3140852..922bfbed 100644 --- a/evals/retrieval/README.md +++ b/evals/retrieval/README.md @@ -164,7 +164,10 @@ Two external services, both configured by env (no models ship with NMG): ## Results -Formal run results are recorded as dated documents under `docs/`, not here: +Formal run results are recorded as dated experiments under `docs/`: + +- [Lexical tie ranking and neural ranker outcome](../../docs/experiments/retrieval-quality/lexical-tie-ranking-2026-09-30.md) + — paired five-dataset lexical evaluation and the failed local neural ranking arm. - [Retrieval-quality baseline 2026-08-16](../../docs/experiments/retrieval-quality-baseline-2026-08-16.md) — first pinned run, lexical arm, the original three-dataset protocol. diff --git a/src/core/store/retrieval.ts b/src/core/store/retrieval.ts index 34877364..26cd9fb2 100644 --- a/src/core/store/retrieval.ts +++ b/src/core/store/retrieval.ts @@ -46,6 +46,7 @@ import { queryOverlapTerms, recallHitTerms, recallReason, + rerankEqualScoresByQueryCoverage, termOverlapScore, type StoreRow as Row, } from "./search-ranking.ts"; @@ -177,6 +178,21 @@ function markDuplicateResults(results: MemorySearchResult[]): void { } } +function rankSelectedEvidence( + query: string, + results: MemorySearchResult[], + retrievalMode: SearchOptions["retrievalMode"], + hasSemanticQuery: boolean, +): MemorySearchResult[] { + if (retrievalMode !== "fts5" || hasSemanticQuery) return results; + return rerankEqualScoresByQueryCoverage( + query, + results, + (result) => result.combinedScore, + (result) => `${result.memory.statement} ${result.evidence.content.slice(0, 500)}`, + ); +} + export function withRetrieval(Base: TBase) { return class extends Base { // Base-class and cross-cluster members (resolved at assembly time) @@ -573,7 +589,16 @@ export function withRetrieval(Base: TBase) { }; perf?.stop(SECTION.secondPass); } - const { results, selectedNodes, estimatedTokens, exhausted } = selection; + const { selectedNodes, estimatedTokens, exhausted } = selection; + // Reorder the evidence already selected by the budget, so lexical ties + // cannot change candidate membership or the second-pass decision. + const results = rankSelectedEvidence( + query, + selection.results, + options.retrievalMode, + semantic !== undefined, + ); + selections = buildSelections(results); const projectedEdges = new Map( edgeProjection.edges.map((edge) => [edge.relationId, edge] as const), ); diff --git a/src/core/store/search-ranking.ts b/src/core/store/search-ranking.ts index a6ec5571..4484fbd5 100644 --- a/src/core/store/search-ranking.ts +++ b/src/core/store/search-ranking.ts @@ -3,6 +3,61 @@ import type { MemoryNode, MemorySearchResult, MemoryType, RecallCue } from "../t export type StoreRow = Record; +const wordSegmenter = new Intl.Segmenter(undefined, { granularity: "word" }); + +function wordTerms(text: string): Set { + const terms = new Set(); + for (const part of wordSegmenter.segment(text.normalize("NFKC").toLowerCase())) { + if (part.isWordLike && Array.from(part.segment).length >= 2) terms.add(part.segment); + } + return terms; +} + +/** Reorder only adjacent equal-score candidates, leaving the selected set intact. */ +export function rerankEqualScoresByQueryCoverage( + query: string, + candidates: readonly T[], + scoreOf: (candidate: T) => number, + textOf: (candidate: T) => string, +): T[] { + if ( + candidates.length < 2 || + !candidates.some( + (candidate, index) => index > 0 && scoreOf(candidate) === scoreOf(candidates[index - 1]!), + ) + ) { + return [...candidates]; + } + const queryTerms = wordTerms(query); + if (queryTerms.size === 0) return [...candidates]; + const candidateTerms = candidates.map((candidate) => wordTerms(textOf(candidate))); + const weights = [...queryTerms].map((term) => { + const frequency = candidateTerms.filter((terms) => terms.has(term)).length; + return { + term, + weight: Math.log(1 + (candidates.length - frequency + 0.5) / (frequency + 0.5)), + }; + }); + const coverage = candidateTerms.map((terms) => + weights.reduce((sum, { term, weight }) => sum + (terms.has(term) ? weight : 0), 0), + ); + const reordered: T[] = []; + for (let start = 0; start < candidates.length;) { + let end = start + 1; + while (end < candidates.length && scoreOf(candidates[end]!) === scoreOf(candidates[start]!)) + end++; + const tied = candidates + .slice(start, end) + .map((candidate, offset) => ({ candidate, index: start + offset })); + tied.sort( + (left, right) => coverage[right.index]! - coverage[left.index]! || left.index - right.index, + ); + reordered.push(...tied.map(({ candidate }) => candidate)); + start = end; + } + return reordered; +} + export function contextUsefulness(query: string, result: MemorySearchResult): number { const normalized = normalize(query); const type = result.memory.memoryType; diff --git a/tests/core/store/retrieval.test.ts b/tests/core/store/retrieval.test.ts index 07b937f3..ad9e7fc4 100644 --- a/tests/core/store/retrieval.test.ts +++ b/tests/core/store/retrieval.test.ts @@ -48,6 +48,40 @@ test("searchContext returns results and relations for a lexical query", () => { }); }); +test("lexical tie-break keeps Active Graph selection order aligned with returned evidence", () => { + withStore((store) => { + const decoy = store.remember({ + statement: "The store uses WAL", + nodeName: "Atlas SQLite", + importance: 1, + }); + const target = store.remember({ + statement: "Atlas SQLite stores project data", + nodeName: "Database choice", + importance: 0.5, + }); + const query = "Atlas SQLite"; + const context = store.searchContext(query, { + retrievalMode: "fts5", + limit: 2, + }); + assert.ok(context.activeGraph); + assert.equal(context.results[0]?.memory.id, target.memory.id); + assert.equal(context.results[1]?.memory.id, decoy.memory.id); + assert.deepEqual( + context.activeGraph.memoryIds, + context.results.map((result) => result.memory.id), + ); + assert.deepEqual( + context.activeGraph.selections.map((selection) => selection.memoryId), + context.activeGraph.memoryIds, + ); + const trace = store.retrievalTrace(context.activeGraph.id); + assert.deepEqual(trace?.resultMemoryIds, context.activeGraph.memoryIds); + assert.deepEqual(trace?.selections?.map((selection) => selection.memoryId), context.activeGraph.memoryIds); + }); +}); + test("relevanceFloor gates disclosed results and may abstain", () => { withStore((store) => { store.remember({ diff --git a/tests/core/store/search-ranking.test.ts b/tests/core/store/search-ranking.test.ts new file mode 100644 index 00000000..b1f5b7c9 --- /dev/null +++ b/tests/core/store/search-ranking.test.ts @@ -0,0 +1,22 @@ +import assert from "node:assert/strict"; +import test from "node:test"; + +import { rerankEqualScoresByQueryCoverage } from "../../../src/core/store/search-ranking.ts"; + +test("query coverage resolves lexical score ties without crossing score boundaries", () => { + const candidates = [ + { id: "partial", score: 4, text: "Atlas project notes" }, + { id: "complete", score: 4, text: "Atlas project uses SQLite" }, + { id: "lower", score: 2, text: "Atlas project uses SQLite" }, + ]; + const rank = (items: typeof candidates) => rerankEqualScoresByQueryCoverage( + "Atlas project SQLite", + items, + (item) => item.score, + (item) => item.text, + ); + const ranked = rank(candidates); + assert.deepEqual(ranked.map((item) => item.id), ["complete", "partial", "lower"]); + assert.deepEqual(rank(ranked), ranked, "reranking is stable on repeated calls"); + assert.deepEqual(candidates.map((item) => item.id), ["partial", "complete", "lower"]); +});