Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
39 changes: 39 additions & 0 deletions docs/decisions/implemented/2026-09-30-lexical-tie-ranking.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,39 @@
# Resolve lexical score ties after evidence selection

[中文](2026-09-30-lexical-tie-ranking.zh-CN.md)

**Status:** implemented
**Approved:** explicit
**Relates to:** [independent relevance training](2026-09-15-independent-relevance-training.md)

## Problem

Equal first-stage lexical scores leave some selected evidence in a weak order.
The tested small neural ranker did not demonstrate a material, reliable gain over
product ranking on real recall traces. The measured outcomes are in
[the retrieval experiment](../../experiments/retrieval-quality/lexical-tie-ranking-2026-09-30.md).

## Decision

The FTS5-only path reorders adjacent, exactly equal-`combinedScore` records after
the Active Graph budget has selected them. The tie-break uses IDF-weighted coverage
of the query's word terms over a selected statement and its first 500 evidence
characters. It preserves the selected set, non-tied order, QPP decisions, and
hybrid retrieval. The small neural ranker is a failed experiment for this task;
no neural ranker or model artifact is added by this change.

## Alternatives considered

- Add the tested MLP to product ranking: rejected because its same-candidate
result was weaker than the existing ranker, and its top-tie gain was not
statistically supported.
- Reorder all candidates or alter candidate generation: outside the measured
rule and would change the evidence budget or recall set.
- Keep existing lexical tie order: loses the measured LoCoMo ranking gain.

## Consequences

The rule uses transient CPU work and no model weights. It improves LoCoMo ranking
on the pinned sample while LongMemEval and BEAM have small regressions; R@20 and
candidate membership do not improve. It is not evidence of better semantic
candidate recovery or answer quality.
34 changes: 34 additions & 0 deletions docs/decisions/implemented/2026-09-30-lexical-tie-ranking.zh-CN.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,34 @@
# 在选出证据后处理词法分数并列

[English](2026-09-30-lexical-tie-ranking.md)

**Status:** implemented
**Approved:** explicit
**Relates to:** [独立相关性训练](2026-09-15-independent-relevance-training.zh-CN.md)

## Problem

首阶段词法分数相同的证据可能以较弱的顺序进入上下文。试过的小型神经网络排序器,
在真实召回轨迹上没有证明比产品现有排序有实质且可靠的收益。具体结果见
[检索实验](../../experiments/retrieval-quality/lexical-tie-ranking-2026-09-30.md)。

## Decision

仅在 FTS5 词法路径中,于 Active Graph 预算选出证据后,对相邻且
`combinedScore` 完全相同的记录排序。并列判定使用查询词在记忆陈述及前 500
个证据字符中的 IDF 加权覆盖度。候选集合、非并列顺序、QPP 决策和混合检索
保持原样。本任务的小型神经网络排序实验失败;此变更不加入神经网络排序器
或模型产物。

## Alternatives considered

- 把试过的 MLP 接入产品:同一候选集的效果弱于现有排序;限制在顶部并列组
后的小幅收益也没有统计证据支持,因此拒绝。
- 重排所有候选或修改候选生成:超出了已测规则,也会改变证据预算或召回集合。
- 保留原词法并列顺序:会失去 LoCoMo 样本上测得的排序收益。

## Consequences

规则消耗少量临时 CPU 时间,不存储模型权重。固定样本上的 LoCoMo 排序改善,
LongMemEval 和 BEAM 小幅回退;R@20 和候选集合没有改善。这不能证明语义
候选召回或最终回答质量提高。
7 changes: 7 additions & 0 deletions docs/design/design.md
Original file line number Diff line number Diff line change
Expand Up @@ -1855,6 +1855,13 @@ Search signals are separated by purpose:
- regular expression only as an advanced/debug fallback over a bounded candidate
set or raw session subset.

On the FTS5-only path, after the Active Graph budget selects the ranked evidence,
adjacent candidates with exactly equal `combinedScore` are ordered by IDF-weighted
coverage of the query's word terms in their statements and bounded evidence
excerpts. The rule does not change candidate membership, QPP expansion, or hybrid
retrieval. Its evaluation and the rejected neural ranking arm are recorded in
[lexical tie ranking](../decisions/implemented/2026-09-30-lexical-tie-ranking.md).

Arbitrary model-generated regex is not a relevance ranker and must not scan the
entire store by default. Surface anchors are admitted beside semantic candidates
before the shared AG budget and projection. Plain prose produces no surface
Expand Down
Original file line number Diff line number Diff line change
@@ -0,0 +1,57 @@
# Lexical tie ranking and neural ranker outcome

## Question and protocol

Can a CPU-only, model-free rule improve the order of evidence already selected by
NMG, under the 100 MB additional resident-memory budget? The comparison calls
the product search once per question and scores the same returned candidates
before and after reordering adjacent records with exactly equal `combinedScore`.
The rule uses `Intl.Segmenter` word terms (NFKC, lowercase, at least two Unicode
code points) and IDF-weighted query-term coverage over each statement plus the
first 500 evidence characters. Query-term document frequency is computed over
the selected ranked prefix. Unranked chain and block supplements are untouched.

The pinned samples contain LoCoMo 1,540 questions from 10 users, LongMemEval
100, BEAM 400, PersonaMem 500, and HaluMem 352. The same Active Graph and appended
budgets are used in both arms. Official gold evidence labels are used only by
the retrieval scorer, never for fitting or selecting this rule. The rule was fixed
before the held-out comparison. The paired unit for the exact sign-flip check is
a user, not a question; a p-value from only two users has little resolution.

## Held-out lexical results

| Dataset | R@1 before → after | R@5 before → after | MRR(Q) before → after | Paired MRR direction |
| --- | ---: | ---: | ---: | --- |
| LoCoMo | 7.71% → 17.42% | 14.86% → 29.09% | 0.1830 → 0.3364 | 441 queries improved, 70 worsened; 10/10 users improved; p=0.00195 |
| LongMemEval | 50.59% → 50.00% | 66.47% → 66.47% | 0.8823 → 0.8800 | 3 improved, 4 worsened |
| BEAM | 2.76% → 2.76% | 6.60% → 6.48% | 0.1352 → 0.1338 | 7 improved, 5 worsened |
| PersonaMem | 1.84% → 2.24% | 8.76% → 8.63% | 0.1246 → 0.1326 | 25 improved, 14 worsened; unadjusted p=0.017 |
| HaluMem | 0.00% → 4.23% | 5.04% → 9.27% | 0.0285 → 0.1012 | 38 improved, 2 worsened; only 2 users, p=0.5 |

R@20 is unchanged on every sample because the rule only permutes selected
evidence. The offline rerank averaged roughly 0.13–2.0 ms per query across the
five samples and requires no model weights. The pre-integration paired report and
the post-integration five-dataset report are in the ignored local evaluation
workspace at `evals/results/retrieval/program-ranking-probe/`. After integration,
all scoring fields in the product lexical arm match the pre-integration candidate
arm; a cached-embedding LongMemEval control retained its original hybrid metrics.

These are ranking metrics, not an answer-quality or candidate-recovery result.
Absolute lexical recall remains low on BEAM, PersonaMem, and HaluMem. The small
LongMemEval and BEAM regressions are real paired outcomes of this fixed rule.

## Neural ranking outcome

A small MLP trained on multilingual SemRel with 17 lexical-overlap features did
not meet the adoption bar on NMG's independent real-recall traces: 19 queries,
43 explicitly useful memories, and no verified negative labels. On the same
candidate sets, the existing ranker had Hit@1 16/19 and MRR(Q) 0.921; the MLP
had MRR(Q) 0.861. Restricting the MLP to the highest equal-usefulness group
gave Hit@1 17/19 and MRR(Q) 0.947, but the paired result was not significant
(exact p=1.0). The available labels also cannot establish precision on irrelevant
memories. Changing a seed is not evidence that the route succeeds.

The neural ranker therefore failed this task's requirement of a demonstrated,
material gain over product ranking. No neural model is included in the lexical
tie-ranking change. This conclusion applies to the tested ranking task and data;
it does not turn untested semantic candidate recovery into a measured result.
5 changes: 4 additions & 1 deletion evals/retrieval/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -164,7 +164,10 @@ Two external services, both configured by env (no models ship with NMG):

## Results

Formal run results are recorded as dated documents under `docs/`, not here:
Formal run results are recorded as dated experiments under `docs/`:

- [Lexical tie ranking and neural ranker outcome](../../docs/experiments/retrieval-quality/lexical-tie-ranking-2026-09-30.md)
— paired five-dataset lexical evaluation and the failed local neural ranking arm.

- [Retrieval-quality baseline 2026-08-16](../../docs/experiments/retrieval-quality-baseline-2026-08-16.md)
— first pinned run, lexical arm, the original three-dataset protocol.
Expand Down
27 changes: 26 additions & 1 deletion src/core/store/retrieval.ts
Original file line number Diff line number Diff line change
Expand Up @@ -46,6 +46,7 @@ import {
queryOverlapTerms,
recallHitTerms,
recallReason,
rerankEqualScoresByQueryCoverage,
termOverlapScore,
type StoreRow as Row,
} from "./search-ranking.ts";
Expand Down Expand Up @@ -177,6 +178,21 @@ function markDuplicateResults(results: MemorySearchResult[]): void {
}
}

function rankSelectedEvidence(
query: string,
results: MemorySearchResult[],
retrievalMode: SearchOptions["retrievalMode"],
hasSemanticQuery: boolean,
): MemorySearchResult[] {
if (retrievalMode !== "fts5" || hasSemanticQuery) return results;
return rerankEqualScoresByQueryCoverage(
query,
results,
(result) => result.combinedScore,
(result) => `${result.memory.statement} ${result.evidence.content.slice(0, 500)}`,
);
}

export function withRetrieval<TBase extends Constructor>(Base: TBase) {
return class extends Base {
// Base-class and cross-cluster members (resolved at assembly time)
Expand Down Expand Up @@ -573,7 +589,16 @@ export function withRetrieval<TBase extends Constructor>(Base: TBase) {
};
perf?.stop(SECTION.secondPass);
}
const { results, selectedNodes, estimatedTokens, exhausted } = selection;
const { selectedNodes, estimatedTokens, exhausted } = selection;
// Reorder the evidence already selected by the budget, so lexical ties
// cannot change candidate membership or the second-pass decision.
const results = rankSelectedEvidence(
query,
selection.results,
options.retrievalMode,
semantic !== undefined,
);
selections = buildSelections(results);
const projectedEdges = new Map(
edgeProjection.edges.map((edge) => [edge.relationId, edge] as const),
);
Expand Down
55 changes: 55 additions & 0 deletions src/core/store/search-ranking.ts
Original file line number Diff line number Diff line change
Expand Up @@ -3,6 +3,61 @@ import type { MemoryNode, MemorySearchResult, MemoryType, RecallCue } from "../t

export type StoreRow = Record<string, string | number | Uint8Array | null>;

const wordSegmenter = new Intl.Segmenter(undefined, { granularity: "word" });

function wordTerms(text: string): Set<string> {
const terms = new Set<string>();
for (const part of wordSegmenter.segment(text.normalize("NFKC").toLowerCase())) {
if (part.isWordLike && Array.from(part.segment).length >= 2) terms.add(part.segment);
}
return terms;
}

/** Reorder only adjacent equal-score candidates, leaving the selected set intact. */
export function rerankEqualScoresByQueryCoverage<T>(
query: string,
candidates: readonly T[],
scoreOf: (candidate: T) => number,
textOf: (candidate: T) => string,
): T[] {
if (
candidates.length < 2 ||
!candidates.some(
(candidate, index) => index > 0 && scoreOf(candidate) === scoreOf(candidates[index - 1]!),
)
) {
return [...candidates];
}
const queryTerms = wordTerms(query);
if (queryTerms.size === 0) return [...candidates];
const candidateTerms = candidates.map((candidate) => wordTerms(textOf(candidate)));
const weights = [...queryTerms].map((term) => {
const frequency = candidateTerms.filter((terms) => terms.has(term)).length;
return {
term,
weight: Math.log(1 + (candidates.length - frequency + 0.5) / (frequency + 0.5)),
};
});
const coverage = candidateTerms.map((terms) =>
weights.reduce((sum, { term, weight }) => sum + (terms.has(term) ? weight : 0), 0),
);
const reordered: T[] = [];
for (let start = 0; start < candidates.length;) {
let end = start + 1;
while (end < candidates.length && scoreOf(candidates[end]!) === scoreOf(candidates[start]!))
end++;
const tied = candidates
.slice(start, end)
.map((candidate, offset) => ({ candidate, index: start + offset }));
tied.sort(
(left, right) => coverage[right.index]! - coverage[left.index]! || left.index - right.index,
);
reordered.push(...tied.map(({ candidate }) => candidate));
start = end;
}
return reordered;
}

export function contextUsefulness(query: string, result: MemorySearchResult): number {
const normalized = normalize(query);
const type = result.memory.memoryType;
Expand Down
34 changes: 34 additions & 0 deletions tests/core/store/retrieval.test.ts
Original file line number Diff line number Diff line change
Expand Up @@ -48,6 +48,40 @@ test("searchContext returns results and relations for a lexical query", () => {
});
});

test("lexical tie-break keeps Active Graph selection order aligned with returned evidence", () => {
withStore((store) => {
const decoy = store.remember({
statement: "The store uses WAL",
nodeName: "Atlas SQLite",
importance: 1,
});
const target = store.remember({
statement: "Atlas SQLite stores project data",
nodeName: "Database choice",
importance: 0.5,
});
const query = "Atlas SQLite";
const context = store.searchContext(query, {
retrievalMode: "fts5",
limit: 2,
});
assert.ok(context.activeGraph);
assert.equal(context.results[0]?.memory.id, target.memory.id);
assert.equal(context.results[1]?.memory.id, decoy.memory.id);
assert.deepEqual(
context.activeGraph.memoryIds,
context.results.map((result) => result.memory.id),
);
assert.deepEqual(
context.activeGraph.selections.map((selection) => selection.memoryId),
context.activeGraph.memoryIds,
);
const trace = store.retrievalTrace(context.activeGraph.id);
assert.deepEqual(trace?.resultMemoryIds, context.activeGraph.memoryIds);
assert.deepEqual(trace?.selections?.map((selection) => selection.memoryId), context.activeGraph.memoryIds);
});
});

test("relevanceFloor gates disclosed results and may abstain", () => {
withStore((store) => {
store.remember({
Expand Down
22 changes: 22 additions & 0 deletions tests/core/store/search-ranking.test.ts
Original file line number Diff line number Diff line change
@@ -0,0 +1,22 @@
import assert from "node:assert/strict";
import test from "node:test";

import { rerankEqualScoresByQueryCoverage } from "../../../src/core/store/search-ranking.ts";

test("query coverage resolves lexical score ties without crossing score boundaries", () => {
const candidates = [
{ id: "partial", score: 4, text: "Atlas project notes" },
{ id: "complete", score: 4, text: "Atlas project uses SQLite" },
{ id: "lower", score: 2, text: "Atlas project uses SQLite" },
];
const rank = (items: typeof candidates) => rerankEqualScoresByQueryCoverage(
"Atlas project SQLite",
items,
(item) => item.score,
(item) => item.text,
);
const ranked = rank(candidates);
assert.deepEqual(ranked.map((item) => item.id), ["complete", "partial", "lower"]);
assert.deepEqual(rank(ranked), ranked, "reranking is stable on repeated calls");
assert.deepEqual(candidates.map((item) => item.id), ["partial", "complete", "lower"]);
});
Loading