feat(retrieval): improve lexical score ties after neural ranker failure - #75
Merged
Merged
Conversation
Apply query-term coverage after Active Graph evidence selection on the FTS5-only path. Keep candidate membership, QPP decisions, and hybrid ranking intact. Record paired retrieval results and the failed small neural ranker experiment. No neural model is added.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
变更描述
What:在纯 FTS5 词法检索中,对 Active Graph 已选出的、相邻且
combinedScore完全相同的候选,按查询词在陈述和前 500 个证据字符中的 IDF 加权覆盖度排序。候选集合、非并列顺序、QPP 决策和混合检索保持原样。Why:#52 后继续探索的小型神经网络排序方案没有达到产品准入门槛:在 19 条独立真实召回查询上,同一候选集 MRR(Q) 为 0.861,低于现有排序的 0.921;限制到顶部并列组后虽为 0.947,但配对结果不显著(p=1.0)。本次神经网络排序实验失败,不加入模型。程序化的词法并列排序在固定的五套召回样本中有可复核收益,尤其 LoCoMo 的 R@1 由 7.71% 升至 17.42%、MRR(Q) 由 0.1830 升至 0.3364,10/10 用户改善(按用户分组的精确符号翻转 p=0.00195)。
Changes:
未验证项
Not verified:没有运行最终回答质量/读者评分或线上 A/B;没有实测新增常驻内存 RSS 的 100 MB 硬上限。规则无需模型权重,离线重排的平均 CPU 开销约 0.13–2.0 毫秒/查询。神经网络实验只有 19 条真实查询和显式正例,没有已验证负例,不能据此估计无关内容的精确率。缓存向量的 LongMemEval 对照指标不变,未对全部混合检索组合重跑端到端回答评估。
完成检查项
本地质量检查
npm run verify:static通过。npm run agent:verify -- <本 PR 的 9 个路径>通过,包含npm run test:product、npm run build和所有阻塞路由检查;41 项定向测试通过。npm run docs:check通过;非平凡改动有双语决策记录和实验记录。npm run complexity:gate通过。RCP(Repository Control Plane)
repo-developmentin-flight goal;创建 PR 后结清。agent:verify结果写入.nmg/verification/latest.json,全部阻塞检查通过。nmg-rcp forge-status --pr 75确认 CI 全绿。CI 完成确认
All checks passed为 SUCCESS。