Research brief

All digests

Research Paper Digest · 2026-09-29

1 paper

01 · When Does Selection Replace Extraction? A Pre-Registered Test of Agent Memory with a Typed Decision Model

何时选择取代抽取?使用类型化决策模型对智能体记忆的预注册测试

Research recordDetails
AuthorsRishabh Sharma, Rishika Lall
Published2026-09-27
Sourcesopenalex
Focusconversational memory selection vs extraction, typed decision model (Jev), reranking and budget trade-offs, latency and cost efficiency
会话记忆:选择与抽取之争, 类型化决策模型(Jev), 重排序与预算权衡, 延迟与成本效率

Reading verdict

Skim · 浏览

English

The pre-registered empirical findings on selection vs extraction and the explicit cost/latency measurements are relevant to designing a low-latency candidate-selection stage, but the paper does not address structured graph memory, entity resolution, or multi-hop graph traversal central to the thesis.

中文

该论文关于选择与抽取的预注册实证结果以及明确的成本/延迟测量,对设计低延迟候选选择阶段有参考价值,但未涉及论文核心的结构化图记忆、实体解析或多跳图遍历问题。

Research synopsis

English

The paper asks whether conversational agent memory must consist of LLM-extracted facts or whether selecting raw prior turns suffices. The authors run a pre-registered empirical comparison on held-out LoCoMo conversations and LongMemEval using Jev, a typed decision model that selects raw turns, versus an LLM-extraction memory and reranking approaches. At a tight budget on LoCoMo, single-call Jev selection is non-inferior to extraction (one-sided 95% bound −3.0 vs a −5-point margin) and is far cheaper to write (reported 3,061× cost reduction). They report that reranking yields large gains at very small budgets but that its marginal benefit shrinks as budget grows; at generous budgets extraction can be more accurate. The paper releases code, data, and a pre-registered plan.

中文

本文探讨会话智能体记忆是否必须依赖由大模型抽取的事实,还是选择原始历史轮次就足够。作者在 LoCoMo 和 LongMemEval 上做了预注册的实证比较,使用 Jev(一种类型化决策模型)以单次调用选择原始轮次,比较其与 LLM 抽取记忆和重排序方法的表现。在 LoCoMo 的紧预算设置下,单次调用的 Jev 选择被证明在统计上不劣于抽取(单边 95% 界 −3.0,门槛 −5 点),且写入成本显著更低(报告为 3,061×)。作者还观察到重排序在极小预算下带来大增益,但随预算增长其边际收益下降;在宽松预算下抽取方法可能更准确。论文同时发布了代码、数据和预注册计划。

Thesis relevance

English

Overlap: both study agent memory design choices and empirical trade-offs among accuracy, latency, and cost, which is relevant to the thesis’s interest in memory strategies for retrieval and reasoning. Differences/limitations: this paper focuses on unstructured raw-turn selection and reranking rather than persistent graph memory, knowledge-graph construction, or entity resolution. Complementarity: its findings on low-latency candidate selection and budget-dependent reranking effects could inform the thesis’s pipeline—for example, as a lightweight filter before triple extraction and graph insertion.

中文

重合点:两者都研究智能体记忆的设计选择,以及准确性、延迟和成本之间的经验权衡,这对论文关注的记忆策略与检索/推理问题具有参考价值。差异/局限:该论文集中于非结构化的原始轮次选择和重排序,而非持久化图记忆、知识图谱构建或实体解析。互补性:其关于低延迟候选选择和重排序随预算变化的结论,可用于为论文设计的流水线提供轻量级预筛选机制(例如在抽取三元组并插入图之前)。

English

The paper is a well-documented, pre-registered empirical contribution that foregrounds cost and latency as first-order concerns; cite it when discussing empirical protocols, pre-registration, or when motivating a candidate-selection stage in a structured memory pipeline. Emphasize its budget-dependent interpretation of prior contradictory results.

中文

该论文为预注册的实证研究,明确将成本与延迟作为一阶考量;在讨论实验证据、预注册流程或为结构化记忆流水线设计候选选择阶段时可引用。着重其关于预算依赖性如何解释以往冲突性结果的论点。

Method and evaluation

English

Consider adopting a Jev-like typed decision model as a low-latency selector to provide candidate passages or turns before performing LLM triple extraction and entity-resolution for graph memory. Empirically compare single-call selection, reranking budgets, and multi-call extraction in your MultiHop-RAG benchmark, measuring downstream effects on bridge-entity identification, graph coherence, and end-to-end latency. Report cost (compute and write) and ablation over budget sizes to reveal when extraction justifies its cost.

中文

可考虑引入类似 Jev 的类型化决策模型,作为低延迟的选择器,在进行 LLM 三元组抽取和实体解析以构建图记忆之前提供候选段或轮次。对比单次调用选择、不同重排序预算和多次调用抽取在 MultiHop-RAG 基准上的表现,衡量其对桥接实体识别、图一致性及端到端延迟的下游影响。报告计算与写入成本并做预算规模消融,以揭示何时抽取的开销是合理的。

Future directions

English

Integrate selection-first pipelines where Jev-like selection reduces the candidate set for triple extraction and entity-resolution, then measure how selection errors propagate into graph fragmentation and multi-hop retrieval. Also test whether budget-adaptive reranking policies improve long-term graph coherence and bridge-entity reuse.

中文

将“先选择后抽取”的流水线集成到图记忆构建中,使用类似 Jev 的选择器缩减候选集,然后衡量选择错误如何传播为图碎片化并影响多跳检索。还可测试自适应预算的重排序策略是否能提升长期图一致性和桥接实体的重用率。