arXiv Artificial Intelligence

DolphinBench: Mapping the Pareto Frontier of Agent Memory

DolphinBench: Mapping the Pareto Frontier of Agent Memory

Quick summary

arXiv:2609.24971v2 Announce Type: cross Abstract: Agents today often take real-world actions that depend on long-term memory and context recall over time. However, most current memory benchmarks are built for a conversational question-answer format, where the question itself signals that some fact must be retrieved, and often which one. Moreover, benchmarks rarely require anything beyond accuracy from submissions, allowing memory systems to make unreasonable cost/time tradeoffs to achieve higher scores. We present DolphinBench, a benchmark that evaluates memory directly through an agent's task

Key takeaways

  • arXiv:2609.24971v2 Announce Type: cross Abstract: Agents today often take real-world actions that depend on long-term memory and context recall over time.
  • However, most current memory benchmarks are built for a conversational question-answer format, where the question itself signals that some fact must be retrieved, and often which one.
  • Moreover, benchmarks rarely require anything beyond accuracy from submissions, allowing memory systems to make unreasonable cost/time tradeoffs to achieve higher scores.

Why it matters

The value of this work lies as much in how it was tested as in the claim itself. Sample design, baselines, uncertainty and replication help separate a laboratory result from real-world impact.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗