arXiv Artificial Intelligence

Era by Eon: Benchmarking Enterprise Agents on Hidden Knowledge

Era by Eon: Benchmarking Enterprise Agents on Hidden Knowledge

Quick summary

arXiv:2609.30055v1 Announce Type: cross Abstract: In the Era by Eon benchmark, each question states the rules for its answer, and code computes the answer from a generated company's data. When agents can run code, the four strongest models each answer 22 to 25 of 27 such questions, so the benchmark barely separates them. We add eight question templates that depend on hidden facts. No question or document states a hidden fact, and the records that seem to hold it show something else. Other data implies it. For example, the sales system says a customer dropped a purchase because of timing. On a

Key takeaways

  • arXiv:2609.30055v1 Announce Type: cross Abstract: In the Era by Eon benchmark, each question states the rules for its answer, and code computes the answer from a generated company's data.
  • When agents can run code, the four strongest models each answer 22 to 25 of 27 such questions, so the benchmark barely separates them.
  • We add eight question templates that depend on hidden facts.

Why it matters

The value of this work lies as much in how it was tested as in the claim itself. Sample design, baselines, uncertainty and replication help separate a laboratory result from real-world impact.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗