arXiv Artificial IntelligenceCorporateBench: Large-Scale Q&A Benchmarking with Temporal Knowledge Bases
arXiv:2608.27391v1 Announce Type: new Abstract: LLMs are increasingly able to answer complex questions about…
Curated from international AI laboratories, specialist publications and technology outlets. Last update: 3 hours ago.
arXiv Artificial IntelligencearXiv:2608.27391v1 Announce Type: new Abstract: LLMs are increasingly able to answer complex questions about…
arXiv Artificial IntelligencearXiv:2608.27364v1 Announce Type: new Abstract: We study how sophistication in generative AI (genAI) use…
arXiv Artificial IntelligencearXiv:2608.27340v1 Announce Type: new Abstract: Steering interventions targeting eval-awareness, a model's…
arXiv Artificial IntelligencearXiv:2608.27311v1 Announce Type: new Abstract: Agent harnesses shape how language-model agents use…
arXiv Artificial IntelligencearXiv:2608.27268v1 Announce Type: new Abstract: Although Large language models (LLMs) mediate access to…
arXiv Artificial IntelligencearXiv:2608.27266v1 Announce Type: new Abstract: Efficiently improving autonomous agents across diverse tasks…
arXiv Artificial IntelligencearXiv:2608.27260v1 Announce Type: new Abstract: LLM agents increasingly rely on generated interaction data to…
arXiv Artificial IntelligencearXiv:2608.27167v1 Announce Type: new Abstract: An LLM agent shown a professional-looking market panel…
arXiv Artificial IntelligencearXiv:2608.27149v1 Announce Type: new Abstract: Conversational AI systems, such as chatbots and virtual…
arXiv Artificial IntelligencearXiv:2608.27147v1 Announce Type: new Abstract: The development of frontier models is commonly perceived to…
arXiv Artificial IntelligencearXiv:2608.27146v1 Announce Type: new Abstract: Tool-augmented LLM agents must rely on untrusted runtime…
arXiv Artificial IntelligencearXiv:2608.27144v1 Announce Type: new Abstract: In recent years, graph anomaly detection (GAD) based on…