arXiv Artificial IntelligenceTRACE: An Evidence-Grounded Benchmark for Safety Evaluation of Large Reasoning Models
arXiv:2608.24232v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) generate intermediate reasoning…
Curated from international AI laboratories, specialist publications and technology outlets. Last update: 3 hours ago.
arXiv Artificial IntelligencearXiv:2608.24232v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) generate intermediate reasoning…
arXiv Artificial IntelligencearXiv:2608.24228v1 Announce Type: new Abstract: Many LLM applications are most useful when they provide…
arXiv Artificial IntelligencearXiv:2608.24218v1 Announce Type: new Abstract: Enterprise entity alignment must handle semi-structured…
arXiv Artificial IntelligencearXiv:2608.24214v1 Announce Type: new Abstract: Agentic retrieval-augmented generation (RAG) requires…
arXiv Artificial IntelligencearXiv:2608.24192v1 Announce Type: new Abstract: Aligning large language models to human preferences is…
arXiv Artificial IntelligencearXiv:2608.24188v1 Announce Type: new Abstract: Coding agents re-send large file reads and tool outputs to a…
arXiv Artificial IntelligencearXiv:2608.24174v1 Announce Type: new Abstract: Recent studies on GUI agents have increasingly focused on…
arXiv Artificial IntelligencearXiv:2608.24160v1 Announce Type: new Abstract: Multimodal understanding models that can jointly judge…
arXiv Artificial IntelligencearXiv:2608.24114v1 Announce Type: new Abstract: Training multi-turn LLM agents with reinforcement learning…
arXiv Artificial IntelligencearXiv:2608.24112v1 Announce Type: new Abstract: Modern text-to-image (T2I) models often have similar total…
arXiv Artificial IntelligencearXiv:2608.24103v1 Announce Type: new Abstract: Commercial design platforms increasingly edit documents…
arXiv Artificial IntelligencearXiv:2608.24099v1 Announce Type: new Abstract: GUI agents often encounter dynamic anomalies when deployed on…