arXiv Artificial IntelligenceStrategyBench: Evaluating Explicit Strategy Induction in Large Language Models
arXiv:2608.23475v1 Announce Type: new Abstract: As large language models are increasingly used in data-scarce…
Curated from international AI laboratories, specialist publications and technology outlets. Last update: 3 hours ago.
arXiv Artificial IntelligencearXiv:2608.23475v1 Announce Type: new Abstract: As large language models are increasingly used in data-scarce…
arXiv Artificial IntelligencearXiv:2608.23417v1 Announce Type: new Abstract: Agent skills are reusable procedural artifacts that extend…
arXiv Artificial IntelligencearXiv:2608.23373v1 Announce Type: new Abstract: Forecasting long-range influenza-like illness (ILI) matters…
arXiv Artificial IntelligencearXiv:2608.23318v1 Announce Type: new Abstract: Hint-based reinforcement learning addresses reward sparsity…
arXiv Artificial IntelligencearXiv:2608.23313v1 Announce Type: new Abstract: Vision-language model safety benchmarks typically evaluate…
arXiv Artificial IntelligencearXiv:2608.23264v1 Announce Type: new Abstract: Although Large Language Models (LLMs) are aligned to optimize…
arXiv Artificial IntelligencearXiv:2608.23263v1 Announce Type: new Abstract: The FAIR Digital Object (FDO) framework mandates that…
arXiv Artificial IntelligencearXiv:2608.23256v1 Announce Type: new Abstract: Recent work proposes next-chunk reasoning RL for leveraging…
arXiv Artificial IntelligencearXiv:2608.23218v1 Announce Type: new Abstract: Advances in neural theorem provers have been impressive, but…
arXiv Artificial IntelligencearXiv:2608.23205v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) have revolutionized reasoning…
arXiv Artificial IntelligencearXiv:2608.23196v1 Announce Type: new Abstract: People increasingly face a novel decision when seeking…
arXiv Artificial IntelligencearXiv:2608.23086v1 Announce Type: new Abstract: Black-box large language models need confidence scores that…