arXiv Artificial IntelligenceDon't Claim Benchmark-Oriented Optimization Improves General Coding Capability -- Diverse Evaluation Is Required
arXiv:2608.13566v1 Announce Type: cross Abstract: Post-training papers, model cards, and blog posts often…
Curated from international AI laboratories, specialist publications and technology outlets. Last update: 3 hours ago.
arXiv Artificial IntelligencearXiv:2608.13566v1 Announce Type: cross Abstract: Post-training papers, model cards, and blog posts often…
arXiv Artificial IntelligencearXiv:2608.13563v1 Announce Type: cross Abstract: Early-stage teams often lack users, time, and budget to run…
arXiv Artificial IntelligencearXiv:2608.14528v1 Announce Type: new Abstract: This study investigates the methodological and theoretical…
arXiv Artificial IntelligencearXiv:2608.14522v1 Announce Type: new Abstract: As AI systems make more morally loaded decisions across…
arXiv Artificial IntelligencearXiv:2608.14509v1 Announce Type: new Abstract: Systems that ask a language model to reach a conclusion from…
arXiv Artificial IntelligencearXiv:2608.14490v1 Announce Type: new Abstract: We present a Test-time World-model Inference (Twin) system…
arXiv Artificial IntelligencearXiv:2608.14456v1 Announce Type: new Abstract: Short-horizon forecasting of fine particulate matter (PM2.5)…
arXiv Artificial IntelligencearXiv:2608.14452v1 Announce Type: new Abstract: Spreadsheets are widely used to organize, analyze, and…
arXiv Artificial IntelligencearXiv:2608.14446v1 Announce Type: new Abstract: In the current artificial intelligence-driven innovation era…
arXiv Artificial IntelligencearXiv:2608.14441v1 Announce Type: new Abstract: Self-evolving agents improve future behavior from interaction…
arXiv Artificial IntelligencearXiv:2608.14425v1 Announce Type: new Abstract: LLM evaluations often use fixed sampling budgets, testing…
arXiv Artificial IntelligencearXiv:2608.14407v1 Announce Type: new Abstract: We present a survey of the past and future of AI Scientists…