arXiv Artificial IntelligenceSkillRet: A Large-Scale Benchmark for Skill Retrieval in LLM Agents
arXiv:2605.05726v3 Announce Type: replace Abstract: As LLM agents are increasingly deployed with large…
Curated from international AI laboratories, specialist publications and technology outlets. Last update: 3 hours ago.
arXiv Artificial IntelligencearXiv:2605.05726v3 Announce Type: replace Abstract: As LLM agents are increasingly deployed with large…
arXiv Artificial IntelligencearXiv:2608.30731v2 Announce Type: replace-cross Abstract: Assessing claim check-worthiness is an essential…
arXiv Artificial IntelligencearXiv:2608.30646v2 Announce Type: replace-cross Abstract: Reliable uncertainty estimation is a crucial…
arXiv Artificial IntelligencearXiv:2608.30621v2 Announce Type: replace-cross Abstract: Collecting natural-language referring expressions…
arXiv Artificial IntelligencearXiv:2608.30092v2 Announce Type: replace-cross Abstract: We present Arkios, a 1.04B-parameter dense…
arXiv Artificial IntelligencearXiv:2608.29899v2 Announce Type: replace-cross Abstract: Dense retrieval over long documents is expensive…
arXiv Artificial IntelligencearXiv:2608.29464v2 Announce Type: replace-cross Abstract: Chain-of-thought (CoT) monitoring assumes that…
arXiv Artificial IntelligencearXiv:2608.31075v2 Announce Type: replace Abstract: Recent advances in large reasoning models (LRMs) have…
arXiv Artificial IntelligencearXiv:2608.30751v2 Announce Type: replace Abstract: Large language models (LLMs) trained only on text and…
arXiv Artificial IntelligencearXiv:2608.30362v2 Announce Type: replace Abstract: As LLM agents take real-world actions through tools…
arXiv Artificial IntelligencearXiv:2608.29249v2 Announce Type: replace Abstract: The online culinary ecosystem is increasingly populated…
arXiv Artificial IntelligencearXiv:2608.29207v2 Announce Type: replace Abstract: Protein structure modeling rests on a single…