arXiv Artificial IntelligenceLifeAgentBench: Benchmarking LLMs for Long-Horizon, Cross-Dimensional Lifestyle Health Reasoning
arXiv:2601.13880v2 Announce Type: replace Abstract: Personalized lifestyle health analysis requires…
Curated from international AI laboratories, specialist publications and technology outlets. Last update: 3 hours ago.
arXiv Artificial IntelligencearXiv:2601.13880v2 Announce Type: replace Abstract: Personalized lifestyle health analysis requires…
arXiv Artificial IntelligencearXiv:2511.15755v3 Announce Type: replace Abstract: Large language models (LLMs) promise to accelerate…
arXiv Artificial IntelligencearXiv:2511.04093v2 Announce Type: replace Abstract: Large language models (LLMs) excel at reasoning but…
arXiv Artificial IntelligencearXiv:2510.15144v4 Announce Type: replace Abstract: Simulating human reasoning in open-ended tasks has long…
arXiv Artificial IntelligencearXiv:2510.14980v3 Announce Type: replace Abstract: Large language models (LLMs) have shown strong abilities…
arXiv Artificial IntelligencearXiv:2510.01030v2 Announce Type: replace Abstract: The human ability to translate diverse perceptual and…
arXiv Artificial IntelligencearXiv:2508.04080v2 Announce Type: replace Abstract: Standard large language model prompting treats geospatial…
arXiv Artificial IntelligencearXiv:2505.13180v3 Announce Type: replace Abstract: Integrating Large Language Models with symbolic planners…
arXiv Artificial IntelligencearXiv:2609.01603v1 Announce Type: cross Abstract: Evaluating software engineering agents on realistic…
arXiv Artificial IntelligencearXiv:2609.01601v1 Announce Type: cross Abstract: The repository-level code generation task requires…
arXiv Artificial IntelligencearXiv:2609.01600v1 Announce Type: cross Abstract: Dynamic agent harnesses let language models change the…
arXiv Artificial IntelligencearXiv:2609.01597v1 Announce Type: cross Abstract: Natural language is emerging as a primary feedback channel…