Evaluating the Retrieval Robustness of Large Language Models
Quick summary
arXiv:2505.21870v2 Announce Type: replace-cross Abstract: Retrieval-augmented generation (RAG) generally enhances large language models' (LLMs) ability to solve knowledge-intensive tasks. But RAG could also lead to performance degradation due to imperfect retrieval and the model's limited ability to leverage retrieved content. In this work, we evaluate the robustness of LLMs in practical RAG setups (henceforth retrieval robustness). We focus on three research questions: (1) whether RAG is always better than non-RAG; (2) whether more retrieved documents always lead to better performance; and (3
Key takeaways
- arXiv:2505.21870v2 Announce Type: replace-cross Abstract: Retrieval-augmented generation (RAG) generally enhances large language models' (LLMs) ability to solve knowledge-intensive tasks.
- But RAG could also lead to performance degradation due to imperfect retrieval and the model's limited ability to leverage retrieved content.
- In this work, we evaluate the robustness of LLMs in practical RAG setups (henceforth retrieval robustness).
Why it matters
The value of this work lies as much in how it was tested as in the claim itself. Sample design, baselines, uncertainty and replication help separate a laboratory result from real-world impact.

Member comments