Reinforcement Learning Improves Traversal of Parametric Knowledge in LLMs
Quick summary
arXiv:2511.05933v3 Announce Type: replace-cross Abstract: Reinforcement learning (RL) is often credited with improving reasoning at the expense of factual knowledge. We instead find that reasoning models outperform their instruction-tuned versions on factual recall by accessing existing parametric knowledge more effectively. Across five model families, structured prompting, which explicitly guides models through hierarchical traversal, recovers most of this gap, suggesting that much of the missing knowledge is latent rather than absent. Controlled RL experiments further support this: training
Key takeaways
- arXiv:2511.05933v3 Announce Type: replace-cross Abstract: Reinforcement learning (RL) is often credited with improving reasoning at the expense of factual knowledge.
- We instead find that reasoning models outperform their instruction-tuned versions on factual recall by accessing existing parametric knowledge more effectively.
- Across five model families, structured prompting, which explicitly guides models through hierarchical traversal, recovers most of this gap, suggesting that much of the missing knowledge is latent rather than absent.
Why it matters
“Reinforcement Learning Improves Traversal of Parametric Knowledge in LLMs” should be evaluated beyond branding and benchmark scores. Its practical importance will emerge in task accuracy, latency, unit cost, safety and integration with real workflows.

Member comments