Probing for Knowledge Attribution in Large Language Models
Quick summary
arXiv:2602.22787v3 Announce Type: replace-cross Abstract: Large language model (LLM) hallucinations, meaning fluent but factually incorrect generations, fall into two types: faithfulness violations, where the model misuses provided context, and factuality violations, where answers reflect errors in internal knowledge. Proper mitigation depends on knowing which source drives each answer. We study contributive attribution, i.e. the classification of the dominant knowledge source behind each output, and show that a simple linear probe trained on hidden representations can reliably identify it. We
Key takeaways
- arXiv:2602.22787v3 Announce Type: replace-cross Abstract: Large language model (LLM) hallucinations, meaning fluent but factually incorrect generations, fall into two types: faithfulness violations, where the model misuses provided context, and factuality violations, where answers reflect errors in internal knowledge.
- Proper mitigation depends on knowing which source drives each answer.
- We study contributive attribution, i.e.
Why it matters
“Probing for Knowledge Attribution in Large Language Models” highlights the need for repeatable measurement rather than a single impressive demonstration. Independent validation across datasets and clearly stated limitations determine whether a result can guide product decisions.

Member comments