Audio LLMs Know When They Can't Hear You
Quick summary
arXiv:2609.30625v1 Announce Type: new Abstract: Audio large language models allow users to interact with the model through speech. When an input recording is too degraded, the model may misinterpret the user's query and respond based on an incorrect transcription. In this paper, we study model-conditional transcription reliability: whether an Audio LLM can recognize when its own transcription is unreliable. We first prompt the Audio LLM to assess whether its own transcription would be reliable, and find that the model is a poor judge of its own transcription reliability: in most cases, it pred
Key takeaways
- arXiv:2609.30625v1 Announce Type: new Abstract: Audio large language models allow users to interact with the model through speech.
- When an input recording is too degraded, the model may misinterpret the user's query and respond based on an incorrect transcription.
- In this paper, we study model-conditional transcription reliability: whether an Audio LLM can recognize when its own transcription is unreliable.
Why it matters
The value of this work lies as much in how it was tested as in the claim itself. Sample design, baselines, uncertainty and replication help separate a laboratory result from real-world impact.

Member comments