arXiv Artificial Intelligence

Audio LLMs Know When They Can't Hear You

Audio LLMs Know When They Can't Hear You

Quick summary

arXiv:2609.30625v1 Announce Type: new Abstract: Audio large language models allow users to interact with the model through speech. When an input recording is too degraded, the model may misinterpret the user's query and respond based on an incorrect transcription. In this paper, we study model-conditional transcription reliability: whether an Audio LLM can recognize when its own transcription is unreliable. We first prompt the Audio LLM to assess whether its own transcription would be reliable, and find that the model is a poor judge of its own transcription reliability: in most cases, it pred

Key takeaways

  • arXiv:2609.30625v1 Announce Type: new Abstract: Audio large language models allow users to interact with the model through speech.
  • When an input recording is too degraded, the model may misinterpret the user's query and respond based on an incorrect transcription.
  • In this paper, we study model-conditional transcription reliability: whether an Audio LLM can recognize when its own transcription is unreliable.

Why it matters

The value of this work lies as much in how it was tested as in the claim itself. Sample design, baselines, uncertainty and replication help separate a laboratory result from real-world impact.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗