arXiv Artificial Intelligence

Diffusion Large Language Models for Visual Speech Recognition

Diffusion Large Language Models for Visual Speech Recognition

Quick summary

arXiv:2605.28456v2 Announce Type: replace Abstract: Existing Visual Speech Recognition (VSR) systems commonly rely on left-to-right autoregressive decoding, which can force premature decisions on visually ambiguous tokens before sufficient context is available. We propose DLLM-VSR, to the best of our knowledge, the first Diffusion Large Language Model (DLLM)-based VSR framework, formulating transcription as iterative masked denoising with flexible-order decoding. With confidence-based unmasking, DLLM-VSR commits high-confidence positions early and uses the committed tokens as bidirectional con

Key takeaways

  • arXiv:2605.28456v2 Announce Type: replace Abstract: Existing Visual Speech Recognition (VSR) systems commonly rely on left-to-right autoregressive decoding, which can force premature decisions on visually ambiguous tokens before sufficient context is available.
  • We propose DLLM-VSR, to the best of our knowledge, the first Diffusion Large Language Model (DLLM)-based VSR framework, formulating transcription as iterative masked denoising with flexible-order decoding.
  • With confidence-based unmasking, DLLM-VSR commits high-confidence positions early and uses the committed tokens as bidirectional con

Why it matters

“Diffusion Large Language Models for Visual Speech Recognition” should be evaluated beyond branding and benchmark scores. Its practical importance will emerge in task accuracy, latency, unit cost, safety and integration with real workflows.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗