Transcribe, Then Reason: Two-Pass Decomposition for Multimodal Review
Quick summary
arXiv:2609.18958v1 Announce Type: cross Abstract: The natural way to review a long recording or document with a multimodal model is to hand it the raw source and ask for a review in one call. We show that this quietly fails: the model satisfices, dropping roughly a third of the content and embellishing the rest. The failure is not perception--almost all of the dropped content reappears when the same model is simply asked to transcribe the source. The bottleneck is generation under load: a single pass cannot perceive, reason over, and write a long faithful review at the same time, because doing
Key takeaways
- arXiv:2609.18958v1 Announce Type: cross Abstract: The natural way to review a long recording or document with a multimodal model is to hand it the raw source and ask for a review in one call.
- We show that this quietly fails: the model satisfices, dropping roughly a third of the content and embellishing the rest.
- The failure is not perception--almost all of the dropped content reappears when the same model is simply asked to transcribe the source.
Why it matters
This model development creates a new option for users and a new testing obligation for developers. A fixed evaluation set comparing quality, cost and failure behavior is more useful than launch claims.

Member comments