arXiv Artificial Intelligence

SP-DocReader: Difference-Aware Self-Play for Precise Document OCR

SP-DocReader: Difference-Aware Self-Play for Precise Document OCR

Quick summary

arXiv:2610.11148v1 Announce Type: cross Abstract: Accurate page transcription remains difficult for vision language models under limited input and training budgets. We present SP-DocReader, a self-play framework for optical character recognition (OCR) that targets residual errors after supervised fine-tuning. Reading Discrepancy Masking aligns reference and generated model tokens through a longest common subsequence, then scores unmatched positions with their full conditioning prefixes. Focused Fidelity Loss adds direct negative log-likelihood supervision at unmatched ground-truth positions. O

Key takeaways

  • arXiv:2610.11148v1 Announce Type: cross Abstract: Accurate page transcription remains difficult for vision language models under limited input and training budgets.
  • We present SP-DocReader, a self-play framework for optical character recognition (OCR) that targets residual errors after supervised fine-tuning.
  • Reading Discrepancy Masking aligns reference and generated model tokens through a longest common subsequence, then scores unmatched positions with their full conditioning prefixes.

Why it matters

“SP-DocReader: Difference-Aware Self-Play for Precise Document OCR” should be evaluated beyond branding and benchmark scores. Its practical importance will emerge in task accuracy, latency, unit cost, safety and integration with real workflows.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗