arXiv Artificial Intelligence

Benchmarking Document Parsers on Mathematical Formula Extraction from PDFs

Benchmarking Document Parsers on Mathematical Formula Extraction from PDFs

Quick summary

arXiv:2512.09874v3 Announce Type: replace-cross Abstract: Correctly parsing mathematical formulas from PDFs is critical for training large language models and building scientific knowledge bases from academic literature, yet existing benchmarks either exclude formulas entirely or lack semantically-aware evaluation metrics. We introduce a benchmarking framework centered on synthetically generated PDFs with precise LaTeX ground truth, enabling systematic control over layout, formulas, and content characteristics. For evaluation, we apply LLM-as-a-judge to assess semantic equivalence of parsed fo

Key takeaways

  • arXiv:2512.09874v3 Announce Type: replace-cross Abstract: Correctly parsing mathematical formulas from PDFs is critical for training large language models and building scientific knowledge bases from academic literature, yet existing benchmarks either exclude formulas entirely or lack semantically-aware evaluation metrics.
  • We introduce a benchmarking framework centered on synthetically generated PDFs with precise LaTeX ground truth, enabling systematic control over layout, formulas, and content characteristics.
  • For evaluation, we apply LLM-as-a-judge to assess semantic equivalence of parsed fo

Why it matters

“Benchmarking Document Parsers on Mathematical Formula Extraction from PDFs” highlights the need for repeatable measurement rather than a single impressive demonstration. Independent validation across datasets and clearly stated limitations determine whether a result can guide product decisions.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗