arXiv Artificial Intelligence

Cross-lingual Biography Enrichment via Claim Extraction and Alignment

Cross-lingual Biography Enrichment via Claim Extraction and Alignment

Quick summary

arXiv:2608.23390v1 Announce Type: cross Abstract: English Wikipedia is often treated as the default encyclopedic source, yet non-English Wikipedia editions can contain richer locally grounded information for long-tail figures. We study cross-lingual biography enrichment: enriching an existing English biography with facts supported by a non-English biography about the same person. Focusing on women from non-English-speaking contexts, we introduce \textsc{CLAW-4L}, a benchmark consisting of 300 Wikipedia biography pairs linking an English biography with its French, Chinese or Azerbaijani counter

Key takeaways

  • arXiv:2608.23390v1 Announce Type: cross Abstract: English Wikipedia is often treated as the default encyclopedic source, yet non-English Wikipedia editions can contain richer locally grounded information for long-tail figures.
  • We study cross-lingual biography enrichment: enriching an existing English biography with facts supported by a non-English biography about the same person.
  • Focusing on women from non-English-speaking contexts, we introduce \textsc{CLAW-4L}, a benchmark consisting of 300 Wikipedia biography pairs linking an English biography with its French, Chinese or Azerbaijani counter

Why it matters

“Cross-lingual Biography Enrichment via Claim Extraction and Alignment” highlights the need for repeatable measurement rather than a single impressive demonstration. Independent validation across datasets and clearly stated limitations determine whether a result can guide product decisions.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗