arXiv Artificial Intelligence

BETA-Labeling for Multilingual Dataset Construction in Low-Resource IR

BETA-Labeling for Multilingual Dataset Construction in Low-Resource IR

Quick summary

arXiv:2602.14488v3 Announce Type: replace-cross Abstract: IR in low-resource languages remains limited by the scarcity of high-quality, task-specific annotated datasets. Manual annotation is expensive and difficult to scale, while using large language models (LLMs) as automated annotators introduces concerns about label reliability, bias, and evaluation validity. This work presents a Bangla IR dataset constructed using a BETA-labeling framework involving multiple LLM annotators from diverse model families. The framework incorporates contextual alignment, consistency checks, and majority agreem

Key takeaways

  • arXiv:2602.14488v3 Announce Type: replace-cross Abstract: IR in low-resource languages remains limited by the scarcity of high-quality, task-specific annotated datasets.
  • Manual annotation is expensive and difficult to scale, while using large language models (LLMs) as automated annotators introduces concerns about label reliability, bias, and evaluation validity.
  • This work presents a Bangla IR dataset constructed using a BETA-labeling framework involving multiple LLM annotators from diverse model families.

Why it matters

The value of this work lies as much in how it was tested as in the claim itself. Sample design, baselines, uncertainty and replication help separate a laboratory result from real-world impact.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗