arXiv Artificial Intelligence

Language Chain in Alignment: Cross-lingual Ranking Preference Optimization

Language Chain in Alignment: Cross-lingual Ranking Preference Optimization

Quick summary

arXiv:2608.23149v2 Announce Type: replace-cross Abstract: The alignment of Large Language Models heavily relies on English-centric high-quality preference data, which often leads to suboptimal performance in other languages. In this paper, we propose Cross-lingual Ranking Preference Optimization~(CRPO), a novel framework that leverages robust preference knowledge from English to facilitate preference alignment in the target language. We design a hierarchical structure within parallel preference pairs across the target language and English to jointly optimize intra- and inter-lingual preference

Key takeaways

  • arXiv:2608.23149v2 Announce Type: replace-cross Abstract: The alignment of Large Language Models heavily relies on English-centric high-quality preference data, which often leads to suboptimal performance in other languages.
  • In this paper, we propose Cross-lingual Ranking Preference Optimization~(CRPO), a novel framework that leverages robust preference knowledge from English to facilitate preference alignment in the target language.
  • We design a hierarchical structure within parallel preference pairs across the target language and English to jointly optimize intra- and inter-lingual preference

Why it matters

The value of this work lies as much in how it was tested as in the claim itself. Sample design, baselines, uncertainty and replication help separate a laboratory result from real-world impact.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗