CroCo: Cross-Lingual Contrastive Preference Tuning on Self-Generations
Quick summary
arXiv:2605.26293v2 Announce Type: replace-cross Abstract: Prior work establishes that controlled contrastiveness between self-generated responses from large language models, set via reward scores, improves downstream preference tuning in English. We extend this method to multiple languages and evaluate two models across a total of 14 high and low-resource languages on a diverse set of tasks. Our central finding is that cross-lingual contrastive preference tuning on self-generations (CroCo) transfers without language-specific preference annotation. A reward model trained on English preferences
Key takeaways
- arXiv:2605.26293v2 Announce Type: replace-cross Abstract: Prior work establishes that controlled contrastiveness between self-generated responses from large language models, set via reward scores, improves downstream preference tuning in English.
- We extend this method to multiple languages and evaluate two models across a total of 14 high and low-resource languages on a diverse set of tasks.
- Our central finding is that cross-lingual contrastive preference tuning on self-generations (CroCo) transfers without language-specific preference annotation.
Why it matters
“CroCo: Cross-Lingual Contrastive Preference Tuning on Self-Generations” should be evaluated beyond branding and benchmark scores. Its practical importance will emerge in task accuracy, latency, unit cost, safety and integration with real workflows.

Member comments