What Does Post-Training Change in Multilingual Reasoning?
Quick summary
arXiv:2609.37104v1 Announce Type: cross Abstract: Open-source reasoning models provide unequal access to reasoning capability across languages. When a model can solve a problem but cannot deliver a complete solution in the user's language, language becomes an access barrier rather than merely a source of performance variation. We audit Qwen3 checkpoints on competition-mathematics tasks in eleven languages. Across the ten non-English languages, only 15.4-17.9% of problems receive a correct, terminating solution with visible reasoning in the requested language in any of 16 samples, compared with
Key takeaways
- arXiv:2609.37104v1 Announce Type: cross Abstract: Open-source reasoning models provide unequal access to reasoning capability across languages.
- When a model can solve a problem but cannot deliver a complete solution in the user's language, language becomes an access barrier rather than merely a source of performance variation.
- We audit Qwen3 checkpoints on competition-mathematics tasks in eleven languages.
Why it matters
“What Does Post-Training Change in Multilingual Reasoning?” should be evaluated beyond branding and benchmark scores. Its practical importance will emerge in task accuracy, latency, unit cost, safety and integration with real workflows.

Member comments