arXiv Artificial Intelligence

InterPol: De-anonymizing LM Arena via Interpolated Preference Learning

InterPol: De-anonymizing LM Arena via Interpolated Preference Learning

Quick summary

arXiv:2603.15220v2 Announce Type: replace Abstract: Strict anonymity of model responses is a key for the reliability of voting-based leaderboards, such as LM Arena. While prior studies have attempted to compromise this assumption using simple statistical features like TF-IDF or bag-ofwords, these methods often lack the discriminative power to distinguish between stylistically similar or within-family models. To overcome these limitations and expose the severity of vulnerability, we introduce INTERPOL, a model-driven identification framework that learns to distinguish target models from others

Key takeaways

  • arXiv:2603.15220v2 Announce Type: replace Abstract: Strict anonymity of model responses is a key for the reliability of voting-based leaderboards, such as LM Arena.
  • While prior studies have attempted to compromise this assumption using simple statistical features like TF-IDF or bag-ofwords, these methods often lack the discriminative power to distinguish between stylistically similar or within-family models.
  • To overcome these limitations and expose the severity of vulnerability, we introduce INTERPOL, a model-driven identification framework that learns to distinguish target models from others

Why it matters

This development is a reminder to test misuse and data-leak scenarios alongside speed and quality. Trust should come from testable controls and clear failure reporting, not protection claims alone.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗