arXiv Artificial Intelligence

Overthinking: Amplifying Reasoning Weights to Extract Learned Secrets

Overthinking: Amplifying Reasoning Weights to Extract Learned Secrets

Quick summary

arXiv:2607.08173v2 Announce Type: replace Abstract: Black box auditing of language models is an essential pre-deployment tool, but it may miss subtle forms of misalignment and hidden information. To better elicit hidden information during an auditing process, we introduce \emph{overthinking}: the process of using reasoning task vectors to amplify the propensity to think out loud of reasoning models. Given the parameters of a non-reasoning instruct model $M$ and reasoning-distilled model $R$, we define the \emph{overthinking model} as $\boldsymbol{\theta}_{\mathcal{O}_\alpha} = \boldsymbol{\the

Key takeaways

  • arXiv:2607.08173v2 Announce Type: replace Abstract: Black box auditing of language models is an essential pre-deployment tool, but it may miss subtle forms of misalignment and hidden information.
  • To better elicit hidden information during an auditing process, we introduce \emph{overthinking}: the process of using reasoning task vectors to amplify the propensity to think out loud of reasoning models.
  • Given the parameters of a non-reasoning instruct model $M$ and reasoning-distilled model $R$, we define the \emph{overthinking model} as $\boldsymbol{\theta}_{\mathcal{O}_\alpha} = \boldsymbol{\the

Why it matters

This model development creates a new option for users and a new testing obligation for developers. A fixed evaluation set comparing quality, cost and failure behavior is more useful than launch claims.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗