arXiv Artificial Intelligence

Reference-Grafting Matches Fine-Tuning at Eliciting Sandbagged Capabilities

Reference-Grafting Matches Fine-Tuning at Eliciting Sandbagged Capabilities

Quick summary

arXiv:2608.29458v1 Announce Type: cross Abstract: Sandbagging, in which a model deliberately underperforms on an evaluation despite retaining the underlying capability, threatens the safety evaluations that frontier-model governance depends on. The Elicitation Game found that fine-tuning elicits hidden capability from sandbagging model organisms whereas additive activation steering fails. We revisit that verdict with reference-grafting, which sets an activation's coordinate along a contrast direction to the value it takes in an honest reference, at a small set of circuits chosen by active lear

Key takeaways

  • arXiv:2608.29458v1 Announce Type: cross Abstract: Sandbagging, in which a model deliberately underperforms on an evaluation despite retaining the underlying capability, threatens the safety evaluations that frontier-model governance depends on.
  • The Elicitation Game found that fine-tuning elicits hidden capability from sandbagging model organisms whereas additive activation steering fails.
  • We revisit that verdict with reference-grafting, which sets an activation's coordinate along a contrast direction to the value it takes in an honest reference, at a small set of circuits chosen by active lear

Why it matters

“Reference-Grafting Matches Fine-Tuning at Eliciting Sandbagged Capabilities” shows why AI risk cannot be reduced to answer accuracy. Access controls, logging, human approval and incident response need to be designed into the workflow from the start.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗