arXiv Artificial Intelligence

Alignment via Training Against Probes Without Losing Monitorability

Alignment via Training Against Probes Without Losing Monitorability

Quick summary

arXiv:2609.38645v1 Announce Type: cross Abstract: Models are usually aligned based on their observed outputs, using demonstrations, preference data, or reward signals. These objectives reward responses that look aligned. More capable models may learn to satisfy them without internalizing the intended behavior, for example by faking compliance during training. Such superficial compliance could be harder when the objective is defined on model internals rather than outputs. Therefore, we study probe-guided fine-tuning, using probes that detect undesired properties in model activations as a direct

Key takeaways

  • arXiv:2609.38645v1 Announce Type: cross Abstract: Models are usually aligned based on their observed outputs, using demonstrations, preference data, or reward signals.
  • These objectives reward responses that look aligned.
  • More capable models may learn to satisfy them without internalizing the intended behavior, for example by faking compliance during training.

Why it matters

The value of this work lies as much in how it was tested as in the claim itself. Sample design, baselines, uncertainty and replication help separate a laboratory result from real-world impact.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗