Trait-space Monitoring for Emergent Misalignment During Supervised Finetuning
Quick summary
arXiv:2606.07631v2 Announce Type: replace-cross Abstract: Emergent misalignment (EM) occurs when narrow finetuning induces dangerous behavior outside the finetuning task. Detecting this shift through repeated behavioral evaluation is costly, motivating our checkpoint-level monitoring from internal representations. We define a fixed coordinate system from seven alignment-relevant activation directions and use it to track representational drift during LoRA finetuning of four open-source 7-9B language models. Finetuning drift in this space exhibits a dominant axis that explains 78.6% of variance
Key takeaways
- arXiv:2606.07631v2 Announce Type: replace-cross Abstract: Emergent misalignment (EM) occurs when narrow finetuning induces dangerous behavior outside the finetuning task.
- Detecting this shift through repeated behavioral evaluation is costly, motivating our checkpoint-level monitoring from internal representations.
- We define a fixed coordinate system from seven alignment-relevant activation directions and use it to track representational drift during LoRA finetuning of four open-source 7-9B language models.
Why it matters
“Trait-space Monitoring for Emergent Misalignment During Supervised Finetuning” highlights the need for repeatable measurement rather than a single impressive demonstration. Independent validation across datasets and clearly stated limitations determine whether a result can guide product decisions.

Member comments