arXiv Artificial Intelligence

Harmful SFT Leaves a Continuous Trace in LLM Checkpoint Updates

Harmful SFT Leaves a Continuous Trace in LLM Checkpoint Updates

Quick summary

arXiv:2610.07518v1 Announce Type: cross Abstract: Safety auditing of post-trained large language models typically relies on model behavior, requiring model execution and depending on the coverage of available evaluations. This work asks a different question: Do the target behaviors optimized during supervised fine-tuning (SFT) leave readable evidence directly in checkpoint updates? We find that harmful-compliance SFT induces a continuous, objective-dependent ordering in checkpoint-update space. Using a reference geometry defined by pure harmful-compliance, safety-targeted, and benign-utility S

Key takeaways

  • arXiv:2610.07518v1 Announce Type: cross Abstract: Safety auditing of post-trained large language models typically relies on model behavior, requiring model execution and depending on the coverage of available evaluations.
  • This work asks a different question: Do the target behaviors optimized during supervised fine-tuning (SFT) leave readable evidence directly in checkpoint updates?
  • We find that harmful-compliance SFT induces a continuous, objective-dependent ordering in checkpoint-update space.

Why it matters

This development is a reminder to test misuse and data-leak scenarios alongside speed and quality. Trust should come from testable controls and clear failure reporting, not protection claims alone.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗