arXiv Artificial Intelligence

Beyond Average Safety: Chance-Constrained LLM Fine-tuning

Beyond Average Safety: Chance-Constrained LLM Fine-tuning

Quick summary

arXiv:2609.29960v1 Announce Type: cross Abstract: Fine-tuning large language models on new objectives can improve helpfulness, instruction following, or domain-specific performance, but it can also induce regressions on safety-critical prompts. Existing safety-preserving fine-tuning methods typically control average safety loss or use weighted auxiliary penalties, which can obscure rare but severe failures. We propose a chance-constrained formulation for safety-preserving fine-tuning that limits the fraction of safety examples whose degradation relative to a reference model exceeds a prescribe

Key takeaways

  • arXiv:2609.29960v1 Announce Type: cross Abstract: Fine-tuning large language models on new objectives can improve helpfulness, instruction following, or domain-specific performance, but it can also induce regressions on safety-critical prompts.
  • Existing safety-preserving fine-tuning methods typically control average safety loss or use weighted auxiliary penalties, which can obscure rare but severe failures.
  • We propose a chance-constrained formulation for safety-preserving fine-tuning that limits the fraction of safety examples whose degradation relative to a reference model exceeds a prescribe

Why it matters

“Beyond Average Safety: Chance-Constrained LLM Fine-tuning” shows why AI risk cannot be reduced to answer accuracy. Access controls, logging, human approval and incident response need to be designed into the workflow from the start.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗