arXiv Artificial Intelligence

AUDITPLAN: Commit, Then Answer for Auditable Safety Alignment

AUDITPLAN: Commit, Then Answer for Auditable Safety Alignment

Quick summary

arXiv:2609.19325v1 Announce Type: cross Abstract: Safety tuning pipelines judge only the final answer, which makes it difficult to distinguish robust refusal from two undesirable shortcuts: blanket refusal on benign requests and polished but unfaithful safety rationales that do not actually constrain the answer. We propose AUDITPLAN, a single-model plan-then-answer approach where the model first emits a compact structured safety plan and then answers conditioned on it. The plan records a threat label, intended action, and explicit constraints, enabling machine-checkable auditing while remainin

Key takeaways

  • arXiv:2609.19325v1 Announce Type: cross Abstract: Safety tuning pipelines judge only the final answer, which makes it difficult to distinguish robust refusal from two undesirable shortcuts: blanket refusal on benign requests and polished but unfaithful safety rationales that do not actually constrain the answer.
  • We propose AUDITPLAN, a single-model plan-then-answer approach where the model first emits a compact structured safety plan and then answers conditioned on it.
  • The plan records a threat label, intended action, and explicit constraints, enabling machine-checkable auditing while remainin

Why it matters

“AUDITPLAN: Commit, Then Answer for Auditable Safety Alignment” shows why AI risk cannot be reduced to answer accuracy. Access controls, logging, human approval and incident response need to be designed into the workflow from the start.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗