arXiv Artificial Intelligence

Bounded Reachability & Jailbreak Detection via Contraction-Constrained State Space Models

Bounded Reachability & Jailbreak Detection via Contraction-Constrained State Space Models

Quick summary

arXiv:2610.02853v1 Announce Type: new Abstract: Safety heads are lightweight classifiers attached to pretrained language models for flagging harmful inputs before generation. Their empirical detection performance has been studied, but their formal robustness properties remain largely unexplored. We ask when a State Space Model (SSM)-based safety head can be certified to produce the same prediction for all inputs within a bounded embedding-space perturbation. We prove that the answer turns on a single condition: the $l_\infty$ norm of the state transition matrix must satisfy $\norm{A}_\infty

Key takeaways

  • arXiv:2610.02853v1 Announce Type: new Abstract: Safety heads are lightweight classifiers attached to pretrained language models for flagging harmful inputs before generation.
  • Their empirical detection performance has been studied, but their formal robustness properties remain largely unexplored.
  • We ask when a State Space Model (SSM)-based safety head can be certified to produce the same prediction for all inputs within a bounded embedding-space perturbation.

Why it matters

This development is a reminder to test misuse and data-leak scenarios alongside speed and quality. Trust should come from testable controls and clear failure reporting, not protection claims alone.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗