arXiv Artificial Intelligence

NeuroBreak: Unveil Internal Jailbreak Mechanisms in Large Language Models

NeuroBreak: Unveil Internal Jailbreak Mechanisms in Large Language Models

Quick summary

arXiv:2509.03985v3 Announce Type: replace-cross Abstract: Jailbreak attacks bypass the safety alignment of large language models (LLMs) to elicit harmful outputs, yet the vast parameter space makes diagnosing the underlying failure mechanisms extremely challenging. We present NeuroBreak, a visual analytics system that helps experts progressively unpack jailbreak mechanisms from layer-level semantics down to neuron-level behaviors. A layer-wise probing pipeline traces how harmful representations evolve across layers, while a dual-dimensional character--behavior categorization reveals each safet

Key takeaways

  • arXiv:2509.03985v3 Announce Type: replace-cross Abstract: Jailbreak attacks bypass the safety alignment of large language models (LLMs) to elicit harmful outputs, yet the vast parameter space makes diagnosing the underlying failure mechanisms extremely challenging.
  • We present NeuroBreak, a visual analytics system that helps experts progressively unpack jailbreak mechanisms from layer-level semantics down to neuron-level behaviors.
  • A layer-wise probing pipeline traces how harmful representations evolve across layers, while a dual-dimensional character--behavior categorization reveals each safet

Why it matters

“NeuroBreak: Unveil Internal Jailbreak Mechanisms in Large Language Models” shows why AI risk cannot be reduced to answer accuracy. Access controls, logging, human approval and incident response need to be designed into the workflow from the start.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗