arXiv Artificial Intelligence

CounterSteer: Suppressing Indirect Prompt Injection with Activation Steering

CounterSteer: Suppressing Indirect Prompt Injection with Activation Steering

Quick summary

arXiv:2609.36570v1 Announce Type: cross Abstract: Indirect prompt injection makes an LLM agent treat untrusted retrieved text as instructions. We present CounterSteer, an inference-time defense that suppresses this behavior inside the model. Per model, a five-step recipe fits a residual-stream direction from paired episodes differing only in whether an embedded instruction is followed, and retains it only if it passes pre-specified causal and capability gates. At deployment, the direction is subtracted from every tool-result token during prefill. The edit is always on--there is no detection de

Key takeaways

  • arXiv:2609.36570v1 Announce Type: cross Abstract: Indirect prompt injection makes an LLM agent treat untrusted retrieved text as instructions.
  • We present CounterSteer, an inference-time defense that suppresses this behavior inside the model.
  • Per model, a five-step recipe fits a residual-stream direction from paired episodes differing only in whether an embedded instruction is followed, and retains it only if it passes pre-specified causal and capability gates.

Why it matters

“CounterSteer: Suppressing Indirect Prompt Injection with Activation Steering” shows why AI risk cannot be reduced to answer accuracy. Access controls, logging, human approval and incident response need to be designed into the workflow from the start.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗