arXiv Artificial Intelligence

Controlled Decoding Attacks on Black-Box LLMs

Controlled Decoding Attacks on Black-Box LLMs

Quick summary

arXiv:2609.36956v1 Announce Type: cross Abstract: Manipulating next-token probabilities during generation can bypass the safety alignment of large language models. Existing approaches, however, rely on access to model weights or numerical token probabilities and therefore do not apply to interfaces that return only sampled text. Reconstructing probabilities from sampled outputs offers a possible alternative, but finite sampling produces sparse and noisy estimates, while repeating this process at every generation step incurs substantial query costs. Our empirical observations suggest that large

Key takeaways

  • arXiv:2609.36956v1 Announce Type: cross Abstract: Manipulating next-token probabilities during generation can bypass the safety alignment of large language models.
  • Existing approaches, however, rely on access to model weights or numerical token probabilities and therefore do not apply to interfaces that return only sampled text.
  • Reconstructing probabilities from sampled outputs offers a possible alternative, but finite sampling produces sparse and noisy estimates, while repeating this process at every generation step incurs substantial query costs.

Why it matters

“Controlled Decoding Attacks on Black-Box LLMs” shows why AI risk cannot be reduced to answer accuracy. Access controls, logging, human approval and incident response need to be designed into the workflow from the start.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗