arXiv Artificial Intelligence

Divide, Consult, Conquer: Capability Laundering Through Aligned LLMs

Divide, Consult, Conquer: Capability Laundering Through Aligned LLMs

Quick summary

arXiv:2609.15383v1 Announce Type: cross Abstract: Language model safety is typically evaluated one interaction at a time. We show that a weaker, unaligned model can split a harmful task into benign-looking subproblems, consult a stronger aligned model independently on each, and combine the answers locally. We call this attack capability laundering. Unlike a jailbreak, no single response is a harmful task. We measure consultation-aided uplift using tasks that a raw frontier model solves, the aligned frontier refuses, and the unassisted orchestrator fails. We evaluate GPT-5.5, Claude Opus 4.8, a

Key takeaways

  • arXiv:2609.15383v1 Announce Type: cross Abstract: Language model safety is typically evaluated one interaction at a time.
  • We show that a weaker, unaligned model can split a harmful task into benign-looking subproblems, consult a stronger aligned model independently on each, and combine the answers locally.
  • We call this attack capability laundering.

Why it matters

“Divide, Consult, Conquer: Capability Laundering Through Aligned LLMs” shows why AI risk cannot be reduced to answer accuracy. Access controls, logging, human approval and incident response need to be designed into the workflow from the start.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗