MADBench: Benchmarking the Security of Multi-Agent Debate
Quick summary
arXiv:2609.39146v1 Announce Type: new Abstract: Multi-agent debate (MAD) can improve large language model (LLM) reasoning by allowing multiple agents to exchange and critique their answers to the same task. However, the interactions that enable agents to correct mistakes can also spread adversarial errors and steer the agents toward an incorrect answer. Although some efforts have been made to examine particular attack types on MAD, systematic evaluation of MAD under diverse attacks remains limited. A central question is whether debate mitigates adversarial influence or amplifies it. In this pa
Key takeaways
- arXiv:2609.39146v1 Announce Type: new Abstract: Multi-agent debate (MAD) can improve large language model (LLM) reasoning by allowing multiple agents to exchange and critique their answers to the same task.
- However, the interactions that enable agents to correct mistakes can also spread adversarial errors and steer the agents toward an incorrect answer.
- Although some efforts have been made to examine particular attack types on MAD, systematic evaluation of MAD under diverse attacks remains limited.
Why it matters
This development is a reminder to test misuse and data-leak scenarios alongside speed and quality. Trust should come from testable controls and clear failure reporting, not protection claims alone.

Member comments