Delegating Authorization to Misaligned Agents: Coalitional Alignment and Safe Control
Quick summary
arXiv:2609.15803v1 Announce Type: cross Abstract: Long-running AI agents create a control problem: each action they take changes the state, which in turn affects the trajectory of future actions. If the agent is not fully aligned, then guaranteeing safety requires approving consequential actions before allowing them to be executed. But requiring human approval at every step makes attention a bottleneck. Delegating review to other AI agents raises the same alignment problem: the reviewers may themselves be misaligned. We identify a condition on a reviewing panel that is weaker than individual a
Key takeaways
- arXiv:2609.15803v1 Announce Type: cross Abstract: Long-running AI agents create a control problem: each action they take changes the state, which in turn affects the trajectory of future actions.
- If the agent is not fully aligned, then guaranteeing safety requires approving consequential actions before allowing them to be executed.
- But requiring human approval at every step makes attention a bottleneck.
Why it matters
“Delegating Authorization to Misaligned Agents: Coalitional Alignment and Safe Control” shows why AI risk cannot be reduced to answer accuracy. Access controls, logging, human approval and incident response need to be designed into the workflow from the start.

Member comments