arXiv Artificial Intelligence

Delegating Authorization to Misaligned Agents: Coalitional Alignment and Safe Control

Delegating Authorization to Misaligned Agents: Coalitional Alignment and Safe Control

Quick summary

arXiv:2609.15803v1 Announce Type: cross Abstract: Long-running AI agents create a control problem: each action they take changes the state, which in turn affects the trajectory of future actions. If the agent is not fully aligned, then guaranteeing safety requires approving consequential actions before allowing them to be executed. But requiring human approval at every step makes attention a bottleneck. Delegating review to other AI agents raises the same alignment problem: the reviewers may themselves be misaligned. We identify a condition on a reviewing panel that is weaker than individual a

Key takeaways

  • arXiv:2609.15803v1 Announce Type: cross Abstract: Long-running AI agents create a control problem: each action they take changes the state, which in turn affects the trajectory of future actions.
  • If the agent is not fully aligned, then guaranteeing safety requires approving consequential actions before allowing them to be executed.
  • But requiring human approval at every step makes attention a bottleneck.

Why it matters

“Delegating Authorization to Misaligned Agents: Coalitional Alignment and Safe Control” shows why AI risk cannot be reduced to answer accuracy. Access controls, logging, human approval and incident response need to be designed into the workflow from the start.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗