arXiv Artificial Intelligence

Many-Tier Instruction Hierarchy in LLM Agents

Many-Tier Instruction Hierarchy in LLM Agents

Quick summary

arXiv:2604.09443v4 Announce Type: replace-cross Abstract: Large language model agents receive instructions from many sources-system messages, user prompts, tool outputs, other agents, and more-each carrying different levels of trust and authority. When these instructions conflict, agents must reliably follow the highest-privilege instruction to remain safe and effective. The dominant paradigm, instruction hierarchy (IH), assumes a fixed, small set of privilege levels (typically fewer than five) defined by rigid role labels (e.g., system > user). This is inadequate for real-world agentic settin

Key takeaways

  • arXiv:2604.09443v4 Announce Type: replace-cross Abstract: Large language model agents receive instructions from many sources-system messages, user prompts, tool outputs, other agents, and more-each carrying different levels of trust and authority.
  • When these instructions conflict, agents must reliably follow the highest-privilege instruction to remain safe and effective.
  • The dominant paradigm, instruction hierarchy (IH), assumes a fixed, small set of privilege levels (typically fewer than five) defined by rigid role labels (e.g., system > user).

Why it matters

“Many-Tier Instruction Hierarchy in LLM Agents” should be evaluated beyond branding and benchmark scores. Its practical importance will emerge in task accuracy, latency, unit cost, safety and integration with real workflows.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗