arXiv Artificial Intelligence

Scepsy: Serving Agentic Workflows Using Aggregate LLM Pipelines

Scepsy: Serving Agentic Workflows Using Aggregate LLM Pipelines

Quick summary

arXiv:2604.15186v2 Announce Type: replace-cross Abstract: Agentic workflows carry out complex tasks by orchestrating multiple large language models (LLMs) and tools. Serving them at a target throughput with low latency is hard because they are written in arbitrary agentic frameworks and their execution times are unpredictable: execution branches, fans out, or recurs in data-dependent ways. Since their LLMs often outnumber the available GPUs, they also oversubscribe GPUs. We describe Scepsy, a serving system that schedules arbitrary multi-LLM agentic workflows onto a GPU cluster. Scepsy exploit

Key takeaways

  • arXiv:2604.15186v2 Announce Type: replace-cross Abstract: Agentic workflows carry out complex tasks by orchestrating multiple large language models (LLMs) and tools.
  • Serving them at a target throughput with low latency is hard because they are written in arbitrary agentic frameworks and their execution times are unpredictable: execution branches, fans out, or recurs in data-dependent ways.
  • Since their LLMs often outnumber the available GPUs, they also oversubscribe GPUs.

Why it matters

“Scepsy: Serving Agentic Workflows Using Aggregate LLM Pipelines” exposes the compute, energy and supply-chain layer behind model competition. Capacity shifts can influence model costs, service availability and the ability of smaller companies to compete.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗