Scepsy: Serving Agentic Workflows Using Aggregate LLM Pipelines
Quick summary
arXiv:2604.15186v2 Announce Type: replace-cross Abstract: Agentic workflows carry out complex tasks by orchestrating multiple large language models (LLMs) and tools. Serving them at a target throughput with low latency is hard because they are written in arbitrary agentic frameworks and their execution times are unpredictable: execution branches, fans out, or recurs in data-dependent ways. Since their LLMs often outnumber the available GPUs, they also oversubscribe GPUs. We describe Scepsy, a serving system that schedules arbitrary multi-LLM agentic workflows onto a GPU cluster. Scepsy exploit
Key takeaways
- arXiv:2604.15186v2 Announce Type: replace-cross Abstract: Agentic workflows carry out complex tasks by orchestrating multiple large language models (LLMs) and tools.
- Serving them at a target throughput with low latency is hard because they are written in arbitrary agentic frameworks and their execution times are unpredictable: execution branches, fans out, or recurs in data-dependent ways.
- Since their LLMs often outnumber the available GPUs, they also oversubscribe GPUs.
Why it matters
“Scepsy: Serving Agentic Workflows Using Aggregate LLM Pipelines” exposes the compute, energy and supply-chain layer behind model competition. Capacity shifts can influence model costs, service availability and the ability of smaller companies to compete.

Member comments