Cascadia: A Control-Plane-Free Alternative to Hyperconverged AI Infrastructure
Quick summary
arXiv:2609.38697v1 Announce Type: cross Abstract: We present Cascadia, a system for serving large language models on fleets of commodity Intel AIPCs using their CPU, integrated-GPU, and NPU resources. Every node embeds ingress, scheduling, and execution; inference requests require no dedicated routing control plane. Nodes join a libp2p QUIC mesh using CA-issued ed25519 admission certificates, gossip signed capabilities, exchange live load over direct peer streams, and route OpenAI-compatible requests to eligible peers. An operator-run certificate authority handles admission and fleet managemen
Key takeaways
- arXiv:2609.38697v1 Announce Type: cross Abstract: We present Cascadia, a system for serving large language models on fleets of commodity Intel AIPCs using their CPU, integrated-GPU, and NPU resources.
- Every node embeds ingress, scheduling, and execution; inference requests require no dedicated routing control plane.
- Nodes join a libp2p QUIC mesh using CA-issued ed25519 admission certificates, gossip signed capabilities, exchange live load over direct peer streams, and route OpenAI-compatible requests to eligible peers.
Why it matters
“Cascadia: A Control-Plane-Free Alternative to Hyperconverged AI Infrastructure” exposes the compute, energy and supply-chain layer behind model competition. Capacity shifts can influence model costs, service availability and the ability of smaller companies to compete.

Member comments