arXiv Artificial Intelligence

ShardMeter: Sharded and Geo-Distributed Training Without the Guesswork

ShardMeter: Sharded and Geo-Distributed Training Without the Guesswork

Quick summary

arXiv:2608.23840v1 Announce Type: cross Abstract: Training large-scale AI models often outgrows a single data center, demanding sharded, multi-cluster, and decentralized training. However, the huge space of resource allocations makes exhaustive benchmarking and manual tuning impractical, while performance depends on tightly coupled factors like model size, GPU memory, batch size, bandwidth, and sharding strategy. We introduce ShardMeter, a lightweight analytical performance model that predicts the end-to-end runtime of transformer-based workloads across arbitrary sharded, distributed, and even

Key takeaways

  • arXiv:2608.23840v1 Announce Type: cross Abstract: Training large-scale AI models often outgrows a single data center, demanding sharded, multi-cluster, and decentralized training.
  • However, the huge space of resource allocations makes exhaustive benchmarking and manual tuning impractical, while performance depends on tightly coupled factors like model size, GPU memory, batch size, bandwidth, and sharding strategy.
  • We introduce ShardMeter, a lightweight analytical performance model that predicts the end-to-end runtime of transformer-based workloads across arbitrary sharded, distributed, and even

Why it matters

AI progress is not only a software story. Chips, data centers and energy decisions help determine which models can operate economically and what end users ultimately pay.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗