Abstract
Reliability in cloud AI infrastructure is crucial for cloud service providers, prompting the widespread use of hardware redundancies. However, these redundancies can inadvertently lead to hidden degradation, known as “gray failure”, for AI workloads, significantly affecting end-to-end performance and concealing performance issues, which complicates root cause analysis for failures and regressions. We introduce SuperBench, a proactive validation system for AI infrastructure that mitigates hidden degradation caused by hardware redundancies and enhances overall reliability. SuperBench features a comprehensive benchmark suite, capable of evaluating individual hardware components and representing most real AI workloads. It comprises a Validator that learns benchmark criteria to pinpoint defective components clearly. Additionally, SuperBench incorporates a Selector to balance validation time and issue-related penalties, enabling optimal timing for validation execution with a tailored subset of benchmarks. Through testbed evaluation and simulation, we demonstrate that SuperBench can increase the mean time between incidents by up to 22.61×. SuperBench has been successfully deployed in Azure production, validating hundreds of thousands of GPUs every year.
| Original language | English |
|---|---|
| Article number | 8 |
| Journal | ACM Transactions on Computer Systems |
| Volume | 44 |
| Issue number | 2 |
| DOIs | |
| State | Published - 24 Apr 2026 |
| Externally published | Yes |
Keywords
- AI infrastructure
- distributed training
- proactive validation
- Reliability
Fingerprint
Dive into the research topics of 'SuperBench: A Proactive Validation System for Improving Reliability of Cloud AI Infrastructure'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver