Thousands of accelerators wired to act as one machine.
A cluster is a fleet of accelerators joined by very high-speed interconnect so a single model can be trained across all of them at once.
Interconnect quality matters as much as chip count: if the parts cannot exchange gradients fast enough, adding more of them stops helping.
Reliability is a design problem at this scale — with tens of thousands of parts, something fails constantly, so training checkpoints frequently.