Teaching a model by adjusting its weights on data.
Training runs data through a model, measures how wrong the output is, and nudges billions of weights to be slightly less wrong. Repeat across trillions of tokens and a general-purpose model emerges.
A frontier training run occupies tens of thousands of accelerators for weeks and is the single largest line item in a lab's budget. It is a one-off cost per model version, unlike inference.
Pre-training produces raw capability; post-training steps such as instruction tuning and reinforcement learning from human feedback make it usable and safe.
Training capacity is bought, not rented: Microsoft — the largest single funder of frontier training compute — trades at 505.41, up 1.14% on the day, with data-centre spending the number analysts watch each quarter.