A standard test — and an imperfect one.
Benchmarks score models on fixed question sets covering reasoning, coding, maths or knowledge, so releases can be compared on a common scale.
They decay: once a benchmark is public it leaks into training data, and scores rise without capability rising with them. This is called contamination.
Treat leaderboard position as a weak signal and your own evaluation on your own tasks as the strong one.