LiveSPX7,620-1.28%NDX29,127-1.41%US10Y4.961%+3.23%BTC76,816-1.24%ETH2,468-2.07%GOLD4,320-1.02%NVDA210.96-8.42%MSFT505.41+1.14%GOOGL349.39+3.23%
Advertisement
AI glossary

Latency

How long the user waits, split into two numbers.

Time to first token is how long before anything appears; tokens per second is how fast the rest streams. Users feel the first number as responsiveness and the second as fluency.

Long prompts push time to first token up because the whole context must be processed before generation starts. Model size and batching push tokens per second down.

Providers trade latency against cost by batching more requests onto one accelerator — cheaper per token, slower per user.

Calculators
Related terms