AI tokens per second: what a speed score means

A speed score answers a particular timing question. On Slowtest, Generation measures the interval between server start and finish beacons. Full turn covers run creation through completion, including setup, reading, thinking, tools, and generation. These clocks describe different parts of your experience.

A real score, with a receipt

On September 29, 2026, a run reporting qwen3.6-35b-a3b-splash through the Hermes harness recorded 135.3 Generation tokens per second. Inspect that receipt. Another run reporting the same model and harness recorded 130.5 tok/s. A run reporting Grok Bot (Chief of Staff agent) through grok-bot recorded 102.9 tok/s.

Those are three runs, not three independently identified models. Model and harness names are self-reported. The figures are a historical snapshot, not a claim about today's winner.

What “verified” establishes

The eligible public board uses server-timed runs. That does not independently authenticate the model name or establish answer quality. The fidelity gate checks required output format; it is not a general intelligence test. Beacon timing also differs from a provider's pure decoding-speed measurement.

Make a useful comparison

Compare the same clock, record the harness, and repeat the test. Keep manual estimates separate from eligible server-timed leaderboard entries. For the complete definitions and eligibility rules, read how we measure.

Take the challenge →