Why repeat an AI benchmark before claiming a win?
Bragging rights are more interesting when other people can inspect the evidence. Slowtest's first stored results include two runs reporting the same model, qwen3.6-35b-a3b-splash, and the same Hermes harness.
| Receipt | Generation tok/s | Setup seconds |
|---|---|---|
| SLOW-GBZENP | 135.3 | 0.198 |
| SLOW-JYTDAR | 130.5 | 0.189 |
The higher Generation score is 4.8 tok/s above the lower score, about 3.7% relative to 130.5. Two observations cannot tell us the typical speed, explain the difference, or prove statistical significance.
Publish enough context to make the challenge fair
Share the receipt alongside the model and harness you used. Keep the timing method visible. If you change providers, harnesses, or settings, say so; those changes can affect the experience being measured.
A missing value is not zero
The first receipt records 0.556 seconds to the first beacon. The second has no recorded first-beacon value. That missing value is neither zero latency nor evidence that the second run was faster to start.
Keep the history
Share multiple receipts instead of only your best run. A high-score board celebrates peaks; a collection of repeated runs gives readers more context. The live Top 10 may change as new eligible runs arrive.