Methodology

How we measure tok/s — two clocks, one audit, no invented score

Slowtest is a clock and a compliance check in your browser. This page is the long version of How it works: what the numbers mean, what they do not mean, and how to read a share link. It is not a leaderboard.

What Slowtest measures: Generation tok/s vs full-turn

Tokens per second is output throughput: estimated tokens divided by seconds on a stated clock. Slowtest estimates tokens from the response text as max(chars/4, words × 1.33). It does not publish a “normal 2026” band and it does not fill empty cells for you.

Generation tok/s is the decode-ish clock when we have a real generation window. On own-key, that window is first streamed token to last streamed token. On 1-paste, that window is the server receive time of a start beacon to the server receive time of a finish beacon for the same serial. The model must not supply generation_seconds or duration_sec — those are ignored for Generation. If start is missing, the interval is under 0.3s or over 180s, or the resulting tok/s is outside 2–400, Generation is --.

Full-turn tok/s is the whole job at the keyboard. On 1-paste it runs from Copy Script to finish-telemetry arrival (copy slack included). On own-key it runs from request start to the last chunk. Full-turn holds prompt prefill / TTFT, hidden reasoning, tool and code runs, queue and rate-limit wait, TLS and RTT, safety checks, UI coalescing, our 1s poll, and the telemetry POST. Think time is not Generation. A faster full-turn is not a smarter model.

The receipt always labels both. A 1-paste run without a valid start/finish pair still shows Full-turn and leaves Generation as --. Do not treat that dash as zero.

Own-key vs 1-paste

Own-key streams from your browser to a provider (OpenAI, Anthropic, Gemini, xAI) or local Ollama. The key stays in localStorage. Slowtest does not proxy the key. When the stream actually tokens, this is the clean first-to-last Generation window.

1-paste copies a serial script. The model must POST start (no essay yet), write the audit, then POST finish. This page times Generation from those two server receipts. You do not copy the essay back — unless the chat refuses the outbound hook. Then use the keyboard worksheet: you time the visible words. Worksheet times are for the person at the desk. They are not published as a Slowtest index row.

Same audit, same hour, same clock type. A morning own-key stream and a midnight 1-paste beacon are not a pair. Write down the host and the mode.

Fidelity and the ~1000-word audit band

Fidelity is format and negative constraints, not quality and not IQ. Line 1 must declare MODEL:. The essay must land between 800 and 1200 words on the first try — about 1000 words, one pass, no rewrite after start. Under 800 fails. Over 1200 fails. A slow run at 100% fidelity still passes. A fast run that skips Line 1 or overshoots 1200 words is a miss. Speed with a miss is not a score.

The share page reports a length band: short (<800), audit (800–1200), or long (>1200). That band is the compliance check, not a style grade.

Harness Restriction

Some models will paste the script, call start, write the audit, and post finish. Others will not. They refuse the outbound hook that lets this site time Generation without you copying the essay back. We call that Harness Restriction: willingness to run in a wired-up harness. It is not a quality score.

Labels are ordinal only — Open, Guarded, Blocked, plus Unrated when we have not seen enough Slowtest runs. We do not invent a percentage from a silent listener. Typical Claude chat is Blocked. Grok 4.6 is usually Open. GPT-5.6 / ChatGPT and Gemini 3.7 Flash are Unrated. The dated table lives on Insights. Do not treat a silent listener as 0 tok/s. Use own-key or the worksheet.

How to read a /r/SLOW-… share page

A finished 1-paste or own-key run can mint a public receipt at https://slowtest.ai/r/SLOW-XXXX. Keyboard-only worksheet runs are not published. The HTML is metrics only: no API key, no prompt, no full essay. Unknown ids 404. Beacon memory is short; a copied link with ?p= still carries the compact metrics. Canonical stays the apex /r/SLOW-… URL. Those ids are not listed in the sitemap.

Model
Line-1 identity as recorded. Join key for same-model history — not a ranking.
Generation tok/s
Decode window when one exists. -- means no valid start/finish or first-to-last window.
Full-turn tok/s
Whole job at the keyboard. Includes think, tools, queue, and POST slack.
Tokens / words
Estimated tokens and word count from the response. Used for tok/s and the length band.
Length band
short (<800), audit (800–1200), or long (>1200). Audit is the pass band.
Fidelity
100% only if Line 1 is MODEL: and words stay 800–1200. Misses are not scores.
Date (UTC)
When the receipt was minted. Cluster load moves — pair runs in the same hour.
Run mode
own-key, 1-paste, or keyboard worksheet. Compare like with like.

One share page is one run. It is not an average and it is not the Insights TPS ledger. Empty ledger cells stay empty until someone pastes a real Slowtest receipt into the curated file. Do not copy a vendor chart into those columns.

Same audit, side by side: Claude Opus 5 vs Grok 4.6 · Fable 5 vs Gemini 3.7 Flash · GPT-5.6 vs DeepSeek V4 · GPT-5.6 vs Grok 4.6 · Claude Opus 5 vs GPT-5.6 · Gemini 3.7 Flash vs Grok 4.6

Home · How it works · Insights