Why Tokens Per Second Matters Less Than You Think
Why a wrong fast answer costs more than a right slow one, and why the ecosystem isn’t ready 2026-10-11 16:47:27 Author: b1twis3.ca(查看原文) 阅读量:8 收藏

Why a wrong fast answer costs more than a right slow one, and why the ecosystem isn’t ready for it anyway

Oct 10, 2026 · @b1twis3 | HM

Do we need faster models?

Every couple of days the community posts a screenshot of a local model running at 40, 50, 60 or even 300 tps. The replies go like “how and what kind of hardware is needed” Less replies go by “What did the extra speed actually buy us?”.

For most engineering based work, inference, running models locally, the honest answer is perhaps little. But past 20 to 40 tps, the bottleneck now leaves the model. It moves to the user reading the decoded text or output, to the workflow/tool that reads the output or built around it, and above all to the work of checking whether the answer is right, not just false positive.

In this article, I make an argument. A wrong fast answer costs you more than a right slow one, because verification, triage, allocating resource for that, dominate the real cost of using a model. However, there might be fields where speed genuinely matters because their ecosystem is.

The token per second numbers shared are usually single stream decode, meaning how fast one request emits tokens after the first one appears. It is one of three speeds that matter, and often the least important of the three.

Prefill speed is how fast the model reads your prompt before it starts answering. It matters when you paste in long files, codebases, or chat history.

Single stream decode is how fast one answer comes out once it starts. It matters when one person is watching one reply.

Aggregate throughput is the total output across many requests running at once. It matters for batch jobs, multi agent setups, and shared servers.

The three can move independently. A four node cluster (Dgx Sparks for instance) serving a 744B MoE might show 29 tok/s single stream, 70 tok/s aggregate across 8 streams, and 550 tok/s prefill. Swap the quantization or the engine and any one of those can jump while the other two stay flat or fall.

Decode or prose speed dominates the conversation because it is easy to observe and measure, it fits in one number, and looks cool. It hides the two things that usually matters and decide whether the session was productive in your field, such as, how long you have waited before anything appeared, and whether what appeared was correct.

This mean, human are the slowest reader or consumer in the loop. Because the average adult reads English nonfiction silently at about 238 words per minute, according to a review of 190 studies (Brysbaert, 2019). At roughly 0.75 words per token, that is about 5.3 tok/s. Fiction runs slightly faster, around 5.8 tok/s.

The same paper notes that reading above 1,000 words per minute with real comprehension is not possible, because the eye can only take in so much per fixation. That ceiling works out to about 22 tok/s. So, this means a model with 25 tok/s prose is already faster than a reader who does not exist (yet).

Test it yourself. You ask for a 600 word explanation, about 800 tokens.

Decode speedGeneration timeReading time at 238 wpmWhen you finish reading
25 tok/s32 s151 s151 s
100 tok/s8 s151 s151 s

Both models write faster than you read, so you finish at the same time either way; the fast one just finishes writing sooner. With code, the effect is even stronger, because people read code much more slowly than prose.

Now consider this claim, which I tested and observed a lot, at work, or during my personal projects. The time a model spends generating is usually the cheapest part of using it. the expensive parts and real cost come after. This includes, reading the answer, triaging it, and paying for it when it turns out to be wrong, not practical, theoretical only or simply wrong.

Example: Reviewing a suspicious function

When you ask an LLM model whether a parser, decoder, deserializer has a memory related vulnerability. The answer might cost ~1K tokens. Below is a mimic or illustrative when you compare a smaller fast (100 tok/s) with a larger slow one (20 tok/s).

Fast setupSlow setup
Decode speed100 tok/s20 tok/s
Generation time10 s50 s
Chance the answer is right60%85% (More parameters)
Your time to verify5 min5 min
Cost when wrong (notice, dig in, redo)15 min15 min
Expected time per task11.2 min8.1 min
Per 100 tasks18.6 h13.5 h

The fast setup saves 40 seconds of generation per task and loses almost 4 minutes to rework. Over 100 tasks it costs about five extra hours, despite decoding five times faster.

And this is the friendly case, where every wrong answer gets caught. An uncaught one is worse, a false positive escalated to a vendor, a missed bug that ships, or a confident claim in a report that someone else builds on. Those costs have no upper bound.

On local hardware, speed is usually bought with something else. The common levers are a smaller model, fewer bits per weight (4 bit down to 2 bit), a smaller context, or a lower reasoning setting such as low or off efforts. Each one can raise decode speed, and each one can raise the chance of a wrong answer.

That trade is often invisible. Benchmarks report tok/s precisely and quality loosely, so a 40% speed gain looks like a clear win while a few points of accuracy loss looks like noise.

What happens when generation is cheap and triaging is not (Bug Bounty Example)

The curl project ended its HackerOne bug bounty at the end of January 2026 after being overwhelmed by low quality, AI generated vulnerability reports (BleepingComputer). Daniel Stenberg has said that until early 2025 only about one in six security reports to curl were real (The New Stack). Each fake report took seconds to generate and hours of a small team’s time to disprove.

That is the formula at scale. Faster generation would only have made it worse, more bug bounty reports per hour, the same handful of maintainers, the same hours per report.

In METR’s 2025 randomized trial, 16 experienced open source developers took 19% longer on real issues when AI tools were allowed. They had predicted a 24% speedup and still believed they had been sped up by 20% afterwards (METR, 2025). METR now treats those numbers as out of date and thinks later tools likely do speed developers up (METR, Feb 2026). The lesson is the perception gap, fluent, fast output feels productive even when review and cleanup eat the gain.

Hence, the more expert the user, the higher the verification cost and the lower the tolerance for errors. A security researcher, a reviewer or a scientist cannot skim and accept. They have to trace every claim back to the code, the log or the paper. For that kind of user, the most valuable property of a model is not how fast it talks. It is how often they can trust what it says without redoing the work themselves.

The ecosystem is not ready to use the speed

Even if speed were free and correctness unchanged, most of it would be wasted. Almost everything around the model runs slower than the model already does. Compilers, test suites, fuzzers, container startup and network calls do not get faster because the model did. A test suite that takes 90 seconds takes 90 seconds whether the patch arrived in 5 seconds or 30. A web fetch still waits on the remote server.

In practice, many agent steps are dominated by these waits. Faster decode only shortens the gaps between them. Real work passes through approvals such as a permission prompt, a code review, a security sign off, a change advisory board. A firewall rule review drafted in 10 seconds or 60 seconds still waits for the same weekly approval meeting.

The same applies downstream. Reviewers, maintainers and analysts absorb output at human speed. The curl example shows what happens when generation outruns review, the queue grows, quality drops, and the people at the end of the pipe burn out.

Speed matters when nobody reads the tokens as they arrive, or when the model is the only slow part of the loop. Notice which metric each case actually needs, most of them are not single stream decode.

Field or useExampleWhy speed mattersMetric that matters
Reasoning models8,000 hidden thinking tokens before the answer5.3 min at 25 tok/s, 1.8 min at 75 tok/s, and you can’t read alongSingle stream decode
Voice and live assistantsSpoken back and forthA pause of a few seconds breaks the conversationTime to first token
Code autocompleteInline suggestions while typingSuggestions must land before you type past themTime to first token
Tight agent loopsFast tools, cached context, many short stepsDecode becomes the largest slice of each stepSingle stream decode
Batch automationAlert triage, log classification, evals, synthetic dataVolume, with no human waiting on any single itemAggregate throughput

Another example could be a a SOC analyst wants 50,000 alerts classified in one batch, about 300 tokens of output each, or 15M tokens. At 25 tok/s single stream that takes about 167 hours. At 70 tok/s aggregate across parallel requests it takes about 60 hours. For volume work you scale out with batching and more parallel requests.

Reasoning models are the strongest counterargument to this article. When most of the output is hidden thinking, decode time is pure waiting. Even there, the cheaper fix is often controlling reasoning effort, so the model stops thinking for 8,000 tokens about a question that needed 800.

What to optimize instead

If the goal is less time to a trusted answer, the levers line up roughly in this order:

  1. Use the largest model and the most bits per weight your memory allows, with enough context headroom for real work.
  2. Long prompts are the real wait in local work. Prefix caching, keeping stable context at the front of the prompt, and fast prefill pay off on every single request.
  3. Match thinking budget to the question. A low setting for lookups and drafting, high only for hard analysis or complex research.
  4. Speculative decoding and faster interconnects improve speed without touching quality. Track acceptance rate, not just peak tok/s.
  5. For evals, triage and data jobs, tune concurrency and batching rather than single stream speed.
  6. Tests, reproducible proofs of concept, citations to exact lines and logs.

Sources


文章来源: http://b1twis3.ca/why-tokens-per-second-matters-less-than-you-think/
如有侵权请联系:admin#unsafe.sh