01 / WHAT THE NUMBER MEANS
“A myriad a second” can average four things.
CONTEXT CAPACITYHow much fitsThe content a example can clasp in one request, including its reply. A size, not a speed.
INPUT PROCESSINGHow accelerated it readsPrompt tokens can be processed mostly in parallel. Huge inputs motionless obtain time, and the input charge is distinct from the output rate.
AGGREGATE THROUGHPUTHow much a scheme serves10,000 streams × 100 tokens/sec = 1,000,000 tokens/sec in total. Toy arithmetic: it says nothing concerning whenever any one stream finishes.
SINGLE-AGENT OUTPUT · OUR DIALHow accelerated one delegate writesOne delegate penning a myriad tokens a second. In norm generation, all token depends on the ones before it.
Anthropic’s Claude Opus 5.5 overview, for example, lists a 1M-token environment window, a much smaller output cap per petition and “moderate” related latency. Memory size is not penning speed, and it does not average a million-token answer fits in one request.
Engineering helps, inside limits. Services attain big totals by batching many requests, as the 2023 vLLM paper describes, and speculative decoding can obtain multiple drafted tokens in one pass. Neither removes a genuine sequence in which stage two needs the outcome of stage one.
A token is a chunk of text, frequently part of a word, so token counts are not term counts.
TRY IT / THE IMAGINED DIAL
Turn the dial.
Pick a workload, afterward glide the imagined speed from 100 to 1,000,000 tokens per second.
Imagined speed for one agent1,000,000 tokens/sec
1001K10K100K1M
OUTPUT TOKENS1,000,0001,000 narrative endings × 1,000 tokens each
WRITING ONLY1 secondat 1,000,000 tokens/sec
At 1,000,000 tokens/sec, 1,000 narrative endings (1,000,000 output tokens) obtain concerning 1 second to write.
All four budgets at the two ends of the dial| Workload | Output tokens | 100 tokens/sec | 1,000,000 tokens/sec |
|---|---|---|---|
| 1,000 narrative endings × 1,000 | 1,000,000 | 2 h 46 min 40 s | 1 second |
| 40 app drafts × 20,000 | 800,000 | 2 h 13 min 20 s | 0.8 seconds |
| 10,000 critiqued candidates × 300 | 3,000,000 | 8 h 20 min | 3 seconds |
| 10,000 rehearsals × 1,000 | 10,000,000 | 27 h 46 min 40 s | 10 seconds |
Imagined comparisons, not measurements of any model. Times figure generated output only: no reading, ranking, tool calls, permissions, tests, deployment, group or experiments. A prosperity is a penning allowance, not a completed product.
Notice
Parallel streams can accelerate these autonomous jobs too. But all outline motionless needs period on its own stream: matching total throughput does not equivalent completion time. One accelerated delegate matters most whenever all stage waits on the last.
02 / FOUR THOUGHT EXPERIMENTS
Cheap drafts move the difficult part.
Each prosperity counts written output only. None is a completed product, a verified finding or a genuine result.
A map of imaginable endings
1,000 × 1,000 = 1,000,000 tokens · 1 second of writing
Ask for an ending to your narrative and get a thousand: hopeful, dark, strange. Nobody says a thousand, so the helpful type maps them for you to explore, afterward blends the two you like.
Still scarce: taste. Only you cognize which ending is yours.
Software for a neighborhood tool library
40 × 20,000 = 800,000 tokens · 0.8 seconds of writing
Describe a lending app for shared tools and forty drafts be before you complete the sentence. Screens could adapt, from borrowing a ladder to a repair-day sign-up. Tests motionless run on their own clock, and if none checks the seven-day due date, all forty can continue during lending ladders for seventy.
Still scarce: a apparent spec, and tests of what “working” means.
Ten genuine experiments from ten thousand ideas
10,000 × 300 = 3,000,000 tokens · 3 seconds of writing
Hunting for a improved catalyst? An delegate could propose and critique ten thousand candidates. Then everything waits at the bench: reactions may run for hours, cultures for days, site trials for a season. Ideas from one example can additionally portion one blind spot.
Still scarce: bodily evidence, and choosing which ten experiments acquire lab time.
Rehearsing a difficult conversation
10,000 × 1,000 = 10,000,000 tokens · 10 seconds of writing
Before talking to your landlord concerning the lease, an delegate could perform the conversation ten thousand ways: stubborn landlord, generous landlord, you whenever tired. Use it akin a escape simulator, for custom and blind spots. Ten thousand rehearsals of the incorrect individual are a assured mistake, not a prophecy.
Still scarce: fidelity to the genuine person, and your own practice.
03 / THE SERIAL BOTTLENECK
10,000× faster penning is not 10,000× faster work.
Take the forty app drafts: 800,000 output tokens. Compare an explanatory 100 tokens per second alongside the imagined million, afterward add one fixed inspect following penning that speed does not touch, specified as a test run.
Fixed inspect following writing60 seconds
01 min1 hour1 day
Illustrative 100 tokens/sec2 h 14 min 20 s
Writing 99.3% · Checking 0.7%
Imagined 1,000,000 tokens/sec60.8 seconds
Writing 1.3% · Checking 98.7%
WritingFixed checkEach bar shows anywhere its own total goes.
WRITING SPEEDUP10,000×1,000,000 ÷ 100 tokens/sec
END-TO-END SPEEDUP132.57×8,060 ÷ 60.8 seconds
With 60 seconds of checking, the lot takes 2 h 14 min 20 s at 100 tokens/sec and 60.8 seconds at 1,000,000 tokens/sec. That is 132.57× faster end to end, and checking is 98.7% of the faster total.
The identical 800,000 tokens alongside four inspect times| Fixed check | 100 tokens/sec | 1,000,000 tokens/sec | End to end |
|---|---|---|---|
| None | 2 h 13 min 20 s | 0.8 seconds | 10,000× |
| 60 seconds | 2 h 14 min 20 s | 60.8 seconds | 132.57× |
| 15 min | 2 h 28 min 20 s | 15 min 1 s | 9.88× |
| 24 h | 26 h 13 min 20 s | 24 h 1 s | 1.09× |
Toy model: total period = output tokens ÷ speed + one fixed check, counted formerly for the entire lot (not per draft) following penning ends. It does not simulate parallel tests, cost, energy or outline quality.
With a one-minute check, the job drops from 8,060 to 60.8 seconds: concerning 132.57 times faster, not 10,000. At zero the complete 10,000× returns; at a day the acquire nearly vanishes. Whatever you do not speed up becomes nearly all the remaining time, the logic of Amdahl’s law. Real requests have additional specified steps: OpenAI’s latency guide notes that extremely ample prompts, tool calls and network trips add delays of their own.
04 / WHAT GETS PRECIOUS
When generation gets cheap, judgement gets precious.
Every test complete ends in the identical place. What stays scarce is judgment, wearing distinct hats:
- TasteWhich choice is yours.
- Specs and testsWhat “working” means.
- PracticeSpeed cannot study it for you.
- Physical evidenceThe earth answers at its own pace.
- FidelityA simulation is lone as fine as its model.
- Cost and energyIs this value operating at all?
Fast does not average cheap: all token runs on hardware person pays to power. More options do not justify a fine one either. Drafts from one example and one set of assumptions can be correlated, or all incorrect together. Volume measures output, not understanding.
05 / USE IT THIS WEEK
Before you ask for more, decide how you volition choose.
No imaginary dial required. Next period you hand activity to an AI agent:
- Write downward what “done” means.One testable sentence, akin “ladders are due rear in seven days,” not “handles loans.”
- Set criteria before study options.Two or three, so the most fluent outline does not win by default.
- Find the dilatory step.Name the inspect that speed volition not touch. Shorten or automate it anywhere possible, without skipping the validation all outcome needs.
- Ask for disagreement, not volume.Request options built on distinct assumptions, afterward ask what would create all of them wrong.
If your dilatory stage is deciding what to measure or which bets to make, that is scheme additional than tooling. Private consulting helps explain the larger scheme and anywhere your next move matters ↗
Sources, dates and boundaries.
This Field Note adapts ideas from echohive’s movie One Million Tokens a Second; it is an edited companion, not a transcript. The speed is imagined, and no origin below measures, claims or predicts it. They assistance lone the distinctions between capacity, reading, serving and writing.
Open the four sources- Anthropic, Claude Opus 5.5 overview. Official documentation. Lists a 1M-token environment window, a distinct maximum output per petition and “moderate” related latency. It does not depict a myriad output tokens per second.
- Anthropic, Context windows. Official documentation. Describes the environment opening as operating recollection that includes the generated response: a capacity, not a speed.
- OpenAI, Latency optimization. Official API guide. Output generation commonly dominates latency; extremely ample prompts motionless matter, and tool or network calls add their own delays.
- Agrawal et al., Taming Throughput-Latency Tradeoff in LLM Inference alongside Sarathi-Serve. arXiv, 2024. A historic foundation, not a current benchmark: parallel immediate handling (prefill), token-by-token decoding and batching.
The 2023 vLLM document connected in division 01 is additionally a historic foundation. All workloads, speeds, budgets and the bottleneck example are explanatory arithmetic, not benchmarks, forecasts or merchandise specifications. Sources checked October 2, 2026.