What if AI worked at 1.000.000 tokens per seconds?

Hacker News by 8 min read 39x views
What if AI worked at 1.000.000 tokens per seconds?

Share Post

01 / WHAT THE NUMBER MEANS

“A myriad a second” can average four things.

CONTEXT CAPACITYHow much fitsThe content a example can clasp in one request, including its reply. A size, not a speed.

INPUT PROCESSINGHow accelerated it readsPrompt tokens can be processed mostly in parallel. Huge inputs motionless obtain time, and the input charge is distinct from the output rate.

AGGREGATE THROUGHPUTHow much a scheme serves10,000 streams × 100 tokens/sec = 1,000,000 tokens/sec in total. Toy arithmetic: it says nothing concerning whenever any one stream finishes.

SINGLE-AGENT OUTPUT · OUR DIALHow accelerated one delegate writesOne delegate penning a myriad tokens a second. In norm generation, all token depends on the ones before it.

Anthropic’s Claude Opus 5.5 overview, for example, lists a 1M-token environment window, a much smaller output cap per petition and “moderate” related latency. Memory size is not penning speed, and it does not average a million-token answer fits in one request.

Engineering helps, inside limits. Services attain big totals by batching many requests, as the 2023 vLLM paper describes, and speculative decoding can obtain multiple drafted tokens in one pass. Neither removes a genuine sequence in which stage two needs the outcome of stage one.

A token is a chunk of text, frequently part of a word, so token counts are not term counts.

TRY IT / THE IMAGINED DIAL

Turn the dial.

Pick a workload, afterward glide the imagined speed from 100 to 1,000,000 tokens per second.

Imagined speed for one agent1,000,000 tokens/sec

1001K10K100K1M

OUTPUT TOKENS1,000,0001,000 narrative endings × 1,000 tokens each

WRITING ONLY1 secondat 1,000,000 tokens/sec

At 1,000,000 tokens/sec, 1,000 narrative endings (1,000,000 output tokens) obtain concerning 1 second to write.

All four budgets at the two ends of the dial
Writing-only period at explanatory speeds, not measurements.
WorkloadOutput tokens100 tokens/sec1,000,000 tokens/sec
1,000 narrative endings × 1,0001,000,0002 h 46 min 40 s1 second
40 app drafts × 20,000800,0002 h 13 min 20 s0.8 seconds
10,000 critiqued candidates × 3003,000,0008 h 20 min3 seconds
10,000 rehearsals × 1,00010,000,00027 h 46 min 40 s10 seconds

Imagined comparisons, not measurements of any model. Times figure generated output only: no reading, ranking, tool calls, permissions, tests, deployment, group or experiments. A prosperity is a penning allowance, not a completed product.

Notice

Parallel streams can accelerate these autonomous jobs too. But all outline motionless needs period on its own stream: matching total throughput does not equivalent completion time. One accelerated delegate matters most whenever all stage waits on the last.

02 / FOUR THOUGHT EXPERIMENTS

Cheap drafts move the difficult part.

Each prosperity counts written output only. None is a completed product, a verified finding or a genuine result.

A map of imaginable endings

1,000 × 1,000 = 1,000,000 tokens · 1 second of writing

Ask for an ending to your narrative and get a thousand: hopeful, dark, strange. Nobody says a thousand, so the helpful type maps them for you to explore, afterward blends the two you like.

Still scarce: taste. Only you cognize which ending is yours.

Software for a neighborhood tool library

40 × 20,000 = 800,000 tokens · 0.8 seconds of writing

Describe a lending app for shared tools and forty drafts be before you complete the sentence. Screens could adapt, from borrowing a ladder to a repair-day sign-up. Tests motionless run on their own clock, and if none checks the seven-day due date, all forty can continue during lending ladders for seventy.

Still scarce: a apparent spec, and tests of what “working” means.

Ten genuine experiments from ten thousand ideas

10,000 × 300 = 3,000,000 tokens · 3 seconds of writing

Hunting for a improved catalyst? An delegate could propose and critique ten thousand candidates. Then everything waits at the bench: reactions may run for hours, cultures for days, site trials for a season. Ideas from one example can additionally portion one blind spot.

Still scarce: bodily evidence, and choosing which ten experiments acquire lab time.

Rehearsing a difficult conversation

10,000 × 1,000 = 10,000,000 tokens · 10 seconds of writing

Before talking to your landlord concerning the lease, an delegate could perform the conversation ten thousand ways: stubborn landlord, generous landlord, you whenever tired. Use it akin a escape simulator, for custom and blind spots. Ten thousand rehearsals of the incorrect individual are a assured mistake, not a prophecy.

Still scarce: fidelity to the genuine person, and your own practice.

03 / THE SERIAL BOTTLENECK

10,000× faster penning is not 10,000× faster work.

Take the forty app drafts: 800,000 output tokens. Compare an explanatory 100 tokens per second alongside the imagined million, afterward add one fixed inspect following penning that speed does not touch, specified as a test run.

Fixed inspect following writing60 seconds

01 min1 hour1 day

Illustrative 100 tokens/sec2 h 14 min 20 s

Writing 99.3% · Checking 0.7%

Imagined 1,000,000 tokens/sec60.8 seconds

Writing 1.3% · Checking 98.7%

WritingFixed checkEach bar shows anywhere its own total goes.

WRITING SPEEDUP10,000×1,000,000 ÷ 100 tokens/sec

END-TO-END SPEEDUP132.57×8,060 ÷ 60.8 seconds

With 60 seconds of checking, the lot takes 2 h 14 min 20 s at 100 tokens/sec and 60.8 seconds at 1,000,000 tokens/sec. That is 132.57× faster end to end, and checking is 98.7% of the faster total.

The identical 800,000 tokens alongside four inspect times
Total period = penning + one fixed check, counted once.
Fixed check100 tokens/sec1,000,000 tokens/secEnd to end
None2 h 13 min 20 s0.8 seconds10,000×
60 seconds2 h 14 min 20 s60.8 seconds132.57×
15 min2 h 28 min 20 s15 min 1 s9.88×
24 h26 h 13 min 20 s24 h 1 s1.09×

Toy model: total period = output tokens ÷ speed + one fixed check, counted formerly for the entire lot (not per draft) following penning ends. It does not simulate parallel tests, cost, energy or outline quality.

With a one-minute check, the job drops from 8,060 to 60.8 seconds: concerning 132.57 times faster, not 10,000. At zero the complete 10,000× returns; at a day the acquire nearly vanishes. Whatever you do not speed up becomes nearly all the remaining time, the logic of Amdahl’s law. Real requests have additional specified steps: OpenAI’s latency guide notes that extremely ample prompts, tool calls and network trips add delays of their own.

04 / WHAT GETS PRECIOUS

When generation gets cheap, judgement gets precious.

Many drafts, one narrow gate Conceptual drawing. A broad grid of outline cards on the remaining flows toward a narrow gate. One cardstock passes through and is marked as chosen. It illustrates the argument, not data.
Conceptual drawing, not data: generation widens the options; judgement decides what passes.

Every test complete ends in the identical place. What stays scarce is judgment, wearing distinct hats:

  • TasteWhich choice is yours.
  • Specs and testsWhat “working” means.
  • PracticeSpeed cannot study it for you.
  • Physical evidenceThe earth answers at its own pace.
  • FidelityA simulation is lone as fine as its model.
  • Cost and energyIs this value operating at all?

Fast does not average cheap: all token runs on hardware person pays to power. More options do not justify a fine one either. Drafts from one example and one set of assumptions can be correlated, or all incorrect together. Volume measures output, not understanding.

05 / USE IT THIS WEEK

Before you ask for more, decide how you volition choose.

No imaginary dial required. Next period you hand activity to an AI agent:

  1. Write downward what “done” means.One testable sentence, akin “ladders are due rear in seven days,” not “handles loans.”
  2. Set criteria before study options.Two or three, so the most fluent outline does not win by default.
  3. Find the dilatory step.Name the inspect that speed volition not touch. Shorten or automate it anywhere possible, without skipping the validation all outcome needs.
  4. Ask for disagreement, not volume.Request options built on distinct assumptions, afterward ask what would create all of them wrong.

If your dilatory stage is deciding what to measure or which bets to make, that is scheme additional than tooling. Private consulting helps explain the larger scheme and anywhere your next move matters ↗

Sources, dates and boundaries.

This Field Note adapts ideas from echohive’s movie One Million Tokens a Second; it is an edited companion, not a transcript. The speed is imagined, and no origin below measures, claims or predicts it. They assistance lone the distinctions between capacity, reading, serving and writing.

Open the four sources
  1. Anthropic, Claude Opus 5.5 overview. Official documentation. Lists a 1M-token environment window, a distinct maximum output per petition and “moderate” related latency. It does not depict a myriad output tokens per second.
  2. Anthropic, Context windows. Official documentation. Describes the environment opening as operating recollection that includes the generated response: a capacity, not a speed.
  3. OpenAI, Latency optimization. Official API guide. Output generation commonly dominates latency; extremely ample prompts motionless matter, and tool or network calls add their own delays.
  4. Agrawal et al., Taming Throughput-Latency Tradeoff in LLM Inference alongside Sarathi-Serve. arXiv, 2024. A historic foundation, not a current benchmark: parallel immediate handling (prefill), token-by-token decoding and batching.

The 2023 vLLM document connected in division 01 is additionally a historic foundation. All workloads, speeds, budgets and the bottleneck example are explanatory arithmetic, not benchmarks, forecasts or merchandise specifications. Sources checked October 2, 2026.

Other Article Hacker News
↑
Close Right Ads
Close Left Ads