Why Spiky Inference Traffic Breaks the Dedicated GPU Math

Aug 04, 2026 07:00 AM - 3 weeks ago 454

Introduction

A dedicated GPU moving spiky LLM conclusion postulation clears a specific, calculable throughput floor, 1,910 billable tokens per 2nd sustained, earlier it thumps per-token billing; the derivation and its utilization-percentage balanced travel later successful this article. This article applies that level to llama3.3-70b-instruct connected DigitalOcean, useful done 3 postulation shapes pinch definitive hourly distributions, and covers the 3 capacity patterns, scheduled capacity, a reserved level pinch serverless overflow, and axenic serverless, that spiky workloads really take between.

Pricing basis. This article uses the DigitalOcean H200 GPU Droplet on-demand complaint effective August 1, 2026, $4.47 per GPU-hour, pursuing the complaint alteration announced successful Upcoming GPU Pricing Updates (DigitalOcean, published July 21, 2026). All dollar figures and crossover percentages beneath bespeak this complaint unless a fig is explicitly branded arsenic the prior, pre-August-1 complaint for comparison.

The modular proposal is that high-volume workloads postgraduate from serverless to dedicated capacity. That proposal is correct for dependable postulation and incomplete for spiky traffic, because it treats measurement arsenic the deciding adaptable erstwhile the deciding adaptable is whether you know, successful advance, which hours the GPU will beryllium busy. A dedicated GPU is worthy renting only for hours you tin support supra the sustained floor, and you tin only rent precisely those hours if you cognize erstwhile they are. Genuinely unpredictable bursts cannot beryllium scheduled around; predictable ones can, and that quality decides whether the dedicated GPU mathematics breaks aliases holds, not the size of the peak-to-trough ratio aliases the monthly token measurement alone.

The clearest measurement to spot this is to clasp the postulation identical and alteration only what you cognize astir it. Take 1 GPU serving an 8-hour regular highest astatine afloat load and thing the remainder of the day. Held astir the clock, that GPU is engaged 33.3% of the period (8 of 24 hours), beneath the 46.9% crossover derived later successful this article, and it costs $3,263.10 a period against a serverless-equivalent measure of $2,318.37 for the aforesaid tokens, a nonaccomplishment of $944.73. Rented only for the 8 hours it is really needed, the aforesaid hardware serving the aforesaid workload costs $1,087.70 for those hours, a redeeming of $1,230.67, 53.1% cheaper than the serverless-equivalent cost. Same hardware, aforesaid workload, different bill. The only adaptable that changed is whether you paid for the 16 hours you did not need, and the only point that lets you skip paying for them is knowing successful beforehand which 8 hours to buy.

DigitalOcean’s ain guidance already points astatine this distinction. From the Dedicated Inference documentation:

Choose serverless conclusion complete dedicated conclusion erstwhile you request to get started quickly without managing immoderate components down an conclusion endpoint, don’t person a civilization exemplary to big aliases optimize, aliases person unpredictable aliases spiky conclusion traffic.

This article operationalizes that guidance alternatively than arguing against it: it gives you the nonstop throughput floor, the archetypes wherever the guidance holds and wherever it gets much nuanced, and the capacity shape to scope for erstwhile you cognize your traffic’s shape. Scope: this article covers llama3.3-70b-instruct FP8 connected a DigitalOcean H200, billed by full tokens (input positive output). Self-hosting decisions independent of postulation shape, quantization tradeoffs, and batch-size tuning are covered successful the companion costs framework and are retired of scope here.

Key Takeaways

  • Effective August 1, 2026, a DigitalOcean H200 GPU Droplet and an H200 Dedicated Inference endpoint costs the aforesaid $4.47/hr, truthful they now stock 1 break-even point: 1,910 sustained billable tokens per second, 46.9% of the measured 4,071.6 tok/s total-token ceiling, to hit DO Serverless Inference astatine $0.65/1M tokens. Before that date, the GPU Droplet ran astatine $3.44/hr pinch a separate, little 36.1% threshold; the prime betwixt the 2 products is now operational, not financial.
  • What breaks the dedicated GPU mathematics is predictability, aliases the deficiency of it: you cannot opportunity successful beforehand which hours the GPU will beryllium busy. One GPU serving an identical 8-hour regular highest loses money held astir the timepiece (33.3% mean utilization, beneath the 46.9% crossover) and wins decisively rented only for those 8 hours (100% utilization of billed time). Predictability, not the size of the spike, decides the outcome.
  • Holding the peak-to-trough ratio fixed astatine 10:1 and changing only the monthly measurement still changes the verdict, but the post-rate-change margins move against dedicated: astatine Archetype 1’s scale, serverless now wins outright by 27.5%, not a near-coin-flip. At astir 3.5x the volume, only the single-GPU reserved level still thumps axenic serverless; the 2-GPU level and axenic dedicated sized to highest some now suffer to serverless. Ratio unsocial ne'er wished the winner; the measurement astatine which highest first exceeds 1 GPU’s ceiling did, and astatine the caller complaint that ceiling has to beryllium cleared by a wider separator to salary off.
  • A user postulation style pinch unpredictable 2-hour bursts cannot usage scheduled capacity astatine all, moreover though its burst is short capable that a scheduled model astatine that load would different beryllium worthy building. Scheduling requires knowing the model successful advance; erstwhile the burst tin onshore astatine immoderate hour, the only measurement a dedicated GPU covers it is to tally continuously, which returns the style to the grounded 14.8%-utilization level test.
  • Scheduled capacity, a GPU rented only for a predictable business-hours model pinch the remainder of the time routed to serverless, saves 37.3% against an all-serverless baseline for a B2B postulation style sitting adjacent the crossover. The redeeming narrows to 34.5% erstwhile realistic per-day provisioning overhead (30 minutes of billed creation and teardown time) is priced in, and DigitalOcean bills GPU Droplets from creation, not from readiness.

The Sustained Floor: The Only Number That Matters

Every comparison successful this article reduces to 1 question: does a fixed hr of GPU clip present much than 1,910 billable tokens per second, connected average? Below that line, DigitalOcean Serverless Inference is cheaper for those tokens. Above it, the dedicated GPU is cheaper, and the separator grows pinch each further constituent of utilization.

Methodology: Why Total Billable Tokens

DigitalOcean prices llama3.3-70b-instruct connected Serverless Inference astatine a level $0.65 per 1 cardinal tokens, input and output, symmetric. Because some sides of the measure are priced identically, full billable tokens (input positive output) is simply a valid communal portion for comparing serverless walk against dedicated GPU throughput, and each number successful this article, connected some sides of each comparison, counts the aforesaid thing. This is the 1 methodology determination the article makes, and it is stated present once.

That ground matters because the measured throughput anchor comes from a benchmark tally astatine a 1,024-input/1,024-output token ratio, not from a ratio needfully typical of your traffic. Total-token throughput astatine that ratio was 4,071.6 tok/s (2,036 tok/s output positive astir the aforesaid successful input, measured connected a azygous H200 moving llama3.3-70b-instruct FP8 pinch vLLM 0.24.0). If you tally that benchmark yourself, return the Total token throughput (tok/s) statement alternatively than the Output token throughput (tok/s) statement erstwhile applying the level look successful this article, since the level is stated per billable token. A different input:output ratio changes prefill-versus-decode equilibrium and does not simply halve aliases double that full figure. If your accumulation ratio departs substantially from 1:1, benchmark your ain configuration pinch the harness successful the cost framework’s methodology section earlier trusting the 1,910 tok/s level astatine look value.

Environment disclaimer. The 2,036 tok/s and 4,071.6 tok/s figures successful this article were measured connected a azygous DigitalOcean H200 GPU Droplet moving llama3.3-70b-instruct FP8 pinch vLLM 0.24.0 astatine 1,024-token inputs and outputs. A different model, quantization level, GPU generation, serving model version, aliases token-length distribution produces different throughput. Treat these arsenic a reproducible reference constituent for this circumstantial configuration, not a cosmopolitan capacity guarantee.

The Floor Identity

This derivation is independent of throughput measurement, it depends only connected the GPU’s hourly complaint and the serverless per-token rate:

floor_tokens_per_second = gpu_hourly_rate / (serverless_rate_per_token * 3600) floor = 4.47 / ((0.65 / 1_000_000) * 3600) print(f"{floor:,.0f} billable tokens per second")

Output

1,910 billable tokens per second

A DigitalOcean H200 GPU Droplet must prolong 1,910 billable tokens per second, averaged complete the billing period, earlier it costs little per token than DO Serverless. Below that, each further 2nd of idle capacity is billed astatine the afloat $4.47/hr complaint and delivers zero tokens against it.

Converting the Floor to a Utilization Percentage

The level becomes a percent erstwhile you disagreement it by the GPU’s measured total-token ceiling:

total_tps = 4071.6 # measured full (input + output) throughput, azygous H200 gpu_hourly = 4.47 # H200 GPU Droplet, effective August 1, 2026; besides the existent Dedicated Inference rate serverless_rate = 0.65 / 1_000_000 floor = gpu_hourly / (serverless_rate * 3600) print(f"H200 crossover (GPU Droplet and Dedicated Inference): {floor / total_tps:.1%}")

Output

H200 crossover (GPU Droplet and Dedicated Inference): 46.9%

Before August 1, 2026, the Dedicated Inference endpoint ($4.47/hr, managed serving stack included) needed a higher sustained level than the earthy GPU Droplet ($3.44/hr), because its $1.03/hr guidance premium had to beryllium earned backmost successful the aforesaid per-token comparison. Effective August 1, 2026, the GPU Droplet complaint roseate to lucifer Dedicated Inference astatine $4.47/hr (DigitalOcean, “Upcoming GPU Pricing Updates,” published July 21, 2026), truthful some products now stock the aforesaid 46.9% crossover and the aforesaid $3,263.10 monthly costs for 1 GPU. The prime betwixt them is now operational alternatively than financial: prime the GPU Droplet to negociate the serving stack yourself, aliases Dedicated Inference to person DigitalOcean negociate it, since neither carries a costs premium complete the different anymore. Every archetype and capacity shape later successful this article checks its utilization against 46.9% unless stated otherwise.

Reference Constants

Every number successful this article traces backmost to this table. Rounding convention: 4,071.6 tok/s and 730 hours per period throughout, truthful that tables work together pinch each different to the cent. Pricing reflects the H200 GPU Droplet complaint effective August 1, 2026, astatine which constituent it converged pinch the existing Dedicated Inference rate.

Constant Value
DO Serverless Inference, llama3.3-70b-instruct $0.65 per 1M tokens, input and output, symmetric
H200 GPU Droplet and Dedicated Inference endpoint $4.47 per GPU-hour, effective August 1, 2026 (both products now stock 1 rate)
Measured saturated output throughput 2,036 tok/s (FP8, vLLM 0.24.0, 1,024 successful / 1,024 out, azygous H200)
Measured saturated full throughput 4,071.6 tok/s (same run, input positive output)
Billing ground utilized passim this article total billable tokens
Sustained floor 1,910 billable tok/s
Crossover, H200 GPU Droplet and Dedicated Inference 46.9%
One H200 monthly (730 hr) $3,263.10
One H200 monthly capacity astatine 100% 10,700,164,800 billable tokens (10.70B)

Why Predictability Decides the Dedicated GPU Math, Not Peak-to-Trough Ratio

It is tempting to scope for the peak-to-trough ratio arsenic the azygous number that should thrust the serverless-versus-dedicated decision: a 3:1 style looks safe for dedicated, a 20:1 style looks evidently incorrect for it. That small heart does not past interaction pinch volume. Hold the ratio fixed astatine 10:1 and alteration only the monthly scale, and the correct architecture changes pinch it.

The B2B business-hours style successful the adjacent conception runs astatine 85% utilization for 8 hours and 8.5% for 16 hours, precisely a 10:1 peak-to-trough ratio, astatine a standard wherever 1 GPU covers the full peak. At that scale, astatine the post-August-1 GPU rate, axenic dedicated ($3,263.10) costs 27.5% much than axenic serverless ($2,364.74), a cleanable serverless triumph alternatively than a near-tossup. Now return a workload astatine the aforesaid 10:1 ratio (12,215 billable tok/s highest for 8 hours, 1,221 tok/s trough for 16 hours) scaled up until the highest exceeds 1 GPU’s 4,071.6 tok/s ceiling. At that scale, reserving 1 GPU and routing overflow to serverless costs $7,899.84 a period against $8,346.13 for axenic serverless, a 5.3% saving, and against $9,789.30 for axenic dedicated sized to peak, a 19.3% saving. Same ratio, larger volume, and the winning architecture is still not the aforesaid arsenic astatine Archetype 1’s scale, but the separator for the shape that wins is thinner than it was earlier the complaint change, and neither the 2-GPU level nor axenic dedicated sized to highest thumps serverless anymore. The ratio told you thing astir which broadside would win; the volume, specifically whether it pushed highest request past a azygous GPU’s ceiling, did.

Nor does spiky postulation arsenic specified make dedicated capacity uneconomic. A predictable spike tin beryllium scheduled around: reserve the GPU only for the hours it is needed, and it captures astir half the measure compared to moving it astir the timepiece (worked done successful the capacity patterns conception below). What really breaks the mathematics is not knowing which hours to reserve. An unpredictable spike, 1 that tin onshore astatine immoderate hr pinch nary beforehand signal, cannot beryllium scheduled astir astatine all, because scheduling requires knowing the model earlier it opens. The only measurement a dedicated GPU covers a genuinely unpredictable burst is to tally continuously and hold for it, which forces the GPU backmost onto the 24-hour mean utilization trial it was trying to avoid.

That is the existent thesis of this article: a dedicated GPU is worthy renting only for hours you tin support supra 1,910 billable tokens per second, and you tin only rent precisely those hours, and only those hours, if you cognize successful beforehand erstwhile they are. Traffic that spikes connected a schedule you tin sanction is simply a scheduling problem pinch a scheduling solution. Traffic that spikes without informing is not, sloppy of really ample aliases mini the spike is comparative to the trough.

Three Traffic Shapes, Three Verdicts

Utilization beneath is the fraction of 1 H200’s total-token capacity (4,071.6 billable tok/s), time-weighted crossed 24 hours. Each archetype has an definitive hourly shape, not conscionable an average, truthful you tin reproduce the number yourself. Traffic shapes are schematic constructions, chosen to bracket the crossover from some sides alternatively than drawn from immoderate measured dataset. They are worked examples showing really the level trial behaves crossed a scope of shapes, not precedent for your ain workload. Profile your ain hourly distribution earlier applying immoderate verdict here.

Hourly GPU utilization for each 3 archetypes plotted against the 46.9% crossover line, pinch Archetype 2's burst shown astatine 1 schematic position among galore imaginable hours.

Archetype Shape Avg utilization Monthly billable Serverless Dedicated Verdict
1. B2B business hours 85% for 8 hr, 8.5% for 16 hr 34.0% 3,638M $2,364.74 $3,263.10 serverless by 27.5%
2. Consumer viral spikes 90% for 2 hr, 8% for 22 hr, burst hours onshore astatine unpredictable times 14.8% 1,587M $1,031.67 $3,263.10 serverless by 68.4%
3. Steady API backend (control) flat 75% 75.0% 8,025M $5,216.33 $3,263.10 dedicated by 37.4%

Monthly serverless and dedicated costs for each of the 3 archetypes.

Archetype 1: B2B Business Hours

avg_util = (85 * 8 + 8.5 * 16) / 24 print(f"{avg_util:.1f}%")

Output

34.0%

This shape’s highest is 3,461 billable tok/s (85% of the 4,072 tok/s ceiling) for 8 hours a day, comfortably nether 1 GPU’s ceiling, truthful a azygous H200 covers each hr without overflow. This shape’s 34.0% mean utilization sits 12.9 percent points beneath the post-August-1 46.9% crossover, and the 27.5% spread betwixt serverless and dedicated is simply a cleanable serverless win, not noise: astatine the existent GPU rate, a business-hours style for illustration this 1 nary longer sits adjacent the threshold. The existent reply for a business-hours shape for illustration this 1 is still scheduled capacity, worked successful the adjacent section, which thumps some axenic options by 37.3%.

Archetype 2: Consumer Viral Spikes

avg_util = (2 * 0.90 + 22 * 0.08) / 24 print(f"{avg_util:.4f}")

Output

0.1483

At 14.8% mean utilization, this style is 68.4% cheaper connected serverless than connected dedicated, the widest spread of the 3 archetypes. That spread unsocial mightiness propose scheduling astir the 2-hour burst the measurement Archetype 1 schedules astir its 8-hour window. It cannot beryllium done here, and the logic why is the article’s halfway point: scheduled capacity cannot thief here. Renting a GPU only for the engaged hours requires knowing which hours those are. When the burst tin onshore astatine immoderate time, the only measurement a dedicated GPU covers it is to tally continuously, which returns you to the level trial this style fails.

A 2-hour burst is short capable that a scheduled model astatine that load would ordinarily beryllium worthy building, the aforesaid measurement Archetype 1’s 8-hour model is. The quality is wholly that Archetype 1’s model has a known commencement and extremity clip and this 1 does not.

Archetype 3: Steady API Backend (Control)

avg_util = 75.0 monthly_tokens = avg_util / 100 * 10_700_164_800 serverless_cost = monthly_tokens * 0.65 / 1_000_000 print(f"{monthly_tokens/1e6:,.0f}M tokens -> ${serverless_cost:,.2f} serverless vs $3,263.10 dedicated")

Output

8,025M tokens -> $5,216.33 serverless vs $3,263.10 dedicated

This is the honesty anchor. At level 75% utilization, good supra the 46.9% crossover, dedicated wins by 37.4%, though the separator has narrowed from earlier the complaint change: a pricier GPU-hour eats into dedicated’s advantage moreover astatine precocious utilization. This archetype supports 1 conclusion: dedicated infrastructure is the correct instrumentality for this shape. It does not support a broader indictment of dedicated infrastructure; thing successful this article argues against dedicated GPUs for steady, high-throughput workloads, it argues that postulation which cannot committedness a sustained level should not beryllium sized against one.

Three Capacity Patterns for Spiky Traffic

Once you cognize a workload’s shape, location are 3 architectures to take between, not two.

Pattern Use it when Fails when
Scheduled capacity Window timing is predictable and load wrong the model exceeds 1,910 billable tok/s Burst timing is unknown, truthful nary model tin beryllium scheduled
Reserved level pinch serverless overflow Peak request exceeds 1 GPU’s ceiling and the marginal reserved GPU’s ain mean utilization clears 46.9% The marginal GPU would tally only during a short highest window, truthful it sits beneath the crossover
Pure serverless Neither information holds Load clears the level astir the clock, wherever dedicated is simply cheaper

Scheduled Capacity, Worked connected Archetype 1

Archetype 1’s business-hours model is predictable, truthful it is the style scheduled capacity is built for: rent the GPU only for the 8-hour highest model and way the remaining 16 hours to serverless.

Strategy Monthly
All serverless $2,364.74
All dedicated $3,263.10
Scheduled capacity (GPU for the 8 hr window, serverless overnight) $1,481.82

Scheduled capacity saves 37.3% against the cheaper of the 2 axenic options (all-serverless). That fig assumes zero billed provisioning time, which is not realistic: DigitalOcean bills a GPU Droplet from the infinitesimal it is created, not from the infinitesimal it finishes booting and is fresh to serve. Publishing the provisioning sensitivity keeps the redeeming fig honest.

Billed provisioning per day Monthly Saving
None $1,481.82 37.3%
10 minutes $1,504.48 36.4%
20 minutes $1,527.14 35.4%
30 minutes $1,549.80 34.5%
gpu_hourly = 4.47 serverless_rate = 0.65 / 1_000_000 total_tps = 4071.6 days_month = 730 / 24 gpu_cost = gpu_hourly * 8 * days_month trough_tokens = 0.085 * total_tps * 3600 * 16 * days_month trough_cost = trough_tokens * serverless_rate scheduled_base = gpu_cost + trough_cost for extra_min in (0, 10, 20, 30): extra_cost = gpu_hourly * (extra_min / 60) * days_month full = scheduled_base + extra_cost redeeming = (2364.74 - total) / 2364.74 * 100 print(f"{extra_min:>2} min provisioning: ${total:,.2f}/mo, {saving:.1f}% redeeming vs. all-serverless")

Output

0 min provisioning: $1,481.82/mo, 37.3% redeeming vs. all-serverless 10 min provisioning: $1,504.48/mo, 36.4% redeeming vs. all-serverless 20 min provisioning: $1,527.14/mo, 35.4% redeeming vs. all-serverless 30 min provisioning: $1,549.80/mo, 34.5% redeeming vs. all-serverless

The array does not value everything a scheduled-capacity architecture costs successful practice: regular create-and-destroy automation to build and maintain, a acold endpoint astatine the commencement of each window, teardown verification against orphaned resources that support billing aft a book fails partway, and nary guarantee that GPU capacity is disposable successful your region the infinitesimal a model opens. At a 34.5% to 37.3% margin, that operational overhead is worthy carrying.

Reserved Floor pinch Serverless Overflow

This shape only becomes chopped from axenic dedicated erstwhile highest request exceeds a azygous GPU’s ceiling. Archetype 1’s highest is 3,461 billable tok/s against a 4,072 tok/s ceiling, truthful 1 GPU covers each hour, nary overflow ever occurs, and the shape collapses into plain axenic dedicated astatine $3,263.10. The shape only becomes a chopped action astatine a standard wherever highest really exceeds the ceiling.

Larger workload: highest 12,215 billable tok/s for 8 hours, trough 1,221 billable tok/s for 16 hours, 12.84B billable tokens per month, the aforesaid 10:1 peak-to-trough ratio arsenic Archetype 1 astatine astir 3.5 times the volume.

Strategy Monthly
All serverless $8,346.13
Pure dedicated (3 GPUs sized to peak) $9,789.30
Reserved floor, 2 GPUs positive overflow $8,844.57
Reserved floor, 1 GPU positive overflow $7,899.84

Hourly utilization of GPU #1 and GPU #2 crossed the aforesaid workload, each plotted against the 46.9% crossover line.

One reserved GPU wins, and the logic why is the determination norm to return from this pattern. A reserved GPU ever serves up to its afloat ceiling whenever request exists, it is ne'er throttled beneath its ceiling to manufacture overflow, truthful the mobility is ne'er really the workload averages retired overall. The mobility is whether that circumstantial GPU’s ain mean utilization clears the 46.9% crossover:

  • The first reserved GPU runs astatine its 4,071.6 tok/s ceiling done each 8 highest hours and still carries the full 1,221 tok/s trough overnight, because the trough ne'er exceeds what 1 GPU tin serve. Its ain mean utilization crossed the period is 53.3%, supra the 46.9% crossover, truthful it earns its rate.
  • A 2nd reserved GPU would only ever tally during the 8-hour peak, because the first GPU already absorbs the afloat trough by itself. That is 33.3% work astatine best, beneath the crossover, truthful a 2nd GPU loses money moreover though the workload arsenic a full is ample capable to look for illustration an evident dedicated-capacity candidate.

This is simply a sharper illustration of the marginal-GPU norm than it was earlier the complaint change. At the anterior $3.44/hr rate, axenic dedicated sized to highest ($7,533.60) and moreover the 2-GPU level ($7,340.77) some hit axenic serverless ($8,346.13); the quality betwixt the options was a matter of degree. At the existent $4.47/hr rate, only the 1-GPU level still wins: axenic dedicated now costs $9,789.30 and the 2-GPU level $8,844.57, some supra axenic serverless. The marginal-utilization figures themselves do not change, GPU #1 still runs astatine 53.3% and GPU #2 still runs astatine 33.3% duty, since those dangle connected postulation style alternatively than price, but the cushion against the crossover has shrunk: GPU #1’s utilization now clears the crossover by 6.4 points, down from 17.2 points earlier the complaint change. Adding a GPU past the 1 that clears the level is now much apt to suffer to serverless than it was before.

Do not publication the trough level itself arsenic the requirement. The trough unsocial does not request to prolong 1,910 tok/s; a reserved GPU fills first and takes the busiest hours disposable to it, truthful what matters is the marginal GPU’s ain mean utilization crossed the afloat month, not whether the trough by itself clears the floor.

Pure Serverless

Pure serverless wins whenever neither of the different 2 conditions holds, meaning the model cannot beryllium scheduled and nary azygous GPU’s marginal utilization would clear the crossover. Archetype 2 is the cleanable example: its burst is short and unpredictable, truthful it cannot beryllium scheduled, and it ne'er generates capable sustained load for moreover 1 reserved GPU to clear 46.9%, truthful a reserved level does not thief either. For this shape, axenic serverless is the correct answer, not simply a fallback.

Cold Starts and Burst Latency connected Serverless

Cost is not the only axis. Serverless trades burst latency for elasticity, and that tradeoff is dedicated capacity’s morganatic counterargument. A reserved GPU that is already lukewarm serves the first petition of a spike astatine the aforesaid latency arsenic the thousandth. A serverless endpoint whitethorn not.

Two chopped effects hide down the building “cold start,” and they request abstracted measurement because they person different causes and different mitigations:

  • Idle acold start. Time to first token connected the first petition aft a quiet period, erstwhile weights whitethorn request staging and capacity whitethorn request scheduling. This is what astir group mean by acold start, and it is measured pinch a azygous petition aft a deliberate idle window.
  • Queue hold nether concurrency. When a burst arrives faster than the serving excavation admits it, requests hold to beryllium scheduled into a batch alternatively than being served connected arrival. Continuous batching makes this businesslike successful aggregate, but an individual request’s clip to first token past reflects its position successful the queue alternatively than the model’s load time. The signature is unmistakable: respective requests person their first token astatine the aforesaid wall-clock instant, truthful measured clip to first token falls arsenic nonstop clip rises.

Conflating the 2 produces misleading numbers. A trial that fires 32 concurrent requests instantly aft an idle model measures some astatine erstwhile and attributes the full to acold start.

Measure them separately against your ain account:

  • Steady-state clip to first token astatine debased concurrency, arsenic the baseline.
  • Time to first token connected a azygous petition aft an idle model of 15 to 30 minutes, which isolates idle acold start.
  • Time to first token crossed a concurrency ramp from a lukewarm state, which isolates queue behavior. Record the nonstop timestamp of each request, not conscionable its latency, because the wall-clock presence shape is what separates queueing from acold start.
  • Recovery clip backmost to the steady-state median erstwhile the burst clears.

Report percentiles alternatively than means, and authorities the ramp style alongside immoderate burst figure, because a number without a ramp is not reproducible. Run the customer successful the aforesaid region arsenic the endpoint. A distant customer adds its information travel to each measurement and inflates the baseline much than the burst, which compresses the very ratio you are trying to observe.

This article does not people a azygous burst latency figure, because that number is circumstantial to relationship tier, region, model, and clip of day. The protocol supra is what makes your ain measurement defensible.

How This Compares to Other DigitalOcean Inference Cost Benchmarks

This article’s level look and its measured throughput anchor travel from the cost model piece, which derives the wide effective-cost-per-token look for dedicated GPU inference. This article applies that look specifically to postulation predictability alternatively than to utilization successful the abstract.

That model portion states its crossover arsenic 72.2% wherever this article now states 46.9%. As of this writing, that spread has 2 independent causes, not one, and repricing unsocial would not adjacent it. The first origin is the token basis: the model derives costs per output token, utilizing the benchmark’s 2,036 tok/s output throughput, while this article derives costs per billable token, utilizing the aforesaid benchmark’s 4,071.6 tok/s full throughput, because DigitalOcean Serverless bills input and output alike for llama3.3-70b-instruct. At the benchmark’s 1:1 input-to-output ratio, full throughput is doubly output throughput, truthful a period expressed per billable token sits astatine half the utilization of the aforesaid period expressed per output token, connected its ain accounting for a facet of two. The 2nd origin is the pricing date: this article uses the H200 GPU Droplet complaint effective August 1, 2026, $4.47/hr, pursuing DigitalOcean’s complaint alteration (DigitalOcean, “Upcoming GPU Pricing Updates,” published July 21, 2026); the model piece’s published 72.2% fig predates that alteration and, arsenic of this writing, has not been repriced, truthful it still reflects the anterior $3.44/hr rate. If the model portion repriced to $4.47/hr without besides correcting its token basis, its crossover would emergence to astir 93.8% (the aforesaid output-token-basis mathematics down the Dedicated Inference fig earlier successful this article), which would widen the spread betwixt the 2 pieces alternatively than adjacent it. Treat the model piece’s 72.2% fig arsenic old connected some counts until it is updated, and recompute your ain crossover from your ain existent complaint and your ain token ground alternatively than reconciling the 2 published numbers against each other.

A related DigitalOcean tutorial, Serverless vs. Dedicated vs. Self-Hosted LLM Inference Cost (published July 10, 2026), measures Qwen3-32B connected MI300X and reports a 22% to 48% duty-cycle break-even. That is not the aforesaid period arsenic this article’s 46.9% crossover, now shared by some the GPU Droplet and Dedicated Inference, and the 2 are not straight comparable: different model, different GPU, different hourly rate, and that portion measures work rhythm successful the absurd wherever this portion measures whether postulation predictability lets you deed a work rhythm astatine all. DigitalOcean’s August 1, 2026 pricing update covers AMD arsenic good arsenic NVIDIA GPU Droplets, truthful the MI300X complaint that portion was priced against besides changes connected that date; corroborate whether its 22-48% fig has been repriced earlier treating it arsenic current, for the aforesaid logic this article flags its ain comparison to the model portion above.

Published crossover figures alteration pinch model, GPU, quantization, serving configuration, and the serverless complaint they are measured against. A duty-cycle break-even measured connected 1 exemplary and accelerator is not straight comparable to a token-floor period measured connected another. Compare methodology earlier comparing thresholds, and measurement your own.

For the mechanics of routing overflow postulation from a reserved GPU to serverless, arsenic utilized successful the reserved-floor shape above, spot How to Use the Inference Router.

This article treats input and output tokens arsenic a azygous billable portion because DigitalOcean prices them identically for llama3.3-70b-instruct. That symmetry is not universal. For models that complaint a premium connected output, the costs ground splits and the input:output ratio of your workload starts to matter independently of its postulation shape, which is covered successful The Hidden Cost of Output Token Pricing for Llama 3.3 70B.

FAQ

Is Serverless aliases Dedicated Cheaper for Bursty LLM Traffic?

It depends connected whether the bursts are predictable, not connected really ample they are. A dedicated GPU needs to prolong 1,910 billable tokens per second, 46.9% of 1 H200’s ceiling, averaged complete the hours it is billed, to hit DO Serverless Inference astatine $0.65 per 1M tokens for llama3.3-70b-instruct. If you tin schedule a GPU for precisely the hours your postulation clears that floor, dedicated wins for those hours. If the engaged hours cannot beryllium predicted successful advance, moving the GPU continuously to drawback them usually fails the level test, and serverless is cheaper.

What Sustained Throughput Justifies a Dedicated GPU?

For a DigitalOcean H200 GPU Droplet moving llama3.3-70b-instruct FP8, 1,910 billable tokens per 2nd sustained, averaged crossed the billing period, is the break-even constituent against DO Serverless Inference. That is 46.9% of the measured 4,071.6 tok/s total-token ceiling, astatine the GPU Droplet complaint of $4.47/hr effective August 1, 2026. The H200 Dedicated Inference endpoint shares that aforesaid $4.47/hr complaint and truthful the aforesaid 46.9% threshold; earlier the complaint change, the GPU Droplet ran astatine $3.44/hr pinch a little 36.1% threshold, and Dedicated Inference carried a separate, higher level for its managed-serving premium.

Can a Dedicated GPU Scale to Zero?

No. DigitalOcean bills GPU Droplets and Dedicated Inference endpoints from creation, and that billing continues whether the assets is actively serving traffic, idle, aliases powered off. Powering disconnected a GPU does not extremity the charge.

Only destroying a GPU Droplet aliases Dedicated Inference deployment stops billing connected it. If your capacity scheme depends connected regular creation and teardown, verify the teardown measurement really completes; an orphaned GPU from a grounded teardown book keeps billing astatine the afloat hourly rate.

Does a Hybrid Setup Always Save Money?

No. A reserved level pinch serverless overflow only saves money erstwhile the marginal reserved GPU’s ain mean utilization, crossed the afloat month, clears the 46.9% crossover. A workload ample capable to request aggregate GPUs astatine highest does not automatically warrant each of them: successful this article’s larger-workload example, the first GPU clears 53.3% utilization and earns its complaint astatine $7,899.84 against $8,346.13 for axenic serverless, while a 2nd GPU astatine 33.3% work does not, and adding it pushes the full to $8,844.57, supra axenic serverless. At the existent rate, moreover axenic dedicated sized to highest ($9,789.30) loses to serverless, truthful the norm cuts harder than it utilized to: use it per GPU, not to the workload’s average, and do not presume that because 1 GPU clears the floor, adding much capacity is free money.

How Do I Calculate the Crossover for My Own Traffic?

Run the level personality pinch your ain GPU’s hourly complaint and your provider’s per-token rate: floor_tokens_per_second = gpu_hourly_rate / (serverless_rate_per_token * 3600). Divide that consequence by your GPU’s measured saturated total-token throughput to get a percentage. Measure your ain saturated throughput pinch the benchmark harness successful the cost framework’s methodology section alternatively than reusing the 4,071.6 tok/s fig successful this article, which is circumstantial to llama3.3-70b-instruct FP8 connected a azygous H200 astatine a 1,024:1,024 input:output ratio.

What Are DigitalOcean Serverless Inference Rate Limits?

Two different pages picture 2 different limits, and some matter. The Inference APIs reference states a level 5,000 requests per hr and 250 requests per infinitesimal per OAuth token. The Inference Limits page documents gradual requests-per-minute and tokens-per-minute quotas that standard pinch your relationship tier, from 120 RPM and 500K to 750K TPM astatine Tier 1 up to 4,500 RPM and 3.5M to 70M TPM astatine Tier 5. Check your account’s tier connected the Resource Limits page successful the Control Panel earlier sizing a burst against either figure.

The astir reliable root is your ain relationship alternatively than either page. Every Serverless Inference consequence carries x-ratelimit-limit-requests, x-ratelimit-limit-tokens-per-minute, and x-ratelimit-limit-tokens-per-day headers, each pinch a matching remaining and reset value. Read your existent quota from those headers earlier sizing a burst, and expect them to beryllium the crushed truth if they disagree pinch the documentation.

Conclusion

This article covered 1 governing number for llama3.3-70b-instruct connected a DigitalOcean H200: a GPU Droplet needs 1,910 sustained billable tokens per second, 46.9% of its measured ceiling, to hit DO Serverless Inference astatine $0.65 per 1M tokens. It applied that level crossed 3 postulation archetypes, business-hours, viral-spike, and steady-state, and 3 capacity patterns, scheduled capacity, a reserved level pinch serverless overflow, and axenic serverless, utilizing the 8-hour regular highest illustration to show the aforesaid GPU serving the aforesaid workload losing money held astir the timepiece and winning decisively rented only for the hours it is needed. Across each case, what decided the result was ne'er the size of the spike; it was whether you knew, successful advance, which hours the GPU would beryllium busy.

Two things travel for sizing your ain traffic. Check measurement against a azygous GPU’s ceiling earlier trusting a peak-to-trough ratio alone: holding a 10:1 ratio fixed and scaling measurement up still moves the correct architecture, but astatine the existent GPU complaint the larger workload’s 2 multi-GPU options, axenic dedicated sized to highest and the 2-GPU reserved floor, some now suffer to axenic serverless; only the single-GPU reserved level still wins, and by a thinner separator than earlier the complaint change. And benignant by predictability earlier size: a predictable business-hours spike and an unpredictable viral spike, akin successful size, onshore connected other sides of the decision, 1 scheduled astir for a 37.3% redeeming and the different incapable to beryllium scheduled astir astatine all, because scheduling requires knowing the model earlier it opens.

Start pinch predictability, past volume. If you tin sanction the hours your load will clear the floor, scheduled capacity aliases a reserved level pinch overflow will hit some axenic options. If you cannot, DigitalOcean’s ain guidance already tells you wherever to start: serverless is built for postulation you cannot predict. Profile your sustained token level against your ain measured throughput, utilizing the look successful this article, earlier committing to reserved GPU capacity for a workload that spikes without warning.

You tin besides mention to the pursuing tutorials from our Inference successful Production bid to get started:

  1. Why Your LLM Bill Is 3× What You Expected
  2. How to Choose the Right LLM Model for Inference Use Case
  3. Prompt Caching successful Practice: From 7% to 74% Hit Rate
  4. Multi-Provider LLM Routing Is Not a Problem, It’s Your Architecture

References

  • Token Economics Across Traffic Profiles connected Dedicated GPUs — Derives the effective-cost-per-token look and benchmark harness this article’s 1,910 tok/s level builds on.
  • DigitalOcean Inference Mode Comparison for Your Each Use Case — Compares serverless, dedicated, batch, and Inference Router modes truthful you tin lucifer a hosting prime to your postulation style earlier moving the level math.
  • Serverless vs. Dedicated vs. Self-Hosted LLM Inference: When Self-Hosting Actually Gets Cheaper — Measures duty-cycle break-even connected a different exemplary and GPU, useful arsenic a cross-check against this article’s 46.9% crossover.
  • Dedicated vs. Serverless Inference arsenic You Scale — Covers erstwhile predictable baselines warrant migrating from serverless to dedicated capacity arsenic measurement grows.
  • Multi-Model API Cost Governance pinch the Inference Router — Shows really to way overflow postulation crossed exemplary tiers, the serverless broadside of the reserved-floor shape worked successful this article.
  • Metrics that Matter pinch Serverless Inference — Explains which latency, throughput, and costs metrics to measurement earlier sizing a burst against serverless complaint limits aliases SLAs.
  • Why Serverless Inference Consistency Varies connected the Same Model — Benchmarks idle acold commencement and queue-wait behaviour that the burst-latency conception supra points you to measurement yourself.
  • The Hidden Cost of Output Token Pricing for Llama 3.3 70B — Covers erstwhile input:output pricing asymmetry changes the level mathematics beyond the total-token billing ground utilized here.
  • Continuous Batching vs. Static Batching successful LLM Inference — Explains really continuous batching affects saturated throughput, the ceiling the 1,910 tok/s level is expressed arsenic a percent of.
  • Upcoming GPU Pricing Updates — Documents the August 1, 2026 H200 complaint alteration to $4.47/hr that repriced each crossover fig successful this article.

Creative CommonsThis activity is licensed nether a Creative Commons Attribution-NonCommercial- ShareAlike 4.0 International License.

More