inferecon.com
THE ECONOMICS OF INFERENCE

v0.3 — defaults captured 2026-07

LCOI calculator

Levelized cost of inference per million tokens. Pick a hardware, model, and region preset to see a plausible $/M-tokens figure split into GPU amortisation, electricity, and cooling. Every numeric field is overridable — presets only fill defaults.

Configuration

New SXM units still $35–40k; used SXM trades $15–28k (some listings as low as $6–15k). Residual cut from 30% to 25% for the 3-year-forward projection — the secondary-market bid-ask has widened materially since April, and Blackwell ramp accelerates Hopper depreciation. Rental cross-check: H100 on-demand has fallen to roughly $2.00–2.64/hr, spot as low as $1.43/hr (Spheron, 2026-06/07) — a rough rent-vs-buy breakeven at these used prices lands in the 17,500–24,500 hour range, i.e. 2–3 years at 30–40% utilisation.

3,000 tok/s prefill / 310 tok/s decode on NVIDIA H100 SXM.

EIA Electricity Monthly Update (industrial, 2025 average)

~15% inter-node overhead.

Share of wall-clock time the GPU is producing tokens.

26% of GPU time on prefill. Affects sessions/yr and cost/session — not the $/M token price.

23% of tokens by count. Changing token counts shifts cost/session, not $/M. See Advanced to move the per-token price.

Cost per million tokens
$0.95

Blended — cost allocated by GPU time (26% prefill / 74% decode).

in$0.31out$3.05
CapExFacilityElectricityCoolingOpEx

Cost breakdown

Component$/yrShare
GPU amortisation$348,34871.7%
Facility overhead$65,19013.4%
Electricity$7,5691.6%
Cooling$3,0270.6%
OpEx$62,00012.8%
Total$486,134100%

OpEx: fixed $30,000 + marginal $32,000.

Annual input tokens
395.6B
Annual output tokens
118.7B
Cost allocation in / out
$124,547 / $361,587
Combined tok/s / GPU
999
Annual sessions
395.6M
Cost per session
$0.0012
Utilisation peak / avg
60% / 60%

Related Resources: Read the main insights of this calculator and see assumptions for sources, dates, and methodology caveats.

← Back to home