v0.3 — defaults captured 2026-07
LCOI calculator
Levelized cost of inference per million tokens. Pick a hardware, model, and region preset to see a plausible $/M-tokens figure split into GPU amortisation, electricity, and cooling. Every numeric field is overridable — presets only fill defaults.
Configuration
New SXM units still $35–40k; used SXM trades $15–28k (some listings as low as $6–15k). Residual cut from 30% to 25% for the 3-year-forward projection — the secondary-market bid-ask has widened materially since April, and Blackwell ramp accelerates Hopper depreciation. Rental cross-check: H100 on-demand has fallen to roughly $2.00–2.64/hr, spot as low as $1.43/hr (Spheron, 2026-06/07) — a rough rent-vs-buy breakeven at these used prices lands in the 17,500–24,500 hour range, i.e. 2–3 years at 30–40% utilisation.
3,000 tok/s prefill / 310 tok/s decode on NVIDIA H100 SXM.
EIA Electricity Monthly Update (industrial, 2025 average)
~15% inter-node overhead.
Share of wall-clock time the GPU is producing tokens.
26% of GPU time on prefill. Affects sessions/yr and cost/session — not the $/M token price.
23% of tokens by count. Changing token counts shifts cost/session, not $/M. See Advanced to move the per-token price.
Blended — cost allocated by GPU time (26% prefill / 74% decode).
Cost breakdown
| Component | $/yr | Share |
|---|---|---|
| GPU amortisation | $348,348 | 71.7% |
| Facility overhead | $65,190 | 13.4% |
| Electricity | $7,569 | 1.6% |
| Cooling | $3,027 | 0.6% |
| OpEx | $62,000 | 12.8% |
| Total | $486,134 | 100% |
OpEx: fixed $30,000 + marginal $32,000.
- Annual input tokens
- 395.6B
- Annual output tokens
- 118.7B
- Cost allocation in / out
- $124,547 / $361,587
- Combined tok/s / GPU
- 999
- Annual sessions
- 395.6M
- Cost per session
- $0.0012
- Utilisation peak / avg
- 60% / 60%
Related Resources: Read the main insights of this calculator and see assumptions for sources, dates, and methodology caveats.