epoch.training← the field guide

Field Guide · Part III — The modern stack · 17

// the cost of a run

What a training run actually costs

The number you see quoted — "$X million to train" — is almost always just the GPU rental. The real bill includes interconnect, storage, power and cooling, and the runs that crashed at 80% and had to start over. Those line items are where budgets die.

Training compute has grown at a staggering pace — Sevilla et al. found that since the early 2010s the compute used in landmark models has doubled roughly every six months, far outstripping Moore's law. That curve is why cost stopped being a rounding error and became the thing that decides who can train at the frontier. But "cost" is more than a headline dollar figure. Let's build it from the ground up.

The obvious part: GPUs × hours × utilization

Start with the compute. You need a certain number of floating-point operations to train a model; divide by your hardware's real throughput to get GPU-hours, then multiply by price. The trap is the word real. Nobody sustains a GPU's peak FLOPS — after communication stalls, memory limits, and pipeline bubbles, large runs often sit at 30–50% Model FLOPS Utilization. Your effective cost per useful FLOP is therefore two to three times the marketing number.

Rough training cost (compute only):

   cost ≈ (GPU-hours) × (price per GPU-hour)

   GPU-hours = total_FLOPs / (peak_FLOPs/s × MFU × 3600)

Toy example — 2,000 GPUs, 30 days, $2/GPU-hr:
   2000 × 30 × 24 × $2  =  $2.88 M   ← rental only

At 40% MFU you paid for 100% of those hours but got 40%
of the compute. Halve MFU and you double the real bill.
The compute line item. Utilization (MFU) is the hidden multiplier — the same run at 20% vs 40% MFU costs twice as much for the same result.

The parts nobody quotes

Buy vs build, honestly

All of this feeds one decision: rent or own. Cloud GPU-hours carry a premium but zero commitment — right for bursty or one-off training. Owning hardware is cheaper per hour only if you keep it busy: a $30M cluster idle half the time costs double per useful hour, and now you own the megawatts, the cooling, the failed-run risk, and a depreciating asset in a market where next year's GPU is faster. The honest model isn't "cloud vs on-prem sticker price." It's total cost per useful FLOP at your real utilization, with power, people, and failed runs included. Do that arithmetic and the answer is usually less flattering to "just buy the GPUs" than the vendor slide suggests.

Why it matters: teams routinely under-budget training by ignoring utilization, power, and failed runs — then blow past the number by 2–3× and can't explain why. If you can build the full stack — GPU-hours ÷ real MFU, plus interconnect, storage, megawatts, and a restart budget — you can size a run before you spend, and make buy-vs-build a calculation instead of a vibe. That model is the single most useful spreadsheet in modern AI infrastructure.

Sources & where to go deeper

D. Patterson, J. Gonzalez, Q. Le, C. Liang, L.-M. Munguia, D. Rothchild, D. So, M. Texier & J. Dean, "Carbon Emissions and Large Neural Network Training", arXiv 2104.10350, 2021 — energy and carbon accounting for T5, GPT-3, and other large models; shows power is first-order.
J. Sevilla, L. Heim, A. Ho, T. Besiroglu, M. Hobbhahn & P. Villalobos, "Compute Trends Across Three Eras of Machine Learning", arXiv 2202.05924, 2022 — training compute doubling ~every 6 months since the early 2010s.

This is one page of twenty.

The workshops go deep on the real thing — scheduling, storage, interconnect, GPUs — hands-on, on real infrastructure, from someone who's run these machines at national scale.

See the trainings →