
$ ./epoch-training --whoami
The machines that train AI
/ explained by someone who ran them.
Supercomputing and AI-infrastructure training for engineers who want to stop treating the cluster like a black box. Taught by someone who architected top-100 supercomputers, handled storage and configuration management behind hurricane forecasts, and helped run the grid that confirmed the Higgs boson.
// the whole discipline is one question: correct, or fast? the answer is yes — and the engineering that makes "both" possible is the entire subject.
Track record
Selected publications
Peer-reviewed, at the venues that matter.
Not a marketing claim — published work on scheduling and petascale storage at the Cray User Group and the Parallel Data Storage Workshop (SC).
-
S. Alam, H. El-Harake, K. Howard, N. Stringfellow, F. Verzelloni6th Parallel Data Storage Workshop (PDSW '11), held with SC11 · 2011 · CSCS
-
G. Renker, N. Stringfellow, K. Howard, S. Alam, S. TrofinoffCray User Group (CUG) · 2011 · CSCS
What you'll learn
How the machines actually work — not how the vendor slide says they do.
The training splits along the axis every HPC engineer eventually meets: the half that keeps the job correct, and the half that makes it fast. Real systems need both.
Storage & the I/O wall
Parallel filesystems (Lustre, VAST), why 10,000 ranks hitting one disk is its own discipline, and why the storage decision quietly decides the whole run.
GPUs & the roofline
Why AI lives on GPUs, memory-bound vs compute-bound, and why your accelerator is "fast" and still 90% idle.
The network & MPI
Interconnects, collectives, and why a $200M machine spends most of its time waiting on the fabric — not computing.
Scaling & its limits
Amdahl's law, "just add nodes" and where it breaks, and how to find the bottleneck you actually have instead of the one you assumed.
Reliability at scale
Checkpointing, node failure, silent data corruption — everything that goes wrong at hour 47, and how the pros plan for it.
Precision & performance
FP8/FP16/FP32, mixed precision, and the throughput-vs-correctness tradeoff that can quietly ruin a model.
Trainings
Workshops & cohorts
Dates for the next cohorts are being scheduled. Join a waitlist and you'll be first to hear — before it opens publicly.
Loading trainings…
Free · no course to buy first
The Supercomputing Field Guide
Twenty concepts every infrastructure engineer should understand about the machines AI runs on — one page each, in plain language. It's the on-ramp to the full training. Drop your email and I'll send it.
One email with the guide, then occasional field notes on HPC/AI infrastructure. No spam. Unsubscribe anytime.

