AI Inference Engineer 2026 — Special Course¶
Phase 5 · ML Systems Engineering · Special Course
From transformer-execution fundamentals to dense-70B on Hopper to MoE-672B on Blackwell — the modern inference stack, end to end.
Up-to-date is not a side requirement. It is the discipline.
The inference layer has moved more in the last twelve months than the previous three years combined. FP4 went from research to native silicon. Disaggregated prefill/decode went from paper to production. MoE went from "interesting" to "the default architecture above 30B." Any course that does not lead with what shipped in 2025–2026 is already mis-training engineers.
This course is structured as four parts that can be read independently or as a sequence. Parts 1–3 stand on their own; together they walk the precision floor down (FP16 → FP8 → FP4), the architecture from dense to sparse (Llama / Qwen → DeepSeek / Qwen3-MoE), and the hardware up (single-GPU → 8× Hopper → GB200 NVL72 Blackwell). Part 4 is a different kind of chapter: one engine, one node, ~96 pull requests, and a measured 60× — the discipline of optimization as it actually happens, with every number traceable to the diff that produced it.
Layer mapping: L3–L8. Runtime / scheduler / kernels / collectives / fabric / observability — the full ML Systems stack as it applies to inference.
Role targets: AI Inference Engineer · GPU Runtime Engineer · LLM Runtime Optimization Engineer · Production Inference Engineer · MLSys Engineer.
Prerequisites:
- Phase 5 — ML Systems Engineering — Stage 0–3 — measurement discipline, runtime foundations, transformer execution internals, GPU kernels.
- Phase 5 — Edge AI — Edge LLM Inference Internals — GEMV vs GEMM, roofline basics, the decode bottleneck.
- Phase 3 — Neural Networks → Transformer Fundamentals — Q/K/V, attention, multi-head, the full block.
- Comfort reading Rust or C++ for kernel work, Python for runtime and benchmark glue.
What comes after: a reproducible inference benchmark repo for a model + runtime + hardware target of your choice, with a parity report against a published reference and a measured $/MTok cost line.
🧠 Interactive companion: LLM Inference Visualizer¶
A 3D, hands-on companion to this course: LLM Inference Visualizer — walk the forward pass of a dense decoder-only transformer (Qwen 2.5 7B/72B, Llama 3.3 70B), see each stage land memory-bound vs compute-bound on an NVIDIA H200 roofline, and slice the model across TP = 1/2/4/8 GPUs to watch the weights shard and the all-reduce cost grow.
It makes the core lessons of this course tangible — especially:
- Part 1 · Lecture 03 — Roofline, bandwidth, and the memory hierarchy — the roofline chart, decode on the memory-bound side.
- Part 2 · Lecture 01 — Anatomy of a 70B-class dense model — GQA, RoPE, RMSNorm, SwiGLU rendered to scale on the Llama/Qwen pair.
- Part 2 · Lecture 04 — Single-node multi-GPU serving (tensor parallelism) — TP sharding and the all-reduce collectives, visualized.
Run it locally: git clone https://github.com/ai-hpc/llm-inference-viz && cd llm-inference-viz && npm install && npm run dev → open http://localhost:3002/llm.
Course Map (4 parts, 27 lectures)¶
🧭 Part 1 — Fundamentals of AI Inference / MLSys (5 lectures)¶
The mental model, the metrics, the math, and the runtime landscape. Anyone who finishes Part 1 can read any model card in 2026 and predict its inference cost shape.
⚙️ Part 2 — Dense Decoder-Only Inference at Hopper (7 lectures)¶
The end-to-end Hopper stack for 70B-class dense models. Anchored on a Llama 3.3 70B ↔ Qwen 2.5 72B comparison so every concept lands on two concrete deployable systems.
🧬 Part 3 — MoE Inference at Blackwell (5 lectures)¶
The Blackwell stack for modern MoE — DeepSeek V3.1 (with MLA + MTP) and Qwen3-MoE (235B-A22B) — at FP4 on GB200 NVL72.
🔬 Part 4 — Optimizing a Real Engine (10 lectures)¶
A measured case study. Parts 1–3 pin teaching anchors; this part pins one engine's history: SparkInfer-K3 running Kimi K3 (2.8T params, 896 experts, hybrid MLA + KDA) on a single 8× H200 node, from 1.01 → 60.17 tok/s decode at 128k across ~96 pull requests — measured against llama.cpp at 18.44 on the same box and the same weights.
The subject is not the numbers, which belong to one model on one box in 2026-08. It is the discipline: build the scoreboard before the optimization, find the binding ceiling before writing kernels, and recognize the failure mode where a bug makes the benchmark better.
Course Outcomes¶
By the end of all three parts you should be able to:
- Read any 2026-era model card and predict its inference cost shape (KV growth, prefill vs decode dominance, dense vs MoE routing, expected dominant precision).
- Pick a runtime (vLLM / SGLang / TRT-LLM / llama.cpp / MLX) for a workload + hardware + SLO and defend the choice.
- Quantize a 70B-class dense or 200B+ MoE model and validate parity against a reference within a defined budget.
- Stand up 8× Hopper TP serving with continuous batching, paged KV, prefix cache, and speculation — and explain which knobs moved which metric.
- Stand up Blackwell EP serving for MoE with token-level routing, MTP, and (where applicable) disaggregated P/D.
- Ship a reproducible benchmark with TTFT, TPOT, throughput, p99, and a defended $/MTok cost line.
- Take any optimization claim apart — name three ways the measurement could be flattering itself before arguing about the technique — and build a gate for your own workload that survives an honest engineer, a hostile contributor, and non-deterministic hardware (Part 4).
Currency / Refresh Discipline¶
Up-to-date is the differentiator of this course. The discipline is baked in:
- Every lecture closes with
## Current as of YYYY-MMstating the date the content was written and the specific model / runtime / hardware versions it pinned. REFRESH-LOG.mdtracks every dated update to lectures and benchmark data.- Versioned benchmark data: when a model or runtime ships a new version, prior benchmark numbers are kept in dated subfiles for archaeology; the latest pinned at the top of each lecture.
- Primary sources only — model cards, technical reports, GitHub releases, official benchmark pages. Blog posts (which silently rot) are last-resort and dated.
- Live benchmark reference: for current cross-stack numbers (tokens/s, perf/$, tokens/MW, interactivity) across hardware (H100 → B200 → GB200/GB300 NVL72, MI355X) and runtimes (vLLM / SGLang / TRT-LLM), use a continuously-updated public benchmark such as SemiAnalysis InferenceX (live dashboard, Apache-2.0) rather than any fixed number printed in a lecture — software-stack gains move these weekly. Treat the lecture numbers as teaching anchors; treat the live dashboard as truth at time of deployment.
- Refresh cadence: six months default; three months if a major model class drops (e.g. DeepSeek V4, Llama 5, Qwen 4) or a hardware generation lands (B300 → Vera Rubin etc.).
What You Should Produce¶
A single repo that, by the end of Part 2, contains:
- A reproducible benchmark harness (parametric over model / runtime / hardware).
- One full pipeline at three quantization levels (FP16 reference → FP8 → AWQ-INT4) with parity report.
- TP-scaling numbers on 8× H100 or H200 with NCCL timing breakdown.
- Continuous-batching + prefix-cache + speculation numbers with each knob isolated.
- 128K-context bench with FP8 KV vs FP16 KV.
- A cost model: $/MTok across configurations on the chosen hardware.
By the end of Part 3, the same harness extends to MoE on Blackwell with EP, MTP speculation, and (where the cluster allows) disaggregated P/D measurements.
Part 4 turns that harness into something auditable: a pinned reference with a drift assertion, a correctness gate ordered before the speed gate, a raise-only frontier whose every value traces to a committed measurement, a diagnosis naming the binding ceiling with a number per candidate, an optimization ladder of at least five changes including the ones that measured nothing, and a deliberate-corruption suite proving the gate catches a bug that makes the engine faster.
Exit Criteria¶
You are done with this course when you can:
- Explain, on a whiteboard, why decode is bandwidth-bound and what makes a workload escape that regime.
- Defend a precision floor (FP16 vs FP8 vs INT4 vs FP4) to a roboticist who does not trust quantization.
- Tell, from a profile trace alone, whether a workload is compute-bound, memory-bound, comm-bound, or scheduler-bound — and what to change to verify.
- Walk through your benchmark repo with another engineer and they reproduce your numbers on the same hardware class within ±5%.
- Hand your harness to someone hostile and have them find only holes you already documented.
If you cannot do all five, you have a notebook of recipes, not a body of inference-engineering work. Re-run the benchmarks.