All courses
Elite tierinfrastructure

LLM Infrastructure & MLOps

Learn to size, serve, train, monitor, and cost real LLM deployments: GPU memory and KV-cache capacity planning, continuous batching and quantization tradeoffs, distributed training with verified gradient accumulation, production drift monitoring, and utilization-aware cost economics.

Builds on: Python 3.12 + numpy 2.4.4 (GPU memory/KV-cache modeling, batching simulation, gradient-accumulation verification, drift-monitoring statistics, cost-per-token economics)

6

modules

~6h

total time

Certificate

on completion

Cancel anytime

Curriculum

6 modules · 11 lessons

Unlock Elite tier

Subscribing to Elite unlocks every course at or below this tier — not just this one.