All courses
Professional tierinfrastructure

SRE at Scale

Learn to run Site Reliability Engineering practices across a large, interdependent service fleet: measuring and capping toil, rolling up composite error budgets across dependency chains, deriving incident metrics correctly at scale, planning capacity against nonlinear queueing behavior, and running guarded chaos experiments.

Builds on: Python 3.12

6

modules

~6h

total time

Certificate

on completion

Cancel anytime

Curriculum

6 modules · 11 lessons

Unlock Professional tier

Subscribing to Professional unlocks every course at or below this tier — not just this one.