Lab: CrashLoopBackOff
A service crash-looping after a routine config change, with logs that show a normal startup and no error anywhere. Work out what is killing it.
Practice
Reading about a failure and diagnosing one are different skills. Labs drop you into an incident with the same partial, noisy information you would have in production, and ask you to work it out before showing you the answer.
Six stages, in the order a real investigation goes — not the order a tutorial would put them in.
You are handed a failure and the context an engineer would actually have. Nothing is pre-diagnosed.
Realistic command output, logs, and service state. Some of it is relevant. Some of it is not.
You decide what to run next. The lab tells you what each command would have returned.
Rank what could produce these signals, and identify which check eliminates the most possibilities.
The explanation, including why the obvious first guess was wrong.
Remediation, the validation that proves it worked, and the rollback if it did not.
Labs
Introductory labs are free. Advanced labs are part of Pro.
A service crash-looping after a routine config change, with logs that show a normal startup and no error anywhere. Work out what is killing it.
Instances stopped building on a cloud with visible free capacity. Work the evidence, form hypotheses, and find why the scheduler has no candidates.
A routine-looking Terraform pull request. Find the changes that would destroy data before you approve it, and work out which one is not what it appears to be.
Newsletter
A practical weekly briefing on AI engineering, infrastructure, production operations, and the technologies powering the AI stack.
One email a week. No sponsorship placements inside the technical sections. Unsubscribe in one click.