Inside The AI Stack

Hands-on lab

Incident Exercise — Every Service Reports Down, Nothing Is Broken

A guided production incident where the whole OpenStack control plane reports down while every workload keeps running. Work the evidence and find what is actually wrong.

advancedOpenStackRabbitMQNovaNeutron
ByJames JoynerPublished 5 min read
~35 minProadvanced

Scenario

At 02:14 the monitoring system reports every nova-compute service down across all eleven compute hosts, and every Neutron agent dead. Nothing has been deployed for six days. Users have not reported anything.

What you will practise

  • Separate a control-plane failure from a data-plane failure before touching anything
  • Recognise the message bus as a shared dependency from the shape of the symptoms
  • Read queue depth and consumer counts to locate the actual fault
  • Choose a recovery order that does not extend the outage

Verification status

This resource has not been executed end to end in a lab environment. Commands and configuration are reviewed by an engineer, but treat them as reference rather than as a tested procedure.

Author

James Joyner

Builds and operates the infrastructure layers underneath production AI systems.

James founded Inside The AI Stack to publish the kind of infrastructure and operations material he wanted while running production systems: specific, tested where it claims to be tested, and written by someone who has had to fix the thing at 3am. He works across AI infrastructure, private cloud, and platform engineering, and reviews every technical resource published here before it is marked as verified.

  • AI infrastructure
  • OpenStack operations
  • Kubernetes
  • Terraform
  • Linux systems engineering
  • Observability

Primary sources

Related resources chosen because they are the next thing you would actually need — not because they share a keyword.

Lab

Lab: Nova NoValidHost

Instances stopped building on a cloud with visible free capacity. Work the evidence, form hypotheses, and find why the scheduler has no candidates.

expert· 35 minOpenStackNova
Learning path

OpenStack Production Engineer

A sequenced learning path for engineers who operate OpenStack in production — architecture, service-by-service depth, deployment, and troubleshooting under pressure.

intermediate· 5 minOpenStackNova

Newsletter

Inside The AI Stack Brief

A practical weekly briefing on AI engineering, infrastructure, production operations, and the technologies powering the AI stack.

One email a week. We send a confirmation link first and never mail an address that has not confirmed. No sponsorship placements inside the technical sections, and an unsubscribe link in every issue.