Hands-on lab
Incident Exercise — The Cloud Is Full at Thirty-One Percent
A guided production incident where instance builds fail for lack of capacity on a cloud that is barely a third used. Find the capacity that does not exist.
Scenario
Instance builds have been failing with NoValidHost since yesterday afternoon. The capacity dashboard shows the cloud at 31 percent utilisation. Two hosts were replaced last week after a power incident.
What you will practise
- Distinguish genuine exhaustion from an accounting error before adding hardware
- Read placement allocations against what is actually running on a host
- Recognise fragmentation as distinct from a shortage of total capacity
- Reconcile stale allocations safely, without causing over-scheduling
Verification status
This resource has not been executed end to end in a lab environment. Commands and configuration are reviewed by an engineer, but treat them as reference rather than as a tested procedure.
Author
James Joyner
Builds and operates the infrastructure layers underneath production AI systems.
James founded Inside The AI Stack to publish the kind of infrastructure and operations material he wanted while running production systems: specific, tested where it claims to be tested, and written by someone who has had to fix the thing at 3am. He works across AI infrastructure, private cloud, and platform engineering, and reviews every technical resource published here before it is marked as verified.
- AI infrastructure
- OpenStack operations
- Kubernetes
- Terraform
- Linux systems engineering
- Observability
Primary sources
- Placement usage and allocation — OpenStack Foundation
- Nova scheduling configuration — OpenStack Foundation
Continue from here
Related resources chosen because they are the next thing you would actually need — not because they share a keyword.
Lab: Nova NoValidHost
Instances stopped building on a cloud with visible free capacity. Work the evidence, form hypotheses, and find why the scheduler has no candidates.
OpenStack Production Engineer
A sequenced learning path for engineers who operate OpenStack in production — architecture, service-by-service depth, deployment, and troubleshooting under pressure.
Newsletter
Inside The AI Stack Brief
A practical weekly briefing on AI engineering, infrastructure, production operations, and the technologies powering the AI stack.
One email a week. We send a confirmation link first and never mail an address that has not confirmed. No sponsorship placements inside the technical sections, and an unsubscribe link in every issue.
