Kolla-Ansible Production Architecture
How Kolla-Ansible actually structures an OpenStack deployment, where configuration comes from, and how to make changes without discovering them during an outage.
Pillar
Cloud infrastructure from the operator side. OpenStack receives particular depth here because production OpenStack knowledge is scarce, hard-won, and rarely written down honestly.
Consolidated guides rather than a page per concept. Each area below is one resource, or will be, rather than a cluster of near-duplicates.
Identity, placement, quota, and the message bus that ties services together — plus the ways a control plane fails without returning an error.
How placement decisions are made, why `NoValidHost` is usually a service-state problem rather than a capacity problem, and how to prove which it is.
Agent state, tunnelling, security groups, volume backends, and the correlation between an agent going quiet and a workload failing to build.
Deployment tooling, configuration management, upgrade sequencing, and the health checks worth running before you believe a cluster is fine.
Resources
Every resource states its author, its review status, and whether the procedures in it were executed or only reviewed.
How Kolla-Ansible actually structures an OpenStack deployment, where configuration comes from, and how to make changes without discovering them during an outage.
Operating OpenStack in production: service state versus status, the message bus, placement disagreements, and the quiet failures that keep dashboards green.
Topics
OpenStack has its own hub because it carries enough depth to justify one — and because honest production OpenStack material is genuinely scarce.
No layer of the stack is operated in isolation. These are the sections you are most likely to need next.
Newsletter
A practical weekly briefing on AI engineering, infrastructure, production operations, and the technologies powering the AI stack.
One email a week. No sponsorship placements inside the technical sections. Unsubscribe in one click.