Research project
AI Infrastructure Operations Survey — 2027 Edition
A planned survey of how teams actually operate AI infrastructure in production — what breaks, what is measured, and where engineering time goes. Not yet conducted.
Why this survey
There is a large amount of published material about what AI infrastructure should look like, and almost none about what teams actually run and what actually breaks.
The specific gaps we want to close:
- Where does operational time actually go? Our working hypothesis, from our own work, is that data-path and scheduling problems consume far more time than model-serving problems — but that is a hypothesis drawn from a small sample.
- What proportion of teams measure accelerator utilisation at all, and of those, what do they see?
- How much capacity is idle, and for what reasons?
- What is the gap between what teams monitor and what causes their incidents?
Intended method
Population. Engineers with direct operational responsibility for infrastructure running AI workloads — training, inference, or both. Not managers reporting on behalf of teams, and not vendors.
Sampling. Self-selected respondents, which is a real limitation and will be stated as such in any published result. Self-selected samples over-represent people who read engineering publications and under-represent teams that are not struggling enough to seek material out.
Instrument. Structured questions with a small number of free-text fields. Published in full alongside the results, so anyone can see exactly what was asked and how the framing might have shaped the answers.
Analysis. Reported as distributions, not averages. Cross-tabulated by deployment scale, since the operational reality at eight accelerators and eight hundred is not the same subject.
Publication. The anonymised dataset published under an open licence with the report, so disagreement can be evidence-based.
What we will not do
- Report a median without the distribution. A median hides the shape, and the shape is the finding.
- Claim causation from a survey. Correlations between practices and outcomes will be reported as correlations, with that limitation stated in the same sentence rather than in a footnote.
- Accept sponsorship that touches the questions. If this is ever sponsored, the sponsor sees the results at publication like everyone else, and that arrangement is disclosed here.
- Publish if the response count is too low to say anything. A survey with thirty responses produces numbers that look like findings and are not. If we do not clear a meaningful threshold, we will say so and the project stays unpublished.
Timeline
Not scheduled. This project is defined and not started.
We will announce it in Inside The AI Stack Brief when the instrument is ready, because a survey is only as good as who responds to it.
Verification status
This resource has not been executed end to end in a lab environment. Commands and configuration are reviewed by an engineer, but treat them as reference rather than as a tested procedure.
Author
James Joyner
Builds and operates the infrastructure layers underneath production AI systems.
James founded Inside The AI Stack to publish the kind of infrastructure and operations material he wanted while running production systems: specific, tested where it claims to be tested, and written by someone who has had to fix the thing at 3am. He works across AI infrastructure, private cloud, and platform engineering, and reviews every technical resource published here before it is marked as verified.
- AI infrastructure
- OpenStack operations
- Kubernetes
- Terraform
- Linux systems engineering
- Observability
Continue from here
Related resources chosen because they are the next thing you would actually need — not because they share a keyword.
GPU Infrastructure for AI
What determines accelerator performance in production: memory capacity versus bandwidth, interconnect topology, and the checks that find a misplaced workload.
Newsletter
Inside The AI Stack Brief
A practical weekly briefing on AI engineering, infrastructure, production operations, and the technologies powering the AI stack.
One email a week. No sponsorship placements inside the technical sections. Unsubscribe in one click.