Inside The AI Stack

Pillar

DevOps & Platform Engineering

Platform engineering for teams that carry a pager: container runtime behaviour, cluster operations, infrastructure as code, Linux fundamentals, and the observability that makes production legible.

What this section covers

Consolidated guides rather than a page per concept. Each area below is one resource, or will be, rather than a cluster of near-duplicates.

Containers and runtime behaviour

What actually happens when a container starts, why it exits, and how to build images that behave predictably under an orchestrator.

Kubernetes operations

Scheduling, resource limits, probes, rollout strategy, and reading cluster events instead of guessing from pod status.

Infrastructure as code

State management, plan review as the primary safety control, module design, and containing the blast radius of an apply.

Observability

Metrics, logs, and traces that answer questions during an incident rather than only after one.

Resources

Resources in this pillar

Every resource states its author, its review status, and whether the procedures in it were executed or only reviewed.

Container Startup Troubleshooting

Exec format errors, permission denied, missing commands, bad entrypoints, immediate exits, and CrashLoopBackOff — diagnosed as one family rather than a page per error string.

intermediate· 4 minDockerContainers

Docker Production Engineering

Building container images for production: reproducible builds, correct signal handling, non-root runtime, layer strategy, and keeping secrets out of image history.

intermediate· 4 minDockerContainers

Kubernetes Workload Troubleshooting

A single diagnostic path for workloads that will not schedule, will not stay up, or keep getting evicted — driven by events and previous logs rather than guesswork.

intermediate· 3 minKubernetes

Production Kubernetes Operations

The operational settings that decide whether a workload survives a node failure or a rollout: requests and limits, the three probes, disruption budgets, and scheduling.

intermediate· 4 minKubernetesContainers

Terraform Production Practices

Running Terraform against production infrastructure: reviewing the plan mechanically, containing blast radius, structuring state, and avoiding the destroys nobody noticed.

intermediate· 4 minTerraformOpenTofu

Terraform State Operations

The state commands that are genuinely dangerous, what each one actually does, and how to run them with a way back.

advanced· 3 minTerraformOpenTofu

Topics

Topic hubs

A topic gets its own hub once enough finished resources sit beneath it to make the page worth opening. Topics below that threshold are reachable from the list above.

Newsletter

Inside The AI Stack Brief

A practical weekly briefing on AI engineering, infrastructure, production operations, and the technologies powering the AI stack.

One email a week. No sponsorship placements inside the technical sections. Unsubscribe in one click.