Incident Diagnosis Method
A repeatable method for diagnosing production failures: form hypotheses the evidence supports, order checks by what they eliminate, and know when to stop and mitigate.
Pillar
Operating systems under real failure conditions: incident intelligence, diagnostic method, configuration review, and the AI-assisted tooling we build to shorten the distance between a symptom and a root cause.
Consolidated guides rather than a page per concept. Each area below is one resource, or will be, rather than a cluster of near-duplicates.
Working from evidence rather than intuition: what to measure first, how to rank hypotheses, and how to kill a wrong one quickly.
Triage order, escalation, communication, and the discipline of writing down what you found while you still remember it.
Instrumentation that answers questions you have not thought of yet, and alerting that pages a human only when a human is needed.
The analyzers we build for our own operations work, published here because they are useful, not because they generate pages.
Resources
Every resource states its author, its review status, and whether the procedures in it were executed or only reviewed.
A repeatable method for diagnosing production failures: form hypotheses the evidence supports, order checks by what they eliminate, and know when to stop and mitigate.
Building observability that shortens incidents rather than producing dashboards: what to instrument, how to alert on symptoms, and why cardinality decides your bill.
Workbench
These run entirely in your browser. Pasted plans, configs, and command output never leave your machine.
Find destructive and high-risk changes in a Terraform plan before you apply it.
Terraform · OpenTofu
Check a Dockerfile against production-readiness rules.
Docker · OCI · Containers
Turn service and agent listings into a ranked view of what is actually broken.
OpenStack · Nova · Neutron
Generate Prometheus alerting rules that will not page you for nothing.
Prometheus · Alertmanager · Grafana
A searchable library of engineering prompts, kept inside the application.
LLM · Incident Response · Kubernetes
No layer of the stack is operated in isolation. These are the sections you are most likely to need next.
Newsletter
A practical weekly briefing on AI engineering, infrastructure, production operations, and the technologies powering the AI stack.
One email a week. No sponsorship placements inside the technical sections. Unsubscribe in one click.