Tag: SRE
The Observability Tax: When Monitoring Costs Exceed Downtime Costs
Observability costs can spiral when teams collect more telemetry than they actually use. A more mature approach prioritizes the data that directly improves incident detection, diagnosis and recovery ...
The Three Tiers of Agentic Incident Response: When to Trust AI Autonomy
A three-tier model for agentic incident response balances AI automation with human oversight, matching autonomy to risk, reversibility, blast radius and diagnostic confidence ...
The Most Dangerous Reliability Failures Aren’t Component Failures
A production provisioning failure shows why reliability problems can emerge from interactions between healthy components, and how system-level constraints and feedback can prevent them ...
Observability’s Gaslighting Problem: “Send Less Data” Isn’t a Strategy
A familiar pattern is emerging in observability conversations. As telemetry volumes grow and costs rise, the default recommendation is often to collect less data: Sample more, retain less, index selectively, filter earlier, ...
Automated Diagnosis Isn’t Automated Understanding: What Postmortems Teach Us About Building Trustworthy Incident AI
AI incident tools can reduce alert noise, but real root-cause diagnosis requires causal reasoning, live dependency context, uncertainty handling and strong postmortem data ...
Beyond Log Search: What We Learned Building a RAG-Based Incident Diagnosis System
A RAG-based AIOps framework can cut incident diagnosis time by grounding LLM reasoning in real runbooks, tickets and postmortems, improving root-cause accuracy while giving SREs source-backed answers they can trust ...
Mezmo Open Sources AI SRE Operations
Site reliability engineering has been quietly buckling under its own success. The scope of what SRE teams are expected to own — observability, incident response, telemetry pipelines, capacity, cost, resilience — keeps ...
Why Your Observability Stack Is Costing You More Than Your Cloud Bill
There's a pattern playing out across engineering teams right now that nobody talks about openly: the tool meant to reduce operational complexity has quietly become one of the biggest line items on ...
When the Structure Becomes the Culture
Why micro teams and rotation reshape culture, not just throughput, in modern SRE. Most SRE leaders design teams around the systems they own. We designed ours around movement. We introduced micro teams ...
The Death of the Four Golden Signals: Designing Telemetry for Non-Deterministic Infrastructure
In complex software systems, our traditional definition of operational health has always been comfortably binary. For over a decade, site reliability engineering (SRE) teams have relied on the industry-standard ‘Four Golden Signals’ ...
Agentic SRE: The Next Frontier of Reliability
Agentic SRE is the evolution of site reliability engineering where AI agents help observe systems, reason over telemetry and take bounded operational actions under human-defined guardrails ...
On-Call: The Silent Force Shaping Engineering Culture
There is a silent force shaping engineering culture inside every technology organization. It affects productivity, team morale, psychological safety, and long-term retention. And yet, it is rarely discussed in executive meetings or ...

