Product Talk · Free post · Building AI products · Execution, roadmaps & process

When AI Becomes Your SRE: How Incident.io Is Automating Incident Response

Teresa TorresNov 6, 2025
SourceProduct Talk
KindFree post
PublishedNov 6, 2025
Originalproducttalk.org ↗
N:

This episode of a product podcast features Incident.io's founding engineer Lawrence Jones and AI product lead Ed Dean discussing how their team is building an AI SRE that helps diagnose and respond to production incidents. The team evolved from simple prototypes that summarized incidents into a multi-agent system that forms hypotheses, tests them, and drafts fixes within Slack. The discussion focuses on which debugging steps can safely be automated, how retrieval and re-ranking surface relevant context, and how post-incident evaluations score AI performance against what actually happened. It also addresses earning human trust through transparent reasoning and expressed uncertainty. The conversation is useful for anyone designing AI systems that must collaborate with experts in high-stakes workflows.

01Key takeaways

  • Target the parts of incident debugging that can be safely automated first, then expand scope gradually.
  • Simple deterministic tagging and re-ranking can beat complicated vector retrieval setups in practice.
  • Score AI diagnoses against known outcomes after an incident to measure real accuracy.
  • Show reasoning and uncertainty to users so they can calibrate trust in AI recommendations.
  • Use Slack as a familiar interface so humans and AI agents collaborate where work already happens.

02Key sections

From incident coordination to AI diagnosis
Incident.io began by helping teams coordinate during outages and is now extending into an AI SRE that assists with diagnosis. The hosts frame the shift as moving from coordination tooling to active reasoning.
Prototyping and retrieval strategy
The team iterated from simple incident summaries to richer retrieval pipelines. They found that deterministic tagging and re-ranking often outperformed complex vector-only setups.
Multi-agent investigation workflow
The system runs parallel checks across data sources, builds hypotheses, and refines findings, with humans collaborating throughout. This shows how agents can divide investigative work.
Evaluating and trusting the AI
Post-incident 'time travel' evals score the AI's conclusions after the true cause is known. Trust depends on exposing reasoning and uncertainty, not just accuracy.
Future directions
The team discusses expanding integrations and capabilities as the AI SRE matures. The focus stays on compressing time-to-diagnosis during real outages.

03From the post

“Listen to this episode on: Spotify | Apple Podcasts When your site goes down, every second counts. For years, Incident.io has helped engineering teams coordinate through chaos—getting the right people in the room, keeping stakeholders informed, and restoring order fast. Now, they’re building something new: an AI SRE”

“AI's biggest impact comes from compressing time—identifying causes minutes instead of hours.”Teresa Torres · Product Talk · 00:00

04Frameworks mentioned

Summary and takeaways written by PM Atlas; quotes are short excerpts. © the original author.