An AI-powered incident response agent for engineering teams that reads alerts, logs, dashboards, and deploy history, then suggests likely root causes and next actions.
Built to demonstrate agentic AI workflows with the ReAct (Reason + Act) pattern and LLM tool use.
┌─────────────────────────────────────────────────────────┐
│ CLI Interface │
│ (investigate / alerts / timeline / postmortem) │
└──────────────────────┬──────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────┐
│ Incident Agent (ReAct Loop) │
│ │
│ ┌──────────┐ ┌──────────┐ ┌──────────────────┐ │
│ │ REASON │──▶│ ACT │──▶│ OBSERVE │ │
│ │ (LLM │ │ (call │ │ (feed results │ │
│ │ thinks) │ │ tools) │ │ back to LLM) │ │
│ └──────────┘ └──────────┘ └──────────────────┘ │
│ ▲ │ │
│ └──────────────────────────────┘ │
│ repeat until diagnosis │
└──────────────────────┬──────────────────────────────────┘
│
┌────────────┼────────────┐
▼ ▼ ▼
┌──────────────┐ ┌──────────┐ ┌──────────────┐
│ Tools │ │ Generators│ │ Integrations │
│ │ │ │ │ │
│ • get_alerts │ │ • timeline│ │ • Datadog │
│ • get_logs │ │ • postmort│ │ • Grafana │
│ • get_metrics│ │ │ │ • CloudWatch │
│ • get_deploys│ └───────────┘ │ • GitLab │
│ • detect_ │ └──────────────┘
│ changes │
│ • suggest_ │
│ actions │
└──────────────┘
The agent follows the ReAct pattern — the same approach used by production AI agents:
- System Prompt: The agent receives an SRE persona with investigation instructions
- User Trigger: "High latency on api-gateway" or "investigate current alerts"
- Reason: The LLM decides which tool to call and why
- Act: The agent calls a tool (e.g.,
get_alerts,get_logs) - Observe: Tool results are fed back to the LLM as context
- Repeat: Steps 3-5 loop until the agent has enough evidence
- Diagnose: The agent produces a structured incident analysis
This is implemented using OpenAI function calling — the LLM sees tool schemas and emits structured tool calls that the agent executes.
cd Incident-Response-agent
python -m venv .venv
source .venv/bin/activate
pip install -e .cp .env.example .env
# Edit .env — set your OPENAI_API_KEY
# Demo mode is on by default (uses mock data, no real integrations needed)# Full agent investigation (requires OpenAI API key)
incident-agent investigate "High latency on the api-gateway, payment errors rising"
# Quick commands (work in demo mode without an API key)
incident-agent alerts # Show active alerts
incident-agent deploys # Show recent deployments
incident-agent changes # Detect recent changes by risk
incident-agent timeline # Generate incident timeline
incident-agent postmortem # Generate postmortem draft
incident-agent status # Show configuration statusThe built-in demo simulates a realistic production incident:
Payment service DB connection pool exhaustion after a bad deploy
T-60m: Deploy of payment-service v2.14.0 (changed DB pool library, reduced pool size)T-45m: Deploy of api-gateway v1.8.3 (rate limiter config update)T-30m: Error rates start rising on payment-serviceT-20m: Latency alert fires on api-gateway (P99 > 4.5s)T-15m: CPU alert fires on payment-service (94%)T-10m: Error rate alert fires (12.4% 5xx)T-8m: DB connection pool exhaustion alert firesT-0: On-call engineer starts investigation
The agent will:
- Pull alerts and notice multiple services are affected
- Check logs and find connection pool exhaustion errors
- Review metrics showing the spike pattern
- Detect the recent deploy as the highest-risk change
- Correlate the timing and recommend a rollback
src/incident_agent/
├── main.py # CLI entry point (Typer)
├── agent.py # Core ReAct agent loop ← the key learning file
├── config.py # Configuration from environment
├── models.py # Pydantic data models
├── tools/ # Tools the LLM can call
│ ├── registry.py # Tool registration & execution
│ ├── alerts.py # Fetch alerts from monitoring
│ ├── logs.py # Query application logs
│ ├── metrics.py # Query metric time-series
│ ├── deploys.py # Query deployment history
│ ├── changes.py # Detect recent changes
│ └── actions.py # Suggest SRE playbook actions
├── integrations/ # Real API clients (optional)
│ ├── datadog_client.py
│ ├── grafana_client.py
│ ├── cloudwatch_client.py
│ └── gitlab_client.py
├── generators/ # Output generators
│ ├── timeline.py # Incident timeline builder
│ └── postmortem.py # Postmortem draft generator
└── demo/
└── mock_data.py # Realistic mock incident data
The core loop where the LLM reasons, calls tools, observes results, and repeats. This is how most production agents work (LangChain, CrewAI, AutoGPT all use variations of this).
Tools are defined with JSON Schema so the LLM knows their parameters. The registry maps tool names to Python functions and handles execution.
The LLM doesn't execute code — it outputs structured tool call requests. The agent framework executes them and feeds results back. This is the "tool use" pattern.
Detecting "what changed recently" by correlating deploy timestamps with alert timestamps. This is the core of most incident investigations.
The agent produces structured artifacts (timeline, postmortem) that are directly useful in incident response workflows.
- Create
src/incident_agent/tools/your_tool.py - Define
TOOL_DEFdict withname,description,parameters,function - Import in
agent.pyand add tobuild_tool_registry()
- Implement the client in
src/incident_agent/integrations/ - Add config fields to
config.py - Update the corresponding tool to use the real client when
config.demo_modeisFalse
Edit the SYSTEM_PROMPT in agent.py to change the agent's investigation strategy, output format, or persona.
- Python 3.11+ — async/await for concurrent operations
- OpenAI API — GPT-4o with function calling for the agent brain
- Pydantic — data validation and serialization
- Typer + Rich — beautiful CLI with tables and panels
- httpx — async HTTP client for integrations