← Back
sunny-crypto444

sunny-crypto444/ai-incident-response-agent

View on GitHub ↗
Stars
25
Forks
7
Watchers
25
Open issues
0
Contributors
1
Language
Python
License
—
Default branch
main
Created Sep 29, 2026Updated Sep 29, 2026

Star growth

Today—
This week—
This month—

Star history will appear here once this repo has been tracked for a couple of days.

README

Incident Response Agent 🚨

An AI-powered incident response agent for engineering teams that reads alerts, logs, dashboards, and deploy history, then suggests likely root causes and next actions.

Built to demonstrate agentic AI workflows with the ReAct (Reason + Act) pattern and LLM tool use.

Architecture

┌─────────────────────────────────────────────────────────┐
│                    CLI Interface                         │
│         (investigate / alerts / timeline / postmortem)   │
└──────────────────────┬──────────────────────────────────┘
                       │
                       ▼
┌─────────────────────────────────────────────────────────┐
│               Incident Agent (ReAct Loop)                │
│                                                          │
│   ┌──────────┐   ┌──────────┐   ┌──────────────────┐   │
│   │  REASON  │──▶│   ACT    │──▶│     OBSERVE      │   │
│   │ (LLM     │   │ (call    │   │ (feed results    │   │
│   │  thinks) │   │  tools)  │   │  back to LLM)    │   │
│   └──────────┘   └──────────┘   └──────────────────┘   │
│        ▲                              │                  │
│        └──────────────────────────────┘                  │
│              repeat until diagnosis                      │
└──────────────────────┬──────────────────────────────────┘
                       │
          ┌────────────┼────────────┐
          ▼            ▼            ▼
┌──────────────┐ ┌──────────┐ ┌──────────────┐
│   Tools      │ │ Generators│ │ Integrations │
│              │ │           │ │              │
│ • get_alerts │ │ • timeline│ │ • Datadog    │
│ • get_logs   │ │ • postmort│ │ • Grafana    │
│ • get_metrics│ │           │ │ • CloudWatch │
│ • get_deploys│ └───────────┘ │ • GitLab     │
│ • detect_    │               └──────────────┘
│   changes    │
│ • suggest_   │
│   actions    │
└──────────────┘

How the Agent Works

The agent follows the ReAct pattern — the same approach used by production AI agents:

  1. System Prompt: The agent receives an SRE persona with investigation instructions
  2. User Trigger: "High latency on api-gateway" or "investigate current alerts"
  3. Reason: The LLM decides which tool to call and why
  4. Act: The agent calls a tool (e.g., get_alerts, get_logs)
  5. Observe: Tool results are fed back to the LLM as context
  6. Repeat: Steps 3-5 loop until the agent has enough evidence
  7. Diagnose: The agent produces a structured incident analysis

This is implemented using OpenAI function calling — the LLM sees tool schemas and emits structured tool calls that the agent executes.

Quick Start

1. Install

cd Incident-Response-agent
python -m venv .venv
source .venv/bin/activate
pip install -e .

2. Configure

cp .env.example .env
# Edit .env — set your OPENAI_API_KEY
# Demo mode is on by default (uses mock data, no real integrations needed)

3. Run

# Full agent investigation (requires OpenAI API key)
incident-agent investigate "High latency on the api-gateway, payment errors rising"

# Quick commands (work in demo mode without an API key)
incident-agent alerts                    # Show active alerts
incident-agent deploys                   # Show recent deployments
incident-agent changes                   # Detect recent changes by risk
incident-agent timeline                  # Generate incident timeline
incident-agent postmortem                # Generate postmortem draft
incident-agent status                    # Show configuration status

Demo Scenario

The built-in demo simulates a realistic production incident:

Payment service DB connection pool exhaustion after a bad deploy

  • T-60m: Deploy of payment-service v2.14.0 (changed DB pool library, reduced pool size)
  • T-45m: Deploy of api-gateway v1.8.3 (rate limiter config update)
  • T-30m: Error rates start rising on payment-service
  • T-20m: Latency alert fires on api-gateway (P99 > 4.5s)
  • T-15m: CPU alert fires on payment-service (94%)
  • T-10m: Error rate alert fires (12.4% 5xx)
  • T-8m: DB connection pool exhaustion alert fires
  • T-0: On-call engineer starts investigation

The agent will:

  1. Pull alerts and notice multiple services are affected
  2. Check logs and find connection pool exhaustion errors
  3. Review metrics showing the spike pattern
  4. Detect the recent deploy as the highest-risk change
  5. Correlate the timing and recommend a rollback

Project Structure

src/incident_agent/
├── main.py              # CLI entry point (Typer)
├── agent.py             # Core ReAct agent loop ← the key learning file
├── config.py            # Configuration from environment
├── models.py            # Pydantic data models
├── tools/               # Tools the LLM can call
│   ├── registry.py      # Tool registration & execution
│   ├── alerts.py        # Fetch alerts from monitoring
│   ├── logs.py          # Query application logs
│   ├── metrics.py       # Query metric time-series
│   ├── deploys.py       # Query deployment history
│   ├── changes.py       # Detect recent changes
│   └── actions.py       # Suggest SRE playbook actions
├── integrations/        # Real API clients (optional)
│   ├── datadog_client.py
│   ├── grafana_client.py
│   ├── cloudwatch_client.py
│   └── gitlab_client.py
├── generators/          # Output generators
│   ├── timeline.py      # Incident timeline builder
│   └── postmortem.py    # Postmortem draft generator
└── demo/
    └── mock_data.py     # Realistic mock incident data

Key Concepts to Study

1. ReAct Pattern (agent.py)

The core loop where the LLM reasons, calls tools, observes results, and repeats. This is how most production agents work (LangChain, CrewAI, AutoGPT all use variations of this).

2. Tool Registration (tools/registry.py)

Tools are defined with JSON Schema so the LLM knows their parameters. The registry maps tool names to Python functions and handles execution.

3. Function Calling (agent.py → OpenAI API)

The LLM doesn't execute code — it outputs structured tool call requests. The agent framework executes them and feeds results back. This is the "tool use" pattern.

4. Data Correlation (tools/changes.py)

Detecting "what changed recently" by correlating deploy timestamps with alert timestamps. This is the core of most incident investigations.

5. Structured Output (generators/)

The agent produces structured artifacts (timeline, postmortem) that are directly useful in incident response workflows.

Extending the Agent

Add a new tool

  1. Create src/incident_agent/tools/your_tool.py
  2. Define TOOL_DEF dict with name, description, parameters, function
  3. Import in agent.py and add to build_tool_registry()

Add a real integration

  1. Implement the client in src/incident_agent/integrations/
  2. Add config fields to config.py
  3. Update the corresponding tool to use the real client when config.demo_mode is False

Customize the agent behavior

Edit the SYSTEM_PROMPT in agent.py to change the agent's investigation strategy, output format, or persona.

Tech Stack

  • Python 3.11+ — async/await for concurrent operations
  • OpenAI API — GPT-4o with function calling for the agent brain
  • Pydantic — data validation and serialization
  • Typer + Rich — beautiful CLI with tables and panels
  • httpx — async HTTP client for integrations