Signals are scattered
Alert labels, Kubernetes state, and runbooks each show a different part of the incident.
SRE Agent turns an Alertmanager notification into a grounded incident analysis by connecting live Kubernetes evidence with the team's runbooks.
01 / THE PROBLEM
An alert can point to a failing workload, but the useful context lives elsewhere: cluster status, recent events, logs, and operational knowledge. This project brings that evidence together in one response an engineer can inspect and act on.
Alert labels, Kubernetes state, and runbooks each show a different part of the incident.
Responders must locate the right procedure and reconcile it with the live system state.
A useful diagnosis should explain its reasoning, cite observations, and suggest checks to verify recovery.
02 / THE SYSTEM
Each component has a clear job. The agent investigates; the manager owns the knowledge and the record.
SRE Agent receives Alertmanager webhooks and turns firing alerts into incident queries for monitored workloads.
Alertmanager → AgentFor configured namespaces, the agent gathers bounded, read-only Kubernetes status, events, and logs.
Kubernetes APISRE Manager searches its runbook knowledge and returns context to the agent over a shared HTTP contract.
Manager → ChromaDB CloudThe agent writes a structured diagnosis and mitigation steps. SRE Manager preserves the result in an issues inbox.
Agent → SQLite CloudWhen cluster access is unavailable, diagnosis can continue with the alert and available runbook context.
03 / IN PRACTICE
The result is organized around a diagnosis, the evidence behind it, and concrete mitigation checks. Repeated analyses are grouped by alert fingerprint in the manager's Issues view.
Illustrative example based on the project's sample application.
The reservation path is failing while the workload reports an unready container. Investigate the application state before treating the alert as a deployment failure.
1. Inspect readiness events and recent application logs.
2. Follow the runbook's recovery procedure.
3. Confirm readiness and the alert condition clear.
04 / SYSTEM DESIGN
Useful incident analysis depends on current knowledge, trustworthy evidence, and a controlled investigation path.
SRE Manager holds the runbooks and issue history. The agent retrieves relevant guidance when it investigates an alert.
The analysis brings together live cluster observations and runbook guidance, then separates supporting evidence from recommended action.
Read-only cluster access, namespace allowlists, and limits on investigation scope keep the response focused. Individual alert failures stay isolated.
BUILT BY ALVARO RIVAS
I’m happy to discuss the problem, the system design, and how it supports incident response.