An SRE portfolio project by Alvaro Rivas

From alert to actionable diagnosis.

SRE Agent turns an Alertmanager notification into a grounded incident analysis by connecting live Kubernetes evidence with the team's runbooks.

Signals → context → next steps
Incident workspace EXAMPLE
ALERT RECEIVED Alertmanager
Reservation flow degraded A firing alert starts the investigation.
KUBERNETES Workload evidence Pod running · container not ready
RUNBOOK Operational guidance Reservation recovery procedure
ACTIONABLE ANALYSIS Investigate readiness events and application logs.
Illustrative flow · Alert → evidence → next step
CORE CAPABILITIES
  • Observability system
  • Kubernetes
  • RAG
  • Runbook
  • LLM

01 / THE PROBLEM

Alerts tell you what happened.
Engineers still need the why.

An alert can point to a failing workload, but the useful context lives elsewhere: cluster status, recent events, logs, and operational knowledge. This project brings that evidence together in one response an engineer can inspect and act on.

01

Signals are scattered

Alert labels, Kubernetes state, and runbooks each show a different part of the incident.

02

Context takes time

Responders must locate the right procedure and reconcile it with the live system state.

03

Action needs evidence

A useful diagnosis should explain its reasoning, cite observations, and suggest checks to verify recovery.

02 / THE SYSTEM

One connected path from signal to next step.

Each component has a clear job. The agent investigates; the manager owns the knowledge and the record.

  1. 01

    Receive the alert

    SRE Agent receives Alertmanager webhooks and turns firing alerts into incident queries for monitored workloads.

    Alertmanager → Agent
  2. 02

    Inspect the cluster

    For configured namespaces, the agent gathers bounded, read-only Kubernetes status, events, and logs.

    Kubernetes API
  3. 03

    Find relevant runbooks

    SRE Manager searches its runbook knowledge and returns context to the agent over a shared HTTP contract.

    Manager → ChromaDB Cloud
  4. 04

    Publish the analysis

    The agent writes a structured diagnosis and mitigation steps. SRE Manager preserves the result in an issues inbox.

    Agent → SQLite Cloud

When cluster access is unavailable, diagnosis can continue with the alert and available runbook context.

03 / IN PRACTICE

A response an engineer can actually use.

The result is organized around a diagnosis, the evidence behind it, and concrete mitigation checks. Repeated analyses are grouped by alert fingerprint in the manager's Issues view.

  • Evidence from alerts and cluster state
  • Relevant runbook context
  • Steps to mitigate and confirm recovery

Illustrative example based on the project's sample application.

SRE Manager
IssuesRunbooks
Issues / Reservation flow
INCIDENT ANALYSIS

Reservation flow degraded

Firing
coffeehouse-webappCriticalLatest analysis
DIAGNOSIS

The reservation path is failing while the workload reports an unready container. Investigate the application state before treating the alert as a deployment failure.

SUPPORTING EVIDENCE
Alert: reservation failure condition is active
Kubernetes: pod is running, container is not ready
Runbook: reservation recovery procedure retrieved
NEXT STEPS

1. Inspect readiness events and recent application logs.

2. Follow the runbook's recovery procedure.

3. Confirm readiness and the alert condition clear.

04 / SYSTEM DESIGN

Designed for real incident work.

Useful incident analysis depends on current knowledge, trustworthy evidence, and a controlled investigation path.

One source of truth

SRE Manager holds the runbooks and issue history. The agent retrieves relevant guidance when it investigates an alert.

Grounded conclusions

The analysis brings together live cluster observations and runbook guidance, then separates supporting evidence from recommended action.

Controlled investigation

Read-only cluster access, namespace allowlists, and limits on investigation scope keep the response focused. Individual alert failures stay isolated.

BUILT BY ALVARO RIVAS

Curious about the decisions behind it?

I’m happy to discuss the problem, the system design, and how it supports incident response.

Visit my portfolio