Building Autonomous Self-Healing OpenShift Clusters
A narrative walkthrough of the same system covered in the mechanical runbook — written so a colleague picking this up cold can understand not just what to type, but why each piece exists and how it fits together.
I.Understanding the Architecture
Read this section before you type a single command. Everything downstream is easier to debug — and easier to explain to someone watching over your shoulder — if you understand why each piece is here.
Most SRE teams have already solved alert detection. Prometheus fires, AlertManager routes it, a pager goes off. What almost nobody has solved is the fifteen minutes that follow: the engineer opening four different dashboards, cross-referencing logs against a deploy timeline, forming a hypothesis, checking it, and finally deciding what to actually do. That gap — triage, correlation, decision, action — is manual almost everywhere, and it's why MTTR numbers haven't moved much even as observability tooling has gotten dramatically better. Better detection didn't shrink the gap after detection.
This system exists to close that specific gap, for a specific, bounded class of incidents, without pretending the problem is easier than it is. It doesn't replace an SRE's judgment — it automates the parts of triage that are mechanical (gather the metrics, pull the logs, find the trace) and hands the judgment call to an LLM only under tight, auditable constraints, with the blast radius of any automated action explicitly capped by how confident and how risky that action is.
Why detection alone isn't diagnosis
Prometheus and AlertManager answer one question well: is something wrong? They cannot answer why, because a metric threshold breach is a symptom, not a cause. "Pod OOMKilled" tells you what happened to one container; it says nothing about whether that container leaked memory on its own or was pushed into resource exhaustion by a slow dependency three services upstream. Getting from symptom to cause has always required a human to pull three or four different data sources into their head at once and reason across them. That reasoning step — not the alerting — is the actual bottleneck, and it's the piece this pipeline automates.
The SRE Intelligence service: the orchestration brain
Sitting at the center of the whole system is a small FastAPI application called the SRE Intelligence service. It has one job: when AlertManager's webhook fires, gather every piece of context a human on-call engineer would gather by hand, hand it to an LLM in a structured prompt, and act on the LLM's answer according to a fixed trust policy. It is deliberately not itself intelligent — it's plumbing. All of the judgment lives in the LLM call; all of the safety lives in what the SRE Intelligence service is and isn't allowed to do with that judgment.
Two different questions: logs vs. traces
The pipeline correlates three independent signals, and it's worth being precise about why each one is there, because they answer genuinely different questions:
- Metrics (Prometheus) answer is something wrong, and how bad — the alert itself, plus a quick read on memory/CPU/restart trends.
- Logs (Loki) answer what actually happened around the time of the alert — error messages, timeouts, warnings, the specific text a human would grep for.
- Traces (Tempo/OpenTelemetry) answer where, in a chain of service-to-service calls, the problem originated — which specific downstream hop was slow, and by how much.
A metric alone can't tell you that order-service's memory spike was caused by a slow call to inventory-service, three network hops away. A log line might hint at it ("slow downstream call"), but only a trace proves it, with a duration attached to the exact span. Real root-cause analysis needs all three lenses at once — which is exactly the kind of multi-source correlation that traditional rule-based alerting was never designed to do, and exactly what an LLM, given the same three inputs a human would use, is well suited to.
Why the LLM sits at the center
The model (Qwen2.5-7B-Instruct in AWQ 4-bit, served locally via KServe and vLLM) receives the firing alert, the metrics snapshot, the last 15 minutes of relevant log lines, a summary of the slowest recent trace for the affected service, and any matching runbook text — and returns a single structured judgment: a probable root cause, a confidence score, a recommended action, and which of three trust tiers that action belongs to. This is the one part of the pipeline that requires judgment rather than rules, because "is this restart safe, or does this alert actually need a human" is not something you can express as a Prometheus threshold. It's also the part of the architecture most people are skeptical of — which is why every action the model can take is drawn from a small, fixed enum (never arbitrary code or commands), and why its output is schema-validated before anything happens with it.
Why a runbook knowledge base at all
An LLM reasoning from nothing but the alert text will guess. Giving it your team's actual, already-written runbook for that specific alert grounds its answer in institutional knowledge instead of general pattern-matching — the same reason a human on-call engineer checks the runbook before improvising. Here, that's a plain ConfigMap: alert name in, runbook text out, no database, no embedding model, no similarity search. It doesn't need to be more sophisticated than that, because the lookup is always an exact match on a known alert name, never a fuzzy search over unknown incident text.
The three-tier trust model
This is the actual design philosophy of the whole system, and it's worth stating plainly: match the blast radius of automation to the confidence and the risk of the action, not to how impressive full autonomy sounds. Concretely:
- Tier 1 — low risk, auto-execute. Pod restarts, config rollbacks, HPA scaling. These are reversible, well-understood, and low-blast-radius even if the LLM's confidence is slightly off. Only auto-executed above a high confidence threshold (0.75 by default in this build's configuration — conservative production deployments should start higher, e.g. 0.85, and lower it only after reviewing a week of audit log entries).
- Tier 2 — medium risk, human-approved. Network policy changes, node drains, PVC resizes. These can cause a different kind of outage if wrong, so a human clicks Approve or Reject in Slack before anything happens — the LLM does the analysis, a person keeps the authority.
- Tier 3 — high severity or low confidence, escalate. Database failovers, region failovers, or simply "the model isn't sure." No automated action at all — the on-call engineer gets paged with a pre-analyzed context pack (root cause, confidence, reasoning) instead of a bare alert, so the fifteen-minute triage gap shrinks even when a human has to make the final call.
Notice what this buys you: the system's autonomy is exactly as wide as its judgment is trustworthy for that specific class of action, and never wider. That's the whole idea.
The test applications: a real cascading failure, deliberately
To demonstrate this honestly, the demo workload needs to produce a genuine incident, not a scripted crash. order-service and inventory-service are two minimal FastAPI services built for exactly one purpose: order-service calls inventory-service on every request with no timeout and no circuit breaker — a classic, well-known SRE failure pattern, and one that's almost always the actual root cause behind "mystery" memory spikes in production. inventory-service exposes a debug endpoint to become artificially slow on command; when it does, order-service's in-flight requests pile up against its deliberately tight memory limit until it's OOMKilled — a real resource-exhaustion cascade, not a simulation. Both services are instrumented with real OpenTelemetry auto-instrumentation, so the slow call genuinely shows up as a long span in Tempo, and the timeout-free client genuinely logs a warning line that Loki surfaces. Nothing about the "cascading failure" in this demo is faked; it's just deliberately small and reliably reproducible.
The request lifecycle, start to finish
Putting it all together, here's what happens the moment something breaks:
- A Prometheus rule fires; AlertManager posts a webhook to the SRE Intelligence service.
- The SRE Intelligence service queries Prometheus for current metrics, queries Loki for recent log lines from the affected pod, queries Tempo for the slowest recent trace involving that pod's owning Deployment, and looks up any matching runbook — all before it ever talks to the LLM.
- All of that context gets assembled into a single structured prompt and sent to the locally-served Qwen2.5-7B-Instruct model.
- The model returns a JSON root-cause analysis: cause, confidence, recommended action, action type, and a tier assignment.
- The trust-tier executor branches: Tier 1 calls the MCP server via a fresh connection opened for that call (no persistent session is kept between actions — see the MCP Server section for why); Tier 2 posts an interactive Slack approval card and waits for a click; Tier 3 posts a Slack escalation with the full context pack.
- Every outcome — automatic or human-approved — gets written to an audit log and incremented in Prometheus metrics, so the loop is inspectable after the fact.
II.Prerequisites
What you need before opening a terminal, and the two items that need a day or two of lead time.
Cluster
- An existing OpenShift cluster (IPI on AWS in this guide, but nothing here is AWS-specific beyond the node-provisioning examples).
- At least one GPU-enabled worker node. This guide doesn't walk through provisioning one via a MachineSet — that's a standard OpenShift Machine API operation and your cluster's existing worker MachineSets are the best template to copy if you need to add one. Concretely, for Qwen2.5-7B-Instruct in AWQ 4-bit quantized form, a single AWS
g4dn.2xlarge(one NVIDIA T4, 16 GB VRAM) is sufficient — quantized weights are ~4.5 GiB, leaving ~10 GiB for KV cache and runtime overhead.
Local tooling
ocCLI, logged in as cluster-adminpython3.11+pippodmanordocker- A HuggingFace account (a token is needed for the model download step)
- ngrok or an equivalent tunnel tool (needed later for the Slack integration)
Confirm cluster access
# Confirm cluster-admin access
oc whoami
oc get clusterversion
Create the working namespaces
Most observability operator namespaces (openshift-operators-redhat, openshift-logging, openshift-tempo-operator, openshift-opentelemetry-operator) are created automatically when you install the operators via OperatorHub following the linked docs later in this guide. The exception is openshift-tempo, which hosts the TempoStack instance and must be created manually:
oc new-project business-workloads oc label namespace business-workloads \ app.kubernetes.io/part-of=sre-platform \ openshift.io/cluster-monitoring=true oc new-project sre-intelligence oc label namespace sre-intelligence \ app.kubernetes.io/part-of=sre-platform \ openshift.io/cluster-monitoring=true oc new-project mcp-server oc label namespace mcp-server \ app.kubernetes.io/part-of=sre-platform oc create namespace openshift-tempo
business-workloads holds the demo application workloads (order-service, inventory-service, OTEL collector). sre-intelligence holds the SRE Intelligence service and its runbook ConfigMap. mcp-server holds the openshift-mcp-server — the component that holds cluster credentials and executes remediation actions on behalf of the SRE Intelligence service.III.Installing OpenShift AI 3.4
This is the platform that gives you KServe and vLLM for model serving. Rather than reproduce operator YAML here that will drift the moment Red Hat ships a point release, follow the official install guide directly:
Installing and deploying OpenShift AI Self-Managed 3.4 →
One thing to get right that isn't obvious from a quick skim of that guide: choose Standard deployment mode for KServe (RHOAI's current name for what upstream KServe calls RawDeployment) when you create the DataScienceCluster. It gives you a plain Kubernetes Deployment/Service/HPA for the model server instead of a Knative-backed Serverless one — which is exactly what the ServingRuntime and InferenceService later in this guide expect (their serving.kserve.io/deploymentMode: Standard annotation is what selects it), and it means the OpenShift Serverless and Service Mesh operators aren't required for KServe itself. You lose scale-to-zero, which doesn't matter for a demo that's running the whole time anyway.
Create the DataScienceProject for LLM workloads
Once the operator and DataScienceCluster are in place per the guide above, create the namespace this pipeline serves its model from:
oc apply -f - <<EOF
apiVersion: v1
kind: Namespace
metadata:
name: llm-serving
labels:
opendatahub.io/dashboard: "true"
modelmesh-enabled: "false" # Forces KServe mode
EOF
IV.GPU & Platform Operators
Two more operators are needed before the GPU node is schedulable, and both are RHOAI platform dependencies rather than something specific to this pipeline's design — so, same as above, follow Red Hat's own install guide rather than a YAML snapshot that will go stale:
OpenShift AI 3.4 platform requirements →
- Node Feature Discovery (NFD) — labels nodes with detected hardware features (including the presence of a GPU), which the NVIDIA GPU Operator and RHOAI's scheduling both depend on.
- NVIDIA GPU Operator — installs and manages the GPU driver, device plugin, and monitoring stack on any node NFD has labeled as GPU-capable.
Install both operators via OperatorHub, then create their instances with default configuration — the Dashboard's "Create instance" flow with no fields changed is sufficient here: create the NodeFeatureDiscovery instance first, then the ClusterPolicy instance for the NVIDIA GPU Operator once nodes are labeled. On OpenShift specifically, confirm the created ClusterPolicy has use_ocp_driver_toolkit enabled (it's on by default) — without it, driver builds fail because they can't compile kernel modules against a standard RHCOS image.
After both instances are created and their pods are healthy, verify the GPU is actually schedulable before moving on:
oc apply -f - <<EOF
apiVersion: v1
kind: Pod
metadata:
name: gpu-verify
namespace: business-workloads
spec:
restartPolicy: Never
containers:
- name: nvidia-smi
image: nvidia/cuda:12.3.1-base-ubi9
command: ["nvidia-smi"]
resources:
limits:
nvidia.com/gpu: 1
EOF
oc logs gpu-verify -n business-workloads
oc delete pod gpu-verify -n business-workloads
You should see a normal nvidia-smi table naming a T4. If the pod stays Pending, check oc describe node <gpu-node> for a GPU taint that nothing is tolerating yet — this resolves itself once the operators finish reconciling, but is worth checking before assuming something's actually broken.
V.A Hardware Profile for the GPU
This step is easy to miss because nothing in the RHOAI install flow forces it on you — the model deployment will simply never schedule.
RHOAI 3.4's default hardware profile only offers CPU and memory as schedulable resources — it has no concept of an accelerator out of the box. If you deploy the InferenceService in the next section without first creating a profile that exposes the GPU, the predictor pod requests nvidia.com/gpu: 1 against a profile that never told the scheduler that resource exists on any node, and it sits Pending indefinitely with no obviously-relevant error.
The fix is a custom HardwareProfile naming the GPU as an accelerator resource:
apiVersion: infrastructure.opendatahub.io/v1
kind: HardwareProfile
metadata:
annotations:
opendatahub.io/dashboard-feature-visibility: '[]'
opendatahub.io/description: Nvidia
opendatahub.io/disabled: "false"
opendatahub.io/display-name: Nvidia
name: nvidia
namespace: redhat-ods-applications
spec:
identifiers:
- defaultCount: 2
displayName: CPU
identifier: cpu
maxCount: 4
minCount: 1
resourceType: CPU
- defaultCount: 4Gi
displayName: Memory
identifier: memory
maxCount: 8Gi
minCount: 2Gi
resourceType: Memory
- defaultCount: 1
displayName: gpu
identifier: nvidia.com/gpu
maxCount: 1
minCount: 1
resourceType: Accelerator
scheduling:
node:
nodeSelector:
nvidia.com/gpu.present: "true" # NFD label set on every node the GPU Operator detects a GPU on --
# works across any number of GPU nodes without pinning a hostname
tolerations: []
type: Node
oc apply -f hardwareprofile-nvidia.yaml
InferenceService in the next section binds to this profile with two annotations on its metadata: opendatahub.io/hardware-profile-name: nvidia and opendatahub.io/hardware-profile-namespace: redhat-ods-applications — naming the profile and the namespace it lives in. That's the full mechanism; nothing else is required to wire the two resources together.VI.Model & Observability Storage
Everything that needs object storage in this pipeline — the model artifact, Loki's chunks, Tempo's traces — uses MinIO running in-cluster as the S3-compatible backend. Three separate buckets are created: one per component. Each operator expects its own dedicated bucket; sharing a bucket causes tenant-detection collisions (particularly in Tempo, which interprets top-level directories as tenant names).
MinIO gives you an S3-compatible API running as a normal in-cluster Deployment, backed by a PVC. Everything downstream talks to it via the same S3-compatible API and uses the internal Service DNS — no external account or egress required.
oc new-project minio oc create secret generic minio-creds \ --from-literal=MINIO_ROOT_USER=admin \ --from-literal=MINIO_ROOT_PASSWORD='minio.yaml' \ -n minio # DEMO ONLY: replace with a strong generated password before any shared or # persistent deployment. Production: use HashiCorp Vault, AWS Secrets Manager, or # OpenShift Sealed Secrets -- never commit plaintext passwords to git.
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: minio-data
namespace: minio
spec:
accessModes:
- ReadWriteOnce
resources:
requests:
storage: 50Gi
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: minio
namespace: minio
spec:
replicas: 1
selector:
matchLabels:
app: minio
template:
metadata:
labels:
app: minio
spec:
containers:
- name: minio
image: quay.io/minio/minio:latest
args:
- server
- /data
- --console-address
- ":9001"
envFrom:
- secretRef:
name: minio-creds
ports:
- containerPort: 9000
name: api
- containerPort: 9001
name: console
volumeMounts:
- name: data
mountPath: /data
securityContext:
allowPrivilegeEscalation: false
runAsNonRoot: true
capabilities:
drop: ["ALL"]
volumes:
- name: data
persistentVolumeClaim:
claimName: minio-data
---
apiVersion: v1
kind: Service
metadata:
name: minio-api
namespace: minio
spec:
selector:
app: minio
ports:
- port: 9000
targetPort: 9000
---
apiVersion: v1
kind: Service
metadata:
name: minio-console
namespace: minio
spec:
selector:
app: minio
ports:
- port: 9001
targetPort: 9001
oc apply -f minio.yaml oc create route edge minio-console --service=minio-console --port=9001 -n minio
Before anything can upload to it, the buckets must exist — MinIO doesn't auto-create them. Open the MinIO console to create all three:
oc get route minio-console -n minio -o jsonpath='{.spec.host}'
# Open https://<that host> in a browser
# Log in: admin /
# Buckets → Create Bucket → create each of the three buckets below
The three buckets and their purposes:
sre-models— LLM model weights (read by KServe's storage-initializer)sre-logs— Loki log chunks and index (read-write by LokiStack)sre-traces— Tempo trace data (read-write by TempoStack)
All three use the same MinIO credentials: access key admin, secret , endpoint http://minio-api.minio.svc.cluster.local:9000. Each component's Kubernetes Secret names a different bucket.
VII.Serving the LLM
This is the piece almost everything else in the pipeline exists to feed context into. Qwen2.5-7B-Instruct (AWQ 4-bit) is chosen for this demo because it fits within a single T4 GPU (4.5 GiB weights, ~10 GiB remaining for KV cache), ships an official AWQ quantization repo requiring no third-party tooling, and properly supports a dedicated system role — making it well-suited for structured JSON output in a constrained single-GPU demo environment. In production, substitute with your organisation's approved model (Claude via API, a Red Hat-supported model, or a fine-tuned model trained on your own incident history).
Download and upload the model
# Start from a clean local dir -- never re-download into a directory that may hold # files from a different model/precision (their leftover shards will ride along) rm -rf ./Qwen2.5-7B-Instruct-AWQ pip install huggingface_hub export HF_TOKEN=<your-hf-token> python3 - <<'PYEOF' import os from huggingface_hub import snapshot_download snapshot_download( repo_id="Qwen/Qwen2.5-7B-Instruct-AWQ", local_dir="./Qwen2.5-7B-Instruct-AWQ", token=os.environ["HF_TOKEN"], ignore_patterns=["*.bin", "*.pt"], # *.bin / *.pt → skip any legacy PyTorch files. # Qwen's official AWQ repo ships a single safetensors shard (~4.5 GB). # No consolidated.safetensors or shard index to worry about. ) PYEOF # Confirm download: expect one model safetensors shard (~4.5 GB) plus tokenizer/config files ls -lh ./Qwen2.5-7B-Instruct-AWQ # Get the MinIO console route oc get route minio-console -n minio -o jsonpath='{.spec.host}' # Open https://<that host> in a browser, log in as admin /# Navigate to: Buckets → sre-models → Upload # Upload the entire contents of the ./Qwen2.5-7B-Instruct-AWQ/ directory # into a path named Qwen2.5-7B-Instruct-AWQ/ inside the bucket. # Expected files: one .safetensors shard (~4.5 GB) plus tokenizer and config files.
ServingRuntime and InferenceService
Apply the ServingRuntime as a standalone manifest — no discovery dance, no copying a pre-installed template. This is the exact runtime this pipeline serves Qwen2.5-7B-Instruct through:
vllm-serving-runtime.yamlapiVersion: serving.kserve.io/v1alpha1
kind: ServingRuntime
metadata:
annotations:
opendatahub.io/apiProtocol: REST
opendatahub.io/recommended-accelerators: '["nvidia.com/gpu"]'
opendatahub.io/runtime-version: v0.18.0
opendatahub.io/serving-runtime-scope: global
opendatahub.io/template-display-name: vLLM NVIDIA GPU ServingRuntime for KServe
opendatahub.io/template-name: vllm-cuda-runtime-template
openshift.io/display-name: vLLM NVIDIA GPU ServingRuntime for KServe
labels:
opendatahub.io/dashboard: "true"
name: qwen2-5-7b-instruct
namespace: llm-serving
spec:
annotations:
opendatahub.io/kserve-runtime: vllm
prometheus.io/path: /metrics
prometheus.io/port: "8080"
containers:
- args:
- --port=8080
- --model=/mnt/models
- --served-model-name={{.Name}}
command:
- python
- -m
- vllm.entrypoints.openai.api_server
env:
- name: HF_HOME
value: /tmp/hf_home
image: registry.redhat.io/rhaii/vllm-cuda-rhel9@sha256:5800e12b2a465f15961fcf34b645d79ed4f91ec9161eab22b1205d12682183c8
name: kserve-container
ports:
- containerPort: 8080
protocol: TCP
multiModel: false
supportedModelFormats:
- autoSelect: true
name: vLLM
The --served-model-name={{.Name}} in the runtime above is a generic template default, templated from whatever resource name uses this runtime. The InferenceService below overrides it — along with everything else that makes this deployment actually work on a T4 — with its own explicit args, which is what vLLM actually starts with:
Before the InferenceService, create the ServiceAccount that carries the model bucket's credentials, and the Data Connection Secret it references:
model-serving-sa.yamlapiVersion: v1
kind: ServiceAccount
metadata:
name: sre-models-sa
namespace: llm-serving
secrets:
- name: sre-models # the Data Connection Secret below -- listing it here is what makes
# KServe's storage-initializer pick up its credentials automatically
data-connection.yaml
apiVersion: v1
kind: Secret
metadata:
name: sre-models
namespace: llm-serving
labels:
opendatahub.io/dashboard: "true"
opendatahub.io/managed: "true"
annotations:
opendatahub.io/connection-type: s3
opendatahub.io/connection-type-protocol: s3
opendatahub.io/connection-type-ref: s3
openshift.io/description: sre-models
openshift.io/display-name: sre-models
stringData:
AWS_ACCESS_KEY_ID: admin
AWS_SECRET_ACCESS_KEY:
AWS_DEFAULT_REGION: us-east-1
AWS_S3_BUCKET: sre-models
AWS_S3_ENDPOINT: http://minio-api.minio.svc.cluster.local:9000
Then the InferenceService itself — the real, working configuration this pipeline runs on:
inference-service.yamlapiVersion: serving.kserve.io/v1beta1
kind: InferenceService
metadata:
annotations:
modelFormat: vLLM
opendatahub.io/connection-path: Qwen2.5-7B-Instruct-AWQ/
opendatahub.io/connections: sre-models
opendatahub.io/hardware-profile-name: nvidia
opendatahub.io/hardware-profile-namespace: redhat-ods-applications
opendatahub.io/model-type: generative
openshift.io/description: qwen2-5-7b-instruct
openshift.io/display-name: qwen2-5-7b-instruct
security.opendatahub.io/enable-auth: "false"
serving.kserve.io/deploymentMode: Standard
labels:
networking.kserve.io/visibility: exposed
opendatahub.io/dashboard: "true"
name: qwen2-5-7b-instruct
namespace: llm-serving
spec:
predictor:
automountServiceAccountToken: false
deploymentStrategy:
type: Recreate
maxReplicas: 1
minReplicas: 1
model:
args:
- --served-model-name=qwen2-5-7b-instruct # must match LLM_MODEL in the SRE Intelligence service --
# without it the endpoint registers under the literal path
# /mnt/models and every inference call 404s
- --quantization=awq # tells vLLM the weights are AWQ 4-bit; without this flag vLLM
# attempts to load them as fp16 and errors on the quant config
- --max-model-len=8192 # leaves ~10 GiB of the T4's 16 GB for KV cache after AWQ weights (~4.5 GiB)
- --enforce-eager # skips CUDA graph capture overhead; the latency cost is irrelevant
# for a single-request RCA call
env:
- name: VLLM_LOGGING_LEVEL
value: DEBUG
modelFormat:
name: vLLM
resources:
limits:
cpu: "4"
memory: 8Gi
nvidia.com/gpu: "1"
requests:
cpu: "2"
memory: 4Gi
nvidia.com/gpu: "1"
runtime: qwen2-5-7b-instruct
storage:
key: sre-models # the Data Connection Secret's name
path: Qwen2.5-7B-Instruct-AWQ/
nodeSelector:
nvidia.com/gpu.present: "true" # matches any GPU node NFD has labeled -- no hostname to maintain
serviceAccountName: sre-models-sa
timeout: 30
serving.kserve.io/deploymentMode: Standard is RHOAI's current name for what the DataScienceCluster setting earlier called RawDeployment mode — same behavior, same no-Knative deployment, just the name the Dashboard and newer InferenceService annotations use. The two hardware-profile annotations are what bind this deployment to the HardwareProfile created in the previous section — that's the confirmed mechanism, not a guess. The networking.kserve.io/visibility: exposed label creates an OpenShift Route in front of the predictor for convenience (it's what populates status.url below with a public HTTPS address) — the SRE Intelligence service itself never uses that Route, only the internal cluster address.
oc apply -f model-serving-sa.yaml
oc apply -f data-connection.yaml
oc apply -f vllm-serving-runtime.yaml
oc apply -f inference-service.yaml
oc get pods -n llm-serving -w # model loading takes 3-5 min
Test it
# .status.address.url is the internal, cluster-only endpoint -- the one the enrichment # service actually calls -- and it already includes the port, unlike the external route. LLM_ENDPOINT=$(oc get inferenceservice qwen2-5-7b-instruct \ -n llm-serving -o jsonpath='{.status.address.url}') # Internal address is plain http:// -- no TLS, no -k needed for this call. (If you instead # derive from .status.url, that's the external https:// Route created by the "exposed" # label above, and DOES need -k against its self-signed cluster cert.) oc run llm-test --rm -i --restart=Never --image=curlimages/curl:latest \ -n llm-serving -- \ curl -s -X POST "${LLM_ENDPOINT}/v1/chat/completions" \ -H "Content-Type: application/json" \ -d '{"model":"qwen2-5-7b-instruct", "messages":[{"role":"user","content":"Reply with OK only"}], "max_tokens":10}'
$LLM_ENDPOINT — it goes into the SRE Intelligence service's Secret as LLM_BASE_URL exactly as-is, not hand-typed. A mismatch there silently breaks every tier at once, since every alert falls into the SRE Intelligence service's exception handler and posts a Slack escalation instead of demonstrating Tier 1/2.VIII.Logs, Traces & Alerting
This is the single largest chunk of new setup time in the whole build — budget roughly two hours and don't try to rush it. Everything here is inline (not linked to Red Hat's docs) because these three operators are specific to this pipeline's own observability design, not general RHOAI platform dependencies.
sre-models, sre-logs, sre-traces). Only the bucketnames / bucket field differs between the two secrets.Loki (logs)
- Block storage (StorageClass) — for internal PVCs: write-ahead log, index cache, compactor working space. Verify a StorageClass exists before installing:
oc get sc. - Object storage (S3-compatible) — for actual log data chunks and indices. This is the MinIO bucket configured in the Storage section.
storageClassName field in the LokiStack CR refers to block storage, not object storage.stable-6.5 for Logging 6.5/6.6). The Loki Operator installs into openshift-operators-redhat (AllNamespaces mode); the Logging Operator installs into openshift-logging. Getting the namespace wrong causes silent install failures.Install both operators via OperatorHub, following the official guides — channel names drift between OpenShift versions, so the upstream docs are the authoritative source for the current channel:
Use exactly these settings:
Loki Operator
- Update channel →
stable-6.5 - Installed Namespace →
openshift-operators-redhat(operator recommended) - Select Enable Operator recommended cluster monitoring on this Namespace
- Update approval → Automatic
Red Hat OpenShift Logging Operator
- Update channel →
stable-6.5 - Installed Namespace →
openshift-logging(operator recommended) - Select Enable Operator recommended cluster monitoring on this Namespace
- Update approval → Automatic
Verify both operators are installed before continuing:
oc get csv -n openshift-operators-redhat | grep loki oc get csv -n openshift-logging | grep cluster-logging
Create the LokiStack instance — small, single-pod, MinIO-backed:
oc create secret generic loki-s3-creds -n openshift-logging \ --from-literal=access_key_id=admin \ --from-literal=access_key_secret='lokistack.yaml' \ --from-literal=bucketnames=sre-logs \ --from-literal=endpoint=http://minio-api.minio.svc.cluster.local:9000 \ --from-literal=region=us-east-1
apiVersion: loki.grafana.com/v1
kind: LokiStack
metadata:
name: sre-observability
namespace: openshift-logging
spec:
size: 1x.demo
storageClassName: gp3-csi # check `oc get sc` -- use your cluster's default StorageClass
storage:
schemas:
- version: v13
effectiveDate: "2024-01-01"
secret:
name: loki-s3-creds
type: s3
tenants:
mode: openshift-logging
# -n openshift-logging is redundant with the YAML's own metadata.namespace, and that's # the point -- apply it explicitly so a stray context switch or console habit can't land # this in openshift-operators-redhat by accident (easy to do via the console's "Create # instance" button if the project selector is still on the operator's namespace). oc apply -f lokistack.yaml -n openshift-logging oc get lokistack sre-observability -n openshift-logging oc get pods -n openshift-logging -w # ~3-5 min
Create the collector service account and grant the required ClusterRoles before creating the ClusterLogForwarder. Grant permissions BEFORE creating the CLF — missing a ClusterRoleBinding causes the Operator to destroy the entire collector DaemonSet and stop all log collection:
# Create the collector service account oc create sa logging-collector -n openshift-logging # Grant the three required roles (MUST be done before creating ClusterLogForwarder) oc adm policy add-cluster-role-to-user logging-collector-logs-writer \ -z logging-collector -n openshift-logging oc adm policy add-cluster-role-to-user collect-application-logs \ -z logging-collector -n openshift-logging oc adm policy add-cluster-role-to-user collect-infrastructure-logs \ -z logging-collector -n openshift-logging # Verify all three bindings exist oc get clusterrolebinding -o wide | grep logging-collector
Now create the ClusterLogForwarder to forward the demo workload's logs into LokiStack:
tls.ca block, the collector cannot verify the gateway certificate — it fails silently at runtime and no logs reach LokiStack, even though the CLF status shows Ready: True.apiVersion: observability.openshift.io/v1 kind: ClusterLogForwarder metadata: name: instance # must be "instance" -- ClusterLogForwarder is a singleton in 6.x namespace: openshift-logging spec: serviceAccount: name: logging-collector inputs: - name: business-workloads-app-logs type: application application: includes: - namespace: business-workloads - namespace: sre-intelligence outputs: - name: loki-app type: lokiStack lokiStack: target: name: sre-observability namespace: openshift-logging authentication: token: {from: serviceAccount} tls: ca: configMapName: sre-observability-gateway-ca-bundle # auto-created by the Loki Operator; name follows pattern <lokistack-name>-gateway-ca-bundle key: service-ca.crt pipelines: - name: app-to-loki inputRefs: [business-workloads-app-logs] outputRefs: [loki-app]
oc apply -f clusterlogforwarder.yaml
# All three conditions must be True before logs flow
oc get clusterlogforwarder instance -n openshift-logging \
-o jsonpath='{.status.conditions}' | python3 -m json.tool
Tempo + OpenTelemetry (traces)
openshift-tempo-operator, not openshift-tempo — the latter is reserved for the actual TempoStack instance, its S3 secret, and the OpenTelemetry Collector.Install the Tempo Operator via OperatorHub, following the official guide:
Once the operator controller pod is healthy, continue with the TempoStack instance below.
oc create secret generic tempo-s3-creds -n openshift-tempo \ --from-literal=access_key_id=admin \ --from-literal=access_key_secret='' \ --from-literal=bucket=sre-traces \ --from-literal=endpoint=http://minio-api.minio.svc.cluster.local:9000
Note the field is bucket here, singular — Tempo's storage secret schema differs slightly from Loki's bucketnames.
apiVersion: tempo.grafana.com/v1alpha1 kind: TempoStack metadata: name: sre-observability namespace: openshift-tempo spec: managementState: Managed # required -- Unmanaged leaves the operator hands-off, nothing reconciles storageSize: 10Gi storage: secret: {name: tempo-s3-creds, type: s3} resources: total: limits: {memory: 2Gi, cpu: "2"} tenants: mode: openshift # required for the Observe → Tracing console tab; # UIPlugin instances without multi-tenancy are not shown authentication: - tenantName: dev tenantId: "1610b0c3-c509-4592-a256-a1871353dbfa" observability: metrics: createServiceMonitors: true # lets user-workload Prometheus scrape Tempo component metrics createPrometheusRules: true # installs alerting rules for Tempo component health template: gateway: enabled: true # gateway is required when mode: openshift -- provides auth/authz for multi-tenancy; # cannot combine gateway with jaegerQuery.ingress (gateway serves the Jaeger UI instead) queryFrontend: jaegerQuery: enabled: true monitorTab: enabled: true # enables the RED-metrics Monitor tab in the Jaeger console prometheusEndpoint: https://thanos-querier.openshift-monitoring.svc.cluster.local:9092 # port 9092 = Thanos Querier for user-workload monitoring (per RH docs); # requires the spanmetrics connector in the OTEL Collector below
oc apply -f tempostack.yaml -n openshift-tempo oc get tempostack sre-observability -n openshift-tempo oc get pods -n openshift-tempo -w
The Tempo Operator does not auto-create tenant RBAC. Create the write and read ClusterRoles manually and bind them:
tempo-tenant-rbac.yamlapiVersion: rbac.authorization.k8s.io/v1 kind: ClusterRole metadata: name: tempo-sre-observability-dev-write rules: - apiGroups: [tempo.grafana.com] resources: [dev] resourceNames: [traces] verbs: [create] --- apiVersion: rbac.authorization.k8s.io/v1 kind: ClusterRole metadata: name: tempo-sre-observability-dev-read rules: - apiGroups: [tempo.grafana.com] resources: [dev] resourceNames: [traces] verbs: [get]
oc apply -f tempo-tenant-rbac.yaml # Bind write to the OTEL collector SA. # Expected warning: "ServiceAccount 'otel-collector-collector' not found" # The SA is created by the OpenTelemetry Operator when otel-collector.yaml is applied # in the next section. The binding is stored now and takes effect automatically once the SA exists. oc adm policy add-cluster-role-to-user \ tempo-sre-observability-dev-write \ -z otel-collector-collector -n business-workloads # Bind read to all authenticated users (console UI + SRE Intelligence service). # Expected warning: "Group 'system:authenticated' not found" # system:authenticated is a virtual OpenShift group (not stored in etcd) representing # all logged-in users. The warning is cosmetic -- the binding is valid and effective. oc adm policy add-cluster-role-to-group \ tempo-sre-observability-dev-read \ system:authenticated
OpenShift Cluster Observability Operator — console UI for traces and logs
The Cluster Observability Operator (COO) provides the Observe → Tracing and Observe → Logs tabs in the OpenShift console via its UIPlugin CRD. Install it first, then create one UIPlugin per capability.
Install the COO via OperatorHub, following the official guide:
Use exactly these settings:
- Update channel →
stable - Version → 1.0.0 or later
- Installation mode → All namespaces on the cluster (default)
- Installed Namespace →
openshift-cluster-observability-operator(operator recommended) - Select Enable Operator recommended cluster monitoring on this Namespace
- Update approval → Automatic
Verify installation: go to Operators → Installed Operators and confirm the Cluster Observability Operator entry appears. Then confirm the operator pod is running:
oc get pods -n openshift-cluster-observability-operator
Once healthy, create the two UIPlugin resources below.
UIPlugin — Distributed Tracing (Observe → Tracing)
The plugin auto-discovers TempoStack instances that have multi-tenancy enabled (tenants.mode: openshift + template.gateway.enabled: true). TempoStacks without a gateway are explicitly not shown in the console. The sre-observability TempoStack qualifies because both are configured in its spec above.
apiVersion: observability.openshift.io/v1alpha1
kind: UIPlugin
metadata:
name: distributed-tracing
namespace: openshift-cluster-observability-operator
spec:
type: DistributedTracing
distributedTracing:
timeout: 30s
oc apply -f uiplugin-tracing.yaml oc get uiplugin distributed-tracing -n openshift-cluster-observability-operator
UIPlugin — Logging (Observe → Logs)
The Logging plugin requires an explicit reference to the LokiStack instance because a cluster can have multiple LokiStack deployments.
uiplugin-logging.yamlapiVersion: observability.openshift.io/v1alpha1
kind: UIPlugin
metadata:
name: logging
namespace: openshift-cluster-observability-operator
spec:
type: Logging
logging:
lokiStack:
name: sre-observability # must match your LokiStack instance name exactly
logsLimit: 800
timeout: 30s
schema: viaq # viaq (default) or otel or select (lets user choose in UI)
oc apply -f uiplugin-logging.yaml oc get uiplugin logging -n openshift-cluster-observability-operator
The COO reconciles each UIPlugin into a ConsolePlugin resource and registers it with the console operator automatically. Wait ~30s for the console to reload, then verify both plugins are active:
oc get uiplugin distributed-tracing logging -n openshift-cluster-observability-operator
- Observe → Tracing — search by service name
order-serviceorinventory-serviceto see spans from the demo workload. - Observe → Logs — filter by namespace
business-workloadsorsre-intelligenceto see application logs from the enrichment pipeline.
Install the OpenTelemetry Operator via OperatorHub, following the official guide:
Once the operator controller pod is healthy, create the Collector instance:
otel-collector.yamlapiVersion: opentelemetry.io/v1beta1
kind: OpenTelemetryCollector
metadata:
name: otel-collector
namespace: business-workloads
spec:
mode: deployment
observability:
metrics:
enableMetrics: true # creates a ServiceMonitor so user-workload Prometheus scrapes this collector
config:
extensions:
bearertokenauth:
filename: /var/run/secrets/kubernetes.io/serviceaccount/token
connectors:
span_metrics: # derives RED (rate/error/duration) metrics from spans and exports them
# in Prometheus format -- required for the Jaeger Monitor tab to work
metrics_flush_interval: 15s
processors:
batch:
timeout: 5s # flush every 5s regardless of batch size
send_batch_max_size: 10000
receivers:
otlp:
protocols:
grpc: {}
http: {}
exporters:
otlp_http/tempo:
endpoint: https://tempo-sre-observability-gateway.openshift-tempo.svc.cluster.local:8080/api/traces/v1/dev
auth: {authenticator: bearertokenauth}
headers: {X-Scope-OrgID: dev} # The otlphttp exporter appends /v1/traces → final URL: /api/traces/v1/dev/v1/traces.
# Uses HTTP+bearer-token on port 8080 (same pattern as SRE Intelligence service reads).
# gRPC port 8090 requires mTLS client certs; HTTP port 8080 accepts bearer token.
# Tempo NetworkPolicy allows only gateway → distributor, so all external writes
# must go through the gateway.
tls:
ca_file: /var/run/secrets/kubernetes.io/serviceaccount/service-ca.crt # OpenShift automatically injects the service-serving CA
# at this path in every pod -- no ConfigMap or volumeMount needed
prometheus:
endpoint: 0.0.0.0:8889
add_metric_suffixes: false
resource_to_telemetry_conversion:
enabled: true # propagates k8s resource attributes (service.name etc.) as Prometheus labels
service:
extensions: [bearertokenauth]
pipelines:
traces:
receivers: [otlp]
processors: [batch]
exporters: [otlp_http/tempo, span_metrics] # spans feed both Tempo storage and the span_metrics connector
metrics:
receivers: [span_metrics]
processors: [batch]
exporters: [prometheus]
oc apply -f otel-collector.yaml
oc get svc -n business-workloads -l app.kubernetes.io/managed-by=opentelemetry-operator
# If the generated Service name differs from otel-collector-collector, update
# OTEL_EXPORTER_OTLP_ENDPOINT on both demo-workload Deployments and roll them.
order-service to be running. After completing section IX — The Demo Workload, run the commands at the end of that section to confirm Tempo is receiving traces and the field names match what tempo_trace_summary() expects.Prometheus alerting rules
business-workloads-alerts.yamlapiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
name: sre-platform-rules
namespace: business-workloads
labels:
openshift.io/prometheus-rule-evaluation-scope: leaf-prometheus
spec:
groups:
- name: business-workloads.pod-health
interval: 10s # tightened from the usual 30s default -- demo speed over production
# noise-avoidance. See the Testing section for the full timing budget this buys.
rules:
- alert: PodCrashLoopingDetected
expr: increase(kube_pod_container_status_restarts_total[5m]) > 2
for: 1m # Threshold of > 2 (3+ restarts in 5 min) prevents Test 1's OOMKill + pipeline pod_restart
# (2 total restarts) from also firing PodCrashLoopingDetected.
# Test 4's rapid crash loop (exponential backoff: 10s, 20s, 40s...) reaches
# 3 restarts within the first ~70 seconds, so it still fires reliably.
labels: {severity: warning, tier: "1"}
annotations:
summary: "Pod {{ $labels.pod }} is crash looping"
runbook: "restart-pod"
- alert: PodOOMKilled
expr: kube_pod_container_status_last_terminated_reason{reason="OOMKilled"} == 1
for: 0m
labels: {severity: warning, tier: "1"}
annotations:
summary: "Pod {{ $labels.pod }} was OOMKilled"
runbook: "oom-increase-limits"
# This is the alert the demo workload's cascade is designed to trip -- the flagship demo.
# for: 0m is deliberate, not a demo shortcut: this expr is a boolean "did this container
# just get OOMKilled" state, not a rate needing debounce -- there's no flapping risk to
# guard against the way there would be for, say, a CPU-threshold alert.
- alert: SuspiciousEgressTraffic
expr: rate(container_network_transmit_bytes_total{namespace="business-workloads",pod=~"inventory-service.*"}[5m]) > 1000
for: 20s
# 1000 bytes/sec threshold -- in-cluster /health responses are tiny (~200 bytes each),
# so even 100 concurrent connections only produce ~3KB/s of transmit. Idle baseline
# is ~120 bytes/sec, so 1000 is comfortably above noise and comfortably below the spike.
# 5m rate window so the spike persists long enough for Slack approval.
labels: {severity: warning, tier: "2"}
annotations:
summary: "Unexpected egress volume from {{ $labels.pod }}"
runbook: "quarantine-pod"
# Genuinely reachable: inventory-service's /simulate-egress-spike endpoint (Demo Workload
# section) drives real transmit bytes past this threshold. 1000 bytes/sec is comfortably above
# idle baseline (a handful of kubelet health probes) and comfortably below what a few
# hundred concurrent request loops produce -- tune it during your dry run if your node's
# actual idle noise floor differs.
- alert: DataIntegrityCheckFailed
expr: increase(app_data_integrity_failures_total{namespace="business-workloads"}[2m]) > 0
for: 15s
labels: {severity: critical, tier: "3"}
annotations:
summary: "Data integrity check failing for {{ $labels.pod }}"
# Deliberately NO runbook key for this alertname -- see the Runbooks section. Requires
# the order-service ServiceMonitor and User Workload Monitoring (Audit Metrics section) to
# actually be scraped; the PrometheusRule itself evaluates fine either way, but the metric
# won't exist as a time series until both of those are in place.
- alert: HPAThrashing
expr: |
abs(kube_horizontalpodautoscaler_status_desired_replicas
- kube_horizontalpodautoscaler_status_current_replicas) > 2
for: 5m
labels: {severity: warning, tier: "1"}
annotations:
summary: "HPA {{ $labels.horizontalpodautoscaler }} is thrashing"
runbook: "hpa-stabilize"
- alert: HighMemoryPressure
expr: |
(container_memory_working_set_bytes{container!=""}
/ kube_pod_container_resource_limits{resource="memory"}) > 0.90
for: 3m
labels: {severity: warning, tier: "2"}
annotations:
summary: "Pod {{ $labels.pod }} memory usage above 90%"
runbook: "memory-pressure"
- alert: UpstreamServiceLatencyHigh
expr: histogram_quantile(0.95, rate(http_request_duration_seconds_bucket[5m])) > 2.0
for: 3m
labels: {severity: critical, tier: "3"}
annotations:
summary: "P95 latency above 2s for {{ $labels.job }}"
runbook: "latency-investigation"
oc apply -f business-workloads-alerts.yaml
IX.The Demo Workload
Two small services whose entire purpose is to fail in a specific, realistic, reproducible way. See the architecture section above for why this shape was chosen instead of a single stock pod.
inventory-service — the slow dependency, and the egress spike
Latency is toggled at runtime via /inject-latency — no redeploy needed when you trigger the fault later. It also carries the Tier 2 scenario's fault: /simulate-egress-spike genuinely floods outbound requests at order-service over the pod's real network interface (not loopback, which cAdvisor's network metrics don't see), so container_network_transmit_bytes_total — an existing, already-scraped platform metric, no new instrumentation needed for it — genuinely spikes. Same concurrent-burst pattern as order-service's /self-load below, just aimed outward instead of at itself.
import os, time, threading, logging
from concurrent.futures import ThreadPoolExecutor
import httpx
from fastapi import FastAPI
logging.basicConfig(level=logging.INFO, format='%(asctime)s %(levelname)s %(message)s')
log = logging.getLogger("inventory-service")
app = FastAPI(title="inventory-service")
_state = {"latency_seconds": 0.0}
_lock = threading.Lock()
ORDER_SERVICE_URL = os.getenv("ORDER_SERVICE_URL", "http://order-service.business-workloads.svc.cluster.local:8080")
@app.get("/health")
def health():
return {"status": "ok"}
@app.get("/work")
def work():
with _lock:
delay = _state["latency_seconds"]
if delay > 0:
log.info(f"injected latency active: sleeping {delay}s before responding")
time.sleep(delay)
return {"status": "done", "delay_applied": delay}
@app.post("/inject-latency")
def inject_latency(seconds: float = 0.0):
with _lock:
_state["latency_seconds"] = seconds
log.warning(f"latency injection set to {seconds}s -- /work will now sleep before responding")
return {"latency_seconds": seconds}
@app.post("/simulate-egress-spike")
def simulate_egress_spike(n: int = 100, duration: int = 60):
# n concurrent loops hammering order-service's /health for `duration` seconds -- genuinely
# spikes this pod's transmit bytes on its real interface. Self-expiring, no reset needed:
# the threads stop on their own once `duration` elapses.
# n=100 (not 400) -- each httpx.Client holds connection buffers; 400 concurrent clients
# OOMKill the pod before the egress alert's `for: 20s` window elapses. 100 is enough to
# cross the 1000 bytes/sec threshold while staying well within 512Mi.
# duration=60 (not 30) -- gives the alert time to fire (2m rate window + 20s for clause).
log.warning(f"egress spike simulation starting: {n} concurrent loops for {duration}s")
threading.Thread(target=_egress_burst, args=(n, duration), daemon=True).start()
return {"started": True, "concurrency": n, "duration_seconds": duration}
def _egress_burst(n: int, duration: int):
end = time.time() + duration
def _hammer():
with httpx.Client(timeout=2) as c:
while time.time() < end:
try:
c.get(f"{ORDER_SERVICE_URL}/health")
except Exception:
pass
with ThreadPoolExecutor(max_workers=n) as pool:
list(pool.map(lambda _: _hammer(), range(n)))
log.info("egress spike simulation finished")
order-service — the service with no timeout, and the data-integrity check
/self-load fires a burst of concurrent requests against its own /request endpoint, in-process — the "load generator" is ~10 lines of stdlib ThreadPoolExecutor, not a separate tool. It also carries the Tier 3 scenario: a background loop that, once /simulate-corruption is toggled on, genuinely fails a periodic self-check and increments a real Prometheus counter — deliberately with no matching runbook entry (see the Runbooks section), so the LLM has nothing to ground an answer in and should reasonably escalate rather than guess.
import os, time, threading, logging
from concurrent.futures import ThreadPoolExecutor
import httpx
from fastapi import FastAPI
from prometheus_client import Counter, make_asgi_app
logging.basicConfig(level=logging.INFO, format='%(asctime)s %(levelname)s %(message)s')
log = logging.getLogger("order-service")
app = FastAPI(title="order-service")
INVENTORY_URL = os.getenv("INVENTORY_URL",
"http://inventory-service.business-workloads.svc.cluster.local:8080")
# -- Data-integrity check (Tier 3 scenario) --------------------------------
data_integrity_failures = Counter('app_data_integrity_failures_total',
'Simulated data integrity check failures')
app.mount("/metrics", make_asgi_app())
_corruption_state = {"enabled": False}
@app.post("/simulate-corruption")
def simulate_corruption(enabled: bool = True):
_corruption_state["enabled"] = enabled
log.warning(f"data integrity corruption simulation set to {enabled}")
return {"corruption_enabled": enabled}
def _integrity_check_loop():
# Runs for the life of the pod, not just while corruption is toggled on -- a real
# periodic self-check that happens to always pass until you flip the switch.
while True:
time.sleep(3)
if _corruption_state["enabled"]:
data_integrity_failures.inc()
log.error("data integrity check FAILED -- simulated corruption active")
threading.Thread(target=_integrity_check_loop, daemon=True).start()
# No timeout, no circuit breaker -- this IS the fault the demo exists to show.
# A production client would set both; skip that here on purpose.
client = httpx.Client(timeout=None)
@app.get("/health")
def health():
return {"status": "ok"}
@app.get("/request")
def make_request():
start = time.time()
resp = client.get(f"{INVENTORY_URL}/work")
elapsed = time.time() - start
if elapsed > 1.0:
log.warning(f"slow inventory-service response: GET {INVENTORY_URL}/work took {elapsed:.2f}s")
return {"downstream_status": resp.status_code, "elapsed_seconds": round(elapsed, 2)}
@app.post("/self-load")
def self_load(n: int = 80):
# n is a calibration knob, not a constant -- tune it during your dry run
# until you reliably see OOMKill within ~20-30s on your actual node sizing.
log.info(f"self-load firing {n} concurrent requests to /request")
with ThreadPoolExecutor(max_workers=n) as pool:
list(pool.map(lambda _: _fire_one(), range(n)))
return {"fired": n}
def _fire_one():
try:
httpx.get("http://localhost:8080/request", timeout=None)
except Exception as e:
log.warning(f"self-load request failed: {e}")
Shared build files
Identical for both services — only main.py differs. opentelemetry-bootstrap -a install auto-installs the FastAPI and httpx instrumentors matching what's actually imported, which is how the slow call ends up as a real span in Tempo rather than a synthetic one.
fastapi==0.110.0
uvicorn[standard]==0.27.1
httpx==0.26.0
opentelemetry-distro==0.45b0
opentelemetry-exporter-otlp==1.24.0
prometheus-client==0.20.0 # only order-service's /metrics uses this -- harmless unused import in inventory-service
Containerfile (both services)
FROM registry.access.redhat.com/ubi9/python-311:latest WORKDIR /app COPY requirements.txt . RUN pip install --no-cache-dir -r requirements.txt && opentelemetry-bootstrap -a install COPY main.py . EXPOSE 8080 CMD ["opentelemetry-instrument", "uvicorn", "main:app", "--host", "0.0.0.0", "--port", "8080"]
Build and push
# The internal registry's external route is NOT enabled by default on OpenShift -- # without it there's no host your local `podman push` can reach (the internal Service # DNS only resolves inside the cluster's pod network). Safe to re-run, no-op if already set. oc patch configs.imageregistry.operator.openshift.io/cluster \ --type=merge -p '{"spec":{"defaultRoute":true}}' oc get route default-route -n openshift-image-registry -w # --tls-verify=false: the route's default cert is signed by the cluster's internal CA, # which your local podman won't trust out of the box. # --platform linux/amd64: REQUIRED if you're building on Apple Silicon (or any arm64 # workstation) -- the cluster's EC2 nodes are x86_64. Without this flag, podman builds # for your host's native arch, the image runs fine locally, and then fails on the # cluster with "exec /usr/bin/container-entrypoint: exec format error." Cross-arch # build is slower (QEMU emulation), not incorrect. REGISTRY=$(oc get route default-route -n openshift-image-registry -o jsonpath='{.spec.host}') podman login --tls-verify=false ${REGISTRY} -u $(oc whoami) -p $(oc whoami -t) for svc in inventory-service order-service; do podman build --platform linux/amd64 -t ${svc}:v1 ./${svc} podman tag ${svc}:v1 ${REGISTRY}/business-workloads/${svc}:v1 podman push --tls-verify=false ${REGISTRY}/business-workloads/${svc}:v1 done
Deploy both services
order-service's memory limit is deliberately tight (96Mi) so a burst of blocked requests can actually exhaust it within the demo window. The OTEL_EXPORTER_OTLP_ENDPOINT below targets the Collector Service created in the observability section — deploy that section before you need traces to appear; until then these pods just log OTLP connection-refused warnings on export (non-fatal, spans dropped, nothing else breaks).
apiVersion: apps/v1
kind: Deployment
metadata:
name: inventory-service
namespace: business-workloads
spec:
replicas: 1
selector:
matchLabels: {app: inventory-service}
template:
metadata:
labels: {app: inventory-service}
spec:
containers:
- name: inventory-service
image: image-registry.openshift-image-registry.svc:5000/business-workloads/inventory-service:v1
ports: [{containerPort: 8080}]
env:
- {name: ORDER_SERVICE_URL, value: "http://order-service.business-workloads.svc.cluster.local:8080"}
- {name: OTEL_SERVICE_NAME, value: "inventory-service"}
- {name: OTEL_EXPORTER_OTLP_ENDPOINT, value: "http://otel-collector-collector.business-workloads.svc.cluster.local:4317"}
- {name: OTEL_EXPORTER_OTLP_INSECURE, value: "true"}
- {name: OTEL_TRACES_EXPORTER, value: "otlp"}
- {name: OTEL_METRICS_EXPORTER, value: "none"}
- {name: OTEL_LOGS_EXPORTER, value: "none"}
resources:
requests: {cpu: 100m, memory: 128Mi}
limits: {cpu: 500m, memory: 512Mi} # must survive the egress burst (400 concurrent httpx
# connections) long enough for SuspiciousEgressTraffic
# to fire (for: 20s) -- 256Mi OOMs before that
readinessProbe: {httpGet: {path: /health, port: 8080}, initialDelaySeconds: 5}
livenessProbe: {httpGet: {path: /health, port: 8080}, initialDelaySeconds: 10}
---
apiVersion: v1
kind: Service
metadata: {name: inventory-service, namespace: business-workloads}
spec:
selector: {app: inventory-service}
ports: [{port: 8080, targetPort: 8080}]
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: order-service
namespace: business-workloads
spec:
replicas: 1
selector:
matchLabels: {app: order-service}
template:
metadata:
labels: {app: order-service}
spec:
containers:
- name: order-service
image: image-registry.openshift-image-registry.svc:5000/business-workloads/order-service:v1
ports: [{containerPort: 8080}]
env:
- {name: INVENTORY_URL, value: "http://inventory-service.business-workloads.svc.cluster.local:8080"}
- {name: OTEL_SERVICE_NAME, value: "order-service"}
- {name: OTEL_EXPORTER_OTLP_ENDPOINT, value: "http://otel-collector-collector.business-workloads.svc.cluster.local:4317"}
- {name: OTEL_EXPORTER_OTLP_INSECURE, value: "true"}
- {name: OTEL_TRACES_EXPORTER, value: "otlp"}
- {name: OTEL_METRICS_EXPORTER, value: "none"}
- {name: OTEL_LOGS_EXPORTER, value: "none"}
resources:
requests: {cpu: 100m, memory: 64Mi}
limits: {cpu: 500m, memory: 96Mi} # deliberately tight -- see note above
readinessProbe: {httpGet: {path: /health, port: 8080}, initialDelaySeconds: 5}
livenessProbe: {httpGet: {path: /health, port: 8080}, initialDelaySeconds: 10}
---
apiVersion: v1
kind: Service
metadata:
name: order-service
namespace: business-workloads
labels: {app: order-service} # ServiceMonitor matches Service labels, not pod selector
spec:
selector: {app: order-service}
ports: [{name: http, port: 8080, targetPort: 8080}] # named -- the Audit Metrics section's
# ServiceMonitor for the data-integrity counter
# matches by port NAME, not number
oc apply -f demo-workload.yaml oc rollout status deployment/inventory-service -n business-workloads --timeout=2m oc rollout status deployment/order-service -n business-workloads --timeout=2m
Before you ever inject a fault, confirm the happy path works:
oc exec -n business-workloads deploy/order-service -- curl -s http://localhost:8080/request
# Expect: {"downstream_status":200,"elapsed_seconds":0.0x}
order-service can't reach inventory-service at all — a DNS or Service issue — the fault-injection step later will look identical to a real cascading failure but for the wrong reason. Confirm the plumbing here first.Verify Tempo is receiving traces
Now that both services are running and generating OTEL spans, confirm Tempo received the trace from the request above:
# Query through the gateway (bearer token auth) -- NOT the query-frontend directly.
# With tenants.mode: openshift and gateway.enabled: true, the query-frontend enforces
# mutual TLS (mTLS) for internal service-to-service communication. Attempting direct
# access with -vk will succeed the TLS handshake but then fail with:
# "tlsv13 alert certificate required"
# because the server requires a client certificate that non-Tempo pods do not have.
# This is expected and correct -- the gateway is the intended external query path.
TOKEN=$(oc exec -n business-workloads deploy/order-service -- \
cat /var/run/secrets/kubernetes.io/serviceaccount/token)
oc exec -n business-workloads deploy/order-service -- curl -sk \
-H "Authorization: Bearer ${TOKEN}" \
-H "X-Scope-OrgID: dev" \
"https://tempo-sre-observability-gateway.openshift-tempo.svc.cluster.local:8080/api/traces/v1/dev/tempo/api/search?limit=5&tags=service.name%3Dorder-service"
traces array confirms Tempo is receiving spans. Compare the field names (traceID, durationMs) against what the SRE Intelligence service's tempo_trace_summary() expects — if they differ, update the parsing logic before relying on this in the pipeline.X.Runbook Knowledge Base
One ConfigMap. No database, no embedding model — see the architecture section for why that's a deliberate simplification, not a shortcut.
DataIntegrityCheckFailed key below, on purpose. That's the Tier 3 scenario in the Testing section — the LLM gets a real alert and a genuine "No runbook found for this alert" from rag_lookup(), exactly the way an actual novel incident would show up with nothing to ground it. If you add a runbook entry for it later, expect the scenario to stop landing in Tier 3 as reliably, since grounded confidence tends to run higher.apiVersion: v1
kind: ConfigMap
metadata:
name: business-workloads-runbooks
namespace: sre-intelligence
data:
runbooks.json: |
{
"PodCrashLoopingDetected": {
"title": "Pod CrashLoopBackOff — Config Regression vs Runtime Failure",
"content": "Symptoms: Pod restarts repeatedly with CrashLoopBackOff. Diagnosis: check logs for the exit reason — is the pod failing on startup (config regression) or during operation (runtime failure)? Key signals: (1) If the pod is newly deployed (crash-looping from first start) AND logs contain a hostname or service name that does not exist in the cluster — this is a config regression. action_type MUST be config_rollback. pod_restart will not help; the pod will crash again with the same error on every restart. (2) If the pod is newly deployed AND logs show connection refused to a known service that exists — escalate; the endpoint is correct but something else is blocking it. (3) If the pod has been running for hours and just started crashing — runtime failure. Use pod_restart. The distinction between case 1 and 2 is critical: only use config_rollback when the hostname itself does not exist."
},
"PodOOMKilled": {
"title": "OOMKilled Container — Check Tempo Trace for Slow Downstream Dependency",
"content": "Symptoms: Container terminated with OOMKilled. Before assuming a memory leak, check the Tempo trace for this service's outbound calls. A slow or hanging downstream dependency with no timeout causes in-flight requests to pile up in memory until the limit is hit — this is the most common cause of sudden OOMKills that have no gradual memory growth trend in Prometheus. If the trace shows a long-running downstream span, the root cause is the dependency, not a memory leak. Remediation (Tier 1): pod_restart clears the pile-up immediately. The trace-confirmed real fix (adding a client timeout or circuit breaker) is a human-scoped follow-up."
},
"SuspiciousEgressTraffic": {
"title": "Unexpected Egress Volume — Quarantine with NetworkPolicy (Tier 2, approval required)",
"content": "Symptoms: pod sending high volume of outbound traffic past threshold. This alert is Tier 2 — action_type MUST be network_policy_change. Do not use pod_restart; restarting does not isolate the pod and the same behaviour resumes immediately. The correct response is to apply a NetworkPolicy that denies all egress except DNS, keeping the pod running for investigation while cutting its network access. A human must approve this before execution because restricting egress can itself cause an outage if the traffic was legitimate. Diagnosis: check Loki logs for destination hostnames and request patterns, cross-reference Tempo traces to determine if the traffic is to known or unknown endpoints."
},
"HPAThrashing": {
"title": "HPA Thrashing — Replica Count Oscillation",
"content": "Symptoms: desired vs actual replica count oscillates. Diagnosis: oc describe hpa for current/desired/min/max, check if the CPU threshold is too close to actual usage, check scaleDown stabilizationWindowSeconds. Remediation: set stabilizationWindowSeconds to 300s, or raise targetCPUUtilizationPercentage by 10%, or raise minReplicas for known-bursty workloads."
},
"HighMemoryPressure": {
"title": "High Memory Pressure — Pre-OOM Intervention",
"content": "Symptoms: container above 90% of memory limit but not yet OOMKilled. Requires human approval (Tier 2). Diagnosis: check memory growth pattern over 30 min, distinguish leak (linear growth) from load spike (plateau). Options: wait and monitor if spike; rolling restart to buy time and file a bug if leak; raise limit temporarily if approaching cap."
},
"UpstreamServiceLatencyHigh": {
"title": "High P95 Latency — Upstream Service Investigation",
"content": "Symptoms: P95 latency above 2s, likely upstream dependency. Requires human escalation (Tier 3). Diagnosis: check which downstream service is slowest, check database query times, check inter-service network latency, check node resource pressure. Always escalate — latency issues often cascade across service owners."
}
}
oc apply -f runbooks-configmap.yaml
XI.OpenShift MCP Server
The SRE Intelligence service delegates all cluster mutations to openshift-mcp-server — a 216-tool MCP server that holds cluster credentials and exposes them over a standard MCP streamable-http transport. The SRE Intelligence service is an MCP client; it calls tools like rollout_restart_deployment and apply_manifest instead of making direct Kubernetes API calls itself. This separation keeps the SRE Intelligence service's RBAC footprint minimal and makes it easy to swap or extend the action layer independently of the pipeline logic.
The SRE Intelligence service opens a fresh MCP connection for each remediation action. Each call does a full TCP handshake and MCP initialize() — this is negligible overhead for an alert-driven system where actions are infrequent and reliability matters more than latency. A persistent session is not used: the MCP library's streamablehttp_client context manager uses anyio cancel scopes that are task-local and cannot be safely shared across asyncio BackgroundTask coroutines.
Build and push the MCP server image
git clone https://github.com/ay-garg/openshift-mcp-server.git
cd openshift-mcp-server
REGISTRY=$(oc get route default-route -n openshift-image-registry -o jsonpath='{.spec.host}')
podman build --platform linux/amd64 -t openshift-mcp-server:v1 .
podman login --tls-verify=false ${REGISTRY} -u $(oc whoami) -p $(oc whoami -t)
podman tag openshift-mcp-server:v1 ${REGISTRY}/mcp-server/openshift-mcp-server:v1
podman push --tls-verify=false ${REGISTRY}/mcp-server/openshift-mcp-server:v1
cd ..
Deploy it
mcp-server-deployment.yamlapiVersion: v1 kind: ServiceAccount metadata: name: mcp-server-sa namespace: mcp-server --- # A long-lived token Secret for the SA -- the MCP server reads OCP_TOKEN from this at startup. # The token is auto-populated and refreshed by the control plane. # Without OCP_API_URL + OCP_TOKEN, the server falls through to kubeconfig auth # and OCP_SKIP_TLS_VERIFY is never applied, causing x509 errors against the API server. apiVersion: v1 kind: Secret metadata: name: mcp-server-sa-token namespace: mcp-server annotations: kubernetes.io/service-account.name: mcp-server-sa type: kubernetes.io/service-account-token --- # Scoped ClusterRole granting only the verbs the MCP server actually uses. # cluster-admin is a security audit failure in production -- an automation bot # should never be able to read secrets in kube-system or delete arbitrary resources. apiVersion: rbac.authorization.k8s.io/v1 kind: ClusterRole metadata: name: mcp-server-role rules: - apiGroups: ["apps"] resources: ["deployments", "replicasets"] verbs: ["get", "list", "patch", "update"] - apiGroups: ["apps"] resources: ["deployments/status"] verbs: ["get"] - apiGroups: ["networking.k8s.io"] resources: ["networkpolicies"] verbs: ["get", "list", "create", "patch", "delete"] - apiGroups: ["autoscaling"] resources: ["horizontalpodautoscalers"] verbs: ["get", "list", "patch", "update"] - apiGroups: [""] resources: ["pods", "pods/log", "services", "namespaces", "events"] verbs: ["get", "list"] --- apiVersion: rbac.authorization.k8s.io/v1 kind: ClusterRoleBinding metadata: name: mcp-server-role-binding subjects: - kind: ServiceAccount name: mcp-server-sa namespace: mcp-server roleRef: kind: ClusterRole name: mcp-server-role apiGroup: rbac.authorization.k8s.io --- apiVersion: apps/v1 kind: Deployment metadata: name: openshift-mcp-server namespace: mcp-server spec: replicas: 1 selector: matchLabels: app: openshift-mcp-server template: metadata: labels: app: openshift-mcp-server spec: serviceAccountName: mcp-server-sa securityContext: runAsNonRoot: true seccompProfile: type: RuntimeDefault containers: - name: mcp-server image: image-registry.openshift-image-registry.svc:5000/mcp-server/openshift-mcp-server:v1 ports: - name: mcp-http containerPort: 8080 protocol: TCP env: - name: MCP_TRANSPORT value: "streamable-http" - name: MCP_HOST value: "0.0.0.0" - name: MCP_PORT value: "8080" - name: OCP_MODE value: "server" - name: OCP_API_URL value: "https://kubernetes.default.svc:443" # explicit URL activates path-2 auth in client.py; # OCP_SKIP_TLS_VERIFY is only read on this path - name: OCP_TOKEN valueFrom: secretKeyRef: name: mcp-server-sa-token key: token - name: OCP_SKIP_TLS_VERIFY value: "true" # DEMO ONLY: skips TLS verification for the Kubernetes API. # Production fix: remove this and instead mount the cluster CA bundle # at /var/run/secrets/kubernetes.io/serviceaccount/ca.crt # (already auto-mounted) and configure the MCP server to use it. resources: requests: {cpu: 100m, memory: 256Mi} limits: {cpu: "1", memory: 512Mi} securityContext: allowPrivilegeEscalation: false capabilities: drop: ["ALL"] readOnlyRootFilesystem: true runAsNonRoot: true startupProbe: # /health does not exist -- tcpSocket is the correct probe type tcpSocket: port: 8080 initialDelaySeconds: 5 periodSeconds: 5 failureThreshold: 12 # 60 s startup window livenessProbe: tcpSocket: port: 8080 initialDelaySeconds: 15 periodSeconds: 30 failureThreshold: 3 readinessProbe: tcpSocket: port: 8080 initialDelaySeconds: 10 periodSeconds: 10 failureThreshold: 3 volumeMounts: - name: tmp mountPath: /tmp - name: home mountPath: /home/ocp-mcp volumes: - name: tmp emptyDir: {} - name: home emptyDir: {} --- apiVersion: v1 kind: Service metadata: name: openshift-mcp-server namespace: mcp-server spec: selector: app: openshift-mcp-server ports: - name: mcp-http port: 8080 targetPort: 8080
oc apply -f mcp-server-deployment.yaml oc rollout status deployment/openshift-mcp-server -n mcp-server --timeout=3m oc logs deployment/openshift-mcp-server -n mcp-server | head -20
| Enrichment action_type | MCP tool called | What it does |
|---|---|---|
pod_restart | rollout_restart_deployment | Rolling restart of the pod's owning Deployment |
config_rollback | rollout_undo_deployment | Rolls back to the previous Deployment revision |
hpa_scale | apply_manifest | Patches HPA minReplicas and scaleDown stabilizationWindow |
network_policy_change | apply_manifest | Applies a NetworkPolicy restricting egress to DNS only (quarantine) |
XII.The SRE Intelligence Service
The orchestration brain described in the architecture section, in full. Everything downstream — metrics, PromQL, Slack — is already baked into this main.py from the start; there's no later "rebuild for observability" step.
fastapi>=0.110.0 uvicorn[standard]>=0.27.1 httpx>=0.27.0 # mcp requires >=0.27; 0.26.x causes ResolutionImpossible openai>=1.14.3 mcp>=1.5.0,<2.0.0 # 1.5+ has streamable-http; 2.x requires pydantic>=2.12 which conflicts pydantic>=2.6.3 python-multipart>=0.0.9 prometheus-client>=0.20.0 # audit/metrics endpoint -- included from the start, no second image build latersre-intelligence-service/main.py
import os, json, logging, time, hmac, hashlib, uuid
from datetime import datetime, timedelta
from typing import Optional
from fastapi import FastAPI, Request, BackgroundTasks
from fastapi.responses import JSONResponse
import httpx
from openai import AsyncOpenAI
from mcp import ClientSession
from mcp.client.streamable_http import streamablehttp_client
from prometheus_client import Counter, Histogram, make_asgi_app
from pydantic import BaseModel as _BaseModel
logging.basicConfig(level=logging.INFO,
format='%(asctime)s %(levelname)s %(message)s')
log = logging.getLogger(__name__)
app = FastAPI(title="SRE Intelligence Service", version="1.0.0")
# ── Pydantic models for strict webhook validation ────────────────────────
class _AlertLabel(_BaseModel):
alertname: str = "unknown"
severity: str = "warning"
namespace: str = ""
pod: str = ""
tier: str = "3"
class _Alert(_BaseModel):
status: str
labels: _AlertLabel = _AlertLabel()
annotations: dict = {}
startsAt: str = ""
class _WebhookPayload(_BaseModel):
alerts: list[_Alert] = []
# ── LLM circuit breaker state ────────────────────────────────────────────
_llm_failures = 0
_llm_circuit_open = False
_LLM_CIRCUIT_THRESHOLD = 3
_LLM_CIRCUIT_RESET_SECONDS = 300
async def _llm_circuit_reset():
"""Reset the LLM circuit breaker after the cool-down window.
Uses asyncio.create_task() + await asyncio.sleep() -- no get_event_loop(),
no globals().update() -- both are fragile in Python 3.10+ async contexts."""
global _llm_circuit_open, _llm_failures
await asyncio.sleep(_LLM_CIRCUIT_RESET_SECONDS)
_llm_circuit_open = False
_llm_failures = 0
log.info("LLM circuit breaker RESET after %ds cool-down", _LLM_CIRCUIT_RESET_SECONDS)
# ── Audit metrics ─────────────────────────────────────────────────────
alerts_processed = Counter('aiops_alerts_total', 'Alerts processed', ['tier', 'action'])
llm_latency = Histogram('aiops_llm_duration_seconds', 'LLM call duration')
actions_executed = Counter('aiops_actions_total', 'Actions executed', ['action_type', 'status'])
app.mount("/metrics", make_asgi_app())
# ── Config from environment variables ──────────────────────────────────
# _require() fails loudly at startup for missing critical config.
# Silent defaults hide misconfiguration for hours; a clear RuntimeError is better.
def _require(name: str) -> str:
val = os.getenv(name, "").strip()
if not val:
raise RuntimeError(
f"Required env var {name!r} is not set. "
f"Add it to sre-intelligence-config (non-sensitive) or sre-intelligence-secrets (credentials)."
)
return val
def _optional(name: str, default: str = "") -> str:
return os.getenv(name, default).strip() or default
SA_TOKEN_PATH = "/var/run/secrets/kubernetes.io/serviceaccount/token"
PROMETHEUS = _require("PROMETHEUS_URL")
LOKI_GATEWAY = _require("LOKI_GATEWAY_URL")
TEMPO_QUERY = _require("TEMPO_QUERY_URL")
LLM_URL = _require("LLM_BASE_URL")
LLM_MODEL = _require("LLM_MODEL")
SLACK_TOKEN = _require("SLACK_BOT_TOKEN")
MCP_URL = _require("MCP_SERVER_URL")
SLACK_CHAN = _optional("SLACK_APPROVAL_CHANNEL", "#sre-approvals")
SLACK_ESCALATE_CHAN = _optional("SLACK_ESCALATION_CHANNEL", "#sre-escalations")
RUNBOOKS_PATH = _optional("RUNBOOKS_PATH", "/etc/business-workloads/runbooks.json")
SLACK_SIGNING_SECRET = _optional("SLACK_SIGNING_SECRET", "") # from Slack app → Basic Information → App Credentials
CONF_T1 = float(_optional("CONFIDENCE_TIER1", "0.75"))
CONF_T2 = float(_optional("CONFIDENCE_TIER2", "0.70"))
LLM_TEMPERATURE = float(_optional("LLM_TEMPERATURE", "0.05"))
LLM_MAX_TOKENS = int(_optional("LLM_MAX_TOKENS", "1024")) # 1024 leaves headroom for reasoning + JSON within the 8192-token context window
llm = AsyncOpenAI(base_url=f"{LLM_URL}/v1", api_key="unused")
with open(RUNBOOKS_PATH) as f:
RUNBOOKS = json.load(f)
SYSTEM_PROMPT = """You are a senior Site Reliability Engineer specializing in OpenShift 4.
You will be given a firing alert, recent metrics, recent log lines (from Loki), and a
trace summary (from Tempo) describing the slowest recent span for the affected service.
Analyze the incident context and respond ONLY with valid JSON matching this schema exactly:
{
"root_cause": "one or two sentences describing the probable root cause",
"confidence": 0.87,
"affected_components": ["namespace/pod-name", "namespace/deployment-name"],
"recommended_action": "specific remediation description",
"action_type": "pod_restart",
"action_params": {"namespace": "...", "deployment": "..."},
"reasoning": "your step-by-step analysis"
}
RULE 1 — TIER-ACTION BINDING (non-negotiable):
The alert's tier label is authoritative. Your action_type MUST come from the same tier.
Alert tier=1 → action_type MUST be one of: pod_restart | config_rollback | hpa_scale
Alert tier=2 → action_type MUST be one of: network_policy_change | node_drain | pvc_resize
Alert tier=3 → action_type MUST be one of: database_failover | region_failover | escalate
Never pick a Tier 1 action for a Tier 2 alert, or vice versa. The tier label encodes
the SRE team's risk policy for that alert class — it overrides your own risk assessment.
RULE 2 — PodOOMKilled (tier=1):
If the Tempo trace shows a slow downstream call and the alert is OOMKilled or memory-related,
use action_type: pod_restart. The slow downstream is causing in-flight requests to pile up;
a restart clears the pile-up immediately. The real fix (adding a timeout on the caller) is
a human-scoped follow-up, not something to auto-remediate.
RULE 3 — PodCrashLoopingDetected (tier=1), apply in order:
Step A — Is the crashing pod a NEWLY deployed revision (crash-looping from its first start)?
Evidence: pod age in minutes, not hours; restart count climbs from zero immediately on creation.
If NO (long-running pod suddenly starts crashing) → use pod_restart; treat as runtime failure.
If YES → proceed to Step B.
Step B — What does the log evidence say about WHY it cannot connect?
Case 1: Logs contain a hostname or service name that does NOT exist in the cluster
(e.g., "nonexistent-service", "no such host") → use action_type: config_rollback.
Rationale: the deployment introduced a wrong endpoint. pod_restart will crash again
immediately. Only reverting to the previous revision fixes the misconfigured value.
Case 2: Logs contain "connection refused" to a hostname that IS a real in-cluster service
→ do NOT use config_rollback. The endpoint is correct but something is blocking it:
a NetworkPolicy, the downstream pod being down, or a recent outage. Use escalate.
Case 3: Logs show the downstream service returning errors or being unavailable
→ the crash is caused by a dependency failure, not the pod's own config.
Use pod_restart as short-term mitigation; note the real fix is restoring the dependency.
Case 4: Log evidence is ambiguous or missing
→ use escalate. A rollback that reverts working code is worse than paging a human.
RULE 4 — No text outside the JSON object."""
# ── Slack signature verification ─────────────────────────────────────────
def _verify_slack_signature(body: bytes, timestamp: str, signature: str) -> bool:
if not timestamp or abs(time.time() - float(timestamp)) > 300:
return False # replay protection: reject requests older than 5 minutes
basestring = f"v0:{timestamp}:{body.decode()}".encode()
expected = "v0=" + hmac.new(
SLACK_SIGNING_SECRET.encode(), basestring, hashlib.sha256
).hexdigest()
return hmac.compare_digest(expected, signature)
# ── Health — checks LLM reachability in addition to pod liveness ────────
@app.get("/health")
async def health():
checks: dict = {"llm": "ok"}
try:
async with httpx.AsyncClient(timeout=3, verify="/var/run/secrets/kubernetes.io/serviceaccount/service-ca.crt") as c: # cluster service CA auto-mounted; verifies internal TLS certs
r = await c.get(f"{LLM_URL}/v1/models")
if r.status_code != 200:
checks["llm"] = f"degraded ({r.status_code})"
except Exception as e:
checks["llm"] = f"fail: {type(e).__name__}"
overall = "ok" if all(v == "ok" for v in checks.values()) else "degraded"
return {"status": overall, "ts": datetime.utcnow().isoformat(), "checks": checks}
# ── AlertManager webhook ────────────────────────────────────────────────
@app.post("/webhook")
async def receive_alert(request: Request, bg: BackgroundTasks):
try:
payload = _WebhookPayload(**(await request.json()))
except Exception as e:
log.warning("Invalid webhook payload: %s", e)
return JSONResponse({"error": "invalid payload"}, status_code=400)
firing = [a for a in payload.alerts if a.status == "firing"]
if not firing:
return JSONResponse({"skipped": "no firing alerts"})
accepted = 0
for alert in firing:
labels = alert.labels
log.info("━━━ Webhook received: %s | severity=%s tier=%s ns=%s pod=%s ━━━",
labels.alertname, labels.severity,
labels.tier, labels.namespace, labels.pod)
bg.add_task(process_alert, alert.model_dump())
accepted += 1
return JSONResponse({"accepted": accepted})
# ── Pipeline ─────────────────────────────────────────────────────────────
async def process_alert(alert: dict):
name = alert["labels"].get("alertname", "unknown")
ns = alert["labels"].get("namespace", "")
pod = alert["labels"].get("pod", "")
corr_id = str(uuid.uuid4())[:8] # short ID to correlate all log lines for this incident
log.info(f"[{corr_id}] Processing: {name} | ns={ns} | pod={pod}")
try:
ctx = await enrich(alert)
if not ctx.get("signals_available", True):
log.warning(f"[{corr_id}] All signal sources unavailable — forcing tier=3, skipping LLM")
await slack_escalate(alert, error=(
"All observability sources (Prometheus, Loki, Tempo) were unreachable. "
"Auto-remediation skipped — incomplete context must not drive automated actions."))
await audit(alert, {"root_cause": "Signal sources unavailable", "confidence": 0.0,
"action_type": "escalate", "tier": 3}, corr_id)
return
rb = await rag_lookup(name)
rca = await llm_analyze(alert, ctx, rb)
await execute(alert, rca)
await audit(alert, rca, corr_id)
except Exception as e:
log.error(f"[{corr_id}] Pipeline error for {name}: {e}")
await slack_escalate(alert, error=str(e))
# ── Enrich: metrics (Prometheus) + logs (Loki) + trace (Tempo) ────────────
async def enrich(alert: dict) -> dict:
ns, pod = alert["labels"].get("namespace",""), alert["labels"].get("pod","")
metrics, logs, trace_summary = {}, [], "no trace data"
with open(SA_TOKEN_PATH) as f:
sa_token = f.read().strip()
prom_headers = {"Authorization": f"Bearer {sa_token}"}
async with httpx.AsyncClient(timeout=10, verify="/var/run/secrets/kubernetes.io/serviceaccount/service-ca.crt") as c: # cluster service CA auto-mounted; verifies Thanos/Loki/Tempo gateway certs
if pod and ns:
for query, key in [
(f'container_memory_working_set_bytes{{pod="{pod}",namespace="{ns}"}}', "mem"),
(f'rate(container_cpu_usage_seconds_total{{pod="{pod}",namespace="{ns}"}}[5m])', "cpu"),
(f'kube_pod_container_status_restarts_total{{pod="{pod}",namespace="{ns}"}}', "restarts"),
]:
try:
r = await c.get(f"{PROMETHEUS}/api/v1/query",
params={"query": query}, headers=prom_headers)
res = r.json()["data"]["result"]
metrics[key] = res[0]["value"][1] if res else "n/a"
except Exception:
metrics[key] = "n/a"
logs = await loki_logs(ns, pod)
log.info("LOKI | fetched %d log lines for %s/%s", len(logs), ns, pod)
parts = pod.split("-")
deploy = "-".join(parts[:-2]) if len(parts) > 2 else pod
trace_summary = await tempo_trace_summary(deploy)
log.info("TEMPO | %s", trace_summary)
log.info("PROM | mem=%s cpu=%s restarts=%s",
metrics.get("mem", "n/a"), metrics.get("cpu", "n/a"),
metrics.get("restarts", "n/a"))
# Guard: if ALL three signal sources are empty, auto-remediation on a blank slate
# is dangerous. Flag this so process_alert() can force escalation instead.
signals_available = (
any(v != "n/a" for v in metrics.values()) or
bool(logs) or
trace_summary not in ("no trace data", "trace lookup failed", "no matching traces in the alert window")
)
return {"metrics": metrics, "logs": logs, "trace": trace_summary, "signals_available": signals_available}
# ── Loki: LogQL query via the LokiStack gateway, tenant-scoped, SA-token auth ──
async def loki_logs(namespace: str, pod: str, minutes: int = 15) -> list:
with open(SA_TOKEN_PATH) as f:
token = f.read().strip()
end = datetime.utcnow()
start = end - timedelta(minutes=minutes)
query = f'{{kubernetes_namespace_name="{namespace}",kubernetes_pod_name="{pod}"}} |~ "(?i)error|timeout|warn|refused|reset|slow"'
params = {
"query": query,
"start": str(int(start.timestamp() * 1e9)),
"end": str(int(end.timestamp() * 1e9)),
"limit": "30",
"direction": "backward",
}
try:
async with httpx.AsyncClient(verify="/var/run/secrets/kubernetes.io/serviceaccount/service-ca.crt", timeout=10) as c: # cluster service CA auto-mounted
r = await c.get(f"{LOKI_GATEWAY}/api/logs/v1/application/loki/api/v1/query_range",
params=params, headers={"Authorization": f"Bearer {token}"})
r.raise_for_status()
streams = r.json()["data"]["result"]
lines = []
for stream in streams:
for entry in stream["values"]:
raw = entry[1]
try:
parsed = json.loads(raw)
msg = parsed.get("message", parsed.get("log", ""))
level = parsed.get("level", "")
ts = (parsed.get("@timestamp") or "")[:19] # trim to seconds
container = (parsed.get("kubernetes") or {}).get("container_name", "")
parts = [p for p in [ts, level.upper(), container, msg] if p]
except Exception:
# CRI format: "2024-01-01T00:00:00Z stdout F " or plain text
if " stdout " in raw or " stderr " in raw:
p = raw.split(" ", 3)
parts = [p[3]] if len(p) > 3 else [raw]
else:
parts = [raw]
line = " ".join(parts)
# prompt injection guard: strip override attempts, cap line length
line = line.replace("Ignore previous instructions", "[FILTERED]")
line = line.replace("system prompt", "[FILTERED]")
lines.append(line[:300])
return lines[:20]
except Exception as e:
log.warning(f"Loki query failed: {e}")
return []
# ── Tempo: find the slowest recent trace for a service, summarize its longest span ──
async def tempo_trace_summary(service_name: str, minutes: int = 15) -> str:
if not service_name:
return "no trace data"
end = int(datetime.utcnow().timestamp())
start = end - minutes * 60
with open(SA_TOKEN_PATH) as f:
tempo_token = f.read().strip()
# Gateway (Observatorium) requires bearer token auth.
# Tenant is encoded in TEMPO_QUERY base URL path (/api/traces/v1/dev/tempo),
# so no X-Scope-OrgID header is needed.
tempo_headers = {
"Authorization": f"Bearer {tempo_token}",
}
try:
async with httpx.AsyncClient(timeout=10, verify="/var/run/secrets/kubernetes.io/serviceaccount/service-ca.crt") as c:
r = await c.get(f"{TEMPO_QUERY}/api/search", params={
"q": f'{{resource.service.name="{service_name}"}}',
"start": start, "end": end, "limit": 5,
}, headers=tempo_headers)
r.raise_for_status()
traces = r.json().get("traces", [])
if not traces:
return "no matching traces in the alert window"
traces.sort(key=lambda t: int(t.get("durationMs", 0)), reverse=True)
slowest = traces[0]
trace_id = slowest["traceID"]
r2 = await c.get(f"{TEMPO_QUERY}/api/traces/{trace_id}", headers=tempo_headers)
r2.raise_for_status()
spans = [s for b in r2.json().get("batches", [])
for s in b.get("scopeSpans", [{}])[0].get("spans", [])]
longest = max(spans, default=None, key=lambda s:
int(s.get("endTimeUnixNano", 0)) - int(s.get("startTimeUnixNano", 0)))
if not longest:
return f"trace {trace_id}: {slowest.get('durationMs')}ms total, no span detail"
dur_ms = (int(longest["endTimeUnixNano"]) - int(longest["startTimeUnixNano"])) / 1e6
return f"slowest span '{longest.get('name')}' took {dur_ms:.0f}ms (trace {trace_id}, total {slowest.get('durationMs')}ms)"
except Exception as e:
log.warning(f"Tempo query failed: {e}")
return "trace lookup failed"
# ── Runbook lookup -- in-memory dict loaded from the ConfigMap ────────────
async def rag_lookup(alert_name: str) -> str:
rb = RUNBOOKS.get(alert_name)
if rb:
log.info("RUNBOOK | matched: %s", rb["title"])
return f"## {rb['title']}\n{rb['content']}"
log.info("RUNBOOK | no match for %s", alert_name)
return "No runbook found for this alert."
# ── LLM call with retry -- open-source instruct models can produce malformed JSON ──
import re
async def _llm_call_with_retry(prompt: str, retries: int = 2) -> dict:
messages = [
{"role": "system", "content": SYSTEM_PROMPT},
{"role": "user", "content": prompt},
]
last_err = None
for attempt in range(1, retries + 1):
resp = await llm.chat.completions.create(
model=LLM_MODEL,
# Qwen2.5 properly supports a dedicated system role -- no need to fold
# the system prompt into the user turn. The system message sets the
# persona and JSON schema; the user message carries the alert context.
messages=messages,
temperature=LLM_TEMPERATURE,
max_tokens=LLM_MAX_TOKENS,
)
raw = resp.choices[0].message.content or ""
raw = raw.strip()
finish = resp.choices[0].finish_reason
if not raw:
log.warning(f"LLM returned empty response (attempt {attempt}/{retries}, finish_reason={finish})")
else:
log.info(f"LLM raw response ({len(raw)} chars, finish_reason={finish}): {raw[:200]}")
if "```" in raw:
raw = raw.split("```")[1].lstrip("json").strip()
# Extract the first top-level JSON object if the model added trailing text.
# If braces are unbalanced (truncated JSON), close them and try anyway.
m = re.search(r'\{', raw)
if m:
depth, end = 0, m.start()
for i, ch in enumerate(raw[m.start():], m.start()):
if ch == '{': depth += 1
elif ch == '}': depth -= 1
if depth == 0:
end = i + 1
break
extracted = raw[m.start():end]
if not extracted or extracted == '{':
# Unbalanced braces -- take everything from first { and close the open braces
extracted = raw[m.start():] + ('}' * depth)
log.warning(f"Unbalanced JSON braces (depth={depth}), attempting repair")
raw = extracted
try:
return json.loads(raw)
except json.JSONDecodeError as e:
last_err = e
log.warning(f"LLM JSON parse failed (attempt {attempt}/{retries}): {e}")
raise last_err
# ── LLM root cause analysis ──────────────────────────────────────────────
async def llm_analyze(alert: dict, ctx: dict, runbooks: str) -> dict:
global _llm_failures, _llm_circuit_open
if _llm_circuit_open:
log.warning("LLM circuit breaker OPEN — escalating without analysis")
return {"tier": 3, "action_type": "escalate", "confidence": 0.0,
"root_cause": "LLM unavailable (circuit breaker open)", "reasoning": ""}
start = time.time()
labels = alert["labels"]
prompt = f"""ALERT:
name={labels.get('alertname')} severity={labels.get('severity')} \
namespace={labels.get('namespace','n/a')} pod={labels.get('pod','n/a')}
description: {alert.get('annotations',{}).get('description','n/a')}
started: {alert.get('startsAt','n/a')}
METRICS:
memory_bytes={ctx['metrics'].get('mem')} cpu_rate={ctx['metrics'].get('cpu')} \
restarts={ctx['metrics'].get('restarts')}
RECENT LOG LINES (Loki, 15 min, error/timeout/warn/slow filter):
{chr(10).join(ctx['logs'][-15:]) if ctx['logs'] else 'none'}
TRACE SUMMARY (Tempo, slowest recent span):
{ctx['trace']}
MATCHING RUNBOOKS:
{runbooks}"""
log.info("─── LLM prompt context for %s ───", labels.get("alertname"))
for line in (ctx["logs"] or [])[-5:]:
log.info(" LOG | %s", line.strip()[:200])
log.info("─── Calling LLM ───")
try:
rca = await _llm_call_with_retry(prompt, retries=2)
_llm_failures = 0
except Exception as e:
_llm_failures += 1
if _llm_failures >= _LLM_CIRCUIT_THRESHOLD:
_llm_circuit_open = True
asyncio.create_task(_llm_circuit_reset()) # proper asyncio pattern -- no get_event_loop(), no globals()
log.error("LLM circuit breaker OPENED after %d failures", _llm_failures)
raise
t1 = {"pod_restart", "config_rollback", "hpa_scale"}
t2 = {"node_drain", "pvc_resize", "network_policy_change"}
all_valid = t1 | t2 | {"database_failover", "region_failover", "escalate"}
conf = float(rca.get("confidence", 0))
atype = rca.get("action_type", "escalate")
# Normalize: smaller instruct models sometimes return the tier label ("Tier 1")
# instead of the enum value. Fall back to the alert's own tier label.
if atype not in all_valid:
tier_hint = alert.get("labels", {}).get("tier", "3")
atype = {"1": "pod_restart", "2": "network_policy_change"}.get(tier_hint, "escalate")
rca["action_type"] = atype
log.warning(f"Normalized invalid action_type to {atype} (tier hint={tier_hint})")
# The alert's tier label is authoritative — the LLM enriches with root
# cause and action but cannot downgrade the tier. Low confidence on a
# Tier 1 alert bumps to Tier 2 (require approval) as a safety net.
alert_tier = int(alert.get("labels", {}).get("tier", "3"))
if alert_tier == 1 and conf < CONF_T1:
rca["tier"] = 2
else:
rca["tier"] = alert_tier
log.info(f"RCA: tier={rca['tier']} conf={conf:.2f} action={atype}")
llm_latency.observe(time.time() - start)
alerts_processed.labels(tier=str(rca["tier"]), action=atype).inc()
return rca
# ── Execute by tier ──────────────────────────────────────────────────────
async def execute(alert: dict, rca: dict):
tier = rca.get("tier", 3)
action = rca.get("action_type", "escalate")
conf = rca.get("confidence", 0)
ns = alert["labels"].get("namespace", "")
pod = alert["labels"].get("pod", "")
if tier == 1:
log.info("ACTION | Tier 1 AUTO-EXECUTE: %s on %s/%s (conf=%.2f)", action, ns, pod, conf)
await run_action(action, {
"namespace": ns,
"pod": pod,
**rca.get("action_params", {}),
})
elif tier == 2:
log.info("ACTION | Tier 2 APPROVAL REQUIRED: %s on %s/%s (conf=%.2f) → Slack", action, ns, pod, conf)
await slack_approval(alert, rca)
else:
log.info("ACTION | Tier 3 ESCALATION: %s on %s/%s (conf=%.2f) → Slack", action, ns, pod, conf)
await slack_escalate(alert, rca=rca)
# ── MCP client -- delegates all cluster mutations to the openshift-mcp-server ──
# Per-call connections: streamablehttp_client uses anyio cancel scopes that are
# task-local and cannot be shared across asyncio tasks. A persistent session
# entered in the startup task cannot be used inside BackgroundTask coroutines.
async def call_mcp_tool(tool_name: str, tool_args: dict) -> str:
async with streamablehttp_client(MCP_URL) as (read, write, _):
async with ClientSession(read, write) as session:
await session.initialize()
result = await session.call_tool(tool_name, tool_args)
msg = str(result.content[0].text) if result.content else "done"
if result.isError:
raise RuntimeError(f"MCP tool {tool_name} failed: {msg}")
return msg
def get_deployment_name(pod: str) -> str:
# Standard Kubernetes pod naming: <deployment>-<rs-hash>-<pod-hash>.
# Strip the last two hash segments to get the deployment name.
# Works correctly for all demo workload names (order-service, inventory-service).
parts = pod.split("-")
return "-".join(parts[:-2]) if len(parts) > 2 else pod
async def run_action(action: str, params: dict):
ns = params.get("namespace", "business-workloads")
pod = params.get("pod", "")
deploy = params.get("deployment") or get_deployment_name(pod)
if action == "pod_restart":
result = await call_mcp_tool("rollout_restart_deployment",
{"namespace": ns, "name": deploy}) # MCP tool uses "name", not "deployment_name"
log.info(f"MCP rollout_restart_deployment {deploy} in {ns}: {result}")
elif action == "config_rollback":
result = await call_mcp_tool("rollout_undo_deployment",
{"namespace": ns, "name": params.get("deployment", deploy)})
log.info(f"MCP rollout_undo_deployment {deploy} in {ns}: {result}")
elif action == "hpa_scale":
hpa_name = params.get("hpa_name", deploy)
min_replicas = int(params.get("min_replicas", 3))
manifest = f"""apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: {hpa_name}
namespace: {ns}
spec:
minReplicas: {min_replicas}
behavior:
scaleDown:
stabilizationWindowSeconds: 300
policies:
- type: Pods
value: 1
periodSeconds: 60
"""
result = await call_mcp_tool("apply_manifest", {"yaml_content": manifest})
log.info(f"MCP apply_manifest HPA {hpa_name} in {ns}: {result}")
elif action == "network_policy_change":
# Tier 2 -- only ever called after a human clicks Approve in Slack.
app_label = deploy # reuse the owner-ref-resolved deploy name
manifest = f"""apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: quarantine-{app_label}
namespace: {ns}
labels:
created-by: sre-intelligencement
spec:
podSelector:
matchLabels:
app: {app_label}
policyTypes: [Egress]
egress:
- ports:
- protocol: UDP
port: 53
"""
result = await call_mcp_tool("apply_manifest", {"yaml_content": manifest})
log.info(f"MCP apply_manifest NetworkPolicy quarantine-{app_label} in {ns}: {result}")
else:
# node_drain, pvc_resize, database_failover, region_failover, escalate:
# valid LLM-facing enum values with no MCP handler here.
# process_alert()'s except block catches this and calls slack_escalate().
# Safe failure: unknown action → human page, not a silent no-op.
raise ValueError(f"No MCP handler for action: {action}")
# ── Slack Tier 2 approval ─────────────────────────────────────────────────
async def slack_approval(alert: dict, rca: dict):
payload = {"action": rca["action_type"],
"alert": alert["labels"], "rca": rca}
blocks = [
{"type":"header","text":{"type":"plain_text",
"text":f"⚠️ AI Ops — Approval Required: {rca['action_type']}"}},
{"type":"section","fields":[
{"type":"mrkdwn","text":f"*Alert:*\n{alert['labels'].get('alertname')}"},
{"type":"mrkdwn","text":f"*Confidence:*\n{rca['confidence']:.0%}"},
{"type":"mrkdwn","text":f"*Namespace:*\n{alert['labels'].get('namespace','n/a')}"},
{"type":"mrkdwn","text":f"*Pod:*\n{alert['labels'].get('pod','n/a')}"},
]},
{"type":"section","text":{"type":"mrkdwn",
"text":f"*Root Cause:*\n{rca['root_cause']}"}},
{"type":"section","text":{"type":"mrkdwn",
"text":f"*Proposed Action:*\n{rca['recommended_action']}"}},
{"type":"actions","elements":[
{"type":"button","style":"primary",
"text":{"type":"plain_text","text":"✅ Approve"},
"action_id":"approve","value":json.dumps(payload)},
{"type":"button","style":"danger",
"text":{"type":"plain_text","text":"❌ Reject"},
"action_id":"reject","value":json.dumps(payload)},
]},
]
async with httpx.AsyncClient() as c:
await c.post("https://slack.com/api/chat.postMessage",
headers={"Authorization": f"Bearer {SLACK_TOKEN}"},
json={"channel": SLACK_CHAN, "blocks": blocks,
"text": f"AI Ops approval needed: {rca['action_type']}"})
# ── Slack interactive actions endpoint ───────────────────────────────────
@app.post("/slack/actions")
async def slack_action(request: Request, bg: BackgroundTasks):
raw_body = await request.body()
ts = request.headers.get("X-Slack-Request-Timestamp", "")
sig = request.headers.get("X-Slack-Signature", "")
if SLACK_SIGNING_SECRET and not _verify_slack_signature(raw_body, ts, sig):
log.warning("Rejected Slack request with invalid signature")
return JSONResponse({"error": "invalid signature"}, status_code=401)
form = await request.form()
body = json.loads(form.get("payload", "{}"))
act = body.get("actions", [{}])[0]
aid = act.get("action_id")
val = json.loads(act.get("value", "{}"))
user = body.get("user", {}).get("name", "unknown")
chan = body.get("channel", {}).get("id", SLACK_CHAN)
msg_ts = body.get("message", {}).get("ts")
rca = val.get("rca", {})
labels = val.get("alert", {})
alert_name = labels.get("alertname", "unknown")
action = rca.get("action_type", "unknown")
async def _update_and_notify(status_emoji, status_text, detail):
async with httpx.AsyncClient() as c:
if msg_ts:
await c.post("https://slack.com/api/chat.update",
headers={"Authorization": f"Bearer {SLACK_TOKEN}"},
json={"channel": chan, "ts": msg_ts,
"blocks": [
{"type":"header","text":{"type":"plain_text",
"text":f"{status_emoji} {alert_name} — {status_text}"}},
{"type":"section","fields":[
{"type":"mrkdwn","text":f"*Action:*\n{action}"},
{"type":"mrkdwn","text":f"*{status_text} by:*\n@{user}"},
{"type":"mrkdwn","text":f"*Namespace:*\n{labels.get('namespace','n/a')}"},
{"type":"mrkdwn","text":f"*Pod:*\n{labels.get('pod','n/a')}"},
]},
{"type":"section","text":{"type":"mrkdwn","text":detail}},
],
"text": f"{alert_name} {status_text.lower()} by {user}"})
if aid == "approve":
log.info("ACTION | %s APPROVED by %s — executing %s", alert_name, user, action)
bg.add_task(run_action, action,
{"namespace": labels.get("namespace"), "pod": labels.get("pod"),
**rca.get("action_params", {})})
await _update_and_notify("✅", "Approved",
f"*Root cause:* {rca.get('root_cause','n/a')}\n"
f"*Confidence:* {rca.get('confidence',0):.0%}\n"
f"Executing `{action}` now.")
return JSONResponse({"text": f"✅ Approved by @{user}. Executing {action}…"})
else:
log.info("ACTION | %s REJECTED by %s — incident remains open", alert_name, user)
await _update_and_notify("❌", "Rejected",
f"*Root cause:* {rca.get('root_cause','n/a')}\n"
f"*Confidence:* {rca.get('confidence',0):.0%}\n"
"Action was rejected. Incident remains open for manual investigation.")
return JSONResponse({"text": f"❌ Rejected by @{user}. Incident remains open."})
# ── Tier 3 escalation -- Slack, not PagerDuty (see the Slack section) ────
async def slack_escalate(alert: dict, rca: Optional[dict]=None, error: Optional[str]=None):
name = alert["labels"].get("alertname", "alert")
fields = [
{"type":"mrkdwn","text":f"*Alert:*\n{name}"},
{"type":"mrkdwn","text":f"*Severity:*\n{alert['labels'].get('severity','critical')}"},
{"type":"mrkdwn","text":f"*Namespace:*\n{alert['labels'].get('namespace','n/a')}"},
{"type":"mrkdwn","text":f"*Pod:*\n{alert['labels'].get('pod','n/a')}"},
]
blocks = [
{"type":"header","text":{"type":"plain_text",
"text":f"🚨 AI Ops — Tier 3 Escalation: {name}"}},
{"type":"section","fields":fields},
]
if rca:
blocks += [
{"type":"section","text":{"type":"mrkdwn",
"text":f"*Root Cause:*\n{rca.get('root_cause','n/a')}"}},
{"type":"section","text":{"type":"mrkdwn",
"text":f"*Confidence:*\n{rca.get('confidence',0):.0%}"}},
{"type":"section","text":{"type":"mrkdwn",
"text":f"*Reasoning:*\n{rca.get('reasoning','n/a')}"}},
{"type":"section","text":{"type":"mrkdwn",
"text":f"*Recommended Action:*\n{rca.get('recommended_action','n/a')}"}},
]
if error:
blocks.append({"type":"section","text":{"type":"mrkdwn",
"text":f"*Processing Error:*\n```{error}```"}})
async with httpx.AsyncClient() as c:
await c.post("https://slack.com/api/chat.postMessage",
headers={"Authorization": f"Bearer {SLACK_TOKEN}"},
json={"channel": SLACK_ESCALATE_CHAN, "blocks": blocks,
"text": f"AI Ops Tier 3 escalation: {name}"})
# ── Audit log ────────────────────────────────────────────────────────────
async def audit(alert: dict, rca: dict, corr_id: str = ""):
log.info(json.dumps({
"event": "AUDIT",
"ts": datetime.utcnow().isoformat(),
"corr_id": corr_id,
"alert": alert["labels"].get("alertname"),
"namespace": alert["labels"].get("namespace"),
"pod": alert["labels"].get("pod"),
"root_cause": rca.get("root_cause"),
"confidence": rca.get("confidence"),
"action": rca.get("action_type"),
"tier": rca.get("tier"),
"auto": rca.get("tier") == 1,
}))
run_action() is written for specific alert names. The pipeline handles any alert AlertManager delivers to the webhook. What is fixed is the set of remediation actions the system is allowed to take: pod_restart, config_rollback, hpa_scale, and network_policy_change. These are intentional safety boundaries — the LLM can reason freely over any alert, but it can only trigger actions from this defined list. For a new alert type, no code change is needed; adding a matching runbook entry in the ConfigMap is enough to give the LLM the grounding it needs to pick the right action.
The tier label on the alert (set in the PrometheusRule) controls what happens next: Tier 1 → run_action() fires automatically; Tier 2 → a Slack message goes to #sre-approvals with Approve/Reject buttons, and run_action() is called only if a human clicks Approve; Tier 3 → a context pack goes to #sre-escalations and run_action() is never called.
If the LLM returns an action_type the code does not recognise (for example node_drain or database_failover, which are valid LLM-facing enum values but have no MCP handler yet), run_action() raises a ValueError. process_alert() catches it and calls slack_escalate(), so a Slack message lands in #sre-escalations with the full context pack — the same outcome as a Tier 3 escalation. The incident is never silently dropped.| Tier | Routing in execute() | When run_action() executes |
|---|---|---|
| 1 — Auto | run_action() called directly | Immediately, if LLM confidence ≥ CONF_T1 (0.75) |
| 2 — Approve | slack_approval() → #sre-approvals | Only when a human clicks Approve |
| 3 — Escalate | slack_escalate() → #sre-escalations | Never — no automated action |
FROM registry.access.redhat.com/ubi9/python-311:latest WORKDIR /app COPY requirements.txt . RUN pip install --no-cache-dir -r requirements.txt COPY main.py . EXPOSE 8080 CMD ["uvicorn", "main:app", "--host", "0.0.0.0", "--port", "8080", "--workers", "2"]
REGISTRY=$(oc get route default-route -n openshift-image-registry -o jsonpath='{.spec.host}')
podman build --platform linux/amd64 -t sre-intelligence-service:v1 .
podman login --tls-verify=false ${REGISTRY} -u $(oc whoami) -p $(oc whoami -t)
podman tag sre-intelligence-service:v1 ${REGISTRY}/sre-intelligence/sre-intelligence-service:v1
podman push --tls-verify=false ${REGISTRY}/sre-intelligence/sre-intelligence-service:v1
Deploy it
All non-sensitive configuration lives in a single ConfigMap so it can be edited and re-applied without touching credentials or rebuilding the image. The Deployment's envFrom loads both the ConfigMap and the Secret, giving the container the full set of env vars it needs.
apiVersion: v1 kind: ConfigMap metadata: name: sre-intelligence-config namespace: sre-intelligence data: PROMETHEUS_URL: "https://thanos-querier.openshift-monitoring.svc.cluster.local:9091" # platform Prometheus requires auth; Thanos Querier accepts the SA bearer token LOKI_GATEWAY_URL: "https://sre-observability-gateway-http.openshift-logging.svc.cluster.local:8080" TEMPO_QUERY_URL: "https://tempo-sre-observability-gateway.openshift-tempo.svc.cluster.local:8080/api/traces/v1/dev/tempo" # Observatorium gateway on port 8080; wraps Tempo paths under /api/traces/v1/{tenant}/tempo/ # The function appends /api/search and /api/traces/{id} to this base URL MCP_SERVER_URL: "http://openshift-mcp-server.mcp-server.svc.cluster.local:8080/mcp" RUNBOOKS_PATH: "/etc/business-workloads/runbooks.json" SLACK_APPROVAL_CHANNEL: "#sre-approvals" SLACK_ESCALATION_CHANNEL: "#sre-escalations" CONFIDENCE_TIER1: "0.75" CONFIDENCE_TIER2: "0.70" LLM_MODEL: "qwen2-5-7b-instruct" # must match --served-model-name in the InferenceService args LLM_TEMPERATURE: "0.05" # low = deterministic JSON; raise if responses feel repetitive LLM_MAX_TOKENS: "512" # JSON response fits in 512 tokens; 1500 risks hitting context limit when Loki logs are verbose
oc apply -f sre-intelligence-config.yaml
oc apply -f sre-intelligence-config.yaml, then roll the deployment — envFrom is snapshotted at pod start, not watched live: oc rollout restart deployment/sre-intelligence-service -n sre-intelligence.# No ClusterRole or ClusterRoleBinding needed here. The SRE Intelligence service makes # no direct Kubernetes API calls -- all cluster mutations go through the MCP server, # and the Loki bearer token is the SA's own auto-mounted credential (file read, # not a k8s API call). The only RBAC the SA needs is the Loki reader grant # applied in the Observability section (cluster-logging-application-view binding). apiVersion: v1 kind: ServiceAccount metadata: name: sre-intelligence-sa namespace: sre-intelligence --- apiVersion: v1 kind: Secret metadata: name: sre-intelligence-secrets namespace: sre-intelligence stringData: # Credentials only -- all non-sensitive config lives in sre-intelligence-config ConfigMap above. # Paste the exact $LLM_ENDPOINT value from .status.address.url -- do not hand-type it. LLM_BASE_URL: "http://qwen2-5-7b-instruct-predictor.llm-serving.svc.cluster.local:8080" SLACK_BOT_TOKEN: "xoxb-your-slack-bot-token" SLACK_SIGNING_SECRET: "<from Slack app → Basic Information → App Credentials → Signing Secret>" --- apiVersion: apps/v1 kind: Deployment metadata: name: sre-intelligence-service namespace: sre-intelligence spec: replicas: 1 selector: matchLabels: app: sre-intelligence-service template: metadata: labels: app: sre-intelligence-service spec: serviceAccountName: sre-intelligence-sa containers: - name: enrichment image: image-registry.openshift-image-registry.svc:5000/sre-intelligence/sre-intelligence-service:v1 ports: - containerPort: 8080 envFrom: - configMapRef: name: sre-intelligence-config # all non-sensitive config: URLs, thresholds, LLM params, Slack channels - secretRef: name: sre-intelligence-secrets # credentials only: LLM_BASE_URL, SLACK_BOT_TOKEN volumeMounts: - name: runbooks mountPath: /etc/business-workloads livenessProbe: httpGet: {path: /health, port: 8080} initialDelaySeconds: 15 periodSeconds: 20 readinessProbe: httpGet: {path: /health, port: 8080} initialDelaySeconds: 10 periodSeconds: 10 resources: requests: {cpu: 250m, memory: 256Mi} limits: {cpu: "1", memory: 512Mi} volumes: - name: runbooks configMap: name: business-workloads-runbooks --- apiVersion: v1 kind: Service metadata: name: sre-intelligence-service namespace: sre-intelligence labels: app: sre-intelligence-service # ServiceMonitor matches the Service's own labels, not its pod selector spec: selector: app: sre-intelligence-service ports: - name: http # named on purpose -- the ServiceMonitor later matches by port NAME, not number port: 8080 targetPort: 8080
oc apply -f sre-intelligence-config.yaml oc apply -f sre-intelligence-deployment.yaml oc rollout status deployment/sre-intelligence-service -n sre-intelligence --timeout=3m oc exec -n sre-intelligence deploy/sre-intelligence-service -- curl -s http://localhost:8080/health # Verify TLS connectivity to each internal endpoint using the mounted service CA. # The service-serving CA is mounted at /var/run/secrets/kubernetes.io/serviceaccount/service-ca.crt (NOT the # system bundle at ca-bundle.crt — the OpenShift service CA is not in that bundle). # Python httpx uses verify= to point to this file explicitly; curl must do the same. # Wrap in sh -c so $(cat ...) is evaluated INSIDE the pod, not on your local shell oc exec -n sre-intelligence deploy/sre-intelligence-service -- \ sh -c 'curl -s --cacert /var/run/secrets/kubernetes.io/serviceaccount/service-ca.crt \ -H "Authorization: Bearer $(cat /var/run/secrets/kubernetes.io/serviceaccount/token)" \ "https://thanos-querier.openshift-monitoring.svc.cluster.local:9091/api/v1/query?query=up"' | head -c 100 oc exec -n sre-intelligence deploy/sre-intelligence-service -- \ sh -c 'curl -s --cacert /var/run/secrets/kubernetes.io/serviceaccount/service-ca.crt \ -H "Authorization: Bearer $(cat /var/run/secrets/kubernetes.io/serviceaccount/token)" \ "https://sre-observability-gateway-http.openshift-logging.svc.cluster.local:8080/api/logs/v1/application/loki/api/v1/labels"' | head -c 100 # Tempo gateway -- X-Scope-OrgID identifies the tenant; response {"traces":[...]} = working oc exec -n sre-intelligence deploy/sre-intelligence-service -- \ sh -c 'curl -s --cacert /var/run/secrets/kubernetes.io/serviceaccount/service-ca.crt \ -H "Authorization: Bearer $(cat /var/run/secrets/kubernetes.io/serviceaccount/token)" \ -H "X-Scope-OrgID: dev" \ "https://tempo-sre-observability-gateway.openshift-tempo.svc.cluster.local:8080/api/traces/v1/dev/tempo/api/search?limit=3&tags=service.name%3Dorder-service"' | head -c 200
Grant the SRE Intelligence service's ServiceAccount all the permissions it needs. Run these after oc apply -f sre-intelligence-deployment.yaml — the SA must exist before the bindings can reference it:
# 1. Query Thanos Querier (platform Prometheus) -- without this, Prometheus returns 403 oc adm policy add-cluster-role-to-user cluster-monitoring-view \ -z sre-intelligence-sa -n sre-intelligence # 2. Read application logs from the LokiStack gateway via SubjectAccessReview. # The Loki Operator does not auto-create this ClusterRole in all versions -- # create it explicitly so the binding always works regardless of operator version. oc apply -f - <<EOF apiVersion: rbac.authorization.k8s.io/v1 kind: ClusterRole metadata: name: cluster-logging-application-view rules: - apiGroups: [loki.grafana.com] resources: [application] verbs: [get] EOF oc adm policy add-cluster-role-to-user cluster-logging-application-view \ -z sre-intelligence-sa -n sre-intelligence
Configure AlertManager to deliver to the SRE Intelligence service
With the SRE Intelligence service running, configure the platform AlertManager to route business-workloads alerts to its webhook — the final wire that completes the pipeline.
AlertManager → SRE Intelligence service webhook
business-workloads namespace has openshift.io/cluster-monitoring=true, so its PrometheusRules are evaluated by the platform Prometheus and alerts route to the platform AlertManager in openshift-monitoring. The platform AlertManager does not pick up namespace-scoped AlertmanagerConfig resources — those only work with the user workload AlertManager. The webhook must be configured directly on the platform AlertManager's alertmanager-main secret instead.Export the current config, add the webhook receiver and route, and apply it back:
# Dump the current config
oc get secret alertmanager-main -n openshift-monitoring \
-o jsonpath='{.data.alertmanager\.yaml}' | base64 -d > alertmanager.yaml
Edit alertmanager.yaml to add the sre-intelligencement receiver and a matching route. The result should look like this (existing receivers and routes preserved, new entries highlighted):
"global":
"http_config":
"proxy_from_environment": true
"inhibit_rules":
- "equal":
- "namespace"
- "alertname"
"source_matchers":
- "severity = critical"
"target_matchers":
- "severity =~ warning|info"
- "equal":
- "namespace"
- "alertname"
"source_matchers":
- "severity = warning"
"target_matchers":
- "severity = info"
"receivers":
- "name": "Default"
- "name": "Watchdog"
- "name": "Critical"
- "name": "sre-intelligencement"
"webhook_configs":
- "url": "http://sre-intelligence-service.sre-intelligence.svc.cluster.local:8080/webhook"
"send_resolved": false
"route":
"group_by":
- "namespace"
"group_interval": "5m"
"group_wait": "30s"
"receiver": "Default"
"repeat_interval": "12h"
"routes":
- "matchers":
- "alertname = Watchdog"
"receiver": "Watchdog"
- "matchers":
- "severity = critical"
"receiver": "Critical"
"continue": true
- "matchers":
- "namespace = business-workloads"
"receiver": "sre-intelligencement"
"group_by":
- "alertname"
- "namespace"
- "pod"
"group_wait": "5s"
"group_interval": "30s"
"repeat_interval": "12h"
"continue": true on the severity = critical route is required — without it, critical-severity business-workloads alerts (like DataIntegrityCheckFailed) match the Critical receiver and never reach the sre-intelligencement webhook. With continue, both receivers fire.# Apply the updated config back to the secret oc create secret generic alertmanager-main \ --from-file=alertmanager.yaml \ -n openshift-monitoring \ --dry-run=client -o yaml | oc apply -f - # Verify the config was loaded (may take ~30s for AlertManager to reload) # grep for the receiver name, webhook URL, and namespace to confirm all three are present oc exec -n openshift-monitoring alertmanager-main-0 -c alertmanager -- \ cat /etc/alertmanager/config_out/alertmanager.env.yaml | \ grep -E "sre-intelligence|business-workloads|webhook"
XIII.Slack: Approval & Escalation
One Slack app handles both Tier 2 (approval) and Tier 3 (escalation) — PagerDuty is deliberately not used here; Tier 3 posts to a second, distinct channel using the same bot and token instead. Re-adding PagerDuty later is a symmetric change confined to slack_escalate(), nothing else in the pipeline needs to move.
1. Get a public HTTPS URL Slack will actually trust
-k/--tls-verify=false equivalent you can hand Slack; it enforces certificate trust with no override, so a self-signed cert means the Interactivity callback will always fail to connect, regardless of how the SRE Intelligence service itself is configured. This needs a URL with a real, publicly-trusted certificate.Fastest path for a public HTTPS URL: tunnel through ngrok, which provisions a trusted HTTPS URL for free and sidesteps the cluster's router certificate entirely.
# Terminal 1 -- keep running through setup and while the pipeline is in use oc port-forward svc/sre-intelligence-service 8080:8080 -n sre-intelligence # Terminal 2 -- also keep running; prints your public URL on startup ngrok http 8080 # Copy the "Forwarding" URL, e.g. https://xxxx.ngrok-free.app
2. Create the Slack App
api.slack.com/apps → Create New App → From scratch. Name it AI Ops Bot, pick your workspace.
3. Add Bot Token Scopes
OAuth & Permissions → Scopes → Bot Token Scopes — add each individually:
| Scope | Why |
|---|---|
chat:write | Post the Tier 2 approval message |
chat:write.public | Post without an explicit bot invite (skip if the channel is private — invite manually in step 7 instead) |
channels:read | Resolve the channel name to an ID |
4. Confirm Socket Mode is OFF
/slack/actions endpoint will never receive. It can default on depending on how the app was created; toggle it off here. If an App-Level Token already got generated, it's harmless to leave unused.5. Turn on Interactivity and set the Request URL
Interactivity & Shortcuts → toggle On → paste https://xxxx.ngrok-free.app/slack/actions into Request URL → Save.
oc logs deployment/sre-intelligence-service -n sre-intelligence.6. Install the app to your workspace
OAuth & Permissions → Install to Workspace → Allow. This is the actual "install a bot" step — Slack only generates the Bot User OAuth Token (xoxb-…) after this. Copy it.
7. Create both channels
Create #sre-approvals (Tier 2) and #sre-escalations (Tier 3) if they don't already exist. If either is private and you skipped chat:write.public, /invite @AI Ops Bot explicitly — without one of those two, the message post fails silently.
8. Wire the token into the cluster
# Patch the bot token into the Secret (credential, not in the ConfigMap) oc patch secret sre-intelligence-secrets -n sre-intelligence \ --type merge \ -p '{"stringData": {"SLACK_BOT_TOKEN": "xoxb-your-actual-token"}}' # If your channels differ from the defaults in sre-intelligence-config.yaml, patch the ConfigMap oc patch configmap sre-intelligence-config -n sre-intelligence \ --type merge \ -p '{"data": { "SLACK_APPROVAL_CHANNEL": "#your-approvals-channel", "SLACK_ESCALATION_CHANNEL": "#your-escalations-channel" }}' oc rollout restart deployment/sre-intelligence-service -n sre-intelligence
XIV.Audit Metrics
The metrics counters and /metrics endpoint already shipped in the SRE Intelligence service above — this section is confirming the endpoint, wiring Prometheus to it, and querying the result. No second build.
oc exec -n sre-intelligence deploy/sre-intelligence-service -- curl -s http://localhost:8080/metrics/ | head -5
# Trailing slash required -- FastAPI's app.mount("/metrics", make_asgi_app()) serves the
# sub-app at /metrics/, and /metrics (no slash) returns a redirect that curl doesn't follow.
# Expect Prometheus-format output: HELP aiops_alerts_total ... / TYPE aiops_alerts_total counter
Enable User Workload Monitoring
openshift-*) namespaces — a ServiceMonitor in business-workloads gets created successfully and silently does nothing: no error, zero scrape targets, zero data in every query below.oc apply -f - <<EOF
apiVersion: v1
kind: ConfigMap
metadata:
name: cluster-monitoring-config
namespace: openshift-monitoring
data:
config.yaml: |
enableUserWorkload: true
EOF
# Wait ~10s for the Cluster Monitoring Operator to reconcile the ConfigMap and
# create the StatefulSet before polling its rollout -- running oc rollout status
# immediately returns "Error: no StatefulSet found" if the object doesn't exist yet.
sleep 10
oc rollout status statefulset/prometheus-user-workload -n openshift-user-workload-monitoring --timeout=3m
Create the ServiceMonitor
oc apply -f - <<EOF
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
name: sre-intelligence-service
namespace: sre-intelligence
spec:
selector:
matchLabels:
app: sre-intelligence-service
endpoints:
- port: http # matches the Service port's NAME, not its number
path: /metrics/ # trailing slash required -- FastAPI's mount() serves the sub-app at /metrics/
interval: 30s
EOF
oc get servicemonitor sre-intelligence-service -n sre-intelligence
order-service's app_data_integrity_failures_total counter (Demo Workload section) is a custom application metric, not a platform cAdvisor one like the metrics the other alert rules use — it needs the same UWM-plus-ServiceMonitor treatment as the SRE Intelligence service, or DataIntegrityCheckFailed will never see a time series to alert on.oc apply -f - <<EOF
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
name: order-service
namespace: business-workloads
spec:
selector:
matchLabels:
app: order-service
endpoints:
- port: http # matches the Service port's NAME, set in the Demo Workload section
path: /metrics/ # trailing slash required -- same FastAPI mount() behavior as the SRE Intelligence service
interval: 15s
EOF
oc get servicemonitor order-service -n business-workloads
Query the metrics
OpenShift's default monitoring stack doesn't support creating custom Grafana dashboards through the console — Observe → Dashboards is a read-only viewer for pre-packaged dashboard ConfigMaps, not an editor. Use Observe → Metrics instead (Administrator perspective): a free-text PromQL box, no setup beyond what's above, updates live.
| What it shows | PromQL |
|---|---|
| Alerts per tier (last 24h) | sum by(tier) (increase(aiops_alerts_total[24h])) |
| Tier 1 auto-resolution rate (last 1h) | sum(rate(aiops_alerts_total{tier="1"}[1h])) / sum(rate(aiops_alerts_total[1h])) |
| LLM median latency | histogram_quantile(0.50, sum by(le) (rate(aiops_llm_duration_seconds_bucket[10m]))) |
| Actions by type (last 24h) | sum by(action_type) (increase(aiops_actions_total[24h])) |
sum() before dividing — Prometheus vector division matches by identical label sets by default, and a tier="1"-filtered vector won't line up cleanly against an unfiltered one carrying varying labels. The latency query aggregates by(le) before histogram_quantile — harmless with one replica, but the technically correct form, and necessary the moment this scales past one pod.XV.Testing & Validation
Test 1 is the flagship — the scenario named in a "cascading failure resolved in under 90 seconds" claim. Tests 2 and 3 demonstrate the other two trust tiers and are optional to run if you're validating the full pipeline; Test 1 alone demonstrates the headline claim.
interval: 10s rule-group evaluation and groupWait: 5s, worst case is roughly 10-15s of Prometheus/AlertManager latency on top of whatever the enrichment/LLM pipeline itself takes (observed in the 10-30s range) — comfortably inside a 90-second window with margin, without a synthetic shortcut.Reset script — run between every test
reset-demo-workload.shoc exec -n business-workloads deploy/inventory-service -- \
curl -s -X POST "http://localhost:8080/inject-latency?seconds=0"
oc rollout status deployment/order-service -n business-workloads --timeout=60s
oc delete networkpolicy -n business-workloads -l created-by=sre-intelligencement --ignore-not-found
# Corruption simulation is a persistent toggle, unlike the egress spike (which self-expires
# after its `duration` seconds) -- explicitly turn it off or DataIntegrityCheckFailed keeps firing.
oc exec -n business-workloads deploy/order-service -- \
curl -s -X POST "http://localhost:8080/simulate-corruption?enabled=false"
Test 1 — the flagship demo: cascading failure, Tier 1, under 90 seconds
# Make the downstream dependency slow -- this causes in-flight requests to pile up in order-service oc exec -n business-workloads deploy/inventory-service -- \ curl -s -X POST "http://localhost:8080/inject-latency?seconds=30" # Burst concurrent load -- exhausts order-service memory and triggers OOMKill oc exec -n business-workloads deploy/order-service -- \ curl -s -m 2 -X POST "http://localhost:8080/self-load?n=150" & # Tail the pipeline -- alert fires naturally via AlertManager, no synthetic webhook oc logs -f deploy/sre-intelligence-service -n sre-intelligence | grep -E "AUDIT|Processing|RCA|Restarted"
Processing: PodOOMKilled | ns=business-workloads | pod=order-service-xxx RCA: tier=1 conf=0.9x action=pod_restart Restarted deployment/order-service in business-workloads (source pod: order-service-xxx) AUDIT ... "root_cause":"...slow downstream call to inventory-service...","tier":1,"auto":true
oc exec -n sre-intelligence deploy/sre-intelligence-service -- curl -s -X POST http://localhost:8080/webhook -H "Content-Type: application/json" -d '{"alerts":[{"status":"firing","labels":{"alertname":"PodOOMKilled","severity":"warning","namespace":"business-workloads","pod":"<POD>"},"annotations":{"description":"Container was OOMKilled"},"startsAt":"2026-06-04T10:00:00Z"}]}'bash reset-demo-workload.sh
Test 2 — Tier 2 (supporting): genuine egress spike, Slack approval
Real fault this time — inventory-service's /simulate-egress-spike genuinely floods outbound requests, genuinely spiking container_network_transmit_bytes_total past the SuspiciousEgressTraffic rule's threshold. No hand-fired webhook.
# Fire the real spike -- 100 concurrent loops for 60s. n=100 (not 400) to stay within # inventory-service's 512Mi memory limit; duration=60 to outlast the alert's 2m rate window # + 20s `for:` clause. Tune both during your dry run if needed. oc exec -n business-workloads deploy/inventory-service -- \ curl -s -X POST "http://localhost:8080/simulate-egress-spike?n=100&duration=60" # After ~30s, confirm the transmit rate crossed the 1000 bytes/sec threshold oc exec -n openshift-monitoring prometheus-k8s-0 -c prometheus -- \ curl -s 'http://localhost:9090/api/v1/query' \ --data-urlencode 'query=rate(container_network_transmit_bytes_total{namespace="business-workloads",pod=~"inventory-service.*"}[2m])' # Wait for the real alert -- same principle as Test 1, no synthetic shortcut oc logs -f deploy/sre-intelligence-service -n sre-intelligence | grep -E "Tier 2|Quarantined|AUDIT"
A Slack message posts to #sre-approvals within seconds of the real alert firing. Review it carefully before clicking Approve.
- network_policy_change — the runbook-recommended action; quarantines the pod by blocking all egress except DNS while it stays running for investigation. Larger or instruction-following models tend to pick this.
- pod_restart — a valid but less precise choice; stops the current traffic spike but doesn't isolate the pod, and the same behaviour can recur immediately on restart. Smaller models (like 7B) may default to this when log lines look healthy (all
200 OK) and don't strongly signal a threat.
pod_restart, click Reject and apply the NetworkPolicy manually instead — that is the correct SRE response for suspicious egress.
# After approving, verify the action was applied oc get networkpolicy -n business-workloads # expect quarantine-inventory-service if network_policy_change was chosen bash reset-demo-workload.sh
Test 3 — Tier 3 (supporting): genuine data-integrity failure, Slack escalation with context pack
Toggle real corruption on, wait for the real alert — no runbook grounds this one, so the escalation's context pack is the LLM reasoning from the actual failure counter and logs alone, not a hand-typed payload.
# Toggle the real failure condition on oc exec -n business-workloads deploy/order-service -- \ curl -s -X POST "http://localhost:8080/simulate-corruption?enabled=true" # Wait for the real alert -- the first failed check lands within ~3s of the toggle, but # the rule still has to stay true through its own 15s `for:` window on top of that, checked # at 10s evaluation ticks. Expect this one to take a bit longer to trip than Tests 1/2 -- # roughly 25-40s before the webhook fires, not the near-instant PodOOMKilled case. oc logs -f deploy/sre-intelligence-service -n sre-intelligence | grep -E "Tier 3|Processing|AUDIT" # Verify a Slack message landed in #sre-escalations with a populated context pack -- # root_cause/confidence/reasoning fields grounded in "No runbook found for this alert", # not a raw alert dump bash reset-demo-workload.sh
Test 4 — Tier 1 (flagship variant): bad deployment auto-rolled back, no human click
This test demonstrates config_rollback — the third distinct remediation action — and is the most compelling scenario for audiences: the AI detects that a deployment introduced a bad configuration, identifies it as a regression rather than a transient runtime error, and automatically reverts to the last known-good revision without any human approval. The service recovers on its own.
Override the order-service container startup command to exit immediately with a config validation error. This simulates what happens when an app validates its configuration on startup and finds a broken dependency — a realistic deployment regression scenario. The pod crash-loops and its logs contain the exact keywords (connection refused, INVENTORY_URL) that guide the LLM toward config_rollback.
order-service resolves INVENTORY_URL lazily — only when /request is called, not on startup. A bad env var alone keeps the pod healthy. Overriding the command is the reliable way to produce a genuine crash loop with meaningful error logs for the demo.# Push a "bad" deployment revision that exits immediately with a config error. # The startup command simulates an app that validates its config before serving. # exit 1 → pod crash-loops; the log line contains keywords the LLM uses for RCA. oc patch deployment order-service -n business-workloads --type=json \ -p='[{"op":"add","path":"/spec/template/spec/containers/0/command", "value":["sh","-c","echo \"FATAL: cannot connect to inventory-service - Connection refused to http://nonexistent-service.business-workloads.svc.cluster.local:8080. Check INVENTORY_URL configuration.\" && exit 1"]}]' # Watch crash loop begin (pod exits immediately on every restart) oc get pods -n business-workloads -w # Tail SRE Intelligence service -- watch for config_rollback decision and execution oc logs -f deploy/sre-intelligence-service -n sre-intelligence | grep -E "AUDIT|Processing|RCA|rollback"
Processing: PodCrashLoopingDetected | ns=business-workloads | pod=order-service-xxx LOKI | fetched N log lines ← "connection refused" to nonexistent-service RCA: tier=1 conf=0.8x action=config_rollback ACTION | Tier 1 AUTO-EXECUTE: config_rollback on business-workloads/order-service-xxx MCP rollout_undo_deployment order-service in business-workloads: ... AUDIT ... "action":"config_rollback","tier":1,"auto":true
PodCrashLoopingDetected explicitly guides the model: if logs show a connection refused or wrong endpoint after a deployment change, prefer config_rollback over pod_restart.# Verify order-service recovered (previous good revision restored) oc rollout history deployment/order-service -n business-workloads oc get pods -n business-workloads -l app=order-service # Confirm the env variable was reverted oc exec -n business-workloads deploy/order-service -- env | grep INVENTORY_URL # Should show: INVENTORY_URL=http://inventory-service.business-workloads.svc.cluster.local:8080 bash reset-demo-workload.sh
Confidence threshold tuning
# All tunable parameters live in sre-intelligence-config — edit and reapply, then roll
oc patch configmap sre-intelligence-config -n sre-intelligence \
--type merge \
-p '{"data": {
"CONFIDENCE_TIER1": "0.80",
"CONFIDENCE_TIER2": "0.65",
"LLM_TEMPERATURE": "0.05",
"LLM_MAX_TOKENS": "1024"
}}'
oc rollout restart deployment/sre-intelligence-service -n sre-intelligence
Validation checklist
- GPU node Ready,
nvidia-smishows a T4 with the correct driver version - LLM endpoint responds to
/v1/chat/completionswithin 5 seconds, under the nameqwen2-5-7b-instruct - AlertManager webhook delivers to
/webhook(check SRE Intelligence service logs) - Loki query returns real log lines for
order-serviceafter a/self-loadburst - Tempo search returns a real trace for
order-servicewith a long span toinventory-service - Runbook ConfigMap lookup returns matching content for
PodOOMKilled - Test 1: run the real-timing validation once and write down the actual number — confirm it lands under 90 seconds before relying on this timing
- Test 1:
order-serviceOOMKills from a real AlertManager-delivered webhook (no hand-fired curl in the normal flow), the RCA cites the slow downstream call, pod recovers without human action - Test 2:
/simulate-egress-spikegenuinely tripsSuspiciousEgressTraffic(confirm viaoc logsthat it was AlertManager, not a curl, that delivered it); Slack message arrives in#sre-approvalswithin seconds; Approve triggersquarantine_pod()and a real NetworkPolicy appears - Test 3:
/simulate-corruptiongenuinely tripsDataIntegrityCheckFailed; Slack message posted to#sre-escalationswith populatedroot_cause/reasoninggrounded in "no runbook found," not a raw dump - Test 4:
oc set envbad downstream URL triggersPodCrashLoopingDetected; SRE Intelligence service auto-executesconfig_rollback(Tier 1, no human click);order-servicerecovers and Loki logs confirm connection-refused evidence was used in the RCA - Both new ServiceMonitors (
order-service, alongside the existingsre-intelligence-serviceone) show as scrape targets before relying on Test 3 - Observe → Metrics shows
aiops_alerts_totalincrementing by tier - AUDIT log entries present for each processed alert
reset-demo-workload.shrun and confirmed clean before your next run
Production Excellence
What makes this pipeline genuinely production-grade — implemented in this guide, not aspirationally.
Security
- Least-privilege RBAC. The MCP server runs with a scoped
ClusterRole— onlyget/list/patchon the exact resource types it uses. Nocluster-admin. - TLS everywhere, verified. Every internal HTTP client (Prometheus, Loki, Tempo) uses the cluster service CA at
/var/run/secrets/kubernetes.io/serviceaccount/service-ca.crt. Zeroverify=False. - Slack signature verification. All action callbacks are validated with HMAC-SHA256 + 5-minute replay window. Unsigned requests return HTTP 401 before any payload is processed.
- Webhook input validation. AlertManager payloads are validated via a Pydantic model at entry — malformed requests return HTTP 400 and never touch the pipeline.
- Prompt injection guards. Log lines are sanitised (instruction-override patterns stripped, hard-capped at 300 chars) before they reach the LLM prompt.
- Fail-loud configuration. All critical infrastructure URLs raise
RuntimeErrorat startup if unset — no silent fallback to stale defaults.
Reliability
- LLM circuit breaker. Three consecutive failures open the circuit and force
tier=3escalation for 5 minutes. Reset via a properasyncio.create_task()coroutine — no fragileget_event_loop(). - Per-call MCP connections. Each remediation action opens a fresh MCP connection. This avoids anyio cancel-scope constraints (the MCP library's
streamablehttp_clientcontext manager must stay alive within the same asyncio task; it cannot be shared acrossBackgroundTaskcoroutines). Per-call overhead is negligible for an alert-driven system; reliability is the priority. - Graceful signal degradation. If all three observability sources (Prometheus, Loki, Tempo) return empty context, the pipeline forces
tier=3escalation rather than letting the LLM reason blind. - Dependency-aware health check.
/healthprobes the LLM endpoint and returns{"status":"degraded"}with per-dependency detail — Kubernetes gets an honest liveness signal, not a hardcodedok.
Trust & Control
- Tier-action binding enforced in the system prompt. The alert's
tierlabel is authoritative: a tier-2 alert cannot select a tier-1 action, regardless of LLM reasoning. Rules encode SRE policy, not suggestions. - Schema-validated LLM output. The
action_typeis parsed against a known enum before anything executes. Unrecognised values escalate to Slack rather than silently failing or acting on bad data. - Runbook grounding. The LLM receives the team's actual runbook for each alert type before forming a hypothesis — the difference between general pattern-matching and institutional knowledge.
- Blast radius bounded by design. Tier 1 auto-executes above a confidence threshold. Tier 2 requires a human click. Tier 3 never auto-executes. This is enforced in code, not documentation.
Observability
- Correlation IDs on every log line. Each alert processed gets a short UUID prefix threaded through every log entry and the AUDIT record — concurrent incidents are traceable without log grepping.
- Immutable structured audit trail. Every pipeline decision — alert, root cause, confidence, action, tier, auto/manual — is emitted as a JSON AUDIT record. Every action is traceable, reviewable, and auditable.
- Prometheus metrics built in from day one.
aiops_alerts_total{tier,action},aiops_llm_duration_seconds, andaiops_actions_totalare emitted without configuration. ServiceMonitor is pre-wired for user-workload Prometheus. - Three-signal RCA context. Every LLM call is grounded in Prometheus metrics, Loki log lines, and a Tempo trace — the same signals a human engineer pulls manually, automated and correlated in one prompt.
Architecture
- Decision separated from execution. The SRE Intelligence service decides; the MCP server executes. The LLM output never directly invokes cluster APIs — it produces a validated enum that maps to a bounded set of MCP tools.
- In-cluster data sovereignty. No telemetry, log data, or alert context leaves the cluster. The LLM runs locally via KServe + vLLM. No data-residency or compliance exposure.
- Async throughout. The SRE Intelligence service uses
async/awaitandBackgroundTaskscorrectly — the webhook returns immediately, processing is non-blocking, and concurrent alerts are handled without thread contention. - Credentials separated from configuration. Sensitive values (LLM URL, Slack token, signing secret) live in a Kubernetes Secret; all non-sensitive configuration lives in a ConfigMap. Rotating a credential does not require re-building the image.
XVI.Evolving the Pipeline
This pipeline works end-to-end today. The following are the most impactful directions to evolve it — each one addresses a real limitation of the current design and has a clear, concrete implementation path.
Alert deduplication — Redis-backed fingerprinting
AlertManager retries webhook deliveries. Without deduplication, the same OOMKill can trigger three pod restarts in rapid succession. An in-memory dict is unsafe across multiple SRE Intelligence service replicas — the right answer is an external store.
How: Store a fingerprint (alertname:pod:startsAt) in Redis with a TTL equal to twice the AlertManager evaluation interval. Before processing, check: if the key exists, skip. This is a single redis.set(key, "1", ex=ttl, nx=True) call — the nx=True (set only if not exists) makes it atomic and idempotent across replicas. Red Hat's Data Grid (based on Infinispan) is the supported in-cluster option.
Durable buffer between AlertManager and the SRE Intelligence service
Direct HTTP webhooks are fire-and-forget: if the enrichment pod restarts during an incident, that alert is gone. A message queue decouples delivery from processing and gives you durability, replay, and backpressure — exactly what you need during an incident storm.
How: AMQ Streams (Kafka) is the Red Hat-supported option. AlertManager POSTs to the Kafka Bridge HTTP API (no Kafka client needed in AlertManager). The SRE Intelligence service runs an aiokafka consumer loop alongside the FastAPI webhook handler. Both call the same process_alert() function — the architecture change is shallow. Consumer group offsets mean a restarted pod picks up exactly where it left off.
Semantic search over incident history — vector DB + embeddings
The current runbook lookup is exact-match by alert name. A novel incident with no matching runbook gets "No runbook found" — the LLM reasons from first principles. The next step is grounding the LLM in your organisation's actual incident history: "here are the 5 most similar past incidents, what was their root cause and how were they resolved?"
How: Embed every AUDIT record using a sentence-transformer model (OpenShift AI supports this via KServe). Store embeddings in a vector database — pgvector on a PostgreSQL instance or a dedicated store like Weaviate. At enrichment time, embed the current alert context and retrieve the top-k most similar past incidents. Inject them into the LLM prompt alongside the runbook. Over time, the system learns your org's specific failure patterns instead of reasoning from generic SRE knowledge.
Policy gate between LLM decision and MCP execution — OPA or Kyverno
The current trust model (tier 1 → auto-execute, tier 2 → approve) is binary and time-invariant. It cannot answer: Is this action allowed for this specific service, right now, given our change freeze? A policy engine between the LLM's recommendation and the MCP call adds a dynamic, auditable enforcement layer.
How: Before call_mcp_tool(), call an OPA sidecar (or Kyverno policy) with the action context as input. The policy evaluates rules like:
- Is there an active change freeze for this namespace? (check a ConfigMap or API)
- Is this service marked as P0/critical in the service catalog? If so, require a Tier 2 approval even for Tier 1 actions.
- Is it outside business hours? Escalate instead of auto-execute.
- Has this deployment been restarted more than N times in the last hour? Block further restarts, escalate.
OPA returns allow: true/false with a reason. The SRE Intelligence service checks the decision before executing — blocked actions get escalated to Slack with the policy reason attached. This makes the automation governance-ready for regulated industries.
Human feedback loop — closing the learning cycle
Every time an on-call engineer clicks Reject on a Tier 2 Slack card, that rejection is lost. The system never learns that its recommendation was wrong. Closing this loop is the difference between an automation tool and one that improves with use.
How: When a Tier 2 action is rejected, prompt the engineer in Slack: "What was wrong? [Wrong action] [Wrong root cause] [False alarm] [Other]". Store the AUDIT record with the rejection reason and the engineer's label. Feed labelled rejections back as few-shot examples in the LLM prompt: "In a similar past incident, pod_restart was rejected because X — consider Y instead." Over time, this becomes fine-tuning data for a model that has learned your organisation's specific remediation preferences.