Incident copilot: retrieval that knows when to distrust itself
Awon Aziz — AI / MLOps engineer. 2026. Advisory only — three builds, 38 tests on the current one

An anomaly detector tells you a metric is strange. It does not tell you why, and it has no memory. This searches a postmortem knowledge base for incidents that resembled this one and puts a drafted root-cause hypothesis in front of a human.

THE HARD PART
The dense half of the hybrid retriever made rankings worse. Below roughly fifty documents the embedder has not learned enough for its ranking to mean anything, and averaging it in actively buried the correct answer. The system had to be able to switch off half of itself.

DECISIONS, AND WHAT EACH COST
1. Two agents with separate contexts, not one prompt asked to retrieve and reason at once
   Why: Keeps each agent's context scoped to one job, and lets a human read what the investigator actually found before the reporter dresses it up.
   Cost: Two model calls instead of one, and a handoff where information can be dropped.
2. Reciprocal rank fusion instead of averaging scores
   Why: A cosine similarity and a TF-IDF score do not average into anything meaningful. RRF sums 1/(k+rank) with k=60, so a document missing from one retriever simply contributes nothing from it.
   Cost: Rank-only fusion discards magnitude — a runaway best match and a marginal one look identical if both rank first.
3. A corpus-size trust gate in front of the dense path
   Why: Below 50 documents the dense signal is untrusted and queries fall back to sparse only. This is the whole result of the evaluation work.
   Cost: A magic number that is specific to this corpus and will need revisiting as the knowledge base grows.
4. TF-IDF first, neural embeddings as a documented one-line swap
   Why: The whole pipeline stands up with zero downloads and zero API keys, so a reviewer can run it.
   Cost: Lexical overlap only — 'CPU spike' and 'processor saturation' do not match, which is exactly what forced the next build.
5. Advisory only. Nothing calls an infrastructure API
   Why: Auto-remediation is a far larger and riskier scope than what is built, and pretending otherwise produces demos that are impressive right up until they delete something.
   Cost: It is not a self-healing system, and the word 'self-healing' appears nowhere in the repository.

DELIBERATELY NOT BUILT
- No agent writes to any infrastructure. Output is a draft for a person.
- No production knowledge base — the corpus is synthetic and labelled as such.
- The neural embedder path was not reachable in the sandbox, so the dense side ran on an LSA fallback and the numbers reflect that.

STACK: Python, Chroma, CrewAI, FastAPI, Kubernetes, scikit-learn

METRICS
- Tests, current build: 38
- Held-out eval cases: 8
- Builds superseded: 3
- Auto-remediation paths: 0

SOURCE
- https://github.com/AwonAziz/Hybrid-retrieval
- https://github.com/AwonAziz/agentic-incident-copilot
- https://github.com/AwonAziz/ai-incident-response-system

Full case study: https://awonaziz.github.io/project/incident-copilot/
Contact: awonaziz786@gmail.com