Retrieval that knows when to distrust itself
Three builds over three months, each one answering something the previous one could not. It ends with a retriever that will switch off half of itself when the evidence says that half is making things worse.
The problem
An anomaly detector tells you a number is strange. It does not tell you why, and it has no memory — the same failure can happen four times and the detector will describe it identically four times, while the engineer who worked the first one has left or forgotten.
The gap is not detection. The gap is between this metric is anomalous and we have seen this before, here is what it turned out to be. Postmortems are exactly that institutional memory, and they sit unread in a wiki because nobody searches a wiki at 3am.
So: when the detector fires, retrieve the past incidents that resemble this one and put a drafted hypothesis in front of a human. Not an automatic fix — a first paragraph.
Three versionsWhat each one added
DecisionsAnd what each one cost
Two agents, not one prompt
Investigation and reporting run as separate agents with separate context rather than one prompt asked to retrieve and reason at once. Each agent's context stays scoped to one job, which mirrors the role-decomposition pattern used in current LLM root-cause-analysis work.
Cost. Two model calls instead of one, and a handoff where information can be dropped. Worth it for the ability to inspect what the investigator actually found before the reporter dresses it up.
Advisory only — nothing auto-remediates
Suggested actions are output for a human to read. No part of this module calls an infrastructure API. Auto-remediation is a meaningfully larger and riskier scope than what is built here, and pretending otherwise would be the kind of demo that is impressive until it deletes something.
Cost. It is not a self-healing system, and the word "self-healing" does not appear anywhere in the repository.
TF-IDF first, neural embeddings second
The first agentic build used TF-IDF embeddings inside Chroma so the whole pipeline would stand up with zero downloads and zero API keys. Chroma still does the real vector-database work — indexing, persistence, nearest-neighbour search — regardless of what feeds it. The upgrade path to a neural embedder is a documented one-line swap.
Cost. Lexical overlap only. "CPU spike" and "processor saturation" do not match. That limitation is what drove the next build.
Reciprocal rank fusion rather than score averaging
Two retrievers produce two rankings on incompatible scales — a cosine
similarity and a TF-IDF score do not average into anything meaningful.
RRF works on ranks instead, summing 1 / (k + rank) across
retrievers with k = 60 from the original paper. A document
missing from one retriever's list simply contributes nothing from it,
so no penalty term is needed.
Cost. Rank-only fusion throws away magnitude: a runaway best match and a marginal one look the same if they are both ranked first.
Fusion is kept dependency-free
fusion.py imports nothing — no Chroma, no embedders, no
models. That is what makes the fusion maths independently testable in
isolation, and it is why the seven fusion tests run in milliseconds
with no fixtures.
Cost. One more module in a small codebase.
The findingWhy the gate exists
The sandbox this was built in could not reach HuggingFace, so the dense retriever fell back to LSA — truncated SVD over the same TF-IDF matrix — instead of the intended pretrained sentence embedder. That fallback was supposed to be a convenience. Measuring it turned it into the most useful result in the project.
LSA fitted on a six-document corpus ranked an unrelated incident first on two of the eight golden cases that sparse retrieval alone got right. Naive fusion made those cases worse, not better. Six documents is nowhere near enough co-occurrence data for latent structure to exist, so the "latent structure" it found was noise.
A pretrained embedder does not have this failure mode — it is trained on external text and is not fitted to the corpus at all. So the problem is specific to the fallback, and the fix is specific too:
IncidentKnowledgeBase.MIN_CORPUS_SIZE_FOR_LSA = 50
Below that document count the dense signal is untrusted and queries run
sparse-only rather than fusing in a ranking known to hurt at this scale.
kb.dense_trusted and kb.active_dense_embedder_name
report which path is live, and every returned incident carries a
matched_by field showing whether sparse, dense or both
found it. Grow the corpus past fifty entries or use the real embedder
and the gate lifts on its own. It is a scale-aware decision, not a
permanent downgrade.
A second finding, about the sparse half
One of the eight golden cases describes an anomaly type not present in the knowledge base at all. TF-IDF scored it at 0.382 similarity — inside the 0.38 to 0.46 band that genuine matches occupy. It shares generic infrastructure vocabulary with the real postmortems while being about something else entirely, and lexical similarity cannot tell the difference. That is the argument for a dense retriever, written down before the dense retriever was added, and it is still the argument for finishing that upgrade on a machine with normal network access.
EvidenceTests, and the thing tests cannot do
Thirty-five tests across four files, all runnable without an LLM key or network access: postmortem parsing, retrieval correctness, the fusion maths in isolation, the LSA trust gate, orchestration output shape, and the evaluation harness's own correctness.
Tests and evaluation are not the same thing
The tests assert the code is correct — pass or fail, run on every change. The eval harness measures the quality of what the pipeline produces, against eight held-out anomaly queries, and is meant to be run on demand and tracked over time. It reports retrieval Hit@1 and Hit@3, a model-graded score from zero to five on whether the hypothesis matches the true root cause, whether the agent's stated confidence actually tracks correctness, and token usage.
Two of the eight cases deliberately have no clean answer
- One describes an anomaly type that is not in the knowledge base at all. The correct behaviour is low confidence, not a confident wrong match, so the harness scores it as appropriate uncertainty rather than forcing a hit or miss on a question with no right answer.
- One is genuinely ambiguous between two postmortems and carries an accepted alternate answer. Real incidents are often ambiguous, and an eval set that pretends otherwise is measuring the wrong thing.
With no LLM key set, the "hypothesis" being graded is the top retrieved root cause returned verbatim. Grading that against itself scores near-perfectly by construction, and confidence is always reported as low, so calibration cannot be measured either. Mock mode proves the harness runs end to end. It is not a quality signal, and the README says exactly that before anyone else can.
Not builtScope, stated
- The postmortem knowledge base is synthetic — six incidents written to match what the project's own Isolation Forest would plausibly flag: CPU saturation, memory leaks, network partitions, disk I/O contention, bad deploys, connection-pool exhaustion. The schema and ingestion path do not change when real incident history replaces it.
- The telemetry the detector reads is simulated, not live cloud metrics. The cloud SDKs sit commented out in
requirements.txtrather than imported and unused. - The human review step is currently Approve and Dismiss buttons with no state behind them. Rebuilding it as an explicit LangGraph — dismissed reports looping back for re-investigation, approved ones written into the knowledge base as new postmortems — is the next planned phase. The harness already accepts a pipeline function, which is exactly what makes the two implementations comparable when it exists.
- The numbers in this case study came from the LSA path. Nothing here has been re-measured with the real neural embedder on a machine with normal internet access, and the repository lists that as an open task rather than assuming the result.
Source
Worth opening in this order:
- Hybrid-retrieval — the current build.
fusion.pyfor the RRF maths,vector_store.pyfor the trust gate,eval/harness.pyandeval/golden_set.pyfor the evaluation. - agentic-incident-copilot — the first agentic layer, with the Kubernetes manifests.
- ai-incident-response-system — where the line started, in June.