Awon Aziz AI & MLOps
Back to the work

Retrieval that knows when to distrust itself

Three builds over three months, each one answering something the previous one could not. It ends with a retriever that will switch off half of itself when the evidence says that half is making things worse.

Built
Jun – Aug 2026
Repositories
3
Tests, current build
35
Evaluation cases
8
Status
Advisory only

The problem

An anomaly detector tells you a number is strange. It does not tell you why, and it has no memory — the same failure can happen four times and the detector will describe it identically four times, while the engineer who worked the first one has left or forgotten.

The gap is not detection. The gap is between this metric is anomalous and we have seen this before, here is what it turned out to be. Postmortems are exactly that institutional memory, and they sit unread in a wiki because nobody searches a wiki at 3am.

So: when the detector fires, retrieve the past incidents that resemble this one and put a drafted hypothesis in front of a human. Not an automatic fix — a first paragraph.

Three versionsWhat each one added

A timeline of three builds: June, an Isolation Forest detector over simulated multi-cloud telemetry; August 7, an agentic retrieval layer with two agents and a Chroma knowledge base; August 9, hybrid sparse and dense retrieval with reciprocal rank fusion and an evaluation harness. 24 Jun 2026 7 Aug 2026 9 Aug 2026 Detector Isolation Forest, unsupervised AWS / Azure / GCP simulated rule engine + triage engine 4 severity tiers, dedupe Slack / PagerDuty routing Rich terminal dashboard Open question: it can say what is strange, not what it probably is. Agentic layer Chroma over 6 postmortems TF-IDF embeddings Investigator + Reporter agents FastAPI + Streamlit review queue Kubernetes manifests runs without an LLM key Open question: is the retrieval any good? Nothing measured it. Hybrid retrieval + evaluation sparse TF-IDF and dense MiniLM, fused with RRF (k = 60) LSA fallback when the embedder cannot be downloaded trust gate: dense dropped below 50 documents every result records which retriever found it eval harness — 8 golden anomalies, Hit@1 / Hit@3, model-graded root cause, confidence calibration What it found: the dense half was hurting. So it gets switched off until the corpus is big enough to support it.
Each version exists because the previous one raised a question it could not answer itself. The third one is the first that can be wrong out loud.

DecisionsAnd what each one cost

Two agents, not one prompt

Investigation and reporting run as separate agents with separate context rather than one prompt asked to retrieve and reason at once. Each agent's context stays scoped to one job, which mirrors the role-decomposition pattern used in current LLM root-cause-analysis work.

Cost. Two model calls instead of one, and a handoff where information can be dropped. Worth it for the ability to inspect what the investigator actually found before the reporter dresses it up.

Advisory only — nothing auto-remediates

Suggested actions are output for a human to read. No part of this module calls an infrastructure API. Auto-remediation is a meaningfully larger and riskier scope than what is built here, and pretending otherwise would be the kind of demo that is impressive until it deletes something.

Cost. It is not a self-healing system, and the word "self-healing" does not appear anywhere in the repository.

TF-IDF first, neural embeddings second

The first agentic build used TF-IDF embeddings inside Chroma so the whole pipeline would stand up with zero downloads and zero API keys. Chroma still does the real vector-database work — indexing, persistence, nearest-neighbour search — regardless of what feeds it. The upgrade path to a neural embedder is a documented one-line swap.

Cost. Lexical overlap only. "CPU spike" and "processor saturation" do not match. That limitation is what drove the next build.

Reciprocal rank fusion rather than score averaging

Two retrievers produce two rankings on incompatible scales — a cosine similarity and a TF-IDF score do not average into anything meaningful. RRF works on ranks instead, summing 1 / (k + rank) across retrievers with k = 60 from the original paper. A document missing from one retriever's list simply contributes nothing from it, so no penalty term is needed.

Cost. Rank-only fusion throws away magnitude: a runaway best match and a marginal one look the same if they are both ranked first.

Fusion is kept dependency-free

fusion.py imports nothing — no Chroma, no embedders, no models. That is what makes the fusion maths independently testable in isolation, and it is why the seven fusion tests run in milliseconds with no fixtures.

Cost. One more module in a small codebase.

The findingWhy the gate exists

The sandbox this was built in could not reach HuggingFace, so the dense retriever fell back to LSA — truncated SVD over the same TF-IDF matrix — instead of the intended pretrained sentence embedder. That fallback was supposed to be a convenience. Measuring it turned it into the most useful result in the project.

Compared sparse-only, dense-only and fused rankings side by side

LSA fitted on a six-document corpus ranked an unrelated incident first on two of the eight golden cases that sparse retrieval alone got right. Naive fusion made those cases worse, not better. Six documents is nowhere near enough co-occurrence data for latent structure to exist, so the "latent structure" it found was noise.

A pretrained embedder does not have this failure mode — it is trained on external text and is not fitted to the corpus at all. So the problem is specific to the fallback, and the fix is specific too:

IncidentKnowledgeBase.MIN_CORPUS_SIZE_FOR_LSA = 50

Below that document count the dense signal is untrusted and queries run sparse-only rather than fusing in a ranking known to hurt at this scale. kb.dense_trusted and kb.active_dense_embedder_name report which path is live, and every returned incident carries a matched_by field showing whether sparse, dense or both found it. Grow the corpus past fifty entries or use the real embedder and the gate lifts on its own. It is a scale-aware decision, not a permanent downgrade.

A second finding, about the sparse half

One of the eight golden cases describes an anomaly type not present in the knowledge base at all. TF-IDF scored it at 0.382 similarity — inside the 0.38 to 0.46 band that genuine matches occupy. It shares generic infrastructure vocabulary with the real postmortems while being about something else entirely, and lexical similarity cannot tell the difference. That is the argument for a dense retriever, written down before the dense retriever was added, and it is still the argument for finishing that upgrade on a machine with normal network access.

EvidenceTests, and the thing tests cannot do

Thirty-five tests across four files, all runnable without an LLM key or network access: postmortem parsing, retrieval correctness, the fusion maths in isolation, the LSA trust gate, orchestration output shape, and the evaluation harness's own correctness.

Tests and evaluation are not the same thing

The tests assert the code is correct — pass or fail, run on every change. The eval harness measures the quality of what the pipeline produces, against eight held-out anomaly queries, and is meant to be run on demand and tracked over time. It reports retrieval Hit@1 and Hit@3, a model-graded score from zero to five on whether the hypothesis matches the true root cause, whether the agent's stated confidence actually tracks correctness, and token usage.

Two of the eight cases deliberately have no clean answer

  • One describes an anomaly type that is not in the knowledge base at all. The correct behaviour is low confidence, not a confident wrong match, so the harness scores it as appropriate uncertainty rather than forcing a hit or miss on a question with no right answer.
  • One is genuinely ambiguous between two postmortems and carries an accepted alternate answer. Real incidents are often ambiguous, and an eval set that pretends otherwise is measuring the wrong thing.
Read the mock-mode numbers sceptically — the repository says so first

With no LLM key set, the "hypothesis" being graded is the top retrieved root cause returned verbatim. Grading that against itself scores near-perfectly by construction, and confidence is always reported as low, so calibration cannot be measured either. Mock mode proves the harness runs end to end. It is not a quality signal, and the README says exactly that before anyone else can.

Not builtScope, stated

  • The postmortem knowledge base is synthetic — six incidents written to match what the project's own Isolation Forest would plausibly flag: CPU saturation, memory leaks, network partitions, disk I/O contention, bad deploys, connection-pool exhaustion. The schema and ingestion path do not change when real incident history replaces it.
  • The telemetry the detector reads is simulated, not live cloud metrics. The cloud SDKs sit commented out in requirements.txt rather than imported and unused.
  • The human review step is currently Approve and Dismiss buttons with no state behind them. Rebuilding it as an explicit LangGraph — dismissed reports looping back for re-investigation, approved ones written into the knowledge base as new postmortems — is the next planned phase. The harness already accepts a pipeline function, which is exactly what makes the two implementations comparable when it exists.
  • The numbers in this case study came from the LSA path. Nothing here has been re-measured with the real neural embedder on a machine with normal internet access, and the repository lists that as an open task rather than assuming the result.

Source

Worth opening in this order:

Back to the other systems · Next: the job funnel