Awon Aziz AI & MLOps
Back to the work

Drift, retrain, and a promotion that can be refused

A model that passes every test and still quietly becomes wrong is the failure mode continuous integration cannot see. The code still runs. The predictions stop meaning anything. This is the machinery for catching that.

Built
Jul 2026
Python
~2,300 lines
Tests
43
Drift measures
3
Data
Synthetic

The problem

Most machine learning projects end at the notebook. The hard part starts after that: a model trained on last quarter's distribution keeps answering confidently when this quarter's data has moved. Nothing throws. No test fails. The service is healthy by every definition the deployment pipeline knows about.

Catching it means watching the data, not the process — comparing what the model is being asked now against what it was trained on, and having a decision ready for when those differ.

A real reportCommitted in the repository

drift_20260713_082618.json severity: severe · mean PSI 1.655
feature_0 PSI 4.875
feature_1 PSI 2.924
feature_2 PSI 0.386
feature_3 PSI 0.058
feature_4 PSI 0.032

Square root scale so the small values stay visible. The tick is the alert threshold, PSI 0.10. This report was produced by deliberately injecting drift, which is what the repository's inject_drift.py exists for.

Note the last two rows. Their PSI is under the threshold and their Kolmogorov–Smirnov p-values are 0.41 and 0.63 — comfortably not significant — yet both are flagged moderate, because Jensen–Shannon divergence crossed 0.05. Three measures, disagreeing, which is the entire reason to run three.

DecisionsAnd what each one cost

Three drift measures, not one

Population stability index catches distribution shift in binned feature values. A Kolmogorov–Smirnov test asks whether the reference and current samples plausibly came from the same distribution. Jensen–Shannon divergence measures distance symmetrically. They fail to agree regularly, and each disagreement is information — a feature can shift in shape without shifting in bin mass, or move slightly in a way that is statistically unremarkable but semantically real.

Cost. Three thresholds to tune instead of one, and a severity rule that has to arbitrate between them. All three are environment variables rather than constants, because the right values are dataset-specific and anyone who claims otherwise is guessing.

Champion and challenger, with a margin

A retrained model does not replace the live one by virtue of being newer. It is trained, logged to MLflow with its full parameter and metric set, and compared against the current champion. It only gets promoted if it beats the champion by the configured margin.

Cost. Drift can be real and the retrain can still produce nothing better, in which case you have burned a pipeline run and are still serving a model you now know is drifting. That is a worse-feeling outcome than automatic promotion, and a correct one.

MLflow on SQLite by default

Tracking runs against a local SQLite file rather than a hosted tracking server, overridable with an environment variable. Someone cloning the repository gets working experiment tracking on the first command, with no infrastructure to stand up first.

Cost. Single-writer. Fine for one machine, wrong for a team, and the override exists for exactly that reason.

A drift injector shipped alongside the detector

scripts/inject_drift.py deliberately corrupts the incoming distribution so the detector can be watched doing its job. A detector nobody has seen fire is a detector nobody should believe.

Cost. Every number in this case study came from synthetic drift on synthetic data. It proves the mechanism works, not that the thresholds are right for any real dataset.

EvidenceWhat is committed, and what it means

Forty-three tests across the drift detector, the training pipeline and the API. The champion metadata file in the repository records a real promoted run:

{
  "model_type": "random_forest",
  "f1":        0.873,
  "auc":       0.937,
  "precision": 0.885,
  "recall":    0.863,
  "cv_f1_mean":0.876,
  "promoted_at": "2026-07-13T08:26:18Z"
}

Those are real artefacts from a real run, and they describe a random forest classifying a dataset the project generated for itself. They say the pipeline trains, evaluates, compares and promotes correctly. They say nothing whatsoever about performance on a real problem, and quoting them without that sentence attached would be dishonest.

The serving surface

Nine endpoints, grouped into inference, operations and drift. /predict and /predict/batch with latency recorded per request; /health shaped as a Kubernetes liveness probe; /model/info reading current champion metadata; /metrics and a separate /metrics/prometheus for scraping; /drift/status and /drift/run; and /retrain. Every prediction is logged, which is what gives the drift detector a current distribution to compare against in the first place.

Not builtStated rather than implied

  • The CI workflows are described in the README but are not in the repository. Retraining and promotion exist and run as scripts — scripts/retrain.py — and are written to be driven by a scheduled job. Wiring them into GitHub Actions is real remaining work, not a formality, and the README is ahead of the code on this point.
  • The data is synthetic throughout. Generator, reference set and injected drift. The 800-row reference file is committed so the drift numbers above can be reproduced exactly, which is the useful thing about synthetic data and also its entire limitation.
  • No feature store and no serving-time schema contract. Drift is measured on the five features the generator produces; a real deployment needs the schema pinned somewhere upstream of the detector.

Source

github.com/AwonAziz/ml-lifecycle-platform — src/drift/detector.py for the three measures, config/settings.py for every threshold in one place, scripts/retrain.py for the promotion logic.

Back to the other systems · Next: the pair engineer