Drift, retrain, and a promotion that can be refused
A model that passes every test and still quietly becomes wrong is the failure mode continuous integration cannot see. The code still runs. The predictions stop meaning anything. This is the machinery for catching that.
The problem
Most machine learning projects end at the notebook. The hard part starts after that: a model trained on last quarter's distribution keeps answering confidently when this quarter's data has moved. Nothing throws. No test fails. The service is healthy by every definition the deployment pipeline knows about.
Catching it means watching the data, not the process — comparing what the model is being asked now against what it was trained on, and having a decision ready for when those differ.
A real reportCommitted in the repository
Square root scale so the small values stay visible. The tick is the
alert threshold, PSI 0.10. This report was produced by
deliberately injecting drift, which is what the repository's
inject_drift.py exists for.
DecisionsAnd what each one cost
Three drift measures, not one
Population stability index catches distribution shift in binned feature values. A Kolmogorov–Smirnov test asks whether the reference and current samples plausibly came from the same distribution. Jensen–Shannon divergence measures distance symmetrically. They fail to agree regularly, and each disagreement is information — a feature can shift in shape without shifting in bin mass, or move slightly in a way that is statistically unremarkable but semantically real.
Cost. Three thresholds to tune instead of one, and a severity rule that has to arbitrate between them. All three are environment variables rather than constants, because the right values are dataset-specific and anyone who claims otherwise is guessing.
Champion and challenger, with a margin
A retrained model does not replace the live one by virtue of being newer. It is trained, logged to MLflow with its full parameter and metric set, and compared against the current champion. It only gets promoted if it beats the champion by the configured margin.
Cost. Drift can be real and the retrain can still produce nothing better, in which case you have burned a pipeline run and are still serving a model you now know is drifting. That is a worse-feeling outcome than automatic promotion, and a correct one.
MLflow on SQLite by default
Tracking runs against a local SQLite file rather than a hosted tracking server, overridable with an environment variable. Someone cloning the repository gets working experiment tracking on the first command, with no infrastructure to stand up first.
Cost. Single-writer. Fine for one machine, wrong for a team, and the override exists for exactly that reason.
A drift injector shipped alongside the detector
scripts/inject_drift.py deliberately corrupts the
incoming distribution so the detector can be watched doing its job.
A detector nobody has seen fire is a detector nobody should believe.
Cost. Every number in this case study came from synthetic drift on synthetic data. It proves the mechanism works, not that the thresholds are right for any real dataset.
EvidenceWhat is committed, and what it means
Forty-three tests across the drift detector, the training pipeline and the API. The champion metadata file in the repository records a real promoted run:
{
"model_type": "random_forest",
"f1": 0.873,
"auc": 0.937,
"precision": 0.885,
"recall": 0.863,
"cv_f1_mean":0.876,
"promoted_at": "2026-07-13T08:26:18Z"
}
Those are real artefacts from a real run, and they describe a random forest classifying a dataset the project generated for itself. They say the pipeline trains, evaluates, compares and promotes correctly. They say nothing whatsoever about performance on a real problem, and quoting them without that sentence attached would be dishonest.
The serving surface
Nine endpoints, grouped into inference, operations and drift.
/predict and /predict/batch with latency
recorded per request; /health shaped as a Kubernetes
liveness probe; /model/info reading current champion
metadata; /metrics and a separate
/metrics/prometheus for scraping;
/drift/status and /drift/run; and
/retrain. Every prediction is logged, which is what gives
the drift detector a current distribution to compare against in the
first place.
Not builtStated rather than implied
- The CI workflows are described in the README but are not in the repository. Retraining and promotion exist and run as scripts —
scripts/retrain.py— and are written to be driven by a scheduled job. Wiring them into GitHub Actions is real remaining work, not a formality, and the README is ahead of the code on this point. - The data is synthetic throughout. Generator, reference set and injected drift. The 800-row reference file is committed so the drift numbers above can be reproduced exactly, which is the useful thing about synthetic data and also its entire limitation.
- No feature store and no serving-time schema contract. Drift is measured on the five features the generator produces; a real deployment needs the schema pinned somewhere upstream of the detector.
Source
github.com/AwonAziz/ml-lifecycle-platform
— src/drift/detector.py for the three measures,
config/settings.py for every threshold in one place,
scripts/retrain.py for the promotion logic.