Model lifecycle: drift, retrain, and a promotion that can be refused
Awon Aziz — AI / MLOps engineer. 2026. Shipped — 43 tests, 3 drift measures, synthetic data

A model that passes every test and still quietly becomes wrong is the failure mode CI cannot see. The service stays healthy, no test fails, and the predictions stop meaning anything. This is the machinery for catching that.

THE HARD PART
The right drift thresholds are dataset-specific, so they ship as environment variables rather than constants. Anyone who claims there is a correct PSI threshold is guessing.

DECISIONS, AND WHAT EACH COST
1. Three drift measures, not one
   Why: PSI catches distribution shift in binned values, a Kolmogorov–Smirnov test asks whether the two samples plausibly share a distribution, and Jensen–Shannon divergence measures distance symmetrically. They fail to agree regularly, and each disagreement is information — a feature can shift in shape without shifting in bin mass.
   Cost: Three thresholds to tune and a severity rule that has to arbitrate between them.
2. Champion and challenger with a promotion margin
   Why: A retrained model does not replace the live one by virtue of being newer. It is logged to MLflow and promoted only if it beats the champion by the configured margin.
   Cost: Drift can be real and the retrain still produce nothing better. You have burned a pipeline run and are still serving a model you now know is drifting — a worse-feeling outcome, and the correct one.
3. A drift injector ships alongside the detector
   Why: scripts/inject_drift.py deliberately corrupts the input so a severe reading can be produced on demand. A detector that has never fired is untested.
   Cost: The committed severe report is manufactured, and the file says so in its header.
4. MLflow on SQLite by default
   Why: Someone cloning the repository gets working experiment tracking on the first command, with no infrastructure to stand up first. Overridable by environment variable.
   Cost: Single-writer. Fine for one machine, wrong for a team, and the override exists for exactly that reason.

DELIBERATELY NOT BUILT
- The dataset is synthetic, generated by the project. F1 0.873 and AUC 0.937 are only readable as 'on the data it made up'.
- No real traffic. Nothing here has ever served a production request.
- No model registry beyond MLflow's own tracking store.

STACK: Python, FastAPI, MLflow, Evidently AI, SciPy, Streamlit, Docker

METRICS
- Tests: 43
- Python: ~2,300 lines
- Drift measures: 3
- Mean PSI, severe: 1.655

SOURCE
- https://github.com/AwonAziz/ml-lifecycle-platform

Full case study: https://awonaziz.github.io/project/model-lifecycle/
Contact: awonaziz786@gmail.com