Awon Aziz AI & MLOps

I build the part that runsafter the model ships.

Entry-level, based in Rawalpindi. My work sits on the operational half of machine learning — drift detection and model promotion, retrieval evaluation, incident triage, and scheduled automation that has to keep working when nobody is watching. Where something didn't work, the repository says so.

Job funnel — scanning right now reading…
10/18 job boards responding
2,067 open roles read this scan
23 matched my filters

10 of 18 sources responding. The rest return 404 for their board token and are shown, not hidden — a wrong token fails quietly and belongs in view.

Level 6 Diploma in AI Operations, Al Nafi International Colleges, 2026 Open to remote work and relocation

SystemsJun – Sep 2026

Four systems, each built to answer something the last one couldn't

They are all public and all runnable. Two of them are simulations and say so on the tin; one of them has been running unattended for a month. Each case study below covers the architecture, the decisions that were contested, what the decision cost, and what is deliberately not built.

3 repositoriesJun – Aug 2026Python, Chroma, CrewAI, FastAPI, Kubernetes

Incident copilot: retrieval that knows when to distrust itself

An anomaly detector flags a metric. The copilot searches a postmortem knowledge base for incidents that looked like this one, and two agents draft a root-cause hypothesis for a human to review. Retrieval is sparse and dense, fused with reciprocal rank fusion. An eval harness scores the whole pipeline against eight held-out anomalies.

What testing turned up

The dense retriever, falling back to LSA because the sandbox couldn't reach HuggingFace, ranked an unrelated incident first on two of the eight cases that sparse retrieval alone got right. Fusing it in made the system worse. The fix was a corpus-size gate: below 50 documents the dense signal is untrusted and queries fall back to sparse only.

Read the case study Hybrid-retrieval repo
Retrieval pipeline: an anomaly fans out to a sparse and a dense retriever, the dense path passes through a corpus-size trust gate, results are fused by reciprocal rank fusion, then two agents produce an advisory report. Anomaly flagged sparse TF-IDF dense MiniLM / LSA trust gate corpus ≥ 50 docs gate open? no RRF top-k postmortems matched_by: sparse / dense / both Investigator Reporter advisory report — human reviews, nothing auto-remediates
The gate is the whole point: when the corpus is too small for the dense embedder to have learned anything, its ranking is dropped rather than averaged in.

Running unattended since 16 AugPython, GitHub Actions, Pages

Job funnel: eighteen boards, every twenty minutes, no scraping

A scheduled workflow reads the same public ATS endpoints that company career pages are built from — Greenhouse, Lever, Ashby, SmartRecruiters — plus three aggregator feeds, filters to the roles worth applying to, dedupes against the previous run and commits the result. A static page renders it. It was built to run one person's job search, and it does.

A boundary, deliberately drawn

LinkedIn, Indeed, Wellfound and Turing are not touched. None publish a readable public API and LinkedIn's terms prohibit automated collection, so the funnel leaves them alone and the README says why. The Ashby and SmartRecruiters parsers were written from documentation rather than a captured live response, and that is flagged in the known-limitations section rather than glossed.

Read the case study Repository
Eleven company boards and three aggregator feeds fan into a scheduled scan, which filters, dedupes against a seen file, and commits JSON that a static page renders. tier 1 — company boards Greenhouse Lever Ashby SmartRecruiters tier 2 — aggregators RemoteOK We Work Remotely Remotive scheduled scan cron */20, workflow_dispatch, concurrency group filter titles + locations dedupe against seen.json commit jobs.json · seen.json · status.json static dashboard on Pages
If every source fails at once the previous jobs.json is left alone rather than overwritten with an empty file.

Simulated data, real machineryJul 2026MLflow, FastAPI, SciPy, Streamlit, Docker

Model lifecycle: drift, retrain, and a promotion that can be refused

A classifier is served over FastAPI while every prediction is logged. A detector compares live feature distributions against a stored reference using population stability index, a Kolmogorov–Smirnov test and Jensen–Shannon divergence. Severe drift triggers a retrain; the challenger only replaces the champion if it beats it by the configured margin.

What the repository actually contains

The committed champion metadata records F1 0.873 and AUC 0.937 — on the synthetic dataset the project generates itself, which is the only honest way to read those numbers. A committed drift report shows what a severe reading looks like: mean PSI 1.66 across five features after deliberate drift injection.

Read the case study Repository
A serving loop: the API logs predictions, the drift detector compares them to a reference, severe drift triggers retraining, and the challenger is promoted only if it beats the champion. FastAPI serving /predict · /health · /drift/status prediction log every request, with latency drift detector PSI ≥ 0.10 · KS p < 0.05 · JS ≥ 0.05 per feature, against stored reference no drift severe — retrain challenger champion promote only if F1 gain > margin
Promotion is a comparison, not a schedule. A retrained model that is no better stays out.

No orchestration frameworkAug 2026Python, Pydantic, Streamlit, OpenRouter

AI pair engineer: four agents, and a budget on what each one is told

A four-stage review pipeline that runs before a human reviewer sees the code: analyse, generate tests, refactor, then compare the original against the refactor and check behaviour survived. Each stage receives only the upstream findings relevant to its own job, and every response is validated against a Pydantic schema before it is trusted.

The constraint that shaped it

The tool accepts arbitrary code from a stranger, so it generates tests but never executes them. Submitted source is treated as untrusted input and is never interpolated into a shell command. Sandboxed execution is listed as future work with the resource limits it would need, rather than shipped and hoped about.

Read the case study Repository
Submitted code goes to an analyzer, whose findings are filtered separately to a test engineer and a refactoring engineer, and a final reviewer compares original against refactored output. submitted code — untrusted Code analyzer static evidence + model findings testing maintainability Test engineer edge + failure paths Refactor behaviour preserving Final reviewer regressions, not style opinions structured report — schema-validated, never executed
Context is scoped on purpose: the refactor stage never sees the testing findings, and the reviewer sees both versions plus the tests.

FoundationWhere the ops habits came from

Ninety-five lab exercises, committed one at a time

Before the systems above there was a year of cloud labs — every exercise with the commands that were actually run and a note on what they did. This is the reason drift thresholds and Kubernetes manifests turn up in the projects rather than being read about.

Lab repositories and smaller projects, with size and dates
Repository Contents Size Worked on
Devops-CICD-labs 53 labs across Jenkins, Argo CD, Buildkite, Concourse, Docker, Helm, Kubernetes, Prometheus, Grafana, Terraform and Ansible 119 commits Mar 2026
Cybersecurity-labs 34 labs across four tracks — SOC workflow, incident response and adversary emulation, digital forensics, Linux hardening 72 commits Mar – Apr 2026
RedHat-Linux-Labs 8 labs — Podman, multi-container apps, SSH hardening, journald analysis, system diagnostics 24 commits Feb 2026
agentic-incident-copilot The first agentic layer over the anomaly detector — Chroma, two CrewAI agents, Kubernetes manifests. Superseded by the hybrid retrieval build 13 tests Aug 2026
ai-incident-response-system Where the incident line started — Isolation Forest over simulated AWS, Azure and GCP telemetry, feeding a triage engine and a terminal dashboard simulated Jun 2026
music-visualizer Web Audio and Canvas, three modes, no framework. Written because React felt like overkill for it HTML / JS 2025
Crossy-roads in three.js A browser game built from memory of playing the original at nine. The repository description opens with “horribly optimised” three.js 2025
Paints-undo-colabs Colab notebooks for running Paints-UNDO, the drawing-process reconstruction model notebook 2025

In progressNo public repository yet

What is open on the other monitor

Listed because they are true, not because they are finished. None of these have a repository to link yet, so take them as intent rather than evidence.

  • LangGraph, specifically human-in-the-loop. The hybrid retrieval build ends at an Approve/Dismiss button with no state behind it. The next version should loop a dismissed report back for re-investigation and write an approved one into the knowledge base as a new postmortem — and the eval harness already accepts a pipeline function, so both versions can be scored against the same eight cases.
  • A LoRA fine-tuning project in PyTorch. Learning the training side properly rather than only the serving side.
  • An indie fighting game with a roster of more than seventy characters. A long-running personal project, and the reason three.js and game loops show up in the archive above.
  • An automated operating system. Early, ambitious, and entirely for its own sake.