Job funnel: eighteen boards, every twenty minutes, no scraping
Awon Aziz — AI / MLOps engineer. 2026. Running unattended since 16 Aug 2026 — 587 automated commits

Nearly every company careers page is a thin client over an applicant tracking system, and those systems publish the listings as JSON at a public endpoint. A scheduled workflow reads those endpoints, filters, dedupes against the previous run and commits the result to a static page.

THE HARD PART
The largest job board in the world is outside the system, on purpose. LinkedIn's terms explicitly prohibit automated collection and they publish no readable public search API, so the funnel does not go near them. The correct cost was paying it.

DECISIONS, AND WHAT EACH COST
1. Two source tiers, labelled differently on the dashboard
   Why: Tier one is company ATS endpoints (Greenhouse, Lever, Ashby, SmartRecruiters) — freshest signal. Tier two is aggregators — coverage of companies not on the list. The badge colour tells you which kind of result you are reading.
   Cost: Tier one only works for the four supported systems. Workday and custom sites need a new parser written for them.
2. Ambiguous locations are kept, not dropped
   Why: A job survives the filter if its location matches the allow list OR if the location is missing or ambiguous. For a job search, wrongly dropping a real match is the expensive error.
   Cost: Noise. A filter that never wrongly drops anything will sometimes keep something irrelevant.
3. Seniority is tagged, not filtered
   Why: Hiding senior roles is a toggle on the dashboard that remembers itself on that device. The pipeline does not get to decide what the reader is allowed to see.
   Cost: One more piece of client state to carry.
4. Failed sources are shown, not hidden
   Why: Eight of eighteen sources currently return 404 for their board token. A wrong token fails quietly, and a failure that fails quietly belongs in view rather than looking like an oversight.
   Cost: The dashboard permanently shows a row of red. It looks like a bug, and that is the point.
5. If every source fails at once, the previous jobs.json is left alone
   Why: An empty successful response is worse than no response. The commit only lands if the scan produced real data.
   Cost: A stale page is possible if every source is genuinely dead for an extended period.

DELIBERATELY NOT BUILT
- No headless browser, no HTML scraping, no login replay.
- No LinkedIn, Indeed, Wellfound or Turing. Those stay manual and the README says why.
- The Ashby and SmartRecruiters parsers were written from documentation rather than a captured live response — flagged in known-limitations, not glossed.

STACK: Python, GitHub Actions, GitHub Pages, requests

METRICS
- Sources polled: 18
- Scan interval: 20 min
- Automated commits: 587
- Human commits, all time: 3
- Runtime dependencies: requests only

SOURCE
- https://github.com/AwonAziz/cleanjobfunnel
- https://awonaziz.github.io/cleanjobfunnel/

Full case study: https://awonaziz.github.io/project/job-funnel/
Contact: awonaziz786@gmail.com