Awon Aziz AI & MLOps
Back to the work

Eighteen boards, every twenty minutes, no scraping

The only system here that has been running continuously. It was built to handle one person's job search and has been doing exactly that, unattended, since the middle of August.

Running since
16 Aug 2026
Scan interval
20 min
Sources configured
18
Automated commits
486
Dependencies
requests
Latest scan reading…
10/18 job boards responding
2,067 open roles read
23 matched the filters

10 of 18 sources responding. The rest return 404 for their board token and are shown, not hidden.

The problem

Job boards are an interface problem pretending to be a search problem. The listing you want is usually on the company's own careers page hours before an aggregator picks it up, and by the time you see it on a board it has four hundred applicants. Checking twenty career pages by hand, several times a day, is not a thing a person does for long.

Nearly every company careers page is a thin client over an applicant tracking system, and those systems publish the listings as JSON at a public endpoint — it is the same data the careers page itself renders. No login, no bypassed authentication, nothing not meant to be read that way. So read that instead.

DecisionsAnd what each one cost

Two tiers, and they are labelled differently on the dashboard

Tier one is direct from the company — Greenhouse, Lever, Ashby, SmartRecruiters — scoped to a named list in config/companies.json. Freshest signal. Tier two is aggregators: RemoteOK, We Work Remotely, Remotive, which catch roles at companies not on the list. The dashboard colours the source badge by tier so it is obvious which kind of result you are reading.

Cost. Tier one only works for companies whose careers page runs on one of the four supported systems. Workday and custom sites cannot be added without writing a new parser.

LinkedIn, Indeed, Wellfound and Turing are deliberately untouched

None of them publish a readable public API for search results, and LinkedIn's terms explicitly prohibit automated collection. So the funnel does not go near them and the README states that plainly rather than leaving the omission to look like an oversight. Those stay manual; everything here just cuts down how often the manual check is needed.

Cost. The largest job board in the world is outside the system. That is the correct cost to pay.

Ambiguous locations are kept, not dropped

A job survives the location filter if its location text matches the allow list or if the location is missing or ambiguous. The filter errs towards showing too much rather than silently discarding a real match.

Cost. Noise. A filter that never wrongly drops anything will sometimes keep something irrelevant, and for a job search that is the cheaper error.

Seniority is tagged, not filtered

Titles matching the seniority patterns are marked rather than removed, and hiding them is a toggle on the dashboard that remembers the choice on that device. The pipeline does not get to decide what the reader is allowed to see.

Cost. One more piece of state to carry.

Recency shown as a border colour, because the real number is not available

The useful signal is "under a hundred applicants", and LinkedIn is the only board that exposes that figure — inside its own app only. The honest substitute is time: a green left border means first seen in the last six hours, cyan means the last twenty-four. A proxy, labelled as a proxy.

Cost. First-seen is when this funnel noticed a posting, not when it went live. Close enough at a twenty-minute interval; not the same thing.

Operating itWhat unattended actually requires

Most of what is interesting about this project is not the parsing. It is the handful of things that have to be true for a workflow to run four hundred and eighty-six times without anyone looking at it.

  • Total failure does not destroy data. If every source fails — no network, a bad deploy, anything — the previous jobs.json is left untouched rather than overwritten with an empty list. There is an integration test for exactly this case.
  • Pushes retry, and back off. Two scans overlapping would race on the same branch. The commit step retries up to five times, resetting onto the latest main and re-committing with a randomised sleep between attempts.
  • Runs cannot overlap. A concurrency group with cancel-in-progress: false lets a running scan finish rather than killing it with the next one.
  • Per-source health is written out, not just logged. Every scan emits status.json with an ok flag, a result count and the error string for each source. That file is what the panel at the top of this page is reading.
  • The schedule is a suggestion. GitHub does not guarantee timing on scheduled workflows and a run can slip under load, so the dashboard treats anything over ninety minutes old as amber rather than pretending the cron is a stopwatch.

486 of the 489 commits were made by the workflow

Worth saying plainly, because a commit count is easy to misread. Three commits are hand-written; the rest are chore: refresh job data from job-funnel-bot. That number is not evidence of a month of coding. It is evidence of a month of uninterrupted running, which is the thing the project was actually for.

Evidence

Tests make no live network calls. Each parser is checked against a sample payload shaped like that provider's documented schema, with requests.get mocked, so the suite cannot go red because a board was having a bad afternoon.

python tests/test_parsers.py       # every source parser, in isolation
python tests/test_integration.py   # dedupe, first_seen persistence,
                                   # and the all-sources-failed safety net

Roughly 690 lines of Python and one third-party dependency. There is no database, no queue, no framework and no server — a scheduled script, a JSON file and a static page. Everything else would have been scaffolding around a problem that did not have it.

Known limitsFrom the repository's own list

  • The Greenhouse and Lever parsers were verified against real live responses. Ashby and SmartRecruiters were built from official documentation rather than a captured payload. If one of those returns zero results for a company that is obviously hiring, the field names probably need adjusting — and the raw response in that run's log is the first place to look.
  • A job disappearing from the dashboard almost always means its source stopped returning it — filled or closed — rather than a bug. Almost always is not always.
  • Board tokens in config/companies.json are partly educated guesses. A wrong guess fails quietly and shows up in the Actions log, which is why the status file surfaces the failures instead of hiding them.

The panel at the top of this page shows those failures live. Eight of the eighteen configured sources are currently returning 404 for their board token. That is what the system actually looks like, so that is what it shows.

Source

github.com/AwonAziz/cleanjobfunnel — start with scripts/fetch_jobs.py and .github/workflows/scan.yml.

Back to the other systems · Next: the model lifecycle