Awon Aziz AI & MLOps
Back to the work

Four agents, and a budget on what each one is told

Code review arrives after the pull request is open, which is the most expensive moment to find out something needed restructuring. This runs the first pass before a human is involved — and refuses to do the one thing that would make it dangerous.

Built
Aug 2026
Stages
4
Tests
20
Orchestration framework
None
Executes your code
Never

The problem

Feedback that arrives after a commit is feedback that costs a context switch to act on. The reviewer has to reconstruct what you were thinking, you have to reconstruct it too, and by then the shape of the code has set. The obvious response is to move the first pass earlier — before the pull request, while the code is still soft.

The less obvious problem is that a single prompt asked to "review this code" produces a paragraph of plausible observations with no structure, no severity, and nothing a tool downstream can act on. Making the output useful turned out to be more work than making it exist.

DecisionsAnd what each one cost

Context is budgeted per stage, not broadcast

Each agent receives only what its own job needs. The analyzer gets the language, the source and the static analysis evidence. The test engineer gets the source plus the analyzer's findings filtered to testing-relevant categories. The refactoring engineer gets the source plus the findings filtered to maintainability and readability. The reviewer gets both versions of the code, the generated tests and the key findings.

Cost. A filter that has to be maintained as finding categories change, and the risk that something genuinely cross-cutting gets withheld from the stage that needed it. Bought: fewer tokens, lower latency, and no upstream noise leaking into a stage's reasoning.

No orchestration framework

No LangChain, no LangGraph, no CrewAI, no AutoGen. Four prompt files, a client, and Pydantic models. For a linear four-stage pipeline a framework would have added a dependency, an abstraction and a new set of failure modes, in exchange for control flow that fits in one module.

Cost. Retries, parsing and error handling are hand-written here rather than inherited. When the pipeline stops being linear, that decision should be revisited rather than defended.

Model output is validated, not trusted

Every finding, test case and proposed change must conform to a Pydantic schema. The parser handles three response shapes — raw JSON, JSON inside a fenced block, and JSON with stray prose around it — and when parsing fails the system retries with an explicit correction prompt rather than guessing. Malformed responses raise an error instead of silently producing a plausible-looking empty result. Temperature sits at 0.0 to 0.1, because code analysis has no use for variety.

Cost. More failure paths to handle and a strictly narrower range of things the model is allowed to say.

The reviewer does not approve style

The final stage compares the original against the refactored version and checks behaviour preservation, correctness, complexity, testability and regressions. It explicitly does not approve a change that is only stylistic. A refactoring agent will always find something to rename; the reviewer exists to stop that counting as an improvement.

Cost. Real readability improvements that happen to be cosmetic get rejected along with the noise.

The constraintThat shaped everything else

From the repository, in capitals, near the top

The user can submit arbitrary code. Do not execute submitted code directly on the host system.

That single line decides the architecture. The pipeline generates tests but never runs them, which means it cannot verify that a refactor actually preserved behaviour — it can only reason about whether it should have. That is a real, acknowledged weakness, and it is the right trade against a tool that runs a stranger's code on the machine hosting it.

Submitted source is treated as untrusted input throughout and is never interpolated into a shell command. Sandboxed execution is listed as future work with the conditions it would require — isolation in a container with strict CPU, memory, filesystem, network and wall-clock limits — rather than as a vague intention. A feature described that precisely has been thought about; a feature described as "add sandboxing later" has not.

Evidence

Twenty tests, all with the model call mocked, so the suite runs with no API key and no network: schema validation, JSON parsing across all three response shapes including malformed ones, empty input handling, Python syntax parsing, and agent behaviour against fixed responses.

The repository ships examples/sample.py — deliberately imperfect but safe Python with nested conditionals, mixed responsibilities, weak validation and a reachable KeyError. A demo that analyses clean code proves nothing, so the demo input has real problems in it.

Not builtFrom the repository's own limitations

  • Single file only. No repository-level context, so anything whose problem lives in the relationship between modules is invisible to it.
  • No test execution, by design, which means no verified behaviour preservation.
  • Security analysis is pattern-based, not runtime. It reads code; it does not observe it.
  • Highly dynamic or metaprogrammed code gets weaker tests, because test generation is grounded in visible structure.
  • Needs an OpenRouter key. Unlike the incident copilot, there is no mock mode here.

Source

github.com/AwonAziz/AI-Pair-Engineer — the four prompt files under prompts/ are the most readable way in, then models/schemas.py for the contract every response has to satisfy.

Back to the other systems · Next: the incident copilot