AI pair engineer: four agents, and a budget on what each is told Awon Aziz — AI / MLOps engineer. 2026. Shipped — 4 stages, 20 tests, never executes your code A four-stage review pipeline that runs before a pull request exists: analyse, generate tests, refactor, then compare the original against the refactor and check behaviour survived. It stops at the report. THE HARD PART The tool accepts arbitrary code from a stranger. It generates tests but never executes them, and submitted source is never interpolated into a shell command. Making this useful meant making it safe enough to point at code you did not write. DECISIONS, AND WHAT EACH COST 1. Context is budgeted per stage, not broadcast to all of them Why: The test engineer gets the source plus the analyzer's findings filtered to testing categories. The refactoring engineer gets findings filtered to maintainability. The reviewer gets both versions plus the tests. Fewer tokens, lower latency, and no upstream noise leaking into a stage's reasoning. Cost: A filter that has to be maintained as categories change, and the risk that something genuinely cross-cutting gets withheld from the stage that needed it. 2. No orchestration framework Why: No LangChain, no LangGraph, no CrewAI, no AutoGen. Four prompt files, a client, and Pydantic models. For a linear pipeline a framework adds a dependency, an abstraction and a new set of failure modes in exchange for control flow that fits in one module. Cost: Retries, parsing and error handling are hand-written. When the pipeline stops being linear this decision should be revisited rather than defended. 3. Model output is validated against a schema, never trusted Why: The parser handles raw JSON, fenced JSON, and JSON with stray prose. When parsing fails it retries with an explicit correction prompt rather than guessing. Malformed responses raise instead of producing a plausible-looking empty result. Cost: More failure paths to handle, and a strictly narrower range of things the model is allowed to say. 4. The final reviewer does not approve style Why: A refactoring agent will always find something to rename. The reviewer checks behaviour preservation, correctness, complexity, testability and regressions — and explicitly rejects a change that is only cosmetic. Cost: Real readability improvements that happen to be cosmetic get rejected along with the noise. DELIBERATELY NOT BUILT - No sandboxed execution. Listed as future work along with the resource limits it would need, rather than shipped and hoped about. - No merge capability. It produces a report, a person decides. - Temperature is pinned to 0.0–0.1 because code analysis has no use for variety. STACK: Python, Pydantic, Streamlit, OpenRouter METRICS - Stages: 4 - Tests: 20 - Orchestration dependencies: 0 - Times it executes your code: 0 SOURCE - https://github.com/AwonAziz/AI-Pair-Engineer Full case study: https://awonaziz.github.io/project/pair-engineer/ Contact: awonaziz786@gmail.com