Building Human-Gated Autonomous Coding Agents
A production architecture for autonomous code work that keeps public actions, identity, and irreversible decisions under explicit human control.
Autonomous coding is easy to demo and hard to operate. A model can edit a file in minutes, but a production system must decide whether the task is real, whether the repository permits the proposed work, whether tests actually prove the change, and whether anyone is authorized to publish it. The safest design treats those questions as separate gates instead of one long prompt.
Start with a control boundary
The worker should own local, reversible actions: clone a repository, create an isolated branch, inspect code, write a patch, run tests, and produce evidence. The owner should retain actions with external consequences: posting a claim, commenting on an issue, pushing a branch, opening a pull request, changing a wallet, or approving a workflow. This split makes failures recoverable and keeps an enthusiastic model from turning uncertainty into public noise.
discover → verify → build → test → independent review → owner approval → submit
Each arrow is a state transition with durable evidence. A dashboard can then show why a task is waiting instead of presenting every delay as a generic error.
Give the worker a task contract
A good task contract contains the repository, issue URL, source text, known acceptance criteria, test command, security constraints, and the exact artifact the worker must leave behind. It must also state that issue text and repository files are untrusted data. A bounty issue can contain prompt injection, requests for environment dumps, or instructions that conflict with repository policy.
The contract should never give the model credentials or permission to publish. The worker only needs a writable checkout and a bounded way to execute repository tools.
Use repository-native tests
A generic syntax check is useful, but it is not the same as the project test suite. The orchestrator should discover the narrowest authoritative test command, run it in a writable test environment, record the exact command and exit status, and prevent the reviewer from claiming broader coverage. If dependencies cannot be installed safely, the result is an environment limitation, not a fabricated pass.
Require independent review
The reviewer should use a different model identity from every model that modified the workspace. It receives the task contract, a trusted inventory of changed files, the full bounded diff, the maintainer-facing summary, and the test receipt. The reviewer runs read-only and returns a structured verdict with concrete defects. A bare failure token is not enough evidence.
Independence matters because the producer is naturally biased toward its own interpretation. A second model can catch scope expansion, unsupported claims, missing edge cases, and tests that do not exercise the changed path.
Make approval a real state
Approval should be stored in the ledger, not inferred from a button click that immediately starts a network request. A robust submission state machine preserves owner approval across transient GitHub failures and reconciles existing forks, branches, and pull requests before creating anything new. This prevents duplicate PRs after restarts.
Treat failure classes differently
- Code rejection: send concrete reviewer feedback through a bounded repair cycle.
- Provider outage: preserve the task in a retry state; do not label the code defective.
- Missing specification: hold for clarification instead of inventing a placeholder diff.
- Policy conflict: stop before implementation and surface the repository rule.
- Public API failure: retry idempotently while keeping approval and evidence.
What this architecture buys you
The system becomes understandable. Operators can see whether work is waiting for compute, scope, tests, review, owner approval, maintainer action, merge, or settlement. That clarity is more valuable than raw task throughput because it prevents false confidence and makes every public action explainable.
How this article was produced
AI tools may assist with research organization, drafting, or editing. Indra Wijaya reviews each article, checks primary sources, verifies technical claims where practical, and remains responsible for the published result.