Model Fallbacks and Circuit Breakers for AI Workers
How to route coding, review, and recovery across multiple model identities without retry storms or false review independence.
Multiple models improve resilience only when the router understands failure. Treating every non-zero result as “try the next model” can burn an entire quota pool, repeat a rejected design, or make the same model review its own work under a different label.
Build a failure taxonomy
Separate implementation rejection from technical unavailability. A review that identifies a concrete defect should stop the reviewer fallback chain and enter a repair cycle. A 429 response, expired authentication, malformed tool call, provider 5xx, or account lock is infrastructure evidence and may justify another provider or a bounded wait.
Persist circuit-breaker state
When a provider reports an account-wide limit, write a trusted state file with provider, status, checked time, retry hint, and the model used for the probe. Every worker process reads that state before reserving a lane. This turns thousands of failing calls into one lightweight probe.
{
"provider": "chatgpt-codex",
"status": "USAGE_LIMIT",
"ready": false,
"retry_hint": "provider supplied timestamp"
}
Track model identity
Producer identity belongs in the task ledger. If a fallback model touched the workspace before failing, it is still a producer and must be excluded from independent review. Review routes should compare exact model identity, not friendly names such as “primary” and “fallback.”
Route by workload
Fast models can handle classification, small patches, and formatting. Strong workhorse models fit normal repository tasks. Frontier models should be reserved for genuinely complex or high-value work. Independent review benefits from a strong but distinct model, while recovery should receive the reviewer feedback and remain bounded.
Fail over providers carefully
A secondary provider should pass a live inference probe, expose the models the router expects, and publish current quota state. Cached “ready” data is not enough after an account lock. If both providers are unavailable, preserve tasks in a quota wait state and stop dispatching.
Fail back deliberately
When the primary returns, allow active fallback tasks to finish. New tasks can move back to the primary after a trusted canary passes. Keep the fallback warm instead of deleting it, and avoid moving a workspace between models mid-turn unless the recovery protocol explicitly supports it.
Measure useful outcomes
Track calls, technical failures, code rejections, successful reviews, recovery cycles, tokens or cost when available, queue age, and tasks completed per provider. Raw call count is not throughput; a provider that returns fast but produces repeated review failures may be the expensive route.
How this article was produced
AI tools may assist with research organization, drafting, or editing. Indra Wijaya reviews each article, checks primary sources, verifies technical claims where practical, and remains responsible for the published result.