Backend

Model Fallbacks and Circuit Breakers for AI Workers

How to route coding, review, and recovery across multiple model identities without retry storms or false review independence.

BACKENDModel Fallbacks and Circuit Breakers for AI WorkersCoreMesh Journal · engineering notes
Reserved advertising placementDisabled until AdSense approval

Multiple models improve resilience only when the router understands failure. Treating every non-zero result as “try the next model” can burn an entire quota pool, repeat a rejected design, or make the same model review its own work under a different label.

Build a failure taxonomy

Separate implementation rejection from technical unavailability. A review that identifies a concrete defect should stop the reviewer fallback chain and enter a repair cycle. A 429 response, expired authentication, malformed tool call, provider 5xx, or account lock is infrastructure evidence and may justify another provider or a bounded wait.

Persist circuit-breaker state

When a provider reports an account-wide limit, write a trusted state file with provider, status, checked time, retry hint, and the model used for the probe. Every worker process reads that state before reserving a lane. This turns thousands of failing calls into one lightweight probe.

{
  "provider": "chatgpt-codex",
  "status": "USAGE_LIMIT",
  "ready": false,
  "retry_hint": "provider supplied timestamp"
}

Track model identity

Producer identity belongs in the task ledger. If a fallback model touched the workspace before failing, it is still a producer and must be excluded from independent review. Review routes should compare exact model identity, not friendly names such as “primary” and “fallback.”

Route by workload

Fast models can handle classification, small patches, and formatting. Strong workhorse models fit normal repository tasks. Frontier models should be reserved for genuinely complex or high-value work. Independent review benefits from a strong but distinct model, while recovery should receive the reviewer feedback and remain bounded.

Fail over providers carefully

A secondary provider should pass a live inference probe, expose the models the router expects, and publish current quota state. Cached “ready” data is not enough after an account lock. If both providers are unavailable, preserve tasks in a quota wait state and stop dispatching.

Fail back deliberately

When the primary returns, allow active fallback tasks to finish. New tasks can move back to the primary after a trusted canary passes. Keep the fallback warm instead of deleting it, and avoid moving a workspace between models mid-turn unless the recovery protocol explicitly supports it.

Measure useful outcomes

Track calls, technical failures, code rejections, successful reviews, recovery cycles, tokens or cost when available, queue age, and tasks completed per provider. Raw call count is not throughput; a provider that returns fast but produces repeated review failures may be the expensive route.


How this article was produced

AI tools may assist with research organization, drafting, or editing. Indra Wijaya reviews each article, checks primary sources, verifies technical claims where practical, and remains responsible for the published result.

IW
ABOUT THE AUTHOR

Indra Wijaya

Open-source engineer and automation builder working on CoreMesh, AI-agent infrastructure, backend systems, CI/CD, and evidence-first software delivery.

Author profile →