Recovering a Production Stack After a Cloud Account Lock
A practical recovery sequence for rebuilding public services, control planes, and worker queues when the original cloud account becomes inaccessible.
An account lock changes the problem from service repair to disaster recovery. The original virtual machine may still exist, but operators cannot assume they can reach it, snapshot it, or retrieve its database. Recovery should begin on infrastructure controlled through an independent account.
Freeze assumptions first
Mark the old provider unavailable even if the last health snapshot said ready. Stop using cached model, billing, or compute state. Do not rotate keys repeatedly or create accounts to evade enforcement. Preserve support evidence and request that the provider retain the original resources.
Inventory what is recoverable
Classify each component by source:
- Versioned source: clone from GitHub and verify the exact commit.
- Static public content: rebuild from source or a known release artifact.
- Configuration: reconstruct from documented environment variables and secret stores.
- Stateful data: restore only from a verified backup; do not invent historical records.
- Identity and DNS: keep at the registrar or DNS provider independent from the failed compute account.
Create a separate staging surface
Use temporary hostnames and block indexing. Bring up health endpoints, application services, and monitoring before changing production DNS. Staging should use the same process manager and routing pattern intended for production, with public writes and automated submissions disabled.
Respect the recovery host
A smaller fallback server may not support the original concurrency. Reduce worker lanes, add memory limits, and preserve the queue rather than overcommitting the host. Static publishing and APIs are cheap; model workers, test suites, browsers, and language toolchains can consume gigabytes.
Rebuild authentication without copying secrets into chat
Reuse an existing trusted secret store when possible, or perform an interactive owner login directly on the host. Keep session secrets, GitHub tokens, model credentials, and wallet metadata out of source control. Start with read-only integrations and enable writes only after reconciliation.
Cut over in layers
- Move the public journal and verify HTTP, TLS, sitemap, and analytics.
- Move stateless APIs and verify readiness against their external dependencies.
- Move the private control plane with authentication and database integrity checks.
- Connect GitHub and model providers interactively.
- Import historical state only from trusted backups.
- Change DNS with a low TTL, monitor, and keep rollback records.
Improve the next recovery
Store application source outside the compute account, encrypt scheduled database backups to a second provider, document DNS ownership, and test a restore regularly. A recovery plan is complete only when another machine can reproduce the service without relying on memory or a single vendor console.
How this article was produced
AI tools may assist with research organization, drafting, or editing. Indra Wijaya reviews each article, checks primary sources, verifies technical claims where practical, and remains responsible for the published result.