CI Failure Triage: Code Bug or Broken Upstream?
A repeatable method for distinguishing a real regression from permissions, dependency outages, flaky runners, and repository-level workflow gates.
A red check is evidence that a workflow did not complete successfully. It is not automatically evidence that the patch is wrong. Effective triage starts by identifying where the failure occurred and which party controls that layer.
Classify the failing layer
- Source: compilation, type errors, assertions, or deterministic behavior caused by the patch.
- Test environment: missing runtime, incompatible dependency, exhausted disk, or unavailable service.
- CI platform: cancelled runner, internal error, queue timeout, or artifact outage.
- Repository policy: workflow approval, branch protection, admission gate, or maintainer-only secret.
- External upstream: package registry, API, container registry, DNS, or rate-limit failure.
Collect the first decisive evidence
Open the failed job, identify the first command that returned non-zero, and preserve the surrounding log. Later stack traces are often cascading effects. Record the workflow run, job, step, commit SHA, runner image, and timestamp. If the workflow never started, the problem is admission or platform state, not test execution.
Reproduce the exact gate
Use the repository's documented runtime and exact command. A local substitute such as a linter cannot prove that a failing integration suite is healthy. Conversely, a reviewer running inside a read-only sandbox should not reinterpret its own cache permission error as failure of a test that already passed in the writable runner.
Compare against the base branch
If the same job fails on the base branch or an unrelated commit with an identical signature, the patch is less likely to be the cause. This is not permission to ignore the result; it is a reason to document the upstream condition and avoid changing correct code to satisfy a broken environment.
Retry only when the failure is retryable
Network timeouts, cancelled runners, transient 5xx responses, and approved workflow waits can justify a retry. Assertion failures, compiler errors, and security scanners with stable findings require a code or configuration change. Blindly rerunning deterministic failures wastes quota and hides the real defect.
Write a maintainer-facing report
A useful report says what command ran, where it failed, whether the base branch reproduces it, what changed locally, and which evidence remains unavailable. Avoid claiming “all tests pass” when only a focused test ran. Precision makes it easier for a maintainer to choose between rerun, repair, or infrastructure escalation.
Automate the classification, not the conclusion
A control plane can recognize exit codes, platform messages, rate limits, and workflow states. It should still preserve logs and uncertainty. The goal is to route work correctly: repair code, wait for infrastructure, request workflow approval, or ask a human for a repository decision.
How this article was produced
AI tools may assist with research organization, drafting, or editing. Indra Wijaya reviews each article, checks primary sources, verifies technical claims where practical, and remains responsible for the published result.