CI success rates have been stable for 30 minutes and the backlog of queued jobs has cleared. This incident is resolved. A brief retro will be linked below once available.
Elevated failure rates for GitHub-hosted CI runs
- Affected
- GitHub Cloud, Developer Portal
- Started
- Resolved
- Duration
- 1h 38m
Updates
- Resolved
- Monitoring
GitHub has applied a fix and success rates have returned to normal. We are monitoring queued and retried jobs to confirm full recovery before closing.
- Identified
The issue has been identified as an upstream GitHub incident affecting runner image pulls in
us-east. Impact is limited to:- Workflows using
ubuntu-latesthosted runners - Backstage software templates that kick off CI on creation
Self-hosted runners are not affected. Teams needing an urgent build can re-run on a self-hosted runner label as a workaround.
- Workflows using
- Investigating
We are investigating elevated failure rates on GitHub-hosted Actions runners. Pipeline jobs that pull container images are timing out. Backstage scaffolder templates that trigger CI are also affected.
Summary
Summary
A roughly 90-minute degradation of GitHub-hosted CI runners caused intermittent job failures, primarily for workflows pulling container images. The root cause was an upstream GitHub incident. No data was lost and self-hosted runners were unaffected.
Follow-ups
- Document the self-hosted runner fallback in the platform runbook.
- Add a dashboard alert for CI success-rate drops below 95%.