Q: How do you troubleshoot, triage, and handle failed workflows in GitHub Actions?
Systematic workflow failure triage: categorizing failures into code, transient infrastructure, or pipeline design defects, using the GitHub CLI (gh) for rapid inspection, re-running failed jobs, and implementing automated retries.
#CI/CD #GitHub Actions #Troubleshooting #gh CLI #Debugging #Retries #Resilience
🎙️ Candidate Opening & Architectural Context
"I treat workflow failures as either code defects, transient infrastructure/external dependency issues, or pipeline design bugs. First I inspect the failed job logs and artifacts using the gh CLI, rerun failed jobs for transient network issues, and correlate with external provider outages. Then I fix the root cause and implement guardrails like step retries, cache optimizations, debug logging, or quarantining flaky tests."
Advertisement
🛠️ Production Runbook & Step-by-Step Resolution
1️⃣
Rapid Triage Using the GitHub CLI (gh)
Diagnose workflow failures quickly without clicking through multiple web UI menus:
# List recent pipeline runs
gh run list --limit 20
# View only the failed log output directly in terminal
gh run view 12345678 --log-failed
# Download diagnostic artifacts
gh run download 12345678 --dir ./artifacts/
# Re-run only the failed jobs in the workflow
gh run rerun 12345678 --failed
- Inspect Failed Logs: Use
gh run view <run-id> --log-failedto output only the exact failing step lines. - Download Artifacts: Retrieve test reports, screenshots, and logs using
gh run download. - Rerun Only Failed Jobs: Avoid re-executing successful 30-minute test jobs by using
gh run rerun --failedto retry only the affected component. - Debug Logging: Enable runner and step diagnostic logging by setting repository secrets
ACTIONS_STEP_DEBUG=trueandACTIONS_RUNNER_DEBUG=true.
2️⃣
Implementing Pipeline Resilience & Automated Retries
Prevent recurring transient failures from blocking development velocity:
# Example automated retry for network-sensitive step
- name: Push container image with exponential retry
uses: nick-fields/retry@v3
with:
timeout_minutes: 5
max_attempts: 3
retry_wait_seconds: 10
command: docker push ghcr.io/org/api:${{ github.sha }}
- Step-Level Retries: Use retry actions (such as
nick-fields/retry) for network-sensitive operations like Docker image pushes or registry downloads. - API Rate Limit Mitigation: Authenticate all external API calls (e.g. GitHub API, Docker Hub) using personal or bot tokens to avoid anonymous IP rate limits.
- Flaky Test Isolation: Quarantine flaky end-to-end tests into a non-blocking job or rerun failed tests once before failing the entire pipeline.
- Notification Webhooks: Configure Slack / PagerDuty alert webhooks triggered only on
if: failure()to notify on-call engineers immediately.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Categorize failures into code vs transient vs configuration bugs. Use the gh CLI for fast log extraction and rerun-failed operations, and add automated step retries for network-sensitive registry and package downloads."
⚡ 60-Second Elevator Pitch Talking Points
- Diagnose pipeline failures rapidly with gh run view --log-failed and re-run only failed jobs.
- Enable ACTIONS_STEP_DEBUG and ACTIONS_RUNNER_DEBUG to uncover hidden runner and network issues.
- Harden pipelines against transient failures with exponential retries and authenticated registry access.
Advertisement