The bill you optimized and the bill you actually pay
I watched a platform team spend a full quarter on CI cost optimization. They right-sized runner instances, moved heavy jobs to spot-backed machines, rebuilt their Docker layer caching, and shaved the median workflow from 31 minutes to 19. Genuinely good work. The GitHub Actions bill dropped 22%.
Then I asked them one question: what percentage of your workflow runs have a run_attempt greater than 1?
Nobody knew. We pulled the data. It was 11%. And because their retry culture was "re-run all jobs" rather than "re-run failed jobs", most of those retries re-billed the entire matrix: unit tests, integration tests, the browser suite, all of it. The retry line item alone was worth more than half of the savings they'd just spent a quarter earning.
That's the thing about retried jobs. They don't show up as a line item anywhere. GitHub's billing page tells you minutes by runner type, not minutes by attempt number. So the waste hides inside numbers that look legitimate, and every cost review walks right past it.
Where retry minutes come from
Retries happen for three reasons, and only one of them is defensible:
- Infrastructure flakiness. Runner evictions, registry timeouts, npm having a bad afternoon. Legitimate retries, and automatic retry-with-backoff is the correct tool.
- Flaky tests. The test failed, the code is fine, everyone knows the drill. Click re-run, go get coffee. This is the bulk of retry spend at most companies I've looked at, and it's pure waste.
- Habitual retries. The failure might be real, but the developer has been trained by months of flaky failures to retry first and read logs never. This is the most corrosive category because it means a real regression gets a free second chance to sneak through on a lucky rerun.
Notice that two of the three are downstream of the same root cause: a CI signal nobody trusts. Once your developers believe red might mean nothing, every red becomes a retry, and every retry becomes billable minutes. If you're in a SOC2 or change-management environment, it's also worse than a cost problem: your required CI gate is now a control that people routinely bypass by rolling the dice again, which is an awkward thing to explain in an audit.
Measure it: your retry rate in one script
You don't need a vendor to see this number. The GitHub API exposes run_attempt on every workflow run, and a run whose latest attempt is greater than 1 was retried at least once. Thirty days of data for a repo:
#!/usr/bin/env bash
REPO="acme/monolith"
SINCE=$(date -u -d '30 days ago' +%Y-%m-%d 2>/dev/null || date -u -v-30d +%Y-%m-%d)
gh api "repos/$REPO/actions/runs?created=>=$SINCE" --paginate \
--jq '.workflow_runs[] | [.run_attempt, .name] | @tsv' \
| awk -F'\t' '
{ total++; if ($1 > 1) { retried++; by_wf[$2]++ } }
END {
printf "runs: %d retried: %d rate: %.1f%%\n", total, retried, 100*retried/total
for (wf in by_wf) printf " %-40s %d\n", wf, by_wf[wf]
}'
Sample output from a real-ish monorepo:
runs: 14210 retried: 1278 rate: 9.0%
CI / integration-tests 611
CI / e2e-playwright 402
CI / unit-tests 148
Deploy / staging 117
Two things jump out in almost every org I've run this against. First, the rate is higher than anyone guessed. Second, retries cluster in the test-heavy workflows, not the build or deploy ones. That's your tell that this is a flaky-test problem wearing an infrastructure costume. If you want billable minutes per attempt rather than run counts, the /actions/runs/{run_id}/timing endpoint will give you the exact figure, but honestly the run-count version is usually enough to start an uncomfortable conversation.
The math: cheaper runners vs. fewer reruns
Let's put real numbers on a mid-sized org. Say 300 engineers, 1,500 workflow runs per weekday, an average of 24 billable runner-minutes per run across all jobs, on GitHub-hosted Linux runners at $0.008 per minute.
Baseline monthly spend, ignoring retries:
1,500 runs/day x 22 days x 24 min x $0.008 = $6,336/month
Now layer in a 9% retry rate. Here's the part people underestimate: a retry is rarely a surgical re-run of one job. Developers click "Re-run all jobs" about half the time (or the failed job is a needs: dependency and drags its downstream jobs with it), so the average retry re-bills roughly 70% of the original run's minutes:
1,500 x 22 x 9% x (24 min x 0.7) x $0.008 = $399/month... per retry attempt
Except retries themselves fail and get retried, because the flaky test is still flaky. With a realistic 25% re-retry rate, you're at roughly $530/month in pure retry minutes. On a $6,336 baseline, that's 8% of your entire Actions bill spent re-running work that already ran.
Now compare two interventions:
- Option A: cut runner unit cost in half. Move to self-hosted or a managed runner provider at roughly half the per-minute price. Savings on the baseline: about $3,170/month. Savings on retry waste: about $265/month, because cheap waste is still waste. Total: ~$3,435/month.
- Option B: cut the retry rate from 9% to 2% by finding and quarantining the flaky tests driving it. Savings: about $410/month in direct compute. Sounds small next to Option A, until you price the second-order effects.
On raw compute, Option A wins, and you should absolutely do it. Faster, cheaper runners are the closest thing to free money in CI (it's why BuildPulse builds GitHub Actions runners alongside flaky-test detection: half the cost per minute, roughly twice the speed). But stopping at Option A is how teams end up with a cheap bill for an untrustworthy pipeline. The retry minutes were never the expensive part.
The costs that don't show up on the Actions invoice
Every retry carries three costs beyond the billable minutes:
Queue contention. Retried jobs compete for the same concurrency slots as first-attempt jobs. At a 9% retry rate with fan-out, you're running the equivalent of 6–8% more CI volume at peak hours, which means everyone's first-attempt jobs queue longer. Teams respond by buying more concurrency. You are now paying extra capacity to host your own waste.
Merge latency. A retry isn't instant. The developer notices the red check (eventually), clicks re-run, and waits another full cycle. On a 24-minute pipeline, each retry adds 30–60 minutes of wall-clock time to a PR once you count the noticing. At a 9% retry rate across 1,500 daily runs, that's dozens of engineer-hours per day spent babysitting checks. Price that at loaded cost and it dwarfs the compute line by an order of magnitude.
Signal decay. This is the one that should worry a platform lead most. Every successful retry teaches your org that red means "try again", not "something broke". Eventually someone retries past a real failure, it merges, and you're doing incident review on a bug your CI caught and your culture waved through. If your CI gate is part of a documented change-management control, "we re-ran it until it passed" is not a sentence you want in the timeline.
Fix it at the source
The playbook for driving retry rate down is not mysterious, it's just unglamorous:
1. Split infrastructure retries from test retries. Automatic retry is fine for genuinely transient failures. Scope it to the steps that talk to networks, not the test run:
- name: Pull base images
uses: nick-fields/retry@v3
with:
max_attempts: 3
retry_wait_seconds: 20
timeout_minutes: 10
command: docker compose pull
- name: Run tests
run: go test ./... # no retry wrapper here, on purpose
Wrapping the test step itself in a retry doesn't fix anything. It converts a visible reliability problem into an invisible billing problem and launders flaky failures into green checkmarks.
2. Find the tests actually driving reruns. Retries cluster. In most suites, fewer than 2% of tests cause more than 80% of retry-driven failures. You need per-test, per-attempt history to find them, which is exactly what BuildPulse flaky-test detection builds from your JUnit output: which tests fail intermittently, how often, and what each one costs you in reruns. You can approximate it yourself by archiving test reports and diffing failures across attempts, and that's a fine way to start.
3. Quarantine, then fix. Pull the worst offenders out of the merge-blocking path immediately (keep running them, stop letting them gate) and work the fix queue by cost, not by vibes. The test that fails 0.5% of the time in a 40-minute required workflow is more expensive than the one that fails 5% of the time in a 3-minute optional one. We've written before about how to actually fix flaky tests once you've found them; quarantine is what buys you the time to do it properly.
4. Make retry rate a first-class metric. Put the script above (or your platform's equivalent) on a dashboard next to cost per merged PR. Review it monthly like you review cloud spend. What gets graphed gets fixed.
Do both, in the right order
If I were handed this problem on Monday, the sequence is: measure retry rate this week, quarantine the top ten flaky tests next week, and run the runner-cost migration in parallel because the two don't conflict. Cheaper, faster runners cut the price of every minute, including the wasted ones. Eliminating reruns cuts the wasted minutes themselves, plus the queue contention, the merge latency, and the slow erosion of trust in your own gate.
The teams that only do the runner migration get a cheaper bill and the same broken feedback loop. The teams that do both get a bill that's smaller than either fix alone predicts, because the savings multiply: fewer minutes times a lower price per minute.
Run the script. If your retry rate comes back under 3%, congratulations, go optimize something else. If it comes back at 9%, you've just found the easiest money in your CI budget, and it was hiding in plain sight the whole time.



