Flaky Tests
7 min read

Flaky test cost: a model you can plug your own headcount into

A 5% CI failure rate from flaky tests looks like a rounding error. Priced out across compute, attention, delays, and trust, it's a headcount-sized budget line.

BuildPulse Team

September 28, 2026

Listen

Flaky test cost: a model for engineering leaders | BuildPulse Blog

The line item that isn't in your budget

Somewhere in your finance system there's a line for CI compute. There is no line for "engineers babysitting red builds," and that's the expensive one.

Here's the scene that prompted this post. A director I know was asked, mid budget review, why CI spend was up 30% year over year while headcount was up 12%. She pulled the data and found the answer in about an hour: roughly one in twenty CI runs failed on tests that passed on rerun, and every one of those failures triggered a full pipeline rerun. The compute was the visible symptom. The invisible cost, once she modeled it, was about eight times larger.

A 5% test failure rate sounds small. Ninety-five percent is an A. But CI isn't graded on a curve, and 5% of every run, every push, every merge, compounds into something that deserves its own budget line. This post is that budget line: four cost buckets, real math, and a model you can rerun with your own numbers.

What 5% actually means

First, a distinction that trips up almost every leadership conversation about this: a 5% run failure rate does not mean your tests are 5% flaky. It means they're barely flaky at all, and it still doesn't matter.

The probability that a run fails is one minus the probability that every test passes. With 2,000 tests in a suite, a per-test flake rate of just 0.0026% produces a 5% run-level failure rate. Push per-test flakiness to a still-tiny 0.05% and the math turns grim: 1 - (1 - 0.0005)^2000 ≈ 63% of runs fail. Nearly two out of three.

This is why "our tests are 99.99% reliable" and "CI is red constantly" are both true at the same company. Individual test reliability compounds across the suite, and the suite runs on every push. If your engineers push four times per PR, a 5% per-run failure rate means roughly 18% of PRs eat at least one flaky failure before merging. At 63% per run, essentially every PR does.

So when we price out "5%," understand that 5% is the good scenario. Plenty of teams I've talked to are living at 15–30% and have simply stopped noticing, the way you stop noticing a smell.

Line item 1: rerun compute, the cheap part

Let's build the model on a concrete org: 250 engineers, each merging about two PRs a week. That's roughly 100 merged PRs per workday. With pushes, updates, and the merge itself, call it four CI runs per PR: 400 runs a day.

At a 5% flake-failure rate, that's 20 failed runs a day. The default remediation everywhere is "rerun the whole thing," so each failure costs you one full pipeline. If your pipeline burns 200 machine-minutes per run (a 25-minute wall clock across eight parallel jobs is typical at this size), that's 4,000 wasted CI minutes a day.

On GitHub-hosted Linux runners at $0.008/minute, that's $32 a day, or about $8,000 a year. We benchmarked rerun waste across real projects and the pattern holds: the raw compute is real money, but it's the smallest number in this post by an order of magnitude.

Which is exactly why flaky tests survive budget season. The only cost that shows up on an invoice is the one that doesn't matter.

Line item 2: engineer attention

A flaky failure is never just a rerun click. Here's the actual sequence: CI goes red, the engineer gets a notification, they context-switch out of whatever they were doing, open the logs, spend a few minutes deciding whether this failure is real (this step is the whole problem; if they could tell instantly, flakiness would be an annoyance instead of a tax), conclude it's probably flaky, click rerun, and then either sit in limbo waiting or switch to something else and get interrupted again 25 minutes later when the rerun finishes.

Research on interrupted work puts the cost of resuming a focused task after an interruption at 20+ minutes. Be conservative and call the total disruption 30 minutes per incident: five minutes of triage, the rerun-and-recheck loop, and the resumption cost.

  • 20 flaky failures a day × 30 minutes = 10 engineer-hours a day
  • At a loaded cost of $110/hour, that's $1,100 a day
  • Across 250 workdays: roughly $275,000 a year

That is more than a senior engineer's fully loaded cost, spent entirely on distinguishing fake failures from real ones. And this assumes every triage reaches the right conclusion in five minutes. It doesn't. Some engineers will spend 45 minutes debugging a failure that was never real, and we covered why retries hide rather than eliminate this cost: auto-retry converts visible triage time into invisible compute and slower feedback, it doesn't refund the attention.

Line item 3: delayed merges and releases

Each flaky failure adds 30–60 minutes to the affected PR's merge time: the rerun itself plus whatever queue time and human latency surrounds it. Annoying for one PR. Structural for the org, because delay compounds in two ways.

First, merge queues amplify it. If you run a merge queue, one flaky failure doesn't cost one PR its slot; it can invalidate every PR batched behind it, forcing the queue to rebuild and re-test the lot. A 5% failure rate against a queue of five means most batches contain at least one flake.

Second, delay changes behavior. When merging takes half a day of ambient friction, engineers batch their work into bigger PRs to amortize the pain. Bigger PRs are harder to review and riskier to deploy, which pushes up your change failure rate, which triggers more process, which slows delivery further. I've watched this loop run at three different companies. Nobody ever traces it back to the test suite, because the test suite failure was six causal steps upstream.

Pricing this bucket precisely requires your own cycle-time data, but a defensible floor is the direct delay: 20 incidents a day × 45 minutes of blocked PR time is 15 hours of daily pipeline latency injected into your delivery flow, before any queue amplification.

Line item 4: the trust discount

This is the bucket that doesn't fit in a spreadsheet and dominates the other three.

When 5% of red builds are false alarms, engineers develop the only rational response: assume red is fake until proven otherwise. Rerun first, investigate never. The moment that habit forms, your CI gate stops being a gate. Real regressions get rerun past. Real regressions get merged. You find out in production, and your change failure rate quietly absorbs the damage while your dashboards insist the tests are passing.

One production incident that a trusted CI signal would have caught typically costs more than the entire compute line item for the year. Most teams running at 5%+ eat several of those annually and never connect them to test flakiness, because by the time the incident review happens, the test that would have caught it did fail, someone reran it, and it passed.

If you operate under SOC2 or similar change-management controls, there's a second-order cost: your CI gate is a documented control, and "we reran until green" is a genuinely uncomfortable sentence in an audit. Flaky tests are a compliance problem, not just an engineering one, and the remediation cost of an auditor pulling that thread lands on your calendar, not your CFO's.

The model, in code you can rerun

Here's the whole thing as a script. Change the first block to your numbers and take the output to your next planning cycle.

# Flaky test cost model. Adjust the inputs, keep the structure.

engineers = 250
prs_per_engineer_per_week = 2
ci_runs_per_pr = 4
run_failure_rate = 0.05          # fraction of runs failing on flaky tests
machine_minutes_per_run = 200    # sum across parallel jobs
cost_per_ci_minute = 0.008       # GitHub-hosted Linux
disruption_minutes = 30          # triage + rerun loop + refocus, per incident
loaded_hourly_cost = 110
workdays = 250

runs_per_day = engineers * prs_per_engineer_per_week / 5 * ci_runs_per_pr
flaky_failures_per_day = runs_per_day * run_failure_rate

compute_annual = (flaky_failures_per_day * machine_minutes_per_run
                  * cost_per_ci_minute * workdays)

attention_annual = (flaky_failures_per_day * disruption_minutes / 60
                    * loaded_hourly_cost * workdays)

delay_hours_daily = flaky_failures_per_day * 0.75  # blocked PR time

print(f"Flaky failures/day:      {flaky_failures_per_day:.0f}")
print(f"Rerun compute/year:      ${compute_annual:,.0f}")
print(f"Engineer attention/year: ${attention_annual:,.0f}")
print(f"Pipeline delay/day:      {delay_hours_daily:.0f} hours")
Flaky failures/day:      20
Rerun compute/year:      $8,000
Engineer attention/year: $275,000
Pipeline delay/day:      15 hours

Add the trust bucket however your org prices incident risk, but even leaving it at zero, this 250-engineer org is spending roughly $283,000 a year to maintain a 5% failure rate. Not to fix it. To have it.

Run the same model at a 15% failure rate, which is where many teams actually sit, and the attention line alone clears $800,000.

What to do with the number

The point of the model isn't precision. It's that flaky tests currently compete for engineering time against features, and features have revenue attached while flakiness has vibes attached. A dollar figure changes the negotiation.

Three moves, in order:

  • Measure your actual rate. Most teams guess low. You need per-test failure history across runs, not anecdotes. This is the core of what BuildPulse does: it ingests your test results and tells you which tests are flaky, how often, and what they're costing you in reruns, so the model above runs on your data instead of my defaults.
  • Quarantine before you fix. A known-flaky test running in the merge-blocking path is paying all four cost buckets daily. Pull it out of the gate, keep it running and reported, and your failure rate drops immediately while the fix waits its turn in the backlog like everything else.
  • Fix in cost order, not annoyance order. The test that fails 4% of the time in a suite that runs 400 times a day outranks the one that fails 30% of the time in a nightly job. Frequency times blast radius, not loudness.

A 5% failure rate feels like the cost of doing business right up until you price it. Then it looks like what it is: a headcount-sized line item you've been paying without ever approving it.

Stop guessing which tests you can trust

BuildPulse finds your flaky tests, ranks them by the engineering time they cost, and lets you quarantine the worst in one click. See results on your first build.

Free to start · No credit card required · Setup is a single CI step