The number on the slide is wrong
Quarterly business review. Slide 14 says change failure rate: 9%, trending down. The VP presents it with the quiet confidence of someone who read the DORA report and picked the elite-performer benchmark as the target. Meanwhile, two floors down (or two Slack channels over), the platform team knows the post-deploy smoke suite fails roughly one run in six for reasons that have nothing to do with the code being deployed. Half those failures trigger an automatic rollback. The other half get rerun until green.
Both halves are lying to slide 14. In opposite directions.
Change failure rate is the DORA metric executives quote most and interrogate least. It sounds like an outcome metric, immune to gaming: either the deploy broke production or it didn't. But CFR is computed from signals, and in most organizations the loudest signal in the chain is a test suite. If that suite is flaky, your CFR is not a measurement. It's a vibe with a decimal point.
What CFR is supposed to measure
The definition is simple enough: the percentage of deployments to production that result in degraded service requiring remediation, where remediation means a hotfix, a rollback, a fix-forward patch, or an incident. I've written before about how almost nobody computes the base metric correctly, so I'll skip the definitional archaeology here and focus on the part that matters for this post: every practical CFR implementation depends on classifying deployments as failed or not, and that classification almost always runs through automated tests somewhere.
Three common places tests enter the CFR pipeline:
- Post-deploy smoke or canary suites that gate promotion and trigger rollback on failure.
- Revert and hotfix detection, where a revert PR or a
hotfix/branch merged shortly after a deploy marks that deploy as failed. - Incident correlation, where an alert or incident within some window after a deploy attributes failure to that deploy.
Flaky tests contaminate all three. The first one directly, the second two through the behavior flakiness trains into your engineers.
Inflation: the rollbacks that never were
Start with the direct path. Your deploy pipeline runs a smoke suite against the canary. A test that depends on an eventually-consistent read, or a hardcoded timeout, or a shared database row, fails 4% of the time regardless of what shipped. Your pipeline does the responsible thing and rolls back.
That rollback is now a remediation event. Your CFR tooling, which counts rollbacks in the numerator, dutifully records a change failure. Nothing failed. No user saw anything. But the metric ticked up, and six weeks later someone in a leadership meeting asks why team Atlas has a 14% change failure rate when team Borealis is at 5%. The honest answer might be that Atlas owns the service with the flakiest smoke tests and the most aggressive auto-rollback policy. That answer never makes it into the meeting.
Do the arithmetic on how fast this compounds. Say you deploy 200 times a quarter and your true change failure rate is 6%: 12 real failures. Your smoke suite has a 4% flake-triggered rollback rate: 8 phantom failures. Your dashboard reads 10%. You are reporting a number that is two-thirds noise on the margin, and the noise is not evenly distributed across teams, so every cross-team comparison you make with it is unfair to whoever inherited the worst tests.
Worse: the teams being punished by phantom rollbacks respond rationally. They loosen the gate.
Deflation: rerun-to-green eats real failures
The loosened gate is the second failure mode, and it's the more dangerous one because it pushes CFR down while pushing actual failures up.
Once a team has been burned by enough phantom rollbacks, the auto-rollback gets replaced with a retry step, or a human with merge rights and a deadline. Red smoke run? Rerun it. Green on attempt two? Ship it. This feels pragmatic, and I've written about what retry culture actually costs, but here's the CFR-specific damage: a rerun-until-green gate cannot distinguish a flaky failure from a real regression that fails intermittently. Race conditions, resource leaks, and load-dependent bugs all present as "failed once, passed on retry." The gate waves them through.
The regression ships. It surfaces two days later as elevated error rates, gets triaged as an ops issue, and gets fixed in a routine PR that nobody labels as a hotfix. No rollback, no incident tagged to the deploy, no revert. Your CFR numerator never hears about it. The deploy that shipped the bug is recorded as a success.
So flaky tests inflate CFR through phantom rollbacks and deflate it through masked regressions, often in the same org, often in the same quarter. The two errors do not cancel out. They just mean the number on the dashboard has an unknown sign of bias, which is a fancy way of saying it's useless for decisions.
If you're audited, this is worse than useless
For teams in SOC2, ISO 27001, or regulated environments, CFR is often more than a dashboard: it's evidence that your change-management controls work. Your auditors see a deploy gate, a test suite, and a low failure rate, and reasonably conclude that changes are verified before release.
Rerun-to-green quietly breaks that story. The control on paper says "tests must pass before promotion." The control in practice says "tests must pass eventually, on some attempt, after a human decided the failure didn't count." That gap between documented and actual control behavior is exactly what makes flaky tests a change-management problem rather than an annoyance. If your CFR is part of the evidence package, you want to be able to explain how it's computed and defend every exclusion. "We rerun failures we believe are flaky" is not a defensible exclusion policy. "Failures matching tests quarantined by our flake-detection system are excluded, and here's the audit log" is.
Cleaning the signal before it reaches the dashboard
You cannot fix CFR by adjusting CFR. You fix it upstream, in roughly this order.
1. Measure your flake rate first
Before you trust any correction, you need to know how big the contamination is. Track, per pipeline stage: how often the same commit produces different test outcomes across runs. That's your flake rate, and it's the error bar on every downstream metric. If your smoke suite flakes 5% of the time and your reported CFR is 8%, you should present CFR as "somewhere between 3% and 8% until we fix the suite," not as 8.0%. Presenting an honest range is uncomfortable exactly once. Presenting a false precision is uncomfortable forever, retroactively.
This is where a flake-detection layer earns its keep: BuildPulse identifies which failures came from known-flaky tests by watching outcomes across runs, which gives you the classification you need to correct the numerator instead of guessing.
2. Define the numerator by remediation, not pipeline color
A change failure should require evidence of production impact plus a remediation action: an incident, a rollback executed because of observed degradation, or a hotfix. A red pipeline stage is not, by itself, a change failure. Encode that distinction in your tooling. If your rollback automation fires, record why: which check failed, and whether that check's failure was corroborated by anything else (error rates, latency, a second independent probe).
3. Tag reverts and rollbacks by cause
Make cause attribution cheap and mandatory. A label taxonomy on revert PRs works fine:
# .github/workflows/revert-labeler.yml
name: Require revert cause
on:
pull_request:
types: [opened, labeled, unlabeled]
jobs:
check:
if: startsWith(github.event.pull_request.title, 'Revert')
runs-on: ubuntu-latest
steps:
- name: Require exactly one cause label
env:
LABELS: ${{ join(github.event.pull_request.labels.*.name, ' ') }}
run: |
count=0
for l in $LABELS; do
case "$l" in
revert:defect|revert:flaky-gate|revert:ops|revert:other) count=$((count+1));;
esac
done
if [ "$count" -ne 1 ]; then
echo "Revert PRs need exactly one revert:* cause label" >&2
exit 1
fi
Now your CFR job can count revert:defect in the numerator and report revert:flaky-gate separately, as its own metric. That second number is arguably more actionable than CFR itself: it's the direct cost of test unreliability, denominated in rollbacks, and it pairs nicely with a cost model your CFO will actually read.
4. Replace rerun-to-green with quarantine
Blanket retries at the gate are how real regressions sneak into production and out of your numerator. The alternative is targeted: detect tests that flake, quarantine them so they run but don't block, and file the fix work. The gate stays strict for every test you actually trust, which means a red gate goes back to meaning something. When the gate means something, rollback events mean something, and CFR starts converging on reality from both directions at once.
5. Recompute history, honestly
Once you have cause tagging and flake classification, re-run the last two or three quarters of CFR with the corrected numerator. Show both curves on the same chart at the next review: reported and corrected. Yes, this means admitting the old number was wrong. That conversation goes better than you'd expect, because "we found systematic measurement error and fixed it" is a story about engineering rigor. "We've been optimizing a noisy number for a year" is the story you tell if you wait for someone else to find it.
The metric was never the point
CFR exists to answer one question: when we ship, does it break? Flaky tests make your CI unable to answer the smaller question underneath it: when the build is red, is something actually broken? You cannot build a trustworthy answer to the first question on top of an untrustworthy answer to the second. Every DORA rollout I've seen skip this dependency ends up in the same place: a dashboard nobody defends in the meeting where it matters.
Fix the flake rate, or at least measure and correct for it, before CFR goes anywhere near an exec deck. The number will move when you do. In which direction is exactly what you're paying to find out.



