The number went up and nothing else did
A team I worked with rolled out an 80% line coverage mandate in Q2. By mid-Q3, repo-wide coverage had climbed from 64% to 81%. Leadership presented the chart. Everyone clapped.
Then we looked at the other charts. Escaped defects: flat. Change failure rate: flat. Time spent babysitting CI: up, because the suite had grown by 2,400 tests and a noticeable fraction of the new ones failed intermittently. The mandate had produced exactly what it measured, which is lines of code executed during the test run, and almost nothing it was supposed to proxy for.
This is not an unusual story. It is the default outcome of a code coverage mandate, and it is worth understanding mechanically, because the failure mode is not laziness or malicious compliance. It is rational engineers responding to the incentive you handed them.
What a coverage number actually measures
Line coverage answers one question: did this line execute while the test process was running? That's it. It does not ask whether any assertion inspected the result. It does not ask whether the test would fail if the line were wrong. A test suite that executes 85% of your code and verifies none of it scores better than a suite that rigorously verifies the 60% of your code where the money lives.
Coverage is genuinely useful as a negative signal. Zero coverage on a payment module tells you something real: nothing is watching that code. But the inverse doesn't hold. High coverage tells you the code ran, not that anyone checked the answer. The moment you attach a gate to the number, Goodhart's law kicks in and people start manufacturing the cheap kind of coverage, because the gate cannot tell the difference.
Under deadline pressure, with a PR blocked at 78.6% against an 80% gate, here is what actually gets written.
Species one: the assertion-free test
The fastest way to turn red coverage green is to execute code without checking anything:
// user-service.test.js
it('creates a user', async () => {
await createUser({ email: 'test@example.com', plan: 'pro' });
});
it('handles the webhook payload', () => {
const result = parseWebhook(fixturePayload);
expect(result).toBeDefined();
});
The first test has no assertions at all. The second has one that is almost impossible to fail: parseWebhook would have to return undefined without throwing. Both tests light up dozens of lines in the coverage report. Neither would catch a bug where the user gets created on the wrong plan or the webhook amounts are parsed in cents instead of dollars.
These tests aren't free, either. They run on every PR, consume CI minutes, and show up as green checkmarks that reviewers reasonably interpret as "this behavior is verified." They are worse than no test, because no test is at least honest about the gap.
Species two: the snapshot dragnet
The second-fastest way to manufacture coverage is a snapshot test over a big object:
it('builds the monthly report', () => {
const report = buildReport(orders, { month: '2025-06' });
expect(report).toMatchSnapshot();
});
One test, one line of assertion, three hundred lines of serialized output in the snapshot file, and every branch of buildReport now counts as covered. On paper this looks like a thorough test. In practice it is a tripwire attached to everything and aimed at nothing.
When the snapshot fails, nobody reads the 300-line diff. They run the update command, commit the new snapshot, and move on. The test has degraded into a change detector that gets silenced on contact. Reviewers rubber-stamp snapshot updates because reviewing them properly would take longer than the feature did.
And here is where it gets expensive: large snapshots are a flakiness factory. If buildReport embeds generatedAt: new Date(), the snapshot fails depending on when CI runs. If it serializes a Map, an object built from a parallel fetch, or anything that came out of a database query without an explicit sort, the field order shifts between runs and the snapshot fails nondeterministically. We've written about the ORDER BY you never wrote and tests that only fail at midnight; snapshot dragnets are the most reliable way to mass-produce both.
So the coverage mandate didn't just buy you weak tests. It bought you flaky tests, which actively corrode trust in the suite. Once developers learn that a red build is probably a snapshot timestamp, they stop investigating red builds. That's the real cost, and it compounds: we've put numbers on it in a cost model you can plug your own headcount into.
Species three: the mock mirror
The third species mocks every collaborator and then asserts that the mocks were called:
it('processes the order', async () => {
const charge = jest.fn().mockResolvedValue({ ok: true });
const notify = jest.fn();
await processOrder(order, { charge, notify });
expect(charge).toHaveBeenCalled();
expect(notify).toHaveBeenCalled();
});
This verifies that the implementation is shaped like the implementation. It will pass if the charge amount is wrong, if the currency is wrong, if notify fires before the charge settles. It will fail the moment someone refactors the internals, even when behavior is preserved. High coverage, zero behavioral guarantee, maximum friction on future change. The mandate smiles on it anyway.
If you're in a regulated shop, this is worse than useless
For SOC2 and ISO 27001 teams, the coverage gate is often written into change-management controls as evidence that changes are tested before release. An auditor samples a PR, sees the gate passed at 83%, and checks the box.
But if a meaningful slice of that 83% is assertion-free tests and auto-updated snapshots, the control exists on paper and not in practice. You are generating evidence of a verification process that isn't verifying anything. That is a materially worse position than having no gate, because you've now documented a control you can't honestly stand behind. We've covered the adjacent problem in flaky tests are a SOC2 problem: a gate that gets rerun or gamed until green is not a control, it's a ritual.
What to set policy on instead
None of this means "stop measuring coverage." It means stop gating on a single repo-wide number and start gating on things that resist gaming. Four policies that hold up:
1. Gate on diff coverage, not repo-wide coverage. The repo-wide number is dominated by history nobody is touching this sprint. What you actually care about is whether new and changed lines ship with tests. A diff coverage gate asks exactly that, and it's hard to game with filler tests because the denominator is small and the reviewer can see the whole thing:
- name: Diff coverage gate
run: |
pip install diff-cover
diff-cover coverage.xml \
--compare-branch origin/main \
--fail-under 85
A PR touching 40 lines with 34 covered is a conversation a reviewer can actually have. "The repo is at 79.4%" is not.
2. Ratchet the global number; never mandate a target. If you want the repo-wide figure to improve, set the threshold to today's actual value and only allow it to rise:
// jest.config.js
module.exports = {
coverageThreshold: {
global: {
// Current reality. Bump when it improves. Never lower.
lines: 64.2,
branches: 51.8,
},
},
};
A ratchet removes the deadline-pressure scramble that breeds filler tests. Nobody needs to manufacture 16 points of coverage by Friday; they just can't make it worse. Progress comes from the diff gate doing its job on every PR. We go deeper on threshold design in our code coverage policy guide.
3. Spot-check test quality with mutation testing. Mutation testing flips the question from "did the code run" to "would the tests notice if the code were wrong." It mutates your source (inverts a conditional, swaps an operator) and checks whether any test fails. Assertion-free tests and rubber-stamped snapshots get exposed immediately, because mutants survive. Running Stryker or mutmut on every PR is too slow for most suites, so don't. Run it weekly against your two or three highest-risk modules and treat a low mutation score as a review finding, not a build failure. One afternoon of reading surviving mutants will tell you more about test coverage quality than a quarter of coverage dashboards.
4. Make red mean something before you gate on anything. A coverage gate is one more reason for CI to go red. If your suite already fails intermittently, adding gates just adds noise to a channel nobody trusts, and the snapshot flakes your old mandate created are part of the problem. Detect flaky tests, quarantine them so they stop blocking merges, and fix them on a tracked queue. This is the problem BuildPulse was built for, and it pairs directly with coverage: a diff coverage gate is only a control if a failure reliably indicates a real gap, not a Date.now() in a snapshot.
The policy conversation to have instead
If you currently have an 80% mandate, here's the honest exercise: pull the fifty most recently added tests and count how many have an assertion that could plausibly fail on a real bug. If the answer embarrasses you, the mandate is working exactly as designed, and the design is the problem.
The goal was never the number. The goal is that when CI is green you ship without flinching, and when it's red someone drops what they're doing. A diff coverage gate, a ratchet, periodic mutation spot checks, and a flaky-test quarantine get you there. A repo-wide 80% threshold gets you a chart for the board deck and a test suite that is slowly learning to lie to you.
Pick the version where the tests are on your side.



