Flaky tests: from 0.5% to 98%
If each of 800 tests fails for no reason 0.5% of the time, the suite is red on almost every run. Why retry isn't a cure, and what catches flakiness?
98%
The bigger the suite, the less reliable
Chance that any test fails on a clean run
A single test that fails for no reason 0.5% of the time is invisible. Put 800 of them together: 1 − 0.995^800 ≈ 98%.
Probabilities compound
False red as the test count grows
at least one test failing for no reason, %
Computed from the formula: 1 − 0.995^n. Not a measurement; the post only gives ≈ 98% for n = 800.
Failures aren't independent either
50 ms locally, a timeout under load.
CI runs on shared, throttled hardware; an async test that finishes in 50 ms on your laptop falls into a timeout under load.
AI didn't invent flakiness, it mass-produced it
Four patterns that look right
- A fixed sleepTo "wait" for async work: time.Sleep(2 * time.Second), await page.waitForTimeout(2000), cy.wait(2000).
- Unseeded randomnessA different value on every run: Math.random(), rand.Intn(100), a faker with no fixed seed.
- A brittle selectorNailed to the DOM structure: nth-child, an auto-generated class name. Breaks when the markup shifts a little.
- A mock for everythingThe test verifies an interaction, not an outcome.
Retry is a treadmill, not a cure
All of it is reactive
- RetryBurns real CI minutes and stretches the feedback loop; you pay twice to learn nothing new.
- QuarantineA silently growing pile of disabled tests. That pile is debt, and it compounds.
- TicketAges in backlog noise and quietly stops mattering.
Catch the anti-pattern, not the failure
When does it step in?
| Retry, quarantine, ticket | Static scan at the commit boundary | |
|---|---|---|
| Timing | After the test has already broken CI | Before the test even runs, in the diff |
| Effect | Hides the symptom, sends you the bill | Flags the known anti-patterns |
| Repeatability | Passes on rerun, then fails again | Same diff, same result; no model, no network |
| Scope | Doesn't make the test deterministic | Stops obvious flakiness before it ever reaches CI |

Detection is the safety net, design is the real fix
Write the test deterministically
- Replace fixed waits with condition-based waits. Wait for a fact, not for the clock. (pending)
- Seed every source of randomness; inject a deterministic clock and ID. (pending)
- Select by role or test id, not DOM position. (pending)
- Mock the boundary, not the logic; keep at least one real integration path. (pending)
When does this pattern stop being the right answer?
Limits of the static gate
- Won't catch a clever raceIt catches the obvious anti-patterns, not a race condition buried inside your own code.
- Keep the rules conservativeA noisy gate gets ignored; an ignored gate is worse than none.
- If the system is flakyA test that is flaky because the system under test isn't deterministic is telling you something about the system.
A green CI you have to rerun isn't green.
The cheapest place to kill a flaky test is the line it's written on: before it costs you a build.

On the line it's written
Source: sade.dev, Flaky Tests: Retrying Isn't a Fix