Skip to content
Muhammet Şafak
tr

Flaky tests: from 0.5% to 98%

If each of 800 tests fails for no reason 0.5% of the time, the suite is red on almost every run. Why retry isn't a cure, and what catches flakiness?

98%

The bigger the suite, the less reliable

Chance that any test fails on a clean run

A single test that fails for no reason 0.5% of the time is invisible. Put 800 of them together: 1 − 0.995^800 ≈ 98%.

Probabilities compound

False red as the test count grows

at least one test failing for no reason, %

False red as the test count grows (at least one test failing for no reason, %): 100 tests: 39.4, 200 tests: 63.3, 300 tests: 77.8, 400 tests: 86.5, 500 tests: 91.8, 600 tests: 95.1, 700 tests: 97, 800 tests: 98.2, 900 tests: 98.9, 1000 tests: 99.3.

Computed from the formula: 1 − 0.995^n. Not a measurement; the post only gives ≈ 98% for n = 800.

Failures aren't independent either

50 ms locally, a timeout under load.

CI runs on shared, throttled hardware; an async test that finishes in 50 ms on your laptop falls into a timeout under load.

AI didn't invent flakiness, it mass-produced it

Four patterns that look right

  • A fixed sleepTo "wait" for async work: time.Sleep(2 * time.Second), await page.waitForTimeout(2000), cy.wait(2000).
  • Unseeded randomnessA different value on every run: Math.random(), rand.Intn(100), a faker with no fixed seed.
  • A brittle selectorNailed to the DOM structure: nth-child, an auto-generated class name. Breaks when the markup shifts a little.
  • A mock for everythingThe test verifies an interaction, not an outcome.

Retry is a treadmill, not a cure

All of it is reactive

  • RetryBurns real CI minutes and stretches the feedback loop; you pay twice to learn nothing new.
  • QuarantineA silently growing pile of disabled tests. That pile is debt, and it compounds.
  • TicketAges in backlog noise and quietly stops mattering.

Catch the anti-pattern, not the failure

When does it step in?

Retry, quarantine, ticketStatic scan at the commit boundary
TimingAfter the test has already broken CIBefore the test even runs, in the diff
EffectHides the symptom, sends you the billFlags the known anti-patterns
RepeatabilityPasses on rerun, then fails againSame diff, same result; no model, no network
ScopeDoesn't make the test deterministicStops obvious flakiness before it ever reaches CI
Tan taking notes in a notebook

Detection is the safety net, design is the real fix

Write the test deterministically

  • Replace fixed waits with condition-based waits. Wait for a fact, not for the clock. (pending)
  • Seed every source of randomness; inject a deterministic clock and ID. (pending)
  • Select by role or test id, not DOM position. (pending)
  • Mock the boundary, not the logic; keep at least one real integration path. (pending)

When does this pattern stop being the right answer?

Limits of the static gate

  • Won't catch a clever raceIt catches the obvious anti-patterns, not a race condition buried inside your own code.
  • Keep the rules conservativeA noisy gate gets ignored; an ignored gate is worse than none.
  • If the system is flakyA test that is flaky because the system under test isn't deterministic is telling you something about the system.

A green CI you have to rerun isn't green.

The cheapest place to kill a flaky test is the line it's written on: before it costs you a build.

Tan giving a thumbs-up

On the line it's written

sade.dev

Source: sade.dev, Flaky Tests: Retrying Isn't a Fix

Share and download

Search the site

Start typing to search posts, projects and pages.

Escto closePowered by Pagefind