18,750 req/s: precise, repeated, and three times wrong
Keyboard: ← → to move, F for full screen, O for overview.

Measurement · Load test
Precise, repeated, wrong
18,750 req/s: three repetitions agreed closely, and the number was three times wrong.
Muhammet Şafak
Same run, same service
Two phases, a threefold gap
| Ladder (open loop) | Tuning phase (closed loop) | |
|---|---|---|
| Go, 4 cores | 18,750 req/s | 55,715 req/s |
| Within itself | Consistent | Consistent |
| In hindsight | Turned out wrong | Turned out right |
Which one is wrong?
Consistency is not accuracy.
Both measurements were consistent within themselves, and nothing said which of them was wrong.
Looking at the data
The repetitions ranged from 1,250 to 48,000.
The median was 18,750, because that is where the middle repetition happened to land.
First diagnosis
It was correct and not enough.
The writeback of the CHECKPOINT issued before every step spilled into the measurement window: at 3,000 req/s, the write p99 was 10.6 ms, the read 3.3 ms.
Second attempt
A fivefold divergence, zero spoiled windows
| Cell | Repetition 1 | Repetition 2 | Repetition 3 |
|---|---|---|---|
| Go, 1 core | 1,250 | 6,500 | 3,500 |
| PHP-FPM, 1 core | 1,750 | 2,750 | 750 |
PHP-FPM, 1 core, 3,000 req/s
oha reports two latencies
| Latency | p50 | p75 | p99 | Slowest |
|---|---|---|---|---|
| Time to first byte | 1.42 ms | 47 ms | 191 ms | 267 ms |
| Latency-corrected | 1.44 ms | 1,539 ms | 3,718 ms | 3,819 ms |
Mechanism
The correction turns the limit into a cliff
- 1Step 1: Counts from the schedule
the moment it should have been sent - 2Step 2: Responses slow down
connections fill up - 3Step 3: The schedule slips
every request inherits it - 4Step 4: Cliff
either passes or misses
What the search could not measure
The cliff moves between repetitions.
Exponential search and bisection assume that pass/fail is monotonic in the rate. Neither was true.
The fix
Drop the search, set fixed rates
- 7Fixed ratesPer cell, the same in every repetition
- VoteMajority ruleIf most repetitions carried it, it is carried
- 115%Probe thresholdBefore every repetition
One more lie
Chromium on the host machine
| Metric | Value |
|---|---|
| Headless Chromium processes | About thirty |
| Total CPU | 1,019% |
| Left for the Docker VM | 94% |
| Go, 4 cores | 17,490 req/s |
Result · Go, 4 cores
Two methods, within a few percent
| Measurement | Req/s | Note |
|---|---|---|
| Tuning phase (closed loop) | 55,715 | Twelve hours earlier |
| Five repetitions (quiet machine) | 55,836–57,321 | Median 56,571; 45/45 repetitions passed the probe |

For your own setup
Three concrete additions
- A second phase that measures in a different loop, and a rule that voids the run when the two diverge (pending)
- A probe before every repetition that runs no application code (pending)
- A record of the environment alongside the repetition (pending)

The measurement itself
The second road
The raw data of the numbers and the discarded attempts is in the research notebook.