Skip to content

18,750 req/s: precise, repeated, and three times wrong

Keyboard: ← → to move, F for full screen, O for overview.

Tan thinking, hand on chin

Measurement · Load test

Precise, repeated, wrong

18,750 req/s: three repetitions agreed closely, and the number was three times wrong.

Muhammet Şafak

Same run, same service

Two phases, a threefold gap

Ladder (open loop)Tuning phase (closed loop)
Go, 4 cores18,750 req/s55,715 req/s
Within itselfConsistentConsistent
In hindsightTurned out wrongTurned out right

Which one is wrong?

Consistency is not accuracy.

Both measurements were consistent within themselves, and nothing said which of them was wrong.

Looking at the data

The repetitions ranged from 1,250 to 48,000.

The median was 18,750, because that is where the middle repetition happened to land.

First diagnosis

It was correct and not enough.

The writeback of the CHECKPOINT issued before every step spilled into the measurement window: at 3,000 req/s, the write p99 was 10.6 ms, the read 3.3 ms.

Second attempt

A fivefold divergence, zero spoiled windows

Req/s measured in three repetitions of the same cell after the fix
CellRepetition 1Repetition 2Repetition 3
Go, 1 core1,2506,5003,500
PHP-FPM, 1 core1,7502,750750

PHP-FPM, 1 core, 3,000 req/s

oha reports two latencies

p50, p75, p99 and slowest values of time to first byte and latency-corrected latency in the same measurement
Latencyp50p75p99Slowest
Time to first byte1.42 ms47 ms191 ms267 ms
Latency-corrected1.44 ms1,539 ms3,718 ms3,819 ms

Mechanism

The correction turns the limit into a cliff

  1. Step 1: Counts from the schedulethe moment it should have been sent
  2. Step 2: Responses slow downconnections fill up
  3. Step 3: The schedule slipsevery request inherits it
  4. Step 4: Cliffeither passes or misses

What the search could not measure

The cliff moves between repetitions.

Exponential search and bisection assume that pass/fail is monotonic in the rate. Neither was true.

The fix

Drop the search, set fixed rates

  • 7Fixed ratesPer cell, the same in every repetition
  • VoteMajority ruleIf most repetitions carried it, it is carried
  • 115%Probe thresholdBefore every repetition

One more lie

Chromium on the host machine

The effect on the measurement of headless Chromium processes on the host machine during one run
MetricValue
Headless Chromium processesAbout thirty
Total CPU1,019%
Left for the Docker VM94%
Go, 4 cores17,490 req/s

Result · Go, 4 cores

Two methods, within a few percent

Go, four cores: the tuning phase in the closed loop versus the result of five repetitions on a quiet machine
MeasurementReq/sNote
Tuning phase (closed loop)55,715Twelve hours earlier
Five repetitions (quiet machine)55,836–57,321Median 56,571; 45/45 repetitions passed the probe
Tan thinking, hand on chin

For your own setup

Three concrete additions

  • A second phase that measures in a different loop, and a rule that voids the run when the two diverge (pending)
  • A probe before every repetition that runs no application code (pending)
  • A record of the environment alongside the repetition (pending)
Tan waving

The measurement itself

The second road

muhammetsafak.com.tr/en/research/capacity-per-core-go-php-oauth2/

The raw data of the numbers and the discarded attempts is in the research notebook.

Share, embed, download