Skip to content
Muhammet Şafak
tr

64 goroutines on four cores: Mutex and channel run at the same speed, but the Mutex's p99 is five times higher

When G goroutines on P cores hammer the same shared counter, how many operations a second do sync.Mutex, an owner-goroutine channel and a buffered channel manage, and how do tail latency and fairness of waiting change?

Finding

With 64 goroutines on four cores, sync.Mutex and a one-way channel run at the same speed (7.69 and 7.94 million operations/s) but the Mutex's p99 wait is 5 times higher. On eight cores the Mutex is 2.4 and 4.0 times faster (15.96 million operations/s; the channels 6.65 and 3.97 million); the ratios come from unflagged cells.

G=64 P=4 · p99 wait, Mutex → one-way channel
491,519 → 98,303 ns
G=64 P=8 · Mutex
15.96 million operations/s
G=64 P=8 · one-way channel
6.65 million operations/s
Cells carrying a quality flag
518 / 845

Method

go-sync-bench measured three patterns: S1 single counter (sync.Mutex, a one-way channel to an owner goroutine, a request-reply channel, atomic), S2 read ratio (Mutex, RWMutex, owner-goroutine channel), S3 producer-consumer queue (channel buffers of 0, 1, 64, 1024 and a queue built on sync.Cond). Each cell, per G goroutines and P core group (1, 2, 4 and 8 via cpuset), ran 5 rounds; each repetition had a 0.5 s warm-up, a 1.5 s throughput window and a 2.0 s latency window. 845 cells, 4,225 valid repetitions. After every round × core group container came a 60 s cooldown, in a network-less, read-only Docker container with 4 GiB of memory. Cell medians were taken from all 5 repetitions; no repetition was discarded. Flag thresholds are 10% for throughput, 25% for p99 and 5% for clock drift, over the full range. When picking rerun candidates, the most extreme repetition was dropped and the remaining range examined (throughput 10%, p99 50%); the first design's trigger selected 294 candidates, which exceeded the 84-cell budget, so this rule was used and 88 cells were rerun (2026-09-30 01:38-02:31 UTC). 518 of the 845 cells are flagged, so confidence is low. Latency percentiles are upper bounds of log buckets and are sampled at 1/16; the floor is 164 ns, 4 times the clock resolution.

Low confidence Single run or uncontrolled environment — indicates direction, not an exact figure.
Measured on

measured 5 days ago

Published

Environment

Go
1.27.1 · linux/arm64
Virtualization
Docker Engine 29.8.0 · 12 vCPU · aarch64
Hardware
Apple M4 Pro · 12 CPU
Core groups
P1 = 4 · P2 = 4-5 · P4 = 4-7 · P8 = 4-11 (cpuset)
Container
network-less · read-only · 4 GiB memory
Image
first pass sha256:640a046f219a… · rerun sha256:eabc1573265a…
Protocol
5 rounds · 845 cells · 0.5 s warm-up + 1.5 s throughput + 2.0 s latency · 60 s cooldown (per container)
Seed
932269622

Technologies

Go Docker

To reproduce

./bench/verify.sh && STAMP=$(date -u +%F) ./bench/run.sh --all

When I pointed 64 goroutines at the same counter on four cores, sync.Mutex and a channel did the same number of operations a second: 7.69 and 7.94 million. In the same cell the p99 wait was 491,519 ns against 98,303 ns. That pair is where I saw that “which one is faster” has no single answer; this record writes down what that answer depends on, as far as I measured it. All the code and the raw runs are in the go-sync-bench repository.

For a practical walk-through of goroutines and channels, see Concurrency in Go: goroutines and channels in practice. Here there is only measurement.

What I measured

The main pattern (S1) is this: G goroutines increment a single shared counter. There are four ways to reach the counter.

Model Access style
Mutex Counter protected by sync.Mutex
One-way channel A single owner goroutine holds the counter; the others send requests to the channel and do not wait for a reply
Request-reply channel The same owner goroutine, but each request carries a reply channel and waits for the answer
atomic Atomic increment; for context only, not part of the headline claim
The request channel is unbuffered. A buffered request channel could change the result; I did not measure that.

Two more patterns ran: in S2 the read ratio varies (50, 90 and 99 percent reads), and in S3 a producer-consumer queue is tried with different buffers (channel 0, 1, 64, 1024 and a queue built on sync.Cond). The core count P is the size of the core group given to the container via cpuset: 1, 2, 4 and 8.

In this record “operations/s” is the median of a throughput window, and “wait” is the time it takes an operation to reach the counter during the latency window. Latency percentiles are reported as upper bounds of log buckets; so numbers like 491,519 are the ceiling of a bucket, not an exact measured value.

64 goroutines, four cores: same speed, different tail

Model Operations/s p50 p99 p99.9
Mutex 7,688,545 639 ns 491,519 ns 1,441,791 ns
One-way channel 7,944,446 175 ns 98,303 ns 114,687 ns
S1, G=64, P=4, workload W=0. Both cells are unflagged. p99 ratio 491,519 / 98,303 = 5.00.

The throughput difference is 3.3 percent in the channel’s favor, so within noise. The wait difference is not noise: 639 against 175 at p50, 491,519 against 98,303 at p99, 1,441,791 against 114,687 ns at p99.9. The Mutex produces the same total with a wider distribution. A benchmark that looks at the average would have seen these two models as equal.

I did not measure why it is this way. I could build an explanation of how the Mutex wakes its waiting goroutines, but it would not be evidence this experiment produced; I am leaving it out of the text.

On eight cores the table turned

G=64: throughput on four and eight cores

The Mutex nearly doubled on eight cores while the one-way channel fell back. The request-reply channel's P=4 cell carries a quality flag and stayed out of the chart; its P=8 value is in the table below.

  • Mutex
  • One-way channel

M Source: go-sync-bench, results/2026-09-29/merged/summary.json, S1 cells

Data table
go-sync-bench, results/2026-09-29/merged/summary.json, S1 cells
Series P=4P=8
Mutex 7.69 M15.96 M
One-way channel 7.94 M6.65 M

The chart is drawn in the browser; the table below carries the same data.

Model Operations/s (P=8) Mutex ratio p99
Mutex 15,962,148 1.00 114,687 ns
One-way channel 6,648,606 2.40 122,879 ns
Request-reply channel 3,970,457 4.02 131,071 ns
S1, G=64, P=8. All three cells are unflagged. The ratio column is how many times faster the Mutex is than that model: 15,962,148 / 6,648,606 = 2.40 and 15,962,148 / 3,970,457 = 4.02.

The Mutex sped up 2.08 times from P=4 to P=8. The one-way channel dropped to 0.84 of its earlier value: a single owner goroutine holds the counter; I don’t know what adding cores contributes to it. The experiment did not measure the cause; I am not even offering a hypothesis.

At P=8 the p99 is close across all three models: 114,687, 122,879 and 131,071 ns. The fivefold gap at four cores is not visible here. The question of which model’s tail is “better” has no single answer; it depends on P.

What I have to leave out of scope: the Mutex cell at P=1 (63.10 million) carries the spread_persistent and observer_effect flags, and the one at P=2 (13.22 million) is flagged. I did not take either into the claim. I also can’t say whether the Mutex curve is monotonic in P, because the P=1 and P=2 cells are flagged.

10,000 goroutines: the Mutex is fast, its unfairness is unexplained

Model Operations/s Jain p99 Flag
Mutex 14,878,958 0.8237 27,262,975 ns fairness_unexplained
One-way channel 2,052,681 0.9999 5,767,167 ns none
Request-reply channel 1,778,745 1.0000 6,291,455 ns none
S1, G=10,000, P=8. Mutex / one-way channel = 7.25; Mutex / request-reply channel = 8.36. The closer Jain gets to 1, the more equal the share each goroutine gets.

Here the Mutex is 7.25 and 8.36 times faster than the channels. But the Jain fairness index is 0.8237 and that value is flagged: fairness_unexplained. The harness can tie unfairness to only one of two sources: parking behavior (sync). The rule is this: if G ≤ P, if the per-goroutine time-slice bound b ≥ 1, or if at least half of the sampled acquisitions exceed the parking threshold, the source is sync, otherwise unexplained. In this cell G=10,000 > P=8, b = 0.12 and the share of slow acquisitions is only 0.0176; since none of the three conditions holds, the unfairness could not be tied to parking behavior. The rule does not look at the Jain value. The expected Jain the harness computes from b is 0.12 (expected_quantum_jain), the measured one 0.8237; I have not verified the quantum model this calculation rests on, so I am not attributing anything to the scheduler. I cannot trace the Mutex’s unfairness to a cause. By my design rule these cells do not enter the fairness comparison across models; in the table I am not saying “the Mutex is more unfair”, I am saying “its unfairness is unexplained”.

The context of the same table: on four cores the Mutex’s Jain value is 0.4064 at G=1,024 and 0.1892 at G=10,000, both fairness_unexplained. In the channel cells the Jain values at G=10,000 and P=8 are 0.9999 and 1; at the same G and P=4 the one-way channel’s Jain is 0.7995 (from the rerun), so the channels’ fairness depends on P as well.

Read ratio and queue buffer

In S2 the Mutex speeds up as the read ratio rises; the owner-goroutine channel barely moves.

S2, G=64, P=4: throughput by read ratio

The Mutex is unflagged. The channel's 90 percent read cell comes from the rerun, and its 99 percent read cell carries the spread_throughput flag. The RWMutex cells are flagged observer_effect and are not in the chart.

  • Mutex
  • Owner-goroutine channel

M Source: go-sync-bench, results/2026-09-29/merged/summary.json, S2 cells

Data table
go-sync-bench, results/2026-09-29/merged/summary.json, S2 cells
Series 50% reads90% reads99% reads
Mutex 9.86 M12.23 M13.12 M
Owner-goroutine channel 4.42 M4.58 M4.5 M

The chart is drawn in the browser; the table below carries the same data.

At 50 percent reads the Mutex is 2.23 times the channel, at 90 percent 2.67 times. I am not writing the ratio for the 99 percent cell because it is flagged. The Mutex’s speed at 99 percent is 1.33 times its speed at 50 percent.

In S3, at G=64, P=4: the unbuffered channel does 5,414,190, the channel with a buffer of one 6,912,398, and the sync.Cond queue with a buffer of 1,024 4,461,277 operations/s. These three cells are unflagged: when the buffer goes up to one, the channel is 1.28 times faster, and 1.55 times faster than the sync.Cond queue. The channels with buffers of 64 and 1,024 and the sync.Cond cell with a buffer of 64 are flagged; numbers like 18.72 and 21.98 million only show direction, and I did not derive ratios.

Conclusion

How I measured

The experiment ran 845 cells; each cell 5 rounds, each round with a 0.5 s warm-up, a 1.5 s throughput window and a 2.0 s latency window per repetition. 4,225 valid repetitions in total. The median duration of a repetition is 4.38 s; eight repetitions exceeded the budget (for example s1-atomic-g10000-p4-w0 at 55 s). After every round × core group container there is a 60 s cooldown; there is no wait between cells inside a container.

The container is network-less and read-only, with 4 GiB of memory. The first pass had no sleep gate; the host machine slept once: in the 11:45-11:59 UTC window, round 3 of the s1-chan_oneway-g64-p1-w0 cell was affected by this and that cell is not part of the claim. Before the rerun, a sleep 0 check was added to run.sh.

The results come from two passes. The first pass started at 2026-09-29T08:34:54Z with the sha256:640a046f219a… image. The first design’s rerun trigger selected 294 candidates (290 had failed the deviation gate, 11 a second gate; one cell can fail both). That exceeded the 84-cell budget. The trigger was changed to a trimmed range: the most extreme repetition is dropped and the remaining range examined (10% for throughput, 50% for p99), and this only selects which cell is rerun. The budget was raised from 84 to 88; the 88 cells were rerun with the sha256:eabc1573265a… image at 2026-09-30 01:38-02:31 UTC. In the merged table the source shows rerun; 36 of them stayed above the threshold even after the rerun. The cell median is still taken from all 5 repetitions. The flag is a separate criterion: over the full range, 10% for throughput, 25% for p99, 5% for clock drift.

Quality flags:

Flag Cell count Meaning
observer_effect 171 The operation rate in the latency window is below 0.9 of the throughput window's; the measurement is affecting itself
fairness_unexplained 235 Unfairness cannot be tied to parking behavior: G > P, b < 1 and fewer than half of the sampled acquisitions are above the parking threshold; the Jain value is not looked at
spread_p99 189 Spread of p99 across repetitions above 0.25
spread_throughput 125 Spread of throughput across repetitions above 0.10
calib_drift 72 Clock calibration drifted by more than 0.05
spread_persistent 36 Spread persists even after the rerun
A cell can carry more than one flag; 518 cells are flagged in total, 61.3% of 845. Because the share exceeds 15%, the confidence level the measurement recommends is low (confidence: low).

Noise floor: for P=2, atomic at G=2 has a Jain of 0.9999, for P=4 at G=4 0.9998, for P=8 at G=4 0.9997. So the measurement setup does not distort fairness on its own. The self-test run passed 40 of its 40 trials. The publishing command is in the reproduce field; the raw data and the license (MIT) are in the repository.

Limits and honesty notes

  1. A majority of flagged cells. 518 of 845 cells are flagged. I took the throughput and p99 ratios in the headline from cells that carry none of these flags; the G=10,000, P=8 Mutex cell carries only the fairness flag. Where flagged cells are mentioned to show direction, they are labeled as such.
  2. vCPUs inside a VM. The measurement ran in a Docker virtual machine; the P/E core split is not visible from the VM and the cluster clock depends on the number of busy cores. P values are not physical cores but groups of vCPUs given via cpuset.
  3. P=2 is noisy. I could not verify why.
  4. The fairness flag’s rule. The flag looks at G, P, b and the parking share; it cannot tell scheduler-caused unfairness apart. I did not verify the quantum model (expected_quantum_jain) and did not use it in the explanations.
  5. Latency resolution. Latency percentiles are upper bounds of log buckets and are sampled at a rate of 1/16. < 164 ns is the floor, 4 times the clock resolution. I did not use the starving_fraction field because the sample count is small.
  6. The request channel is unbuffered. A buffered request channel could change the result of the channel models.
  7. Two goroutines at G=1. The owner-goroutine models run 2 goroutines at G=1.
  8. Out of scope. sync.Map was not measured at all; atomic was not measured in S2 either. A synthetic counter was measured, not a production workload.

To produce your own number, run the go-sync-bench repository on your own hardware; choose P according to your own target, because the clearest result of this record is that the answer depends on P.

Related posts

Other records

All records
Service & load Measurement

One core carries 14,330 OAuth2 requests in Go and 5,152 in PHP-FPM

The same API verifies an RS256 bearer token on every request and then reads or writes one row in PostgreSQL — on one, two and four cores, how much mixed traffic does it carry in Go, PHP-FPM and FrankenPHP worker mode?

Finding

On four cores Go carried 57,321 mixed requests a second, FrankenPHP 25,659 and php-fpm 20,606 — 14,330, 6,415 and 5,152 per core. The number that goes into a capacity plan is not that one but application CPU per request: 66.8, 110.7 and 187.7 microseconds. At saturation FrankenPHP uses only 2.84 of its four cores, against Go's 3.83 and php-fpm's 3.87. The database is not the constraint: on the same four cores PostgreSQL alone writes 68,212 rows a second, above the mixed ceiling of the fastest candidate.

measured 17 days ago

Medium confidence

Laravel's preload curve: 123 files buy eight times what the last 1,912 do

How far can a curated preload take Laravel, and what does each slice cost in start-up time?

Finding

The curve is not proportional to volume. The first 1,592 files — Laravel's own framework — buy 30 ms and add 1.2 seconds to start-up. The next 1,094 Symfony files buy 9.5 ms for free. The **123 files** after that (psr, carbon) buy 15.7 ms, more than the 1,094 before them. And the last 1,912 buy 1.8 ms while adding another 1.2 seconds. So the blanket preload the earlier record measured as a ceiling is the worst point on the curve that is not the origin: stopping at 2,809 files gives 12.77 ms for 1,514 ms of start-up, while 4,721 files ask 2,691 ms to reach 10.96 ms.

measured 43 days ago

High confidence

opcache preload cuts the deploy bill by up to fourteen times — but five of seven frameworks do not hand it to you

With `opcache.preload` on, how long is the first request seven PHP frameworks serve after a deploy, what does the gain cost, and who can actually have it?

Finding

Preload shortens the cold first request by between 3.5× and 14.2× on the six candidates whose classes are PHP source: Symfony drops from 35.58 ms to 2.50 ms, down to Phalcon's bare figure. But only two of the seven candidates — Symfony and CodeIgniter — publish a preload file of their own; for the other five the gain sits on the table waiting for the user to write one. Writing one is not as easy as it looks: a preload generated blindly from the classmap never brings Symfony up at all, and on CodeIgniter it does worse (5.29 ms) than the hand-picked official file (3.13 ms). And the cost does not vanish: Laravel's classmap preload takes the 62 ms it saves each visitor and writes it back as 2,340 ms of php-fpm start-up.

measured 43 days ago

High confidence

Search the site

Start typing to search posts, projects and pages.

Esc to close Powered by Pagefind