# 64 goroutines on four cores: Mutex and channel run at the same speed, but the Mutex's p99 is five times higher

> sync.Mutex vs channels in 845 cells: 64 goroutines on 4 cores run at the same speed, but the Mutex's p99 is 5x higher; on 8 cores it is 2.4x faster.

- Kind: Measurement
- Question: When G goroutines on P cores hammer the same shared counter, how many operations a second do sync.Mutex, an owner-goroutine channel and a buffered channel manage, and how do tail latency and fairness of waiting change?
- Finding: With 64 goroutines on four cores, sync.Mutex and a one-way channel run at the same speed (7.69 and 7.94 million operations/s) but the Mutex's p99 wait is 5 times higher. On eight cores the Mutex is 2.4 and 4.0 times faster (15.96 million operations/s; the channels 6.65 and 3.97 million); the ratios come from unflagged cells.
- Method: go-sync-bench measured three patterns: S1 single counter (sync.Mutex, a one-way channel to an owner goroutine, a request-reply channel, atomic), S2 read ratio (Mutex, RWMutex, owner-goroutine channel), S3 producer-consumer queue (channel buffers of 0, 1, 64, 1024 and a queue built on sync.Cond). Each cell, per G goroutines and P core group (1, 2, 4 and 8 via cpuset), ran 5 rounds; each repetition had a 0.5 s warm-up, a 1.5 s throughput window and a 2.0 s latency window. 845 cells, 4,225 valid repetitions. After every round × core group container came a 60 s cooldown, in a network-less, read-only Docker container with 4 GiB of memory. Cell medians were taken from all 5 repetitions; no repetition was discarded. Flag thresholds are 10% for throughput, 25% for p99 and 5% for clock drift, over the full range. When picking rerun candidates, the most extreme repetition was dropped and the remaining range examined (throughput 10%, p99 50%); the first design's trigger selected 294 candidates, which exceeded the 84-cell budget, so this rule was used and 88 cells were rerun (2026-09-30 01:38-02:31 UTC). 518 of the 845 cells are flagged, so confidence is low. Latency percentiles are upper bounds of log buckets and are sampled at 1/16; the floor is 164 ns, 4 times the clock resolution.
- Metrics: G=64 P=4 · p99 wait, Mutex → one-way channel: 491,519 → 98,303 ns · G=64 P=8 · Mutex: 15.96 million operations/s · G=64 P=8 · one-way channel: 6.65 million operations/s · Cells carrying a quality flag: 518 / 845
- Measured on: 2026-09-29
- Confidence: Low confidence
- Status: Current
- Programme: Language & runtime
- Environment: Go 1.27.1 · linux/arm64 · Virtualization Docker Engine 29.8.0 · 12 vCPU · aarch64 · Hardware Apple M4 Pro · 12 CPU · Core groups P1 = 4 · P2 = 4-5 · P4 = 4-7 · P8 = 4-11 (cpuset) · Container network-less · read-only · 4 GiB memory · Image first pass sha256:640a046f219a… · rerun sha256:eabc1573265a… · Protocol 5 rounds · 845 cells · 0.5 s warm-up + 1.5 s throughput + 2.0 s latency · 60 s cooldown (per container) · Seed 932269622
- Technologies: Go, Docker
- To reproduce: ./bench/verify.sh && STAMP=$(date -u +%F) ./bench/run.sh --all
- Source code: https://github.com/muhammetsafak/go-sync-bench
- Raw data: https://github.com/muhammetsafak/go-sync-bench/tree/main/results/2026-09-29
- Raw data licence: https://github.com/muhammetsafak/go-sync-bench/blob/main/LICENSE
- Published: 2026-09-30
- Source: https://muhammetsafak.com/research/mutex-vs-channel-under-contention/
- Language: en-US
- Author: Muhammet Şafak

---
When I pointed 64 goroutines at the same counter on four cores, `sync.Mutex`
and a channel did the same number of operations a second: 7.69 and 7.94
million. In the same cell the p99 wait was 491,519 ns against 98,303 ns. That
pair is where I saw that "which one is faster" has no single answer; this
record writes down what that answer depends on, as far as I measured it. All
the code and the raw runs are in the
[go-sync-bench](https://github.com/muhammetsafak/go-sync-bench) repository.

For a practical walk-through of goroutines and channels, see [Concurrency in
Go: goroutines and channels in
practice](/blog/concurrency-in-go-goroutines-and-channels-in-practice/).
Here there is only measurement.

## What I measured

The main pattern (S1) is this: G goroutines increment a single shared counter.
There are four ways to reach the counter.

| Model | Access style |
| --- | --- |
| Mutex | Counter protected by sync.Mutex |
| One-way channel | A single owner goroutine holds the counter; the others send requests to the channel and do not wait for a reply |
| Request-reply channel | The same owner goroutine, but each request carries a reply channel and waits for the answer |
| atomic | Atomic increment; for context only, not part of the headline claim |

The request channel is unbuffered. A buffered request channel could change the result; I did not measure that.

Two more patterns ran: in S2 the read ratio varies (50, 90 and 99 percent
reads), and in S3 a producer-consumer queue is tried with different buffers
(channel 0, 1, 64, 1024 and a queue built on `sync.Cond`). The core count P is
the size of the core group given to the container via cpuset: 1, 2, 4 and 8.

In this record "operations/s" is the median of a throughput window, and "wait"
is the time it takes an operation to reach the counter during the latency
window. Latency percentiles are reported as upper bounds of log buckets; so
numbers like 491,519 are the ceiling of a bucket, not an exact measured value.

## 64 goroutines, four cores: same speed, different tail

| Model | Operations/s | p50 | p99 | p99.9 |
| --- | --- | --- | --- | --- |
| Mutex | 7,688,545 | 639 ns | 491,519 ns | 1,441,791 ns |
| One-way channel | 7,944,446 | 175 ns | 98,303 ns | 114,687 ns |

S1, G=64, P=4, workload W=0. Both cells are unflagged. p99 ratio 491,519 / 98,303 = 5.00.

The throughput difference is 3.3 percent in the channel's favor, so within
noise. The wait difference is not noise: 639 against 175 at p50, 491,519
against 98,303 at p99, 1,441,791 against 114,687 ns at p99.9. The Mutex
produces the same total with a wider distribution. A benchmark that looks at
the average would have seen these two models as equal.

I did not measure why it is this way. I could build an explanation of how the
Mutex wakes its waiting goroutines, but it would not be evidence this
experiment produced; I am leaving it out of the text.

## On eight cores the table turned

**G=64: throughput on four and eight cores**

The Mutex nearly doubled on eight cores while the one-way channel fell back. The request-reply channel's P=4 cell carries a quality flag and stayed out of the chart; its P=8 value is in the table below.

Source: go-sync-bench, results/2026-09-29/merged/summary.json, S1 cells

|  | P=4 | P=8 |
| --- | --- | --- |
| Mutex | 7.69 M | 15.96 M |
| One-way channel | 7.94 M | 6.65 M |

| Model | Operations/s (P=8) | Mutex ratio | p99 |
| --- | --- | --- | --- |
| Mutex | 15,962,148 | 1.00 | 114,687 ns |
| One-way channel | 6,648,606 | 2.40 | 122,879 ns |
| Request-reply channel | 3,970,457 | 4.02 | 131,071 ns |

S1, G=64, P=8. All three cells are unflagged. The ratio column is how many times faster the Mutex is than that model: 15,962,148 / 6,648,606 = 2.40 and 15,962,148 / 3,970,457 = 4.02.

The Mutex sped up 2.08 times from P=4 to P=8. The one-way channel dropped to
0.84 of its earlier value: a single owner goroutine holds the counter; I don't
know what adding cores contributes to it. The experiment did not measure the
cause; I am not even offering a hypothesis.

At P=8 the p99 is close across all three models: 114,687, 122,879 and 131,071
ns. The fivefold gap at four cores is not visible here. The question of which
model's tail is "better" has no single answer; it depends on P.

What I have to leave out of scope: the Mutex cell at P=1 (63.10 million)
carries the `spread_persistent` and `observer_effect` flags, and the one at
P=2 (13.22 million) is flagged. I did not take either into the claim. I also
can't say whether the Mutex curve is monotonic in P, because the P=1 and P=2
cells are flagged.

## 10,000 goroutines: the Mutex is fast, its unfairness is unexplained

| Model | Operations/s | Jain | p99 | Flag |
| --- | --- | --- | --- | --- |
| Mutex | 14,878,958 | 0.8237 | 27,262,975 ns | fairness_unexplained |
| One-way channel | 2,052,681 | 0.9999 | 5,767,167 ns | none |
| Request-reply channel | 1,778,745 | 1.0000 | 6,291,455 ns | none |

S1, G=10,000, P=8. Mutex / one-way channel = 7.25; Mutex / request-reply channel = 8.36. The closer Jain gets to 1, the more equal the share each goroutine gets.

Here the Mutex is 7.25 and 8.36 times faster than the channels. But the Jain
fairness index is 0.8237 and that value is flagged: `fairness_unexplained`.
The harness can tie unfairness to only one of two sources: parking behavior
(`sync`). The rule is this: if G ≤ P, if the per-goroutine time-slice bound b
≥ 1, or if at least half of the sampled acquisitions exceed the parking
threshold, the source is `sync`, otherwise `unexplained`. In this cell
G=10,000 > P=8, b = 0.12 and the share of slow acquisitions is only 0.0176; since none
of the three conditions holds, the unfairness could not be tied to parking
behavior. The rule does not look at the Jain value. The expected Jain the
harness computes from b is 0.12 (`expected_quantum_jain`), the measured one
0.8237; I have not verified the quantum model this calculation rests on, so I
am not attributing anything to the scheduler. I cannot trace the Mutex's
unfairness to a cause. By my design rule these cells do not enter the
fairness comparison across models; in the table I am not saying "the Mutex is
more unfair", I am saying "its unfairness is unexplained".

The context of the same table: on four cores the Mutex's Jain value is 0.4064
at G=1,024 and 0.1892 at G=10,000, both `fairness_unexplained`. In the channel
cells the Jain values at G=10,000 and P=8 are 0.9999 and 1; at the same G and
P=4 the one-way channel's Jain is 0.7995 (from the rerun), so the channels'
fairness depends on P as well.

## Read ratio and queue buffer

In S2 the Mutex speeds up as the read ratio rises; the owner-goroutine channel
barely moves.

**S2, G=64, P=4: throughput by read ratio**

The Mutex is unflagged. The channel's 90 percent read cell comes from the rerun, and its 99 percent read cell carries the spread_throughput flag. The RWMutex cells are flagged observer_effect and are not in the chart.

Source: go-sync-bench, results/2026-09-29/merged/summary.json, S2 cells

|  | 50% reads | 90% reads | 99% reads |
| --- | --- | --- | --- |
| Mutex | 9.86 M | 12.23 M | 13.12 M |
| Owner-goroutine channel | 4.42 M | 4.58 M | 4.5 M |

At 50 percent reads the Mutex is 2.23 times the channel, at 90 percent 2.67
times. I am not writing the ratio for the 99 percent cell because it is
flagged. The Mutex's speed at 99 percent is 1.33 times its speed at 50
percent.

In S3, at G=64, P=4: the unbuffered channel does 5,414,190, the channel with a
buffer of one 6,912,398, and the `sync.Cond` queue with a buffer of 1,024
4,461,277 operations/s. These three cells are unflagged: when the buffer goes
up to one, the channel is 1.28 times faster, and 1.55 times faster than the
`sync.Cond` queue. The channels with buffers of 64 and 1,024 and the
`sync.Cond` cell with a buffer of 64 are flagged; numbers like 18.72 and 21.98
million only show direction, and I did not derive ratios.

## Conclusion

> **Result**
>
> At G=64 and P=4 the Mutex and the one-way channel run at the same speed (7.69
> and 7.94 million operations/s) but the Mutex's p99 is 5 times higher. At P=8
> the Mutex is 2.4 and 4.0 times faster. At G=10,000 and P=8 the Mutex is 7.25
> and 8.36 times faster, but its unfairness cannot be tied to a cause. The
> answer to "which one is faster?" depends on G and P.

## How I measured

The experiment ran 845 cells; each cell 5 rounds, each round with a 0.5 s
warm-up, a 1.5 s throughput window and a 2.0 s latency window per repetition.
4,225 valid repetitions in total. The median duration of a repetition is 4.38
s; eight repetitions exceeded the budget (for example
`s1-atomic-g10000-p4-w0` at 55 s). After every round × core group container
there is a 60 s cooldown; there is no wait between cells inside a container.

The container is network-less and read-only, with 4 GiB of memory. The first
pass had no sleep gate; the host machine slept once: in the 11:45-11:59 UTC
window, round 3 of the `s1-chan_oneway-g64-p1-w0` cell was affected by this
and that cell is not part of the claim. Before the rerun, a `sleep 0` check
was added to `run.sh`.

The results come from two passes. The first pass started at 2026-09-29T08:34:54Z
with the `sha256:640a046f219a…` image. The first design's rerun trigger
selected 294 candidates (290 had failed the deviation gate, 11 a second gate;
one cell can fail both). That exceeded the 84-cell budget. The trigger was
changed to a trimmed range: the most extreme repetition is dropped and the
remaining range examined (10% for throughput, 50% for p99), and this only
selects which cell is rerun. The budget was raised from 84 to 88; the 88 cells
were rerun with the `sha256:eabc1573265a…` image at 2026-09-30 01:38-02:31
UTC. In the merged table the source shows `rerun`; 36 of them stayed above the
threshold even after the rerun. The cell median is still taken from all 5
repetitions. The flag is a separate criterion: over the full range, 10% for
throughput, 25% for p99, 5% for clock drift.

Quality flags:

| Flag | Cell count | Meaning |
| --- | --- | --- |
| observer_effect | 171 | The operation rate in the latency window is below 0.9 of the throughput window's; the measurement is affecting itself |
| fairness_unexplained | 235 | Unfairness cannot be tied to parking behavior: G > P, b < 1 and fewer than half of the sampled acquisitions are above the parking threshold; the Jain value is not looked at |
| spread_p99 | 189 | Spread of p99 across repetitions above 0.25 |
| spread_throughput | 125 | Spread of throughput across repetitions above 0.10 |
| calib_drift | 72 | Clock calibration drifted by more than 0.05 |
| spread_persistent | 36 | Spread persists even after the rerun |

A cell can carry more than one flag; 518 cells are flagged in total, 61.3% of 845. Because the share exceeds 15%, the confidence level the measurement recommends is low (confidence: low).

Noise floor: for P=2, atomic at G=2 has a Jain of 0.9999, for P=4 at G=4
0.9998, for P=8 at G=4 0.9997. So the measurement setup does not distort
fairness on its own. The self-test run passed 40 of its 40 trials. The
publishing command is in the `reproduce` field; the raw data and the license
(MIT) are in the repository.

> **Tan:** This record does not say which one is faster; it says when each one is faster. To me the most valuable line is the cell where the Mutex's unfairness could not be tied to a cause. Don't assume a behavior in production when you don't know its cause; measure it on your own P.

## Limits and honesty notes

1. **A majority of flagged cells.** 518 of 845 cells are flagged. I took the
   throughput and p99 ratios in the headline from cells that carry none of
   these flags; the G=10,000, P=8 Mutex cell carries only the fairness flag.
   Where flagged cells are mentioned to show direction, they are labeled as
   such.
2. **vCPUs inside a VM.** The measurement ran in a Docker virtual machine; the
   P/E core split is not visible from the VM and the cluster clock depends on
   the number of busy cores. P values are not physical cores but groups of
   vCPUs given via cpuset.
3. **P=2 is noisy.** I could not verify why.
4. **The fairness flag's rule.** The flag looks at G, P, b and the parking
   share; it cannot tell scheduler-caused unfairness apart. I did not verify
   the quantum model (`expected_quantum_jain`) and did not use it in the
   explanations.
5. **Latency resolution.** Latency percentiles are upper bounds of log buckets
   and are sampled at a rate of 1/16. `< 164 ns` is the floor, 4 times the
   clock resolution. I did not use the `starving_fraction` field because the
   sample count is small.
6. **The request channel is unbuffered.** A buffered request channel could
   change the result of the channel models.
7. **Two goroutines at G=1.** The owner-goroutine models run 2 goroutines at
   G=1.
8. **Out of scope.** `sync.Map` was not measured at all; atomic was not
   measured in S2 either. A synthetic counter was measured, not a production
   workload.

To produce your own number, run the
[go-sync-bench](https://github.com/muhammetsafak/go-sync-bench) repository on
your own hardware; choose P according to your own target, because the clearest
result of this record is that the answer depends on P.
