Skip to content

5 breaking points in data-intensive systems and the Laravel production stack

Keyboard: ← → to move, F for full screen, O for overview.

Tan in front of a dashboard

Muhammet Şafak — Presentations

5 breaking points in data-intensive systems

From the cheapest fix to the costliest, and a Laravel stack with a reason for every part.

Muhammet Şafak

Section 01

Five breaking points

From the working set to a separate data tier: from the cheapest intervention to the most expensive.

What sets the scale

Not requests: the data profile.

A system can take 50 requests per minute and still be drowned by its data; another handles 5,000 requests per second comfortably while its data tier sits untouched.

Data profile

Five axes

  • Working set / RAM ratioDoes the frequently read data fit in memory?
  • Read / write ratioWhich way does the load tilt?
  • Single table sizeHow big has a single table grown?
  • Write throughputHow much does the primary write per second?
  • Consistency requirementWhich reads must see the freshest data?

Starting architecture

Starting with a single primary

PHP-FPM reaches PostgreSQL through pgBouncer and Redis directly. Enough for a few tens of GB of data and thousands of requests per minute, as long as the working set fits in RAM.

  1. PHP-FPMapplication
  2. pgBouncerconnection pool
  3. PostgreSQL 16single primary
  4. Redis 7cache · session · lock

Breaking point 01 · Working set

Signal: cache hit ratio

The frequently read data doesn’t fit in RAM. Above 99% is healthy. A drop toward 95%, read together with disk IOPS, shows the working set has spilled out of RAM.

pg_stat_database.sql
SELECTsum(blks_hit) * 100  / nullif(sum(blks_hit) + sum(blks_read), 0)  AS cache_hit_ratioFROM pg_stat_database;

Order of intervention

Exhaust the cheap one first

  1. Step 1: Index disciplinecut unnecessary scans
  2. Step 2: Separate dead datashrink the hot data
  3. Step 3: Vertical scalelast, more RAM

Breaking point 02 · Read load

Reads to the replica, writes to the primary

Signal: CPU is high, SELECT dominates, cl_waiting piles up in pgBouncer. sticky sends reads that follow a write within the same request to the primary.

config/database.php
'pgsql' => [  'driver'  => 'pgsql',  'read'  => ['host' => ['10.0.0.2']],   // replica  'write' => ['host' => ['10.0.0.1']],   // primary  'sticky' => true,  // ...shared settings],

The cost

The new problem the replica brings

  • LagReplication lag: the replica falls behind; the read-after-write guarantee is lost.
  • Monitor itIn pg_stat_replication, read the replay_lsn difference as lag_bytes.
  • PrimaryReads that need consistency, such as balance, stock and permissions, stay on the primary.
  • LimitA read replica adds nothing to write capacity.

Breaking point 03 · Single table

A table split by time

Signal: n_dead_tup keeps piling up, last_autovacuum falls behind. Instead of creating partitions by hand, you can use pg_partman.

events.sql
CREATE TABLE events (  id          bigserial,  occurred_at timestamptz NOT NULL,  payload     jsonb       NOT NULL) PARTITION BY RANGE (occurred_at);CREATE TABLE events_2026_05 PARTITION OF events  FOR VALUES FROM ('2026-05-01') TO ('2026-06-01');

Gain and condition

What does partitioning buy you?

  • Partition pruningThe query looks only at the relevant partition.
  • Cheap deletionOld data goes away with DROP TABLE events_2026_01.
  • Bounded vacuumVacuum scans the partition, not the whole table.
  • Primary keyIt must include the partition key: (id, occurred_at).
  • BRINA small, cheap index for time-ordered data.

Breaking point 04 · Write load

Signal: I/O wait, checkpoint and WAL pressure

  1. Step 1: Reduce amplificationdrop idx_scan = 0 indexes, fillfactor/HOT, batch insert
  2. Step 2: Take out append dataappend-heavy data is separated from the primary
  3. Step 3: Split the data domainlast: functional sharding

Realistic expectation

Most systems never reach point 4.

Don't rush: most systems never reach point 4, and most of those that do arrive early because they skipped point 1 (a missing index, a bloated working set).

Breaking point 05

A separate data tier

  • The moveThe data-heavy domain is moved to a separate PostgreSQL instance.
  • MigrationpgBackRest restore, logical replication and a short cutover.
  • CostA second backup, monitoring and upgrade; no JOIN across the two data sets.
  • RuleThe decision to separate is made by the data profile, not by the org chart.

Wrong diagnosis

Wrong breaking points

A read replica scales reads; a cache reduces reads; partitioning tames a single table. None of them increases write throughput.
What you seeWhat it really isThe right move
"The database is slow"N+1: the app fires 200 queriesEager loading; count queries
"We need a replica"One query lacks an indexEXPLAIN ANALYZE, add the index
"We need sharding"Replicas, partitioning untriedExhaust points 2 and 3 first
"Move to NoSQL"jsonb, BRIN, partitioning unusedUse PostgreSQL to the fullest
"Cache the writes"A cache cuts reads, not writesLook at point 4

The common error pattern

None of them adds write capacity.

A read replica scales reads. A cache reduces reads. Partitioning tames a single table. None of them increases write throughput.

jsonb · LISTEN/NOTIFY · partitioning · BRIN · logical replication

You used 20%. That's not a limit.

Most teams use 20% of PostgreSQL and conclude that "PostgreSQL isn't enough".

Summary

From cheapest to most expensive

  1. Step 1: Working setindex + RAM
  2. Step 2: Read loadread replica
  3. Step 3: Single tablepartitioning
  4. Step 4: Write loadamplification + separation
  5. Step 5: Data-heavy domainseparate data tier

Section 02

Laravel stack

I'm not saying "everything is necessary" — I'm someone who argues for keeping it minimal until you understand it's needed — but what each component buys you is concrete.

Anatomy

Why is each component there?

Requests reach PHP-FPM through Cloudflare. PHP-FPM connects to three backing services separately: PostgreSQL through pgBouncer, Redis and RabbitMQ. Workers run with Horizon under Supervisor.

  1. CloudflareDDoS, WAF, rate limiting, TLS, CDN
  2. NginxTLS, static assets, fastcgi in one config
  3. PHP-FPMopcache + preload
  4. pgBouncertransaction mode connection pool
  5. PostgreSQLpgBackRest: full + incremental + WAL PITR
  6. Redis · RabbitMQcache/session/lock · queue

Redis

One instance, three hats

Cache (Cache::remember), session (SESSION_DRIVER=redis) and lock (Cache::lock) all run safely on a single Redis, as long as these three settings are in place.

redis.conf
maxmemory-policy volatile-lruappendonly nosave 900 1 300 10 60 10000

Queue (optional)

RabbitMQ versus the Redis queue

RedisRabbitMQ
PersistenceServes from memory; goes to disk via RDB or AOFWrites persistent messages to disk
RedeliveryDepends on the retry_after timeoutDepends on the consumer's connection
RoutingThe routing decision stays in the applicationExchanges decide
MonitoringHorizon; limited to what the application dispatchesA broker-level management panel

What happens when removed

What do you lose if you remove each piece?

What do you lose if you remove each piece?
ComponentWhat happens when removed
pgBouncerIn connection storms PostgreSQL connections run out and requests throw 500
SupervisorWorkers don't come back after a crash, you have to set up alarms
HorizonVisibility collapses, you answer "why is the queue slow" blind
opcacheEvery request reads and parses the PHP file from disk, ~5x slowdown
pgBackRestOnly pg_dump remains: no PITR, a day's backup as your error margin
Redis lockYou're open to race conditions, you're left without a distributed-safe mutex

Boring stack

Measurably useful, easy to change.

The strength of a boring stack is exactly this: every piece is quantitatively valuable and easy to change.

Tan smiling

Thank you

sade.dev

Measure the data profile first, then pick the cheapest intervention.

Share, embed, download