Book II of III The Unreliable Systems Trilogy
Designing Reliable Event-Driven Systems
Messaging, retries, idempotency, outbox, ordering and schema evolution
A practical guide to the patterns that keep event-driven systems correct when networks, brokers, deploys and people misbehave.
- 333 pages
- 1st edition
- Free
Page of the Türkçe edition: /tr/kitaplar/designing-reliable-event-driven-systems/
No sign-up, no email. License: © 2026 Muhammet Şafak — free to download and read
The book's thesis
message brokers
failure semantics
- 01
Failure semantics
Define what happens when a message is late, lost, duplicated or out of order.
- 02
Delivery & duplication
Treat at-least-once as a fact and make handlers safe to run twice.
- 03
Recovery & replay
Plan how the system finds its way back after things went wrong.
About the book
One instance committed an order and shut down before it published the event; no alert fired. By the next morning a buyer was looking at an order marked as placed, while Payments had authorized no card and Inventory had reserved nothing. The scene is set at Parcelly, a fictional marketplace, and nobody wrote a bug in the usual sense: every line did what it said, and the system failed between a database commit and a network call.
The book’s claim is that the problem is not message brokers but failure semantics: deciding in advance what happens when a message is late, lost, duplicated or out of order. The gap in the opening scene is the dual-write problem, which the book closes with the transactional outbox before moving on to delivery, duplication, recovery and replay.
It is broker-agnostic and does not promise that following it will make your system correct. Reliability is not a property you add; it is a set of failures you have decided how to handle and a set you have decided to accept. Each numbered chapter states the pattern’s costs, its common mistakes and when not to use it.
Contents
- 01 Preface
- 02 The Lie of the Happy Path
- 03 Messaging Semantics
- 04 A Practical Failure Model
- 05 The Dual-Write Problem and the Transactional Outbox
- 06 Relays, Polling and Change Data Capture
- 07 Retries Done Right
- 08 Idempotency
- 09 Poison Messages and Dead Letters
- 10 Ordering
- 11 Sagas and Long-Running Workflows
- 12 Event Design
- 13 Schema Evolution
- 14 Observability and Reconciliation
- 15 Recovery, Disasters and Broker Migration
- 16 Testing Event-Driven Systems
- 17 A Reliability Checklist and Closing Thoughts
- 18 Appendix: Glossary and Pattern Quick Reference
Frequently asked
3 questions
-
Who is this book for?
Backend and platform engineers and tech leads who build or operate systems that communicate through events.
-
Which message broker or language do I need?
None in particular. The book is not a tutorial for any specific broker; it is broker-agnostic and uses language-independent pseudocode and Mermaid diagrams.
-
Will following the book make my system correct?
The book does not promise that. It treats reliability as a set of failures you have decided how to handle and a set you have decided to accept, and it asks you to treat at-least-once delivery as a fact and make handlers safe to run twice.
Continue the trilogy