# Designing Reliable Event-Driven Systems

> A practical guide to the patterns that keep event-driven systems correct when networks, brokers, deploys and people misbehave.

- Series: Book II of III · The Unreliable Systems Trilogy (https://muhammetsafak.com/books/series/the-unreliable-systems-trilogy/)
- Edition: 1st edition · 333 pages
- License: © 2026 Muhammet Şafak — free to download and read
- Who it's for: Backend and platform engineers, Tech leads
- Editions: English (Available), Türkçe (Available), Español (Available), 简体中文 (soon), Português (Available), Deutsch (soon), Français (soon)
- Source: https://muhammetsafak.com/books/designing-reliable-event-driven-systems/
- Language: en-US
- Author: Muhammet Şafak

---
One instance committed an order and shut down before it published the event; no alert fired. By the next morning a buyer was looking at an order marked as placed, while Payments had authorized no card and Inventory had reserved nothing. The scene is set at Parcelly, a fictional marketplace, and nobody wrote a bug in the usual sense: every line did what it said, and the system failed between a database commit and a network call.

The book's claim is that the problem is not message brokers but failure semantics: deciding in advance what happens when a message is late, lost, duplicated or out of order. The gap in the opening scene is the dual-write problem, which the book closes with the transactional outbox before moving on to delivery, duplication, recovery and replay.

It is broker-agnostic and does not promise that following it will make your system correct. Reliability is not a property you add; it is a set of failures you have decided how to handle and a set you have decided to accept. Each numbered chapter states the pattern's costs, its common mistakes and when not to use it.

## The book's thesis

- Not message brokers
- But failure semantics

## Contents

1. Preface
2. The Lie of the Happy Path
3. Messaging Semantics
4. A Practical Failure Model
5. The Dual-Write Problem and the Transactional Outbox
6. Relays, Polling and Change Data Capture
7. Retries Done Right
8. Idempotency
9. Poison Messages and Dead Letters
10. Ordering
11. Sagas and Long-Running Workflows
12. Event Design
13. Schema Evolution
14. Observability and Reconciliation
15. Recovery, Disasters and Broker Migration
16. Testing Event-Driven Systems
17. A Reliability Checklist and Closing Thoughts
18. Appendix: Glossary and Pattern Quick Reference

## Downloads

- English · EPUB: https://files.muhammetsafak.com/books/designing-reliable-event-driven-systems/en/designing-reliable-event-driven-systems-en.epub
- English · PDF: https://files.muhammetsafak.com/books/designing-reliable-event-driven-systems/en/designing-reliable-event-driven-systems-en.pdf
- Türkçe · EPUB: https://files.muhammetsafak.com/books/designing-reliable-event-driven-systems/tr/designing-reliable-event-driven-systems-tr.epub
- Türkçe · PDF: https://files.muhammetsafak.com/books/designing-reliable-event-driven-systems/tr/designing-reliable-event-driven-systems-tr.pdf
- Español · EPUB: https://files.muhammetsafak.com/books/designing-reliable-event-driven-systems/es/designing-reliable-event-driven-systems-es.epub
- Español · PDF: https://files.muhammetsafak.com/books/designing-reliable-event-driven-systems/es/designing-reliable-event-driven-systems-es.pdf
- Português · EPUB: https://files.muhammetsafak.com/books/designing-reliable-event-driven-systems/pt/designing-reliable-event-driven-systems-pt.epub
- Português · PDF: https://files.muhammetsafak.com/books/designing-reliable-event-driven-systems/pt/designing-reliable-event-driven-systems-pt.pdf

## Frequently asked

### Who is this book for?

Backend and platform engineers and tech leads who build or operate systems that communicate through events.

### Which message broker or language do I need?

None in particular. The book is not a tutorial for any specific broker; it is broker-agnostic and uses language-independent pseudocode and Mermaid diagrams.

### Will following the book make my system correct?

The book does not promise that. It treats reliability as a set of failures you have decided how to handle and a set you have decided to accept, and it asks you to treat at-least-once delivery as a fact and make handlers safe to run twice.
