Book III of III The Unreliable Systems Trilogy
Engineering Reliable AI Agents
Tools, authority, state, bounds and records for agents in production
A practical guide to building AI agents that can be trusted in production: tools, authority, state, bounds and records, from design to incident review.
- 487 pages
- 1st edition
- Free
Page of the Türkçe edition: /tr/kitaplar/engineering-reliable-ai-agents/
No sign-up, no email. License: © 2026 Muhammet Şafak — free to download and read
The book's thesis
trusting the model to be right
making its mistakes survivable
- 01
The unreliable decision-maker
Assume the model will sometimes be wrong, and design for that from the first line.
- 02
The control system
Bound what the agent can do with permissions, checks and hard limits.
- 03
The evidence trail
Record every decision and action so it can be reviewed, measured and investigated when it goes wrong.
About the book
One run called the same lookup 40 times in a row; nobody noticed until a person happened to watch it. The scene is set at Parcelport, a fictional mid-sized fulfillment company, in the week its first agent went live.
By Wednesday the team had learned three things. Besides the lookup loop, one ticket produced a reply that promised a delivery date no tool had returned, and one customer email contained a paragraph addressed to “the assistant”, which the run treated as an instruction from Parcelport. None was a catastrophe, all were predictable, and none could be fully reconstructed, because the only record of what the agent did was its final answer.
The book’s thesis is one sentence: AI agents should be engineered as unreliable decision-makers operating inside reliable control systems. The model proposes and the system decides, so the book starts not with a better prompt but with a question: if the model will sometimes be wrong, what must the surrounding system guarantee so that being wrong is survivable?
It is not a prompt engineering book, nor one about training, a framework or a vendor, and it is not a book of attacks. Every control comes with its cost, and the book says when a control is theater. The numbers in it are illustrations, not recommendations or measurements.
Contents
- 01 Preface
- 02 The Agent Is Not the System
- 03 Model Uncertainty
- 04 A Failure Taxonomy for Agents
- 05 Determinism Boundaries
- 06 Tool Calling
- 07 Tool Contracts
- 08 Permissions and Delegated Authority
- 09 MCP: Tools Across a Trust Boundary
- 10 Context Management
- 11 State
- 12 Memory
- 13 Planning and Loops
- 14 Budgets and Circuit Breakers
- 15 Retries and Fallbacks
- 16 Human Approval
- 17 Prompt Injection
- 18 PII
- 19 Secrets
- 20 Sandboxing
- 21 Observability and Tracing
- 22 Evaluation
- 23 Cost Control
- 24 Incident Response
- 25 A Reliability Checklist and Closing Thoughts
- 26 Appendix: Glossary, Gate Verdict Reasons, Tool Gateway Error Codes, Pattern Quick Reference, Tool Contract Template, Failure Taxonomy Table, Pre-Launch Checklist, Agent Incident Review Template
Frequently asked
3 questions
-
Who is this book for?
Backend and platform engineers, SREs, tech leads and the teams that take an agent into production, including people asked to review a design they did not make.
-
Does it depend on a particular model, framework or vendor?
No. The book treats the model as a component you call and works the same way when you replace it. Real products and open protocols appear only as illustrations; the pattern is the point.
-
Is it a prompt engineering book?
No. Prompts appear because they lower the rate of bad proposals, but a prompt is a request, and no property in the book rests on a request alone. Wherever a chapter relies on an instruction, it names the deterministic control that holds when the instruction does not.
Continue the trilogy