# The Backup Didn't Come Up: An Hour-by-Hour Log of an Outage

> Our whole system went down in four minutes and the Ankara backup never came up. What we prioritized at the crisis table, and what was left behind.

- Published: 2026-10-09
- Category: Journal
- Tags: Resilience, Reliability, Infrastructure
- Reading time: 7 min read
- Source: https://muhammetsafak.com/blog/the-backup-didnt-come-up-an-hour-by-hour-log-of-an-outage/
- Language: en-US
- Author: Muhammet Şafak

---
> **TL;DR:** On the morning of October 8, 2026, all our systems, internal ones included, went down in four minutes. The Istanbul setup's backup was in Ankara, but it didn't come up. It took about seven hours; the inventory, the priority definition, and the company working as one team made the difference.

On the morning of October 8, all of our systems, internal processes included, crashed within four minutes. The setup in Istanbul had a backup in Ankara, but when we needed it, that backup didn't come up. By around four in the afternoon, the minimum set of applications was back on its feet; this post is a log of that day, hour by hour, from the desk where I was sitting.

## Between 08:56 and 09:01

At 8:56 I was using the systems. I had connected to the database and set up my environments for the work I had planned for the day. Everything was fine, nothing was slow. I remember thinking, "What could possibly happen in four minutes?"

At 9:00 in the morning, the phones of technical support and the customer representatives froze at the same moment. The news reached the software team at the speed of sound, because phones freezing at this scale wasn't normal. A few people tried opening the site and the app: none of it worked. Between 9:00 and 9:01, the whole system, including what we use for our internal processes, had gone down one piece after another.

## First: is it an attack or not

The first thing that came to my mind was that access to the servers might have been cut off by an attack. So I looked at the network traffic logs first. There was no unusual spike; what I saw was the ordinary morning rush that happens every day.

Once the attack possibility was out, we started tracing the whole flow, beginning with the DNS records. Everything was as it should be, and nobody had changed anything. This elimination paid off: all the arrows now pointed to the servers and the infrastructure side.

I have to note my own assumption here. It hadn't crossed my mind that, on a system monitored 24/7, the team might not know about the outage. When no word came from the infrastructure side, I chose to rule out our own side rather than walk straight over to them.

## 09:10: not the servers, the backbone

At first contact, the picture became clear right away. The infrastructure team had already noticed the situation and started working. The problem wasn't individual servers; the entire backbone had collapsed. This meant it was not something that would be fixed in a short time.

At 09:15, a few colleagues informed management, sales, and support. The message was clear: throughout this whole process, the company had to act as a single team. The software and infrastructure teams were the ones who had to solve the problem, but the real burden would be carried by sales and support, who talk to the customers. At the same time, other colleagues checked every development made in the last 72 hours. Whatever the cause was, it wasn't in the application code.

For a broader piece on how a company behaves in a crisis, you can read my post [Crisis Is Where a Company's Character Shows](/blog/crisis-is-where-a-companys-character-shows/).

## 10:00: waiting for plan B

Our servers were in a data center in Istanbul. For situations like this, natural disasters included, there was a clone of the systems in the data center in Ankara. It was a setup that had never been needed so far, kept as plan B. The infrastructure team said that bringing everything up properly could take one to two hours, and got to work.

Waiting was the hardest part. The tension grew by the minute, and we had nothing in sight to work on. I couldn't stand it any longer, turned to the colleague with full access to production, and asked two questions.

- "Is the database up?" Answer: "No."
- "Are the replicas up, and when was the last backup?" After a quick check, the answer: "The replicas are running."

These two answers were the first step toward being able to do something other than wait: we had data in hand that we could build preparations on.

## 10:30: the crisis table

When I learned the replicas were up, I suggested we prepare to bring the applications up on another server. A minute later, the whole team had gathered around the table. The more experienced member of the team took the lead at the board.

The first job was the inventory: a complete list of projects and services. Then we picked the priority ones. After that, we put together a full dependency list for our main projects: the Message Broker, the Worker, and all the connected microservices. We roughly sketched the relationships between them on the board.

In my view, what helps most in a crisis isn't a clever idea; it's everyone looking at the same board.

## 11:30: the crisis grows

The infrastructure team reported that they couldn't bring up the servers in Ankara. Customers were rightly furious, and that fury was flowing into the support and sales teams. The load was on the shoulders of our colleagues at the other end of the phone.

As the software team, I started looking for a suitable data center to put our own plan into action. AWS or GCP wasn't an option because of regulations. My priority was a short-term solution, since the damage had already grown as much as it could. At the same time, most of the team took on another job: analyzing whether anything was forgotten in some corner that a microservice depended on.

## 12:30: defining the minimum product

A meeting request came from the infrastructure team, and we joined without wasting time. A few executives were in the meeting too. The infrastructure team had set up new servers completely from scratch and was trying to bring up the applications and microservices. The reason we were there was clear: to report in real time which systems and services had to be up at the minimum level, and to verify live that the application was working correctly.

The priority definition was clear as well. Customers being able to use everything fully came first. The services we use for our own internal processes were not urgent.

Around 16:00, the minimum-level applications were fully up. About seven hours had passed since 9:00 in the morning.

> **Tan:** In this incident, the backup had its first trial on the most expensive day. Trying a switch-over on a calm day takes time too, but there the cost of failing is only that time. I can say a plan really exists only once someone other than the person who wrote it can run it too.

## A backup that has never been used is an unproven backup

Problems like this are in fact normal for internet applications. Because technical teams know this, they mostly prepare with fallback plans, like a plan B and a plan C. What we went through that day wasn't a hack or an attack; it was a perfectly ordinary problem.

What wasn't ordinary was that plans B and C, in other words our fallbacks, didn't work. I don't want to blame anyone here. The backup system existed; but it had never kicked in before, and on the first day it was supposed to, it didn't come up. What I'm criticizing isn't a team or a person; it's a practice: trusting a plan that has never been put into action as if it had been.

That day also showed how right I was about the three claims I've long defended.

1. **A backup existing doesn't mean a backup working.** "We have a clone in Ankara" is a statement of existence, not a statement of capability.
2. **The inventory and the dependency list must be ready before the crisis starts.** I had always said this; that day we had to put the list together in the middle of the crisis, at the board.
3. **The minimum product must be defined in advance.** The decision "customers can use everything fully, internal services wait" shouldn't be made in the moment of crisis; it should already be written on a page. That day we had to make this decision in the middle of a meeting.

Think about the last time your own backup was actually brought up. If your answer is "never" or "I don't remember", that backup is, for you right now, an assumption.
