# How do I do graceful shutdown on SIGTERM during autoscaling scale-in?

> On SIGTERM drain the load balancer first, finish in-flight work inside the server.Shutdown(ctx) deadline, then align the grace period with that drain.

- Asked: 2026-05-18
- Answered: 2026-05-21
- Asked by: Ozan
- Tags: performans, olcekleme, go
- Source: https://muhammetsafak.com/just-ask/graceful-shutdown-on-sigterm-for-autoscaled-apis/
- Language: en-US
- Author: Muhammet Şafak

---
**Question:** On AWS, Autoscaling spins up new EC2 instances when CPU goes above 70% and shuts some down when it drops below 30% (scale-in). When an instance is terminated, long-running HTTP requests or queue jobs in flight get cut off mid-way and leave inconsistency behind.

How do I make the app (Laravel Octane / Go) catch `SIGTERM`, finish the in-flight requests, and shut down gracefully by refusing new ones?


Short answer: scale-in cuts work off mid-flight because the app doesn't drain on `SIGTERM` — what you need to fix is the process's shutdown lifecycle.

## Short answer

The real issue isn't autoscaling, it's the shutdown protocol: when the platform sends SIGTERM, your app either dies instantly or gets cut off before finishing. The right order is: stop the traffic first, finish in-flight work next, die last. I covered why Octane's persistent-process model changes this behaviour in the [Laravel Octane post](/blog/laravel-octane-persistent-process-performance/).

## Why

1. **Shutdown is a protocol, not an event.** Three things have to happen when the "die" signal arrives, and the order matters; skip two and go straight to the third and you corrupt data.

2. **The platform's patience is finite.** SIGKILL follows SIGTERM once the grace period expires; if your drain is longer than that window, work still gets cut off.

3. **No drain window is a 100% guarantee.** The process can be hard-killed, so the last line of defence has to live in the code itself.

## What to do

1. **Catch SIGTERM and stop accepting new requests.** Fail your readiness probe or **deregister** from the load balancer so it stops routing new requests to you. Don't close existing connections yet.

2. **In Go, use `server.Shutdown(ctx)`.** Catch SIGTERM with `signal.NotifyContext` and hand `http.Server` a context with a deadline. It refuses new connections and lets open requests finish until the timeout.

3. **In Octane, use a graceful stop/reload.** Workers finish the current request and stop claiming new ones. On the queue side, `queue:work` handles SIGTERM itself as long as you don't `kill -9` it.

4. **Align the platform's grace period with your drain.** Use an ASG lifecycle hook or k8s `terminationGracePeriodSeconds` to match the platform's wait to your longest acceptable drain time.

5. **Make long jobs idempotent/resumable.** Every step idempotent, so an interrupted job can safely restart or resume.

**Bottom line:** order matters — drain the LB first, then stop new intake, then finish in-flight work, then die. In Go that's `server.Shutdown(ctx)`; in Octane it's a graceful reload plus `queue:work`'s native SIGTERM behavior; on the platform side, align the grace period with your real drain time. Make jobs idempotent on top of that and even a hard kill won't corrupt data.

## Related Reading

- [Laravel Octane: the performance that comes with a persistent process](/blog/laravel-octane-persistent-process-performance/) — Blog
- [How do I cut GC pressure in Go with sync.Pool and escape analysis under load?](https://muhammetsafak.com/just-ask/go-gc-pressure-sync-pool-and-escape-analysis-under-load/) — Just Ask
- [Does my channel buffer size choice hide backpressure problems, and how do I decide it?](https://muhammetsafak.com/just-ask/channel-buffer-size-choice-hide-backpressure-problems-decide/) — Just Ask
- [My Go worker in a container never receives SIGTERM and won't shut down gracefully—is this a PID 1 problem?](https://muhammetsafak.com/just-ask/go-worker-never-receives-sigterm-pid-1-problem/) — Just Ask
