Graceful Shutdown in Node.js: Why Requests Drop Every Time You Deploy
If in-flight requests get cut off whenever you deploy or restart, you're not handling the shutdown signal (SIGTERM). How to finish in-flight requests with server.close() before exiting, the keep-alive trap that keeps the process alive, and signals that never reach Node in Docker and PM2, all verified by actually running it.
A few errors on every deploy
Everything's fine normally, but every time you deploy or restart the server, users hit a handful of errors. Slow requests, like payments or saving a post, are getting cut off midway.
Let's reproduce it. I sent a request that takes 2 seconds, then sent the server a shutdown signal (SIGTERM) half a second later.
[test] SIGTERM sent (570ms)
[test] server exited signal=SIGTERM (577ms) ← exits the instant the signal arrives
[test] in-flight request: FAILED (UND_ERR_SOCKET) ← the 2-second request is cut off
Without code to handle it, Node.js exits immediately on SIGTERM. In-flight requests, files being written, and uncommitted DB transactions are simply dropped.
A graceful shutdown means that on a shutdown signal, the server stops accepting new requests, finishes the ones in flight, cleans up connections and resources, and then exits.
Where shutdown signals come from
Almost every tool that "stops" a server first sends SIGTERM (a polite request to exit), then SIGKILL (forced) if the process hasn't exited in time. A process can't intercept SIGKILL.
| Situation | Signal sent | Wait before forced kill |
|---|---|---|
| Ctrl + C in a terminal | SIGINT | - |
| docker stop | SIGTERM | 10 seconds by default |
| Kubernetes pod termination | SIGTERM | 30 seconds by default |
| pm2 stop / pm2 reload | SIGINT | 1.6 seconds by default (kill_timeout) |
| Serverless, PaaS redeploys | Usually SIGTERM | Varies by platform |
So our server's job is to shut itself down cleanly within that grace period.
The basic implementation: server.close()
import http from "node:http";
const server = http.createServer(app);
server.listen(3000);
let shuttingDown = false;
function shutdown(signal) {
if (shuttingDown) return; // only once, even if signals repeat
shuttingDown = true;
console.log(`${signal} received, shutting down`);
// 1. Stop accepting new connections; the callback runs when in-flight requests finish
server.close(async () => {
// 2. Clean up external connections like DB and Redis
await db.end();
console.log("Clean exit");
process.exit(0);
});
// 3. If it still hasn't finished, force exit (shorter than the platform's grace period)
setTimeout(() => {
console.error("Timed out, forcing exit");
process.exit(1);
}, 10_000).unref();
}
process.on("SIGTERM", shutdown);
process.on("SIGINT", shutdown);
Running the same experiment again:
[test] SIGTERM sent (567ms)
[server] SIGTERM received, shutting down
[test] new request during shutdown: REFUSED (ECONNREFUSED) ← no new requests
[test] in-flight request: OK "slow done" ← the in-flight one finishes
[server] all connections closed
.unref() on the force-exit timer means "don't keep the process alive just for this timer." If cleanup finishes early, the process exits without waiting for it.
The trap: the request finished, so why is it still running?
With the code above, the last response went out at 2 seconds, but the process didn't exit until after 5 seconds.
[test] in-flight request: OK (2083ms)
[server] all connections closed (5094ms) ← waited 3 more seconds
The cause is HTTP keep-alive. Clients keep the connection open after a response so they can reuse it. server.close() only closes connections that are idle at the moment you call it, and this one was busy handling a request at that moment. So after the response, the now-idle connection lingered until the server's keep-alive timeout (5 seconds by default in Node.js).
In environments with a short grace period (PM2's default is 1.6 seconds), those few seconds can get you force-killed before your cleanup code finishes. The fix: during shutdown, close idle connections as soon as each response finishes.
// During shutdown, clean up idle connections whenever a response finishes
server.on("request", (req, res) => {
res.on("finish", () => {
if (shuttingDown) setImmediate(() => server.closeIdleConnections());
});
});
function shutdown(signal) {
if (shuttingDown) return;
shuttingDown = true;
server.close(() => process.exit(0));
server.closeIdleConnections(); // close connections that are idle right now
setTimeout(() => {
server.closeAllConnections(); // last resort: drop every remaining connection
process.exit(1);
}, 10_000).unref();
}
[test] in-flight request: OK (2088ms)
[server] all connections closed (2091ms) ← exits right after the response
closeIdleConnections() and closeAllConnections() are available from Node.js 18.2.
Docker: when the signal never arrives
If your shutdown code is solid but docker stop always takes 10 seconds and ends in a forced kill, the signal isn't reaching Node.js.
# ❌ Shell form: runs as /bin/sh -c "node server.js"
CMD node server.js
# ✅ Exec form: node becomes the container's PID 1 and receives signals directly
CMD ["node", "server.js"]
With the shell form, the shell (/bin/sh) is PID 1 and Node.js is its child. The shell doesn't forward the SIGTERM it receives, so Node.js never gets the signal and is killed by SIGKILL 10 seconds later.
| How it's started | Signal delivery |
|---|---|
| CMD node server.js (shell form) | ❌ The shell receives it and doesn't forward it |
| CMD ["node", "server.js"] (exec form) | ✅ |
| CMD ["npm", "start"] | ⚠️ npm sits in between. The official Node.js Docker guidance recommends running node directly |
| docker run --init or tini | ✅ Handles signal forwarding and zombie process reaping for you |
PM2: SIGINT and a short grace period
When PM2 stops a process it sends SIGINT, not SIGTERM, and force-kills after 1.6 seconds by default. So take care of two things:
- Handle SIGINT too (like process.on("SIGINT", shutdown) above).
- If requests or cleanup take longer, increase kill_timeout.
// ecosystem.config.js
module.exports = {
apps: [
{
name: "api",
script: "server.js",
kill_timeout: 10000, // wait 10 seconds before force-killing
},
],
};
Behind a load balancer
If several servers sit behind a load balancer, it also helps to make the health check fail as soon as shutdown starts, so the load balancer stops sending traffic.
app.get("/health", (req, res) => {
res.status(shuttingDown ? 503 : 200).end();
});
Summary: why this is worth knowing
| Problem | Fix |
|---|---|
| Dies instantly on the signal, cutting off requests | Handle SIGTERM/SIGINT + server.close() |
| Never exits because cleanup never finishes | A force-exit timer shorter than the grace period (.unref()) |
| Requests are done but it lingers for seconds | Clean up keep-alive connections with closeIdleConnections() |
| DB connections and files left open | Clean up external resources in the server.close callback |
| Always force-killed after 10s in Docker | Exec-form CMD ["node", "server.js"] or --init |
| Force-killed mid-cleanup under PM2 | Handle SIGINT + raise kill_timeout |
| Traffic keeps arriving at a server that's shutting down | Return 503 from the health check |
Deploys happen several times a day. The few requests dropped each time barely show up in the logs, but users remember it as "the service where payments sometimes fail." Put as much care into the code that stops your server as into the code that starts it.