The NestJS docs handle shutdown with a fixed delay. On my deploys, every fixed wait lost requests. What reached zero was waiting until the traffic actually stopped. A later bug taught me to time the exit itself.
The NestJS docs handle shutdown with a fixed delay. On my deploys, every fixed wait lost requests. What reached zero was waiting until the traffic actually stopped. A later bug taught me to time the exit itself.
As of September 2026, my NestJS APIs shut down in a set order, and it took five tries to find it. Most of this post is the four that failed. The numbers come from a probe sending about 40 requests a second through real deploys, on Coolify and on k3s.
The NestJS answer is in the Terminus recipeNestJS docs. Set gracefulShutdownTimeoutMs, and on SIGTERM the health check reports shutting_down with a 503. The orchestrator stops sending new traffic, and the app finishes what is in flight during the delay. The docs suggest a delay slightly longer than the readiness check interval. On Kubernetes, the usual companion advice is a preStop hook that sleeps for a few seconds.
Both are a fixed wait. You guess how long the proxy needs to notice, and hope the guess holds on every deploy. My health endpoints are hand-written, not Terminus, so I have not measured its delay itself. I measured the same idea in other forms.
Before any of this, Coolify's rolling deploy on its own lost 3 requests on each of four deploys, all within about a second of the old container's last response. About 99.6% of requests got through.
Keep serving for N seconds, then close. The textbook drain, and it still lost about 3 requests a deploy. Traefik, the proxy in front, only dropped the old container when it left Docker. So it kept routing there for a second or two after the listener closed, and those requests did not fail fast. A dial into a container that is about to vanish hung for 5 seconds or more.
Send Connection: close while draining. No change. The lost requests were new connections, not reused ones. I kept the header because it is correct, not because it helped.
Close the listener as soon as SIGTERM arrives. Better: a fast refusal instead of a 5-second hang. Still 13 fast 502s over about 2 seconds, while a healthy container beside it had been serving for 9.6 seconds. Traefik kept sending a share of traffic at the closed one and did not retry it on the other.
Point proxy health checks at readiness. Across six deploys, Traefik stopped routing 1.2 to 2.0 seconds after readiness went false on three. On the other three it never stopped. A deploy fires a burst of Docker events, each one rebuilds the proxy config, and a rebuilt service starts out healthy again.
Every one of these picks a moment to close. The moment I needed depended on when the proxy's health checker happened to fire, and the process has no way to know that in advance.
So the process stopped predicting and started watching. When the proxy has moved on, the traffic stops, and the process can see that from where it sits.
On SIGTERM, readiness answers 503 and liveness keeps answering 200. They are separate endpoints on purpose: a liveness probe wired to the readiness answer is how a draining pod gets killed mid-request. The process keeps serving and notes when the last real request arrived. Health probes are left out of that clock, because they never stop.
Once 2 seconds pass with no real request, it closes the listener and drops idle keep-alive sockets. Requests already running finish. After a tail of at most 2 seconds, Nest re-raises the signal and the process exits. If the traffic never goes quiet, it closes anyway at 15 seconds, because a drain that waits forever is a hung shutdown.
On Kubernetes the rest is config. Readiness is polled every 2 seconds and fails after 2 misses. The rollout uses maxUnavailable: 0 and maxSurge: 1, and the grace period is 60 seconds, well above the 15-second deadline. There is no preStop sleep: the drain already waits as long as it needs, and no longer.
Kubernetes has the same race Coolify has, because kube-proxy removes an endpoint asynchronously. The same drain closed it on both, which is why I think of this as application work, not platform work.
The next failure came later, in a different Nest API of mine. An agent added process.on('SIGTERM', ...) next to Nest's enableShutdownHooks(). Nest's own handler runs once, removes itself, and re-raises the signal so Node's default action can end the process. The extra listener was permanent. It caught the re-raised signal, and Node only exits on a signal by default when nothing is listening for it.
The process stopped exiting on SIGTERM. Under node --watch it sat at "Waiting for graceful termination" forever, and a pod waits out its whole grace period before it is killed. All 4,531 tests stayed green, because the agent's proof checked that the port closed, and it did. Nothing checked that the process exited. The orchestrating agent caught it by booting the merged code and timing the exit by hand: 4 seconds before the change, still alive at 25 seconds after it.
The fix was one word, once instead of on, plus a test that process.listenerCount('SIGTERM') is 0 after one signal. I also keep a small script now that boots a process, sends SIGTERM, and prints how many seconds it took to exit, or that it hung.
preStop sleep on either platform. My numbers are about fixed waits in general, not about those two.Unready first, serve until quiet, close, then exit. And I judge a shutdown change by one number: seconds from SIGTERM to the process being gone.
On SIGTERM, return 503 from readiness and keep serving. Liveness stays 200 on a separate endpoint.
Close the listener only after no non-health request has arrived for 2 s. Keep a hard deadline below the orchestrator's grace period.
Never add a second process.on('SIGTERM') beside Nest's enableShutdownHooks(). If you need one, use once, and test that listenerCount('SIGTERM') is 0 after one signal.
Prove a shutdown change by timing the process exit after SIGTERM, not by checking that a handler ran or a port closed.