A WebSocket that works on a laptop and fails in production almost always fails at a hop in the middle: a load balancer, an ingress proxy, a NAT gateway or a corporate proxy. The symptoms are frustratingly vague. The handshake returns 400 or 502, connections close every sixty seconds, users are disconnected during each deploy, or one server ends up holding most of the sockets. The client usually sees close code 1006, which RFC 6455 reserves for a connection that ended without a close frame, so the code says only that something in the path hung up.

This page is a field guide organised by symptom. For each one it explains the mechanism, how to prove which hop is responsible, and the configuration that fixes it on the common proxies. The architectural choices, layer 4 or layer 7, algorithms, draining and reconnect storms, are in load balancing WebSockets. Everything here assumes those choices are made and something is still going wrong.

Symptom: the handshake never upgrades

A WebSocket starts as an HTTP/1.1 GET with Upgrade: websocket and Connection: Upgrade, and the server answers 101 Switching Protocols. Upgrade and Connection are hop-by-hop headers: a proxy consumes them and does not forward them unless told to. nginx documents exactly this, so a plain proxy_pass delivers an ordinary GET to the backend, which replies 400 or 426, or serves the page. Before nginx 1.29.7 the upstream connection also defaulted to HTTP/1.0, which cannot upgrade at all, hence the proxy_http_version line in every older example.

The second cause is protocol mismatch on the backend leg. WebSockets over HTTP/2 use a different mechanism, extended CONNECT from RFC 8441, and a proxy that speaks HTTP/2 to the browser may still need HTTP/1.1 to the backend for the upgrade to work. When a handshake fails, test it hop by hop with curl, first against the backend directly, then against each proxy in turn; the first hop that returns something other than 101 is the culprit.

# Handshake test through the whole path: expect "HTTP/1.1 101 Switching Protocols"
curl -i -N --http1.1 \
  -H "Connection: Upgrade" -H "Upgrade: websocket" \
  -H "Sec-WebSocket-Version: 13" -H "Sec-WebSocket-Key: dGhlIHNhbXBsZSBub25jZQ==" \
  https://chat.example.com/ws/

# AWS ALB idle timeout (default 60, range 1-4000)
aws elbv2 modify-load-balancer-attributes --load-balancer-arn "$ALB_ARN" \
  --attributes Key=idle_timeout.timeout_seconds,Value=900

# AWS NLB TCP idle timeout per listener (default 350, range 60-6000)
aws elbv2 modify-listener-attributes --listener-arn "$LISTENER_ARN" \
  --attributes Key=tcp.idle_timeout.seconds,Value=900

Symptom: drops at a fixed interval

Every hop between client and server keeps its own idle timer, which closes a connection that has carried no bytes for some number of seconds. The effective limit of the path is the minimum over all hops, including hops you do not control, such as an office NAT. The diagram shows the chain and the two different limits that matter.

Every hop has its own idle clock; the shortest one decidesBrowserping every N sCorp proxy / NATunknown, measureCloud balancerALB idle 60 sIngress proxynginx read 60 sApp serversends pingsTwo different limits per hopIdle timeout: closes a connection that carried no bytes for T seconds. Pings reset it.Maximum lifetime: closes the connection at T seconds even if busy. Pings do not help.Effective idle limit = min over hops. Ping interval should be well under half of it.Symptom: drops every 60 s exactlyA hop idle timeout with no trafficFix: pings below the minimum,then raise timeouts you controlSymptom: drops at a fixed age, busy or notA maximum lifetime or a timeoutthat ignores activityFix: raise it, and reconnect gracefullyClose code 1006 at the client means the TCP connection ended without a WebSocket close frame
Idle timeouts are reset by any traffic, including ping frames. Maximum lifetimes are not, so the second column of symptoms needs a different fix.

The defaults are short. AWS Application Load Balancers default to a 60 second idle timeout, configurable from 1 to 4000 seconds. AWS Network Load Balancers default to a 350 second TCP idle timeout, configurable per listener from 60 to 6000 seconds since September 2024. nginx closes an upgraded connection if the backend sends nothing for 60 seconds, the proxy_read_timeout default. HAProxy uses timeout tunnel once a connection is upgraded, replacing the client and server timeouts that applied during the handshake.

The fix has two halves and you need both. First, send traffic often enough that no hop ever sees an idle connection. Ping frames from the server are the cheapest option, because they do not touch application code on the client; their design is covered in WebSocket ping and pong keepalives. Second, raise the timeouts you control comfortably above the ping interval, so a delayed ping does not kill a healthy connection. A small script makes the arithmetic explicit and keeps it in version control next to the proxy configuration.

HOPS = [  # (name, idle_timeout_s or None, max_lifetime_s or None)
    ("office NAT (measured)", 300, None),
    ("AWS ALB idle_timeout", 60, None),
    ("nginx proxy_read_timeout", 60, None),
    ("app server", None, None),
]

def plan(hops, safety=3):
    idle = [t for _, t, _ in hops if t]
    life = [(n, m) for n, _, m in hops if m]
    floor = min(idle)
    ping = max(5, floor // safety)
    print(f"effective idle limit {floor}s -> ping every {ping}s")
    for name, m in life:
        print(f"{name}: connections end at {m}s regardless; client must reconnect")
    return ping

plan(HOPS)   # effective idle limit 60s -> ping every 20s

Symptom: drops at a fixed age, even when busy

If connections die at the same age whether they are idle or carrying messages every second, pings will not help, because the limit is a maximum lifetime or a timeout that ignores activity. Google Cloud is the clearest example, and its three Application Load Balancer families behave differently. On the classic Application Load Balancer, WebSocket connections, idle or active, close when the backend service timeout expires, and that timeout defaults to 30 seconds. On the global external Application Load Balancer, active WebSocket connections ignore the backend service timeout and are closed after 24 hours. The regional external Application Load Balancer closes idle connections at the backend service timeout, but not active ones.

So a chat app moved behind a classic Google Cloud balancer with default settings disconnects every 30 seconds no matter how many pings it sends. Raising the backend service timeout fixes that, but every platform has some ceiling, and proxies also restart during their own maintenance. Treat a connection's end as normal: reconnect with jittered backoff, resume from the last acknowledged message, and spread planned reconnects rather than letting every client hit the lifetime limit at once. Reconnection design is covered in WebSocket reconnection strategies.

Measure the path, do not guess it

Documentation tells you about hops you know about. A probe tells you about the path as it really is, including the hotel Wi-Fi and the antivirus proxy. Run the script below once silently and once with traffic every ten seconds. A connection that dies at the same age in both runs has hit a lifetime limit; one that survives the busy run but dies in the quiet one has hit an idle timeout, and the age tells you which hop.

import asyncio, sys, time
import websockets
from websockets.exceptions import ConnectionClosed

async def probe(url, send_every=None):
    """Hold one connection open and report how long it lived and how it ended."""
    t0 = time.monotonic()
    # ping_interval=None disables the library's own keepalive so we measure the path
    async with websockets.connect(url, ping_interval=None) as ws:
        try:
            while True:
                if send_every:
                    await ws.send("tick")
                    await asyncio.sleep(send_every)
                else:
                    await ws.recv()
        except ConnectionClosed as e:
            code = e.rcvd.code if e.rcvd else None   # None: no close frame, reported as 1006
            print(f"closed after {time.monotonic() - t0:.0f}s close_code={code}")

if __name__ == "__main__":
    # idle run, then busy run: if both die at the same age it is a lifetime limit
    asyncio.run(probe(sys.argv[1]))
    asyncio.run(probe(sys.argv[1], send_every=10))

Run it from the places your users are, not just from inside the VPC: a corporate network, a mobile carrier, a home router. The answer is often a hop you never configured, which is exactly why the ping interval has to sit below every timeout you can find, not just the balancer's. For packet-level work, WebSocket debugging tools covers browser devtools, Wireshark and proxies that show frames.

Symptom: users drop on every deploy

When a target is removed from a balancer, existing connections are either cut immediately or allowed to drain for a deregistration delay. On an AWS target group that delay defaults to 300 seconds, after which remaining connections are closed. For HTTP requests five minutes is generous; for WebSockets it is an arbitrary cut-off that disconnects every remaining user at the same instant, producing a reconnect wave on the surviving servers.

The robust pattern is to let the application close connections itself, a few at a time with close code 1001, Going Away, during a window shorter than the balancer's delay, so the balancer's cut only catches stragglers. Make sure the process supervisor gives the server long enough before sending SIGKILL. A Kubernetes terminationGracePeriodSeconds of 30, the default, ends a drain that was planned to take two minutes.

Symptom: one server holds most connections

Round robin balances new connections, not open ones. After a deploy or a scale-out, the newest server receives a fair share of new handshakes but none of the thousands of connections that already exist, and it can stay underloaded for hours while older servers run hot. Use least connections at the balancer where it is available, nginx least_conn and HAProxy leastconn below, and shed load deliberately from hot servers by closing a small fraction of their connections so they reconnect elsewhere.

Sticky sessions are a separate question. A pure WebSocket needs no stickiness, because after the upgrade all traffic rides one TCP connection to one server. Libraries that start with HTTP long polling and upgrade later, such as Socket.IO, do need it, because the polling requests and the upgrade must reach the same server. Turning on stickiness for a pure WebSocket service only makes rebalancing harder.

Symptom: health checks pass, sockets fail

A balancer health check usually requests an HTTP path and expects 200. It says nothing about whether the server can accept another WebSocket: the event loop may be saturated, file descriptors exhausted or the upgrade path broken by a bad deploy. Make the health endpoint report the things that actually limit sockets, such as open connections against a configured maximum and event loop lag, and fail it on purpose while draining.

Exhaustion can also happen at the proxy. A proxy opening connections to one backend address and port is limited by its ephemeral port range for that pair; Linux defaults to 32768 to 60999, about 28,000 ports, so a single nginx talking to a single backend address tops out near that many concurrent WebSockets no matter how much memory it has. More backend addresses, more proxy source addresses or a wider ip_local_port_range raise the ceiling.

Proxy configurations that work

The nginx block below forwards the upgrade only when the client asked for one, balances by connection count, and raises the read and send timeouts above any sensible ping interval.

map $http_upgrade $connection_upgrade {
    default upgrade;
    ''      close;
}
upstream ws_backend {
    least_conn;                       # long-lived connections: count, not requests
    server 10.0.1.10:8080;
    server 10.0.1.11:8080;
}
server {
    location /ws/ {
        proxy_pass http://ws_backend;
        proxy_http_version 1.1;       # needed before nginx 1.29.7
        proxy_set_header Upgrade $http_upgrade;
        proxy_set_header Connection $connection_upgrade;
        proxy_set_header Host $host;
        proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
        proxy_read_timeout 3600s;     # default 60s
        proxy_send_timeout 3600s;
    }
}

The HAProxy equivalent relies on timeout tunnel. Keep timeout client-fin short so half-closed sockets from vanished clients do not sit for the full tunnel timeout.

defaults
    mode http
    timeout connect 5s
    timeout client  30s       # applies during the HTTP handshake
    timeout server  30s
    timeout tunnel  1h        # replaces both once the connection is upgraded
    timeout client-fin 30s    # do not hold half-closed sockets for an hour

backend ws
    balance leastconn
    option httpchk GET /healthz
    server ws1 10.0.1.10:8080 check
    server ws2 10.0.1.11:8080 check

On Kubernetes, note that the community ingress-nginx controller was retired: the Kubernetes project ended best-effort maintenance in March 2026 and recommends migrating to Gateway API. The underlying rules do not change with the controller; whichever one you move to, find its equivalents of the read timeout, the upgrade handling and the load balancing algorithm before migrating a WebSocket service.

Trade-offs

Long timeouts keep quiet connections alive but also keep dead ones around: a client that vanished without a FIN occupies a slot until the timeout or a failed ping notices. Short ping intervals detect dead peers quickly and survive aggressive middleboxes, but every ping is a wake-up for a mobile radio and a packet per connection per interval at the server; at a million connections and a 20 second interval that is 50,000 pings per second. Layer 4 balancing avoids most header and timeout surprises but loses path routing and TLS termination at the balancer. Choose the shortest hop you must survive, ping at a third of it, and set every hop you control to several times that.

What to do next

  1. Draw your real path, every hop from browser to process, and write down each idle timeout and maximum lifetime from documentation.
  2. Run the probe from at least three outside networks, idle and busy, and record the age and close code of each drop.
  3. Set the ping interval to a third of the shortest idle timeout you found, and raise every timeout you control well above it.
  4. Check for lifetime limits, especially on Google Cloud classic balancers, and make the client reconnect and resume cleanly.
  5. Drain from the application with close code 1001 inside the balancer deregistration delay and the pod grace period.
  6. Switch to least connections, add connection count to the health check, and alert on proxy ephemeral port usage.
Key takeaway: WebSocket failures behind load balancers come from a few mechanisms: proxies that do not forward the upgrade, idle timeouts on every hop, maximum lifetimes that ignore activity, deregistration delays that cut connections during deploys, balancing by new connections only, and port or descriptor exhaustion. Measure the path with a probe, ping at a third of the shortest idle timeout, raise the timeouts you control, and design the client to reconnect and resume because some limit always exists.