Authenticating a WebSocket is awkward for one structural reason: the browser's WebSocket constructor accepts a URL and a list of subprotocols and nothing else, so the familiar Authorization header is not available to browser code. Every pattern in use is a way of working around that constraint, and each moves the credential to a different layer with different logging, replay and expiry behaviour.

The bidi security article covers the threat model: cross-site hijacking, origin checks, per-message authorisation, revocation and limits. This article is the implementation companion. It answers which layer should authenticate for each kind of client, then builds the main patterns properly: a one-time ticket service, first-frame authentication, and an in-band refresh protocol so long-lived connections do not outlive their credentials. It closes with how four common stacks implement the same ideas and a worked migration away from long-lived tokens in URLs.

Where authentication can live

A WebSocket connection passes through up to four points where a credential can be checked. The edge or gateway sees the TLS session, client certificates, IP address and request headers. The upgrade handler in your application sees the HTTP GET that asks to switch protocols, including cookies, the query string and the Origin header. After the switch, the first application frame can carry credentials in your own protocol. Finally, every later message must be authorised against the identity established earlier.

Four places a WebSocket can be authenticated, and what each one can seeClientbrowser, app, service1. Edge / gatewayTLS, mTLS, authorizer2. Upgrade handlerHTTP GET + Upgrade3. First frameCONNECT / auth message4. Every messageauthorise action + resource with the bound identitysees: cert, IP, headerssees: cookies, query, Originsees: app payloadConnection recorduser id, tenant, scopescredential expiry (exp)auth method, ticket idnode id for revocation fan-outRefresh loopclient sends reauth before expserver validates, moves exp forwardtimer fires at exp: close 4001revocation event: close 4003bind onceAuthenticate once, as early as the client allows; keep the result on the connection; re-check it on a clock
Authentication happens once, at the earliest layer that can see a credential; the result is stored on the connection and re-checked on a timer and on revocation events.

The design rule is to authenticate at the earliest layer the client can support, because each later layer means holding an unauthenticated socket open for longer. Rejecting at the upgrade costs one HTTP response. Rejecting after the first frame means you have already allocated a connection, a buffer and probably a file descriptor for an anonymous peer, which is why first-frame authentication needs a strict deadline.

Choosing a pattern by client type

ClientRecommended patternWhy
Browser, same site as the APISession cookie at upgrade + exact Origin allowlistCookie is sent automatically; Origin check stops cross-site hijacking
Browser SPA with bearer tokensOne-time ticket in the query stringNothing long-lived appears in URLs or logs
Browser via a framework (Socket.IO)Token in the framework's connect payloadCredential travels inside the protocol, not the URL
Mobile or desktop appAuthorization header at upgradeNative clients can set headers; reuse API validation
Service to servicemTLS at the edge, or a bearer headerWorkload identity without user tokens
Devices behind a gatewayGateway authorizer on connectReject before your servers spend anything

Two patterns are missing on purpose. Long-lived access tokens in the query string end up in load balancer logs, proxy logs, browser history and error trackers, so the credential outlives its purpose in a dozen places. Smuggling a token through Sec-WebSocket-Protocol works but misuses a negotiation header that servers echo and proxies log; WebSocket subprotocols explains the trade-offs if you inherit it.

Building a one-time ticket service

A ticket is a short random string that stands in for the user's real credential for a few seconds. The page calls an ordinary HTTPS endpoint, authenticated the way the rest of the API is, and receives a ticket. It then opens wss://rt.example.com/ws?ticket=.... The upgrade handler redeems the ticket and attaches the stored identity to the connection. Because the ticket is single-use and expires in about thirty seconds, a copy in a log is useless by the time anyone reads it.

// Ticket issue and redemption (Node, node-redis v4). GETDEL needs Redis 6.2+.
import crypto from "node:crypto";

const TICKET_TTL_S = 30;

// POST /ws-ticket -- behind your normal session or bearer authentication
export async function issueTicket(req, res) {
  const ticket = crypto.randomBytes(32).toString("base64url");
  const claims = { sub: req.user.id, tenant: req.user.tenant, scopes: req.user.scopes,
                   origin: req.get("Origin"), exp: req.user.tokenExp };
  await redis.set(`wst:${ticket}`, JSON.stringify(claims), { EX: TICKET_TTL_S });
  res.set("Cache-Control", "no-store").json({ ticket });
}

// Called from the HTTP server's "upgrade" event, before handleUpgrade
export async function redeemTicket(req) {
  const ticket = new URL(req.url, "http://x").searchParams.get("ticket");
  if (!ticket || ticket.length > 64) return null;
  const raw = await redis.getDel(`wst:${ticket}`);         // atomic: a second use finds nothing
  if (!raw) return null;
  const claims = JSON.parse(raw);
  if (claims.origin !== req.headers.origin) return null;   // bound to the page that asked for it
  return claims;
}

Four details make the difference between a ticket and a weaker token. Redemption is atomic: GETDEL reads and deletes in one command, so two concurrent upgrades cannot both succeed. The ticket is bound to the origin of the page that requested it, so a ticket stolen by a script on another site fails the comparison. The claims carry the expiry of the underlying access token, so the connection inherits a deadline instead of living forever. And the issuing response is no-store so no cache keeps it.

The stateless alternative is a signed ticket: an HMAC or a short JWT with a few seconds of lifetime and a unique id. It removes the Redis round trip but cannot be single-use without storing the ids it has seen, at which point you have rebuilt the stateful version. Use signed tickets when the issuing and redeeming services cannot share a store, keep their lifetime under ten seconds, and accept that a replay inside that window is possible.

First-frame authentication

Some environments cannot do anything at the upgrade: a gateway that does not forward query strings, or a protocol such as STOMP or a GraphQL-over-WebSocket protocol that defines its own connect message. In those cases the connection opens anonymously and the first frame carries the token. The connection is a state machine with two states, unauthenticated and authenticated, and the unauthenticated state must be tiny.

  • Arm a deadline at accept time, typically five seconds, and close with code 1008 (policy violation) if no valid auth frame arrives.
  • Accept exactly one frame type while unauthenticated; anything else closes the connection.
  • Apply a small maximum frame size before authentication, a few kilobytes, separate from the larger limit after.
  • Do not subscribe the socket to any topic, room or broadcast until the auth frame has been validated.
  • Count unauthenticated sockets per IP and globally, and shed load when the count spikes.

The cost is that the origin check and the credential check now happen at different times. Keep the Origin allowlist at the upgrade regardless, so browsers on other sites are rejected before they can even attempt the first frame.

Keeping long-lived connections honest: in-band refresh

Access tokens usually live for minutes; dashboard and chat connections live for hours. Without a refresh path you either let the connection outlive its credential or force a reconnect every few minutes, which drops in-flight state and multiplies handshake load. An in-band refresh lets the client present a new token over the existing socket. The protocol has three messages and one timer.

// In-band refresh on the server: the connection keeps its identity, only exp moves.
function armExpiry(ws) {
  clearTimeout(ws.expiryTimer);
  const ms = ws.identity.exp * 1000 - Date.now();
  ws.expiryTimer = setTimeout(() => ws.close(4001, "credential expired"), Math.max(ms, 0));
}

ws.on("message", async (data) => {
  const msg = JSON.parse(data);
  if (msg.type === "reauth") {
    const claims = await verifyAccessToken(msg.token);      // same validation as your HTTP API
    if (!claims || claims.sub !== ws.identity.sub) return ws.close(4003, "identity mismatch");
    ws.identity.exp = claims.exp;
    ws.identity.scopes = claims.scopes;                      // scopes may shrink; honour it
    armExpiry(ws);
    return ws.send(JSON.stringify({ type: "reauth_ok", exp: claims.exp }));
  }
  if (!authorise(ws.identity, msg)) return ws.send(JSON.stringify({ type: "error", code: "forbidden" }));
  handle(ws, msg);
});

The client schedules a refresh at about 80 percent of the token lifetime, obtains a new access token through its normal refresh flow, and sends {type: "reauth", token}. The server validates it exactly as the HTTP API would, JWT validation included, insists that the subject has not changed, replaces the scopes and moves the timer. If the refresh never arrives the timer closes the socket with a private code in the 4000-4999 range, and the client knows to obtain a fresh ticket rather than retry blindly. Note the scope replacement: if an administrator removed a permission, the next refresh is where the connection loses it.

How common stacks implement it

Frameworks wrap these ideas in their own vocabulary. The mapping matters because each one places the credential at a different layer, which changes what appears in logs and when an invalid token is detected.

// --- Socket.IO: the auth object travels in the CONNECT packet, not the URL ---
const socket = io("https://rt.example.com", { auth: (cb) => cb({ token: getAccessToken() }) });
io.use(async (socket, next) => {
  const claims = await verifyAccessToken(socket.handshake.auth.token);
  if (!claims) return next(new Error("unauthorized"));     // client receives connect_error
  socket.data.identity = claims;
  next();
});

// --- ASP.NET Core SignalR: browsers send access_token in the query string ---
// client: new HubConnectionBuilder().withUrl("/hubs/chat", { accessTokenFactory: getAccessToken })
options.Events = new JwtBearerEvents {
  OnMessageReceived = ctx => {
    var token = ctx.Request.Query["access_token"];
    if (!string.IsNullOrEmpty(token) && ctx.HttpContext.Request.Path.StartsWithSegments("/hubs"))
      ctx.Token = token;
    return Task.CompletedTask;
  }
};

// --- Spring STOMP over WebSocket: authenticate the CONNECT frame ---
registration.interceptors(new ChannelInterceptor() {
  @Override public Message<?> preSend(Message<?> message, MessageChannel channel) {
    StompHeaderAccessor acc = MessageHeaderAccessor.getAccessor(message, StompHeaderAccessor.class);
    if (StompCommand.CONNECT.equals(acc.getCommand()))
      acc.setUser(authenticate(acc.getFirstNativeHeader("Authorization")));
    return message;
  }
});

  • Socket.IO sends the auth object inside its CONNECT packet once the transport is open, so the token stays out of URLs. Passing a function, as above, means every reconnect fetches a current token.
  • SignalR documents sending the token as access_token in the query string for WebSockets and Server-Sent Events because browsers cannot set headers. Restrict the query lookup to hub paths, keep token lifetimes short, and make sure request logging does not record full URLs.
  • Spring STOMP authenticates the STOMP CONNECT frame with a ChannelInterceptor on the client inbound channel and sets the user on the session, which later @MessageMapping methods and destination security rules use.
  • AWS API Gateway WebSocket APIs run authorizers only on the $connect route, using IAM or a Lambda authorizer that can read headers and query string parameters. Persist the resulting identity against the connection id during connect so later routes can look it up.

Many servers, one identity

Once you run more than one connection server, identity becomes distributed state. Keep the authoritative record on the connection object in memory and index it by user id on each node, so a revocation event can be fanned out and each node closes its own matching sockets. Tickets must be redeemable on any node, which is why the store is shared; sticky routing at the load balancer is not an authentication mechanism. Load balancer pitfalls covers idle timeouts and header handling that silently break the upgrade path.

Reconnection interacts with authentication in a way that bites at scale. After a deploy or network blip, every client reconnects, and with tickets every reconnect is first an HTTPS call to the issuing endpoint. Rate limit issuance per user, add jitter to reconnects as described in reconnection strategies, and make sure the issuing service scales with the connection fleet rather than with ordinary API traffic.

Worked example: removing tokens from URLs

A team runs a trading dashboard that connects with wss://rt.example.com/ws?token=<15-minute JWT>. A log review finds those JWTs in the CDN access logs, retained for 90 days. The migration runs in four steps. First, add the ticket endpoint and the redemption path while still accepting token, and emit a metric for each method. Second, ship the client change: fetch a ticket, connect with it, and send reauth at 12 minutes. Third, once the metric shows under one percent of connections using token, reject it at the upgrade with HTTP 401; browser clients only observe a failed connection, commonly close code 1006, so the client falls back to fetching a ticket. Fourth, purge or redact the historical logs and add a log-pipeline rule that strips query strings on the WebSocket path. The result: no credential in any URL lasts longer than 30 seconds, and every connection has a deadline tied to a real token.

Failure modes

SymptomCauseFix
Clients loop reconnecting with no error detailBrowser hides the HTTP status of a failed upgradeExpose a pre-flight auth check endpoint; log rejection reasons server-side
Ticket accepted twiceGET then DEL as two commandsAtomic GETDEL or a Lua script
Connections survive account disableNo expiry timer or revocation fan-outArm a timer at exp; subscribe nodes to revocation events
Spike of anonymous socketsFirst-frame auth without a deadlineFive-second deadline, tiny pre-auth frame limit, per-IP caps
Tokens found in logsQuery-string tokens and default URL loggingTickets, URL redaction, shorter token lifetimes
Permissions removed but still effectiveScopes captured at connect, never refreshedReplace scopes on reauth; push revocation events

What to do next

  1. List every WebSocket entry point and record which layer authenticates it and what credential it accepts.
  2. Remove long-lived tokens from URLs: introduce a ticket endpoint with atomic, origin-bound, 30-second tickets.
  3. Store the credential expiry on each connection, arm a close timer, and implement an in-band reauth message.
  4. For first-frame protocols, enforce a connect deadline, a pre-auth frame size limit and per-IP unauthenticated caps.
  5. Confirm request logs on the WebSocket path strip query strings, at the CDN, load balancer and application.
  6. Fan out revocation events to every connection node and test disabling a user with an open connection.
  7. Re-read bidi security to confirm per-message authorisation and Origin checks are in place.
Key takeaway: Browsers cannot send an Authorization header on a WebSocket, so pick the earliest layer each client can use: cookies plus an Origin allowlist for same-site pages, one-time origin-bound tickets for token-based SPAs, headers or mTLS for native and service clients, and a strict deadline when only the first frame can carry credentials. Store the identity and its expiry on the connection, refresh it in-band, close on expiry and revocation, and keep credentials out of URLs and logs.