Published on

WebSockets From the Wire Up

Authors
  • Mehdi Akiki avatar
    Name
    Mehdi Akiki
    Twitter

Most WebSocket explanations start with new WebSocket("ws://...") and a diagram of two arrows. That's useful for getting started but it leaves a gap: you can use the API without understanding why it works the way it does, and that gap shows up every time something breaks at the protocol level.

This article starts from the physical layer and works up. By the end, tracing what happens when you call ws.send("hello") from a browser should feel mechanical, not magical.

Contents

  1. The wire: analog signals becoming bits
  2. The kernel wakes up: NIC interrupt → socket buffer
  3. What listen() and connect() actually mean
  4. TCP: a byte stream, not a message protocol
  5. HTTP/1.1 upgrade: turning a TCP connection into a WebSocket
  6. WebSocket framing: the binary format
  7. Tracing ws.send("hello") from browser to server
  8. A working server and client
  9. Connection lifecycle
  10. The close handshake
  11. Ping, pong, and dead connection detection
  12. TLS and WSS
  13. Proxies, nginx, and the Upgrade header problem
  14. Backpressure: what happens when the sender is faster than the receiver
  15. TCP gotchas: Nagle's algorithm and delayed ACK
  16. Comparison: WebSocket vs the alternatives
  17. Security checklist
  18. WebSocket over HTTP/2 (RFC 8441)
  19. Mental model and common failures
  20. Experiments — see it yourself

1. The wire: analog signals becoming bits

Ethernet (the most common physical layer for servers) encodes bits as voltage changes on a copper cable, or as light pulses on fiber. The exact encoding depends on the standard and link speed — 100BASE-TX uses MLT-3, 1000BASE-T uses PAM5, 25G/100G+ links use PAM4 — but the principle is the same: a signal transition at a specific time maps to a 1 or a 0.

The NIC (Network Interface Card) on the receiving end has a clock recovery circuit that samples the incoming signal and produces a stream of bits. This is the boundary between analog and digital. Below this line: physics. Above it: bytes.

Those bytes form an Ethernet frame:

[ 6 bytes: destination MAC ]
[ 6 bytes: source MAC      ]
[ 2 bytes: EtherType       ]  ← 0x0800 = IPv4, 0x86DD = IPv6
[ N bytes: payload         ]  ← an IP packet
[ 4 bytes: CRC checksum    ]

The NIC verifies the CRC, strips the Ethernet header, and needs to get the IP packet to the kernel. It does this via DMA (Direct Memory Access): the NIC writes the packet bytes directly into a pre-allocated ring buffer in kernel memory, then raises a hardware interrupt.


2. The kernel wakes up: NIC interrupt → socket buffer

The interrupt fires. The CPU pauses whatever it was doing, saves its state, and jumps to the NIC's interrupt handler in the kernel. The handler sees that a packet arrived in the DMA ring buffer.

The kernel's networking stack takes over:

  1. IP layer: checks the destination IP address, verifies the header checksum, extracts the protocol field (6 = TCP, 17 = UDP).
  2. TCP layer: finds the right socket by matching (src_ip, src_port, dst_ip, dst_port) — a 4-tuple that uniquely identifies a TCP connection. Each socket has two buffers: a receive buffer (sk_rcvbuf) where incoming bytes pile up, and a send buffer (sk_sndbuf) where outgoing bytes wait to be transmitted.
  3. The payload bytes are appended to the socket's receive buffer. If a userspace process is blocked in read() or recv() on that socket's file descriptor, it is woken up.

What a socket actually is in memory

This is the part most explanations skip. A socket is not an abstraction — it is a concrete data structure in kernel memory.

In Linux, every socket is a struct sock (defined in include/net/sock.h). The fields that matter for understanding WebSocket behavior:

struct sock {
    // The 4-tuple (src_ip, src_port, dst_ip, dst_port) lives here
    struct sock_common  __sk_common;

    // Receive buffer: a linked list of sk_buff (one per TCP segment)
    struct sk_buff_head sk_receive_queue;
    int                 sk_rcvbuf;       // max bytes allowed in receive queue
                                         // default: ~87 KB (net.core.rmem_default)

    // Send buffer: bytes waiting to go out on the wire
    struct sk_buff_head sk_write_queue;
    int                 sk_sndbuf;       // max bytes allowed in send queue
                                         // default: ~212 KB (net.core.wmem_default)

    // Current TCP state: TCP_ESTABLISHED, TCP_LISTEN, TCP_CLOSE_WAIT, etc.
    unsigned int        sk_state;

    // Wait queue: processes sleeping on this socket (blocked in read/write)
    wait_queue_head_t   sk_wq;
};

Each packet in the kernel is a struct sk_buff (skb). It's not a copy of the packet — it's a descriptor with four pointers (head, data, tail, end) into a contiguous memory region. As a packet moves up the network stack, headers are stripped by advancing data forward. As it moves down, headers are prepended by moving data backward. In the common path, no copying happens — the kernel does pointer arithmetic rather than memcpy. Copies do occur in specific cases: cloning for multicast delivery, retransmission of TCP segments, and some hardware offload paths.

head → [ headroom for headers ]
data → [ IP header ][ TCP header ][ payload bytes ]
tail →
end  → [ tailroom ]

The socket is exposed to userspace as a file descriptor — just an integer. The chain is:

process fd table
  └─ int fd (e.g. 5)
       └─ struct file  (kernel VFS object)
            └─ struct socket  (POSIX socket layer)
                 └─ struct sock  (protocol-specific: TCP state, buffers, 4-tuple)

Every socket is a file in Linux. read(), write(), select(), epoll() all work on file descriptors — they don't know or care that fd 5 is a TCP socket. The VFS layer dispatches read() to the socket's receive queue. This is not an abstraction convenience; it's the actual kernel design, which is why you can do:

import socket, os
s = socket.socket()
s.connect(("example.com", 80))
# s.fileno() returns the raw integer fd
os.write(s.fileno(), b"GET / HTTP/1.0\r\n\r\n")
data = os.read(s.fileno(), 4096)

The os.write and os.read calls on a socket fd go through the exact same kernel path as writing to a file.


3. What listen() and connect() actually mean

Before TCP, before WebSocket — before any of it — you need to understand what these two syscalls do at the kernel level. These are words engineers use every day without a concrete picture. Here is the picture.

listen() — opening a door in kernel memory

When a server calls listen(fd, backlog), the kernel does two things:

First, it changes the socket's state from CLOSED to LISTEN and registers the 4-tuple (0.0.0.0, port, *, *) in the kernel's hash table of listening sockets. This is the "door". Any TCP SYN packet arriving on that port will be matched against this table.

Second, it allocates two queues in kernel memory:

Listening socket
├── SYN queue (incomplete connections)
│   └── half-open connections: SYN received, SYN+ACK sent, waiting for ACK
│       [ slot ] [ slot ] [ slot ] ...   ← size: /proc/sys/net/ipv4/tcp_max_syn_backlog
│
└── Accept queue (complete connections)
    └── full 3-way handshake done, waiting for the app to call accept()
        [ conn ] [ conn ] [ conn ] ...   ← size: the backlog argument to listen()

The kernel manages the 3-way handshake entirely by itself — no application code runs during it. A SYN arrives → kernel puts it in the SYN queue and sends SYN+ACK. The client's ACK arrives → kernel moves the connection from the SYN queue to the accept queue. The application hasn't done anything yet.

The backlog in listen(fd, 128) is the maximum size of the accept queue — how many fully established connections can sit waiting before the application calls accept(). If the accept queue is full when a new connection completes its handshake, the kernel's default behavior is to drop the final ACK from the client. The client-side connect() returns success (the client believes the connection is established), but the server-side accept queue never logs it. The client's first data send will eventually time out or be RST'd. Setting net.ipv4.tcp_abort_on_overflow=1 makes the kernel send an explicit RST instead of silently dropping. This is where servers get overwhelmed: not because they can't handle connections, but because the application calls accept() too slowly.

import socket

# AF_INET = IPv4, SOCK_STREAM = TCP
server = socket.socket(socket.AF_INET, socket.SOCK_STREAM)
server.setsockopt(socket.SOL_SOCKET, socket.SO_REUSEADDR, 1)
server.bind(("0.0.0.0", 8080))

# At this point: no kernel queue, no connection possible
server.listen(128)
# NOW: two queues exist in kernel memory
# The kernel starts answering SYN packets by itself
# Your application is not involved yet

conn, addr = server.accept()
# THIS is when your application gets involved.
# accept() pops one entry from the accept queue.
# It returns a NEW socket fd — a connected socket with its own sk_rcvbuf/sk_sndbuf.
# The listening socket keeps listening.

The listening socket and the connected socket are two different struct sock objects. The listening socket lives on port 8080 forever. Each connected socket has a unique 4-tuple — (server_ip, 8080, client_ip, client_port) — and its own independent buffers. This is how thousands of connections can exist on the same port simultaneously: they are different sockets sharing the same listening port.

connect() — the active open

When a client calls connect(fd, address), the kernel:

  1. Picks an ephemeral source port (typically from the range 32768–60999, configurable via net.ipv4.ip_local_port_range)
  2. Records the 4-tuple (client_ip, ephemeral_port, server_ip, server_port) in the socket
  3. Sends a TCP SYN packet
  4. Blocks the calling thread/process in the SYN_SENT state

The connect() call does not return until the 3-way handshake completes (or times out). When it returns without error, the TCP connection is ESTABLISHED — the kernel has already exchanged SYN, SYN+ACK, and ACK. The application gets back a fd to a working byte pipe.

// Go: connect to a server
conn, err := net.Dial("tcp", "example.com:8080")
// If err == nil: 3-way handshake is complete.
// conn is a net.Conn wrapping a socket fd in ESTABLISHED state.
// The kernel has already allocated sk_rcvbuf and sk_sndbuf for this connection.

The mental model that doesn't leave

Think of it this way:

  • listen() tells the kernel: "build a waiting room for me on this port, handle the handshakes yourself, and put finished connections in a queue — I'll pick them up when I'm ready."
  • accept() is you walking into the waiting room and picking up the next finished connection.
  • connect() is the client knocking on the door — executing the handshake — and getting a direct line once the door opens.

The key insight: the kernel, not your application, does the TCP handshake. Your code is never in the loop during SYN/SYN-ACK/ACK. By the time your code runs, the connection is already established. accept() just hands you the key to an already-open channel.

This is also why a SYN flood attack works: it fills the SYN queue with half-open connections, preventing real connections from completing, without the application being able to do anything about it. The kernel is under attack, not the app.


4. TCP: a byte stream, not a message protocol

This is the most important thing to internalize before talking about WebSockets.

TCP delivers an ordered, reliable stream of bytes. It has no concept of messages.

When you call send("hello") followed by send(" world") on a TCP socket, the receiver might get:

  • "hello world" in one read
  • "hel" then "lo world" then "" then more
  • "hello" then " world" in two reads

TCP's job is to get all the bytes there in order. It knows nothing about where one message ends and the next begins. That's the application's problem.

The TCP state machine manages the connection lifecycle. The relevant states for a WebSocket connection:

Server calls listen()  → LISTEN
Client calls connect() → SYN sent → SYN_SENT
Server receives SYN   → SYN+ACK sent → SYN_RECEIVED
Client receives SYN+ACK → ACK sent → ESTABLISHED
Server receives ACK   → ESTABLISHED

The 3-way handshake (SYN → SYN+ACK → ACK) is happening below the application. By the time accept() returns on the server, the TCP connection is already ESTABLISHED and the kernel has already allocated the socket buffers. The application gets a file descriptor to an already-working byte pipe.

"Client" and "server" at the TCP level mean only this: the client called connect() (active open), the server called listen() then accept() (passive open). Nothing more. There is no inherent asymmetry in data flow after the connection is established.


5. HTTP/1.1 upgrade: turning a TCP connection into a WebSocket

A WebSocket connection starts as a plain HTTP/1.1 connection. The client sends a normal HTTP GET request with a few extra headers:

GET /chat HTTP/1.1
Host: server.example.com
Upgrade: websocket
Connection: Upgrade
Sec-WebSocket-Key: dGhlIHNhbXBsZSBub25jZQ==
Sec-WebSocket-Version: 13

Sec-WebSocket-Key is 16 random bytes, base64-encoded. It's not a security mechanism — it's a handshake confirmation. The server proves it understood the upgrade by computing:

SHA-1(key + "258EAFA5-E914-47DA-95CA-C5AB0DC85B11")

then base64-encoding the result and sending it back:

HTTP/1.1 101 Switching Protocols
Upgrade: websocket
Connection: Upgrade
Sec-WebSocket-Accept: s3pPLMBiTxaQ9kYGzzhZRbK+xOo=

101 Switching Protocols is the response code. After this exchange, both sides stop speaking HTTP. The TCP connection stays open, but the protocol running over it is now WebSocket. The HTTP headers are gone; what follows is raw WebSocket frames.

This is the upgrade: one HTTP request, one HTTP response, then the connection is repurposed. No new TCP connection. No new port. The same file descriptor, the same kernel socket buffers, just a different framing protocol on top of the byte stream.

"Client" and "server" at the HTTP/WebSocket level inherit the TCP roles. The one that initiated the TCP connection sent the HTTP upgrade request. But after 101, the WebSocket protocol is symmetric — both sides can send frames at any time without the other side asking.

Subprotocol negotiation

The client can declare which application-level protocols it speaks:

Sec-WebSocket-Protocol: chat, superchat

The server picks one and echoes it back:

Sec-WebSocket-Protocol: chat

If the server doesn't include Sec-WebSocket-Protocol, no subprotocol is active — the connection still works, but both sides must agree out-of-band on message format. Common subprotocols you'll encounter in the wild: STOMP (message brokers), MQTT (IoT), and the various GraphQL-over-WebSocket protocols (graphql-ws, subscriptions-transport-ws).


6. WebSocket framing: the binary format

TCP is a byte stream. WebSocket needs to deliver messages. It solves this the same way gRPC does: a fixed-size header that tells the receiver how many bytes to read.

Each WebSocket frame:

 0                   1                   2                   3
 0 1 2 3 4 5 6 7 8 9 0 1 2 3 4 5 6 7 8 9 0 1 2 3 4 5 6 7 8 9 0 1
+-+-+-+-+-------+-+-------------+-------------------------------+
|F|R|R|R| opcode|M| payload len |    Extended payload length    |
|I|S|S|S|  (4)  |A|     (7)     |            (16 or 64)        |
|N|V|V|V|       |S|             |                               |
| |1|2|3|       |K|             |                               |
+-+-+-+-+-------+-+-------------+-------------------------------+
|     Extended payload length (if payload len == 126 or 127)   |
+---------------------------------------------------------------+
|               Masking key (4 bytes, if MASK=1)               |
+---------------------------------------------------------------+
|                        Payload data                          |
+---------------------------------------------------------------+

FIN bit: if 1, this is the last (or only) frame of a message. If 0, more frames follow (fragmented message).

Opcode (4 bits):

  • 0x0 — continuation frame
  • 0x1 — text frame (UTF-8 payload)
  • 0x2 — binary frame
  • 0x8 — close
  • 0x9 — ping
  • 0xA — pong

MASK bit: if 1, the payload is XOR-masked with the 4-byte masking key. Clients must always mask frames they send. Servers must never mask frames they send. This is not optional — a server receiving an unmasked client frame must close the connection. The masking prevents a specific class of cache-poisoning attack against HTTP proxies that don't understand WebSocket.

Payload length uses a variable encoding:

  • 0–125: length is in the 7-bit field itself
  • 126: the actual length follows as a 16-bit big-endian integer (max 65535 bytes)
  • 127: the actual length follows as a 64-bit big-endian integer

This is why a WebSocket frame has a 2-byte minimum header for small messages, 4-byte for medium, and 10-byte for large — the length field expands as needed.

RSV bits and extensions

The three RSV bits (RSV1, RSV2, RSV3) must be 0 in every frame unless the client and server negotiated an extension that assigns them meaning. A receiver that sees non-zero RSV bits without a corresponding negotiated extension must close the connection with code 1002.

The most common extension is permessage-deflate (RFC 7692): it compresses each message payload with DEFLATE and uses RSV1=1 to signal that a frame is compressed. Negotiate it during the upgrade:

Sec-WebSocket-Extensions: permessage-deflate; client_max_window_bits

The server responds with what it accepts:

Sec-WebSocket-Extensions: permessage-deflate

After that, RSV1=1 on any data frame means "decompress before delivering." For text-heavy payloads (JSON, XML) this typically cuts wire size by 60–80% with negligible CPU cost.

Control frame rules

Opcodes 0x8 (close), 0x9 (ping), and 0xA (pong) are control frames. The spec imposes constraints that don't apply to data frames:

  • Cannot be fragmented — control frames must always have FIN=1. Sending a ping with FIN=0 is a protocol error.
  • Payload limited to 125 bytes — the 7-bit payload length field only; extended length is forbidden.
  • Can interleave with fragmented data messages — a ping can arrive in the middle of a fragmented message sequence and the receiver must respond immediately, without waiting for the fragmented message to finish.

The 125-byte payload limit matters for close frames: the 2-byte status code takes the first two bytes, leaving only 123 bytes for the human-readable reason string. Reason strings that exceed this get truncated or omitted.

Fragmentation and reassembly

A message that's too large to send in one frame — or that the application wants to stream — can be split across multiple frames using the FIN bit and the continuation opcode (0x0):

Frame 1: FIN=0, opcode=0x1 (text)          ← opens a new text message
Frame 2: FIN=0, opcode=0x0 (continuation)  ← middle fragment
Frame 3: FIN=1, opcode=0x0 (continuation)  ← final fragment, deliver now

The receiver accumulates payloads and delivers the assembled message when it sees FIN=1. The original opcode (0x1 text or 0x2 binary) only appears in the first frame; continuation frames always use 0x0.

Control frames can interrupt a fragmented sequence without disturbing it. A ping arriving between fragment 1 and fragment 2 is answered immediately; the partially received message stays in the reassembly buffer.

Most WebSocket libraries handle fragmentation invisibly — splitting large sends, reassembling on receipt. Where it surfaces: frame-level captures in Wireshark, or writing a custom WebSocket parser from scratch.


7. Tracing ws.send("hello") from browser to server

Put it all together. You call ws.send("hello") in a browser.

Browser (userspace):

  1. The WebSocket API builds a frame: FIN=1, opcode=0x1 (text), MASK=1, payload length=5, masking key=4 random bytes, payload="hello" XOR'd with the masking key.
  2. The frame bytes are passed to the OS via the write() or send() syscall on the socket file descriptor.

Kernel (send path): 3. The bytes land in the TCP send buffer (sk_sndbuf). 4. The TCP layer segments them (5 bytes of WebSocket data fits in one TCP segment with plenty of room), adds the TCP header (src_port, dst_port, sequence number, ACK number, flags), and passes the segment down to IP. 5. IP adds its header (src/dst IP, TTL, protocol=6). IPv4 computes a header checksum over the IP header itself; IPv6 omits it entirely, relying on lower-layer (Ethernet CRC) and upper-layer (TCP checksum) integrity checks instead. 6. The Ethernet layer adds the MAC addresses (looked up via ARP or the routing table) and the frame goes into the NIC's DMA transmit ring buffer. 7. The NIC reads from the ring buffer, serializes the bytes, drives the electrical or optical signal onto the wire.

Wire: 8. Voltage transitions. Fiber pulses. Physics.

Server NIC → kernel: 9. The NIC on the server end recovers the bits from the analog signal, verifies the Ethernet CRC, DMA-writes the frame into the kernel ring buffer, raises an interrupt. 10. The kernel's TCP layer receives the segment, verifies the TCP checksum, finds the socket by 4-tuple, appends the payload bytes to the socket's receive buffer, sends a TCP ACK back.

Server (userspace): 11. The server's event loop (epoll, kqueue, io_uring — depending on the runtime) was notified that the socket is readable. It calls read() on the fd. 12. The WebSocket library reads the 2-byte frame header, sees MASK=1 and payload length=5, reads the 4-byte masking key, reads 5 bytes, XOR's them with the masking key to recover "hello". 13. Your application code receives the string "hello".

Thirteen steps. Every one of them is real; none of them is hidden by the abstraction above it. The abstraction just gives you a shorter name for all thirteen.


8. A working server and client

Go

// go get github.com/gorilla/websocket
package main

import (
    "fmt"
    "log"
    "net/http"

    "github.com/gorilla/websocket"
)

var upgrader = websocket.Upgrader{
    // In production: check r.Header.Get("Origin") here
    CheckOrigin: func(r *http.Request) bool { return true },
}

func handle(w http.ResponseWriter, r *http.Request) {
    // This performs the HTTP → WebSocket upgrade (the 101 exchange)
    conn, err := upgrader.Upgrade(w, r, nil)
    if err != nil {
        log.Println("upgrade error:", err)
        return
    }
    defer conn.Close()

    // Each connection gets its own goroutine — this IS the read loop.
    // conn.ReadMessage() blocks until a full WebSocket frame arrives,
    // parses the header, unmasks the payload, and returns the message.
    for {
        msgType, msg, err := conn.ReadMessage()
        if err != nil {
            // websocket.IsCloseError checks for clean close frames (1000, 1001, ...)
            if websocket.IsCloseError(err, websocket.CloseNormalClosure, websocket.CloseGoingAway) {
                log.Println("client closed connection cleanly")
            } else {
                log.Println("read error:", err)
            }
            return
        }
        fmt.Printf("received [type=%d]: %s\n", msgType, msg)

        // Echo back. WriteMessage builds the frame (FIN=1, no mask — server never masks),
        // writes it to sk_sndbuf, the kernel handles the rest.
        if err := conn.WriteMessage(msgType, msg); err != nil {
            log.Println("write error:", err)
            return
        }
    }
}

func main() {
    http.HandleFunc("/ws", handle)
    log.Println("listening on :8080")
    log.Fatal(http.ListenAndServe(":8080", nil))
}

The Go model: one goroutine per connection. Goroutines start with a small stack (around 2–8 KB, depending on Go version) that grows on demand — far cheaper than OS threads (1–8 MB each). This makes tens of thousands of concurrent connections achievable on a typical server; actual limits depend on available memory, OS file descriptor limits (ulimit -n), and your specific workload. Each goroutine blocks in ReadMessage() — which internally calls read() on the fd — while the Go scheduler multiplexes goroutines across a small pool of OS threads. Blocking in a goroutine is not the same as blocking an OS thread.

Python

# pip install websockets
import asyncio
import websockets

async def handle(websocket):
    print(f"client connected: {websocket.remote_address}")
    try:
        async for message in websocket:
            # 'message' is already unmasked, decoded if text frame
            print(f"received: {message}")
            await websocket.send(f"echo: {message}")
    except websockets.exceptions.ConnectionClosedOK:
        print("client closed connection cleanly")
    except websockets.exceptions.ConnectionClosedError as e:
        print(f"connection dropped: {e}")

async def main():
    # websockets.serve() does the HTTP upgrade for each incoming connection
    async with websockets.serve(handle, "localhost", 8080):
        await asyncio.Future()  # run forever

asyncio.run(main())

The Python model: one event loop, all connections share it. async for message in websocket is a coroutine that yields control back to the event loop while waiting for a frame. When the kernel signals the fd is readable (via epoll), the event loop resumes the coroutine. No threads. No blocking.

Both models sit on top of the same OS primitives: epoll_wait() to know when an fd has data, read() to pull bytes from sk_receive_queue, write() to push bytes into sk_write_queue.


9. Connection lifecycle

A WebSocket connection passes through four states:

CONNECTING → OPEN → CLOSING → CLOSED

CONNECTING: the TCP connection exists and the HTTP upgrade request has been sent, but 101 hasn't been received yet. You cannot call send().

OPEN: 101 was received. Both sides can send frames freely. This is the only state where data frames are valid.

CLOSING: one side has sent a close frame. The connection is in the close handshake. You should not send new data frames, but you must still read to process the close echo.

CLOSED: the close handshake completed (or the TCP connection dropped). The fd is gone. Any further reads/writes return an error.


10. The close handshake

Closing a WebSocket is not the same as dropping the TCP connection. The protocol defines a two-step close:

Side A sends:  [close frame, code=1000]
Side B reads:  close frame → echoes back [close frame, code=1000]
Side A reads:  the echo
Both sides:    call TCP close() — FIN/FIN-ACK exchange

Close status codes you'll encounter:

CodeMeaning
1000Normal closure
1001Going away (server restarting, browser navigating away)
1002Protocol error
1003Unsupported data type
1006Abnormal — no close frame was received; TCP dropped
1011Server encountered an unexpected error

Code 1006 is important: it means the TCP connection was terminated without a WebSocket close frame. The library synthesizes this code — it never appears on the wire. Seeing 1006 in your logs means a router dropped the connection, a proxy timed out, or the process crashed.

In Go:

// Send a clean close frame before closing
conn.WriteMessage(
    websocket.CloseMessage,
    websocket.FormatCloseMessage(websocket.CloseNormalClosure, ""),
)

In Python:

# websockets handles the close echo automatically
await websocket.close(code=1000, reason="done")

11. Ping, pong, and dead connection detection

TCP keepalive exists at the OS level, but it has a problem: it only detects that the TCP path is alive, not that the application at the other end is still running. A proxy between client and server can be alive while the client behind it is gone.

WebSocket has its own keepalive: ping and pong frames (opcodes 0x9 and 0xA). Either side can send a ping at any time. The other side must respond with a pong containing the same payload. If no pong arrives within a timeout, the connection is considered dead.

// Go: send a ping every 30 seconds, close if no pong within 10 seconds
conn.SetPongHandler(func(data string) error {
    // reset the read deadline on every pong
    conn.SetReadDeadline(time.Now().Add(60 * time.Second))
    return nil
})

go func() {
    ticker := time.NewTicker(30 * time.Second)
    defer ticker.Stop()
    for range ticker.C {
        if err := conn.WriteMessage(websocket.PingMessage, nil); err != nil {
            return
        }
    }
}()
# Python: websockets handles ping/pong automatically
# ping_interval=20 sends a ping every 20s
# ping_timeout=10 closes if no pong within 10s
async with websockets.serve(handle, "localhost", 8080,
                            ping_interval=20, ping_timeout=10):
    await asyncio.Future()

The rule: always implement heartbeat. Default TCP keepalive is often 2 hours. Your dead connections will pile up long before the OS notices.


12. TLS and WSS

wss:// is WebSocket over TLS. The stack order is:

WebSocket framing
    └─ TLS record layer  (encryption, MAC, handshake)
         └─ TCP byte stream
              └─ IP / Ethernet / wire

TLS sits between TCP and WebSocket. The TLS handshake happens before the HTTP upgrade. By the time the GET /ws HTTP/1.1 Upgrade: websocket request is sent, TLS is already established and the HTTP bytes are already encrypted.

From the kernel's perspective, TLS changes nothing about the socket: sk_receive_queue still holds TCP segments, the fd still works the same way. TLS is handled entirely in userspace (by OpenSSL, BoringSSL, or rustls). The kernel delivers encrypted bytes; the TLS library decrypts them before the WebSocket library ever sees them.

Proxies treat wss:// differently from ws://. An HTTP proxy that handles ws:// can inspect and potentially interfere with the frames. With wss://, the proxy sees only encrypted bytes — it can tunnel the connection but not inspect it. This is why enterprise proxies often block wss:// on non-standard ports: they can't tell if it's WebSocket or something else.


13. Proxies, nginx, and the Upgrade header problem

WebSocket breaks behind reverse proxies by default because HTTP proxies were designed for request/response cycles, not long-lived bidirectional streams.

The specific problem: HTTP/1.1 proxies forward the Upgrade header by name, but the HTTP spec says hop-by-hop headers (including Upgrade and Connection) must not be forwarded. So a naive proxy strips them. The server never sees the upgrade request and responds with a normal HTTP response instead of 101.

The nginx fix:

location /ws {
    proxy_pass http://backend:8080;

    # Tell nginx to forward the Upgrade header
    proxy_http_version 1.1;
    proxy_set_header Upgrade $http_upgrade;
    proxy_set_header Connection "upgrade";

    # Disable buffering — WebSocket frames must be forwarded immediately
    proxy_buffering off;

    # Long timeout: WebSocket connections can be idle for minutes
    proxy_read_timeout 3600s;
    proxy_send_timeout 3600s;
}

Without proxy_read_timeout, nginx will close the connection after 60 seconds of inactivity (the default). Without proxy_buffering off, nginx buffers the stream and your frames won't be forwarded until the buffer fills up — which defeats the purpose of WebSocket.

Load balancing adds another problem: WebSocket connections are stateful and long-lived. If a load balancer routes connection A to server 1 and later routes a new connection from the same client to server 2, server 2 has no state for that client. Solutions: sticky sessions (route the same client to the same server), or move state to a shared store (Redis pub/sub is common for broadcasting).

Intermediary timeouts and the idle-connection problem

Every hop between client and server has an idle timeout — a timer that fires if no bytes cross the TCP connection for N seconds:

IntermediaryTypical default
nginx60 s (proxy_read_timeout)
AWS ALB60 s (configurable up to 4000 s)
Cloudflare100 s
Enterprise firewall5 min (varies widely)

When the timer fires, the intermediary closes the TCP connection without sending a WebSocket close frame. Both endpoints may not notice for seconds or minutes (until the next send or read). The client sees close code 1006 — the library-synthesized "no close frame received" code.

This is the real reason heartbeat is mandatory in production. You're not asking "is my peer alive?" — you're asking "is the entire network path between us still forwarding packets?" A proxy doesn't know about WebSocket keepalive; it counts seconds since the last byte on the TCP stream.

Diagnostic signatures: connections drop on suspiciously round intervals (60 s, 300 s). Code 1006. Only affects users behind specific corporate proxies or VPNs. Never reproduces when connecting directly to the server.

Fix: set ping interval below the lowest timeout in your infrastructure. 25–45 seconds covers most environments. Confirm nginx proxy_read_timeout is set to at least 2× your ping interval.


14. Backpressure: what happens when the sender is faster than the receiver

The sk_rcvbuf field in struct sock is the maximum number of bytes allowed in the receive queue — typically 87 KB by default, configurable up to net.core.rmem_max. When the receive buffer fills up, TCP advertises a window size of zero to the sender. The sender stops sending and waits for the window to open.

This is TCP flow control, and it propagates all the way up to your application:

app writes fast → sk_sndbuf fills → write() blocks
                                         ↓
                              TCP sends data until
                              remote sk_rcvbuf fills
                                         ↓
                              remote TCP sends window=0
                                         ↓
                              your TCP stops sending
                                         ↓
                              sk_sndbuf stays full
                                         ↓
                              your write() stays blocked

In Go, conn.WriteMessage() will block if the send buffer is full. In Python, await websocket.send() will block the coroutine (returning control to the event loop) until the send buffer has room.

This is correct behavior — it means the application is being naturally throttled to the speed the receiver can handle. The failure mode is when you don't respect backpressure: queuing messages in memory faster than they can be sent, until your process runs out of RAM. Always check for write errors and respect blocking writes.


15. TCP gotchas: Nagle's algorithm and delayed ACK

Two TCP features that interact poorly with WebSocket's small-frame pattern.

Nagle's algorithm (enabled by default on every TCP socket) buffers small writes. When your application calls write() with a small payload — a ping, a short event, a control frame — the kernel holds the bytes and waits until either a full-size segment accumulates or a TCP ACK arrives from the remote end. On interactive workloads this adds 40–200 ms of latency per small message.

The fix is TCP_NODELAY, which disables Nagle:

import socket
s = socket.socket()
s.setsockopt(socket.IPPROTO_TCP, socket.TCP_NODELAY, 1)
tc, _ := conn.(*net.TCPConn)
tc.SetNoDelay(true)

Most WebSocket libraries enable TCP_NODELAY automatically. If you're seeing latency spikes on small messages that go away under load, check that it's actually set (ss -tin will show nonagle in the options).

Delayed ACK holds acknowledgments for up to 40 ms on the receiving side, hoping to piggyback them on outbound data. When the sender has Nagle enabled and the receiver uses delayed ACK, they deadlock: the sender waits for an ACK before sending the next segment; the receiver waits 40 ms before sending the ACK, hoping data will arrive to piggyback it on. The result is a guaranteed ~40 ms stall on every request/response exchange that involves small messages in both directions.

TCP_NODELAY on the sender breaks the deadlock — the sender stops waiting for ACKs before sending. The combination of TCP_NODELAY on both ends eliminates this entirely.


16. Comparison: WebSocket vs the alternatives

WebSocketServer-Sent EventsLong pollingHTTP/2 streams
DirectionBidirectionalServer → Client onlyServer → Client (simulated)Bidirectional
ProtocolWebSocket (after upgrade)HTTPHTTPHTTP/2
FramingBinary framesText, \n\n delimitedFull HTTP response per messageHTTP/2 DATA frames
Proxy friendlyNeeds configYesYesNeeds HTTP/2 support
LatencyLowLowHigh (round-trip per message)Low
MultiplexingNo (one stream per connection)N/AN/AYes (many streams per TCP conn)
Browser supportUniversalUniversalUniversalUniversal
ComplexityMediumLowLowMedium

Use SSE when: server pushes events to the browser and the browser never needs to send data back (live feeds, notifications, progress bars). It's HTTP, so it works through every proxy without configuration.

Use long polling when: you need server push but can't change infrastructure. It's wasteful (a new HTTP request per message) but works everywhere.

Use WebSocket when: you need low-latency bidirectional messaging (chat, multiplayer games, collaborative editing, live terminals).

Use HTTP/2 streams when: you're building a service API and want streaming responses without leaving HTTP. gRPC is the most common case.


17. Security checklist

A WebSocket connection bypasses most of the protections that come with HTTP:

  • Validate the Origin header — browsers always send it; non-browser clients don't. If your server only serves browsers, reject connections from unexpected origins. Don't rely on this alone.
  • Authenticate on connection or in the first message — HTTP session cookies ride along with the upgrade request, but most WebSocket frameworks don't check them automatically. Pass a signed token in a query parameter (WSS only — never WS) or authenticate via a first-message handshake.
  • Rate-limit at the message level — a single WebSocket connection can send thousands of messages per second. HTTP rate limiters only see one request: the upgrade. Track messages-per-connection-per-second in your handler.
  • Enforce a maximum message size — read the frame header, check the declared payload length before allocating memory, and close with code 1009 (message too big) if it exceeds your limit. Don't buffer first and validate later.
  • Validate every message — WebSocket has no schema enforcement. Parse and validate before acting on anything.
  • CSRF via WebSocket — a malicious page can open a WebSocket connection to your server from a user's browser, inheriting their session cookies. Origin validation is the primary defense. SameSite=Strict cookies also help.

18. WebSocket over HTTP/2 (RFC 8441)

Classic WebSocket requires HTTP/1.1. The Upgrade header doesn't exist in HTTP/2 — HTTP/2 has its own multiplexing model and doesn't use connection-level upgrades.

RFC 8441 adds a WebSocket tunnel over HTTP/2 using an extended CONNECT method with a :protocol pseudo-header:

:method    = CONNECT
:protocol  = websocket
:scheme    = https
:path      = /ws
:authority = server.example.com

The HTTP/2 stream acts as a byte pipe: WebSocket framing sits inside it instead of on a raw TCP connection. The benefit is that a single TCP connection can multiplex both regular HTTP/2 requests and WebSocket streams simultaneously — something classic WebSocket can't do.

Browser support: Chrome 91+, Firefox 89+. Server library support is uneven. If you're deploying HTTP/2 everywhere and need WebSocket, it's worth checking RFC 8441 support in your stack. The pragmatic fallback — and what most production systems do today — is to serve WebSocket on a separate HTTP/1.1 endpoint.


19. Mental model and common failures

Layer-by-layer recap

Every section of this article covers one rung of the same ladder. Here it is compressed:

LayerKey conceptWhere you see it
Physical (§1)Voltage → bits; NIC clock recoveryEthernet CRC errors, fiber dB budget
Kernel (§2)struct sock, sk_buff pointer model, fd chainstrace, /proc/net/tcp, ss -tnp
listen/connect (§3)Kernel does the 3-way handshake; app calls accept()SYN floods, accept queue overflow
TCP stream (§4)No message concept; ordered reliable bytesPartial reads, need for framing
HTTP upgrade (§5)One 101 exchange; same fd, new protocolnginx Upgrade stripping, subprotocols
WS framing (§6)FIN+opcode+MASK+variable lengthWireshark WebSocket dissector, frame builder
Application (§8–11)Goroutines / event loops; close handshake; heartbeatCode 1006, dead connection detection
Infrastructure (§12–13)Proxies need config; TCP flow control = backpressureproxy_buffering off, blocked writes
Transport quirks (§15)Nagle + delayed ACK deadlockss -tin, latency on small frames

Common failure modes

SymptomMost likely causeHow to verifyFix
Close code 1006, no patternTCP dropped without close frametcpdump, check kernel logsInvestigate network path
Code 1006 every ~60 snginx or LB idle timeoutTime the drops exactlyproxy_read_timeout 3600s + heartbeat
Code 1006 only for VPN/corp usersEnterprise firewall timeoutReproduce on VPNLower ping_interval to 25–30 s
"426 Upgrade Required" from serverProxy stripped Upgrade headercurl -v through proxyproxy_http_version 1.1, proxy_set_header Upgrade
Latency spikes on small sends, gone under loadNagle's algorithmss -tin for nonagleTCP_NODELAY on socket
~40 ms stall on every request/responseNagle + delayed ACK deadlocktcpdump timestampsTCP_NODELAY on both ends
Server OOM under spikeUnbounded message queue (backpressure ignored)ss -tnp Send-Q depthRespect blocking writes, add queue limits
Client connects; server never sees itAccept queue fullss -tnlp Recv-QCall accept() faster; increase backlog
Frame rejected with code 1002RSV bits set without extension negotiatedstrace or WiresharkCheck extension negotiation in handshake
Close reason truncated125-byte control frame limitFrame hex dumpShorten reason string; log code, not reason

20. Experiments — see it yourself

Theory is only half the work. Every concept in this article is observable from your terminal right now. Run these, look at the output, and the model becomes permanent.

Experiment 1: Watch the exact syscalls with strace

Start a Python WebSocket server in one terminal:

pip install websockets
python3 -c "
import asyncio, websockets

async def h(ws):
    await ws.send('hello')

async def main():
    async with websockets.serve(h, 'localhost', 8765):
        await asyncio.Future()

asyncio.run(main())
"

In a second terminal, attach strace to it:

PID=$(pgrep -f "websockets.serve")
strace -p $PID -e trace=network,read,write -s 128 2>&1

Then connect a client. What you'll see in strace:

accept4(3, {sa_family=AF_INET, sin_port=htons(54321), ...}) = 5
read(5, "GET / HTTP/1.1\r\nUpgrade: websocket\r\n...", 65536) = 187
write(5, "HTTP/1.1 101 Switching Protocols\r\n...", 129) = 129
write(5, "\x81\x05hello", 7) = 7

Four lines: accept queue pop, HTTP upgrade read, 101 write, WebSocket frame write. Decode that last write:

  • \x81 = 10000001 = FIN=1, opcode=0x1 (text frame)
  • \x05 = payload length 5
  • hello = unmasked payload (server never masks)

The entire framing spec, visible in one line of strace output.


Experiment 2: Watch socket states and queue sizes live

# All established TCP connections with process info
ss -tnp state established

# Watch it update every second
watch -n1 'ss -tnp'

# Show just the listening socket
ss -tnlp sport = 8765

Expected output while a connection is open:

State   Recv-Q  Send-Q  Local Address:Port  Peer Address:Port  Process
ESTAB   0       0       127.0.0.1:8765      127.0.0.1:54321    users:(("python3",pid=1234,fd=5))
LISTEN  0       128     0.0.0.0:8765        0.0.0.0:*          users:(("python3",pid=1234,fd=3))

Recv-Q is bytes sitting in sk_receive_queue not yet read by the app. Send-Q is bytes in sk_write_queue not yet sent. Now create backpressure: slow down the server's reading loop and flood it with messages. Watch Recv-Q climb in real time. When it hits sk_rcvbuf (~87 KB), TCP advertises window=0 and the sender stalls.


Experiment 3: The accept queue — watch it fill up

# slow_server.py
import socket, time

s = socket.socket()
s.setsockopt(socket.SOL_SOCKET, socket.SO_REUSEADDR, 1)
s.bind(("0.0.0.0", 9000))
s.listen(5)           # backlog=5: accept queue holds max 5 connections
print("listening — sleeping 5s before first accept()")
time.sleep(5)         # kernel handles handshakes; app does nothing

while True:
    conn, addr = s.accept()
    print(f"accepted: {addr}")
python3 slow_server.py &

# Fire 10 simultaneous connections
for i in $(seq 10); do
  python3 -c "import socket; s=socket.socket(); s.connect(('localhost',9000))" &
done

# Check the accept queue depth while the server is sleeping
ss -tnlp sport = 9000

Output:

State   Recv-Q  Send-Q  Local Address:Port
LISTEN  5       5       0.0.0.0:9000

Recv-Q on a LISTEN socket = connections currently queued, waiting for accept(). It is capped at 5 (the backlog). The other 5 clients had their 3-way handshake completed by the kernel — they think they're connected — but accept() will never be called for them. They were silently dropped. This is how a slow application causes connection loss under load, with no error visible at the network level.


Experiment 4: Do the WebSocket handshake by hand with nc

No library. Just netcat and your keyboard. Proves the handshake is plain HTTP text.

# Start a server
python3 -c "
import asyncio, websockets

async def h(ws):
    msg = await ws.recv()
    await ws.send(f'got: {msg}')

async def main():
    async with websockets.serve(h, 'localhost', 8765):
        await asyncio.Future()

asyncio.run(main())
" &

# Connect with nc
nc localhost 8765

Type this exactly (blank line at the end):

GET / HTTP/1.1
Host: localhost:8765
Upgrade: websocket
Connection: Upgrade
Sec-WebSocket-Key: dGhlIHNhbXBsZSBub25jZQ==
Sec-WebSocket-Version: 13

You get back:

HTTP/1.1 101 Switching Protocols
Upgrade: websocket
Connection: Upgrade
Sec-WebSocket-Accept: s3pPLMBiTxaQ9kYGzzhZRbK+xOo=

The connection is now in WebSocket mode. The terminal is live. To send a frame you'd need to type raw binary — but you've proved the upgrade is nothing more than HTTP text.


Experiment 5: Capture frames with tcpdump

# Capture loopback, show ASCII content, capture everything
sudo tcpdump -i lo -A -s 0 port 8765

What you'll see — first the TCP 3-way handshake (flags only, no payload):

12:00:01.000001 IP 127.0.0.1.54321 > 127.0.0.1.8765: Flags [S], seq 0, win 65495, length 0
12:00:01.000012 IP 127.0.0.1.8765 > 127.0.0.1.54321: Flags [S.], seq 0, ack 1, win 65495, length 0
12:00:01.000021 IP 127.0.0.1.54321 > 127.0.0.1.8765: Flags [.], ack 1, win 65495, length 0

Then the HTTP upgrade request (plain ASCII — you can read every byte):

12:00:01.001000 IP 127.0.0.1.54321 > 127.0.0.1.8765: Flags [P.], seq 1:188, length 187
GET / HTTP/1.1
Host: localhost:8765
Upgrade: websocket
Connection: Upgrade
Sec-WebSocket-Key: dGhlIHNhbXBsZSBub25jZQ==
Sec-WebSocket-Version: 13

Then the 101 response (also plain ASCII):

12:00:01.002000 IP 127.0.0.1.8765 > 127.0.0.1.54321: Flags [P.], seq 1:130, length 129
HTTP/1.1 101 Switching Protocols
Upgrade: websocket
Connection: Upgrade
Sec-WebSocket-Accept: s3pPLMBiTxaQ9kYGzzhZRbK+xOo=

Then the WebSocket frame — the ASCII dump looks garbled because it's binary:

12:00:01.003000 IP 127.0.0.1.8765 > 127.0.0.1.54321: Flags [P.], seq 130:137, length 7
.....hello

The leading ..... is the 2-byte frame header (\x81\x05) rendered as non-printable characters. hello appears because the payload is plain ASCII and tcpdump's -A flag prints whatever bytes it finds. The frame is 7 bytes total: 2 header + 5 payload. That matches exactly.

For a proper field-by-field decode: open Wireshark, capture on lo, filter with websocket. The WebSocket dissector will label every bit:

WebSocket
  Frame: Text (fin=True, opcode=1, mask=False, len=5)
    FIN: True
    Reserved: 0x0
    Opcode: Text (1)
    Mask: False
    Payload length: 5
    Payload: hello

This is the same frame. Wireshark just has a dissector that understands the binary layout — it's doing the same arithmetic you'd do by hand on the raw bytes.


Experiment 6: Read the kernel's socket table directly

# No tools — raw kernel data
cat /proc/net/tcp
sl  local_address rem_address   st tx_queue rx_queue
0:  0100007F:225D 00000000:0000 0A 00000000:00000000
1:  0100007F:225D 0100007F:D431 01 00000000:00000000

Decode: 0100007F = 127.0.0.1 (little-endian hex). 225D = port 8765. 0A = LISTEN, 01 = ESTABLISHED. tx_queue:rx_queue = send/recv buffer depth in bytes. Same numbers ss shows — but here you're reading the kernel's own data structure with no abstraction in between.


Experiment 7: Build and send a WebSocket frame manually

No library on the send side. This is the full protocol in 25 lines:

import socket, base64, os, struct

def handshake(sock, host, port):
    key = base64.b64encode(os.urandom(16)).decode()
    sock.sendall((
        f"GET / HTTP/1.1\r\nHost: {host}:{port}\r\n"
        f"Upgrade: websocket\r\nConnection: Upgrade\r\n"
        f"Sec-WebSocket-Key: {key}\r\nSec-WebSocket-Version: 13\r\n\r\n"
    ).encode())
    assert b"101" in sock.recv(4096)

def send_text(sock, msg: str):
    payload = msg.encode()
    mask = os.urandom(4)
    # XOR each payload byte with mask[i % 4] — that's the entire masking algorithm
    masked = bytes(b ^ mask[i % 4] for i, b in enumerate(payload))
    # FIN=1, opcode=1 (text), MASK=1, 7-bit length
    sock.sendall(bytes([0x81, 0x80 | len(payload)]) + mask + masked)

def recv_text(sock):
    h = sock.recv(2)
    length = h[1] & 0x7F          # server never masks, no masking key to read
    return sock.recv(length).decode()

s = socket.socket()
s.connect(("localhost", 8765))
handshake(s, "localhost", 8765)
send_text(s, "hello from raw socket")
print("server replied:", recv_text(s))
s.close()

Run this against the Python server from Experiment 4. Expected output:

server replied: got: hello from raw socket

One line. That's the full round-trip: raw TCP socket → manual HTTP upgrade → manually built and masked WebSocket frame → server receives it, echoes it back → manually parsed server frame → decoded string. No library touched on the client side. The masking loop (b ^ mask[i % 4]) is the only "complex" part — it's XOR with a cycling 4-byte key. Everything else in a production WebSocket library is error handling and edge cases layered on top of this exact core.


Experiment 8: Trigger a protocol error — send an unmasked client frame

RFC 6455 §5.1: a server MUST close the connection upon receiving an unmasked frame from a client. Let's verify that.

Start the echo server from Experiment 4, then run:

import socket, base64, os

s = socket.socket()
s.connect(("localhost", 8765))

key = base64.b64encode(os.urandom(16)).decode()
s.sendall((
    f"GET / HTTP/1.1\r\nHost: localhost:8765\r\n"
    f"Upgrade: websocket\r\nConnection: Upgrade\r\n"
    f"Sec-WebSocket-Key: {key}\r\nSec-WebSocket-Version: 13\r\n\r\n"
).encode())
assert b"101" in s.recv(4096)

# 0x81 = FIN=1 + opcode 1 (text). 0x05 = MASK=0 + length 5. No masking key.
s.sendall(bytes([0x81, 0x05]) + b"hello")

raw = s.recv(4096)
print(f"response hex: {raw.hex()}")
print(f"opcode: 0x{raw[0] & 0x0F:x}  (8 = close)")
print(f"close code: {int.from_bytes(raw[2:4], 'big')}")
s.close()

Expected output:

response hex: 880203ea
opcode: 0x8  (8 = close)
close code: 1002

Decode 880203ea: 88 = FIN=1 + opcode 8 (close). 02 = payload length 2. 03ea = 0x03EA = 1002. The server detected the masking violation and closed with the correct code. The entire detection and response path is in the WebSocket library's frame parser, two bytes into the frame.


Experiment 9: Send a fragmented message with raw sockets

Prove that the server reassembles three frames — FIN=0, FIN=0, FIN=1 — into one message before calling your handler.

import socket, base64, os

def hs(sock):
    key = base64.b64encode(os.urandom(16)).decode()
    sock.sendall((
        f"GET / HTTP/1.1\r\nHost: localhost:8765\r\n"
        f"Upgrade: websocket\r\nConnection: Upgrade\r\n"
        f"Sec-WebSocket-Key: {key}\r\nSec-WebSocket-Version: 13\r\n\r\n"
    ).encode())
    assert b"101" in sock.recv(4096)

def frame(opcode, payload, fin):
    mask = os.urandom(4)
    masked = bytes(b ^ mask[i % 4] for i, b in enumerate(payload))
    b0 = (0x80 if fin else 0x00) | opcode
    return bytes([b0, 0x80 | len(payload)]) + mask + masked

s = socket.socket()
s.connect(("localhost", 8765))
hs(s)

s.sendall(frame(0x1, b"hel",   fin=False))  # text, FIN=0 — opens the message
s.sendall(frame(0x0, b"lo ",   fin=False))  # continuation, FIN=0
s.sendall(frame(0x0, b"world", fin=True))   # continuation, FIN=1 — delivers the message

r = s.recv(4096)
n = r[1] & 0x7F
print("server received:", r[2:2+n].decode())
s.close()

Expected:

server received: got: hello world

The echo server called your handler exactly once with "hello world". It never saw three separate messages. The fragmentation was completely transparent to the application.

Capture this with Wireshark (websocket filter) to see three DATA frames: the first with opcode=1 and FIN=0, the second and third with opcode=0. Only the third has FIN=1.


Experiment 10: Control frame too large — trigger code 1002

A control frame payload over 125 bytes is a protocol violation (RFC 6455 §5.5). Using extended length (payload length field = 126) in a control frame is explicitly forbidden. The server must close with code 1002.

import socket, base64, os

s = socket.socket()
s.connect(("localhost", 8765))

key = base64.b64encode(os.urandom(16)).decode()
s.sendall((
    f"GET / HTTP/1.1\r\nHost: localhost:8765\r\n"
    f"Upgrade: websocket\r\nConnection: Upgrade\r\n"
    f"Sec-WebSocket-Key: {key}\r\nSec-WebSocket-Version: 13\r\n\r\n"
).encode())
assert b"101" in s.recv(4096)

# Ping with 126-byte payload using extended length encoding
# 0x89 = FIN=1 + opcode 9 (ping). 0xFE = MASK=1 + 126 (extended length trigger).
# Two more bytes give the actual length: 0x00 0x7E = 126.
payload = b"x" * 126
mask = os.urandom(4)
masked = bytes(b ^ mask[i % 4] for i, b in enumerate(payload))
s.sendall(bytes([0x89, 0xFE, 0x00, 0x7E]) + mask + masked)

r = s.recv(4096)
print(f"close code: {int.from_bytes(r[2:4], 'big')}")  # 1002
s.close()

Expected:

close code: 1002

Two ways to violate the control frame size rule in one frame: the payload itself is 126 bytes (over the 125-byte limit), and you used extended length encoding (forbidden for control frames). Either one is sufficient.


Experiment 11: TCP_NODELAY — see nonagle with ss -tin

Connect to the echo server with TCP_NODELAY set, then inspect the socket with ss -tin. The nonagle flag is explicit proof that Nagle is disabled.

Connect a client and hold the connection open:

import socket, base64, os, time

s = socket.socket()
s.setsockopt(socket.IPPROTO_TCP, socket.TCP_NODELAY, 1)   # disable Nagle
s.connect(("localhost", 8765))

key = base64.b64encode(os.urandom(16)).decode()
s.sendall((
    f"GET / HTTP/1.1\r\nHost: localhost:8765\r\n"
    f"Upgrade: websocket\r\nConnection: Upgrade\r\n"
    f"Sec-WebSocket-Key: {key}\r\nSec-WebSocket-Version: 13\r\n\r\n"
).encode())
s.recv(4096)
print(f"connected, sleeping 30s (fd={s.fileno()})")
time.sleep(30)
s.close()

While it's sleeping, in another terminal:

ss -tin dst :8765

With TCP_NODELAY set:

ESTAB 0 0 127.0.0.1:54321 127.0.0.1:8765
     cubic wscale:7,7 rto:201 rtt:0.044/0.022 ato:40 mss:65483 pmtu:65535
     cwnd:10 bytes_sent:187 segs_out:3 segs_in:2 send 119Gbps nonagle

Without it (remove the setsockopt line):

ESTAB 0 0 127.0.0.1:54322 127.0.0.1:8765
     cubic wscale:7,7 rto:201 rtt:0.038/0.019 ato:40 mss:65483 pmtu:65535
     cwnd:10 bytes_sent:187 segs_out:3 segs_in:2 send 119Gbps

nonagle is present in the first run, absent in the second. Also notice ato:40 in both — that's the delayed ACK timer set to 40 ms. When Nagle is active (second run), these two interact: the sender won't send until it gets an ACK; the receiver delays the ACK for 40 ms. That's the deadlock.


Experiment 12: Simulate an intermediary — SIGKILL the server

A clean close sends a WebSocket close frame and TCP FIN. SIGKILL forces the OS to tear down the socket immediately, sending a TCP RST with no cleanup. The client receives the RST, gets no close frame, and the library synthesizes close code 1006.

Terminal 1 — start the echo server:

python3 server.py &
echo "server PID: $!"

Terminal 2 — connect a client and block waiting for a message:

import asyncio, websockets

async def main():
    async with websockets.connect("ws://localhost:8765") as ws:
        print("connected")
        try:
            msg = await ws.recv()
        except websockets.exceptions.ConnectionClosed as e:
            print(f"connection closed: code={e.rcvd.code if e.rcvd else 'none'}, synthesized=1006")

asyncio.run(main())

Terminal 1 — kill the server without letting it clean up:

kill -9 <server_PID>

Client output:

connected
connection closed: code=none, synthesized=1006

No close frame arrived — the OS sent a TCP RST, the client's recv() got an error, and the library synthesized 1006. This is exactly what an intermediary timeout looks like: no warning, no close code on the wire, just a dropped TCP connection. The 1006 is a library convention, not a frame that traveled over the network.


Experiment 13: permessage-deflate — watch RSV1 in Wireshark

The Python websockets library negotiates permessage-deflate automatically when both sides support it. You can see it working in two places: the negotiated extension header, and RSV1=1 on compressed frames.

import asyncio, websockets

async def main():
    async with websockets.connect("ws://localhost:8765") as ws:
        print("negotiated extensions:", ws.extensions)
        msg = "hello " * 100   # 600 bytes — repetitive, compresses well
        await ws.send(msg)
        r = await ws.recv()
        print(f"sent {len(msg.encode())} bytes, received {len(r.encode())} bytes back")

asyncio.run(main())

The extensions field shows what was negotiated:

negotiated extensions: [PerMessageDeflate(client_no_context_takeover=False, ...)]

Now open Wireshark, capture on lo, filter websocket. Find the outbound data frame and expand it:

WebSocket
  Frame: Text (fin=True, opcode=1, mask=True, len=18)
    FIN: True
    Reserved: 0x4   ← RSV1=1 (compressed)
    Opcode: Text (1)
    Mask: True
    Payload length: 18   ← 600 bytes compressed to 18
    Masking-Key: ...
    Masked payload: ...

Reserved: 0x4 means bit 6 is set — that's RSV1. The 600-byte payload compressed to 18 bytes (97% reduction). This is why JSON-heavy WebSocket APIs almost always negotiate permessage-deflate: the compression ratio on structured text is dramatic. The RSV1=1 is the flag that tells the receiver to run DEFLATE decompression before passing the bytes to the application.


Experiment 14: Backpressure — watch the TCP receive window shrink

Set up a slow-reading TCP server and a fast-writing client. Watch ss -tin as the receive window drops to zero and the sender stalls.

Slow reader (terminal 1):

# slow_reader.py
import socket, time

s = socket.socket()
s.setsockopt(socket.SOL_SOCKET, socket.SO_REUSEADDR, 1)
s.bind(("0.0.0.0", 9999))
s.listen(1)
conn, addr = s.accept()
print(f"accepted {addr}")
while True:
    data = conn.recv(512)    # read only 512 bytes at a time
    if not data:
        break
    time.sleep(0.05)          # artificial slowdown: ~10 KB/s read rate

Fast writer (terminal 2):

# fast_writer.py
import socket, time

s = socket.socket()
s.connect(("localhost", 9999))
chunk = b"x" * 8192   # 8 KB chunks
sent = 0
while True:
    n = s.send(chunk)
    sent += n
    print(f"\rsent: {sent // 1024} KB", end="", flush=True)

Observer (terminal 3, run while the above two are running):

watch -n0.3 'ss -tin dst :9999'

What you'll see over time:

# Early — window is open, data flows freely:
ESTAB 0  0  127.0.0.1:55001  127.0.0.1:9999
      rcv_wnd:65280 ...

# Mid — receive buffer filling, window shrinking:
ESTAB 0  87040  127.0.0.1:55001  127.0.0.1:9999
      rcv_wnd:16384 ...

# Late — receive buffer full, window=0, sender stalled:
ESTAB 0  212992  127.0.0.1:55001  127.0.0.1:9999
      rcv_wnd:0 ...

rcv_wnd:0 means the receiver has advertised a zero window — it has no room left in sk_rcvbuf. The sender's s.send() is now blocking inside the kernel. Send-Q:212992 is the bytes waiting in the sender's sk_sndbuf, going nowhere. This is the backpressure chain made visible: slow reader → full recv buffer → window=0 → blocked sender. For a WebSocket server, the same chain runs whenever conn.WriteMessage() blocks.


WebSocket is not HTTP. After the 101 response, HTTP is done. There are no HTTP headers, no request/response cycles, no status codes. The protocol is WebSocket.

WebSocket is not TCP. TCP is the byte stream underneath. WebSocket adds framing (message boundaries), opcodes (text/binary/ping/pong/close), and the masking requirement. You could implement WebSocket on top of any ordered reliable byte stream.

WebSocket is not the same as Server-Sent Events (SSE). SSE is still HTTP — the server sends a long-lived HTTP response with Content-Type: text/event-stream and streams chunks. It's unidirectional (server → client only), works over HTTP/1.1 or HTTP/2 without a protocol upgrade, and is much simpler to implement. WebSocket is bidirectional and requires the upgrade.

WebSocket is not HTTP/2 server push. HTTP/2 server push lets the server push resources (CSS, JS, images) into the client cache ahead of a request. It's not a messaging channel. It was deprecated in Chrome in 2022.


References

RFCs — all freely available at rfc-editor.org:

  • RFC 6455 — The WebSocket Protocol. 71 pages, but unusually readable. §4 (handshake) and §5 (framing) cover everything in this article in precise terms.
  • RFC 7692 — WebSocket Per-Message Compression (permessage-deflate).
  • RFC 8441 — Bootstrapping WebSockets with HTTP/2.

Linux kernel source (browsable at elixir.bootlin.com):

  • include/net/sock.h — struct sock definition (sk_rcvbuf, sk_sndbuf, sk_receive_queue, sk_wq)
  • include/linux/skbuff.h — struct sk_buff and the head/data/tail/end pointer model
  • net/ipv4/tcp.c — TCP send/receive paths
  • net/ipv4/tcp_input.c — SYN queue, accept queue, tcp_abort_on_overflow

Man pages:

  • man 7 socket — SO_REUSEADDR, sk_rcvbuf/sk_sndbuf sysctl knobs
  • man 7 tcp — TCP_NODELAY, tcp_abort_on_overflow, keepalive parameters
  • man 8 ss — interpreting Recv-Q/Send-Q, -tin options

I build and scale reliable production systems. Open to full-time and freelance work with U.S.-based teams that value ownership and execution.

Got something in mind?

Book a Discovery Call