Skip to main content

WebSockets at the OS Level: File Descriptors, Epoll, Buffers, and Kernel Tuning

· 10 min read
Victor Gazotti
Laravel FullStack Developer

When software engineers talk about WebSockets, the conversation usually focuses on client-side APIs (new WebSocket(url)), backend message handlers, or high-level Pub/Sub abstractions. We treat connections as lightweight objects in our code that magically stay alive and receive real-time updates.

But what actually happens underneath your application runtime? What is the Operating System kernel doing when your server maintains 100,000 or 1,000,000 concurrent WebSocket connections?

In this article, we are stepping away from high-level application frameworks to explore WebSockets at the OS level. We will inspect how Linux manages persistent TCP sockets, how event demultiplexing like epoll enables high concurrency, how kernel socket buffers consume system RAM, and how sysctl kernel tuning prevents your real-time infrastructure from falling over under load.

Understanding these OS-level mechanics changes how you architect, debug, and scale real-time applications.


1. Everything is a File: Sockets as File Descriptors​

In Unix-like operating systems, almost every I/O resource is abstracted as a file. A WebSocket connection is no exception. At the OS level, a WebSocket is simply an established TCP connection represented by a integer handle called a File Descriptor (FD).

When a client completes the HTTP 101 Switching Protocols handshake, the underlying OS socket remains in the ESTABLISHED state within the kernel's TCP stack. From that moment on, the operating system tracks the connection in its internal file descriptor table.

+-----------------------------------------------------------------+
| User Space Application |
| (Node.js / Swoole / Go / Rust Event Loop reading FD 12) |
+-----------------------------------------------------------------+
| syscall: read(12, buf, len)
v
+-----------------------------------------------------------------+
| Linux Kernel |
| FD Table -> [FD 12] -> struct socket -> struct sock |
| | |
| +--> Receive Buffer (rmem) |
| +--> Send Buffer (wmem) |
+-----------------------------------------------------------------+
^
| Network Packets (TCP/IP)
+-----------------------------------------------------------------+
| Network Interface (NIC) |
+-----------------------------------------------------------------+

The File Descriptor Bottleneck​

Because every open WebSocket occupies a file descriptor, your OS limits directly cap your maximum concurrent connections.

By default, Linux limits a single process to 1,024 open file descriptors (ulimit -n). If your application attempts to accept the 1,025th connection, the OS responds with an immediate error: EMFILE: Too many open files.

To scale WebSockets at the OS level, you must adjust both process-level and system-wide file descriptor limits:

# Check current soft and hard open file limits
ulimit -Sn
ulimit -Hn

# Set process limit in /etc/security/limits.conf for your app user
appuser soft nofile 1048576
appuser hard nofile 1048576

# Set system-wide maximum open files in /etc/sysctl.conf
fs.file-max = 2097152

Without raising these kernel limits, no amount of CPU or RAM optimization will let your server scale beyond default boundaries.


2. From Blocking I/O to Non-Blocking Event Loops: select, poll, and epoll​

To understand how modern servers handle high-concurrency WebSockets, we must examine how the OS notifies an application that a WebSocket frame has arrived on a socket.

The Problem with Thread-per-Connection​

In traditional HTTP models, a web server could assign one thread per request. The thread issues a blocking read() call on the socket and waits for data.

With WebSockets, connections stay open for minutes or hours, and clients spend 95% of their time idle. If you assign one OS thread per WebSocket connection, 100,000 connections would require 100,000 kernel threads. The memory footprint of thread stacks alone (often 2MB to 8MB per thread by default) would consume hundreds of gigabytes of RAM before processing a single byte of application payload, not to mention the massive CPU overhead of continuous context switching.

The Evolution of Event Demultiplexing​

To solve the C10K (and now C1000K) problem, operating systems introduced I/O multiplexing system calls:

  1. select() (Legacy): The application passes an array of file descriptors to the kernel. The kernel scans the array sequentially to check if any FD is readable. This is an $O(N)$ operation. At 50,000 connections, scanning the entire list on every tick destroys CPU performance.
  2. poll() (Legacy): Solves the fixed file descriptor limit of select(), but still relies on an $O(N)$ linear scan of all descriptors.
  3. epoll (Linux) / kqueue (BSD/macOS): Instead of passing all FDs every time, the application registers interest in specific file descriptors once with the kernel using epoll_ctl(). When network packets arrive at the Network Interface Card (NIC), hardware interrupts trigger the kernel TCP stack, which places ready sockets onto a red-black tree and ready list. The application calls epoll_wait(), which returns only the file descriptors that actually have data ready to read in $O(1)$ time.
Traditional Polling (select/poll):
App -> Kernel: "Are any of these 100,000 FDs ready?"
Kernel: [Scans FD 1, FD 2, FD 3 ... FD 100,000] -> $O(N)$ overhead

Event-Driven Demultiplexing (epoll):
App -> Kernel: "Tell me which FDs became active."
Kernel: "Only FD 42 and FD 891 have incoming data." -> $O(1)$ lookup

Modern high-performance WebSocket runtimes—whether Node.js (libuv), Go (netpoll), Swoole/ReactPHP (epoll wrapper), or Rust (tokio/mio)—all build their event loops on top of epoll or the newer Linux kernel subsystem io_uring.


3. Anatomy of a WebSocket Socket: Kernel Buffers and Backpressure​

A common misconception is that when an application calls socket.send(data), the data immediately flows across the wire to the client.

In reality, the data flows from User Space into Kernel Space socket buffers.

Receive (rmem) and Send (wmem) Buffers​

For every open TCP socket, the Linux kernel allocates memory structures:

  • rmem (Receive Buffer): Holds incoming TCP segments received from the network interface until the user-space event loop calls read() or recv().
  • wmem (Send Buffer): Holds outgoing WebSocket frames produced by your application until the TCP engine transmits them and receives an ACK from the client.
[ User Space Application ]
| write("Hello WebSocket")
v
[ Kernel Send Buffer (wmem) ] <--- Backpressure builds here if client is slow!
| TCP Segments / ACKs
v
[ Network Interface (NIC) ]

OS-Level Backpressure Mechanism​

What happens if a connected WebSocket client is on a slow mobile network (high latency, packet loss) and your server broadcasts 500 messages per second?

  1. Your application writes WebSocket frames to the socket.
  2. The kernel places frames into the socket's wmem buffer.
  3. Because the client is slow to acknowledge (ACK) TCP segments, the wmem buffer fills up.
  4. When wmem reaches its configured maximum capacity, the socket becomes unwriteable.
  5. In non-blocking mode, calling write() returns EAGAIN or EWOULDBLOCK.

If your application ignores this kernel signal and keeps buffering messages in user-space RAM, your process will suffer an Out-Of-Memory (OOM) crash. Respecting OS-level backpressure—by pausing writes when the socket buffer is full—is essential for building resilient real-time backend systems.


4. Kernel Tuning for High-Concurrency WebSockets (sysctl.conf)​

Out of the box, Linux kernel network defaults are tuned for short-lived, high-throughput HTTP/1.1 requests, not long-lived persistent WebSocket connections.

To run high-density WebSocket servers, you must tune specific kernel sysctl parameters in /etc/sysctl.conf.

A. Socket Buffer Memory Allocation​

If you maintain 500,000 WebSocket connections and each socket allocates a default 128KB buffer for reading and writing, your kernel will attempt to allocate 128 GB of RAM just for socket buffers!

By default, Linux autotunes TCP buffers, but you should constrain the minimum, default, and maximum buffer sizes for WebSockets where individual message frames are usually small:

# /etc/sysctl.conf

# Structure: min default max (in bytes)
# Restrain receive buffer size for high socket density
net.ipv4.tcp_rmem = 4096 87380 4194304

# Restrain send buffer size (prevents memory exhaustion on slow clients)
net.ipv4.tcp_wmem = 4096 16384 4194304

# System-wide TCP memory limits (measured in 4KB memory pages)
net.ipv4.tcp_mem = 786432 1048576 1572864

By lowering the initial socket memory footprint (tcp_wmem default around 16KB), 100,000 idle WebSockets will consume only ~1.6 GB of socket memory instead of 12+ GB.

B. TCP Connection Backlog​

When thousands of WebSocket clients reconnect simultaneously (for instance, after a network partition or deployment restart), the OS connection queue can overflow.

# Maximum number of connections allowed in the TCP listen backlog
net.core.somaxconn = 65535

# Maximum number of remembered connection requests (SYN flood prevention queue)
net.ipv4.tcp_max_syn_backlog = 65535

C. Port Exhaustion and Ephemeral Ports​

If your server acts as a WebSocket proxy, gateway, or load balancer forwarding traffic to upstream nodes, it can run out of local ephemeral ports.

# Expand the local ephemeral port range
net.ipv4.ip_local_port_range = 1024 65535

# Allow reuse of TIME_WAIT sockets for new connections when safe
net.ipv4.tcp_tw_reuse = 1

5. Context Switching and User-Space vs. Kernel-Space Boundaries​

Every time your WebSocket server receives a frame or broadcasts a message, data traverses the boundary between Kernel Space and User Space.

+-------------------------------------------------------------+
| USER SPACE: App parses WebSocket frame header (Opcode, Mask)|
+-------------------------------------------------------------+
^ |
Syscall: recv() Syscall: send()
| v
+-------------------------------------------------------------+
| KERNEL SPACE: TCP reassembly, checksum validation, buffers |
+-------------------------------------------------------------+
  1. Inbound Data Flow: Network packet arrives -> NIC interrupt -> Kernel TCP stack reassembles payload into socket receive buffer -> epoll_wait() notifies user event loop -> recv() system call copies bytes from kernel memory to user-space buffer -> App parses WebSocket frame header (masking key, opcode, payload length).
  2. Outbound Data Flow: App constructs WebSocket frame -> send() system call copies frame into kernel send buffer -> Kernel breaks frame into TCP segments -> NIC transmits packets across network.

The cost of frequent read/write system calls combined with user-to-kernel context switches adds CPU latency under high message throughput. This is why modern low-level WebSocket engines utilize socket batching (recvmmsg/sendmmsg) or zero-copy abstractions to minimize system call overhead.


6. How OS-Level Awareness Informs Real-Time Architecture​

When you understand WebSockets at the operating system level, many architectural decisions become clear:

  • Why horizontal scaling is compulsory: You cannot simply stack unlimited TCP sockets on a single server instance; kernel file descriptor limits, socket buffer memory, and CPU core interrupt distribution set hard physical boundaries.
  • Why sticky sessions or distributed Pub/Sub are needed: Because socket descriptors are bound to specific kernel network stacks on specific physical machines, broadcasting a message to 500,000 clients requires a distributed pub/sub backbone (like Redis or NATS) to route events to the correct machine owning that specific OS file descriptor.
  • Why operating WebSocket infrastructure yourself is painful: Keeping OS kernel parameters tuned, managing buffer backpressure, preventing memory leaks on slow clients, and orchestrating zero-downtime rolling deploys without terminating established file descriptors requires significant DevOps overhead.

Where Ressonance Fits In​

This operational complexity is precisely why we built Ressonance.

Ressonance abstracts the low-level Linux networking complexities—file descriptor tuning, epoll event loop management, socket memory bounds, pub/sub fanout, and backpressure—into a reliable, developer-friendly WebSocket as a Service platform.

Whether you are using Laravel Echo, Node.js, or any frontend stack, Ressonance handles the underlying low-level socket infrastructure so you can focus on building your application instead of tuning Linux kernel parameters.


Conclusion​

When we view WebSockets through the lens of the Operating System, we realize they are much more than simple JavaScript objects or application-level events. They are persistent TCP sockets governed by OS file descriptors, multiplexed by kernel event loops like epoll, buffered inside kernel memory, and constrained by TCP flow control.

By understanding how file descriptor limits (ulimit), socket memory buffers (tcp_rmem/tcp_wmem), backpressure handling, and kernel tuning (sysctl) operate under the hood, you can diagnose performance bottlenecks, eliminate random disconnects, and design systems capable of supporting millions of concurrent real-time connections.

If you want the power of high-density, production-ready WebSockets without spending weeks tuning Linux kernel parameters and managing socket servers, Ressonance provides a robust infrastructure out of the box.

👉 Create your free Ressonance account today and scale your real-time features with confidence.