← Low-latency networking

Ch 6: End systems

Source: Sterbenz & Touch 2001, Ch 6 (pp. 285–342), read in full. Page numbers are the book’s printed pages. For a trading engineer this chapter is the theory behind kernel bypass: copies, context switches, user/kernel crossings, interrupts vs polling, and what belongs on the NIC.

Numbering note: the chapter prints some IDs differently from Appendix A (Context Switch Avoidance is “E-II.6a” in the chapter, E-II.6c in the appendix; the ILP principle is “E-4E” vs E-4D; the chapter’s closing list swaps the labels of copy minimization and remapping). These notes use the Appendix A IDs.

The chapter in one paragraph

Without a fast, short path between the NIC and application memory, the host is the bottleneck however fast the network is (E-II). Find the real bottleneck by measuring: in the late 1980s people blamed the transport protocol, but the costs were in the operating system, in per-byte work (copies, checksums) and in timers. The target is a zero-copy path, about one context switch per application data unit, few user/kernel crossings, polling where arrivals are predictable, all protocol passes done in one loop (ILP), a bypass path for the common case, a nonblocking path from NIC to memory, and a NIC whose hardware/software split is set by the packet interarrival time.

Find the real bottleneck first (pp. 285–291)

Why traditional hosts were slow (pp. 291–294)

The ideal: zero copy (pp. 294–296)

Protocol software (pp. 296–301)

The operating system (pp. 301–309)

Protocol optimizations (pp. 309–313)

Host organization (pp. 313–326)

The network interface (pp. 326–340)

Rate, packet size Time per packet 100 MHz CPU 1 GHz CPU
1 Gb/s, 32 B 250 ns 25 250
1 Gb/s, 1 KB 8 µs 800 8,000
10 Gb/s, 128 B 100 ns 10 100
10 Gb/s, 1 KB 800 ns 80 800

A budget of 25 instructions is “marginal, at best; 80 instructions is more likely to be feasible” (footnote 27). Below that, the processor saturates (Fig. 6.19) and the work must move into hardware.

Trading-network lens (my mapping)

Kernel bypass is this chapter in one product. AMD’s Onload is “a high performance user-level network stack, which accelerates TCP and UDP network I/O for applications using the BSD sockets on Linux”. It is “a user-level shared library that intercepts network-related system calls and implements the protocol stack, and supporting kernel modules”, and it uses the ef_vi interface of Solarflare NICs (checked: https://github.com/Xilinx-CNS/onload). In the book’s terms: a protocol bypass with a fallback to the normal stack, no copy through the kernel, and no user/kernel crossing per packet. The book judged user-space protocols slow because each packet needed system calls (Fig. 6.8a); bypass stacks avoid that by giving the process direct access to NIC queues. No single source for that last point.

Polling, as Linux does it. NAPI is the book’s hybrid: the device interrupts, then the driver keeps interrupts masked while it polls until the work is done. Busy polling “allows a user process to check for incoming packets before the device interrupt fires” and “trades off CPU cycles for lower latency”; it is enabled per socket with SO_BUSY_POLL or system-wide with the net.core.busy_poll and net.core.busy_read sysctls (checked: https://docs.kernel.org/networking/napi.html). Trading systems go further and dedicate whole cores to spinning on the NIC: principle 4H applied with 2026 prices.

From one context switch per ADU to none. Pinning the hot thread to an isolated core, away from the scheduler and from interrupts, is E-II.6c carried to its limit. No single source.

Cache discipline. The book’s code advice (loops inside the I-cache, data aligned to cache lines, one off-chip read can cost more than the rest of the path) is tick-to-trade coding practice today. Its idea of the NIC writing straight into the cache (p. 322) exists in current server CPUs (for example Intel’s Data Direct I/O). No single source.

NUMA (E-II.4m). Keep the NIC, its queues and interrupts, the buffers and the trading thread on the same CPU socket; crossing sockets adds latency. No single source.

The instruction budget is still the limit. At 10 Gb/s a minimum 64-byte frame arrives every 67 ns, about 200 cycles on a 3 GHz core: the book’s Table 6.1 problem, one generation later. That is why FPGA NICs parse and filter market data in hardware and hand software only the messages it needs (E-1Ch: interarrival time decides). Arithmetic plus No single source.

Benchmark with the application running (E-I). A NIC ping-pong number says as little about tick-to-trade as “this TCP runs at 1 Gb/s” said about applications. Measure under the real workload. My mapping.

Acting before the check. The book insists data must not be used before its check passes (p. 339). Hardware that starts decoding a message before the Ethernet FCS has arrived must be able to throw that work away if the FCS turns out bad. My mapping, No single source.

Self-check

  1. What were the three end-system conjectures of the late 1980s, and what did analysis of TCP/IP find instead?
  2. What forces a one-copy transmit path with TCP and sockets?
  3. How expensive is a context switch according to the book, what is the target per application data unit, and how do threads help?
  4. When is polling better than interrupts? What goes wrong if the poll interval is too short or too long? What hybrid does the book suggest, and which Linux mechanism works that way?
  5. How does a protocol bypass decide which packets take the fast path?
  6. Using Table 6.1, how many instructions does a 1 GHz processor have per 128-byte packet at 10 Gb/s? What does that imply for the NIC design?
  7. Why does the book say a 1 µs NIC is not worth building for ordinary LAN/WAN use, and why does trading disagree?
  8. What problem does a single NIC create in a NUMA multiprocessor?
Answers
  1. EC1 a new transport protocol, EC2 protocols on the network interface, EC3 protocol functions in hardware. Clark’s analysis found the costs in the OS, per-byte operations (copying, checksumming) and timers (p. 290).
  2. The sender must keep data until it is acknowledged, and socket semantics let the application reuse its buffer as soon as the call returns, so the stack keeps its own copy (pp. 295, 330).
  3. Hundreds of RISC instructions; at most one per application data unit; threads share an address space, so switching between them involves no memory-management work (pp. 302–303).
  4. When the protocol knows when data will arrive. Too short wastes cycles; too long delays data and needs more buffer. One interrupt per burst, polling within the burst (p. 304). Linux NAPI.
  5. Send and receive filters compare each packet with a template set up per connection or flow; matches take the bypass, everything else goes through the normal stack (pp. 309–310).
  6. 100 instructions in 100 ns. Header processing at that rate is marginal for an embedded processor, so the per-packet work tends to move into hardware (E-1Ch).
  7. With milliseconds of network latency and a 100 ms user budget, a 1 µs NIC is a second-order improvement (1A). Trading is the book’s “control-feedback” case, where the budget is microseconds.
  8. The processor the NIC is attached to becomes the bottleneck for distributing data to the others; with a NIC per processor, data must still arrive at the right one (pp. 324–325).

Source: knowledge base note low-latency/06-ch06-end-systems.md — own-words notes with sources, projected at build time.