Low-latency networking
Design principles from Sterbenz & Touch applied to trading networks: delay budgets, cut-through, kernel bypass, polling, open-loop recovery.
51 questions ·quiz yourself → · knowledge base projected 2026-10-03
low-latency/02-ch01-02-fundamentals.md · read the full note
Write the delay equation. Which term does cut-through switching attack, and which one does a zero-copy host stack attack?
D = (1 + h + c)·b/r + t_p. Cut-through removes most of the h·b/r term (each hop waits for the header, not the whole frame). Zero copy removes the c·b/r term.
Path segments of 50, 2, 25 and 10 µs: which one should you leave alone, and which principle says so?
The 2 µs segment: 2/87 ≈ 2% of the total. Second-Order Effect Corollary (1A).
Why can an operation that hits only 1% of packets still be on the critical path?
If order must be kept, every later packet waits behind the slow one (Critical Path Corollary, 1B, third observation).
Why should a checksum live in the trailer rather than the header?
Fields that steer processing must be decoded first; a checksum computed over the data is only ready at the end. With the checksum in the trailer, a pipeline shorter than the packet can compute or check it as the bytes stream through (2.4.5.1).
When is polling better than interrupts, according to the book? Why do trading hosts take it further?
When the protocol knows when data will arrive, because interrupts are expensive (4H). Trading hosts dedicate whole cores to spinning on the NIC because a core is cheap compared with the microseconds saved (2A).
Give one open-loop and one closed-loop way to recover lost market data.
Open loop: A/B feeds (every packet sent twice on separate paths) or FEC. Closed loop: a retransmission request or snapshot recovery after detecting a sequence gap.
A path has a one-way delay of 1 ms at 10 Gb/s. How much data is in flight, and why does that hurt feedback control?
rd = 10¹⁰ b/s × 10⁻³ s = 10⁷ bits ≈ 1.25 MB in flight one way; a feedback loop sees about 2rd ≈ 2.5 MB go by before its action takes effect, so it reacts to stale conditions (p. 5, 5D).
low-latency/03-ch03-topology.md · read the full note
Name the delay components the book uses in Ch 3. Which one changes with load?
Propagation t_p, forwarding t_f, queuing t_q, and transmission t_b = b/r (plus t_b again in store-and-forward routers). Queuing t_q changes with load.
Using the book's 0.7c for fiber, what is the one-way propagation delay over 1,200 km?
0.7 × 300,000 km/s = 210,000 km/s; 1,200 km ÷ 210,000 km/s ≈ 5.7 ms.
A path has nine 10G links and one 1G link. What can it deliver, and which principle says so?
1 Gb/s: R = min(r_i), the Network Bandwidth Principle (N-1Ab).
Why, according to the book, are links bit-serial instead of striping a flow over parallel links? Which modern feature follows the same logic?
Skew between parallel links reorders packets and would force lock-step switch planes or resequencing (p. 89). ECMP and LAG hashing per flow keep each flow on one link for the same reason.
In Fig. 3.7 the lowest-latency path and the highest-bandwidth path differ. How does a trading network deal with that?
It sends small latency-critical traffic over the fast low-bandwidth path and bulk traffic over the high-bandwidth path.
Ten receivers behind one link want the same 1 Gb/s feed. What does that link carry with unicast, and with multicast?
Unicast: up to 10 Gb/s (n·r). Multicast: 1 Gb/s (Example 3.4).
What did Example 3.3 teach about counting hops?
A "hop" can hide a whole network: provider POPs built from small routers held about 10 router hops each (p. 109).
low-latency/04-ch04-control.md · read the full note
Write D₁ for store-and-forward and d₁ for cut-through. With h = 3 and a 1500-byte frame at 10 Gb/s, how much serialization does each pay (ignore t_p, t_f, t_q)?
D₁ = h·t_b + Σt_p + (h − 1)(t_f + t_q); d₁ = t_b + Σt_p + (h − 1)·t_s. t_b = 1500 × 8 / 10¹⁰ = 1.2 µs, so store-and-forward pays 3.6 µs and cut-through 1.2 µs.
How many end-to-end latencies does a connection-oriented transfer need before the data arrives, and why?
Three: SETUP out, CONNECT back, then the data out (p. 131), one round trip more than a datagram.
Why does message switching hurt short messages even on a lightly loaded network?
Each node stores whole messages, so a short message queued behind a long one waits for the long one's full transmission at every hop (p. 126).
Root-initiated vs leaf-initiated join: which does IP multicast use, and what does the book say is missing from IP multicast addresses?
Leaf-initiated (receiver-initiated). The group address lives in a separate address space that does not help a joining node find a route to the tree (p. 145).
Forward vs backward congestion notification: which reacts faster, and why?
Backward: the congested node signals the source directly, about one end-to-end latency, with no turnaround at the receiver; forward notification needs a full round trip plus the receiver's reaction (pp. 154–155).
Why does the book want queues kept nearly empty even when buffer space is available?
A packet in a queue cannot cut through and waits behind everything ahead of it; queues should absorb only transients (N-II.4, p. 156).
Check the book's monitoring arithmetic: how many 40-byte packets per second fit in 10 Gb/s?
10¹⁰ ÷ 320 ≈ 31 million per second, about 10× the book's figure.
low-latency/05-ch05-links-and-switches.md · read the full note
From Table 5.1, what is the propagation delay per km in fiber and in air, and what does a straight radio path save over 1,000 km?
Fiber about 5 µs/km, air about 3.3 µs/km; about 1.7 ms one way over 1,000 km.
What does the input queue of a cut-through fast packet switch need to be, according to the book?
A per-byte shift register just long enough to cover the label lookup (p. 205).
Why does the book say trailers are essential for cut-through?
Values computed over the payload (CRC) can be appended or checked as the data streams by; otherwise the whole packet has to be held while it is processed (pp. 214–215).
What throughput limit does head-of-line blocking impose, and what are the ways around it?
2 − √2 ≈ 58.6%. Output queuing via speedup, internal buffering or internal expansion (Clos), or virtual output queues with a matching scheduler.
What is the processing budget for a 128-byte packet at 10 Gb/s? For a minimum Ethernet frame including preamble and gap?
100 ns (Table 5.4). (64 + 20) × 8 = 672 bits, so 67.2 ns.
Why must serial pipeline stages be sized for minimum-size packets, while parallel engines can use the average? What does the parallel approach cost?
A serial stage that is slow for one small packet delays every packet behind it; parallel engines can average out, but they reorder packets and add jitter (pp. 253–254).
Two feeds each average 4 Gb/s but burst at line rate into one 10 Gb/s egress port. What does Fig. 5.25 predict?
The bursts will overlap even though 8 Gb/s fits on average; the switch has to buffer (adding latency) or drop.
What does the book get wrong about SONET protection switching?
It prints 50 µs; the GR-253 requirement is 50 ms.
low-latency/06-ch06-end-systems.md · read the full note
What were the three end-system conjectures of the late 1980s, and what did analysis of TCP/IP find instead?
EC1 a new transport protocol, EC2 protocols on the network interface, EC3 protocol functions in hardware. Clark's analysis found the costs in the OS, per-byte operations (copying, checksumming) and timers (p. 290).
What forces a one-copy transmit path with TCP and sockets?
The sender must keep data until it is acknowledged, and socket semantics let the application reuse its buffer as soon as the call returns, so the stack keeps its own copy (pp. 295, 330).
How expensive is a context switch according to the book, what is the target per application data unit, and how do threads help?
Hundreds of RISC instructions; at most one per application data unit; threads share an address space, so switching between them involves no memory-management work (pp. 302–303).
When is polling better than interrupts? What goes wrong if the poll interval is too short or too long? What hybrid does the book suggest, and which Linux mechanism works that way?
When the protocol knows when data will arrive. Too short wastes cycles; too long delays data and needs more buffer. One interrupt per burst, polling within the burst (p. 304). Linux NAPI.
How does a protocol bypass decide which packets take the fast path?
Send and receive filters compare each packet with a template set up per connection or flow; matches take the bypass, everything else goes through the normal stack (pp. 309–310).
Using Table 6.1, how many instructions does a 1 GHz processor have per 128-byte packet at 10 Gb/s? What does that imply for the NIC design?
100 instructions in 100 ns. Header processing at that rate is marginal for an embedded processor, so the per-packet work tends to move into hardware (E-1Ch).
Why does the book say a 1 µs NIC is not worth building for ordinary LAN/WAN use, and why does trading disagree?
With milliseconds of network latency and a 100 ms user budget, a 1 µs NIC is a second-order improvement (1A). Trading is the book's "control-feedback" case, where the budget is microseconds.
What problem does a single NIC create in a NUMA multiprocessor?
The processor the NIC is attached to becomes the bottleneck for distributing data to the others; with a NIC per processor, data must still arrive at the right one (pp. 324–325).
low-latency/07-ch07-transport.md · read the full note
What are the two common misreadings of the end-to-end arguments, and what is the right question to ask instead?
E2E-Only and Everything-E2E. Ask whether a hop-by-hop copy of the function improves end-to-end performance (T-3A).
Why does reliable delivery approach 1.5 RTT at very high data rates, and what does one retransmission do?
Sending time shrinks to zero while the handshake, data and final ACK still need about 1.5 RTT; one retransmission adds another round trip, about 2.5 RTT (pp. 360–362).
What control is possible for a datagram transport such as UDP, and where should retransmission logic live for periodically synchronized state?
Only framing, multiplexing and open-loop measures; no closed-loop control. Retransmission belongs in the application's state machines (p. 365).
List the four conditions under which the book recommends open-loop error control. Does market data meet them?
Loss tolerance expressed statistically; real-time latency needs; long or high-BDP paths; receivers that can absorb the redundancy (p. 398). Market data fits the latency and redundancy conditions; its loss tolerance comes from recovery mechanisms behind the open-loop layer.
Why does the book say pure repetition is rarely useful, and why do A/B feeds use it anyway?
An erasure code gives the same protection for less overhead (1.5× vs 2×), and back-to-back repeats do not survive bursts (p. 400). A/B uses separate paths, so it survives path failures, adds no decode latency, and does not load the congested link.
Why can't the receivers of a multicast stream simply acknowledge the source, and what do NAK-based schemes need instead?
ACK implosion: messages from every receiver converge on the links near the root (Fig. 7.18). NAKs need a liveness mechanism, since silence is ambiguous (p. 393).
What does grouping small messages into one packet trade, and why do order-entry sockets avoid it?
Less per-packet overhead against the delay of waiting to fill the packet (p. 381); orders cannot wait.
Where does the bandwidth a bursty flow needs lie, and what makes it rise toward the peak?
Between the average and the peak rate (r_a < R < r_p); correlated bursts push it toward the peak (p. 407).
low-latency/08-ch08-09-applications-future.md · read the full note
Name the book's latency utility classes. Which is trading closest to, and how does it differ?
Best effort, interactive, real-time (hard and soft), deadline. Trading is closest to hard real-time (a step in utility), but the step is set relative to competitors and moves.
Write the response-time formula. Which terms can a trading firm reduce?
T_r = d_c + 2[(1 + h + c)·b/r + t_p] + d_s. The firm controls its own processing d_c and the network terms (hops, copies, rate, path length), not the exchange's processing d_s.
What is a datacycle, when does it beat request/response, and what is its market data equivalent?
Repeatedly broadcasting the whole data set; it wins when the cycle time is short compared with a request round trip (p. 464). Snapshot or refresh channels in market data feeds.
When does compression reduce end-to-end delay?
When the encode and decode time is less than the transmission time saved (p. 465).
Why should a location-independent interface still expose latency?
Because applications that could adapt to latency need to know it; hiding it makes performance unpredictable (A-4Fl, pp. 483–484).
What practical advice does Ch 9 give for preparing for a future you cannot predict?
Keep re-checking the resource trade-offs and question current traffic assumptions; design protocols and systems that can adapt (Ø4, 2A, pp. 491–495).
The lessons that matter for trading networks
Trading-network lens
The book's design principles, mapped onto trading networks in the note writer's own words.
Reference notes
- low-latency/00-book-guide.md
- low-latency/01-principles-map.md
- low-latency/02-ch01-02-fundamentals.md
- low-latency/03-ch03-topology.md
- low-latency/04-ch04-control.md
- low-latency/05-ch05-links-and-switches.md
- low-latency/06-ch06-end-systems.md
- low-latency/07-ch07-transport.md
- low-latency/08-ch08-09-applications-future.md