← All topics

Low-latency networking

Design principles from Sterbenz & Touch applied to trading networks: delay budgets, cut-through, kernel bypass, polling, open-loop recovery.

51 questions ·quiz yourself → · knowledge base projected 2026-10-03

low-latency/02-ch01-02-fundamentals.md · read the full note

Write the delay equation. Which term does cut-through switching attack, and which one does a zero-copy host stack attack?

D = (1 + h + c)·b/r + t_p. Cut-through removes most of the h·b/r term (each hop waits for the header, not the whole frame). Zero copy removes the c·b/r term.

From low-latency/02-ch01-02-fundamentals.md · link
Path segments of 50, 2, 25 and 10 µs: which one should you leave alone, and which principle says so?

The 2 µs segment: 2/87 ≈ 2% of the total. Second-Order Effect Corollary (1A).

From low-latency/02-ch01-02-fundamentals.md · link
Why can an operation that hits only 1% of packets still be on the critical path?

If order must be kept, every later packet waits behind the slow one (Critical Path Corollary, 1B, third observation).

From low-latency/02-ch01-02-fundamentals.md · link
Why should a checksum live in the trailer rather than the header?

Fields that steer processing must be decoded first; a checksum computed over the data is only ready at the end. With the checksum in the trailer, a pipeline shorter than the packet can compute or check it as the bytes stream through (2.4.5.1).

From low-latency/02-ch01-02-fundamentals.md · link
When is polling better than interrupts, according to the book? Why do trading hosts take it further?

When the protocol knows when data will arrive, because interrupts are expensive (4H). Trading hosts dedicate whole cores to spinning on the NIC because a core is cheap compared with the microseconds saved (2A).

From low-latency/02-ch01-02-fundamentals.md · link
Give one open-loop and one closed-loop way to recover lost market data.

Open loop: A/B feeds (every packet sent twice on separate paths) or FEC. Closed loop: a retransmission request or snapshot recovery after detecting a sequence gap.

From low-latency/02-ch01-02-fundamentals.md · link
A path has a one-way delay of 1 ms at 10 Gb/s. How much data is in flight, and why does that hurt feedback control?

rd = 10¹⁰ b/s × 10⁻³ s = 10⁷ bits ≈ 1.25 MB in flight one way; a feedback loop sees about 2rd ≈ 2.5 MB go by before its action takes effect, so it reacts to stale conditions (p. 5, 5D).

From low-latency/02-ch01-02-fundamentals.md · link

low-latency/03-ch03-topology.md · read the full note

Name the delay components the book uses in Ch 3. Which one changes with load?

Propagation t_p, forwarding t_f, queuing t_q, and transmission t_b = b/r (plus t_b again in store-and-forward routers). Queuing t_q changes with load.

From low-latency/03-ch03-topology.md · link
Using the book's 0.7c for fiber, what is the one-way propagation delay over 1,200 km?

0.7 × 300,000 km/s = 210,000 km/s; 1,200 km ÷ 210,000 km/s ≈ 5.7 ms.

From low-latency/03-ch03-topology.md · link
A path has nine 10G links and one 1G link. What can it deliver, and which principle says so?

1 Gb/s: R = min(r_i), the Network Bandwidth Principle (N-1Ab).

From low-latency/03-ch03-topology.md · link
Why, according to the book, are links bit-serial instead of striping a flow over parallel links? Which modern feature follows the same logic?

Skew between parallel links reorders packets and would force lock-step switch planes or resequencing (p. 89). ECMP and LAG hashing per flow keep each flow on one link for the same reason.

From low-latency/03-ch03-topology.md · link
In Fig. 3.7 the lowest-latency path and the highest-bandwidth path differ. How does a trading network deal with that?

It sends small latency-critical traffic over the fast low-bandwidth path and bulk traffic over the high-bandwidth path.

From low-latency/03-ch03-topology.md · link
Ten receivers behind one link want the same 1 Gb/s feed. What does that link carry with unicast, and with multicast?

Unicast: up to 10 Gb/s (n·r). Multicast: 1 Gb/s (Example 3.4).

From low-latency/03-ch03-topology.md · link
What did Example 3.3 teach about counting hops?

A "hop" can hide a whole network: provider POPs built from small routers held about 10 router hops each (p. 109).

From low-latency/03-ch03-topology.md · link

low-latency/04-ch04-control.md · read the full note

Write D₁ for store-and-forward and d₁ for cut-through. With h = 3 and a 1500-byte frame at 10 Gb/s, how much serialization does each pay (ignore t_p, t_f, t_q)?

D₁ = h·t_b + Σt_p + (h − 1)(t_f + t_q); d₁ = t_b + Σt_p + (h − 1)·t_s. t_b = 1500 × 8 / 10¹⁰ = 1.2 µs, so store-and-forward pays 3.6 µs and cut-through 1.2 µs.

From low-latency/04-ch04-control.md · link
How many end-to-end latencies does a connection-oriented transfer need before the data arrives, and why?

Three: SETUP out, CONNECT back, then the data out (p. 131), one round trip more than a datagram.

From low-latency/04-ch04-control.md · link
Why does message switching hurt short messages even on a lightly loaded network?

Each node stores whole messages, so a short message queued behind a long one waits for the long one's full transmission at every hop (p. 126).

From low-latency/04-ch04-control.md · link
Root-initiated vs leaf-initiated join: which does IP multicast use, and what does the book say is missing from IP multicast addresses?

Leaf-initiated (receiver-initiated). The group address lives in a separate address space that does not help a joining node find a route to the tree (p. 145).

From low-latency/04-ch04-control.md · link
Forward vs backward congestion notification: which reacts faster, and why?

Backward: the congested node signals the source directly, about one end-to-end latency, with no turnaround at the receiver; forward notification needs a full round trip plus the receiver's reaction (pp. 154–155).

From low-latency/04-ch04-control.md · link
Why does the book want queues kept nearly empty even when buffer space is available?

A packet in a queue cannot cut through and waits behind everything ahead of it; queues should absorb only transients (N-II.4, p. 156).

From low-latency/04-ch04-control.md · link
Check the book's monitoring arithmetic: how many 40-byte packets per second fit in 10 Gb/s?

10¹⁰ ÷ 320 ≈ 31 million per second, about 10× the book's figure.

From low-latency/04-ch04-control.md · link

low-latency/05-ch05-links-and-switches.md · read the full note

From Table 5.1, what is the propagation delay per km in fiber and in air, and what does a straight radio path save over 1,000 km?

Fiber about 5 µs/km, air about 3.3 µs/km; about 1.7 ms one way over 1,000 km.

From low-latency/05-ch05-links-and-switches.md · link
What does the input queue of a cut-through fast packet switch need to be, according to the book?

A per-byte shift register just long enough to cover the label lookup (p. 205).

From low-latency/05-ch05-links-and-switches.md · link
Why does the book say trailers are essential for cut-through?

Values computed over the payload (CRC) can be appended or checked as the data streams by; otherwise the whole packet has to be held while it is processed (pp. 214–215).

From low-latency/05-ch05-links-and-switches.md · link
What throughput limit does head-of-line blocking impose, and what are the ways around it?

2 − √2 ≈ 58.6%. Output queuing via speedup, internal buffering or internal expansion (Clos), or virtual output queues with a matching scheduler.

From low-latency/05-ch05-links-and-switches.md · link
What is the processing budget for a 128-byte packet at 10 Gb/s? For a minimum Ethernet frame including preamble and gap?

100 ns (Table 5.4). (64 + 20) × 8 = 672 bits, so 67.2 ns.

From low-latency/05-ch05-links-and-switches.md · link
Why must serial pipeline stages be sized for minimum-size packets, while parallel engines can use the average? What does the parallel approach cost?

A serial stage that is slow for one small packet delays every packet behind it; parallel engines can average out, but they reorder packets and add jitter (pp. 253–254).

From low-latency/05-ch05-links-and-switches.md · link
Two feeds each average 4 Gb/s but burst at line rate into one 10 Gb/s egress port. What does Fig. 5.25 predict?

The bursts will overlap even though 8 Gb/s fits on average; the switch has to buffer (adding latency) or drop.

From low-latency/05-ch05-links-and-switches.md · link
What does the book get wrong about SONET protection switching?

It prints 50 µs; the GR-253 requirement is 50 ms.

From low-latency/05-ch05-links-and-switches.md · link

low-latency/06-ch06-end-systems.md · read the full note

What were the three end-system conjectures of the late 1980s, and what did analysis of TCP/IP find instead?

EC1 a new transport protocol, EC2 protocols on the network interface, EC3 protocol functions in hardware. Clark's analysis found the costs in the OS, per-byte operations (copying, checksumming) and timers (p. 290).

From low-latency/06-ch06-end-systems.md · link
What forces a one-copy transmit path with TCP and sockets?

The sender must keep data until it is acknowledged, and socket semantics let the application reuse its buffer as soon as the call returns, so the stack keeps its own copy (pp. 295, 330).

From low-latency/06-ch06-end-systems.md · link
How expensive is a context switch according to the book, what is the target per application data unit, and how do threads help?

Hundreds of RISC instructions; at most one per application data unit; threads share an address space, so switching between them involves no memory-management work (pp. 302–303).

From low-latency/06-ch06-end-systems.md · link
When is polling better than interrupts? What goes wrong if the poll interval is too short or too long? What hybrid does the book suggest, and which Linux mechanism works that way?

When the protocol knows when data will arrive. Too short wastes cycles; too long delays data and needs more buffer. One interrupt per burst, polling within the burst (p. 304). Linux NAPI.

From low-latency/06-ch06-end-systems.md · link
How does a protocol bypass decide which packets take the fast path?

Send and receive filters compare each packet with a template set up per connection or flow; matches take the bypass, everything else goes through the normal stack (pp. 309–310).

From low-latency/06-ch06-end-systems.md · link
Using Table 6.1, how many instructions does a 1 GHz processor have per 128-byte packet at 10 Gb/s? What does that imply for the NIC design?

100 instructions in 100 ns. Header processing at that rate is marginal for an embedded processor, so the per-packet work tends to move into hardware (E-1Ch).

From low-latency/06-ch06-end-systems.md · link
Why does the book say a 1 µs NIC is not worth building for ordinary LAN/WAN use, and why does trading disagree?

With milliseconds of network latency and a 100 ms user budget, a 1 µs NIC is a second-order improvement (1A). Trading is the book's "control-feedback" case, where the budget is microseconds.

From low-latency/06-ch06-end-systems.md · link
What problem does a single NIC create in a NUMA multiprocessor?

The processor the NIC is attached to becomes the bottleneck for distributing data to the others; with a NIC per processor, data must still arrive at the right one (pp. 324–325).

From low-latency/06-ch06-end-systems.md · link

low-latency/07-ch07-transport.md · read the full note

What are the two common misreadings of the end-to-end arguments, and what is the right question to ask instead?

E2E-Only and Everything-E2E. Ask whether a hop-by-hop copy of the function improves end-to-end performance (T-3A).

From low-latency/07-ch07-transport.md · link
Why does reliable delivery approach 1.5 RTT at very high data rates, and what does one retransmission do?

Sending time shrinks to zero while the handshake, data and final ACK still need about 1.5 RTT; one retransmission adds another round trip, about 2.5 RTT (pp. 360–362).

From low-latency/07-ch07-transport.md · link
What control is possible for a datagram transport such as UDP, and where should retransmission logic live for periodically synchronized state?

Only framing, multiplexing and open-loop measures; no closed-loop control. Retransmission belongs in the application's state machines (p. 365).

From low-latency/07-ch07-transport.md · link
List the four conditions under which the book recommends open-loop error control. Does market data meet them?

Loss tolerance expressed statistically; real-time latency needs; long or high-BDP paths; receivers that can absorb the redundancy (p. 398). Market data fits the latency and redundancy conditions; its loss tolerance comes from recovery mechanisms behind the open-loop layer.

From low-latency/07-ch07-transport.md · link
Why does the book say pure repetition is rarely useful, and why do A/B feeds use it anyway?

An erasure code gives the same protection for less overhead (1.5× vs 2×), and back-to-back repeats do not survive bursts (p. 400). A/B uses separate paths, so it survives path failures, adds no decode latency, and does not load the congested link.

From low-latency/07-ch07-transport.md · link
Why can't the receivers of a multicast stream simply acknowledge the source, and what do NAK-based schemes need instead?

ACK implosion: messages from every receiver converge on the links near the root (Fig. 7.18). NAKs need a liveness mechanism, since silence is ambiguous (p. 393).

From low-latency/07-ch07-transport.md · link
What does grouping small messages into one packet trade, and why do order-entry sockets avoid it?

Less per-packet overhead against the delay of waiting to fill the packet (p. 381); orders cannot wait.

From low-latency/07-ch07-transport.md · link
Where does the bandwidth a bursty flow needs lie, and what makes it rise toward the peak?

Between the average and the peak rate (r_a < R < r_p); correlated bursts push it toward the peak (p. 407).

From low-latency/07-ch07-transport.md · link

low-latency/08-ch08-09-applications-future.md · read the full note

Name the book's latency utility classes. Which is trading closest to, and how does it differ?

Best effort, interactive, real-time (hard and soft), deadline. Trading is closest to hard real-time (a step in utility), but the step is set relative to competitors and moves.

From low-latency/08-ch08-09-applications-future.md · link
Write the response-time formula. Which terms can a trading firm reduce?

T_r = d_c + 2[(1 + h + c)·b/r + t_p] + d_s. The firm controls its own processing d_c and the network terms (hops, copies, rate, path length), not the exchange's processing d_s.

From low-latency/08-ch08-09-applications-future.md · link
What is a datacycle, when does it beat request/response, and what is its market data equivalent?

Repeatedly broadcasting the whole data set; it wins when the cycle time is short compared with a request round trip (p. 464). Snapshot or refresh channels in market data feeds.

From low-latency/08-ch08-09-applications-future.md · link
When does compression reduce end-to-end delay?

When the encode and decode time is less than the transmission time saved (p. 465).

From low-latency/08-ch08-09-applications-future.md · link
Why should a location-independent interface still expose latency?

Because applications that could adapt to latency need to know it; hiding it makes performance unpredictable (A-4Fl, pp. 483–484).

From low-latency/08-ch08-09-applications-future.md · link
What practical advice does Ch 9 give for preparing for a future you cannot predict?

Keep re-checking the resource trade-offs and question current traffic assumptions; design protocols and systems that can adapt (Ø4, 2A, pp. 491–495).

From low-latency/08-ch08-09-applications-future.md · link

The lessons that matter for trading networks

Delay is a sum; optimize the biggest term
D = (1 + h + c)·b/r + t_p, and a part that contributes 2% is not worth touching (1A, N-1Al). Measure every hop of tick-to-trade before buying anything. [02](/low-latency/ch01-02-fundamentals/), [03](/low-latency/ch03-topology/)
In a colocation, distance is not the problem
Propagation over metres of fiber is nanoseconds, so switch hops, serialization, queuing, copies and the application dominate. Over long routes it is the reverse, and only a straighter or faster medium helps (fiber about 5 µs/km, radio about 3.3 µs/km). [03](/low-latency/ch03-topology/), [05](/low-latency/ch05-links-and-switches/)
Never store and forward, never copy
Cut-through switching pays the serialization time once instead of at every hop; layer-1 switches go further; zero-copy kernel-bypass stacks remove the copies in the host (II.3, S-II.3, E-II.3). [04](/low-latency/ch04-control/), [05](/low-latency/ch05-links-and-switches/), [06](/low-latency/ch06-end-systems/)
A queue is latency, and bursts collide
Keep queues nearly empty and links left of the knee; bursty flows collide at an output port even when the average fits, and correlated bursts defeat statistical multiplexing (N-II.4, Fig. 5.25, Fig. 7.22, p. 407). [05](/low-latency/ch05-links-and-switches/), [07](/low-latency/ch07-transport/)
Packet rate, not bandwidth, sets the design
The time per packet decides what can be done in software and what must be hardware; serial stages must handle the worst case, not the average (S-I.3, E-1Ch, Tables 5.4 and 6.1). [05](/low-latency/ch05-links-and-switches/), [06](/low-latency/ch06-end-systems/)
Poll when you know data is coming
Interrupts, context switches and user/kernel crossings each cost microseconds; with cheap cores, spinning on the NIC is the right trade (4H, E-II.6c, E-II.6k, 2A). [06](/low-latency/ch06-end-systems/)
Multicast is a resource trade-off, and its tree is state
One copy per link instead of n; switches replicate in hardware; group joins and trees are state that must be set up before it is needed and kept alive (2C, Example 3.4, N-6B, p. 147). [03](/low-latency/ch03-topology/), [04](/low-latency/ch04-control/), [05](/low-latency/ch05-links-and-switches/)
Prefer open loop when round trips are expensive
Redundant A/B paths, snapshots (a datacycle) and FEC recover without waiting; any recovery that needs a round trip costs at least an RTT (6D, T-6De, Fig. 7.6, p. 464). [07](/low-latency/ch07-transport/), [08](/low-latency/ch08-09-applications-future/)
Late is as bad as lost
For real-time data, do not trade timeliness for reliability; a packet far out of order is no better than a lost one (T-I.2, p. 371). [07](/low-latency/ch07-transport/)
Pay setup costs before the open, and re-check trade-offs as prices change
Sessions, joins and connections belong outside the critical path (5A, 6A); the right design moves as bandwidth, processing and memory change in relative cost (Ø4, 2A). [02](/low-latency/ch01-02-fundamentals/), [08](/low-latency/ch08-09-applications-future/)

Trading-network lens

The book's design principles, mapped onto trading networks in the note writer's own words.

The delay model with today's numbersNo single source
Serialization b/r for one frame (frame bytes only; preamble and inter-frame gap add 20 bytes on the wire): | Frame | 10 Gb/s | 25 Gb/s | |---|---|---| | 64 B | 51 ns | 20 ns | | 100 B | 80 ns | 32 ns | | 1500 B | 1.2 µs | 480 ns | Propagation: light in fiber covers about 1 km in 4.9–5 µs (refractive index ≈ 1.47), in air about 1 km in 3.3 µs. So a 100 m cross-connect costs about 0.5 µs one way, and over 1000 km a straight-line radio path beats fiber by more than 1.5 ms one way, before counting that fiber routes are longer than the straight line. What this means: - In a colocation the t_p propagation term is small, so the (1 + h + c) terms and the time inside hosts and applications dominate. The book's WAN-centric advice (cut propagation) turns into: remove hops, store-and-forward and copies. - With 1500-byte frames, three store-and-forward hops add 3.6 µs of serialization alone at 10 Gb/s; cut-through leaves roughly the header time per hop. With small market data frames the gap is smaller, so measure before buying. - Cut-through has limits that match the book's Fig. 2.16: it only works when the output port is free. A queued frame waits; and switches generally store and forward when the egress port is faster than the ingress port (an underrun risk). **No single source**.
Critical path dependency (1B, third observation)
A feed handler must process messages in sequence order. A rare message type handled on a slow path, or a gap waiting for recovery, delays every message behind it. That rare path is on the critical path.
Measure before optimizing (1A)
Break tick-to-trade into wire → NIC → host stack → application → host stack → NIC → wire and time each part with hardware timestamps. Optimize the largest term first.
Header first, checksum last (2.4.5.1)No single source
Ethernet already follows this ordering: destination address first, FCS at the end. That is what lets a cut-through switch forward after reading the header, and it is also why a cut-through switch cannot drop a corrupt frame: the FCS arrives after forwarding has started, so the error only shows up downstream. **No single source**.
Interrupts vs polling (4H) and resource trade-offs (2A)
Trading hosts dedicate CPU cores to busy-polling the NIC: cores became cheap enough that burning one to save the interrupt and wake-up latency is a good trade. The book's principle, with today's prices.
Separate control mechanisms (6E) and loss characterization (L-4F)
UDP market data has no congestion control, so loss comes from buffer overflow (microbursts) or line errors. Check which: output discards on the switch point to bursts, CRC errors point to a bad link. See [../multicast/02-microburst-buffer.md](/multicast/microburst-buffer/).
Hard state amortized (5A) and minimize round trips (6A)
Order-entry sessions are opened and logged in before the market opens, and feed handlers join their multicast groups at startup, so setup round trips stay off the critical path during trading.
Latency is a sum, and only t_q movesNo single source
In a colocation, model the path as NIC → switch → switch → exchange handoff and assign each term: propagation (metres of fiber), per-switch forwarding, serialization, queuing. All but the queuing term are roughly constant; queuing is where jitter and microburst loss come from. **No single source**.
DiameterNo single source
The book's 10-hop target is for a WAN. For the latency-critical path in a trading site the target is one or two switch hops, with layer-1 switches for pure fan-out, and that drives flat designs. **No single source**.
Straight paths are the productUnverified
The book's Boston–New York-via-Chicago example (12 ms vs 1.5 ms) is the trading-route problem in miniature: for long routes such as Chicago–New Jersey, firms pay for the straightest path they can get. **Unverified** (route-specific figures not checked yet).
The fastest path is not the fattest (Fig. 3.7)No single source
A low-latency radio route has less bandwidth than fiber, so trading networks split traffic: small latency-critical messages on the fast, thin path, bulk traffic on fiber. **No single source**.
Striping and per-flow hashingNo single source
The book's skew argument is why ECMP and link aggregation hash per flow: all packets of one flow stay on one member link and stay in order. A single market data feed therefore uses one member link, however many there are. **No single source**.
Bottlenecks and fan-in
R = min(r_i) also applies in time: several feeds merged onto one 10G link make that link the bottleneck during bursts, which is where microburst drops happen. See [../multicast/02-microburst-buffer.md](/multicast/microburst-buffer/).
Multicast (Example 3.4)
One feed to many consumers costs r per link instead of n·r, which is why market data is distributed by multicast. See [../multicast/00-glossary.md](/multicast/glossary/).
Overlays hide the path (N-IIo)No single source
Tunnels hide the physical route from routing and add encapsulation, so keep the latency-critical path on an underlay you can see and control. **No single source**.
"When is a hop not a hop?"No single source
Know what is inside each hop you pay for: a single "cross-connect" can contain patch panels, a media converter or a provider switch. **No single source**.
Latency gets relatively more important (p. 114)
The book's 2001 observation is the economics of the low-latency trading race: bandwidth kept getting cheaper, distance did not.
Read switch latency specs with the formulas
Store-and-forward latency grows with frame size (the h·t_b term); cut-through latency does not. Benchmarks also measure in two ways. RFC 1242 (July 1991, section 3.8) defines latency for a store-and-forward device as last bit in → first bit out (LIFO), and for a bit-forwarding device as first bit in → first bit out (FIFO). It adds that a bridge which starts sending before the frame has fully arrived (a "cut through" device) is still measured last-bit-in to first-bit-out, "even though the value would be negative". The gap between the two measures is exactly the frame's serialization time t_b, so always ask which one a datasheet quotes and at which frame size. Checked against .
Big frames delay small ones (p. 126)No single source
A 9,000-byte jumbo frame holds a 10 Gb/s port for 7.2 µs; an order that arrives just behind it waits. Keep latency-critical traffic off links and queues that carry bulk transfers, or give it strict priority. **No single source**.
Pay setup costs before the open
Order-entry TCP sessions are opened and logged in early, and feed handlers join their multicast groups at startup and stay joined, the book's advice to reuse expensive connections such as multicast trees (p. 147). The IGMP querier trap in [../multicast/03-failure-modes.md](/multicast/failure-modes/) is what happens when that state quietly times out.
Why PIM-SM needs an RP
The book complains that an IP multicast group address says nothing about where the tree is (p. 145). PIM-SM fixes this with a rendezvous point that everyone can find; SSM fixes it by having receivers name the source (S,G) themselves. See [../multicast/00-glossary.md](/multicast/glossary/). My mapping.
No closed loop for market data
UDP multicast market data has no congestion control at all, so the only tools are the open-loop ones: engineer enough capacity and keep queues short (N-II.4). The book's 110% OC-192 example is the microburst case: past 100%, the excess is simply lost. See [../multicast/02-microburst-buffer.md](/multicast/microburst-buffer/).
Rerouting reordersNo single source
When a path changes (routing convergence, a link in a LAG fails), packets of one feed can arrive out of order or with a gap; feed handlers need a reordering window, and A/B arbitration helps fill the gap. **No single source**.
Monitoring in hardware (N-1B)
At tens of millions of packets per second per port, latency and microburst monitoring has to be done by hardware timestamping and capture, not by polling counters. See the BurstRadar notes in [../multicast/02-microburst-buffer.md](/multicast/microburst-buffer/).
The medium is a latency choice
From the book's Table 5.1: fiber about 5 µs/km, radio about 3.3 µs/km, so a straight radio path saves about 1.7 ms per 1,000 km one way. Inside a building, fiber costs about 5 ns per metre, so cable lengths matter only at the nanosecond level. Arithmetic from the book's table.
Layer-1 switches are the book's unbuffered crossbar
Arista's 7130 Connect series forwards port to port in 4 ns with "full signal recovery and regeneration", does "not buffer or queue data", and allows one-to-many connections with the same latency (vendor product page, checked: ). That is the book's electronic crosspoint with its free duplicate state (p. 239), used for market data fan-out. The price of having no buffer: two inputs cannot be merged onto one output at layer 1; merging needs a device with buffering and arbitration.
Burst collision is the microburst
Fig. 5.25 is the fan-in problem described by Arista and Pico: feeds that each fit the egress link on average still collide in time, and the switch must queue (latency) or drop (gaps). See [../multicast/02-microburst-buffer.md](/multicast/microburst-buffer/).
Packet rate, not bandwidthNo single source
Market data messages are small, and bursts arrive as packet-rate spikes. At 10 Gb/s a minimum 64-byte frame plus 20 bytes of preamble and gap is 67 ns, or 14.88 million frames per second. Whatever is serial in a NIC, kernel or feed handler must keep up with the worst case (S-I.3); spreading flows over cores, like the book's parallel forwarding engines, raises throughput but must hash per flow so each feed stays in order. Arithmetic is **No single source**; the core-spreading comparison is my mapping.
Worst case, not average (p. 256)
The book's warning about an occasional slow lookup backing up the queue applies to any per-message step on the hot path, such as an order-book update that sometimes allocates memory.
Multicast replication has its own scheduling
When a switch replicates a market data packet to many egress ports, a busy output delays only its own copy if the switch splits the fanout, or holds everything if it does not. Latency can therefore differ between subscriber ports of the same switch. My mapping; check the specific switch's documentation.
Classify at ingress (S-II.4c)
QoS marking and ACLs must act at line rate before queuing, so a strict-priority queue for orders actually protects them. My mapping.
FPGAs (p. 224)No single source
The book already notes FPGAs as a fast, reprogrammable middle ground for packet processing; trading later put feed handling and order entry on FPGA NICs and switches. **No single source**.
Protection switching vs redundant feeds
SONET protection takes up to 50 ms and IP reconvergence takes longer, while redundant A/B feeds lose nothing during a single failure. See the JPX and MOEX cases in [../multicast/01-incidents.md](/multicast/incidents/). My mapping.
Kernel bypass is this chapter in one productNo single source
AMD's Onload is "a high performance user-level network stack, which accelerates TCP and UDP network I/O for applications using the BSD sockets on Linux". It is "a user-level shared library that intercepts network-related system calls and implements the protocol stack, and supporting kernel modules", and it uses the ef_vi interface of Solarflare NICs (checked: ). In the book's terms: a protocol bypass with a fallback to the normal stack, no copy through the kernel, and no user/kernel crossing per packet. The book judged user-space protocols slow because each packet needed system calls (Fig. 6.8a); bypass stacks avoid that by giving the process direct access to NIC queues. **No single source** for that last point.
Polling, as Linux does it
NAPI is the book's hybrid: the device interrupts, then the driver keeps interrupts masked while it polls until the work is done. Busy polling "allows a user process to check for incoming packets before the device interrupt fires" and "trades off CPU cycles for lower latency"; it is enabled per socket with `SO_BUSY_POLL` or system-wide with the `net.core.busy_poll` and `net.core.busy_read` sysctls (checked: ). Trading systems go further and dedicate whole cores to spinning on the NIC: principle 4H applied with 2026 prices.
From one context switch per ADU to noneNo single source
Pinning the hot thread to an isolated core, away from the scheduler and from interrupts, is E-II.6c carried to its limit. **No single source**.
Cache disciplineNo single source
The book's code advice (loops inside the I-cache, data aligned to cache lines, one off-chip read can cost more than the rest of the path) is tick-to-trade coding practice today. Its idea of the NIC writing straight into the cache (p. 322) exists in current server CPUs (for example Intel's Data Direct I/O). **No single source**.
NUMA (E-II.4m)No single source
Keep the NIC, its queues and interrupts, the buffers and the trading thread on the same CPU socket; crossing sockets adds latency. **No single source**.
The instruction budget is still the limitNo single source
At 10 Gb/s a minimum 64-byte frame arrives every 67 ns, about 200 cycles on a 3 GHz core: the book's Table 6.1 problem, one generation later. That is why FPGA NICs parse and filter market data in hardware and hand software only the messages it needs (E-1Ch: interarrival time decides). Arithmetic plus **No single source**.
Benchmark with the application running (E-I)
A NIC ping-pong number says as little about tick-to-trade as "this TCP runs at 1 Gb/s" said about applications. Measure under the real workload. My mapping.
Acting before the checkNo single source
The book insists data must not be used before its check passes (p. 339). Hardware that starts decoding a message before the Ethernet FCS has arrived must be able to throw that work away if the FCS turns out bad. My mapping, **No single source**.
Market data is the datagram case
UDP multicast feeds have no connection and no closed loop at all (p. 365). Reliability moves into the application: the feed's own sequence numbers, gap detection and recovery logic are Application Layer Framing (T-4C) in practice.
A/B feeds are the book's spatial redundancy
Sending every packet twice over separate networks is repetition, which the book calls "rarely useful" next to erasure codes. A/B pays 2× bandwidth anyway for three reasons the book's own principles explain: it survives the loss of a whole path, not just of packets; the first copy to arrive wins, so it adds no decode delay; and the redundant copy does not add load to the congested link, which FEC on the same path would (p. 399). See [../multicast/03-failure-modes.md](/multicast/failure-modes/) for how RPF problems can silently remove the B side.
Snapshots are periodic updates
Snapshot or refresh channels resend the current state periodically; a lost update only delays convergence (p. 400). Recovering by snapshot is open loop; asking a retransmission server for the missing range is closed loop and costs at least a round trip.
Recovery is NAK-based, so liveness needs heartbeatsNo single source
Multicast receivers cannot acknowledge (ACK implosion, Fig. 7.18), so recovery is receiver-driven, the book's NAK case. NAKs need a liveness check, which is what feed heartbeats provide: silence is only meaningful if the feed promises to say something regularly. **No single source**.
Late is as bad as lost
"A packet that is significantly out of order is no better than a lost packet" (p. 371), and the book's real-time rule says not to trade timeliness for reliability (T-I.2). A feed handler that waits too long for a gap fill delivers stale prices, the same point as the Arista Primer's buffering remark in [../multicast/02-microburst-buffer.md](/multicast/microburst-buffer/).
Round trips are what remainNo single source
Connection shortening (Fig. 7.6) is why any recovery that needs a round trip is so expensive at today's rates: the data takes nanoseconds to send, a retransmission takes at least an RTT plus detection time. On sparse order-entry TCP sessions there may be no later packets to produce duplicate ACKs, so a lost segment waits for the retransmission timer. **No single source** for the order-entry part.
Do not wait to fill packetsNo single source
Grouping small writes saves per-packet overhead but delays the first message (p. 381); order-entry sockets usually disable Nagle's algorithm (TCP_NODELAY) for this reason. **No single source**.
Correlated bursts break statistical multiplexingNo single source
A market-wide event makes every feed burst at the same moment, which is exactly the correlated case of p. 407. Links that carry several feeds must be sized near the sum of the peaks, not the averages. **No single source**.
Stay left of the kneeNo single source
Queuing delay grows well before any loss (Fig. 7.22), so latency-critical links are run at low average utilization. **No single source**.
Trading's utility function is relativeNo single source
The book's utility curves are functions of absolute latency. A trading decision behaves like hard real-time, a step from full value to almost nothing, but the step sits wherever the fastest competitor is, and it moves. That is why "fast enough" is never a fixed number in trading. **No single source**.
Split the order round trip with the response-time formula
T_r = d_c + 2[(1 + h + c)·b/r + t_p] + d_s maps onto an order: d_c is your decision time, the bracket is the network both ways, d_s is the exchange's processing, which you cannot change. Only d_c and the network terms are yours to optimize. My mapping.
Snapshot channels are a datacycle
A feed that keeps rebroadcasting the full book state over multicast is the book's datacycle (p. 464): a late joiner or a handler recovering from a gap waits at most one cycle instead of making a request round trip, and multicasting the cycle keeps network load flat however many clients there are. My mapping.
Market data is push, done properly
The exchange knows when its state changes, so it pushes on events (A-6B, 4H); consumers never poll. My mapping.
Compression rarely pays on fast linksNo single source
A 100-byte message takes about 80 ns to send at 10 Gb/s, so any encode/decode step longer than the transmission time saved makes delivery slower (p. 465 formula). Compact binary formats with fixed, byte-aligned fields (8B) are the trade-off that usually wins. Arithmetic plus **No single source**.
Ask for the varianceNo single source
Like NTP's variance bounds (p. 482), a timestamp or latency figure is only usable together with its error bound. Hardware timestamping and PTP clocks report this; a single latency number without percentiles or clock accuracy says little. **No single source**.
Do not let interfaces hide latency (A-4Fl)
An aggregated or consolidated data source hides the extra hops and processing between you and the original publisher; know which path your data took. See the Nasdaq SIP case in [../multicast/01-incidents.md](/multicast/incidents/). My mapping.

Reference notes