← Low-latency networking

Ch 5: Network components (links and switches)

Source: Sterbenz & Touch 2001, Ch 5 (pp. 165–284), read in full. Page numbers are the book’s printed pages. This is the chapter that matters most for trading networks: how a switch can forward without store-and-forward, where buffers sit, what head-of-line blocking costs, and why packet rate matters more than bandwidth.

Numbering note: the packet-rate principle is printed as “S-II.4p” on p. 254 but listed as S-I.3 in Appendix A. These notes use the Appendix A IDs.

The chapter in one paragraph

Links: propagation speed depends on the medium; link protocols should scale with rate; the Ethernet story shows a protocol surviving by turning into a switched point-to-point mesh. Switches: old routers were slow because every packet crossed a shared bus into a shared CPU and memory and was stored before forwarding. Fast packet switches fixed this with per-connection state, hardware pipelines and cut-through. The fabric is where blocking happens: buffer placement, head-of-line blocking, virtual output queues, crossbars and multistage networks, and how to replicate multicast. IP switches then had to do the same at line rate without connection state, which turns lookup, classification and scheduling into the hard problems.

Propagation speed (pp. 167–169): v = c/n. Fiber glass has n ≈ 1.45–1.48, so light travels at about 0.68c. The fastest media are vacuum (1.0c), air and some coax (very near 1.0c); the slowest are fiber and twisted pair (about two-thirds of c). The propagation speed has nothing to do with a link’s “speed”, which is its symbol rate (p. 167).

Medium (Table 5.1, p. 169) Velocity Delay per km
Twisted pair 0.67c ≈ 5 µs
Coax 0.66–0.95c ≈ 4 µs
Optical fiber 0.68c ≈ 5 µs
Wireless (radio, infrared, visible) 1.0c ≈ 3.3 µs

What a switch does (pp. 195–201)

Granularity Functions (Fig. 5.12)
Per packet, data manipulation Input processing, switch fabric, output processing, packet buffers, link decapsulation and framing
Per packet, transfer control Filtering and classification, congestion control (policing, shaping, discard), forwarding table, output scheduling
Per flow or longer Management, signaling, topology and link state, routing, traffic management (reservation, admission control)

Why traditional routers were slow (pp. 199–201): a computer with network interfaces on a bus. Every packet paid the store-and-forward time t_b; packets competed for one CPU and one memory, which raised the forwarding delay t_f; and every packet crossed the bus twice, which capped the number of interfaces. Moving processing onto the interfaces with bus-master transfers (NSFNET routers, mid-1990s) removed the CPU bottleneck but kept a store-and-forward hop on each interface and the single bus.

The ideal switch (p. 201): R = ∞, D = 0 and unlimited ports. “In the ideal case, nodes should pipeline and cut through packets with zero per packet delays” (S-II.3).

Fast packet switches (pp. 202–225)

Rate 32 B 128 B 1 KB
1 Gb/s 250 ns 1 µs 8 µs
10 Gb/s 25 ns 100 ns 800 ns

Switch fabrics (pp. 226–249)

Fast datagram (IP) switches (pp. 249–274)

Trading-network lens (my mapping)

The medium is a latency choice. From the book’s Table 5.1: fiber about 5 µs/km, radio about 3.3 µs/km, so a straight radio path saves about 1.7 ms per 1,000 km one way. Inside a building, fiber costs about 5 ns per metre, so cable lengths matter only at the nanosecond level. Arithmetic from the book’s table.

Layer-1 switches are the book’s unbuffered crossbar. Arista’s 7130 Connect series forwards port to port in 4 ns with “full signal recovery and regeneration”, does “not buffer or queue data”, and allows one-to-many connections with the same latency (vendor product page, checked: https://www.arista.com/en/products/7130-connect). That is the book’s electronic crosspoint with its free duplicate state (p. 239), used for market data fan-out. The price of having no buffer: two inputs cannot be merged onto one output at layer 1; merging needs a device with buffering and arbitration.

Burst collision is the microburst. Fig. 5.25 is the fan-in problem described by Arista and Pico: feeds that each fit the egress link on average still collide in time, and the switch must queue (latency) or drop (gaps). See ../multicast/02-microburst-buffer.md.

Packet rate, not bandwidth. Market data messages are small, and bursts arrive as packet-rate spikes. At 10 Gb/s a minimum 64-byte frame plus 20 bytes of preamble and gap is 67 ns, or 14.88 million frames per second. Whatever is serial in a NIC, kernel or feed handler must keep up with the worst case (S-I.3); spreading flows over cores, like the book’s parallel forwarding engines, raises throughput but must hash per flow so each feed stays in order. Arithmetic is No single source; the core-spreading comparison is my mapping.

Worst case, not average (p. 256). The book’s warning about an occasional slow lookup backing up the queue applies to any per-message step on the hot path, such as an order-book update that sometimes allocates memory.

Multicast replication has its own scheduling. When a switch replicates a market data packet to many egress ports, a busy output delays only its own copy if the switch splits the fanout, or holds everything if it does not. Latency can therefore differ between subscriber ports of the same switch. My mapping; check the specific switch’s documentation.

Classify at ingress (S-II.4c). QoS marking and ACLs must act at line rate before queuing, so a strict-priority queue for orders actually protects them. My mapping.

FPGAs (p. 224). The book already notes FPGAs as a fast, reprogrammable middle ground for packet processing; trading later put feed handling and order entry on FPGA NICs and switches. No single source.

Protection switching vs redundant feeds. SONET protection takes up to 50 ms and IP reconvergence takes longer, while redundant A/B feeds lose nothing during a single failure. See the JPX and MOEX cases in ../multicast/01-incidents.md. My mapping.

Self-check

  1. From Table 5.1, what is the propagation delay per km in fiber and in air, and what does a straight radio path save over 1,000 km?
  2. What does the input queue of a cut-through fast packet switch need to be, according to the book?
  3. Why does the book say trailers are essential for cut-through?
  4. What throughput limit does head-of-line blocking impose, and what are the ways around it?
  5. What is the processing budget for a 128-byte packet at 10 Gb/s? For a minimum Ethernet frame including preamble and gap?
  6. Why must serial pipeline stages be sized for minimum-size packets, while parallel engines can use the average? What does the parallel approach cost?
  7. Two feeds each average 4 Gb/s but burst at line rate into one 10 Gb/s egress port. What does Fig. 5.25 predict?
  8. What does the book get wrong about SONET protection switching?
Answers
  1. Fiber about 5 µs/km, air about 3.3 µs/km; about 1.7 ms one way over 1,000 km.
  2. A per-byte shift register just long enough to cover the label lookup (p. 205).
  3. Values computed over the payload (CRC) can be appended or checked as the data streams by; otherwise the whole packet has to be held while it is processed (pp. 214–215).
  4. 2 − √2 ≈ 58.6%. Output queuing via speedup, internal buffering or internal expansion (Clos), or virtual output queues with a matching scheduler.
  5. 100 ns (Table 5.4). (64 + 20) × 8 = 672 bits, so 67.2 ns.
  6. A serial stage that is slow for one small packet delays every packet behind it; parallel engines can average out, but they reorder packets and add jitter (pp. 253–254).
  7. The bursts will overlap even though 8 Gb/s fits on average; the switch has to buffer (adding latency) or drop.
  8. It prints 50 µs; the GR-253 requirement is 50 ms.

Source: knowledge base note low-latency/05-ch05-links-and-switches.md — own-words notes with sources, projected at build time.