← Low-latency networking

Ch 7: End-to-end protocols

Source: Sterbenz & Touch 2001, Ch 7 (pp. 343–429), read in full. Page numbers are the book’s printed pages. This chapter gives the vocabulary for how market data and order traffic are protected: open-loop vs closed-loop control, ARQ vs forward error correction vs plain repetition, and why stale data is no better than lost data.

The chapter in one paragraph

The transport layer turns hop-by-hop services into an end-to-end path: framing, multiplexing, connection state, error control, flow control and congestion control. At high data rates, data transfer shrinks but round trips do not, so every mechanism that waits for feedback (setup, acknowledgment, retransmission) dominates the delay. The book’s answer: use open-loop control where you know the path (rate control, FEC, redundancy), keep closed-loop control for what really needs it, decouple error control from flow and congestion control, and let the application frame its own data when it can handle loss and reordering better than the transport.

The end-to-end arguments, used properly (pp. 345–352)

Mechanisms and service models (pp. 352–357)

Open-loop vs closed-loop control (pp. 357–359)

What high speed does to state (pp. 360–364)

Transfer modes (pp. 364–373)

Keeping state cheap (pp. 373–378)

Framing and multiplexing (pp. 378–386)

Error control (pp. 386–400)

Errors (pp. 386–390): bit errors are rare on fiber (around 10⁻¹²); on wired networks most loss is congestion, often as bursts. Misordering comes from multiple paths, striping, multi-path fabrics, retransmissions and reroutes; “in high-speed networks it is better to avoid the latency of retransmission with receiver reordering” (p. 388). Losing one fragment means resending the whole packet, which can drive congestion collapse (p. 389).

Sequence numbers (p. 391): the space must not wrap while old packets can still be alive: 2ⁿ > (t_rexmit + 2·t_MPL + t_ACK)·r_pkt [Watson 1981]. Fields tied to the bandwidth–delay product need room or a scale factor (T-8C).

Closed-loop retransmission (ARQ) (pp. 391–397):

Open-loop error control (pp. 398–400) fits when all of these hold: the application tolerates some statistical loss; it has (near) real-time latency needs that a retransmission round trip would break; the path is long or has a high bandwidth–delay product; and the receivers can absorb the extra data and decode it.

Flow and congestion control (pp. 400–422)

Security (pp. 422–427)

Trading-network lens (my mapping)

Market data is the datagram case. UDP multicast feeds have no connection and no closed loop at all (p. 365). Reliability moves into the application: the feed’s own sequence numbers, gap detection and recovery logic are Application Layer Framing (T-4C) in practice.

A/B feeds are the book’s spatial redundancy. Sending every packet twice over separate networks is repetition, which the book calls “rarely useful” next to erasure codes. A/B pays 2× bandwidth anyway for three reasons the book’s own principles explain: it survives the loss of a whole path, not just of packets; the first copy to arrive wins, so it adds no decode delay; and the redundant copy does not add load to the congested link, which FEC on the same path would (p. 399). See ../multicast/03-failure-modes.md for how RPF problems can silently remove the B side.

Snapshots are periodic updates. Snapshot or refresh channels resend the current state periodically; a lost update only delays convergence (p. 400). Recovering by snapshot is open loop; asking a retransmission server for the missing range is closed loop and costs at least a round trip.

Recovery is NAK-based, so liveness needs heartbeats. Multicast receivers cannot acknowledge (ACK implosion, Fig. 7.18), so recovery is receiver-driven, the book’s NAK case. NAKs need a liveness check, which is what feed heartbeats provide: silence is only meaningful if the feed promises to say something regularly. No single source.

Late is as bad as lost. “A packet that is significantly out of order is no better than a lost packet” (p. 371), and the book’s real-time rule says not to trade timeliness for reliability (T-I.2). A feed handler that waits too long for a gap fill delivers stale prices, the same point as the Arista Primer’s buffering remark in ../multicast/02-microburst-buffer.md.

Round trips are what remain. Connection shortening (Fig. 7.6) is why any recovery that needs a round trip is so expensive at today’s rates: the data takes nanoseconds to send, a retransmission takes at least an RTT plus detection time. On sparse order-entry TCP sessions there may be no later packets to produce duplicate ACKs, so a lost segment waits for the retransmission timer. No single source for the order-entry part.

Do not wait to fill packets. Grouping small writes saves per-packet overhead but delays the first message (p. 381); order-entry sockets usually disable Nagle’s algorithm (TCP_NODELAY) for this reason. No single source.

Correlated bursts break statistical multiplexing. A market-wide event makes every feed burst at the same moment, which is exactly the correlated case of p. 407. Links that carry several feeds must be sized near the sum of the peaks, not the averages. No single source.

Stay left of the knee. Queuing delay grows well before any loss (Fig. 7.22), so latency-critical links are run at low average utilization. No single source.

Self-check

  1. What are the two common misreadings of the end-to-end arguments, and what is the right question to ask instead?
  2. Why does reliable delivery approach 1.5 RTT at very high data rates, and what does one retransmission do?
  3. What control is possible for a datagram transport such as UDP, and where should retransmission logic live for periodically synchronized state?
  4. List the four conditions under which the book recommends open-loop error control. Does market data meet them?
  5. Why does the book say pure repetition is rarely useful, and why do A/B feeds use it anyway?
  6. Why can’t the receivers of a multicast stream simply acknowledge the source, and what do NAK-based schemes need instead?
  7. What does grouping small messages into one packet trade, and why do order-entry sockets avoid it?
  8. Where does the bandwidth a bursty flow needs lie, and what makes it rise toward the peak?
Answers
  1. E2E-Only and Everything-E2E. Ask whether a hop-by-hop copy of the function improves end-to-end performance (T-3A).
  2. Sending time shrinks to zero while the handshake, data and final ACK still need about 1.5 RTT; one retransmission adds another round trip, about 2.5 RTT (pp. 360–362).
  3. Only framing, multiplexing and open-loop measures; no closed-loop control. Retransmission belongs in the application’s state machines (p. 365).
  4. Loss tolerance expressed statistically; real-time latency needs; long or high-BDP paths; receivers that can absorb the redundancy (p. 398). Market data fits the latency and redundancy conditions; its loss tolerance comes from recovery mechanisms behind the open-loop layer.
  5. An erasure code gives the same protection for less overhead (1.5× vs 2×), and back-to-back repeats do not survive bursts (p. 400). A/B uses separate paths, so it survives path failures, adds no decode latency, and does not load the congested link.
  6. ACK implosion: messages from every receiver converge on the links near the root (Fig. 7.18). NAKs need a liveness mechanism, since silence is ambiguous (p. 393).
  7. Less per-packet overhead against the delay of waiting to fill the packet (p. 381); orders cannot wait.
  8. Between the average and the peak rate (r_a < R < r_p); correlated bursts push it toward the peak (p. 407).

Source: knowledge base note low-latency/07-ch07-transport.md — own-words notes with sources, projected at build time.