← Low-latency networking

Principles map (Sterbenz & Touch, Appendix A, in my words)

Every axiom and principle in the book, paraphrased in one line, with the page where the book states it. Source: Appendix A (pp. 555–574), checked against Chapter 2. Use the IDs when citing the book in other notes, for example “S-II.3 (p. 201)”.

Prefixes: none = general (Ch 2), N network (Ch 3–4), L link (Ch 5), S switch (Ch 5), E end system (Ch 6), T transport (Ch 7), A application (Ch 8).

The five axioms

Ø Know the past, present and future

ID In my words Page
Ø1 Genuinely new ideas are rare. Most “new” ideas have a history whose lessons you can learn or ignore. 14
Ø2 An old idea looks different in a new context. The difference tells you which old lessons still apply. 14
Ø3 The future will contain at least one surprise that changes everything. 15, 489
Ø4 Knowing the past is not enough: keep re-checking trade-offs and assumptions as technology changes. 20, 491
Ø-A, Ø-B “Not invented here”: operating systems and networking did not begin with Unix and TCP/IP, nor with Windows and the PC. 17, 18

I Application primacy

ID In my words Page
I The only reason to build a high-performance network is the distributed applications that need it. 20
E-I Host-side communication optimizations must not slow down the applications themselves. 286
I.1 Chicken and egg: the next killer app needs infrastructure first, and infrastructure is hard to justify without the app. 21
I.2 The metric is interapplication delay: the total delay to move data between applications. For users it includes the time spent inside the application. 21
T-I.2 Real-time media: do not trade timeliness for reliability; use a playout buffer to reorder and absorb jitter. 371
A-I.2i Interactive response: 100 ms is the ideal, 1 s the upper target, and a consistent response beats a variable one. 441
I.3 Bandwidth and latency are the two network metrics that make up interapplication delay. 22
A-I.3 Latency needs come from user expectations; for data-heavy applications, delay sensitivity is what drives bandwidth needs. 434
S-I.3 Packet rate (pps) is a key switch throughput measure. Shared resources must sustain the average pps; anything in the serial critical path must handle the worst-case pps, or packets queue behind it. 254
I.4, E-I.4 Networking must be a first-class part of system design, on a par with memory or graphics in a host. 22, 294

II High-performance paths goal

ID In my words Page
II Network and hosts together must give the applications a low-latency, high-bandwidth path. 24
N-II, L-II, L-IIc, S-II, E-II The same goal per component: network paths, links (link processing must not add significant latency), link components (sustain line rate), nodes, and the host path between NIC and application memory. 80, 167, 188, 195, 285
N-IIo Overlays must keep the performance of the physical path; keep the number of overlay layers small. 92
II.1, N-II.1 Path establishment: routing and signaling must find and set up good paths fast enough for the application. 24, 134
II.2 Path protection: when resources are scarce, reserve and arbitrate them so other traffic cannot congest your path. 25
N-II.2, L-II.2, S-II.2, E-II.2 The same per component: QoS with bandwidth and latency bounds, fair MAC arbitration, switch admission control and policing, reserved CPU and memory in the host. 151, 182, 215, 308
A-II.2b Without a policy saying otherwise, share resources fairly among best-effort applications. 439
II.3 Store-and-forward avoidance: store-and-forward and copying cost so much latency that they should be avoided wherever possible. 25
S-II.3 Switches: avoid store-and-forward, minimize per-packet queuing; ideally pipeline and cut through with zero per-packet delay. 201
E-II.3, E-II.3m Hosts: avoid copies and any extra pass over the bytes; ideally zero copy; remap memory instead of copying. 295, 306
II.4 Blocking avoidance: avoid blocking from paths that collide and from queues that build. 26
N-II.4 Avoid congestion (engineering, reservations, dropping). Keep buffers nearly empty and queue only for transients, so packets can cut through. 156
S-II.4f A nonblocking fabric is the core of a fast switch: space-division parallelism, internal speedup, pipelined buffering with cut-through. 227
S-II.4q Avoid head-of-line blocking: output queuing needs internal speedup; virtual output queues need more queues; combine as needed. 232
S-II.4c Classify packets before any input queuing, with a classification time bounded by the strictest service class. 269
S-II.4s Output scheduling must run at line rate. Finer-grained queues isolate flows better but cost complexity. 273
S-II.4a Extra (active) processing must not slow the fast path for ordinary packets. 277
E-II.4, E-II.4m The host’s memory–NIC interconnect must be nonblocking and not interfere with CPU–memory traffic or other I/O. 318, 325
T-II.4c Congestion avoidance: act before queues build; run just left of the knee of the throughput curve. 411
II.5, N-II.5 Avoid contention on shared media; meshes scale better than shared-medium links. 26, 82
II.6 Control on which the critical path depends must be cheap; avoid costly hand-offs between protocol modules. 26
E-II.6c Minimize context switches: aim for about one per application data unit. 302
E-II.6k Minimize user/kernel crossings: each one costs checks, buffer copies and indirection. 305
II.7 Reliability and security requirements may have to be traded against performance. 26

III Limiting constraints

ID In my words Page
III Real-world constraints make high-performance paths hard to provide. 27
III.1 Speed of light: propagation delay is physics and cannot be optimized directly. 27
III.2 Channel capacity is limited by physics; multiplexing and spatial reuse help but do not remove the limit. 27
III.3 Component switching speed is limited by process technology and, finally, physics. 28
III.4, S-III.4 Cost and scaling complexity decide what gets deployed; build large switches from regular, recursive fabrics whose cost grows logarithmically. 28, 241
III.5, A-III.5 The network is heterogeneous: applications, hosts, nodes and links all differ. 29, 453
III.6, N-III Policy and administration prevent the optimal topology and constrain paths, which makes principled design more important, not less. 29, 108
III.7 Backward compatibility blocks radical change; improvements are incremental, and hacks become institutions. 30
E-III.7, A-III.7 So optimize and extend the protocols already deployed, and give legacy applications a sensible default behavior. 291, 484
III.8 Standards help interoperability, but too early or too specific blocks progress, and too late or too vague is useless. 30

IV Systemic optimization

ID In my words Page
IV A network is a system of systems; analyze and optimize the parts together. 31
E-IV In a host, the organization, CPU–memory interconnect, memory, OS, protocol stack and NIC must be optimized together. 289
T-IV Transport protocols must deliver the services applications need, with options that do not degrade each other. 355
IV1, E-IV1, A-IV1 Optimizations have side effects; analyze them before you trust the gain. 32, 298, 456
IV2 Keep it simple and open: complex systems are hard to optimize, closed ones nearly impossible. 32
IV3, A-IV3 Decide carefully where functions live; bad partitioning can ruin performance, and some “high-speed” applications are just badly partitioned slow ones. 33, 437
IV4 Leave fields, hooks and options so the design can adapt when trade-offs change. 34

The eight design principles (all under IV)

1 Selective optimization

ID In my words Page
1 You cannot optimize everything; spend effort and money on the biggest contributors. 34
1A Second-order effect: do not optimize a part that contributes only about 10% or less. 34
N-1Al, A-1Al Latency along a path is the sum of its parts; fixing one part helps only in proportion to its share. User-perceived delay is the final measure. 83, 456
N-1Ah The number of per-hop delays is bounded by the network diameter; keep the diameter small. 87
N-1Ab, A-1Ab Path bandwidth is the minimum along the path (the bottleneck); upgrading anything else is wasted. 89, 470
1B Critical path: optimize what the data flow and its control actually depend on. 35
N-1B Build monitoring into the critical path so it filters and aggregates without intruding. 161
S-1Bc, S-1Bd Simplify per-byte and per-packet work for hardware; apply fast-packet-switching techniques to datagram (IP) switching. 203, 250
E-1B, T-1B The same for host protocol processing, and for per-packet encryption and authentication. 297, 426
1C Decide carefully what goes into scarce or expensive technology (hardware vs software, cache vs memory). 36
S-1C, S-1Ch In switches: hardware vs embedded software for input and output processing. 223, 267
E-1Ci, E-1Ch In hosts: what goes on the NIC rather than in host software, and what goes into NIC hardware rather than its embedded controller. Packet interarrival time decides it. 329, 331
A-1C Not every function belongs in the application; keep the network impact of design choices in mind. 476

2 Resource trade-offs

ID In my words Page
2, N-2 A network is bandwidth, processing and memory (B, P, M), with latency as the constraint; balance them for cost and performance. 37, 110
A-2 Bandwidth can buy latency: prefetch and presend. 462
2A, N-2A The relative costs change over time, so the right design changes too. 37, 115
A-2A “High speed” is relative; demand has always grown to the limit of what links can do. 434
2B, N-2B, S-2B Weigh optimal utilization (and its complex algorithms) against simply overprovisioning. 38, 152, 222
N-2Bc, N-2Br Fairness vs complexity in congestion control; rerouting overhead vs path optimality. 156, 159
A-2B Compression saves bandwidth but costs processing and delay. 466
2C Support multicast: it saves link bandwidth and sender work, and control protocols depend on it. 39
L-2C, S-2C Links should support broadcast and multicast; switches should replicate natively, to save bandwidth and avoid the delay of sending copies one after another. 194, 246

3 End-to-end arguments

ID In my words Page
3, T-3 A function can be done correctly and completely only with the endpoints’ knowledge; the network alone cannot provide it (Saltzer, Reed, Clark). 39, 346
3A, T-3A, N-3A Duplicating an end-to-end function hop by hop is worth it if it improves end-to-end performance (e.g. link retransmission on lossy links, network help with congestion control). 40, 347, 153
3B What is hop-by-hop in one context is end-to-end in another; the argument applies recursively. 41

4 Protocol layering

ID In my words Page
4 Layering is a useful way to think about and organize protocols. 42
L-4f Early filtering: drop traffic not meant for this node as early as possible, at every layer. 194
4A, E-4A Layered architecture does not mean a layered implementation; a process per layer performs badly. 45, 294
T-4A Avoid multiplexing at several layers; demultiplex once for all layers. 384
4B Do not put a function in a layer if a higher layer must repeat it anyway, unless it clearly helps. 45
4C, E-4C Layers’ formats and control mechanisms must not interfere; PDUs should translate cheaply between layers. 46, 329
T-4C, A-4C Application Layer Framing (ALF): match transport units to application data units, so the application can deal with loss and reordering itself. 383, 476
A-4Ck, A-4Cd, A-4Cv Knobs must map to real actions, dials to real observable properties; give a variance with any approximate value. 480–482
4D Hourglass: the network layer (IP) is the common point; addressing must be common. 46
4E, E-4D Integrated Layer Processing (ILP): when one component handles several layers, do all the passes over the data at once. 47, 311
4F, A-4Fh Abstraction must not hide what upper layers need to know. 47, 483
L-4F Tell higher layers why loss happened, so they react correctly. 193
N-4F, A-4Ff, A-4Fl Let sessions see network parameters; adaptive applications need path feedback; location-independent interfaces must still expose latency. 147, 455, 483
4G Offer several interface styles: synchronous and asynchronous, interrupt-driven and polled. 48
4H, E-4H, A-4H Interrupts vs polling: interrupts handle asynchronous events but are expensive; poll when you know when data will arrive. 48, 304, 481
4I, A-4I Interfaces must scale: values need an exponent, a precision and a variance. 48, 478

5 State management

ID In my words Page
5 Balance fast, approximate, coarse state against slow, accurate, fine state. 50
5A Hard state (set up by signaling, removed explicitly) vs soft state (refreshed, times out) vs stateless (decide per packet): setup latency against per-packet cost. 51
N-5A, S-5A Connection setup costs latency but can make forwarding faster (label swapping, no store-and-forward). 132, 205
T-5A Hard state is deterministic and stable; soft state adapts and survives failures. 374
5B, T-5B, A-5B Aggregate state to cut memory and update traffic, at the cost of granularity. “State shared is fate shared.” 51, 374, 475
5C, N-5C, N-5Cb, N-5Cl, N-5Cw, A-5C Use hierarchy and clustering to manage scale, aggregate bandwidth, keep the diameter (and latency) low, and place caches. 53, 100–106, 460
5D Scope of information: decide quickly on local information; global state is stale by the time you collect it. 53
A-5D Quick partial information beats slow complete information, especially if refined over time. 442
5E, T-5E Start new state from past history so it converges quickly. 53, 378
5F Minimize control overhead: control traffic is a tax on data capacity. 54
T-5F Know the path MTU; avoid fragmentation. 382

6 Control mechanism latency

ID In my words Page
6 Control needs current information and must converge as fast as the network changes. 54
6A Minimize round trips (hop-by-hop acknowledgments, parameter ranges, overlapping control with data). 55
N-6A, T-6A Overlap signaling with data; combine connection setup with the request and the data. 137, 370
6B Control should be exercised by whoever has the knowledge. 57
N-6B Root-controlled multicast when one entity knows the group; leaf-controlled (receivers join) for large dynamic groups. 145
A-6B, A-6Bu Client pull vs server push; let users steer adaptive applications. 451, 455
6C Anticipate future state; act before damage needs repair (congestion avoidance). 59
S-6Cc, S-6Cd Bound offered load (admission, policing, shaping); throttle sources as queues build; drop to keep queues from building in steady state. 221, 274
6D, T-6D Open-loop vs closed-loop control: use known path properties to act without feedback; use feedback to react to change. 59, 359
T-6Da, T-6Dc, T-6Do Closed-loop reaction fast but not oscillating; closed loop is needed without reservations; use open loop where you know the path. 359, 411, 403
T-6De Forward error correction: low-latency, open-loop error control when bandwidth is available and some statistical loss is acceptable. 399
6E, T-6E Keep control mechanisms separate and their signals explicit (loss is not always congestion); decouple error, flow and congestion control. 60, 395

7 Distributed data

ID In my words Page
7 Choose and organize exchanged data to minimize volume and latency, and to allow incremental processing. 60
7A, A-7A, A-7Ac Partition and structure data to cut transfer size and latency; show something useful early. 61, 442, 461
7B, A-7Bs, A-7Bt Put data close to the application (replication, caching, spatial and temporal locality). 61, 457, 459

8 Protocol data units

ID In my words Page
8 PDU size and structure are critical for bandwidth and latency. 61
8A, S-8A, T-8A Small packets multiplex well but cut the time per packet; large packets are efficient but delay others. 62, 209, 380
8B, N-8B, S-8B Header fields: simple encoding, byte aligned, fixed length; variable fields carry a length first. 62, 133, 215
8C, L-8C, T-8C Fields tied to rate or bandwidth–delay product (sequence numbers, timers, windows) must scale or carry a scale factor. 63, 176, 391

Where the principles show up in trading networks (my mapping)

The book predates electronic trading’s low-latency arms race and never mentions it. This section is my own mapping, No single source unless linked; chapter notes go deeper.

Source: knowledge base note low-latency/01-principles-map.md — own-words notes with sources, projected at build time.