Principles map (Sterbenz & Touch, Appendix A, in my words)
Every axiom and principle in the book, paraphrased in one line, with the page where the book states it. Source: Appendix A (pp. 555–574), checked against Chapter 2. Use the IDs when citing the book in other notes, for example “S-II.3 (p. 201)”.
Prefixes: none = general (Ch 2), N network (Ch 3–4), L link (Ch 5), S switch (Ch 5), E end system (Ch 6), T transport (Ch 7), A application (Ch 8).
The five axioms
Ø Know the past, present and future
| ID | In my words | Page |
|---|---|---|
| Ø1 | Genuinely new ideas are rare. Most “new” ideas have a history whose lessons you can learn or ignore. | 14 |
| Ø2 | An old idea looks different in a new context. The difference tells you which old lessons still apply. | 14 |
| Ø3 | The future will contain at least one surprise that changes everything. | 15, 489 |
| Ø4 | Knowing the past is not enough: keep re-checking trade-offs and assumptions as technology changes. | 20, 491 |
| Ø-A, Ø-B | “Not invented here”: operating systems and networking did not begin with Unix and TCP/IP, nor with Windows and the PC. | 17, 18 |
I Application primacy
| ID | In my words | Page |
|---|---|---|
| I | The only reason to build a high-performance network is the distributed applications that need it. | 20 |
| E-I | Host-side communication optimizations must not slow down the applications themselves. | 286 |
| I.1 | Chicken and egg: the next killer app needs infrastructure first, and infrastructure is hard to justify without the app. | 21 |
| I.2 | The metric is interapplication delay: the total delay to move data between applications. For users it includes the time spent inside the application. | 21 |
| T-I.2 | Real-time media: do not trade timeliness for reliability; use a playout buffer to reorder and absorb jitter. | 371 |
| A-I.2i | Interactive response: 100 ms is the ideal, 1 s the upper target, and a consistent response beats a variable one. | 441 |
| I.3 | Bandwidth and latency are the two network metrics that make up interapplication delay. | 22 |
| A-I.3 | Latency needs come from user expectations; for data-heavy applications, delay sensitivity is what drives bandwidth needs. | 434 |
| S-I.3 | Packet rate (pps) is a key switch throughput measure. Shared resources must sustain the average pps; anything in the serial critical path must handle the worst-case pps, or packets queue behind it. | 254 |
| I.4, E-I.4 | Networking must be a first-class part of system design, on a par with memory or graphics in a host. | 22, 294 |
II High-performance paths goal
| ID | In my words | Page |
|---|---|---|
| II | Network and hosts together must give the applications a low-latency, high-bandwidth path. | 24 |
| N-II, L-II, L-IIc, S-II, E-II | The same goal per component: network paths, links (link processing must not add significant latency), link components (sustain line rate), nodes, and the host path between NIC and application memory. | 80, 167, 188, 195, 285 |
| N-IIo | Overlays must keep the performance of the physical path; keep the number of overlay layers small. | 92 |
| II.1, N-II.1 | Path establishment: routing and signaling must find and set up good paths fast enough for the application. | 24, 134 |
| II.2 | Path protection: when resources are scarce, reserve and arbitrate them so other traffic cannot congest your path. | 25 |
| N-II.2, L-II.2, S-II.2, E-II.2 | The same per component: QoS with bandwidth and latency bounds, fair MAC arbitration, switch admission control and policing, reserved CPU and memory in the host. | 151, 182, 215, 308 |
| A-II.2b | Without a policy saying otherwise, share resources fairly among best-effort applications. | 439 |
| II.3 | Store-and-forward avoidance: store-and-forward and copying cost so much latency that they should be avoided wherever possible. | 25 |
| S-II.3 | Switches: avoid store-and-forward, minimize per-packet queuing; ideally pipeline and cut through with zero per-packet delay. | 201 |
| E-II.3, E-II.3m | Hosts: avoid copies and any extra pass over the bytes; ideally zero copy; remap memory instead of copying. | 295, 306 |
| II.4 | Blocking avoidance: avoid blocking from paths that collide and from queues that build. | 26 |
| N-II.4 | Avoid congestion (engineering, reservations, dropping). Keep buffers nearly empty and queue only for transients, so packets can cut through. | 156 |
| S-II.4f | A nonblocking fabric is the core of a fast switch: space-division parallelism, internal speedup, pipelined buffering with cut-through. | 227 |
| S-II.4q | Avoid head-of-line blocking: output queuing needs internal speedup; virtual output queues need more queues; combine as needed. | 232 |
| S-II.4c | Classify packets before any input queuing, with a classification time bounded by the strictest service class. | 269 |
| S-II.4s | Output scheduling must run at line rate. Finer-grained queues isolate flows better but cost complexity. | 273 |
| S-II.4a | Extra (active) processing must not slow the fast path for ordinary packets. | 277 |
| E-II.4, E-II.4m | The host’s memory–NIC interconnect must be nonblocking and not interfere with CPU–memory traffic or other I/O. | 318, 325 |
| T-II.4c | Congestion avoidance: act before queues build; run just left of the knee of the throughput curve. | 411 |
| II.5, N-II.5 | Avoid contention on shared media; meshes scale better than shared-medium links. | 26, 82 |
| II.6 | Control on which the critical path depends must be cheap; avoid costly hand-offs between protocol modules. | 26 |
| E-II.6c | Minimize context switches: aim for about one per application data unit. | 302 |
| E-II.6k | Minimize user/kernel crossings: each one costs checks, buffer copies and indirection. | 305 |
| II.7 | Reliability and security requirements may have to be traded against performance. | 26 |
III Limiting constraints
| ID | In my words | Page |
|---|---|---|
| III | Real-world constraints make high-performance paths hard to provide. | 27 |
| III.1 | Speed of light: propagation delay is physics and cannot be optimized directly. | 27 |
| III.2 | Channel capacity is limited by physics; multiplexing and spatial reuse help but do not remove the limit. | 27 |
| III.3 | Component switching speed is limited by process technology and, finally, physics. | 28 |
| III.4, S-III.4 | Cost and scaling complexity decide what gets deployed; build large switches from regular, recursive fabrics whose cost grows logarithmically. | 28, 241 |
| III.5, A-III.5 | The network is heterogeneous: applications, hosts, nodes and links all differ. | 29, 453 |
| III.6, N-III | Policy and administration prevent the optimal topology and constrain paths, which makes principled design more important, not less. | 29, 108 |
| III.7 | Backward compatibility blocks radical change; improvements are incremental, and hacks become institutions. | 30 |
| E-III.7, A-III.7 | So optimize and extend the protocols already deployed, and give legacy applications a sensible default behavior. | 291, 484 |
| III.8 | Standards help interoperability, but too early or too specific blocks progress, and too late or too vague is useless. | 30 |
IV Systemic optimization
| ID | In my words | Page |
|---|---|---|
| IV | A network is a system of systems; analyze and optimize the parts together. | 31 |
| E-IV | In a host, the organization, CPU–memory interconnect, memory, OS, protocol stack and NIC must be optimized together. | 289 |
| T-IV | Transport protocols must deliver the services applications need, with options that do not degrade each other. | 355 |
| IV1, E-IV1, A-IV1 | Optimizations have side effects; analyze them before you trust the gain. | 32, 298, 456 |
| IV2 | Keep it simple and open: complex systems are hard to optimize, closed ones nearly impossible. | 32 |
| IV3, A-IV3 | Decide carefully where functions live; bad partitioning can ruin performance, and some “high-speed” applications are just badly partitioned slow ones. | 33, 437 |
| IV4 | Leave fields, hooks and options so the design can adapt when trade-offs change. | 34 |
The eight design principles (all under IV)
1 Selective optimization
| ID | In my words | Page |
|---|---|---|
| 1 | You cannot optimize everything; spend effort and money on the biggest contributors. | 34 |
| 1A | Second-order effect: do not optimize a part that contributes only about 10% or less. | 34 |
| N-1Al, A-1Al | Latency along a path is the sum of its parts; fixing one part helps only in proportion to its share. User-perceived delay is the final measure. | 83, 456 |
| N-1Ah | The number of per-hop delays is bounded by the network diameter; keep the diameter small. | 87 |
| N-1Ab, A-1Ab | Path bandwidth is the minimum along the path (the bottleneck); upgrading anything else is wasted. | 89, 470 |
| 1B | Critical path: optimize what the data flow and its control actually depend on. | 35 |
| N-1B | Build monitoring into the critical path so it filters and aggregates without intruding. | 161 |
| S-1Bc, S-1Bd | Simplify per-byte and per-packet work for hardware; apply fast-packet-switching techniques to datagram (IP) switching. | 203, 250 |
| E-1B, T-1B | The same for host protocol processing, and for per-packet encryption and authentication. | 297, 426 |
| 1C | Decide carefully what goes into scarce or expensive technology (hardware vs software, cache vs memory). | 36 |
| S-1C, S-1Ch | In switches: hardware vs embedded software for input and output processing. | 223, 267 |
| E-1Ci, E-1Ch | In hosts: what goes on the NIC rather than in host software, and what goes into NIC hardware rather than its embedded controller. Packet interarrival time decides it. | 329, 331 |
| A-1C | Not every function belongs in the application; keep the network impact of design choices in mind. | 476 |
2 Resource trade-offs
| ID | In my words | Page |
|---|---|---|
| 2, N-2 | A network is bandwidth, processing and memory (B, P, M), with latency as the constraint; balance them for cost and performance. | 37, 110 |
| A-2 | Bandwidth can buy latency: prefetch and presend. | 462 |
| 2A, N-2A | The relative costs change over time, so the right design changes too. | 37, 115 |
| A-2A | “High speed” is relative; demand has always grown to the limit of what links can do. | 434 |
| 2B, N-2B, S-2B | Weigh optimal utilization (and its complex algorithms) against simply overprovisioning. | 38, 152, 222 |
| N-2Bc, N-2Br | Fairness vs complexity in congestion control; rerouting overhead vs path optimality. | 156, 159 |
| A-2B | Compression saves bandwidth but costs processing and delay. | 466 |
| 2C | Support multicast: it saves link bandwidth and sender work, and control protocols depend on it. | 39 |
| L-2C, S-2C | Links should support broadcast and multicast; switches should replicate natively, to save bandwidth and avoid the delay of sending copies one after another. | 194, 246 |
3 End-to-end arguments
| ID | In my words | Page |
|---|---|---|
| 3, T-3 | A function can be done correctly and completely only with the endpoints’ knowledge; the network alone cannot provide it (Saltzer, Reed, Clark). | 39, 346 |
| 3A, T-3A, N-3A | Duplicating an end-to-end function hop by hop is worth it if it improves end-to-end performance (e.g. link retransmission on lossy links, network help with congestion control). | 40, 347, 153 |
| 3B | What is hop-by-hop in one context is end-to-end in another; the argument applies recursively. | 41 |
4 Protocol layering
| ID | In my words | Page |
|---|---|---|
| 4 | Layering is a useful way to think about and organize protocols. | 42 |
| L-4f | Early filtering: drop traffic not meant for this node as early as possible, at every layer. | 194 |
| 4A, E-4A | Layered architecture does not mean a layered implementation; a process per layer performs badly. | 45, 294 |
| T-4A | Avoid multiplexing at several layers; demultiplex once for all layers. | 384 |
| 4B | Do not put a function in a layer if a higher layer must repeat it anyway, unless it clearly helps. | 45 |
| 4C, E-4C | Layers’ formats and control mechanisms must not interfere; PDUs should translate cheaply between layers. | 46, 329 |
| T-4C, A-4C | Application Layer Framing (ALF): match transport units to application data units, so the application can deal with loss and reordering itself. | 383, 476 |
| A-4Ck, A-4Cd, A-4Cv | Knobs must map to real actions, dials to real observable properties; give a variance with any approximate value. | 480–482 |
| 4D | Hourglass: the network layer (IP) is the common point; addressing must be common. | 46 |
| 4E, E-4D | Integrated Layer Processing (ILP): when one component handles several layers, do all the passes over the data at once. | 47, 311 |
| 4F, A-4Fh | Abstraction must not hide what upper layers need to know. | 47, 483 |
| L-4F | Tell higher layers why loss happened, so they react correctly. | 193 |
| N-4F, A-4Ff, A-4Fl | Let sessions see network parameters; adaptive applications need path feedback; location-independent interfaces must still expose latency. | 147, 455, 483 |
| 4G | Offer several interface styles: synchronous and asynchronous, interrupt-driven and polled. | 48 |
| 4H, E-4H, A-4H | Interrupts vs polling: interrupts handle asynchronous events but are expensive; poll when you know when data will arrive. | 48, 304, 481 |
| 4I, A-4I | Interfaces must scale: values need an exponent, a precision and a variance. | 48, 478 |
5 State management
| ID | In my words | Page |
|---|---|---|
| 5 | Balance fast, approximate, coarse state against slow, accurate, fine state. | 50 |
| 5A | Hard state (set up by signaling, removed explicitly) vs soft state (refreshed, times out) vs stateless (decide per packet): setup latency against per-packet cost. | 51 |
| N-5A, S-5A | Connection setup costs latency but can make forwarding faster (label swapping, no store-and-forward). | 132, 205 |
| T-5A | Hard state is deterministic and stable; soft state adapts and survives failures. | 374 |
| 5B, T-5B, A-5B | Aggregate state to cut memory and update traffic, at the cost of granularity. “State shared is fate shared.” | 51, 374, 475 |
| 5C, N-5C, N-5Cb, N-5Cl, N-5Cw, A-5C | Use hierarchy and clustering to manage scale, aggregate bandwidth, keep the diameter (and latency) low, and place caches. | 53, 100–106, 460 |
| 5D | Scope of information: decide quickly on local information; global state is stale by the time you collect it. | 53 |
| A-5D | Quick partial information beats slow complete information, especially if refined over time. | 442 |
| 5E, T-5E | Start new state from past history so it converges quickly. | 53, 378 |
| 5F | Minimize control overhead: control traffic is a tax on data capacity. | 54 |
| T-5F | Know the path MTU; avoid fragmentation. | 382 |
6 Control mechanism latency
| ID | In my words | Page |
|---|---|---|
| 6 | Control needs current information and must converge as fast as the network changes. | 54 |
| 6A | Minimize round trips (hop-by-hop acknowledgments, parameter ranges, overlapping control with data). | 55 |
| N-6A, T-6A | Overlap signaling with data; combine connection setup with the request and the data. | 137, 370 |
| 6B | Control should be exercised by whoever has the knowledge. | 57 |
| N-6B | Root-controlled multicast when one entity knows the group; leaf-controlled (receivers join) for large dynamic groups. | 145 |
| A-6B, A-6Bu | Client pull vs server push; let users steer adaptive applications. | 451, 455 |
| 6C | Anticipate future state; act before damage needs repair (congestion avoidance). | 59 |
| S-6Cc, S-6Cd | Bound offered load (admission, policing, shaping); throttle sources as queues build; drop to keep queues from building in steady state. | 221, 274 |
| 6D, T-6D | Open-loop vs closed-loop control: use known path properties to act without feedback; use feedback to react to change. | 59, 359 |
| T-6Da, T-6Dc, T-6Do | Closed-loop reaction fast but not oscillating; closed loop is needed without reservations; use open loop where you know the path. | 359, 411, 403 |
| T-6De | Forward error correction: low-latency, open-loop error control when bandwidth is available and some statistical loss is acceptable. | 399 |
| 6E, T-6E | Keep control mechanisms separate and their signals explicit (loss is not always congestion); decouple error, flow and congestion control. | 60, 395 |
7 Distributed data
| ID | In my words | Page |
|---|---|---|
| 7 | Choose and organize exchanged data to minimize volume and latency, and to allow incremental processing. | 60 |
| 7A, A-7A, A-7Ac | Partition and structure data to cut transfer size and latency; show something useful early. | 61, 442, 461 |
| 7B, A-7Bs, A-7Bt | Put data close to the application (replication, caching, spatial and temporal locality). | 61, 457, 459 |
8 Protocol data units
| ID | In my words | Page |
|---|---|---|
| 8 | PDU size and structure are critical for bandwidth and latency. | 61 |
| 8A, S-8A, T-8A | Small packets multiplex well but cut the time per packet; large packets are efficient but delay others. | 62, 209, 380 |
| 8B, N-8B, S-8B | Header fields: simple encoding, byte aligned, fixed length; variable fields carry a length first. | 62, 133, 215 |
| 8C, L-8C, T-8C | Fields tied to rate or bandwidth–delay product (sequence numbers, timers, windows) must scale or carry a scale factor. | 63, 176, 391 |
Where the principles show up in trading networks (my mapping)
The book predates electronic trading’s low-latency arms race and never mentions it. This section is my own mapping, No single source unless linked; chapter notes go deeper.
- I.2 Interapplication delay is what trading calls tick-to-trade: from a market data packet arriving to the order leaving, application time included.
- 1A / N-1Al (latency is a sum): in a colocation the fiber is metres long, so per-hop switch latency, the host stack and the application dominate. Measure each term before optimizing one.
- II.3 / S-II.3 (cut through): cut-through switches, and layer-1 switches for fan-out, exist to remove the store-and-forward term.
- II.3 / E-II.3, II.6 / E-II.6c–k, 4H (copies, context switches, kernel crossings, polling): this is the case for kernel-bypass NIC stacks and busy-polling cores in trading hosts.
- 2C / S-2C (native multicast): market data is distributed by multicast, and switches replicate it in hardware.
- 6D / T-6De (open-loop error control): A/B feed arbitration is open-loop redundancy (every packet sent twice on separate paths); retransmission and snapshot recovery are the closed-loop fallback. See ../multicast/03-failure-modes.md.
- N-II.4 / S-6Cd (keep queues short): buffering is latency; see ../multicast/02-microburst-buffer.md.
- 7B / 2.4.1 (move data and applications closer): colocation is this principle taken to the limit.