Ch 8–9: Networked applications, future directions
Source: Sterbenz & Touch 2001, Ch 8 (pp. 431–488) and Ch 9 (pp. 489–500), read in full. Page numbers are the book’s printed pages. Both chapters are marked “skim” in the book guide; the parts worth keeping for trading networks are the latency utility classes, the response-time formula, latency masking (especially the datacycle), and the reminder that resource trade-offs keep changing.
Numbering note: the chapter prints the variance principle as “A-5Av” (p. 482); Appendix A lists it as A-4Cv.
Ch 8: what makes an application “high speed” (pp. 431–447)
- The user sees the whole delay, including the application’s own (A-I.2). For live media, quality (frame rate, resolution) is the metric instead (A-I.2c).
- Bandwidth classes (pp. 434–437): individual vs aggregate bandwidth (one 20 Mb/s HDTV stream is moderate, 75 million of them is about 1.5 Pb/s); and how an application scales: bandwidth enhanced (works at any rate, better when faster, e.g. streaming), challenged (needs restructuring at higher rates, e.g. distributed computing), enabled (only works above a threshold, e.g. telemedicine imaging).
- “Some ‘high-speed’ applications are simply poorly partitioned low-speed applications” (A-IV3, p. 437): where you cut an application decides how much it must send.
- Latency utility functions (Figs. 8.2–8.3, pp. 438–444):
| Class | Utility as latency grows | Example |
|---|---|---|
| Best effort | Slowly decreasing; high variance tolerated | Email, printing |
| Interactive | High up to about 100 ms, falls toward 1 s; consistency matters | Web browsing |
| Real-time (hard) | Step from 1 to 0 at the bound T_rt | Process control, telemetry |
| Real-time (soft) | Degrades but survives occasional late or lost data | Teleconferencing (audio delay over 30 ms is heard as echo, over 500 ms breaks conversation) |
| Deadline | Real-time with a much longer bound; a minimum rate b/T_dl suffices | Remote backup |
- Interactive response time (p. 440): T_r = d_c + 2[(1 + h + c)·b/r + t_p] + d_s, the client’s processing, the network round trip and the server’s processing. Users prefer consistent response to a fast one with spikes: occasional 800 ms responses among 100 ms ones frustrate more than a steady 500 ms (pp. 441–442).
- Quick partial information beats slow complete information, refined over time (A-5D); structure data so the first useful piece arrives within the bound (A-7A).
- Table 8.1 (p. 446) gives order-of-magnitude budgets: distributed computing and process control 1 µs–10 s and loss-intolerant; live voice 30 ms; Web 100 ms–1 s; email minutes to an hour.
- Application categories (pp. 447–454): information access (client/server, asymmetric, response time matters), push (the server knows when data changes), telepresence (peer to peer, real-time sync), distributed computing (arbitrary exchange). Footnote 10 (p. 451) notes that many “server push” services of the late 1990s, including financial information, were really browser timers polling.
Ch 8: adapting to latency and bandwidth (pp. 454–478)
- Latency is the sum of everything that cannot be parallelized (A-1Al, p. 456), the application included.
- Mirroring and caching (pp. 457–462): put copies close to users (spatial locality, A-7Bs), keep what will be reused (temporal locality, A-7Bt), and build cache hierarchies that balance bandwidth against storage (A-5C). Ways to reach the best copy: manual choice, proxy, anycast, server redirection, active redirection.
- Anticipation trades bandwidth for latency (A-2, pp. 462–464): prefetching (fetching three levels of 10 links each turns a 10 Mb/s page into 1.11 Gb/s of traffic), application prediction of remote state, server presending (the server pushes on events instead of clients polling, which is principle 4H again).
- Datacycle (p. 464, Fig. 8.12): broadcast the whole data set over and over. If the cycle time T_c is short compared with a request round trip, a client waits less by “sipping” from the stream than by asking. Multicasting the cycle can even lower backbone load; erasure (Tornado) coding turns it into a “digital fountain”.
- Compression (pp. 465–466): with compression ratio q and encode/decode times t_e, t_d, the one-way delay becomes D = (d_c + t_e) + [(1 + h + c)·bq/r + t_p] + (t_d + d_s). It reduces delay only when t_e + t_d is smaller than the transmission time saved (A-2B).
- Pipelining at the application layer only helps if stages really run in parallel (p. 466).
- Bandwidth techniques (pp. 470–474): caching, anticipation, compression, multicast (application-layer multicast when the network will not do it), layered coding with the network dropping fine layers first, ADU size, striping (needs skew below one stripe unit).
- Aggregate operations to reduce per-request costs (A-5B, p. 475); use ALF (A-4C, p. 476).
Ch 8: knobs and dials (pp. 478–484)
- Applications need knobs to ask the network for behavior and dials to read its state (Fig. 8.17). Knobs must map to real actions (A-4Ck), dials to observable properties (A-4Cd); offer both polled and asynchronous dials (A-4H).
- Values should carry a scale (exponent), a precision and a variance (A-4I, A-4Cv): NTP gained variance bounds on its time reports, and TCP’s RTT estimate carries a variance (p. 482).
- Location-independent interfaces must not hide latency (A-4Fl): a URL says nothing about how far away or how big the page is; distributed shared memory that hides a 100 ms access cannot give predictable performance (pp. 483–484).
- Legacy (p. 484): Unix send() copies the buffer so the call can return synchronously; an asynchronous interface that releases buffers later would avoid that copy.
Ch 9: future directions (pp. 489–500)
- The only certain thing about the future is surprise (Ø3). Ten years before the book, nobody expected universal consumer Internet, the Web or ubiquitous laptops and phones (pp. 489–491).
- Resource trade-offs keep shifting (pp. 491–493): first bandwidth was scarce (protocols saved bits at the cost of host processing), then fiber made it cheap and the bottleneck moved to routers and hosts, then the Web turned traffic into short transactions (“the majority of traffic was on the order of 10 packets, of which nearly half are connection setup and teardown”, p. 492), then streaming came back. Design protocols that can adapt; that is hardest at the IP waist of the hourglass.
- If one resource became effectively free (p. 493), caching would still be needed for latency, and prefetching would still be limited by memory: one resource is only as useful as the others allow.
- Wrong assumptions are expensive (pp. 494–495): video-on-demand trials failed for lack of infrastructure, not demand; providers built asymmetric access and were surprised when Napster turned consumers into content providers. Questioning the traffic patterns of the day would have been enough to prepare.
- Technologies to watch in 2001: all-optical networking, mobile wireless, active networking; application scenarios: ubiquitous computing, teleimmersion, distributed computing, the interplanetary Internet (pp. 495–499).
- Conclusion (pp. 499–500): technologies change, the design principles do not.
Trading-network lens (my mapping)
Trading’s utility function is relative. The book’s utility curves are functions of absolute latency. A trading decision behaves like hard real-time, a step from full value to almost nothing, but the step sits wherever the fastest competitor is, and it moves. That is why “fast enough” is never a fixed number in trading. No single source.
Split the order round trip with the response-time formula. T_r = d_c + 2[(1 + h + c)·b/r + t_p] + d_s maps onto an order: d_c is your decision time, the bracket is the network both ways, d_s is the exchange’s processing, which you cannot change. Only d_c and the network terms are yours to optimize. My mapping.
Snapshot channels are a datacycle. A feed that keeps rebroadcasting the full book state over multicast is the book’s datacycle (p. 464): a late joiner or a handler recovering from a gap waits at most one cycle instead of making a request round trip, and multicasting the cycle keeps network load flat however many clients there are. My mapping.
Market data is push, done properly. The exchange knows when its state changes, so it pushes on events (A-6B, 4H); consumers never poll. My mapping.
Compression rarely pays on fast links. A 100-byte message takes about 80 ns to send at 10 Gb/s, so any encode/decode step longer than the transmission time saved makes delivery slower (p. 465 formula). Compact binary formats with fixed, byte-aligned fields (8B) are the trade-off that usually wins. Arithmetic plus No single source.
Ask for the variance. Like NTP’s variance bounds (p. 482), a timestamp or latency figure is only usable together with its error bound. Hardware timestamping and PTP clocks report this; a single latency number without percentiles or clock accuracy says little. No single source.
Do not let interfaces hide latency (A-4Fl). An aggregated or consolidated data source hides the extra hops and processing between you and the original publisher; know which path your data took. See the Nasdaq SIP case in ../multicast/01-incidents.md. My mapping.
Self-check
- Name the book’s latency utility classes. Which is trading closest to, and how does it differ?
- Write the response-time formula. Which terms can a trading firm reduce?
- What is a datacycle, when does it beat request/response, and what is its market data equivalent?
- When does compression reduce end-to-end delay?
- Why should a location-independent interface still expose latency?
- What practical advice does Ch 9 give for preparing for a future you cannot predict?
Answers
- Best effort, interactive, real-time (hard and soft), deadline. Trading is closest to hard real-time (a step in utility), but the step is set relative to competitors and moves.
- T_r = d_c + 2[(1 + h + c)·b/r + t_p] + d_s. The firm controls its own processing d_c and the network terms (hops, copies, rate, path length), not the exchange’s processing d_s.
- Repeatedly broadcasting the whole data set; it wins when the cycle time is short compared with a request round trip (p. 464). Snapshot or refresh channels in market data feeds.
- When the encode and decode time is less than the transmission time saved (p. 465).
- Because applications that could adapt to latency need to know it; hiding it makes performance unpredictable (A-4Fl, pp. 483–484).
- Keep re-checking the resource trade-offs and question current traffic assumptions; design protocols and systems that can adapt (Ø4, 2A, pp. 491–495).