Ch 4: Network control and signaling
Source: Sterbenz & Touch 2001, Ch 4 (pp. 119–164), read in full. Page numbers are the book’s printed pages. The formulas below were checked on the rendered PDF pages, because text extraction drops the operators.
Numbering note: as in Ch 3, some principle boxes inside Ch 4 differ from Appendix A (for example “Optimal Resource Utilization versus Overengineering” is printed as N-5A on p. 152 but is N-2B in the appendix; “Dynamic Path Rerouting” is N-2B on p. 159 but N-2Br in the appendix). These notes use the Appendix A IDs.
The chapter in one paragraph
Control works on connections, flows and sessions, which last far longer than packets. The recurring question is: how much setup latency do you pay up front, and what does it buy you per packet? Store-and-forward datagrams need no setup but pay the serialization time at every hop. Connection-oriented fast packet switching pays one extra round trip, then cuts through every switch. In-between schemes try to get both. The chapter then covers multicast tree setup, session control, traffic management (reserve vs overprovision, congestion signaling, dropping early to keep queues short), route changes, and why monitoring must be built into the hardware.
Delay formulas (pp. 121–132)
Variables: t_p = x/(k·c) propagation per hop (k ≈ 0.7 for fiber); t_b = b/r time to send one packet; t_f forwarding (lookup or label swap); t_q queuing; t_g gap between packets; t_sig signaling processing per node; h hops; n packets.
| Scheme | Delay | Page |
|---|---|---|
| Store-and-forward datagram, one packet | D₁ = h·t_b + Σt_p + (h − 1)(t_f + t_q) | 128 |
| Store-and-forward, n packets | D_n = (n + h − 1)·t_b + (n − 1)·max[(t_g − t_b), (t_f + t_q)] + Σt_p + (h − 1)(t_f + t_q) | 128 |
| Connection setup | t_setup = 2Σt_p + (2h − 1)·t_sig | 131 |
| Cut-through switch, one packet | d₁ = t_b + Σt_p + (h − 1)·t_s, with t_s = t_f + t_q | 131 |
| Cut-through, n packets, with setup | D_n = n·t_b + (n − 1)·t_g + 3Σt_p + (2h − 1)·t_sig + (h − 1)·t_s | 131 |
The key difference: store-and-forward pays h·t_b, cut-through pays t_b once. Connection setup costs one extra round trip.
What the book draws from the formulas:
- Store-and-forward (p. 128): bigger packets hurt, because each hop stores the whole packet (h·t_b). Long flows dilute the fixed costs. In WANs propagation dominates; in LANs forwarding (and maybe queuing) dominates. If forwarding plus queuing take longer than the packet interarrival time, forwarding dominates and every hop counts.
- Connection-oriented (pp. 131–132): the transfer takes three end-to-end latencies (SETUP, CONNECT, data), one round trip more than datagrams. For a single-packet transaction the signaling dominates; for a long stream it is amortized. A permanent mesh of virtual connections gives datagram service without per-datagram setup.
Circuit switching (pp. 122–124): one round trip to set up, then almost no delay in the switches (no buffering, no store-and-forward), but unused capacity is wasted on bursty traffic.
Message switching (pp. 124–126): no setup, but each node stores the whole message, and large messages delay small ones queued behind them, even in a lightly loaded network. “Large message cross traffic is detrimental to short message traffic” (p. 126). Packet switching bounds packet size to limit this (p. 126).
Signaling (pp. 129–134)
- Each node on the SETUP path decodes the message, routes it, checks resources (admission control), returns a hop-by-hop PROCEEDING and forwards the SETUP. The per-hop acknowledgment lets a lost message be resent after a short per-hop timer instead of an end-to-end one: the effect of a three-way handshake with one end-to-end round trip (pp. 129–130).
- Keep signaling messages simply encoded and small enough for one packet, and make the protocol survive lost messages (N-8B, p. 133).
- Example 4.1, ATM signaling (p. 133): the flow was sensible, but the messages inherited SS7’s complexity and did not fit in a 48-byte cell, so a reliable segmentation layer (S-AAL) was needed. Mid-1990s ATM switches handled “only a few hundred to a thousand connection establishments per second”.
Between datagrams and connections (pp. 134–143)
The spectrum (p. 134): message switching → datagram forwarding → data-driven soft state → control-driven soft state → optimistic connection setup → fast reservation → explicit virtual connections → physical circuits.
- Soft state for datagrams (pp. 135–137): make datagrams switchable with a flow ID (IPv6), labels (MPLS) or a source route (rarely a good trade: the sender must know the topology). The label state is installed either control-driven (signaled together with routing per forwarding equivalence class, independent of traffic) or data-driven (the router spots a flow, then switches it; the first packets are routed normally, Fig. 4.8).
- Overlap control and data (N-6A, pp. 137–138): send data optimistically right after the SETUP (best effort until a COMMIT confirms the reservation), or reserve tentatively at each hop (fast reservation).
- Burst switching for optical networks, whose switches are too slow for per-packet switching (pp. 139–143): tell-and-go (TAG), in-band terminator (IBT), reserve-a-fixed-duration (RFD) and just-enough-time (JET); open-ended reservations need a RELEASE, closed-ended ones carry the burst length. Skim unless you work on optical networks.
Multicast trees and sessions (pp. 143–148)
- A multicast tree needs state even in a datagram network. Finding the optimal tree is NP-complete, so heuristics are used (p. 143).
- Root-initiated join: the root adds leaves (ADD messages flow down). Simple; suits small groups whose membership the root knows; used first for ATM point-to-multipoint; does not scale to large, changing groups (p. 144).
- Leaf-initiated join: the new member sends a JOIN toward the tree. IP multicast is receiver-initiated, “but unfortunately IP multicast uses a distinct address space that does not assist nodes in finding an optimal route to the group” (p. 145). A hierarchy of sub-roots is the middle ground (N-6B).
- Sessions (several flows for one multi-party application) need extra round trips for negotiation. Ways to cut them: set up flows in parallel; start the connection SETUP as soon as the session request passes; send parameter ranges; cache user profiles; and reuse expensive connections such as multicast trees from one session to the next (pp. 146–147). Place session resources (e.g. a transcoder) near the users (p. 148).
Traffic management (pp. 148–157)
- Ways to carry mixed traffic: separate networks, virtual partitions of one network (“probably the worst of all worlds”, p. 149), coarse classes (diffserv), per-flow reservations (intserv, ATM).
- Example 4.2 (p. 150): per-flow RSVP did not scale to backbones; diffserv scales but gives no per-flow guarantee. ATM’s many traffic classes made switches complex; “recognizing that bandwidth is cheap enough to allow some overengineering, and simplifying the traffic management … would probably have sped the deployment”.
- Reservations serve four purposes (pp. 151–152): commit a service level; protect flows from each other (shaping to the contract, policing excess); admission control; and building only the capacity customers pay for. The book doubts bandwidth will ever be cheap enough to skip all of this, so it recommends some overprovisioning plus tractable QoS (N-2B).
- Congestion still happens: guarantees are statistical, best-effort traffic exists, classes are coarse (p. 153). Network help speeds recovery (N-3A).
- Speed of reaction (pp. 154–155): an OC-192 link at 110% load loses about 1 Gb/s. On a 5,000 km WAN, forward notification (mark packets, receiver turns it around) loses over 500 Mb before the source reacts; backward notification straight to the source loses about 250 Mb.
- Multicast flow control: if every receiver ran feedback with the root, the root would drown in messages; use open-loop rate control, with network nodes taking part (p. 155).
- Congestion avoidance (p. 156): when queues start to build, drop early (RED); in cell networks drop whole frames (early or partial packet discard). “Buffers should be kept as empty as possible, with queuing only for transient situations, to allow cut-through and avoid the latency of FIFO queuing” (N-II.4).
- Hop-by-hop control (pp. 156–157): per-hop rate negotiation reacts faster on short hops; credit-based flow control guarantees zero congestion loss through backpressure, but per-connection credit accounting was too complex for ATM switches of the day.
Route changes (pp. 157–160)
- Multicast trees drift away from optimal as members join and leave: prune branches with no listeners and reroute long stubs (Fig. 4.16).
- Rerouting a flow to a shorter path reorders packets: new packets overtake those still on the old path (p. 159).
- Example 4.3: mobile IP’s triangle routing and cellular calls that stay anchored to the first switching center.
Monitoring (pp. 161–162)
- Things change faster than a management interval, let alone a human, so control such as congestion reaction must run inside the nodes in real time.
- Statistics volume: the book says that per-packet records for 40-byte packets on a 10 Gb/s link come to “approximately 3 million records/s per link” and over 3×10⁹/s for a 1K×1K switch (p. 161). That figure looks 10× too low: 10¹⁰ b/s ÷ (40 × 8 b) ≈ 31 million packets/s per link. The conclusion only gets stronger: sample and reduce locally, in hardware.
- Monitoring must be built into the critical path, filtering and aggregating without intruding (N-1B).
Trading-network lens (my mapping)
Read switch latency specs with the formulas. Store-and-forward latency grows with frame size (the h·t_b term); cut-through latency does not. Benchmarks also measure in two ways. RFC 1242 (July 1991, section 3.8) defines latency for a store-and-forward device as last bit in → first bit out (LIFO), and for a bit-forwarding device as first bit in → first bit out (FIFO). It adds that a bridge which starts sending before the frame has fully arrived (a “cut through” device) is still measured last-bit-in to first-bit-out, “even though the value would be negative”. The gap between the two measures is exactly the frame’s serialization time t_b, so always ask which one a datasheet quotes and at which frame size. Checked against https://www.rfc-editor.org/rfc/rfc1242.txt.
Big frames delay small ones (p. 126). A 9,000-byte jumbo frame holds a 10 Gb/s port for 7.2 µs; an order that arrives just behind it waits. Keep latency-critical traffic off links and queues that carry bulk transfers, or give it strict priority. No single source.
Pay setup costs before the open. Order-entry TCP sessions are opened and logged in early, and feed handlers join their multicast groups at startup and stay joined, the book’s advice to reuse expensive connections such as multicast trees (p. 147). The IGMP querier trap in ../multicast/03-failure-modes.md is what happens when that state quietly times out.
Why PIM-SM needs an RP. The book complains that an IP multicast group address says nothing about where the tree is (p. 145). PIM-SM fixes this with a rendezvous point that everyone can find; SSM fixes it by having receivers name the source (S,G) themselves. See ../multicast/00-glossary.md. My mapping.
No closed loop for market data. UDP multicast market data has no congestion control at all, so the only tools are the open-loop ones: engineer enough capacity and keep queues short (N-II.4). The book’s 110% OC-192 example is the microburst case: past 100%, the excess is simply lost. See ../multicast/02-microburst-buffer.md.
Rerouting reorders. When a path changes (routing convergence, a link in a LAG fails), packets of one feed can arrive out of order or with a gap; feed handlers need a reordering window, and A/B arbitration helps fill the gap. No single source.
Monitoring in hardware (N-1B). At tens of millions of packets per second per port, latency and microburst monitoring has to be done by hardware timestamping and capture, not by polling counters. See the BurstRadar notes in ../multicast/02-microburst-buffer.md.
Self-check
- Write D₁ for store-and-forward and d₁ for cut-through. With h = 3 and a 1500-byte frame at 10 Gb/s, how much serialization does each pay (ignore t_p, t_f, t_q)?
- How many end-to-end latencies does a connection-oriented transfer need before the data arrives, and why?
- Why does message switching hurt short messages even on a lightly loaded network?
- Root-initiated vs leaf-initiated join: which does IP multicast use, and what does the book say is missing from IP multicast addresses?
- Forward vs backward congestion notification: which reacts faster, and why?
- Why does the book want queues kept nearly empty even when buffer space is available?
- Check the book’s monitoring arithmetic: how many 40-byte packets per second fit in 10 Gb/s?
Answers
- D₁ = h·t_b + Σt_p + (h − 1)(t_f + t_q); d₁ = t_b + Σt_p + (h − 1)·t_s. t_b = 1500 × 8 / 10¹⁰ = 1.2 µs, so store-and-forward pays 3.6 µs and cut-through 1.2 µs.
- Three: SETUP out, CONNECT back, then the data out (p. 131), one round trip more than a datagram.
- Each node stores whole messages, so a short message queued behind a long one waits for the long one’s full transmission at every hop (p. 126).
- Leaf-initiated (receiver-initiated). The group address lives in a separate address space that does not help a joining node find a route to the tree (p. 145).
- Backward: the congested node signals the source directly, about one end-to-end latency, with no turnaround at the receiver; forward notification needs a full round trip plus the receiver’s reaction (pp. 154–155).
- A packet in a queue cannot cut through and waits behind everything ahead of it; queues should absorb only transients (N-II.4, p. 156).
- 10¹⁰ ÷ 320 ≈ 31 million per second, about 10× the book’s figure.