Multicast market data in practice (exchange feeds, A/B arbitration, recovery, low latency)
Merged on 2026-10-03 from the Mac session’s research note 2026-10-03-multicast-quant/02-multicast-in-quant-industry.md (AI-written, Chinese), translated and condensed.
Verification status. ✅ = checked against the source named. ❌ / ⚠️ = the Mac note was wrong or imprecise and is corrected here (also in the group verification log). Everything else is Unverified: exchange specs change, so check the current spec before relying on a field layout. SGX and HKEX feeds are covered, checked, in ../exchanges/sg-hk-exchanges.md.
Why exchanges send market data as UDP multicast
- Fan-out costs the sender nothing extra: the matching engine sends one packet and switches copy it in hardware. With TCP the sender keeps a socket, buffer and retransmission state per subscriber, so cost grows with N.
- Fairness: every colocated member should see the event at the same moment. Writing to N TCP sockets in turn creates a queue order; hardware replication does not.
- No head-of-line blocking: with TCP, one slow receiver pushes back on the sender. With UDP a slow receiver only drops its own packets and recovers out of band.
- Small overhead: an 8-byte UDP header, no handshake, no ACKs. Message bodies are fixed-layout binary (SBE, ITCH, MITCH, EOBI) that can be parsed in place.
| UDP multicast | TCP (replay, SoupBinTCP) | REST / WebSocket | |
|---|---|---|---|
| Fan-out cost | O(1), hardware copies | O(N), sender writes each | O(N) plus HTTP |
| Arrival order across receivers | Same for everyone | Staggered per subscriber | Worse |
| Slow consumer | Drops its own packets | Back-pressures the sender | Same as TCP |
| Reliability | None; needs gap recovery | In the transport, paid for in latency | In the transport |
The design core: reliability and ordering move out of the transport into the application, as sequence numbers plus recovery channels. WebSocket and REST appear only at the far end (retail apps, data vendors), never in the exchange’s core distribution.
Exchange feeds
CME Globex MDP 3.0
- Messages in SBE (Simple Binary Encoding, FIX-based schema). Products are split into channels. Each channel has its own multicast groups for the incremental feed (
35=X), the snapshot / market recovery feed (35=W) and the instrument definition feed (35=d), each sent on feed A and B. The channel list gives the IPs and ports. - Packet = a 12-byte binary packet header (uint32 packet sequence number per channel + uint64 sending time in ns), then SBE messages, each with the standard 8-byte SBE header (block length, template ID, schema ID, version). CME also puts a 2-byte message size before each message (No single source; check the spec).
- MDP 3.0 is market data; iLink 3 is order entry (FIXP/SBE over TCP). There is no “iLink 300” market data protocol. You receive MDP 3.0 by joining the channel’s multicast groups, not through an order session.
- Recovery ⚠️: CME still runs TCP replay/recovery: a 2026 CME notice changes its IP and port for channels 326, 327 and 329. But CME’s client wiki describes it as not a performance-based recovery option, for small-scale recovery, used only if the other options are unavailable, with at most 2000 packets per request. The main paths are the other feed (A/B) and the market recovery snapshot loops. (Seen in search-result summaries of CME’s wiki page “MDP 3.0 - TCP Recovery”; the page did not render for the fetcher. The Mac note presented TCP replay as the normal small-gap path.)
- Connectivity requirements (PIM-SM, RP, BGP) are in 03 §3.
Nasdaq TotalView-ITCH 5.0 over MoldUDP64
- ITCH 5.0 binary messages (system event, stock directory, add order, executed, cancel, delete, replace, trade, cross …) inside MoldUDP64 packets.
- MoldUDP64 header, 20 bytes ✅: Session (10), Sequence Number (8, the number of the first message in the packet), Message Count (2).
- ❌ A message count of 0 is a heartbeat; 0xFFFF is end of session ✅ MoldUDP64 spec. The Mac note had 0xFFFF as the heartbeat. Heartbeats are typically sent once a second and carry the next expected sequence number, so receivers notice loss even when the market is quiet ✅ (same spec).
- Small gaps: send a request packet to the re-request server for a sequence range ✅ (the spec has a “Request Packet” section). Large gaps or a late start: take a GLIMPSE snapshot over TCP, then follow the live feed from the snapshot’s sequence number.
- The SGX note confirms the same pattern from SGX’s own MoldUDP64 spec: heartbeat = count 0, retransmissions are unicast answers to the requesting socket (sg-hk-exchanges.md).
- ITCH 5.0 spec: https://www.nasdaqtrader.com/content/technicalsupport/specifications/dataproducts/NQTVITCHSpecification.pdf
Other venues (all Unverified unless marked)
- NYSE Pillar Integrated Feed: order-by-order, UDP multicast, line A and line B. https://www.nyse.com/market-data/pillar
- Eurex T7 EOBI: order-by-order, multicast services A and B. The Mac note says Eurex alternates which service sends first by partition (odd/even). Not found in Eurex’s public pages, so Unverified.
- LSE Millennium MITCH: market by order, channels A and B. ⚠️ MITCH has its own 8-byte Unit Header, not MoldUDP64 framing as the Mac note said: Length (2), Message Count (1), Market Data Group (1), Sequence Number (4) ✅, checked in NSE Kenya’s MITCH UDP spec v1.22 §7.6, another Millennium venue. LSE’s own spec was not checked.
- SGX and HKEX: see sg-hk-exchanges.md (MoldUDP64 + Rewinder + Glimpse at SGX; OMD-C lines A/B, retransmission and refresh services at HKEX).
The A/B pattern
- Two independent paths carrying the same messages with the same sequence numbers. Packet boundaries may differ: HKEX says the lines can package messages differently, so arbitrate on message sequence numbers, not packets (sg-hk-exchanges.md).
- Take the first copy of each sequence number from either line and drop the later copy. If losses on the two lines are independent, loss probability falls from p to about p².
- Arbitrate, do not fail over: listen to both lines with equal priority. HKEX calls “use line B only when line A has a gap” incorrect.
- A gap exists only when a sequence number is missing from both lines (after a short reorder window). Only then start recovery.
Receiver architecture: fast path multicast, slow path recovery
┌─ Feed A (multicast) ─> [line handler A] ─┐
Exchange ───┤ ├─> [arbiter] ─> [order book] ─> strategy
└─ Feed B (multicast) ─> [line handler B] ─┘ │
│ missing on both lines
v
[recovery client] ─unicast/TCP─> re-request · replay · snapshot
- Line handler (one per line): receive (often via kernel bypass), validate the header, extract (session, sequence). Busy-poll on a pinned core with no allocation on the hot path.
- Arbiter: track the next expected sequence number; pass the first copy downstream, drop duplicates; tolerate small reordering between lines.
- Gap detection: when sequence N is followed by N+k (k > 1) on the merged stream, record the gap and keep processing later messages. Never stall the live stream waiting for a retransmission.
- Recovery: small gaps go to re-request or replay. Large gaps, late joins and restarts go to a snapshot first, then you follow the live stream from the snapshot’s sequence. Merging recovered messages back in order is the part most often written wrong.
- Isolation: recovery uses separate connections, bandwidth and rate limits, so it can never slow the live feed. Exchanges build their retransmission services the same way.
Checklist (Unverified, engineering practice from the Mac note): pinned threads per line with a lock-free SPSC queue into the arbiter; process each (session, seq) once; a gap timer that escalates to snapshot rebuild; heartbeats also confirm the sequence position; after a reconnect, take a snapshot before trusting the stream; timestamp at every stage (NIC hardware timestamp, then line handler, then book) for later replay and audit.
Build it yourself: the 2027 prep plan portfolio project is a Python A/B publisher and receiver with gap recovery. The book view of the same mechanisms is in low-latency/07.
Open-source feed handlers to read (repositories exist as of 2026-10-03; code quality not reviewed): penberg/helix (C++ ITCH/MoldUDP64), harris2001/UltraLowLatencyFeedHandler, Ashutosh0x/rust-finance, Lunyn-HFT/lunary. ❌ The Mac note’s cmegroup/CME-Application-Samples returns 404.
Reliable multicast and middleware
- Exchange feeds use NACK-style recovery out of band (re-request, replay over TCP), not NACKs on the multicast tree, which avoids NACK storms.
- Application-level FEC is uncommon on exchange feeds: A/B arbitration already makes loss rare (the Mac note’s finding; Unverified). FEC shows up in messaging middleware (29West LBT-RM optionally), in NORM (RFC 5740), and at the Ethernet PHY.
- ⚠️ PHY FEC correction: 25G NRZ lanes (including 100G built as 4×25G) use RS(528,514) “KR4”; RS(544,514) “KP4” is for 50G PAM4 lanes. IEEE 802.3 task-force material puts them at about 169 ns and 198 ns of added latency (seen in search results, slides not opened). The Mac note gave RS(544,514) for 25G/100G.
| Middleware | What it is | Notes (Unverified unless marked) |
|---|---|---|
| 29West LBM / Informatica Ultra Messaging | Commercial low-latency messaging over UDP multicast (LBT-RM, NAK-based, optional FEC) | ❌ Informatica completed the acquisition in March 2010, not 2011 (DBTA and other news reports, via search) |
| UMDS | UM’s “last mile” to desktops over TCP/WebSocket | |
| OpenMAMA | Middleware-neutral market data API (FINOS) | Bridges to several transports |
| Aeron | Open-source UDP unicast/multicast and shared-memory IPC, with Archive for record/replay | |
| Solace PubSub+ | Broker-based fan-out (not IP multicast) with replay | |
| TIBCO Rendezvous | Older UDP-multicast messaging | Largely replaced |
Low latency: network and hardware
- Colocation: cross-connects inside the exchange data centre; fibre adds about 5 ns per metre (No single source, physics: light in glass travels at roughly 2/3 c).
- Joining groups: a dynamic join (IGMP report, snooping, PIM joins hop by hop) is control-plane work and takes milliseconds. Common practice:
- join every needed group at process start and stay joined all day;
- use static groups on the switch for the known feeds;
- use fast-leave or immediate-leave on host ports;
- let the application, not the network, decide which data to read.
- Kernel bypass (latency figures in the Mac note were unsourced and are left out):
- OpenOnload: LD_PRELOAD socket acceleration with no code change, Xilinx-CNS/onload.
- ef_vi: direct NIC queue access.
- DPDK: polling drivers, so you build your own UDP/IP handling.
- AF_XDP: kernel-assisted fast path.
- ❌ ExaNIC is now Cisco Nexus SmartNIC; the Nexus 3550 series is the former ExaLINK Fusion/Hydra switches ✅ (Cisco’s Exablaze acquisition pages, via search). The Mac note had ExaNIC as “now the Nexus 3550”.
- Switch side: IGMP snooping always on; an explicit querier in pure L2; PIM-SM between VLANs and subnets, with static or Anycast RP to avoid RP election churn; PIM snooping where several PIM routers share a VLAN. Troubleshooting commands are in 03. ⚠️ The Mac note had an NX-OS config snippet. It mixes IOS-style and NX-OS-style syntax and was not checked, so it is left out; take lab configs from Cisco’s NX-OS multicast guide.
- FPGA: parse MDP 3.0 or ITCH on the NIC, filter symbols, keep top of book in hardware, and hardware-timestamp at start of frame. A common split is FPGA pre-filtering with CPU strategy logic.
Internal distribution
[exchange links, A/B multicast, kernel-bypass NIC]
└─ edge ticker plant (normalize to an internal format, one timestamp base)
├─ internal multicast (LBM / Aeron / own) → latency-critical strategy hosts
├─ shared-memory SPSC ring → processes on the same host
├─ RDMA (RoCE) → high-throughput consumers
└─ TCP / WebSocket → risk, research, desktops
Why some firms shrink internal multicast: IGMP/PIM/snooping state is hard to debug; join/leave or RP storms can hurt a whole segment; receivers still pay per-packet costs, which shared memory avoids; explicit receiver lists (TCP/RDMA) are easier to control than IGMP membership; RDMA matured. The Mac note cites Jane Street’s Signals & Threads, “Multicast and the Markets” (Brian Nigito; not re-listened). ❌ The other cited episode, “Ep.15 Reliable Multicast with Doug Patti”, does not match: Doug Patti’s episode is “State Machine Replication, and Why You Should Care”.
Timestamps and clocks
- Why: latency attribution (exchange event → we received → we sent), backtests, audit.
- Where to timestamp: kernel (
SO_TIMESTAMPNS, µs and jittery), NIC hardware (SO_TIMESTAMPING, check withethtool -T, ns), FPGA at start of frame. The earlier the stamp, the cleaner the data. - PTP (IEEE 1588): a GNSS-fed grandmaster, switches as boundary or transparent clocks, and
linuxptp(ptp4l+phc2sys) disciplining NIC and system clocks. - ❌ MiFID II RTS 25 for the high-frequency algorithmic trading technique: maximum divergence from UTC 100 µs, timestamp granularity 1 µs or finer (Meinberg’s MiFID II FAQ and other summaries via search; EUR-Lex blocked automated fetching). The Mac note said “98 µs, 100 µs granularity”. Regulation: Commission Delegated Regulation (EU) 2017/574, https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX%3A32017R0574.
Interview angles
Topics the Mac note collected (they can seed the planned question bank in the prep plan):
- Why do exchanges use UDP multicast and not TCP? Answer with: fan-out cost, fairness, no head-of-line blocking, overhead, and reliability through sequence numbers plus recovery.
- How do you detect loss? (Sequence gaps on the merged A/B stream.)
- How do you arbitrate A/B? (First copy wins, reorder window, no failover.)
- How do you merge recovered data with the live stream? (Bounded buffering, escalate to a snapshot.)
- “Joined but no data”, why? Check:
- no querier, or snooping aged out;
- wrong NIC joined;
- TTL 1;
- RPF failure;
- RP or PIM config;
- static group missing.
- How do you cut receive latency? (Kernel bypass, pinning and isolcpus, busy polling, zero-copy parsing, SPSC queues, NUMA, hardware timestamps.)
- ARQ vs FEC, and why exchanges keep NACKs out of band.
- Clocks: NIC/FPGA timestamps, PTP vs NTP, RTS 25’s 100 µs.
TL;DR
- Multicast solves “the same event to everyone at the same moment”: O(1) fan-out, fairness, and slow consumers that hurt only themselves.
- Reliability lives in the application: A/B arbitration on the fast path, re-request/replay and snapshots on the slow path.
- Feeds differ in detail (MoldUDP64 64-bit message sequence, CME 32-bit packet sequence, MITCH’s own unit header) but share the pattern.
- Receiver latency: kernel bypass, pinned zero-copy hot path, hardware timestamps with PTP.