← Multicast & market data

Multicast market data in practice (exchange feeds, A/B arbitration, recovery, low latency)

Merged on 2026-10-03 from the Mac session’s research note 2026-10-03-multicast-quant/02-multicast-in-quant-industry.md (AI-written, Chinese), translated and condensed.

Verification status. ✅ = checked against the source named. ❌ / ⚠️ = the Mac note was wrong or imprecise and is corrected here (also in the group verification log). Everything else is Unverified: exchange specs change, so check the current spec before relying on a field layout. SGX and HKEX feeds are covered, checked, in ../exchanges/sg-hk-exchanges.md.

Why exchanges send market data as UDP multicast

  1. Fan-out costs the sender nothing extra: the matching engine sends one packet and switches copy it in hardware. With TCP the sender keeps a socket, buffer and retransmission state per subscriber, so cost grows with N.
  2. Fairness: every colocated member should see the event at the same moment. Writing to N TCP sockets in turn creates a queue order; hardware replication does not.
  3. No head-of-line blocking: with TCP, one slow receiver pushes back on the sender. With UDP a slow receiver only drops its own packets and recovers out of band.
  4. Small overhead: an 8-byte UDP header, no handshake, no ACKs. Message bodies are fixed-layout binary (SBE, ITCH, MITCH, EOBI) that can be parsed in place.
UDP multicast TCP (replay, SoupBinTCP) REST / WebSocket
Fan-out cost O(1), hardware copies O(N), sender writes each O(N) plus HTTP
Arrival order across receivers Same for everyone Staggered per subscriber Worse
Slow consumer Drops its own packets Back-pressures the sender Same as TCP
Reliability None; needs gap recovery In the transport, paid for in latency In the transport

The design core: reliability and ordering move out of the transport into the application, as sequence numbers plus recovery channels. WebSocket and REST appear only at the far end (retail apps, data vendors), never in the exchange’s core distribution.

Exchange feeds

CME Globex MDP 3.0

Nasdaq TotalView-ITCH 5.0 over MoldUDP64

Other venues (all Unverified unless marked)

The A/B pattern

Receiver architecture: fast path multicast, slow path recovery

            ┌─ Feed A (multicast) ─> [line handler A] ─┐
Exchange ───┤                                          ├─> [arbiter] ─> [order book] ─> strategy
            └─ Feed B (multicast) ─> [line handler B] ─┘        │
                                                                │ missing on both lines
                                                                v
                                   [recovery client] ─unicast/TCP─> re-request · replay · snapshot
  1. Line handler (one per line): receive (often via kernel bypass), validate the header, extract (session, sequence). Busy-poll on a pinned core with no allocation on the hot path.
  2. Arbiter: track the next expected sequence number; pass the first copy downstream, drop duplicates; tolerate small reordering between lines.
  3. Gap detection: when sequence N is followed by N+k (k > 1) on the merged stream, record the gap and keep processing later messages. Never stall the live stream waiting for a retransmission.
  4. Recovery: small gaps go to re-request or replay. Large gaps, late joins and restarts go to a snapshot first, then you follow the live stream from the snapshot’s sequence. Merging recovered messages back in order is the part most often written wrong.
  5. Isolation: recovery uses separate connections, bandwidth and rate limits, so it can never slow the live feed. Exchanges build their retransmission services the same way.

Checklist (Unverified, engineering practice from the Mac note): pinned threads per line with a lock-free SPSC queue into the arbiter; process each (session, seq) once; a gap timer that escalates to snapshot rebuild; heartbeats also confirm the sequence position; after a reconnect, take a snapshot before trusting the stream; timestamp at every stage (NIC hardware timestamp, then line handler, then book) for later replay and audit.

Build it yourself: the 2027 prep plan portfolio project is a Python A/B publisher and receiver with gap recovery. The book view of the same mechanisms is in low-latency/07.

Open-source feed handlers to read (repositories exist as of 2026-10-03; code quality not reviewed): penberg/helix (C++ ITCH/MoldUDP64), harris2001/UltraLowLatencyFeedHandler, Ashutosh0x/rust-finance, Lunyn-HFT/lunary. ❌ The Mac note’s cmegroup/CME-Application-Samples returns 404.

Reliable multicast and middleware

Middleware What it is Notes (Unverified unless marked)
29West LBM / Informatica Ultra Messaging Commercial low-latency messaging over UDP multicast (LBT-RM, NAK-based, optional FEC) ❌ Informatica completed the acquisition in March 2010, not 2011 (DBTA and other news reports, via search)
UMDS UM’s “last mile” to desktops over TCP/WebSocket
OpenMAMA Middleware-neutral market data API (FINOS) Bridges to several transports
Aeron Open-source UDP unicast/multicast and shared-memory IPC, with Archive for record/replay
Solace PubSub+ Broker-based fan-out (not IP multicast) with replay
TIBCO Rendezvous Older UDP-multicast messaging Largely replaced

Low latency: network and hardware

Internal distribution

[exchange links, A/B multicast, kernel-bypass NIC]
        └─ edge ticker plant (normalize to an internal format, one timestamp base)
             ├─ internal multicast (LBM / Aeron / own)  → latency-critical strategy hosts
             ├─ shared-memory SPSC ring                → processes on the same host
             ├─ RDMA (RoCE)                            → high-throughput consumers
             └─ TCP / WebSocket                        → risk, research, desktops

Why some firms shrink internal multicast: IGMP/PIM/snooping state is hard to debug; join/leave or RP storms can hurt a whole segment; receivers still pay per-packet costs, which shared memory avoids; explicit receiver lists (TCP/RDMA) are easier to control than IGMP membership; RDMA matured. The Mac note cites Jane Street’s Signals & Threads, “Multicast and the Markets” (Brian Nigito; not re-listened). ❌ The other cited episode, “Ep.15 Reliable Multicast with Doug Patti”, does not match: Doug Patti’s episode is “State Machine Replication, and Why You Should Care”.

Timestamps and clocks

Interview angles

Topics the Mac note collected (they can seed the planned question bank in the prep plan):

  1. Why do exchanges use UDP multicast and not TCP? Answer with: fan-out cost, fairness, no head-of-line blocking, overhead, and reliability through sequence numbers plus recovery.
  2. How do you detect loss? (Sequence gaps on the merged A/B stream.)
  3. How do you arbitrate A/B? (First copy wins, reorder window, no failover.)
  4. How do you merge recovered data with the live stream? (Bounded buffering, escalate to a snapshot.)
  5. “Joined but no data”, why? Check:
    • no querier, or snooping aged out;
    • wrong NIC joined;
    • TTL 1;
    • RPF failure;
    • RP or PIM config;
    • static group missing.
  6. How do you cut receive latency? (Kernel bypass, pinning and isolcpus, busy polling, zero-copy parsing, SPSC queues, NUMA, hardware timestamps.)
  7. ARQ vs FEC, and why exchanges keep NACKs out of band.
  8. Clocks: NIC/FPGA timestamps, PTP vs NTP, RTS 25’s 100 µs.

TL;DR

  1. Multicast solves “the same event to everyone at the same moment”: O(1) fan-out, fairness, and slow consumers that hurt only themselves.
  2. Reliability lives in the application: A/B arbitration on the fast path, re-request/replay and snapshots on the slow path.
  3. Feeds differ in detail (MoldUDP64 64-bit message sequence, CME 32-bit packet sequence, MITCH’s own unit header) but share the pattern.
  4. Receiver latency: kernel bypass, pinned zero-copy hot path, hardware timestamps with PTP.

Source: knowledge base note multicast/05-market-data-in-practice.md — own-words notes with sources, projected at build time.