Multicast & market data
IGMP, snooping and queriers, PIM-SM, A/B feed arbitration, public incident postmortems and exchange feed handling.
20 questions ·quiz yourself → · knowledge base projected 2026-10-03
multicast/05-market-data-in-practice.md · read the full note
Why do exchanges use UDP multicast and not TCP?
Answer with: fan-out cost, fairness, no head-of-line blocking, overhead, and reliability through sequence numbers plus recovery.
How do you detect loss?
(Sequence gaps on the merged A/B stream.)
How do you arbitrate A/B?
(First copy wins, reorder window, no failover.)
How do you merge recovered data with the live stream?
(Bounded buffering, escalate to a snapshot.)
"Joined but no data", why?
Check:
- no querier, or snooping aged out;
- wrong NIC joined;
- TTL 1;
- RPF failure;
- RP or PIM config;
- static group missing.
How do you cut receive latency?
(Kernel bypass, pinning and isolcpus, busy polling, zero-copy parsing, SPSC queues, NUMA, hardware timestamps.)
ARQ vs FEC for market data recovery — and why do exchanges keep NACK-style recovery out of band?
ARQ (retransmit on detected loss) is closed-loop and costs at least a round trip; FEC adds repair packets up front, an open-loop trade of bandwidth for no round trip. Exchange feeds keep recovery out of band: NACK-style re-request and replay run over separate unicast/TCP connections, never on the multicast tree, which avoids NACK storms. A/B arbitration already makes loss rare, so application-level FEC is uncommon on exchange feeds.
How do you handle timestamps and clocks in a trading stack?
Timestamp as early as possible: NIC hardware (ns, SO_TIMESTAMPING, check with ethtool -T) beats kernel (SO_TIMESTAMPNS, µs) beats the application. Sync with PTP (IEEE 1588): a GNSS-fed grandmaster, switches as boundary or transparent clocks, linuxptp (ptp4l + phc2sys) disciplining NIC and system clocks. MiFID II RTS 25 requires a maximum 100 µs divergence from UTC and 1 µs timestamp granularity for high-frequency algorithmic trading.
Extra practice
What is a microburst, and why do published duration figures differ by orders of magnitude?
A microburst is a very short burst whose instantaneous rate overruns an egress port's bandwidth or buffer, causing queuing or drops even when the average utilization is low.
Published durations disagree because sources measure at different granularities in different settings: a colo glossary says typically 1–100 milliseconds, a switch vendor's HFT primer says microseconds, and the BurstRadar paper measures tens to hundreds of microseconds. That is two to three orders of magnitude apart — say which time scale you mean before discussing one.
Why can a minute-averaged bandwidth graph look healthy while the link is dropping packets?
Averaging hides sub-interval bursts. Pico's example: minute-averaged bandwidth well below link capacity, yet the feed was already dropping packets and gapping. Arista's number from a BATS cross-connect: 86 Mb/s averaged over one second, but up to 382 Mb/s over one millisecond — about 4.4×.
Market data runs over UDP with no retransmission, so packets dropped because a buffer filled are simply gone. The monitoring interval decides whether you can even see the problem: a per-minute graph can show a link at 10% while it drops packets.
What is the buffer trade-off on a low-latency trading switch?
Buffers are the main source of switch latency, but under congestion you need them to avoid drops. Two causes of congestion make buffers necessary: ingress faster than egress (speed mismatch), and several ingress ports converging on one egress port (many-to-one) — the classic market-data fan-in.
The vendor framing is worth quoting in an interview: if a switch delays traffic by up to 2 ms during a microburst, the feed handler may mark the data stale by its timestamp and discard it anyway — a deep buffer avoids drops, but data that arrives late may be just as useless for trading.
How do you detect a microburst? Why can't NetFlow or sFlow see them?
Microbursts live tens to hundreds of microseconds — sampling tools like NetFlow and sFlow simply never sample during one. Switches that detect that a burst happened (Cisco Nexus 5600/6000, Arista 7150S) can't tell you which flows caused it; INT adds a telemetry header to every packet, costing roughly 10% extra bandwidth.
The BurstRadar idea: a microburst lives in one egress port's queue, so everything needed to describe it is on a single switch — capture telemetry only for the packets involved in the burst and export it in courier packets cloned on demand (P4 on a Tofino). Results: with a burst every 200 µs it processes 10× less telemetry than INT and detects a burst within tens of µs of its start.
A multicast feed works at the open and dies a few minutes later, with no config change. What do you check first?
Whether the VLAN still has an IGMP querier. With snooping on but no querier (the only multicast router removed or disconnected), nobody sends Queries, so hosts stop sending Reports, and the switch's group entry ages out.
Where "a few minutes" comes from: the Group Membership Interval = robustness × query interval + query response interval = 2 × 125 + 10 = 260 s with RFC defaults. The fix: ip igmp snooping querier on the switch (best), or PIM on the SVI. Verify with show ip igmp snooping querier, show ip igmp snooping groups, show ip igmp snooping mrouter.
Feed B has been silently dead for weeks and you only find out when feed A fails. What happened?
An RPF failure on the B feed. The A and B sources arrive on different uplinks; if unicast routing is asymmetric — the route to the B source points at the A uplink — every B packet fails the Reverse Path Forwarding check and is dropped. The config has A and B, but only A works, and B's absence only becomes visible when A dies.
Check show ip rpf <source>, show ip mroute <group>, and whether the RPF failed counter climbs in show ip mroute count. Fix the unicast routing, or pin the RPF interface with a static mroute (ip mroute <source> <mask> <rpf-neighbor>). Day to day: monitor reception on A and B separately, not just the arbitrated output.
Why can PIM-SM's SPT switchover cause brief duplicates or loss, and what are the two ways to avoid it?
In ASM, receivers first get traffic down the RP's shared tree (RPT, *,G). On Cisco, as soon as the first packet from a new source arrives, the last-hop router joins (S,G) toward the source and prunes the source off the shared tree — during the move, both trees may carry traffic (duplicates) or neither (loss), and the path change can reorder packets.
Two escapes, each with a cost:
- SSM (
ip pim ssm default): receivers ask for (S,G) directly with IGMPv3 — no RP, no switchover. Receivers must know the source address. - Never switch (
ip pim spt-threshold infinity): everything stays on the shared tree; the RP becomes a bandwidth, latency and failure bottleneck.
And read the exchange's docs first: not every venue uses SSM — CME's Aurora hub requires PIM-SM with an RP.
Why do 32 IPv4 multicast groups share one MAC address, and why does it matter on a trading LAN?
An IPv4 multicast MAC is 01:00:5e + a zero bit + the low 23 bits of the group address. The group has 28 variable bits, so 5 bits are lost: 32 groups map to each MAC (239.1.1.5, 239.129.1.5, 224.1.1.5 … all → 01:00:5e:01:01:05).
A switch that snoops by MAC, or a NIC hardware filter, therefore lets 31 unwanted groups through; the host stack drops them by IP — but only after paying the interrupts and CPU, which is noise on a latency-sensitive feed host. Mitigations: plan groups so only the low 23 bits vary (e.g. stay inside 239.0.0.0–239.127.255.255), or use IP-based snooping where the switch supports it.
ASM vs SSM: what is the difference, and which do exchanges actually use?
ASM (*,G): the network must discover sources, so PIM-SM needs an RP (and MSDP across domains). SSM (S,G "channels", RFC 4607): the receiver names the source with IGMPv3/MLDv2 — no RP, no source discovery, and strangers cannot inject traffic into the channel.
SSM fits market data in principle: publishers are known and fixed. But exchanges choose the model, not you. SGX is exploring separate RP addresses for its market data (FAQ 12.9) — PIM-SM with a rendezvous point; CME's Aurora hub also requires PIM-SM with an RP. Read the connectivity spec before designing the network.
Does Linux load-balance multicast across SO_REUSEPORT sockets?Corrected vs source
No. Every matching socket gets its own copy of each multicast packet — SO_REUSEPORT lets several sockets bind the same group port, and the kernel delivers multicasts and broadcasts to each listener (__udp4_lib_mcast_deliver in net/ipv4/udp.c). The hash-based load balancing applies to unicast only.
If several processes on one host must share a feed, they each get the full stream and must deduplicate upstream, or one process reads and fans out internally (shared memory / SPSC rings).
Multicast packets are being lost. How do you find where on a Linux host?
Work from the socket down to the network, checking each drop counter:
- First: did the host actually join?
ip maddr show dev eth0(note:ip maddr addonly manages link-layer addresses — it cannot join an IP group),/proc/net/igmp. - Socket queue:
netstat -su— receive buffer errors mean the application is too slow; raiseSO_RCVBUFandnet.core.rmem_max. - NIC:
ethtool -S eth0 | grep -i drop. - Switch: interface counters — this is where microbursts show up.
For capture: tcpdump -ni eth0 'dst host 239.1.1.5', and ip[8] = 1 catches multicast still at TTL 1 (scoping mistakes). One multicast flow hashes to a single RX queue — steer it with ethtool -N flow steering.
Summarize the IGMP-snooping-without-querier failure in one breath.
"IGMP snooping constrains forwarding using membership state, and membership state is only maintained by query/report cycles — remove the querier and snooping either floods or silently ages out. It's the 'works for five minutes, then the market-data feed is gone' incident."
Fix ranking: snooping querier > PIM on the SVI > static mrouter port > static CAM entry > disabling snooping. Numbers to quote: GMI = robustness × QI + QRI; RFC defaults → 260 s; IOS snooping-querier default (60 s query interval) → ≈130 s — measure it in the lab rather than trusting either.
Glossary
Public incidents worth retelling
Montreal Exchange, 2019-06-13: a network change stopped multicast from leaving
| Time | 2019-06-13, 1:30–6:00 a.m. |
| What happened | The trading system was working, but multicast data for HSVF (High Speed Vendor Feed) and OBF (Order Book Feed) was not forwarded externally |
| How it was found | Around 2:00 a.m. a participant called to say they were not getting trade confirmation details on their feeds. **The exchange's own monitoring did not catch it** |
| Root cause | A network change the previous evening, which stopped multicast data from reaching external participants |
| Response | At 3:30 a.m. all instruments were put into pre-open. After the fix they reopened at 6:00 a.m., so participants could remove, modify or add orders first |
| Source |
JPX / TSE FLEX, 2022-06-07: Line 1 down all day, Line 2 carried the feed
| Time | From about 07:00 JST on 2022-06-07, all day |
| Classification | Network incident. System: FLEX (market information system) |
| Impact | Line 1 of FLEX Standard (100M) and FLEX Full (100M), at Co-Location and access points AP1–AP4 |
| Groups | Line 1 of Standard group 014 and Full group 064 could not be distributed all day |
| User action | Keep running on the FLEX messages from Line 2. Standard (WB) and Full (WB) were not affected |
| Source |
Moscow Exchange, 2015-06-29: the exchange switched in 15 seconds, clients took 30 seconds to 4 minutes
| Time | 2015-06-29, 12:28–12:32 MSK |
| Root cause | A network infrastructure component, the load balancer, malfunctioned |
| Exchange side | Switched to the backup infrastructure within 15 seconds |
| Participant side | No market data until they switched to the backup system, which took 30 seconds to 4 minutes depending on how they connected |
| Source |
Nasdaq SIP, 2013-08-22: market data distribution is a single point for the whole market
| Nature | A SIP capacity and software problem, **not a multicast configuration fault** |
| Nasdaq's account | That morning Arca disconnected and reconnected to the SIP more than 20 times, each time using significant resources. It also sent a stream of invalid stock symbols, which caused many reject messages. Arca's traffic was more than double the capacity of 10,000 messages per second per data port. A latent software flaw kept the system from failing over properly |
| Duration | 30 minutes to fix the technical problem, then nearly 3 hours of testing before trading resumed |
| Source 1 | |
| Source 2 |
Takeaways
- Multicast solves "the same event to everyone at the same moment": O(1) fan-out, fairness, and slow consumers that hurt only themselves.
- Reliability lives in the application: A/B arbitration on the fast path, re-request/replay and snapshots on the slow path.
- Feeds differ in detail (MoldUDP64 64-bit message sequence, CME 32-bit packet sequence, MITCH's own unit header) but share the pattern.
- Receiver latency: kernel bypass, pinned zero-copy hot path, hardware timestamps with PTP.