← All topics

Multicast & market data

IGMP, snooping and queriers, PIM-SM, A/B feed arbitration, public incident postmortems and exchange feed handling.

20 questions ·quiz yourself → · knowledge base projected 2026-10-03

multicast/05-market-data-in-practice.md · read the full note

Why do exchanges use UDP multicast and not TCP?

Answer with: fan-out cost, fairness, no head-of-line blocking, overhead, and reliability through sequence numbers plus recovery.

From multicast/05-market-data-in-practice.md · link
How do you detect loss?

(Sequence gaps on the merged A/B stream.)

From multicast/05-market-data-in-practice.md · link
How do you arbitrate A/B?

(First copy wins, reorder window, no failover.)

From multicast/05-market-data-in-practice.md · link
How do you merge recovered data with the live stream?

(Bounded buffering, escalate to a snapshot.)

From multicast/05-market-data-in-practice.md · link
"Joined but no data", why?

Check:

  • no querier, or snooping aged out;
  • wrong NIC joined;
  • TTL 1;
  • RPF failure;
  • RP or PIM config;
  • static group missing.
From multicast/05-market-data-in-practice.md · link
How do you cut receive latency?

(Kernel bypass, pinning and isolcpus, busy polling, zero-copy parsing, SPSC queues, NUMA, hardware timestamps.)

From multicast/05-market-data-in-practice.md · link
ARQ vs FEC for market data recovery — and why do exchanges keep NACK-style recovery out of band?

ARQ (retransmit on detected loss) is closed-loop and costs at least a round trip; FEC adds repair packets up front, an open-loop trade of bandwidth for no round trip. Exchange feeds keep recovery out of band: NACK-style re-request and replay run over separate unicast/TCP connections, never on the multicast tree, which avoids NACK storms. A/B arbitration already makes loss rare, so application-level FEC is uncommon on exchange feeds.

From multicast/05-market-data-in-practice.md · link
How do you handle timestamps and clocks in a trading stack?

Timestamp as early as possible: NIC hardware (ns, SO_TIMESTAMPING, check with ethtool -T) beats kernel (SO_TIMESTAMPNS, µs) beats the application. Sync with PTP (IEEE 1588): a GNSS-fed grandmaster, switches as boundary or transparent clocks, linuxptp (ptp4l + phc2sys) disciplining NIC and system clocks. MiFID II RTS 25 requires a maximum 100 µs divergence from UTC and 1 µs timestamp granularity for high-frequency algorithmic trading.

From multicast/05-market-data-in-practice.md · link

Extra practice

What is a microburst, and why do published duration figures differ by orders of magnitude?

A microburst is a very short burst whose instantaneous rate overruns an egress port's bandwidth or buffer, causing queuing or drops even when the average utilization is low.

Published durations disagree because sources measure at different granularities in different settings: a colo glossary says typically 1–100 milliseconds, a switch vendor's HFT primer says microseconds, and the BurstRadar paper measures tens to hundreds of microseconds. That is two to three orders of magnitude apart — say which time scale you mean before discussing one.

Derived from multicast/02-microburst-buffer.md; the three source figures are cited there. · site-written answer · link
Why can a minute-averaged bandwidth graph look healthy while the link is dropping packets?

Averaging hides sub-interval bursts. Pico's example: minute-averaged bandwidth well below link capacity, yet the feed was already dropping packets and gapping. Arista's number from a BATS cross-connect: 86 Mb/s averaged over one second, but up to 382 Mb/s over one millisecond — about 4.4×.

Market data runs over UDP with no retransmission, so packets dropped because a buffer filled are simply gone. The monitoring interval decides whether you can even see the problem: a per-minute graph can show a link at 10% while it drops packets.

Derived from multicast/02-microburst-buffer.md (Pico and Arista figures). · site-written answer · link
What is the buffer trade-off on a low-latency trading switch?

Buffers are the main source of switch latency, but under congestion you need them to avoid drops. Two causes of congestion make buffers necessary: ingress faster than egress (speed mismatch), and several ingress ports converging on one egress port (many-to-one) — the classic market-data fan-in.

The vendor framing is worth quoting in an interview: if a switch delays traffic by up to 2 ms during a microburst, the feed handler may mark the data stale by its timestamp and discard it anyway — a deep buffer avoids drops, but data that arrives late may be just as useless for trading.

Derived from multicast/02-microburst-buffer.md (Arista HFT primer concepts). · site-written answer · link
How do you detect a microburst? Why can't NetFlow or sFlow see them?

Microbursts live tens to hundreds of microseconds — sampling tools like NetFlow and sFlow simply never sample during one. Switches that detect that a burst happened (Cisco Nexus 5600/6000, Arista 7150S) can't tell you which flows caused it; INT adds a telemetry header to every packet, costing roughly 10% extra bandwidth.

The BurstRadar idea: a microburst lives in one egress port's queue, so everything needed to describe it is on a single switch — capture telemetry only for the packets involved in the burst and export it in courier packets cloned on demand (P4 on a Tofino). Results: with a burst every 200 µs it processes 10× less telemetry than INT and detects a burst within tens of µs of its start.

Derived from multicast/02-microburst-buffer.md (BurstRadar paper, NUS/APSys). · site-written answer · link
A multicast feed works at the open and dies a few minutes later, with no config change. What do you check first?

Whether the VLAN still has an IGMP querier. With snooping on but no querier (the only multicast router removed or disconnected), nobody sends Queries, so hosts stop sending Reports, and the switch's group entry ages out.

Where "a few minutes" comes from: the Group Membership Interval = robustness × query interval + query response interval = 2 × 125 + 10 = 260 s with RFC defaults. The fix: ip igmp snooping querier on the switch (best), or PIM on the SVI. Verify with show ip igmp snooping querier, show ip igmp snooping groups, show ip igmp snooping mrouter.

Derived from multicast/03-failure-modes.md trap #1 and the Lab 1 note. · site-written answer · link
Feed B has been silently dead for weeks and you only find out when feed A fails. What happened?

An RPF failure on the B feed. The A and B sources arrive on different uplinks; if unicast routing is asymmetric — the route to the B source points at the A uplink — every B packet fails the Reverse Path Forwarding check and is dropped. The config has A and B, but only A works, and B's absence only becomes visible when A dies.

Check show ip rpf <source>, show ip mroute <group>, and whether the RPF failed counter climbs in show ip mroute count. Fix the unicast routing, or pin the RPF interface with a static mroute (ip mroute <source> <mask> <rpf-neighbor>). Day to day: monitor reception on A and B separately, not just the arbitrated output.

Derived from multicast/03-failure-modes.md trap #2. · site-written answer · link
Why can PIM-SM's SPT switchover cause brief duplicates or loss, and what are the two ways to avoid it?

In ASM, receivers first get traffic down the RP's shared tree (RPT, *,G). On Cisco, as soon as the first packet from a new source arrives, the last-hop router joins (S,G) toward the source and prunes the source off the shared tree — during the move, both trees may carry traffic (duplicates) or neither (loss), and the path change can reorder packets.

Two escapes, each with a cost:

  • SSM (ip pim ssm default): receivers ask for (S,G) directly with IGMPv3 — no RP, no switchover. Receivers must know the source address.
  • Never switch (ip pim spt-threshold infinity): everything stays on the shared tree; the RP becomes a bandwidth, latency and failure bottleneck.

And read the exchange's docs first: not every venue uses SSM — CME's Aurora hub requires PIM-SM with an RP.

Derived from multicast/03-failure-modes.md trap #3. · site-written answer · link
Why do 32 IPv4 multicast groups share one MAC address, and why does it matter on a trading LAN?

An IPv4 multicast MAC is 01:00:5e + a zero bit + the low 23 bits of the group address. The group has 28 variable bits, so 5 bits are lost: 32 groups map to each MAC (239.1.1.5, 239.129.1.5, 224.1.1.5 … all → 01:00:5e:01:01:05).

A switch that snoops by MAC, or a NIC hardware filter, therefore lets 31 unwanted groups through; the host stack drops them by IP — but only after paying the interrupts and CPU, which is noise on a latency-sensitive feed host. Mitigations: plan groups so only the low 23 bits vary (e.g. stay inside 239.0.0.0–239.127.255.255), or use IP-based snooping where the switch supports it.

Derived from multicast/04-protocol-fundamentals.md. · site-written answer · link
ASM vs SSM: what is the difference, and which do exchanges actually use?

ASM (*,G): the network must discover sources, so PIM-SM needs an RP (and MSDP across domains). SSM (S,G "channels", RFC 4607): the receiver names the source with IGMPv3/MLDv2 — no RP, no source discovery, and strangers cannot inject traffic into the channel.

SSM fits market data in principle: publishers are known and fixed. But exchanges choose the model, not you. SGX is exploring separate RP addresses for its market data (FAQ 12.9) — PIM-SM with a rendezvous point; CME's Aurora hub also requires PIM-SM with an RP. Read the connectivity spec before designing the network.

Derived from multicast/04-protocol-fundamentals.md and exchanges/sg-hk-exchanges.md. · site-written answer · link
Does Linux load-balance multicast across SO_REUSEPORT sockets?Corrected vs source

No. Every matching socket gets its own copy of each multicast packet — SO_REUSEPORT lets several sockets bind the same group port, and the kernel delivers multicasts and broadcasts to each listener (__udp4_lib_mcast_deliver in net/ipv4/udp.c). The hash-based load balancing applies to unicast only.

If several processes on one host must share a feed, they each get the full stream and must deduplicate upstream, or one process reads and fans out internally (shared memory / SPSC rings).

Derived from multicast/04-protocol-fundamentals.md; checked against ip(7) and net/ipv4/udp.c. · site-written answer · link
Multicast packets are being lost. How do you find where on a Linux host?

Work from the socket down to the network, checking each drop counter:

  • First: did the host actually join? ip maddr show dev eth0 (note: ip maddr add only manages link-layer addresses — it cannot join an IP group), /proc/net/igmp.
  • Socket queue: netstat -su — receive buffer errors mean the application is too slow; raise SO_RCVBUF and net.core.rmem_max.
  • NIC: ethtool -S eth0 | grep -i drop.
  • Switch: interface counters — this is where microbursts show up.

For capture: tcpdump -ni eth0 'dst host 239.1.1.5', and ip[8] = 1 catches multicast still at TTL 1 (scoping mistakes). One multicast flow hashes to a single RX queue — steer it with ethtool -N flow steering.

Derived from multicast/04-protocol-fundamentals.md (Linux tools section). · site-written answer · link
Summarize the IGMP-snooping-without-querier failure in one breath.

"IGMP snooping constrains forwarding using membership state, and membership state is only maintained by query/report cycles — remove the querier and snooping either floods or silently ages out. It's the 'works for five minutes, then the market-data feed is gone' incident."

Fix ranking: snooping querier > PIM on the SVI > static mrouter port > static CAM entry > disabling snooping. Numbers to quote: GMI = robustness × QI + QRI; RFC defaults → 260 s; IOS snooping-querier default (60 s query interval) → ≈130 s — measure it in the lab rather than trusting either.

Derived from multicast/06-lab-1-igmp-snooping-no-querier.md (the lab's own interview one-liner and fix ranking). · site-written answer · link

Glossary

Multicast group (G)
A class D address (224.0.0.0/4). Packets sent to it are copied to every receiver that has joined
Source (S)
The host sending the multicast. For market data, the exchange's feed publisher
IGMP
How hosts tell the local router or switch which groups they want. v2 names only the group; v3 can also name the source
IGMP Report / Query
A Report says "I am in this group"; a Query is the querier asking every so often "who is still here?"
Querier
The device that sends Queries on a segment: usually a PIM router, or a switch running the snooping querier
IGMP snooping
A layer-2 switch listens to IGMP and sends multicast only to ports that have receivers, instead of flooding
mrouter port
The port on a snooping switch that leads to a multicast router
PIM-SM / ASM
Sparse mode: receivers first get traffic through the RP, from any source (Any-Source Multicast)
RP
Rendezvous Point: the router where sources and receivers meet in PIM-SM
RPT / shared tree (*,G)
The distribution tree rooted at the RP
SPT (S,G)
The shortest-path tree rooted at the source
SSM
Source-Specific Multicast: receivers ask for (S,G) directly, with no RP. Default range 232/8
RPF
Reverse Path Forwarding check: a multicast packet must arrive on the interface the router would use to reach the source by unicast, or it is dropped
A/B feeds, Line 1/2
The exchange sends the same data over two independent paths
A/B arbitration
The feed handler merges both feeds by sequence number: the first copy wins, and a gap on one side is filled from the other
Gap
A break in the sequence numbers, meaning lost packets. UDP does not retransmit; recovery goes through the exchange's retransmission or snapshot service
Microburst
A very short burst that overruns egress bandwidth or the buffer and causes queuing or drops; see [02](/multicast/microburst-buffer/)
Fan-in / fan-out
Many-to-one / one-to-many. Multicast itself is fan-out; several feeds converging on one egress port is fan-in
Feed handler
The program that receives market data, decodes it, arbitrates A/B and handles gaps
Line handler
The part of a feed handler that reads one line (A or B) and extracts sequence numbers; see [05](/multicast/market-data-in-practice/)
MLD
IPv6's version of IGMP (MLDv1 ≈ IGMPv2, MLDv2 ≈ IGMPv3), carried in ICMPv6
MoldUDP64
Nasdaq's framing for sequenced messages over UDP: session, 64-bit sequence number, message count (0 = heartbeat, 0xFFFF = end of session)
ITCH
Nasdaq's order-by-order market data message format; many venues use variants (SGX ITCH, Millennium's MITCH)
SBE
Simple Binary Encoding: fixed-layout binary messages that can be read in place; used by CME MDP 3.0 and iLink 3
Snapshot / market recovery
A feed or service that sends the current book state so a receiver can rebuild after a large gap or a late start
Re-request / replay
Asking the exchange to resend a range of sequence numbers (unicast UDP or TCP); for small gaps
Kernel bypass
Receiving packets in user space without the kernel network stack (Onload, ef_vi, DPDK, AF_XDP)
PTP
Precision Time Protocol (IEEE 1588): clock sync to sub-microsecond accuracy using hardware timestamps

Public incidents worth retelling

Montreal Exchange, 2019-06-13: a network change stopped multicast from leaving
Time2019-06-13, 1:30–6:00 a.m.
What happenedThe trading system was working, but multicast data for HSVF (High Speed Vendor Feed) and OBF (Order Book Feed) was not forwarded externally
How it was foundAround 2:00 a.m. a participant called to say they were not getting trade confirmation details on their feeds. **The exchange's own monitoring did not catch it**
Root causeA network change the previous evening, which stopped multicast data from reaching external participants
ResponseAt 3:30 a.m. all instruments were put into pre-open. After the fix they reopened at 6:00 a.m., so participants could remove, modify or add orders first
Source
- Multicast changes need a change window and a rollback plan. After the change, confirm from the **outside** that the multicast really arrives. A healthy matching engine does not mean the feed is getting out. - Watching internal system state is not enough. Monitor whether the feed can be received externally.
Sources: <https://www.m-x.ca/f_avis_tech_en/19-007_en.pdf>
JPX / TSE FLEX, 2022-06-07: Line 1 down all day, Line 2 carried the feed
TimeFrom about 07:00 JST on 2022-06-07, all day
ClassificationNetwork incident. System: FLEX (market information system)
ImpactLine 1 of FLEX Standard (100M) and FLEX Full (100M), at Co-Location and access points AP1–AP4
GroupsLine 1 of Standard group 014 and Full group 064 could not be distributed all day
User actionKeep running on the FLEX messages from Line 2. Standard (WB) and Full (WB) were not affected
Source (WebFetch gets 403; curl with a browser user agent works)
Lesson: A/B lines (Line 1 / Line 2) are not decoration. The feed handler must arbitrate **per multicast group**: if one group dies on the A side, use the B side for that group only. Do not mark the whole feed dead, and do not keep waiting for the A side.
Sources: <https://www.jpx.co.jp/english/systems/system-status/archive/20220607.html> (WebFetch gets 403; curl with a browser user agent works)
Moscow Exchange, 2015-06-29: the exchange switched in 15 seconds, clients took 30 seconds to 4 minutes
Time2015-06-29, 12:28–12:32 MSK
Root causeA network infrastructure component, the load balancer, malfunctioned
Exchange sideSwitched to the backup infrastructure within 15 seconds
Participant sideNo market data until they switched to the backup system, which took 30 seconds to 4 minutes depending on how they connected
Source
Lesson: however good the exchange's redundancy is, how long **you** are down depends on your own failover design. Do you subscribe to primary and backup at the same time? Does it switch automatically, or does a person have to step in?
Sources: <https://www.moex.com/n9836>
Nasdaq SIP, 2013-08-22: market data distribution is a single point for the whole market
NatureA SIP capacity and software problem, **not a multicast configuration fault**
Nasdaq's accountThat morning Arca disconnected and reconnected to the SIP more than 20 times, each time using significant resources. It also sent a stream of invalid stock symbols, which caused many reject messages. Arca's traffic was more than double the capacity of 10,000 messages per second per data port. A latent software flaw kept the system from failing over properly
Duration30 minutes to fix the technical problem, then nearly 3 hours of testing before trading resumed
Source 1
Source 2 (WebFetch gets 403; curl with a browser user agent works)
Where the source material differs: - It said NYSE disputed the attribution. The CNBC article does **not** report any NYSE response. It says Nasdaq accepted part of the responsibility, and it cites an independent analysis by Nanex (Eric Scott Hunsader): the quote burst "had to originate with the SIP itself", and neither Arca nor anything else outside the SIP could have caused it. So the accurate version is that an independent analysis disagreed with Nasdaq blaming Arca. - It said trading halted market-wide. As usually reported, the halt covered Nasdaq-listed stocks on all exchanges, while NYSE-listed stocks kept trading. **Unverified**: general knowledge, not checked against the two sources above. Lesson: the market data path is a single point for the whole market. Plan capacity for the extreme case (a reconnect storm on top of a flood of bad messages), and test failover for real.
Sources: <https://www.businessinsurance.com/nasdaq-says-software-bug-caused-trading-outage/> · <https://www.cnbc.com/2013/08/29/nasdaq-takes-some-responsibility-for-flash-freeze.html> (WebFetch gets 403; curl with a browser user agent works)

Takeaways

  1. Multicast solves "the same event to everyone at the same moment": O(1) fan-out, fairness, and slow consumers that hurt only themselves.
  2. Reliability lives in the application: A/B arbitration on the fast path, re-request/replay and snapshots on the slow path.
  3. Feeds differ in detail (MoldUDP64 64-bit message sequence, CME 32-bit packet sequence, MITCH's own unit header) but share the pattern.
  4. Receiver latency: kernel bypass, pinned zero-copy hot path, hardware timestamps with PTP.

Reference notes