Three classic market data multicast failure modes
These three are well-known problems in network engineering, but no matching public incident report was found. Do not invent a source. The mechanisms and commands below follow the RFCs and Cisco’s official docs. Commands are Cisco IOS / IOS XE syntax; for other vendors, check their own docs.
1. IGMP snooping with no querier: fine at the open, feed gone a few minutes later
Mechanism
- With IGMP snooping on, the switch forwards multicast only to ports that have sent an IGMP Report.
- A host sends a Report on its own when it joins a group, so traffic flows at first.
- After that, Reports are refreshed only when a querier sends Queries. With no querier on the segment (no PIM router, no snooping querier), nobody sends Queries and hosts stop sending Reports.
- When the entry ages out, the switch prunes the traffic and the server loses the feed.
Where 260 seconds comes from: the Group Membership Interval in RFC 2236 (IGMPv2) and RFC 3376 (IGMPv3) = Robustness Variable × Query Interval + Query Response Interval = 2 × 125 + 10 = 260 seconds (all defaults).
- The actual snooping aging time and behavior vary by vendor: some switches flood unknown multicast when there is no mrouter port, others drop it. Check the device docs and test it.
Fix: put a querier in the VLAN, either a layer-3 interface running PIM or the snooping querier on the switch:
ip igmp snooping querier
ip igmp snooping querier address <IP> ! with no address set and no IP on the device, no general queries are sent
Troubleshooting commands
show ip igmp snooping querier
show ip igmp snooping groups
show ip igmp snooping mrouter
References
- RFC 2236 §8.4: https://www.rfc-editor.org/rfc/rfc2236#section-8.4
- RFC 3376 §8.4: https://www.rfc-editor.org/rfc/rfc3376#section-8.4
- Cisco Catalyst 3850 IGMP Snooping configuration guide (includes the snooping querier): https://www.cisco.com/c/en/us/td/docs/switches/lan/catalyst3850/software/release/16-6/configuration_guide/ip_mcast_rtng/b_166_ip_mcast_rtng_3850_cg/configuring_igmp_snooping.html
2. RPF failure on the A/B feeds: the redundancy looks fine but is gone
Mechanism
- When a router receives a multicast packet, it runs the RPF (Reverse Path Forwarding) check: the packet must arrive on the interface that the router’s unicast routing uses to reach the source. Otherwise it is dropped. This prevents loops.
- The A and B feed sources come in on different uplinks. If unicast routing is asymmetric (the route to the B source points at the A uplink), B’s multicast packets fail RPF and are dropped.
- Result: the config has A and B, but only one of them works. You find out that B was broken all along when A fails.
Fix
- Fix unicast routing so that each source’s RPF interface is the one its traffic actually arrives on.
- Or set the RPF interface or neighbor with a static mroute:
ip mroute <source network> <mask> <RPF neighbor or interface>. - If several hops are asymmetric, a GRE tunnel plus static mroutes also works.
Troubleshooting commands
show ip rpf <source address> ! RPF interface and RPF neighbor
show ip mroute <group address> ! is the incoming interface correct?
show ip mroute count ! is the RPF failed counter climbing?
Day to day: monitor reception on the A and B feeds separately, not just the arbitrated output.
References
- Cisco IP Multicast Troubleshooting Guide (worked RPF failure cases): https://www.cisco.com/c/en/us/support/docs/ip/ip-multicast/16450-mcastguide0.html
3. PIM-SM SPT switchover: brief duplicates or loss
Mechanism
- In PIM-SM (ASM), receivers first get traffic down the RP’s shared tree (RPT, (*,G)).
- By default on Cisco, as soon as the first packet from a new source arrives, the last-hop router sends an (S,G) Join toward the source, moves to the shortest-path tree (SPT), and prunes that source off the shared tree.
- During the switch, both trees may carry traffic for a moment (duplicates), or neither may (loss), and the path change can reorder packets.
Two common approaches
| Approach | How | Cost |
|---|---|---|
| SSM | Receivers use IGMPv3 to ask for (S,G) directly. No RP and no shared tree, so no switchover. On Cisco: ip pim ssm default (default range 232.0.0.0/8) |
Receivers must know the source address; hosts and switches must support IGMPv3 |
| Never switch to SPT | ip pim spt-threshold infinity [group-list <ACL>]: every source for those groups stays on the shared tree |
All traffic goes through the RP, which becomes a bandwidth and latency bottleneck and a single point of failure |
Note: not every exchange uses SSM. CME Globex Hub - Aurora requires PIM Sparse Mode: the customer configures the RP’s IP address, the routers terminating CME connections must run BGP, and each router needs a fixed path to its CME data center. So read the exchange’s connectivity docs first.
References
- Cisco IOS IP Multicast Command Reference (
ip pim spt-threshold): https://www.cisco.com/c/en/us/td/docs/ios-xml/ios/ipmulti/command/imc-cr-book/imc_i3.html - Cisco “Optimizing PIM Sparse Mode in a Large IP Multicast Deployment” (Catalyst 9500, 17.3): https://www.cisco.com/c/en/us/td/docs/switches/lan/catalyst9500/software/release/17-3/configuration_guide/ip_mcast_rtng/b_173_ip_mcast_rtng_9500_cg/ip_multicast_optimization__optimizing_pim_sparse_mode_in_a_large_ip_multicast_deployment.html
- CME Globex Hub - Aurora: https://cmegroupclientsite.atlassian.net/wiki/spaces/EPICSANDBOX/pages/457088573/CME+Globex+Hub+-+Aurora
- CME MDP 3.0 Dissemination: https://cmegroupclientsite.atlassian.net/wiki/display/EPICSANDBOX/MDP+3.0+-+Dissemination (not read in detail yet)