← Multicast & market data

Three classic market data multicast failure modes

These three are well-known problems in network engineering, but no matching public incident report was found. Do not invent a source. The mechanisms and commands below follow the RFCs and Cisco’s official docs. Commands are Cisco IOS / IOS XE syntax; for other vendors, check their own docs.

1. IGMP snooping with no querier: fine at the open, feed gone a few minutes later

Mechanism

Where 260 seconds comes from: the Group Membership Interval in RFC 2236 (IGMPv2) and RFC 3376 (IGMPv3) = Robustness Variable × Query Interval + Query Response Interval = 2 × 125 + 10 = 260 seconds (all defaults).

Fix: put a querier in the VLAN, either a layer-3 interface running PIM or the snooping querier on the switch:

ip igmp snooping querier
ip igmp snooping querier address <IP>   ! with no address set and no IP on the device, no general queries are sent

Troubleshooting commands

show ip igmp snooping querier
show ip igmp snooping groups
show ip igmp snooping mrouter

References

2. RPF failure on the A/B feeds: the redundancy looks fine but is gone

Mechanism

Fix

Troubleshooting commands

show ip rpf <source address>    ! RPF interface and RPF neighbor
show ip mroute <group address>  ! is the incoming interface correct?
show ip mroute count            ! is the RPF failed counter climbing?

Day to day: monitor reception on the A and B feeds separately, not just the arbitrated output.

References

3. PIM-SM SPT switchover: brief duplicates or loss

Mechanism

Two common approaches

Approach How Cost
SSM Receivers use IGMPv3 to ask for (S,G) directly. No RP and no shared tree, so no switchover. On Cisco: ip pim ssm default (default range 232.0.0.0/8) Receivers must know the source address; hosts and switches must support IGMPv3
Never switch to SPT ip pim spt-threshold infinity [group-list <ACL>]: every source for those groups stays on the shared tree All traffic goes through the RP, which becomes a bandwidth and latency bottleneck and a single point of failure

Note: not every exchange uses SSM. CME Globex Hub - Aurora requires PIM Sparse Mode: the customer configures the RP’s IP address, the routers terminating CME connections must run BGP, and each router needs a fixed path to its CME data center. So read the exchange’s connectivity docs first.

References

Source: knowledge base note multicast/03-failure-modes.md — own-words notes with sources, projected at build time.