Multicast protocol fundamentals (addressing, IGMP, snooping, PIM, sockets)
Merged on 2026-10-03 from the Mac session’s research note 2026-10-03-multicast-quant/01-multicast-fundamentals.md (AI-written, Chinese), translated and condensed.
Verification status. ✅ = checked against the source named. ❌ / ⚠️ = the Mac note was wrong or imprecise and is corrected here (also logged in the group verification log). Everything else is Unverified. It reads as standard textbook material, but nobody has checked it line by line. The failure stories (no querier, RPF, SPT switchover) are in 03 and are not repeated here.
Why it matters for trading
- Exchanges send market data as UDP multicast inside data centres and colocation: Nasdaq ITCH over MoldUDP64, CME MDP 3.0, and most European venues. Details in 05.
- Firms often re-multicast normalized data internally, so one send reaches every strategy, risk and storage process.
- Multicast across the public internet is close to unusable (see “Limits” below), so in practice it is a technology for one data centre, campus or colo.
Delivery models
| Model | Who gets the packet | One-to-many cost | Typical use |
|---|---|---|---|
| Unicast | One receiver | N receivers = N copies from the sender | Client/server |
| Broadcast | Every host in the L2 domain | Stops at the router; everyone must process it | ARP, DHCP discover |
| Multicast | Hosts that joined the group | Sender sends once; the network copies at branch points, so each link carries a packet once | IPTV, market data, conferencing |
| Anycast | The nearest of several nodes (by routing metric) | One-to-one, but the target varies | DNS root servers, CDN edges |
The model comes from RFC 1112 (Host Extensions for IP Multicasting, August 1989 ✅; the Mac note said 1986, which is the date of its predecessor RFC 988). Three ideas:
- The source sends once to a group address and does not track receivers.
- The network replicates packets along a distribution tree.
- It is receiver-driven: hosts join with IGMP, and the tree grows from the last-hop routers.
The cost: multicast runs over UDP, so there is no reliability, ordering or congestion control. Applications add those (see “Reliable multicast” below and 05).
Addresses
IPv4: 224.0.0.0/4 (old class D)
| Block | Name | Notes |
|---|---|---|
| 224.0.0.0/24 | Local Network Control | Never forwarded off the subnet. 224.0.0.1 all hosts, .2 all routers, .5/.6 OSPF, .9 RIPv2, .22 IGMPv3 reports, .251 mDNS, .252 LLMNR |
| 224.0.1.0/24 | Internetwork Control | Routable control traffic; 224.0.1.1 = NTP ✅ IANA |
| 224.2.0.0/16 | SDP/SAP block | Old MBone session announcements |
| 232.0.0.0/8 | SSM | Source-Specific Multicast ✅ IANA (RFC 4607) |
| 233.0.0.0–233.251.255.255 | GLOP | Static per-AS allocation (16-bit ASN in the middle octets) ✅ IANA (RFC 3180). ⚠️ Not all of 233/8 as the Mac note said: 233.252.0.0/24 is MCAST-TEST-NET |
| 234.0.0.0/8 | Unicast-prefix-based | ✅ IANA (RFC 6034) |
| 239.0.0.0/8 | Administratively scoped (organization-local) | ✅ IANA (RFC 2365). The private space, like RFC 1918 for unicast. 239.255.0.0/16 = local scope, 239.192.0.0/14 = organization-local scope |
Practice: plan internal market-data groups inside 239/8 (for example by site / asset class / service), not in low 224.x, and avoid overlap between sites.
Source: IANA IPv4 Multicast Address Space Registry, https://www.iana.org/assignments/multicast-addresses/multicast-addresses.xhtml (checked 2026-10-03).
IPv6: ff00::/8
- Format:
ff+ 4 flag bits + 4 scope bits + 112-bit group ID. - ⚠️ Flags are
0RPT: T (0x1) = transient (not well-known), P (0x2) = prefix-based (RFC 3306), R (0x4) = embedded RP (RFC 3956). The SSM range ff3x::/32 has P=1 and T=1. The Mac note called 0x2 an “S flag”; there is no such flag. (No single source: corrected from general knowledge; the RFCs were not opened.) - Scope: 1 interface-local, 2 link-local, 4 admin-local, 5 site-local, 8 organization-local, e global.
- Well-known groups ✅ IANA IPv6 registry: ff02::1 all nodes, ff02::2 all routers, ff02::1:2 All_DHCP_Relay_Agents_and_Servers, ff05::1:3 All_DHCP_Servers. ❌ The Mac note listed ff02::1:3 as “all DHCP servers”; ff02::1:3 is LLMNR (RFC 4795).
- Group membership in IPv6 is MLD (MLDv1 = RFC 2710, MLDv2 = RFC 3810), carried in ICMPv6 (next header 58). IGMP is IP protocol 2.
Multicast MAC addresses and the 32:1 overlap
- IPv4: MAC =
01:00:5e+ a 0 bit + the low 23 bits of the group address. Example: 239.1.1.5 →01:00:5e:01:01:05. - IPv6:
33:33+ the low 32 bits of the group (RFC 2464). - An IPv4 group has 28 variable bits but only 23 reach the MAC, so 32 groups share each MAC. 239.1.1.5 and 239.129.1.5 (and 224.1.1.5, 225.1.1.5 …) all map to
01:00:5e:01:01:05. - Effect: a switch that snoops by MAC, or a NIC filter, lets unwanted groups through. The host stack drops them by IP, but only after spending interrupts and CPU, which is noise on a latency-sensitive feed host.
- Fixes: plan groups so that only the low 23 bits vary (keep the first octet and the top bit of the second octet fixed, e.g. stay inside 239.0.0.0–239.127.255.255), or use IP-based snooping where the switch supports it.
IGMP: hosts and the first-hop router
| Version | RFC | What it adds |
|---|---|---|
| IGMPv1 | RFC 1112 (1989) | Query and Report. No Leave: membership just times out |
| IGMPv2 | RFC 2236 (1997) | Leave Group (sent to 224.0.0.2), Group-Specific Query, Max Response Time, querier election (lowest IP wins) |
| IGMPv3 | RFC 3376 (2002) | Source filtering: INCLUDE (only from S) or EXCLUDE (all but S). Reports go to 224.0.0.22. RFC 4604 covers its use for SSM |
How it works:
- Querier: one device per segment sends Queries, normally a router; in a pure L2 network, a switch’s snooping querier (03 §1).
- Join: the host sends an unsolicited Report straight away. v1/v2 hosts suppress their Report if they hear another member’s; v3 drops suppression so the router can track every listener.
- Leave (v2+): the host sends Leave, the router sends a Group-Specific Query, and traffic stops only if nobody answers.
- Timers people tune: query interval (125 s in the RFC; Cisco IOS defaults to 60 s), robustness variable (2), last-member query interval (1 s). Shorter timers make joins and leaves take effect faster.
ASM vs SSM
- ASM (Any-Source, (*,G)): the network must discover sources. PIM-SM does it with an RP, and MSDP between domains.
- SSM (RFC 4607, (S,G) “channels”): the receiver names the source with IGMPv3 or MLDv2. No RP, no source discovery, and strangers cannot inject traffic into the channel. IPv4 range 232/8, IPv6 ff3x::/32.
- SSM fits market data well in principle: the publisher is known and fixed. But exchanges choose the model, not you. CME’s Aurora hub requires PIM-SM with an RP (03 §3).
Layer 2: keeping multicast from flooding
- A switch floods frames whose destination MAC it has not learned, and multicast MACs are never learned. So a switch with no multicast config floods every group to every port in the VLAN: every host takes the interrupts.
- IGMP snooping: the switch listens to IGMP and builds a group → ports table. It is a behaviour, not a protocol, so the only document is the informational RFC 4541. It says 224.0.0.x (link-local) traffic should still be flooded. Vendors differ.
- No querier, no reliable snooping: entries age out. This is the classic trap in 03 §1.
- Proxy reporting / report suppression on the switch hides real membership from the router. Remember this layer when troubleshooting.
- CGMP: Cisco’s pre-snooping protocol in which the router tells the switch which ports want a MAC. Legacy; replaced by snooping.
- MVR (Multicast VLAN Registration): one multicast VLAN shared by receivers in many VLANs (common in IPTV).
- Typical market-data LAN: a dedicated VLAN, IGMP snooping, an explicit querier, and storm control on access ports.
Layer 3: the PIM family
PIM is “protocol independent” because it uses whatever unicast routing table exists (OSPF, IS-IS, BGP) for its RPF check.
- PIM-DM (RFC 3973): flood and prune, then re-flood periodically. Simple, chatty, only for dense small networks. DVMRP (RFC 1075) used the same idea.
- PIM-SM (RFC 7761): explicit joins, no flooding.
- Each group has an RP. The receiver-side DR sends (*,G) Joins toward the RP, building the shared tree (RPT).
- The source-side DR Registers (encapsulates) the first packets to the RP so the RP learns the source.
- The last-hop router can then join (S,G) toward the source and move to the SPT.
- RPF check: forward a multicast packet only if it arrived on the interface this router would use to reach the source (or the RP) by unicast; otherwise drop it. It prevents loops, and it is the most common cause of multicast loss when unicast routing is asymmetric or reconverging (03 §2). ❌ The Mac note cited RFC 7761 §4.4.2 for RPF. That section is “Receiving Register Messages at the RP”. RPF is defined with “RPF Neighbor” among the RFC’s definitions and applied in §4.2, “Data Packet Forwarding Rules” ✅ (RFC 7761 text).
- SPT switchover, two views to reconcile. The Mac note says trading networks often keep Cisco’s default immediate switchover, to avoid the RP detour. 03 §3 says networks avoid the switchover glitch with SSM or
spt-threshold infinity. Both are real trade-offs and neither came with a source: the default gives the shortest path but a brief glitch when each new source appears; infinity keeps everything on the RP path; SSM removes the question. No single source for which one is “common”. - Bidir-PIM (RFC 5015): shared trees only, no (S,G) state; for many-to-many.
- PIM-SSM (RFC 4607, overview RFC 3569): (S,G) trees only.
- MBGP: BGP carrying multicast routes (SAFI 2), so that the multicast RPF topology can differ from unicast. ⚠️ The Mac note cited RFC 2858; that RFC is obsoleted by RFC 4760 ✅ (rfc-editor.org).
- MSDP (RFC 3618): RPs in different PIM domains exchange Source-Active messages so ASM can work across domains. Little used today; new designs prefer SSM.
Reliable multicast (overview)
- ACK-based: every receiver acknowledges, which causes ACK implosion as the group grows.
- NACK-based: receivers report only gaps. Random timer back-off means one NACK covers everyone missing the same packet, and one retransmission serves them all. Examples: PGM (RFC 3208), NORM (RFC 5740).
- FEC: the sender adds repair packets (Reed-Solomon, RaptorQ RFC 6330), so receivers rebuild losses with no round trip. The cost is extra bandwidth plus coding delay. Framework: RFC 3453. The Mac note’s “quant firms use 5–20% redundancy” had no source and is dropped. Exchange feeds generally use A/B feeds plus recovery services instead (05).
- The design trade-offs are covered in low-latency/07 (ARQ vs FEC vs repetition, ACK implosion).
Limits and operations
- Scoping: the old MBone TTL convention (1 = subnet, under 32 = site, and so on) is coarse. Prefer administrative scoping (239/8, RFC 2365) with boundaries on routers.
- Why the public internet has no multicast: no billing model for one-to-many; interdomain ASM needs PIM-SM + RP + MBGP + MSDP; patchy support in access networks; abuse worries. One-to-many on the internet is done with unicast CDNs instead.
- Operational pitfalls: RPF failures when unicast routing changes; snooping aging with no querier (flooding or black holes); 32:1 MAC overlap; no congestion control, so a drop hits every receiver at once; unfamiliar tooling (snooping tables, mroutes,
ip maddr).
Sockets and Linux tools
struct ip_mreq mreq;
mreq.imr_multiaddr.s_addr = inet_addr("239.1.1.5"); /* group */
mreq.imr_interface.s_addr = inet_addr("10.0.0.5"); /* local NIC IP; INADDR_ANY lets the kernel pick */
int on = 1;
setsockopt(fd, SOL_SOCKET, SO_REUSEADDR, &on, sizeof on); /* before bind() */
/* bind() to the group port, then join: */
setsockopt(fd, IPPROTO_IP, IP_ADD_MEMBERSHIP, &mreq, sizeof mreq);
/* sender side */
setsockopt(fd, IPPROTO_IP, IP_MULTICAST_IF, &ifaddr, sizeof ifaddr);
unsigned char ttl = 4; setsockopt(fd, IPPROTO_IP, IP_MULTICAST_TTL, &ttl, sizeof ttl); /* default 1 */
unsigned char loop = 0; setsockopt(fd, IPPROTO_IP, IP_MULTICAST_LOOP, &loop, sizeof loop); /* default 1 */
| Option | Default | Notes |
|---|---|---|
IP_ADD_MEMBERSHIP / IP_DROP_MEMBERSHIP |
— | Join or leave. On multi-NIC hosts name the interface; with INADDR_ANY the kernel may join on the wrong NIC (a classic “joined but no data”) |
SO_REUSEADDR |
off | Several sockets or processes on the same group port; set it before bind() |
SO_REUSEPORT |
off | ❌ The Mac note said Linux load-balances multicast across SO_REUSEPORT sockets. It does not: every matching socket gets its own copy (“Multicasts and broadcasts go to each listener”, __udp4_lib_mcast_deliver in Linux net/ipv4/udp.c ✅). Hash load-balancing applies to unicast |
IP_MULTICAST_TTL |
1 | 1 means the traffic never leaves the subnet; raise it when the feed crosses routers |
IP_MULTICAST_LOOP |
1 | The sending host also receives its own packets; feed publishers usually turn it off |
IP_MULTICAST_IF |
— | Choose the sending NIC; needed on multi-NIC hosts |
SO_RCVBUF |
OS default | Raise it (and net.core.rmem_max) for bursts, or the socket queue overflows and drops |
IPv6 equivalents: IPV6_JOIN_GROUP (struct ipv6_mreq), IPV6_MULTICAST_HOPS, IPV6_MULTICAST_LOOP, IPV6_MULTICAST_IF. Man page: ip(7).
Linux tools:
ip maddr show [dev eth0]: multicast addresses on each interface. Check this first: did the host really join?- ❌
ip maddr add/deldoes not join an IP group. Per ip-maddress(8) ✅ it only manages link-layer (MAC) addresses and cannot join protocol multicast groups. Join with a socket, or with a tool that opens one. netstat -g(net-tools) lists group memberships. ❌ss -gdoes not exist; ss(8) ✅ has no such option./proc/net/igmp: kernel IGMP state per interface.- tcpdump filters (No single source, standard pcap syntax):
tcpdump -ni eth0 'ip multicast' # all IPv4 multicast tcpdump -ni eth0 'dst host 239.1.1.5' # one group tcpdump -ni eth0 'ether[0] & 1 = 1' # all multicast/broadcast frames tcpdump -ni eth0 'ip multicast and ip[8] = 1' # multicast with TTL 1 (scoping mistakes) tcpdump -ni eth0 igmp # joins, leaves, queries - Where were packets dropped?
netstat -su(receive buffer errors = the socket queue),ethtool -S eth0 | grep -i drop(the NIC), switch interface counters (the network). - RSS: one multicast flow hashes to a single RX queue. Steer it to a chosen queue or core with
ethtool -N(flow steering). - smcroute: static multicast routing daemon for small setups without PIM.
References
- RFC 1112 (IGMPv1, host model, MAC mapping), 2236 (IGMPv2), 3376 (IGMPv3), 4604 (IGMPv3/MLDv2 for SSM), 4541 (snooping, informational), 2365 (admin scoping), 4607 (SSM), 3569 (SSM overview), 7761 (PIM-SM), 3973 (PIM-DM), 5015 (Bidir-PIM), 3618 (MSDP), 4760 (multiprotocol BGP), 3208 (PGM), 5740 (NORM), 6330 (RaptorQ), 3453 (FEC in reliable multicast). All at
https://www.rfc-editor.org/rfc/rfcNNNN. - IANA registries: IPv4 multicast, IPv6 multicast.
- Linux source:
net/ipv4/udp.c(https://github.com/torvalds/linux/blob/master/net/ipv4/udp.c).