Microbursts and buffers
What a microburst is (three sources, three time scales)
| Source | Duration | Other points |
|---|---|---|
| Beeks glossary | Typically 1–100 milliseconds | The instantaneous rate can be tens or hundreds of times the average, or even above port bandwidth; it can overflow switch and NIC buffers and drop packets |
| Arista HFT Primer | Typically microseconds | Packets headed for an egress port overrun its bandwidth for a very short time. Over a longer period the same thing is just congestion or oversubscription |
| BurstRadar paper | Tens to hundreds of microseconds | Queuing microbursts in datacenter networks |
When studying: the “microburst” time scale differs by two to three orders of magnitude between sources, because they measure at different granularity and look at different settings. Say which time scale you mean before discussing one.
Links:
- Beeks: https://guides.beeksgroup.com/glossary/Microburst.html
- Arista: https://www.arista.com/assets/data/pdf/HFT/HFTTradingNetworkPrimer.pdf
- BurstRadar: https://www.comp.nus.edu.sg/~chanmc/papers-by-year/2018-apsys-burstradar.pdf
Why average bandwidth graphs do not show the problem
Pico (formerly Corvil):
- The minute-averaged bandwidth was well below link capacity (50 Mbps in their example), yet the feed was already dropping packets and gapping. Only the per-microburst maximum shows it.
- Market data runs over UDP, so there is no retransmission. Packets dropped because the buffer filled up are gone, and so is the market data in them.
- Numbers: some of the largest feeds burst to 20–24 Gbps, while many trading firms run 10 Gbps links with a steady 6–8 Gbps.
- Source: https://pico.net/blog/why-quality-of-market-data-matters-more-in-volatile-markets
Arista gives a number too: on a BATS cross-connect, 86 Mb/s averaged over one second, but up to 382 Mb/s measured over one millisecond (about 4.4×).
Takeaway: the monitoring interval decides whether you can see the problem. A per-minute graph can show a link at 10% while it is dropping packets.
The buffer trade-off (Arista Primer)
- Two main causes of congestion: ingress faster than egress (speed mismatch), and several ingress ports sending to one egress port at the same time (many-to-one).
- Buffering and queuing are the main sources of latency in an Ethernet switch, but under congestion you need buffers to avoid drops. Buffer size and allocation have to balance no-loss against ultra-low latency.
- Arista asks: if a switch delays traffic by up to 2 ms during a microburst, won’t the host or feed handler mark the data stale by its timestamp and maybe drop it? The point: a deep buffer avoids drops, but data that arrives late may be just as useless for trading.
Reading the Arista Primer with care (differs from the source material)
The source material said Arista “discusses how fan-in causes congestion”. The original says something else:
- It is a vendor solution brief that defends Arista’s own low-latency switches. The text says RFC 2544 was written “eleven years ago”, so it dates from about 2010; the PDF footer says © 2016.
- It discusses fan-in in order to criticize competitors’ tests that send multicast from 23 or 47 ports into one port at once. Arista calls this many-to-one multicast contrived and never seen in real trading networks, and stresses that multicast is one-to-many (fan-out).
- It says 10G links offer far more bandwidth than real market data needs. Pico’s later numbers (feeds bursting to 20–24 Gbps) contradict this, so that conclusion is out of date.
- The 2 ms line is a rhetorical question, not an incident report.
So use the Primer for concepts (buffers, microbursts, fan-in vs fan-out), not as a neutral source of facts.
Detecting microbursts: BurstRadar
- Paper: “BurstRadar: Practical Real-time Microburst Monitoring for Datacenter Networks” (NUS). The file name says APSys 2018; the PDF is the anonymized submission (“Paper #27”).
- Problem: microbursts last only tens to hundreds of µs, so sampling tools like NetFlow and sFlow cannot see them at all. Commercial switches (the paper names Cisco Nexus 5600/6000 and Arista 7150S) can detect that a microburst happened but not which flows caused it. INT (In-band Telemetry) adds a telemetry header to every packet, which is costly (the paper estimates about 10% extra bandwidth).
- Idea: a microburst lives in one port’s egress queue, so everything needed to describe it is on a single switch. Capture telemetry only for the packets involved in the microburst, and export it in courier packets cloned on demand.
- Implementation: written in P4 on a Barefoot Tofino switch.
- Results: even with a microburst every 200 µs, it processes 10× less telemetry data than INT; it detects a burst within tens of µs of its start; and it handles simultaneous bursts on several egress ports.