← Latest papers
💰 quantitative finance

Packets, Transactions and Queues: Design Principles for HFT Systems from a Measurement Study of CME Market Data

By analyzing over a year of CME market data, this paper challenges the conventional single-threaded HFT design by demonstrating that while one thread suffices for sub-period packet processing, a two-stage threaded architecture can significantly reduce queuing tails caused by transaction bursts, provided the split shortens the system's slowest stage.

Original authors: Vincent Maciejewski

Published 2026-09-29✓ Author reviewed ⓘ
📖 6 min read🧠 Deep dive

Original authors: Vincent Maciejewski

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the world of high-frequency trading, where computers buy and sell stocks in fractions of a second, speed is not just an advantage; it is the entire game. These systems operate on a simple premise: if you can process market information faster than anyone else, you can profit from tiny price differences before they disappear. To do this, engineers build specialized software that listens to a constant stream of data from stock exchanges, decodes it, and makes decisions in microseconds. For years, the industry has operated under a strict rule: keep the most critical part of this software on a single processor core. The logic was that moving data between different cores, or "threads," was too slow and risky, adding delays that would ruin the system's speed. This approach treated the software like a single, focused worker who never hands off a task, believing that any interruption would cost more than the work itself.

However, this long-held belief relied on an assumption about how market data arrives: that it comes in a steady, random stream, like raindrops falling at unpredictable intervals. If that were true, the single-worker approach would indeed be the fastest. But what if the data doesn't fall randomly? What if it arrives in sudden, intense bursts, where thousands of updates hit the system in the blink of an eye? A new measurement study of real market data from the Chicago Mercantile Exchange suggests that the old rule might be wrong for the busiest moments. By tracking billions of data packets over more than a year, researchers discovered that market data does not arrive randomly. Instead, it arrives in tight clusters, where one event triggers a rapid succession of others, creating a "self-exciting" pattern. This discovery changes the math of speed. It turns out that when data arrives in these specific, clustered bursts, splitting the work across multiple processors can actually make the system faster and more reliable, provided the system is designed to handle the rhythm of the bursts correctly.

The researchers began by looking at the raw data stream as it travels from the exchange's matching engine, where orders are processed, to the traders' computers. They followed every single packet of data, noting exactly when it left the exchange and when it arrived. They found that the exchange's system acts like a gatekeeper with a fixed speed limit. Even when the matching engine processes orders incredibly fast—sometimes within a fraction of a microsecond of each other—the exchange's data publisher cannot send them all at once. It sends them one by one, with a minimum gap of about 7.5 microseconds between each packet. This creates a train of data packets that arrive at the trader's computer with a steady, rhythmic spacing, regardless of how chaotic the activity was at the source.

This rhythmic arrival is the key to the new findings. The researchers built a computer simulation to test how different software designs would handle this specific rhythm. They compared the traditional single-threaded approach, where one processor does all the work, against a multi-stage pipeline, where the work is split among several processors working in sequence. In their simulation, they fed the system the exact timing of the real data packets. The results were clear: for tasks that take longer than the 7.5-microsecond gap between packets, the single-threaded approach creates a massive backlog. When a burst of data arrives, the single processor gets overwhelmed, and the delay for the last few packets in the burst grows to be dozens of times longer than the task itself. This delay is the "tail" that traders fear, as it means their decisions are made too late.

In contrast, the multi-stage pipeline handled these bursts with ease. By splitting the work, the system could process the incoming train of packets in parallel. While the first processor was decoding the first packet, the second was already working on the second, and so on. This allowed the system to drain the backlog much faster, keeping the delay for every packet low and consistent. The simulation showed that for tasks taking 16 microseconds or longer, splitting the work reduced the worst-case delays by a factor of ten or more, with only a tiny penalty for the typical, non-bursty moments. The researchers confirmed that this improvement was not due to the sheer volume of data, but specifically to the clustered, bursty nature of the arrival times. When they simulated the same amount of data arriving randomly, the multi-stage system offered no advantage, and the single-threaded system remained efficient.

The study also ruled out several other potential causes for the delays. They found that the size of the data packets or the number of messages inside them was not the primary driver of the slowdown. Even when they rearranged the data to remove the bursts but kept the same number of packets, the massive delays disappeared. This proved that the problem was purely about the timing of the arrivals. The researchers also looked at the exchange itself to understand why the data arrived in these clusters. They found that the exchange's matching engine often processes multiple orders almost simultaneously, likely because many traders are reacting to the same market event at the same time. However, the exchange's publisher then spaces these out, creating the rhythmic train that the traders' systems must handle.

For the designers of these trading systems, the paper offers a clear, data-driven guide. If a system's processing time is shorter than the 7.5-microsecond gap between packets, the old rule still applies: keep it on a single thread. There is no benefit to splitting the work, and it only adds unnecessary complexity. But if the processing time is longer than that gap, the single-threaded approach will fail during bursts, and the system should be split into multiple stages. The researchers emphasize that the goal is not to use as many processors as possible, but to ensure that the slowest part of the process is fast enough to keep up with the exchange's rhythm. They also found that the specific arrangement of the processors matters less than ensuring that the slowest stage is handled efficiently.

This work does not claim to have solved every problem in high-speed trading, nor does it suggest that the single-threaded approach is obsolete. It simply provides a precise measurement of when that approach stops working and when a different design becomes necessary. By measuring the real world rather than relying on theoretical models, the researchers have given engineers a concrete threshold to measure against. They have shown that the nature of the data stream—specifically its tendency to arrive in self-exciting bursts—dictates the best way to build the software that consumes it. The lesson is that in the high-speed world of finance, understanding the rhythm of the data is just as important as the speed of the computer.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →