Achieving Single-Digit Microsecond Latency Inference for Capital Markets

Originally published at: Achieving Single-Digit Microsecond Latency Inference for Capital Markets | NVIDIA Technical Blog

In algorithmic trading, reducing response times to market events is crucial. To keep pace with high-speed electronic markets, latency-sensitive firms often use specialized hardware like FPGAs and ASICs. Yet, as markets grow more efficient, traders increasingly depend on advanced models such as deep neural networks to enhance profitability. Because implementing these complex models on low-level…

Worth noting that this benchmark has moved on since this was published.

In late April 2026, STAC published audited results for a stack featuring myrtle.ai’s VOLLO running on AMD Versal Premium silicon - with p99 latency of less than 2 microseconds for LSTM_A on the Tacana suite. That’s more than 2× faster than the GH200 figures referenced here.

The STAC article is available here: STAC-ML™ Pack for myrtle.ai VOLLO™ (Rev C) with Silicom Artena (AMD VP1802 FPGA) on Supermicro AS-2015CS-TNR | STAC - Insight for the Algorithmic Enterprise | STAC , or you can see it here on myrtle.ai’s website: Microsecond AI Inference for Trading

For anyone evaluating inference infrastructure for latency-sensitive trading workloads, the current record is worth factoring in.