Quantitative tradingSparse attention
Hours of market context at tick-level speed
A quantitative trading firm wanted short-horizon equity price prediction models that could see far more market history than their transformer stack allowed.
The challenge
Tick-level signals depend on context that builds over hours: order flow accumulating through a session, shifts in liquidity, and volatility clustering. The firm's dense-attention models were capped at minutes of ticks. At a 64K-tick window, roughly a morning session, a naive four-head layer needs about 64 GB just to store its attention scores, far more than a 16 GB inference GPU holds. IO-aware kernels such as FlashAttention avoid storing the scores, but they still perform about 4.4 trillion operations per layer, which rules out real-time inference.
Our approach
Market data has strong temporal locality: a tick is informed most by the ticks just before it. We replaced dense attention with sliding-window sparse attention, where each tick attends to a window of w recent ticks (32–64). That cuts cost from O(n²) to O(n·w). Stacking layers widens the receptive field, so information from early in the session still reaches the prediction.
Dense attention O(n²)
Every tick attends to every earlier tick. Memory and compute grow with the square of the window.
Sliding-window sparse O(n·w)
Each tick attends to its w most recent neighbors. Cost grows linearly with the window.
Modeling
A microstructure-aware sparse transformer
- Causal sliding-window attention tuned for tick data, with dilated and learned sparsity patterns evaluated as alternatives
- Microstructure features as inputs: order book imbalance, VWAP, and realized volatility
- Targets for short-horizon order book imbalance and price movement
- Session-length windows that capture the regime of the day, not just the last few minutes
Engineering
A custom CUDA inference path
- Sparse attention kernels in C++/CUDA, exposed to PyTorch as extensions
- Fused QKV projection, attention, and feed-forward layers to cut kernel launches and memory traffic
- Feature engineering on the GPU, which removes CPU–GPU round trips from the hot path
- A circular-buffer KV cache, so each new tick updates the model incrementally instead of reprocessing the window
- End-to-end profiling with NVIDIA Nsight Systems and Nsight Compute
- 64.2 GB → 0.31 GB
- attention memory at a 64K-tick window
- 4,398 → 4.3 GFLOPs
- attention compute at the same window
- 16 GB GPU
- enough for a 64K-tick window (3–4 hours) on a single inference card
Theoretical estimates for one attention layer: fp32, batch 1, four heads, head dim 64, window 64; dense assumes the full score matrix is stored (16 GB per head). Our open-source reference implementation is described under Research, and the theory is covered in our sparse attention whitepaper.