From our writing
The “Quadratic Tax”: why financial foundation models struggle to scale
The NYSE now processes more than 1.2 trillion messages a day, and catching a regime shift or volatility spike can mean looking back across 64,000+ ticks. Dense attention charges a quadratic tax on that context: double the sequence and cost quadruples.
The article makes the case for sub-quadratic attention built on hardware-aware sparsity. It covers sliding windows for local microstructure and dilated patterns for longer trends, implemented as fused, memory-coalesced CUDA kernels so near-linear scaling shows up as real throughput on NVIDIA Hopper and Blackwell GPUs.
The goal isn’t just to make models “faster.” The goal is to make them economically scalable.