Home / Case studies

Selected client work

Two engagements, one theme: getting state-of-the-art models to work on data they were never built for.

Quantitative tradingSparse attention

Hours of market context at tick-level speed

A quantitative trading firm wanted short-horizon equity price prediction models that could see far more market history than their transformer stack allowed.

The challenge

Tick-level signals depend on context that builds over hours: order flow accumulating through a session, shifts in liquidity, and volatility clustering. The firm's dense-attention models were capped at minutes of ticks. At a 64K-tick window, roughly a morning session, a naive four-head layer needs about 64 GB just to store its attention scores, far more than a 16 GB inference GPU holds. IO-aware kernels such as FlashAttention avoid storing the scores, but they still perform about 4.4 trillion operations per layer, which rules out real-time inference.

Our approach

Market data has strong temporal locality: a tick is informed most by the ticks just before it. We replaced dense attention with sliding-window sparse attention, where each tick attends to a window of w recent ticks (32–64). That cuts cost from O(n²) to O(n·w). Stacking layers widens the receptive field, so information from early in the session still reaches the prediction.

Dense attention O(n²)

Every tick attends to every earlier tick. Memory and compute grow with the square of the window.

Sliding-window sparse O(n·w)

Each tick attends to its w most recent neighbors. Cost grows linearly with the window.

Rows are ticks being predicted; columns are earlier ticks they can attend to.

Modeling

A microstructure-aware sparse transformer

  • Causal sliding-window attention tuned for tick data, with dilated and learned sparsity patterns evaluated as alternatives
  • Microstructure features as inputs: order book imbalance, VWAP, and realized volatility
  • Targets for short-horizon order book imbalance and price movement
  • Session-length windows that capture the regime of the day, not just the last few minutes

Engineering

A custom CUDA inference path

  • Sparse attention kernels in C++/CUDA, exposed to PyTorch as extensions
  • Fused QKV projection, attention, and feed-forward layers to cut kernel launches and memory traffic
  • Feature engineering on the GPU, which removes CPU–GPU round trips from the hot path
  • A circular-buffer KV cache, so each new tick updates the model incrementally instead of reprocessing the window
  • End-to-end profiling with NVIDIA Nsight Systems and Nsight Compute
64.2 GB → 0.31 GB
attention memory at a 64K-tick window
4,398 → 4.3 GFLOPs
attention compute at the same window
16 GB GPU
enough for a 64K-tick window (3–4 hours) on a single inference card

Theoretical estimates for one attention layer: fp32, batch 1, four heads, head dim 64, window 64; dense assumes the full score matrix is stored (16 GB per head). Our open-source reference implementation is described under Research, and the theory is covered in our sparse attention whitepaper.

EngineeringVision-language models

Teaching vision-language models to read engineering drawings

An engineering client asked us to adapt open-weight vision-language models (VLMs) so they could use engineering drawings as a primary source of information.

The challenge

The client's drawings arrive as PDFs in two very different forms: vector files exported from CAD, and raster scans of legacy and marked-up sheets. General-purpose VLMs can read the text on a drawing, but they miss what an engineer sees. They lose track of which dimension belongs to which feature, read past tolerances and GD&T frames, and do not connect a section view back to the part it cuts through.

Our approach

We borrowed a pattern from electronic design automation (EDA). EDA tools take raw layout geometry and step it up through extracted connectivity, hierarchy, and finally function. We applied the same layered approach to drawings: each level of understanding is trained on top of the one below, so the model infers higher levels of abstraction from grounded evidence rather than guessing them.

Engineering drawing layer EDA analogue
  1. L4
    Design intentPart and assembly function, manufacturing and inspection meaning
    Functional specification
  2. L3
    Views & featuresOrthographic, section, and detail views linked into holes, threads, and other features
    Schematic & hierarchy
  3. L2
    AnnotationsDimensions, tolerances, GD&T, notes, title blocks, and tables bound to geometry
    Netlist extraction
  4. L1
    PrimitivesLines, arcs, text, and symbols from vector paths or raster pixels
    Layout geometry

Phase 1 · Fine-tuning

Adapt an open-weight VLM with PEFT

  • Built a drawing corpus that covers both vector and raster PDFs, with task data defined for each layer of the ladder
  • Used vector PDFs, where geometry and text are exact, to ground labels for the harder raster cases
  • Trained LoRA adapters with SFT for layer-by-layer extraction into structured outputs
  • Applied DPO on engineer-reviewed preference pairs to penalize invented dimensions and mis-attributed annotations
  • Delivered an evaluation suite for each layer, so progress could be measured rather than eyeballed

Phase 2 · Pretraining

Extend the base model on drawings

  • Continued pretraining of the VLM on a large corpus of drawings, so drafting conventions and symbols become native to the model instead of learned per task
  • Mixed drawing data with general data to keep broad visual and language skills intact
  • Re-applied the Phase 1 fine-tuning recipes on the stronger base, pushing reliable inference further up the ladder toward design intent
  • Kept data and weights entirely on the client's infrastructure using open-weight models

Why phase it this way? Phase 1 delivers a working model quickly and produces the labeled data and evaluations that Phase 2 depends on. It also shows exactly where fine-tuning stops being enough, which makes the case for pretraining with evidence instead of assumption.

Contact

Have a model that needs to be faster?

We are taking on new consulting engagements. Tell us what you are building and where it is stuck, and we will reply within two business days.

contact@deepsolute.com
Deepsolute LLC 1530 PB Ln F4201 Wichita Falls, TX 76302 Timothy Fox on LinkedIn →