SubQ 1.1 Small: Sparse attention cuts long-context compute by 64.5x at 1M tokens
Subquadratic – Introducing SubQ 1.1 Small
Subquadratic released the model card for SubQ 1.1 Small. It replaces quadratic dense attention with Subquadratic Sparse Attention (SSA) that routes based on content relevance, scaling linearly with context length. At 1M tokens, SSA uses 64.5x less compute than dense attention and runs 56x faster than FlashAttention-2. The model scores near-perfect on needle-in-a-haystack from 1M to 12M tokens and 99.12% on RULER at 128K. General reasoning holds: GPQA Diamond 85.4%, LiveCodeBench pass@4 89.7%, AutomationBench Finance 13%. Training started from an open-weight frontier model, replaced attention with SSA, then ran staged context extension up to 2M and ~1T tokens of continued pretraining on books, documents, and repo-scale code. The post does not name the base model. SubQ 1.1 Small is deploying with select design partners; a broader lineup from 2M to 12M tokens is planned later this year.
Why it matters: SubQ 1.1 Small ships a deployable sparse-attention model with a 64.5x compute reduction and near-perfect 12M-token retrieval. Held below 85 because it's still a model card + design-partner deployment — no open weights or public API yet, so the production story is incomplete.