black and white lenovo laptop

For years, the AI world has been locked in a love-hate relationship with the Transformer architecture. We love the intelligence, but we hate the 'quadratic bottleneck.' In simple terms, if you double the length of your prompt, the computational cost doesn't just double—it quadruples. This is why processing massive documents or entire codebases has remained prohibitively expensive. But a Miami-based startup called Subquadratic is looking to change the math entirely with the release of SubQ 1.1 Small.

Linear Scaling in a Quadratic World

The breakthrough behind SubQ 1.1 Small lies in its fully subquadratic architecture. Unlike standard models that waste compute by analyzing every possible relationship between words, SubQ uses a sparse-attention mechanism. It intelligently selects a small subset of relevant tokens to attend to based on content rather than fixed patterns.

The result? Scaling that is $O(n)$ or $O(n \log n)$ instead of $O(n^2)$. In plain English, doubling your input now only doubles your cost. While the research version of the model has successfully handled a staggering 12 million tokens—roughly 9 million words—the production API is currently optimized for a 1-million-token context window.

Why This Changes the Game for Agents and RAG

Until now, developers have relied on Retrieval-Augmented Generation (RAG) to handle large datasets. RAG acts like a search engine, pulling small chunks of data to feed into a model because the model itself couldn't handle the full corpus. SubQ 1.1 Small suggests a future where these complex chunking strategies and retrieval pipelines become obsolete.

By allowing for a massive, native context window at 1,000x lower cost than traditional Transformers, SubQ is tailor-made for reasoning-heavy agent workflows. Whether it's analyzing a 500-page legal contract or navigating a massive software repository, the model processes the entire context in a single pass. This reduces the risk of 'forgetting' or missing critical information hidden in the middle of a long document.

The Road Ahead

Of course, the AI community remains cautiously optimistic. While previous attempts at subquadratic models like Mamba or Hyena showed promise, they often struggled to match Transformer-level quality at scale. Subquadratic claims they’ve broken that pattern, but researchers are already calling for independent benchmarks to verify these massive efficiency gains. If the claims hold up, SubQ 1.1 Small won't just be a new model—it will be the blueprint for the next era of sustainable, long-context AI.

Sources

Media