AI & Models
Smaller models could disrupt the economics of AI labs
Brian Armstrong predicts 80% of AI workloads will shift to 99% cheaper models within 12-18 months, a trend that could disrupt the financial outlook for major AI labs.
The artificial intelligence industry is facing a potential shift in how enterprises deploy technology, moving away from a reliance on the largest available systems. Coinbase co-founder Brian Armstrong predicts that the vast majority of tasks will soon transition to significantly cheaper alternatives. According to Armstrong, “[D]emand for intelligence is near infinite, but 80% of workloads will be running on 99% cheaper models within 12-18 months.” Under this forecast, only 20% of workloads will still run on the latest generation of models where IQ maxing—defined as prioritizing the highest possible intelligence in a model—is important.
Early enterprise testing supports the feasibility of this transition. Harvey, a legal AI startup, recently tested cost reductions in partnership with the inference platform Fireworks AI. In the test, Harvey was able to reduce inference costs—the cost of running an AI model to generate output—by 3x without reducing quality. Gabe Pereyra, co-founder of Harvey, noted that while quality remains the primary focus in the legal sector, the definition of quality is shifting. Instead of defaulting to the most powerful model for every task, companies are looking for the most efficient model that can deliver the correct answer.
This shift toward smaller, more efficient models represents a departure from the scaling-first approach that has dominated the industry’s recent development. For years, major artificial intelligence labs have focused on training increasingly compute-intensive frontier models, which are the most advanced and compute-intensive models available. However, as enterprise users face rising cost pressures, the demand for these expensive options may decrease. The trend is often discussed in the context of competition between major labs and providers of open-weight models, such as DeepSeek, where the weights of the models are publicly available. But the primary shift is from large models to smaller, more efficient ones. If businesses successfully transition their workloads to these cheaper alternatives, it could deal a financial blow to major labs like OpenAI and Anthropic, which have relied on the assumption that the most powerful and expensive models would consistently command the market.
Why it matters
The move toward smaller, cheaper models challenges the assumption that the most powerful models will always dominate the market. This shift could force a major economic reset for frontier model labs that have heavily invested in training massive, compute-intensive systems.