On Monday, Decagon CEO Jesse Zhang published a provocative new theory, posted under the title “Everyone is wrong about open source AI in the enterprise.” His essay grapples with one of the most interesting contradictions of today’s rapidly evolving AI economy: While more mature AI deployments are increasingly switching to lighter, often open-source models, the overall enterprise spend on expensive state-of-the-art (SOTA) frontier models has barely budged. This paradox, Zhang suggests, paints a nuanced picture of the relationship between cutting-edge research and practical application.
Key Takeaways
- **The Two-Phase AI Lifecycle:** Frontier models are primarily for “discovery” – proving out new use cases – while open-source models increasingly handle “production” – running mature, cost-optimized applications.
- **Paradox of Spend vs. Volume:** Despite a clear shift in token volume towards cheaper open-source models, enterprise spending on premium frontier models remains stubbornly high, driven by the expanding scope of AI applications and strategic pricing.
- **A Stable, Two-Tiered Economy:** The AI market is evolving into a stable, two-tiered structure where both frontier labs and open-source projects thrive by serving distinct, yet interconnected, stages of the AI adoption journey.
Zhang’s theory challenges the common perception that open-source models are direct competitors to their frontier counterparts, poised to usurp their market share. Instead, he argues they represent two distinct, yet complementary, phases of the same AI adoption lifecycle. In this model, expensive frontier models initially serve to prove out complex use cases and establish new capabilities. Once these applications mature and their value is validated, enterprises often transition to more cost-effective, adaptable open-source alternatives for scaled production.
The AI Lifecycle: From Discovery to Production
The premise is straightforward: when an enterprise embarks on a new AI initiative, particularly one pushing the boundaries of what’s possible, they naturally gravitate towards the most capable models available. These are typically the large, proprietary models developed by frontier labs like Anthropic, OpenAI, or Google. Their superior performance, broader general knowledge, and often easier integration via APIs make them ideal for rapid prototyping, concept validation, and tackling novel, complex problems where accuracy and breadth of understanding are paramount. This initial phase is about “discovery” – exploring potential, understanding limitations, and demonstrating value.
However, as an AI application moves from experimental validation to full-scale deployment, the economic calculus shifts. Operational costs become a significant factor. This is where open-source models shine. Once a use case is proven and understood, enterprises can often port the logic to a smaller, fine-tuned, or more specialized open-source model. These models, while perhaps not as universally capable as their frontier cousins, can be significantly cheaper to run, offer greater customization, and provide more control over data and deployment environments. This transition marks the “production” phase – optimizing for efficiency, scalability, and cost-effectiveness.
The Data-Driven Paradox: Volume Shifts, Spend Endures
Zhang doesn’t provide extensive data in his post, but the evidence supporting this intriguing dynamic is readily available across the industry. Examining platforms that aggregate AI model usage reveals a fascinating split between token volume and overall expenditure.
Consider Vercel’s AI gateway dashboard, a valuable window into real-world enterprise AI traffic. In recent weeks, we’ve seen a significant surge in token volumes for models like DeepSeek, which has climbed to lead, processing over a third of the tokens flowing through Vercel’s infrastructure. Z.ai, the lab behind the popular GLM-5.2 model, has also made impressive gains, securing a respectable fourth place. These shifts clearly indicate a growing adoption of lighter, often open-source-leaning, models for routine tasks and scaled applications.
Yet, when we scroll down to analyze overall token spend, the picture changes dramatically. Anthropic, despite a relative dip in its share of raw token volume, still accounts for more than half of the total AI expenditure on the platform. While Anthropic’s own rising prices have contributed to maintaining this share over the past month, the underlying trend is undeniable: enterprises are willing to pay a premium for the capabilities that frontier models offer, even as they diversify their model usage.
OpenRouter, another prominent platform capturing a significant segment of the AI market (albeit with a slightly less enterprise-focused user base), tells a similar story. DeepSeek V4 Flash is a clear winner in terms of raw usage, processing an astounding 5.3 trillion tokens weekly. In contrast, Opus 4.8, a leading frontier model, handles just over 2 trillion tokens. The sheer volume difference is striking. However, OpenRouter also registers the average token cost for Opus 4.8 as approximately 23 times higher than that of V4 Flash ($1.37 per million tokens compared to a mere 6 cents). This enormous price differential strongly suggests that, despite the lower token volume, Opus 4.8 is likely still capturing a disproportionately large share of total spending on the platform.
And the landscape continues to evolve. We haven’t even fully accounted for newer arrivals like Nvidia’s Nemotron, which is poised to make a significant impact. With Nvidia’s deep connections in the enterprise and data center space, coupled with Nemotron’s inherent adaptability, it’s expected to quickly capture a substantial share of future frontier model deployments, further reinforcing the continued demand for high-end, proprietary solutions.
Why Frontier Labs Persist: Deconstructing the Resilience
These figures, while not providing absolute proof of Zhang’s explicit AI lifecycle model, certainly underscore his core observation: frontier labs like Anthropic are not suffering unduly from the rise of open-source models – at least not yet. There are several compelling explanations for this resilience:
One primary reason is the explosive growth of the overall market for AI-addressable tasks. The landscape of problems that AI can solve is expanding so rapidly that frontier models can maintain their premium position simply by dominating the early-stage deployments for these newly identified, often complex, use cases. As Zhang succinctly puts it, “The frontier labs will keep owning discovery. Open source will increasingly own production.” This implies a continuous pipeline of new, challenging problems where the cutting edge is always in demand.
Another crucial factor is the inherent difficulty of certain use cases. Even as clients transition to open-source models for many applications, there remain specific tasks that are so intricate, require such extensive knowledge, or demand such nuanced understanding that they cannot be entirely replicated or replaced by cheaper, less sophisticated alternatives. For these “hard problems,” the superior performance and broad capabilities of frontier models justify the higher cost.
Furthermore, the “discovery” phase isn’t just about initial prototyping. It encompasses continuous innovation, rapid iteration, and the ability to benchmark against the best. Frontier models often come with robust API documentation, dedicated support, and faster development cycles for new features, which can significantly accelerate an enterprise’s ability to explore and implement novel AI functionalities. For enterprises focused on competitive advantage and innovation speed, these benefits often outweigh the higher per-token cost.
Conversely, the “production” phase isn’t merely about cost-cutting. It’s about control, customization, and long-term operational stability. Open-source models allow companies to fine-tune models on proprietary data, ensuring better performance for specific domain tasks and addressing data privacy or security concerns that might arise from sending sensitive information to third-party APIs. This strategic control becomes paramount as AI systems become embedded in critical business processes.
Rethinking the “Coffee Beans” Analogy
As recently as last September, I contemplated the possibility that foundation labs would effectively become “coffee bean” suppliers to Starbucks – that is, serving as commodity inputs while the application layer reaped the bulk of the profits. Some aspects of that prediction have indeed materialized: we’ve seen vertical AI plays successfully switch to lighter, specialized models, and the economics of many “GPT wrapper” startups have remained relatively stable, leveraging the underlying models as a service.
However, what we’re also witnessing is that, on a token-for-token basis, frontier providers have been remarkably effective at holding onto the most desirable part of the marketplace – the premium token price. This isn’t just about superior technology; it’s about strategic positioning in a rapidly expanding market, where the cutting edge is constantly redefined, and new, complex problems always emerge that only the most advanced models can tackle effectively. This ability to command premium pricing for “discovery” and high-stakes “hard problems” doesn’t seem likely to change anytime soon.
This two-tiered economy of models, where frontier labs drive innovation and open-source models drive optimized production, appears to be shaping up as a relatively stable and enduring feature of the AI economy. It suggests a collaborative, rather than purely competitive, dynamic between these two powerful forces in the AI landscape.
The Bottom Line
Jesse Zhang’s theory, strongly supported by recent market data, recalibrates our understanding of the AI ecosystem. Rather than a zero-sum game, the relationship between frontier and open-source models is a symbiotic one, reflecting a natural evolution from initial discovery and validation to scaled production and optimization. Enterprises will continue to leverage the cutting-edge capabilities of premium frontier models to explore new possibilities and solve the toughest problems, while simultaneously optimizing costs and gaining control with open-source alternatives for mature applications. This dynamic ensures that innovation continues at both ends of the spectrum, fostering a robust and diversified AI market for the foreseeable future.
When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.
Source: {feed_title}

