In the six months following the February 6, 2026 baseline, AI agents overtook human users in compute consumption.
According to The Decoder, OpenRouter analyst Peter Walker disclosed data showing that on a seven-day average basis as of August 10, 2026, AI agent token consumption on the platform reached 7.3 trillion, up approximately 14x from 0.51 trillion on February 6. Over the same period, human user consumption rose from roughly 0.5 trillion to 1.4 trillion, a 2.8x increase. Total agent consumption is now more than five times that of humans.
Why 14x and Not 5x
Agent consumption is growing faster than its absolute share because of structural differences in per-request patterns. OpenRouter platform data shows that a typical agent task consumes roughly 13 to 15 times more tokens than a human conversation. Agent workloads must carry tool definitions, MCP gateway configurations, skill preamble instructions, and multi-turn reasoning context; a single autonomous programming or research task can consume as many tokens as hundreds of ordinary Q&A sessions.
The proliferation of reasoning models has amplified this effect. OpenRouter data points to an "overthinking" phenomenon among reasoning models: even for requests that require no deep reasoning, the model spends substantial tokens on internal deliberation before responding.
Caching Partially Dilutes Cost Inflation, but Structural Pressure Remains
Of the 7.3 trillion tokens generated by agents, approximately 70% to 85% come from cache hits. Cache hits are billed at discounted rates, so actual cost growth is lower than the raw token figures suggest. Agent workflows repeatedly reuse system prompts, tool definitions, and historical context, forming a naturally cache-friendly structure.
Even after discounts, consumption volume seven times that of humans remains a major variable in inference providers' revenue structure. As agent task complexity rises, net new output tokens that cannot be covered by caching will accumulate bills faster.
Chinese Models' Changing Share of the Agent Market
The OpenRouter official blog shows that DeepSeek's token share on the platform rose from 9% in January 2026 to 18% in June, and Chinese models' overall token consumption has surpassed that of US models. After DeepSeek V4 was released on April 24, 2026, its lightweight version, V4-Flash, accounted for 70% of DeepSeek's agent token traffic by the end of May.
DeepSeek V4 Flash is priced at $0.09 per million input tokens and $0.18 per million output tokens, while GPT-5.5 is priced at $5 input and $30 output—a roughly 55x price gap. Human users during the same period still use DeepSeek V3.2, while agent traffic has switched to V4-Flash, showing that agent and human preferences for models have diverged.
Infrastructure Business Models in Transition
OpenRouter co-founder and COO Chris Clark said in a SaaStr interview that the platform is expected to process approximately 28 trillion tokens per week, accounting for about 1% of global inference volume—more than Salesforce has processed over its entire lifetime.
Agent traffic brings more predictable batch loads rather than random human spikes, affecting compute scheduling and pricing strategies. Agents are less tolerant of service quality issues: a single tool-call parsing error can bring down an entire automation pipeline. Performance differences of the same open-source model across different providers come from middleware software, not model weights.
Enterprise Budgeting Models No Longer Apply
Blockonomi data shows that among enterprise users, the top 10% "frontier enterprises" generate 8.3 times more output tokens per active user than average enterprises, compared to just 2.6 times in January of this year. Weekly active users of agent tools in the legal industry have grown 108x since February of this year, with tasks including contract review, due diligence, and regulatory comparison—all long-context scenarios.
The Impact of This Shift
The February 6, 2026 baseline marks the inference market's core buyer shifting from humans to programs. OpenRouter's coverage is clearly biased toward open-source models, and no public data exists on closed-source API agent usage. When a single agent task costs 15x more than a human conversation, and Chinese models—leveraging a 55x price advantage—dominate agent traffic, the pricing logic, capacity planning, and competitive moats of the AI infrastructure industry must be recalculated.
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接