AI agents don't use the internet like people do — and their appetite for computing resources is beginning to show.
Autonomous AI workloads are consuming roughly five times as many LLM tokens as human-driven usage, according to a recent analysis from Andreessen Horowitz using OpenRouter data. Much of that consumption involves cached inputs, which can cost substantially less than repeatedly processing the same context from scratch.
The figures offer a glimpse of an internet increasingly used not only by people, but by software capable of continuously calling models, accessing services and completing tasks on its own.
That shift is also beginning to influence how internet infrastructure is built. In May 2025, Coinbase introduced x402, a payments protocol built around the long-dormant HTTP 402 “Payment Required” status code and designed in part to let AI agents pay for digital services automatically.
What the data says
Andreessen Horowitz based its analysis on data from OpenRouter, a service that routes LLM requests across hundreds of models.
The figures indicate that agent-driven workloads consume roughly five times as many tokens as human-driven usage. That difference makes sense given how agents operate: instead of sending a single prompt and waiting for an answer, autonomous systems can repeatedly call models as they plan, use tools, evaluate results, and continue working toward a goal.
However, the raw token volume does not translate directly into an equivalent increase in cost. According to the data cited by a16z, approximately 85% of the usage involves cached inputs.
For comparison, OpenAI's latest frontier model, GPT-5.6 Sol, currently costs $4.00 per million input tokens for a standard context window and only $0.40 for cached inputs. This 90% cost reduction makes high-frequency agentic usage more scalable.
The agent-usage trend is emerging alongside a broader shift in technology investment. A16z's analysis also points to growing investor interest in AI, nuclear energy, space, defense, and infrastructure, with AI accounting for a large share of current thematic investment.
Those capital flows reflect how closely the AI boom has become tied to the physical infrastructure needed to support growing compute demand.
Additionally, AI represents approximately 50% of the current thematic investment. This massive investment in AI can explain the massive data centre expansion and the consequent LLM usage growth.
Another interesting detail to note is that OpenRouter - the company whose data is used by a16z - was recently acquired by the payment giant Stripe for a reported $7.5 billion.
OpenRouter's growing strategic value also illustrates the expanding infrastructure layer around AI agents. Stripe reportedly acquired the model-routing company for $7.5 billion, adding a business that sits directly between applications, agents, and the models they call.
The enterprise takeaway
The a16z data suggests enterprises may need to rethink how they measure AI consumption as autonomous agents become more common.
A human employee may submit a handful of prompts during a task. An AI agent can make repeated model calls while planning, retrieving information, invoking tools, and evaluating its own progress. That means token consumption can increase rapidly even when the number of employees using AI remains unchanged.
Caching can reduce those costs, but organizations running agentic systems will still need visibility into token usage, model calls, and long-running workflows. Budget controls designed around individual chatbot sessions may not translate cleanly to autonomous software capable of operating continuously.
The bigger shift is that companies may increasingly be building for two kinds of users: people interacting directly with software and AI agents interacting with services on their behalf.
Other News: Google is reportedly considering a deal worth more than $1.5 billion for Mechanize technology and talent as it intensifies its push into autonomous AI coding.


