AI Agents Use 5 Times More LLM Tokens Than Humans, a16z Data Shows

AI letters.

AI agents are consuming more LLM tokens than human users as autonomous workloads reshape how the web handles AI traffic, costs, and infrastructure. Image: Igor Omilaev/Unsplash

Written By
Eric Mboizi
Eric Mboizi
Aug 26, 2026
3 minute read
eWeek content and product recommendations are editorially independent. We may make money when you click on links to our partners. Learn More

AI agents don't use the internet like people do — and their appetite for computing resources is beginning to show.

Autonomous AI workloads are consuming roughly five times as many LLM tokens as human-driven usage, according to a recent analysis from Andreessen Horowitz using OpenRouter data. Much of that consumption involves cached inputs, which can cost substantially less than repeatedly processing the same context from scratch.

The figures offer a glimpse of an internet increasingly used not only by people, but by software capable of continuously calling models, accessing services and completing tasks on its own.

That shift is also beginning to influence how internet infrastructure is built. In May 2025, Coinbase introduced x402, a payments protocol built around the long-dormant HTTP 402 “Payment Required” status code and designed in part to let AI agents pay for digital services automatically.

What the data says

Andreessen Horowitz based its analysis on data from OpenRouter, a service that routes LLM requests across hundreds of models.

The figures indicate that agent-driven workloads consume roughly five times as many tokens as human-driven usage. That difference makes sense given how agents operate: instead of sending a single prompt and waiting for an answer, autonomous systems can repeatedly call models as they plan, use tools, evaluate results, and continue working toward a goal.

However, the raw token volume does not translate directly into an equivalent increase in cost. According to the data cited by a16z, approximately 85% of the usage involves cached inputs.

For comparison, OpenAI's latest frontier model, GPT-5.6 Sol, currently costs $4.00 per million input tokens for a standard context window and only $0.40 for cached inputs. This 90% cost reduction makes high-frequency agentic usage more scalable.

The agent-usage trend is emerging alongside a broader shift in technology investment. A16z's analysis also points to growing investor interest in AI, nuclear energy, space, defense, and infrastructure, with AI accounting for a large share of current thematic investment. 

Advertisement

Those capital flows reflect how closely the AI boom has become tied to the physical infrastructure needed to support growing compute demand.

Additionally, AI represents approximately 50% of the current thematic investment. This massive investment in AI can explain the massive data centre expansion and the consequent LLM usage growth.

Another interesting detail to note is that OpenRouter - the company whose data is used by a16z - was recently acquired by the payment giant Stripe for a reported $7.5 billion. 

OpenRouter's growing strategic value also illustrates the expanding infrastructure layer around AI agents. Stripe reportedly acquired the model-routing company for $7.5 billion, adding a business that sits directly between applications, agents, and the models they call.

The enterprise takeaway

The a16z data suggests enterprises may need to rethink how they measure AI consumption as autonomous agents become more common.

A human employee may submit a handful of prompts during a task. An AI agent can make repeated model calls while planning, retrieving information, invoking tools, and evaluating its own progress. That means token consumption can increase rapidly even when the number of employees using AI remains unchanged.

Caching can reduce those costs, but organizations running agentic systems will still need visibility into token usage, model calls, and long-running workflows. Budget controls designed around individual chatbot sessions may not translate cleanly to autonomous software capable of operating continuously.

The bigger shift is that companies may increasingly be building for two kinds of users: people interacting directly with software and AI agents interacting with services on their behalf.

Advertisement

Other News: Google is reportedly considering a deal worth more than $1.5 billion for Mechanize technology and talent as it intensifies its push into autonomous AI coding. 

Eric Mboizi

Eric Mboizi is a technology news writer covering software development, emerging technologies, and the evolving digital landscape for TechRepublic and eWeek. He holds a bachelor’s degree in software engineering from Makerere University and has more than five years of experience creating technical content for developers and technology professionals. In addition to his work as a journalist, Eric is an Ethereum developer with more than four years of experience in blockchain technology. His hands-on development background gives him a practical perspective on software engineering, decentralized technologies, and the real-world implications of new technology trends.

eWeek Logo

eWeek has the latest technology news and analysis, buying guides, and product reviews for IT professionals and technology buyers. The site's focus is on innovative solutions and covering in-depth technical content. eWeek stays on the cutting edge of technology news and IT trends through interviews and expert analysis. Gain insight from top innovators and thought leaders in the fields of IT, business, enterprise software, startups, and more.

Property of TechnologyAdvice. © 2026 TechnologyAdvice. All Rights Reserved

Advertiser Disclosure: Some of the products that appear on this site are from companies from which TechnologyAdvice receives compensation. This compensation may impact how and where products appear on this site including, for example, the order in which they appear. TechnologyAdvice does not include all companies or all types of products available in the marketplace.