“Token Prices Are Falling, but Bills Are Rising”: As AI Advances, Token Management Becomes a Competitive Edge
Authored On
Modified
LLM token prices fall, yet overall costs continue to rise Companies that once pushed for wider AI use shift toward tighter controls Spread of AI agents makes cost management increasingly urgent

The price of tokens used to run artificial intelligence (AI) models has fallen to a record low. An aggressive pricing push by Chinese AI companies and intensifying competition among global technology giants have rapidly driven down per-unit charges. Yet the proliferation of high-performance models and AI agents has sharply increased both the total number of tokens consumed and the frequency of model calls, raising the likelihood that companies’ actual AI operating costs will climb. The ability to manage spending by improving token efficiency, rather than simply expanding usage, is emerging as a critical factor in competitive AI adoption.
AI Token Costs Trend Downward
According to market research firm SiliconData on the 28th, its large language model (LLM) token spending index fell to $0.97 on the 31st of last month (local time). That was its lowest level since the index was introduced late last year and less than half its peak this summer. Rather than tracking an individual provider’s list price, the LLM token spending index measures the volume-weighted average price per million tokens actually consumed in the market. A token is a basic unit into which AI breaks down information such as a user’s question, text or image for processing; consumption generally rises as AI services see heavier use and tasks become more complex.
Chinese AI companies’ low-price push has been a major driver of the decline. As Chinese open-source models, including Moonshot AI’s Kimi K3, have expanded their global presence by offering competitive performance at low prices, leading U.S. AI companies have responded with successive price cuts. OpenAI, for example, cut application programming interface (API) prices for GPT-5.6 Terra and Luna by 20% and 80%, respectively, in July, then reduced Sol’s price by 20% in August. Last month, Google introduced Gemini 3.7 Flash at roughly half the price of its predecessor, charging $0.75 per million input tokens and $3.75 per million output tokens. The Economy has also examined how model pricing and token efficiency are reshaping competition among AI providers.
Token Consumption Rises Sharply
Lower prices, however, have not translated directly into lower AI costs. Most recently released high-performance models undertake longer reasoning processes to solve difficult problems, while autonomous AI agents repeatedly call models and verify results as they carry out a single task. If token consumption grows faster than token prices fall, the overall cost burden can increase. Anthropic’s latest model, Claude Fable 5.1, for example, retained the same input and output prices as its predecessor, Fable 5—$10 and $50 per million tokens, respectively—while cutting the cache-read price by 75%, from $1 to $0.25. Yet in the Artificial Analysis Intelligence Index (AAII), Fable 5.1 cost $3.69 per task at maximum effort, 18% more than Fable 5’s $3.14. That was because Fable 5.1 used approximately 1.7 times as many output tokens as its predecessor to achieve higher performance.
This growth in token-related spending appears likely to accelerate. On the 16th, Huawei released its research report, “Intelligent World 2035: Turning Visions into Action,” forecasting that annual global token consumption in 2035 would be 100,000 times its current level. The assessment suggests that an internet centered on AI agents could increase computing and network demand in ways distinct from today’s application-centered internet. Market research firm Gartner likewise projects that AI agents will drive inference costs per workflow up by more than fivefold by 2028. The Swiss Institute of Artificial Intelligence has examined the same divergence between falling token prices and rising total bills.
Companies’ Initial AI Adoption Strategies
Companies are steadily revising their AI operating strategies in response to changes in token cost structures. In the early days of generative AI, they focused more on increasing usage than on cost efficiency. Because they still needed to establish how useful AI would be in day-to-day work, encouraging employees to use it and making AI use a routine organizational practice took priority. One market expert observed, “Before generative AI became an established production tool, companies had little need to scrutinize token costs or model-specific prices. Their main goal was to drive ‘innovation’ by expanding AI use cases across the organization as quickly as possible and encouraging repeated use.”
Some companies, including Amazon, Microsoft, Shopify and Accenture, previously encouraged active AI use by incorporating employees’ adoption into work metrics and performance reviews or by operating usage leaderboards and incentive programs. Spanish banking giant BBVA provided ChatGPT Enterprise licenses to 3,000 employees and encouraged them to create a range of customized GPTs. Optum, the healthcare services and healthcare IT company owned by UnitedHealth Group, one of the largest U.S. health insurers, also directly managed AI usage frequency by tracking whether some employees used ChatGPT or Microsoft (MS) Copilot at least once a day.
Table 1. Changes in Corporate AI Operating Strategies
| Category | Early AI adoption | Recent approach |
|---|---|---|
| Operating objective | Increase employee AI use and identify workplace applications | Optimize token costs against productivity gains |
| Usage management | Track usage frequency and operate leaderboards and incentives | Set per-employee usage limits and curb unnecessary or duplicate calls |
| Model selection | Favor high-specification models on the basis of performance | Select lower-cost or high-performance models according to task difficulty |
| Cost control | Commit substantial budgets to establish regular AI use | Manage efficiency by linking work outcomes, prompts and token costs |
| System design | Allow employees to choose models and methods of use directly | Automatically assign an appropriate model according to task complexity |
Cost-Cutting Strategies Gain Traction
Recently, however, companies have begun seeking more efficient returns on their AI spending instead of indiscriminately increasing token consumption. After Uber exhausted its AI coding budget for the year in just four months, it capped employees’ AI model spending at $1,500 per month each. Walmart instructed employees to use relatively inexpensive models for simple tasks and introduced a system to reduce costs arising when multiple employees submit the same question or task. Amazon discontinued the employee usage leaderboard it had operated to encourage AI adoption. The move was intended to prevent employees from using AI for unnecessary tasks merely to compete in the rankings.
U.S. human resources software company Rippling has also built an “AI Spend Console” that compares work output, prompt counts and token costs, while introducing an AI gateway that automatically selects a model in an appropriate price range according to task difficulty. U.S. data and AI platform company Databricks adopted an approach that changes the model itself to suit the complexity of the work, rather than merely setting usage limits for individual employees. After internal tests found that Chinese company Z.ai’s open-weight GLM 5.2 model delivered task performance statistically similar to Anthropic’s high-performance Claude Opus 4.8, Databricks expanded use of the GLM model family in developers’ work and introduced a “smart routing” system that assesses task complexity and automatically selects a suitable model. In Databricks’ tests, GLM 5.2 cost $1.28 per task, approximately 34% less than Opus 4.8’s $1.94. Research from the Swiss Institute of Artificial Intelligence likewise identifies model routing and cost per completed task as central to managing AI consumption.
Token-Efficiency Technologies Gain Prominence
The ability to use tokens efficiently is expected to become increasingly important as the spread of AI agents makes cost management more urgent. Among the leading ways to reduce token consumption today are minimizing model calls and repeated computation through prompt and context caching. Instead of processing recurring inputs—such as system prompts, long documents and conversation histories—from scratch each time, these methods reuse previously computed results. Another approach eliminates the model call altogether. AI gateway providers use methods including in-memory caches that immediately return a previous response to an identical request, Redis caches shared across multiple servers, and semantic caches that identify questions with similar meanings despite different wording and reuse existing answers. Because these approaches do not call the model API, the requests incur no token charges. The Economy has detailed the role of caching and model routing in controlling API costs.
Prompt and context compression are also major cost-saving measures. They reduce the number of input tokens by summarizing long documents or conversation histories, removing less important sentences and tokens, and using retrieval-augmented generation (RAG) to select the passages most relevant to a query. A prominent implementation is Microsoft Research’s LLMLingua. It uses a small language model to assess the importance of input tokens and remove unnecessary material, compressing prompts by up to 20 times while limiting performance loss to approximately 1.5 percentage points.