Skip to main content
  • Home
  • Tech
  • “Token Prices Are Falling, but Bills Are Rising”: As AI Advances, Token Management Becomes a Competitive Edge

“Token Prices Are Falling, but Bills Are Rising”: As AI Advances, Token Management Becomes a Competitive Edge

Picture

Member for

1 year 1 month
Real name
Oliver Griffin
Bio
[email protected]

Oliver Griffin is a policy and tech reporter at The Economy, focusing on the intersection of artificial intelligence, government regulation, and macroeconomic strategy. Based in Dublin, Oliver has reported extensively on European Union policy shifts and their ripple effects across global markets. Prior to joining The Economy, he covered technology policy for an international think tank, producing research cited by major institutions, including the OECD and IMF. Oliver studied political economy at Trinity College Dublin and later completed a master’s in data journalism at Columbia University. His reporting blends field interviews with rigorous statistical analysis, offering readers a nuanced understanding of how policy decisions shape industries and everyday lives. Beyond his newsroom work, Oliver contributes op-eds on ethics in AI and has been a guest commentator on BBC World and CNBC Europe.

Modified

LLM token prices fall, yet overall costs continue to rise
Companies that once pushed for wider AI use shift toward tighter controls
Spread of AI agents makes cost management increasingly urgent

The price of tokens used to run artificial intelligence (AI) models has fallen to a record low. An aggressive pricing push by Chinese AI companies and intensifying competition among global technology giants have rapidly driven down per-unit charges. Yet the proliferation of high-performance models and AI agents has sharply increased both the total number of tokens consumed and the frequency of model calls, raising the likelihood that companies’ actual AI operating costs will climb. The ability to manage spending by improving token efficiency, rather than simply expanding usage, is emerging as a critical factor in competitive AI adoption.

AI Token Costs Trend Downward

According to market research firm SiliconData on the 28th, its large language model (LLM) token spending index fell to $0.97 on the 31st of last month (local time). That was its lowest level since the index was introduced late last year and less than half its peak this summer. Rather than tracking an individual provider’s list price, the LLM token spending index measures the volume-weighted average price per million tokens actually consumed in the market. A token is a basic unit into which AI breaks down information such as a user’s question, text or image for processing; consumption generally rises as AI services see heavier use and tasks become more complex.

Chinese AI companies’ low-price push has been a major driver of the decline. As Chinese open-source models, including Moonshot AI’s Kimi K3, have expanded their global presence by offering competitive performance at low prices, leading U.S. AI companies have responded with successive price cuts. OpenAI, for example, cut application programming interface (API) prices for GPT-5.6 Terra and Luna by 20% and 80%, respectively, in July, then reduced Sol’s price by 20% in August. Last month, Google introduced Gemini 3.7 Flash at roughly half the price of its predecessor, charging $0.75 per million input tokens and $3.75 per million output tokens. The Economy has also examined how model pricing and token efficiency are reshaping competition among AI providers.

Token Consumption Rises Sharply

Lower prices, however, have not translated directly into lower AI costs. Most recently released high-performance models undertake longer reasoning processes to solve difficult problems, while autonomous AI agents repeatedly call models and verify results as they carry out a single task. If token consumption grows faster than token prices fall, the overall cost burden can increase. Anthropic’s latest model, Claude Fable 5.1, for example, retained the same input and output prices as its predecessor, Fable 5—$10 and $50 per million tokens, respectively—while cutting the cache-read price by 75%, from $1 to $0.25. Yet in the Artificial Analysis Intelligence Index (AAII), Fable 5.1 cost $3.69 per task at maximum effort, 18% more than Fable 5’s $3.14. That was because Fable 5.1 used approximately 1.7 times as many output tokens as its predecessor to achieve higher performance.

This growth in token-related spending appears likely to accelerate. On the 16th, Huawei released its research report, “Intelligent World 2035: Turning Visions into Action,” forecasting that annual global token consumption in 2035 would be 100,000 times its current level. The assessment suggests that an internet centered on AI agents could increase computing and network demand in ways distinct from today’s application-centered internet. Market research firm Gartner likewise projects that AI agents will drive inference costs per workflow up by more than fivefold by 2028. The Swiss Institute of Artificial Intelligence has examined the same divergence between falling token prices and rising total bills.

Companies’ Initial AI Adoption Strategies

Companies are steadily revising their AI operating strategies in response to changes in token cost structures. In the early days of generative AI, they focused more on increasing usage than on cost efficiency. Because they still needed to establish how useful AI would be in day-to-day work, encouraging employees to use it and making AI use a routine organizational practice took priority. One market expert observed, “Before generative AI became an established production tool, companies had little need to scrutinize token costs or model-specific prices. Their main goal was to drive ‘innovation’ by expanding AI use cases across the organization as quickly as possible and encouraging repeated use.”

Some companies, including Amazon, Microsoft, Shopify and Accenture, previously encouraged active AI use by incorporating employees’ adoption into work metrics and performance reviews or by operating usage leaderboards and incentive programs. Spanish banking giant BBVA provided ChatGPT Enterprise licenses to 3,000 employees and encouraged them to create a range of customized GPTs. Optum, the healthcare services and healthcare IT company owned by UnitedHealth Group, one of the largest U.S. health insurers, also directly managed AI usage frequency by tracking whether some employees used ChatGPT or Microsoft (MS) Copilot at least once a day.

Table 1. Changes in Corporate AI Operating Strategies

CategoryEarly AI adoptionRecent approach
Operating objectiveIncrease employee AI use and identify workplace applicationsOptimize token costs against productivity gains
Usage managementTrack usage frequency and operate leaderboards and incentivesSet per-employee usage limits and curb unnecessary or duplicate calls
Model selectionFavor high-specification models on the basis of performanceSelect lower-cost or high-performance models according to task difficulty
Cost controlCommit substantial budgets to establish regular AI useManage efficiency by linking work outcomes, prompts and token costs
System designAllow employees to choose models and methods of use directlyAutomatically assign an appropriate model according to task complexity
Source: Compilation of international media reports

Cost-Cutting Strategies Gain Traction

Recently, however, companies have begun seeking more efficient returns on their AI spending instead of indiscriminately increasing token consumption. After Uber exhausted its AI coding budget for the year in just four months, it capped employees’ AI model spending at $1,500 per month each. Walmart instructed employees to use relatively inexpensive models for simple tasks and introduced a system to reduce costs arising when multiple employees submit the same question or task. Amazon discontinued the employee usage leaderboard it had operated to encourage AI adoption. The move was intended to prevent employees from using AI for unnecessary tasks merely to compete in the rankings.

U.S. human resources software company Rippling has also built an “AI Spend Console” that compares work output, prompt counts and token costs, while introducing an AI gateway that automatically selects a model in an appropriate price range according to task difficulty. U.S. data and AI platform company Databricks adopted an approach that changes the model itself to suit the complexity of the work, rather than merely setting usage limits for individual employees. After internal tests found that Chinese company Z.ai’s open-weight GLM 5.2 model delivered task performance statistically similar to Anthropic’s high-performance Claude Opus 4.8, Databricks expanded use of the GLM model family in developers’ work and introduced a “smart routing” system that assesses task complexity and automatically selects a suitable model. In Databricks’ tests, GLM 5.2 cost $1.28 per task, approximately 34% less than Opus 4.8’s $1.94. Research from the Swiss Institute of Artificial Intelligence likewise identifies model routing and cost per completed task as central to managing AI consumption.

Token-Efficiency Technologies Gain Prominence

The ability to use tokens efficiently is expected to become increasingly important as the spread of AI agents makes cost management more urgent. Among the leading ways to reduce token consumption today are minimizing model calls and repeated computation through prompt and context caching. Instead of processing recurring inputs—such as system prompts, long documents and conversation histories—from scratch each time, these methods reuse previously computed results. Another approach eliminates the model call altogether. AI gateway providers use methods including in-memory caches that immediately return a previous response to an identical request, Redis caches shared across multiple servers, and semantic caches that identify questions with similar meanings despite different wording and reuse existing answers. Because these approaches do not call the model API, the requests incur no token charges. The Economy has detailed the role of caching and model routing in controlling API costs.

Prompt and context compression are also major cost-saving measures. They reduce the number of input tokens by summarizing long documents or conversation histories, removing less important sentences and tokens, and using retrieval-augmented generation (RAG) to select the passages most relevant to a query. A prominent implementation is Microsoft Research’s LLMLingua. It uses a small language model to assess the importance of input tokens and remove unnecessary material, compressing prompts by up to 20 times while limiting performance loss to approximately 1.5 percentage points.

Picture

Member for

1 year 1 month
Real name
Oliver Griffin
Bio
[email protected]

Oliver Griffin is a policy and tech reporter at The Economy, focusing on the intersection of artificial intelligence, government regulation, and macroeconomic strategy. Based in Dublin, Oliver has reported extensively on European Union policy shifts and their ripple effects across global markets. Prior to joining The Economy, he covered technology policy for an international think tank, producing research cited by major institutions, including the OECD and IMF. Oliver studied political economy at Trinity College Dublin and later completed a master’s in data journalism at Columbia University. His reporting blends field interviews with rigorous statistical analysis, offering readers a nuanced understanding of how policy decisions shape industries and everyday lives. Beyond his newsroom work, Oliver contributes op-eds on ethics in AI and has been a guest commentator on BBC World and CNBC Europe.