Skip to main content
  • Home
  • Tech
  • “Efficiency Over Performance”: Rising Token Costs Reshape AI Competition as Google, OpenAI, and Anthropic Bet on Lightweight Models

“Efficiency Over Performance”: Rising Token Costs Reshape AI Competition as Google, OpenAI, and Anthropic Bet on Lightweight Models

Picture

Member for

1 year
Real name
Siobhán Delaney
Bio
Siobhán Delaney is a Dublin-based writer for The Economy, focusing on culture, education, and international affairs. With a background in media and communication from University College Dublin, she contributes to cross-regional coverage and translation-based commentary. Her work emphasizes clarity and balance, especially in contexts shaped by cultural difference and policy translation.

Modified

Google launches lightweight AI models focused on reducing operating costs
OpenAI and Anthropic also prioritize cost efficiency as the token-maxxing era fades
AI tokenomics reshapes the market while low-cost Chinese models gain ground
Source: Google

Google has introduced a new generation of lightweight artificial intelligence models designed to reduce the cost of operating AI services. As AI tokenomics becomes a central management concern for companies around the world, the models focus on consuming fewer tokens, lowering output prices, and reducing the financial burden on customers.

This trend toward greater model segmentation and lighter architectures is not limited to Google, with OpenAI, Anthropic, and other major US AI companies increasingly adopting similar strategies in their product portfolios.

Google’s Latest Lightweight Models

On July 21, local time, Google launched Gemini 3.6 Flash and Gemini 3.5 Flash-Lite, two models designed to improve the token efficiency of AI agents performing real-world business tasks while reducing operating costs. Gemini 3.6 Flash improves coding, multimodal processing, document analysis, and knowledge-based task performance compared with the previous Gemini 3.5 Flash while substantially lowering token consumption. Independent AI evaluation organization Artificial Analysis estimated that Gemini 3.6 Flash used 17% fewer output tokens than its predecessor and reduced token consumption by as much as 65% on DeepSWE, a software engineering benchmark.

Under Google’s application programming interface pricing, Gemini 3.6 Flash costs $1.50 per million input tokens and $7.50 per million output tokens, compared with an output price of approximately $9 for Gemini 3.5 Flash. Google emphasized that customers could substantially reduce the operating cost of AI agents through the combined effects of lower token consumption and a reduced output price. The model’s overall capabilities, however, remained below those of other frontier systems. Artificial Analysis assigned Gemini 3.6 Flash an Artificial Analysis Intelligence Index score of 50, placing it slightly behind Muse Spark 1.1, which received a score of 51 after Meta entered the paid enterprise AI market with an aggressive pricing strategy.

Gemini 3.5 Flash-Lite, released alongside the new model, is designed for extremely rapid processing and large volumes of requests. Developers can select among Minimal, Low, and Higher reasoning levels depending on the task. For high-throughput workloads such as large-scale document processing and AI-powered search, users can choose a setting that minimizes cost and latency, while more complicated multistep tasks can be assigned a higher level of reasoning. According to Artificial Analysis, Gemini 3.5 Flash-Lite can generate approximately 350 output tokens per second and is priced at $0.30 per million input tokens and $2.50 per million output tokens.

AI Companies Shift Their Focus Toward Cost

Google is not the only company accelerating the release of lightweight models. OpenAI’s latest GPT-5.6 family, for example, consists of Sol, which targets maximum performance; Terra, a lower-cost general-purpose model; and Luna, which emphasizes speed and cost efficiency. The structure encourages users to assign models according to the difficulty of each task rather than relying on a single high-performance system for every problem. Under standard API pricing, GPT-5.6 Terra costs $2.50 per million input tokens and $15 per million output tokens, while GPT-5.6 Luna is priced more affordably at $1 per million input tokens and $6 per million output tokens.

Anthropic is also responding to demand for cost-effective models through Claude Haiku. Haiku 4.5 is designed for real-time tasks requiring a moderate level of reasoning, including customer service, coding, and research support, and was developed on the assumption that it would share responsibilities with larger models. A higher-level Sonnet model may plan a complicated assignment, for example, while multiple Haiku agents simultaneously perform narrower tasks such as searching code, reviewing source material, and monitoring data flows. Haiku 4.5 is priced at $1 per million input tokens and $5 per million output tokens.

AI companies have begun focusing more heavily on cost efficiency because the token expenses generated in proportion to generative AI usage are becoming increasingly burdensome. Until 2025, a practice known as “token maxxing,” in which companies used the most powerful models available and deliberately maximized token consumption, was common in Silicon Valley. Some major technology companies, including Meta and Amazon, even ranked their developers according to token usage in an effort to encourage greater adoption of AI. That environment has recently reversed. Although technological progress has sharply reduced the price of individual tokens, the spread of AI throughout corporate operations has caused total enterprise expenditure to rise. According to the Linux Foundation, AI token prices have fallen by approximately 98% over the past three years, while companies’ overall AI-related costs have increased by 320% during the same period.

Competition Shifts from Performance to Efficiency

AI tokenomics has emerged as a central concept in this new environment. The term refers to analyzing and managing the production and consumption of tokens from an economic perspective and has become an important foundation for efforts to improve AI cost efficiency across global industries. Companies increasingly recognize token consumption as a direct operating expense and are developing more disciplined model-management strategies. One representative approach is to assign complicated reasoning and final decision-making to high-performance models while transferring repetitive and standardized tasks to inexpensive lightweight systems. This is precisely the market targeted by the new models released by Google, OpenAI, and Anthropic.

An increasing number of companies are also building model-routing systems that analyze the complexity of each request, the required response speed, and the acceptable cost before automatically connecting the task to the most appropriate model. According to Ara Karajian, chief economist at corporate expense-management platform Ramp, the proportion of companies using model routers increased from 1% in 2025 to 5% in 2026. The shift suggests that corporate AI adoption is moving away from the simple selection of a single preferred provider and toward the dynamic allocation of workloads across multiple systems.

Companies are also increasingly adopting inexpensive Chinese AI models to reduce costs. On AI model aggregation platform OpenRouter, DeepSeek’s V4 Flash has recorded the highest token usage since mid-May, while other Chinese models such as Tencent’s Hy3 and MiniMax’s M3 have also ranked near the top. An analysis of June usage found that Chinese models accounted for approximately 60% of all developer token traffic on the platform. OpenRouter lists DeepSeek V4 Flash at $0.09 per million input tokens and $0.18 per million output tokens, while Tencent Hy3 costs $0.132 per million input tokens and $0.528 per million output tokens.

Chinese AI companies are further reducing barriers to adoption by releasing model weights without significant restrictions. Developers can deploy Chinese open-source or open-weight systems on their own servers, retain control over corporate data, and reduce their exposure to changes in API prices or usage policies imposed by outside providers. This combination of extremely low prices, greater deployment flexibility, and improving performance is expanding the presence of Chinese models in international markets.

Against this backdrop, the release of lightweight models by US AI companies appears less like an optional product expansion than a necessary survival strategy. If American providers fail to offer affordable alternatives to Chinese systems and rely exclusively on superior performance, customers may transfer routine workloads elsewhere, eliminating opportunities to sell them more advanced models, cloud infrastructure, and agent-development tools. Lightweight models therefore serve both as independent revenue sources and as a defensive layer protecting the broader commercial ecosystems of leading AI companies. This is why the outcome of the next phase of AI competition is likely to depend not only on which company builds the most powerful model, but also on which can combine high-performance and lightweight systems most efficiently.

Picture

Member for

1 year
Real name
Siobhán Delaney
Bio
Siobhán Delaney is a Dublin-based writer for The Economy, focusing on culture, education, and international affairs. With a background in media and communication from University College Dublin, she contributes to cross-regional coverage and translation-based commentary. Her work emphasizes clarity and balance, especially in contexts shaped by cultural difference and policy translation.