China Lowers Barriers with ‘Open Weights’ as US Counters with ‘Token Efficiency,’ Per-Task Costs to Decide AI Leadership
Authored On
Modified
Chinese AI rapidly expands global market share on low costs and open access US models can offer lower real-world task costs depending on token efficiency and evaluation methodology US Big Tech joins price-cutting drive, shifting AI competition from unit pricing to overall efficiency

Chinese artificial intelligence (AI) models are making inroads into the global market by leveraging low prices and openness. Their expansion has been particularly pronounced in Africa, where adoption is growing primarily across the financial and agricultural sectors because sensitive data can be processed on in-house servers and models can be readily tailored to local languages and industry-specific needs. Some real-world evaluations, however, have found that proprietary US models can produce higher-quality responses at lower costs than Chinese open-weight models. OpenAI, Anthropic and Google are also cutting prices for their midrange and lightweight models in rapid succession, putting further pressure on the price competitiveness of Chinese models.
Chinese Open-Weight Models Make Inroads into Africa
According to the South China Morning Post (SCMP) on August 25, Curacel, Africa’s largest insurtech company and headquartered in Nigeria, recently added Zhipu AI’s Chinese-developed “GLM-5.3” model to its AI-powered insurance claims and fraud-detection system alongside the US models it already used. “For high-volume, repetitive tasks such as coding, data extraction, document classification and customer support, leading Chinese models have nearly closed the performance gap with proprietary Western systems while delivering the same outputs at a far lower cost,” Curacel CEO Henry Mascot said. “By using a multimodel architecture that flexibly assigns tasks based on quality and cost, we are avoiding the risk of becoming locked into a particular Big Tech vendor.”
The principal reason African technology companies are turning to Chinese models—including Alibaba’s “Qwen,” Moonshot AI’s “Kimi,” DeepSeek and Zhipu’s “GLM”—is their open-weight architecture, which allows businesses to download model weights directly and run them independently on their own servers. Open-weight models can process bank transaction records and know-your-customer (KYC) documents on corporate servers without transmitting them to external APIs, reducing regulatory burdens in the financial, healthcare and public sectors. “The ability to host and control highly sensitive financial data, including bank transaction records and KYC documents, within a company’s own infrastructure rather than transmitting it to an overseas third-party API is decisive,” said Bernard Momanyi Nyagaka, co-founder of Nairobi-based enterprise AI consultancy Sanifu AI. “In finance, data sovereignty should be a core design principle, not a feature offered by a particular vendor.”
Linguistic and regional suitability has also helped drive the spread of Chinese models. “Sunflower,” developed by Ugandan researchers, selected Qwen3 as its foundation model after comparing US and Chinese alternatives and provides farmers with weather information and agricultural advice in more than 30 local languages. According to research published in the international academic journal PMLR, the Sunflower 14B and 32B models outperformed existing models in assessments of comprehension across Uganda’s principal languages. By combining a Chinese-developed foundation model with African-language data and industry-specific use cases, the project effectively lowered localization costs and accelerated service deployment.
Low Token Prices as a Competitive Advantage
Price is another factor behind Chinese AI’s growing global reach. US-developed AI provides state-of-the-art performance, but most offerings are proprietary paid services that incur ongoing costs based on usage. Chinese AI models, by contrast, are released as open source, allowing anyone to download and modify them free of charge without separate approval. Companies can build their own AI services by covering only the cost of servers. “A Toyota hatchback is enough to take your child to school—why use an expensive Ferrari?” said Moses Kemibaro, who runs a digital marketing business in Kenya. “Chinese AI can reduce total deployment costs by as much as 90%.”
According to the Financial Times, an analysis by US corporate payments platform Ramp of payment records from 70,000 companies found that spending on Anthropic’s flagship model, “Claude Fable 5,” accounted for only 11.4% of total spending on the company’s models. The less capable “Opus 4.8” and “Opus 5” accounted for 34.9% and 19.1%, respectively. Fable 5 is priced at $10 per million input tokens and $50 per million output tokens, roughly twice the price of OpenAI’s “GPT-5.6 Sol.”
Table 1. Cost and Usage Comparison of US and Chinese AI Models
| Category | US Models | Chinese Models | Comparison |
|---|---|---|---|
| Cost of building an e-commerce website | Anthropic Fable 5 $48.99 | Alibaba Qwen 3.7 Max $4.08 | Fable 5 is approximately 12 times more expensive |
| Functional accuracy | 100% for most models | Overall completion quality was broadly similar | |
| Share of global monthly token usage | - | Less than 10% in early 2025 → More than 60% in July 2026 | Chinese models’ share surged |
| Weekly usage ranking (August 3–9, 2026) | - | Ranked first through fourth Six of the top 10 | Chinese models dominated the top rankings |
Global AI Demand Prioritizes Value over Best-in-Class Performance
The price disparity was also evident in real-world tasks. In an experiment conducted by Bloomberg and independent AI evaluation platform Vals AI, seven leading US and Chinese models were instructed to build the same e-commerce website, with most achieving 100% functional accuracy. Anthropic’s Fable 5, however, incurred a charge of $48.99, while Alibaba’s Qwen 3.7 Max completed the task for $4.08. DeepSeek V4 cost even less. Costs differed by as much as twelvefold for outputs of comparable quality.
This price disparity has also been reflected in usage metrics. Chinese models’ share of global monthly token usage on AI model aggregation platform OpenRouter jumped from less than 10% early last year to more than 60% in July this year. Chinese models swept the top four positions in the weekly usage rankings for August 3–9 and accounted for six of the top 10. US AI assistant startup Lindy switched its primary model from an Anthropic product to DeepSeek in June, reducing costs by 90%. Workflow automation startup Palcia also uses Anthropic for real-time customer service while assigning overnight or non-urgent tasks to a Chinese MiniMax model.
The Chinese government is also seeking to capitalize on this trend as an opportunity to expand its soft power through AI. In May, China launched Huawei Cloud’s model-as-a-service (MaaS) offering—which provides access to DeepSeek, Qwen and GLM—in South Africa, followed in July by the establishment of the World AI Cooperation Organization (WAICO), with 29 participating countries. It also unveiled plans to provide 5,000 AI training and seminar opportunities to developing countries over the next five years and establish international AI application cooperation centers with major multilateral organizations, including the African Union (AU). The strategy seeks to secure an early lead in the Global South’s AI ecosystem by combining low-cost open-weight models with cloud services, education and data-center development.
US Models Prove “Cheaper and More Accurate” in Real-World Comparisons
However, substantial technical variables make it difficult to equate the rapid adoption of Chinese models directly with market dominance. AI-powered search and market-intelligence company AlphaSense released research earlier this month showing that OpenAI’s and Anthropic’s proprietary models were actually cheaper than Chinese open-weight models for certain tasks. In the experiment, OpenAI’s “GPT-5.6 Sol” and Anthropic’s “Opus 4.8” delivered higher-quality responses at lower costs than Moonshot AI’s “Kimi K3” and Zhipu AI’s “GLM-5.2.”
The cost reversal stemmed from differences in token efficiency among models. Kimi K3’s output price of $15 per million tokens was lower than those of Opus 4.8 at $25 and GPT-5.6 Sol at $30, but it required more tokens to complete an answer. The median total cost per query for GPT-5.6 Sol was 13% lower than for Kimi K3, while its quality score was 20% higher. Opus 4.8 also achieved a 13% higher quality score at half the cost of Kimi K3, while GLM-5.2 cost approximately twice as much as Kimi K3 and received a lower score.
Rankings varied depending on the evaluation methodology. In a separate comparison, AI model benchmarking company Artificial Analysis (AA) estimated that Kimi K3 and GLM-5.2 had lower per-task costs than US models. This is because the cost per task changes with the test questions, reasoning intensity, answer length and number of retries. In AA’s latest metrics, DeepSeek V4 Flash recorded the lowest cost among high-performance models at $0.03, whereas Qwen 3.8 Max cost $1.13—more than Opus 5 in medium-reasoning mode at $0.72 and GPT-5.6 Sol in high-reasoning mode at $0.81.
US AI Companies Pursue Both High Performance and Low Costs
Price cuts by US AI companies are also intensifying pressure on the cost competitiveness of Chinese models. On July 30, OpenAI reduced API prices for “Terra,” the midrange model in its GPT-5.6 series, and the lightweight, high-speed “Luna” model by 20% and 80%, respectively. The reductions came about three weeks after the two models were released. Per-million-token input and output prices fell to $2 and $12, respectively, for Terra and $0.20 and $1.20 for Luna, while the flagship “Sol” model retained its previous pricing.
The price reductions focused on midrange and lightweight models designed to handle high-volume, repetitive workloads. OpenAI said improvements to model architecture, reasoning systems and context-management technology had reduced the computing power and number of tokens required for the same task. In customer evaluations released by the company, Notion found that Terra delivered quality comparable to GPT-5.5 while cutting per-task costs by half and processing time by 60%. Software company Blitzy said Luna processed 2.2 times more context than GPT-5.4 Mini while using one-eighth as many output tokens and costing 87% less.
Anthropic has also strengthened its portfolio of low-cost, high-efficiency products. Last month, the company released Claude Opus 5 at the same price as its predecessor, Opus 4.8—$5 per million input tokens and $25 per million output tokens. The company said the model costs half as much as its flagship Fable 5 while scoring within 0.5% of Fable 5’s best result in coding benchmarks. Google likewise priced “Gemini 3.6 Flash,” which uses 17% fewer output tokens than its predecessor, at $1.50 per million input tokens and $7.50 per million output tokens, while setting input and output prices for the high-volume “3.5 Flash-Lite” at $0.30 and $2.50, respectively.
- Previous “Domestic Price War Unleashes Flood of New Models”: China’s Automakers Race Ahead as Safety Testing and After-Sales Networks Falter, Impeding Global Expansion
- Next “Building AI and Rare-Earth Supply Chains Without China”: U.S. Champions Pax Silica as Japan Expands Cooperation With Resource-Rich Nations and Accelerates Indigenous Technology Development