Skip to main content
  • Home
  • Tech
  • "From Cost Cutting to Capital Raising": DeepSeek Goes All In on New Model, In-House Chip and IPO—Will It Reshape the U.S.-China AI Rivalry?

"From Cost Cutting to Capital Raising": DeepSeek Goes All In on New Model, In-House Chip and IPO—Will It Reshape the U.S.-China AI Rivalry?

Picture

Member for

1 year 1 month
Real name
Aoife Brennan
Bio
[email protected]

Aoife Brennan is a contributing writer for The Economy, with a focus on education, youth, and societal change. Based in Limerick, she holds a degree in political communication from Queen’s University Belfast. Aoife’s work draws connections between cultural narratives and public discourse in Europe and Asia.

Modified

DeepSeek Launches V4.1 Flash Model With Lower Output Costs
Pushes Ahead With In-House Inference Chip Ahead of Mainland China IPO
Need for Public-Market Funding Comes Into Focus as U.S.-China AI Battle Intensifies

Chinese artificial intelligence (AI) startup DeepSeek has once again made an aggressive move built around its cost advantage. The company is striving to reduce service operating expenses by releasing a new model priced below its existing flagship and accelerating the development of an in-house inference chip. Market analysts view these improvements to its cost structure as groundwork for reining in rapidly mounting infrastructure spending and facilitating a smooth initial public offering (IPO).

DeepSeek’s New Model

On Sept. 10, local time, DeepSeek unveiled DeepSeek-V4.1-Flash, which carries lower output costs than its flagship DeepSeek V4 Pro model. Of V4.1 Flash’s 552 billion total parameters, only 8 billion are activated when processing a query and 16 billion when generating a response. The architecture minimizes parameter activation to reduce computational workloads and processing costs. The key-value (KV) cache used for long-context processing has also been reduced. Cache capacity stands at 890 bytes per token, while high-bandwidth memory (HBM) usage is one-quarter that of the previous generation and solid-state drive (SSD) storage requirements are one-eighth as large. The model supports context and image inputs of up to 1 million tokens.

V4.1 Flash’s application programming interface (API) fees during peak hours are set at $0.30 per million input tokens and $1.20 per million output tokens, falling sharply during off-peak periods to $0.15 for input and $0.60 for output. GPT-5.6 Luna costs $0.20 per million input tokens and $1.20 per million output tokens, while Claude Sonnet 5 costs $2 and $10, respectively. DeepSeek plans to route V4 Pro API requests to V4.1 Flash. The company said it would phase out the existing model after V4.1 Flash outperformed V4 Pro across multiple evaluations covering performance, cost, speed and total task-completion time. This arrangement will remain in place until the release of V4.1 Pro, the company’s next flagship model.

Further Cost Reductions Through an In-House AI Chip

DeepSeek’s cost-cutting drive extends beyond model development to efforts to secure its own chips. According to Reuters, DeepSeek has been pursuing a project to develop an AI inference chip since last year. The company’s decision to prioritize chips used for inference rather than training reflects the AI industry’s distinctive cost structure. During AI model training, enormous computing demand is concentrated at specific points when a model is developed or updated. Inference workloads, by contrast, are performed repeatedly whenever users submit requests throughout the period in which a service remains operational. As the number of users and queries rises, the required number of servers, volume of AI chips and electricity consumption increase in tandem.

An in-house inference chip could reduce the costs associated with graphics processing units (GPUs) during this process. A purpose-built chip designed exclusively for inference on a specific AI model can allocate hardware resources around the computational and data-movement patterns repeatedly used in real-world services. Unlike general-purpose GPUs sold by Nvidia and other suppliers, such a chip can concentrate computing performance and memory architecture on inference workloads, lowering both power consumption and equipment acquisition costs. Optimizing the model architecture and chip in tandem could substantially reduce the hardware resources required to generate responses of comparable quality, ultimately delivering a meaningful reduction in the unit cost of processing each query. DeepSeek’s chip project remains at an early stage, however, and the company has yet to disclose detailed product specifications, the manufacturing process or the chipmaker involved.

Table 1. DeepSeek’s Cost-Reduction Strategy

AreaCore StrategyExpected Effect
Model ComputationSelectively activate only a portion of total parametersReduce computational workloads and processing costs
Memory ArchitectureReduce the KV cache and lower HBM and SSD usageCut server memory and storage costs
API PricingCut input and output fees by 50% during off-peak hoursDistribute usage more evenly and strengthen price competitiveness
Model OperationsRedirect V4 Pro requests to the more efficient V4.1 FlashReduce the operating costs of the existing model
In-House SemiconductorDevelop a dedicated inference chip optimized for the model architectureReduce dependence on general-purpose GPUs and electricity while lowering inference costs
Source: DeepSeek, Reuters

Plans for a Shanghai IPO

Market analysts have suggested that DeepSeek’s efforts are groundwork for an IPO. Reports that the company was moving forward with an IPO first began emerging in earnest in July. At the time, Bloomberg reported, citing multiple people familiar with the matter, that DeepSeek had begun preparing for a listing on a mainland Chinese stock exchange. The company was reportedly discussing listing options with external advisers, including accounting firms and investment banks, and considering submitting an IPO application as early as this year before making its market debut in 2027. Reuters also reported that month that DeepSeek was seeking fresh funding at a valuation of approximately $73.9 billion while considering a listing on the Shanghai Stock Exchange’s Science and Technology Innovation Board, commonly known as the STAR Market.

The listing preparations have advanced further this month. Reuters recently reported that DeepSeek had selected CITIC Securities, one of China’s largest securities firms, to guide its preparations for the offering. Mainland Chinese companies typically undergo preliminary procedures through a brokerage—including pre-listing counseling, internal-control reviews and financial restructuring—before submitting a formal listing application to an exchange. DeepSeek has not yet been confirmed to have filed a formal IPO application, however, and the offering size, listing valuation and detailed timetable have yet to be finalized. Securities-industry observers increasingly expect the company to begin the application process by the end of this year and pursue a stock-market debut in 2027.

Spending Continues to Mount

DeepSeek’s expanding user base lies behind its push to secure large-scale funding. The company’s cost burden has risen rapidly this year as its user base and service traffic have grown simultaneously. According to Chinese market research firm QuestMobile, the DeepSeek application had approximately 129 million monthly active users (MAUs) in June. Model usage among developers has also surged. Data from AI model aggregation platform OpenRouter show that DeepSeek’s share of token usage doubled from about 9% at the beginning of the year to roughly 18% in early June. In mid-July, DeepSeek models accounted for 16.7% of all tokens processed through OpenRouter, briefly placing the company first among model developers.

This increase in computing demand has translated into higher infrastructure spending. According to the Financial Times (FT), DeepSeek has spent $1.6 billion on AI infrastructure so far this year. That is nearly 10 times its expenditure for all of last year and more than three times its annualized recurring revenue (ARR) of $500 million as of last month. Personnel expenses are also rising. DeepSeek announced in June that it would at least double its headcount across all divisions, including research, engineering and product development. This month, it also began recruiting approximately 150 additional engineers to work on backend development and computing systems for AI agents. The company cited sharp increases in data-processing volumes, server numbers, model-training workloads, active users and request traffic as the reasons for the recruitment drive. It said its existing backend infrastructure was approaching its limits and required a sweeping overhaul.

The U.S.-China AI Competitive Landscape

If DeepSeek completes its listing and successfully taps public markets, competition with U.S. AI companies is expected to intensify further. Until now, the confrontation between the United States and China in AI has centered on allegations of unauthorized distillation by Chinese companies. Anthropic said in February that an internal investigation had found DeepSeek, Moonshot AI and MiniMax used approximately 24,000 fraudulent accounts to generate more than 16 million conversations with Claude. OpenAI also submitted a document to the U.S. House Select Committee on China that month, saying it had detected signs that accounts believed to be associated with DeepSeek employees had gained access through intermediaries. The U.S. National Security Agency, Cybersecurity and Infrastructure Security Agency and Federal Bureau of Investigation also concluded that six Chinese AI companies—DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun and Z.AI—had acquired the capabilities of cutting-edge U.S. AI models on an industrial scale.

Distillation lies at the heart of the controversy. The method entails submitting large volumes of queries to a high-performing AI model, collecting its responses and using those outputs as training data to enhance the performance of another model. U.S. authorities acknowledge that distillation itself is a legitimate and useful research technique but argue that the methods and scale employed by Chinese companies are problematic. Critics say information amounting to billions of tokens flowed into China through millions of queries from late 2024 onward and that the Chinese government was likely aware of the activity. “The competitiveness of DeepSeek and other Chinese AI companies has so far stemmed from their ability to use limited capital and computing resources to drive costs to exceptionally low levels,”

one market expert said. “They have reduced development costs and lead times by distilling high-performance external models while compensating for their funding disadvantage with low-cost APIs and efficient model architectures.” The expert added, “Once DeepSeek begins raising substantial capital through an IPO, it could narrow the resource gap with U.S. companies by investing directly in proprietary semiconductors, data centers and research and development (R&D). It would also face pressure from the market to demonstrate technological independence and profitability. The listing could therefore create the conditions for direct competition with leading U.S. companies such as OpenAI and Anthropic across capital, infrastructure and technological capabilities.”

Picture

Member for

1 year 1 month
Real name
Aoife Brennan
Bio
[email protected]

Aoife Brennan is a contributing writer for The Economy, with a focus on education, youth, and societal change. Based in Limerick, she holds a degree in political communication from Queen’s University Belfast. Aoife’s work draws connections between cultural narratives and public discourse in Europe and Asia.