Model launches

AI-generated text

DeepSeek launches V4.1 Flash, sharpens price competition in AI models

Chinese company DeepSeek introduced the V4.1 Flash model, which it says can operate at a fraction of one cent per million tokens and delivers near-competitive performance versus major US and Chinese rivals.

DeepSeek launches V4.1 Flash, sharpens price competition in AI models

Chinese company DeepSeek has unveiled its V4.1 Flash language model, a release that has re-ignited price competition among AI service providers. DeepSeek says the new model can in some cases run at only a fraction of one cent per one million tokens while delivering performance close to several U.S. and Chinese rivals.

What DeepSeek claims

According to the company, V4.1 Flash contains 552 billion parameters, but its architecture activates only a small portion of that capacity at a time, lowering runtime resource requirements. DeepSeek reports that the model outperforms its previous V4-Pro on programming tasks and agent-related workloads, and in its internal tests it exceeds the performance of some known Chinese rivals, including the Moonshot Kimi K3. On public benchmarks, however, V4.1 Flash still trails the flagship models from Anthropic and OpenAI.

Pricing strategy and market dynamics

DeepSeek’s approach focuses on offering sufficiently good performance for a large share of tasks at a much lower cost rather than competing for absolute top-end accuracy. That positioning targets AI agents in particular: these systems can run autonomously for hours, consume large numbers of tokens, and thus are highly sensitive to per-token pricing differences.

Market observers such as Artificial Analysis have labeled the situation a "DeepSeek death zone," arguing that rivals face a difficult choice: match DeepSeek on price or deliver a performance advantage large enough to justify higher fees. Alongside DeepSeek, Tencent, ByteDance and Alibaba are also pursuing aggressive pricing, indicating the Chinese price war is likely to continue.

Market and semiconductor reactions

The announcement produced rapid market moves: Hong Kong-listed MiniMax and Z.AI shares fell by more than 8 percent, while Alibaba’s stock dropped by over 2 percent. Semiconductor stocks also declined: SK Hynix’s U.S.-traded shares slid 5.8 percent and Micron fell 5.3 percent on Thursday. Investors worry that more efficient model operation could reduce the hardware demand needed to run AI models over time.

Technical changes behind V4.1 Flash

DeepSeek reduced the size of the KV cache in V4.1 Flash, which lowers the need for high-bandwidth HBM memory and SSD storage when operating the model. That optimization is notable amid ongoing shortages and high demand for certain memory chips used in AI infrastructure.

Deployment and next steps

DeepSeek will retire the V4-Pro model on September 14 and automatically migrate workloads running on V4-Pro to V4.1 Flash at the new, lower pricing. The company is betting that cost competitiveness will play at least as important a role as raw model performance in the next growth phase of AI adoption—an argument that has particular weight for cost-sensitive enterprise use cases.

Why this matters

Widespread adoption of lower-cost, high-performing models could reshape which vendors and solutions remain competitive. If price becomes the dominant decision factor, some enterprise applications that were previously cost-prohibitive with higher-priced U.S. models may become viable with cheaper alternatives.

This article is not investment advice or a recommendation.

Tags: China, artificial intelligence, DeepSeek, price war, AI agents, Moonshot, Anthropic, OpenAI, Alibaba, Tencent, ByteDance, SK Hynix, Micron