Model launches

DeepSeek releases 304B-parameter DeepSeek-V4-Flash-0731, touts strong cost-to-intelligence ratio

DeepSeek introduced DeepSeek-V4-Flash-0731, a 304 billion-parameter model positioned as a cost-effective member of its V4 family with enhanced agentic capabilities.

DeepSeek releases 304B-parameter DeepSeek-V4-Flash-0731, touts strong cost-to-intelligence ratio

DeepSeek has released a new member of its V4 family, DeepSeek-V4-Flash-0731, which the company describes as having "substantially enhanced agentic capabilities." The model contains 304 billion parameters and is distributed as a 167 GB package on Hugging Face.

Performance and pricing

Early comparisons highlight a favorable cost-to-intelligence ratio for DeepSeek-V4-Flash-0731. Artificial Analysis ranks the model ahead of MiniMax M3, a 428 billion-parameter model, on its intelligence metrics. DeepSeek's published pricing is $0.14 per million input tokens and $0.27 per million output tokens, positioning the model as a potentially strong value option relative to larger models.

Visuals and user tests

The model reportedly performs well on charts that compare an Intelligence Index versus Cost per Intelligence Index Task, suggesting a competitive intelligence-per-cost metric. Community testing also shows that output quality can depend on the chosen reasoning effort: a pelican-themed prompt produced disappointing output with the default reasoning level, but yielded noticeably better results when the reasoning level was raised.

Example command used through OpenRouter:

llm -m openrouter/deepseek/deepseek-v4-flash-0731 -t pelican -o reasoning_effort high

Availability and context

DeepSeek-V4-Flash-0731 is available via OpenRouter and identifiable on Hugging Face by its 167 GB package size. Community reports and third-party analyses suggest the model is a competitive alternative to larger-parameter models, particularly where cost-efficiency matters more than raw parameter count.

Why this matters

In the LLM market, operating cost and resource requirements are critical for practical deployments. If DeepSeek-V4-Flash-0731 delivers on its implied intelligence-per-cost advantage, it could lower the cost of deploying capable models in use cases where sheer parameter count does not directly translate to better real-world outcomes.

Tags and sources

The release and early tests have been discussed across community channels (including Hacker News) and in analyst summaries; the model sits within the generative AI, LLM and OpenRouter ecosystems.