DeepSeek's new V4.1 Flash model claims to shrink the memory footprint of AI inference to a quarter of its predecessor's, a change that would cut the cost of running long-lived software agents and reset the price floor for Chinese AI cloud services.
The Hangzhou-based developer said the 552-billion-parameter model, released Sept 10, uses a Causal-Encoder-Decoder architecture that projects the decoder's key-value cache directly from encoder hidden states rather than deriving it from every decoder layer. The company said high-bandwidth memory demand falls to one-fourth and solid-state drive demand to one-eighth of the prior generation, with the cache compressed to 890 bytes per token — 1/437th of the original DeepSeek V1.
"The cache is where agent economics are won or lost," said Wei Sun, a senior analyst covering China cloud infrastructure at a Shanghai research firm, who reviewed the published specifications. "If those numbers hold under production load, the marginal cost of a 20-step tool-calling loop drops by more than half."
DeepSeek priced the model at 0.02 yuan per million cache-hit input tokens in off-peak hours and 0.04 yuan at peak, a 60% cut from V4 Flash. Cache-miss input fell about 33.3% to 1 yuan off-peak, and output fell about 11.1% to 4 yuan. Peak hours run 9:00-12:00 and 14:00-18:00 Beijing time Monday through Friday. On a fixed workload of 1 million cache-hit tokens, 100,000 cache-miss tokens and 10,000 output tokens, the off-peak bill drops to 0.16 yuan from 0.245 yuan, a 34.7% decline. Concurrency limits rise to 2,500 requests from 500.
The architecture is asymmetric by design: the 552B MoE backbone activates 8 billion parameters during prefill and 16 billion during decode. That split matters because agents read far more than they write — code repositories, tool descriptions and conversation history are ingested repeatedly, while the model emits comparatively little. DeepSeek said the structure scales to larger parameter counts, which sets the baseline for a future V4.1 Pro.
Benchmarks put V4.1 Flash ahead of the outgoing V4 Pro on agentic tasks. It scored 90.6 on Terminal-Bench 2.1 against V4 Pro's 87.9, and 88.1 on CyberGym against 83.3. It also posted 90.9 on GPQA Diamond, a Codeforces rating of 3471 and 65.6 on MathArena Apex. On pure reasoning it still trails Anthropic's Opus at 93.4 and OpenAI's GPT-5.6 Sol at 94.1 on GPQA Diamond. DeepSeek did not disclose the test conditions behind the comparisons.
Tencent's coding tools move first, and V4 Pro gets retired
Tencent's WorkBuddy and CodeBuddy, along with the open-source coding agent OpenCode, have integrated V4.1 Flash as official partners. Tencent shares fell 1.982% to HK$8.60 lower on the day, with short selling of $813.87 million representing 15.975% of turnover as of 12:25 local time, according to AASTOCKS data.
DeepSeek will route all requests to the retiring deepseek-v4-pro model to V4.1 Flash at the lower unit price after 12:00 Beijing time on Sept 14, until V4.1 Pro ships. The older deepseek-v4-flash and deepseek-v4-flash-vision-exp names remain callable but now resolve to the new model, folding text and image understanding into a single API entry. The model supports a 1 million-token context window, 384K maximum output, JSON output, tool calling and the Responses API, with thinking enabled by default across low, high and max effort levels.
The compression claim carries a second edge. If replicated across the industry, a fourfold reduction in HBM demand per unit of inference would temper the incremental memory consumption that has underpinned forecasts for Samsung, SK Hynix and Micron. DeepSeek's own framing is narrower — it describes the gain as lowering cache-hit costs in agent scenarios, not as reducing total accelerator count. The two are not the same, and the company published no data on aggregate memory purchases.
UOB Kay Hian kept a constructive view on Greater China internet and AI cloud, naming Alibaba, Trip.com, NetEase and Z.AI among its top picks, while flagging e-commerce as soft and gaming as resilient. Cheaper agent inference supports that call by widening the margin on automation workloads that platforms are already selling. Alibaba's Qwen models and Z.AI's GLM line compete directly with DeepSeek for the same enterprise budgets, so a price cut at the low end compresses the whole domestic market's pricing power even as it expands the volume of work that is economic to automate.
DeepSeek is preparing a listing on Shanghai's STAR Market and has engaged CITIC Securities for counselling, with a reported valuation of 500 billion yuan. The company has not responded to requests for comment on the listing arrangements.
This article is for informational purposes only and does not constitute investment advice.