DeepSeek has released DeepSeek-V4.1-Flash, a 552-billion-parameter model described as the smallest model in a new architecture family. The launch includes lower API prices and a plan to discontinue the more expensive V4 Pro tier.

The model uses DeepSeek’s asymmetric Causal Encoder-Decoder structure with a mixture-of-experts design. It activates 8 billion parameters when processing input and 16 billion parameters when generating output. DeepSeek designed this split to help keep inference costs low.

V4.1 Flash also includes native multimodal visual understanding. Image support is now part of the model rather than requiring the separate experimental model previously used for visual tasks.

DeepSeek V4.1 Flash Architecture

DeepSeek-V4.1-Flash has 552 billion parameters in total, though only a portion is active at a given stage of processing.

Its architecture uses different active-parameter counts for input and output:

  • Input processing: 8 billion active parameters
  • Output generation: 16 billion active parameters

DeepSeek refers to the design as an asymmetric Causal Encoder-Decoder structure built on a mixture-of-experts approach. The configuration is intended to reduce inference costs while supporting both text and visual understanding.

V4.1 Flash Benchmarks

DeepSeek’s benchmark results show V4.1 Flash ahead of its V4 Pro predecessor on several agentic and coding evaluations. The company’s table also places the model alongside OpenAI’s GPT-5.6 Sol and Anthropic’s Claude Opus 5 on several tasks.

On the DeepSWE v1.1 software engineering benchmark, the reported scores were:

 

Model

 

 

DeepSWE v1.1 Score

 

 

DeepSeek-V4.1-Flash

 

 

74.2

 

 

Claude Opus 5

 

 

74.0

 

 

GPT-5.6 Sol

 

 

73.0

 

The results are based on DeepSeek’s own benchmark table. Independent benchmark boards had not yet published scores for V4.1 Flash.

Performance varies by evaluation. On knowledge-focused tests such as HLE, Claude Opus 5 still has a clear lead.

DeepSeek V4.1 Flash API Pricing

New V4.1 Flash pricing took effect on September 10 at noon Beijing time.

 

API Usage Type

 

 

Price per Million Tokens

 

 

Off-peak output

 

 

$0.60

 

 

Uncached input

 

 

$0.15

 

The updated rates represent an approximate 9% reduction in output pricing and a 32% reduction in uncached input pricing compared with the post-increase V4 Flash rates introduced on August 16.

That August adjustment raised certain rates by more than ten times. The increase drew criticism from developers and faced pressure from open-source clones offering lower-priced alternatives to DeepSeek’s official API.

V4 Pro Requests Will Move to V4.1 Flash

Beginning September 14 at noon Beijing time, API requests sent to V4 Pro will be automatically routed to V4.1 Flash. Those requests will be charged at Flash-series prices.

The routing change effectively retires the V4 Pro tier until a future V4.1 Pro model is released.

DeepSeek said its internal and external testing found that V4.1 Flash surpassed V4 Pro across performance, cost, speed, and total completion time.

DeepSeek IPO Preparation

The V4.1 Flash release comes as DeepSeek prepares for an initial public offering on Shanghai’s STAR Market.

Reuters reported that the company hired CITIC Securities to underwrite the listing and plans to begin the IPO process this year. The timing, size, and target valuation have not been decided.

A recent pre-IPO financing round valued the Hangzhou-based company at about $74 billion to $75 billion.