DeepSeek-V4-Flash-0731 is the officially updated version of the DeepSeek-V4-Flash model. The model is built on the same Mixture-of-Experts (MoE) architecture with 284 billion parameters, of which only 13 billion are activated per token, delivering high performance at relatively modest inference compute costs. The model supports a context length of up to one million tokens. It employs the signature engineering innovation of the V4 family — a hybrid attention mechanism that radically reduces the cost of long-context processing. Instead of uniformly handling all tokens, the model uses three modes simultaneously: Compressed Sparse Attention (CSA) compresses the context at a 1:4 ratio, after which the Lightning Indexer selects only the most relevant blocks for computation; Heavily Compressed Attention (HCA) applies extreme 1:128 compression and performs full global attention over ultra-compact representations of the entire history; in parallel, a sliding window of 128 tokens processes the immediate context without compression, preserving detailed local awareness.
In version 0731, a built-in DSpark module has been added that implements speculative decoding: the model can generate several tokens “ahead” and verify them in a single step, noticeably accelerating inference without quality loss (it increases memory footprint and is disabled by default; it can be enabled via the speculative parameter differently across inference engines). However, the key innovation is the improvement in post-training, as a result of which the model’s agentic capabilities have increased several-fold: compared to the preview version, scores on Terminal Bench 2.1 rose from 61.8% to 82.7%, on DeepSWE from 7.3% to 54.4%, and on Cybergym from 38.7% to 76.7%. Furthermore, 0731 outperforms much larger models, notably DeepSeek-V4-Pro (Preview) and the popular GLM-5.2, on most agentic benchmarks, and comes close to the strongest proprietary models such as Opus-4.8.
Use cases span a wide range of tasks where both context length and deployment cost-efficiency are critical: analysis and summarization of large documents — legal dossiers, scientific articles, corporate reports up to a million tokens; search across extensive corporate knowledge bases and information extraction from voluminous archives; multi-step agent work with long sessions and numerous tools; programming and mathematical modeling tasks that demand both reasoning accuracy and the ability to keep large codebases or specifications in view; building chatbots, automatic code generation systems, intelligent document assistants, and research agents. The MIT license allows unrestricted commercial use.
| Model Name | Context | Type | GPU | Status | Link |
|---|
There are no public endpoints for this model yet.
Rent your own physically dedicated instance with hourly or long-term monthly billing.
We recommend deploying private instances in the following scenarios:
| Name | GPU | TPS | Max Concurrency | |||
|---|---|---|---|---|---|---|
1,048,576.0 pipeline |
3 | $7.37 | 7.950 | Launch | ||
1,048,576.0 tensor |
2 | $8.19 | 1.591 | Launch | ||
1,048,576.0 tensor |
4 | $9.17 | 4.724 | Launch | ||
1,048,576.0 tensor |
2 | $9.42 | 7.527 | Launch | ||
1,048,576.0 tensor |
2 | $9.42 | 7.527 | Launch | ||
1,048,576.0 tensor |
4 | $9.50 | 4.724 | Launch | ||
1,048,576.0 tensor |
8 | $11.58 | 1.163 | Launch | ||
1,048,576.0 pipeline |
3 | $11.74 | 7.924 | Launch | ||
1,048,576.0 tensor |
4 | $14.99 | 4.715 | Launch | ||
| Name | GPU | TPS | Max Concurrency | |||
|---|---|---|---|---|---|---|
1,048,576.0 pipeline |
3 | $7.37 | 6.058 | Launch | ||
1,048,576.0 tensor |
4 | $9.17 | 4.248 | Launch | ||
1,048,576.0 tensor |
2 | $9.42 | 6.575 | Launch | ||
1,048,576.0 tensor |
2 | $9.42 | 6.575 | Launch | ||
1,048,576.0 tensor |
4 | $9.50 | 4.248 | Launch | ||
1,048,576.0 pipeline |
3 | $11.74 | 6.032 | Launch | ||
1,048,576.0 pipeline |
3 | $12.38 | 11.069 | Launch | ||
1,048,576.0 tensor |
4 | $14.99 | 4.239 | Launch | ||
1,048,576.0 tensor |
4 | $16.23 | 6.007 | Launch | ||
Contact our dedicated neural networks support team at support@immers.cloud or send your request to the sales department at sale@immers.cloud.