DeepSeek-V4-Flash-0731

reasoning
coding

DeepSeek-V4-Flash-0731 is the officially updated version of the DeepSeek-V4-Flash model. The model is built on the same Mixture-of-Experts (MoE) architecture with 284 billion parameters, of which only 13 billion are activated per token, delivering high performance at relatively modest inference compute costs. The model supports a context length of up to one million tokens. It employs the signature engineering innovation of the V4 family — a hybrid attention mechanism that radically reduces the cost of long-context processing. Instead of uniformly handling all tokens, the model uses three modes simultaneously: Compressed Sparse Attention (CSA) compresses the context at a 1:4 ratio, after which the Lightning Indexer selects only the most relevant blocks for computation; Heavily Compressed Attention (HCA) applies extreme 1:128 compression and performs full global attention over ultra-compact representations of the entire history; in parallel, a sliding window of 128 tokens processes the immediate context without compression, preserving detailed local awareness.

In version 0731, a built-in DSpark module has been added that implements speculative decoding: the model can generate several tokens “ahead” and verify them in a single step, noticeably accelerating inference without quality loss (it increases memory footprint and is disabled by default; it can be enabled via the speculative parameter differently across inference engines). However, the key innovation is the improvement in post-training, as a result of which the model’s agentic capabilities have increased several-fold: compared to the preview version, scores on Terminal Bench 2.1 rose from 61.8% to 82.7%, on DeepSWE from 7.3% to 54.4%, and on Cybergym from 38.7% to 76.7%. Furthermore, 0731 outperforms much larger models, notably DeepSeek-V4-Pro (Preview) and the popular GLM-5.2, on most agentic benchmarks, and comes close to the strongest proprietary models such as Opus-4.8.

Use cases span a wide range of tasks where both context length and deployment cost-efficiency are critical: analysis and summarization of large documents — legal dossiers, scientific articles, corporate reports up to a million tokens; search across extensive corporate knowledge bases and information extraction from voluminous archives; multi-step agent work with long sessions and numerous tools; programming and mathematical modeling tasks that demand both reasoning accuracy and the ability to keep large codebases or specifications in view; building chatbots, automatic code generation systems, intelligent document assistants, and research agents. The MIT license allows unrestricted commercial use.


Announce Date: 31.07.2026
Parameters: 305B
Experts: 256
Activated at inference: 13B
Context: 1049K
Layers: 43, using full attention: 21
Attention Type: Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA)
Developer: DeepSeek
Transformers Version: 4.57.1
vLLM Version: >=0.25.0
License: MIT

Public endpoint

Use our pre-built public endpoints for free to test inference and explore DeepSeek-V4-Flash-0731 capabilities. You can obtain an API access token on the token management page after registration and verification.
Model Name Context Type GPU Status Link
There are no public endpoints for this model yet.

Private server

Rent your own physically dedicated instance with hourly or long-term monthly billing.

We recommend deploying private instances in the following scenarios:

  • maximize endpoint performance,
  • enable full context for long sequences,
  • ensure top-tier security for data processing in an isolated, dedicated environment,
  • use custom weights, such as fine-tuned models or LoRA adapters.

Recommended server configurations for hosting DeepSeek-V4-Flash-0731

Prices:
Name GPU Price, hour TPS Max Concurrency
teslaa100-3.32.384.320
1,048,576.0
pipeline
3 $7.37 7.950 Launch
h100nvl-2.24.192.480
1,048,576.0
tensor
2 $8.19 1.591 Launch
teslaa100-4.16.256.480
1,048,576.0
tensor
4 $9.17 4.724 Launch
h200-2.24.256.320
1,048,576.0
tensor
2 $9.42 7.527 Launch
h200-2.24.256.320.nvlink
1,048,576.0
tensor
2 $9.42 7.527 Launch
teslaa100-4.32.384.320.nvlink
1,048,576.0
tensor
4 $9.50 4.724 Launch
rtx5090-8.44.256.480
1,048,576.0
tensor
8 $11.58 1.163 Launch
h100-3.32.384.320
1,048,576.0
pipeline
3 $11.74 7.924 Launch
h100-4.16.256.480
1,048,576.0
tensor
4 $14.99 4.715 Launch
Prices:
Name GPU Price, hour TPS Max Concurrency
teslaa100-3.32.384.320
1,048,576.0
pipeline
3 $7.37 6.058 Launch
teslaa100-4.16.256.480
1,048,576.0
tensor
4 $9.17 4.248 Launch
h200-2.24.256.320
1,048,576.0
tensor
2 $9.42 6.575 Launch
h200-2.24.256.320.nvlink
1,048,576.0
tensor
2 $9.42 6.575 Launch
teslaa100-4.32.384.320.nvlink
1,048,576.0
tensor
4 $9.50 4.248 Launch
h100-3.32.384.320
1,048,576.0
pipeline
3 $11.74 6.032 Launch
h100nvl-3.24.384.480
1,048,576.0
pipeline
3 $12.38 11.069 Launch
h100-4.16.256.480
1,048,576.0
tensor
4 $14.99 4.239 Launch
h100nvl-4.32.384.480
1,048,576.0
tensor
4 $16.23 6.007 Launch
Prices:
Name GPU Price, hour TPS Max Concurrency
h200-6.52.896.960
1,048,576.0
pipeline
6 $28.39 12.242 Launch
h200-8.52.1024.960
1,048,576.0
tensor
8 $37.37 7.527 Launch
h200-8.52.1024.960.nvlink
1,048,576.0
tensor
8 $37.37 7.527 Launch

Related models

Need help?

Contact our dedicated neural networks support team at support@immers.cloud or send your request to the sales department at sale@immers.cloud.

We use cookies and web analytics services to ensure the proper functioning of the website, analyze web-site traffic, and improve the quality of our services.
By continuing to use the website, you consent to the Privacy Policy and consent to the processing of cookie files and technical data.