DeepSeek-V4-Pro-0813

reasoning
coding

DeepSeek-V4-Pro-0813 is built on a Mixture-of-Experts (MoE) architecture with 1.6 trillion total parameters, of which only 49 billion are activated per token. The configuration includes FP8 quantization for activations and weights, as well as FP4 for expert layers. The model works efficiently with long sequences and can process up to 1,048,576 tokens. Unlike the preview version, the 0813 release is equipped with a DSpark speculative decoding module, which accelerates generation without loss of quality.

Like the DeepSeek-V4-Pro version, the model uses a hybrid attention mechanism combining Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA). CSA first compresses the KV cache along the sequence, merging the entries of every four tokens into one, and then applies DeepSeek Sparse Attention (DSA), in which each query token attends only to compressed KV entries. HCA goes further: it aggressively compresses the KV cache of every 128 tokens into one entry while maintaining dense attention. Alternating CSA and HCA across different layers makes it possible to radically reduce the computational load and memory requirements for the KV cache compared with DeepSeek-V3.2. Additionally, the architecture is reinforced by Manifold-Constrained Hyper-Connections (mHC), a modification of residual connections that ensures numerical stability. The Muon optimizer was used for training, accelerating convergence and improving training quality.

The model demonstrates strong results on a number of key benchmarks. On Terminal Bench 2.1 (testing the ability to solve tasks in the terminal, including working with the command line and file system), the model scores 87.9, ahead of DeepSeek-V4-Flash-0731 (82.7) and DeepSeek-V4-Pro Preview (72.1), approaching Kimi K3 (88.3) and Opus-4.8 (85.0). On NL2Repo (generating code repositories from natural-language descriptions), the result is 61.5 — a significant lead over the preview version (38.5) and the Flash version (54.2). On Cybergym (cybersecurity tasks), the model shows 83.3, beating Kimi K3 (80.0), Opus-4.8 (78.3), and Fable-5 (83.1). On DeepSWE (software engineering) — 62.7, which is higher than Flash (54.4) and Kimi K3 (67.5), but lower than Opus-4.8 (58.0) and Fable-5 (70.0). In Toolathlon-Verified (multi-step tool-use tasks), the model scores 74.1, behind only Kimi K3 (76.5), Opus-4.8 (76.2), and Fable-5 (77.9). Finally, on HLE (Humanity's Last Exam), without tools the model shows 42.7; with tools — 60.0, beating the preview version (37.7 / 48.2) and the Flash version (34.8 / 45.1). Thus, compared with DeepSeek-V4-Pro (Preview), the 0813 release provides a dramatic increase in agentic capabilities and is a full-fledged production solution optimized for real-world scenarios.

Thanks to the million-token context and the high efficiency of CSA + HCA, the model is primarily intended for analyzing ultra-long documents — legal, scientific, financial reports, codebases, and multi-volume archives. Agentic benchmarks confirm its suitability for autonomous agents performing multi-step tasks: terminal management, file-system navigation, calling external tools and APIs. The model is also an effective tool for full-fledged software development — from generating repositories from descriptions to solving complex full-stack-level tasks. In cybersecurity, the model can be used for vulnerability analysis, attack simulation, and automated security auditing. In addition, support for reasoning_effort with three levels (low, medium, high) allows flexible tuning of the balance between speed and depth of reasoning depending on the task.


Announce Date: 13.08.2026
Parameters: 2T
Experts: 384
Activated at inference: 49B
Context: 1049K
Layers: 61, using full attention: 30
Attention Type: Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA)
Developer: DeepSeek
Transformers Version: 4.57.1
vLLM Version: >=0.25.0
License: MIT

Public endpoint

Use our pre-built public endpoints for free to test inference and explore DeepSeek-V4-Pro-0813 capabilities. You can obtain an API access token on the token management page after registration and verification.
Model Name Context Type GPU Status Link
There are no public endpoints for this model yet.

Private server

Rent your own physically dedicated instance with hourly or long-term monthly billing.

We recommend deploying private instances in the following scenarios:

  • maximize endpoint performance,
  • enable full context for long sequences,
  • ensure top-tier security for data processing in an isolated, dedicated environment,
  • use custom weights, such as fine-tuned models or LoRA adapters.

Recommended server configurations for hosting DeepSeek-V4-Pro-0813

Prices:
Name GPU Price, hour TPS Max Concurrency
h200-8.52.1024.1280
1,048,576.0
tensor
8 $37.41 1.085 Launch
h200-8.52.1024.1280.nvlink
1,048,576.0
tensor
8 $37.41 1.085 Launch
Prices:
Name GPU Price, hour TPS Max Concurrency
h200-8.52.1024.960
1,048,576.0
tensor
8 $37.37 1.475 Launch
h200-8.52.1024.960.nvlink
1,048,576.0
tensor
8 $37.37 1.475 Launch

Related models

Need help?

Contact our dedicated neural networks support team at support@immers.cloud or send your request to the sales department at sale@immers.cloud.

We use cookies and web analytics services to ensure the proper functioning of the website, analyze web-site traffic, and improve the quality of our services.
By continuing to use the website, you consent to the Privacy Policy and consent to the processing of cookie files and technical data.