DeepSeek-V4-Flash-Vision-Exp

reasoning
multimodal
coding

DeepSeek-V4-Flash-Vision-Exp is the first open multimodal model from DeepSeek AI (not counting OCR models), built on the DeepSeek-V4-Flash language architecture and augmented with a visual encoder. The model belongs to hybrid expert systems (MoE) with a total parameter count of about 305 billion (up from 284 due to DSpark), of which only 13 billion are activated per token during inference—this ensures high speed and computational efficiency. It supports a context window of up to 1 million tokens and uses mixed FP4+FP8 precision: expert parameters are stored in FP4, while most other weights are in FP8, which significantly reduces GPU memory requirements without loss of quality.

The architectural foundation of the model is a combination of two sparse attention mechanisms—CSA (Compressed Sparse Attention) and HCA (Highly Compressed Attention). CSA compresses context tokens by a factor of 4, turning them into compact vectors, and then uses the Lightning Indexer to select the most relevant compressed blocks for attention computation. In HCA, the ratio reaches 1:128, enabling full global attention on a highly compressed sequence. Additionally, each layer retains a sliding window of 128 uncompressed tokens for accurate perception of local context. This combination reduces KV cache memory consumption by approximately 90% compared to classical full attention. In addition, the model includes the DSpark mechanism, which accelerates generation by predicting several tokens ahead (speculative decoding), increasing inference speed by 60–85%. Manifold-Constrained Hyper-Connections (mHC) are also integrated, enhancing signal propagation between layers.

The model's visual encoder is implemented on the basis of a classic Vision Transformer (ViT)—a standard industry solution that has already been proven in many multimodal models. It does not use experimental approaches, such as the "DeepEncoder V2" in the specialized DeepSeek-OCR model, which employs a special causal processing of images from coarse to fine. The developers chose a reliable ViT with 32 layers, a hidden size of 1024, and 16 attention heads, supplemented by a two-layer adapter for interfacing with the text part. This approach ensures fast integration of visual understanding into the language model.

According to benchmark results, the model demonstrates impressive performance. In text agentic tests, it scores 83.9 on Terminal Bench 2.1 (testing the ability to work with a terminal and command line) and 75.9 on Toolathlon-Verified (assessing the ability to use external tools). In multimodal agentic tasks, Vision-Exp shows an increase of almost 40% compared to the previous version DeepSeek-V4-Flash-0731 on the ApexBench benchmark (Pass@1)—36.5 versus 26.2. On the comprehensive test Agents' Last Exam, the model scores 27.3 points, surpassing Opus-4.8 (25.7). It also achieves 64.3 on Chartography (a test of understanding charts and diagrams) and 35.0 on ZeroBench (Pass@5)—a benchmark that checks the ability to solve new tasks without fine-tuning. Thus, the model not only improves strong text skills but also significantly surpasses its predecessors, as it is able to solve tasks that require simultaneous processing of visual and textual information.

Thanks to its architecture and performance, the model is ideal for automatic image analysis—diagrams, interface screenshots, scientific graphs, documents with illustrations. Its strong agentic capabilities allow creating AI assistants that "see" the screen and perform actions in graphical applications (e.g., for software testing automation or web scraping). The long context window makes it possible to process entire books, project codebases, or legal cases in a single query. The model also performs excellently as an intelligent developer assistant, helping with code writing, debugging, and refactoring, and in business environments—with extracting data from complex reports and presentations. It is a universal tool for researchers, developers, and analysts working with multimodal data.


Announce Date: 31.08.2026
Parameters: 305B
Experts: 256
Activated at inference: 13B
Context: 1049K
Layers: 43, using full attention: 21
Attention Type: Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA)
Developer: DeepSeek
Transformers Version: 5.0.0
vLLM Version: >=0.29.0
License: MIT

Public endpoint

Use our pre-built public endpoints for free to test inference and explore DeepSeek-V4-Flash-Vision-Exp capabilities. You can obtain an API access token on the token management page after registration and verification.
Model Name Context Type GPU Status Link
There are no public endpoints for this model yet.

Private server

Rent your own physically dedicated instance with hourly or long-term monthly billing.

We recommend deploying private instances in the following scenarios:

  • maximize endpoint performance,
  • enable full context for long sequences,
  • ensure top-tier security for data processing in an isolated, dedicated environment,
  • use custom weights, such as fine-tuned models or LoRA adapters.

Recommended server configurations for hosting DeepSeek-V4-Flash-Vision-Exp

Prices:
Name GPU Price, hour TPS Max Concurrency
h200-2.24.256.320
1,048,576.0
tensor
2 $9.42 6.499 Launch
h200-2.24.256.320.nvlink
1,048,576.0
tensor
2 $9.42 6.499 Launch
h100-3.32.384.320
1,048,576.0
pipeline
3 $11.74 5.557 Launch
h100nvl-3.24.384.480
1,048,576.0
pipeline
3 $12.38 10.328 Launch
h100-4.16.256.480
1,048,576.0
tensor
4 $14.99 4.193 Launch
h100nvl-4.32.384.480
1,048,576.0
tensor
4 $16.23 5.960 Launch

Related models

Need help?

Contact our dedicated neural networks support team at support@immers.cloud or send your request to the sales department at sale@immers.cloud.

We use cookies and web analytics services to ensure the proper functioning of the website, analyze web-site traffic, and improve the quality of our services.
By continuing to use the website, you consent to the Privacy Policy and consent to the processing of cookie files and technical data.