DeepSeek-V4-Flash-Vision-Exp is the first open multimodal model from DeepSeek AI (not counting OCR models), built on the DeepSeek-V4-Flash language architecture and augmented with a visual encoder. The model belongs to hybrid expert systems (MoE) with a total parameter count of about 305 billion (up from 284 due to DSpark), of which only 13 billion are activated per token during inference—this ensures high speed and computational efficiency. It supports a context window of up to 1 million tokens and uses mixed FP4+FP8 precision: expert parameters are stored in FP4, while most other weights are in FP8, which significantly reduces GPU memory requirements without loss of quality.
The architectural foundation of the model is a combination of two sparse attention mechanisms—CSA (Compressed Sparse Attention) and HCA (Highly Compressed Attention). CSA compresses context tokens by a factor of 4, turning them into compact vectors, and then uses the Lightning Indexer to select the most relevant compressed blocks for attention computation. In HCA, the ratio reaches 1:128, enabling full global attention on a highly compressed sequence. Additionally, each layer retains a sliding window of 128 uncompressed tokens for accurate perception of local context. This combination reduces KV cache memory consumption by approximately 90% compared to classical full attention. In addition, the model includes the DSpark mechanism, which accelerates generation by predicting several tokens ahead (speculative decoding), increasing inference speed by 60–85%. Manifold-Constrained Hyper-Connections (mHC) are also integrated, enhancing signal propagation between layers.
The model's visual encoder is implemented on the basis of a classic Vision Transformer (ViT)—a standard industry solution that has already been proven in many multimodal models. It does not use experimental approaches, such as the "DeepEncoder V2" in the specialized DeepSeek-OCR model, which employs a special causal processing of images from coarse to fine. The developers chose a reliable ViT with 32 layers, a hidden size of 1024, and 16 attention heads, supplemented by a two-layer adapter for interfacing with the text part. This approach ensures fast integration of visual understanding into the language model.
According to benchmark results, the model demonstrates impressive performance. In text agentic tests, it scores 83.9 on Terminal Bench 2.1 (testing the ability to work with a terminal and command line) and 75.9 on Toolathlon-Verified (assessing the ability to use external tools). In multimodal agentic tasks, Vision-Exp shows an increase of almost 40% compared to the previous version DeepSeek-V4-Flash-0731 on the ApexBench benchmark (Pass@1)—36.5 versus 26.2. On the comprehensive test Agents' Last Exam, the model scores 27.3 points, surpassing Opus-4.8 (25.7). It also achieves 64.3 on Chartography (a test of understanding charts and diagrams) and 35.0 on ZeroBench (Pass@5)—a benchmark that checks the ability to solve new tasks without fine-tuning. Thus, the model not only improves strong text skills but also significantly surpasses its predecessors, as it is able to solve tasks that require simultaneous processing of visual and textual information.
Thanks to its architecture and performance, the model is ideal for automatic image analysis—diagrams, interface screenshots, scientific graphs, documents with illustrations. Its strong agentic capabilities allow creating AI assistants that "see" the screen and perform actions in graphical applications (e.g., for software testing automation or web scraping). The long context window makes it possible to process entire books, project codebases, or legal cases in a single query. The model also performs excellently as an intelligent developer assistant, helping with code writing, debugging, and refactoring, and in business environments—with extracting data from complex reports and presentations. It is a universal tool for researchers, developers, and analysts working with multimodal data.
| Model Name | Context | Type | GPU | Status | Link |
|---|
There are no public endpoints for this model yet.
Rent your own physically dedicated instance with hourly or long-term monthly billing.
We recommend deploying private instances in the following scenarios:
| Name | GPU | TPS | Max Concurrency | |||
|---|---|---|---|---|---|---|
1,048,576.0 tensor |
2 | $9.42 | 6.499 | Launch | ||
1,048,576.0 tensor |
2 | $9.42 | 6.499 | Launch | ||
1,048,576.0 pipeline |
3 | $11.74 | 5.557 | Launch | ||
1,048,576.0 pipeline |
3 | $12.38 | 10.328 | Launch | ||
1,048,576.0 tensor |
4 | $14.99 | 4.193 | Launch | ||
1,048,576.0 tensor |
4 | $16.23 | 5.960 | Launch | ||
Contact our dedicated neural networks support team at support@immers.cloud or send your request to the sales department at sale@immers.cloud.