PaddleOCR-VL-1.6

multimodal

PaddleOCR-VL-1.6 is a compact document processing model developed by the PaddlePaddle team (Baidu). Its architecture is fully inherited from PaddleOCR-VL-1.5 and consists of four key components. The visual encoder is implemented in the NaViT style (Native Resolution Visual Encoder)—a transformer with dynamic resolution that processes image patches at their original scale without forced resizing, which is critical for preserving fine details in documents. The connector is an Adaptive MLP Connector that projects visual features into the language model's embedding space. The text model is the compact ERNIE-4.5-0.3B from Baidu, ensuring the generation of structured text output. The full document processing pipeline additionally utilizes a fourth element, PP-DocLayoutV3, a layout analysis model that performs multi-point region localization, supporting perspective distortion, curved pages, and non-standard layouts.

The key idea behind version 1.6 is not an increase in parameters or a change in architecture, but rather targeted work on "Under-Optimized Regions"—data areas where the previous model demonstrated unstable predictions. To address this, a specialized Data Engine was developed to identify three types of problematic zones: boundary-fragile (unstable predictions with minimal visual changes), coverage-sparse (rare scenarios with insufficient representation in the training data), and unreliable-supervision (erroneous annotations in the training set). Subsequently, a progressive post-training recipe (CPT → SFT → RL/GRPO) is applied, where each stage works with a specific data subset depending on its reliability and training value.

Unlike universal multimodal LLMs (Gemini, GPT, Qwen-VL), PaddleOCR-VL-1.6 is narrowly specialized in OCR and document processing tasks. As a result, with only 0.9B parameters, it outperforms models that are 80–260 times larger, achieving the best result on OmniDocBench v1.6 (a comprehensive evaluation of PDF document parsing: text, formulas, tables, reading order) at 96.33%, surpassing MinerU2.5-Pro (95.75%), GLM-OCR (95.22%), Gemini 3 Pro (92.91%), Qwen3-VL-235B (89.78%), and GPT-5.2 (86.59%). It also secures 1st place on Real5-OmniDocBench (parsing under real-world conditions: scanning, deformation, screen captures, lighting variations, tilt) with a score of 93.19%.

PaddleOCR-VL-1.6 is designed for a wide range of document-related tasks and is most effective where high accuracy is required under limited computational resources. The basic scenario involves converting PDFs, scans, and photos into Markdown/JSON while preserving structure, reading order, tables, and formulas. It further excels in processing complex tables and financial documents—recognizing tables with merged cells, nested structures, or no borders, including financial reports, invoices, registration forms, and statistical tables, with support for table stitching across page breaks and restoration of header hierarchies. Additionally, it handles specialized elements such as recognizing mathematical formulas (outputting in LaTeX), parsing charts and diagrams into tabular data, and reading round, oval, or rectangular stamps and seals even with text overlap and low contrast. Furthermore, it supports end-to-end text spotting—simultaneous localization and recognition of text on IDs, ancient manuscripts, advertising posters, dialogue screenshots, signs, and multilingual images. All of this is available under the open-source Apache 2.0 license.

To achieve the most effective results, it is recommended to use the model in conjunction with PP-DocLayout. For more details on this, as well as specifics on how to properly format and direct queries to the model, please refer to the following link: https://docs.vllm.ai/projects/recipes/en/latest/PaddlePaddle/PaddleOCR-VL.html


Announce Date: 27.05.2026
Parameters: 959M
Context: 132K
Layers: 18
Attention Type: Full Attention
Developer: Baidu, Inc.
Transformers Version: 4.55.0
vLLM Version: >=0.11.1
License: Apache 2.0

Public endpoint

Use our pre-built public endpoints for free to test inference and explore PaddleOCR-VL-1.6 capabilities. You can obtain an API access token on the token management page after registration and verification.
Model Name Context Type GPU Status Link
There are no public endpoints for this model yet.

Private server

Rent your own physically dedicated instance with hourly or long-term monthly billing.

We recommend deploying private instances in the following scenarios:

  • maximize endpoint performance,
  • enable full context for long sequences,
  • ensure top-tier security for data processing in an isolated, dedicated environment,
  • use custom weights, such as fine-tuned models or LoRA adapters.

Recommended server configurations for hosting PaddleOCR-VL-1.6

Prices:
Name GPU Price, hour TPS Max Concurrency
teslaa2-1.16.32.160
131,072.0
1 $0.38 4.305 Launch
teslaa10-1.16.32.160
131,072.0
1 $0.53 7.535 Launch
rtx3080-1.16.32.160
131,072.0
1 $0.57 2.264 Launch
rtx3090-1.16.24.160
131,072.0
1 $0.83 7.968 Launch
rtx4090-1.16.32.160
131,072.0
1 $1.02 7.952 Launch
rtxa5000-2.16.64.160.nvlink
131,072.0
tensor
2 $1.23 15.268 Launch
rtx5090-1.16.64.160
131,072.0
1 $1.59 11.154 Launch
teslaa100-1.16.64.160
131,072.0
1 $2.37 30.739 Launch
h100-1.16.64.160
131,072.0
1 $3.83 30.711 Launch
h100nvl-1.16.96.160
131,072.0
1 $4.11 36.398 Launch
teslaa100-2.24.96.160.nvlink
131,072.0
tensor
2 $4.61 61.677 Launch
h200-1.16.128.160
131,072.0
1 $4.74 55.498 Launch
h200-2.24.256.160.nvlink
131,072.0
tensor
2 $9.40 111.194 Launch
Prices:
Name GPU Price, hour TPS Max Concurrency
teslaa2-1.16.32.160
131,072.0
1 $0.38 4.106 Launch
teslaa10-1.16.32.160
131,072.0
1 $0.53 7.337 Launch
rtx3080-1.16.32.160
131,072.0
1 $0.57 2.066 Launch
rtx3090-1.16.24.160
131,072.0
1 $0.83 7.770 Launch
rtx4090-1.16.32.160
131,072.0
1 $1.02 7.754 Launch
rtxa5000-2.16.64.160.nvlink
131,072.0
tensor
2 $1.23 15.070 Launch
rtx5090-1.16.64.160
131,072.0
1 $1.59 10.955 Launch
teslaa100-1.16.64.160
131,072.0
1 $2.37 30.541 Launch
h100-1.16.64.160
131,072.0
1 $3.83 30.512 Launch
h100nvl-1.16.96.160
131,072.0
1 $4.11 36.200 Launch
teslaa100-2.24.96.160.nvlink
131,072.0
tensor
2 $4.61 61.479 Launch
h200-1.16.128.160
131,072.0
1 $4.74 55.299 Launch
h200-2.24.256.160.nvlink
131,072.0
tensor
2 $9.40 110.995 Launch
Prices:
Name GPU Price, hour TPS Max Concurrency
teslaa2-1.16.32.160
131,072.0
1 $0.38 3.710 Launch
teslaa10-1.16.32.160
131,072.0
1 $0.53 6.940 Launch
rtx3080-1.16.32.160
131,072.0
1 $0.57 1.669 Launch
rtx3090-1.16.24.160
131,072.0
1 $0.83 7.373 Launch
rtx4090-1.16.32.160
131,072.0
1 $1.02 7.357 Launch
rtxa5000-2.16.64.160.nvlink
131,072.0
tensor
2 $1.23 14.673 Launch
rtx5090-1.16.64.160
131,072.0
1 $1.59 10.558 Launch
teslaa100-1.16.64.160
131,072.0
1 $2.37 30.144 Launch
h100-1.16.64.160
131,072.0
1 $3.83 30.116 Launch
h100nvl-1.16.96.160
131,072.0
1 $4.11 35.803 Launch
teslaa100-2.24.96.160.nvlink
131,072.0
tensor
2 $4.61 61.082 Launch
h200-1.16.128.160
131,072.0
1 $4.74 54.902 Launch
h200-2.24.256.160.nvlink
131,072.0
tensor
2 $9.40 110.598 Launch

Related models

Need help?

Contact our dedicated neural networks support team at nn@immers.cloud or send your request to the sales department at sale@immers.cloud.

We use cookies and web analytics services to ensure the proper functioning of the website, analyze web-site traffic, and improve the quality of our services.
By continuing to use the website, you consent to the Privacy Policy and consent to the processing of cookie files and technical data.