Qwen3.8-27B is a flagship compact dense multimodal language model in the new Qwen family, capable of effectively working with text, images, and video. Despite its relatively small size, the model delivers results comparable to much larger closed and open frontier systems, making it one of the best candidates for deployment in private infrastructure.
In Qwen3.8-27B, thinking mode is enabled by default but can be disabled to speed up responses. Reasoning depth is flexibly controlled by the `reasoning_effort` parameter with three levels — `xhigh` (default, for complex analysis), `medium` (balance), and `low` (speed). A `preserve_thinking` mechanism is also implemented, preserving the reasoning chain between dialogue turns, which is critical for multi-step agentic scenarios and KV-cache reuse. The model supports Multi-Token Prediction (MTP) to accelerate generation.
The Qwen3.8-27B architecture retains the hybrid structure introduced in Qwen3.5: 64 layers organized into 16 repeating blocks, each containing three Gated DeltaNet blocks and one Gated Attention block. Gated DeltaNet is an efficient linear attention architecture that provides high throughput when processing long contexts. Gated Attention is full attention responsible for establishing important dependencies in the data and maintaining overall global understanding. This interleaving allows the model to efficiently process sequences up to 262,144 tokens natively, with the ability to expand to 1,000,000 tokens.
The key innovation of Qwen3.8-27B is not so much an architectural change — it remains a successor to Qwen3.5 — as large-scale post-training that delivered a significant leap in programming, professional tasks, research work, and long-horizon agentic scenarios. The model’s success is largely due to a shared post-training pipeline with the top model in the lineup — the 2.4-trillion-parameter Qwen3.8-2.4T-A95B: both models underwent scalable reinforcement learning (Scalable RL) in “million-scale agentic environments” with gradually increasing task complexity, allowing the “smaller” 27-billion model to acquire long-term planning, external tool use, and self-correction skills usually available only to trillion-parameter giants. As a result, compared with Qwen3.6-27B, the model shows substantial progress on nearly all key benchmarks: on the SWE-bench Pro agentic coding benchmark it scores 61.7% vs 53.5%; on the Terminal Bench 2.1 terminal coding benchmark — 73.0% vs 63.4%; on the specialized engineering benchmark QwenSWEBench — 79.0% vs 49.3%; on the JobBench professional benchmark — 33.4% vs 21.8%; and on the CoWorkBench long-horizon office benchmark — 70.7% vs 61.0%.
Thanks to its combination of advanced agentic capabilities, multimodality, and long context, Qwen3.8-27B opens up broad application opportunities. In agentic systems and automation, the model is suitable for building AI agents that can autonomously plan and execute complex multi-step tasks — from terminal and web browser control to working with mobile apps and desktop environments. In software development, outstanding results on SWE-bench Pro and LiveCodeBench make it an excellent choice for code completion, refactoring, repository generation, and other software engineering tasks. For scientific research and professional work, high scores on GPQA Diamond and JobBench allow the model to be used for analyzing scientific literature and solving complex problems in law, finance, medicine, and other fields. Native image and video support opens up opportunities for multimodal applications — analyzing charts, documents, long videos, visual web interface development, and other tasks at the intersection of text and visual information. The model is released under the Apache 2.0 license and easily integrates into existing stacks through support for popular frameworks: Transformers, vLLM, SGLang, TokenSpeed.
| Model Name | Context | Type | GPU | Status | Link |
|---|
There are no public endpoints for this model yet.
Rent your own physically dedicated instance with hourly or long-term monthly billing.
We recommend deploying private instances in the following scenarios:
| Name | GPU | TPS | Max Concurrency | |||
|---|---|---|---|---|---|---|
262,144.0 tensor |
4 | $1.26 | 1.307 | Launch | ||
262,144.0 pipeline |
3 | $1.34 | 1.765 | Launch | ||
262,144.0 tensor |
2 | $1.56 | 1.069 | Launch | ||
262,144.0 tensor |
2 | $1.56 | 1.069 | Launch | ||
262,144.0 tensor |
4 | $1.57 | 3.120 | Launch | ||
262,144.0 tensor |
2 | $1.92 | 1.064 | Launch | ||
262,144.0 |
1 | $2.37 | 3.100 | Launch | ||
262,144.0 tensor |
2 | $2.93 | 1.961 | Launch | ||
262,144.0 |
1 | $3.83 | 3.096 | Launch | ||
262,144.0 |
1 | $4.11 | 3.889 | Launch | ||
262,144.0 tensor |
2 | $4.61 | 7.445 | Launch | ||
262,144.0 |
1 | $4.74 | 6.551 | Launch | ||
262,144.0 tensor |
2 | $9.40 | 14.378 | Launch | ||
| Name | GPU | TPS | Max Concurrency | |||
|---|---|---|---|---|---|---|
262,144.0 pipeline |
3 | $1.34 | 1.243 | Launch | ||
262,144.0 tensor |
4 | $1.57 | 2.548 | Launch | ||
262,144.0 tensor |
6 | $1.65 | 1.334 | Launch | ||
262,144.0 pipeline |
3 | $2.29 | 1.405 | Launch | ||
262,144.0 |
1 | $2.37 | 2.532 | Launch | ||
262,144.0 pipeline |
3 | $2.83 | 1.399 | Launch | ||
262,144.0 tensor |
4 | $2.89 | 2.791 | Launch | ||
262,144.0 tensor |
2 | $2.93 | 1.390 | Launch | ||
262,144.0 tensor |
4 | $3.01 | 2.791 | Launch | ||
262,144.0 tensor |
4 | $3.60 | 2.782 | Launch | ||
262,144.0 |
1 | $3.83 | 2.528 | Launch | ||
262,144.0 |
1 | $4.11 | 3.321 | Launch | ||
262,144.0 tensor |
2 | $4.61 | 6.874 | Launch | ||
262,144.0 |
1 | $4.74 | 5.983 | Launch | ||
262,144.0 tensor |
2 | $9.40 | 13.807 | Launch | ||
| Name | GPU | TPS | Max Concurrency | |||
|---|---|---|---|---|---|---|
262,144.0 tensor |
4 | $1.75 | 1.113 | Launch | ||
262,144.0 |
1 | $2.50 | 1.107 | Launch | ||
262,144.0 tensor |
4 | $2.97 | 1.357 | Launch | ||
262,144.0 tensor |
4 | $3.01 | 1.357 | Launch | ||
262,144.0 tensor |
4 | $3.68 | 1.347 | Launch | ||
262,144.0 |
1 | $3.95 | 1.103 | Launch | ||
262,144.0 |
1 | $4.11 | 1.896 | Launch | ||
262,144.0 pipeline |
3 | $4.34 | 1.282 | Launch | ||
262,144.0 tensor |
2 | $4.61 | 5.443 | Launch | ||
262,144.0 |
1 | $4.74 | 4.558 | Launch | ||
262,144.0 tensor |
4 | $5.74 | 3.144 | Launch | ||
262,144.0 tensor |
2 | $9.40 | 12.376 | Launch | ||
Contact our dedicated neural networks support team at support@immers.cloud or send your request to the sales department at sale@immers.cloud.