Qwen3.8-27B

reasoning
multimodal
coding

Qwen3.8-27B is a flagship compact dense multimodal language model in the new Qwen family, capable of effectively working with text, images, and video. Despite its relatively small size, the model delivers results comparable to much larger closed and open frontier systems, making it one of the best candidates for deployment in private infrastructure.

In Qwen3.8-27B, thinking mode is enabled by default but can be disabled to speed up responses. Reasoning depth is flexibly controlled by the `reasoning_effort` parameter with three levels — `xhigh` (default, for complex analysis), `medium` (balance), and `low` (speed). A `preserve_thinking` mechanism is also implemented, preserving the reasoning chain between dialogue turns, which is critical for multi-step agentic scenarios and KV-cache reuse. The model supports Multi-Token Prediction (MTP) to accelerate generation.

The Qwen3.8-27B architecture retains the hybrid structure introduced in Qwen3.5: 64 layers organized into 16 repeating blocks, each containing three Gated DeltaNet blocks and one Gated Attention block. Gated DeltaNet is an efficient linear attention architecture that provides high throughput when processing long contexts. Gated Attention is full attention responsible for establishing important dependencies in the data and maintaining overall global understanding. This interleaving allows the model to efficiently process sequences up to 262,144 tokens natively, with the ability to expand to 1,000,000 tokens.

The key innovation of Qwen3.8-27B is not so much an architectural change — it remains a successor to Qwen3.5 — as large-scale post-training that delivered a significant leap in programming, professional tasks, research work, and long-horizon agentic scenarios. The model’s success is largely due to a shared post-training pipeline with the top model in the lineup — the 2.4-trillion-parameter Qwen3.8-2.4T-A95B: both models underwent scalable reinforcement learning (Scalable RL) in “million-scale agentic environments” with gradually increasing task complexity, allowing the “smaller” 27-billion model to acquire long-term planning, external tool use, and self-correction skills usually available only to trillion-parameter giants. As a result, compared with Qwen3.6-27B, the model shows substantial progress on nearly all key benchmarks: on the SWE-bench Pro agentic coding benchmark it scores 61.7% vs 53.5%; on the Terminal Bench 2.1 terminal coding benchmark — 73.0% vs 63.4%; on the specialized engineering benchmark QwenSWEBench — 79.0% vs 49.3%; on the JobBench professional benchmark — 33.4% vs 21.8%; and on the CoWorkBench long-horizon office benchmark — 70.7% vs 61.0%.

Thanks to its combination of advanced agentic capabilities, multimodality, and long context, Qwen3.8-27B opens up broad application opportunities. In agentic systems and automation, the model is suitable for building AI agents that can autonomously plan and execute complex multi-step tasks — from terminal and web browser control to working with mobile apps and desktop environments. In software development, outstanding results on SWE-bench Pro and LiveCodeBench make it an excellent choice for code completion, refactoring, repository generation, and other software engineering tasks. For scientific research and professional work, high scores on GPQA Diamond and JobBench allow the model to be used for analyzing scientific literature and solving complex problems in law, finance, medicine, and other fields. Native image and video support opens up opportunities for multimodal applications — analyzing charts, documents, long videos, visual web interface development, and other tasks at the intersection of text and visual information. The model is released under the Apache 2.0 license and easily integrates into existing stacks through support for popular frameworks: Transformers, vLLM, SGLang, TokenSpeed.


Announce Date: 05.08.2026
Parameters: 28B
Context: 263K
Layers: 64, using full attention: 16
Attention Type: Hybrid Attention
Mamba Type: Gated Delta Net
Developer: Qwen
Transformers Version: 5.8.0.dev0
vLLM Version: >=0.17.0
License: Apache 2.0

Public endpoint

Use our pre-built public endpoints for free to test inference and explore Qwen3.8-27B capabilities. You can obtain an API access token on the token management page after registration and verification.
Model Name Context Type GPU Status Link
There are no public endpoints for this model yet.

Private server

Rent your own physically dedicated instance with hourly or long-term monthly billing.

We recommend deploying private instances in the following scenarios:

  • maximize endpoint performance,
  • enable full context for long sequences,
  • ensure top-tier security for data processing in an isolated, dedicated environment,
  • use custom weights, such as fine-tuned models or LoRA adapters.

Recommended server configurations for hosting Qwen3.8-27B

Prices:
Name GPU Price, hour TPS Max Concurrency
teslaa2-4.32.128.160
262,144.0
tensor
4 $1.26 1.307 Launch
teslaa10-3.16.96.160
262,144.0
pipeline
3 $1.34 1.765 Launch
rtx3090-2.16.64.160
262,144.0
tensor
2 $1.56 1.069 Launch
rtx3090-2.16.64.160.nvlink
262,144.0
tensor
2 $1.56 1.069 Launch
teslaa10-4.12.48.160
262,144.0
tensor
4 $1.57 3.120 Launch
rtx4090-2.16.64.160
262,144.0
tensor
2 $1.92 1.064 Launch
teslaa100-1.16.64.160
262,144.0
1 $2.37 3.100 Launch
rtx5090-2.16.64.160
262,144.0
tensor
2 $2.93 1.961 Launch
h100-1.16.64.160
262,144.0
1 $3.83 3.096 Launch
h100nvl-1.16.96.160
262,144.0
1 $4.11 3.889 Launch
teslaa100-2.24.96.160.nvlink
262,144.0
tensor
2 $4.61 7.445 Launch
h200-1.16.128.160
262,144.0
1 $4.74 6.551 Launch
h200-2.24.256.160.nvlink
262,144.0
tensor
2 $9.40 14.378 Launch
Prices:
Name GPU Price, hour TPS Max Concurrency
teslaa10-3.16.96.160
262,144.0
pipeline
3 $1.34 1.243 Launch
teslaa10-4.12.48.160
262,144.0
tensor
4 $1.57 2.548 Launch
teslaa2-6.32.128.160
262,144.0
tensor
6 $1.65 1.334 Launch
rtx3090-3.16.96.160
262,144.0
pipeline
3 $2.29 1.405 Launch
teslaa100-1.16.64.160
262,144.0
1 $2.37 2.532 Launch
rtx4090-3.16.96.160
262,144.0
pipeline
3 $2.83 1.399 Launch
rtx3090-4.16.64.160
262,144.0
tensor
4 $2.89 2.791 Launch
rtx5090-2.16.64.160
262,144.0
tensor
2 $2.93 1.390 Launch
rtx3090-4.16.128.160.nvlink
262,144.0
tensor
4 $3.01 2.791 Launch
rtx4090-4.16.64.160
262,144.0
tensor
4 $3.60 2.782 Launch
h100-1.16.64.160
262,144.0
1 $3.83 2.528 Launch
h100nvl-1.16.96.160
262,144.0
1 $4.11 3.321 Launch
teslaa100-2.24.96.160.nvlink
262,144.0
tensor
2 $4.61 6.874 Launch
h200-1.16.128.160
262,144.0
1 $4.74 5.983 Launch
h200-2.24.256.160.nvlink
262,144.0
tensor
2 $9.40 13.807 Launch
Prices:
Name GPU Price, hour TPS Max Concurrency
teslaa10-4.16.128.160
262,144.0
tensor
4 $1.75 1.113 Launch
teslaa100-1.16.128.160
262,144.0
1 $2.50 1.107 Launch
rtx3090-4.16.96.320
262,144.0
tensor
4 $2.97 1.357 Launch
rtx3090-4.16.128.160.nvlink
262,144.0
tensor
4 $3.01 1.357 Launch
rtx4090-4.16.96.320
262,144.0
tensor
4 $3.68 1.347 Launch
h100-1.16.128.160
262,144.0
1 $3.95 1.103 Launch
h100nvl-1.16.96.160
262,144.0
1 $4.11 1.896 Launch
rtx5090-3.16.96.160
262,144.0
pipeline
3 $4.34 1.282 Launch
teslaa100-2.24.96.160.nvlink
262,144.0
tensor
2 $4.61 5.443 Launch
h200-1.16.128.160
262,144.0
1 $4.74 4.558 Launch
rtx5090-4.16.128.160
262,144.0
tensor
4 $5.74 3.144 Launch
h200-2.24.256.160.nvlink
262,144.0
tensor
2 $9.40 12.376 Launch

Related models

Need help?

Contact our dedicated neural networks support team at support@immers.cloud or send your request to the sales department at sale@immers.cloud.

We use cookies and web analytics services to ensure the proper functioning of the website, analyze web-site traffic, and improve the quality of our services.
By continuing to use the website, you consent to the Privacy Policy and consent to the processing of cookie files and technical data.