Qwen3.8-2.4T-A95B is the first open model at the Qwen-Max level from the Qwen team. It is built on an MoE architecture with a total of 2.4 trillion parameters, yet only about 95 billion parameters are activated at each inference step. This allows the model to deliver outstanding performance while using computational resources efficiently. The model's native context is 262,144 tokens, but it can be extended to 1 million tokens. The model is designed for complex programming tasks, office work, research, and multi-step agentic scenarios.
At the core of the architecture is a hybrid attention mechanism introduced in the Qwen3.5 series. The model consists of 92 layers grouped into 23 repeating blocks. In each block, the first three layers are Gated DeltaNet (linear attention) followed by MoE, then one Gated Attention layer (full attention) followed by MoE. Each MoE layer contains 512 experts, of which 10 routed experts are activated at each step. This combination of linear and full attention allows the model to work effectively with long contexts while maintaining deep understanding.
On key benchmarks, the model demonstrates outstanding results. In PaperBench (evaluation of reproducing scientific papers), Qwen3.8-2.4T-A95B scored 93.0 points, outperforming GPT 5.6 Sol (90.5) and Claude Opus 4.8 (80.3), taking first place. In PLawBench (legal tasks), the result of 73.2 is also the best among the presented models. The model leads in FrontierSWE (73.5), WideSearch (81.9), HealthBench (60.2), and IFBench (82.8), surpassing both closed flagship models and the previous version Qwen3.7-Max. In long-context MRCR v2 256K tasks (needle-in-a-haystack search), the model achieved 92.9%, close to the leader and significantly outperforming Qwen3.7-Max (86.7%). In programming, excellent results in Terminal Bench 2.1 (86.6), AndroidBench (75.1), and the internal benchmarks QwenSWEBench (80.7) and QwenReactBench (1724 Elo) confirm its status as one of the strongest models for development.
A key advantage of the model is its ability to solve complex, end-to-end tasks from start to finish. It excels at autonomous planning, processing environmental feedback, and executing long multi-step scenarios. The model also supports flexible reasoning depth control via the reasoning_effort parameter and preservation of reasoning history (preserve_thinking), allowing computational costs to be adapted to the complexity of a given task. It is worth noting, however, that unlike the open models of the 3.5 and 3.6 series, this model is not multimodal and processes text only.
Use cases span a wide range of domains: from programming and software development — including the full cycle of creating projects from scratch, autonomous code and test generation, work with pull requests and CI checks, where the model shows outstanding results in agentic tasks — to scientific research, where it helps reproduce and improve experiments from scientific papers, as well as conduct analysis and synthesis of new ideas. In addition, the model is effective in long-term agentic tasks requiring planning, execution, and iterative refinement of solutions in complex dynamic environments, and in professional office work, including answering complex document-based questions (law, finance, medicine), report generation, analytics, and expert reasoning across various subject areas.
| Model Name | Context | Type | GPU | Status | Link |
|---|
There are no public endpoints for this model yet.
Rent your own physically dedicated instance with hourly or long-term monthly billing.
We recommend deploying private instances in the following scenarios:
There are no configurations for this model, context and quantization yet.
There are no configurations for this model, context and quantization yet.
There are no configurations for this model, context and quantization yet.
Contact our dedicated neural networks support team at support@immers.cloud or send your request to the sales department at sale@immers.cloud.