Qwen3-Coder-30B-A3B-Instruct is an outstanding example of a high-quality large language model with advanced specialization in programming. This Mixture-of-Experts model has 30.5 billion total parameters, of which only 3.3 billion are activated per token, and out of 128 experts, only 8 are activated per token. The model comprises 48 hidden layers with grouped query attention (32 heads for Q and 4 for KV), delivering exceptional processing efficiency with minimal computational resource consumption. Native support for a 262,144-token context window—expandable up to 1 million tokens via Yarn—makes the model ideal for working with large code repositories within complex projects.
The key unique feature of Qwen3-Coder-30B-A3B-Instruct lies in its superior agent capabilities. The model does not merely generate code; it autonomously interacts with development tools, executes multi-step programming tasks, and is capable of solving complex problems without human intervention. On the LiveCodeBench v6 benchmark, the model achieves an impressive 66.0%, significantly outperforming the base version Qwen3-30B-A3B (57.4%). In AIME25 tasks (advanced mathematics for programming), it demonstrates 85.0% accuracy, surpassing Gemini-2.5-Flash-Thinking (72.0%) and confidently competing with much larger models. The model outperforms DeepSeek V3 on most coding tasks and delivers agent workflow performance comparable to Claude Sonnet 4, a remarkable achievement for an open-source solution.
Qwen3-Coder-30B-A3B-Instruct unlocks entirely new possibilities in software development. The model is integrated with popular agent-based programming platforms, including Qwen Code, CLINE, Roo Code, and Kilo Code, offering a unified function-calling format for seamless operation within CI/CD pipelines. Support for 358 programming languages makes it a universal reference tool for developers. The model particularly excels in repository-scale understanding scenarios, where it can analyze and modify massive codebases, automatically refactor legacy code, and create complex full-stack applications with minimal developer intervention.
| Model Name | Context | Type | GPU | Status | Link |
|---|
There are no public endpoints for this model yet.
Rent your own physically dedicated instance with hourly or long-term monthly billing.
We recommend deploying private instances in the following scenarios:
| Name | GPU | TPS | Max Concurrency | |||
|---|---|---|---|---|---|---|
262,144.0 tensor |
4 | $1.26 | 1.194 | Launch | ||
262,144.0 pipeline |
3 | $1.34 | 1.627 | Launch | ||
262,144.0 tensor |
4 | $1.57 | 2.405 | Launch | ||
262,144.0 pipeline |
3 | $2.29 | 138.17 | 1.761 | Launch | |
262,144.0 |
1 | $2.37 | 281.00 | 2.274 | Launch | |
262,144.0 pipeline |
3 | $2.83 | 1.744 | Launch | ||
262,144.0 tensor |
4 | $2.89 | 2.568 | Launch | ||
262,144.0 tensor |
2 | $2.93 | 1.528 | Launch | ||
262,144.0 tensor |
4 | $3.01 | 2.568 | Launch | ||
262,144.0 tensor |
4 | $3.60 | 2.561 | Launch | ||
262,144.0 |
1 | $3.83 | 152.96 | 2.244 | Launch | |
262,144.0 |
1 | $4.11 | 202.68 | 2.777 | Launch | |
262,144.0 tensor |
2 | $4.61 | 5.200 | Launch | ||
262,144.0 |
1 | $4.74 | 4.568 | Launch | ||
262,144.0 tensor |
2 | $9.40 | 9.842 | Launch | ||
| Name | GPU | TPS | Max Concurrency | |||
|---|---|---|---|---|---|---|
262,144.0 pipeline |
3 | $1.34 | 1.138 | Launch | ||
262,144.0 tensor |
4 | $1.62 | 1.962 | Launch | ||
262,144.0 pipeline |
6 | $1.65 | 1.792 | Launch | ||
262,144.0 pipeline |
3 | $2.29 | 154.00 | 1.261 | Launch | |
262,144.0 |
1 | $2.37 | 156.50 | 1.755 | Launch | |
262,144.0 pipeline |
3 | $2.83 | 148.67 | 1.269 | Launch | |
262,144.0 tensor |
4 | $2.89 | 2.124 | Launch | ||
262,144.0 tensor |
4 | $3.01 | 2.124 | Launch | ||
262,144.0 tensor |
4 | $3.60 | 2.118 | Launch | ||
262,144.0 |
1 | $3.83 | 147.58 | 1.664 | Launch | |
262,144.0 |
1 | $4.11 | 196.88 | 2.276 | Launch | |
262,144.0 pipeline |
3 | $4.34 | 2.156 | Launch | ||
262,144.0 tensor |
2 | $4.61 | 4.666 | Launch | ||
262,144.0 |
1 | $4.74 | 3.988 | Launch | ||
262,144.0 tensor |
4 | $5.74 | 3.319 | Launch | ||
262,144.0 tensor |
2 | $9.40 | 9.308 | Launch | ||
| Name | GPU | TPS | Max Concurrency | |||
|---|---|---|---|---|---|---|
262,144.0 |
1 | $4.11 | 196.26 | 1.113 | Launch | |
262,144.0 tensor |
2 | $4.61 | 142.52 | 3.564 | Launch | |
262,144.0 |
1 | $4.74 | 2.904 | Launch | ||
262,144.0 tensor |
2 | $4.93 | 142.52 | 3.564 | Launch | |
262,144.0 tensor |
4 | $5.74 | 2.103 | Launch | ||
262,144.0 pipeline |
6 | $5.83 | 2.540 | Launch | ||
262,144.0 tensor |
8 | $6.04 | 2.095 | Launch | ||
262,144.0 tensor |
8 | $7.51 | 2.089 | Launch | ||
262,144.0 tensor |
2 | $7.84 | 148.03 | 3.570 | Launch | |
262,144.0 tensor |
2 | $9.40 | 8.180 | Launch | ||
Contact our dedicated neural networks support team at support@immers.cloud or send your request to the sales department at sale@immers.cloud.