Qwen2.5-32B features 32 billion parameters, 64 layers, and a 40/8 attention head architecture, representing a significant leap in computational power and model capabilities. With support for a 128K-token context window and 8K-token generation capacity, the model can handle exceptionally complex and large-scale tasks.
Qwen2.5-32B reintroduces the 32B parameter size to the Qwen series after its absence in Qwen2, offering users a powerful alternative to the flagship 72B model with lower resource requirements. Trained on 18 trillion high-quality tokens, the model demonstrates robust performance with large datasets, expert-level knowledge in specialized domains, superior abstract reasoning capabilities, and the ability to solve problems requiring deep contextual understanding and multi-step analysis.
Qwen2.5-32B is designed for organizations and research teams that need frontier-model capabilities without the full cost of the largest models. Ideal applications include scientific research, complex software development, high-quality content creation, expert support systems in medicine and law, and as a foundation for building highly specialized AI systems.
| Model Name | Context | Type | GPU | Status | Link |
|---|
There are no public endpoints for this model yet.
Rent your own physically dedicated instance with hourly or long-term monthly billing.
We recommend deploying private instances in the following scenarios:
| Name | GPU | TPS | Max Concurrency | |||
|---|---|---|---|---|---|---|
32,768.0 tensor |
2 | $0.93 | 2.508 | Launch | ||
32,768.0 pipeline |
3 | $1.06 | 2.037 | Launch | ||
32,768.0 tensor |
4 | $1.26 | 3.654 | Launch | ||
32,768.0 tensor |
2 | $1.56 | 58.66 | 2.751 | Launch | |
32,768.0 tensor |
2 | $1.56 | 2.751 | Launch | ||
32,768.0 |
1 | $1.59 | 1.135 | Launch | ||
32,768.0 tensor |
4 | $1.82 | 1.358 | Launch | ||
32,768.0 tensor |
2 | $1.92 | 72.85 | 2.742 | Launch | |
32,768.0 |
1 | $2.37 | 59.41 | 6.644 | Launch | |
32,768.0 |
1 | $3.83 | 59.02 | 6.636 | Launch | |
32,768.0 |
1 | $4.11 | 8.235 | Launch | ||
32,768.0 tensor |
2 | $4.61 | 15.560 | Launch | ||
32,768.0 |
1 | $4.74 | 13.607 | Launch | ||
32,768.0 tensor |
2 | $9.40 | 29.487 | Launch | ||
| Name | GPU | TPS | Max Concurrency | |||
|---|---|---|---|---|---|---|
32,768.0 tensor |
4 | $1.26 | 1.329 | Launch | ||
32,768.0 pipeline |
3 | $1.34 | 2.497 | Launch | ||
32,768.0 tensor |
4 | $1.57 | 4.963 | Launch | ||
32,768.0 pipeline |
3 | $2.29 | 20.58 | 3.118 | Launch | |
32,768.0 |
1 | $2.37 | 39.51 | 4.709 | Launch | |
32,768.0 pipeline |
3 | $2.83 | 24.10 | 3.098 | Launch | |
32,768.0 tensor |
4 | $2.89 | 5.450 | Launch | ||
32,768.0 tensor |
2 | $2.93 | 2.478 | Launch | ||
32,768.0 tensor |
4 | $3.01 | 5.450 | Launch | ||
32,768.0 tensor |
4 | $3.60 | 5.432 | Launch | ||
32,768.0 |
1 | $3.83 | 41.38 | 4.641 | Launch | |
32,768.0 |
1 | $4.11 | 58.78 | 6.239 | Launch | |
32,768.0 tensor |
2 | $4.61 | 13.495 | Launch | ||
32,768.0 |
1 | $4.74 | 11.672 | Launch | ||
32,768.0 tensor |
2 | $9.40 | 27.422 | Launch | ||
| Name | GPU | TPS | Max Concurrency | |||
|---|---|---|---|---|---|---|
32,768.0 tensor |
4 | $1.76 | 1.555 | Launch | ||
32,768.0 |
1 | $2.51 | 22.26 | 1.193 | Launch | |
32,768.0 tensor |
4 | $2.97 | 2.043 | Launch | ||
32,768.0 tensor |
4 | $3.68 | 2.024 | Launch | ||
32,768.0 |
1 | $3.96 | 1.184 | Launch | ||
32,768.0 |
1 | $4.12 | 2.784 | Launch | ||
32,768.0 pipeline |
3 | $4.35 | 2.011 | Launch | ||
32,768.0 |
1 | $4.74 | 8.156 | Launch | ||
32,768.0 tensor |
2 | $4.94 | 10.015 | Launch | ||
32,768.0 tensor |
4 | $5.76 | 5.626 | Launch | ||
32,768.0 tensor |
2 | $9.41 | 23.942 | Launch | ||
Contact our dedicated neural networks support team at support@immers.cloud or send your request to the sales department at sale@immers.cloud.