Qwen3-1.7B

reasoning

Qwen3-1.7B is a model with 1.7 billion parameters that retains an architecture of 28 layers and supports a context window of 40K tokens. It uses the same attention configuration (16/8 heads) and architectural design as the smaller model in the series, but offers significantly enhanced capabilities thanks to its increased parameter count. Trained on 36 trillion tokens and supporting 119 languages, it delivers a substantial improvement in text understanding and generation quality compared to similar-sized models.

The model demonstrates strong performance in tasks requiring reasoning and contextual understanding, while maintaining high computational efficiency. Its built-in thinking and non-thinking modes allow it to adapt to different query types — from quick information retrieval to moderately complex analytical tasks. The thinking budget mechanism is especially useful for optimizing performance under variable workloads.

Qwen3-1.7B is ideally suited for desktop applications, cloud services with moderate resource requirements, and simple enterprise solutions. The model excels at document analysis, multilingual customer support, and educational applications where a balance between output quality and processing speed is essential.


Announce Date: 29.04.2025
Parameters: 2B
Context: 41K
Layers: 28
Attention Type: Full or Sliding Window Attention
Developer: Qwen
Transformers Version: 4.51.0
License: Apache 2.0

Public endpoint

Use our pre-built public endpoints for free to test inference and explore Qwen3-1.7B capabilities. You can obtain an API access token on the token management page after registration and verification.
Model Name Context Type GPU Status Link
There are no public endpoints for this model yet.

Private server

Rent your own physically dedicated instance with hourly or long-term monthly billing.

We recommend deploying private instances in the following scenarios:

  • maximize endpoint performance,
  • enable full context for long sequences,
  • ensure top-tier security for data processing in an isolated, dedicated environment,
  • use custom weights, such as fine-tuned models or LoRA adapters.

Recommended server configurations for hosting Qwen3-1.7B

Prices:
Name GPU Price, hour TPS Max Concurrency
teslaa2-1.16.32.160
40,960.0
1 $0.38 38.03 2.430 Launch
teslaa10-1.16.32.160
40,960.0
1 $0.53 163.01 4.091 Launch
rtx3080-1.16.32.160
40,960.0
1 $0.57 225.80 1.381 Launch
rtx3090-1.16.24.160
40,960.0
1 $0.83 215.17 4.314 Launch
rtx4090-1.16.32.160
40,960.0
1 $1.02 4.306 Launch
rtx3090-2.16.64.160.nvlink
40,960.0
tensor
2 $1.56 8.944 Launch
rtx5090-1.16.64.160
40,960.0
1 $1.59 5.952 Launch
teslaa100-1.16.64.160
40,960.0
1 $2.37 16.025 Launch
h100-1.16.64.160
40,960.0
1 $3.83 248.98 16.010 Launch
h100nvl-1.16.96.160
40,960.0
1 $4.11 18.935 Launch
teslaa100-2.24.96.160.nvlink
40,960.0
tensor
2 $4.61 32.366 Launch
h200-1.16.128.160
40,960.0
1 $4.74 28.758 Launch
h200-2.24.256.160.nvlink
40,960.0
tensor
2 $9.40 57.831 Launch
Prices:
Name GPU Price, hour TPS Max Concurrency
teslaa2-1.16.32.160
40,960.0
1 $0.38 59.61 2.368 Launch
teslaa10-1.16.32.160
40,960.0
1 $0.53 156.01 4.030 Launch
rtx3080-1.16.32.160
40,960.0
1 $0.57 223.50 1.319 Launch
rtx3090-1.16.24.160
40,960.0
1 $0.83 4.234 Launch
rtx4090-1.16.32.160
40,960.0
1 $1.02 4.244 Launch
rtx3090-2.16.64.160.nvlink
40,960.0
tensor
2 $1.56 8.944 Launch
rtx5090-1.16.64.160
40,960.0
1 $1.59 5.891 Launch
teslaa100-1.16.64.160
40,960.0
1 $2.37 253.91 15.963 Launch
h100-1.16.64.160
40,960.0
1 $3.83 223.52 15.949 Launch
h100nvl-1.16.96.160
40,960.0
1 $4.11 18.874 Launch
teslaa100-2.24.96.160.nvlink
40,960.0
tensor
2 $4.61 32.366 Launch
h200-1.16.128.160
40,960.0
1 $4.74 28.696 Launch
h200-2.24.256.160.nvlink
40,960.0
tensor
2 $9.40 57.831 Launch
Prices:
Name GPU Price, hour TPS Max Concurrency
teslaa2-1.16.32.160
40,960.0
1 $0.38 40.80 1.984 Launch
teslaa10-1.16.32.160
40,960.0
1 $0.53 105.73 3.646 Launch
rtx3090-1.16.24.160
40,960.0
1 $0.83 174.60 3.869 Launch
rtx3080-2.16.32.160
40,960.0
tensor
2 $0.97 2.606 Launch
rtx4090-1.16.32.160
40,960.0
1 $1.02 206.69 3.860 Launch
rtx3090-2.16.64.160.nvlink
40,960.0
tensor
2 $1.56 8.473 Launch
rtx5090-1.16.64.160
40,960.0
1 $1.59 5.507 Launch
teslaa100-1.16.64.160
40,960.0
1 $2.37 228.70 15.579 Launch
h100-1.16.64.160
40,960.0
1 $3.83 15.565 Launch
h100nvl-1.16.96.160
40,960.0
1 $4.11 18.490 Launch
teslaa100-2.24.96.160.nvlink
40,960.0
tensor
2 $4.61 31.895 Launch
h200-1.16.128.160
40,960.0
1 $4.74 28.312 Launch
h200-2.24.256.160.nvlink
40,960.0
tensor
2 $9.40 57.360 Launch

Related models

Need help?

Contact our dedicated neural networks support team at support@immers.cloud or send your request to the sales department at sale@immers.cloud.

We use cookies and web analytics services to ensure the proper functioning of the website, analyze web-site traffic, and improve the quality of our services.
By continuing to use the website, you consent to the Privacy Policy and consent to the processing of cookie files and technical data.