Gemma 3 1B — a compact language model developed by Google DeepMind, achieving an impressive balance between size and capabilities. With just 1 billion parameters and a context window of 32,000 tokens, the model is highly efficient and capable of running on devices with limited resources. Thanks to its architecture, Gemma 3 1B optimizes memory usage, making it an ideal choice for embedded systems and mobile applications.
It's important to note that the 1B version is a text-only model and does not support image processing, unlike the larger variants in the Gemma 3 series (4B, 12B, and 27B), which offer multimodal capabilities.
Additionally, the model has limited language support and is primarily optimized for tasks in English.
Gemma 3 1B is available with open weights, making it easy to fine-tune and adapt to specific use cases. It comes in multiple quantization levels, ranging from 32-bit down to 4-bit, providing added flexibility when deploying on different hardware platforms.
| Model Name | Context | Type | GPU | Status | Link |
|---|
There are no public endpoints for this model yet.
Rent your own physically dedicated instance with hourly or long-term monthly billing.
We recommend deploying private instances in the following scenarios:
| Name | GPU | TPS | Max Concurrency | |||
|---|---|---|---|---|---|---|
32,768.0 |
1 | $0.38 | 52.19 | 48.327 | Launch | |
32,768.0 |
1 | $0.53 | 70.13 | 81.206 | Launch | |
32,768.0 |
1 | $0.83 | 85.10 | 85.605 | Launch | |
32,768.0 |
1 | $1.02 | 95.84 | 85.439 | Launch | |
32,768.0 tensor |
2 | $1.56 | 48.251 | Launch | ||
32,768.0 |
1 | $1.59 | 94.31 | 117.983 | Launch | |
32,768.0 |
1 | $2.37 | 105.64 | 146.432 | Launch | |
32,768.0 |
1 | $3.83 | 103.88 | 140.625 | Launch | |
32,768.0 |
1 | $4.11 | 205.74 | 154.676 | Launch | |
32,768.0 tensor |
2 | $4.61 | 129.163 | Launch | ||
32,768.0 |
1 | $4.74 | 244.302 | Launch | ||
32,768.0 tensor |
2 | $9.40 | 245.345 | Launch | ||
| Name | GPU | TPS | Max Concurrency | |||
|---|---|---|---|---|---|---|
32,768.0 |
1 | $0.38 | 51.26 | 44.176 | Launch | |
32,768.0 |
1 | $0.53 | 62.69 | 77.055 | Launch | |
32,768.0 |
1 | $0.83 | 80.56 | 81.455 | Launch | |
32,768.0 |
1 | $1.02 | 89.07 | 81.289 | Launch | |
32,768.0 tensor |
2 | $1.56 | 46.198 | Launch | ||
32,768.0 |
1 | $1.59 | 83.91 | 113.833 | Launch | |
32,768.0 |
1 | $2.37 | 110.85 | 144.513 | Launch | |
32,768.0 |
1 | $3.83 | 116.27 | 138.727 | Launch | |
32,768.0 |
1 | $4.11 | 260.99 | 152.778 | Launch | |
32,768.0 tensor |
2 | $4.61 | 128.214 | Launch | ||
32,768.0 |
1 | $4.74 | 242.405 | Launch | ||
32,768.0 tensor |
2 | $9.40 | 244.396 | Launch | ||
Contact our dedicated neural networks support team at support@immers.cloud or send your request to the sales department at sale@immers.cloud.