Gemma 3 12B is a well-balanced mid-sized multimodal language model developed by Google DeepMind, designed to tackle narrow, specialized professional tasks. With 12 billion parameters, the model combines high performance with computational efficiency and supports a wide range of capabilities—from text analysis to image processing. Gemma 3 12B converts visual data into tokens, enabling deep understanding of images. The "Pan&Scan" technology allows adaptive processing of images with any aspect ratio, preserving detail when scaling up to a resolution of 896×896.
Another key feature is the expanded context window of up to 128K tokens. This enables the model to process lengthy legal documents and scientific articles in a single request without losing context. Multilingual support covers more than 140 languages, including Russian, while the enhanced tokenizer from Gemini 2.0 ensures high-quality translation, text generation, and cross-lingual analysis. Additionally, developer-supported quantization makes it possible to run the model even on consumer-grade GPUs with minimal loss in quality.
As a result, Gemma 3 12B is a versatile tool for data analysis, document processing, and information extraction from visual sources—with the ability to run locally and scalable integration into modern AI infrastructures.
| Model Name | Context | Type | GPU | Status | Link |
|---|
There are no public endpoints for this model yet.
Rent your own physically dedicated instance with hourly or long-term monthly billing.
We recommend deploying private instances in the following scenarios:
| Name | GPU | TPS | Max Concurrency | |||
|---|---|---|---|---|---|---|
131,072.0 tensor |
2 | $0.93 | 2.272 | Launch | ||
131,072.0 pipeline |
3 | $1.06 | 8.87 | 1.575 | Launch | |
131,072.0 tensor |
4 | $1.26 | 3.185 | Launch | ||
131,072.0 tensor |
2 | $1.56 | 92.00 | 2.950 | Launch | |
131,072.0 tensor |
2 | $1.56 | 2.950 | Launch | ||
131,072.0 |
1 | $1.59 | 1.147 | Launch | ||
131,072.0 tensor |
4 | $1.82 | 1.265 | Launch | ||
131,072.0 tensor |
2 | $1.92 | 124.25 | 2.468 | Launch | |
131,072.0 |
1 | $2.37 | 106.97 | 4.361 | Launch | |
131,072.0 |
1 | $3.83 | 94.06 | 4.142 | Launch | |
131,072.0 |
1 | $4.11 | 5.091 | Launch | ||
131,072.0 tensor |
2 | $4.61 | 9.476 | Launch | ||
131,072.0 |
1 | $4.74 | 8.319 | Launch | ||
131,072.0 tensor |
2 | $9.40 | 17.845 | Launch | ||
| Name | GPU | TPS | Max Concurrency | |||
|---|---|---|---|---|---|---|
131,072.0 tensor |
2 | $0.93 | 54.50 | 2.504 | Launch | |
131,072.0 pipeline |
3 | $1.06 | 1.761 | Launch | ||
131,072.0 tensor |
4 | $1.26 | 3.418 | Launch | ||
131,072.0 tensor |
2 | $1.56 | 80.29 | 2.708 | Launch | |
131,072.0 tensor |
2 | $1.56 | 2.708 | Launch | ||
131,072.0 |
1 | $1.59 | 1.378 | Launch | ||
131,072.0 tensor |
4 | $1.82 | 1.498 | Launch | ||
131,072.0 tensor |
2 | $1.92 | 86.72 | 2.700 | Launch | |
131,072.0 |
1 | $2.37 | 83.57 | 4.300 | Launch | |
131,072.0 |
1 | $3.83 | 96.29 | 4.295 | Launch | |
131,072.0 |
1 | $4.11 | 5.256 | Launch | ||
131,072.0 tensor |
2 | $4.61 | 9.642 | Launch | ||
131,072.0 |
1 | $4.74 | 8.484 | Launch | ||
131,072.0 tensor |
2 | $9.40 | 18.011 | Launch | ||
| Name | GPU | TPS | Max Concurrency | |||
|---|---|---|---|---|---|---|
131,072.0 tensor |
2 | $0.93 | 1.437 | Launch | ||
131,072.0 tensor |
4 | $1.26 | 16.66 | 2.337 | Launch | |
131,072.0 tensor |
2 | $1.56 | 52.27 | 1.641 | Launch | |
131,072.0 tensor |
2 | $1.56 | 1.641 | Launch | ||
131,072.0 tensor |
2 | $1.92 | 1.633 | Launch | ||
131,072.0 |
1 | $2.37 | 3.539 | Launch | ||
131,072.0 tensor |
2 | $2.93 | 3.140 | Launch | ||
131,072.0 |
1 | $3.83 | 61.70 | 3.534 | Launch | |
131,072.0 |
1 | $4.11 | 4.495 | Launch | ||
131,072.0 tensor |
2 | $4.61 | 8.876 | Launch | ||
131,072.0 |
1 | $4.74 | 7.723 | Launch | ||
131,072.0 tensor |
2 | $9.40 | 17.245 | Launch | ||
Contact our dedicated neural networks support team at support@immers.cloud or send your request to the sales department at sale@immers.cloud.