The Granite-4.1-30B model advances IBM’s approach to building dense language models. Unlike the hybrid architectures of the previous generation, it uses a classic decoder-only transformer with full attention, which ensures deterministic computation and high reproducibility of results—a key requirement for industrial deployment. Engineering choices include Grouped-Query Attention (GQA) to reduce memory consumption, RoPE positional encodings for correct handling of long sequences, SwiGLU activation function, RMSNorm normalization, and tied input/output embeddings (actually the original says "разделённые эмбеддинги" which means "separated embeddings", likely meaning untied; I'll translate as "untied input and output embeddings"). The thirty-billion-parameter version features a deepened architecture: the number of layers reaches 64, endowing it with increased capacity for complex dependencies.
The model was trained in five stages on a corpus of approximately 15 trillion tokens: the first two phases provided basic language understanding on a mix of internet corpora (CommonCrawl – 59%), code (20%), mathematical texts (7%), technical manuals (10.5%), multilingual sources (2%), and specialized materials (1.5%). The third and fourth cycles involved intermediate fine-tuning on curated data with reduced learning rates, and the fifth stage was dedicated to adapting to long sequences up to 512k tokens. Post-training included supervised fine-tuning on 4.1 million examples selected with an LLM-based judge, and multiple reinforcement learning iterations using GRPO with DAPO loss, which significantly improved performance in mathematics, coding, instruction following, and dialogue scenarios.
Benchmarks show very strong results. On MMLU (5-shot) it scores 80.16%, on MMLU-Pro (5-shot, CoT) – 64.09%, on GSM8K (8-shot) – 94.16%, on DeepMind Math (0-shot, CoT) – 81.93%, on HumanEval pass@1 – 89.63%, on MBPP pass@1 – 83.33%, on BBH (3-shot, CoT) – 83.74%, on AGI EVAL – 77.80%. These figures substantially surpass the results of the Granite 4.0 line, clearly illustrating the effectiveness of the new training strategy.
Practical application of Granite 4.1-30B covers the most resource-intensive enterprise tasks. Support for tool calling combined with the ability to precisely follow instructions opens the door to building multi-step autonomous agents that interact with corporate APIs and databases. The extended context enables the model to be used in RAG systems, and its strengths in mathematics and programming allow it to serve as an intelligent assistant.
Importantly, the model is distributed under the Apache 2.0 license, is accompanied by cryptographic signatures, and is certified according to ISO standards, guaranteeing a high level of reliability and transparency.
| Model Name | Context | Type | GPU | Status | Link |
|---|
There are no public endpoints for this model yet.
Rent your own physically dedicated instance with hourly or long-term monthly billing.
We recommend deploying private instances in the following scenarios:
| Name | GPU | TPS | Max Concurrency | |||
|---|---|---|---|---|---|---|
131,072.0 tensor |
4 | $1.62 | 1.582 | Launch | ||
131,072.0 pipeline |
6 | $1.65 | 1.242 | Launch | ||
131,072.0 pipeline |
3 | $2.29 | 1.081 | Launch | ||
131,072.0 |
1 | $2.37 | 1.593 | Launch | ||
131,072.0 pipeline |
3 | $2.83 | 1.078 | Launch | ||
131,072.0 tensor |
4 | $2.89 | 1.703 | Launch | ||
131,072.0 tensor |
2 | $2.93 | 1.010 | Launch | ||
131,072.0 tensor |
4 | $3.01 | 1.703 | Launch | ||
131,072.0 tensor |
4 | $3.60 | 1.699 | Launch | ||
131,072.0 |
1 | $3.83 | 1.591 | Launch | ||
131,072.0 |
1 | $4.11 | 1.991 | Launch | ||
131,072.0 tensor |
2 | $4.61 | 3.765 | Launch | ||
131,072.0 |
1 | $4.74 | 3.334 | Launch | ||
131,072.0 tensor |
2 | $9.40 | 7.246 | Launch | ||
| Name | GPU | TPS | Max Concurrency | |||
|---|---|---|---|---|---|---|
131,072.0 tensor |
4 | $1.62 | 1.284 | Launch | ||
131,072.0 |
1 | $2.37 | 1.295 | Launch | ||
131,072.0 tensor |
4 | $2.89 | 1.406 | Launch | ||
131,072.0 tensor |
4 | $3.01 | 1.406 | Launch | ||
131,072.0 tensor |
4 | $3.60 | 1.401 | Launch | ||
131,072.0 |
1 | $3.83 | 1.293 | Launch | ||
131,072.0 |
1 | $4.11 | 1.693 | Launch | ||
131,072.0 pipeline |
3 | $4.34 | 1.435 | Launch | ||
131,072.0 tensor |
2 | $4.61 | 3.467 | Launch | ||
131,072.0 |
1 | $4.74 | 3.036 | Launch | ||
131,072.0 tensor |
4 | $5.74 | 2.301 | Launch | ||
131,072.0 tensor |
2 | $9.40 | 6.948 | Launch | ||
| Name | GPU | TPS | Max Concurrency | |||
|---|---|---|---|---|---|---|
131,072.0 tensor |
2 | $4.61 | 2.663 | Launch | ||
131,072.0 |
1 | $4.74 | 2.232 | Launch | ||
131,072.0 tensor |
2 | $4.93 | 2.663 | Launch | ||
131,072.0 tensor |
4 | $5.74 | 1.498 | Launch | ||
131,072.0 pipeline |
6 | $5.83 | 1.632 | Launch | ||
131,072.0 tensor |
8 | $6.04 | 2.884 | Launch | ||
131,072.0 tensor |
8 | $7.51 | 2.874 | Launch | ||
131,072.0 tensor |
2 | $7.84 | 2.659 | Launch | ||
131,072.0 tensor |
2 | $8.17 | 3.459 | Launch | ||
131,072.0 tensor |
2 | $9.40 | 6.145 | Launch | ||
Contact our dedicated neural networks support team at support@immers.cloud or send your request to the sales department at sale@immers.cloud.