Granite-4.1-8B is an innovative language model from IBM built on a dense decoder-only transformer architecture with full attention. Unlike the previous Granite 4.0 generation, which used a Mixture of Experts (MoE) architecture and a hybrid attention mechanism with mamba2, the new lineup demonstrates that carefully curated data quality and a well-designed training process can outperform more complex architectural solutions, eliminating performance gaps through generation quality and predictable model behavior—critical for enterprise applications.
The model employs architectural solutions that have already become classics: Grouped Query Attention (GQA) for efficient memory usage, Rotary Position Embeddings (RoPE) for handling long sequences, SwiGLU activations and RMSNorm, as well as tied input and output embeddings. The quality of Granite-4.1-8B is achieved through a multi-stage pretraining process on approximately 15 trillion tokens, comprising five sequential phases. The first two phases build fundamental language understanding on a broad dataset including CommonCrawl (59%), code (20%), mathematical data (7%), technical documentation (10.5%), multilingual data (2%), and domain-specific content (1.5%). The third and fourth phases represent "mid-training" with annealing on high-quality data, while the fifth phase introduces long-context training, extending the window to 512K tokens. The post-training process includes supervised fine-tuning on 4.1 million carefully curated samples using an LLM-as-Judge framework and multi-stage reinforcement learning through on-policy GRPO with DAPO loss, systematically strengthening performance in mathematics, programming, instruction following, and general conversational scenarios.
On benchmarks, Granite-4.1-8B demonstrates impressive results. In mathematical reasoning tasks, the model achieves 92.49% on GSM8K (8-shot) and 80.07% on DeepMind Math (0-shot, CoT). In coding tasks—85.37% on HumanEval pass@1 and 87.30% on MBPP pass@1. In alignment—87.06% on IFEval and 68.98% on ArenaHard. Tool calling is evaluated at 68.27% on BFCL v3. It is particularly noteworthy that the 8B model consistently outperforms IBM's previous flagship, Granite 4.0-H-Small (32B parameters, 9B active in MoE architecture)—a clear confirmation of the effectiveness of the new approach to training and architecture.
Use cases for Granite 4.1-8B are focused on building efficient enterprise solutions. Thanks to advanced tool calling and strict instruction following, the model is ideal for serving as the core of autonomous AI agents that automate complex multi-step business processes through integration with external services and databases. The long context makes it an optimal choice for RAG systems, legal analysis, and working with large volumes of technical documentation. Meanwhile, strong results in mathematics, programming, and other knowledge domains allow it to be deployed as an intelligent assistant.
Released under the Apache 2.0 license with cryptographic signatures and ISO certification, the model provides the necessary level of trust and transparency for mission-critical enterprise applications.
| Model Name | Context | Type | GPU | Status | Link |
|---|
There are no public endpoints for this model yet.
Rent your own physically dedicated instance with hourly or long-term monthly billing.
We recommend deploying private instances in the following scenarios:
| Name | GPU | TPS | Max Concurrency | |||
|---|---|---|---|---|---|---|
131,072.0 tensor |
2 | $0.93 | 1.475 | Launch | ||
131,072.0 pipeline |
3 | $1.06 | 1.177 | Launch | ||
131,072.0 tensor |
4 | $1.26 | 1.750 | Launch | ||
131,072.0 tensor |
2 | $1.56 | 1.573 | Launch | ||
131,072.0 tensor |
2 | $1.56 | 1.573 | Launch | ||
131,072.0 |
1 | $1.59 | 1.018 | Launch | ||
131,072.0 tensor |
2 | $1.92 | 1.569 | Launch | ||
131,072.0 |
1 | $2.37 | 3.222 | Launch | ||
131,072.0 |
1 | $3.83 | 3.218 | Launch | ||
131,072.0 |
1 | $4.11 | 3.858 | Launch | ||
131,072.0 tensor |
2 | $4.61 | 6.696 | Launch | ||
131,072.0 |
1 | $4.74 | 6.007 | Launch | ||
131,072.0 tensor |
2 | $9.40 | 12.267 | Launch | ||
| Name | GPU | TPS | Max Concurrency | |||
|---|---|---|---|---|---|---|
131,072.0 tensor |
2 | $0.93 | 1.280 | Launch | ||
131,072.0 tensor |
4 | $1.26 | 1.555 | Launch | ||
131,072.0 tensor |
2 | $1.56 | 1.378 | Launch | ||
131,072.0 tensor |
2 | $1.56 | 1.378 | Launch | ||
131,072.0 tensor |
2 | $1.92 | 1.374 | Launch | ||
131,072.0 |
1 | $2.37 | 3.027 | Launch | ||
131,072.0 tensor |
2 | $2.93 | 2.094 | Launch | ||
131,072.0 |
1 | $3.83 | 3.023 | Launch | ||
131,072.0 |
1 | $4.11 | 3.663 | Launch | ||
131,072.0 tensor |
2 | $4.61 | 6.501 | Launch | ||
131,072.0 |
1 | $4.74 | 5.812 | Launch | ||
131,072.0 tensor |
2 | $9.40 | 12.072 | Launch | ||
| Name | GPU | TPS | Max Concurrency | |||
|---|---|---|---|---|---|---|
131,072.0 tensor |
4 | $1.26 | 1.184 | Launch | ||
131,072.0 pipeline |
3 | $1.34 | 1.650 | Launch | ||
131,072.0 tensor |
2 | $1.56 | 1.007 | Launch | ||
131,072.0 tensor |
2 | $1.56 | 1.007 | Launch | ||
131,072.0 tensor |
4 | $1.57 | 2.637 | Launch | ||
131,072.0 tensor |
2 | $1.92 | 1.003 | Launch | ||
131,072.0 |
1 | $2.37 | 2.656 | Launch | ||
131,072.0 tensor |
2 | $2.93 | 1.723 | Launch | ||
131,072.0 |
1 | $3.83 | 2.652 | Launch | ||
131,072.0 |
1 | $4.11 | 3.292 | Launch | ||
131,072.0 tensor |
2 | $4.61 | 6.130 | Launch | ||
131,072.0 |
1 | $4.74 | 5.441 | Launch | ||
131,072.0 tensor |
2 | $9.40 | 11.701 | Launch | ||
Contact our dedicated neural networks support team at support@immers.cloud or send your request to the sales department at sale@immers.cloud.