granite-4.1-8b

coding

Granite-4.1-8B is an innovative language model from IBM built on a dense decoder-only transformer architecture with full attention. Unlike the previous Granite 4.0 generation, which used a Mixture of Experts (MoE) architecture and a hybrid attention mechanism with mamba2, the new lineup demonstrates that carefully curated data quality and a well-designed training process can outperform more complex architectural solutions, eliminating performance gaps through generation quality and predictable model behavior—critical for enterprise applications.

The model employs architectural solutions that have already become classics: Grouped Query Attention (GQA) for efficient memory usage, Rotary Position Embeddings (RoPE) for handling long sequences, SwiGLU activations and RMSNorm, as well as tied input and output embeddings. The quality of Granite-4.1-8B is achieved through a multi-stage pretraining process on approximately 15 trillion tokens, comprising five sequential phases. The first two phases build fundamental language understanding on a broad dataset including CommonCrawl (59%), code (20%), mathematical data (7%), technical documentation (10.5%), multilingual data (2%), and domain-specific content (1.5%). The third and fourth phases represent "mid-training" with annealing on high-quality data, while the fifth phase introduces long-context training, extending the window to 512K tokens. The post-training process includes supervised fine-tuning on 4.1 million carefully curated samples using an LLM-as-Judge framework and multi-stage reinforcement learning through on-policy GRPO with DAPO loss, systematically strengthening performance in mathematics, programming, instruction following, and general conversational scenarios.

On benchmarks, Granite-4.1-8B demonstrates impressive results. In mathematical reasoning tasks, the model achieves 92.49% on GSM8K (8-shot) and 80.07% on DeepMind Math (0-shot, CoT). In coding tasks—85.37% on HumanEval pass@1 and 87.30% on MBPP pass@1. In alignment—87.06% on IFEval and 68.98% on ArenaHard. Tool calling is evaluated at 68.27% on BFCL v3. It is particularly noteworthy that the 8B model consistently outperforms IBM's previous flagship, Granite 4.0-H-Small (32B parameters, 9B active in MoE architecture)—a clear confirmation of the effectiveness of the new approach to training and architecture.

Use cases for Granite 4.1-8B are focused on building efficient enterprise solutions. Thanks to advanced tool calling and strict instruction following, the model is ideal for serving as the core of autonomous AI agents that automate complex multi-step business processes through integration with external services and databases. The long context makes it an optimal choice for RAG systems, legal analysis, and working with large volumes of technical documentation. Meanwhile, strong results in mathematics, programming, and other knowledge domains allow it to be deployed as an intelligent assistant.

Released under the Apache 2.0 license with cryptographic signatures and ISO certification, the model provides the necessary level of trust and transparency for mission-critical enterprise applications.


Announce Date: 06.04.2026
Parameters: 9B
Context: 132K
Layers: 40
Attention Type: Full Attention
Developer: IBM
Transformers Version: 4.53.3
vLLM Version: >=0.17.0
License: Apache 2.0

Public endpoint

Use our pre-built public endpoints for free to test inference and explore granite-4.1-8b capabilities. You can obtain an API access token on the token management page after registration and verification.
Model Name Context Type GPU Status Link
There are no public endpoints for this model yet.

Private server

Rent your own physically dedicated instance with hourly or long-term monthly billing.

We recommend deploying private instances in the following scenarios:

  • maximize endpoint performance,
  • enable full context for long sequences,
  • ensure top-tier security for data processing in an isolated, dedicated environment,
  • use custom weights, such as fine-tuned models or LoRA adapters.

Recommended server configurations for hosting granite-4.1-8b

Prices:
Name GPU Price, hour TPS Max Concurrency
teslaa10-2.16.64.160
131,072.0
tensor
2 $0.93 1.475 Launch
teslaa2-3.32.128.160
131,072.0
pipeline
3 $1.06 1.177 Launch
teslaa2-4.32.128.160
131,072.0
tensor
4 $1.26 1.750 Launch
rtx3090-2.16.64.160
131,072.0
tensor
2 $1.56 1.573 Launch
rtx3090-2.16.64.160.nvlink
131,072.0
tensor
2 $1.56 1.573 Launch
rtx5090-1.16.64.160
131,072.0
1 $1.59 1.018 Launch
rtx4090-2.16.64.160
131,072.0
tensor
2 $1.92 1.569 Launch
teslaa100-1.16.64.160
131,072.0
1 $2.37 3.222 Launch
h100-1.16.64.160
131,072.0
1 $3.83 3.218 Launch
h100nvl-1.16.96.160
131,072.0
1 $4.11 3.858 Launch
teslaa100-2.24.96.160.nvlink
131,072.0
tensor
2 $4.61 6.696 Launch
h200-1.16.128.160
131,072.0
1 $4.74 6.007 Launch
h200-2.24.256.160.nvlink
131,072.0
tensor
2 $9.40 12.267 Launch
Prices:
Name GPU Price, hour TPS Max Concurrency
teslaa10-2.16.64.160
131,072.0
tensor
2 $0.93 1.280 Launch
teslaa2-4.32.128.160
131,072.0
tensor
4 $1.26 1.555 Launch
rtx3090-2.16.64.160
131,072.0
tensor
2 $1.56 1.378 Launch
rtx3090-2.16.64.160.nvlink
131,072.0
tensor
2 $1.56 1.378 Launch
rtx4090-2.16.64.160
131,072.0
tensor
2 $1.92 1.374 Launch
teslaa100-1.16.64.160
131,072.0
1 $2.37 3.027 Launch
rtx5090-2.16.64.160
131,072.0
tensor
2 $2.93 2.094 Launch
h100-1.16.64.160
131,072.0
1 $3.83 3.023 Launch
h100nvl-1.16.96.160
131,072.0
1 $4.11 3.663 Launch
teslaa100-2.24.96.160.nvlink
131,072.0
tensor
2 $4.61 6.501 Launch
h200-1.16.128.160
131,072.0
1 $4.74 5.812 Launch
h200-2.24.256.160.nvlink
131,072.0
tensor
2 $9.40 12.072 Launch
Prices:
Name GPU Price, hour TPS Max Concurrency
teslaa2-4.32.128.160
131,072.0
tensor
4 $1.26 1.184 Launch
teslaa10-3.16.96.160
131,072.0
pipeline
3 $1.34 1.650 Launch
rtx3090-2.16.64.160
131,072.0
tensor
2 $1.56 1.007 Launch
rtx3090-2.16.64.160.nvlink
131,072.0
tensor
2 $1.56 1.007 Launch
teslaa10-4.12.48.160
131,072.0
tensor
4 $1.57 2.637 Launch
rtx4090-2.16.64.160
131,072.0
tensor
2 $1.92 1.003 Launch
teslaa100-1.16.64.160
131,072.0
1 $2.37 2.656 Launch
rtx5090-2.16.64.160
131,072.0
tensor
2 $2.93 1.723 Launch
h100-1.16.64.160
131,072.0
1 $3.83 2.652 Launch
h100nvl-1.16.96.160
131,072.0
1 $4.11 3.292 Launch
teslaa100-2.24.96.160.nvlink
131,072.0
tensor
2 $4.61 6.130 Launch
h200-1.16.128.160
131,072.0
1 $4.74 5.441 Launch
h200-2.24.256.160.nvlink
131,072.0
tensor
2 $9.40 11.701 Launch

Related models

Need help?

Contact our dedicated neural networks support team at support@immers.cloud or send your request to the sales department at sale@immers.cloud.

We use cookies and web analytics services to ensure the proper functioning of the website, analyze web-site traffic, and improve the quality of our services.
By continuing to use the website, you consent to the Privacy Policy and consent to the processing of cookie files and technical data.