granite-4.1-30b

coding

The Granite-4.1-30B model advances IBM’s approach to building dense language models. Unlike the hybrid architectures of the previous generation, it uses a classic decoder-only transformer with full attention, which ensures deterministic computation and high reproducibility of results—a key requirement for industrial deployment. Engineering choices include Grouped-Query Attention (GQA) to reduce memory consumption, RoPE positional encodings for correct handling of long sequences, SwiGLU activation function, RMSNorm normalization, and tied input/output embeddings (actually the original says "разделённые эмбеддинги" which means "separated embeddings", likely meaning untied; I'll translate as "untied input and output embeddings"). The thirty-billion-parameter version features a deepened architecture: the number of layers reaches 64, endowing it with increased capacity for complex dependencies.

The model was trained in five stages on a corpus of approximately 15 trillion tokens: the first two phases provided basic language understanding on a mix of internet corpora (CommonCrawl – 59%), code (20%), mathematical texts (7%), technical manuals (10.5%), multilingual sources (2%), and specialized materials (1.5%). The third and fourth cycles involved intermediate fine-tuning on curated data with reduced learning rates, and the fifth stage was dedicated to adapting to long sequences up to 512k tokens. Post-training included supervised fine-tuning on 4.1 million examples selected with an LLM-based judge, and multiple reinforcement learning iterations using GRPO with DAPO loss, which significantly improved performance in mathematics, coding, instruction following, and dialogue scenarios.

Benchmarks show very strong results. On MMLU (5-shot) it scores 80.16%, on MMLU-Pro (5-shot, CoT) – 64.09%, on GSM8K (8-shot) – 94.16%, on DeepMind Math (0-shot, CoT) – 81.93%, on HumanEval pass@1 – 89.63%, on MBPP pass@1 – 83.33%, on BBH (3-shot, CoT) – 83.74%, on AGI EVAL – 77.80%. These figures substantially surpass the results of the Granite 4.0 line, clearly illustrating the effectiveness of the new training strategy.

Practical application of Granite 4.1-30B covers the most resource-intensive enterprise tasks. Support for tool calling combined with the ability to precisely follow instructions opens the door to building multi-step autonomous agents that interact with corporate APIs and databases. The extended context enables the model to be used in RAG systems, and its strengths in mathematics and programming allow it to serve as an intelligent assistant.

Importantly, the model is distributed under the Apache 2.0 license, is accompanied by cryptographic signatures, and is certified according to ISO standards, guaranteeing a high level of reliability and transparency.


Announce Date: 06.04.2026
Parameters: 29B
Context: 132K
Layers: 64
Attention Type: Full Attention
Developer: IBM
Transformers Version: 4.53.3
vLLM Version: >=0.17.0
License: Apache 2.0

Public endpoint

Use our pre-built public endpoints for free to test inference and explore granite-4.1-30b capabilities. You can obtain an API access token on the token management page after registration and verification.
Model Name Context Type GPU Status Link
There are no public endpoints for this model yet.

Private server

Rent your own physically dedicated instance with hourly or long-term monthly billing.

We recommend deploying private instances in the following scenarios:

  • maximize endpoint performance,
  • enable full context for long sequences,
  • ensure top-tier security for data processing in an isolated, dedicated environment,
  • use custom weights, such as fine-tuned models or LoRA adapters.

Recommended server configurations for hosting granite-4.1-30b

Prices:
Name GPU Price, hour TPS Max Concurrency
teslaa10-4.16.64.160
131,072.0
tensor
4 $1.62 1.582 Launch
teslaa2-6.32.128.160
131,072.0
pipeline
6 $1.65 1.242 Launch
rtx3090-3.16.96.160
131,072.0
pipeline
3 $2.29 1.081 Launch
teslaa100-1.16.64.160
131,072.0
1 $2.37 1.593 Launch
rtx4090-3.16.96.160
131,072.0
pipeline
3 $2.83 1.078 Launch
rtx3090-4.16.64.160
131,072.0
tensor
4 $2.89 1.703 Launch
rtx5090-2.16.64.160
131,072.0
tensor
2 $2.93 1.010 Launch
rtx3090-4.16.128.160.nvlink
131,072.0
tensor
4 $3.01 1.703 Launch
rtx4090-4.16.64.160
131,072.0
tensor
4 $3.60 1.699 Launch
h100-1.16.64.160
131,072.0
1 $3.83 1.591 Launch
h100nvl-1.16.96.160
131,072.0
1 $4.11 1.991 Launch
teslaa100-2.24.96.160.nvlink
131,072.0
tensor
2 $4.61 3.765 Launch
h200-1.16.128.160
131,072.0
1 $4.74 3.334 Launch
h200-2.24.256.160.nvlink
131,072.0
tensor
2 $9.40 7.246 Launch
Prices:
Name GPU Price, hour TPS Max Concurrency
teslaa10-4.16.64.160
131,072.0
tensor
4 $1.62 1.284 Launch
teslaa100-1.16.64.160
131,072.0
1 $2.37 1.295 Launch
rtx3090-4.16.64.160
131,072.0
tensor
4 $2.89 1.406 Launch
rtx3090-4.16.128.160.nvlink
131,072.0
tensor
4 $3.01 1.406 Launch
rtx4090-4.16.64.160
131,072.0
tensor
4 $3.60 1.401 Launch
h100-1.16.64.160
131,072.0
1 $3.83 1.293 Launch
h100nvl-1.16.96.160
131,072.0
1 $4.11 1.693 Launch
rtx5090-3.16.96.160
131,072.0
pipeline
3 $4.34 1.435 Launch
teslaa100-2.24.96.160.nvlink
131,072.0
tensor
2 $4.61 3.467 Launch
h200-1.16.128.160
131,072.0
1 $4.74 3.036 Launch
rtx5090-4.16.128.160
131,072.0
tensor
4 $5.74 2.301 Launch
h200-2.24.256.160.nvlink
131,072.0
tensor
2 $9.40 6.948 Launch
Prices:
Name GPU Price, hour TPS Max Concurrency
teslaa100-2.24.96.160.nvlink
131,072.0
tensor
2 $4.61 2.663 Launch
h200-1.16.128.160
131,072.0
1 $4.74 2.232 Launch
teslaa100-2.24.256.160
131,072.0
tensor
2 $4.93 2.663 Launch
rtx5090-4.16.128.160
131,072.0
tensor
4 $5.74 1.498 Launch
rtx4090-6.44.256.160
131,072.0
pipeline
6 $5.83 1.632 Launch
dedicated-rtx3090-8.64.128.960-1
131,072.0
tensor
8 $6.04 2.884 Launch
rtx4090-8.44.256.160
131,072.0
tensor
8 $7.51 2.874 Launch
h100-2.24.256.160
131,072.0
tensor
2 $7.84 2.659 Launch
h100nvl-2.24.192.240
131,072.0
tensor
2 $8.17 3.459 Launch
h200-2.24.256.160.nvlink
131,072.0
tensor
2 $9.40 6.145 Launch

Related models

Need help?

Contact our dedicated neural networks support team at support@immers.cloud or send your request to the sales department at sale@immers.cloud.

We use cookies and web analytics services to ensure the proper functioning of the website, analyze web-site traffic, and improve the quality of our services.
By continuing to use the website, you consent to the Privacy Policy and consent to the processing of cookie files and technical data.