granite-4.1-3b

coding

The Granite-4.1-3B model is the smallest member of the 4.1 family of dense neural networks from IBM. It uses a proven decoder-only architecture with full attention, which, combined with high-quality data and a multi-stage training pipeline, helps compensate for its small size and routing. Unlike larger versions, the 3-billion-parameter modification is focused on maximum efficiency under constrained resources, while still remaining competitive in key metrics. The architectural elements are typical of the lineup: GQA for memory savings, RoPE for long contexts, SwiGLU, RMSNorm, and separate embeddings.

The pre-training stages follow the overall strategy: five phases on a corpus of 15 trillion tokens with source distribution (CommonCrawl ~59%, code ~20%, mathematics ~7%, technical documentation ~10.5%, multilingual ~2%, subject-specific data ~1.5%). The third and fourth phases are intermediate continual training on cleaned samples with gradually decaying learning rate, and the fifth is extending the context window to 512K. Post-training includes supervised fine-tuning on 4.1 million samples with evaluation by an LLM-as-a-judge and reinforcement learning via GRPO with DAPO loss, which significantly improves performance in mathematics, coding, and dialogue.

Despite its compact size, the benchmark results look quite respectable: MMLU (5-shot) – 67.02%, MMLU-Pro (5-shot, CoT) – 49.83%, GSM8K (8-shot) – 86.88%, DeepMind Math (0-shot, CoT) – 64.64%, HumanEval pass@1 – 79.27%, MBPP pass@1 – 61.64%, BBH (3-shot, CoT) – 75.83%, AGI EVAL – 65.16%. These achievements are comparable to models twice its size, confirming the successful choice of training methodology.

Use cases for Granite 4.1-3B are centered around deployment in environments with strict memory and compute constraints. Support for FP8 quantization and the model's inherently small size allow it to run locally on devices, in mobile applications, and even in IoT gateways. At the same time, it retains the ability to integrate with external tools and follow instructions accurately, making it suitable for building simple agents and RAG systems, or for use as a conversational assistant and in pipelines of various corporate procedures.


Announce Date: 06.04.2026
Parameters: 4B
Context: 132K
Layers: 40
Attention Type: Full Attention
Developer: IBM
Transformers Version: 4.53.3
vLLM Version: >=0.17.0
License: Apache 2.0

Public endpoint

Use our pre-built public endpoints for free to test inference and explore granite-4.1-3b capabilities. You can obtain an API access token on the token management page after registration and verification.
Model Name Context Type GPU Status Link
There are no public endpoints for this model yet.

Private server

Rent your own physically dedicated instance with hourly or long-term monthly billing.

We recommend deploying private instances in the following scenarios:

  • maximize endpoint performance,
  • enable full context for long sequences,
  • ensure top-tier security for data processing in an isolated, dedicated environment,
  • use custom weights, such as fine-tuned models or LoRA adapters.

Recommended server configurations for hosting granite-4.1-3b

Prices:
Name GPU Price, hour TPS Max Concurrency
teslaa10-1.16.32.160
131,072.0
1 $0.53 1.511 Launch
teslaa2-2.16.32.160
131,072.0
tensor
2 $0.57 1.785 Launch
rtx3090-1.16.24.160
131,072.0
1 $0.83 1.608 Launch
rtx4090-1.16.32.160
131,072.0
1 $1.02 1.605 Launch
rtx3080-3.16.64.160
131,072.0
pipeline
3 $1.43 1.332 Launch
rtx3090-2.16.64.160.nvlink
131,072.0
tensor
2 $1.56 3.434 Launch
rtx5090-1.16.64.160
131,072.0
1 $1.59 2.325 Launch
rtx3080-4.16.64.160
131,072.0
tensor
4 $1.82 1.951 Launch
teslaa100-1.16.64.160
131,072.0
1 $2.37 6.732 Launch
h100-1.16.64.160
131,072.0
1 $3.83 6.725 Launch
h100nvl-1.16.96.160
131,072.0
1 $4.11 8.005 Launch
teslaa100-2.24.96.160.nvlink
131,072.0
tensor
2 $4.61 13.681 Launch
h200-1.16.128.160
131,072.0
1 $4.74 12.302 Launch
h200-2.24.256.160.nvlink
131,072.0
tensor
2 $9.40 24.822 Launch
Prices:
Name GPU Price, hour TPS Max Concurrency
teslaa10-1.16.32.160
131,072.0
1 $0.53 1.339 Launch
teslaa2-2.16.32.160
131,072.0
tensor
2 $0.57 1.613 Launch
rtx3090-1.16.24.160
131,072.0
1 $0.83 1.437 Launch
rtx4090-1.16.32.160
131,072.0
1 $1.02 1.433 Launch
rtx3080-3.16.64.160
131,072.0
pipeline
3 $1.43 1.160 Launch
rtx3090-2.16.64.160.nvlink
131,072.0
tensor
2 $1.56 3.262 Launch
rtx5090-1.16.64.160
131,072.0
1 $1.59 2.153 Launch
rtx3080-4.16.64.160
131,072.0
tensor
4 $1.82 1.780 Launch
teslaa100-1.16.64.160
131,072.0
1 $2.37 6.560 Launch
h100-1.16.64.160
131,072.0
1 $3.83 6.554 Launch
h100nvl-1.16.96.160
131,072.0
1 $4.11 7.833 Launch
teslaa100-2.24.96.160.nvlink
131,072.0
tensor
2 $4.61 13.509 Launch
h200-1.16.128.160
131,072.0
1 $4.74 12.131 Launch
h200-2.24.256.160.nvlink
131,072.0
tensor
2 $9.40 24.650 Launch
Prices:
Name GPU Price, hour TPS Max Concurrency
teslaa10-1.16.32.160
131,072.0
1 $0.53 1.094 Launch
teslaa2-2.16.32.160
131,072.0
tensor
2 $0.57 1.369 Launch
rtx3090-1.16.24.160
131,072.0
1 $0.83 1.192 Launch
rtx4090-1.16.32.160
131,072.0
1 $1.02 1.188 Launch
rtx3090-2.16.64.160.nvlink
131,072.0
tensor
2 $1.56 3.017 Launch
rtx5090-1.16.64.160
131,072.0
1 $1.59 1.908 Launch
rtx3080-4.16.64.160
131,072.0
tensor
4 $1.82 1.535 Launch
teslaa100-1.16.64.160
131,072.0
1 $2.37 6.315 Launch
h100-1.16.64.160
131,072.0
1 $3.83 6.309 Launch
h100nvl-1.16.96.160
131,072.0
1 $4.11 7.588 Launch
teslaa100-2.24.96.160.nvlink
131,072.0
tensor
2 $4.61 13.264 Launch
h200-1.16.128.160
131,072.0
1 $4.74 11.886 Launch
h200-2.24.256.160.nvlink
131,072.0
tensor
2 $9.40 24.405 Launch

Related models

Need help?

Contact our dedicated neural networks support team at support@immers.cloud or send your request to the sales department at sale@immers.cloud.

We use cookies and web analytics services to ensure the proper functioning of the website, analyze web-site traffic, and improve the quality of our services.
By continuing to use the website, you consent to the Privacy Policy and consent to the processing of cookie files and technical data.