Phi-4

Phi-4 is a modern open-source language model with 14 billion parameters. Although the architectural changes compared to previous versions are minimal, the model demonstrates significant progress in tasks requiring logical thinking and analytical skills, made possible by an innovative training approach. Unlike traditional language model training methods, Phi-4 focuses not on the quantity of data but on its quality. The training of Phi-4 utilized diverse sources, including synthetic data specifically designed to develop reasoning skills, filtered documents from public sources, as well as acquired academic books and question-answer knowledge bases. This allows the model to achieve high performance even with a relatively small size.  

Phi-4 works exclusively with textual data. Its context window is relatively small at 16K tokens, but it supports more than 50 languages, including Russian.  

Overall, Phi-4 is a versatile lightweight model, but according to its developers, it is particularly effective in environments with limited memory and computational resources, as well as for tasks requiring instant response.


Announce Date: 12.12.2024
Parameters: 15B
Context: 17K
Layers: 40
Attention Type: Full Attention
Developer: Microsoft
Transformers Version: 4.47.0
License: MIT

Public endpoint

Use our pre-built public endpoints for free to test inference and explore Phi-4 capabilities. You can obtain an API access token on the token management page after registration and verification.
Model Name Context Type GPU Status Link
There are no public endpoints for this model yet.

Private server

Rent your own physically dedicated instance with hourly or long-term monthly billing.

We recommend deploying private instances in the following scenarios:

  • maximize endpoint performance,
  • enable full context for long sequences,
  • ensure top-tier security for data processing in an isolated, dedicated environment,
  • use custom weights, such as fine-tuned models or LoRA adapters.

Recommended server configurations for hosting Phi-4

Prices:
Name GPU Price, hour TPS Max Concurrency
teslaa10-1.16.32.160
16,384.0
1 $0.53 32.100 2.471 Launch
teslat4-2.16.32.160
16,384.0
tensor
2 $0.54 3.364 Launch
teslaa2-2.16.32.160
16,384.0
tensor
2 $0.57 9.150 3.387 Launch
rtx2080ti-2.12.64.160
16,384.0
tensor
2 $0.69 1.009 Launch
rtx3090-1.16.24.160
16,384.0
1 $0.83 2.783 Launch
rtx4090-1.16.32.160
16,384.0
1 $1.02 2.771 Launch
rtxa5000-2.16.64.160.nvlink
16,384.0
tensor
2 $1.23 8.039 Launch
rtx3080-3.16.64.160
16,384.0
pipeline
3 $1.43 2.068 Launch
rtx5090-1.16.64.160
16,384.0
1 $1.59 5.076 Launch
teslaa100-1.16.64.160
16,384.0
1 $2.37 53.030 19.178 Launch
h100-1.16.64.160
16,384.0
1 $3.83 63.450 19.158 Launch
h100nvl-1.16.96.160
16,384.0
1 $4.11 91.220 23.253 Launch
teslaa100-2.24.96.160.nvlink
16,384.0
tensor
2 $4.61 41.453 Launch
h200-1.16.128.160
16,384.0
1 $4.74 37.004 Launch
h200-2.24.256.160.nvlink
16,384.0
tensor
2 $9.40 77.105 Launch
Prices:
Name GPU Price, hour TPS Max Concurrency
teslat4-2.16.32.160
16,384.0
tensor
2 $0.54 1.784 Launch
teslaa2-2.16.32.160
16,384.0
tensor
2 $0.57 1.808 Launch
rtx3090-1.16.24.160
16,384.0
1 $0.83 1.204 Launch
rtx2080ti-3.12.24.120
16,384.0
pipeline
3 $0.84 1.248 Launch
teslaa10-2.16.64.160
16,384.0
tensor
2 $0.93 6.459 Launch
rtx4090-1.16.32.160
16,384.0
1 $1.02 1.192 Launch
rtxa5000-2.16.64.160.nvlink
16,384.0
tensor
2 $1.23 6.459 Launch
rtx5090-1.16.64.160
16,384.0
1 $1.59 3.497 Launch
rtx3080-4.16.64.160
16,384.0
pipeline
4 $1.82 2.416 Launch
teslaa100-1.16.64.160
16,384.0
1 $2.37 17.599 Launch
h100-1.16.64.160
16,384.0
1 $3.83 17.578 Launch
h100nvl-1.16.96.160
16,384.0
1 $4.11 21.673 Launch
teslaa100-2.24.96.160.nvlink
16,384.0
tensor
2 $4.61 39.874 Launch
h200-1.16.128.160
16,384.0
1 $4.74 35.425 Launch
h200-2.24.256.160.nvlink
16,384.0
tensor
2 $9.40 75.526 Launch
Prices:
Name GPU Price, hour TPS Max Concurrency
teslaa10-2.16.64.160
16,384.0
tensor
2 $0.93 29.000 2.398 Launch
teslat4-4.16.64.160
16,384.0
pipeline
4 $0.96 4.184 Launch
rtxa5000-2.16.64.160.nvlink
16,384.0
tensor
2 $1.23 2.398 Launch
teslaa2-4.32.128.160
16,384.0
pipeline
4 $1.26 4.231 Launch
rtx3090-2.16.64.160
16,384.0
tensor
2 $1.56 46.560 3.022 Launch
rtx4090-2.16.64.160
16,384.0
tensor
2 $1.92 2.999 Launch
teslaa100-1.16.64.160
16,384.0
1 $2.37 45.550 13.537 Launch
rtx5090-2.16.64.160
16,384.0
tensor
2 $2.93 7.609 Launch
h100-1.16.64.160
16,384.0
1 $3.83 53.090 13.517 Launch
h100nvl-1.16.96.160
16,384.0
1 $4.11 87.210 17.612 Launch
teslaa100-2.24.96.160.nvlink
16,384.0
tensor
2 $4.61 35.813 Launch
h200-1.16.128.160
16,384.0
1 $4.74 31.363 Launch
h200-2.24.256.160.nvlink
16,384.0
tensor
2 $9.40 71.464 Launch

Related models

Need help?

Contact our dedicated neural networks support team at nn@immers.cloud or send your request to the sales department at sale@immers.cloud.