Phi-4

Phi-4 is a modern open-source language model with 14 billion parameters. Although the architectural changes compared to previous versions are minimal, the model demonstrates significant progress in tasks requiring logical thinking and analytical skills, made possible by an innovative training approach. Unlike traditional language model training methods, Phi-4 focuses not on the quantity of data but on its quality. The training of Phi-4 utilized diverse sources, including synthetic data specifically designed to develop reasoning skills, filtered documents from public sources, as well as acquired academic books and question-answer knowledge bases. This allows the model to achieve high performance even with a relatively small size.  

Phi-4 works exclusively with textual data. Its context window is relatively small at 16K tokens, but it supports more than 50 languages, including Russian.  

Overall, Phi-4 is a versatile lightweight model, but according to its developers, it is particularly effective in environments with limited memory and computational resources, as well as for tasks requiring instant response.


Announce Date: 12.12.2024
Parameters: 15B
Context: 17K
Layers: 40
Attention Type: Full Attention
Developer: Microsoft
Transformers Version: 4.47.0
License: MIT

Public endpoint

Use our pre-built public endpoints for free to test inference and explore Phi-4 capabilities. You can obtain an API access token on the token management page after registration and verification.
Model Name Context Type GPU Status Link
There are no public endpoints for this model yet.

Private server

Rent your own physically dedicated instance with hourly or long-term monthly billing.

We recommend deploying private instances in the following scenarios:

  • maximize endpoint performance,
  • enable full context for long sequences,
  • ensure top-tier security for data processing in an isolated, dedicated environment,
  • use custom weights, such as fine-tuned models or LoRA adapters.

Recommended server configurations for hosting Phi-4

Prices:
Name GPU Price, hour TPS Max Concurrency
teslaa10-1.16.32.160
16,384.0
1 $0.53 32.100 2.727 Launch
teslat4-2.16.32.160
16,384.0
tensor
2 $0.54 3.876 Launch
teslaa2-2.16.32.160
16,384.0
tensor
2 $0.57 9.150 3.899 Launch
rtx2080ti-2.12.64.160
16,384.0
tensor
2 $0.69 1.521 Launch
rtx3090-1.16.24.160
16,384.0
1 $0.83 3.039 Launch
rtx4090-1.16.32.160
16,384.0
1 $1.02 3.027 Launch
rtxa5000-2.16.64.160.nvlink
16,384.0
tensor
2 $1.23 8.551 Launch
rtx3080-3.16.64.160
16,384.0
pipeline
3 $1.43 2.836 Launch
rtx5090-1.16.64.160
16,384.0
1 $1.59 5.332 Launch
teslaa100-1.16.64.160
16,384.0
1 $2.37 53.030 19.434 Launch
h100-1.16.64.160
16,384.0
1 $3.83 63.450 19.414 Launch
h100nvl-1.16.96.160
16,384.0
1 $4.11 23.509 Launch
teslaa100-2.24.96.160.nvlink
16,384.0
tensor
2 $4.61 41.965 Launch
h200-1.16.128.160
16,384.0
1 $4.74 37.260 Launch
h200-2.24.256.160.nvlink
16,384.0
tensor
2 $9.40 77.617 Launch
Prices:
Name GPU Price, hour TPS Max Concurrency
teslaa10-1.16.32.160
16,384.0
1 $0.53 1.443 Launch
teslat4-2.16.32.160
16,384.0
tensor
2 $0.54 2.592 Launch
teslaa2-2.16.32.160
16,384.0
tensor
2 $0.57 2.616 Launch
rtx3090-1.16.24.160
16,384.0
1 $0.83 1.755 Launch
rtx2080ti-3.12.24.120
16,384.0
pipeline
3 $0.84 2.327 Launch
rtx4090-1.16.32.160
16,384.0
1 $1.02 1.743 Launch
rtxa5000-2.16.64.160.nvlink
16,384.0
tensor
2 $1.23 7.267 Launch
rtx3080-3.16.64.160
16,384.0
pipeline
3 $1.43 1.488 Launch
rtx5090-1.16.64.160
16,384.0
1 $1.59 4.048 Launch
teslaa100-1.16.64.160
16,384.0
1 $2.37 18.150 Launch
h100-1.16.64.160
16,384.0
1 $3.83 18.130 Launch
h100nvl-1.16.96.160
16,384.0
1 $4.11 22.225 Launch
teslaa100-2.24.96.160.nvlink
16,384.0
tensor
2 $4.61 40.681 Launch
h200-1.16.128.160
16,384.0
1 $4.74 35.976 Launch
h200-2.24.256.160.nvlink
16,384.0
tensor
2 $9.40 76.333 Launch
Prices:
Name GPU Price, hour TPS Max Concurrency
teslat4-3.32.64.160
16,384.0
pipeline
3 $0.88 1.285 Launch
teslaa10-2.16.64.160
16,384.0
tensor
2 $0.93 29.000 2.910 Launch
teslaa2-3.32.128.160
16,384.0
pipeline
3 $1.06 5.770 1.320 Launch
rtxa5000-2.16.64.160.nvlink
16,384.0
tensor
2 $1.23 2.910 Launch
rtx3090-2.16.64.160
16,384.0
tensor
2 $1.56 46.560 3.534 Launch
rtx4090-2.16.64.160
16,384.0
tensor
2 $1.92 3.511 Launch
teslaa100-1.16.64.160
16,384.0
1 $2.37 45.550 13.793 Launch
rtx5090-2.16.64.160
16,384.0
tensor
2 $2.93 8.121 Launch
h100-1.16.64.160
16,384.0
1 $3.83 53.090 13.773 Launch
h100nvl-1.16.96.160
16,384.0
1 $4.11 87.210 17.868 Launch
teslaa100-2.24.96.160.nvlink
16,384.0
tensor
2 $4.61 36.325 Launch
h200-1.16.128.160
16,384.0
1 $4.74 31.619 Launch
h200-2.24.256.160.nvlink
16,384.0
tensor
2 $9.40 71.976 Launch

Related models

Need help?

Contact our dedicated neural networks support team at nn@immers.cloud or send your request to the sales department at sale@immers.cloud.