Phi-4

Phi-4 is a modern open-source language model with 14 billion parameters. Although the architectural changes compared to previous versions are minimal, the model demonstrates significant progress in tasks requiring logical thinking and analytical skills, made possible by an innovative training approach. Unlike traditional language model training methods, Phi-4 focuses not on the quantity of data but on its quality. The training of Phi-4 utilized diverse sources, including synthetic data specifically designed to develop reasoning skills, filtered documents from public sources, as well as acquired academic books and question-answer knowledge bases. This allows the model to achieve high performance even with a relatively small size.  

Phi-4 works exclusively with textual data. Its context window is relatively small at 16K tokens, but it supports more than 50 languages, including Russian.  

Overall, Phi-4 is a versatile lightweight model, but according to its developers, it is particularly effective in environments with limited memory and computational resources, as well as for tasks requiring instant response.


Announce Date: 12.12.2024
Parameters: 15B
Context: 17K
Layers: 40
Attention Type: Full Attention
Developer: Microsoft
Transformers Version: 4.47.0
License: MIT

Public endpoint

Use our pre-built public endpoints for free to test inference and explore Phi-4 capabilities. You can obtain an API access token on the token management page after registration and verification.
Model Name Context Type GPU Status Link
There are no public endpoints for this model yet.

Private server

Rent your own physically dedicated instance with hourly or long-term monthly billing.

We recommend deploying private instances in the following scenarios:

  • maximize endpoint performance,
  • enable full context for long sequences,
  • ensure top-tier security for data processing in an isolated, dedicated environment,
  • use custom weights, such as fine-tuned models or LoRA adapters.

Recommended server configurations for hosting Phi-4

Prices:
Name GPU Price, hour TPS Max Concurrency
teslaa10-1.16.32.160
16,384.0
1 $0.53 32.82 2.752 Launch
teslaa2-2.16.32.160
16,384.0
tensor
2 $0.57 9.15 4.456 Launch
rtx3090-1.16.24.160
16,384.0
1 $0.83 50.09 3.064 Launch
rtx4090-1.16.32.160
16,384.0
1 $1.02 69.85 3.052 Launch
rtx3080-3.16.64.160
16,384.0
pipeline
3 $1.43 2.627 Launch
rtx3090-2.16.64.160.nvlink
16,384.0
tensor
2 $1.56 9.066 Launch
rtx5090-1.16.64.160
16,384.0
1 $1.59 83.03 5.361 Launch
teslaa100-1.16.64.160
16,384.0
1 $2.37 53.63 19.459 Launch
h100-1.16.64.160
16,384.0
1 $3.83 63.45 19.272 Launch
h100nvl-1.16.96.160
16,384.0
1 $4.11 91.22 23.358 Launch
teslaa100-2.24.96.160.nvlink
16,384.0
tensor
2 $4.61 41.856 Launch
h200-1.16.128.160
16,384.0
1 $4.74 37.109 Launch
h200-2.24.256.160.nvlink
16,384.0
tensor
2 $9.40 77.508 Launch
Prices:
Name GPU Price, hour TPS Max Concurrency
teslaa10-1.16.32.160
16,384.0
1 $0.53 1.446 Launch
teslaa2-2.16.32.160
16,384.0
tensor
2 $0.57 3.278 Launch
rtx3090-1.16.24.160
16,384.0
1 $0.83 50.17 1.912 Launch
rtx4090-1.16.32.160
16,384.0
1 $1.02 43.48 1.673 Launch
rtx3080-3.16.64.160
16,384.0
pipeline
3 $1.43 43.33 2.527 Launch
rtx3090-2.16.64.160.nvlink
16,384.0
tensor
2 $1.56 8.563 Launch
rtx5090-1.16.64.160
16,384.0
1 $1.59 4.205 Launch
teslaa100-1.16.64.160
16,384.0
1 $2.37 80.23 18.141 Launch
h100-1.16.64.160
16,384.0
1 $3.83 96.69 18.040 Launch
h100nvl-1.16.96.160
16,384.0
1 $4.11 143.66 22.132 Launch
teslaa100-2.24.96.160.nvlink
16,384.0
tensor
2 $4.61 41.354 Launch
h200-1.16.128.160
16,384.0
1 $4.74 36.133 Launch
h200-2.24.256.160.nvlink
16,384.0
tensor
2 $9.40 77.005 Launch
Prices:
Name GPU Price, hour TPS Max Concurrency
teslaa10-2.16.64.160
16,384.0
tensor
2 $0.93 29.00 3.290 Launch
teslaa2-3.32.128.160
16,384.0
pipeline
3 $1.06 5.77 2.053 Launch
rtx3090-2.16.64.160
16,384.0
tensor
2 $1.56 46.56 4.138 Launch
rtx3090-2.16.64.160.nvlink
16,384.0
tensor
2 $1.56 4.138 Launch
rtx4090-2.16.64.160
16,384.0
tensor
2 $1.92 55.00 4.057 Launch
teslaa100-1.16.64.160
16,384.0
1 $2.37 46.21 13.984 Launch
rtx5090-2.16.64.160
16,384.0
tensor
2 $2.93 73.39 8.468 Launch
h100-1.16.64.160
16,384.0
1 $3.83 53.09 13.797 Launch
h100nvl-1.16.96.160
16,384.0
1 $4.11 87.21 17.137 Launch
teslaa100-2.24.96.160.nvlink
16,384.0
tensor
2 $4.61 36.868 Launch
h200-1.16.128.160
16,384.0
1 $4.74 31.876 Launch
h200-2.24.256.160.nvlink
16,384.0
tensor
2 $9.40 72.520 Launch

Related models

Need help?

Contact our dedicated neural networks support team at support@immers.cloud or send your request to the sales department at sale@immers.cloud.

We use cookies and web analytics services to ensure the proper functioning of the website, analyze web-site traffic, and improve the quality of our services.
By continuing to use the website, you consent to the Privacy Policy and consent to the processing of cookie files and technical data.