GLM-5.3-Flash

reasoning
multimodal
coding

GLM-5.3-Flash is the first natively multimodal model in the GLM-5 series. Before its official release, the model was tested anonymously as ox-alpha on OpenCode and OpenRouter and quickly became the most popular model of the week by volume of processed tokens. The model uses a sparse mixture of experts (MoE) with 288 routed experts and 8 active per token plus one shared expert — thus, out of 320B total parameters, only about 18B are activated per token. The language component contains 45 layers and works in tandem with a 24-layer vision encoder for image and video processing. The model supports a configurable reasoning budget via the reasoning_effort parameter with three levels: low, high, and max (default is max).

GLM-5.3-Flash is built on a fundamentally new architecture designed around inference efficiency. The main architectural innovation is hybrid attention, first applied in the GLM series. It combines two types of layers: 34 Kimi Delta Attention (KDA) linear attention layers and 11 DeepSeek Sparse Attention (DSA) sparse attention layers, alternating in a fixed 3:1 pattern. The point of the hybrid is that local linear attention processes the sequence with constant computational complexity, without requiring storage of a growing KV cache, while sparse attention analyzes context globally, maintaining long-range accuracy. In practice, this reduces attention compute cost by about 3.01x and KV cache size by 4.44x compared with the GLM-5.3 architecture.

Additionally, the model uses Manifold-Constrained Hyper-Connections (mHC) — a replacement for the standard residual stream — which helps stabilize training at scale. Another detail: the Multi-head Latent Attention layers operate in NoPE mode (without positional encoding) — position tracking is delegated to the linear attention layers.

The model demonstrates strong results on agentic and coding benchmarks. DeepSWE v1.1 (assessing an agent’s ability to autonomously solve tasks in a real development environment) grew from 46.2 for GLM-5.2 to 63.4 — one of the largest improvements in the series. AutomationBench (testing automation of routine workflows) rose from 26.2 to 48.8. Terminal-Bench 2.1 (working in the terminal and command line) reached 84.3 — coming very close to Claude Opus 4.8’s result of 85.0. On GDPval-AA v2 (evaluating the practical value of an agent in real tasks), the model leads with 1773 Elo.

Thanks to its native multimodality and hybrid architecture, GLM-5.3-Flash is suitable for a wide range of tasks. First and foremost are agentic scenarios and workflow automation. Next, another strong practical application area is software development and programming. Multimodal analysis makes it possible to integrate the model into processes that require understanding images, tables, charts, and video. Developers separately highlight full-cycle Office scenarios, from reviewing financial, legal, and any other documents to building reports, charts, Excel spreadsheets, and presentations based on them. One should also not forget the context length, which makes it possible to handle the above tasks much more efficiently and without additional tools or iterations.


Announce Date: 25.08.2026
Parameters: 322B
Experts: 288
Activated at inference: 2B
Context: 1049K
Layers: 45, using full attention: 11
Attention Type: Kimi Delta Attention (KDA) + DeepSeek Spare Attention (DSA)
Developer: Z.ai
Transformers Version: 5.16.0
vLLM Version: >=0.29.0
License: MIT

Public endpoint

Use our pre-built public endpoints for free to test inference and explore GLM-5.3-Flash capabilities. You can obtain an API access token on the token management page after registration and verification.
Model Name Context Type GPU Status Link
There are no public endpoints for this model yet.

Private server

Rent your own physically dedicated instance with hourly or long-term monthly billing.

We recommend deploying private instances in the following scenarios:

  • maximize endpoint performance,
  • enable full context for long sequences,
  • ensure top-tier security for data processing in an isolated, dedicated environment,
  • use custom weights, such as fine-tuned models or LoRA adapters.

Recommended server configurations for hosting GLM-5.3-Flash

Prices:
Name GPU Price, hour TPS Max Concurrency
h200-2.24.256.320
1,048,576.0
tensor
2 $9.42 2.865 Launch
h200-2.24.256.320.nvlink
1,048,576.0
tensor
2 $9.42 2.865 Launch
h100nvl-4.32.384.480
1,048,576.0
tensor
4 $16.23 3.533 Launch
dedicated-h100-8.96.768.5760-1.nvlink
1,048,576.0
tensor
8 $31.39 5.124 Launch
Prices:
Name GPU Price, hour TPS Max Concurrency
teslaa100-8.44.512.480.nvlink
1,048,576.0
tensor
8 $18.35 14.319 Launch
h200-4.32.768.480
1,048,576.0
tensor
4 $19.23 13.095 Launch
h200-4.32.768.480.nvlink
1,048,576.0
tensor
4 $19.23 13.095 Launch
dedicated-h100-8.96.768.5760-1.nvlink
1,048,576.0
tensor
8 $31.39 14.275 Launch
Prices:
Name GPU Price, hour TPS Max Concurrency
h200-8.52.1024.960
1,048,576.0
tensor
8 $37.37 27.298 Launch
h200-8.52.1024.960.nvlink
1,048,576.0
tensor
8 $37.37 27.298 Launch

Need help?

Contact our dedicated neural networks support team at support@immers.cloud or send your request to the sales department at sale@immers.cloud.

We use cookies and web analytics services to ensure the proper functioning of the website, analyze web-site traffic, and improve the quality of our services.
By continuing to use the website, you consent to the Privacy Policy and consent to the processing of cookie files and technical data.