GLM-5.3-Flash is the first natively multimodal model in the GLM-5 series. Before its official release, the model was tested anonymously as ox-alpha on OpenCode and OpenRouter and quickly became the most popular model of the week by volume of processed tokens. The model uses a sparse mixture of experts (MoE) with 288 routed experts and 8 active per token plus one shared expert — thus, out of 320B total parameters, only about 18B are activated per token. The language component contains 45 layers and works in tandem with a 24-layer vision encoder for image and video processing. The model supports a configurable reasoning budget via the reasoning_effort parameter with three levels: low, high, and max (default is max).
GLM-5.3-Flash is built on a fundamentally new architecture designed around inference efficiency. The main architectural innovation is hybrid attention, first applied in the GLM series. It combines two types of layers: 34 Kimi Delta Attention (KDA) linear attention layers and 11 DeepSeek Sparse Attention (DSA) sparse attention layers, alternating in a fixed 3:1 pattern. The point of the hybrid is that local linear attention processes the sequence with constant computational complexity, without requiring storage of a growing KV cache, while sparse attention analyzes context globally, maintaining long-range accuracy. In practice, this reduces attention compute cost by about 3.01x and KV cache size by 4.44x compared with the GLM-5.3 architecture.
Additionally, the model uses Manifold-Constrained Hyper-Connections (mHC) — a replacement for the standard residual stream — which helps stabilize training at scale. Another detail: the Multi-head Latent Attention layers operate in NoPE mode (without positional encoding) — position tracking is delegated to the linear attention layers.
The model demonstrates strong results on agentic and coding benchmarks. DeepSWE v1.1 (assessing an agent’s ability to autonomously solve tasks in a real development environment) grew from 46.2 for GLM-5.2 to 63.4 — one of the largest improvements in the series. AutomationBench (testing automation of routine workflows) rose from 26.2 to 48.8. Terminal-Bench 2.1 (working in the terminal and command line) reached 84.3 — coming very close to Claude Opus 4.8’s result of 85.0. On GDPval-AA v2 (evaluating the practical value of an agent in real tasks), the model leads with 1773 Elo.
Thanks to its native multimodality and hybrid architecture, GLM-5.3-Flash is suitable for a wide range of tasks. First and foremost are agentic scenarios and workflow automation. Next, another strong practical application area is software development and programming. Multimodal analysis makes it possible to integrate the model into processes that require understanding images, tables, charts, and video. Developers separately highlight full-cycle Office scenarios, from reviewing financial, legal, and any other documents to building reports, charts, Excel spreadsheets, and presentations based on them. One should also not forget the context length, which makes it possible to handle the above tasks much more efficiently and without additional tools or iterations.
| Model Name | Context | Type | GPU | Status | Link |
|---|
There are no public endpoints for this model yet.
Rent your own physically dedicated instance with hourly or long-term monthly billing.
We recommend deploying private instances in the following scenarios:
| Name | GPU | TPS | Max Concurrency | |||
|---|---|---|---|---|---|---|
1,048,576.0 tensor |
2 | $9.42 | 2.865 | Launch | ||
1,048,576.0 tensor |
2 | $9.42 | 2.865 | Launch | ||
1,048,576.0 tensor |
4 | $16.23 | 3.533 | Launch | ||
1,048,576.0 tensor |
8 | $31.39 | 5.124 | Launch | ||
| Name | GPU | TPS | Max Concurrency | |||
|---|---|---|---|---|---|---|
1,048,576.0 tensor |
8 | $18.35 | 14.319 | Launch | ||
1,048,576.0 tensor |
4 | $19.23 | 13.095 | Launch | ||
1,048,576.0 tensor |
4 | $19.23 | 13.095 | Launch | ||
1,048,576.0 tensor |
8 | $31.39 | 14.275 | Launch | ||
Contact our dedicated neural networks support team at support@immers.cloud or send your request to the sales department at sale@immers.cloud.