GLM-4.6 is built on a Mixture-of-Experts (MoE) architecture with a total of 355 billion parameters, of which 32 billion are actively used per forward pass. GLM-4.6 (like its 4.5 version) employs a "deeper but narrower" strategy: the model features more layers with a smaller number of experts and a smaller hidden dimension compared to DeepSeek-V3 and Kimi K2. This architecture delivers superior performance on reasoning tasks. The model uses Grouped-Query Attention with partial RoPE, 96 attention heads for a hidden dimension of 5120 in 92 layers, QK normalization to stabilize attention logits, and the Muon optimizer for accelerated convergence.
GLM-4.6 offers several significant improvements over its predecessor: an increased context window from 128K to 200K tokens, enhanced programming capabilities, advanced reasoning, and efficiency—the model completes tasks using approximately 15% fewer tokens compared to GLM-4.5.According to the official release, GLM-4.6 was tested on eight public benchmarks covering agent tasks, reasoning, and programming. The results demonstrate the model's ability to confidently compete with leading models such as DeepSeek-V3.2-Exp and Claude Sonnet 4. For example: AIME 25 (Mathematical Reasoning) - 98.6%, significantly outperforming Claude Sonnet 4 (74.3%) and DeepSeek-V3.2-Exp (89.3%), LiveCodeBench v6 (Real-World Programming) - 84.5%, substantially ahead of GLM-4.5 (63.3%) and DeepSeek-V3.2-Exp (70.1%), BrowseComp (Agent Tasks with Web Search) - 45.1%, significantly surpassing GLM-4.5 (26.4%) and DeepSeek-V3.2-Exp (40.1%). In practical programming tasks, according to an extended CC-Bench test conducted by the developers, GLM-4.6 achieves practical parity with Claude Sonnet 4, showing a 48.6%-win rate in head-to-head comparisons when performing real-world tasks in frontend development, tool creation, data analysis, testing, and algorithms.
Thanks to its unique characteristics, GLM-4.6 is optimally suited for creating autonomous AI agents, professional software development (from frontend work to refactoring legacy code), analyzing large volumes of documents, creating educational content, and, finally, scientific research.
| Model Name | Context | Type | GPU | Status | Link |
|---|
There are no public endpoints for this model yet.
Rent your own physically dedicated instance with hourly or long-term monthly billing.
We recommend deploying private instances in the following scenarios:
| Name | GPU | TPS | Max Concurrency | |||
|---|---|---|---|---|---|---|
202,752.0 tensor |
4 | $9.17 | 46.290 | 1.337 | Launch | |
202,752.0 tensor |
4 | $9.50 | 1.337 | Launch | ||
202,752.0 pipeline |
3 | $14.36 | 2.680 | Launch | ||
202,752.0 tensor |
4 | $14.99 | 1.334 | Launch | ||
202,752.0 tensor |
4 | $16.23 | 2.053 | Launch | ||
202,752.0 tensor |
4 | $19.23 | 4.469 | Launch | ||
202,752.0 tensor |
4 | $19.23 | 4.469 | Launch | ||
| Name | GPU | TPS | Max Concurrency | |||
|---|---|---|---|---|---|---|
202,752.0 tensor |
8 | $18.35 | 3.097 | Launch | ||
202,752.0 tensor |
4 | $19.23 | 2.316 | Launch | ||
202,752.0 tensor |
4 | $19.23 | 2.316 | Launch | ||
202,752.0 tensor |
8 | $31.39 | 3.090 | Launch | ||
Contact our dedicated neural networks support team at nn@immers.cloud or send your request to the sales department at sale@immers.cloud.