← All models

GLM 5.3 Flash (320B-A18B, IQ2_XXS)

reasoningMITtools

antirez's routed 2-bit GGUF of Z.ai's GLM 5.3 Flash — a natively multimodal 320B-parameter Mixture-of-Experts with 18B active parameters, a 1M-token context, recurrent KDA layers, sparse DSA attention, hyper-connections, and a built-in MTP draft block. The Q2 file is about 90 GiB and targets a 128 GB Mac or DGX Spark with a modest initial context. Its separately packaged vision encoder adds image understanding to the same text model when Gezel launches ds4 with `--vision`.

At a glance

Parameters
320B
Quantization
n/a
Context window
1049K tokens
Approx. size
97.6 GB
Engines
n/a
License
MIT (open)
Version
v1.0.0 , released 2026-09-06
Family
glm

Tuning

Defaults a gezel client applies out of the box, from real evaluation runs against this exact quantization.

sampling.temperature
1
sampling.topP
0.95
sampling.maxTokens
16384
reasoning.thinkingBudget
4096
reasoning.enableThinking
true
profiles
thinking-general, thinking-coding, thinking-precise, instruct, creative

Model behaviors

Client-side behavior modules the catalog enables for this model (reasoning-tag handling, fabrication detection, prompt shaping).

Sources

Upstream
https://huggingface.co/zai-org/GLM-5.3-Flash
Manifest
View on GitHub

Tags