GLM 5.3 Flash (320B-A18B, IQ2_XXS)
antirez's routed 2-bit GGUF of Z.ai's GLM 5.3 Flash — a natively multimodal 320B-parameter Mixture-of-Experts with 18B active parameters, a 1M-token context, recurrent KDA layers, sparse DSA attention, hyper-connections, and a built-in MTP draft block. The Q2 file is about 90 GiB and targets a 128 GB Mac or DGX Spark with a modest initial context. Its separately packaged vision encoder adds image understanding to the same text model when Gezel launches ds4 with `--vision`.
At a glance
- Parameters
- 320B
- Quantization
- n/a
- Context window
- 1049K tokens
- Approx. size
- 97.6 GB
- Engines
- n/a
- License
- MIT (open)
- Version
- v1.0.0 , released 2026-09-06
- Family
- glm
Tuning
Defaults a gezel client applies out of the box, from real evaluation runs against this exact quantization.
- sampling.temperature
1- sampling.topP
0.95- sampling.maxTokens
16384- reasoning.thinkingBudget
4096- reasoning.enableThinking
true- profiles
- thinking-general, thinking-coding, thinking-precise, instruct, creative
Model behaviors
Client-side behavior modules the catalog enables for this model (reasoning-tag handling, fabrication detection, prompt shaping).
- reasoning.strip-think-tags
- prompt.tool-cookbook-condensed
- prompt.meester-build-prelude
- mcp.compact-tool-schemas
- fabrication.detect-past-tense-no-tools
- turn.preamble-folding
Sources
- Upstream
- https://huggingface.co/zai-org/GLM-5.3-Flash
- Manifest
- View on GitHub
Tags
- glm
- zai
- agentic
- coding
- multimodal
- vision
- reasoning
- moe
- tools
- 2-bit
- iq2
- long-context
- large