Muse Glimmer (30B, Q4)
Meta Superintelligence Labs' Muse Glimmer at 30B parameters, distilled from the larger Muse Spark and built for long-horizon tool use on everyday hardware. Multimodal (text + image), 131K context, native tool calling, and adjustable reasoning effort. Meta's own build, sized for machines with 24GB+ memory.
At a glance
- Parameters
- 30B
- Quantization
- Q4_K_M
- Context window
- 131K tokens
- Approx. size
- 16.8 GB
- Engines
- llama.cpp
- License
- Apache-2.0 (open)
- Version
- v1.0.2 , released 2026-08-23
- Family
- muse
Tuning
Defaults a gezel client applies out of the box, from real evaluation runs against this exact quantization.
- sampling.temperature
1- sampling.topP
0.95- sampling.topK
64- sampling.maxTokens
8192- reasoning.thinkingBudget
4096- reasoning.enableThinking
true- reasoning.templateKwargs
{"reasoning_strength":"high"}- profiles
- thinking-general, thinking-coding, thinking-precise, instruct, creative
Sources
- llama.cpp (GGUF)
- meta-models/Muse-Glimmer-30B-GGUF ·
muse-glimmer-30B-kquant-17gb.gguf· Q4_K_M - Upstream
- https://huggingface.co/meta-models/Muse-Glimmer-30B
- Manifest
- View on GitHub
Tags
- meta
- agentic
- multimodal
- vision
- tools
- long-context
- reasoning