← All models

Muse Glimmer (30B, Q4)

generalApache 2.0tools

Meta Superintelligence Labs' Muse Glimmer at 30B parameters, distilled from the larger Muse Spark and built for long-horizon tool use on everyday hardware. Multimodal (text + image), 131K context, native tool calling, and adjustable reasoning effort. Meta's own build, sized for machines with 24GB+ memory.

At a glance

Parameters
30B
Quantization
Q4_K_M
Context window
131K tokens
Approx. size
16.8 GB
Engines
llama.cpp
License
Apache-2.0 (open)
Version
v1.0.2 , released 2026-08-23
Family
muse

Tuning

Defaults a gezel client applies out of the box, from real evaluation runs against this exact quantization.

sampling.temperature
1
sampling.topP
0.95
sampling.topK
64
sampling.maxTokens
8192
reasoning.thinkingBudget
4096
reasoning.enableThinking
true
reasoning.templateKwargs
{"reasoning_strength":"high"}
profiles
thinking-general, thinking-coding, thinking-precise, instruct, creative

Sources

llama.cpp (GGUF)
meta-models/Muse-Glimmer-30B-GGUF · muse-glimmer-30B-kquant-17gb.gguf · Q4_K_M
Upstream
https://huggingface.co/meta-models/Muse-Glimmer-30B
Manifest
View on GitHub

Tags