Ternary Bonsai 27B (2-bit)
PrismML's ternary 2-bit adaptation of Qwen 3.6 27B. Multimodal (text + image + video), 262K native context, thinking and tool-calling support, packaged for Apple Silicon in roughly 8.5GB. A promising 27B-class on-device option for Macs with 16GB+ unified memory; 24GB+ is recommended for useful context headroom.
At a glance
- Parameters
- 27B
- Quantization
- 2bit
- Context window
- 262K tokens
- Approx. size
- 8.5 GB
- Engines
- MLX
- License
- Apache-2.0 (open)
- Version
- v1.0.0 , released 2026-07-14
- Family
- qwen
Tuning
Defaults a gezel client applies out of the box, from real evaluation runs against this exact quantization.
- sampling.temperature
0.7- sampling.topP
0.95- sampling.topK
20- sampling.minP
0- sampling.maxTokens
8192- reasoning.thinkingBudget
4096- reasoning.enableThinking
true- profiles
- thinking-general, thinking-coding, thinking-precise, instruct, creative
Model behaviors
Client-side behavior modules the catalog enables for this model (reasoning-tag handling, fabrication detection, prompt shaping).
- reasoning.strip-think-tags
- provider.merge-system-messages
- mcp.compact-tool-schemas
- provider.compact-write-transcript
- fabrication.detect-past-tense-no-tools
- turn.ollama-num-predict-bumped
- turn.preamble-folding
- turn.ramble-detection
- tools.mlx-grammar
- prompt.derive-by-execution
Sources
- MLX
- prism-ml/Ternary-Bonsai-27B-mlx-2bit
- Upstream
- https://huggingface.co/prism-ml/Ternary-Bonsai-27B-mlx-2bit
- Manifest
- View on GitHub
Tags
- prismml
- qwen
- ternary
- 2-bit
- multimodal
- vision
- video
- tools
- reasoning
- long-context