← All models

DeepSeek V4 Flash (MXFP4)

reasoningMITtools

antirez's MXFP4 community GGUF of DeepSeek V4 Flash 0731 — the July 2026 re-post-trained 284B-parameter Mixture-of-Experts (~13B active per token) with 1M-token context and Compressed Sparse Attention. Routed experts use the MXFP4 micro-scaled 4-bit block format while attention projections, shared experts, router, embeddings, and output head stay at Q8_0 / F16, landing at ~156 GB — close to the native FP4 build's footprint and the highest-fidelity single-file ds4 variant. Needs a workstation or Mac Studio-class machine with ~200 GB of usable memory. Produced for antirez's `ds4` inference engine.

At a glance

Parameters
284B
Quantization
MXFP4
Context window
1000K tokens
Approx. size
156.0 GB
Engines
llama.cpp
License
MIT (open)
Version
v1.0.0 , released 2026-08-01
Family
deepseek

Tuning

Defaults a gezel client applies out of the box, from real evaluation runs against this exact quantization.

sampling.maxTokens
16384
reasoning.thinkingBudget
4096
reasoning.enableThinking
true
profiles
thinking-general, thinking-coding, thinking-precise, instruct, creative

Model behaviors

Client-side behavior modules the catalog enables for this model (reasoning-tag handling, fabrication detection, prompt shaping).

Sources

llama.cpp (GGUF)
antirez/deepseek-v4-gguf · DeepSeek-V4-Flash-MXFP4Experts-F16HC-F16Compressor-F16Indexer-Q8Attn-Q8Shared-Q8Out-chat-v2-mxfp4-0731.gguf · MXFP4
Upstream
https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash
Manifest
View on GitHub

Tags