Model overview
- Family
- DeepSeek V4
- Provider
- DeepSeek
- Architecture
- DeepSeek V4 MoE
- Parameters
- 284B
- Formats
- gguf, safetensors
- Runtimes
- llama.cpp, ollama
- License
- MIT
Available quantizations
Lower-bit quantizations generally use less memory. The compatibility calculator uses these candidates and its transparent approximate memory assumptions; it does not make a performance guarantee.
UD-Q4_K_XL
The conversion repository lists the UD-Q4_K_XL directory at 155 GB across five shards: DeepSeek-V4-Flash-0731-UD-Q4_K_XL-00001-of-00005.gguf, DeepSeek-V4-Flash-0731-UD-Q4_K_XL-00002-of-00005.gguf, DeepSeek-V4-Flash-0731-UD-Q4_K_XL-00003-of-00005.gguf, DeepSeek-V4-Flash-0731-UD-Q4_K_XL-00004-of-00005.gguf, and DeepSeek-V4-Flash-0731-UD-Q4_K_XL-00005-of-00005.gguf. The rounded total is converted to GiB for planning.
- Bits per weight
- 4.5
- Model size
- 144.35 GiB
Compatibility notes
Estimates use the cataloged model-weight size (or a parameter-based estimate), apply weight overhead, and include standard runtime overhead. They do not include context/KV-cache memory. Real requirements vary with runtime, drivers, and other system use. A model page cannot determine compatibility without your hardware profile; use the calculator for that assessment.
DeepSeek V4 Flash 0731 publisher model card; verified value, last verified 2026-09-29. DeepSeek identifies the 0731 checkpoint as its official V4 Flash release and licenses the model under MIT. The publisher card lists 284B total parameters. Weight planning uses that total count; no context limit is added here. View source