Model overview
- Family
- Qwen3 Coder
- Provider
- Qwen
- Architecture
- Qwen3-Next MoE
- Parameters
- 80B
- Formats
- gguf, safetensors
- Runtimes
- llama.cpp
- License
- Apache-2.0
- Maximum context
- 262,144 tokens
Available quantizations
Lower-bit quantizations generally use less memory. The compatibility calculator uses these candidates and its transparent approximate memory assumptions; it does not make a performance guarantee.
Q4_K_M
Qwen's GGUF repository lists four Q4_K_M shards totaling 48.4 GB: Qwen3-Coder-Next-Q4_K_M-00001-of-00004.gguf, Qwen3-Coder-Next-Q4_K_M-00002-of-00004.gguf, Qwen3-Coder-Next-Q4_K_M-00003-of-00004.gguf, and Qwen3-Coder-Next-Q4_K_M-00004-of-00004.gguf. The rounded decimal total is converted to GiB for planning; context and runtime memory are not included.
- Bits per weight
- 4.5
- Model size
- 45.08 GiB
Compatibility notes
Estimates use the cataloged model-weight size (or a parameter-based estimate), apply weight overhead, and include standard runtime overhead. They do not include context/KV-cache memory. Real requirements vary with runtime, drivers, and other system use. A model page cannot determine compatibility without your hardware profile; use the calculator for that assessment.
Qwen3-Coder-Next publisher model card; verified value, last verified 2026-09-30. Qwen documents 80B total parameters, 3B activated parameters, Apache-2.0 licensing, and 256K native context. Weight planning uses total parameters; GGUF files come from Qwen's separately hosted conversion repository. View source