Model overview
- Family
- Qwen3.5
- Provider
- Qwen
- Architecture
- Qwen3.5 MoE
- Parameters
- 122B
- Formats
- gguf, safetensors
- Runtimes
- llama.cpp
- License
- Apache-2.0
- Maximum context
- 262,144 tokens
Available quantizations
Lower-bit quantizations generally use less memory. The compatibility calculator uses these candidates and its transparent approximate memory assumptions; it does not make a performance guarantee.
Q4_K_M
The conversion repository lists two Q4_K_M shards totaling 77.62 GB: Qwen_Qwen3.5-122B-A10B-Q4_K_M-00001-of-00002.gguf and Qwen_Qwen3.5-122B-A10B-Q4_K_M-00002-of-00002.gguf. The rounded decimal total is converted to GiB for weight planning; multimodal runtime and context memory are not included.
- Bits per weight
- 4.5
- Model size
- 72.29 GiB
Compatibility notes
Estimates use the cataloged model-weight size (or a parameter-based estimate), apply weight overhead, and include standard runtime overhead. They do not include context/KV-cache memory. Real requirements vary with runtime, drivers, and other system use. A model page cannot determine compatibility without your hardware profile; use the calculator for that assessment.
Qwen3.5-122B-A10B publisher model card; verified value, last verified 2026-09-30. Qwen documents 122B total parameters, 10B activated parameters, Apache-2.0 licensing, multimodal image input, and native 262,144-token context. Weight planning uses total parameters; multimodal runtime memory is not separately estimated. View source
Multimodal memory scope: The catalog has not verified whether the selected GGUF weights include all multimodal components or require separate files. Multimodal runtime memory is not estimated and may be missing from this result. bartowski Qwen3.5-122B-A10B GGUF repository