Q4_K_M
The conversion repository lists GLM-4.7-Flash-Q4_K_M.gguf at 18.1 GB; its rounded decimal file size is converted to GiB for planning. This does not include context or runtime memory.
- Bits per weight
- 4.5
- Model size
- 16.86 GiB
Model catalog
Z.ai's 30B MoE model activates 3B parameters per token. Weight planning uses the 30B total; active parameters are descriptive only.
Lower-bit quantizations generally use less memory. The compatibility calculator uses these candidates and its transparent approximate memory assumptions; it does not make a performance guarantee.
The conversion repository lists GLM-4.7-Flash-Q4_K_M.gguf at 18.1 GB; its rounded decimal file size is converted to GiB for planning. This does not include context or runtime memory.
Estimates use the cataloged model-weight size (or a parameter-based estimate), apply weight overhead, and include standard runtime overhead. They do not include context/KV-cache memory. Real requirements vary with runtime, drivers, and other system use. A model page cannot determine compatibility without your hardware profile; use the calculator for that assessment.
Z.ai GLM-4.7-Flash publisher model card; verified value, last verified 2026-09-30. Z.ai identifies GLM-4.7-Flash as a 30B-A3B MoE model under MIT. No context limit is recorded because the publisher's current model card does not clearly establish one for this variant. View source