Q4_K_M
A commonly available lower-memory GGUF variant; the listed memory input is an estimate, not the repository file size.
- Bits per weight
- 4.5
Model catalog
The 12B Gemma 3 instruction model adds a larger text-and-image family option; this estimate covers text weights only.
Lower-bit quantizations generally use less memory. The compatibility calculator uses these candidates and its transparent approximate memory assumptions; it does not make a performance guarantee.
A commonly available lower-memory GGUF variant; the listed memory input is an estimate, not the repository file size.
Weight estimates are approximate and include a runtime overhead. Real requirements vary with context length, runtime, drivers, and other system use. A model page cannot determine compatibility without your hardware profile; use the calculator for that assessment.
Google Gemma 3 12B IT model card; verified value, last verified 2026-09-23. Publisher metadata; image encoder, image tokens, and multimodal runtime memory are not modeled. The ggml-org repository provides a separate GGUF conversion. View source