Model overview
- Family
- Mistral Small 4
- Provider
- Mistral AI
- Architecture
- Mistral MoE
- Parameters
- 119B
- Formats
- gguf, safetensors
- Runtimes
- llama.cpp
- License
- Apache-2.0
- Maximum context
- 262,144 tokens
Available quantizations
Lower-bit quantizations generally use less memory. The compatibility calculator uses these candidates and its transparent approximate memory assumptions; it does not make a performance guarantee.
Q4_K_M
The conversion repository lists two Q4_K_M shards totaling 72.64 GB: mistralai_Mistral-Small-4-119B-2603-Q4_K_M-00001-of-00002.gguf and mistralai_Mistral-Small-4-119B-2603-Q4_K_M-00002-of-00002.gguf. The rounded decimal total is converted to GiB for weight planning; vision-related runtime memory is not included.
- Bits per weight
- 4.5
- Model size
- 67.68 GiB
Compatibility notes
Estimates use the cataloged model-weight size (or a parameter-based estimate), apply weight overhead, and include standard runtime overhead. They do not include context/KV-cache memory. Real requirements vary with runtime, drivers, and other system use. A model page cannot determine compatibility without your hardware profile; use the calculator for that assessment.
Mistral Small 4 119B publisher model card; verified value, last verified 2026-10-01. Mistral documents 119B total parameters, 6.5B active parameters, Apache-2.0, multimodal image input, and up to 256K context. Weight planning uses the total count; vision-related runtime memory is not separately estimated. View source
Multimodal memory scope: The catalog has not verified whether the selected GGUF weights include all multimodal components or require separate files. Multimodal runtime memory is not estimated and may be missing from this result. bartowski Mistral Small 4 GGUF repository