Model overview
- Family
- Devstral Small 2
- Provider
- Mistral AI
- Architecture
- Mistral
- Parameters
- 24B
- Formats
- gguf, safetensors
- Runtimes
- llama.cpp
- License
- Apache-2.0
- Maximum context
- 262,144 tokens
Available quantizations
Lower-bit quantizations generally use less memory. The compatibility calculator uses these candidates and its transparent approximate memory assumptions; it does not make a performance guarantee.
Q4_K_M
The conversion repository lists Devstral-Small-2-24B-Instruct-2512-Q4_K_M.gguf at 14.3 GB; its rounded decimal file size is converted to GiB for planning. This does not include context or runtime memory.
- Bits per weight
- 4.5
- Model size
- 13.32 GiB
Compatibility notes
Estimates use the cataloged model-weight size (or a parameter-based estimate), apply weight overhead, and include standard runtime overhead. They do not include context/KV-cache memory. Real requirements vary with runtime, drivers, and other system use. A model page cannot determine compatibility without your hardware profile; use the calculator for that assessment.
Devstral Small 2 24B Instruct 2512 publisher model card; verified value, last verified 2026-10-01. Mistral documents this 24B vision-capable instruct model under Apache-2.0 with a 256K context. The catalog estimates main model weights only; multimodal runtime memory is not separately estimated. View source
Multimodal memory scope: This estimate excludes the separate vision/projector file and its runtime memory. lmstudio-community Devstral Small 2 GGUF repository