Model catalog

Qwen3 4B

Qwen's compact 4B general-purpose model; a separate GGUF conversion is listed for llama.cpp workflows.

Model overview

Available quantizations

Lower-bit quantizations generally use less memory. The compatibility calculator uses these candidates and its transparent approximate memory assumptions; it does not make a performance guarantee.

Q4_K_M

A commonly available lower-memory GGUF variant; the listed memory input is an estimate, not the repository file size.

Bits per weight
4.5

Q8_0

A higher-memory GGUF variant where listed by the conversion repository; the listed memory input remains approximate.

Bits per weight
8

Compatibility notes

Weight estimates are approximate and include a runtime overhead. Real requirements vary with context length, runtime, drivers, and other system use. A model page cannot determine compatibility without your hardware profile; use the calculator for that assessment.

Qwen3 4B publisher model card; verified value, last verified 2026-09-23. Publisher model metadata and license; the separate Qwen GGUF repository is community conversion data. The catalog uses the documented native 32K context and does not assume extended-context settings. View source