Model catalog

NVIDIA Nemotron 3 Nano 30B-A3B

NVIDIA's text MoE has 30B total and 3.5B active parameters; weight planning uses total parameters.

Model overview

Available quantizations

Lower-bit quantizations generally use less memory. The compatibility calculator uses these candidates and its transparent approximate memory assumptions; it does not make a performance guarantee.

Q4_K_M

The conversion repository lists NVIDIA-Nemotron-3-Nano-30B-A3B-Q4_K_M.gguf at 22.4 GB; its rounded decimal file size is converted to GiB for planning. This does not include context or runtime memory.

Bits per weight
4.5
Model size
20.86 GiB

Compatibility notes

Estimates use the cataloged model-weight size (or a parameter-based estimate), apply weight overhead, and include standard runtime overhead. They do not include context/KV-cache memory. Real requirements vary with runtime, drivers, and other system use. A model page cannot determine compatibility without your hardware profile; use the calculator for that assessment.

NVIDIA Nemotron 3 Nano 30B-A3B publisher model card; verified value, last verified 2026-10-01. NVIDIA documents 30B total and 3.5B active parameters and the NVIDIA Nemotron Open Model License. Weight planning uses the total count; no context limit is recorded. View source