Model catalog

NVIDIA Nemotron 3 Super 120B-A12B

NVIDIA's text MoE has 120B total and 12B active parameters; weight planning uses total parameters.

Model overview

Available quantizations

Lower-bit quantizations generally use less memory. The compatibility calculator uses these candidates and its transparent approximate memory assumptions; it does not make a performance guarantee.

Q4_K

The conversion repository lists Nemotron-3-Super-120B-Q4_K.gguf at 69.9 GB; its rounded decimal file size is converted to GiB for planning. This does not include context or runtime memory.

Bits per weight
4.5
Model size
65.08 GiB

Compatibility notes

Estimates use the cataloged model-weight size (or a parameter-based estimate), apply weight overhead, and include standard runtime overhead. They do not include context/KV-cache memory. Real requirements vary with runtime, drivers, and other system use. A model page cannot determine compatibility without your hardware profile; use the calculator for that assessment.

NVIDIA Nemotron 3 Super 120B-A12B publisher model card; verified value, last verified 2026-10-01. NVIDIA documents 120B total and 12B active parameters, the NVIDIA Nemotron Open Model License, and up to 1M context. Weight planning uses the total count. View source