Q4_K
The conversion repository lists Nemotron-3-Super-120B-Q4_K.gguf at 69.9 GB; its rounded decimal file size is converted to GiB for planning. This does not include context or runtime memory.
- Bits per weight
- 4.5
- Model size
- 65.08 GiB
Model catalog
NVIDIA's text MoE has 120B total and 12B active parameters; weight planning uses total parameters.
Lower-bit quantizations generally use less memory. The compatibility calculator uses these candidates and its transparent approximate memory assumptions; it does not make a performance guarantee.
The conversion repository lists Nemotron-3-Super-120B-Q4_K.gguf at 69.9 GB; its rounded decimal file size is converted to GiB for planning. This does not include context or runtime memory.
Estimates use the cataloged model-weight size (or a parameter-based estimate), apply weight overhead, and include standard runtime overhead. They do not include context/KV-cache memory. Real requirements vary with runtime, drivers, and other system use. A model page cannot determine compatibility without your hardware profile; use the calculator for that assessment.
NVIDIA Nemotron 3 Super 120B-A12B publisher model card; verified value, last verified 2026-10-01. NVIDIA documents 120B total and 12B active parameters, the NVIDIA Nemotron Open Model License, and up to 1M context. Weight planning uses the total count. View source