Model catalog

OpenAI gpt-oss-120b

OpenAI's text reasoning MoE has 116.83B total and 5.13B active parameters per token. The mixed-precision MXFP4 GGUF uses its sourced file size for approximate weight planning; context memory is not estimated.

Model overview

Available quantizations

Lower-bit quantizations generally use less memory. The compatibility calculator uses these candidates and its transparent approximate memory assumptions; it does not make a performance guarantee.

MXFP4

Main-model gpt-oss-120b-MXFP4.gguf: 63,387,346,208 bytes, converted with bytes / 2^30. MoE weights use MXFP4 while other tensors retain higher precision; a whole-model bits-per-weight value is not documented. Separate Eagle speculative-decoding files are excluded. File size is not runtime RAM/VRAM usage.

Bits per weight
Unavailable (whole-model value not documented)
Model size
59.03 GiB

Compatibility notes

Estimates use the cataloged model-weight size (or a parameter-based estimate), apply weight overhead, and include standard runtime overhead. They do not include context/KV-cache memory. Real requirements vary with runtime, drivers, and other system use. A model page cannot determine compatibility without your hardware profile; use the calculator for that assessment.

OpenAI gpt-oss-120b publisher model card; verified value, last verified 2026-10-01. OpenAI documents Apache-2.0 and Harmony formatting. The linked publisher model card (https://cdn.openai.com/pdf/419b6906-9da6-406c-a19d-1bb078ac7637/oai_gpt-oss_model_card.pdf), Table 1 and section 2.2, gives 116.83B total, 5.13B active parameters and 131,072-token maximum context. Active parameters are descriptive only; no runtime default context is assumed. View source

Runtime prerequisite: Use a current llama.cpp build with GPT-OSS/MXFP4 support and the model's Harmony chat template (the llama.cpp guide uses --jinja). Memory fit does not guarantee runtime or backend compatibility. Context/KV-cache and runtime-specific buffers can require additional memory and are not separately estimated. Memory compatibility does not confirm runtime-load compatibility. llama.cpp GPT-OSS runtime guide