Model overview
- Family
- gpt-oss
- Provider
- OpenAI
- Architecture
- GPT-OSS MoE
- Parameters
- 116.83B
- Formats
- gguf, safetensors
- Runtimes
- llama.cpp
- License
- Apache-2.0
- Maximum context
- 131,072 tokens
Available quantizations
Lower-bit quantizations generally use less memory. The compatibility calculator uses these candidates and its transparent approximate memory assumptions; it does not make a performance guarantee.
MXFP4
Main-model gpt-oss-120b-MXFP4.gguf: 63,387,346,208 bytes, converted with bytes / 2^30. MoE weights use MXFP4 while other tensors retain higher precision; a whole-model bits-per-weight value is not documented. Separate Eagle speculative-decoding files are excluded. File size is not runtime RAM/VRAM usage.
- Bits per weight
- Unavailable (whole-model value not documented)
- Model size
- 59.03 GiB
Compatibility notes
Estimates use the cataloged model-weight size (or a parameter-based estimate), apply weight overhead, and include standard runtime overhead. They do not include context/KV-cache memory. Real requirements vary with runtime, drivers, and other system use. A model page cannot determine compatibility without your hardware profile; use the calculator for that assessment.
OpenAI gpt-oss-120b publisher model card; verified value, last verified 2026-10-01. OpenAI documents Apache-2.0 and Harmony formatting. The linked publisher model card (https://cdn.openai.com/pdf/419b6906-9da6-406c-a19d-1bb078ac7637/oai_gpt-oss_model_card.pdf), Table 1 and section 2.2, gives 116.83B total, 5.13B active parameters and 131,072-token maximum context. Active parameters are descriptive only; no runtime default context is assumed. View source
Runtime prerequisite: Use a current llama.cpp build with GPT-OSS/MXFP4 support and the model's Harmony chat template (the llama.cpp guide uses --jinja). Memory fit does not guarantee runtime or backend compatibility. Context/KV-cache and runtime-specific buffers can require additional memory and are not separately estimated. Memory compatibility does not confirm runtime-load compatibility. llama.cpp GPT-OSS runtime guide