LLMGauge · GPU compatibility

Which GPUs can run this model?

Choose a catalog model and enter your system RAM to compare its catalogued quantizations against discrete GPUs in the LLMGauge catalog.

01

Choose a model and enter RAM

System RAM is held constant across every GPU result.

Only models in the curated catalog are available.

Enter the same available system RAM value for all GPU comparisons.

Advanced settings

Optional runtime planning

These inputs can establish a cache-memory placement assumption. They do not verify drivers, backend support, actual allocation, or speed.

Choose llama.cpp when that is the runtime you plan to use.
This is an assumption, not a capability check.
The preference is advisory and does not override the result.
Optional planning target for the FP16 KV-cache estimate. Choose a common value below or enter a custom number. Cache memory is included in fit only when its placement is explicit; otherwise it remains advisory.

The representative quantization follows LLMGauge’s existing fit and metadata policy. It is not a model-quality or performance recommendation. Estimates are approximate; see the single-model calculator or hardware recommendations for other workflows.