Model catalog

Bonsai 2 27B

PrismML's ternary 27B reasoning model derived from Qwen3.8-27B; the memory estimate covers its text model and excludes the optional vision projector.

Model overview

Available quantizations

Lower-bit quantizations generally use less memory. The compatibility calculator uses these candidates and its transparent approximate memory assumptions; it does not make a performance guarantee.

PTQ1_0

Publisher-documented 5.95 GB packed text-model file, converted to GiB for approximate memory planning; excludes the optional vision projector and runtime/context memory.

Bits per weight
1.75
Model size
5.54 GiB

PQ2_0

Publisher-documented 7.21 GB packed text-model file, converted to GiB for approximate memory planning; excludes the optional vision projector and runtime/context memory.

Bits per weight
2.13
Model size
6.72 GiB

Compatibility notes

Weight estimates are approximate and include a runtime overhead. Real requirements vary with context length, runtime, drivers, and other system use. A model page cannot determine compatibility without your hardware profile; use the calculator for that assessment.

PrismML Bonsai 2 27B GGUF model card; verified value, last verified 2026-09-26. Publisher metadata identifies a 27.36B model derived from Qwen3.8-27B, Apache-2.0 licensing, and 262,144-token maximum context. Memory inputs cover text-model packing only; the optional vision projector is excluded. The publisher documents a dedicated compatible runtime requirement separately. View source

Runtime prerequisite: PTQ1_0 and PQ2_0 require PrismML's llama.cpp fork with its custom ternary and Hadamard kernels. Stock llama.cpp is not safe for these packings. Memory compatibility does not confirm runtime-load compatibility. PrismML Bonsai demo runtime guide