PrismML: Ternary Bonsai 2 27B
About this model
PrismML Ternary Bonsai 2 27B is an Apache 2.0 open-weight multimodal language model derived from Qwen3.8-27B. Its language-model matrices are stored as native ternary weights in {-1, 0, +1} with FP16 group-wise scaling and blockwise Hadamard rotation, yielding about 1.76 effective bits per weight and roughly 5.95 GB on disk for the GGUF PTQ1_0 language weights—about nine times smaller than the FP16 baseline while PrismML reports roughly 98.2% retention on its primary thinking-mode benchmark suites.
The architecture keeps the Qwen3.8 hybrid-attention stack (~75% linear / ~25% full attention), SwiGLU MLP, RoPE, and RMSNorm, with a 262,144-token context window and optional vision input via a separate mmproj bundle. Weights ship as GGUF (requiring the PrismML llama.cpp fork for ternary kernels) and MLX 2-bit packs for Apple Silicon and CUDA. Thinking mode is the default evaluation and deployment posture; published scores use EvalScope with vLLM on NVIDIA H100 in matched thinking-mode settings.
Reported strengths cluster in mathematics, coding, and instruction following (for example GSM8K, MATH-500, AIME 2025/2026, LiveCodeBench, HumanEval+, and IFEval), with GPQA Diamond and multimodal tasks somewhat below the FP16 parent. Long-horizon agentic benchmarks such as SWE-Bench Verified and Terminal-Bench 2.1 show larger gaps, which PrismML attributes to compounding errors in extended tool-use loops rather than single-shot QA.
Benchmark Scores
Technical Specs
- Parameters: 27.36B
- Architecture: Hybrid-attention Transformer
- Context Window: 262,144 tokens
- Input Modalities: text, image
Hardware Requirements
- VRAM: 8.0 GB
- Compute: Single GPU with 8GB+ VRAM (e.g. NVIDIA RTX 4090) or Apple Silicon via MLX
Pricing
| Input | Output | Currency |
|---|---|---|
| 0.07 / 1M tokens | 0.50 / 1M tokens | USD |