AI Agent Hub
Back to models
🤖

PrismML: Ternary Bonsai 2 27B

Open Source prism-ml Released 2026-09-17
-- 27.36B params 262.1K context Proprietary

About this model

PrismML Ternary Bonsai 2 27B is an Apache 2.0 open-weight multimodal language model derived from Qwen3.8-27B. Its language-model matrices are stored as native ternary weights in {-1, 0, +1} with FP16 group-wise scaling and blockwise Hadamard rotation, yielding about 1.76 effective bits per weight and roughly 5.95 GB on disk for the GGUF PTQ1_0 language weights—about nine times smaller than the FP16 baseline while PrismML reports roughly 98.2% retention on its primary thinking-mode benchmark suites.

The architecture keeps the Qwen3.8 hybrid-attention stack (~75% linear / ~25% full attention), SwiGLU MLP, RoPE, and RMSNorm, with a 262,144-token context window and optional vision input via a separate mmproj bundle. Weights ship as GGUF (requiring the PrismML llama.cpp fork for ternary kernels) and MLX 2-bit packs for Apple Silicon and CUDA. Thinking mode is the default evaluation and deployment posture; published scores use EvalScope with vLLM on NVIDIA H100 in matched thinking-mode settings.

Reported strengths cluster in mathematics, coding, and instruction following (for example GSM8K, MATH-500, AIME 2025/2026, LiveCodeBench, HumanEval+, and IFEval), with GPQA Diamond and multimodal tasks somewhat below the FP16 parent. Long-horizon agentic benchmarks such as SWE-Bench Verified and Terminal-Bench 2.1 show larger gaps, which PrismML attributes to compounding errors in extended tool-use loops rather than single-shot QA.

Benchmark Scores

MBPP
83.07
GSM8K
96.66
IFEval
91.31
MATH-500
98.8
MMMU-Pro
75.49
AIME-2025
95.0
AIME-2026
95.83
HumanEval
95.12
GPQA-Diamond
85.76
LiveCodeBench
90.07
SWE-Bench-Verified
60.8
Terminal-Bench-2.1
52.8

Technical Specs

  • Parameters: 27.36B
  • Architecture: Hybrid-attention Transformer
  • Context Window: 262,144 tokens
  • Input Modalities: text, image

Hardware Requirements

  • VRAM: 8.0 GB
  • Compute: Single GPU with 8GB+ VRAM (e.g. NVIDIA RTX 4090) or Apple Silicon via MLX

Pricing

Input Output Currency
0.07 / 1M tokens 0.50 / 1M tokens USD