AI Agent Hub

Compare LLM Models

Select 2-5 models to compare side by side.

Select models to compare 0 / 5 selected
Select 2-5 models
Attribute
llmfan46
Company llmfan46
Release Date 2026-04-03
Parameters (B) 31
Architecture Transformer (Gemma 4 hybrid attention)
Context Window 256000
Input Modalities text, image
Open Source Yes
License
Score 8.4
VRAM (GB) 24
Compute Minimum 24GB VRAM GPU (RTX 3090/4090) for Q4_K_M GGUF via llama.cpp or Ollama; 32GB+ recommended for Q8_0 or 8k+ context; BF16/full precision requires ~62GB VRAM. CPU offloading possible but significantly slower.
Benchmark: GSM8K 84.5
Benchmark: HumanEval 84.5
Benchmark: MMLU 85.9
Pricing (Input)
Pricing (Output)