Cohere: Command A+
About this model
Command A+ (command-a-plus-05-2026) is Cohere’s Apache 2.0 open-weight flagship in the Command family, released May 20, 2026. It is a decoder-only sparse mixture-of-experts transformer with 218 billion total parameters and 25 billion active parameters per token (128 experts, eight routed experts plus one shared expert per token). The model unifies agentic tool use, multilingual text generation, and vision understanding in one stack, with 128K input context and up to 64K output tokens, native reasoning traces, and conversational function calling across 48 languages.
Cohere positions Command A+ for sovereign enterprise deployment: it ships on Hugging Face in BF16, FP8, and W4A4 quantizations with near-lossless quality, runs on as little as one NVIDIA B200 or two H100 GPUs at W4A4, and integrates with Transformers and vLLM. Reported strengths include multimodal document reasoning (MMMU/MMMU-Pro), math (AIME 2025), graduate science QA (GPQA Diamond), coding (LiveCodeBench, SWE-Bench), and agentic terminal workflows, with large gains over Command A Reasoning on telecom agent benchmarks and Terminal-Bench Hard.
The architecture extends Command A with interleaved sliding-window and global attention (3:1), dropless token-choice MoE routing, additive-bias load balancing, and a normalized sigmoid router over top-k experts. It is available via Cohere’s API and playground, Model Vault, and self-hosted inference for regulated and on-prem workloads.
Benchmark Scores
Technical Specs
- Parameters: 218.0B
- Architecture: Sparse Mixture-of-Experts Transformer
- Context Window: 128,000 tokens
- Input Modalities: text, image
Hardware Requirements
- VRAM: 160.0 GB
- Compute: 2x NVIDIA H100 80GB (W4A4 recommended)
Pricing
| Input | Output | Currency |
|---|---|---|
| 0.30 / 1M tokens | 1.50 / 1M tokens | USD |