AI Agent Hub
Back to models
tiny-aya-base-32K logo

tiny-aya-base-32K

Open Source CohereLabs Released 2026-09-08
-- 3.35B params 32K context Open Source

About this model

Tiny Aya Base 32K is an open-weights pretrained multilingual language model from Cohere Labs, released on Hugging Face as part of the Tiny Aya family. At 3.35 billion parameters it targets efficient deployment while covering more than 70 languages, with emphasis on balanced representation across high- and lower-resource languages. The checkpoint extends the original 8K Tiny Aya Base context to 32K tokens combined across input and output, making it suitable for continued pretraining, instruction tuning, and long-context research rather than turnkey chat use.

Architecturally the model follows the Tiny Aya Cohere2 design: a decoder-only transformer with sliding-window local attention and a global attention layer, implemented as Cohere2ForCausalLM in the Hugging Face ecosystem. It is not instruction-tuned or preference-aligned, so behavior is completion-style rather than assistant-style. Cohere Labs documents it as the base used to build Tiny Aya L2-Thinker and Tiny Aya En-Thinker, which inherit the same 32K window.

The weights are gated under CC-BY-NC 4.0 and Cohere Labs Acceptable Use Policy. The public model card does not publish standard English-centric leaderboard numbers; the Tiny Aya technical report benchmarks instruction-tuned variants (e.g., Tiny Aya Global) on multilingual suites such as Global MMLU, mDolly, and Flores rather than this base 32K checkpoint. Users should expect strong open-ended multilingual generation potential after fine-tuning, with weaker out-of-the-box chain-of-thought or instruction following compared to tuned siblings.

Technical Specs

  • Parameters: 3.35B
  • Architecture: Cohere2 Transformer
  • Context Window: 32,000 tokens
  • Input Modalities: text

Hardware Requirements

  • VRAM: 16.0 GB
  • Compute: Single NVIDIA GPU with 16GB+ VRAM (8GB+ for weights-only; 32K context needs extra KV cache)