LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
A 15B-parameter token-mixer supernet derived from Apriel-1.6 via stochastic distillation. Every decoder layer exposes four trained mixer options —Full Attention, Sliding Window Attention, Gated DeltaNet, and Kimi Delta Attention—enabling flexible architecture selection from a sin…
Our Swallow model has undergone continual pre-training from the Llama 2 family, primarily with the addition of Japanese language data. The tuned versions use supervised fine-tuning (SFT). Links to other models can be found in the index.
Llama 3.1 Swallow is a series of large language models (8B, 70B) that were built by continual pre-training on the Meta Llama 3.1 models. Llama 3.1 Swallow enhanced the Japanese language capabilities of the original Llama 3.1 while retaining the English language capabilities. We u…
BFS-Prover-V2: Scaling up Multi-Turn Off-Policy RL and Multi-Agent Tree Search for LLM Step-Provers
Overview The model is the instruction-tuned version of rinna/youri-7b . It adopts a chat-style input format.
This is a Japanese finetuned model based on deepseek-ai/DeepSeek-R1-Distill-Qwen-14B.
Authors : Yizhe Zhang, Navdeep Jaitly (Apple)
Our Swallow model has undergone continual pre-training from the Llama 2 family, primarily with the addition of Japanese language data. The tuned versions use supervised fine-tuning (SFT). Links to other models can be found in the index.
Model Summary DiaFill is a Japanese dialogue script generation model designed to produce natural, spoken-style dialogue scripts rich in fillers and brief utternaces. Unlike typical assistant models that respond to users, this model is fine-tuned to generate a multi-turn dialogue…
Overview The model is the instruction-tuned version of rinna/nekomata-7b . It adopts the Alpaca input format.
A 15B-parameter hybrid reasoning model combining Transformer attention and Mamba State Space layers for high efficiency and scalability. Derived from Apriel-Nemotron-15B-Thinker through progressive distillation, Apriel-H1 replaces less critical attention layers with linear Mamba…
AHN: Artificial Hippocampus Networks for Efficient Long-Context Modeling