LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
A 15B-parameter token-mixer supernet derived from Apriel-1.6 via stochastic distillation. Every decoder layer exposes four trained mixer options —Full Attention, Sliding Window Attention, Gated DeltaNet, and Kimi Delta Attention—enabling flexible architecture selection from a sin…
BFS-Prover-V2: Scaling up Multi-Turn Off-Policy RL and Multi-Agent Tree Search for LLM Step-Provers
AHN: Artificial Hippocampus Networks for Efficient Long-Context Modeling
[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.
Our Swallow model has undergone continual pre-training from the Llama 2 family, primarily with the addition of Japanese language data. The tuned versions use supervised fine-tuning (SFT). Links to other models can be found in the index.
AHN: Artificial Hippocampus Networks for Efficient Long-Context Modeling
Our Swallow model has undergone continual pre-training from the Llama 2 family, primarily with the addition of Japanese language data. The tuned versions use supervised fine-tuning (SFT). Links to other models can be found in the index.
[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.
🦉GitHub 💬WeChat 百川API支持搜索增强和192K长窗口,新增百川搜索增强知识库、限时免费! 🚀 百川大模型在线对话平台 已正式向公众开放 🎉
Our Swallow model has undergone continual pre-training from the Llama 2 family, primarily with the addition of Japanese language data. The tuned versions use supervised fine-tuning (SFT). Links to other models can be found in the index.
AHN: Artificial Hippocampus Networks for Efficient Long-Context Modeling
Our Swallow model has undergone continual pre-training from the Llama 2 family, primarily with the addition of Japanese language data. The tuned versions use supervised fine-tuning (SFT). Links to other models can be found in the index.