LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
An alternate Meta-Llama-3-8B Repo for the Hermes Tokenizer
--- license: apache-2.0 library name: transformers base model: - Qwen/Qwen2.5-7B pipeline tag: text-generation ---
AquilaMoE: Efficient Training for MoE Models with Scale-Up and Scale-Out Strategies Language Foundation Model & Software Team Beijing Academy of Artificial Intelligence (BAAI) [Paper(released soon)] [Code] [github]
WARNING: The checkpoints on this repo are not fully trained model. Evaluations of intermediary checkpoints and the final model will be added when conducted (see below).
AHN: Artificial Hippocampus Networks for Efficient Long-Context Modeling
Our Swallow model has undergone continual pre-training from the Llama 2 family, primarily with the addition of Japanese language data. The tuned versions use supervised fine-tuning (SFT). Links to other models can be found in the index.
Model Summary Chat with the model at: https://huggingface.co/spaces/HuggingFaceTB/instant-smol
AHN: Artificial Hippocampus Networks for Efficient Long-Context Modeling
[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.
[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.
AHN: Artificial Hippocampus Networks for Efficient Long-Context Modeling
🦉GitHub 💬WeChat 百川API支持搜索增强和192K长窗口,新增百川搜索增强知识库、限时免费! 🚀 百川大模型在线对话平台 已正式向公众开放 🎉