LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
🐙 GitHub • 👾 Discord • 🐤 Twitter • 💬 WeChat 📝 Paper • 💪 Tech Blog • 🙌 FAQ • 📗 Learning Hub
Falcon-40B is a 40B parameters causal decoder-only model built by TII and trained on 1,000B tokens of RefinedWeb enhanced with curated corpora. It is made available under the Apache 2.0 license.
Model card for FalconMamba Instruct model
InternVL3-14B Transformers 🤗 Implementation
CodeGeeX4: Open Multilingual Code Generation Model
[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 Mini-InternVL\]](https://arxiv.org/abs/2410.16261) [\[📜 InternVL 2.5\]](https://huggingface.co/…
This is the base version of the Jamba model. We’ve since released a better, instruct-tuned version, Jamba-1.5-Mini. For even greater performance, check out the scaled-up Jamba-1.5-Large.
2024/08/12, 本仓库代码已更新并使用 transformers =4.44.0 , 请及时更新依赖。
InternVL3-2B Transformers 🤗 Implementation
LFM2.5 is a new family of hybrid models designed for on-device deployment . It builds on the LFM2 architecture with extended pre-training and reinforcement learning.
OLMo 2 1B DPO April 2025 is post-trained variant of the allenai/OLMo-2-0425-1B-SFT model, which has undergone supervised finetuning on an OLMo-specific variant of the Tülu 3 dataset and further DPO training on this dataset. Tülu 3 is designed for state-of-the-art performance on a…