LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
This is the base version of the Jamba model. We’ve since released a better, instruct-tuned version, Jamba-1.5-Mini. For even greater performance, check out the scaled-up Jamba-1.5-Large.
2024/08/12, 本仓库代码已更新并使用 transformers =4.44.0 , 请及时更新依赖。
InternVL3-2B Transformers 🤗 Implementation
LFM2.5 is a new family of hybrid models designed for on-device deployment . It builds on the LFM2 architecture with extended pre-training and reinforcement learning.
OLMo 2 1B DPO April 2025 is post-trained variant of the allenai/OLMo-2-0425-1B-SFT model, which has undergone supervised finetuning on an OLMo-specific variant of the Tülu 3 dataset and further DPO training on this dataset. Tülu 3 is designed for state-of-the-art performance on a…
🤗 Hugging Face      🤖 ModelScope      🐙 Experience Now
Meet 10.7B Solar: Elevating Performance with Upstage Depth UP Scaling!
CodeGen is a family of autoregressive language models for program synthesis from the paper: A Conversational Paradigm for Program Synthesis by Erik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu, Huan Wang, Yingbo Zhou, Silvio Savarese, Caiming Xiong. The models are originally releas…
BLOOM LM BigScience Large Open-science Open-access Multilingual Language Model Model Card
📃 License • 💻 Code • 📑 Technical Report • 📊 Benchmarks • 🚀 Getting Started • 💡 Highlights
👋 Join our Discord community. 📖 Check out the GLM-4.5 technical blog , technical report , and Zhipu AI technical documentation . 📍 Use GLM-4.5 API services on Zhipu AI Open Platform . 👉 One click to GLM-4.5 .