LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
👋 Hi, everyone! We are ByteDance Seed Team.
Sachin Mehta, Mohammad Hossein Sekhavat, Qingqing Cao, Maxwell Horton, Yanzi Jin, Chenfan Sun, Iman Mirzadeh, Mahyar Najibi, Dmitry Belenko, Peter Zatloukal, Mohammad Rastegari
InternLM has open-sourced a 7 billion parameter base model tailored for practical scenarios. The model has the following characteristics: - It leverages trillions of high-quality tokens for training to establish a powerful knowledge base. - It provides a versatile toolset for use…
MiniCPM is an End-Size LLM developed by ModelBest Inc. and TsinghuaNLP, with only 2.4B parameters excluding embeddings. MiniCPM-2B-128k is a long context extension trial of MiniCPM-2B. To our best knowledge, MiniCPM-2B-128k is the first long context( =128k) SLM smaller than 3B。 I…
[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…
🤗 [LongWriter Dataset] • 💻 [Github Repo] • 📃 [LongWriter Paper]
0. TL;DR 1. Model Details 2. Training Details 3. Usage 4. Evaluation 5. Citation
0. TL;DR 1. Model Details 2. Training Details 3. Usage 4. Evaluation 5. Citation
Weights for a sparse model from Gao et al. 2025, used for the qualitative results from the paper (related to bracket counting and variable binding). All weights for the other models used in the paper, as well as lightweight inference code, are present in https://github.com/openai…
1. Model Summary 2. Use 3. Limitations 4. Training 5. Evaluation 7. Citation
RedPajama-INCITE-7B-Chat was developed by Together and leaders from the open-source AI community including Ontocord.ai, ETH DS3Lab, AAI CERC, Université de Montréal, MILA - Québec AI Institute, Stanford Center for Research on Foundation Models (CRFM), Stanford Hazy Research resea…
0. TL;DR 1. Model Details 2. Training Details 3. Usage 4. Evaluation 5. Citation