LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
Sachin Mehta, Mohammad Hossein Sekhavat, Qingqing Cao, Maxwell Horton, Yanzi Jin, Chenfan Sun, Iman Mirzadeh, Mahyar Najibi, Dmitry Belenko, Peter Zatloukal, Mohammad Rastegari
Towards a Recursively Self-Improving Agent for Deep Research
InternLM has open-sourced a 7 billion parameter base model tailored for practical scenarios. The model has the following characteristics: - It leverages trillions of high-quality tokens for training to establish a powerful knowledge base. - It provides a versatile toolset for use…
MiniCPM is an End-Size LLM developed by ModelBest Inc. and TsinghuaNLP, with only 2.4B parameters excluding embeddings. MiniCPM-2B-128k is a long context extension trial of MiniCPM-2B. To our best knowledge, MiniCPM-2B-128k is the first long context( =128k) SLM smaller than 3B。 I…
[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…
CapRL 📖 Paper 🏠 Github 🤗 CapRL Collection 🤗 Daily Paper
🤗 [LongWriter Dataset] • 💻 [Github Repo] • 📃 [LongWriter Paper]
0. TL;DR 1. Model Details 2. Training Details 3. Usage 4. Evaluation 5. Citation
0. TL;DR 1. Model Details 2. Training Details 3. Usage 4. Evaluation 5. Citation
Weights for a sparse model from Gao et al. 2025, used for the qualitative results from the paper (related to bracket counting and variable binding). All weights for the other models used in the paper, as well as lightweight inference code, are present in https://github.com/openai…
1. Model Summary 2. Use 3. Limitations 4. Training 5. Evaluation 7. Citation