LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
Beijing Academy of Artificial Intelligence (BAAI) [Paper][Code][🤗] (would be released soon)
👋 Wechat · 💡 Online Demo · 🎈 Github Page · 📑 Paper 📍Experience the larger-scale CogVLM model on the ZhipuAI Open Platform .
Beijing Academy of Artificial Intelligence (BAAI) [Paper][Code][🤗] (would be released soon)
Tulu is a series of language models that are trained to act as helpful assistants. Tulu 2 7B is a fine-tuned version of Llama 2 that was trained on a mix of publicly available, synthetic and human datasets.
⚠️ DEPRECATION WARNING ⚠️ ⚠️ NOT RECOMMENDED FOR USE IN NEW PROJECTS ⚠️
Play with the model on the SantaCoder Space Demo.
Upon the initial release of OLMo-2 models, we realized the post-trained models did not share the pre-tokenization logic that the base models use. As a result, we have trained new post-trained models. The new models are available under the same names as the original models, but we…
Model Summary: Granite-3.2-2B-Instruct is an 2-billion-parameter, long-context AI model fine-tuned for thinking capabilities. Built on top of Granite-3.1-2B-Instruct, it has been trained using a mix of permissively licensed open-source datasets and internally generated synthetic…
We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1. DeepSeek-R1-Zero, a model trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary step, demonstrated remarkable performance on reasoning. With R…
SmolLM is a series of small language models available in three sizes: 135M, 360M, and 1.7B parameters.
[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 Mini-InternVL\]](https://arxiv.org/abs/2410.16261) [\[📜 InternVL 2.5\]](https://huggingface.co/…