LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
We expand on our Olmo model series by introducing Olmo Hybrid, a new 7B hybrid RNN model in the Olmo family. Olmo Hybrid dramatically outperforms Olmo 3 in final performance, consistently showing roughly 2x data efficiency on core evals over the course of our pretraining run. We…
Upon the initial release of OLMo-2 models, we realized the post-trained models did not share the pre-tokenization logic that the base models use. As a result, we have trained new post-trained models. The new models are available under the same names as the original models, but we…
[🏠Homepage] [🤖 Chat with DeepSeek Coder] [Discord] [Wechat(微信)]
This is the model card of a 🤗 transformers model that has been pushed on the Hub. This model card has been automatically generated.
Intern-Decision-4B is a multimodal structured decision model fine-tuned from Qwen3.5-4B . It accepts a shared state, a schema of named questions, and optional images, and returns an answer distribution for every question in one model forward pass.
This is the model card of a 🤗 transformers model that has been pushed on the Hub. This model card has been automatically generated.
LLaMA-2-7B-32K is an open-source, long context language model developed by Together, fine-tuned from Meta's original Llama-2 7B model. This model represents our efforts to contribute to the rapid progress of the open-source ecosystem for large language models. The model has been…
GitHub Repo Technical Report 👋 Join us on Discord and WeChat
GitHub Repo Technical Report 👋 Join us on Discord and WeChat
GitHub Repo Technical Report 👋 Join us on Discord and WeChat
GitHub Repo Technical Report 👋 Join us on Discord and WeChat
Llama 2 Llama 2 is a collection of pretrained and fine-tuned generative text models ranging in scale from 7 billion to 70 billion parameters. This is the repository for the 13B fine-tuned model, optimized for dialogue use cases and converted for the Hugging Face Transformers form…