LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
Mage-VL An Efficient Codec-Native Streaming Multimodal Foundation Model
InternVL3-1B Transformers 🤗 Implementation
[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…
1. Model Summary 2. How to use 3. Evaluation 4. Training 5. Limitations 6. License
Meta developed and released the Meta Llama 3 family of large language models (LLMs), a collection of pretrained and instruction tuned generative text models in 8 and 70B sizes. The Llama 3 instruction tuned models are optimized for dialogue use cases and outperform many of the av…
[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…
Description: The NVIDIA DeepSeek-R1-0528-FP4 v2 model is the quantized version of the DeepSeek AI's DeepSeek R1 0528 model, which is an auto-regressive language model that uses an optimized transformer architecture. For more information, please check here. The NVIDIA DeepSeek R1…
💻Github Repo • 🤔Reporting Issues • 📜Technical Report
LFM2.5 is a family of hybrid models designed for on-device deployment . It builds on the LFM2 architecture with extended pre-training and reinforcement learning.
InternVL3-8B Transformers 🤗 Implementation
🤗 Hugging Face 🖥️ Official Website 🕖 HunyuanAPI 🕹️ Demo 🤖 ModelScope