LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…
This is the base model for SmolVLM2-2.2B, a lightweight multimodal model designed to analyze video content. The model processes videos, images, and text inputs to generate text outputs - whether answering questions about media files, comparing visual content, or transcribing text…
💻Github Repo • 🤗Model Collections • 📜Technical Report
0. TL;DR 1. Model Details 2. Training Details 3. Usage 4. Evaluation 5. Citation
[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.
CogVLM 是一个强大的开源视觉语言模型(VLM)。CogVLM-17B 拥有 100 亿视觉参数和 70 亿语言参数,在 10 个经典跨模态基准测试上取得了 SOTA 性能,包括 NoCaps、Flicker30k captioning、RefCOCO、RefCOCO+、RefCOCOg、Visual7W、GQA、ScienceQA、VizWiz VQA 和 TDIUC,而在 VQAv2、OKVQA、TextVQA、COCO captioning 等方面则排名第二,超越或与 PaLI-X 55B 持平。您可以通过线上 demo 体验 CogVLM 多…
Hy-Embodied-VLM-1.0 Efficient Physical-World Agents Tencent Robotics X × Hy Vision Team × Futian Laboratory
CogAgent is an open-source visual language model improved based on CogVLM .
HY-Embodied-0.5-X An Enhanced Embodied Foundation Model for Real-World Agents Tencent Robotics X × HY Vision Team
UI-TARS-72B-DPO UI-TARS-2B-SFT UI-TARS-7B-SFT UI-TARS-7B-DPO (Recommended) UI-TARS-72B-SFT UI-TARS-72B-DPO (Recommended) Introduction
  GITHUB      🖥️   official website   |  🕖   HunyuanAPI |  🐳   Gitee Technical Report   |   Demo    |   Tencent Cloud TI    
[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…