LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
- 🔜 Upcoming: The End-to-End GUI Agent Model is currently under active development and will be released in a subsequent update. Stay tuned! - 🚀 2026.02.06: We are happy to present POINTS-GUI-G , our specialized GUI Grounding Model. To facilitate reproducible evaluation, we provid…
Penguin-VL Exploring the Efficiency Limits of VLM with LLM-based Vision Encoders
UI-Venus This repository contains the UI-Venus model from the report UI-Venus Technical Report: Building High-performance UI Agents with RFT.
💻Github Repo • 🤗HF Model Collections • 🤖ModelScope Collections • 💬Online Chat
ZwZ-4B is a fine-grained multimodal perception model built upon Qwen3-VL-4B. It is trained using Region-to-Image Distillation (R2I) combined with reinforcement learning, enabling superior fine-grained visual understanding in a single forward pass — no inference-time zooming or to…
Llama-SEA-LION-v2-8B-IT SEA-LION is a collection of Large Language Models (LLMs) which have been pretrained and instruct-tuned for the Southeast Asia (SEA) region.
Penguin-VL Exploring the Efficiency Limits of VLM with LLM-based Vision Encoders
📖 Paper 🏠 Github 🤗 Spatial-SSRL-7B Model 🤗 Spatial-SSRL-3B Model 🤗 Spatial-SSRL-Qwen3VL-4B Model 🤗 Spatial-SSRL-81k Dataset 📰 Daily Paper
PaCoRe: Learning to Scale Test-Time Compute with Parallel Coordinated Reasoning
PaCoRe: Learning to Scale Test-Time Compute with Parallel Coordinated Reasoning
[\[📂 GitHub\]](https://github.com/bytedance/SAIL) [\[📜 paper\]](https://arxiv.org/abs/2504.10462) [\[🚀 Quick Start\]]( quick-start)
Falcon2-11B-vlm is an 11B parameters causal decoder-only model built by TII and trained on over 5,000B tokens of RefinedWeb enhanced with curated corpora. To bring vision capabilities, we integrate the pretrained CLIP ViT-L/14 vision encoder with our Falcon2-11B chat-finetuned mo…