LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.
Qianfan-VL: Domain-Enhanced Universal Vision-Language Models
UI-TARS-72B-SFT UI-TARS-2B-SFT UI-TARS-7B-SFT UI-TARS-7B-DPO (Recommended) UI-TARS-72B-SFT UI-TARS-72B-DPO (Recommended) Introduction
Qwen-SEA-LION-v4-32B-IT-4BIT (GPTQ model)
💻Github Repo • 🤗Model Collections • 📜Technical Report • 💬Online Chat
Reinforcement learning (RL) (e.g., GRPO) helps with grounding because of its inherent objective alignment—rewarding successful clicks—rather than encouraging long textual Chain-of-Thought (CoT) reasoning. Unlike approaches that rely heavily on verbose CoT reasoning, GRPO directly…
SmolVLM is a compact open multimodal model that accepts arbitrary sequences of image and text inputs to produce text outputs. Designed for efficiency, SmolVLM can answer questions about images, describe visual content, create stories grounded on multiple images, or function as a…
Qianfan-VL: Domain-Enhanced Universal Vision-Language Models
[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.
ZwZ-8B is a fine-grained multimodal perception model built upon Qwen3-VL-8B. It is trained using Region-to-Image Distillation (R2I) combined with reinforcement learning, enabling superior fine-grained visual understanding in a single forward pass — no inference-time zooming or to…
Note: Users are permitted to use this model in accordance with the Llama 3.1 Community License Agreement.
This repository provides Japanese language models trained by SB Intuitions.