LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
NOSA: Native and Offloadable Sparse Attention
💻Github Repo • 🤗Model Collections • 📜Technical Report • 💬Online Chat
💻Github Repo • 🤔Reporting Issues • 📜Technical Report
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ Unlocking the Reasoning Potential of Language Model From Pretraining to Posttraining ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
Llama 3.1 Swallow is a series of large language models (8B, 70B) that were built by continual pre-training on the Meta Llama 3.1 models. Llama 3.1 Swallow enhanced the Japanese language capabilities of the original Llama 3.1 while retaining the English language capabilities. We u…
UI-TARS-72B-DPO UI-TARS-2B-SFT UI-TARS-7B-SFT UI-TARS-7B-DPO (Recommended) UI-TARS-72B-SFT UI-TARS-72B-DPO (Recommended) Introduction
Penguin-VL Exploring the Efficiency Limits of VLM with LLM-based Vision Encoders
This is a Japanese finetuned model based on deepseek-ai/DeepSeek-R1-Distill-Qwen-32B.
CoDA: Coding LM via Diffusion Adaptation
ArmorOCR is a two-stage framework for grounded adversarial OCR perception built on Qwen3-VL-8B-Instruct. It enables single-pass inference on the original image, without any inference-time visual transformations or tool assistance.
Llama-SEA-LION-v2-8B SEA-LION is a collection of Large Language Models (LLMs) which have been pretrained and instruct-tuned for the Southeast Asia (SEA) region.