LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
1. Model Summary 2. Evaluation 3. Intended Use 4. Limitations 5. Security and Responsible Use 6. License 7. Citation
Penguin-VL Exploring the Efficiency Limits of VLM with LLM-based Vision Encoders
SEA-LION is a collection of Large Language Models (LLMs) which has been pretrained and instruct-tuned for the Southeast Asia (SEA) region. The size of the models range from 3 billion to 7 billion parameters. This is the card for the SEA-LION 7B base model.
This model is an example of the Simple Self-Distillation (SimpleSD) method that improves code generation by fine-tuning a language model on its own sampled outputs—without rewards, verifiers, teacher models, or reinforcement learning. Please see the paper below for more informati…
ZwZ-4B is a fine-grained multimodal perception model built upon Qwen3-VL-4B. It is trained using Region-to-Image Distillation (R2I) combined with reinforcement learning, enabling superior fine-grained visual understanding in a single forward pass — no inference-time zooming or to…
Nous-Yarn-Llama-2-13b-64k is a state-of-the-art language model for long context, further pretrained on long context data for 400 steps. This model is the Flash Attention 2 patched version of the original model: https://huggingface.co/conceptofmind/Yarn-Llama-2-13b-64k
Llama 3.1 Swallow is a series of large language models (8B, 70B) that were built by continual pre-training on the Meta Llama 3.1 models. Llama 3.1 Swallow enhanced the Japanese language capabilities of the original Llama 3.1 while retaining the English language capabilities. We u…
--- library name: transformers tags: - bitnet - falcon3 base model: tiiuae/Falcon3-10B-Base license: other license name: falcon-llm-license license link: https://falconllm.tii.ae/falcon-terms-and-conditions.html ---
AgentCPM-Report: Gemini-2.5-pro-DeepResearch Level Local DeepResearch
Penguin-VL Exploring the Efficiency Limits of VLM with LLM-based Vision Encoders
📖 Paper 🏠 Github 🤗 Spatial-SSRL-7B Model 🤗 Spatial-SSRL-3B Model 🤗 Spatial-SSRL-Qwen3VL-4B Model 🤗 Spatial-SSRL-81k Dataset 📰 Daily Paper
━━━━━━━━━━━━━━━━━━━━━━━━━ Unlocking the Reasoning Potential of Language Model From Pretraining to Posttraining ━━━━━━━━━━━━━━━━━━━━━━━━━