LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
Step 3.5 Flash (visit website) is our most capable open-source foundation model, engineered to deliver frontier reasoning and agentic capabilities with exceptional efficiency. We also open-sourced the training codebase (SteptronOss), with support for continue pretrain, SFT, RL (W…
Authors: Erik Nijkamp\ , Hiroaki Hayashi\ , Yingbo Zhou, Caiming Xiong
Medical-Qwen3-Swallow-30B-A3B is a medical-domain language model based on tokyotech-llm/Qwen3-Swallow-30B-A3B-RL-v0.2 . It is designed to support research and development toward safe and trustworthy AI for Japanese clinical settings.
Developed by : Upstage Backbone Model : LLaMA-2 Language(s) : English Library : HuggingFace Transformers License : Fine-tuned checkpoints is licensed under the Non-Commercial Creative Commons license (CC BY-NC-4.0) Where to send comments : Instructions on how to provide feedback…
🤗 Hugging Face      🤖 ModelScope 🐙 Experience Now
💻Github Repo • 🤗Model Collections • 📜Technical Report • 💬Online Chat
💻Github Repo • 🤔Reporting Issues • 📜Technical Report
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ Unlocking the Reasoning Potential of Language Model From Pretraining to Posttraining ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
Penguin-VL Exploring the Efficiency Limits of VLM with LLM-based Vision Encoders
This is a Japanese finetuned model based on deepseek-ai/DeepSeek-R1-Distill-Qwen-32B.
CoDA: Coding LM via Diffusion Adaptation
ArmorOCR is a two-stage framework for grounded adversarial OCR perception built on Qwen3-VL-8B-Instruct. It enables single-pass inference on the original image, without any inference-time visual transformations or tool assistance.