LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
A cute robot wearing a kimono writes calligraphy with one single brush — Stable Diffusion XL
Qwen-SEA-LION-v4-32B-IT-4BIT (GPTQ model)
"A parrot able to speak Japanese, ukiyoe, edo period" — Stable Diffusion XL
AquilaMoE: Efficient Training for MoE Models with Scale-Up and Scale-Out Strategies Language Foundation Model & Software Team Beijing Academy of Artificial Intelligence (BAAI) [Paper(released soon)] [Code] [github]
Medical-GPT-OSS-Swallow-120B is a medical-domain language model based on tokyotech-llm/GPT-OSS-Swallow-120B-RL-v0.1. It is designed to support research and development toward safe and trustworthy AI for Japanese clinical settings.
Developed by : Upstage Backbone Model : LLaMA Variations : It has different model parameter sizes and sequence lengths: 30B/1024, 30B/2048, 65B/1024 Language(s) : English Library : HuggingFace Transformers License : This model is under a Non-commercial Bespoke License and governe…
💻Github Repo • 🤔Reporting Issues • 📜Technical Report
1. Model Summary 2. Use 3. Training 4. License 5. Citation
Ring-lite is a lightweight, fully open-sourced MoE (Mixture of Experts) LLM designed for complex reasoning tasks. It is built upon the publicly available Ling-lite-1.5 model, which has 16.8B parameters with 2.75B activated parameters.. We use a joint training pipeline combining k…
This model is an example of the Simple Self-Distillation (SimpleSD) method that improves code generation by fine-tuning a language model on its own sampled outputs—without rewards, verifiers, teacher models, or reinforcement learning. Please see the paper below for more informati…
AHN: Artificial Hippocampus Networks for Efficient Long-Context Modeling
Step 3.5 Flash (visit website) is our most capable open-source foundation model, engineered to deliver frontier reasoning and agentic capabilities with exceptional efficiency. We also open-sourced the training codebase (SteptronOss), with support for continue pretrain, SFT, RL (W…