LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
One of the focus areas at Together Research is new architectures for long context, improved training, and inference performance over the Transformer architecture. Spinning out of a research program from our team and academic collaborators, with roots in signal processing-inspired…
💻Github Repo • 🤗Model Collections • 📜Technical Report
💻Github Repo • 🤗Model Collections • 📜Technical Report • 💬Online Chat
Qwen-SEA-LION-v4-32B-IT-4BIT (GPTQ model)
"A parrot able to speak Japanese, ukiyoe, edo period" — Stable Diffusion XL
AquilaMoE: Efficient Training for MoE Models with Scale-Up and Scale-Out Strategies Language Foundation Model & Software Team Beijing Academy of Artificial Intelligence (BAAI) [Paper(released soon)] [Code] [github]
Medical-GPT-OSS-Swallow-120B is a medical-domain language model based on tokyotech-llm/GPT-OSS-Swallow-120B-RL-v0.1. It is designed to support research and development toward safe and trustworthy AI for Japanese clinical settings.
Developed by : Upstage Backbone Model : LLaMA Variations : It has different model parameter sizes and sequence lengths: 30B/1024, 30B/2048, 65B/1024 Language(s) : English Library : HuggingFace Transformers License : This model is under a Non-commercial Bespoke License and governe…
💻Github Repo • 🤔Reporting Issues • 📜Technical Report
Ring-lite is a lightweight, fully open-sourced MoE (Mixture of Experts) LLM designed for complex reasoning tasks. It is built upon the publicly available Ling-lite-1.5 model, which has 16.8B parameters with 2.75B activated parameters.. We use a joint training pipeline combining k…
Step 3.5 Flash (visit website) is our most capable open-source foundation model, engineered to deliver frontier reasoning and agentic capabilities with exceptional efficiency. We also open-sourced the training codebase (SteptronOss), with support for continue pretrain, SFT, RL (W…
This is the base model for SmolVLM2-2.2B, a lightweight multimodal model designed to analyze video content. The model processes videos, images, and text inputs to generate text outputs - whether answering questions about media files, comparing visual content, or transcribing text…