LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
- 🔜 Upcoming: The End-to-End GUI Agent Model is currently under active development and will be released in a subsequent update. Stay tuned! - 🚀 2026.02.06: We are happy to present POINTS-GUI-G , our specialized GUI Grounding Model. To facilitate reproducible evaluation, we provid…
0. TL;DR 1. Model Details 2. Training Details 3. Usage 4. Evaluation 5. Citation
We opensource our Aquila2 series, now including Aquila2 , the base language models, namely Aquila2-7B and Aquila2-34B , as well as AquilaChat2 , the chat models, namely AquilaChat2-7B and AquilaChat2-34B , as well as the long-text chat models, namely AquilaChat2-7B-16k and Aquila…
This model is an example of the Simple Self-Distillation (SimpleSD) method that improves code generation by fine-tuning a language model on its own sampled outputs—without rewards, verifiers, teacher models, or reinforcement learning. Please see the paper below for more informati…
This repository is primarily for the Transformers framework. If you're using other open-source frameworks, please use the alternative repository: MiniMax-M1-80k
[2025-06-03] 📄📄📄 We have released the technical report of AgentCPM-GUI! Check it out here. [2025-05-13] 🚀🚀🚀 We have open-sourced AgentCPM-GUI , an on-device GUI agent capable of operating Chinese & English apps and equipped with RFT-enhanced reasoning abilities.
This is a 7B-parameter decoder-only language model with a focus on maximizing Japanese language modeling performance and Japanese downstream task performance. We conducted continued pretraining using Japanese data on the English language model, Mistral-7B-v0.1, to transfer the mo…
SEA-LION is a collection of Large Language Models (LLMs) which have been pretrained and instruct-tuned for the Southeast Asia (SEA) region.
Official research release for the family of XGen models ( 7B ) by Salesforce AI Research:
Introduction We are thrilled to introduce Stable-DiffCoder, which is a strong code diffusion large language model. Built directly on the Seed-Coder architecture, data, and training pipeline, it introduces a block diffusion continual pretraining (CPT) stage with a tailored warmup…
This is an EAGLE3 draft model trained from scratch (random initialization) using the Aurora inference-time training framework for speculative decoding. Unlike traditional approaches that fine-tune pre-trained models, this model is built entirely through Aurora's online training p…
SEA-LION is a collection of Large Language Models (LLMs) which have been pretrained and instruct-tuned for the Southeast Asia (SEA) region.