LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
RedPajama-INCITE-Chat-3B-v1 was developed by Together and leaders from the open-source AI community including Ontocord.ai, ETH DS3Lab, AAI CERC, Université de Montréal, MILA - Québec AI Institute, Stanford Center for Research on Foundation Models (CRFM), Stanford Hazy Research re…
0. TL;DR 1. Model Details 2. Training Details 3. Usage 4. Evaluation 5. Citation
DeepHermes 3 Preview is the latest version of our flagship Hermes series of LLMs by Nous Research, and one of the first models in the world to unify Reasoning (long chains of thought that improve answer accuracy) and normal LLM response modes into one model. We have also improved…
OpenCALM is a suite of decoder-only language models pre-trained on Japanese datasets, developed by CyberAgent, Inc.
SEA-LION is a collection of Large Language Models (LLMs) which has been pretrained and instruct-tuned for the Southeast Asia (SEA) region. The size of the models range from 3 billion to 7 billion parameters. This is the card for the SEA-LION 7B base model.
🤗 HuggingFace 🤖 ModelScope 🪡 AngelSlim
CyberAgentLM2 is a decoder-only language model pre-trained on the 1.3T tokens of publicly available Japanese and English datasets.
Qwen3-Swallow v0.2 is a family of large language models available in 8B , 30B-A3B , and 32B parameter sizes. Built as bilingual Japanese-English models, they were developed through Continual Pre-Training (CPT), Supervised Fine-Tuning (SFT), and Reinforcement Learning with Verifia…
VISTA-9B are GUI-grounding vision-language models trained from Qwen3.5 9B backbones with VISTA: View-Consistent Self-Verified Training for GUI Grounding .
CodeGen2 is a family of autoregressive language models for program synthesis , introduced in the paper:
🚀 LLaDA2.1-flash is now live on ZenmuxAI ! Try it via API 🛠️ or Chat 💬: https://zenmux.ai/inclusionai/llada2.1-flash
🤗 Hugging Face 🤖 ModelScope Tech Report 💻 GitHub