LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
  GITHUB      🖥️   official website   |  🕖   HunyuanAPI |  🐳   Gitee Technical Report   |   Demo    |   Tencent Cloud TI    
🤗 Hugging Face • 🤖 ModelScope • 🟣 wisemodel
Nous-Hermes-13b is a state-of-the-art language model fine-tuned on over 300,000 instructions. This model was fine-tuned by Nous Research, with Teknium and Karan4D leading the fine tuning process and dataset curation, Redmond AI sponsoring the compute, and several other contributo…
[Homepage] [Paper] [Discord] [Dataset] [Github]
Introduction We are thrilled to introduce Seed-Coder, a powerful, transparent, and parameter-efficient family of open-source code models at the 8B scale, featuring base, instruct, and reasoning variants. Seed-Coder contributes to promote the evolution of open code models through…
CapRL 📖 Paper 🏠 Github 🤗 CapRL Collection 🤗 Daily Paper
Introduction We are thrilled to introduce Stable-DiffCoder, which is a strong code diffusion large language model. Built directly on the Seed-Coder architecture, data, and training pipeline, it introduces a block diffusion continual pretraining (CPT) stage with a tailored warmup…
Building the Next Generation of Open-Source and Bilingual LLMs
GPT-OSS-Swallow v0.1 is a family of large language models available in 20B and 120B parameter sizes. Built as bilingual Japanese-English models, they were developed through Continual Pre-Training (CPT), Supervised Fine-Tuning (SFT), and Reinforcement Learning with Verifiable Rewa…
🤗 HuggingFace 🤖 ModelScope 🪡 AngelSlim
We release Nomos 1 , a specialization of Qwen/Qwen3-30B-A3B-Thinking-2507 for mathematical problem-solving and proof-writing in natural language. Nomos-1 was trained in collaboration with Hillclimb AI.
The Shanghai Artificial Intelligence Laboratory, in collaboration with SenseTime Technology, the Chinese University of Hong Kong, and Fudan University, has officially released the 20 billion parameter pretrained model, InternLM-20B. InternLM-20B was pre-trained on over 2.3T Token…