LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
The GLM family welcomes a new generation of open-source models, the GLM-4-32B-0414 series, featuring 32 billion parameters. Its performance is comparable to OpenAI's GPT series and DeepSeek's V3/R1 series, and it supports very user-friendly local deployment features. GLM-4-32B-Ba…
🤗 HuggingFace 🤖 ModelScope 🪡 AngelSlim
📢 DISCLAIMER : The StableLM-Base-Alpha models have been superseded. Find the latest versions in the Stable LM Collection here.
A small ~110M parameter language model implementing the DeepSeek-V4 architecture , fine-tuned for chat/instruction following. Trained from scratch — no weights from DeepSeek-V4 were used.
[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.
This is the RL initialization checkpoint used in the blog JustRL II: Scaling Small LLMs to 128K Reasoning with a Critic (中文版).
Japanese-StableLM-Instruct-JAVocab-Beta-7B
👋 Join our Discord community. 📖 Check out the GLM-4.5 technical blog , technical report , and Zhipu AI technical documentation . 📍 Use GLM-4.5 API services on Z.ai API Platform (Global) or Zhipu AI Open Platform (Mainland China) . 👉 One click to GLM-4.5 .
A cute robot wearing a kimono writes calligraphy with one single brush — Stable Diffusion XL
HiLS-Attention is a chunk-wise sparse attention mechanism that learns chunk selection end-to-end under the language-modeling loss, enabling native sparse training for efficient long-context modeling. This repository hosts the 7B checkpoint continued-trained on top of an OLMo3-sty…
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ Unlocking the Reasoning Potential of Language Model From Pretraining to Posttraining ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
🤗 HuggingFace 🤖 ModelScope 🪡 AngelSlim