LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
[!Note] This repository contains FP8-quantized model weights and configuration files for the post-trained model in the Hugging Face Transformers format. These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc. The quantization method is fin…
[!Note] This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc.
Description: The NVIDIA Qwen3.6-27B NVFP4 model is the quantized version of Alibaba's Qwen3.6-27B model, which is an auto-regressive language model that uses an optimized transformer architecture. For more information, please check here. The NVIDIA Qwen3.6-27B NVFP4 model is quan…
:--- :--- Total Parameters 550B (55B active) Architecture LatentMoE - Mamba-2 + MoE + Attention hybrid with Multi-Token Prediction (MTP) Context Length Up to 1M tokens Minimum GPU Requirement 4xGB200, 4xB200, 4x GB300, 4x B300, 8xH100 Supported Languages English, French, Spanish,…
🤗 HuggingFace 📰 Blog 🎨 Xiaomi MiMo API Platform 🗨️ Xiaomi MiMo Studio
:--- :--- Total Parameters 550B (55B active) Architecture LatentMoE - Mamba-2 + MoE + Attention hybrid with Multi-Token Prediction (MTP) Context Length Up to 1M tokens Minimum GPU Requirement 8x GB200/B200/GB300/B300, 16x H100, 8x H200 Supported Languages English, French, Spanish…
🤗 Hugging Face 🤖 ModelScope 🐙 OpenRouter
🤗 Hugging Face 🤖 ModelScope 🐙 OpenRouter
🤗 Hugging Face 🤖 ModelScope 🐙 OpenRouter
🤗 Hugging Face 🤖 ModelScope 🐙 OpenRouter
A DSpark speculator for Ling3. DSpark extends DFlash with target-model auxiliary features and a confidence head that dynamically chooses the number of draft tokens. The model was trained with SpecForge and is served with SGLang.