LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
Description: The NVIDIA GLM-5.3-Flash NVFP4 model is the quantized version of ZAI's GLM-5.3-Flash model, which is an auto-regressive language model that uses an optimized transformer architecture. GLM-5.3-Flash is a natively multimodal Mixture-of-Experts (MoE) model for reasoning…
👋 Join our WeChat or Discord community. 📖 Check out the GLM-5.3-Flash blog and GLM-5 Technical report . 📍 Use GLM-5.3-Flash API services on Z.ai API Platform.
Description: The NVIDIA Qwen3.8-Flash-Next NVFP4 model is the quantized version of Alibaba's Qwen3.8-Flash-Next model, which is an auto-regressive language model that uses an optimized transformer architecture. Qwen3.8-Flash-Next is a causal language model with a vision encoder,…
DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression
Description: The NVIDIA Qwen3.6-27B NVFP4 model is the quantized version of Alibaba's Qwen3.6-27B model, which is an auto-regressive language model that uses an optimized transformer architecture. For more information, please check here. The NVIDIA Qwen3.6-27B NVFP4 model is quan…
DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
Description: The NVIDIA Qwen3.8-27B NVFP4 model is a quantized version of Alibaba's Qwen3.8-27B model, which is an auto-regressive language model that uses an optimized transformer architecture. For more information on the model, please check here. The model is quantized with Mod…
We introduce Olmo 3, a new family of 7B and 32B models both Instruct and Think variants. Long chain-of-thought thinking improves reasoning tasks like math and coding.
We introduce Olmo 3, a new family of 7B and 32B models both Instruct and Think variants. Long chain-of-thought thinking improves reasoning tasks like math and coding.