LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
py from mistral common.tokens.tokenizers.mistral import MistralTokenizer from mistral common.protocol.instruct.messages import UserMessage from mistral common.protocol.instruct.request import ChatCompletionRequest
Join Our 💬 WeChat 🧩 Discord community. MiniMax Agent ⚡️ API CLI MiniMax Website 🤗 Hugging Face 🐙 GitHub 🤖️ ModelScope 📄 LICENSE
We introduce the updated version of the Qwen3-4B-FP8 non-thinking mode , named Qwen3-4B-Instruct-2507-FP8 , featuring the following key enhancements:
Qwen3-Coder is available in multiple sizes. Today, we're excited to introduce Qwen3-Coder-30B-A3B-Instruct-FP8 . This streamlined model maintains impressive performance and efficiency, featuring the following key enhancements:
DeepSeek-V3-0324 demonstrates notable improvements over its predecessor, DeepSeek-V3, in several key aspects.
We present DeepSeek-V3, a strong Mixture-of-Experts (MoE) language model with 671B total parameters with 37B activated for each token. To achieve efficient inference and cost-effective training, DeepSeek-V3 adopts Multi-head Latent Attention (MLA) and DeepSeekMoE architectures, w…
Qwen2.5 is the latest series of Qwen large language models. For Qwen2.5, we release a number of base language models and instruction-tuned language models ranging from 0.5 to 72 billion parameters. Qwen2.5 brings the following improvements upon Qwen2:
The DeepSeek R1 model has undergone a minor version upgrade, with the current version being DeepSeek-R1-0528. In the latest update, DeepSeek R1 has significantly improved its depth of reasoning and inference capabilities by leveraging increased computational resources and introdu…
Qwen2 is the new series of Qwen large language models. For Qwen2, we release a number of base language models and instruction-tuned language models ranging from 0.5 to 72 billion parameters, including a Mixture-of-Experts model. This repo contains the instruction-tuned 1.5B Qwen2…
Qwen2 is the new series of Qwen large language models. For Qwen2, we release a number of base language models and instruction-tuned language models ranging from 0.5 to 72 billion parameters, including a Mixture-of-Experts model. This repo contains the 0.5B Qwen2 base language mod…