LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
NVIDIA-Nemotron-Nano-VL-12B-V2-FP4-QAD is the quantized version of the NVIDIA Nemotron Nano VL V2 model, which is an auto-regressive vision language model that uses an optimized transformer architecture. For more information, please check here. The NVIDIA Nemotron Nano VL FP4 QAD…
We introduce Olmo 3, a new family of 7B and 32B models both Instruct and Think variants. Long chain-of-thought thinking improves reasoning tasks like math and coding.
[!Note] This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc. In light of its parameter scale, the intende…
We introduce Olmo 3, a new family of 7B and 32B models both Instruct and Think variants. Long chain-of-thought thinking improves reasoning tasks like math and coding.
A Pocket-Sized MLLM for Ultra-Efficient Image and Video Understanding on Your Phone
Model Summary: Granite-4.1-8B is a 8B parameter long-context instruct model finetuned from Granite-4.1-8B-Base using a combination of open source instruction datasets with permissive license and internally collected synthetic datasets. Granite 4.1 models have gone through an impr…
This repository hosts the GPTQ (W4A16, GPTQModel) quantized version of MiniCPM-V 4.6 Thinking. For the original BF16 weights and the full model card, please refer to openbmb/MiniCPM-V-4.6-Thinking.
A GPT-4V Level MLLM for Single Image, Multi Image and Video on Your Phone
This repository hosts the bitsandbytes (NF4, 4-bit) quantized version of MiniCPM-V 4.6. For the original BF16 weights and the full model card, please refer to openbmb/MiniCPM-V-4.6.
This repository hosts the GPTQ (W4A16, GPTQModel) quantized version of MiniCPM-V 4.6. For the original BF16 weights and the full model card, please refer to openbmb/MiniCPM-V-4.6.
Molmo is a family of open vision-language models developed by the Allen Institute for AI. Molmo models are trained on PixMo, a dataset of 1 million, highly-curated image-text pairs. It has state-of-the-art performance among multimodal models with a similar size while being fully…
📣 Update [10-07-2025]: Added a default system prompt to the chat template to guide the model towards more professional, accurate, and safe responses.