LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
The pretraining data has a cutoff date of September 2024.
The pretraining data has a cutoff date of September 2024\.
Model Summary: Granite-4.1-30B is a 30B parameter long-context instruct model finetuned from Granite-4.1-30B-Base using a combination of open source instruction datasets with permissive license and internally collected synthetic datasets. Granite 4.1 models have gone through an i…
The pretraining data has a cutoff date of September 2024.
- 🚀 Online Demo : Explore Step3-VL-10B on Hugging Face Spaces ! - 📢 [Notice] FP8 Quantization Support : FP8 quantized weights are now available. (Download link) - 📢 [Notice] vLLM Support: vLLM integration is now officially supported! (PR 32329) - ✅ [Fixed] HF Inference: Resolved…
[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.
Hermes 4 405B is a frontier, hybrid-mode reasoning model based on Llama-3.1-405B by Nous Research that is aligned to you .
[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.
Hermes 4 405B is a frontier, hybrid-mode reasoning model based on Llama-3.1-405B by Nous Research that is aligned to you .
- 🚀 Online Demo : Explore Step3-VL-10B on Hugging Face Spaces ! - 📢 [Notice] FP8 Quantization Support : FP8 quantized weights are now available. (Download link) - 📢 [Notice] vLLM Support: vLLM integration is now officially supported! (PR 32329) - ✅ [Fixed] HF Inference: Resolved…
STEP3-VL-10B is a lightweight open-source foundation model designed to redefine the trade-off between compact efficiency and frontier-level multimodal intelligence. Despite its compact 10B parameter footprint , STEP3-VL-10B excels in visual perception , complex reasoning , and hu…
[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.