LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
We introduce EXAONE Deep, which exhibits superior capabilities in various reasoning tasks including math and coding benchmarks, ranging from 2.4B to 32B parameters developed and released by LG AI Research. Evaluation results show that 1) EXAONE Deep 2.4B outperforms other models…
🤗 Hugging Face 💻 Github Repository 📑 Technique Report 💬 Issues & Discussions
SEA-Safeguard is a collection of safety-focused Large Language Models (LLMs) built upon the SEA-LION family, designed specifically for the Southeast Asia (SEA) region.
WARNING: This is an intermediary checkpoint and WIP project. It is not fully trained yet. You might want to use Bloom-1B3 if you want a model that has completed training. This model is a distilled version of Bloom-1B3 (10x distillation)
🤗 Hugging Face 💻 Github Repository 📑 Technique Report 💬 Issues & Discussions
WARNING: This is an intermediary checkpoint and WIP project. It is not fully trained yet. You might want to use Bloom-1B3 if you want a model that has completed training. This model is a distilled version of Bloom-1B3
[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.
We introduce EXAONE Deep, which exhibits superior capabilities in various reasoning tasks including math and coding benchmarks, ranging from 2.4B to 32B parameters developed and released by LG AI Research. Evaluation results show that 1) EXAONE Deep 2.4B outperforms other models…
SmolVLM is a compact open multimodal model that accepts arbitrary sequences of image and text inputs to produce text outputs. Designed for efficiency, SmolVLM can answer questions about images, describe visual content, create stories grounded on multiple images, or function as a…
This is a 8bit quantized version of upstage/SOLAR-0-70b-16bit
Llama 3 Youko 8B Instruct GPTQ (rinna/llama-3-youko-8b-instruct-gptq)