LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
🤗 Hugging Face 💻 Github Repository 📑 Technique Report 💬 Issues & Discussions
WARNING: This is an intermediary checkpoint and WIP project. It is not fully trained yet. You might want to use Bloom-1B3 if you want a model that has completed training. This model is a distilled version of Bloom-1B3
[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.
We introduce EXAONE Deep, which exhibits superior capabilities in various reasoning tasks including math and coding benchmarks, ranging from 2.4B to 32B parameters developed and released by LG AI Research. Evaluation results show that 1) EXAONE Deep 2.4B outperforms other models…
SmolVLM is a compact open multimodal model that accepts arbitrary sequences of image and text inputs to produce text outputs. Designed for efficiency, SmolVLM can answer questions about images, describe visual content, create stories grounded on multiple images, or function as a…
This is a 8bit quantized version of upstage/SOLAR-0-70b-16bit
WangchanLION is a joint effort between VISTEC and AI Singapore to develop a Thai-specific collection of Large Language Models (LLMs), pre-trained for Southeast Asian (SEA) languages, and instruct-tuned specifically for the Thai language.
Llama 3 Youko 70B GPTQ (rinna/llama-3-youko-70b-gptq)
Llama 3 Youko 8B Instruct GPTQ (rinna/llama-3-youko-8b-instruct-gptq)
Llama 3 Youko 70B Instruct GPTQ (rinna/llama-3-youko-70b-instruct-gptq)
Llama 3 Youko 8B GPTQ (rinna/llama-3-youko-8b-gptq)
RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning