LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
AHN: Artificial Hippocampus Networks for Efficient Long-Context Modeling
Qianfan-VL: Domain-Enhanced Universal Vision-Language Models
[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.
Model Summary DiaFill is a Japanese dialogue script generation model designed to produce natural, spoken-style dialogue scripts rich in fillers and brief utternaces. Unlike typical assistant models that respond to users, this model is fine-tuned to generate a multi-turn dialogue…
[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.
[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.
Model Introduction We introduce LongCat-Flash-Omni , a state-of-the-art open-source omni-modal model with 560 billion parameters (with 27B activated), excelling at real-time audio-visual interaction, which is attained by leveraging LongCat-Flash's high-performance Shortcut-connec…
We introduce EXAONE Deep, which exhibits superior capabilities in various reasoning tasks including math and coding benchmarks, ranging from 2.4B to 32B parameters developed and released by LG AI Research. Evaluation results show that 1) EXAONE Deep 2.4B outperforms other models…
🤗 Hugging Face 💻 Github Repository 📑 Technique Report 💬 Issues & Discussions
SEA-Safeguard is a collection of safety-focused Large Language Models (LLMs) built upon the SEA-LION family, designed specifically for the Southeast Asia (SEA) region.
WARNING: This is an intermediary checkpoint and WIP project. It is not fully trained yet. You might want to use Bloom-1B3 if you want a model that has completed training. This model is a distilled version of Bloom-1B3 (10x distillation)
SEA-LION is a collection of Large Language Models (LLMs) which has been pretrained and instruct-tuned for the Southeast Asia (SEA) region. The sizes of the models range from 3 billion to 7 billion parameters.