LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.
Our Swallow model has undergone continual pre-training from the Llama 2 family, primarily with the addition of Japanese language data. The tuned versions use supervised fine-tuning (SFT). Links to other models can be found in the index.
[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.
[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.
🤗 Hugging Face 💻 Github Repository 📑 Technique Report 💬 Issues & Discussions
Jamba2 Mini is an open source small language model built for enterprise reliability. With 12B active parameters (52B total), it delivers precise question answering without the computational overhead of reasoning models. The model's SSM-Transformer architecture provides a memory-e…
Llama 3 Youko 70B Instruct (rinna/llama-3-youko-70b-instruct)
💻Github Repo • 🤗HF Model Collections • 🤖ModelScope Collections • 💬Online Chat
🤗 Hugging Face 💻 Github Repository 📑 Technique Report 💬 Issues & Discussions
From Inquiry to Decision: Building Trustworthy Medical AI
Cola DLM ( Co ntinuous La tent D iffusion L anguage M odel) is a hierarchical continuous latent-space diffusion language model. It combines a Text VAE with a block-causal Diffusion Transformer (DiT) prior: the VAE maps text into continuous latent sequences and decodes latents bac…
This repository contains the NVFP4 (4-bit floating point, E2M1) quantized weights for aisingapore/Nemotron-SEA-LION-v4.8-30B-A3B .