LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
CyberAgentLM3 is a decoder-only language model pre-trained on 2.0 trillion tokens from scratch. CyberAgentLM3-Chat is a fine-tuned model specialized for dialogue use cases.
LLaDA2.0-flash is a diffusion language model featuring a 100BA6B Mixture-of-Experts (MoE) architecture. As an enhanced, instruction-tuned iteration of the LLaDA2.0 series, it is optimized for practical applications.
Apriel-1.5-15b-Thinker - Mid training is all you need!
0. TL;DR 1. Model Details 2. Training Details 3. Usage 4. Evaluation 5. Citation
Nous-Hermes-13b is a state-of-the-art language model fine-tuned on over 300,000 instructions. This model was fine-tuned by Nous Research, with Teknium and Karan4D leading the fine tuning process and dataset curation, Redmond AI sponsoring the compute, and several other contributo…
A small ~110M parameter language model implementing the DeepSeek-V4 architecture , fine-tuned for chat/instruction following. Trained from scratch — no weights from DeepSeek-V4 were used.
[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.
A cute robot wearing a kimono writes calligraphy with one single brush — Stable Diffusion XL
A cute robot wearing a kimono writes calligraphy with one single brush — Stable Diffusion XL
A cute robot wearing a kimono writes calligraphy with one single brush — Stable Diffusion XL
[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.
With a new decentralized training algorithm, we fine-tuned GPT-J (6B) on 3.53 billion tokens, resulting in GPT-JT (6B), a model that outperforms many 100B+ parameter models on classification benchmarks.