AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

1,877 models Compare
🤖
bloomz-mt
bigscience

1. Model Summary 2. Use 3. Limitations 4. Training 5. Evaluation 7. Citation

Open Source ↓ 119
🤖
Llama-3-Swallow-8B-v0.1
tokyotech-llm

Llama3 Swallow - Built with Meta Llama 3

Open Source 8.0B ↓ 118
🤖
Qwen-SEA-LION-v4-32B-IT-OV-8BIT
aisingapore

- Model creator: AI Singapore - Original model: Qwen-SEA-LION-v4-32B-IT

Open Source 32.0B ↓ 118
🤖
bloom-3b-intermediate
bigscience

WARNING: The checkpoints on this repo are not fully trained model. Evaluations of intermediary checkpoints and the final model will be added when conducted (see below).

Open Source 3.0B ↓ 117
🤖
Qwen-SEA-Guard-4B-2602
aisingapore

SEA-Guard is a collection of safety-focused Large Language Models (LLMs) built upon the SEA-LION family, designed specifically for the Southeast Asia (SEA) region.

Open Source 4.0B ↓ 115
🤖
Gemma-2-Llama-Swallow-2b-pt-v0.1
tokyotech-llm

Gemma-2-Llama-Swallow series was built by continual pre-training on the gemma-2 models. Gemma 2 Swallow enhanced the Japanese language capabilities of the original Gemma 2 while retaining the English language capabilities. We use approximately 200 billion tokens that were sampled…

Open Source 2.0B ↓ 114
🤖
Cola-DLM
ByteDance-Seed

Cola DLM ( Co ntinuous La tent D iffusion L anguage M odel) is a hierarchical continuous latent-space diffusion language model. It combines a Text VAE with a block-causal Diffusion Transformer (DiT) prior: the VAE maps text into continuous latent sequences and decodes latents bac…

Open Source ↓ 114
🤖
internlm2_5-step-prover
internlm

InternLM2.5-Step-Prover is a 7B language model which achieves state-of-the-art performances on MiniF2F, ProofNet, and Putnam math benchmarks, showing its formal math proving ability in multiple domains.

Open Source ↓ 113
🤖
TinySolar-248m-4k-py
upstage

Open Source ↓ 113
🤖
bilingual-gpt-neox-4b-instruction-sft
rinna

- 2023/08/02 We uploaded the newly trained rinna/bilingual-gpt-neox-4b-instruction-sft with the MIT license. - Please refrain from using the previous model released on 2023/07/31 for commercial purposes if you have already downloaded it. - The new model released on 2023/08/02 is…

Open Source 4.0B ↓ 112
🤖
Swallow-7b-instruct-v0.1
tokyotech-llm

Our Swallow model has undergone continual pre-training from the Llama 2 family, primarily with the addition of Japanese language data. The tuned versions use supervised fine-tuning (SFT). Links to other models can be found in the index.

Open Source 7.0B ↓ 112
🤖
bloom-1b7-intermediate
bigscience

WARNING: The checkpoints on this repo are not fully trained model. Evaluations of intermediary checkpoints and the final model will be added when conducted (see below).

Open Source ↓ 111