AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

440 models for "Base Model" Compare
SingGuard-NSFA-0.8B logo
SingGuard-NSFA-0.8B
inclusionAI

SingGuard-NSFA: Extensible Guardrails for Agentic AI via Generative Reasoning and Real-Time Classification

Open Source 0.8B ↓ 558
Baichuan-M3-235B logo
Baichuan-M3-235B
baichuan-inc

From Inquiry to Decision: Building Trustworthy Medical AI

Open Source 235.0B ↓ 547
Swallow-13b-hf logo
Swallow-13b-hf
tokyotech-llm

Our Swallow model has undergone continual pre-training from the Llama 2 family, primarily with the addition of Japanese language data. The tuned versions use supervised fine-tuning (SFT). Links to other models can be found in the index.

Open Source 13.0B ↓ 537
GLM-4-32B-Base-0414 logo
GLM-4-32B-Base-0414
zai-org

The GLM family welcomes new members, the GLM-4-32B-0414 series models, featuring 32 billion parameters. Its performance is comparable to OpenAI’s GPT series and DeepSeek’s V3/R1 series. It also supports very user-friendly local deployment features. GLM-4-32B-Base-0414 was pre-tra…

Open Source 32.0B ↓ 532
cogagent-9b-20241220 logo
cogagent-9b-20241220
zai-org

🌐 Github 🤗 Huggingface Space 📄 Technical Report 📜 arxiv paper

Multimodal 9.0B ↓ 496
xgen-mm-phi3-mini-instruct-r-v1 logo
xgen-mm-phi3-mini-instruct-r-v1
Salesforce

📣 News 📌 [08/19/2024] xGen-MM-v1.5 released: - 🤗 xgen-mm-phi3-mini-instruct-interleave-r-v1.5 - 🤗 xgen-mm-phi3-mini-base-r-v1.5 - 🤗 xgen-mm-phi3-mini-instruct-singleimg-r-v1.5 - 🤗 xgen-mm-phi3-mini-instruct-dpo-r-v1.5

Multimodal 5.0B ↓ 472
japanese-stablelm-base-gamma-7b logo
japanese-stablelm-base-gamma-7b
stabilityai

This is a 7B-parameter decoder-only language model with a focus on maximizing Japanese language modeling performance and Japanese downstream task performance. We conducted continued pretraining using Japanese data on the English language model, Mistral-7B-v0.1, to transfer the mo…

Open Source 7.0B ↓ 465
BFS-Prover-V2-7B logo
BFS-Prover-V2-7B
ByteDance-Seed

BFS-Prover-V2: Scaling up Multi-Turn Off-Policy RL and Multi-Agent Tree Search for LLM Step-Provers

Reasoning 7.0B ↓ 463
Gemma-2-Llama-Swallow-9b-pt-v0.1 logo
Gemma-2-Llama-Swallow-9b-pt-v0.1
tokyotech-llm

Gemma-2-Llama-Swallow series was built by continual pre-training on the gemma-2 models. Gemma 2 Swallow enhanced the Japanese language capabilities of the original Gemma 2 while retaining the English language capabilities. We use approximately 200 billion tokens that were sampled…

Open Source 9.0B ↓ 462
Yarn-Llama-2-13b-128k logo
Yarn-Llama-2-13b-128k
NousResearch

Nous-Yarn-Llama-2-13b-128k is a state-of-the-art language model for long context, further pretrained on long context data for 600 steps. This model is the Flash Attention 2 patched version of the original model: https://huggingface.co/conceptofmind/Yarn-Llama-2-13b-128k

Open Source 13.0B ↓ 460
Yarn-Llama-2-7b-64k logo
Yarn-Llama-2-7b-64k
NousResearch

Nous-Yarn-Llama-2-7b-64k is a state-of-the-art language model for long context, further pretrained on long context data for 400 steps. This model is the Flash Attention 2 patched version of the original model: https://huggingface.co/conceptofmind/Yarn-Llama-2-7b-64k

Open Source 7.0B ↓ 459
internlm2_5-7b-chat-1m logo
internlm2_5-7b-chat-1m
internlm

💻Github Repo • 🤔Reporting Issues • 📜Technical Report

Open Source 7.0B ↓ 459