AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

133 models for "RLHF" Compare
japanese-gpt-neox-3.6b-instruction-sft-v2 logo
japanese-gpt-neox-3.6b-instruction-sft-v2
rinna

japanese-gpt-neox-3.6b-instruction-sft-v2

Open Source 3.6B ↓ 903
Llama-2-70b-hf logo
Llama-2-70b-hf
NousResearch

Llama 2 Llama 2 is a collection of pretrained and fine-tuned generative text models ranging in scale from 7 billion to 70 billion parameters. This is the repository for the 70B pretrained model, converted for the Hugging Face Transformers format. Links to other models can be foun…

Open Source 70.0B ↓ 899
Nous-Hermes-2-Mistral-7B-DPO logo
Nous-Hermes-2-Mistral-7B-DPO
NousResearch

Nous Hermes 2 on Mistral 7B DPO is the new flagship 7B Hermes! This model was DPO'd from Teknium/OpenHermes-2.5-Mistral-7B and has improved across the board on all benchmarks tested - AGIEval, BigBench Reasoning, GPT4All, and TruthfulQA.

Open Source 7.0B ↓ 848
Ring-1T logo
Ring-1T
inclusionAI

🤗 Hugging Face      🤖 ModelScope      🐙 Experience Now

Open Source ↓ 794
stablelm-2-12b-chat logo
stablelm-2-12b-chat
stabilityai

Stable LM 2 12B Chat is a 12 billion parameter instruction tuned language model trained on a mix of publicly available datasets and synthetic datasets, utilizing Direct Preference Optimization (DPO).

Open Source 12.0B ↓ 676
MiniCPM-MoE-8x2B logo
MiniCPM-MoE-8x2B
openbmb

The MiniCPM-MoE-8x2B is a decoder-only transformer-based generative language model.

Open Source 13.6B ↓ 670
Ring-mini-2.0 logo
Ring-mini-2.0
inclusionAI

🤗 Hugging Face &nbsp&nbsp &nbsp&nbsp🤖 ModelScope   🐙 Experience Now

Open Source ↓ 610
japanese-stablelm-instruct-ja_vocab-beta-7b logo
japanese-stablelm-instruct-ja_vocab-beta-7b
stabilityai

Japanese-StableLM-Instruct-JAVocab-Beta-7B

Open Source 7.0B ↓ 582
japanese-stablelm-instruct-beta-70b logo
japanese-stablelm-instruct-beta-70b
stabilityai

A cute robot wearing a kimono writes calligraphy with one single brush — Stable Diffusion XL

Open Source 70.0B ↓ 572
japanese-gpt-neox-3.6b-instruction-ppo logo
japanese-gpt-neox-3.6b-instruction-ppo
rinna

Overview This repository provides a Japanese GPT-NeoX model of 3.6 billion parameters. The model is based on rinna/japanese-gpt-neox-3.6b-instruction-sft-v2 and has been aligned to serve as an instruction-following conversational agent.

Open Source 3.6B ↓ 562
japanese-stablelm-instruct-beta-7b logo
japanese-stablelm-instruct-beta-7b
stabilityai

A cute robot wearing a kimono writes calligraphy with one single brush — Stable Diffusion XL

Open Source 7.0B ↓ 554
Swallow-13b-hf logo
Swallow-13b-hf
tokyotech-llm

Our Swallow model has undergone continual pre-training from the Llama 2 family, primarily with the addition of Japanese language data. The tuned versions use supervised fine-tuning (SFT). Links to other models can be found in the index.

Open Source 13.0B ↓ 537