LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
japanese-gpt-neox-3.6b-instruction-sft-v2
Llama 2 Llama 2 is a collection of pretrained and fine-tuned generative text models ranging in scale from 7 billion to 70 billion parameters. This is the repository for the 70B pretrained model, converted for the Hugging Face Transformers format. Links to other models can be foun…
Nous Hermes 2 on Mistral 7B DPO is the new flagship 7B Hermes! This model was DPO'd from Teknium/OpenHermes-2.5-Mistral-7B and has improved across the board on all benchmarks tested - AGIEval, BigBench Reasoning, GPT4All, and TruthfulQA.
🤗 Hugging Face 🤖 ModelScope 🐙 Experience Now
Stable LM 2 12B Chat is a 12 billion parameter instruction tuned language model trained on a mix of publicly available datasets and synthetic datasets, utilizing Direct Preference Optimization (DPO).
The MiniCPM-MoE-8x2B is a decoder-only transformer-based generative language model.
🤗 Hugging Face      🤖 ModelScope 🐙 Experience Now
Japanese-StableLM-Instruct-JAVocab-Beta-7B
A cute robot wearing a kimono writes calligraphy with one single brush — Stable Diffusion XL
Overview This repository provides a Japanese GPT-NeoX model of 3.6 billion parameters. The model is based on rinna/japanese-gpt-neox-3.6b-instruction-sft-v2 and has been aligned to serve as an instruction-following conversational agent.
A cute robot wearing a kimono writes calligraphy with one single brush — Stable Diffusion XL
Our Swallow model has undergone continual pre-training from the Llama 2 family, primarily with the addition of Japanese language data. The tuned versions use supervised fine-tuning (SFT). Links to other models can be found in the index.