LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
We release Nomos 1 , a specialization of Qwen/Qwen3-30B-A3B-Thinking-2507 for mathematical problem-solving and proof-writing in natural language. Nomos-1 was trained in collaboration with Hillclimb AI.
Falcon-RW-7B is a 7B parameters causal decoder-only model built by TII and trained on 350B tokens of RefinedWeb. It is made available under the Apache 2.0 license.
The Shanghai Artificial Intelligence Laboratory, in collaboration with SenseTime Technology, the Chinese University of Hong Kong, and Fudan University, has officially released the 20 billion parameter pretrained model, InternLM-20B. InternLM-20B was pre-trained on over 2.3T Token…
Overview This repository provides a Japanese GPT-NeoX model of 3.6 billion parameters. The model is based on rinna/japanese-gpt-neox-3.6b-instruction-sft-v2 and has been aligned to serve as an instruction-following conversational agent.
1. Model Summary 2. Use 3. Training 4. Citation
Apriel-1.5-15b-Thinker - Mid training is all you need!
Sachin Mehta, Mohammad Hossein Sekhavat, Qingqing Cao, Maxwell Horton, Yanzi Jin, Chenfan Sun, Iman Mirzadeh, Mahyar Najibi, Dmitry Belenko, Peter Zatloukal, Mohammad Rastegari
━━━━━━━━━━━━━━━━━━━━━━━━━ Unlocking the Reasoning Potential of Language Model From Pretraining to Posttraining ━━━━━━━━━━━━━━━━━━━━━━━━━
Nous-Yarn-Llama-2-13b-128k is a state-of-the-art language model for long context, further pretrained on long context data for 600 steps. This model is the Flash Attention 2 patched version of the original model: https://huggingface.co/conceptofmind/Yarn-Llama-2-13b-128k
This is a 7B-parameter decoder-only Japanese language model fine-tuned on instruction-following datasets, built on top of the base model Japanese Stable LM Base Gamma 7B.
Building the Next Generation of Open-Source and Bilingual LLMs