LLM Models · NousResearch
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
Nous-Yarn-Llama-2-13b-64k is a state-of-the-art language model for long context, further pretrained on long context data for 400 steps. This model is the Flash Attention 2 patched version of the original model: https://huggingface.co/conceptofmind/Yarn-Llama-2-13b-64k
The authors would like to thank LAION AI for their support of compute for this model. It was trained on the JUWELS supercomputer.
OLMo-Bitnet-1B is a 1B parameter model trained using the method described in The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits.
Redmond-Hermes-Coder 15B is a state-of-the-art language model fine-tuned on over 300,000 instructions. This model was fine-tuned by Nous Research, with Teknium and Karan4D leading the fine tuning process and dataset curation, Redmond AI sponsoring the compute, and several other c…
An alternate Meta-Llama-3-8B Repo for the Hermes Tokenizer