LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
MiniMax-Text-01 is a powerful language model with 456 billion total parameters, of which 45.9 billion are activated per token. To better unlock the long context capabilities of the model, MiniMax-Text-01 adopts a hybrid architecture that combines Lightning Attention, Softmax Atte…
Nous Hermes 2 Mixtral 8x7B SFT is the supervised finetune only version of our new flagship Nous Research model trained over the Mixtral 8x7B MoE LLM.
--- license: mit pipeline tag: text-generation library name: transformers ---
Hermes 3 is the latest version of our flagship Hermes series of LLMs by Nous Research.
MolmoPoint-8B MolmoPoint-8B is a fully-open VLM developed by the Allen Institute for AI (Ai2) that support image, video and multi-image understanding and grounding. It has new pointing mechansim that improves image pointing, video pointing, and video tracking, see our technical r…
This repository provides a tiny 16M parameters language model for debugging and testing purposes. This is created by tuning sbintuitions/tiny-lm with oasset1 datasets in Japanese and English.
💻Github Repo • 🤔Reporting Issues • 📜Technical Report
[🏠Homepage] [🤖 Chat with DeepSeek LLM] [Discord] [Wechat(微信)]
Kimi Linear: An Expressive, Efficient Attention Architecture
Hermes-2 Θ (Theta) 70B is the continuation of our experimental merged model released by Nous Research, in collaboration with Charles Goddard and Arcee AI, the team behind MergeKit.
🤗 Hugging Face      🤖 ModelScope