LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
🦉GitHub 💬WeChat 百川API支持搜索增强和192K长窗口,新增百川搜索增强知识库、限时免费! 🚀 百川大模型在线对话平台 已正式向公众开放 🎉
[!IMPORTANT] ❗ This repo requires the use of the macOS Sequoia (15) Developer Beta to utilize the latest and greatest CoreML has to offer! Sign up for the Apple Beta Software Program here to get access. Check out the companion blog post to learn more about what's new in iOS 18 &…
NOTE: This "delta model" cannot be used directly. Users have to apply it on top of the original LLaMA weights to get actual vicuna-13b-finetuned-langchain-MRKL weights. See https://github.com/rinnakk/vicuna-13b-delta-finetuned-langchain-MRKL model-weights for instructions.
[2026.1.28] We have provided the TileLang kernels supporting prefill (chunked-prefill as well) and decode (multi-token prediction as well). The full attention version is placed at flash mla interface.py while the streaming sparse attention version is placed at streaming sparse at…
Gemma 2 Baku 2B Instruct (rinna/gemma-2-baku-2b-it)
Model Introduction We introduce LongCat-Flash, a powerful and efficient language model with 560 billion total parameters, featuring an innovative Mixture-of-Experts (MoE) architecture. The model incorporates a dynamic computation mechanism that activates 18.6B∼31.3B parameters (a…
Model Card for "calm2-7b-chat-dpo-experimental"
Llama 3.1 Swallow is a series of large language models (8B, 70B) that were built by continual pre-training on the Meta Llama 3.1 models. Llama 3.1 Swallow enhanced the Japanese language capabilities of the original Llama 3.1 while retaining the English language capabilities. We u…
Overview The model is the instruction-tuned version of rinna/youri-7b . It adopts a chat-style input format.
This is a Japanese finetuned model based on deepseek-ai/DeepSeek-R1-Distill-Qwen-14B.
We introduce an updated version of LongCat-Flash-Thinking, named LongCat-Flash-Thinking-2601, a powerful and efficient Large Reasoning Model (LRM) with 560 billion total parameters, built upon an innovative Mixture-of-Experts (MoE) architecture.
Authors : Yizhe Zhang, Navdeep Jaitly (Apple)