LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
1. Model Summary 2. Evaluation 3. Intended Use 4. Limitations 5. Security and Responsible Use 6. License 7. Citation
Overview The model is the instruction-tuned version of rinna/nekomata-14b . It adopts the Alpaca input format.
Our Swallow-MS-7b-v0.1 model has undergone continual pre-training from the Mistral-7B-v0.1, primarily with the addition of Japanese language data.
Introduction We are thrilled to introduce Seed-Coder, a powerful, transparent, and parameter-efficient family of open-source code models at the 8B scale, featuring base, instruct, and reasoning variants. Seed-Coder contributes to promote the evolution of open code models through…
🦉GitHub 💬WeChat 百川API支持搜索增强和192K长窗口,新增百川搜索增强知识库、限时免费! 🚀 百川大模型在线对话平台 已正式向公众开放 🎉
1. Model Summary 2. Use 3. Limitations 4. Training 5. Evaluation 7. Citation
We conduct continual pre-training of google/gemma-2-2b on 80B tokens from a mixture of Japanese and English datasets. The continual pre-training improves the model's performance on Japanese tasks.
[!IMPORTANT] ❗ This repo requires the use of the macOS Sequoia (15) Developer Beta to utilize the latest and greatest CoreML has to offer! Sign up for the Apple Beta Software Program here to get access. Check out the companion blog post to learn more about what's new in iOS 18 &…
NOTE: This "delta model" cannot be used directly. Users have to apply it on top of the original LLaMA weights to get actual vicuna-13b-finetuned-langchain-MRKL weights. See https://github.com/rinnakk/vicuna-13b-delta-finetuned-langchain-MRKL model-weights for instructions.
[2026.1.28] We have provided the TileLang kernels supporting prefill (chunked-prefill as well) and decode (multi-token prediction as well). The full attention version is placed at flash mla interface.py while the streaming sparse attention version is placed at streaming sparse at…
Overview This repository provides an English-Japanese bilingual GPT-NeoX model of 3.8 billion parameters.
Gemma 2 Baku 2B Instruct (rinna/gemma-2-baku-2b-it)