LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
Our Swallow model has undergone continual pre-training from the Llama 2 family, primarily with the addition of Japanese language data. The tuned versions use supervised fine-tuning (SFT). Links to other models can be found in the index.
1. Model Summary 2. Use 3. Limitations 4. Training 5. Evaluation 7. Citation
WARNING: The checkpoints on this repo are not fully trained model. Evaluations of intermediary checkpoints and the final model will be added when conducted (see below).
SEA-Guard is a collection of safety-focused Large Language Models (LLMs) built upon the SEA-LION family, designed specifically for the Southeast Asia (SEA) region.
Cola DLM ( Co ntinuous La tent D iffusion L anguage M odel) is a hierarchical continuous latent-space diffusion language model. It combines a Text VAE with a block-causal Diffusion Transformer (DiT) prior: the VAE maps text into continuous latent sequences and decodes latents bac…
- 2023/08/02 We uploaded the newly trained rinna/bilingual-gpt-neox-4b-instruction-sft with the MIT license. - Please refrain from using the previous model released on 2023/07/31 for commercial purposes if you have already downloaded it. - The new model released on 2023/08/02 is…
Our Swallow model has undergone continual pre-training from the Llama 2 family, primarily with the addition of Japanese language data. The tuned versions use supervised fine-tuning (SFT). Links to other models can be found in the index.
WARNING: The checkpoints on this repo are not fully trained model. Evaluations of intermediary checkpoints and the final model will be added when conducted (see below).
BaichuanMed-OCR-72B is a model fine-tuned from the Qwen2.5-VL-72B-Instruct with our constructed and curated medical report datasets consists of medical report images and related questions and answers (QAs). It has been specifically adapted to perform Optical Character Recognition…
One of the focus areas at Together Research is new architectures for long context, improved training, and inference performance over the Transformer architecture. Spinning out of a research program from our team and academic collaborators, with roots in signal processing-inspired…
CAT-Paws is an agentic LLM that thinks in Japanese (e.g., reasoning trace is in Japanese). The model is based on Qwen3-Swallow-v0.2 which is a continual pretraining model based on Qwen3 to read and write fluently in Japanese.
Our Swallow-MS-7b-v0.1 model has undergone continual pre-training from the Mistral-7B-v0.1, primarily with the addition of Japanese language data.