LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
This model card introduces a moderation model, a GPT-JT model fine-tuned on Ontocord.ai's OIG-moderation dataset v0.1.
WARNING: The checkpoints on this repo are not fully trained model. Evaluations of intermediary checkpoints and the final model will be added when conducted (see below).
[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.
Qwen-SEA-LION-v4-32B-IT-4BIT (GPTQ model)
[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.
Model Introduction We introduce LongCat-Flash, a powerful and efficient language model with 560 billion total parameters, featuring an innovative Mixture-of-Experts (MoE) architecture. The model incorporates a dynamic computation mechanism that activates 18.6B∼31.3B parameters (a…
Redmond-Hermes-Coder 15B is a state-of-the-art language model fine-tuned on over 300,000 instructions. This model was fine-tuned by Nous Research, with Teknium and Karan4D leading the fine tuning process and dataset curation, Redmond AI sponsoring the compute, and several other c…
This is a 8bit quantized version of upstage/SOLAR-0-70b-16bit
AquilaMoE: Efficient Training for MoE Models with Scale-Up and Scale-Out Strategies Language Foundation Model & Software Team Beijing Academy of Artificial Intelligence (BAAI) [Paper(released soon)] [Code] [github]
BaichuanMed-OCR-7B is a model fine-tuned from the Qwen2.5-VL-7B-Instruct with our constructed and curated medical report datasets consists of medical report images and related questions and answers (QAs). It has been specifically adapted to perform Optical Character Recognition (…
[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.
WARNING: The checkpoints on this repo are not fully trained model. Evaluations of intermediary checkpoints and the final model will be added when conducted (see below).