LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
Note: Users are permitted to use this model in accordance with the Llama 3.1 Community License Agreement. Additionally, due to the licensing restrictions of the dataset used to train this model, which prohibits commercial use, the Dragonfly-Med model is restricted to non-commerci…
AHN: Artificial Hippocampus Networks for Efficient Long-Context Modeling
Our Swallow model has undergone continual pre-training from the Llama 2 family, primarily with the addition of Japanese language data. The tuned versions use supervised fine-tuning (SFT). Links to other models can be found in the index.
Note: Users are permitted to use this model in accordance with the Llama 3.1 Community License Agreement.
Qianfan-VL: Domain-Enhanced Universal Vision-Language Models
Notice: This model requires transformers =4.31.0 to work properly.
Overview The model is the instruction-tuned version of rinna/youri-7b . It adopts the Alpaca input format.
Our Swallow model has undergone continual pre-training from the Llama 2 family, primarily with the addition of Japanese language data. The tuned versions use supervised fine-tuning (SFT). Links to other models can be found in the index.
This is an EAGLE3 draft model trained from scratch (random initialization) using the Aurora inference-time training framework for speculative decoding. Unlike traditional approaches that fine-tune pre-trained models, this model is built entirely through Aurora's online training p…
1. Model Summary 2. Use 3. Limitations 4. Training 5. Evaluation 7. Citation
[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.