LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
This is the base model for SmolVLM2-2.2B, a lightweight multimodal model designed to analyze video content. The model processes videos, images, and text inputs to generate text outputs - whether answering questions about media files, comparing visual content, or transcribing text…
--- license: apache-2.0 library name: transformers base model: - Qwen/Qwen2.5-7B pipeline tag: text-generation ---
This is a 16k context version of SmolLM2-1.7B-Instruct, which originnaly only supported 8k context. We finetune the model on 15k samples consisting of a subset of SmolTalk, LongAlign and SeaLong datasets and increase RoPE from 100k to 500k. This improves the evaluation on HELMET…
With a new decentralized training algorithm, we fine-tuned GPT-J (6B) on 3.53 billion tokens, resulting in GPT-JT (6B), a model that outperforms many 100B+ parameter models on classification benchmarks.
Llama 3 Youko 8B Instruct (rinna/llama-3-youko-8b-instruct)
StepFun-Formalizer: Unlocking the Autoformalization Potential of LLMs through Knowledge-Reasoning Fusion
Used dataset for fine-tuning - sahil2801/CodeAlpaca-20k - m-a-p/CodeFeedback-Filtered-Instruction
Qwen3-Swallow v0.2 is a family of large language models available in 8B , 30B-A3B , and 32B parameter sizes. Built as bilingual Japanese-English models, they were developed through Continual Pre-Training (CPT), Supervised Fine-Tuning (SFT), and Reinforcement Learning with Verifia…
Overview We conduct continual pre-training of qwen-14b on 66B tokens from a mixture of Japanese and English datasets. The continual pre-training significantly improves the model's performance on Japanese tasks. It also enjoys the following great features provided by the original…
State-of-the-art bilingual open-sourced Math reasoning LLMs. A solver , prover , verifier , augmentor .
A cute robot wearing a kimono writes calligraphy with one single brush — Stable Diffusion XL