LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
--- license: apache-2.0 library name: transformers base model: - Qwen/Qwen2.5-7B pipeline tag: text-generation ---
This is a 16k context version of SmolLM2-1.7B-Instruct, which originnaly only supported 8k context. We finetune the model on 15k samples consisting of a subset of SmolTalk, LongAlign and SeaLong datasets and increase RoPE from 100k to 500k. This improves the evaluation on HELMET…
Llama 3 Youko 8B Instruct (rinna/llama-3-youko-8b-instruct)
StepFun-Formalizer: Unlocking the Autoformalization Potential of LLMs through Knowledge-Reasoning Fusion
Qwen3-Swallow v0.2 is a family of large language models available in 8B , 30B-A3B , and 32B parameter sizes. Built as bilingual Japanese-English models, they were developed through Continual Pre-Training (CPT), Supervised Fine-Tuning (SFT), and Reinforcement Learning with Verifia…
State-of-the-art bilingual open-sourced Math reasoning LLMs. A solver , prover , verifier , augmentor .
Model Summary Chat with the model at: https://huggingface.co/spaces/HuggingFaceTB/instant-smol
💻Github Repo • 🤔Reporting Issues • 📜Technical Report
We introduce LongCat-Flash-Prover , a flagship $560$-billion-parameter open-source Mixture-of-Experts (MoE) model that advances Native Formal Reasoning in Lean4 through agentic tool-integrated reasoning (TIR). We decompose the native formal reasoning task into three independent f…
SmolVLM is a compact open multimodal model that accepts arbitrary sequences of image and text inputs to produce text outputs. Designed for efficiency, SmolVLM can answer questions about images, describe visual content, create stories grounded on multiple images, or function as a…
State-of-the-art bilingual open-sourced Math reasoning LLMs. A solver , prover , verifier , augmentor .
CapRL: Stimulating Dense Image Caption Capabilities via Reinforcement Learning