LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
Llama 3 Youko 8B Instruct (rinna/llama-3-youko-8b-instruct)
StepFun-Formalizer: Unlocking the Autoformalization Potential of LLMs through Knowledge-Reasoning Fusion
Overview We conduct continual pre-training of qwen-14b on 66B tokens from a mixture of Japanese and English datasets. The continual pre-training significantly improves the model's performance on Japanese tasks. It also enjoys the following great features provided by the original…
State-of-the-art bilingual open-sourced Math reasoning LLMs. A solver , prover , verifier , augmentor .
A cute robot wearing a kimono writes calligraphy with one single brush — Stable Diffusion XL
1. Model Summary 2. Use 3. Limitations 4. Training 5. Evaluation 7. Citation
Model Summary Chat with the model at: https://huggingface.co/spaces/HuggingFaceTB/instant-smol
AHN: Artificial Hippocampus Networks for Efficient Long-Context Modeling
💻Github Repo • 🤔Reporting Issues • 📜Technical Report
This is the model card of a 🤗 transformers model that has been pushed on the Hub. This model card has been automatically generated.
We introduce LongCat-Flash-Prover , a flagship $560$-billion-parameter open-source Mixture-of-Experts (MoE) model that advances Native Formal Reasoning in Lean4 through agentic tool-integrated reasoning (TIR). We decompose the native formal reasoning task into three independent f…
SmolVLM is a compact open multimodal model that accepts arbitrary sequences of image and text inputs to produce text outputs. Designed for efficiency, SmolVLM can answer questions about images, describe visual content, create stories grounded on multiple images, or function as a…