LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
1. Model Summary 2. Use 3. Limitations 4. Training 5. Evaluation 7. Citation
Model Summary Chat with the model at: https://huggingface.co/spaces/HuggingFaceTB/instant-smol
PaCoRe: Learning to Scale Test-Time Compute with Parallel Coordinated Reasoning
AHN: Artificial Hippocampus Networks for Efficient Long-Context Modeling
💻Github Repo • 🤔Reporting Issues • 📜Technical Report
This is the model card of a 🤗 transformers model that has been pushed on the Hub. This model card has been automatically generated.
We introduce LongCat-Flash-Prover , a flagship $560$-billion-parameter open-source Mixture-of-Experts (MoE) model that advances Native Formal Reasoning in Lean4 through agentic tool-integrated reasoning (TIR). We decompose the native formal reasoning task into three independent f…
SmolVLM is a compact open multimodal model that accepts arbitrary sequences of image and text inputs to produce text outputs. Designed for efficiency, SmolVLM can answer questions about images, describe visual content, create stories grounded on multiple images, or function as a…
🚀 BFS-Prover: Scalable Best-First Tree Search for LLM-based Automatic Theorem Proving State-of-the-art tactic generation model in Lean4
Overview We conduct continual pre-training of qwen-7b on 30B tokens from a mixture of Japanese and English datasets. The continual pre-training significantly improves the model's performance on Japanese tasks. It also enjoys the following great features provided by the original Q…