LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
GPT-OSS-Swallow v0.1 is a family of large language models available in 20B and 120B parameter sizes. Built as bilingual Japanese-English models, they were developed through Continual Pre-Training (CPT), Supervised Fine-Tuning (SFT), and Reinforcement Learning with Verifiable Rewa…
WARNING: The checkpoints on this repo are not fully trained model. Evaluations of intermediary checkpoints and the final model will be added when conducted (see below).
We introduce and release LongCat-Flash-Thinking , which is a powerful and efficient large reasoning model (LRM) with 560 billion total parameters, featuring an innovative Mixture-of-Experts (MoE) architecture. The model incorporates a dynamic computation mechanism that activates…
A 15B-parameter token-mixer supernet with 8 optimized deployment presets spanning 1.0× to 10.7× decode throughput at 32K sequence length, all from a single checkpoint. Derived from Apriel-1.6 through stochastic distillation and targeted supervised fine-tuning.
This model card introduces a moderation model, a GPT-JT model fine-tuned on Ontocord.ai's OIG-moderation dataset v0.1.
StepFun-Prover-Preview-32B is a theorem proving model developed by StepFun Team. It can iteratively refine the proof sketch via interacting with Lean4, and achieve 70.0% accuracy with Pass@1 on MiniF2F-test. Advanced usage examples can be seen in github.
💻Github Repo • 🤗Model Collections • 📜Technical Report
💻Github Repo • 🤗HF Model Collections • 🤖ModelScope Collections • 💬Online Chat
StepFun-Prover-Preview-7B is a theorem proving model developed by StepFun Team. It can iteratively refine the proof sketch via interacting with Lean4, and achieve 66.0% accuracy with Pass@1 on MiniF2F-test. Advanced usage examples can be seen in github.
- Arxiv - Github - Model Collection - Data
State-of-the-art bilingual open-sourced Math reasoning LLMs. A solver , prover , verifier , augmentor .
Our Swallow model has undergone continual pre-training from the Llama 2 family, primarily with the addition of Japanese language data. The tuned versions use supervised fine-tuning (SFT). Links to other models can be found in the index.