LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…
Qwen3-Swallow v0.2 is a family of large language models available in 8B , 30B-A3B , and 32B parameter sizes. Built as bilingual Japanese-English models, they were developed through Continual Pre-Training (CPT), Supervised Fine-Tuning (SFT), and Reinforcement Learning with Verifia…
1. Model Summary 2. Evaluation 3. Intended Use 4. Limitations 5. Security and Responsible Use 6. License 7. Citation
Ling is a MoE LLM provided and open-sourced by InclusionAI. We introduce two different sizes, which are Ling-Lite and Ling-Plus. Ling-Lite has 16.8 billion parameters with 2.75 billion activated parameters, while Ling-Plus has 290 billion parameters with 28.8 billion activated pa…
GPT-OSS-Swallow v0.1 is a family of large language models available in 20B and 120B parameter sizes. Built as bilingual Japanese-English models, they were developed through Continual Pre-Training (CPT), Supervised Fine-Tuning (SFT), and Reinforcement Learning with Verifiable Rewa…
Qwen3-Swallow v0.2 is a family of large language models available in 8B , 30B-A3B , and 32B parameter sizes. Built as bilingual Japanese-English models, they were developed through Continual Pre-Training (CPT), Supervised Fine-Tuning (SFT), and Reinforcement Learning with Verifia…
[🏠Homepage] [🤖 Chat with DeepSeek Coder] [Discord] [Wechat(微信)]
This repository contains the model presented in Baichuan-M2: Scaling Medical Capability with Large Verifier System.
BFS-Prover-V2: Scaling up Multi-Turn Off-Policy RL and Multi-Agent Tree Search for LLM Step-Provers
This is a Japanese continually pre-trained model based on mistralai/Mistral-Nemo-Instruct-2407.
SingGuard-NSFA: Extensible Guardrails for Agentic AI via Generative Reasoning and Real-Time Classification
LLaDA2.2-flash is an agent-oriented diffusion language model in the LLaDA2 series. By introducing Levenshtein Editing (with DELETE and INSERT control tokens) to diffusion language modeling, it represents the LLaDA2 series' first step in agentic applications, including long-contex…