LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
- Original model: MiniCPM-1B-sft-bf16 - Model creator and fine-tuned by: ModelBest, OpenBMB, and THUNLP - Paper: link (Note: MiniCPM-S-1B is denoted as ProSparse-1B in the paper.) - Adapted LLaMA version: MiniCPM-S-1B-sft-llama-format - Adapted PowerInfer version: MiniCPM-S-1B-sf…
RedPajama-INCITE-7B-Chat was developed by Together and leaders from the open-source AI community including Ontocord.ai, ETH DS3Lab, AAI CERC, Université de Montréal, MILA - Québec AI Institute, Stanford Center for Research on Foundation Models (CRFM), Stanford Hazy Research resea…
0. TL;DR 1. Model Details 2. Training Details 3. Usage 4. Evaluation 5. Citation
The GLM family welcomes a new generation of open-source models, the GLM-4-32B-0414 series, featuring 32 billion parameters. Its performance is comparable to OpenAI's GPT series and DeepSeek's V3/R1 series, and it supports very user-friendly local deployment features. GLM-4-32B-Ba…
StableCode-Completion-Alpha-3B-4K is a 3 billion parameter decoder-only code completion model pre-trained on diverse set of programming languages that topped the stackoverflow developer survey.
This model is a mixed-precision quantized version of DeepSeek-V3.1-Terminus, with dense layer keep the FP8 quantization of the original model, while MoE layers uses INT4 weights and FP8 activation, also called W4AFP8.
RedPajama-INCITE-7B-Instruct was developed by Together and leaders from the open-source AI community including Ontocord.ai, ETH DS3Lab, AAI CERC, Université de Montréal, MILA - Québec AI Institute, Stanford Center for Research on Foundation Models (CRFM), Stanford Hazy Research r…
Medical-Qwen3-Swallow-8B is a medical-domain language model based on tokyotech-llm/Qwen3-Swallow-8B-RL-v0.2. It is designed to support research and development toward safe and trustworthy AI for Japanese clinical settings.
We introduce EXAONE Deep, which exhibits superior capabilities in various reasoning tasks including math and coding benchmarks, ranging from 2.4B to 32B parameters developed and released by LG AI Research. Evaluation results show that 1) EXAONE Deep 2.4B outperforms other models…
0. TL;DR 1. Model Details 2. Training Details 3. Usage 4. Evaluation 5. Citation
[2026-01-12]🚀🚀🚀 We have open-sourced AgentCPM-Explore , an agent foundation model with only 4B parameters , together with its entire training and inference infrastructure . AgentCPM-Explore has successfully entered 8 classic long-horizon agent benchmarks , including GAIA,HLE, and…
UI2Code^N: A Visual Language Model for Test-Time Scalable Interactive UI-to-Code Generation