LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…
CapRL 📖 Paper 🏠 Github 🤗 CapRL Collection 🤗 Daily Paper
🤗 [LongWriter Dataset] • 💻 [Github Repo] • 📃 [LongWriter Paper]
0. TL;DR 1. Model Details 2. Training Details 3. Usage 4. Evaluation 5. Citation
- Original model: MiniCPM-1B-sft-bf16 - Model creator and fine-tuned by: ModelBest, OpenBMB, and THUNLP - Paper: link (Note: MiniCPM-S-1B is denoted as ProSparse-1B in the paper.) - Adapted LLaMA version: MiniCPM-S-1B-sft-llama-format - Adapted PowerInfer version: MiniCPM-S-1B-sf…
RedPajama-INCITE-7B-Chat was developed by Together and leaders from the open-source AI community including Ontocord.ai, ETH DS3Lab, AAI CERC, Université de Montréal, MILA - Québec AI Institute, Stanford Center for Research on Foundation Models (CRFM), Stanford Hazy Research resea…
0. TL;DR 1. Model Details 2. Training Details 3. Usage 4. Evaluation 5. Citation
The GLM family welcomes a new generation of open-source models, the GLM-4-32B-0414 series, featuring 32 billion parameters. Its performance is comparable to OpenAI's GPT series and DeepSeek's V3/R1 series, and it supports very user-friendly local deployment features. GLM-4-32B-Ba…
RedPajama-INCITE-7B-Instruct was developed by Together and leaders from the open-source AI community including Ontocord.ai, ETH DS3Lab, AAI CERC, Université de Montréal, MILA - Québec AI Institute, Stanford Center for Research on Foundation Models (CRFM), Stanford Hazy Research r…
Medical-Qwen3-Swallow-8B is a medical-domain language model based on tokyotech-llm/Qwen3-Swallow-8B-RL-v0.2. It is designed to support research and development toward safe and trustworthy AI for Japanese clinical settings.
We introduce EXAONE Deep, which exhibits superior capabilities in various reasoning tasks including math and coding benchmarks, ranging from 2.4B to 32B parameters developed and released by LG AI Research. Evaluation results show that 1) EXAONE Deep 2.4B outperforms other models…
UI2Code^N: A Visual Language Model for Test-Time Scalable Interactive UI-to-Code Generation