LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
Model Summary: Granite-3.3-2B-Instruct is a 2-billion parameter 128K context length language model fine-tuned for improved reasoning and instruction-following capabilities. Built on top of Granite-3.3-2B-Base, the model delivers significant gains on benchmarks for measuring gener…
⚠️ DEPRECATION WARNING ⚠️ ⚠️ NOT RECOMMENDED FOR USE IN NEW PROJECTS ⚠️
Command A+ is an open source model with 25 billion active parameters and 218B total parameters model optimized for agentic, multilingual, and reasoning-heavy tasks with a focus on enterprise performance, while also providing support for vision inputs for processing image inputs.
This is a 9B model whose architecture is deepseek-v3, trained from scratch using 350B+ tokens from fully open-source, English-only datasets. It is designed for development and debugging purposes within the open-source community.
📰 Tech Blog 📄 Paper
[2025.01.14] 🔥🔥 We open source MiniCPM-o 2.6 , with significant performance improvement over MiniCPM-V 2.6 , and support real-time speech-to-speech conversation and multimodal live streaming. Try it now.
GitHub Repo Technical Report Join Us 👋 Contact us in Discord and WeChat
📰 Step3 Model Blog 📄 Step3 System Blog
Model Summary: Granite-3.0-1B-A400M-Base is a decoder-only language model to support a variety of text-to-text generation tasks. It is trained from scratch following a two-stage training strategy. In the first stage, it is trained on 8 trillion tokens sourced from diverse domains…
LFM2‑VL is Liquid AI's first series of multimodal models, designed to process text and images with variable resolutions. Built on the LFM2 backbone, it is optimized for low-latency and edge AI applications.
💻Github Repo • 🤔Reporting Issues • 📜Technical Report
[🏠Homepage] [🤖 Chat with DeepSeek LLM] [Discord] [Wechat(微信)]