LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
This is the model is trained using paper, M1: Towards Scalable Test-Time Compute with Mamba Reasoning Models.
🚀 Introducing ERNIE-4.5-VL-28B-A3B-Thinking: A Breakthrough in Multimodal AI
📢 DISCLAIMER : The StableLM-Base-Alpha models have been superseded. Find the latest versions in the Stable LM Collection here.
1. Model Summary 2. Use 3. Limitations 4. Training 5. Evaluation 7. Citation
💻Github Repo • 🤔Reporting Issues • 📜Technical Report
RedPajama-INCITE-Base-3B-v1 was developed by Together and leaders from the open-source AI community including Ontocord.ai, ETH DS3Lab, AAI CERC, Université de Montréal, MILA - Québec AI Institute, Stanford Center for Research on Foundation Models (CRFM), Stanford Hazy Research re…
1. Model Summary 2. Use 3. Limitations 4. Training 5. Evaluation 7. Citation
0. TL;DR 1. Model Details 2. Training Details 3. Usage 4. Evaluation 5. Citation
[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…
The GLM family welcomes new members, the GLM-4-32B-0414 series models, featuring 32 billion parameters. Its performance is comparable to OpenAI’s GPT series and DeepSeek’s V3/R1 series. It also supports very user-friendly local deployment features. GLM-4-32B-Base-0414 was pre-tra…
japanese-gpt-neox-3.6b-instruction-sft-v2