LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…
[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…
This model is a mixed-precision quantized version of DeepSeek-V3.1-Terminus, with dense layer keep the FP8 quantization of the original model, while MoE layers uses INT4 weights and FP8 activation, also called W4AFP8.
[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.
Hy-Embodied-VLM-1.0 Efficient Physical-World Agents Tencent Robotics X × Hy Vision Team × Futian Laboratory
🤗 HuggingFace 📰 Blog 🎨 Xiaomi MiMo API Platform 🗨️ Xiaomi MiMo Studio
This model is based on the principles described in the paper Large Language Diffusion Models.
🤗 HuggingFace 📔 Technical Report
  GITHUB      🖥️   official website   |  🕖   HunyuanAPI |  🐳   Gitee Technical Report   |   Demo    |   Tencent Cloud TI    
Introduction We are thrilled to introduce Stable-DiffCoder, which is a strong code diffusion large language model. Built directly on the Seed-Coder architecture, data, and training pipeline, it introduces a block diffusion continual pretraining (CPT) stage with a tailored warmup…
🤗 HuggingFace 🤖 ModelScope 🪡 AngelSlim
🤗 HuggingFace 🤖 ModelScope 🪡 AngelSlim