LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
UI-Mate: Advancing Open-Weight Foundation GUI Agents with In-Context Demonstrations
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ Unlocking the Reasoning Potential of Language Model From Pretraining to Posttraining ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…
FastVLM: Efficient Vision Encoding for Vision Language Models
This repository provides large language models trained by SB Intuitions.
👋 Join our Discord community. 📖 Check out the GLM-4.5 technical blog , technical report , and Zhipu AI technical documentation . 📍 Use GLM-4.5 API services on Z.ai API Platform (Global) or Zhipu AI Open Platform (Mainland China) . 👉 One click to GLM-4.5 .
We opensource our Aquila2 series, now including Aquila2 , the base language models, namely Aquila2-7B and Aquila2-34B , as well as AquilaChat2 , the chat models, namely AquilaChat2-7B and AquilaChat2-34B , as well as the long-text chat models, namely AquilaChat2-7B-16k and Aquila…
We introduce DeepSeek-Prover-V2, an open-source large language model designed for formal theorem proving in Lean 4, with initialization data collected through a recursive theorem proving pipeline powered by DeepSeek-V3. The cold-start training procedure begins by prompting DeepSe…
A small ~110M parameter language model implementing the DeepSeek-V4 architecture from scratch. This is the pretrained base model — see HuggingFaceTB/nanowhale-100m for the SFT/chat version.
Towards a Recursively Self-Improving Agent for Deep Research
DeepSeek-V2.5-1210 is an upgraded version of DeepSeek-V2.5, with improvements across various capabilities: