LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
A GPT-4o Level MLLM for Single Image, Multi Image and High-FPS Video Understanding on Your Phone
SEA-LION is a collection of Large Language Models (LLMs) which have been pretrained and instruct-tuned for the Southeast Asia (SEA) region.
HiLS-Attention is a chunk-wise sparse attention mechanism that learns chunk selection end-to-end under the language-modeling loss, enabling native sparse training for efficient long-context modeling. This repository hosts the 7B checkpoint continued-trained on top of an OLMo3-sty…
🔥 News : The new version CogAgent-9B-20241220 has been released! Welcome to visit CogAgent GitHub and Technical Report to explore and use our latest model.
Llama-SEA-LION-v2-8B-IT SEA-LION is a collection of Large Language Models (LLMs) which have been pretrained and instruct-tuned for the Southeast Asia (SEA) region.
ZwZ-4B is a fine-grained multimodal perception model built upon Qwen3-VL-4B. It is trained using Region-to-Image Distillation (R2I) combined with reinforcement learning, enabling superior fine-grained visual understanding in a single forward pass — no inference-time zooming or to…
🤗 Hugging Face 🤖 ModelScope 🐙 Experience Link Coming Soon~
StableLM-Base-Alpha-7B-v2 is a 7 billion parameter decoder-only language model pre-trained on diverse English datasets. This model is the successor to the first StableLM-Base-Alpha-7B model, addressing previous shortcomings through the use of improved data sources and mixture rat…
Fine-tuned version of HuggingFaceTB/SmolLM3-3B-Base optimized for grade school math (GSM8K benchmark).
1. Model Summary 2. Use 3. Limitations 4. Training 5. Evaluation 7. Citation
WARNING: The checkpoints on this repo are not fully trained model. Evaluations of intermediary checkpoints and the final model will be added when conducted (see below).
Experimental frankenmerge of Kimi K2-07, 09 and Base