LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
Overview This repository provides a Japanese GPT-NeoX model of 3.6 billion parameters. The model is based on rinna/japanese-gpt-neox-3.6b-instruction-sft-v2 and has been aligned to serve as an instruction-following conversational agent.
Apriel-1.5-15b-Thinker - Mid training is all you need!
━━━━━━━━━━━━━━━━━━━━━━━━━ Unlocking the Reasoning Potential of Language Model From Pretraining to Posttraining ━━━━━━━━━━━━━━━━━━━━━━━━━
Building the Next Generation of Open-Source and Bilingual LLMs
A GPT-4o Level MLLM for Single Image, Multi Image and High-FPS Video Understanding on Your Phone
🔥 News : The new version CogAgent-9B-20241220 has been released! Welcome to visit CogAgent GitHub and Technical Report to explore and use our latest model.
ZwZ-4B is a fine-grained multimodal perception model built upon Qwen3-VL-4B. It is trained using Region-to-Image Distillation (R2I) combined with reinforcement learning, enabling superior fine-grained visual understanding in a single forward pass — no inference-time zooming or to…
🤗 Hugging Face 🤖 ModelScope 🐙 Experience Link Coming Soon~
Fine-tuned version of HuggingFaceTB/SmolLM3-3B-Base optimized for grade school math (GSM8K benchmark).
WARNING: The checkpoints on this repo are not fully trained model. Evaluations of intermediary checkpoints and the final model will be added when conducted (see below).
This is UltraLM-65b delta weights, a chat language model trained upon UltraChat
Large Action Models (LAMs) are advanced language models designed to enhance decision-making by translating user intentions into executable actions. As the brains of AI agents , LAMs autonomously plan and execute tasks to achieve specific goals, making them invaluable for automati…