LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
StableCode-Completion-Alpha-3B-4K is a 3 billion parameter decoder-only code completion model pre-trained on diverse set of programming languages that topped the stackoverflow developer survey.
RedPajama-INCITE-7B-Instruct was developed by Together and leaders from the open-source AI community including Ontocord.ai, ETH DS3Lab, AAI CERC, Université de Montréal, MILA - Québec AI Institute, Stanford Center for Research on Foundation Models (CRFM), Stanford Hazy Research r…
Medical-Qwen3-Swallow-8B is a medical-domain language model based on tokyotech-llm/Qwen3-Swallow-8B-RL-v0.2. It is designed to support research and development toward safe and trustworthy AI for Japanese clinical settings.
We introduce EXAONE Deep, which exhibits superior capabilities in various reasoning tasks including math and coding benchmarks, ranging from 2.4B to 32B parameters developed and released by LG AI Research. Evaluation results show that 1) EXAONE Deep 2.4B outperforms other models…
0. TL;DR 1. Model Details 2. Training Details 3. Usage 4. Evaluation 5. Citation
UI2Code^N: A Visual Language Model for Test-Time Scalable Interactive UI-to-Code Generation
Our Swallow model has undergone continual pre-training from the Llama 2 family, primarily with the addition of Japanese language data. The tuned versions use supervised fine-tuning (SFT). Links to other models can be found in the index.
Sachin Mehta, Mohammad Hossein Sekhavat, Qingqing Cao, Maxwell Horton, Yanzi Jin, Chenfan Sun, Iman Mirzadeh, Mahyar Najibi, Dmitry Belenko, Peter Zatloukal, Mohammad Rastegari
SEA-LION is a collection of Large Language Models (LLMs) which have been pretrained and instruct-tuned for the Southeast Asia (SEA) region.
--- inference: false language: - en tags: - instruction-finetuning pretty name: JudgeLM-100K task categories: - text-generation ---
From Inquiry to Decision: Building Trustworthy Medical AI
OLMo-Bitnet-1B is a 1B parameter model trained using the method described in The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits.