LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
SEA-LION is a collection of Large Language Models (LLMs) which have been pretrained and instruct-tuned for the Southeast Asia (SEA) region.
This is trained on the Yi-34B model with 200K context length, for 3 epochs on the Capybara dataset!
Overview We conduct continual pre-training of llama2-7b on 40B tokens from a mixture of Japanese and English datasets. The continual pre-training significantly improves the model's performance on Japanese tasks.
Llama 3.1 Swallow is a series of large language models (8B, 70B) that were built by continual pre-training on the Meta Llama 3.1 models. Llama 3.1 Swallow enhanced the Japanese language capabilities of the original Llama 3.1 while retaining the English language capabilities. We u…
Step 3.5 Flash (visit website) is our most capable open-source foundation model, engineered to deliver frontier reasoning and agentic capabilities with exceptional efficiency. We also open-sourced the training codebase (SteptronOss), with support for continue pretrain, SFT, RL (W…
Model Card for Gemma-SEA-LION-v4-27B-IT-NVFP4
Authors: Erik Nijkamp\ , Hiroaki Hayashi\ , Yingbo Zhou, Caiming Xiong
🤗 Hugging Face 🤖 ModelScope 🐙 Experience Now
Medical-Qwen3-Swallow-30B-A3B is a medical-domain language model based on tokyotech-llm/Qwen3-Swallow-30B-A3B-RL-v0.2 . It is designed to support research and development toward safe and trustworthy AI for Japanese clinical settings.
💻Github Repo • 🤗HF Model Collections • 🤖ModelScope Collections • 💬Online Chat
Developed by : Upstage Backbone Model : LLaMA-2 Language(s) : English Library : HuggingFace Transformers License : Fine-tuned checkpoints is licensed under the Non-Commercial Creative Commons license (CC BY-NC-4.0) Where to send comments : Instructions on how to provide feedback…
🤗 Hugging Face      🤖 ModelScope 🐙 Experience Now