LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
AHN: Artificial Hippocampus Networks for Efficient Long-Context Modeling
Llama 3 Youko 70B (rinna/llama-3-youko-70b)
Our Swallow model has undergone continual pre-training from the Llama 2 family, primarily with the addition of Japanese language data. The tuned versions use supervised fine-tuning (SFT). Links to other models can be found in the index.
We introduce an updated version of LongCat-Flash-Thinking-2601, a powerful and efficient Large Reasoning Model (LRM) with 560 billion total parameters, built upon an innovative Mixture-of-Experts (MoE) architecture. Beyond inheriting the domain-parallel training recipe in our pre…
AHN: Artificial Hippocampus Networks for Efficient Long-Context Modeling
AHN: Artificial Hippocampus Networks for Efficient Long-Context Modeling
[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.
Gemma-2-Llama-Swallow series was built by continual pre-training on the gemma-2 models. Gemma 2 Swallow enhanced the Japanese language capabilities of the original Gemma 2 while retaining the English language capabilities. We use approximately 200 billion tokens that were sampled…
This repository provides Japanese language models trained by SB Intuitions.
[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.
This repository provides Japanese language models trained by SB Intuitions.