AI Agent Hub
Back to models
Nemotron-SEA-LION-v4.8-120B-A12B logo

Nemotron-SEA-LION-v4.8-120B-A12B

Open Source aisingapore Released 2026-09-18
-- 120.0B params 262.1K context Open Source

About this model

Banner!

Technical Report 👁️

Nemotron-SEA-LION-v4.8-120B-A12B

Last updated: 2026-09-18

SEA-LION is a collection of Large Language Models (LLMs) which have been pretrained and instruct-tuned for the Southeast Asian (SEA) region. Nemotron-SEA-LION-v4.8-120B-A12B is built upon the aisingapore/Nemotron-SEA-LION-v4.8-120B-A12B-Base, by fine-tuning with a mixture of SFT and online on-policy distillation (OPD).

Model Details

Model Description

SEA-LION stands for Southeast Asian Languages In One Network.

Nemotron-SEA-LION-v4.8-120B-A12B is post-trained on English and the 7 SEA languages, across various tasks.

For tokenization, the model employs the default tokenizer used in nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16.

  • Developed by: AI Products Pillar, AI Singapore
  • Funded by: National Research Foundation Singapore
  • Shared by: AI Products Pillar, AI Singapore
  • Model type: Instruction-tuned language model
  • Architecture: Mamba2-Transformer Hybrid MoE
  • Context length: 262,144 tokens
  • Language(s): Burmese, English, Indonesian, Filipino, Malay, Tamil, Thai, and Vietnamese
  • License: MIT
  • Parent model: aisingapore/Nemotron-SEA-LION-v4.8-120B-A12B-Base

Model Sources

Usage

Transformers

Use the code below to get started with the model with 🤗 Transformers libraries.

from transformers import AutoTokenizer, AutoModelForCausalLM
import torch

model_name = "aisingapore/Nemotron-SEA-LION-v4.8-120B-A12B"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
    model_name,
    torch_dtype=torch.bfloat16,
    trust_remote_code=True,
    device_map="auto"
)

messages = [
    {"role": "user", "content": "What is Nasi goreng?"},
]

tokenized_chat = tokenizer.apply_chat_template(
    messages,
    tokenize=True,
    add_generation_prompt=True,
    return_tensors="pt"
).to(model.device)

if not isinstance(tokenized_chat, torch.Tensor):
    input_ids = tokenized_chat["input_ids"]
else:
    input_ids = tokenized_chat

outputs = model.generate(
    input_ids,
    max_new_tokens=50,
    temperature=1.0,
    top_p=0.95,
    eos_token_id=tokenizer.eos_token_id
)

print(tokenizer.decode(outputs[0]))

Training Details

Training Data

The training data is a subset of aisingapore/SEA-Instruct-2602.

Evaluation

Testing Data, Factors & Metrics:

We evaluated Nemotron-SEA-LION-v4.8 on general language, multi-turn chat and instruction-following capabilities. For the evaluation used in SEA-LION-v4.8, the new version of SEA-HELM (Southeast Asian Holistic Evaluation of Language Models) is adopted as the evaluation framework*.

Testing Data

General language capabilities:

For the evaluation of general language capabilities, we employed the SEA-HELM evaluation benchmark across a variety of tasks.

These tasks include Question Answering (QA), Sentiment Analysis (Sentiment), Metaphor Understanding, Toxicity Detection (Toxicity), SEA-SafeguardBench, Translation in both directions (Eng>Lang & Lang>Eng), Abstractive Summarisation (Abssum), Causal Reasoning (Causal), Natural Language Inference (NLI), Linguistic Diagnostics (LINDSEA), SEA-NLI, Cultural Knowledge (Kalahi), Thai Exam, and Global MMLU Lite. Language coverage within SEA-HELM includes Filipino, Indonesian, Tamil, Thai, and Vietnamese, and is now further extended to Malay and Burmese.

Instruction-following and Multi-turn Chat:

We evaluated the models on instruction-following and multi-turn chat capabilities with SEA-IFEval (based on IFEval) and SEA-MTBench (based on MT-Bench) respectively. The two datasets were originally in English, the linguists and native speakers in the team worked together to filter, localise and translate the datasets into the respective target languages to ensure that the examples remained reasonable, meaningful and natural, so as to avoid the translationese and cultural erasure associated with machine-translated benchmarks, native speakers of the target languages participate at each stage of dataset planning and construction.

Factors

All evaluations ran with the model-specific generation parameters defined in the model config. Each evaluation comprised of 8 runs with different seeds and the average score for each prompt was then calculated.

Bootstrapping with replacement for each prompt was done to obtain the task score. The final result is then the aggregation of the various task scores.

For all tasks, the model was expected to provide an answer tag from which the answer was automatically extracted. For tasks where options were provided, the answer should comprise one of the pre-defined options.

The evaluation was done zero-shot with native prompts on a sample of 100-1000 instances for each dataset. Task prompts are issued in the target language itself, on the view that full support for a language requires a model both to interpret native instructions and to respond coherently in that language.

  • SEA-IFEval: SEA-IFEval evaluates a model's ability to adhere to constraints provided in the prompt, for example, beginning a response with a specific word/phrase or answering with a certain number of sections. Additionally, accuracy is normalised by the proportion of responses in the correct language (if the model performs the task correctly but responds in the wrong language, it is judged to have failed the task).

  • SEA-MTBench: SEA-MTBench evaluates a model's ability to engage in multi-turn (2 turns) conversations and respond in ways that align with human needs. We use gpt-oss-120b as the judge model. The metric used is the average criteria score where model responses were judged based on a set of specific criteria for each prompt.

Metrics

The following metrics were used for text capabilities:

Task Metric
Sentiment Analysis Accuracy
Extractive QA (ID, VI, TH, TA) ChrF++
MCQ-QA (TL, MY, MS) Accuracy
Metaphor Accuracy
Abstractive Summarisation Rouge-L
Translations MetricX-24 score (with reference)
Toxicity Detection Accuracy
SEA-Safeguard Accuracy
Causal Reasoning Accuracy
Natural Language Inference Accuracy
LINDSEA Accuracy
LINDSEA syntax (LLM-as-a-Judge) Average criteria score
Global MMLU Lite Accuracy
ThaiExam Accuracy
Kalahi Accuracy
Kalahi (LLM-as-a-Judge) Average criteria score
SEA-NLI Accuracy
SEA-IFEval Accuracy
SEA-MTBench (LLM-as-a-Judge) Average criteria score

Results

For details on Nemotron-SEA-LION-v4.8 performance, please refer to the SEA-HELM Leaderboard.

Environmental Impact

  • Hardware type: H200
  • GPU-hours: approximately 2688
  • Cloud provider: SMC H200
  • Compute region: Singapore
  • Carbon emissions: approximately 0.073 - 1.468 MT

Technical Specifications

Technical Report

For training details, see the SEA-LION-v4.8 Technical Report.

Model Architecture

The architecture is based on the highly efficient Nemotron-3-Super foundation. The detailed architecture can be found at nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16 documentation.

Uses

Out-of-Scope Use

The model has not been aligned for safety. Developers and users should perform their own safety fine-tuning and related security measures. In no event shall the authors be held liable for any claims, damages, or other liabilities arising from the use of the released weights and codes.

Bias, Risks, and Limitations

The model was not tested for robustness against adversarial prompting. It is important for users to be aware that our model exhibits certain limitations that warrant consideration. Like many LLMs, the model can hallucinate and occasionally generates irrelevant content, introducing fictional elements that are not grounded in the provided context. Users should also exercise caution in interpreting and validating the model's responses due to the potential inconsistencies.

Citation

BibTeX:

@misc{aisingapore2026sealionv48technicalreport,
      title={SEA-LION-v4.8: A Technical Report},
      author={Adila Aulia and Ahmed Dabeer and Ahn Jeongmi and Antonyrex Sajeban and Chan Hok Teng Adwin and Cheng Zi Yi Nicholas and Choa Hsueh Mei Esther and Heng Jonathan and Jann Railey Estrada Montalan and Lee Chwan Ren and Leong Wai Yi and Leong Wei Qi and Liew Rachel and Limkonchotiwat Peerat and Muhammad Ridzuan Bin Mokhtar and Nagarajan Karthik and Ng Boon Cheong Raymond and Ngee Chia Tai and Ngui Jian Gang and Nguyen Thanh Ngan and Ong Tat-Wee David and Pereira Mark and Phang Shi Wei Benjamin and Poon Joseph and Rengarajan Hamsawardhini and Susanto Yosephine and Sutaveephamochanon Anocha and Tan Choon Meng and Tan Chor Phin Evelyn and Tan Le Min Sheryl and Tan Siao Wei Jessica and Tan Yixian and Tasawong Panuthep and Tee Jun Yun and Teng Kok Wai Walter and Teo Eng Sipp Leslie and Tjhi William and Tuchinda Pume and Wu Donghang and Yong Xianbin and Zhang Zhou},
      year={2026},
      eprint={2609.18310},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2609.18310},
}

Team

AI Products Pillar, AI Singapore

Acknowledgement

This project is supported by the National Research Foundation Singapore and Infocomm Media Development Authority (IMDA), Singapore under its National Large Language Model Funding Initiative.

Contact

sealion@aisingapore.org

Technical Specs

  • Parameters: 120.0B
  • Architecture: Mamba2-Transformer Hybrid MoE
  • Context Window: 262,144 tokens
  • Input Modalities: text

Hardware Requirements

  • VRAM: 80.4 GB
  • Compute: 2x NVIDIA H100/H200 80GB (NVFP4); BF16 requires 4x H100/H200