OTel-2.0-LLM-31B-IT
About this model
OTel-2.0-LLM-31B-IT is an open-source, Apache 2.0-licensed instruction-tuned language model developed by Farbod Tavakkoli and collaborators across AT&T, GSMA, MLCommons, Microsoft, AMD, Dell, Red Hat, and Pleias. It is post-trained from Google's Gemma 4 31B-IT on approximately 440 billion telecom-domain tokens, sourced from a ~15B-token raw corpus covering standards from 3GPP, ETSI, ITU, GSMA, CAMARA, O-RAN, and TM Forum. The model targets telco-grade workflows including RAG over technical documentation, standards interpretation, network operations support, product development, and telecom-specific Q&A.
Released in July 2026, OTel 2.0 is the flagship model in the Open Telco AI (OTel 2.0) initiative and is positioned as the top-performing open-source model on the GSMA Open Telco AI leaderboard, with claimed average scores of ~90.3% across seven telecom benchmarks (TeleQnA, TeleTables, TeleMath, TeleLogs, 3GPP-TSG, ORANBench, srsRANBench). GSMA and AT&T report it outperforms general-purpose frontier models on domain-specific telecom tasks while remaining deployable on right-sized on-premises hardware. Day-0 inference is available via Microsoft Foundry, Featherless AI, and Red Hat, with weights hosted on Hugging Face.
General-purpose benchmarks (MMLU, HumanEval, GSM8K) have not been independently published for OTel 2.0; scores reflect the base Gemma 4 31B-IT model from Google's official model card (MMLU Pro 85.2%, LiveCodeBench v6 80.0%). Independent verification of the public checkpoint (TelcoAIBench, August 2026) measured lower telecom benchmark scores (~62.5% average) than leaderboard claims, noting the model receives weekly weight updates and checkpoints should be pinned for reproducible evaluation. The model supports English text only and is not multimodal.
Benchmark Scores
Technical Specs
- Parameters: 31.0B
- Architecture: Transformer (Gemma 4 hybrid sliding-window and global attention)
- Context Window: 256,000 tokens
- Input Modalities: text
Hardware Requirements
- VRAM: 70.0 GB
- Compute: FP16 inference needs ~70 GB VRAM (e.g. A100 80GB); INT4 quantization runs on ~17–24 GB (RTX 4090). Training used on-prem AMD MI355X GPUs with Dell infrastructure; data processing used ~530 AMD MI300X GPUs on Microsoft Azure.