AI Agent Hub
Back to models
CapsFus-LLaMA logo

CapsFus-LLaMA

Open Source BAAI Released 2023-11-29
-- 13.0B params 4.1K context Open Source

About this model

CapsFus-LLaMA is an open-source text generation model released by BAAI (Beijing Academy of Artificial Intelligence) as part of the CapsFusion project (CVPR 2024). It is a fine-tuned LLaMA-2-13B model trained specifically to fuse raw web captions and synthetic captions from vision captioners into unified CapsFusion captions, replacing costly ChatGPT-based fusion at scale for multimodal pretraining data construction.

The model is initialized from LLaMA-2-13B and trained on roughly one million triplets (raw caption, synthetic caption, ChatGPT-produced fused target) for two epochs with AdamW and a cosine learning-rate schedule. In the CapsFusion paper, CapsFus-LLaMA is evaluated against ChatGPT on 100 held-out fusion cases via human judgment; it matches or exceeds ChatGPT on 80 of 100 samples while enabling large-scale caption generation over the CapsFusion-120M dataset.

Weights and distributed inference code are published on Hugging Face and GitHub for researchers building or curating image-text training data. The model is not a general-purpose chat assistant and is not reported on standard academic NLP or reasoning leaderboards; its intended use is caption refinement and fusion in vision-language model data pipelines.

Technical Specs

  • Parameters: 13.0B
  • Architecture: Transformer
  • Context Window: 4,096 tokens
  • Input Modalities: text

Hardware Requirements

  • VRAM: 26.0 GB
  • Compute: Single NVIDIA A100 40GB