CapsFus-LLaMA
About this model
CapsFus-LLaMA is an open-source text generation model released by BAAI (Beijing Academy of Artificial Intelligence) as part of the CapsFusion project (CVPR 2024). It is a fine-tuned LLaMA-2-13B model trained specifically to fuse raw web captions and synthetic captions from vision captioners into unified CapsFusion captions, replacing costly ChatGPT-based fusion at scale for multimodal pretraining data construction.
The model is initialized from LLaMA-2-13B and trained on roughly one million triplets (raw caption, synthetic caption, ChatGPT-produced fused target) for two epochs with AdamW and a cosine learning-rate schedule. In the CapsFusion paper, CapsFus-LLaMA is evaluated against ChatGPT on 100 held-out fusion cases via human judgment; it matches or exceeds ChatGPT on 80 of 100 samples while enabling large-scale caption generation over the CapsFusion-120M dataset.
Weights and distributed inference code are published on Hugging Face and GitHub for researchers building or curating image-text training data. The model is not a general-purpose chat assistant and is not reported on standard academic NLP or reasoning leaderboards; its intended use is caption refinement and fusion in vision-language model data pipelines.
Technical Specs
- Parameters: 13.0B
- Architecture: Transformer
- Context Window: 4,096 tokens
- Input Modalities: text
Hardware Requirements
- VRAM: 26.0 GB
- Compute: Single NVIDIA A100 40GB