We propose a video captioning model to generate a caption for a short video clip. The model includes vision (green) and textual (blue) branches to benefit video captioning by both video and text inputs. We release the checkpoint trained on Panda-70M.

xiankgx
/
panda-70m-video-captioning
Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers