Datasets Activity Stream About Order by Relevance Name Ascending Name Descending Last Modified Go 2 datasets found Groups: Multimodal Learning Organizations: No Organization Filter Results MSR-VTT The dataset used in the paper is MSR-VTT, a large video description dataset for bridging video and language. The dataset contains 10k video clips with length varying from 10 to... Dataset JSON Video Captioning Dataset A video captioning dataset generated by pseudolabeling videos with image captioning models. Dataset JSON