2 datasets found

Groups: Multimodal Learning

Filter Results
  • MSR-VTT

    The dataset used in the paper is MSR-VTT, a large video description dataset for bridging video and language. The dataset contains 10k video clips with length varying from 10 to...
  • Video Captioning Dataset

    A video captioning dataset generated by pseudolabeling videos with image captioning models.