3 datasets found

Tags: text-to-video retrieval

Filter Results
  • Condensed Movies

    The dataset used for text-to-video retrieval and video classification tasks.
  • ActivityNet Captions

    The ActivityNet Captions is a benchmark dataset proposed for dense video captioning. There are 20K untrimmed videos in total, and each video has several annotated segments with...
  • MSR-VTT

    The dataset used in the paper is MSR-VTT, a large video description dataset for bridging video and language. The dataset contains 10k video clips with length varying from 10 to...
You can also access this registry using the API (see API Docs).