1 dataset found

Tags: image-audio retrieval

Filter Results
  • Places

    The dataset used in the paper is Places, a large dataset of 400k pairs of images from the Places 205 dataset and corresponding spoken audio captions.
You can also access this registry using the API (see API Docs).