85 datasets found

Formats: JSON

Filter Results
  • Self-Supervised Image Captioning with CLIP

    Image captioning, a fundamental task in vision-language understanding, seeks to generate accurate natural language descriptions for provided images. Current image captioning...
  • COCO

    Large scale datasets [18, 17, 27, 6] boosted text conditional image generation quality. However, in some domains it could be difficult to make such datasets and usually it could...
  • NoCaps

    NoCaps is a large benchmark dataset for image captioning, containing 30,000 image captions.
  • MSCOCO

    Human Pose Estimation (HPE) aims to estimate the position of each joint point of the human body in a given image. HPE tasks support a wide range of downstream tasks such as...
  • MeaCap: Memory-Augmented Zero-shot Image Captioning

    Zero-shot image captioning without well-paired image-text data can be divided into two categories, training-free and text-only-training. Generally, these two types of methods...