Dataset - LDM

COCO-QA-OBJ

The COCO-QA-OBJ dataset is used for object counting task. It consists of 123,287 images and 78,736 train and 38,948 test questions.
- Dataset
- JSON
COCO-QA-LOC

The COCO-QA-LOC dataset is used for location identification task. It consists of 123,287 images and 78,736 train and 38,948 test questions.
- Dataset
- JSON
COCO-QA-ID

The COCO-QA-ID dataset is used for color identification task. It consists of 123,287 images and 78,736 train and 38,948 test questions.
- Dataset
- JSON
COCO-QA

The COCO-QA dataset is used for visual question answering task. It consists of 123,287 images and 78,736 train and 38,948 test questions.
- Dataset
- JSON
Show and tell: A neural image caption generator

Show and tell: A neural image caption generator.
- Dataset
- JSON
From show to tell: A survey on deep learning-based image captioning

From show to tell: A survey on deep learning-based image captioning.
- Dataset
- JSON
Microsoft COCO

The Microsoft COCO dataset was used for training and evaluating the CNNs because it has become a standard benchmark for testing algorithms aimed at scene understanding and...
- Dataset
- JSON
Self-Supervised Image Captioning with CLIP

Image captioning, a fundamental task in vision-language understanding, seeks to generate accurate natural language descriptions for provided images. Current image captioning...
- Dataset
- JSON
COCO

Large scale datasets [18, 17, 27, 6] boosted text conditional image generation quality. However, in some domains it could be difficult to make such datasets and usually it could...
- Dataset
- JSON
NoCaps

NoCaps is a large benchmark dataset for image captioning, containing 30,000 image captions.
- Dataset
- JSON
MSCOCO

Human Pose Estimation (HPE) aims to estimate the position of each joint point of the human body in a given image. HPE tasks support a wide range of downstream tasks such as...
- Dataset
- JSON
MeaCap: Memory-Augmented Zero-shot Image Captioning

Zero-shot image captioning without well-paired image-text data can be divided into two categories, training-free and text-only-training. Generally, these two types of methods...
- Dataset
- JSON

You can also access this registry using the API (see API Docs).

72 datasets found