-
Self-Supervised Image Captioning with CLIP
Image captioning, a fundamental task in vision-language understanding, seeks to generate accurate natural language descriptions for provided images. Current image captioning... -
MeaCap: Memory-Augmented Zero-shot Image Captioning
Zero-shot image captioning without well-paired image-text data can be divided into two categories, training-free and text-only-training. Generally, these two types of methods...