LatteGAN: Visually Guided Language Attention for Multi-Turn Text-Conditioned Image Manipulation

Text-guided image manipulation tasks have recently gained attention in the vision-and-language community. The GeNeVA task is a multi-turn text-conditioned image generation (MTIM) task. It involves two participants: a Teller that instructs how to modify the image, and a Drawer that draws the image according to the Teller’s instructions.

BibTex: