Connecting Vision and Language with Localized Narratives

Jordi Pont-Tuset,Jasper Uijlings,Soravit Changpinyo,Radu Soricut,Vittorio Ferrari

Connecting Vision and Language with Localized Narratives

2019

Jordi Pont-Tuset
Jasper Uijlings
Soravit Changpinyo
Radu Soricut
Vittorio Ferrari

We propose Localized Narratives, an efficient way to collect image captions with dense visual grounding. We ask annotators to describe an image with their voice while simultaneously hovering their mouse over the region they are describing. Since the voice and the mouse pointer are synchronized, we can localize every single word in the description. This dense visual grounding takes the form of a mouse trace segment per word and is unique to our data. We annotate 500k images with Localized Narratives: the whole COCO dataset and 380k images of the Open Images dataset. We provide an extensive analysis of these annotations, which we will release early 2020. Moreover, we demonstrate the utility of our data on two applications which benefit from our mouse trace: controlled image captioning and image generation.

Keywords:

Pointer (user interface)
Narrative
Natural language processing
Pattern recognition
image generation
Closed captioning
Ask price
Computer science
Mouseover
Artificial intelligence
multimodal image

Correction
Source
Cite
Save
Machine Reading By IdeaReader

References

Citations