Flickr30K swMATH ID: 36502 Software Authors: Bryan A. Plummer, Liwei Wang, Chris M. Cervantes, Juan C. Caicedo, Julia Hockenmaier, Svetlana Lazebnik Description: The Flickr30K dataset has become a standard benchmark for sentence-based image description. This paper presents Flickr30K Entities, which augments the 158k captions from Flickr30k with 244k coreference chains, linking mentions of the same entities across different captions for the same image, and associating them with 276k manually annotated bounding boxes. Such annotations are essential for continued progress in automatic image description and grounded language understanding. They enable us to define a new benchmark for localization of textual entity mentions in an image. We present a strong baseline for this task that combines an image-text embedding, detectors for common objects, a color classifier, and a bias towards selecting larger objects. While our baseline rivals in accuracy more complex state-of-the-art models, we show that its gains cannot be easily parlayed into improvements on such tasks as image-sentence retrieval, thus underlining the limitations of current methods and the need for further research. Homepage: http://bryanplummer.com/Flickr30kEntities/ Source Code: https://github.com/BryanPlummer/flickr30k_entities Keywords: Computer Vision; Pattern Recognition; arXiv_cs.CV; arXiv_cs.CL; image description; Language; Region Phrase Correspondence; Datasets; Crowdsourcing Related Software: MS-COCO; ImageNet; BLEU; VQA; DenseCap; Faster R-CNN; Fashion-MNIST; ViLBERT; CIDEr; Im2Text; LXMERT; GloVe; PointNet; MNIST; CamNet; MVSNet; SynSin; Make3D; PIFuHD; BRISK Cited in: 6 Publications Standard Articles 1 Publication describing the Software Year Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models Bryan A. Plummer, Liwei Wang, Chris M. Cervantes, Juan C. Caicedo, Julia Hockenmaier, Svetlana Lazebnik 2015 all top 5 Cited by 18 Authors 1 Arachie, Chidubem 1 Caglayan, Ozan 1 Chanda, Bhabatosh 1 Dey, Moni Shankar 1 Haralampieva, Veneta 1 Hu, Changhua 1 Huang, Bert 1 Huang, Feicheng 1 Li, Zhixin 1 Ma, Huifang 1 Mondal, Ranjan 1 Si, XiaoSheng 1 Specia, Lucia 1 Szeliski, Richard 1 Wei, Haiyang 1 Yu, Yong 1 Zhang, Canlong 1 Zhang, Jianxun all top 5 Cited in 6 Serials 1 Machine Learning 1 Neural Computation 1 The Journal of Artificial Intelligence Research (JAIR) 1 Journal of Machine Learning Research (JMLR) 1 Mathematical Morphology. Theory and Applications 1 Texts in Computer Science Cited in 2 Fields 6 Computer science (68-XX) 1 Biology and other natural sciences (92-XX) Citations by Year