Image Captioning Algorithm Based on Sufficient Visual Information and Text Information

Yongqiang Zhao,Rao Yuan,Wu Lianwei,Cong Feng

Image Captioning Algorithm Based on Sufficient Visual Information and Text Information

2020

Most existing attention-based methods on image captioning focus on the current visual information and text information at each step to generate the next word, without considering the coherence between the visual information and the text information itself. We propose sufficient visual information (SVI) module to supplement the existing visual information contained in the network, and propose sufficient text information (STI) module to predict more text Words to supplement the text information contained in the network. Sufficient visual information module embed the attention value from the past two steps into the current attention to adapt to human visual coherence. Sufficient text information module can predict the next three words in one step, and jointly use their probabilities for inference. Finally, this paper combines these two modules to form an image captioning algorithm based on sufficient visual information and text information model (SVITI) to further integrate existing visual information and future text information in the network, thereby improving the image captioning performance of the model. These three methods are used in the classic image captioning algorithm, and have achieved achieve significant performance improvement compared to the latest method on the MS COCO dataset.

Keywords:

Correction
Source
Cite
Save
Machine Reading By IdeaReader

References

Citations