Visual image caption generation for service robotics and industrial applications
Journal
Proceedings - 2019 IEEE International Conference on Industrial Cyber Physical Systems, ICPS 2019
Pages
827-832
Date Issued
2019
Author(s)
Abstract
Image caption generation is a task that generates a sentence from a raw image, which is mimicking the intelligence of human that can acquire knowledge from the view. The difficulty of this task is the combination of multimodal knowledge learning, i.e. recognition of objects, actions, scenes, human, etc. In order to perform semantic understanding for service robotics or other industrial applications, the caption must be enhanced for recognition of the objects in the confined environment. We propose a template-based augmentation method for improving the capability of object recognition while retaining the other capability of the image caption model. This work opens a new era of image caption generation training procedure that the caption dataset and the classification dataset can be combined to train the deep captioning model. We show in our experiments that our improved model outperforms the original model in SPICE metrics by 4 times.
SDGs
Type
conference paper
