Redirigiendo al acceso original de articulo en 17 segundos...
Inicio  /  Applied Sciences  /  Vol: 9 Par: 13 (2019)  /  Artículo
ARTÍCULO
TITULO

Variational Autoencoder-Based Multiple Image Captioning Using a Caption Attention Map

Boeun Kim    
Saim Shin and Hyedong Jung    

Resumen

Image captioning is a promising research topic that is applicable to services that search for desired content in a large amount of video data and a situation explanation service for visually impaired people. Previous research on image captioning has been focused on generating one caption per image. However, to increase usability in applications, it is necessary to generate several different captions that contain various representations for an image. We propose a method to generate multiple captions using a variational autoencoder, which is one of the generative models. Because an image feature plays an important role when generating captions, a method to extract a Caption Attention Map (CAM) of the image is proposed, and CAMs are projected to a latent distribution. In addition, methods for the evaluation of multiple image captioning tasks are proposed that have not yet been actively researched. The proposed model outperforms in the aspect of diversity compared with the base model when the accuracy is comparable. Moreover, it is verified that the model using CAM generates detailed captions describing various content in the image.

 Artículos similares

       
 
Jianying Li and Xiangjun Shao    
Image captioning is a challenging task, which generates a sentence for a given image. The earlier captioning methods mainly decode the visual features to generate caption sentences for the image. However, the visual features lack the context semantic inf... ver más
Revista: Information

 
Oscar Ondeng, Heywood Ouma and Peter Akuon    
Visual understanding is a research area that bridges the gap between computer vision and natural language processing. Image captioning is a visual understanding task in which natural language descriptions of images are automatically generated using visio... ver más
Revista: Applied Sciences

 
Fariba Lotfi, Amin Beheshti, Helia Farhood, Matineh Pooshideh, Mansour Jamzad and Hamid Beigy    
In our digital age, data are generated constantly from public and private sources, social media platforms, and the Internet of Things. A significant portion of this information comes in the form of unstructured images and videos, such as the 95 million d... ver más
Revista: Algorithms

 
Viktar Atliha and Dmitrij ?e?ok    
Image captioning is a very important task, which is on the edge between natural language processing (NLP) and computer vision (CV). The current quality of the captioning models allows them to be used for practical tasks, but they require both large compu... ver más
Revista: Applied Sciences

 
Deepika Kumar, Varun Srivastava, Daniela Elena Popescu and Jude D. Hemanth    
Image captioning is oriented towards describing an image with the best possible use of words that can provide a semantic, relatable meaning of the scenario inscribed. Different models can be used to accomplish this arduous task depending on the context a... ver más
Revista: Applied Sciences