Javascript must be enabled to continue!
Image caption extraction to aid visual learning
View through CrossRef
An image caption generator is essential for social media enthusiasts or visually impaired individuals. It can be used as a plugin in popular social media platforms to recommend suitable captions or to assist visually impaired people in comprehending the image content on the web, thereby eliminating ambiguity in image meaning and ensuring accurate knowledge acquisition. This research describes an image caption generator that utilizes a Convolutional Neural Network (CNN) and a Long Short-Term Memory (LSTM) model to generate natural language descriptions of images. The CNN was employed to extract features from the input image, which were then fed into the LSTM to generate the corresponding caption. The model is trained on a large dataset of image-caption pairs, using a combination of supervised and reinforcement learning techniques. The model's performance is evaluated using several metrics, and the results demonstrate that the proposed CNN LSTM model outperforms existing state-of-the-art approaches in generating accurate and diverse image captions. This model has the potential to be used in various applications, including image retrieval, content-based image search, and assisting visually impaired individuals to understanding their surroundings. It also discusses about the structure and functions of the various neural networks involved.
i-manager Publications
Title: Image caption extraction to aid visual learning
Description:
An image caption generator is essential for social media enthusiasts or visually impaired individuals.
It can be used as a plugin in popular social media platforms to recommend suitable captions or to assist visually impaired people in comprehending the image content on the web, thereby eliminating ambiguity in image meaning and ensuring accurate knowledge acquisition.
This research describes an image caption generator that utilizes a Convolutional Neural Network (CNN) and a Long Short-Term Memory (LSTM) model to generate natural language descriptions of images.
The CNN was employed to extract features from the input image, which were then fed into the LSTM to generate the corresponding caption.
The model is trained on a large dataset of image-caption pairs, using a combination of supervised and reinforcement learning techniques.
The model's performance is evaluated using several metrics, and the results demonstrate that the proposed CNN LSTM model outperforms existing state-of-the-art approaches in generating accurate and diverse image captions.
This model has the potential to be used in various applications, including image retrieval, content-based image search, and assisting visually impaired individuals to understanding their surroundings.
It also discusses about the structure and functions of the various neural networks involved.
Related Results
Foreign aid mix and manufactured exports performance in sub-Saharan Africa
Foreign aid mix and manufactured exports performance in sub-Saharan Africa
This study aims at finding out effects of foreign aid mix on manufactured exports performance in Sub-Saharan Africa. This is important as the region has lagged behind on promotion ...
CREATING LEARNING MEDIA IN TEACHING ENGLISH AT SMP MUHAMMADIYAH 2 PAGELARAN ACADEMIC YEAR 2020/2021
CREATING LEARNING MEDIA IN TEACHING ENGLISH AT SMP MUHAMMADIYAH 2 PAGELARAN ACADEMIC YEAR 2020/2021
The pandemic Covid-19 currently demands teachers to be able to use technology in teaching and learning process. But in reality there are still many teachers who have not been able ...
Double Exposure
Double Exposure
I. Happy Endings
Chaplin’s Modern Times features one of the most subtly strange endings in Hollywood history. It concludes with the Tramp (Chaplin) and the Gamin (Paulette Godda...
A Neuro-symbolic Framework for Hallucination Detection in Image Captioning Using Knuth– Morris–Pratt Pattern Matching
A Neuro-symbolic Framework for Hallucination Detection in Image Captioning Using Knuth– Morris–Pratt Pattern Matching
The research on image captioning has improved significantly in terms of describe of a scene appears on an image. Thisparticular task of image caption generation is considered as mo...
Image Caption Generator using Deep Learning Approach
Image Caption Generator using Deep Learning Approach
The creation of captions for a picture is the focus of the image caption generator. The image's semantic meaning is extracted and translated into plain language. Also, there are bu...
Strengthening the Effectiveness of Aid Delivery in Teacher Education: A Fiji Case Study
Strengthening the Effectiveness of Aid Delivery in Teacher Education: A Fiji Case Study
<p dir="ltr">As a result of increasing development challenges and higher aid allocations to the Pacific, questions of aid effectiveness have become increasingly important. Ef...
An Essential Image Augmentation Processes for Pattern Based Image Retrieval System
An Essential Image Augmentation Processes for Pattern Based Image Retrieval System
An image retrieval system is an image search engine is most useful for human day to day life. But still, the image retrieval systems are working in the traditional manner. Nowadays...
Pengaruh Tri-n-Oktil Posfin Oksida dan Tingkat Ekstraksi pada Pemurnian Konsentrat Thorium
Pengaruh Tri-n-Oktil Posfin Oksida dan Tingkat Ekstraksi pada Pemurnian Konsentrat Thorium
Telah dilakukan ekstraksi konsentrat thorium oksalat hasil olah monasit memakai ekstraktan Tri – n - Oktil Posfin Oksida (TOPO). Pengotor yang paling banyak terkandung dalam kon...

