Search engine for discovering works of Art, research articles, and books related to Art and Culture
ShareThis
Javascript must be enabled to continue!

Multi-Modal Amharic Movie Genre Classification Using a Deep Learning Approach

View through CrossRef
The rapid growth of Ethiopia’s film industry, particularly in Amharic cinema, has created a demand for an intelligent model that can automatically classify movies by genre. Traditional manual methods are inefficient, inconsistent, and not scalable. Although various studies have explored movie genre classification using machine learning and deep learning on multimodal data, none have focused on Amharic films. This study proposes a novel deep learning-based multimodal model that combines video frames and audio data from Amharic movie trailers. It uses 10,152 video frame sequences and 13,169 audio chunks for training. Preprocessing includes Wiener filtering and spectral subtraction for audio enhancement, and CLAHE and Gaussian filtering for frame improvement. Frame features were extracted using a pre-trained I3D ConvNet, while audio features combined handcrafted descriptors (MFCC, ZCR, Chroma, Spectral Roll-off, and RMSE) with deep features from a BiLSTM network. Different neural network architectures, I3D, CNN, BiLSTM, and CNN-BiLSTM, were trained for action, comedy, drama, and romance genres using a multiclass classification approach and early fusion for feature integration. The models achieved accuracies of 89% (I3D), 77% (CNN), 82% (BiLSTM), and 92% (CNN-BiLSTM).In summary, the CNN-BiLSTM with early fusion outperformed unimodal models, demonstrating the effectiveness of combining visual and audio features for Amharic movie genre classification, thereby enhancing film organization, retrieval, and recommendation systems. The study demonstrates the effectiveness of multimodal fusion for low-resource languages and contributes to content organization, recommendation, and accessibility in Ethiopian cinema.
Title: Multi-Modal Amharic Movie Genre Classification Using a Deep Learning Approach
Description:
The rapid growth of Ethiopia’s film industry, particularly in Amharic cinema, has created a demand for an intelligent model that can automatically classify movies by genre.
Traditional manual methods are inefficient, inconsistent, and not scalable.
Although various studies have explored movie genre classification using machine learning and deep learning on multimodal data, none have focused on Amharic films.
This study proposes a novel deep learning-based multimodal model that combines video frames and audio data from Amharic movie trailers.
It uses 10,152 video frame sequences and 13,169 audio chunks for training.
Preprocessing includes Wiener filtering and spectral subtraction for audio enhancement, and CLAHE and Gaussian filtering for frame improvement.
Frame features were extracted using a pre-trained I3D ConvNet, while audio features combined handcrafted descriptors (MFCC, ZCR, Chroma, Spectral Roll-off, and RMSE) with deep features from a BiLSTM network.
Different neural network architectures, I3D, CNN, BiLSTM, and CNN-BiLSTM, were trained for action, comedy, drama, and romance genres using a multiclass classification approach and early fusion for feature integration.
The models achieved accuracies of 89% (I3D), 77% (CNN), 82% (BiLSTM), and 92% (CNN-BiLSTM).
In summary, the CNN-BiLSTM with early fusion outperformed unimodal models, demonstrating the effectiveness of combining visual and audio features for Amharic movie genre classification, thereby enhancing film organization, retrieval, and recommendation systems.
The study demonstrates the effectiveness of multimodal fusion for low-resource languages and contributes to content organization, recommendation, and accessibility in Ethiopian cinema.

Related Results

MULTI-MODAL AMHARIC MOVIE GENRE CLASSIFICATION USING A DEEP LEARNING APPROACH
MULTI-MODAL AMHARIC MOVIE GENRE CLASSIFICATION USING A DEEP LEARNING APPROACH
ABSTRACTThe rapid growth of Ethiopia's film industry, especially in Amharic cinema, has created a significant need for intelligent systems that can automatically categoriz...
Developing an audio search engine for Amharic speech web resources
Developing an audio search engine for Amharic speech web resources
Abstract While general-purpose search engines primarily serve English-language content, the web has seen enormous growth in non-resource-rich languages like Amhar...
Amharic Adhoc Information Retrieval System Based on Morphological Features
Amharic Adhoc Information Retrieval System Based on Morphological Features
Information retrieval (IR) is one of the most important research and development areas due to the explosion of digital data and the need of accessing relevant information from huge...
Developing Amharic Sign Language Recognition Model for Amharic Characters Using Deep Learning Approach
Developing Amharic Sign Language Recognition Model for Amharic Characters Using Deep Learning Approach
Abstract Hearing-impaired people use Sign Language to communicate with each other as well as with other communities. Usually, they are unable to communicate with normal peo...
ANALISIS MODAL KERJA PADA KOPERASI SERBA USAHA DI KOTA METRO
ANALISIS MODAL KERJA PADA KOPERASI SERBA USAHA DI KOTA METRO
Modal kerja merupakan suatu kekayaan yang digunakan untuk membelanjai perusahaan sehari-hari. Modal kerja biasanya berbentuk uang kas, piutang, persediaan barang yang kesemuanya it...
Figurative Language Found in “Wolf Town” Movie
Figurative Language Found in “Wolf Town” Movie
Abstract          This study entitled “figurative language found in “Wolf Town” movie. The purposes of this study are to identify the types of figurative language and to analyze t...

Back to Top