Javascript must be enabled to continue!
A Multimodal Approach to Language Identification in Sotho-Tswana Musical Videos
View through CrossRef
Language plays a crucial role in Sotho-Tswana musical videos, as it helps determine the sentiment and genre. The Sotho-Tswana languages, spoken in parts of Southern Africa, are used to compose many indigenous songs and music. However, speakers of one of the Sotho-Tswana languages may not understand other Sotho-Tswana languages. Given the widespread availability of these musical videos on social media platforms, there is a need for appropriate recommendations for users based on the language used in the videos. While traditional language identification in music has focused on audio, music information for identifying the singing language can also be embedded in other modalities, such as visual and text. This study employs a multimodal approach to identify the singing language in Sotho-Tswana musical videos. The multimodal approach focuses on three modalities, visual, audio, and textual/lyrics. A multimodal dataset of Sotho-Tswana musical videos is used to train deep learning and language models, for each of the modalities. After the independent training, for each of the modalities, a decision-level (late) fusion method is used to combine the results of the training from the three modalities. The results demonstrate that a multimodal approach outperforms single-modality methods, such as those relying solely on lyrics or textual information.
Title: A Multimodal Approach to Language Identification in Sotho-Tswana Musical Videos
Description:
Language plays a crucial role in Sotho-Tswana musical videos, as it helps determine the sentiment and genre.
The Sotho-Tswana languages, spoken in parts of Southern Africa, are used to compose many indigenous songs and music.
However, speakers of one of the Sotho-Tswana languages may not understand other Sotho-Tswana languages.
Given the widespread availability of these musical videos on social media platforms, there is a need for appropriate recommendations for users based on the language used in the videos.
While traditional language identification in music has focused on audio, music information for identifying the singing language can also be embedded in other modalities, such as visual and text.
This study employs a multimodal approach to identify the singing language in Sotho-Tswana musical videos.
The multimodal approach focuses on three modalities, visual, audio, and textual/lyrics.
A multimodal dataset of Sotho-Tswana musical videos is used to train deep learning and language models, for each of the modalities.
After the independent training, for each of the modalities, a decision-level (late) fusion method is used to combine the results of the training from the three modalities.
The results demonstrate that a multimodal approach outperforms single-modality methods, such as those relying solely on lyrics or textual information.
Related Results
Social Media Use in Neurology: An Analysis of Alzheimer's Information on TikTok with Emphasis on Role of Healthcare Professionals
Social Media Use in Neurology: An Analysis of Alzheimer's Information on TikTok with Emphasis on Role of Healthcare Professionals
Abstract
Introduction
Alzheimer's disease (AD) is the most common neurodegenerative cause of dementia. Social media has become a major source of information for patients and famili...
Hubungan Perilaku Pola Makan dengan Kejadian Anak Obesitas
Hubungan Perilaku Pola Makan dengan Kejadian Anak Obesitas
<p><em><span style="font-size: 11.0pt; font-family: 'Times New Roman',serif; mso-fareast-font-family: 'Times New Roman'; mso-ansi-language: EN-US; mso-fareast-langua...
Assessment of Genetic Diversity and Relationship of the Two Sanga Type Cattle of Botswana Based on Microsatellite Markers
Assessment of Genetic Diversity and Relationship of the Two Sanga Type Cattle of Botswana Based on Microsatellite Markers
Abstract
The study was performed to evaluate genetic variation on two Sanga type cattle found in Botswana; Tswana and Tuli using twelve microsatellite markers. All amplifie...
The significance of word formation techniques utilised in Northern Sotho plant names
The significance of word formation techniques utilised in Northern Sotho plant names
Word formation is a system of rules that can produce new words based on existing lexical items. Word formation strategies are employed by languages to develop their terminology or ...
Učinak poučavanja razrednomu jeziku u izobrazbi nastavnika njemačkoga
Učinak poučavanja razrednomu jeziku u izobrazbi nastavnika njemačkoga
The actual use of classroom language is principally limited to the classroom environment. As far as foreign language learning is concerned, the classroom often turns out to be the ...
Validation of the Tswana Versions of the Roland-Morris Disability Questionnaire, Quebec Disability Scale and Waddell Disability Index
Validation of the Tswana Versions of the Roland-Morris Disability Questionnaire, Quebec Disability Scale and Waddell Disability Index
The use of reliable and valid outcome measures inclinical research as well as clinical practice is very important. Selfreported questionnaires are widely used as outcome measures t...
Multimodal Emotion Recognition and Human Computer Interaction for AI-Driven Mental Health Support (Preprint)
Multimodal Emotion Recognition and Human Computer Interaction for AI-Driven Mental Health Support (Preprint)
BACKGROUND
Mental health has become one of the most urgent global health issues of the twenty-first century. The World Health Organization (WHO) reports tha...
Analyzing Quality of YouTube Videos about Premature Ovarian Failure in the Past Decade
Analyzing Quality of YouTube Videos about Premature Ovarian Failure in the Past Decade
Abstract
Background
To determine the quality of YouTube videos about premature ovarian failure (POF), and variations in quality of professional YouTube videos about POF.
M...

