Javascript must be enabled to continue!
T5 for Hate Speech, Augmented Data, and Ensemble
View through CrossRef
We conduct relatively extensive investigations of automatic hate speech (HS) detection using different State-of-The-Art (SoTA) baselines across 11 subtasks spanning six different datasets. Our motivation is to determine which of the recent SoTA models is best for automatic hate speech detection and what advantage methods, such as data augmentation and ensemble, may have on the best model, if any. We carry out six cross-task investigations. We achieve new SoTA results on two subtasks—macro F1 scores of 91.73% and 53.21% for subtasks A and B of the HASOC 2020 dataset, surpassing previous SoTA scores of 51.52% and 26.52%, respectively. We achieve near-SoTA results on two others—macro F1 scores of 81.66% for subtask A of the OLID 2019 and 82.54% for subtask A of the HASOC 2021, in comparison to SoTA results of 82.9% and 83.05%, respectively. We perform error analysis and use two eXplainable Artificial Intelligence (XAI) algorithms (Integrated Gradient (IG) and SHapley Additive exPlanations (SHAP)) to reveal how two of the models (Bi-Directional Long Short-Term Memory Network (Bi-LSTM) and Text-to-Text-Transfer Transformer (T5)) make the predictions they do by using examples. Other contributions of this work are: (1) the introduction of a simple, novel mechanism for correcting Out-of-Class (OoC) predictions in T5, (2) a detailed description of the data augmentation methods, and (3) the revelation of the poor data annotations in the HASOC 2021 dataset by using several examples and XAI (buttressing the need for better quality control). We publicly release our model checkpoints and codes to foster transparency.
Title: T5 for Hate Speech, Augmented Data, and Ensemble
Description:
We conduct relatively extensive investigations of automatic hate speech (HS) detection using different State-of-The-Art (SoTA) baselines across 11 subtasks spanning six different datasets.
Our motivation is to determine which of the recent SoTA models is best for automatic hate speech detection and what advantage methods, such as data augmentation and ensemble, may have on the best model, if any.
We carry out six cross-task investigations.
We achieve new SoTA results on two subtasks—macro F1 scores of 91.
73% and 53.
21% for subtasks A and B of the HASOC 2020 dataset, surpassing previous SoTA scores of 51.
52% and 26.
52%, respectively.
We achieve near-SoTA results on two others—macro F1 scores of 81.
66% for subtask A of the OLID 2019 and 82.
54% for subtask A of the HASOC 2021, in comparison to SoTA results of 82.
9% and 83.
05%, respectively.
We perform error analysis and use two eXplainable Artificial Intelligence (XAI) algorithms (Integrated Gradient (IG) and SHapley Additive exPlanations (SHAP)) to reveal how two of the models (Bi-Directional Long Short-Term Memory Network (Bi-LSTM) and Text-to-Text-Transfer Transformer (T5)) make the predictions they do by using examples.
Other contributions of this work are: (1) the introduction of a simple, novel mechanism for correcting Out-of-Class (OoC) predictions in T5, (2) a detailed description of the data augmentation methods, and (3) the revelation of the poor data annotations in the HASOC 2021 dataset by using several examples and XAI (buttressing the need for better quality control).
We publicly release our model checkpoints and codes to foster transparency.
Related Results
Mapping the scientific knowledge and approaches to defining and measuring hate crime, hate speech, and hate incidents: A systematic review
Mapping the scientific knowledge and approaches to defining and measuring hate crime, hate speech, and hate incidents: A systematic review
Abstract
Background
The difficulties in defining hate crime, hate incidents and hate speech, and in finding a common conc...
Detection of Hate Speech in COVID-19–Related Tweets in the Arab Region: Deep Learning and Topic Modeling Approach
Detection of Hate Speech in COVID-19–Related Tweets in the Arab Region: Deep Learning and Topic Modeling Approach
Background
The massive scale of social media platforms requires an automatic solution for detecting hate speech. These automatic solutions will help reduce the ...
Detection of Hate Speech in COVID-19–Related Tweets in the Arab Region: Deep Learning and Topic Modeling Approach (Preprint)
Detection of Hate Speech in COVID-19–Related Tweets in the Arab Region: Deep Learning and Topic Modeling Approach (Preprint)
BACKGROUND
The massive scale of social media platforms requires an automatic solution for detecting hate speech. These automatic solutions will help reduce ...
Vihapuheen kohteet ja teemat sekä lajit ja muodot ennen ja nyt
Vihapuheen kohteet ja teemat sekä lajit ja muodot ennen ja nyt
Tässä artikkelissa on analysoitu vihapuheen olemusta ja puhunnan muotoja 1930- ja 2000-luvuilla. Tavoitteena on ollut etsiä niitä yhtäläisyyksiä ja eroja, joita kahdella eri aikaka...
A Tale of Two Hates: Testing of the Duplex Theory
A Tale of Two Hates: Testing of the Duplex Theory
This research study focuses on a central premise of Sternberg's (2003) Duplex Theory of Hate. It explored the relevance of attitudes, exposure to hate speech, and behavior as media...
Forensic Linguistics of Hate Speech on Social Media against President Joko Widodo by Chairman of UGM’s Student Executive Board
Forensic Linguistics of Hate Speech on Social Media against President Joko Widodo by Chairman of UGM’s Student Executive Board
This research discusses the hate speech delivered by the chairman of BEM UGM against President Joko Widodo, uploaded on social media. This research uses a forensic linguistic appro...
Hate Speech Detection Using Textual and User Features
Hate Speech Detection Using Textual and User Features
Social media platforms provide users with a powerful platform to share their ideas. Using one’s right to expression to incite
hatred toward a particular group of people ...
Pola Komunikasi Interpersonal Terapis Wicara pada Anak Telambat Bicara (Speech Delay)
Pola Komunikasi Interpersonal Terapis Wicara pada Anak Telambat Bicara (Speech Delay)
Abstract. The tittle of this study is Interpersonal Communication Patterns of Speech Therapists in Children with Speech Delay. Children with speech delay disorders are also social ...

