Search engine for discovering works of Art, research articles, and books related to Art and Culture
ShareThis
Javascript must be enabled to continue!

The ‘Sessiz İstila’ Dataset: Emotion-Annotated Turkish Tweets on Anti-Refugee Discourse and Emoji Lexicon

View through CrossRef
This article describes an emotion-annotated corpus of Turkish-language posts published on X (formerly Twitter) that discuss ‘sessiz istila’ (‘silent invasion’), a contested phrase used in Turkish public debate to frame the presence of refugees, particularly Syrian refugees, as a threat. We retrieved posts using the academictwitteR interface to the Twitter Academic Research API with the keyword ‘sessiz istila’, restricted to publicly available Turkish-language posts published between 1 June 2021 and 31 December 2022, yielding 47,024 posts. We normalized every post with a rule set implemented in R that replaces retweet markers, URLs, user mentions, and hashtags with Turkish placeholder tokens, converts text to lowercase, and substitutes each emoji with a Turkish description drawn from a purpose-built emoji lexicon of 3,664 entries released alongside the corpus. Emotion labels were assigned automatically by an Emotion Recognition Model obtained by fine-tuning BERTurk (bert-base-turkish-cased) on a class-balanced subset of the TREMO dataset containing 18,018 validated entries, 3,003 for each of the six Ekman categories (happiness, fear, sadness, disgust, surprise and anger). We retained predictions with softmax confidence below 0.6 but flagged them as ‘ambiguous’ rather than discarding them so that users can apply their own confidence policy. The released corpus therefore contains 40,880 confidently labeled posts and 6,144 ambiguous posts, each with its creation timestamp, normalized text, date, year, and emotion label. The dataset supports secondary research on online xenophobia and anti-immigrant sentiment, on event-driven emotional dynamics in social media, and on benchmarking and error analysis of emotion classifiers for Turkish, an agglutinative and comparatively low-resource language. The balanced TREMO training file and the Turkish emoji lexicon are reusable across corpora.
Title: The ‘Sessiz İstila’ Dataset: Emotion-Annotated Turkish Tweets on Anti-Refugee Discourse and Emoji Lexicon
Description:
This article describes an emotion-annotated corpus of Turkish-language posts published on X (formerly Twitter) that discuss ‘sessiz istila’ (‘silent invasion’), a contested phrase used in Turkish public debate to frame the presence of refugees, particularly Syrian refugees, as a threat.
We retrieved posts using the academictwitteR interface to the Twitter Academic Research API with the keyword ‘sessiz istila’, restricted to publicly available Turkish-language posts published between 1 June 2021 and 31 December 2022, yielding 47,024 posts.
We normalized every post with a rule set implemented in R that replaces retweet markers, URLs, user mentions, and hashtags with Turkish placeholder tokens, converts text to lowercase, and substitutes each emoji with a Turkish description drawn from a purpose-built emoji lexicon of 3,664 entries released alongside the corpus.
Emotion labels were assigned automatically by an Emotion Recognition Model obtained by fine-tuning BERTurk (bert-base-turkish-cased) on a class-balanced subset of the TREMO dataset containing 18,018 validated entries, 3,003 for each of the six Ekman categories (happiness, fear, sadness, disgust, surprise and anger).
We retained predictions with softmax confidence below 0.
6 but flagged them as ‘ambiguous’ rather than discarding them so that users can apply their own confidence policy.
The released corpus therefore contains 40,880 confidently labeled posts and 6,144 ambiguous posts, each with its creation timestamp, normalized text, date, year, and emotion label.
The dataset supports secondary research on online xenophobia and anti-immigrant sentiment, on event-driven emotional dynamics in social media, and on benchmarking and error analysis of emotion classifiers for Turkish, an agglutinative and comparatively low-resource language.
The balanced TREMO training file and the Turkish emoji lexicon are reusable across corpora.

Related Results

The Language of Emoji in Social Media
The Language of Emoji in Social Media
The very fast development of information technology which is characterized by an influx of industry 4.0 has changed the way of human and behavior in language. The grammar which i...
Sessiz İstifa Kavramının Teorik İncelemesi
Sessiz İstifa Kavramının Teorik İncelemesi
Sessiz istifa kavramı işletme yönetimi literatüründe son yıllarda artan bir öneme sahip kavramlardan biridir. Örgütlerin istihdam ettikleri bireylerden en temel beklentilerinden bi...
Sessiz İstifa ve Sessiz İşten Çıkarma
Sessiz İstifa ve Sessiz İşten Çıkarma
Dünyada ve yönetim alanında meydana gelen değişim ve gelişmeler yeni konuların ortaya çıkmasına sebep olmaktadır. Yeni çıkan ve son zamanlarda popüler olan kavramlardan olan sessiz...
How Good Is GPT’s “Emojinal Intelligence”? Investigating Emoji Patterns in LLM-Generated Social Media Text
How Good Is GPT’s “Emojinal Intelligence”? Investigating Emoji Patterns in LLM-Generated Social Media Text
Recent advancement in Large Language Models (LLMs) has opened the prospect of generating text for social media content that mimics human writing. The misuse of these tools presents...
Multimodal Emotion Recognition and Human Computer Interaction for AI-Driven Mental Health Support (Preprint)
Multimodal Emotion Recognition and Human Computer Interaction for AI-Driven Mental Health Support (Preprint)
BACKGROUND Mental health has become one of the most urgent global health issues of the twenty-first century. The World Health Organization (WHO) reports tha...
The Role Of Memes And Emoji In Enhancing Customer Engagement
The Role Of Memes And Emoji In Enhancing Customer Engagement
The prevalent trend of using memes and emoji on social media has changed the landscape of communication across online platforms. Memes and emoji, despite being a popular online lan...

Back to Top