Search engine for discovering works of Art, research articles, and books related to Art and Culture
ShareThis
Javascript must be enabled to continue!

IRMA: the 335-million-word Italian coRpus for studying MisinformAtion

View through CrossRef
The dissemination of false information on the internet has received considerable attention over the last decade. Misinformation often spreads faster than mainstream news, thus making manual fact checking inefficient or, at best, labor-intensive. Therefore, there is an increasing need to develop methods for automatic detection of misinformation. Although resources for creating such methods are available in English, other languages are often underrepresented in this effort. With this contribution, we present IRMA, a corpus containing over 600,000 Italian news articles (335+ million tokens) collected from 56 websites classified as ‘untrustworthy’ by professional factcheckers. The corpus is freely available and comprises a rich set of text- and website-level data, representing a turnkey resource to test hypotheses and develop automatic detection algorithms. It contains texts, titles, and dates (from 2004 to 2022), along with three types of semantic measures (i.e., keywords, topics at three different resolutions, and LIWC lexical features). IRMA also includes domain specific information such as source type (e.g., political, health, conspiracy, etc.), quality, and higher-level metadata, including several metrics of website incoming traffic that allow to investigate user online behavior. IRMA constitutes the largest corpus of misinformation available today in Italian, making it a valid tool for advancing quantitative research on untrustworthy news detection and ultimately helping limit the spread of misinformation.
Title: IRMA: the 335-million-word Italian coRpus for studying MisinformAtion
Description:
The dissemination of false information on the internet has received considerable attention over the last decade.
Misinformation often spreads faster than mainstream news, thus making manual fact checking inefficient or, at best, labor-intensive.
Therefore, there is an increasing need to develop methods for automatic detection of misinformation.
Although resources for creating such methods are available in English, other languages are often underrepresented in this effort.
With this contribution, we present IRMA, a corpus containing over 600,000 Italian news articles (335+ million tokens) collected from 56 websites classified as ‘untrustworthy’ by professional factcheckers.
The corpus is freely available and comprises a rich set of text- and website-level data, representing a turnkey resource to test hypotheses and develop automatic detection algorithms.
It contains texts, titles, and dates (from 2004 to 2022), along with three types of semantic measures (i.
e.
, keywords, topics at three different resolutions, and LIWC lexical features).
IRMA also includes domain specific information such as source type (e.
g.
, political, health, conspiracy, etc.
), quality, and higher-level metadata, including several metrics of website incoming traffic that allow to investigate user online behavior.
IRMA constitutes the largest corpus of misinformation available today in Italian, making it a valid tool for advancing quantitative research on untrustworthy news detection and ultimately helping limit the spread of misinformation.

Related Results

Žanrovska analiza pomorskopravnih tekstova i ostvarenje prijevodnih univerzalija u njihovim prijevodima s engleskoga jezika
Žanrovska analiza pomorskopravnih tekstova i ostvarenje prijevodnih univerzalija u njihovim prijevodima s engleskoga jezika
Genre implies formal and stylistic conventions of a particular text type, which inevitably affects the translation process. This „force of genre bias“ (Prieto Ramos, 2014) has been...
Who is susceptible to online health misinformation? A test of four psychosocial hypotheses
Who is susceptible to online health misinformation? A test of four psychosocial hypotheses
ABSTRACTObjective: Health misinformation on social media threatens public health. One question that could lend insight into how and through whom misinformation spreads is whether c...
The Discussions of Monkeypox Misinformation on Social Media
The Discussions of Monkeypox Misinformation on Social Media
The global outbreak of the monkeypox virus was declared a health emergency by the World Health Organization (WHO). During such emergencies, misinformation about health suggestions ...

Back to Top