Javascript must be enabled to continue!

IRMA: the 335-million-word Italian coRpus for studying MisinformAtion

The dissemination of false information on the internet has received considerable attention over the last decade. Misinformation often spreads faster than mainstream news, thus making manual fact checking inefficient or, at best, labor-intensive. Therefore, there is an increasing need to develop methods for automatic detection of misinformation. Although resources for creating such methods are available in English, other languages are often underrepresented in this effort. With this contribution, we present IRMA, a corpus containing over 600,000 Italian news articles (335+ million tokens) collected from 56 websites classified as ‘untrustworthy’ by professional factcheckers. The corpus is freely available and comprises a rich set of text- and website-level data, representing a turnkey resource to test hypotheses and develop automatic detection algorithms. It contains texts, titles, and dates (from 2004 to 2022), along with three types of semantic measures (i.e., keywords, topics at three different resolutions, and LIWC lexical features). IRMA also includes domain specific information such as source type (e.g., political, health, conspiracy, etc.), quality, and higher-level metadata, including several metrics of website incoming traffic that allow to investigate user online behavior. IRMA constitutes the largest corpus of misinformation available today in Italian, making it a valid tool for advancing quantitative research on untrustworthy news detection and ultimately helping limit the spread of misinformation.

Center for Open Science

Fabio Carrella Alessandro Miani Stephan Lewandowsky

2023

Title: IRMA: the 335-million-word Italian coRpus for studying MisinformAtion

Description:

The dissemination of false information on the internet has received considerable attention over the last decade.

Misinformation often spreads faster than mainstream news, thus making manual fact checking inefficient or, at best, labor-intensive.

Therefore, there is an increasing need to develop methods for automatic detection of misinformation.

Although resources for creating such methods are available in English, other languages are often underrepresented in this effort.

With this contribution, we present IRMA, a corpus containing over 600,000 Italian news articles (335+ million tokens) collected from 56 websites classified as ‘untrustworthy’ by professional factcheckers.

The corpus is freely available and comprises a rich set of text- and website-level data, representing a turnkey resource to test hypotheses and develop automatic detection algorithms.

It contains texts, titles, and dates (from 2004 to 2022), along with three types of semantic measures (i.

, keywords, topics at three different resolutions, and LIWC lexical features).

IRMA also includes domain specific information such as source type (e.

, political, health, conspiracy, etc.

), quality, and higher-level metadata, including several metrics of website incoming traffic that allow to investigate user online behavior.

IRMA constitutes the largest corpus of misinformation available today in Italian, making it a valid tool for advancing quantitative research on untrustworthy news detection and ultimately helping limit the spread of misinformation.

Back

Background Although at present there is broad agreement among researchers, health professionals, and policy makers on the need to control and combat health misi...

Prevalence of Health Misinformation on Social Media: Systematic Review (Preprint)

BACKGROUND Although at present there is broad agreement among researchers, health professionals, and policy makers on the need to control and combat health ...

Žanrovska analiza pomorskopravnih tekstova i ostvarenje prijevodnih univerzalija u njihovim prijevodima s engleskoga jezika

Genre implies formal and stylistic conventions of a particular text type, which inevitably affects the translation process. This „force of genre bias“ (Prieto Ramos, 2014) has been...

Who is susceptible to online health misinformation? A test of four psychosocial hypotheses

ABSTRACTObjective: Health misinformation on social media threatens public health. One question that could lend insight into how and through whom misinformation spreads is whether c...

The Discussions of Monkeypox Misinformation on Social Media

The global outbreak of the monkeypox virus was declared a health emergency by the World Health Organization (WHO). During such emergencies, misinformation about health suggestions ...

A Technique for Constructing <span class="changedDisabl

To solve the problem of constructing the frequency responses (FR) of filters on switched capacitors, which belong to the class of electronic circuits with a periodically changing s...

Does X Mark the Spot? Investigating discussions about cancer screening programs on X/Twitter through corpus analysis (Preprint)

BACKGROUND While cancer screening is proven to be effective in the early detection of the disease and early detection enables better treatment options, screening ...

Successful Replacement Therapy After <span c

Background. Vitamin D has recognized immunomodulatory, anti-proliferative, and differentiation-regulating effects primarily mediated through its genomic effects via the vitamin D r...

Email:
Password:

Email:

IRMA: the 335-million-word Italian coRpus for studying MisinformAtion

Related Results