IRMA: the 335-million-word Italian coRpus for studying MisinformAtion

Research output: Chapter in Book/Report/Conference proceedingConference Contribution (Conference Proceeding)

3 Citations (Scopus)

Abstract

The dissemination of false information on the internet has received considerable attention over the last decade. Misinformation often spreads faster than mainstream news, thus making manual fact checking inefficient or, at best, labor-intensive. Therefore, there is an increasing need to develop methods for automatic detection of misinformation. Although resources for creating such methods are available in English, other languages are often underrepresented in this effort. With this contribution, we present IRMA, a corpus containing over 600,000 Italian news articles (335+ million tokens) collected from 56 websites classified as 'untrustworthy' by professional fact-checkers. The corpus is freely available and comprises a rich set of text- and website-level data, representing a turnkey resource to test hypotheses and develop automatic detection algorithms. It contains texts, titles, and dates (from 2004 to 2022), along with three types of semantic measures (i.e., keywords, topics at three different resolutions, and LIWC lexical features). IRMA also includes domain-specific information such as source type (e.g., political, health, conspiracy, etc.), quality, and higher-level metadata, including several metrics of website incoming traffic that allow to investigate user online behavior. IRMA constitutes the largest corpus of misinformation available today in Italian, making it a valid tool for advancing quantitative research on untrustworthy news detection and ultimately helping limit the spread of misinformation.

Original languageEnglish
Title of host publicationProceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics
EditorsAndreas Vlachos, Isabelle Augenstein
PublisherAssociation for Computational Linguistics (ACL)
Pages2339–2349
Number of pages11
ISBN (Electronic)9781959429449
DOIs
Publication statusPublished - 6 May 2023
Event17th Conference of the European Chapter of the Association for Computational Linguistics, EACL 2023 - Dubrovnik, Croatia
Duration: 2 May 20236 May 2023

Publication series

Name
ISSN (Print)1525-2450
ISSN (Electronic)1525-2450

Conference

Conference17th Conference of the European Chapter of the Association for Computational Linguistics, EACL 2023
Country/TerritoryCroatia
CityDubrovnik
Period2/05/236/05/23

Bibliographical note

Publisher Copyright:
© 2023 Association for Computational Linguistics.

Research Groups and Themes

  • TeDCog

Fingerprint

Dive into the research topics of 'IRMA: the 335-million-word Italian coRpus for studying MisinformAtion'. Together they form a unique fingerprint.

Cite this