Restricted Dataset for "Large Scale Crowdsourcing and Characterization of Twitter Abusive Behavior"

Antigoni-Maria Founta, Aristotle University of Thessaloniki
Constantinos Djouvas, Cyprus University of Technology
Despoina Chatzakou, Aristotle University of Thessaloniki
Ilias Leontiadis, Telefonica Research
Jeremy Blackburn, University of Alabama at Birmingham
Gianluca Stringhini, University College London
Athena Vakali, Aristotle University of Thessaloniki
Michael Sirivianos, Cyprus University of Technology
Nicolas Kourtellis, Telefonica Research

Publication Date

2-21-2020

Abstract

Restricted Dataset for the "Large Scale Crowdsourcing and Characterization of Twitter Abusive Behavior" paper, published in ICWSM 2018. The full text of the paper can be found here. The Public version of the dataset can be found here

hatespeech_text_label_vote_RESTRICTED_100K.csv: contains ~100K raws with tweet text, the associated majority label, and the number of votes for the majority label. The tweets are shuffled so that there is no connection between tweet IDs and texts (in order to be in line with the T&C of Twitter).
retweets.csv: contains ~2K rows, where every row consists of the row number in the hatespeech_text_label_vote_RESTRICTED_100K.csv file which is the first occurrence of a Tweet text followed by comma-separated row numbers of all other occurrences of the same Tweet text in the same file. There are ~8K other occurrences due to retweets.

Please cite the paper in any published work that uses any of these resources.

@inproceedings{founta2018large,
    title={Large Scale Crowdsourcing and Characterization of Twitter Abusive Behavior},
    author={Founta, Antigoni-Maria and Djouvas, Constantinos and Chatzakou, Despoina and Leontiadis, Ilias and Blackburn, Jeremy and Stringhini, Gianluca and Vakali, Athena and Sirivianos, Michael and Kourtellis, Nicolas},
    booktitle={11th International Conference on Web and Social Media, ICWSM 2018},
    year={2018},
    organization={AAAI Press}
}

For any further questions contact a.m.founta at gmail dot com AND markos.charalambous at eecei dot cut dot ac dot cy

Repository

Zenodo

Access Instructions

Access to this data is restricted.

Funder

Funder: European Commission
Funder DOI: 10.13039/501100000780
EnhaNcing seCurity And privacy in the Social wEb: a user centered approach for the protection of minors
691025

Link to Dataset

COinS

Restricted Dataset for "Large Scale Crowdsourcing and Characterization of Twitter Abusive Behavior"

Publication Date

Abstract

Repository

Access Instructions

Funder

Search

Browse

Author Corner

Research Data Catalog

Restricted Dataset for "Large Scale Crowdsourcing and Characterization of Twitter Abusive Behavior"

Authors

Publication Date

Abstract

Repository

Access Instructions

Funder

Share

Search

Browse

Author Corner