Social science · Paper · July 2023

Inferring Hate Speech Trends for Contemporary Tweets Using a Novel Machine Learning Approach from Supervised Learning Algorithms

Aryan Singhal

Abstract

Social media platforms such as Twitter have become ubiquitous in our contemporary society, providing a platform for individuals to express their opinions and engage in discussions on a wide range of topics including those that are neutral and controversial. However, the growing popularity of Twitter has also led to an increase in the prevalence of hate speech, which raises concerns about its impact on individuals and society. This research investigates hate speech trends on Twitter by utilizing supervised machine learning classification, specifically Naive Bayes, and employing Natural Language Processing (NLP) features such as on a range of neutral and controversial topics. The study compares the prevalence of hate speech in these topics and tracks such trends from January 2022 to January 2023. The results show that hate speech was nearly 400% more prevalent in controversial topics than in neutral topics over the course of the year. In addition, this research finds that controversial topics are consistently more vulnerable to hate speech throughout the course of the year when compared to neutral topics. To conduct this study, a Multinomial Naive Bayes classification model was trained on a publicly available Twitter dataset that was specifically labeled for semantic hate speech and achieved an accuracy rate of 94.46%. Ultimately, the higher vulnerability of controversial topics should necessitate policymakers to introduce stricter warnings or frequent policy reminders to platform users. Such changes will foster a respectful and inclusive platform for users, preserving their freedom of expression and encouraging constructive discussions.