This project builds a text classification model to automatically distinguish between hate speech, offensive language, and neutral tweets using NLP preprocessing techniques and machine learning.
Dataset: ~24.8K tweets labeled as hate speech (0), offensive language (1), or neither (2), sourced from Davidson et al.'s hate speech study. Goal: Clean raw, noisy tweet text (mentions, RTs, punctuation, slang) and train a classifier that generalizes well despite class imbalance.