Mohammad Atari
2021
Improving Counterfactual Generation for Fair Hate Speech Detection
Aida Mostafazadeh Davani
|
Ali Omrani
|
Brendan Kennedy
|
Mohammad Atari
|
Xiang Ren
|
Morteza Dehghani
Proceedings of the 5th Workshop on Online Abuse and Harms (WOAH 2021)
Bias mitigation approaches reduce models’ dependence on sensitive features of data, such as social group tokens (SGTs), resulting in equal predictions across the sensitive features. In hate speech detection, however, equalizing model predictions may ignore important differences among targeted social groups, as hate speech can contain stereotypical language specific to each SGT. Here, to take the specific language about each SGT into account, we rely on counterfactual fairness and equalize predictions among counterfactuals, generated by changing the SGTs. Our method evaluates the similarity in sentence likelihoods (via pre-trained language models) among counterfactuals, to treat SGTs equally only within interchangeable contexts. By applying logit pairing to equalize outcomes on the restricted set of counterfactuals for each instance, we improve fairness metrics while preserving model performance on hate speech detection.
2019
Reporting the Unreported: Event Extraction for Analyzing the Local Representation of Hate Crimes
Aida Mostafazadeh Davani
|
Leigh Yeh
|
Mohammad Atari
|
Brendan Kennedy
|
Gwenyth Portillo Wightman
|
Elaine Gonzalez
|
Natalie Delong
|
Rhea Bhatia
|
Arineh Mirinjian
|
Xiang Ren
|
Morteza Dehghani
Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)
Official reports of hate crimes in the US are under-reported relative to the actual number of such incidents. Further, despite statistical approximations, there are no official reports from a large number of US cities regarding incidents of hate. Here, we first demonstrate that event extraction and multi-instance learning, applied to a corpus of local news articles, can be used to predict instances of hate crime. We then use the trained model to detect incidents of hate in cities for which the FBI lacks statistics. Lastly, we train models on predicting homicide and kidnapping, compare the predictions to FBI reports, and establish that incidents of hate are indeed under-reported, compared to other types of crimes, in local press.
Search
Co-authors
- Aida Mostafazadeh Davani 2
- Brendan Kennedy 2
- Xiang Ren 2
- Morteza Dehghani 2
- Leigh Yeh 1
- show all...