Nora Lewis

dblp:367/0226 · DBLP profile ↗
← Back
2ranked-venue papers in the field
1as first author
2since 2021 · last 2024
—ORCID · none

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 2 (1 first)
YearPublicationVenuePosition
2024 On the Effectiveness of Text and Image Embeddings in Multimodal Hate Speech Detection
abstract
Social media content is increasingly subject to hate speech towards specific demographic groups. In this context, machine learning approaches appear relevant to detect malicious contents and support moderators in mitigation initiatives. However, many of the existing approaches are exclusively focused on the analysis of textual contents. On the other hand, studies that address multiple data modalities often rely on a single feature representation and a fully supervised learning setting. In this paper, we tackle multimodal hate speech detection resorting to different learning settings (one-class learning and binary classification). We also investigate the effectiveness of multiple deep learning model backbones and language models to extract embedding feature representations for text and image modalities. Our experiments with a real-world hate speech dataset show that there is a significant performance gap between one-class learning and binary classification, and that the choice of embedding representations for image and text modalities can impact the detection performance for different predictive models.
Nora Lewis, Charles C. Cavalcante, Zois Boukouvalas, Roberto Corizzo
IEEE Big Data1
2023 Multimodal One-class Learning for Malicious Online Content Detection
abstract
Social media content can present a number of threats, including misinformation and hate speech towards specific demographic groups. One challenge is to effectively discriminate between benign and malicious posts, given the massive amount of available content. In this context, predictive models for malicious content detection can be extremely valuable, leading to the automatic removal of posts and user accounts or content being flagged for subsequent moderation. However, some of the existing detection models are limited to the analysis of a single data modality. At the same time, most multi-modal approaches operate in a fully supervised learning setting that assumes the availability of labeled data for both benign and malicious content. In this paper, we fill this gap by proposing a multimodal one-class learning approach for malicious online content detection. Our approach leverages feature extraction, dimensionality reduction, and one-class learning models to analyze text and image data in online posts simultaneously. Models learn their decision function in the challenging scenario where only benign online content is used as training data, overcoming the limitations of a fully supervised setting. Our experiments with two real-world datasets containing misinformation and hate speech posts reveal the effectiveness of different combinations of one-class learning models and dimensionality reduction techniques.
Roberto Corizzo, Nora Lewis, Lucas P. Damasceno, Allison Shafer, Charles C. Cavalcante, Zois Boukouvalas
IEEE Big Data2