Sebastian Leal-Arenas

dblp:353/6791 · DBLP profile ↗
← Back
2ranked-venue papers in the field
0as first author
2since 2021 · last 2024
0009-0007-3549-3270ORCID · corroborated

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 2
YearPublicationVenuePosition
2024 Multimodal Deep Learning for Online Meme Classification
abstract
Memes possess a humorous intent, yet they can also be used for malicious purposes. Analysing meme data has the potential to enhance content monitoring, identify emerging topics, and support content moderation in online platforms. Memes also represent an interesting use case for multimodal machine learning, as they combine text and image data. In this study, we explored the linguistic characteristics and analysed the convergent themes of five meme classes through common word extraction. Moreover, we compared the effectiveness of various machine learning models, i.e., unimodal (text or image) and multimodal (early fusion, late fusion) in binary and multiclass meme classification tasks. Our results on a large meme dataset showed that memes heavily adhered to current affairs, demonstrated by the high frequency of topical words across meme classes. Regarding model accuracy, early fusion achieved superior accuracy over late fusion in meme classification. Binary models outperformed multi-class classification methods. However, fusion models did not consistently surpass the accuracy of independent text or image-based models.
Stephanie Han, Sebastian Leal-Arenas, Eftim Zdravevski, Charles C. Cavalcante, Zois Boukouvalas, Roberto Corizzo
IEEE Big Data2
2023 One-GPT: A One-Class Deep Fusion Model for Machine-Generated Text Detection
abstract
On the brink of the one-year anniversary since the public release of ChatGPT, scholarly research has directed their interest toward detection methodologies for machine-generated text. Different models have been proposed, including feature-based classification and detection approaches, as well as deep learning architectures, with a small portion of them integrating contextual information to enhance accurate predictions. Moreover, detection approaches explored thus far have focused primarily on English datasets, with limited attention given to the examination of similar methods in other languages. As a result, the applicability and efficacy of these methods in linguistically diverse contexts remains underexplored. In this paper, we present a one-class deep fusion model that considers both contextual text features derived from word embeddings and linguistic features to detect machine-generated texts in English and Spanish. Experimental results indicated that our model outperformed popular baseline one-class learning models in the detection task, presenting higher accuracy scores in the English dataset. Results are discussed in comparison to competing classifiers as well as the language biases found in detection models.
Roberto Corizzo, Sebastian Leal-Arenas
IEEE Big Data2