Aymé Arango

dblp:245/1776 · DBLP profile ↗
← Back
4ranked-venue papers
4as first author
2since 2021 · last 2024
0000-0003-2310-8084ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4 · 4 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
2 papers
Web and social media mining · 86% Information retrieval · 14%
Software engineering, system software, and programming languages
1 paper
Empirical software engineering · 100%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Web and social media mining › hate speech
hate speech detection
0.822020
Language Agnostic Hate Speech Detection · SIGIR 2020
Hate Speech Detection is Not as Easy as You May Think: A Closer Look at Model Validation · SIGIR 2019
Empirical software engineering
experimental methodology
0.412019
Hate Speech Detection is Not as Easy as You May Think: A Closer Look at Model Validation · SIGIR 2019
Empirical software engineering › software evaluation
model validation
0.412019
Hate Speech Detection is Not as Easy as You May Think: A Closer Look at Model Validation · SIGIR 2019
Information retrieval
cross-language information retrieval
0.112020
Language Agnostic Hate Speech Detection · SIGIR 2020

Methods — techniques the papers use, named apart from their topics

supervised classification · 1.2
YearPublicationVenuePosition
2024 MultiFOLD: Multi-source Domain Adaption for Offensive Language Detection
abstract
Automatic offensive language detection remains challenging, and is a crucial part of preserving the openness of digital spaces, which are an integral part of our everyday experi- ence. The ever-growing forms of offensive online content makes traditional supervised approaches harder to scale due to the financial and psychological costs incurred by collect- ing human annotations. In this work, we propose a domain adaptation framework for offensive language detection, Mul- tiFOLD, which learns and adapts from multiple existing data sets (or source domains) to an unlabeled target domain. Under the hood, a curriculum learning algorithm is employed that kicks off learning with the instances most similar to the target domain while gradually expanding to more distant instances. The proposed model is trained with a standard task-specific loss and a domain adversarial objective which aims to min- imize the language distinctions across the multiple sources and the target, allowing the classifier to distinguish offen- siveness rather than domain. Our experiments on six pub- licly available data sets demonstrate the effectiveness of Mul- tiFOLD. Relative improvement in F1 of 0.5% (WOAH) to 29.7% (ICWSM) is found across five out of the six datasets compared to the state-of-the-art domain adaptation baseline BERT-DAA, resulting in an average of 6% relative F1-score gain.
Aymé Arango, Parisa Kaghazgaran, Sheikh Muhammad Sarwar, Vanessa Murdock 0001, C. J. Lee
ICWSM1
2022 Hate speech detection is not as easy as you may think: A closer look at model validation (extended version)
Aymé Arango, Jorge Pérez 0001, Barbara Poblete
Inf. Syst.1
2020 Language Agnostic Hate Speech Detection
abstract
The growth in social Web platforms in the past years has brought an increase in displays of online hate speech. This subject is considered as a critical matter in the Web community, since it can be related to potentially dangerous actions that affect individuals and groups in the physical world. The automatic detection of this type of expressions has been the center of several investigations over the past few years. However, most research on this subject has been done for the English language and on rather limited datasets. In addition, although some works approach the problem from a multilingual perspective, analyzing different language separately, across-lingual perspective of this problem has not been used so far.
Aymé Arango
SIGIR1
2019 Hate Speech Detection is Not as Easy as You May Think: A Closer Look at Model Validation
abstract
Hate speech is an important problem that is seriously affecting the dynamics and usefulness of online social communities. Large scale social platforms are currently investing important resources into automatically detecting and classifying hateful content, without much success. On the other hand, the results reported by state-of-the-art systems indicate that supervised approaches achieve almost perfect performance but only within specific datasets. In this work, we analyze this apparent contradiction between existing literature and actual applications. We study closely the experimental methodology used in prior work and their generalizability to other datasets. Our findings evidence methodological issues, as well as an important dataset bias. As a consequence, performance claims of the current state-of-the-art have become significantly overestimated. The problems that we have found are mostly related to data overfitting and sampling issues. We discuss the implications for current research and re-conduct experiments to give a more accurate picture of the current state-of-the art methods.
Aymé Arango, Jorge Pérez 0001, Barbara Poblete
SIGIR1