EDBT 2026 Demo / reviewers in the wild / expert
Aymé Arango
dblp:245/1776
· DBLP profile ↗
4ranked-venue papers
4as first author
2since 2021 · last 2024
0000-0003-2310-8084ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 4 · 4 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
2 papers |
Web and social media mining · 86% Information retrieval · 14% | |
| Software engineering, system software, and programming languages
1 paper |
Empirical software engineering · 100% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Web and social media mining › hate speech
hate speech detection |
0.8 | 2 | 2020 | Language Agnostic Hate Speech Detection · SIGIR 2020 Hate Speech Detection is Not as Easy as You May Think: A Closer Look at Model Validation · SIGIR 2019 |
Empirical software engineering
experimental methodology |
0.4 | 1 | 2019 | Hate Speech Detection is Not as Easy as You May Think: A Closer Look at Model Validation · SIGIR 2019 |
Empirical software engineering › software evaluation
model validation |
0.4 | 1 | 2019 | Hate Speech Detection is Not as Easy as You May Think: A Closer Look at Model Validation · SIGIR 2019 |
Information retrieval
cross-language information retrieval |
0.1 | 1 | 2020 | Language Agnostic Hate Speech Detection · SIGIR 2020 |
Methods — techniques the papers use, named apart from their topics
supervised classification · 1.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | MultiFOLD: Multi-source Domain Adaption for Offensive Language DetectionabstractAutomatic offensive language detection remains challenging, and is a crucial part of preserving the openness of digital spaces, which are an integral part of our everyday experi- ence. The ever-growing forms of offensive online content makes traditional supervised approaches harder to scale due to the financial and psychological costs incurred by collect- ing human annotations. In this work, we propose a domain adaptation framework for offensive language detection, Mul- tiFOLD, which learns and adapts from multiple existing data sets (or source domains) to an unlabeled target domain. Under the hood, a curriculum learning algorithm is employed that kicks off learning with the instances most similar to the target domain while gradually expanding to more distant instances. The proposed model is trained with a standard task-specific loss and a domain adversarial objective which aims to min- imize the language distinctions across the multiple sources and the target, allowing the classifier to distinguish offen- siveness rather than domain. Our experiments on six pub- licly available data sets demonstrate the effectiveness of Mul- tiFOLD. Relative improvement in F1 of 0.5% (WOAH) to 29.7% (ICWSM) is found across five out of the six datasets compared to the state-of-the-art domain adaptation baseline BERT-DAA, resulting in an average of 6% relative F1-score gain. Aymé Arango, Parisa Kaghazgaran, Sheikh Muhammad Sarwar, Vanessa Murdock 0001, C. J. Lee |
ICWSM | 1 |
| 2022 | Hate speech detection is not as easy as you may think: A closer look at model validation (extended version)
Aymé Arango, Jorge Pérez 0001, Barbara Poblete |
Inf. Syst. | 1 |
| 2020 | Language Agnostic Hate Speech DetectionabstractThe growth in social Web platforms in the past years has brought an increase in displays of online hate speech. This subject is considered as a critical matter in the Web community, since it can be related to potentially dangerous actions that affect individuals and groups in the physical world. The automatic detection of this type of expressions has been the center of several investigations over the past few years. However, most research on this subject has been done for the English language and on rather limited datasets. In addition, although some works approach the problem from a multilingual perspective, analyzing different language separately, across-lingual perspective of this problem has not been used so far. Aymé Arango |
SIGIR | 1 |
| 2019 | Hate Speech Detection is Not as Easy as You May Think: A Closer Look at Model ValidationabstractHate speech is an important problem that is seriously affecting the dynamics and usefulness of online social communities. Large scale social platforms are currently investing important resources into automatically detecting and classifying hateful content, without much success. On the other hand, the results reported by state-of-the-art systems indicate that supervised approaches achieve almost perfect performance but only within specific datasets. In this work, we analyze this apparent contradiction between existing literature and actual applications. We study closely the experimental methodology used in prior work and their generalizability to other datasets. Our findings evidence methodological issues, as well as an important dataset bias. As a consequence, performance claims of the current state-of-the-art have become significantly overestimated. The problems that we have found are mostly related to data overfitting and sampling issues. We discuss the implications for current research and re-conduct experiments to give a more accurate picture of the current state-of-the art methods. Aymé Arango, Jorge Pérez 0001, Barbara Poblete |
SIGIR | 1 |