EDBT 2026 Demo / reviewers in the wild / expert
Peiling Yi
dblp:300/2932
· DBLP profile ↗
4ranked-venue papers
4as first author
4since 2021 · last 2025
0000-0001-6680-6320ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 4 · 4 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ID-XCB: Data-Independent Debiasing for Fair and Accurate Transformer-Based Cyberbullying DetectionabstractThe use of swear words is a common proxy to collect datasets with cyberbullying incidents, which increases the chances of collecting such events that are otherwise hard to find. However, datasets collected through this means also have a risk of introducing biases in cyberbullying detection models which can learn spurious associations between swear words and the presence of incidents. In this study, we undertake a pioneering study of measuring and mitigating swearing bias in cyberbullying detection tasks. Initially, we employ word-level bias measures to demonstrate the distinctive features related to swearing biases in transformer-based cyberbullying detection models. Subsequently, we introduce ID-XCB, the first data-independent debiasing technique that combines adversarial training, bias constraints and a debias fine-tuning approach aimed at alleviating model attention to bias-inducing words without impacting overall model performance. Lastly, we explore ID-XCB on two popular session-based cyberbullying detection datasets along with a comprehensive set of ablation studies and model generalisation studies. Our findings show that ID-XCB learns robust cyberbullying detection capabilities while mitigating biases tied to swear word usage. It consistently outperforms state-of-the-art debiasing methods in terms of both performance improvement and bias mitigation. In addition, by combining quantitative and qualitative analyses, we demonstrate the potential for generalisability of our approach when tackling unseen data. Peiling Yi, Arkaitz Zubiaga |
ICWSM | 1 |
| 2025 | Detecting Harassment and Defamation in Cyberbullying with Emotion-Adaptive TrainingabstractExisting research on detecting cyberbullying incidents on social media has primarily concentrated on harassment and is typically approached as a binary classification task. However, cyberbullying encompasses various forms, such as denigration and harassment, which celebrities frequently face. Furthermore, suitable training data for these diverse forms of cyberbullying remains scarce. In this study, we first develop a celebrity cyberbullying dataset that encompasses two distinct types of incidents: harassment and defamation. We investigate various types of transformer-based models, namely masked (RoBERTa, Bert and DistilBert), replacing (Electra), autoregressive (XLnet), masked&permuted (Mp-net), text-text (T5) and large language models (Llama2 and Llama3) under low source settings. We find that they perform competitively on explicit harassment binary detection, however, their performance is substantially lower on harassment and denigration multi-classification tasks. Therefore, we propose an emotion-adaptive training framework (EAT) that helps transfer knowledge from the domain of emotion detection to the domain of cyberbullying detection to help detect indirect cyberbullying events. EAT consistently improves the average macro F1, precision and recall by 20% in cyberbullying detection tasks across nine transformer-based models under low-resource settings. Our claims are supported by intuitive theoretical insights and extensive experiments. Peiling Yi, Arkaitz Zubiaga |
ICWSM | 1 |
| 2023 | Learning like human annotators: Cyberbullying detection in lengthy social media sessionsabstractThe inherent characteristic of cyberbullying of being a recurrent attitude calls for the investigation of the problem by looking at social media sessions as a whole, beyond just isolated social media posts. However, the lengthy nature of social media sessions challenges the applicability and performance of session-based cyberbullying detection models. This is especially true when one aims to use state-of-the-art Transformer-based pre-trained language models, which only take inputs of a limited length. In this paper, we address this limitation of transformer models by proposing a conceptually intuitive framework called LS-CB, which enables cyberbullying detection from lengthy social media sessions. LS-CB relies on the intuition that we can effectively aggregate the predictions made by transformer models on smaller sliding windows extracted from lengthy social media sessions, leading to an overall improved performance. Our extensive experiments with six transformer models on two session-based datasets show that LS-CB consistently outperforms three types of competitive baselines including state-of-the-art cyberbullying detection models. In addition, we conduct a set of qualitative analyses to validate the hypotheses that cyberbullying incidents can be detected through aggregated analysis of smaller chunks derived from lengthy social media sessions (H1), and that cyberbullying incidents can occur at different points of the session (H2), hence positing that frequently used text truncation strategies are suboptimal compared to relying on holistic views of sessions. Our research in turn opens an avenue for fine-grained cyberbullying detection within sessions in future work. Peiling Yi, Arkaitz Zubiaga |
WWW | 1 |
| 2022 | Cyberbullying Detection across Social Media Platforms via Platform-Aware Adversarial Encoding
Peiling Yi, Arkaitz Zubiaga |
ICWSM | 1 |