EDBT 2026 Demo / reviewers in the wild / expert
Xun Yao
dblp:148/3846
· DBLP profile ↗
9ranked-venue papers
7as first author
7since 2021 · last 2026
0000-0001-5633-8113ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 4 first-author · 5 since 2021Databases, data management, data science and information retrieval · 5 · 4 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DENOTER: Dual cascadEd iNformatiOn filTERing for robust sequential recommendationabstractSequential recommendation systems aim to predict users’ next interactions by analyzing historical sequences of their behaviors. However, these systems are vulnerable to adversarial attacks that disrupt input sequences through the injection of noisy or irrelevant interactions. Accordingly, this paper introduces the D ual cascad E d i N formati O n fil TER ing ( DENOTER ) algorithm designed to mitigate the impact of adversarial attacks. Unlike existing recommendation methods, which often overlook the varying importance of different components ( i.e ., items and their associated features) within behavioral sequences, DENOTER employs a two-tier adaptive filtering mechanism. Specifically, at the feature level, a dimension-adjustable, forward-only stochastic model is introduced to dynamically select and preserve essential features, thereby effectively filtering out malicious perturbations. At the item level, a Gumbel-driven sampling strategy is leveraged to selectively retain salient items. Through multi-layer refinement with alternating feature- and item-level filtering, DENOTER progressively reduces the influence of adversarially perturbed items and features within input sequences, enabling robust recommendation without specialized data augmentation or model retraining. Extensive empirical evaluations on five datasets and four encoders demonstrate that DENOTER consistently outperforms existing baselines under diverse adversarial conditions, achieving superior robustness and accuracy even in low-resource and out-of-domain settings. Xun Yao, Wangli Yang, Xinrong Hu, Jie Yang 0009, Yi Guo 0001 |
Pattern Recognit. | 1 |
| 2024 | Masking the Unknown: Leveraging Masked Samples for Enhanced Data AugmentationabstractData Augmentation (DA) has become a widely adopted strategy for addressing data scarcity in numerous NLP tasks, especially in scenarios with limited resources or imbalanced classes. However, many existing augmentation techniques rely on randomness or additional resources, presenting challenges in both performance and practical implementation. Furthermore, there is a lack of exploration into what constitutes effective augmentation. In this paper, we systematically evaluate existing DA methods across a comprehensive range of text-classification benchmarks. The empirical analysis highlights that the most significant change resulting from augmentation is observed in the data variance. This observation inspires the proposed approach, termed Mask-for-Data Augmentation (M4DA), which strategically masks tokens from original samples for augmentation. Specifically, M4DA consists of a Variance-Oriented Masker Module (VMM), which ensures an increase in data variances, and a Complexity-Enhanced Selection Module (CSM), designed to select the augmented sample with the highest semantic complexity. The effectiveness of the proposed method is empirically validated across various text-classification benchmarks, including scenarios with limited or full resources and imbalanced classes. Experimental results demonstrate considerable improvements over state-of-the-arts. Xun Yao, Zijian Huang 0012, Xinrong Hu, Jie Yang 0009, Yi Guo 0001 |
UAI | 1 |
| 2024 | SCAD: Subspace Clustering based Adversarial DetectorabstractAdversarial examples pose significant challenges for Natural Language Processing (NLP) model robustness, often causing notable performance degradation. While various detection methods have been proposed with the aim of differentiating clean and adversarial inputs, they often require fine-tuning with ample data, which is problematic for low-resource scenarios. To alleviate this issue, a Subspace Clustering based Adversarial Detector (termed SCAD) is proposed in this paper, leveraging a union of subspaces to model the clean data distribution. Specifically, SCAD estimates feature distribution across semantic subspaces, assigning unseen examples to the nearest one for effective discrimination. The construction of semantic subspaces does not require many observations and hence ideal for the low-resource setting. Xinrong Hu, Wushuan Chen, Jie Yang 0009, Yi Guo 0001, Xun Yao, Bangchao Wang, Junping Liu |
WSDM | 5 |
| 2024 | COTER: Conditional Optimal Transport meets Table RetrievalabstractAd hoc table retrieval refers to the task of performing semantic matching between given queries and candidate tables. In recent years, the approach to addressing this retrieval task has undergone significant shifts, transitioning from utilizing hand-crafted features to leveraging the power of Pre-trained Language Models (PLMs). However, key challenges arise when candidate tables contain shared items, and/or queries may refer to only a subset of table items rather than the entire one. Existing models often struggle to distinguish the most informative items and fail to accurately identify the relevant items required to match with the query. Xun Yao, Xinrong Hu, Jie Yang 0009, Yi Guo 0001, Daniel (Dianliang) Zhu |
WSDM | 1 |
| 2023 | Improving Adversarially Robust Sequential Recommendation through Generalizable PerturbationsabstractSequential recommendation is of great importance for a variety of purposes, such as application engineering, resource optimization, and marketing. Yet, existing sequence-based recommendation models are susceptible to adversarial attacks, which aim to perturb input sequences and mislead trained models, resulting in incorrect predictions. Defense methods are accordingly adopted to enhance model robustness. Nevertheless, these methods encounter challenges, such as error propagation (from the model output to generate adversarial samples), the high system complexity, and the difficulty of maintaining the model generalizability. To bridge this gap, this paper introduces a simple yet effective adversarial defense algorithm, termed Perturbation-Driven Sequential Recommendation (PDSR). In the training process, PDSR leverages a simple perturbation-generation module to create adversarial samples, eliminating the need for gradient estimation, thus streamlining the process. Additionally, it also incorporates a robust encoder designed to increase tolerance towards representation variations by ensuring alignment between original and perturbed representations, thereby boosting model generalizability. Comprehensive experiments are conducted based on a combination of five benchmark datasets, two attack methods, and four sequential recommendation models. When compared to four state-of-the-art defense baselines, PDSR demonstrates notable improvements in defense performance. Xun Yao, Ruyi He, Xinrong Hu, Jie Yang 0009, Yi Guo 0001, Zijian Huang 0012 |
IEEE Big Data | 1 |
| 2023 | Towards Robust Token Embeddings for Extractive Question Answering
Xun Yao, Junlong Ma, Xinrong Hu, Jie Yang 0009, Yi Guo 0001, Junping Liu |
WISE | 1 |
| 2023 | CREAM: Named Entity Recognition with Concise query and REgion-Aware Minimization
Xun Yao, Xinrong Hu, Jie Yang 0009, Yi Guo 0001 |
WISE | 1 |
| 2012 | An Integration of Medical Information Retrieval into undergraduate PBL curriculum
Ping Qing, Xun Yao, Xuehong Wan |
AMIA | 2 |
| 2012 | One-year Outcome of An EHR-based Case Learning System for Medical Students, Residents and Doctors
Xun Yao, Xiaoyan Lv |
AMIA | 1 |