VLDB 2026 Research / reviewers in the wild / expert
Wenchang Ma
dblp:277/0953
· DBLP profile ↗
7ranked-venue papers
1as first author
7since 2021 · last 2024
0009-0006-8264-4775ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | General Debiasing for Graph-based Collaborative Filtering via Adversarial Graph DropoutabstractGraph neural networks (GNNs) have shown impressive performance in recommender systems, particularly in collaborative filtering (CF). The key lies in aggregating neighborhood information on a user-item interaction graph to enhance user/item representations. However, we have discovered that this aggregation mechanism comes with a drawback - it amplifies biases present in the interaction graph. For instance, a user's interactions with items can be driven by both unbiased true interest and various biased factors like item popularity or exposure. However, the current aggregation approach combines all information, both biased and unbiased, leading to biased representation learning. Consequently, graph-based recommenders can learn distorted views of users/items, hindering the modeling of their true preferences and generalizations. An Zhang 0003, Wenchang Ma, Leheng Sheng, Xiang Wang 0010 |
WWW | 2 |
| 2024 | Robust Collaborative Filtering to Popularity Distribution ShiftabstractIn leading collaborative filtering (CF) models, representations of users and items are prone to learn popularity bias in the training data as shortcuts. The popularity shortcut tricks are good for in-distribution (ID) performance but poorly generalized to out-of-distribution (OOD) data, i.e., when popularity distribution of test data shifts w.r.t. the training one. To close the gap, debiasing strategies try to assess the shortcut degrees and mitigate them from the representations. However, there exist two deficiencies: (1) when measuring the shortcut degrees, most strategies only use statistical metrics on a single aspect (i.e., item frequency on item and user frequency on user aspect), failing to accommodate the compositional degree of a user–item pair; (2) when mitigating shortcuts, many strategies assume that the test distribution is known in advance. This results in low-quality debiased representations. Worse still, these strategies achieve OOD generalizability with a sacrifice on ID performance. In this work, we present a simple yet effective debiasing strategy, PopGo , which quantifies and reduces the interaction-wise popularity shortcut without any assumptions on the test data. It first learns a shortcut model, which yields a shortcut degree of a user–item pair based on their popularity representations. Then, it trains the CF model by adjusting the predictions with the interaction-wise shortcut degrees. By taking both causal- and information-theoretical looks at PopGo, we can justify why it encourages the CF model to capture the critical popularity-agnostic features while leaving the spurious popularity-relevant patterns out. We use PopGo to debias two high-performing CF models (matrix factorization [ 28 ] and LightGCN [ 19 ]) on four benchmark datasets. On both ID and OOD test sets, PopGo achieves significant gains over the state-of-the-art debiasing strategies (e.g., DICE [ 71 ] and MACR [ 58 ]). Codes and datasets are available at https://github.com/anzhang314/PopGo . An Zhang 0003, Wenchang Ma, Jingnan Zheng, Xiang Wang 0010, Tat-Seng Chua |
ACM Trans. Inf. Syst. | 2 |
| 2023 | Boosting Causal Discovery via Adaptive Sample Reweighting
An Zhang 0003, Fangfu Liu, Wenchang Ma, Zhibo Cai, Xiang Wang 0010, Tat-Seng Chua |
ICLR | 3 |
| 2023 | Discovering Dynamic Causal Space for DAG Structure LearningabstractDiscovering causal structure from purely observational data (i.e., causal discovery), aiming to identify causal relationships among variables, is a fundamental task in machine learning.The recent invention of differentiable score-based DAG learners is a crucial enabler, which reframes the combinatorial optimization problem into a differentiable optimization with a DAG constraint over directed graph space. Despite their great success, these cutting-edge DAG learners incorporate DAG-ness independent score functions to evaluate the directed graph candidates, lacking in considering graph structure. As a result, measuring the data fitness alone regardless of DAG-ness inevitably leads to discovering suboptimal DAGs and model vulnerabilities. Fangfu Liu, Wenchang Ma, An Zhang 0003, Xiang Wang 0010, Yueqi Duan, Tat-Seng Chua |
KDD | 2 |
| 2022 | Incorporating Bias-aware Margins into Contrastive Loss for Collaborative FilteringabstractCollaborative filtering (CF) models easily suffer from popularity bias, which makes recommendation deviate from users’ actual preferences. However, most current debiasing strategies are prone to playing a trade-off game between head and tail performance, thus inevitably degrading the overall recommendation accuracy. To reduce the negative impact of popularity bias on CF models, we incorporate Bias-aware margins into Contrastive loss and propose a simple yet effective BC Loss, where the margin tailors quantitatively to the bias degree of each user-item interaction. We investigate the geometric interpretation of BC loss, then further visualize and theoretically prove that it simultaneously learns better head and tail representations by encouraging the compactness of similar users/items and enlarging the dispersion of dissimilar users/items. Over six benchmark datasets, we use BC loss to optimize two high-performing CF models. In various evaluation settings (i.e., imbalanced/balanced, temporal split, fully-observed unbiased, tail/head test evaluations), BC loss outperforms the state-of-the-art debiasing and non-debiasing methods with remarkable improvements. Considering the theoretical guarantee and empirical success of BC loss, we advocate using it not just as a debiasing strategy, but also as a standard loss in recommender models. Codes are available at https://github.com/anzhang314/BC-Loss. An Zhang 0003, Wenchang Ma, Xiang Wang 0010, Tat-Seng Chua |
NeurIPS | 2 |
| 2021 | CR-Walker: Tree-Structured Graph Reasoning and Dialog Acts for Conversational RecommendationabstractGrowing interests have been attracted in Conversational Recommender Systems (CRS), which explore user preference through conversational interactions in order to make appropriate recommendation.However, there is still a lack of ability in existing CRS to ( 1) traverse multiple reasoning paths over background knowledge to introduce relevant items and attributes, and (2) arrange selected entities appropriately under current system intents to control response generation.To address these issues, we propose CR-Walker in this paper, a model that performs tree-structured reasoning on a knowledge graph, and generates informative dialog acts to guide language generation.The unique scheme of tree-structured reasoning views traversed entity at each hop as part of dialog acts to facilitate language generation, which links how entities are selected and expressed.Automatic and human evaluations show that CR-Walker can arrive at more accurate recommendation, and generate more informative and engaging responses. Wenchang Ma, Ryuichi Takanobu, Minlie Huang |
EMNLP (1) | 1 |
| 2021 | Extract, Denoise and Enforce: Evaluating and Improving Concept Preservation for Text-to-Text GenerationabstractPrior studies on text-to-text generation typically assume that the model could figure out what to attend to in the input and what to include in the output via seq2seq learning, with only the parallel training data and no additional guidance.However, it remains unclear whether current models can preserve important concepts in the source input, as seq2seq learning does not have explicit focus on the concepts and commonly used evaluation metrics also treat concepts equally important as other tokens.In this paper, we present a systematic analysis that studies whether current seq2seq models, especially pre-trained language models, are good enough for preserving important input concepts and to what extent explicitly guiding generation with the concepts as lexical constraints is beneficial.We answer the above questions by conducting extensive analytical experiments on four representative text-to-text generation tasks.Based on the observations, we then propose a simple yet effective framework to automatically extract, denoise, and enforce important input concepts as lexical constraints.This new method performs comparably or better than its unconstrained counterpart on automatic metrics, demonstrates higher coverage for concept preservation, and receives better ratings in the human evaluation. 1 Yuning Mao, Wenchang Ma, Deren Lei, Jiawei Han 0001, Xiang Ren 0001 |
EMNLP (1) | 2 |