Chang Gong 0001

dblp:215/2767-1 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
7since 2021 · last 2026
0000-0001-7941-5011ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 PORCA: Root Cause Analysis with Partially Observed Data
Chang Gong 0001, Di Yao 0001, Jin Wang 0007, Wenbin Li 0012, Lanting Fang, Yongtao Xie, Kaiyu Feng, Peng Han 0005, Jingping Bi
ICDE1
2026 DPBL: Denoised Player Behavior Representation Learning
abstract
The video game industry has emerged as a significant economic force, driving extensive research on optimizing the gaming environment and improving gaming experiences. Among these endeavors, player behavior representation learning has become a critical way to model valuable player properties and is beneficial for a wide range of downstream tasks. However, some common factors, such as login rewards and daily tasks, can trigger similar behaviors among different players, which are informative and noisy for learning high-quality player behavior representations. Existing methods ignore the low signal-to-noise ratio in player behavior data and waste too much modeling capacity on less informative behaviors, resulting in their learned representations being noisy. In this paper, we propose a novel model for Denoised Player Behavior representation Learning, namely DPBL, which consists of two key modules. The first module extracts various player behavior patterns and isolates them from less informative noise. The second module utilizes the extracted patterns to refine the embedding of each behavior and eliminates noise. To optimize DPBL, two contrastive learning strategies are proposed to identify the noise that should be eliminated and to learn distinguishable representations, respectively. With the above design, DPBL is capable of mitigating the impact of noise in the data and learning high-quality representations that effectively capture player characteristics. We conducted extensive experiments on two real-world datasets, and DPBL outperforms all baselines on various downstream tasks with an improvement of 1.4% ∼ 18.1%. The results also show that DPBL achieves an improvement of 5.6% ∼ 23.0% in the denoising experiments, which proves that DPBL is more robust to noisy behaviors. Code is available at https://github.com/LwbXc/DPBL.
Wenbin Li 0012, Di Yao 0001, Zijie Xu 0006, Chang Gong 0001, Quanliang Jing, Runze Wu 0001, Haining Tan, Zhipeng Hu, Tangjie Lv, Changjie Fan, Jingping Bi
IEEE Trans. Games4
2024 CausalTAD: Causal Implicit Generative Model for Debiased Online Trajectory Anomaly Detection
abstract
Trajectory anomaly detection, aiming to estimate the anomaly risk of trajectories given the Source-Destination (SD) pairs, has become a critical problem for many real-world applications. Existing solutions directly train a generative model for observed trajectories and calculate the conditional generative probability$P(T \vert C)$as the anomaly risk, where$T$and$C$represent the trajectory and SD pair respectively. However, we argue that the observed trajectories are confounded by road network preference which is a common cause of both SD distribution and trajectories. Existing methods ignore this issue limiting their generalization ability on out-of-distribution trajectories. In this paper, we define the debiased trajectory anomaly detection problem and propose a causal implicit generative model, namely CausalTAD, to solve it. CausalTAD adopts do-calculus to eliminate the confounding bias of road network preference and estimates$P(T\vert do(C))$as the anomaly criterion. Extensive experiments show that CausalTadcan not only achieve superior performance on trained trajectories but also generally improve the performance of out-of-distribution data, with improvements of 2.1% ~ 5.7% and 10.6% ~ 32.7% respectively.
Wenbin Li 0012, Di Yao 0001, Chang Gong 0001, Xiaokai Chu, Quanliang Jing, Yunxia Fan, Jingping Bi
ICDE3
2024 CausalMMM: Learning Causal Structure for Marketing Mix Modeling
abstract
In online advertising, marketing mix modeling (MMM) is employed to predict the gross merchandise volume (GMV) of brand shops and help decision-makers to adjust the budget allocation of various advertising channels. Traditional MMM methods leveraging regression techniques can fail in handling the complexity of marketing. Although some efforts try to encode the causal structures for better prediction, they have the strict restriction that causal structures are prior-known and unchangeable. In this paper, we define a new causal MMM problem that automatically discovers the interpretable causal structures from data and yields better GMV predictions. To achieve causal MMM, two essential challenges should be addressed: (1) Causal Heterogeneity. The causal structures of different kinds of shops vary a lot. (2) Marketing Response Patterns. Various marketing response patterns i.e., carryover effect and shape effect, have been validated in practice. We argue that causal MMM needs dynamically discover specific causal structures for different shops and the predictions should comply with the prior known marketing response patterns. Thus, we propose CausalMMM that integrates Granger causality in a variational inference framework to measure the causal relationships between different channels and predict the GMV with the regularization of both temporal and saturation marketing response patterns. Extensive experiments show that CausalMMM can not only achieve superior performance of causal structure learning on synthetic datasets with improvements of 5.7%\sim 7.1%, but also enhance the GMV prediction results on a representative E-commerce platform.
Chang Gong 0001, Di Yao 0001, Lei Zhang 0206, Wenbin Li 0012, Yueyang Su, Jingping Bi
WSDM1
2023 Causal Discovery from Temporal Data
abstract
Temporal data representing chronological observations of complex systems can be ubiquitously collected in smart industry, medicine, finance and etc. In the last decade, many tasks have been studied for mining temporal data and offered significant value for various applications. Among these tasks, causal discovery aims to understand the underlying generation mechanism of temporal data and has attracted much research attention. According to whether the data is calibrated, existing causal discovery approaches can be divided into two subtasks, i.e., multivariate time-series causal discovery, and event sequence causal discovery. Previous tutorials or surveys have primarily focused on causal discovery from time-series data and disregarded the second ones. In this tutorial, we elucidate the correlation between the two subtasks and provide a comprehensive review of the existing solutions. Moreover, we offer some potential applications and summarize new perspectives for discovering causal relations from temporal data. We hope the audiences can obtain a systematic overview of this topic and inspire some new ideas for their own research.
Chang Gong 0001, Di Yao 0001, Chuzhe Zhang, Wenbin Li 0012, Jingping Bi, Lun Du, Jin Wang 0007
KDD1
2022 CausalMTA: Eliminating the User Confounding Bias for Causal Multi-touch Attribution
abstract
Multi-touch attribution (MTA), aiming to estimate the contribution of each advertisement touchpoint in conversion journeys, is essential for budget allocation and automatically advertising. Existing methods first train a model to predict the conversion probability of the advertisement journeys with historical data and calculate the attribution of each touchpoint by using the results counterfactual predictions. An assumption of these works is the conversion prediction model is unbiased. It can give accurate predictions on any randomly assigned journey, including both the factual and counterfactual ones. Nevertheless, this assumption does not always hold as the user preferences act as the common cause for both ad generation and user conversion, involving the confounding bias and leading to an out-of-distribution (OOD) problem in the counterfactual prediction. In this paper, we define the causal MTA task and propose CausalMTA to solve this problem. It systemically eliminates the confounding bias from both static and dynamic perspectives and learn an unbiased conversion prediction model using historical data. We also provide a theoretical analysis to prove the effectiveness of CausalMTA with sufficient ad journeys. Extensive experiments on both synthetic and real data in Alibaba advertising platform show that CausalMTA can not only achieve better prediction performance than the state-of-the-art method but also generate meaningful attribution credits across different advertising channels.
Di Yao 0001, Chang Gong 0001, Lei Zhang 0206, Jingping Bi
KDD2
2021 TrajCross: Trajecotry Cross-Modal Retrieval with Contrastive Learning
abstract
In this paper, we propose a new task namely trajectory cross-modal retrieval which achieves the cross-modal search between coordinate trajectories and images containing trajectories. Nevertheless, trajectory cross-modal retrieval is rather challenging in learning the representations of each modality and reduce the cross-domain discrepancy caused by the inconsistent data distribution at the same time. we proposes a cross-modal retrieval model TrajCross based on multi-level representation for trajectory cross-modal retrieval. Specifically, TrajCross extracts the location features and the shape information respectively for the represention of multi-modal data. we adopt a contrastive learning method to achieve semantic preservation among similar multi-modal data. Extensive experiments show that TrajCross significantly outperforms state-of-the-art cross-modal retrieval methods.
Quanliang Jing, Di Yao 0001, Chang Gong 0001, Xinxin Fan, Haining Tan, Jingping Bi
IEEE BigData3