EDBT 2026 Demo / reviewers in the wild / expert
Shixiang Zhu
dblp:133/3853
· DBLP profile ↗
4ranked-venue papers in the field
1as first author
4since 2021 · last 2025
0000-0002-2241-6096ORCID · corroborated
Domains — venue-derived; a paper can count in several
Data Mining & Knowledge Discovery · 3Database Systems & Data Management · 1 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Conditional Generative Modeling for High-dimensional Marked Temporal Point Processes
Zheng Dong 0005, Zekai Fan, Shixiang Zhu |
KDD (1) | 3 |
| 2025 | Counterfactual Fairness Through Transforming Data Orthogonal to BiasabstractMachine learning models have demonstrated exceptional capabilities in solving complex problems across a variety of domains. However, these models can sometimes exhibit biased decision-making, leading to unequal treatment of different groups. Despite substantial research on counterfactual fairness, existing methods remain underdeveloped in addressing the impact of multivariate and continuous sensitive variables on decision-making outcomes. To tackle this gap, we propose a novel data pre-processing algorithm, Orthogonal to Bias (OB), which is designed to eliminate the influence of a group of continuous sensitive variables, thereby promoting counterfactual fairness in machine learning applications. Our approach, based on the assumption of an elliptical distribution within a structural causal model (SCM), shows that counterfactual fairness can be achieved by ensuring the data is orthogonal to the observed sensitive variables. The OB algorithm is model-agnostic, making it applicable to a wide range of machine learning models and tasks. To enhance numerical stability, we also introduce a sparse variant that incorporates regularization. Empirical evaluations on both simulated and real-world datasets-spanning scenarios with both discrete and continuous sensitive variables-demonstrate that our method effectively promotes fairer outcomes without compromising predictive accuracy. Shixiang Zhu |
KDD (2) | 2 |
| 2024 | Counterfactual Generative Models for Time-Varying TreatmentsabstractEstimating the counterfactual outcome of treatment is essential for decision-making in public health and clinical science, among others. Often, treatments are administered in a sequential, time-varying manner, leading to an exponentially increased number of possible counterfactual outcomes. Furthermore, in modern applications, the outcomes are high-dimensional and conventional average treatment effect estimation fails to capture disparities in individuals. To tackle these challenges, we propose a novel conditional generative framework capable of producing counterfactual samples under time-varying treatment, without the need for explicit density estimation. Our method carefully addresses the distribution mismatch between the observed and counterfactual distributions via a loss function based on inverse probability re-weighting, and supports integration with state-of-the-art conditional generative models such as the guided diffusion and conditional variational autoencoder. We present a thorough evaluation of our method using both synthetic and real-world data. Our results demonstrate that our method is capable of generating high-quality counterfactual samples and outperforms the state-of-the-art baselines. Shenghao Wu, Wenbin Zhou 0002, Minshuo Chen, Shixiang Zhu |
KDD | 4 |
| 2022 | Imitation Learning of Neural Spatio-Temporal Point ProcessesabstractWe present a novel Neural Embedding Spatio-Temporal (NEST) point process model for spatio-temporal discrete event data and develop an efficient imitation learning (a type of reinforcement learning) based approach for model fitting. Despite the rapid development of one-dimensional temporal point processes for discrete event data, the study of spatial-temporal aspects of such data is relatively scarce. Our model captures complex spatio-temporal dependence between discrete events by carefully design a mixture of heterogeneous Gaussian diffusion kernels, whose parameters are parameterized by neural networks. This new kernel is the key that our model can capture intricate spatial dependence patterns and yet still lead to interpretable results as we examine maps of Gaussian diffusion kernel parameters. The imitation learning model fitting for the NEST is more robust than the maximum likelihood estimate. It directly measures the divergence between the empirical distributions between the training data and the model-generated data. Moreover, our imitation learning-based approach enjoys computational efficiency due to the explicit characterization of the reward function related to the likelihood function; furthermore, the likelihood function under our model enjoys tractable expression due to Gaussian kernel parameterization. Experiments based on real data show our method’s good performance relative to the state-of-the-art and the good interpretability of NEST’s result. Shixiang Zhu, Shuang Li 0002, Zhigang Peng, Yao Xie 0002 |
IEEE Trans. Knowl. Data Eng. | 1 |