Shixiang Zhu

dblp:133/3853 · DBLP profile ↗
← Back
23ranked-venue papers
9as first author
15since 2021 · last 2025
0000-0002-2241-6096ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 3 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 3 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Computer networks · 2Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Black-Box Optimization with Implicit Constraints for Public Policy
abstract
Black-box optimization (BBO) has become increasingly relevant for tackling complex decision-making problems, especially in public policy domains such as police redistricting. However, its broader application in public policymaking is hindered by the complexity of defining feasible regions and the high-dimensionality of decisions. This paper introduces a novel BBO framework, termed as the Conditional And Generative Black-box Optimization (CageBO). This approach leverages a conditional variational autoencoder to learn the distribution of feasible decisions, enabling a two-way mapping between the original decision space and a simplified, constraint-free latent space. The CageBO efficiently handles the implicit constraints often found in public policy applications, allowing for optimization in the latent space while evaluating objectives in the original space. We validate our method through a case study on large-scale police redistricting problems in Atlanta, Georgia. Our results reveal that our CageBO offers notable improvements in performance and efficiency compared to the baselines.
Wenqian Xing, Shixiang Zhu
AAAI4
2025 New User Event Prediction Through the Lens of Causal Inference
abstract
Modeling and analysis for event series generated by users of heterogeneous behavioral patterns are closely involved in our daily lives, including credit card fraud detection, online platform user recommendation, and social network analysis. The most commonly adopted approach to this task is to assign users to behavior-based categories and analyze each of them separately. However, this requires extensive data to fully understand the user behavior, presenting challenges in modeling newcomers without significant historical knowledge. In this work, we propose a novel discrete event prediction framework for new users with limited history, without needing to know the user’s category. We treat the user event history as the "treatment" for future events and the user category as the key confounder. Thus, the prediction problem can be framed as counterfactual outcome estimation, where each event is re-weighted by its inverse propensity score. We demonstrate the improved performance of the proposed framework with a numerical simulation study and two real-world applications, including Netflix rating prediction and seller contact prediction for customer support at Amazon.
Henry Shaowu Yuchi, Shixiang Zhu, Yigit M. Arisoy, Matthew C. Spencer
AISTATS2
2025 Recurrent Neural Goodness-of-Fit Test for Time Series
abstract
Time series data are crucial across diverse domains such as finance and healthcare, where accurate forecasting and decision-making rely on advanced modeling techniques. While generative models have shown great promise in capturing the intricate dynamics inherent in time series, evaluating their performance remains a major challenge. Traditional evaluation metrics fall short due to the temporal dependencies and potential high dimensionality of the features. In this paper, we propose the REcurrent NeurAL (RENAL) Goodness-of-Fit test, a novel and statistically rigorous framework for evaluating generative time series models. By leveraging recurrent neural networks, we transform the time series into conditionally independent data pairs, enabling the application of a chi-square-based goodness-of-fit test to the temporal dependencies within the data. This approach offers a robust, theoretically grounded solution for assessing the quality of generative models, particularly in settings with limited time sequences. We demonstrate the efficacy of our method across both synthetic and real-world datasets, outperforming existing methods in terms of reliability and accuracy. Our method fills a critical gap in the evaluation of time series generative models, offering a tool that is both practical and adaptable to high-stakes applications.
Wenbin Zhou 0002, Liyan Xie, Shixiang Zhu
AISTATS4
2025 Uncertainty-Aware Robust Learning on Noisy Graphs
abstract
Graph neural networks (GNNs) have excelled in various graph learning tasks, particularly node classification. However, their performance is often hampered by noisy measurements in real-world graphs, which can corrupt critical patterns in the data. To address this, we propose a novel uncertainty-aware graph learning framework inspired by distributionally robust optimization. Specifically, we use a graph neural network-based encoder to embed the node features and find the optimal node embeddings by minimizing the worst-case risk through a minimax formulation. Such an uncertainty-aware learning process leads to improved node representations and a more robust graph predictive model that effectively mitigates the impact of uncertainty arising from data noise. Our experimental results demonstrate superior predictive performance over baselines across noisy scenarios.
Kaize Ding, Shixiang Zhu
ICASSP3
2025 Conditional Generative Modeling for High-dimensional Marked Temporal Point Processes
Zheng Dong 0005, Zekai Fan, Shixiang Zhu
KDD (1)3
2025 Counterfactual Fairness Through Transforming Data Orthogonal to Bias
abstract
Machine learning models have demonstrated exceptional capabilities in solving complex problems across a variety of domains. However, these models can sometimes exhibit biased decision-making, leading to unequal treatment of different groups. Despite substantial research on counterfactual fairness, existing methods remain underdeveloped in addressing the impact of multivariate and continuous sensitive variables on decision-making outcomes. To tackle this gap, we propose a novel data pre-processing algorithm, Orthogonal to Bias (OB), which is designed to eliminate the influence of a group of continuous sensitive variables, thereby promoting counterfactual fairness in machine learning applications. Our approach, based on the assumption of an elliptical distribution within a structural causal model (SCM), shows that counterfactual fairness can be achieved by ensuring the data is orthogonal to the observed sensitive variables. The OB algorithm is model-agnostic, making it applicable to a wide range of machine learning models and tasks. To enhance numerical stability, we also introduce a sparse variant that incorporates regularization. Empirical evaluations on both simulated and real-world datasets-spanning scenarios with both discrete and continuous sensitive variables-demonstrate that our method effectively promotes fairer outcomes without compromising predictive accuracy.
Shixiang Zhu
KDD (2)2
2025 Topology-Aware Conformal Prediction for Stream Networks
abstract
Stream networks, a unique class of spatiotemporal graphs, exhibit complex directional flow constraints and evolving dependencies, making uncertainty quantification a critical yet challenging task. Traditional conformal prediction methods struggle in this setting due to the need for joint predictions across multiple interdependent locations and the intricate spatio-temporal dependencies inherent in stream networks. Existing approaches either neglect dependencies, leading to overly conservative predictions, or rely solely on data-driven estimations, failing to capture the rich topological structure of the network. To address these challenges, we propose Spatio-Temporal Adaptive Conformal Inference (STACI), a novel framework that integrates network topology and temporal dynamics into the conformal prediction framework. STACI introduces a topology-aware nonconformity score that respects directional flow constraints and dynamically adjusts prediction sets to account for temporal distributional shifts. We provide theoretical guarantees on the validity of our approach and demonstrate its superior performance on both synthetic and real-world datasets. Our results show that STACI effectively balances prediction efficiency and coverage, outperforming existing conformal prediction methods for stream networks.
Jifan Zhang, Fangxin Wang 0003, Zihe Song 0001, Philip S. Yu, Kaize Ding, Shixiang Zhu
NeurIPS6
2024 Counterfactual Generative Models for Time-Varying Treatments
abstract
Estimating the counterfactual outcome of treatment is essential for decision-making in public health and clinical science, among others. Often, treatments are administered in a sequential, time-varying manner, leading to an exponentially increased number of possible counterfactual outcomes. Furthermore, in modern applications, the outcomes are high-dimensional and conventional average treatment effect estimation fails to capture disparities in individuals. To tackle these challenges, we propose a novel conditional generative framework capable of producing counterfactual samples under time-varying treatment, without the need for explicit density estimation. Our method carefully addresses the distribution mismatch between the observed and counterfactual distributions via a loss function based on inverse probability re-weighting, and supports integration with state-of-the-art conditional generative models such as the guided diffusion and conditional variational autoencoder. We present a thorough evaluation of our method using both synthetic and real-world data. Our results demonstrate that our method is capable of generating high-quality counterfactual samples and outperforms the state-of-the-art baselines.
Shenghao Wu, Wenbin Zhou 0002, Minshuo Chen, Shixiang Zhu
KDD4
2022 Neural Spectral Marked Point Processes
Shixiang Zhu, Haoyun Wang, Zheng Dong 0005, Xiuyuan Cheng, Yao Xie 0002
ICLR1
2022 Distributionally robust weighted k-nearest neighbors
abstract
Learning a robust classifier from a few samples remains a key challenge in machine learning. A major thrust of research has been focused on developing k-nearest neighbor (k-NN) based algorithms combined with metric learning that captures similarities between samples. When the samples are limited, robustness is especially crucial to ensure the generalization capability of the classifier. In this paper, we study a minimax distributionally robust formulation of weighted k-nearest neighbors, which aims to find the optimal weighted k-NN classifiers that hedge against feature uncertainties. We develop an algorithm, Dr.k-NN, that efficiently solves this functional optimization problem and features in assigning minimax optimal weights to training samples when performing classification. These weights are class-dependent, and are determined by the similarities of sample features under the least favorable scenarios. When the size of the uncertainty set is properly tuned, the robust classifier has a smaller Lipschitz norm than the vanilla k-NN, and thus improves the generalization capability. We also couple our framework with neural-network-based feature embedding. We demonstrate the competitive performance of our algorithm compared to the state-of-the-art in the few-training-sample setting with various real-data experiments.
Shixiang Zhu, Liyan Xie, Minghe Zhang, Rui Gao 0001, Yao Xie 0002
NeurIPS1
2022 Spatio-Temporal Point Processes With Attention for Traffic Congestion Event Modeling
abstract
We present a novel framework for modeling traffic congestion events over road networks. Using multi-modal data by combining count data from traffic sensors with police reports that report traffic incidents, we aim to capture two types of triggering effect for congestion events. Current traffic congestion at one location may cause future congestion over the road network, and traffic incidents may cause spread traffic congestion. To model the non-homogeneous temporal dependence of the event on the past, we use a novel attention-based mechanism based on neural networks embedding for point processes. To incorporate the directional spatial dependence induced by the road network, we adapt the “tail-up” model from the context of spatial statistics to the traffic network setting. We demonstrate our approach’s superior performance compared to the state-of-the-art methods for both synthetic and real data.
Shixiang Zhu, Ruyi Ding, Minghe Zhang, Pascal Van Hentenryck, Yao Xie 0002
IEEE Trans. Intell. Transp. Syst.1
2022 Imitation Learning of Neural Spatio-Temporal Point Processes
abstract
We present a novel Neural Embedding Spatio-Temporal (NEST) point process model for spatio-temporal discrete event data and develop an efficient imitation learning (a type of reinforcement learning) based approach for model fitting. Despite the rapid development of one-dimensional temporal point processes for discrete event data, the study of spatial-temporal aspects of such data is relatively scarce. Our model captures complex spatio-temporal dependence between discrete events by carefully design a mixture of heterogeneous Gaussian diffusion kernels, whose parameters are parameterized by neural networks. This new kernel is the key that our model can capture intricate spatial dependence patterns and yet still lead to interpretable results as we examine maps of Gaussian diffusion kernel parameters. The imitation learning model fitting for the NEST is more robust than the maximum likelihood estimate. It directly measures the divergence between the empirical distributions between the training data and the model-generated data. Moreover, our imitation learning-based approach enjoys computational efficiency due to the explicit characterization of the reward function related to the likelihood function; furthermore, the likelihood function under our model enjoys tractable expression due to Gaussian kernel parameterization. Experiments based on real data show our method’s good performance relative to the state-of-the-art and the good interpretability of NEST’s result.
Shixiang Zhu, Shuang Li 0002, Zhigang Peng, Yao Xie 0002
IEEE Trans. Knowl. Data Eng.1
2021 Goodness-of-Fit Test for Mismatched Self-Exciting Processes
abstract
Recently there have been many research efforts in developing generative models for self-exciting point processes, partly due to their broad applicability for real-world applications. However, rarely can we quantify how well the generative model captures the nature or ground-truth since it is usually unknown. The challenge typically lies in the fact that the generative models typically provide, at most, good approximations to the ground-truth (e.g., through the rich representative power of neural networks), but they cannot be precisely the ground-truth. We thus cannot use the classic goodness-of-fit (GOF) test framework to evaluate their performance. In this paper, we develop a GOF test for generative models of self-exciting processes by making a new connection to this problem with the classical statistical theory of Quasi-maximum-likelihood estimator (QMLE). We present a non-parametric self-normalizing statistic for the GOF test: the Generalized Score (GS) statistics, and explicitly capture the model misspecification when establishing the asymptotic distribution of the GS statistic. Numerical simulation and real-data experiments validate our theory and demonstrate the proposed GS test’s good performance.
Song Wei, Shixiang Zhu, Minghe Zhang, Yao Xie 0002
AISTATS2
2021 Deep Fourier Kernel for Self-Attentive Point Processes
abstract
We present a novel attention-based model for discrete event data to capture complex non-linear temporal dependence structures. We borrow the idea from the attention mechanism and incorporate it into the point processes’ conditional intensity function. We further introduce a novel score function using Fourier kernel embedding, whose spectrum is represented using neural networks, which drastically differs from the traditional dot-product kernel and can capture a more complex similarity structure. We establish our approach’s theoretical properties and demonstrate our approach’s competitive performance compared to the state-of-the-art for synthetic and real data.
Shixiang Zhu, Minghe Zhang, Ruyi Ding, Yao Xie 0002
AISTATS1
2021 Sequential Adversarial Anomaly Detection with Deep Fourier Kernel
abstract
We present a novel adversarial detector for the anomalous sequence when there are only one-class training samples. The detector is developed by finding the best detector that can discriminate against the worst-case, which statistically mimics the training sequences. We explicitly capture the dependence in sequential events using the marked point process with a deep Fourier kernel. The detector evaluates a test sequence and compares it with an optimal time-varying threshold, which is also learned from data. Using numerical experiments on simulations and real-world datasets, we demonstrate the superior performance of our proposed method.
Shixiang Zhu, Henry Shaowu Yuchi, Minghe Zhang, Yao Xie 0002
ICASSP1
2020 Adversarial Anomaly Detection for Marked Spatio-Temporal Streaming Data
abstract
Spatio-temporal event data are becoming increasingly commonplace in a wide variety of applications, such as electronic transaction records, social network data, and crime incident reports. How to efficiently detect anomalies in these dynamic systems using these streaming event data? This work proposes a novel anomaly detection framework for such event data combining the Long Short-Term Memory (LSTM) and marked spatio-temporal point processes. The detection procedure can be computed in an online and distributed fashion via feeding the streaming data through an LSTM and a neural network-based discriminator. This work studies the false-alarm-rate and detection delay using theory and simulation and shows that it can achieve weak signal detection by aggregating local statistics over time and networks. Finally, we demonstrate the good performance using real-world data sets.
Shixiang Zhu, Henry Shaowu Yuchi, Yao Xie 0002
ICASSP1
2019 Crime Event Embedding with Unsupervised Feature Selection
abstract
We present a novel event embedding algorithm for crime data that can jointly capture time, location, and the complex free-text component of each event. The embedding is achieved by regularized Restricted Boltzmann Machines (RBMs), and we introduce a new way to regularize by imposing a ℓ1penalty on the conditional distributions of the observed variables of RBMs. This choice of regularization performs feature selection and it also leads to efficient computation since the gradient can be computed in a closed form. The feature selection forces embedding to be based on the most important keywords, which captures the common modus operandi (M.O.) in crime series. Using numerical experiments on a large-scale crime dataset, we show that our regularized RBMs can achieve better event embedding and the selected features are highly interpretable from human understanding.
Shixiang Zhu, Yao Xie 0002
ICASSP1
2018 Sequential Adaptive Detection for In-Situ Transmission Electron Microscopy (TEM)
abstract
We develop new efficient online algorithms for detecting transient sparse signals in TEM video sequences, by adopting the recently developed framework for sequential detection jointly with online convex optimization [1]. We cast the problem as detecting an unknown sparse mean shift of Gaussian observations, and develop adaptive CUSUM and adaptive SSRS procedures, which are based on likelihood ratio statistics with post-change mean vector being online maximum likelihood estimators with l1. We demonstrate the meritorious performance of our algorithms for TEM imaging using real data.
Yang Cao 0013, Shixiang Zhu, Yao Xie 0002, Jordan Key, Josh Kacher, Raymond R. Unocic, Christopher M. Rouleau
ICASSP2
2018 Crime Incidents Embedding Using Restricted Boltzmann Machines
abstract
We present a new approach for detecting related crime series, by unsupervised learning of the latent feature embeddings from narratives of crime record via the Gaussian-Bernoulli Restricted Boltzmann Machine (GBRBM). This is a drastically different approach from prior work on crime analysis, which typically considers only time and location and at most category information. After the embedding, related cases are closer to each other in the Euclidean feature space, and the unrelated cases are far apart, which is a good property can enable subsequent analysis such as detection and clustering of related cases. Experiments over several series of related crime incidents hand labeled by the Atlanta Police Department reveal the promise of our embedding methods.
Shixiang Zhu, Yao Xie 0002
ICASSP1
2018 Learning Temporal Point Processes via Reinforcement Learning
abstract
Social goods, such as healthcare, smart city, and information networks, often produce ordered event data in continuous time. The generative processes of these event data can be very complex, requiring flexible models to capture their dynamics. Temporal point processes offer an elegant framework for modeling event data without discretizing the time. However, the existing maximum-likelihood-estimation (MLE) learning paradigm requires hand-crafting the intensity function beforehand and cannot directly monitor the goodness-of-fit of the estimated model in the process of training. To alleviate the risk of model-misspecification in MLE, we propose to generate samples from the generative model and monitor the quality of the samples in the process of training until the samples and the real data are indistinguishable. We take inspiration from reinforcement learning (RL) and treat the generation of each event as the action taken by a stochastic policy. We parameterize the policy as a flexible recurrent neural network and gradually improve the policy to mimic the observed event distribution. Since the reward function is unknown in this setting, we uncover an analytic and nonparametric form of the reward function using an inverse reinforcement learning formulation. This new RL framework allows us to derive an efficient policy gradient algorithm for learning flexible point process models, and we show that it performs well in both synthetic and real data.
Shuang Li 0002, Shixiang Zhu, Nan Du 0002, Yao Xie 0002
NeurIPS3
2017 Senz: A Context Awareness Middleware System Used in Mobile Devices
abstract
With the continuing penetration of sensor devices and the development of wireless communication techniques, increasing number of applications involving context awareness in ubiquitous computing have been used in daily life. How to collect data from mobile devices at a low energy cost and to mine contextual habits of users remains a key challenge for ubiquitous computing. We proposed an efficient context- awareness computing middleware system, Senz. By leveraging high-efficiency mobile data transmission method, using high-performance context recognition algorithms, and combining mobile data with online third-party data, this middleware system can recognize various user behavior patterns reliably, accurately and efficiently. Experiments show that the Senz recognition accuracy of context activity is above 83% on average and the energy cost is relatively low. By integrating Senz SDK and cloud computing engine, developers can build rich user experience apps with better understanding of users' behavior data and providing various personalized contextual services.
Hengyang Zhang, Tao Huang 0005, Yunjie Liu 0001, Shixiang Zhu, Yuanying Chi
VTC Spring4
2014 Enabling Smartphone Based HD Video Chats by Cooperative Transmissions in CRNs
Xuewei Cui, Wei Cheng 0001, Shixiang Zhu, Yan Huo 0001
WASA4
2013 Cooperative relay selection in cognitive radio networks
abstract
The benefits of cognitive radio networks have been well recognized with the dramatic development of the wireless applications in recent years. While many existing works assume that the secondary transmissions are negative interference to the primary users (PUs), in this paper, we take secondary users (SUs) as positive potential cooperators for the primary users. In particular, we consider the problem of cooperative relay selection, in which the PUs actively select appropriate SUs as relay nodes to enhance their transmission performance. The most critical challenge for such a problem of cooperative relay selection is how to select a relay efficiently. But due to the potentially large number of secondary users, it is infeasible for a PU transmitter to first scan all the SUs and then pick the best one. Basically, the PU transmitter intends to observe the SUs sequentially. After observing a SU, the PU needs to make a decision on whether to terminate its observation and use the current SU as its relay or to skip it and observe the next SU. We address this problem by using the optimal stopping theory, and derive the optimal stopping rule. To evaluate the performance of our proposed scheme, we conduct an extensive simulation study. The results reveal the impact of different parameters on the system performance, which can be adjusted to satisfy specific system requirements.
Shixiang Zhu, Hongjuan Li, Xiuzhen Cheng, Yan Huo 0001
INFOCOM2