VLDB 2026 Research / reviewers in the wild / expert
Xueru Zhang
dblp:51/8761
· DBLP profile ↗
41ranked-venue papers
13as first author
27since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 28 · 5 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 4 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 1 since 2021Security and privacy · 2 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Stabilizing Self-Consuming Diffusion Models with Latent Space FilteringabstractAs synthetic data proliferates across the Internet, it is often reused to train successive generations of generative models. This creates a "self-consuming loop" that can lead to training instability or *model collapse*. Common strategies to address the issue---such as accumulating historical training data or injecting fresh real data---either increase computational cost or require expensive human annotation. In this paper, we empirically analyze the latent space dynamics of self-consuming diffusion models and observe that the low-dimensional structure of latent representations extracted from synthetic data degrade over generations. Based on this insight, we propose *Latent Space Filtering* (LSF), a novel approach that mitigates model collapse by filtering out less realistic synthetic data from mixed datasets. Theoretically, we present a framework that connects latent space degradation to empirical observations. Experimentally, we show that LSF consistently outperforms existing baselines across multiple real-world datasets, effectively mitigating model collapse without increasing training cost or relying on human annotation. Zhongteng Cai, Yaxuan Wang, Xueru Zhang |
AAAI | 4 |
| 2026 | Addressing Polarization and Unfairness in Performative PredictionabstractIn many real-world applications of machine learning—such as recommendations, hiring, and lending—deployed models influence the data they are trained on, leading to feedback loops between predictions and data distribution. The performative prediction (PP) framework captures this phenomenon by modeling the data distribution as a function of the deployed model. While prior work has focused on finding performative stable (PS) solutions for robustness, their societal impacts, particularly regarding fairness, remain underexplored. We show that PS solutions can lead to severe polarization and prediction performance disparities, and that conventional fairness interventions in previous works often fail under model-dependent distribution shifts due to failing the PS criteria. To address these challenges in PP, we introduce novel fairness mechanisms that provably ensure both stability and fairness, validated by theoretical analysis and empirical results. Xueru Zhang |
AAAI | 4 |
| 2026 | Achieving Fairness Without Harm via Selective Demographic ExpertsabstractAs machine learning systems become increasingly integrated into human-centered domains such as healthcare, ensuring fairness while maintaining high predictive performance is critical. Existing bias mitigation techniques often impose a trade-off between fairness and accuracy, inadvertently degrading performance for certain demographic groups. In high-stakes domains like clinical diagnosis, such trade-offs are ethically and practically unacceptable. In this study, we propose a fairness-without-harm approach by learning distinct representations for different demographic groups and selectively applying demographic experts consisting of group-specific representations and personalized classifiers through a no-harm constrained selection. We evaluate our approach on three real-world medical datasets—covering eye disease, skin cancer, and X-ray diagnosis—as well as two face datasets. Extensive empirical results demonstrate the effectiveness of our approach in achieving fairness without harm. Xuwei Tan, Yuanlong Wang 0004, Thai-Hoang Pham, Ping Zhang 0016, Xueru Zhang |
AAAI | 5 |
| 2026 | Understanding Structured Financial Data with LLMs: A Case Study on Fraud DetectionabstractDetecting fraud in financial transactions typically relies on tabular models that demand heavy feature engineering to handle highdimensional data and offer limited interpretability, making it difficult for humans to understand predictions.Large Language Models (LLMs), in contrast, can produce human-readable explanations and facilitate feature analysis, potentially reducing the manual workload of fraud analysts and informing system refinements.However, they perform poorly when applied directly to tabular fraud detection due to the difficulty of reasoning over many features, the extreme class imbalance, and the absence of contextual information.To bridge this gap, we introduce FinFRE-RAG, a two-stage approach that applies importance-guided feature reduction to serialize a compact subset of numeric/categorical attributes into natural language and performs retrieval-augmented in-context learning over label-aware, instance-level exemplars.Across four public fraud datasets and three families of open-weight LLMs, FinFRE-RAG substantially improves F1/MCC over direct prompting and is competitive with strong tabular baselines in several settings.Although these LLMs still lag behind specialized classifiers, they narrow the performance gap and provide interpretable rationales, highlighting their value as assistive tools in fraud analysis.* Work done during internship at Coinbase.This paper contains the author's personal opinions and does not constitute a company policy or statement.These opinions are not endorsed by or affiliated with Coinbase, Inc. or its subsidiaries. Xuwei Tan, Xueru Zhang |
ACL (1) | 3 |
| 2026 | Observations and Remedies for Large Language Model Bias in Self-Consuming Performative LoopabstractThe rapid advancement of large language models (LLMs) has led to growing interest in using synthetic data to train future models.However, this creates a self-consuming retraining loop, where models are trained on their own outputs and may cause performance drops and induce emerging biases.In real-world applications, previously deployed LLMs may influence the data they generate, leading to a dynamic system driven by user feedback.For example, if a model continues to underserve users from a group, less query data will be collected from this particular demographic of users.In this study, we introduce the concept of Self-Consuming Performative Loop (SCPL) and investigate the role of synthetic data in shaping bias during these dynamic iterative training processes under controlled performative feedback.This controlled setting is motivated by the inaccessibility of real-world user preference data from dynamic production systems, and enables us to isolate and analyze feedback-driven bias evolution in a principled manner.We focus on two types of loops, including the typical retraining setting and the incremental finetuning setting, which is largely underexplored.Through experiments on three real-world tasks, we find that the performative loop increases preference bias and decreases disparate bias.We design a reward-based rejection sampling strategy to mitigate the bias, moving towards more trustworthy self-improving systems. Yaxuan Wang, Zhongteng Cai, Yujia Bao, Xueru Zhang, Yang Liu 0018 |
ACL (1) | 4 |
| 2026 | An ECC-Based Three-Factor Authentication and Key Management Protocol for Session Key Leakage Resilience in Internet of DronesabstractAs a network with dynamic network topology and limited on-board resource, the Internet of Drones(IoD) suffers various cyber attacks, such as real-time data tampering, clogging, jamming and etc.. The real-time data tampering, especially the session key leakage, in the IoD can distort time-sensitive information and disrupt system operations. Many existing works based on authentication and key agreement perform low efficient and not secure enough for real-time data transmitting in the system. In this manuscript, we proposed a three-factor authentication and key agreement protocol (HEAT) to solve these issues. The HEAT can solve session key leakage during data transmission by refining the user authentication mechanism and introducing the Elliptic Curve Cryptography into communications between drones and users. Compared with existing protocols, the HEAT achieves less authentication delay and higher data efficiency. The HEAT is proven secure with numerous experiments under the extended Canetti-Krawczyk model. Xueru Zhang, Di Wu 0042, Rui Zhang 0083 |
IEEE Internet Things J. | 1 |
| 2025 | Open-Set Heterogeneous Domain Adaptation: Theoretical Analysis and AlgorithmabstractDomain adaptation (DA) tackles the issue of distribution shift by learning a model from a source domain that generalizes to a target domain. However, most existing DA methods are designed for scenarios where the source and target domain data lie within the same feature space, which limits their applicability in real-world situations. Recently, heterogeneous DA (HeDA) methods have been introduced to address the challenges posed by heterogeneous feature space between source and target domains. Despite their successes, current HeDA techniques fall short when there is a mismatch in both feature and label spaces. To address this, this paper explores a new DA scenario called open-set HeDA (OSHeDA). In OSHeDA, the model must not only handle heterogeneity in feature space but also identify samples belonging to novel classes. To tackle this challenge, we first develop a novel theoretical framework that constructs learning bounds for prediction error on target domain. Guided by this framework, we propose a new DA method called Representation Learning for OSHeDA (RL-OSHeDA). This method is designed to simultaneously transfer knowledge between heterogeneous data sources and identify novel classes. Experiments across text, image, and clinical data demonstrate the effectiveness of our algorithm. Model implementation is available at https://github.com/pth1993/OSHeDA. Thai-Hoang Pham, Yuanlong Wang 0004, Changchang Yin, Xueru Zhang, Ping Zhang 0016 |
AAAI | 4 |
| 2025 | DroughtSet: Understanding Drought Through Spatial-Temporal LearningabstractDrought is one of the most destructive and expensive natural disasters, severely impacting natural resources and risks by depleting water resources and diminishing agricultural yields. Under climate change, accurately predicting drought is critical for mitigating drought-induced risks. However, the intricate interplay among the physical and biological drivers that regulate droughts limits the predictability and understanding of drought, particularly at a subseasonal to seasonal (S2S) time scale. While deep learning has demonstrated the potential to address climate forecasting challenges, its application to drought prediction has received relatively less attention. In this work, we propose a new dataset, DroughtSet, which integrates relevant predictive features and three drought indices from multiple remote sensing and reanalysis datasets across the contiguous United States (CONUS). DroughtSet specifically provides the machine learning community with a new real-world dataset to benchmark drought prediction models and more generally, time-series forecasting methods. Furthermore, we propose a spatial-temporal model SPDrought to predict and interpret S2S droughts. Our model learns from the spatial and temporal information of physical and biological features to predict three types of droughts simultaneously. Multiple strategies are employed to quantify the importance of physical and biological features for drought prediction. Our results provide insights for researchers to better understand the predictability and sensitivity of drought to biological and physical conditions. We aim to contribute to the climate field by proposing a new tool to predict and understand the occurrence of droughts and provide the AI community with a new benchmark to study deep learning applications in climate science. Xuwei Tan, Yanlan Liu, Xueru Zhang |
AAAI | 4 |
| 2025 | Self-Consuming Generative Models with Adversarially Curated DataabstractRecent advances in generative models have made it increasingly difficult to distinguish real data from model-generated synthetic data. Using synthetic data for successive training of future model generations creates “self-consuming loops,” which may lead to model collapse or training instability. Furthermore, synthetic data is often subject to human feedback and curated by users based on their preferences. Ferbach et al. (2024) recently showed that when data is curated according to user preferences, the self-consuming retraining loop drives the model to converge toward a distribution that optimizes those preferences. However, in practice, data curation is often noisy or adversarially manipulated. For example, competing platforms may recruit malicious users to adversarially curate data and disrupt rival models. In this paper, we study how generative models evolve under self-consuming retraining loops with noisy and adversarially curated data. We theoretically analyze the impact of such noisy data curation on generative models and identify conditions for the robustness and stability of the retraining process. Building on this analysis, we design attack algorithms for competitive adversarial scenarios, where a platform with a limited budget employs malicious users to misalign a rival’s model from actual user preferences. Experiments on both synthetic and real-world datasets demonstrate the effectiveness of the proposed algorithms. Xiukun Wei, Xueru Zhang |
ICML | 2 |
| 2025 | The Boundaries of Fair AI in Medical Image Prognosis: A Causal PerspectiveabstractAs machine learning (ML) algorithms are increasingly used in medical image analysis, concerns have emerged about their potential biases against certain social groups. Although many approaches have been proposed to ensure the fairness of ML models, most existing works focus only on medical image diagnosis tasks, such as image classification and segmentation, and overlooked prognosis scenarios, which involve predicting the likely outcome or progression of a medical condition over time. To address this gap, we introduce FairTTE, the first comprehensive framework for assessing fairness in time-to-event (TTE) prediction in medical imaging. FairTTE encompasses a diverse range of imaging modalities and TTE outcomes, integrating cutting-edge TTE prediction and fairness algorithms to enable systematic and fine-grained analysis of fairness in medical image prognosis. Leveraging causal analysis techniques, FairTTE uncovers and quantifies distinct sources of bias embedded within medical imaging datasets. Our large-scale evaluation reveals that bias is pervasive across different imaging modalities and that current fairness methods offer limited mitigation. We further demonstrate a strong association between underlying bias sources and model disparities, emphasizing the need for holistic approaches that target all forms of bias. Notably, we find that fairness becomes increasingly difficult to maintain under distribution shifts, underscoring the limitations of existing solutions and the pressing need for more robust, equitable prognostic models. Thai-Hoang Pham, Jiayuan Chen 0003, Seungyeon Lee 0002, Yuanlong Wang 0004, Sayoko Moroi, Xueru Zhang, Ping Zhang 0016 |
NeurIPS | 6 |
| 2024 | Performative Federated Learning: A Solution to Model-Dependent and Heterogeneous Distribution ShiftsabstractWe consider a federated learning (FL) system consisting of multiple clients and a server, where the clients aim to collaboratively learn a common decision model from their distributed data. Unlike the conventional FL framework that assumes the client's data is static, we consider scenarios where the clients' data distributions may be reshaped by the deployed decision model. In this work, we leverage the idea of distribution shift mappings in performative prediction to formalize this model-dependent data distribution shift and propose a performative FL framework. We first introduce necessary and sufficient conditions for the existence of a unique performative stable solution and characterize its distance to the performative optimal solution. Then we propose the performative FedAvg algorithm and show that it converges to the performative stable solution at a rate of O(1/T) under both full and partial participation schemes. In particular, we use novel proof techniques and show how the clients' heterogeneity influences the convergence. Numerical results validate our analysis and provide valuable insights into real-world applications. Tongxin Yin, Zhongzhu Chen, Xueru Zhang, Yang Liu 0018, Mingyan Liu |
AAAI | 5 |
| 2024 | Algorithmic Decision-Making under Agents with Persistent ImprovementabstractThis paper studies algorithmic decision-making under human strategic behavior, where a decision-maker uses an algorithm to make decisions about human agents, and the latter with information about the algorithm may exert effort strategically and improve to receive favorable decisions. Unlike prior works that assume agents benefit from their efforts immediately, we consider realistic scenarios where the impacts of these efforts are persistent and agents benefit from efforts by making improvements gradually. We first develop a dynamic model to characterize persistent improvements and based on this construct a Stackelberg game to model the interplay between agents and the decision-maker. We analytically characterize the equilibrium strategies and identify conditions under which agents have incentives to invest efforts to improve their qualifications. With the dynamics, we then study how the decision-maker can design an optimal policy to incentivize the largest improvements inside the agent population. We also extend the model to settings where 1) agents may be dishonest and game the algorithm into making favorable but erroneous decisions; 2) honest efforts are forgettable and not sufficient to guarantee persistent improvements. With the extended models, we further examine conditions under which agents prefer honest efforts over dishonest behavior and the impacts of forgettable efforts. Xuwei Tan, Xueru Zhang |
AIES (1) | 3 |
| 2024 | Non-linear Welfare-Aware Strategic LearningabstractThis paper studies algorithmic decision-making in the presence of strategic individual behaviors, where an ML model is used to make decisions about human agents and the latter can adapt their behavior strategically to improve their future data. Existing results on strategic learning have largely focused on the linear setting where agents with linear labeling functions best respond to a (noisy) linear decision policy. Instead, this work focuses on general non-linear settings where agents respond to the decision policy with only "local information" of the policy. Moreover, we simultaneously consider objectives of maximizing decision-maker welfare (model prediction accuracy), social welfare (agent improvement caused by strategic behaviors), and agent welfare (the extent that ML underestimates the agents). We first generalize the agent best response model in previous works to the non-linear setting and then investigate the compatibility of welfare objectives. We show the three welfare can attain the optimum simultaneously only under restrictive conditions which are challenging to achieve in non-linear settings. The theoretical results imply that existing works solely maximizing the welfare of a subset of parties usually diminish the welfare of others. We thus claim the necessity of balancing the welfare of each party in non-linear settings and propose an irreducible optimization algorithm suitable for general strategic learning. Experiments on synthetic and real data validate the proposed algorithm. Xueru Zhang |
AIES (1) | 2 |
| 2024 | Automating Data Annotation under Strategic Human Agents: Risks and Potential SolutionsabstractAs machine learning (ML) models are increasingly used in social domains to make consequential decisions about humans, they often have the power to reshape data distributions. Humans, as strategic agents, continuously adapt their behaviors in response to the learning system. As populations change dynamically, ML systems may need frequent updates to ensure high performance. However, acquiring high-quality *human-annotated* samples can be highly challenging and even infeasible in social domains. A common practice to address this issue is using the model itself to annotate unlabeled data samples. This paper investigates the long-term impacts when ML models are retrained with *model-annotated* samples when they incorporate human strategic responses. We first formalize the interactions between strategic agents and the model and then analyze how they evolve under such dynamic interactions. We find that agents are increasingly likely to receive positive decisions as the model gets retrained, whereas the proportion of agents with positive labels may decrease over time. We thus propose a *refined retraining process* to stabilize the dynamics. Last, we examine how algorithmic fairness can be affected by these retraining processes and find that enforcing common fairness constraints at every round may not benefit the disadvantaged group in the long run. Experiments on (semi-)synthetic and real data validate the theoretical findings. Xueru Zhang |
NeurIPS | 2 |
| 2024 | Privacy-Aware Randomized Quantization via Linear ProgrammingabstractDifferential privacy mechanisms such as the Gaussian or Laplace mechanism have been widely used in data analytics for preserving individual privacy. However, they are mostly designed for continuous outputs and are unsuitable for scenarios where discrete values are necessary. Although various quantization mechanisms were proposed recently to generate discrete outputs under differential privacy, the outcomes are either biased or have an inferior accuracy-privacy trade-off. In this paper, we propose a family of quantization mechanisms that is unbiased and differentially private. It has a high degree of freedom and we show that some existing mechanisms can be considered as special cases of ours. To find the optimal mechanism, we formulate a linear optimization that can be solved efficiently using linear programming tools. Experiments show that our proposed mechanism can attain a better privacy-accuracy trade-off compared to baselines. Zhongteng Cai, Xueru Zhang, Mohammad Mahdi Khalili |
UAI | 2 |
| 2024 | Non-stationary Domain Generalization: Theory and AlgorithmabstractAlthough recent advances in machine learning have shown its success to learn from independent and identically distributed (IID) data, it is vulnerable to out-of-distribution (OOD) data in an open world. Domain generalization (DG) deals with such an issue and it aims to learn a model from multiple source domains that can be generalized to unseen target domains. Existing studies on DG have largely focused on stationary settings with homogeneous source domains. However, in many applications, domains may evolve along a specific direction (e.g., time, space). Without accounting for such non-stationary patterns, models trained with existing methods may fail to generalize on OOD data. In this paper, we study domain generalization in non-stationary environment. We first examine the impact of environmental non-stationarity on model performance and establish the theoretical upper bounds for the model error at target domains. Then, we propose a novel algorithm based on adaptive invariant representation learning, which leverages the non-stationary pattern to train a model that attains good performance on target domains. Experiments on both synthetic and real data validate the proposed algorithm. Thai-Hoang Pham, Xueru Zhang, Ping Zhang 0016 |
UAI | 2 |
| 2024 | Generating synthetic computed tomography for radiotherapy: SynthRAD2023 challenge reportabstractRadiation therapy plays a crucial role in cancer treatment, necessitating precise delivery of radiation to tumors while sparing healthy tissues over multiple days. Computed tomography (CT) is integral for treatment planning, offering electron density data crucial for accurate dose calculations. However, accurately representing patient anatomy is challenging, especially in adaptive radiotherapy, where CT is not acquired daily. Magnetic resonance imaging (MRI) provides superior soft-tissue contrast. Still, it lacks electron density information, while cone beam CT (CBCT) lacks direct electron density calibration and is mainly used for patient positioning. Adopting MRI-only or CBCT-based adaptive radiotherapy eliminates the need for CT planning but presents challenges. Synthetic CT (sCT) generation techniques aim to address these challenges by using image synthesis to bridge the gap between MRI, CBCT, and CT. The SynthRAD2023 challenge was organized to compare synthetic CT generation methods using multi-center ground truth data from 1080 patients, divided into two tasks: (1) MRI-to-CT and (2) CBCT-to-CT. The evaluation included image similarity and dose-based metrics from proton and photon plans. The challenge attracted significant participation, with 617 registrations and 22/17 valid submissions for tasks 1/2. Top-performing teams achieved high structural similarity indices (≥0.87/0.90) and gamma pass rates for photon (≥98.1%/99.0%) and proton (≥97.3%/97.0%) plans. However, no significant correlation was found between image similarity metrics and dose accuracy, emphasizing the need for dose evaluation when assessing the clinical applicability of sCT. SynthRAD2023 facilitated the investigation and benchmarking of sCT generation techniques, providing insights for developing MRI-only and CBCT-based adaptive radiotherapy. It showcased the growing capacity of deep learning to produce high-quality sCT, reducing reliance on conventional CT for treatment planning. Evi M. C. Huijben, Maarten L. Terpstra, Arthur Jr Galapon, Suraj Pai, Adrian Thummerer, Peter J. Koopmans, Manya Afonso, Maureen van Eijnatten, Oliver J. Gurney-Champion, Zeli Chen, Kaiyi Zheng, Chuanpu Li, Haowen Pang, Chuyang Ye, Runqi Wang, Fuxin Fan, Jingna Qiu, Yixing Huang, Juhyung Ha, Jong Sung Park, Alexandra Alain-Beaudoin, Silvain Bériault, Pengxin Yu, Zhanyao Huang, Gengwan Li, Xueru Zhang, Yubo Fan, Bowen Xin, Aaron Nicolson, Lujia Zhong, Zhiwei Deng, Gustav Mueller-Franzes, Firas Khader, Xia Li 0005, Ye Zhang 0039, Cédric Hémon, Valentin Boussot, Shaobin Wang, Derk Mus, Bram Kooiman, Chelsea A. H. Sargeant, Edward G. A. Henderson, Satoshi Kondo, Satoshi Kasai, Reza Karimzadeh, Bulat Ibragimov, Thomas Helfer, Jessica Dafflon, Enpei Wang, Zoltán Perkó, Matteo Maspero |
Medical Image Anal. | 29 |
| 2023 | Fairness and Accuracy under Domain Generalization
Thai-Hoang Pham, Xueru Zhang, Ping Zhang 0016 |
ICLR | 2 |
| 2023 | Loss Balancing for Fair Supervised LearningabstractSupervised learning models have been used in various domains such as lending, college admission, face recognition, natural language processing, etc. However, they may inherit pre-existing biases from training data and exhibit discrimination against protected social groups. Various fairness notions have been proposed to address unfairness issues. In this work, we focus on Equalized Loss (EL), a fairness notion that requires the expected loss to be (approximately) equalized across different groups. Imposing EL on the learning process leads to a non-convex optimization problem even if the loss function is convex, and the existing fair learning algorithms cannot properly be adopted to find the fair predictor under the EL constraint. This paper introduces an algorithm that can leverage off-the-shelf convex programming tools (e.g., CVXPY (Diamond and Boyd, 2016; Agrawal et al., 2018)) to efficiently find the global optimum of this non-convex optimization. In particular, we propose the ELminimizer algorithm, which finds the optimal fair predictor under EL by reducing the non-convex optimization to a sequence of convex optimization problems. We theoretically prove that our algorithm finds the global optimal solution under certain conditions. Then, we support our theoretical results through several empirical studies Mohammad Mahdi Khalili, Xueru Zhang, Mahed Abroshan |
ICML | 2 |
| 2023 | Counterfactually Fair RepresentationabstractThe use of machine learning models in high-stake applications (e.g., healthcare, lending, college admission) has raised growing concerns due to potential biases against protected social groups. Various fairness notions and methods have been proposed to mitigate such biases. In this work, we focus on Counterfactual Fairness (CF), a fairness notion that is dependent on an underlying causal graph and first proposed by Kusner $\textit{et al.}$; it requires that the outcome an individual perceives is the same in the real world as it would be in a "counterfactual" world, in which the individual belongs to another social group.
Learning fair models satisfying CF can be challenging. It was shown in (Kusner $\textit{et al.}$) that a sufficient condition for satisfying CF is to $\textbf{not}$ use features that are descendants of sensitive attributes in the causal graph. This implies a simple method that learns CF models only using non-descendants of sensitive attributes while eliminating all descendants. Although several subsequent works proposed methods that use all features for training CF models, there is no theoretical guarantee that they can satisfy CF. In contrast, this work proposes a new algorithm that trains models using all the available features. We theoretically and empirically show that models trained with this method can satisfy CF. Zhiqun Zuo, Mohammad Mahdi Khalili, Xueru Zhang |
NeurIPS | 3 |
| 2023 | A fair and interpretable network for clinical risk prediction: a regularized multi-view multi-task learning approach
Thai-Hoang Pham, Changchang Yin, Laxmi Mehta, Xueru Zhang, Ping Zhang 0016 |
Knowl. Inf. Syst. | 4 |
| 2023 | Differentially Private Real-Time Release of Sequential DataabstractMany data analytics applications rely on temporal data, generated (and possibly acquired) sequentially for online analysis. How to release this type of data in a privacy-preserving manner is of great interest and more challenging than releasing one-time, static data. Because of the (potentially strong) temporal correlation within the data sequence, the overall privacy loss can accumulate significantly over time; an attacker with statistical knowledge of the correlation can be particularly hard to defend against. An idea that has been explored in the literature to mitigate this problem is to factor this correlation into the perturbation/noise mechanism. Existing work, however, either focuses on the offline setting (where perturbation is designed and introduced after the entire sequence has become available), or requires a priori information on the correlation in generating perturbation. In this study we propose an approach where the correlation is learned as the sequence is generated, and is used for estimating future data in the sequence. This estimate then drives the generation of the noisy released data. This method allows us to design better perturbation and is suitable for real-time operations. Using the notion of differential privacy, we show this approach achieves high accuracy with lower privacy loss compared to existing methods. Xueru Zhang, Mohammad Mahdi Khalili, Mingyan Liu |
ACM Trans. Priv. Secur. | 1 |
| 2022 | Fairness Interventions as (Dis)Incentives for Strategic ManipulationabstractAlthough machine learning (ML) algorithms are widely used to make decisions about individuals in various domains, concerns have arisen that (1) these algorithms are vulnerable to strategic manipulation and "gaming the algorithm"; and (2) ML decisions may exhibit bias against certain social groups. Existing works have largely examined these as two separate issues, e.g., by focusing on building ML algorithms robust to strategic manipulation, or on training a fair ML algorithm. In this study, we set out to understand the impact they each have on the other, and examine how to characterize fair policies in the presence of strategic behavior. The strategic interaction between a decision maker and individuals (as decision takers) is modeled as a two-stage (Stackelberg) game; when designing an algorithm, the former anticipates the latter may manipulate their features in order to receive more favorable decisions. We analytically characterize the equilibrium strategies of both, and examine how the algorithms and their resulting fairness properties are affected when the decision maker is strategic (anticipates manipulation), as well as the impact of fairness interventions on equilibrium strategies. In particular, we identify conditions under which anticipation of strategic behavior may mitigate/exacerbate unfairness, and conditions under which fairness interventions can serve as (dis)incentives for strategic manipulation. Xueru Zhang, Mohammad Mahdi Khalili, Parinaz Naghizadeh Ardabili, Mingyan Liu |
ICML | 1 |
| 2022 | Incentive Mechanisms for Strategic Classification and Regression ProblemsabstractWe study the design of a class of incentive mechanisms that can effectively prevent cheating in a strategic classification and regression problem. A conventional strategic classification or regression problem is modeled as a Stackelberg game, or a principal-agent problem between the designer of a classifier (the principal) and individuals subject to the classifier's decisions (the agents), potentially from different demographic groups. The former benefits from the accuracy of its decisions, whereas the latter may have an incentive to game the algorithm into making favorable but erroneous decisions. While prior works tend to focus on how to design an algorithm to be more robust to such strategic maneuvering, this study focuses on an alternative, which is to design incentive mechanisms to shape the utilities of the agents and induce effort that genuinely improves their skills, which in turn benefits both parties in the Stackelberg game. Specifically, the principal and the mechanism provider (which could also be the principal itself) move together in the first stage, publishing and committing to a classifier and an incentive mechanism. The agents are (simultaneous) second movers and best respond to the published classifier and incentive mechanism. When an agent's strategic action merely changes its observable features, it hurts the performance of the algorithm. However, if the action leads to improvement in the agent's true label, it not only helps the agent achieve better decision outcomes, but also preserves the performance of the algorithm. We study how a subsidy mechanism can induce improvement actions, positively impact a number of social well-being metrics, such as the overall skill levels of the agents (efficiency) and positive or true positive rate differences between different demographic groups (fairness). Xueru Zhang, Mohammad Mahdi Khalili, Parinaz Naghizadeh Ardabili, Mingyan Liu |
EC | 2 |
| 2021 | Improving Fairness and Privacy in Selection ProblemsabstractSupervised learning models have been increasingly used for making decisions about individuals in applications such as hiring, lending, and college admission. These models may inherit pre-existing biases from training datasets and discriminate against protected attributes (e.g., race or gender). In addition to unfairness, privacy concerns also arise when the use of models reveals sensitive personal information. Among various privacy notions, differential privacy has become popular in recent years. In this work, we study the possibility of using a differentially private exponential mechanism as a post-processing step to improve both fairness and privacy of supervised learning models. Unlike many existing works, we consider a scenario where a supervised model is used to select a limited number of applicants as the number of available positions is limited. This assumption is well-suited for various scenarios, such as job application and college admission. We use ``equal opportunity'' as the fairness notion and show that the exponential mechanisms can make the decision-making process perfectly fair. Moreover, the experiments on real-world datasets show that the exponential mechanism can improve both privacy and fairness, with a slight decrease in accuracy compared to the model without post-processing. Mohammad Mahdi Khalili, Xueru Zhang, Mahed Abroshan, Somayeh Sojoudi |
AAAI | 2 |
| 2021 | Cardiac Complication Risk Profiling for Cancer Survivors via Multi-View Multi-Task LearningabstractComplication risk profiling is a key challenge in the healthcare domain due to the complex interaction between heterogeneous entities (e.g., visit, disease, medication) in clinical data. With the availability of real-world clinical data such as electronic health records and insurance claims, many deep learning methods are proposed for complication risk profiling. However, these existing methods face two open challenges. First, data heterogeneity relates to those methods leveraging clinical data from a single view only while the data can be considered from multiple views (e.g., sequence of clinical visits, set of clinical features). Second, generalized prediction relates to most of those methods focusing on single-task learning, whereas each complication onset is predicted independently, leading to suboptimal models. We propose a multi-view multi-task network (MuViTaNet) for predicting the onset of multiple complications to tackle these issues. In particular, MuViTaNet complements patient representation by using a multi-view encoder to effectively extract information by considering clinical data as both sequences of clinical visits and sets of clinical features. In addition, it leverages additional information from both related labeled and unlabeled datasets to generate more generalized representations by using a new multi-task learning scheme for making more accurate predictions. The experimental results show that MuViTaNet outperforms existing methods for profiling the development of cardiac complications in breast cancer survivors. Furthermore, thanks to its multi-view multi-task architecture, MuViTaNet also provides an effective mechanism for interpreting its predictions in multiple perspectives, thereby helping clinicians discover the underlying mechanism triggering the onset and for making better clinical treatments in real-world scenarios. Thai-Hoang Pham, Changchang Yin, Laxmi Mehta, Xueru Zhang, Ping Zhang 0016 |
ICDM | 4 |
| 2021 | Fair Sequential Selection Using Supervised Learning ModelsabstractWe consider a selection problem where sequentially arrived applicants apply for a limited number of positions/jobs. At each time step, a decision maker accepts or rejects the given applicant using a pre-trained supervised learning model until all the vacant positions are filled. In this paper, we discuss whether the fairness notions (e.g., equal opportunity, statistical parity, etc.) that are commonly used in classification problems are suitable for the sequential selection problems. In particular, we show that even with a pre-trained model that satisfies the common fairness notions, the selection outcomes may still be biased against certain demographic groups. This observation implies that the fairness notions used in classification problems are not suitable for a selection problem where the applicants compete for a limited number of positions. We introduce a new fairness notion, ``Equal Selection (ES),'' suitable for sequential selection problems and propose a post-processing approach to satisfy the ES fairness notion. We also consider a setting where the applicants have privacy concerns, and the decision maker only has access to the noisy version of sensitive attributes. In this setting, we can show that the \textit{perfect} ES fairness can still be attained under certain conditions. Mohammad Mahdi Khalili, Xueru Zhang, Mahed Abroshan |
NeurIPS | 2 |
| 2020 | How do fair decisions fare in long-term qualification?abstractAlthough many fairness criteria have been proposed for decision making, their long-term impact on the well-being of a population remains unclear. In this work, we study the dynamics of population qualification and algorithmic decisions under a partially observed Markov decision problem setting. By characterizing the equilibrium of such dynamics, we analyze the long-term impact of static fairness constraints on the equality and improvement of group well-being. Our results show that static fairness constraints can either promote equality or exacerbate disparity depending on the driving factor of qualification transitions and the effect of sensitive attributes on feature distributions. We also consider possible interventions that can effectively improve group qualification or promote equality of group qualification. Our theoretical results and experiments on static real-world datasets with simulated dynamics show that our framework can be used to facilitate social science studies. Xueru Zhang, Ruibo Tu, Yang Liu 0018, Mingyan Liu, Hedvig Kjellström, Kun Zhang 0001, Cheng Zhang 0005 |
NeurIPS | 1 |
| 2020 | Recycled ADMM: Improving the Privacy and Accuracy of Distributed AlgorithmsabstractAlternating direction method of multiplier (ADMM) is a powerful method to solve decentralized convex optimization problems. In distributed settings, each node performs computation with its local data and the local results are exchanged among neighboring nodes in an iterative fashion. During this iterative process the leakage of data privacy arises and can accumulate significantly over many iterations, making it difficult to balance the privacy-accuracy tradeoff. We propose Recycled ADMM (R-ADMM), where a linear approximation is applied to every even iteration, its solution directly calculated using only results from the previous, odd iteration. It turns out that under such a scheme, half of the updates incur no privacy loss and require much less computation compared to the conventional ADMM. Moreover, R-ADMM can be further modified (MR-ADMM) such that each node independently determines its own penalty parameter over iterations. We obtain a sufficient condition for the convergence of both algorithms and provide the privacy analysis based on objective perturbation. It can be shown that the privacy-accuracy tradeoff can be improved significantly compared with conventional ADMM. Xueru Zhang, Mohammad Mahdi Khalili, Mingyan Liu |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2019 | Object-Oriented Automatic and Accurate Shadow Detection for Very High Spatial Resolution Satellite ImagesabstractSeveral existing shadow detection methods cannot keep the balance between accuracy and automaticity well. To overcome the weakness, we present a novel method to detect shadow in very high spatial resolution satellite images. First, a new shadow detection index is developed to obtain the shadow ratio map. The initial shadow mask map is then obtained by utilizing the Gaussian mixture mode and the Otsu's method automatically. Finally, the initial shadow mask map is refined by jointly using the object spectral characteristics and the spatial-correlation relationship between objects. The experimental results performed on different images show that the accuracy and automation of the proposed method are over several state-of-the-art methods. Yuwei Jin, Wenbo Xu 0004, Donghang Shao, Xixu He, Xueru Zhang |
IGARSS | 5 |
| 2019 | Retrieval of Fraction of Absorbed Photosynthetically Active Radiation (FPAR) Based on FengYun-3C /MERSI DataabstractFPAR is one of the key climate parameters for the Global Climate Observing System (GCOS) and the Global Terrestrial Observing System (GTOS). An accurate assessment to FPAR is particularly important to understand the global climate change. In this paper, the FY-3C Medium Resolution Imaging Spectrometer (MERIS) data combined with the improved PROSAIL model and look-up table algorithm to retrieve the FPAR at the Hulunber Grassland. The cross-validation results show that FPAR retrieved from FY-3C data has good consistency with the FPAR products of MODIS, GEOV1 and GLASS, with 0.6654, 0.6893 and 0.6192 at overall Pearson’s correlation coefficients (R), respectively. This study provides the potential for the generating the large-area and long-term sequence FPAR products from FengYun-3C (FY-3C) data. Wenbo Xu 0004, Xueru Zhang, Chunliang Zhao |
IGARSS | 4 |
| 2019 | Surface Albedo Inversion of FY-3C MERSI DataabstractThe surface albedo indicates the ability of the Earth's surface to reflect solar radiation, and it is an important land surface characteristic parameter that affects the radiation and energy balance of the Earth system. This paper presents a surface albedo inversion method for FY-3C MERSI. First, we estimated the narrow-band surface albedo for different surface cover types by using constrained least squares and the RossThick-LiTransit model. Then, we calculated the conversion coefficient of the FY-3C MERSI narrow-band surface albedo to the broad-band surface albedo conversion by using the FY-3C MERSI spectral response function, the 6S radiation transmission model and the USGS spectral library. Finally, we obtained the surface black and white albedo of the four narrow-band and visible-light bands of the FY-3C MERSI, and the spatial resolution of the data is 250m. In the comparative verification experiment, we cross-validated the surface albedo of FY-3C with the surface albedo of MODIS (moderate-resolution imaging spectroradiometer), CGLS (Copernicus Global Land Service) and GLASS (Global LAnd Surface Satellite). The results show that the correlation coefficient is mostly above 0.75, the overall absolute deviation is 0.068 on average and the minimum root mean square error is 0.02 between the albedo products obtained in this paper and the above three products. Therefore, FY-3C MERSI surface albedo products and MODIS, CGLS, GLASS surface albedo products have good consistency in four narrowband and visible light bands. Chunliang Zha, Wenbo Xu 0004, Xueru Zhang, Yantong Wu |
IGARSS | 3 |
| 2019 | Extracting Land Surface Water from FY/MERSI Image Based On Spectral Matching Of Discrete Particle Swarm Optimization and Linear Feature EnhancementabstractLand surface water is one of the most important components of surface cover and global water cycle. In this study, the standard water spectrum selected from FY/MERSI image was firstly used to calculate the water probability. Then, based on water probability, the image was roughly classified into four classes: homogeneous ground, junction of land cover, minor tributaries and other. To extract land surface water from small tributaries, the Duda's Road Operator (DRO) was introduced to enhance the linear features, while Discrete Particle Swarm Optimization (DPSO) was applied to extract land surface water from other three classes. The results show that the method could effectively extract land surface water, especially from small tributaries, and overall accuracy (OA) and Kappa coefficient are improved compared to DPSO algorithm based on spectral matching (SMDPSO). Xueru Zhang, Wenbo Xu 0004, Jinsheng Ren, Xixu He, Yuwei Jin |
IGARSS | 1 |
| 2019 | Group Retention when Using Machine Learning in Sequential Decision Making: the Interplay between User Dynamics and FairnessabstractMachine Learning (ML) models trained on data from multiple demographic groups can inherit representation disparity (Hashimoto et al., 2018) that may exist in the data: the model may be less favorable to groups contributing less to the training process; this in turn can degrade population retention in these groups over time, and exacerbate representation disparity in the long run. In this study, we seek to understand the interplay between ML decisions and the underlying group representation, how they evolve in a sequential framework, and how the use of fairness criteria plays a role in this process. We show that the representation disparity can easily worsen over time under a natural user dynamics (arrival and departure) model when decisions are made based on a commonly used objective and fairness criteria, resulting in some groups diminishing entirely from the sample pool in the long run. It highlights the fact that fairness criteria have to be defined while taking into consideration the impact of decisions on user dynamics. Toward this end, we explain how a proper fairness criterion can be selected based on a general user dynamics model. Xueru Zhang, Mohammadmahdi Khaliligarekani, Cem Tekin, Mingyan Liu |
NeurIPS | 1 |
| 2018 | Improving the Privacy and Accuracy of ADMM-Based Distributed AlgorithmsabstractAlternating direction method of multiplier (ADMM) is a popular method used to design distributed versions of a machine learning algorithm, whereby local computations are performed on local data with the output exchanged among neighbors in an iterative fashion. During this iterative process the leakage of data privacy arises. A differentially private ADMM was proposed in prior work (Zhang & Zhu, 2017) where only the privacy loss of a single node during one iteration was bounded, a method that makes it difficult to balance the tradeoff between the utility attained through distributed computation and privacy guarantees when considering the total privacy loss of all nodes over the entire iterative process. We propose a perturbation method for ADMM where the perturbed term is correlated with the penalty parameters; this is shown to improve the utility and privacy simultaneously. The method is based on a modified ADMM where each node independently determines its own penalty parameter in every iteration and decouples it from the dual updating step size. The condition for convergence of the modified ADMM and the lower bound on the convergence rate are also derived. Xueru Zhang, Mohammad Mahdi Khalili, Mingyan Liu |
ICML | 1 |
| 2013 | Rapid speaker adaptation in latent speaker space with non-negative matrix factorization
Xueru Zhang, Kris Demuynck, Hugo Van hamme |
Speech Commun. | 1 |
| 2012 | Latent variable speaker adaptation of Gaussian mixture weights and meansabstractWe describe a novel fast speaker adaptation algorithm for large vocabulary speech recognition systems, which adapts both the Gaussian means and the mixture weights. Gaussian means are expressed as a linear combination of eigenvoices estimated with principal component analysis. The non-negative Gaussian mixture weights are expressed as a linear combination of a set of latent vectors estimated with non-negative matrix factorization. Experiments on the Wall Street Journal database show that the combination of weight and mean adaptation consistently improves the performance compared to eigenvoice adaptation only. Improvements up to 5.8% relative word error rate reduction were observed with 40 eigenvoices and 40 latent weight vectors. Furthermore, combining weight and mean adaptation outperformed both weight and mean adaptation on itself, even if the latter uses more latent vectors. Xueru Zhang, Kris Demuynck, Hugo Van hamme |
ICASSP | 1 |
| 2011 | Rapid speaker adaptation with speaker adaptive training and non-negative matrix factorizationabstractIn this paper, we describe a novel speaker adaptation algorithm based on Gaussian mixture weight adaptation. A small number of latent speaker vectors are estimated with non-negative matrix factorization (NMF). These base vectors encode the correlations between Gaussian activations as learned from the train data. Expressing the speaker dependent Gaussian mixture weights as a linear combination of a small number of base vectors, reduces the number of parameters that must be estimated from the enrollment data. In order to learn meaningful correlations between Gaussian activations from the train data, the NMF-based weight adaptation was combined with vocal tract length normalization (VTLN) and feature-space maximum likelihood linear regression (fMLLR) based speaker adaptive training based. Evaluation on the 5k closed and 20k open vocabulary Wall Street Journal tasks shows a 4% relative word error rate reduction over the speaker independent recognition system which already incorporates VTLN. The proposed fast adaptation algorithm, using a single enrollment sentence only, results in similar performance as fMLLR adapting on 40 enrollment sentences. Xueru Zhang, Kris Demuynck, Hugo Van hamme |
ICASSP | 1 |
| 2010 | Histogram equalization and noise masking for robust speech recognitionabstractMismatch between training and test conditions deteriorates the performance of speech recognizers. This paper investigates the combination of parametric histogram equalization (pHEQ) and noise masking to compensate for the mismatch caused by additive noise. The proposed front-end maps the distribution of the observed power spectrum vectors to a target distribution. The target distribution matches the distribution of the noise free training data except for an artificially reduced signal-to-noise ratio. Different power spectrum estimation algorithms are used to estimate the noise distribution as used internally by pHEQ more reliably under non-stationary noise conditions. The proposed front-end is evaluated on the Aurora4 database and shows a significant improvement w.r.t. mean-normalized Mel-frequency spectral coefficients. Moreover, the performance could be further improved if better estimates of the instantaneous noise power spectrum were available. Xueru Zhang, Kris Demuynck, Hugo Van hamme |
ICASSP | 1 |
| 2010 | Feature versus model based noise robustnessabstractOver the years, the focus in noise robust speech recognition has shifted from noise robust features to model based techniques such as parallel model combination and uncertainty decoding. In this paper, we contrast prime examples of both approaches in the context of large vocabulary recognition systems such as used for automatic audio indexing and transcription. We look at the approximations the techniques require to keep the computational load reasonable, the resulting computational cost, and the accuracy measured on the Aurora4 benchmark. The results show that a well designed feature based scheme is capable of providing recognition accuracies at least as good as the model based approaches at a substantially lower computational cost. © 2010 ISCA. Kris Demuynck, Xueru Zhang, Dirk Van Compernolle, Hugo Van hamme |
INTERSPEECH | 2 |
| 2009 | Bayesian periodogram smoothing for speech enhancement
Xueru Zhang, Alexander Ypma, Bert de Vries |
ESANN | 1 |