VLDB 2026 Research / reviewers in the wild / expert
Lu Zhang 0021
dblp:71/2843-21
· DBLP profile ↗
39ranked-venue papers
15as first author
15since 2021 · last 2025
0000-0002-8972-8799ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 24 · 4 first-author · 13 since 2021Databases, data management, data science and information retrieval · 12 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 4 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 4 since 2021Systems, architecture and hardware · 4 · 4 first-authorComputer networks · 2 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Root Cause Analysis of Anomalies in Multivariate Time Series through Granger Causal DiscoveryabstractIdentifying the root causes of anomalies in multivariate time series is challenging due to the complex dependencies among the series. In this paper, we propose a comprehensive approach called AERCA that inherently integrates Granger causal discovery with root cause analysis. By defining anomalies as interventions on the exogenous variables of time series, AERCA not only learns the Granger causality among time series but also explicitly models the distributions of exogenous variables under normal conditions. AERCA then identifies the root causes of anomalies by highlighting exogenous variables that significantly deviate from their normal states. Experiments on multiple synthetic and real-world datasets demonstrate that AERCA can accurately capture the causal relationships among time series and effectively identify the root causes of anomalies. Xiao Han 0008, Saima Absar, Lu Zhang 0021, Shuhan Yuan |
ICLR | 3 |
| 2025 | A Causal Lens for Learning Long-term Fair PoliciesabstractFairness-aware learning studies the development of algorithms that avoid discriminatory decision outcomes despite biased training data. While most studies have concentrated on immediate bias in static contexts, this paper highlights the importance of investigating long-term fairness in dynamic decision-making systems while simultaneously considering instantaneous fairness requirements. In the context of reinforcement learning, we propose a general framework where long-term fairness is measured by the difference in the average expected qualification gain that individuals from different groups could obtain. Then, through a causal lens, we decompose this metric into three components that represent the direct impact, the delayed impact, as well as the spurious effect the policy has on the qualification gain. We analyze the intrinsic connection between these components and an emerging fairness notion called benefit fairness that aims to control the equity of outcomes in decision-making. Finally, we develop a simple yet effective approach for balancing various fairness notions. Jacob Lear, Lu Zhang 0021 |
ICLR | 2 |
| 2025 | CAMFeND: Credibility-Aware Multimodal Fake News Detection with Rotational AttentionabstractIn the evolving digital landscape, fake news is a significant challenge, influencing public perception and decision-making. Traditional detection approaches focus on single-modal data or simple multimodal fusion, often overlooking deeper interactions and news credibility. We propose a novel model addressing these limitations by introducing rotational attention and news domain information as a feature. Unlike static attention mechanisms, our rotational attention dynamically shifts query, key, and value roles across text and image inputs, enabling richer cross-modal interaction. Incorporating news domain information further enhances the model’s reliability by associating news posts with top domains extracted from Google search results, reducing false detections. This approach assesses both the content and the broader web context in which the news is discussed. Our model outperforms existing state-of-the-art methods by providing deeper, layered multimodal integration and domain information analysis, resulting in a more robust and adaptive fake news detection system. Nidhi Gupta, Lu Zhang 0021 |
IJCNN | 3 |
| 2024 | Long-Term Fair Decision Making through Deep Generative ModelsabstractThis paper studies long-term fair machine learning which aims to mitigate group disparity over the long term in sequential decision-making systems. To define long-term fairness, we leverage the temporal causal graph and use the 1-Wasserstein distance between the interventional distributions of different demographic groups at a sufficiently large time step as the quantitative metric. Then, we propose a three-phase learning framework where the decision model is trained on high-fidelity data generated by a deep generative model. We formulate the optimization problem as a performative risk minimization and adopt the repeated gradient descent algorithm for learning. The empirical evaluation shows the efficacy of the proposed method using both synthetic and semi-synthetic datasets. Yaowei Hu 0001, Yongkai Wu, Lu Zhang 0021 |
AAAI | 3 |
| 2024 | Time Series Causal Discovery Using a Hybrid MethodabstractIn this paper, we introduce a novel framework, Neural-HATS for inferring causal structures in time series data using a hybrid method. Neural-HATS uniquely combines conditional independence (CI) testing with continuous optimization-based learning methods to enhance causal discovery. Specifically, it leverages an attention-based encoder-decoder architecture with Kernel Conditional Independence (KCI) testing to enable direct CI tests between time series. These CI test results are integrated into continuous optimization algorithms, enhancing both causal inference accuracy and the effectiveness of continuous optimization models. Experimental evaluations demonstrate that Neural-HATS achieves improved causal graph accuracy. Saima Absar, Lu Zhang 0021 |
IEEE Big Data | 2 |
| 2024 | Fair Weak-Supervised Learning: A Multiple-Instance Learning ApproachabstractWith the prevalence of machine learning in many high-stakes decision-making processes, e.g., hiring and admission, it is important to take fairness into account when practitioners design and deploy machine learning models, especially in scenarios with imperfectly labeled data. Multiple-Instance Learning (MIL) is a weakly supervised approach where instances are grouped in labeled bags, each containing several instances sharing the same label. However, current fairness-centric methods in machine learning often fall short when applied to MIL due to their reliance on instance-level labels. In this work, we introduce a Fair Multiple-Instance Learning (FMIL) framework to ensure fairness in weakly supervised learning. In particular, our method bridges the gap between bag-level and instance-level labeling by leveraging the bag labels, inferring high-confidence instance labels to improve both accuracy and fairness in MIL classifiers. Comprehensive experiments underscore that our FMIL framework substantially reduces biases in MIL without compromising accuracy. Yucong Dai, Xiangyu Jiang, Yaowei Hu 0001, Lu Zhang 0021, Yongkai Wu |
IJCNN | 4 |
| 2023 | Striking a Balance in Fairness for Dynamic Systems Through Reinforcement LearningabstractWhile significant advancements have been made in the field of fair machine learning, the majority of studies focus on scenarios where the decision model operates on a static population. In this paper, we study fairness in dynamic systems where sequential decisions are made. Each decision may shift the underlying distribution of features or user behavior. We model the dynamic system through a Markov Decision Process (MDP). By acknowledging that traditional fairness notions and long-term fairness are distinct requirements that may not necessarily align with one another, we propose an algorithmic framework to integrate various fairness considerations with reinforcement learning using both pre-processing and in-processing approaches. Three case studies show that our method can strike a balance between traditional fairness notions, long-term fairness, and utility. Yaowei Hu 0001, Jacob Lear, Lu Zhang 0021 |
IEEE Big Data | 3 |
| 2023 | On Root Cause Localization and Anomaly Mitigation through Causal InferenceabstractDue to a wide spectrum of applications in the real world, such as security, financial surveillance, and health risk, various deep anomaly detection models have been proposed and achieved state-of-the-art performance. However, besides being effective, in practice, the practitioners would further like to know what causes the abnormal outcome and how to further fix it. In this work, we propose RootCLAM, which aims to achieve Root Cause Localization and Anomaly Mitigation from a causal perspective. Especially, we formulate anomalies caused by external interventions on the normal causal mechanism and aim to locate the abnormal features with external interventions as root causes. After that, we further propose an anomaly mitigation approach that aims to recommend mitigation actions on abnormal features to revert the abnormal outcomes such that the counterfactuals guided by the causal mechanism are normal. Experiments on three datasets show that our approach can locate the root causes and further flip the abnormal labels. Xiao Han 0008, Lu Zhang 0021, Yongkai Wu, Shuhan Yuan |
CIKM | 2 |
| 2023 | Neural Time-Invariant Causal Discovery from Time Series DataabstractCausal structure learning from observational data is an active field of research over the past decades. Although many approaches exist, such as constrained-based methods and score-based methods including the emerging deep learning-based methods, most of them address the static, non-dynamic setting. In this paper, we propose a score-based causal discovery algorithm named Neural Time-invariant Causal Discovery (NTiCD), which learns summary causal graphs from multivariate time series data based on the principle of Granger causality. NTiCD is a continuous optimization-based technique that leverages the power of deep neural networks to compute the score values. To this end, we use an LSTM to obtain the hidden non-linear representations of temporal variables in the time series data. Then, these features are aggregated using graph convolutional networks and decoded using an MLP that outputs the forecast of the future data values in the time series. The model is optimized based on a score function subject to regularized loss. The final output is a summary causal graph that captures the time-invariant causal relations within and between time series. We evaluate the performance of our algorithm on several synthetic and real datasets. The result analysis over a number of different datasets demonstrates the improvement in the accuracy of causal structure discovery of temporal data compared to other state-of-the-art methods. Saima Absar, Yongkai Wu, Lu Zhang 0021 |
IJCNN | 3 |
| 2023 | Achieving Counterfactual Fairness for Anomaly Detection
Xiao Han 0008, Lu Zhang 0021, Yongkai Wu, Shuhan Yuan |
PAKDD (1) | 2 |
| 2022 | Achieving Long-Term Fairness in Sequential Decision MakingabstractIn this paper, we propose a framework for achieving long-term fair sequential decision making. By conducting both the hard and soft interventions, we propose to take path-specific effects on the time-lagged causal graph as a quantitative tool for measuring long-term fairness. The problem of fair sequential decision making is then formulated as a constrained optimization problem with the utility as the objective and the long-term and short-term fairness as constraints. We show that such an optimization problem can be converted to a performative risk optimization. Finally, repeated risk minimization (RRM) is used for model training, and the convergence of RRM is theoretically analyzed. The empirical evaluation shows the effectiveness of the proposed algorithm on synthetic and semi-synthetic temporal datasets. Yaowei Hu 0001, Lu Zhang 0021 |
AAAI | 2 |
| 2022 | Achieving Counterfactual Fairness for Causal BanditabstractIn online recommendation, customers arrive in a sequential and stochastic manner from an underlying distribution and the online decision model recommends a chosen item for each arriving individual based on some strategy. We study how to recommend an item at each step to maximize the expected reward while achieving user-side fairness for customers, i.e., customers who share similar profiles will receive a similar reward regardless of their sensitive attributes and items being recommended. By incorporating causal inference into bandits and adopting soft intervention to model the arm selection strategy, we first propose the d-separation based UCB algorithm (D-UCB) to explore the utilization of the d-separation set in reducing the amount of exploration needed to achieve low cumulative regret. Based on that, we then propose the fair causal bandit (F-UCB) for achieving the counterfactual individual fairness. Both theoretical analysis and empirical evaluation demonstrate effectiveness of our algorithms. Wen Huang 0003, Lu Zhang 0021, Xintao Wu |
AAAI | 2 |
| 2022 | Coded Hate Speech Detection via Contextual Information
Depeng Xu 0001, Shuhan Yuan, Angela Uchechukwu Nwude, Lu Zhang 0021, Anna Zajicek, Xintao Wu |
PAKDD (1) | 5 |
| 2021 | A Generative Adversarial Framework for Bounding Confounded Causal EffectsabstractCausal inference from observational data is receiving wide applications in many fields. However, unidentifiable situations, where causal effects cannot be uniquely computed from observational data, pose critical barriers to applying causal inference to complicated real applications. In this paper, we develop a bounding method for estimating the average causal effect (ACE) under unidentifiable situations due to hidden confounding based on Pearl's structural causal model. We propose to parameterize the unknown exogenous random variables and structural equations of a causal model using neural networks and implicit generative models. Then, using an adversarial learning framework, we search the parameter space to explicitly traverse causal models that agree with the given observational distribution, and find those that minimize or maximize the ACE to obtain its lower and upper bounds. The proposed method does not make assumption about the type of structural equations and variables. Experiments using both synthetic and real-world datasets are conducted. Yaowei Hu 0001, Yongkai Wu, Lu Zhang 0021, Xintao Wu |
AAAI | 3 |
| 2021 | Discovering Time-invariant Causal Structure from Temporal DataabstractDiscovering causal structure from temporal data is an important problem in many fields in science. Existing methods usually suffer from several limitations such as assuming linear dependencies among features, limiting to discrete time series, and/or assuming stationarity, i.e., causal dependencies are repeated with the same time lag and strength at all time points. In this paper, we propose an algorithm called the μ-PC that addresses these limitations. It is based on the theory of μ-separation and extends the well-known PC algorithm to the time domain. To be applicable to both discrete and continuous time series, we develop a conditional independence testing technique for time series by leveraging the Recurrent Marked Temporal Point Process (RMTPP) model. Experiments using both synthetic and real-world datasets demonstrate the effectiveness of the proposed algorithm. Saima Absar, Lu Zhang 0021 |
CIKM | 2 |
| 2020 | Fair Multiple Decision Making Through Soft InterventionsabstractPrevious research in fair classification mostly focuses on a single decision model. In reality, there usually exist multiple decision models within a system and all of which may contain a certain amount of discrimination. Such realistic scenarios introduce new challenges to fair classification: since discrimination may be transmitted from upstream models to downstream models, building decision models separately without taking upstream models into consideration cannot guarantee to achieve fairness. In this paper, we propose an approach that learns multiple classifiers and achieves fairness for all of them simultaneously, by treating each decision model as a soft intervention and inferring the post-intervention distributions to formulate the loss function as well as the fairness constraints. We adopt surrogate functions to smooth the loss function and constraints, and theoretically show that the excess risk of the proposed loss function can be bounded in a form that is the same as that for traditional surrogated loss functions. Experiments using both synthetic and real-world datasets show the effectiveness of our approach. Yaowei Hu 0001, Yongkai Wu, Lu Zhang 0021, Xintao Wu |
NeurIPS | 3 |
| 2019 | FairGAN+: Achieving Fair Data Generation and Classification through Generative Adversarial NetsabstractHow to achieve fairness is important for next generation machine learning. Two tasks that are equally important in fair machine learning are how to obtain fair datasets and how to build fair classifiers. In this work, we propose a new generative adversarial network (GAN) model for fair machine learning, named FairGAN+. FairGAN+contains a generator to generate close-to-real samples, a classifier to predict class labels and three discriminators to assist adversarial learning. FairGAN+simultaneously achieves fair data generation and classification by co-training the generative model and the classifier through joint adversarial games with the discriminators. Evaluations on real world data show the effectiveness of FairGAN+on both fair data generation and fair classification. Depeng Xu 0001, Shuhan Yuan, Lu Zhang 0021, Xintao Wu |
IEEE BigData | 3 |
| 2019 | Counterfactual Fairness: Unidentification, Bound and AlgorithmabstractFairness-aware learning studies the problem of building machine learning models that are subject to fairness requirements. Counterfactual fairness is a notion of fairness derived from Pearl's causal model, which considers a model is fair if for a particular individual or group its prediction in the real world is the same as that in the counterfactual world where the individual(s) had belonged to a different demographic group. However, an inherent limitation of counterfactual fairness is that it cannot be uniquely quantified from the observational data in certain situations, due to the unidentifiability of the counterfactual quantity. In this paper, we address this limitation by mathematically bounding the unidentifiable counterfactual quantity, and develop a theoretically sound algorithm for constructing counterfactually fair classifiers. We evaluate our method in the experiments using both synthetic and real-world datasets, as well as compare with existing methods. The results validate our theory and show the effectiveness of our method. Yongkai Wu, Lu Zhang 0021, Xintao Wu |
IJCAI | 2 |
| 2019 | Achieving Causal Fairness through Generative Adversarial NetworksabstractAchieving fairness in learning models is currently an imperative task in machine learning. Meanwhile, recent research showed that fairness should be studied from the causal perspective, and proposed a number of fairness criteria based on Pearl's causal modeling framework. In this paper, we investigate the problem of building causal fairness-aware generative adversarial networks (CFGAN), which can learn a close distribution from a given dataset, while also ensuring various causal fairness criteria based on a given causal graph. CFGAN adopts two generators, whose structures are purposefully designed to reflect the structures of causal graph and interventional graph. Therefore, the two generators can respectively simulate the underlying causal model that generates the real data, as well as the causal model after the intervention. On the other hand, two discriminators are used for producing a close-to-real distribution, as well as for achieving various fairness criteria based on causal quantities simulated by generators. Experiments on a real-world dataset show that CFGAN can generate high quality fair data. Depeng Xu 0001, Yongkai Wu, Shuhan Yuan, Lu Zhang 0021, Xintao Wu |
IJCAI | 4 |
| 2019 | PC-Fairness: A Unified Framework for Measuring Causality-based FairnessabstractA recent trend of fair machine learning is to define fairness as causality-based notions which concern the causal connection between protected attributes and decisions. However, one common challenge of all causality-based fairness notions is identifiability, i.e., whether they can be uniquely measured from observational data, which is a critical barrier to applying these notions to real-world situations. In this paper, we develop a framework for measuring different causality-based fairness. We propose a unified definition that covers most of previous causality-based fairness notions, namely the path-specific counterfactual fairness (PC fairness). Based on that, we propose a general method in the form of a constrained optimization problem for bounding the path-specific counterfactual fairness under all unidentifiable situations. Experiments on synthetic and real-world datasets show the correctness and effectiveness of our method. Yongkai Wu, Lu Zhang 0021, Xintao Wu, Hanghang Tong |
NeurIPS | 2 |
| 2019 | On Convexity and Bounds of Fairness-aware ClassificationabstractIn this paper, we study the fairness-aware classification problem by formulating it as a constrained optimization problem. Several limitations exist in previous works due to the lack of a theoretical framework for guiding the formulation. We propose a general fairness-aware framework to address previous limitations. Our framework provides: (1) various fairness metrics that can be incorporated into classic classification models as constraints; (2) the convex constrained optimization problem that can be solved efficiently; and (3) the lower and upper bounds of real-world fairness measures that are established using surrogate functions, providing a fairness guarantee for constrained classifiers. Within the framework, we propose a constraint-free criterion under which any learned classifier is guaranteed to be fair in terms of the specified fairness metric. If the constraint-free criterion fails to satisfy, we further develop the method based on the bounds for constructing fair classifiers. The experiments using real-world datasets demonstrate our theoretical results and show the effectiveness of the proposed framework. Yongkai Wu, Lu Zhang 0021, Xintao Wu |
WWW | 2 |
| 2019 | Bayesian Network Construction and Genotype-Phenotype Inference Using GWAS StatisticsabstractGenome-wide association studies (GWASs) have received increasing attention to understand how genetic variation affects different human traits. In this paper, we study whether and to what extent exploiting the GWAS statistics can be used for inferring private information about a human individual. We first provide a method to construct a three-layered Bayesian network explicitly revealing the conditional dependency between single-nucleotide polymorphisms (SNPs) and traits from public GWAS catalog. The key challenge in building a Bayesian network from GWAS statistics is the specification of the conditional probability table of a variable with multiple parent variables. We employ the models of independence of causal influences which assume that the causal mechanism of each parent variable is mutually independent. We then formulate three inference problems based on the dependency relationship captured in the Bayesian network, namely trait inference given SNP genotype, genotype inference given trait, and trait inference given known traits, and develop efficient formulas and algorithms. Different from previous work, the possible target of these inference problems we study may be any individual, not limited to GWAS participants. Empirical evaluations show the effectiveness of our proposed methods. In summary, our work implies that meaningful information can be inferred from modeling GWAS statistics, and appropriate privacy protection mechanisms need to be developed to protect genetic privacy not only of GWAS participants but also regular individuals. Lu Zhang 0021, Qiuping Pan, Yue Wang 0009, Xintao Wu, Xinghua Shi |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2019 | Causal Modeling-Based Discrimination Discovery and Removal: Criteria, Bounds, and AlgorithmsabstractAnti-discrimination is an increasingly important task in data science. In this paper, we investigate the problem of discovering both direct and indirect discrimination from the historical data, and removing the discriminatory effects before the data are used for predictive analysis (e.g., building classifiers). The main drawback of existing methods is that they cannot distinguish the part of influence that is really caused by discrimination from all correlated influences. In our approach, we make use of the causal graph to capture the causal structure of the data. Then, we model direct and indirect discrimination as the path-specific effects, which accurately identify the two types of discrimination as the causal effects transmitted along different paths in the graph. For certain situations where indirect discrimination cannot be exactly measured due to the unidentifiability of some path-specific effects, we develop an upper bound and a lower bound to the effect of indirect discrimination. Based on the theoretical results, we propose effective algorithms for discovering direct and indirect discrimination, as well as algorithms for precisely removing both types of discrimination while retaining good data utility. Experiments using the real dataset show the effectiveness of our approaches. Lu Zhang 0021, Yongkai Wu, Xintao Wu |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2018 | FairGAN: Fairness-aware Generative Adversarial NetworksabstractFairness-aware learning is increasingly important in data mining. Discrimination prevention aims to prevent discrimination in the training data before it is used to conduct predictive analysis. In this paper, we focus on fair data generation that ensures the generated data is discrimination free. Inspired by generative adversarial networks (GAN), we present fairness-aware generative adversarial networks, called FairGAN, which are able to learn a generator producing fair data and also preserving good data utility. Compared with the naive fair data generation models, FairGAN further ensures the classifiers which are trained on generated data can achieve fair classification on real data. Experiments on a real dataset show the effectiveness of FairGAN. Depeng Xu 0001, Shuhan Yuan, Lu Zhang 0021, Xintao Wu |
IEEE BigData | 3 |
| 2018 | Achieving Non-Discrimination in PredictionabstractIn discrimination-aware classification, the pre-process methods for constructing a discrimination-free classifier first remove discrimination from the training data, and then learn the classifier from the cleaned data. However, they lack a theoretical guarantee for the potential discrimination when the classifier is deployed for prediction. In this paper, we fill this gap by mathematically bounding the discrimination in prediction. We adopt the causal model for modeling the data generation mechanism, and formally defining discrimination in population, in a dataset, and in prediction. We obtain two important theoretical results: (1) the discrimination in prediction can still exist even if the discrimination in the training data is completely removed; and (2) not all pre-process methods can ensure non-discrimination in prediction even though they can achieve non-discrimination in the modified training data. Based on the results, we develop a two-phase framework for constructing a discrimination-free classifier with a theoretical guarantee. The experiments demonstrate the theoretical results and show the effectiveness of our two-phase framework. Lu Zhang 0021, Yongkai Wu, Xintao Wu |
IJCAI | 1 |
| 2018 | On Discrimination Discovery and Removal in Ranked Data using Causal GraphabstractPredictive models learned from historical data are widely used to help companies and organizations make decisions. However, they may digitally unfairly treat unwanted groups, raising concerns about fairness and discrimination. In this paper, we study the fairness-aware ranking problem which aims to discover discrimination in ranked datasets and reconstruct the fair ranking. Existing methods in fairness-aware ranking are mainly based on statistical parity that cannot measure the true discriminatory effect since discrimination is causal. On the other hand, existing methods in causal-based anti-discrimination learning focus on classification problems and cannot be directly applied to handle the ranked data. To address these limitations, we propose to map the rank position to a continuous score variable that represents the qualification of the candidates. Then, we build a causal graph that consists of both the discrete profile attributes and the continuous score. The path-specific effect technique is extended to the mixed-variable causal graph to identify both direct and indirect discrimination. The relationship between the path-specific effects for the ranked data and those for the binary decision is theoretically analyzed. Finally, algorithms for discovering and removing discrimination from a ranked dataset are developed. Experiments using the real-world dataset show the effectiveness of our approaches. Yongkai Wu, Lu Zhang 0021, Xintao Wu |
KDD | 2 |
| 2017 | STIP: An SNP-trait inference platformabstractGenome-wide association studies (GWASs) have received increasing attention to understand how a genetic variation affects different human traits. Recent works show that the Bayesian network is powerful in modeling the conditional dependency between single-nucleotide polymorphisms (SNPs) and traits using only the GWAS statistics. In this paper, we present STIP, a web-based SNP-trait inference platform capable of a variety of inference tasks, such as trait inference given SNP genotypes and genotype inference given traits. The core of STIP is two Bayesian networks which model the SNP-categorical trait associations and SNP-quantitative trait associations, respectively. Both Bayesian networks are derived from the public GWAS catalog. The inference tasks are based on the dependency relationship captured in the Bayesian networks. The current version of STIP provides three services which are SNP-trait inference, Top-k trait prediction and GWAS catalog exploration. Qiuping Pan, Lu Zhang 0021, Xintao Wu |
BIBM | 2 |
| 2017 | Modeling SNP and quantitative trait association from GWAS catalog using CLG Bayesian networkabstractGenome-wide association studies (GWAS) are a type of genetic methods that have recently received intensive attention. In this paper, we study the construction of the Bayesian network from the GWAS catalog for modeling SNP and quantitative trait associations. Existing methods in the literature can only deal with categorical traits. We address this limitation by leveraging the Conditional Linear Gaussian (CLG) Bayesian network, which can handle a mixture of discrete and continuous variables. A two-layered CLG Bayesian network is built where the SNPs are represented as discrete variables in one layer and quantitative traits are represented as continuous variables in another layer. We propose the method for specifying the CLG Bayesian network, focusing on the specification of the CLG distribution for quantitative traits. We empirically evaluate the construction method, and results demonstrate the effectiveness of our method. Lu Zhang 0021, Qiuping Pan, Xintao Wu |
BIBM | 1 |
| 2017 | A Causal Framework for Discovering and Removing Direct and Indirect DiscriminationabstractIn this paper, we investigate the problem of discovering both direct and indirect discrimination from the historical data, and removing the discriminatory effects before the data is used for predictive analysis (e.g., building classifiers). The main drawback of existing methods is that they cannot distinguish the part of influence that is really caused by discrimination from all correlated influences. In our approach, we make use of the causal network to capture the causal structure of the data. Then we model direct and indirect discrimination as the path-specific effects, which accurately identify the two types of discrimination as the causal effects transmitted along different paths in the network. Based on that, we propose an effective algorithm for discovering direct and indirect discrimination, as well as an algorithm for precisely removing both types of discrimination while retaining good data utility. Experiments using the real dataset show the effectiveness of our approaches. Lu Zhang 0021, Yongkai Wu, Xintao Wu |
IJCAI | 1 |
| 2017 | Achieving Non-Discrimination in Data ReleaseabstractDiscrimination discovery and prevention/removal are increasingly important tasks in data mining. Discrimination discovery aims to unveil discriminatory practices on the protected attribute (e.g., gender) by analyzing the dataset of historical decision records, and discrimination prevention aims to remove discrimination by modifying the biased data before conducting predictive analysis. In this paper, we show that the key to discrimination discovery and prevention is to find the meaningful partitions that can be used to provide quantitative evidences for the judgment of discrimination. With the support of the causal graph, we present a graphical condition for identifying a meaningful partition. Based on that, we develop a simple criterion for the claim of non-discrimination, and propose discrimination removal algorithms which accurately remove discrimination while retaining good data utility. Experiments using real datasets show the effectiveness of our approaches. Lu Zhang 0021, Yongkai Wu, Xintao Wu |
KDD | 1 |
| 2017 | Analysis of Minimum Interaction Time for Continuous Distributed Interactive ComputingabstractDistributed interactive computing allows participants at different locations to interact with each other in real time. In this paper, we study the interaction times of continuous Distributed Interactive Applications (DIAs) in which the application states change due to not only user-initiated operations but also time passing. Given the clients and servers of a continuous DIA, its interaction time is directly affected by how the clients are assigned to the servers as well as the simulation time settings of the servers. We formulate the Minimum Interaction Time (MIT) problem as a combinatorial problem of these two tuning knobs and prove that it is NP-hard. We then approximate the problem by fixing the client assignment or the simulation time offsets among the servers. When the client assignment is fixed, we show that finding the minimum achievable interaction time can be reduced to a weighted bipartite matching problem. We further show that this approach establishes a tight approximation factor of 3 to the MIT problem if each client is assigned to its nearest server. When the simulation time offsets among the servers are fixed, we show that finding the minimum achievable interaction time is still NP-hard. This approach can approximate the MIT problem by a factor within 2 if the simulation times of all servers are synchronized. A mix of the above two approaches better approximates the MIT problem within a factor of 5/3. We further conduct experimental evaluation of these approaches with three real Internet latency datasets. Lu Zhang 0021, Xueyan Tang, Bingsheng He |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2016 | Building Bayesian networks from GWAS statistics based on Independence of Causal InfluenceabstractGenome-wide association studies (GWASs) have received an increasing attention to understand genotype-phenotype relationships. In this paper, we study how to build Bayesian networks from publicly released GWAS statistics to explicitly reveal the conditional dependency between single-nucleotide polymorphisms (SNPs) and traits. The key challenge in building a Bayesian network is the specification of the conditional probability table (CPT) of an variable with multiple parent variables. We employ the Independence of Causal Influences (ICI) which assumes that the causal mechanism of each parent variable is mutually independent. Specifically, we derive a formulation from the Noisy-or model, one of the ICI models, to specify the CPT using the released GWAS statistics. We prove that the specified CPT is accurate as long as the underlying individual-level genotype and phenotype profile data follows the Noisy-or model. We empirically evaluate the Noisy-or model and its derived formulation using data from openSNP. Experimental results demonstrate the effectiveness of our approach. Lu Zhang 0021, Qiuping Pan, Xintao Wu, Xinghua Shi |
BIBM | 1 |
| 2016 | Situation Testing-Based Discrimination Discovery: A Causal Inference Approach
Lu Zhang 0021, Yongkai Wu, Xintao Wu |
IJCAI | 1 |
| 2014 | The Client Assignment Problem for Continuous Distributed Interactive Applications: Analysis, Algorithms, and EvaluationabstractInteractivity is a primary performance measure for distributed interactive applications (DIAs) that enable participants at different locations to interact with each other in real time. Wide geographical spreads of participants in large-scale DIAs necessitate distributed deployment of servers to improve interactivity. In a distributed server architecture, the interactivity performance depends on not only client-to-server network latencies but also interserver network latencies, as well as synchronization delays to meet the consistency and fairness requirements of DIAs. All of these factors are directly affected by how the clients are assigned to the servers. In this paper, we investigate the problem of effectively assigning clients to servers for maximizing the interactivity of DIAs. We focus on continuous DIAs that changes their states not only in response to user operations but also due to the passing of time. We analyze the minimum achievable interaction time for DIAs to preserve consistency and provide fairness among clients, and formulate the client assignment problem as a combinatorial optimization problem. We prove that this problem is NP-complete. Three heuristic assignment algorithms are proposed and their approximation ratios are theoretically analyzed. The performance of the algorithms is also experimentally evaluated using real Internet latency data. The experimental results show that our proposed Greedy Assignment and Distributed-Modify Assignment algorithms generally produce near optimal interactivity and significantly reduce the interaction time between clients compared to the intuitive algorithm that assigns each client to its nearest server. Lu Zhang 0021, Xueyan Tang |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2013 | Sesame: A new bioinformatics semantic workflow design systemabstractBiologists have become increasingly dependent on bioinformatics tools to analyze and interpret their datasets. The number, variety and complexity of these bioinformatics tools have increased dramatically and they have become more and more computationally complex, expensive and resource intensive. Powerful workflow design systems have been developed to automate the execution of set of tools for a specific task. However, designing a complex executable workflow using such tools still requires considerable computational expertise or the help from a bioinformatics expert. In this paper, we developed Sesame, a bioinformatics semantic workflow design system. We have designed a new ontology for bioinformatics tools and services (OBTS) and proposed an ontology driven semantic workflow design mechanism, using this new OBTS. Compared to an executable workflow, the semantic workflow is at the level of biological concepts that are closer to the scientific research. Biologists will greatly benefit from the decoupling of semantic workflow design from executable workflow design with computational implementation details. Currently, a prototype version of Sesame system has been implemented and deployed. Sesame will allow biologists to efficiently perform complex data analysis to address scientific questions. Lu Zhang 0021, Pengfei Xuan, Alexander Duvall, Jonathan Lowe, Arvind Subramanian, Pradip K. Srimani, Feng Luo 0001, Yongping Duan |
BIBM | 1 |
| 2013 | Brief announcement: on minimum interaction time for continuous distributed interactive computingabstractIn this paper, we study the interaction times of continuous distributed interactive computing in which the application states change due to not only user-initiated operations but also time passing. We formulate the Minimum Interaction Time problem as a combinatorial problem of how the clients are assigned to the servers and the simulation time settings of the servers. We also outline two approaches to approximate the problem. Lu Zhang 0021, Xueyan Tang, Bingsheng He |
PODC | 1 |
| 2012 | Optimizing client assignment for enhancing interactivity in distributed interactive applicationsabstractDistributed interactive applications (DIAs) are networked systems that allow multiple participants at different locations to interact with each other. Wide spreads of client locations in large-scale DIAs often require geographical distribution of servers to meet the latency requirements of the applications. In the distributed server architecture, the network latencies involved in the interactions between clients are directly affected by how the clients are assigned to the servers. In this paper, we focus on the problem of assigning clients to appropriate servers in DIAs to enhance their interactivity. We formulate the problem as a combinational optimization problem and prove that it is NP-complete. Then, we propose several heuristic algorithms for fast computation of good client assignments and theoretically analyze their approximation ratios. The proposed algorithms are also experimentally evaluated with real Internet latency data. The results show that the proposed algorithms are efficient and effective in reducing the interaction time between clients, and our proposed Distributed-Modify-Assignment adapts well to the dynamics of client participation and network conditions. For the special case of tree network topologies, we develop a polynomial-time algorithm to compute the optimal client assignment. Lu Zhang 0021, Xueyan Tang |
IEEE/ACM Trans. Netw. | 1 |
| 2011 | The Client Assignment Problem for Continuous Distributed Interactive ApplicationsabstractInteractivity is a primary performance measure for distributed interactive applications (DIAs) that enable participants at different locations to interact with each other in real time. Wide geographical spreads of participants in large-scale DIAs necessitate distributed deployment of servers to improve interactivity. In a distributed server architecture, the interactivity performance depends on not only client-to-server network latencies but also inter-server network latencies as well as synchronization delays to meet the consistency and fairness requirements of DIAs. All of these factors are directly affected by how the clients are assigned to the servers. In this paper, we investigate the problem of effectively assigning clients to servers for maximizing the interactivity of DIAs. We focus on continuous DIAs that change their states not only in response to user operations but also due to the passing of time. We analyze the minimum achievable interaction time for DIAs to preserve consistency and provide fairness among clients, and formulate the client assignment problem as a combinational optimization problem. We prove that this problem is NP-complete. Four heuristic assignment algorithms are proposed and evaluated using real Internet latency data. The experimental results show that our proposed greedy algorithm generally produces near optimal interactivity and significantly reduces the interaction time between clients compared to the intuitive algorithm that assigns each client to its nearest server. Lu Zhang 0021, Xueyan Tang |
ICDCS | 1 |
| 2011 | Client assignment for improving interactivity in distributed interactive applicationsabstractDistributed Interactive Applications (DIAs) are networked systems that allow multiple participants to interact with one another in real time. Wide spreads of client locations in larges-cale DIAs often require geographical distribution of servers to meet the latency requirements of the applications. In the distributed server architecture, how the clients are assigned to the servers directly affects the network latency involved in the interactions between clients. This paper focuses on the client assignment problem for enhancing the interactivity performance of DIAs. We formulate the problem as a combinational optimization problem on graphs and prove that it is NP-complete. Several heuristic algorithms are proposed for fast computation of good client assignments and are experimentally evaluated. The experimental results show that the proposed greedy algorithms perform close to the optimal assignment and generally outperform the Nearest-Assignment algorithm that assigns each client to its nearest server. Lu Zhang 0021, Xueyan Tang |
INFOCOM | 1 |