Yongkai Wu

dblp:183/0976 · DBLP profile ↗
← Back
28ranked-venue papers
5as first author
16since 2021 · last 2025
0000-0002-7313-9439ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 24 · 4 first-author · 14 since 2021Databases, data management, data science and information retrieval · 10 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Theory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2025 Fair Graph U-Net: A Fair Graph Learning Framework Integrating Group and Individual Awareness
abstract
Learning high-level representations for graphs is crucial for tasks like node classification, where graph pooling aggregates node features to provide a holistic view that enhances predictive performance. Despite numerous methods that have been proposed in this promising and rapidly developing research field, most efforts to generalize the pooling operation to graphs are primarily performance-driven, with fairness issues largely overlooked: i) the process of graph pooling could exacerbate disparities in distribution among various subgroups; ii) the resultant graph structure augmentation may inadvertently strengthen intra-group connectivity, leading to unintended inter-group isolation. To this end, this paper extends the initial effort on fair graph pooling to the development of fair graph neural networks, while also providing a unified framework to collectively address group and individual graph fairness. Our experimental evaluations on multiple datasets demonstrate that the proposed method not only outperforms state-of-the-art baselines in terms of fairness but also achieves comparable predictive performance.
Zichong Wang, Zhibo Chu, Thang Viet Doan, Shaowei Wang 0002, Yongkai Wu, Vasile Palade, Wenbin Zhang 0002
AAAI5
2025 Towards counterfactual fairness through auxiliary variables
abstract
The challenge of balancing fairness and predictive accuracy in machine learning models, especially when sensitive attributes such as race, gender, or age are considered, has motivated substantial research in recent years. Counterfactual fairness ensures that predictions remain consistent across counterfactual variations of sensitive attributes, which is a crucial concept in addressing societal biases. However, existing counterfactual fairness approaches usually overlook intrinsic information about sensitive features, limiting their ability to achieve fairness while simultaneously maintaining performance. To tackle this challenge, we introduce EXOgenous Causal reasoning (EXOC), a novel causal reasoning framework motivated by exogenous variables. It leverages auxiliary variables to uncover intrinsic properties that give rise to sensitive attributes. Our framework explicitly defines an auxiliary node and a control node that contribute to counterfactual fairness and control the information flow within the model. Our evaluation, conducted on synthetic and real-world datasets, validates EXOC's superiority, showing that it outperforms state-of-the-art approaches in achieving counterfactual fairness without sacrificing accuracy. Our code is available at https://github.com/CASE-Lab-UMD/counterfactual_fairness_2025.
Bowei Tian, Shwai He, Wanghao Ye, Guoheng Sun, Yucong Dai, Yongkai Wu, Ang Li 0005
ICLR7
2024 Long-Term Fair Decision Making through Deep Generative Models
abstract
This paper studies long-term fair machine learning which aims to mitigate group disparity over the long term in sequential decision-making systems. To define long-term fairness, we leverage the temporal causal graph and use the 1-Wasserstein distance between the interventional distributions of different demographic groups at a sufficiently large time step as the quantitative metric. Then, we propose a three-phase learning framework where the decision model is trained on high-fidelity data generated by a deep generative model. We formulate the optimization problem as a performative risk minimization and adopt the repeated gradient descent algorithm for learning. The empirical evaluation shows the efficacy of the proposed method using both synthetic and semi-synthetic datasets.
Yaowei Hu 0001, Yongkai Wu, Lu Zhang 0021
AAAI2
2024 Learning Causally Disentangled Representations via the Principle of Independent Causal Mechanisms
Aneesh Komanduri, Yongkai Wu, Feng Chen 0001, Xintao Wu
IJCAI2
2024 Fair Weak-Supervised Learning: A Multiple-Instance Learning Approach
abstract
With the prevalence of machine learning in many high-stakes decision-making processes, e.g., hiring and admission, it is important to take fairness into account when practitioners design and deploy machine learning models, especially in scenarios with imperfectly labeled data. Multiple-Instance Learning (MIL) is a weakly supervised approach where instances are grouped in labeled bags, each containing several instances sharing the same label. However, current fairness-centric methods in machine learning often fall short when applied to MIL due to their reliance on instance-level labels. In this work, we introduce a Fair Multiple-Instance Learning (FMIL) framework to ensure fairness in weakly supervised learning. In particular, our method bridges the gap between bag-level and instance-level labeling by leveraging the bag labels, inferring high-confidence instance labels to improve both accuracy and fairness in MIL classifiers. Comprehensive experiments underscore that our FMIL framework substantially reduces biases in MIL without compromising accuracy.
Yucong Dai, Xiangyu Jiang, Yaowei Hu 0001, Lu Zhang 0021, Yongkai Wu
IJCNN5
2024 Achieving Equalized Explainability Through Data Reconstruction
abstract
Recent progress in machine learning has placed a growing emphasis on explainability and fairness. However, many studies have confined efforts to leveraging explanatory techniques to promote model fairness, overlooking the essential fairness of the explanations. This study addresses this gap by proposing a novel principle of equalized explainability, which fulfills fair and uniform explanations across various demographic groups. To this end, we introduce a quantitative measure for assessing the explanation disparity leveraging Explainable Artificial Intelligence (XAI) tools. To achieve equalized explainability, we propose a reconstruction framework including modules for data reconstruction, equalization of explanations, and performance preservation. Experiments using real-world datasets demonstrate this framework’s effectiveness in securing equitable and consistent explanations across different groups, as well as achieving trade-offs between fairness and explanation.
Shuang Wang 0010, Yongkai Wu
IJCNN2
2024 Achieving Fairness through Constrained Recourse
abstract
Data-driven decision-making systems are progressively deployed in high-risk scenarios, raising significant social concerns about their potential to perpetuate inequalities related to demographic characteristics. Recent research efforts have focused on ensuring equal decision-making for individuals, primarily through model adjustments and data modification techniques. However, these approaches often rely on distance-based formulation, overlooking the practical aspects of achieving equity. To address this gap, our study introduces a novel method that leverages actionable recourse to reflect the feasibility of attaining fairness in decision-making. This method leverages constrained optimization to achieve fairness within limited budgets, thereby balancing equity with practical constraints. We present experimental results that demonstrate the superiority of our approach over traditional distance-based methods. These results underscore our method's potential in ensuring equitable decisions and maintaining feasibility and efficiency in real-world applications.
Shuang Wang 0010, Yongkai Wu
IJCNN2
2024 SHED: Shapley-Based Automated Dataset Refinement for Instruction Fine-Tuning
abstract
The pre-trained Large Language Models (LLMs) can be adapted for many downstream tasks and tailored to align with human preferences through fine-tuning. Recent studies have discovered that LLMs can achieve desirable performance with only a small amount of high-quality data, suggesting that a large portion of the data in these extensive datasets is redundant or even harmful. Identifying high-quality data from vast datasets to curate small yet effective datasets has emerged as a critical challenge. In this paper, we introduce SHED, an automated dataset refinement framework based on Shapley value for instruction fine-tuning. SHED eliminates the need for human intervention or the use of commercial LLMs. Moreover, the datasets curated through SHED exhibit transferability, indicating they can be reused across different LLMs with consistently high performance. We conduct extensive experiments to evaluate the datasets curated by SHED. The results demonstrate SHED's superiority over state-of-the-art methods across various tasks and LLMs; notably, datasets comprising only 10% of the original data selected by SHED achieve performance comparable to or surpassing that of the full datasets.
Yexiao He, Zheyu Shen, Guoheng Sun, Yucong Dai, Yongkai Wu, Hongyi Wang 0001, Ang Li 0005
NeurIPS6
2024 Local Differential Privacy in Graph Neural Networks: a Reconstruction Approach
abstract
Graph Neural Networks have achieved tremendous success in modeling complex graph data in a variety of applications. However, there are limited studies investigating privacy protection in GNNs. In this work, we propose a learning framework that can provide local node privacy for users, while incurring low utility loss. We focus on a decentralized notion of Differential Privacy, namely Local Differential Privacy, and apply randomization mechanisms to perturb both feature and label data at the node level before they are collected by a server for model training. Specifically, we investigate the application of randomization mechanisms in high-dimensional feature settings and propose an LDP protocol with strict privacy guarantees. Based on frequency estimation in statistical analysis of randomized data, we develop reconstruction methods to approximate features and labels from perturbed data. We also formulate this learning framework to utilize frequency estimates of graph clusters to supervise the training procedure at a sub-graph level. Extensive experiments on real-world and semi-synthetic datasets demonstrate the validity of our proposed model.
Karuna Bhaila, Wen Huang 0003, Yongkai Wu, Xintao Wu
SDM3
2023 On Root Cause Localization and Anomaly Mitigation through Causal Inference
abstract
Due to a wide spectrum of applications in the real world, such as security, financial surveillance, and health risk, various deep anomaly detection models have been proposed and achieved state-of-the-art performance. However, besides being effective, in practice, the practitioners would further like to know what causes the abnormal outcome and how to further fix it. In this work, we propose RootCLAM, which aims to achieve Root Cause Localization and Anomaly Mitigation from a causal perspective. Especially, we formulate anomalies caused by external interventions on the normal causal mechanism and aim to locate the abnormal features with external interventions as root causes. After that, we further propose an anomaly mitigation approach that aims to recommend mitigation actions on abnormal features to revert the abnormal outcomes such that the counterfactuals guided by the causal mechanism are normal. Experiments on three datasets show that our approach can locate the root causes and further flip the abnormal labels.
Xiao Han 0008, Lu Zhang 0021, Yongkai Wu, Shuhan Yuan
CIKM3
2023 Neural Time-Invariant Causal Discovery from Time Series Data
abstract
Causal structure learning from observational data is an active field of research over the past decades. Although many approaches exist, such as constrained-based methods and score-based methods including the emerging deep learning-based methods, most of them address the static, non-dynamic setting. In this paper, we propose a score-based causal discovery algorithm named Neural Time-invariant Causal Discovery (NTiCD), which learns summary causal graphs from multivariate time series data based on the principle of Granger causality. NTiCD is a continuous optimization-based technique that leverages the power of deep neural networks to compute the score values. To this end, we use an LSTM to obtain the hidden non-linear representations of temporal variables in the time series data. Then, these features are aggregated using graph convolutional networks and decoded using an MLP that outputs the forecast of the future data values in the time series. The model is optimized based on a score function subject to regularized loss. The final output is a summary causal graph that captures the time-invariant causal relations within and between time series. We evaluate the performance of our algorithm on several synthetic and real datasets. The result analysis over a number of different datasets demonstrates the improvement in the accuracy of causal structure discovery of temporal data compared to other state-of-the-art methods.
Saima Absar, Yongkai Wu, Lu Zhang 0021
IJCNN2
2023 Fair Selection through Kernel Density Estimation
abstract
With the prevalence of machine learning in many high-stakes decision-making processes, e.g., hiring and admission, it is important to take fairness into consideration when practitioners design and deploy machine learning models. Although many approaches have been developed for fair machine learning, most of them focus on classification. In this paper, we target a notable but under-explored task, selection, where the number of selected individuals cannot exceed a pre-defined budget, such as employee hiring or university admission with limited positions or capabilities. In the selection task, existing fairness notions designed for classification are not suitable. In particular, our experimental results show that the selection models subject to common fairness notions may still make biased predictions against the underrepresented group. Hence, we propose a novel fairness notion, Selection Parity, which captures the demographic diversity among the selected groups in this restricted selection problem. Since the selection of qualified individuals with a fixed budget is non-differentiable, existing fairness regularization terms cannot be directly integrated with the selection task. To close the gap, we develop a novel in-processing framework named Fair Selection with the Differentiable Distribution Difference constraint (FS-DD), which incorporates a differentiable constraint into the training process and produces fair decisions for selection problems. Our theoretical analysis shows that common fairness metrics are bounded by the proposed Distribution Difference measurement. In other words, the FS-DD framework can guarantee fairness with regard to the common existing fairness metrics. We evaluate the performance of our method as well as several baselines on four real-world datasets. The experimental results demonstrate that the proposed method achieves fairness in various selection settings. In addition, the proposed method has a better fairness-accuracy trade-off compared with existing baseline methods.
Xiangyu Jiang, Yucong Dai, Yongkai Wu
IJCNN3
2023 Achieving Counterfactual Fairness for Anomaly Detection
Xiao Han 0008, Lu Zhang 0021, Yongkai Wu, Shuhan Yuan
PAKDD (1)3
2022 Fair Collective Classification in Networked Data
abstract
Collective classification utilizes network structure information via label propagation to improve prediction accuracy for node classification tasks. Because these models use information from previously labeled nodes which often contain historical bias, they may result in predictions that are biased w.r.t. the sensitive attributes of nodes such as race and gender. Throughout inference, this bias may even be amplified due to propagation especially for networks characterized by homophily. Despite past and ongoing research on fair classification, research to ensure fair collective classification s till remains unexplored. In this paper, we present a fair collective classification framework (denoted as FairCC) and formulate various heuristic methodologies, including node reweighting, threshold adjustment, and postprocessing, to achieve fair prediction. We also implement and test several naive methodologies for fair collective classification. Experiments on semi-synthetic datasets highlight the insufficiency of the naive methodologies and demonstrate the effectiveness of the proposed heuristics in significantly reducing prediction bias.
Karuna Bhaila, Yongkai Wu, Xintao Wu
IEEE Big Data2
2022 SCM-VAE: Learning Identifiable Causal Representations via Structural Knowledge
abstract
The goal of causal representation learning is to map low-level observations to high-level causal concepts to learn interpretable and robust representations for various downstream tasks. Latent variable models such as the variational autoencoder (VAE) are frequently leveraged to learn disentangled representations. However, there are often complex non-linear causal relationships underlying the observed data that cannot be captured through disentangled representations or linear dependence assumptions. Further, an independent conditional prior assumption can make learning causal dependencies in the latent space more challenging. We propose a framework, coined SCM-VAE, which uses apriori causal knowledge, a structural causal prior, and a non-linear additive noise structural causal model (SCM) to learn independent causal mechanisms and identifiable causal representations. We conduct theoretical analysis and perform experiments on synthetic and real-world datasets to show the improved quality of learned causal representations and robustness under interventions.
Aneesh Komanduri, Yongkai Wu, Wen Huang 0003, Feng Chen 0001, Xintao Wu
IEEE Big Data2
2021 A Generative Adversarial Framework for Bounding Confounded Causal Effects
abstract
Causal inference from observational data is receiving wide applications in many fields. However, unidentifiable situations, where causal effects cannot be uniquely computed from observational data, pose critical barriers to applying causal inference to complicated real applications. In this paper, we develop a bounding method for estimating the average causal effect (ACE) under unidentifiable situations due to hidden confounding based on Pearl's structural causal model. We propose to parameterize the unknown exogenous random variables and structural equations of a causal model using neural networks and implicit generative models. Then, using an adversarial learning framework, we search the parameter space to explicitly traverse causal models that agree with the given observational distribution, and find those that minimize or maximize the ACE to obtain its lower and upper bounds. The proposed method does not make assumption about the type of structural equations and variables. Experiments using both synthetic and real-world datasets are conducted.
Yaowei Hu 0001, Yongkai Wu, Lu Zhang 0021, Xintao Wu
AAAI2
2020 Fair Multiple Decision Making Through Soft Interventions
abstract
Previous research in fair classification mostly focuses on a single decision model. In reality, there usually exist multiple decision models within a system and all of which may contain a certain amount of discrimination. Such realistic scenarios introduce new challenges to fair classification: since discrimination may be transmitted from upstream models to downstream models, building decision models separately without taking upstream models into consideration cannot guarantee to achieve fairness. In this paper, we propose an approach that learns multiple classifiers and achieves fairness for all of them simultaneously, by treating each decision model as a soft intervention and inferring the post-intervention distributions to formulate the loss function as well as the fairness constraints. We adopt surrogate functions to smooth the loss function and constraints, and theoretically show that the excess risk of the proposed loss function can be bounded in a form that is the same as that for traditional surrogated loss functions. Experiments using both synthetic and real-world datasets show the effectiveness of our approach.
Yaowei Hu 0001, Yongkai Wu, Lu Zhang 0021, Xintao Wu
NeurIPS2
2019 Counterfactual Fairness: Unidentification, Bound and Algorithm
abstract
Fairness-aware learning studies the problem of building machine learning models that are subject to fairness requirements. Counterfactual fairness is a notion of fairness derived from Pearl's causal model, which considers a model is fair if for a particular individual or group its prediction in the real world is the same as that in the counterfactual world where the individual(s) had belonged to a different demographic group. However, an inherent limitation of counterfactual fairness is that it cannot be uniquely quantified from the observational data in certain situations, due to the unidentifiability of the counterfactual quantity. In this paper, we address this limitation by mathematically bounding the unidentifiable counterfactual quantity, and develop a theoretically sound algorithm for constructing counterfactually fair classifiers. We evaluate our method in the experiments using both synthetic and real-world datasets, as well as compare with existing methods. The results validate our theory and show the effectiveness of our method.
Yongkai Wu, Lu Zhang 0021, Xintao Wu
IJCAI1
2019 Achieving Causal Fairness through Generative Adversarial Networks
abstract
Achieving fairness in learning models is currently an imperative task in machine learning. Meanwhile, recent research showed that fairness should be studied from the causal perspective, and proposed a number of fairness criteria based on Pearl's causal modeling framework. In this paper, we investigate the problem of building causal fairness-aware generative adversarial networks (CFGAN), which can learn a close distribution from a given dataset, while also ensuring various causal fairness criteria based on a given causal graph. CFGAN adopts two generators, whose structures are purposefully designed to reflect the structures of causal graph and interventional graph. Therefore, the two generators can respectively simulate the underlying causal model that generates the real data, as well as the causal model after the intervention. On the other hand, two discriminators are used for producing a close-to-real distribution, as well as for achieving various fairness criteria based on causal quantities simulated by generators. Experiments on a real-world dataset show that CFGAN can generate high quality fair data.
Depeng Xu 0001, Yongkai Wu, Shuhan Yuan, Lu Zhang 0021, Xintao Wu
IJCAI2
2019 PC-Fairness: A Unified Framework for Measuring Causality-based Fairness
abstract
A recent trend of fair machine learning is to define fairness as causality-based notions which concern the causal connection between protected attributes and decisions. However, one common challenge of all causality-based fairness notions is identifiability, i.e., whether they can be uniquely measured from observational data, which is a critical barrier to applying these notions to real-world situations. In this paper, we develop a framework for measuring different causality-based fairness. We propose a unified definition that covers most of previous causality-based fairness notions, namely the path-specific counterfactual fairness (PC fairness). Based on that, we propose a general method in the form of a constrained optimization problem for bounding the path-specific counterfactual fairness under all unidentifiable situations. Experiments on synthetic and real-world datasets show the correctness and effectiveness of our method.
Yongkai Wu, Lu Zhang 0021, Xintao Wu, Hanghang Tong
NeurIPS1
2019 On Convexity and Bounds of Fairness-aware Classification
abstract
In this paper, we study the fairness-aware classification problem by formulating it as a constrained optimization problem. Several limitations exist in previous works due to the lack of a theoretical framework for guiding the formulation. We propose a general fairness-aware framework to address previous limitations. Our framework provides: (1) various fairness metrics that can be incorporated into classic classification models as constraints; (2) the convex constrained optimization problem that can be solved efficiently; and (3) the lower and upper bounds of real-world fairness measures that are established using surrogate functions, providing a fairness guarantee for constrained classifiers. Within the framework, we propose a constraint-free criterion under which any learned classifier is guaranteed to be fair in terms of the specified fairness metric. If the constraint-free criterion fails to satisfy, we further develop the method based on the bounds for constructing fair classifiers. The experiments using real-world datasets demonstrate our theoretical results and show the effectiveness of the proposed framework.
Yongkai Wu, Lu Zhang 0021, Xintao Wu
WWW1
2019 Causal Modeling-Based Discrimination Discovery and Removal: Criteria, Bounds, and Algorithms
abstract
Anti-discrimination is an increasingly important task in data science. In this paper, we investigate the problem of discovering both direct and indirect discrimination from the historical data, and removing the discriminatory effects before the data are used for predictive analysis (e.g., building classifiers). The main drawback of existing methods is that they cannot distinguish the part of influence that is really caused by discrimination from all correlated influences. In our approach, we make use of the causal graph to capture the causal structure of the data. Then, we model direct and indirect discrimination as the path-specific effects, which accurately identify the two types of discrimination as the causal effects transmitted along different paths in the graph. For certain situations where indirect discrimination cannot be exactly measured due to the unidentifiability of some path-specific effects, we develop an upper bound and a lower bound to the effect of indirect discrimination. Based on the theoretical results, we propose effective algorithms for discovering direct and indirect discrimination, as well as algorithms for precisely removing both types of discrimination while retaining good data utility. Experiments using the real dataset show the effectiveness of our approaches.
Lu Zhang 0021, Yongkai Wu, Xintao Wu
IEEE Trans. Knowl. Data Eng.2
2018 Achieving Non-Discrimination in Prediction
abstract
In discrimination-aware classification, the pre-process methods for constructing a discrimination-free classifier first remove discrimination from the training data, and then learn the classifier from the cleaned data. However, they lack a theoretical guarantee for the potential discrimination when the classifier is deployed for prediction. In this paper, we fill this gap by mathematically bounding the discrimination in prediction. We adopt the causal model for modeling the data generation mechanism, and formally defining discrimination in population, in a dataset, and in prediction. We obtain two important theoretical results: (1) the discrimination in prediction can still exist even if the discrimination in the training data is completely removed; and (2) not all pre-process methods can ensure non-discrimination in prediction even though they can achieve non-discrimination in the modified training data. Based on the results, we develop a two-phase framework for constructing a discrimination-free classifier with a theoretical guarantee. The experiments demonstrate the theoretical results and show the effectiveness of our two-phase framework.
Lu Zhang 0021, Yongkai Wu, Xintao Wu
IJCAI2
2018 On Discrimination Discovery and Removal in Ranked Data using Causal Graph
abstract
Predictive models learned from historical data are widely used to help companies and organizations make decisions. However, they may digitally unfairly treat unwanted groups, raising concerns about fairness and discrimination. In this paper, we study the fairness-aware ranking problem which aims to discover discrimination in ranked datasets and reconstruct the fair ranking. Existing methods in fairness-aware ranking are mainly based on statistical parity that cannot measure the true discriminatory effect since discrimination is causal. On the other hand, existing methods in causal-based anti-discrimination learning focus on classification problems and cannot be directly applied to handle the ranked data. To address these limitations, we propose to map the rank position to a continuous score variable that represents the qualification of the candidates. Then, we build a causal graph that consists of both the discrete profile attributes and the continuous score. The path-specific effect technique is extended to the mixed-variable causal graph to identify both direct and indirect discrimination. The relationship between the path-specific effects for the ranked data and those for the binary decision is theoretically analyzed. Finally, algorithms for discovering and removing discrimination from a ranked dataset are developed. Experiments using the real-world dataset show the effectiveness of our approaches.
Yongkai Wu, Lu Zhang 0021, Xintao Wu
KDD1
2017 A Causal Framework for Discovering and Removing Direct and Indirect Discrimination
abstract
In this paper, we investigate the problem of discovering both direct and indirect discrimination from the historical data, and removing the discriminatory effects before the data is used for predictive analysis (e.g., building classifiers). The main drawback of existing methods is that they cannot distinguish the part of influence that is really caused by discrimination from all correlated influences. In our approach, we make use of the causal network to capture the causal structure of the data. Then we model direct and indirect discrimination as the path-specific effects, which accurately identify the two types of discrimination as the causal effects transmitted along different paths in the network. Based on that, we propose an effective algorithm for discovering direct and indirect discrimination, as well as an algorithm for precisely removing both types of discrimination while retaining good data utility. Experiments using the real dataset show the effectiveness of our approaches.
Lu Zhang 0021, Yongkai Wu, Xintao Wu
IJCAI2
2017 Achieving Non-Discrimination in Data Release
abstract
Discrimination discovery and prevention/removal are increasingly important tasks in data mining. Discrimination discovery aims to unveil discriminatory practices on the protected attribute (e.g., gender) by analyzing the dataset of historical decision records, and discrimination prevention aims to remove discrimination by modifying the biased data before conducting predictive analysis. In this paper, we show that the key to discrimination discovery and prevention is to find the meaningful partitions that can be used to provide quantitative evidences for the judgment of discrimination. With the support of the causal graph, we present a graphical condition for identifying a meaningful partition. Based on that, we develop a simple criterion for the claim of non-discrimination, and propose discrimination removal algorithms which accurately remove discrimination while retaining good data utility. Experiments using real datasets show the effectiveness of our approaches.
Lu Zhang 0021, Yongkai Wu, Xintao Wu
KDD2
2016 Using Loglinear Model for Discrimination Discovery and Prevention
abstract
Discrimination discovery and prevention has received intensive attention recently. Discrimination generally refers to an unjustified distinction of individuals based on their membership, or perceived membership, in a certain group, and often occurs when the group is treated less favorably than others. However, existing discrimination discovery and prevention approaches are often limited to examining the relationship between one decision attribute and one protected attribute and do not sufficiently incorporate the effects due to other non-protected attributes. In this paper we develop a single unifying framework that aims to capture and measure discriminations between multiple decision attributes and protected attributes in addition to a set of non-protected attributes. Our approach is based on loglinear modeling. The coefficient values of the fitted loglinear model provide quantitative evidence of discrimination in decision making. The conditional independence graph derived from the fitted graphical loglinear model can be effectively used to capture the existence of discrimination patterns based on Markov properties. We further develop an algorithm to remove discrimination. The idea is modifying those significant coefficients from the fitted loglinear model and using the modified model to generate new data. Our empirical evaluation results show effectiveness of our proposed approach.
Yongkai Wu, Xintao Wu
DSAA1
2016 Situation Testing-Based Discrimination Discovery: A Causal Inference Approach
Lu Zhang 0021, Yongkai Wu, Xintao Wu
IJCAI2