EDBT 2026 Demo / reviewers in the wild / expert
Jieyu Zhao 0001
dblp:59/2379-1
· DBLP profile ↗
25ranked-venue papers
4as first author
15since 2021 · last 2026
0009-0003-9956-5481ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 25 · 4 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | WildFeedback: Aligning LLMs With In-situ User Interactions And FeedbackabstractTaiwei Shi, Zhuoer Wang, Longqi Yang, Ying-Chun Lin, Zexue He, Mengting Wan, Pei Zhou, Sujay Kumar Jauhar, Sihao Chen, Shan Xia, Hongfei Zhang, Jieyu Zhao, Xiaofeng Xu, Xia Song, Jennifer Neville. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Taiwei Shi, Zhuoer Wang, Longqi Yang 0001, Ying-Chun Lin, Zexue He, Mengting Wan, Sujay Kumar Jauhar, Shan Xia, Jieyu Zhao 0001, Jennifer Neville |
ACL (1) | 12 |
| 2025 | Cross-Lingual Pitfalls: Automatic Probing Cross-Lingual Weakness of Multilingual Large Language ModelsabstractZixiang Xu, Yanbo Wang, Yue Huang, Xiuying Chen, Jieyu Zhao, Meng Jiang, Xiangliang Zhang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Zixiang Xu, Yanbo Wang 0005, Yue Huang 0001, Xiuying Chen, Jieyu Zhao 0001, Meng Jiang 0001, Xiangliang Zhang 0001 |
ACL (1) | 5 |
| 2025 | AI Sees Your Location - But With A Bias Toward The Wealthy WorldabstractVisual-Language Models (VLMs) have shown remarkable performance across various tasks, particularly in recognizing geographic information from images.However, VLMs still show regional biases in this task.To systematically evaluate these issues, we introduce a benchmark consisting of 1,200 images paired with detailed geographic metadata.Evaluating four VLMs, we find that while these models demonstrate the ability to recognize geographic information from images, achieving up to 53.8% accuracy in city prediction, they exhibit significant biases.Specifically, performance is substantially higher for economically developed and densely populated regions compared to less developed (-12.5%)and sparsely populated (-17.0%)areas.Moreover, regional biases of frequently over-predicting certain locations remain.For instance, they consistently predict Sydney for images taken in Australia, shown by the low entropy scores for these countries.The strong performance of VLMs also raises privacy concerns, particularly for users who share images online without the intent of being identified.Our code and dataset are publicly available at https://github.com/uscnlp-lime/ FairLocator. Jen-tse Huang 0001, Wenxuan Wang 0001, Jieyu Zhao 0001 |
EMNLP | 6 |
| 2025 | VisBias: Measuring Explicit and Implicit Social Biases in Vision Language ModelsabstractThis research investigates both explicit and implicit social biases exhibited by Vision-Language Models (VLMs).The key distinction between these bias types lies in the level of awareness: explicit bias refers to conscious, intentional biases, while implicit bias operates subconsciously.To analyze explicit bias, we directly pose questions to VLMs related to gender and racial differences: (1) Multiplechoice questions based on a given image (e.g., "What is the education level of the person in the image?")(2) Yes-No comparisons using two images (e.g., "Is the person in the first image more educated than the person in the second image?")For implicit bias, we design tasks where VLMs assist users but reveal biases through their responses: (1) Image description tasks: Models are asked to describe individuals in images, and we analyze disparities in textual cues across demographic groups.(2) Form completion tasks: Models draft a personal information collection form with 20 attributes, and we examine correlations among selected attributes for potential biases.We evaluate Gemini-1.5,GPT-4V, GPT-4o, LLaMA-3.2-Vision and LLaVA-v1.6.Our code and data are publicly available at https: //github.com/uscnlp-lime/VisBias.Note: This paper includes examples of potentially offensive texts generated by VLMs. Jen-tse Huang 0001, Jiantong Qin, Jianping Zhang 0002, Youliang Yuan, Wenxuan Wang 0001, Jieyu Zhao 0001 |
EMNLP | 6 |
| 2025 | MUSE: Machine Unlearning Six-Way Evaluation for Language ModelsabstractLanguage models (LMs) are trained on vast amounts of text data, which may include private and copyrighted content. Data owners may request the removal of their data from a trained model due to privacy or copyright concerns. However, exactly unlearning only these datapoints (i.e., retraining with the data removed) is intractable in modern-day models. This has led to the development of many approximate unlearning algorithms. The evaluation of the efficacy of these algorithms has traditionally been narrow in scope, failing to precisely quantify the success and practicality of the algorithm from the perspectives of both the model deployers and the data owners. We address this issue by proposing MUSE, a comprehensive machine unlearning evaluation benchmark that enumerates six diverse desirable properties for unlearned models: (1) no verbatim memorization, (2) no knowledge memorization, (3) no privacy leakage, (4) utility preservation on data not intended for removal, (5) scalability with respect to the size of removal requests, and (6) sustainability over sequential unlearning requests. Using these criteria, we benchmark how effectively eight popular unlearning algorithms on 7B-parameter LMs can unlearn Harry Potter books and news articles. Our results demonstrate that most algorithms can prevent verbatim memorization and knowledge memorization to varying degrees, but only one algorithm does not lead to severe privacy leakage. Furthermore, existing algorithms fail to meet deployer's expectations because they often degrade general model utility and also cannot sustainably accommodate successive unlearning requests or large-scale content removal. Our findings identify key issues with the practicality of existing unlearning algorithms on language models. Jaechan Lee, Yangsibo Huang, Sadhika Malladi, Jieyu Zhao 0001, Ari Holtzman, Daogao Liu, Luke Zettlemoyer, Noah A. Smith, Chiyuan Zhang |
ICLR | 5 |
| 2024 | InterIntent: Investigating Social Intelligence of LLMs via Intention Understanding in an Interactive Game ContextabstractLarge language models (LLMs) have demonstrated the potential to mimic human social intelligence.However, most studies focus on simplistic and static self-report or performancebased tests, which limits the depth and validity of the analysis.In this paper, we developed a novel framework, INTERINTENT, to assess LLMs' social intelligence by mapping their ability to understand and manage intentions in a game setting.We focus on four dimensions of social intelligence: situational awareness, selfregulation, self-awareness, and theory of mind.Each dimension is linked to a specific game task: intention selection, intention following, intention summarization, and intention guessing.Our findings indicate that while LLMs exhibit high proficiency in selecting intentions, achieving an accuracy of 88%, their ability to infer the intentions of others is significantly weaker, trailing human performance by 20%.Additionally, game performance correlates with intention understanding, highlighting the importance of the four components towards success in this game.These findings underline the crucial role of intention understanding in evaluating LLMs' social intelligence and highlight the potential of using social deduction games as a complex testbed to enhance LLM evaluation.INTERINTENT contributes a structured approach to bridging the evaluation gap in social intelligence within multiplayer games. 1 Game ContextRound: 2 Previous round summary: All players vote "agree" to the team proposal including Player 1 and Player 4 and the quest is successful.Roles: Servant does not have any information; Merlin knows who are evil players but they cannot reveal their identity ... Current round discussion: Player 2: I propose a team including Player 1, Player 2, and Player 3. Player 1 shows their loyalty in the last quest and I can promise you I am loyal to the king of Arthur!Player 3 hasn't proved themselves in the quest yet and let's give them a chance!(1) Situational Awareness Intention Selection I am a servant and I do not have any information.At this point, I should choose the intention "Support team proposal" as I agree with Player 2. Abhishek Anand, Jen-tse Huang 0001, Jieyu Zhao 0001 |
EMNLP | 5 |
| 2024 | "You Gotta be a Doctor, Lin" : An Investigation of Name-Based Bias of Large Language Models in Employment RecommendationsabstractSocial science research has shown that candidates with names indicative of certain races or genders often face discrimination in employment practices.Similarly, Large Language Models (LLMs) have demonstrated racial and gender biases in various applications.In this study, we utilize GPT-3.5-Turbo and Llama 3-70B-Instruct to simulate hiring decisions and salary recommendations for candidates with 320 first names that strongly signal their race and gender, across over 750,000 prompts.Our empirical results indicate a preference among these models for hiring candidates with White female-sounding names over other demographic groups across 40 occupations.Additionally, even among candidates with identical qualifications, salary recommendations vary by as much as 5% between different subgroups.A comparison with real-world labor data reveals inconsistent alignment with U.S. labor market characteristics, underscoring the necessity of risk investigation of LLM-powered systems. Huy Nghiem, John Prindle, Jieyu Zhao 0001, Hal Daumé III |
EMNLP | 3 |
| 2024 | Images Speak Louder than Words: Understanding and Mitigating Bias in Vision-Language Model from a Causal Mediation PerspectiveabstractVision-language models (VLMs) pre-trained on extensive datasets can inadvertently learn biases by correlating gender information with specific objects or scenarios.Current methods, which focus on modifying inputs and monitoring changes in the model's output probability scores, often struggle to comprehensively understand bias from the perspective of model components.We propose a framework that incorporates causal mediation analysis to measure and map the pathways of bias generation and propagation within VLMs.Our framework is applicable to a wide range of vision-language and multimodal tasks.In this work, we apply it to the object detection task and implement it on the GLIP model.This approach allows us to identify the direct effects of interventions on model bias and the indirect effects of interventions on bias mediated through different model components.Our results show that image features are the primary contributors to bias, with significantly higher impacts than text features, specifically accounting for 32.57% and 12.63% of the bias in the MSCOCO and PASCAL-SENTENCE datasets, respectively.Notably, the image encoder's contribution surpasses that of the text encoder and the deep fusion encoder.Further experimentation confirms that contributions from both language and vision modalities are aligned and non-conflicting.Consequently, focusing on blurring gender representations within the image encoder which contributes most to the model bias, reduces bias efficiently by 22.03% and 9.04% in the MSCOCO and PASCAL-SENTENCE datasets, respectively, with minimal performance loss or increased computational demands. 1 Zhaotian Weng, Zijun Gao, Jerone Theodore Alexander Andrews, Jieyu Zhao 0001 |
EMNLP | 4 |
| 2024 | Position: TrustLLM: Trustworthiness in Large Language ModelsabstractLarge language models (LLMs) have gained considerable attention for their excellent natural language processing capabilities. Nonetheless, these LLMs present many challenges, particularly in the realm of trustworthiness. This paper introduces TrustLLM, a comprehensive study of trustworthiness in LLMs, including principles for different dimensions of trustworthiness, established benchmark, evaluation, and analysis of trustworthiness for mainstream LLMs, and discussion of open challenges and future directions. Specifically, we first propose a set of principles for trustworthy LLMs that span eight different dimensions. Based on these principles, we further establish a benchmark across six dimensions including truthfulness, safety, fairness, robustness, privacy, and machine ethics. We then present a study evaluating 16 mainstream LLMs in TrustLLM, consisting of over 30 datasets. Our findings firstly show that in general trustworthiness and capability (i.e., functional effectiveness) are positively related. Secondly, our observations reveal that proprietary LLMs generally outperform most open-source counterparts in terms of trustworthiness, raising concerns about the potential risks of widely accessible open-source LLMs. However, a few open-source LLMs come very close to proprietary ones, suggesting that open-source models can achieve high levels of trustworthiness without additional mechanisms like moderator, offering valuable insights for developers in this field. Thirdly, it is important to note that some LLMs may be overly calibrated towards exhibiting trustworthiness, to the extent that they compromise their utility by mistakenly treating benign prompts as harmful and consequently not responding. Besides these observations, we’ve uncovered key insights into the multifaceted trustworthiness in LLMs. We emphasize the importance of ensuring transparency not only in the models themselves but also in the technologies that underpin trustworthiness. We advocate that the establishment of an AI alliance between industry, academia, the open-source community to foster collaboration is imperative to advance the trustworthiness of LLMs. Yue Huang 0001, Lichao Sun 0001, Haoran Wang 0005, Siyuan Wu 0001, Qihui Zhang, Chujie Gao, Wenhan Lyu, Yixuan Zhang 0001, Xiner Li, Hanchi Sun, Zhengliang Liu, Yixin Liu 0002, Yijue Wang, Bertie Vidgen, Bhavya Kailkhura, Caiming Xiong, Chaowei Xiao, Chunyuan Li, Eric P. Xing, Furong Huang, Heng Ji 0001, Hongyi Wang 0001, Huan Zhang 0001, Huaxiu Yao, Manolis Kellis, Marinka Zitnik, Meng Jiang 0001, Mohit Bansal, James Zou 0001, Jian Pei 0001, Jianfeng Gao 0001, Jiawei Han 0001, Jieyu Zhao 0001, Jiliang Tang, Jindong Wang 0001, Joaquin Vanschoren, John C. Mitchell, Kai Shu, Kaidi Xu, Kai-Wei Chang 0001, Lifang He 0001, Lifu Huang, Michael Backes 0001, Neil Zhenqiang Gong, Philip S. Yu, Quanquan Gu, Ran Xu 0001, Rex Ying, Shuiwang Ji, Suman Jana, Tianlong Chen 0001, Tianming Liu 0001, Tianyi Zhou 0001, William Yang Wang, Xiang Li 0001, Xiangliang Zhang 0001, Xiao Wang 0012, Xing Xie 0001, Xuyu Wang, Yan Liu 0002, Yanfang Ye 0001, Yinzhi Cao, Yong Chen 0016, Yue Zhao 0016 |
ICML | 38 |
| 2024 | Adapting Static Fairness to Sequential Decision-Making: Bias Mitigation Strategies towards Equal Long-term Benefit RateabstractDecisions made by machine learning models can have lasting impacts, making long-term fairness a critical consideration. It has been observed that ignoring the long-term effect and directly applying fairness criterion in static settings can actually worsen bias over time. To address biases in sequential decision-making, we introduce a long-term fairness concept named Equal Long-term Benefit Rate (ELBERT). This concept is seamlessly integrated into a Markov Decision Process (MDP) to consider the future effects of actions on long-term fairness, thus providing a unified framework for fair sequential decision-making problems. ELBERT effectively addresses the temporal discrimination issues found in previous long-term fairness notions. Additionally, we demonstrate that the policy gradient of Long-term Benefit Rate can be analytically simplified to standard policy gradients. This simplification makes conventional policy optimization methods viable for reducing bias, leading to our bias mitigation approach ELBERT-PO. Extensive experiments across various diverse sequential decision-making environments consistently reveal that ELBERT-PO significantly diminishes bias while maintaining high utility. Code is available at https://github.com/umd-huang-lab/ELBERT. Yuancheng Xu, Chenghao Deng, Yanchao Sun, Ruijie Zheng, Jieyu Zhao 0001, Furong Huang |
ICML | 6 |
| 2024 | Fair Abstractive Summarization of Diverse PerspectivesabstractYusen Zhang, Nan Zhang, Yixin Liu, Alexander Fabbri, Junru Liu, Ryo Kamoi, Xiaoxin Lu, Caiming Xiong, Jieyu Zhao, Dragomir Radev, Kathleen McKeown, Rui Zhang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Yusen Zhang 0001, Yixin Liu 0003, Alexander R. Fabbri, Junru Liu, Ryo Kamoi, Xiaoxin Lu, Caiming Xiong, Jieyu Zhao 0001, Dragomir R. Radev, Kathy McKeown, Rui Zhang 0037 |
NAACL-HLT | 9 |
| 2023 | SODAPOP: Open-Ended Discovery of Social Biases in Social Commonsense Reasoning ModelsabstractA common limitation of diagnostic tests for detecting social biases in NLP models is that they may only detect stereotypic associations that are pre-specified by the designer of the test.Since enumerating all possible problematic associations is infeasible, it is likely these tests fail to detect biases that are present in a model but not pre-specified by the designer.To address this limitation, we propose SODAPOP 1 (SOcial bias Discovery from Answers about PeOPle), an approach for automatic social bias discovery in social commonsense question-answering.The SODAPOP pipeline generates modified instances from the Social IQa dataset (Sap et al., 2019b) by ( 1) substituting names associated with different demographic groups, and (2) generating many distractor answers from a masked language model.By using a social commonsense model to score the generated distractors, we are able to uncover the model's stereotypic associations between demographic groups and an open set of words.We also test SODAPOP on debiased models and show the limitations of multiple state-of-the-art debiasing algorithms. Haozhe An, Zongxia Li, Jieyu Zhao 0001, Rachel Rudinger |
EACL | 3 |
| 2023 | A Rose by Any Other Name would not Smell as Sweet: Social Bias in Names MistranslationabstractWe ask the question: Are there widespread disparities in machine translations of names across race/ethnicity, and gender?We hypothesize that the translation quality of names and surrounding context will be lower for names associated with US racial and ethnic minorities due to these systems' tendencies to standardize language to predominant language patterns.We develop a dataset of names that are strongly demographically aligned and propose a translation evaluation procedure based on round-trip translation.We analyze the effect of name demographics on translation quality using generalized linear mixed effects models and find that the ability of translation systems to correctly translate female-associated names is significantly lower than male-associated names.This effect is particularly pronounced for femaleassociated names that are also associated with racial (Black) and ethnic (Hispanic) minorities.This disparity in translation quality between social groups for something as personal as someone's name has significant implications for people's professional, personal and cultural identities, self-worth and ease of communication.Our findings suggest that more MT research is needed to improve the translation of names and to provide high-quality service for users regardless of gender, race, and ethnicity. Sandra Sandoval, Jieyu Zhao 0001, Marine Carpuat, Hal Daumé III |
EMNLP | 2 |
| 2023 | TACO: Temporal Latent Action-Driven Contrastive Loss for Visual Reinforcement Learning
Ruijie Zheng, Yanchao Sun, Jieyu Zhao 0001, Huazhe Xu, Hal Daumé III, Furong Huang |
NeurIPS | 5 |
| 2021 | Double Perturbation: On the Robustness of Robustness and Counterfactual Bias EvaluationabstractChong Zhang, Jieyu Zhao, Huan Zhang, Kai-Wei Chang, Cho-Jui Hsieh. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Jieyu Zhao 0001, Huan Zhang 0001, Kai-Wei Chang 0001, Cho-Jui Hsieh |
NAACL-HLT | 2 |
| 2020 | Towards Understanding Gender Bias in Relation ExtractionabstractAndrew Gaut, Tony Sun, Shirlyn Tang, Yuxin Huang, Jing Qian, Mai ElSherief, Jieyu Zhao, Diba Mirza, Elizabeth Belding, Kai-Wei Chang, William Yang Wang. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. Andrew Gaut, Tony Sun, Shirlyn Tang, Mai ElSherief, Jieyu Zhao 0001, Diba Mirza, Elizabeth M. Belding, Kai-Wei Chang 0001, William Yang Wang |
ACL | 7 |
| 2020 | Mitigating Gender Bias Amplification in Distribution by Posterior RegularizationabstractAdvanced machine learning techniques have boosted the performance of natural language processing.Nevertheless, recent studies, e.g., Zhao et al. (2017) show that these techniques inadvertently capture the societal bias hidden in the corpus and further amplify it.However, their analysis is conducted only on models' top predictions.In this paper, we investigate the gender bias amplification issue from the distribution perspective and demonstrate that the bias is amplified in the view of predicted probability distribution over labels.We further propose a bias mitigation approach based on posterior regularization.With little performance loss, our method can almost remove the bias amplification in the distribution.Our study sheds the light on understanding the bias amplification.* Both authors contributed equally to this work and are listed in alphabetical order. Shengyu Jia, Jieyu Zhao 0001, Kai-Wei Chang 0001 |
ACL | 3 |
| 2020 | Gender Bias in Multilingual Embeddings and Cross-Lingual TransferabstractMultilingual representations embed words from many languages into a single semantic space such that words with similar meanings are close to each other regardless of the language.These embeddings have been widely used in various settings, such as cross-lingual transfer, where a natural language processing (NLP) model trained on one language is deployed to another language.While the crosslingual transfer techniques are powerful, they carry gender bias from the source to target languages.In this paper, we study gender bias in multilingual embeddings and how it affects transfer learning for NLP applications.We create a multilingual dataset for bias analysis and propose several ways for quantifying bias in multilingual representations from both the intrinsic and extrinsic perspectives.Experimental results show that the magnitude of bias in the multilingual representations changes differently when we align the embeddings to different target spaces and that the alignment direction can also have an influence on the bias in transfer learning.We further provide recommendations for using the multilingual word representations for downstream tasks. Jieyu Zhao 0001, Subhabrata Mukherjee, Saghar Hosseini, Kai-Wei Chang 0001, Ahmed Awadallah 0001 |
ACL | 1 |
| 2020 | "The Boating Store Had Its Best Sail Ever": Pronunciation-attentive Contextualized Pun RecognitionabstractHumor plays an important role in human languages and it is essential to model humor when building intelligence systems. Among different forms of humor, puns perform wordplay for humorous effects by employing words with double entendre and high phonetic similarity. However, identifying and modeling puns are challenging as puns usually involved implicit semantic or phonological tricks. In this paper, we propose Pronunciation-attentive Contextualized Pun Recognition (PCPR) to perceive human humor, detect if a sentence contains puns and locate them in the sentence. PCPR derives contextualized representation for each word in a sentence by capturing the association between the surrounding context and its corresponding phonetic symbols. Extensive experiments are conducted on two benchmark datasets. Results demonstrate that the proposed approach significantly outperforms the state-of-the-art methods in pun detection and location tasks. In-depth analyses verify the effectiveness and robustness of PCPR. Yichao Zhou 0001, Jyun-Yu Jiang, Jieyu Zhao 0001, Kai-Wei Chang 0001, Wei Wang 0010 |
ACL | 3 |
| 2020 | LOGAN: Local Group Bias Detection by ClusteringabstractMachine learning techniques have been widely used in natural language processing (NLP). However, as revealed by many recent studies, machine learning models often inherit and amplify the societal biases in data. Various metrics have been proposed to quantify biases in model predictions. In particular, several of them evaluate disparity in model performance between protected groups and advantaged groups in the test corpus. However, we argue that evaluating bias at the corpus level is not enough for understanding how biases are embedded in a model. In fact, a model with similar aggregated performance between different groups on the entire data may behave differently on instances in a local region. To analyze and detect such local bias, we propose LOGAN, a new bias detection technique based on clustering. Experiments on toxicity classification and object classification tasks show that LOGAN identifies bias in a local region and allows us to better analyze the biases in model predictions. Jieyu Zhao 0001, Kai-Wei Chang 0001 |
EMNLP (1) | 1 |
| 2019 | Mitigating Gender Bias in Natural Language Processing: Literature ReviewabstractTony Sun, Andrew Gaut, Shirlyn Tang, Yuxin Huang, Mai ElSherief, Jieyu Zhao, Diba Mirza, Elizabeth Belding, Kai-Wei Chang, William Yang Wang. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019. Tony Sun, Andrew Gaut, Shirlyn Tang, Mai ElSherief, Jieyu Zhao 0001, Diba Mirza, Elizabeth M. Belding, Kai-Wei Chang 0001, William Yang Wang |
ACL (1) | 6 |
| 2019 | Examining Gender Bias in Languages with Grammatical GenderabstractPei Zhou, Weijia Shi, Jieyu Zhao, Kuan-Hao Huang, Muhao Chen, Ryan Cotterell, Kai-Wei Chang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Jieyu Zhao 0001, Kuan-Hao Huang, Muhao Chen 0001, Ryan Cotterell, Kai-Wei Chang 0001 |
EMNLP/IJCNLP (1) | 3 |
| 2019 | Balanced Datasets Are Not Enough: Estimating and Mitigating Gender Bias in Deep Image RepresentationsabstractIn this work, we present a framework to measure and mitigate intrinsic biases with respect to protected variables -such as gender- in visual recognition tasks. We show that trained models significantly amplify the association of target labels with gender beyond what one would expect from biased datasets. Surprisingly, we show that even when datasets are balanced such that each label co-occurs equally with each gender, learned models amplify the association between labels and gender, as much as if data had not been balanced! To mitigate this, we adopt an adversarial approach to remove unwanted features corresponding to protected variables from intermediate representations in a deep neural network - and provide a detailed analysis of its effectiveness. Experiments on two datasets: the COCO dataset (objects), and the imSitu dataset (actions), show reductions in gender bias amplification while maintaining most of the accuracy of the original models. Jieyu Zhao 0001, Mark Yatskar, Kai-Wei Chang 0001, Vicente Ordonez |
ICCV | 2 |
| 2018 | Learning Gender-Neutral Word EmbeddingsabstractWord embedding models have become a fundamental component in a wide range of Natural Language Processing (NLP) applications.However, embeddings trained on human-generated corpora have been demonstrated to inherit strong gender stereotypes that reflect social constructs.To address this concern, in this paper, we propose a novel training procedure for learning gender-neutral word embeddings.Our approach aims to preserve gender information in certain dimensions of word vectors while compelling other dimensions to be free of gender influence.Based on the proposed method, we generate a Gender-Neutral variant of GloVe (GN-GloVe).Quantitative and qualitative experiments demonstrate that GN-GloVe successfully isolates gender information without sacrificing the functionality of the embedding model. Jieyu Zhao 0001, Yichao Zhou 0001, Zeyu Li 0001, Wei Wang 0010, Kai-Wei Chang 0001 |
EMNLP | 1 |
| 2017 | Men Also Like Shopping: Reducing Gender Bias Amplification using Corpus-level ConstraintsabstractLanguage is increasingly being used to define rich visual recognition problems with supporting image collections sourced from the web.Structured prediction models are used in these tasks to take advantage of correlations between co-occurring labels and visual input but risk inadvertently encoding social biases found in web corpora.In this work, we study data and models associated with multilabel object classification and visual semantic role labeling.We find that (a) datasets for these tasks contain significant gender bias and (b) models trained on these datasets further amplify existing bias.For example, the activity cooking is over 33% more likely to involve females than males in a training set, and a trained model further amplifies the disparity to 68% at test time.We propose to inject corpus-level constraints for calibrating existing structured prediction models and design an algorithm based on Lagrangian relaxation for collective inference.Our method results in almost no performance loss for the underlying recognition task but decreases the magnitude of bias amplification by 47.5% and 40.5% for multilabel classification and visual semantic role labeling, respectively. Jieyu Zhao 0001, Mark Yatskar, Vicente Ordonez, Kai-Wei Chang 0001 |
EMNLP | 1 |