Angelina Wang

dblp:210/1014 · DBLP profile ↗
← Back
11ranked-venue papers
6as first author
10since 2021 · last 2026
0000-0001-9140-3523ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 6 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Trustworthy machine learning · 74% Language models and text generation · 8% Image recognition and object detection · 7%
Human-computer interaction and pervasive computing
1 paper
Human-AI interaction · 100%

Topics — the 14 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
fairness
4.982025
Fairness through Difference Awareness: Measuring Desired Group Discrimination in LLMs · ACL (1) 2025
Overwriting Pretrained Bias with Finetuning Data · ICCV 2023
Gender Artifacts in Visual Datasets · ICCV 2023
Machine learning › Trustworthy machine learning › fairness
bias mitigation
1.432023
Gender Artifacts in Visual Datasets · ICCV 2023
REVISE: A Tool for Measuring and Mitigating Bias in Visual Datasets · Int. J. Comput. Vis. 2022
Taxonomizing and Measuring Representational Harms: A Look at Image Tagging · AAAI 2023
Machine learning › Trustworthy machine learning › fairness › fairness evaluation
fairness benchmarking
0.912025
Fairness through Difference Awareness: Measuring Desired Group Discrimination in LLMs · ACL (1) 2025
Natural language and speech › Language models and text generation
large language model evaluation
0.912025
Fairness through Difference Awareness: Measuring Desired Group Discrimination in LLMs · ACL (1) 2025
Machine learning › Trustworthy machine learning › dataset bias
visual dataset bias
0.722022
REVISE: A Tool for Measuring and Mitigating Bias in Visual Datasets · Int. J. Comput. Vis. 2022
REVISE: A Tool for Measuring and Mitigating Bias in Visual Datasets · ECCV (3) 2020
Machine learning › Transfer learning and domain adaptation
fine-tuning
0.712023
Overwriting Pretrained Bias with Finetuning Data · ICCV 2023
Computer vision › Image recognition and object detection
image annotation
0.712023
Taxonomizing and Measuring Representational Harms: A Look at Image Tagging · AAAI 2023
Machine learning › Trustworthy machine learning › fairness › algorithmic bias
bias amplification
0.512021
Directional Bias Amplification · ICML 2021
Computer vision › Vision and language
image captioning
0.512021
Understanding and Evaluating Racial Biases in Image Captioning · ICCV 2021
Computing education › broadening participation in computing
culturally responsive computing
0.312026
Whose Knowledge Counts? Co-Designing Community-Centered AI Auditing Tools with Educators in Hawai'i · CHI 2026
Human-AI interaction › generative AI
generative AI in education
0.312026
Whose Knowledge Counts? Co-Designing Community-Centered AI Auditing Tools with Educators in Hawai'i · CHI 2026
Computer vision › Image recognition and object detection
object recognition
0.212023
Gender Artifacts in Visual Datasets · ICCV 2023
Machine learning › Trustworthy machine learning › robustness
spurious correlation
0.212023
Overwriting Pretrained Bias with Finetuning Data · ICCV 2023
Computer vision › Face, body and person analysis › facial attribute analysis
demographic estimation
0.112021
Understanding and Evaluating Racial Biases in Image Captioning · ICCV 2021

Methods — techniques the papers use, named apart from their topics

content analysis · 2.0co-design workshops · 2.0benchmark suite construction · 0.9interpretability analysis · 0.7image classifier · 0.7finetuning data curation · 0.7fairness measurement taxonomy · 0.7dataset auditing · 0.6sentiment analysis · 0.5sensitive attribute prediction · 0.5manual annotation · 0.5confidence intervals · 0.5
YearPublicationVenuePosition
2026 Whose Knowledge Counts? Co-Designing Community-Centered AI Auditing Tools with Educators in Hawai'i
abstract
Although generative AI is being deployed into classrooms with promises of aiding teachers, educators caution that these tools can have unintended pedagogical repercussions, including cultural misrepresentation and bias. These concerns are heightened in low-resource language and Indigenous education settings, where AI systems frequently underperform. We investigate these challenges in Hawai‘i, where public schools operate under a statewide mandate to integrate Hawaiian language and culture into education. Through four co-design workshops with 22 public school educators, we surfaced concerns about using generative AI in educational settings, particularly around cultural misrepresentation, and corresponding designs for auditing tools that address these issues. We find that educators envision tools grounded in specific Hawaiian cultural values and practices, such as tracing the genealogy of knowledge in source materials. Building on these insights, we conceptualize AI auditing as a community-oriented process rather than the work of isolated individuals, and discuss implications for designing auditing tools.
Dora Zhao, Hannah Cha, Michael J. Ryan, Angelina Wang, Rachel Baker-Ramos, Evyn-Bree Helekahi-Kaiwi, Rebecca Diego, Josiah D. Hester, Diyi Yang
CHI4
2025 Fairness through Difference Awareness: Measuring Desired Group Discrimination in LLMs
abstract
Algorithmic fairness has conventionally adopted the mathematically convenient perspective of racial color-blindness (i.e., difference unaware treatment).However, we contend that in a range of important settings, group difference awareness matters.For example, differentiating between groups may be necessary in legal contexts (e.g., the U.S. compulsory draft applies to men but not women) and harm assessments (e.g., referring to girls as "terrorists" may be less harmful than referring to Muslim people as such).Thus, in contrast to most fairness work, we study fairness through the perspective of treating people differently -when it is contextually appropriate to.We first introduce an important distinction between descriptive (fact-based), normative (value-based), and correlation (association-based) benchmarks.This distinction is significant because each category requires separate interpretation and mitigation tailored to its specific characteristics.Then, we present a benchmark suite composed of eight different scenarios for a total of 16k questions that enables us to assess difference awareness.Finally, we show results across ten models that demonstrate difference awareness is a distinct dimension to fairness where existing bias mitigation strategies may backfire.
Angelina Wang, Michelle Phan, Daniel E. Ho, Oluwasanmi Koyejo
ACL (1)1
2025 Rigor in AI: Doing Rigorous AI Work Requires a Broader, Responsible AI-Informed Conception of Rigor
abstract
In AI research and practice, rigor remains largely understood in terms of methodological rigor---such as whether mathematical, statistical, or computational methods are correctly applied. We argue that this narrow conception of rigor has contributed to the concerns raised by the responsible AI community, including overblown claims about the capabilities of AI systems. Our position is that a broader conception of what rigorous AI research and practice should entail is needed. We believe such a conception---in addition to a more expansive understanding of 1) methodological rigor---should include aspects related to 2) what background knowledge informs what to work on (epistemic rigor); 3) how disciplinary, community, or personal norms, standards, or beliefs influence the work (normative rigor); 4) how clearly articulated the theoretical constructs under use are (conceptual rigor); 5) what is reported and how (reporting rigor); and 6) how well-supported the inferences from existing evidence are (interpretative rigor). In doing so, we also provide useful language and a framework for much needed dialogue about the AI community's work by researchers, policymakers, journalists, and other stakeholders.
Alexandra Olteanu, Su Lin Blodgett, Agathe Balayn, Angelina Wang, Fernando Diaz 0001, Flávio P. Calmon, Margaret Mitchell, Michael D. Ekstrand, Reuben Binns, Solon Barocas
NeurIPS4
2024 Strategies for Increasing Corporate Responsible AI Prioritization
abstract
Responsible artificial intelligence (RAI) is increasingly recognized as a critical concern. However, the level of corporate RAI prioritization has not kept pace. In this work, we conduct 16 semi-structured interviews with practitioners to investigate what has historically motivated companies to increase the prioritization of RAI. What emerges is a complex story of conflicting and varied factors, but we bring structure to the narrative by highlighting the different strategies available to employ, and point to the actors with access to each. While there are no guaranteed steps for increasing RAI prioritization, we paint the current landscape of motivators so that practitioners can learn from each other, and put forth our own selection of promising directions forward.
Angelina Wang, Teresa Datta, John Dickerson 0001
AIES (1)1
2023 Taxonomizing and Measuring Representational Harms: A Look at Image Tagging
abstract
In this paper, we examine computational approaches for measuring the "fairness" of image tagging systems, finding that they cluster into five distinct categories, each with its own analytic foundation. We also identify a range of normative concerns that are often collapsed under the terms "unfairness," "bias," or even "discrimination" when discussing problematic cases of image tagging. Specifically, we identify four types of representational harms that can be caused by image tagging systems, providing concrete examples of each. We then consider how different computational measurement approaches map to each of these types, demonstrating that there is not a one-to-one mapping. Our findings emphasize that no single measurement approach will be definitive and that it is not possible to infer from the use of a particular measurement approach which type of harm was intended to be measured. Lastly, equipped with this more granular understanding of the types of representational harms that can be caused by image tagging systems, we show that attempts to mitigate some of these types of harms may be in tension with one another.
Jared Katzman, Angelina Wang, Morgan Klaus Scheuerman, Su Lin Blodgett, Kristen Laird, Hanna M. Wallach, Solon Barocas
AAAI2
2023 Gender Artifacts in Visual Datasets
abstract
Gender biases are known to exist within large-scale visual datasets and can be reflected or even amplified in downstream models. Many prior works have proposed methods for mitigating gender biases, often by attempting to remove gender expression information from images. To understand the feasibility and practicality of these approaches, we investigate what "gender artifacts" exist in large-scale visual datasets. We define a "gender artifact" as a visual cue correlated with gender, focusing specifically on cues that are learnable by a modern image classifier and have an interpretable human corollary. Through our analyses, we find that gender artifacts are ubiquitous in the COCO and OpenImages datasets, occurring everywhere from low-level information (e.g., the mean value of the color channels) to higher-level image composition (e.g., pose and location of people). Further, bias mitigation methods that attempt to remove gender actually remove more information from the scene than the person. Given the prevalence of gender artifacts, we claim that attempts to remove these artifacts from such datasets are largely infeasible as certain removed artifacts may be necessary for the downstream task of object recognition. Instead, the responsibility lies with researchers and practitioners to be aware that the distribution of images within datasets is highly gendered and hence develop fairness-aware methods which are robust to these distributional shifts across groups.
Nicole Meister, Dora Zhao, Angelina Wang, Vikram V. Ramaswamy, Ruth Fong, Olga Russakovsky
ICCV3
2023 Overwriting Pretrained Bias with Finetuning Data
abstract
Transfer learning is beneficial by allowing the expressive features of models pretrained on large-scale datasets to be finetuned for the target task of smaller, more domain-specific datasets. However, there is a concern that these pretrained models may come with their own biases which would propagate into the finetuned model. In this work, we investigate bias when conceptualized as both spurious correlations between the target task and a sensitive attribute as well as underrepresentation of a particular group in the dataset. Under both notions of bias, we find that (1) models finetuned on top of pretrained models can indeed inherit their biases, but (2) this bias can be corrected for through relatively minor interventions to the finetuning dataset, and often with a negligible impact to performance. Our findings imply that careful curation of the finetuning dataset is important for reducing biases on a downstream task, and doing so can even compensate for bias in the pretrained model.
Angelina Wang, Olga Russakovsky
ICCV1
2022 REVISE: A Tool for Measuring and Mitigating Bias in Visual Datasets
Angelina Wang, Ryan Zhang, Anat Kleiman, Leslie Kim, Dora Zhao, Iroha Shirai, Arvind Narayanan, Olga Russakovsky
Int. J. Comput. Vis.1
2021 Understanding and Evaluating Racial Biases in Image Captioning
abstract
Image captioning is an important task for benchmarking visual reasoning and for enabling accessibility for people with vision impairments. However, as in many machine learning settings, social biases can influence image captioning in undesirable ways. In this work, we study bias propagation pathways within image captioning, focusing specifically on the COCO dataset. Prior work has analyzed gender bias in captions using automatically-derived gender labels; here we examine racial and intersectional biases using manual annotations. Our first contribution is in annotating the perceived gender and skin color of 28,315 of the depicted people after obtaining IRB approval. Using these annotations, we compare racial biases present in both manual and automatically-generated image captions. We demonstrate differences in caption performance, sentiment, and word choice between images of lighter versus darker-skinned people. Further, we find the magnitude of these differences to be greater in modern captioning systems compared to older ones, thus leading to concerns that without proper consideration and mitigation these differences will only become increasingly prevalent. Code and data is available at https://princetonvisualai.github.io/imagecaptioning-bias/.
Dora Zhao, Angelina Wang, Olga Russakovsky
ICCV2
2021 Directional Bias Amplification
abstract
Mitigating bias in machine learning systems requires refining our understanding of bias propagation pathways: from societal structures to large-scale data to trained models to impact on society. In this work, we focus on one aspect of the problem, namely bias amplification: the tendency of models to amplify the biases present in the data they are trained on. A metric for measuring bias amplification was introduced in the seminal work by Zhao et al. (2017); however, as we demonstrate, this metric suffers from a number of shortcomings including conflating different types of bias amplification and failing to account for varying base rates of protected attributes. We introduce and analyze a new, decoupled metric for measuring bias amplification, $BiasAmp_{\rightarrow}$ (Directional Bias Amplification). We thoroughly analyze and discuss both the technical assumptions and normative implications of this metric. We provide suggestions about its measurement by cautioning against predicting sensitive attributes, encouraging the use of confidence intervals due to fluctuations in the fairness of models across runs, and discussing the limitations of what this metric captures. Throughout this paper, we work to provide an interrogative look at the technical measurement of bias amplification, guided by our normative ideas of what we want it to encompass. Code is located at https://github.com/princetonvisualai/directional-bias-amp.
Angelina Wang, Olga Russakovsky
ICML1
2020 REVISE: A Tool for Measuring and Mitigating Bias in Visual Datasets
Angelina Wang, Arvind Narayanan, Olga Russakovsky
ECCV (3)1