VLDB 2026 Research / reviewers in the wild / expert
Angelina Wang
dblp:210/1014
· DBLP profile ↗
11ranked-venue papers
6as first author
10since 2021 · last 2026
0000-0001-9140-3523ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 6 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Trustworthy machine learning · 74% Language models and text generation · 8% Image recognition and object detection · 7% | |
| Human-computer interaction and pervasive computing
1 paper |
Human-AI interaction · 100% |
Topics — the 14 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
fairness |
4.9 | 8 | 2025 | Fairness through Difference Awareness: Measuring Desired Group Discrimination in LLMs · ACL (1) 2025 Overwriting Pretrained Bias with Finetuning Data · ICCV 2023 Gender Artifacts in Visual Datasets · ICCV 2023 |
Machine learning › Trustworthy machine learning › fairness
bias mitigation |
1.4 | 3 | 2023 | Gender Artifacts in Visual Datasets · ICCV 2023 REVISE: A Tool for Measuring and Mitigating Bias in Visual Datasets · Int. J. Comput. Vis. 2022 Taxonomizing and Measuring Representational Harms: A Look at Image Tagging · AAAI 2023 |
Machine learning › Trustworthy machine learning › fairness › fairness evaluation
fairness benchmarking |
0.9 | 1 | 2025 | Fairness through Difference Awareness: Measuring Desired Group Discrimination in LLMs · ACL (1) 2025 |
Natural language and speech › Language models and text generation
large language model evaluation |
0.9 | 1 | 2025 | Fairness through Difference Awareness: Measuring Desired Group Discrimination in LLMs · ACL (1) 2025 |
Machine learning › Trustworthy machine learning › dataset bias
visual dataset bias |
0.7 | 2 | 2022 | REVISE: A Tool for Measuring and Mitigating Bias in Visual Datasets · Int. J. Comput. Vis. 2022 REVISE: A Tool for Measuring and Mitigating Bias in Visual Datasets · ECCV (3) 2020 |
Machine learning › Transfer learning and domain adaptation
fine-tuning |
0.7 | 1 | 2023 | Overwriting Pretrained Bias with Finetuning Data · ICCV 2023 |
Computer vision › Image recognition and object detection
image annotation |
0.7 | 1 | 2023 | Taxonomizing and Measuring Representational Harms: A Look at Image Tagging · AAAI 2023 |
Machine learning › Trustworthy machine learning › fairness › algorithmic bias
bias amplification |
0.5 | 1 | 2021 | Directional Bias Amplification · ICML 2021 |
Computer vision › Vision and language
image captioning |
0.5 | 1 | 2021 | Understanding and Evaluating Racial Biases in Image Captioning · ICCV 2021 |
Computing education › broadening participation in computing
culturally responsive computing |
0.3 | 1 | 2026 | Whose Knowledge Counts? Co-Designing Community-Centered AI Auditing Tools with Educators in Hawai'i · CHI 2026 |
Human-AI interaction › generative AI
generative AI in education |
0.3 | 1 | 2026 | Whose Knowledge Counts? Co-Designing Community-Centered AI Auditing Tools with Educators in Hawai'i · CHI 2026 |
Computer vision › Image recognition and object detection
object recognition |
0.2 | 1 | 2023 | Gender Artifacts in Visual Datasets · ICCV 2023 |
Machine learning › Trustworthy machine learning › robustness
spurious correlation |
0.2 | 1 | 2023 | Overwriting Pretrained Bias with Finetuning Data · ICCV 2023 |
Computer vision › Face, body and person analysis › facial attribute analysis
demographic estimation |
0.1 | 1 | 2021 | Understanding and Evaluating Racial Biases in Image Captioning · ICCV 2021 |
Methods — techniques the papers use, named apart from their topics
content analysis · 2.0co-design workshops · 2.0benchmark suite construction · 0.9interpretability analysis · 0.7image classifier · 0.7finetuning data curation · 0.7fairness measurement taxonomy · 0.7dataset auditing · 0.6sentiment analysis · 0.5sensitive attribute prediction · 0.5manual annotation · 0.5confidence intervals · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Whose Knowledge Counts? Co-Designing Community-Centered AI Auditing Tools with Educators in Hawai'iabstractAlthough generative AI is being deployed into classrooms with promises of aiding teachers, educators caution that these tools can have unintended pedagogical repercussions, including cultural misrepresentation and bias. These concerns are heightened in low-resource language and Indigenous education settings, where AI systems frequently underperform. We investigate these challenges in Hawai‘i, where public schools operate under a statewide mandate to integrate Hawaiian language and culture into education. Through four co-design workshops with 22 public school educators, we surfaced concerns about using generative AI in educational settings, particularly around cultural misrepresentation, and corresponding designs for auditing tools that address these issues. We find that educators envision tools grounded in specific Hawaiian cultural values and practices, such as tracing the genealogy of knowledge in source materials. Building on these insights, we conceptualize AI auditing as a community-oriented process rather than the work of isolated individuals, and discuss implications for designing auditing tools. Dora Zhao, Hannah Cha, Michael J. Ryan, Angelina Wang, Rachel Baker-Ramos, Evyn-Bree Helekahi-Kaiwi, Rebecca Diego, Josiah D. Hester, Diyi Yang |
CHI | 4 |
| 2025 | Fairness through Difference Awareness: Measuring Desired Group Discrimination in LLMsabstractAlgorithmic fairness has conventionally adopted the mathematically convenient perspective of racial color-blindness (i.e., difference unaware treatment).However, we contend that in a range of important settings, group difference awareness matters.For example, differentiating between groups may be necessary in legal contexts (e.g., the U.S. compulsory draft applies to men but not women) and harm assessments (e.g., referring to girls as "terrorists" may be less harmful than referring to Muslim people as such).Thus, in contrast to most fairness work, we study fairness through the perspective of treating people differently -when it is contextually appropriate to.We first introduce an important distinction between descriptive (fact-based), normative (value-based), and correlation (association-based) benchmarks.This distinction is significant because each category requires separate interpretation and mitigation tailored to its specific characteristics.Then, we present a benchmark suite composed of eight different scenarios for a total of 16k questions that enables us to assess difference awareness.Finally, we show results across ten models that demonstrate difference awareness is a distinct dimension to fairness where existing bias mitigation strategies may backfire. Angelina Wang, Michelle Phan, Daniel E. Ho, Oluwasanmi Koyejo |
ACL (1) | 1 |
| 2025 | Rigor in AI: Doing Rigorous AI Work Requires a Broader, Responsible AI-Informed Conception of RigorabstractIn AI research and practice, rigor remains largely understood in terms of methodological rigor---such as whether mathematical, statistical, or computational methods are correctly applied. We argue that this narrow conception of rigor has contributed to the concerns raised by the responsible AI community, including overblown claims about the capabilities of AI systems. Our position is that a broader conception of what rigorous AI research and practice should entail is needed. We believe such a conception---in addition to a more expansive understanding of 1) methodological rigor---should include aspects related to 2) what background knowledge informs what to work on (epistemic rigor); 3) how disciplinary, community, or personal norms, standards, or beliefs influence the work (normative rigor); 4) how clearly articulated the theoretical constructs under use are (conceptual rigor); 5) what is reported and how (reporting rigor); and 6) how well-supported the inferences from existing evidence are (interpretative rigor). In doing so, we also provide useful language and a framework for much needed dialogue about the AI community's work by researchers, policymakers, journalists, and other stakeholders. Alexandra Olteanu, Su Lin Blodgett, Agathe Balayn, Angelina Wang, Fernando Diaz 0001, Flávio P. Calmon, Margaret Mitchell, Michael D. Ekstrand, Reuben Binns, Solon Barocas |
NeurIPS | 4 |
| 2024 | Strategies for Increasing Corporate Responsible AI PrioritizationabstractResponsible artificial intelligence (RAI) is increasingly recognized as a critical concern. However, the level of corporate RAI prioritization has not kept pace. In this work, we conduct 16 semi-structured interviews with practitioners to investigate what has historically motivated companies to increase the prioritization of RAI. What emerges is a complex story of conflicting and varied factors, but we bring structure to the narrative by highlighting the different strategies available to employ, and point to the actors with access to each. While there are no guaranteed steps for increasing RAI prioritization, we paint the current landscape of motivators so that practitioners can learn from each other, and put forth our own selection of promising directions forward. Angelina Wang, Teresa Datta, John Dickerson 0001 |
AIES (1) | 1 |
| 2023 | Taxonomizing and Measuring Representational Harms: A Look at Image TaggingabstractIn this paper, we examine computational approaches for measuring the "fairness" of image tagging systems, finding that they cluster into five distinct categories, each with its own analytic foundation. We also identify a range of normative concerns that are often collapsed under the terms "unfairness," "bias," or even "discrimination" when discussing problematic cases of image tagging. Specifically, we identify four types of representational harms that can be caused by image tagging systems, providing concrete examples of each. We then consider how different computational measurement approaches map to each of these types, demonstrating that there is not a one-to-one mapping. Our findings emphasize that no single measurement approach will be definitive and that it is not possible to infer from the use of a particular measurement approach which type of harm was intended to be measured. Lastly, equipped with this more granular understanding of the types of representational harms that can be caused by image tagging systems, we show that attempts to mitigate some of these types of harms may be in tension with one another. Jared Katzman, Angelina Wang, Morgan Klaus Scheuerman, Su Lin Blodgett, Kristen Laird, Hanna M. Wallach, Solon Barocas |
AAAI | 2 |
| 2023 | Gender Artifacts in Visual DatasetsabstractGender biases are known to exist within large-scale visual datasets and can be reflected or even amplified in downstream models. Many prior works have proposed methods for mitigating gender biases, often by attempting to remove gender expression information from images. To understand the feasibility and practicality of these approaches, we investigate what "gender artifacts" exist in large-scale visual datasets. We define a "gender artifact" as a visual cue correlated with gender, focusing specifically on cues that are learnable by a modern image classifier and have an interpretable human corollary. Through our analyses, we find that gender artifacts are ubiquitous in the COCO and OpenImages datasets, occurring everywhere from low-level information (e.g., the mean value of the color channels) to higher-level image composition (e.g., pose and location of people). Further, bias mitigation methods that attempt to remove gender actually remove more information from the scene than the person. Given the prevalence of gender artifacts, we claim that attempts to remove these artifacts from such datasets are largely infeasible as certain removed artifacts may be necessary for the downstream task of object recognition. Instead, the responsibility lies with researchers and practitioners to be aware that the distribution of images within datasets is highly gendered and hence develop fairness-aware methods which are robust to these distributional shifts across groups. Nicole Meister, Dora Zhao, Angelina Wang, Vikram V. Ramaswamy, Ruth Fong, Olga Russakovsky |
ICCV | 3 |
| 2023 | Overwriting Pretrained Bias with Finetuning DataabstractTransfer learning is beneficial by allowing the expressive features of models pretrained on large-scale datasets to be finetuned for the target task of smaller, more domain-specific datasets. However, there is a concern that these pretrained models may come with their own biases which would propagate into the finetuned model. In this work, we investigate bias when conceptualized as both spurious correlations between the target task and a sensitive attribute as well as underrepresentation of a particular group in the dataset. Under both notions of bias, we find that (1) models finetuned on top of pretrained models can indeed inherit their biases, but (2) this bias can be corrected for through relatively minor interventions to the finetuning dataset, and often with a negligible impact to performance. Our findings imply that careful curation of the finetuning dataset is important for reducing biases on a downstream task, and doing so can even compensate for bias in the pretrained model. Angelina Wang, Olga Russakovsky |
ICCV | 1 |
| 2022 | REVISE: A Tool for Measuring and Mitigating Bias in Visual Datasets
Angelina Wang, Ryan Zhang, Anat Kleiman, Leslie Kim, Dora Zhao, Iroha Shirai, Arvind Narayanan, Olga Russakovsky |
Int. J. Comput. Vis. | 1 |
| 2021 | Understanding and Evaluating Racial Biases in Image CaptioningabstractImage captioning is an important task for benchmarking visual reasoning and for enabling accessibility for people with vision impairments. However, as in many machine learning settings, social biases can influence image captioning in undesirable ways. In this work, we study bias propagation pathways within image captioning, focusing specifically on the COCO dataset. Prior work has analyzed gender bias in captions using automatically-derived gender labels; here we examine racial and intersectional biases using manual annotations. Our first contribution is in annotating the perceived gender and skin color of 28,315 of the depicted people after obtaining IRB approval. Using these annotations, we compare racial biases present in both manual and automatically-generated image captions. We demonstrate differences in caption performance, sentiment, and word choice between images of lighter versus darker-skinned people. Further, we find the magnitude of these differences to be greater in modern captioning systems compared to older ones, thus leading to concerns that without proper consideration and mitigation these differences will only become increasingly prevalent. Code and data is available at https://princetonvisualai.github.io/imagecaptioning-bias/. Dora Zhao, Angelina Wang, Olga Russakovsky |
ICCV | 2 |
| 2021 | Directional Bias AmplificationabstractMitigating bias in machine learning systems requires refining our understanding of bias propagation pathways: from societal structures to large-scale data to trained models to impact on society. In this work, we focus on one aspect of the problem, namely bias amplification: the tendency of models to amplify the biases present in the data they are trained on. A metric for measuring bias amplification was introduced in the seminal work by Zhao et al. (2017); however, as we demonstrate, this metric suffers from a number of shortcomings including conflating different types of bias amplification and failing to account for varying base rates of protected attributes. We introduce and analyze a new, decoupled metric for measuring bias amplification, $BiasAmp_{\rightarrow}$ (Directional Bias Amplification). We thoroughly analyze and discuss both the technical assumptions and normative implications of this metric. We provide suggestions about its measurement by cautioning against predicting sensitive attributes, encouraging the use of confidence intervals due to fluctuations in the fairness of models across runs, and discussing the limitations of what this metric captures. Throughout this paper, we work to provide an interrogative look at the technical measurement of bias amplification, guided by our normative ideas of what we want it to encompass. Code is located at https://github.com/princetonvisualai/directional-bias-amp. Angelina Wang, Olga Russakovsky |
ICML | 1 |
| 2020 | REVISE: A Tool for Measuring and Mitigating Bias in Visual Datasets
Angelina Wang, Arvind Narayanan, Olga Russakovsky |
ECCV (3) | 1 |