Robert Wolfe

dblp:50/689 · DBLP profile ↗
← Back
26ranked-venue papers
8as first author
26since 2021 · last 2026
0000-0001-7133-695XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 21 · 5 first-author · 21 since 2021Artificial intelligence and machine learning · 12 · 8 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Toys that listen, talk, and play: Understanding Children's Sensemaking and Interactions with AI Toys
abstract
Generative AI (genAI) is increasingly being integrated into children’s everyday lives, not only through screens but also through so-called “screen-free” AI toys. These toys can simulate emotions, personalize responses, and recall prior interactions, creating the illusion of an ongoing social connection. Such capabilities raise important questions about how children understand boundaries, agency, and relationships when interacting with AI toys. To investigate this, we conducted two participatory design sessions with eight children ages 6 - 11 where they engaged with three different AI toys, shifting between play, experimentation, and reflection. Our findings reveal that children approached AI toys with genuine curiosity, profiling them as social beings. However, frequent interaction breakdowns and mismatches between apparent intelligence and toy-like form disrupted expectations around play and led to adversarial play. We conclude with implications and design provocations to navigate children’s encounters with AI toys in more transparent, developmentally appropriate, and responsible ways.
Aayushi Dangol, Meghna Gupta, Daeun Yoo, Robert Wolfe, Jason C. Yip 0001, Franziska Roesner, Julie A. Kientz
IDC4
2026 Parent Perspectives on Future Designs of AAC for Children with Speech and Language Difficulties
Aayushi Dangol, Aaleyah Lewis, Hyewon Suh, Robert Wolfe, James Fogarty, Julie A. Kientz
IDC4
2026 Where Does AI Leave a Footprint? Children's Reasoning About AI's Environmental Costs
abstract
Two of the most socially consequential issues facing today’s children are the rise of artificial intelligence (AI) and the rapid changes to the earth’s climate. Both issues are complex and contested, and they are linked through the notable environmental costs of AI use. Using a systems thinking framework, we developed an interactive system called Ecoprompt to help children reason about the environmental impact of AI. EcoPrompt combines a prompt-level environmental footprint calculator with a simulation game that challenges players to reason about the impact of AI use on natural resources that the player manages. We evaluated the system through two participatory design sessions with 16 children ages 6–12. Our findings surfaced children’s perspectives on societal and environmental tradeoffs of AI use, as well as their sense of agency and responsibility. Taken together, these findings suggest opportunities for broadening AI literacy to include systems-level reasoning about AI’s environmental impact.
Aayushi Dangol, Robert Wolfe, Nisha Devasia, Mitsuka Kiyohara, Jason C. Yip 0001, Julie A. Kientz
IDC2
2026 Relief or displacement? How teachers are negotiating generative AI's role in their professional practice
abstract
As generative AI (genAI) rapidly enters classrooms, accompanied by district-level policy rollouts and industry-led teacher trainings, it is important to rethink the canonical “adopt and train” playbook. Decades of educational technology research show that tools promising personalization and access often deepen inequities due to uneven resources, training, and institutional support. Against this backdrop, we conducted semi-structured interviews with 22 teachers from a large U.S. school district that was an early adopter of genAI. Our findings reveal the motivations driving adoption, the factors underlying resistance, and the boundaries teachers negotiate to align genAI use with their values. We further contribute by unpacking the sociotechnical dynamics—including district policies, professional norms, and relational commitments—that shape how teachers navigate the promises and risks of these tools.
Aayushi Dangol, Smriti Kotiyal, Robert Wolfe, Alex J. Bowers, Antonio Vigil, Jason C. Yip 0001, Julie A. Kientz, Suleman Shahid, Tom Yeh, Vincent Cho 0002, Katie Davis 0001
CHI3
2026 SusBench: An Online Benchmark for Evaluating Dark Pattern Susceptibility of Computer-Use Agents
abstract
As LLM-based computer-use agents (CUAs) begin to autonomously interact with real-world interfaces, understanding their vulnerability to manipulative interface designs becomes increasingly critical. We introduce SusBench, an online benchmark for evaluating the susceptibility of CUAs to UI dark patterns, designs that aim to manipulate or deceive users into taking unintentional actions. Drawing nine common dark pattern types from existing taxonomies, we developed a method for constructing believable dark patterns on real-world consumer websites through code injections, and designed 313 evaluation tasks across 55 websites. Our study with 29 participants showed that humans perceived our dark pattern injections to be highly realistic, with the vast majority of participants not noticing that these had been injected by the research team. We evaluated five state-of-the-art CUAs on the benchmark. We found that both human participants and agents are particularly susceptible to the dark patterns of Preselection, Trick Wording, and Hidden Information, while being resilient to other overt dark patterns. Our findings inform the development of more trustworthy CUAs, their use as potential human proxies in evaluating deceptive designs, and the regulation of an online environment increasingly navigated by autonomous agents.
Longjie Guo, Chenjie Yuan, Mingyuan Zhong 0001, Robert Wolfe, Ruican Zhong, Bingbing Wen, Hua Shen 0005, Lucy Lu Wang, Alexis Hiniker
IUI4
2025 If anybody finds out you are in BIG TROUBLE": Understanding Children's Hopes, Fears, and Evaluations of Generative AI
abstract
As generative artificial intelligence (genAI) increasingly mediates how children learn, communicate, and engage with digital content, understanding children's hopes and fears about this emerging technology is crucial.In a pilot study with 37 fifth-graders, we explored how children (ages 9-10) envision genAI and the roles they believe it should play in their daily life.Our findings reveal three key ways children envision genAI: as a companion providing guidance, a collaborator working alongside them, and a task automator that offloads responsibilities.However, alongside these hopeful views, children expressed fears about overreliance, particularly in academic settings, linking it to fears of diminished learning, disciplinary consequences, and long-term failure.This study highlights the need for child-centric AI design that balances these tensions, empowering children with the skills to critically engage with and navigate their evolving relationships with digital technologies.
Aayushi Dangol, Robert Wolfe, Daeun Yoo, Arya Thiruvillakkat, Ben Chickadel, Julie A. Kientz
IDC2
2025 Children's Mental Models of AI Reasoning: Implications for AI Literacy Education
Aayushi Dangol, Robert Wolfe, Runhua Zhao, Jaewon Kim 0002, Trushaa Ramanan, Katie Davis 0001, Julie A. Kientz
IDC2
2025 "AI just keeps guessing": Using ARC Puzzles to Help Children Identify Reasoning Errors in Generative AI
Aayushi Dangol, Runhua Zhao, Robert Wolfe, Trushaa Ramanan, Julie A. Kientz, Jason C. Yip 0001
IDC3
2025 Building the Beloved Community: Designing Technologies for Neighborhood Safety
abstract
Thesis (Ph.D.)--University of Washington, 2024
Ishita Chordia, Robert Wolfe, Jason C. Yip 0001, Alexis Hiniker
CHI2
2025 Fragments to Facts: Partial-Information Fragment Inference from LLMs
abstract
Large language models (LLMs) can leak sensitive training data through memorization and membership inference attacks. Prior work has primarily focused on strong adversarial assumptions, including attacker access to entire samples or long, ordered prefixes, leaving open the question of how vulnerable LLMs are when adversaries have only partial, unordered sample information. For example, if an attacker knows a patient has "hypertension," under what conditions can they query a model fine-tuned on patient data to learn the patient also has "osteoarthritis?" In this paper, we introduce a more general threat model under this weaker assumption and show that fine-tuned LLMs are susceptible to these fragment-specific extraction attacks. To systematically investigate these attacks, we propose two data-blind methods: (1) a likelihood ratio attack inspired by methods from membership inference, and (2) a novel approach, PRISM, which regularizes the ratio by leveraging an external prior. Using examples from medical and legal settings, we show that both methods are competitive with a data-aware baseline classifier that assumes access to labeled in-distribution data, underscoring their robustness.
Lucas Rosenblatt, Bin Han 0011, Robert Wolfe, Bill Howe
ICML3
2025 Trust-Enabled Privacy: Social Media Designs to Support Adolescent User Boundary Regulation
Jaewon Kim 0002, Robert Wolfe, Ramya Bhagirathi Subramanian, Mei-Hsuan Lee, Jessica Colnago, Alexis Hiniker
SOUPS2
2025 Privacy as Social Norm: Systematically Reducing Dysfunctional Privacy Concerns on Social Media
abstract
Through co-design interviews (N=19) and a design evaluation survey (N=136) with U.S. teens ages 13-18, we investigated teens' privacy management on social media. Our study revealed that 28.1% of teens with public accounts and 15.3% with private accounts experience dysfunctional fear, that is, fear that diminishes their quality of life or paralyzes them from taking necessary precautions. These fears fall into three categories: fear of uncontrolled audience reach, fear of online hostility, and fear of personal privacy missteps. While current approaches often emphasize individual vigilance and restrictive measures, our findings show this can paradoxically lead teens to either withdraw from beneficial social interactions or resign themselves to accept privacy violations, viewing them as inevitable. Drawing on teen input, we developed and evaluated ten design prototypes that emphasize empowerment over fear, system-wide explicit emphasis on privacy, clear privacy norms, and flexible controls. Survey results indicate teens perceive these approaches as effectively reducing privacy concerns while preserving social benefits. Our findings suggest that platforms will be more likely to protect teens' privacy and less likely to manufacture unnecessary fear if they include designs that minimize the impact on other users, have low trade-offs with existing features, require minimal user effort, and function independently of community behavior. Such designs include: 1) alerting users about potentially unintentional personal information disclosure and 2) following up on user reports.
Jaewon Kim 0002, Soobin Cho, Robert Wolfe, Jishnu Hari Nair, Alexis Hiniker
Proc. ACM Hum. Comput. Interact.3
2025 Reading AI and Reading the World: Using an Interactive AI System to Promote Children's Understanding of AI Bias
abstract
AI technologies, despite having well-documented biases and shortcomings, are becoming increasingly pervasive across various aspects of society. AI biases often reflect and interact with broader societal biases, underscoring the need to support children in understanding these biases so that they can identify when they (or others) are being discriminated against by an AI-based system. To explore this learning through a new methodology, we built an interactive system called CLIP4KIDS. We conducted four classroom sessions with 28 fifth graders in the United States and examined our data using qualitative thematic analysis. Students frequently described AI biases in terms of “assumptions” and “stereotypes” and drew connections between historical injustices and present biases in AI models. This work contributes a novel tool for learning about AI biases, an empirical account of children’s experiences, and a theoretical analysis incorporating Vossoughi and Gutiérrez’s framework of critical pedagogy and sociocultural theory.
Aayushi Dangol, Robert Wolfe, Akeiylah DeWitt, Ben Chickadel, Julie A. Kientz, Sayamindu Dasgupta
ACM Trans. Comput. Hum. Interact.2
2024 Mediating Culture: Cultivating Socio-cultural Understanding of AI in Children through Participatory Design
abstract
The surge in access to and awareness of Generative Artificial Intelligence (GenAI) such as ChatGPT has sparked discussion over the necessary technological literacies and competencies needed to effectively engage with these systems. In this context, we explore AI as a tool that mediates cultural understanding and remediates human values – that are often influenced by biases and inequities. Using participatory design for learning with a group of 13 children (ages 8-13), we engaged in five co-design sessions featuring different modalities for socio-cultural approaches to AI literacy. We found that children were more aware of the cultural mediation aspect of AI when the content of the interaction aligned with their cultural background and context. This underscored the significance of aligning the representation of culture in these GenAI systems with people’s socio-cultural ecosystems in modern technological literacies. We conclude with design principles for a more critical and holistic approach to AI literacy.
Aayushi Dangol, Michele Newman, Robert Wolfe, Jin Ha Lee 0001, Julie A. Kientz, Jason C. Yip 0001, Caroline Pitt
Conference on Designing Interactive Systems3
2024 Representation Bias of Adolescents in AI: A Bilingual, Bicultural Study
abstract
Popular and news media often portray teenagers with sensationalism, as both a risk to society and at risk from society. As AI begins to absorb some of the epistemic functions of traditional media, we study how teenagers in two countries speaking two languages: 1) are depicted by AI, and 2) how they would prefer to be depicted. Specifically, we study the biases about teenagers learned by static word embeddings (SWEs) and generative language models (GLMs), comparing these with the perspectives of adolescents living in the U.S. and Nepal. We find English-language SWEs associate teenagers with societal problems, and more than 50% of the 1,000 words most associated with teenagers in the pretrained GloVe SWE reflect such problems. Given prompts about teenagers, 30% of outputs from GPT2-XL and 29% from LLaMA-2-7B GLMs discuss societal problems, most commonly violence, but also drug use, mental illness, and sexual taboo. Nepali models, while not free of such associations, are less dominated by social problems. Data from workshops with N=13 U.S. adolescents and N=18 Nepalese adolescents show that AI presentations are disconnected from teenage life, which revolves around activities like school and friendship. Participant ratings of how well 20 trait words describe teens are decorrelated from SWE associations, with Pearson's rho=.02, n.s. in English FastText and rho=.06, n.s. GloVe; and rho=.06, n.s. in Nepali FastText and rho=-.23, n.s. in GloVe. U.S. participants suggested AI could fairly present teens by highlighting diversity, while Nepalese participants centered positivity. Participants were optimistic that, if it learned from adolescents, rather than media sources, AI could help mitigate stereotypes. Our work offers an understanding of the ways SWEs and GLMs misrepresent a developmentally vulnerable group and provides a template for less sensationalized characterization.
Robert Wolfe, Aayushi Dangol, Bill Howe, Alexis Hiniker
AIES (1)1
2024 Dataset Scale and Societal Consistency Mediate Facial Impression Bias in Vision-Language AI
abstract
Multimodal AI models capable of associating images and text hold promise for numerous domains, ranging from automated image captioning to accessibility applications for blind and low-vision users. However, uncertainty about bias has in some cases limited their adoption and availability. In the present work, we study 43 CLIP vision-language models to determine whether they learn human-like facial impression biases, and we find evidence that such biases are reflected across three distinct CLIP model families. We show for the first time that the the degree to which a bias is shared across a society predicts the degree to which it is reflected in a CLIP model. Human-like impressions of visually unobservable attributes, like trustworthiness and sexuality, emerge only in models trained on the largest dataset, indicating that a better fit to uncurated cultural data results in the reproduction of increasingly subtle social biases. Moreover, we use a hierarchical clustering approach to show that dataset size predicts the extent to which the underlying structure of facial impression bias resembles that of facial impression bias in humans. Finally, we show that Stable Diffusion models employing CLIP as a text encoder learn facial impression biases, and that these biases intersect with racial biases in Stable Diffusion XL-Turbo. While pretrained CLIP models may prove useful for scientific studies of bias, they will also require significant dataset curation when intended for use as general-purpose models in a zero-shot setting.
Robert Wolfe, Aayushi Dangol, Alexis Hiniker, Bill Howe
AIES (1)1
2024 ML-EAT: A Multilevel Embedding Association Test for Interpretable and Transparent Social Science
abstract
This research introduces the Multilevel Embedding Association Test (ML-EAT), a method designed for interpretable and transparent measurement of intrinsic bias in language technologies. The ML-EAT addresses issues of ambiguity and difficulty in interpreting the traditional EAT measurement by quantifying bias at three levels of increasing granularity: the differential association between two target concepts with two attribute concepts; the individual effect size of each target concept with two attribute concepts; and the association between each individual target concept and each individual attribute concept. Using the ML-EAT, this research defines a taxonomy of EAT patterns describing the nine possible outcomes of an embedding association test, each of which is associated with a unique EAT-Map, a novel four-quadrant visualization for interpreting the ML-EAT. Empirical analysis of static and diachronic word embeddings, GPT-2 language models, and a CLIP language-and-image model shows that EAT patterns add otherwise unobservable information about the component biases that make up an EAT; reveal the effects of prompting in zero-shot models; and can also identify situations when cosine similarity is an ineffective metric, rendering an EAT unreliable. Our work contributes a method for rendering bias more observable and interpretable, improving the transparency of computational investigations into human minds and societies.
Robert Wolfe, Alexis Hiniker, Bill Howe
AIES (1)1
2024 The Implications of Open Generative Models in Human-Centered Data Science Work: A Case Study with Fact-Checking Organizations
abstract
Calls to use open generative language models in academic research have highlighted the need for reproducibility and transparency in scientific research. However, the impact of generative AI extends well beyond academia, as corporations and public interest organizations have begun integrating these models into their data science pipelines. We expand this lens to include the impact of open models on organizations, focusing specifically on fact-checking organizations, which use AI to observe and analyze large volumes of circulating misinformation, yet must also ensure the reproducibility and impartiality of their work. We wanted to understand where fact-checking organizations use open models in their data science pipelines; what motivates their use of open models or proprietary models; and how their use of open or proprietary models can inform research on the societal impact of generative AI. To answer these questions, we conducted an interview study with N=24 professionals at 20 fact-checking organizations on six continents. Based on these interviews, we offer a five-component conceptual model of where fact-checking organizations employ generative AI to support or automate parts of their data science pipeline, including Data Ingestion, Data Analysis, Data Retrieval, Data Delivery, and Data Sharing. We then provide taxonomies of fact-checking organizations' motivations for using open models and the limitations that prevent them for further adopting open models, finding that they prefer open models for Organizational Autonomy, Data Privacy and Ownership, Application Specificity, and Capability Transparency. However, they nonetheless use proprietary models due to perceived advantages in Performance, Usability, and Safety, as well as Opportunity Costs related to participation in emerging generative AI ecosystems. Finally, we propose a research agenda to address limitations of both open and proprietary models. Our research provides novel perspective on open models in data-driven organizations.
Robert Wolfe, Tanushree Mitra
AIES (1)1
2024 Label-Efficient Group Robustness via Out-of-Distribution Concept Curation
abstract
Deep neural networks are prone to capture correlations between spurious attributes and class labels, leading to low accuracy on some combinations of class labels and spurious attribute values. When a spurious attribute represents a protected class, these low-accuracy groups can manifest discriminatory bias. Existing methods attempting to improve worst-group accuracy assume the training data, validation data, or both are reliably labeled by the spurious attribute. But a model may be perceived to be biased towards a concept that is not represented by pre-existing labels on the training data. In these situations, the spurious attribute must be defined with external information. We propose Concept Correction, a framework that represents a concept as a curated set of images from any source, then labels each training sample by its similarity to the concept set to control spurious correlations. For example, concept sets representing gender can be used to measure and control gender bias even without explicit labels. We demonstrate and evaluate an instance of the framework as Concept DRO, which uses concept sets to estimate group labels, then uses these labels to train with a state of the art distributively robust optimization objective. We show that Concept DRO outperforms existing methods that do not require labels of spurious attributes by up to 33.1 % on three image classification datasets and is competitive with the best methods that assume access to labels. We consider how the size and quality of the concept set influences performance and find that even smaller, manually curated sets of noisy AI-generated images are effective at controlling spurious correlations, suggesting that high-quality, reusable concept sets are easy to create and effective in reducing bias.
Yiwei Yang 0009, Anthony Z. Liu, Robert Wolfe, Aylin Caliskan, Bill Howe
CVPR3
2024 "Sharing, Not Showing Off": How BeReal Approaches Authentic Self-Presentation on Social Media Through Its Design
abstract
Adolescents are particularly vulnerable to the pressures created by social media, such as heightened self-consciousness and the need for extensive self-presentation. In this study, we investigate how BeReal, a social media platform designed to counter some of these pressures, influences adolescents' self-presentation behaviors. We interviewed 29 users aged 13-18 to understand their experiences with BeReal. We found that BeReal's design focuses on spontaneous sharing, including randomly timed daily notifications and reciprocal posting, discourages staged posts, encourages careful curation of the audience, and reduces pressure on self-presentation. The space created by BeReal offers benefits such as validating an unfiltered life and reframing social comparison, but its approach to self-presentation is sometimes perceived as limited or unappealing and, at times, even toxic. Drawing on this empirical data, we propose design guidelines for platforms that support authentic self-presentation while fostering reciprocity and expanding beyond spontaneous photo-sharing. These guidelines aim to enable users to portray themselves more comprehensively and accurately, ultimately supporting teens' developmental needs, particularly in building authentic relationships.
Jaewon Kim 0002, Robert Wolfe, Ishita Chordia, Katie Davis 0001, Alexis Hiniker
Proc. ACM Hum. Comput. Interact.2
2023 Evaluating Biased Attitude Associations of Language Models in an Intersectional Context
abstract
Language models are trained on large-scale corpora that embed implicit biases documented in psychology. Valence associations (pleasantness/unpleasantness) of social groups determine the biased attitudes towards groups and concepts in social cognition. Building on this established literature, we quantify how social groups are valenced in English language models using a sentence template that provides an intersectional context. We study biases related to age, education, gender, height, intelligence, literacy, race, religion, sex, sexual orientation, social class, and weight. We present a concept projection approach to capture the valence subspace through contextualized word embeddings of language models. Adapting the projection-based approach to embedding association tests that quantify bias, we find that language models exhibit the most biased attitudes against gender identity, social class, and sexual orientation signals in language. We find that the largest and better-performing model that we study is also more biased as it effectively captures bias embedded in sociocultural data. We validate the bias evaluation method by overperforming on an intrinsic valence evaluation task. The approach enables us to measure complex intersectional biases as they are known to manifest in the outputs and applications of language models that perpetuate historical biases. Moreover, our approach contributes to design justice as it studies the associations of groups underrepresented in language such as transgender and homosexual individuals.
Shiva Omrani Sabbaghi, Robert Wolfe, Aylin Caliskan
AIES2
2022 VAST: The Valence-Assessing Semantics Test for Contextualizing Language Models
abstract
We introduce VAST, the Valence-Assessing Semantics Test, a novel intrinsic evaluation task for contextualized word embeddings (CWEs). Despite the widespread use of contextualizing language models (LMs), researchers have no intrinsic evaluation task for understanding the semantic quality of CWEs and their unique properties as related to contextualization, the change in the vector representation of a word based on surrounding words; tokenization, the breaking of uncommon words into subcomponents; and LM-specific geometry learned during training. VAST uses valence, the association of a word with pleasantness, to measure the correspondence of word-level LM semantics with widely used human judgments, and examines the effects of contextualization, tokenization, and LM-specific geometry. Because prior research has found that CWEs from OpenAI's 2019 English-language causal LM GPT-2 perform poorly on other intrinsic evaluations, we select GPT-2 as our primary subject, and include results showing that VAST is useful for 7 other LMs, and can be used in 7 languages. GPT-2 results show that the semantics of a word are more similar to the semantics of context in layers closer to model output, such that VAST scores diverge between our contextual settings, ranging from Pearson’s rho of .55 to .77 in layer 11. We also show that multiply tokenized words are not semantically encoded until layer 8, where they achieve Pearson’s rho of .46, indicating the presence of an encoding process for multiply tokenized words which differs from that of singly tokenized words, for which rho is highest in layer 0. We find that a few neurons with values having greater magnitude than the rest mask word-level semantics in GPT-2’s top layer, but that word-level semantics can be recovered by nullifying non-semantic principal components: Pearson’s rho in the top layer improves from .32 to .76. Downstream POS tagging and sentence classification experiments indicate that the GPT-2 uses these principal components for non-semantic purposes, such as to represent sentence-level syntax relevant to next-word prediction. After isolating semantics, we show the utility of VAST for understanding LM semantics via improvements over related work on four word similarity tasks, with a score of .50 on SimLex-999, better than the previous best of .45 for GPT-2. Finally, we show that 8 of 10 WEAT bias tests, which compare differences in word embedding associations between groups of words, exhibit more stereotype-congruent biases after isolating semantics, indicating that non-semantic structures in LMs also mask social biases.
Robert Wolfe, Aylin Caliskan
AAAI1
2022 Contrastive Visual Semantic Pretraining Magnifies the Semantics of Natural Language Representations
abstract
We examine the effects of contrastive visual semantic pretraining by comparing the geometry and semantic properties of contextualized English language representations formed by GPT-2 and CLIP, a zero-shot multimodal image classifier which adapts the GPT-2 architecture to encode image captions.We find that contrastive visual semantic pretraining significantly mitigates the anisotropy found in contextualized word embeddings from GPT-2, such that the intra-layer self-similarity (mean pairwise cosine similarity) of CLIP word embeddings is under .25 in all layers, compared to greater than .95 in the top layer of GPT-2.CLIP word embeddings outperform GPT-2 on wordlevel semantic intrinsic evaluation tasks, and achieve a new corpus-based state of the art for the RG65 evaluation, at .88.CLIP also forms fine-grained semantic representations of sentences, and obtains Spearman's ρ = .73on the SemEval-2017 Semantic Textual Similarity Benchmark with no fine-tuning, compared to no greater than ρ = .45 in any layer of GPT-2.Finally, intra-layer self-similarity of CLIP sentence embeddings decreases as the layer index increases, finishing at .25 in the top layer, while the self-similarity of GPT-2 sentence embeddings formed using the EOS token increases layer-over-layer and never falls below .97.Our results indicate that high anisotropy is not an inevitable consequence of contextualization, and that visual semantic pretraining is beneficial not only for ordering visual representations, but also for encoding useful semantic representations of language, both on the word level and the sentence level.
Robert Wolfe, Aylin Caliskan
ACL (1)1
2022 Gender Bias in Word Embeddings: A Comprehensive Analysis of Frequency, Syntax, and Semantics
abstract
Word embeddings are numeric representations of meaning derived from word co-occurrence statistics in corpora of human-produced texts. The statistical regularities in language corpora encode well-known social biases into word embeddings (e.g., the word vector for family is closer to the vector women than to men). Although efforts have been made to mitigate bias in word embeddings, with the hope of improving fairness in downstream Natural Language Processing (NLP) applications, these efforts will remain limited until we more deeply understand the multiple (and often subtle) ways that social biases can be reflected in word embeddings. Here, we focus on gender to provide a comprehensive analysis of group-based biases in widely-used static English word embeddings trained on internet corpora (GloVe 2014, fastText 2017). While some previous research has helped uncover biases in specific semantic associations between a group and a target domain (e.g., women - family), using the Single-Category Word Embedding Association Test, we demonstrate the widespread prevalence of gender biases that also show differences in: (1) frequencies of words associated with men versus women; (b) part-of-speech tags in gender-associated words; (c) semantic categories in gender-associated words; and (d) valence, arousal, and dominance in gender-associated words. We leave the analysis of non-binary gender to future work due to the challenges in accurate group representation caused by limitations inherent in data.
Aylin Caliskan, Pimparkar Parth Ajay, Tessa Charlesworth, Robert Wolfe, Mahzarin R. Banaji
AIES4
2022 American == White in Multimodal Language-and-Image AI
abstract
Three state-of-the-art language-and-image AI models, CLIP, SLIP, and BLIP, are evaluated for evidence of a bias previously observed in social and experimental psychology: equating American identity with being White. Embedding association tests (EATs) using standardized images of self-identified Asian, Black, Latina/o, and White individuals from the Chicago Face Database (CFD) reveal that White individuals are more associated with collective in-group words than are Asian, Black, or Latina/o individuals, with effect sizes >.4 for White vs. Asian comparisons across all models. In assessments of three core aspects of American identity reported by social psychologists, single-category EATs reveal that images of White individuals are more associated with patriotism and with being born in America, but that, consistent with prior findings in psychology, White individuals are associated with being less likely to treat people of all races and backgrounds equally. Additional tests reveal that the number of images of Black individuals returned by an image ranking task is more strongly correlated with state-level implicit bias scores for White individuals (Pearson's ρ=.63 in CLIP, ρ=.69 in BLIP) than are state demographics (ρ=.60), suggesting a relationship between regional prototypicality and implicit bias. Three downstream machine learning tasks demonstrate biases associating American with White. In a visual question answering task using BLIP, 97% of White individuals are identified as American, compared to only 3% of Asian individuals. When asked in what state the individual depicted lives in, the model responds China 53% of the time for Asian individuals, but always with an American state for White individuals. In an image captioning task, BLIP remarks upon the race of Asian individuals as much as 36% of the time, and the race of Black individuals as much as 18% of the time, but never remarks upon race for White individuals. Finally, when provided with an initialization image of individuals from the CFD and the text "an American person," a synthetic image generator (VQGAN) using the text-based guidance of CLIP consistently lightens the skin tone of individuals of all races (by 35% for Black individuals, based on mean pixel brightness), and generates output images of White individuals with blonde hair. The results indicate that societal biases equating American identity with being White are learned by multimodal language-and-image AI, and that these biases propagate to downstream applications of such models.
Robert Wolfe, Aylin Caliskan
AIES1
2021 Low Frequency Names Exhibit Bias and Overfitting in Contextualizing Language Models
abstract
We use a dataset of U.S. first names with labels based on predominant gender and racial group to examine the effect of training corpus frequency on tokenization, contextualization, similarity to initial representation, and bias in BERT, GPT-2, T5, and XLNet.We show that predominantly female and non-white names are less frequent in the training corpora of these four language models.We find that infrequent names are more self-similar across contexts, with Spearman's ρ between frequency and self-similarity as low as -.763.Infrequent names are also less similar to initial representation, with Spearman's ρ between frequency and linear centered kernel alignment (CKA) similarity to initial representation as high as .702.Moreover, we find Spearman's ρ between racial bias and name frequency in BERT of .492,indicating that lower-frequency minority group names are more associated with unpleasantness.Representations of infrequent names undergo more processing, but are more self-similar, indicating that models rely on less context-informed representations of uncommon and minority names which are overfit to a lower number of observed contexts.
Robert Wolfe, Aylin Caliskan
EMNLP (1)1