Jerone Theodore Alexander Andrews

dblp:222/2713 · also Jerone T. A. Andrews · DBLP profile ↗
← Back
11ranked-venue papers
2as first author
10since 2021 · last 2026
0000-0002-8552-1213ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 2 first-author · 9 since 2021Security and privacy · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
10 papers
Trustworthy machine learning · 64% Face, body and person analysis · 8% Deep learning architectures and training · 5%
Human-computer interaction and pervasive computing
1 paper
Human-AI interaction · 100%

Topics — the 20 heaviest of 22, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
fairness
4.462024
A Taxonomy of Challenges to Curating Fair Datasets · NeurIPS 2024
Position: Measure Dataset Diversity, Don't Just Claim It · ICML 2024
Resampled Datasets Are Not Enough: Mitigating Societal Bias Beyond Single Attributes · EMNLP 2024
Machine learning › Trustworthy machine learning › fairness
bias mitigation
1.422024
Efficient Bias Mitigation Without Privileged Information · ECCV (72) 2024
Men Also Do Laundry: Multi-Attribute Bias Amplification · ICML 2023
Human-AI interaction
algorithmic transparency
1.012026
Treading the Transparency Tightrope: A Taxonomy of Risks and Benefits of Foundation Model Data Transparency for Transparency Advocates · CHI 2026
Machine learning › Deep learning architectures and training
data augmentation
0.912025
GenDataAgent: On-the-fly Dataset Augmentation with Synthetic Data · ICLR 2025
Machine learning › Efficient and distributed learning
data-efficient learning
0.912025
GenDataAgent: On-the-fly Dataset Augmentation with Synthetic Data · ICLR 2025
Machine learning › Transfer learning and domain adaptation
synthetic data augmentation
0.912025
GenDataAgent: On-the-fly Dataset Augmentation with Synthetic Data · ICLR 2025
Machine learning › Trustworthy machine learning › fairness
bias evaluation
0.812024
Resampled Datasets Are Not Enough: Mitigating Societal Bias Beyond Single Attributes · EMNLP 2024
Machine learning › Trustworthy machine learning
dataset bias
0.812024
Position: Measure Dataset Diversity, Don't Just Claim It · ICML 2024
Machine learning › Trustworthy machine learning › fairness
gender bias
0.812024
Images Speak Louder than Words: Understanding and Mitigating Bias in Vision-Language Model from a Causal Mediation Perspective · EMNLP 2024
Machine learning › Probabilistic and Bayesian machine learning › causal inference › causal effect estimation
mediation analysis
0.812024
Images Speak Louder than Words: Understanding and Mitigating Bias in Vision-Language Model from a Causal Mediation Perspective · EMNLP 2024
Machine learning › Trustworthy machine learning › fairness › social bias
social bias in vision-language models
0.812024
Images Speak Louder than Words: Understanding and Mitigating Bias in Vision-Language Model from a Causal Mediation Perspective · EMNLP 2024
Machine learning › Trustworthy machine learning › fairness › bias mitigation
social bias mitigation
0.812024
Resampled Datasets Are Not Enough: Mitigating Societal Bias Beyond Single Attributes · EMNLP 2024
Computer vision › Vision and language
vision-language model
0.812024
Images Speak Louder than Words: Understanding and Mitigating Bias in Vision-Language Model from a Causal Mediation Perspective · EMNLP 2024
Machine learning › Trustworthy machine learning › fairness › algorithmic bias
bias amplification
0.712023
Men Also Do Laundry: Multi-Attribute Bias Amplification · ICML 2023
Computer vision › Face, body and person analysis › face recognition
face representation
0.712023
A View From Somewhere: Human-Centric Face Representations · ICLR 2023
Computer vision › Face, body and person analysis
human-centric computer vision
0.712023
Ethical Considerations for Responsible Data Curation · NeurIPS 2023
Machine learning › Time series and sequential data
anomaly detection
0.412019
"Unexpected Item in the Bagging Area": Anomaly Detection in X-Ray Security Images · IEEE Trans. Inf. Forensics Secur. 2019
Computer vision › Image recognition and object detection
image classification
0.312025
GenDataAgent: On-the-fly Dataset Augmentation with Synthetic Data · ICLR 2025
Empirical software engineering
developer studies
0.212024
A Taxonomy of Challenges to Curating Fair Datasets · NeurIPS 2024
Computer vision › Image recognition and object detection › image classification
object classification
0.112019
"Unexpected Item in the Bagging Area": Anomaly Detection in X-Ray Security Images · IEEE Trans. Inf. Forensics Secur. 2019

Methods — techniques the papers use, named apart from their topics

taxonomy development · 2.0document analysis · 2.0taxonomy · 1.5interview study · 1.5gradient variance sampling · 0.9generative agents · 0.9multi-attribute bias mitigation · 0.8measurement theory · 0.8dataset resampling · 0.8causal mediation analysis · 0.8bias analysis · 0.7bias amplification metric · 0.7
YearPublicationVenuePosition
2026 Treading the Transparency Tightrope: A Taxonomy of Risks and Benefits of Foundation Model Data Transparency for Transparency Advocates
abstract
Data powering AI is often opaque. Researchers, NGOs, and law and policy leaders have called for greater transparency about how data is used for training, fine-tuning, and evaluation. While data transparency is often championed as crucial, what it concretely enables is largely implicit. Similarly, the concerns developers seem to have about transparency go unstated. This lack of clarity has led some researchers to critique transparency demands as disconnected from the actual benefits—or risks—to specific stakeholders. We analyze documentation from four stakeholder groups to create a taxonomy of the risks and benefits of dataset transparency. Data transparency is perceived as either a risk or a benefit given a stakeholder’s position, rather than wholesale. We also propose data availability and data documentation as two lenses through which to consider transparency. We discuss how best to strategically promote situational data transparency that takes into account the relationship between stakeholder position, transparency modality, and benefits/risks.
Morgan Klaus Scheuerman, Wiebke Hutiri, Aida Rahmattalabi, Victoria Matthews, Alice Xiang, Jerone Theodore Alexander Andrews
CHI6
2025 GenDataAgent: On-the-fly Dataset Augmentation with Synthetic Data
abstract
We propose a generative agent that augments training datasets with synthetic data for model fine-tuning. Unlike prior work, which uniformly samples synthetic data, our agent iteratively generates relevant samples on-the-fly, aligning with the target distribution. It prioritizes synthetic data that complements difficult training samples, focusing on those with high variance in gradient updates. Experiments across several image classification tasks demonstrate the effectiveness of our approach.
Zhiteng Li, Jerone Theodore Alexander Andrews, Yunhao Ba, Yulun Zhang 0001, Alice Xiang
ICLR3
2024 Efficient Bias Mitigation Without Privileged Information
Mateo Espinosa Zarlenga, Swami Sankaranarayanan, Jerone Theodore Alexander Andrews, Zohreh Shams, Mateja Jamnik, Alice Xiang
ECCV (72)3
2024 Resampled Datasets Are Not Enough: Mitigating Societal Bias Beyond Single Attributes
abstract
Yusuke Hirota, Jerone Andrews, Dora Zhao, Orestis Papakyriakopoulos, Apostolos Modas, Yuta Nakashima, Alice Xiang. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Yusuke Hirota, Jerone Theodore Alexander Andrews, Dora Zhao, Orestis Papakyriakopoulos, Apostolos Modas, Yuta Nakashima, Alice Xiang
EMNLP2
2024 Images Speak Louder than Words: Understanding and Mitigating Bias in Vision-Language Model from a Causal Mediation Perspective
abstract
Vision-language models (VLMs) pre-trained on extensive datasets can inadvertently learn biases by correlating gender information with specific objects or scenarios.Current methods, which focus on modifying inputs and monitoring changes in the model's output probability scores, often struggle to comprehensively understand bias from the perspective of model components.We propose a framework that incorporates causal mediation analysis to measure and map the pathways of bias generation and propagation within VLMs.Our framework is applicable to a wide range of vision-language and multimodal tasks.In this work, we apply it to the object detection task and implement it on the GLIP model.This approach allows us to identify the direct effects of interventions on model bias and the indirect effects of interventions on bias mediated through different model components.Our results show that image features are the primary contributors to bias, with significantly higher impacts than text features, specifically accounting for 32.57% and 12.63% of the bias in the MSCOCO and PASCAL-SENTENCE datasets, respectively.Notably, the image encoder's contribution surpasses that of the text encoder and the deep fusion encoder.Further experimentation confirms that contributions from both language and vision modalities are aligned and non-conflicting.Consequently, focusing on blurring gender representations within the image encoder which contributes most to the model bias, reduces bias efficiently by 22.03% and 9.04% in the MSCOCO and PASCAL-SENTENCE datasets, respectively, with minimal performance loss or increased computational demands. 1
Zhaotian Weng, Zijun Gao, Jerone Theodore Alexander Andrews, Jieyu Zhao 0001
EMNLP3
2024 Position: Measure Dataset Diversity, Don't Just Claim It
abstract
Machine learning (ML) datasets, often perceived as neutral, inherently encapsulate abstract and disputed social constructs. Dataset curators frequently employ value-laden terms such as diversity, bias, and quality to characterize datasets. Despite their prevalence, these terms lack clear definitions and validation. Our research explores the implications of this issue by analyzing "diversity" across 135 image and text datasets. Drawing from social sciences, we apply principles from measurement theory to identify considerations and offer recommendations for conceptualizing, operationalizing, and evaluating diversity in datasets. Our findings have broader implications for ML research, advocating for a more nuanced and precise approach to handling value-laden properties in dataset construction.
Dora Zhao, Jerone Theodore Alexander Andrews, Orestis Papakyriakopoulos, Alice Xiang
ICML2
2024 A Taxonomy of Challenges to Curating Fair Datasets
abstract
Despite extensive efforts to create fairer machine learning (ML) datasets, there remains a limited understanding of the practical aspects of dataset curation. Drawing from interviews with 30 ML dataset curators, we present a comprehensive taxonomy of the challenges and trade-offs encountered throughout the dataset curation lifecycle. Our findings underscore overarching issues within the broader fairness landscape that impact data curation. We conclude with recommendations aimed at fostering systemic changes to better facilitate fair dataset curation practices.
Dora Zhao, Morgan Klaus Scheuerman, Pooja Chitre, Jerone Theodore Alexander Andrews, Georgia Panagiotidou 0001, Shawn Walker, Kathleen H. Pine, Alice Xiang
NeurIPS4
2023 A View From Somewhere: Human-Centric Face Representations
Jerone Theodore Alexander Andrews, Przemyslaw Joniak, Alice Xiang
ICLR1
2023 Men Also Do Laundry: Multi-Attribute Bias Amplification
abstract
The phenomenon of $\textit{bias amplification}$ occurs when models amplify training set biases at test time. Existing metrics measure bias amplification with respect to single annotated attributes (e.g., $\texttt{computer}$). However, large-scale datasets typically consist of instances with multiple attribute annotations (e.g., $\{\texttt{computer}, \texttt{keyboard}\}$). We demonstrate models can learn to exploit correlations with respect to multiple attributes, which are not accounted for by current metrics. Moreover, we show that current metrics can give the erroneous impression that little to no bias amplification has occurred as they aggregate positive and negative bias scores. Further, these metrics lack an ideal value, making them difficult to interpret. To address these shortcomings, we propose a new metric: $\textit{Multi-Attribute Bias Amplification}$. We validate our metric's utility through a bias amplification analysis on the COCO, imSitu, and CelebA datasets. Finally, we benchmark bias mitigation methods using our proposed metric, suggesting possible avenues for future bias mitigation efforts.
Dora Zhao, Jerone Theodore Alexander Andrews, Alice Xiang
ICML2
2023 Ethical Considerations for Responsible Data Curation
abstract
Human-centric computer vision (HCCV) data curation practices often neglect privacy and bias concerns, leading to dataset retractions and unfair models. HCCV datasets constructed through nonconsensual web scraping lack crucial metadata for comprehensive fairness and robustness evaluations. Current remedies are post hoc, lack persuasive justification for adoption, or fail to provide proper contextualization for appropriate application. Our research focuses on proactive, domain-specific recommendations, covering purpose, privacy and consent, and diversity, for curating HCCV evaluation datasets, addressing privacy and bias concerns. We adopt an ante hoc reflective perspective, drawing from current practices, guidelines, dataset withdrawals, and audits, to inform our considerations and recommendations.
Jerone Theodore Alexander Andrews, Dora Zhao, William Thong, Apostolos Modas, Orestis Papakyriakopoulos, Alice Xiang
NeurIPS1
2019 "Unexpected Item in the Bagging Area": Anomaly Detection in X-Ray Security Images
abstract
The role of anomaly detection in X-ray security imaging, as a supplement to targeted threat detection, is described, and a taxonomy of anomaly types in this domain is presented. Algorithms are described for detecting appearance anomalies of shape, texture, and density, and semantic anomalies of object category presence. The anomalies are detected on the basis of representations extracted from a convolutional neural network pre-trained to identify object categories in photographs, from the final pooling layer for appearance anomalies, and from the logit layer for semantic anomalies. The distribution of representations in normal data is modeled using high-dimensional, full-covariance, Gaussians, and anomalies are scored according to their likelihood relative to those models. The algorithms are tested on X-ray parcel images using stream-of-commerce data as the normal class, and parcels with firearms present the examples of anomalies to be detected. Despite the representations being learned for photographic images and the varied contents of stream-of-commerce parcels, the system, trained on stream-of-commerce images only, is able to detect 90% of firearms as anomalies, while raising false alarms on 18% of stream-of-commerce.
Lewis D. Griffin, Matthew Caldwell, Jerone Theodore Alexander Andrews, Helene Bohler
IEEE Trans. Inf. Forensics Secur.3