Remi Denton

dblp:117/4224 · also Emily Denton, Emily L. Denton · DBLP profile ↗
← Back
20ranked-venue papers
5as first author
9since 2021 · last 2025
0000-0003-4915-0512ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 5 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 8 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2025 Amplifying Trans and Nonbinary Voices: A Community-Centred Harm Taxonomy for LLMs
abstract
We explore large language model (LLM) responses that may negatively impact the transgender and nonbinary (TGNB) community and introduce the Transing Transformers Toolkit, T^3, which provides resources for identifying such harmful response behaviors. The heart of T^3 is a community-centred taxonomy of harms, developed in collaboration with the TGNB community, which we complement with, amongst other guidance, suggested heuristics for evaluation. To develop the taxonomy, we adopted a multi-method approach that included surveys and focus groups with community experts. The contribution highlights the importance of community-centred approaches in mitigating harm, and outlines pathways for LLM developers to improve how their models handle TGNB-related topics.
Eddie L. Ungless, Sunipa Dev, Cynthia L. Bennett, Rebecca Gulotta, Jasmijn Bastings, Remi Denton
ACL (1)6
2025 AI and Non-Western Art Worlds: Reimagining Critical AI Futures through Artistic Inquiry and Situated Dialogue
Rida Qadri, Piotr Mirowski, Remi Denton
CHI3
2025 Towards Equitable Community-Industry Collaborations: Understanding the Experiences of Nonprofits' Collaborations with Tech Companies
abstract
Community-based partnerships are essential to creating inclusive and equitable technologies and design practices. Though recent scholarship in HCI focuses on equitable design practices, there is less focus on understanding the experiences of community-based nonprofit organizations (CBOs) when partnering with technology companies. In this paper, we focus on understanding the perspectives of CBOs by answering the following research question: What are the experiences of CBOs that have collaborated with technology companies? Through a series of design workshops with 18 participants who work at community-based nonprofits that have collaborated with technology firms, we identified four elements of community-industry collaborations that collectively shape the overall experience: divergences in cultural and organizational norms, ''setting the table,'' project relationship dynamics, and affective qualities. We conclude by discussing the power structures that impact community-industry collaboration and suggest reflective practices to guide equitable collaborations between CBOs and tech companies.
Sheena Lewis Erete, Eric Corbett, Natasha Smith-Walker, Jay L. Cunningham, Erin Gatz, Tina M. Park, Tam Perry, Lauren Wilcox, Remi Denton
Proc. ACM Hum. Comput. Interact.9
2024 ViSAGe: A Global-Scale Analysis of Visual Stereotypes in Text-to-Image Generation
abstract
Akshita Jha, Vinodkumar Prabhakaran, Remi Denton, Sarah Laszlo, Shachi Dave, Rida Qadri, Chandan K. Reddy, Sunipa Dev. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Akshita Jha, Vinodkumar Prabhakaran, Remi Denton, Sarah Laszlo, Shachi Dave, Rida Qadri, Chandan K. Reddy, Sunipa Dev
ACL (1)3
2024 SoUnD Framework: Analyzing (So)cial Representation in (Un)structured (D)ata
abstract
Decisions about how to responsibly collect, use and document data often rely upon understanding how people are represented in data. Yet, the unlabeled nature and scale of data used in foundation model development poses a direct challenge to systematic analyses of downstream risks, such as representational harms. We provide a framework designed to help RAI practitioners more easily plan and structure analyses of how people are represented in unstructured data and identify downstream risks. The framework is organized into groups of analyses that map to 3 basic questions: 1) Who is represented in the data, 2) What content is in the data, and 3) How are the two associated. We use the framework to analyze human representation in two commonly used datasets: the Common Crawl web corpus (C4) of 356 billion tokens, and the LAION-400M dataset of 400 million text-image pairs, both developed in the English language. We illustrate how the framework informs action steps for hypothetical teams faced with data use, development, and documentation decisions. Ultimately, the framework structures human representation analyses and maps out analysis planning considerations, goals, and risk mitigation actions at different stages of dataset and model development.
Mark Diaz, Sunipa Dev, Emily Reif, Remi Denton, Vinodkumar Prabhakaran
AIES (1)4
2024 "They only care to show us the wheelchair": disability representation in text-to-image AI models
abstract
This paper reports on disability representation in images output from text-to-image (T2I) generative AI systems. Through eight focus groups with 25 people with disabilities, we found that models repeatedly presented reductive archetypes for different disabilities. Often these representations reflected broader societal stereotypes and biases, which our participants were concerned to see reproduced through T2I. Our participants discussed further challenges with using these models including the current reliance on prompt engineering to reach satisfactorily diverse results. Finally, they offered suggestions for how to improve disability representation with solutions like showing multiple, heterogeneous images for a single prompt and including the prompt with images generated. Our discussion reflects on tensions and tradeoffs we found among the diverse perspectives shared to inform future research on representation-oriented generative AI system evaluation metrics and development processes.
Kelly Mack, Rida Qadri, Remi Denton, Shaun K. Kane, Cynthia L. Bennett
CHI3
2023 From Human to Data to Dataset: Mapping the Traceability of Human Subjects in Computer Vision Datasets
abstract
Computer vision is a "data hungry" field. Researchers and practitioners who work on human-centric computer vision, like facial recognition, emphasize the necessity of vast amounts of data for more robust and accurate models. Humans are seen as a data resource which can be converted into datasets. The necessity of data has led to a proliferation of gathering data from easily available sources, including "public" data from the web. Yet the use of public data has significant ethical implications for the human subjects in datasets. We bridge academic conversations on the ethics of using publicly obtained data with concerns about privacy and agency associated with computer vision applications. Specifically, we examine how practices of dataset construction from public data-not only from websites, but also from public settings and public records-make it extremely difficult for human subjects to trace their images as they are collected, converted into datasets, distributed for use, and, in some cases, retracted. We discuss two interconnected barriers current data practices present to providing an ethics of traceability for human subjects: awareness and control. We conclude with key intervention points for enabling traceability for data subjects. We also offer suggestions for an improved ethics of traceability to enable both awareness and control for individual subjects in dataset curation practices.
Morgan Klaus Scheuerman, Katherine Weathington, Tarun Mugunthan, Remi Denton, Casey Fiesler
Proc. ACM Hum. Comput. Interact.4
2022 Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding
abstract
We present Imagen, a text-to-image diffusion model with an unprecedented degree of photorealism and a deep level of language understanding. Imagen builds on the power of large transformer language models in understanding text and hinges on the strength of diffusion models in high-fidelity image generation. Our key discovery is that generic large language models (e.g., T5), pretrained on text-only corpora, are surprisingly effective at encoding text for image synthesis: increasing the size of the language model in Imagen boosts both sample fidelity and image-text alignment much more than increasing the size of the image diffusion model. Imagen achieves a new state-of-the-art FID score of 7.27 on the COCO dataset, without ever training on COCO, and human raters find Imagen samples to be on par with the COCO data itself in image-text alignment. To assess text-to-image models in greater depth, we introduce DrawBench, a comprehensive and challenging benchmark for text-to-image models. With DrawBench, we compare Imagen with recent methods including VQ-GAN+CLIP, Latent Diffusion Models, and DALL-E 2, and find that human raters prefer Imagen over other models in side-by-side comparisons, both in terms of sample quality and image-text alignment.
Chitwan Saharia, Saurabh Saxena, Lala Li, Jay Whang, Remi Denton, Seyed Kamyar Seyed Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, Jonathan Ho, David J. Fleet, Mohammad Norouzi 0002
NeurIPS6
2021 Do Datasets Have Politics? Disciplinary Values in Computer Vision Dataset Development
abstract
Data is a crucial component of machine learning. The field is reliant on data to train, validate, and test models. With increased technical capabilities, machine learning research has boomed in both academic and industry settings, and one major focus has been on computer vision. Computer vision is a popular domain of machine learning increasingly pertinent to real-world applications, from facial recognition in policing to object detection for autonomous vehicles. Given computer vision's propensity to shape machine learning research and impact human life, we seek to understand disciplinary practices around dataset documentation - how data is collected, curated, annotated, and packaged into datasets for computer vision researchers and practitioners to use for model tuning and development. Specifically, we examine what dataset documentation communicates about the underlying values of vision data and the larger practices and goals of computer vision as a field. To conduct this study, we collected a corpus of about 500 computer vision datasets, from which we sampled 114 dataset publications across different vision tasks. Through both a structured and thematic content analysis, we document a number of values around accepted data practices, what makes desirable data, and the treatment of humans in the dataset construction process. We discuss how computer vision datasets authors value efficiency at the expense of care; universality at the expense of contextuality; impartiality at the expense of positionality; and model work at the expense of data work. Many of the silenced values we identify sit in opposition with social computing practices. We conclude with suggestions on how to better incorporate silenced values into the dataset creation and curation process.
Morgan Klaus Scheuerman, Alex Hanna, Remi Denton
Proc. ACM Hum. Comput. Interact.3
2020 Social Biases in NLP Models as Barriers for Persons with Disabilities
abstract
Building equitable and inclusive NLP technologies demands consideration of whether and how social attitudes are represented in ML models.In particular, representations encoded in models often inadvertently perpetuate undesirable social biases from the data on which they are trained.In this paper, we present evidence of such undesirable biases towards mentions of disability in two different English language models: toxicity prediction and sentiment analysis.Next, we demonstrate that the neural embeddings that are the critical first step in most NLP pipelines similarly contain undesirable biases towards mentions of disability.We end by highlighting topical biases in the discourse about disability which may contribute to the observed model biases; for instance, gun violence, homelessness, and drug addiction are over-represented in texts discussing mental illness.
Ben Hutchinson, Vinodkumar Prabhakaran, Remi Denton, Kellie Webster, Stephen Denuyl
ACL3
2020 Diversity and Inclusion Metrics in Subset Selection
abstract
The ethical concept of fairness has recently been applied in machine learning (ML) settings to describe a wide range of constraints and objectives. When considering the relevance of ethical concepts to subset selection problems, the concepts of diversity and inclusion are additionally applicable in order to create outputs that account for social power and access differentials. We introduce metrics based on these concepts, which can be applied together, separately, and in tandem with additional fairness constraints. Results from human subject experiments lend support to the proposed criteria. Social choice methods can additionally be leveraged to aggregate and choose preferable sets, and we detail how these may be applied.
Margaret Mitchell, Dylan K. Baker, Nyalleng Moorosi, Remi Denton, Ben Hutchinson, Alex Hanna, Timnit Gebru, Jamie Morgenstern
AIES4
2020 Saving Face: Investigating the Ethical Concerns of Facial Recognition Auditing
abstract
Although essential to revealing biased performance, well intentioned attempts at algorithmic auditing can have effects that may harm the very populations these measures are meant to protect. This concern is even more salient while auditing biometric systems such as facial recognition, where the data is sensitive and the technology is often used in ethically questionable manners. We demonstrate a set of fiveethical concerns in the particular case of auditing commercial facial processing technology, highlighting additional design considerations and ethical tensions the auditor needs to be aware of so as not exacerbate or complement the harms propagated by the audited system. We go further to provide tangible illustrations of these concerns, and conclude by reflecting on what these concerns mean for the role of the algorithmic audit and the fundamental product limitations they reveal.
Inioluwa Deborah Raji, Timnit Gebru, Margaret Mitchell, Joy Buolamwini, Joonseok Lee, Remi Denton
AIES6
2018 Stochastic Video Generation with a Learned Prior
abstract
Generating video frames that accurately predict future world states is challenging. Existing approaches either fail to capture the full distribution of outcomes, or yield blurry generations, or both. In this paper we introduce a video generation model with a learned prior over stochastic latent variables at each time step. Video frames are generated by drawing samples from this prior and combining them with a deterministic estimate of the future frame. The approach is simple and easily trained end-to-end on a variety of datasets. Sample generations are both varied and sharp, even many frames into the future, and compare favorably to those from existing approaches.
Remi Denton, Rob Fergus
ICML1
2018 Modeling Others using Oneself in Multi-Agent Reinforcement Learning
abstract
We consider the multi-agent reinforcement learning setting with imperfect information. The reward function depends on the hidden goals of both agents, so the agents must infer the other players’ goals from their observed behavior in order to maximize their returns. We propose a new approach for learning in these domains: Self Other-Modeling (SOM), in which an agent uses its own policy to predict the other agent’s actions and update its belief of their hidden goal in an online manner. We evaluate this approach on three different tasks and show that the agents are able to learn better policies using their estimate of the other players’ goals, in both cooperative and competitive settings.
Roberta Raileanu, Remi Denton, Arthur Szlam, Rob Fergus
ICML2
2017 Unsupervised Learning of Disentangled Representations from Video
abstract
We present a new model DRNET that learns disentangled image representations from video. Our approach leverages the temporal coherence of video and a novel adversarial loss to learn a representation that factorizes each frame into a stationary part and a temporally varying component. The disentangled representation can be used for a range of tasks. For example, applying a standard LSTM to the time-vary components enables prediction of future frames. We evaluating our approach on a range of synthetic and real videos. For the latter, we demonstrate the ability to coherently generate up to several hundred steps into the future.
Remi Denton, Vighnesh Birodkar
NIPS1
2015 User Conditional Hashtag Prediction for Images
abstract
Understanding the content of user's image posts is a particularly interesting problem in social networks and web settings. Current machine learning techniques focus mostly on curated training sets of image-label pairs, and perform image classification given the pixels within the image. In this work we instead leverage the wealth of information available from users: firstly, we employ user hashtags to capture the description of image content; and secondly, we make use of valuable contextual information about the user. We show how user metadata (age, gender, etc.) combined with image features derived from a convolutional neural network can be used to perform hashtag prediction. We explore two ways of combining these heterogeneous features into a learning framework: (i) simple concatenation; and (ii) a 3-way multiplicative gating, where the image model is conditioned on the user metadata. We apply these models to a large dataset of de-identified Facebook posts and demonstrate that modeling the user can significantly improve the tag prediction quality over current state-of-the-art methods.
Remi Denton, Jason Weston, Manohar Paluri, Lubomir D. Bourdev, Rob Fergus
KDD1
2015 Deep Generative Image Models using a Laplacian Pyramid of Adversarial Networks
abstract
In this paper we introduce a generative model capable of producing high quality samples of natural images. Our approach uses a cascade of convolutional networks (convnets) within a Laplacian pyramid framework to generate images in a coarse-to-fine fashion. At each level of the pyramid a separate generative convnet model is trained using the Generative Adversarial Nets (GAN) approach. Samples drawn from our model are of significantly higher quality than existing models. In a quantitive assessment by human evaluators our CIFAR10 samples were mistaken for real images around 40% of the time, compared to 10% for GAN samples. We also show samples from more diverse datasets such as STL10 and LSUN.
Remi Denton, Soumith Chintala, Arthur Szlam, Rob Fergus
NIPS1
2014 Exploiting Linear Structure Within Convolutional Networks for Efficient Evaluation
Remi Denton, Wojciech Zaremba, Joan Bruna, Yann LeCun, Rob Fergus
NIPS1
2014 ChromoHub V2: cancer genomics
abstract
SUMMARY: Cancer genomics data produced by next-generation sequencing support the notion that epigenetic mechanisms play a central role in cancer. We have previously developed Chromohub, an open access online interface where users can map chemical, structural and biological data from public repositories on phylogenetic trees of protein families involved in chromatin mediated-signaling. Here, we describe a cancer genomics interface that was recently added to Chromohub; the frequency of mutation, amplification and change in expression of chromatin factors across large cohorts of cancer patients is regularly extracted from The Cancer Genome Atlas and the International Cancer Genome Consortium and can now be mapped on phylogenetic trees of epigenetic protein families. Explorators of chromatin signaling can now easily navigate the cancer genomics landscape of writers, readers and erasers of histone marks, chromatin remodeling complexes, histones and their chaperones. AVAILABILITY AND IMPLEMENTATION: http://www.thesgc.org/chromohub/.
Muhammad A. Shah, Remi Denton, Matthieu Schapira
Bioinform.2
2012 ChromoHub: a data hub for navigators of chromatin-mediated signalling
abstract
UNLABELLED: The rapidly increasing research activity focused on chromatin-mediated regulation of epigenetic mechanisms is generating waves of data on writers, readers and erasers of the histone code, such as protein methyltransferases, bromodomains or histone deacetylases. To make these data easily accessible to communities of research scientists coming from diverse horizons, we have created ChromoHub, an online resource where users can map on phylogenetic trees disease associations, protein structures, chemical inhibitors, histone substrates, chromosomal aberrations and other types of data extracted from public repositories and the published literature. The interface can be used to define the structural or chemical coverage of a protein family, highlight domain architectures, interrogate disease relevance or zoom in on specific genes for more detailed information. This open-access resource should serve as a hub for cell biologists, medicinal chemists, structural biologists and other navigators that explore the biology of chromatin signalling. AVAILABILITY: http://www.thesgc.org/chromohub/.
Xi Ting Zhen, Remi Denton, Brian D. Marsden, Matthieu Schapira
Bioinform.3