VLDB 2026 Research / reviewers in the wild / expert
Candace Ross
dblp:228/5496
· DBLP profile ↗
10ranked-venue papers
3as first author
9since 2021 · last 2025
0009-0002-8285-9224ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Trustworthy machine learning · 42% Generative modeling · 25% Vision and language · 11% |
Topics — the 16 heaviest of 20, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
fairness |
2.0 | 3 | 2024 | Leveraging Diffusion Perturbations for Measuring Fairness in Computer Vision · AAAI 2024 FACET: Fairness in Computer Vision Evaluation Benchmark · ICCV 2023 Perturbation Augmentation for Fairer NLP · EMNLP 2022 |
Machine learning › Generative modeling › generative model evaluation
diversity evaluation |
0.9 | 1 | 2025 | DIMCIM: A Quantitative Evaluation Framework for Default-Mode Diversity and Generalization in Text-to-Image Generative Models · ICCV 2025 |
Machine learning › Trustworthy machine learning
hallucination |
0.9 | 1 | 2025 | What's in Common? Multimodal Models Hallucinate When Reasoning Across Scenes · NeurIPS 2025 |
Natural language and speech › Language models and text generation
multimodal language model |
0.9 | 1 | 2025 | What's in Common? Multimodal Models Hallucinate When Reasoning Across Scenes · NeurIPS 2025 |
Machine learning › Trustworthy machine learning
robustness |
0.9 | 1 | 2025 | What's in Common? Multimodal Models Hallucinate When Reasoning Across Scenes · NeurIPS 2025 |
Machine learning › Generative modeling › diffusion model
text-to-image generation |
0.9 | 1 | 2025 | DIMCIM: A Quantitative Evaluation Framework for Default-Mode Diversity and Generalization in Text-to-Image Generative Models · ICCV 2025 |
Machine learning › Generative modeling
diffusion model |
0.8 | 1 | 2024 | Leveraging Diffusion Perturbations for Measuring Fairness in Computer Vision · AAAI 2024 |
Machine learning › Trustworthy machine learning › fairness
fairness evaluation |
0.8 | 1 | 2024 | Leveraging Diffusion Perturbations for Measuring Fairness in Computer Vision · AAAI 2024 |
Machine learning › Generative modeling
image generation |
0.8 | 1 | 2024 | Improving Geo-Diversity of Generated Images with Contextualized Vendi Score Guidance · ECCV (87) 2024 |
Machine learning › Trustworthy machine learning › fairness › bias mitigation
demographic bias mitigation |
0.6 | 1 | 2022 | Perturbation Augmentation for Fairer NLP · EMNLP 2022 |
Computer vision › Vision and language › vision-language model
vision-language model evaluation |
0.6 | 1 | 2022 | Winoground: Probing Vision and Language Models for Visio-Linguistic Compositionality · CVPR 2022 |
Computer vision › Vision and language
grounded language learning |
0.3 | 1 | 2018 | Grounding language acquisition by training semantic parsers using captioned videos · EMNLP 2018 |
Natural language and speech › Information extraction and text analysis
semantic parsing |
0.3 | 1 | 2018 | Grounding language acquisition by training semantic parsers using captioned videos · EMNLP 2018 |
Machine learning › Generative modeling
generative model evaluation |
0.3 | 1 | 2025 | DIMCIM: A Quantitative Evaluation Framework for Default-Mode Diversity and Generalization in Text-to-Image Generative Models · ICCV 2025 |
Computer vision › Segmentation and scene understanding
scene understanding |
0.3 | 1 | 2025 | What's in Common? Multimodal Models Hallucinate When Reasoning Across Scenes · NeurIPS 2025 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › semantic representation
logical form |
0.1 | 1 | 2018 | Grounding language acquisition by training semantic parsers using captioned videos · EMNLP 2018 |
Methods — techniques the papers use, named apart from their topics
reference-free evaluation · 0.9multimodal language model · 0.9large language model augmentation · 0.9benchmark construction · 0.9vendi score guidance · 0.8inpainting · 0.8diffusion model · 0.8classifier guidance · 0.8intersectional analysis · 0.7fine-grained annotation analysis · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DIMCIM: A Quantitative Evaluation Framework for Default-Mode Diversity and Generalization in Text-to-Image Generative ModelsabstractRecent advances in text-to-image (T2I) models have achieved impressive quality and consistency. However, this has come at the cost of representation diversity. While automatic evaluation methods exist for benchmarking model diversity, they either require reference image datasets or lack specificity about the kind of diversity measured, limiting their adaptability and interpretability. To address this gap, we introduce the Does-it/Can-it framework, DIM-CIM, a reference-free measurement of default-mode diversity ("Does" the model generate images with expected attributes?) and generalization capacity ("Can" the model generate diverse attributes for a particular concept?). We construct the COCO-DIMCIM benchmark, which is seeded with COCO concepts and captions and augmented by a large language model. With COCO-DIMCIM, we find that widely-used models improve in generalization at the cost of default-mode diversity when scaling from 1.5B to 8.1B parameters. DIMCIM also identifies fine-grained failure cases, such as attributes that are generated with generic prompts but are rarely generated when explicitly requested. Finally, we use DIMCIM to evaluate the training data of a T2I model and observe a correlation of 0.85 between diversity in training images and default-mode diversity. Our work provides a flexible and interpretable framework for assessing T2I model diversity and generalization, enabling a more comprehensive understanding of model performance. Revant Teotia, Candace Ross, Karen Ullrich, Sumit Chopra, Adriana Romero-Soriano, Melissa Hall, Matthew J. Muckley |
ICCV | 2 |
| 2025 | Improving Model Evaluation using SMART Filtering of Benchmark DatasetsabstractVipul Gupta, Candace Ross, David Pantoja, Rebecca J. Passonneau, Megan Ung, Adina Williams. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Candace Ross, David Pantoja, Rebecca J. Passonneau, Megan Ung, Adina Williams |
NAACL (Long Papers) | 2 |
| 2025 | What's in Common? Multimodal Models Hallucinate When Reasoning Across ScenesabstractMultimodal language models possess a remarkable ability to handle an open-vocabulary worth of objects. Yet the best models still suffer from hallucinations when reasoning about scenes in the real world, revealing a gap between their seemingly strong performance on existing perception benchmarks that are saturating and their reasoning in the real world. To address this gap, we build a novel benchmark of in-the-wild scenes that we call Common-O Bench with more than 10.5k examples using exclusively new images not found in web training data to avoid contamination, Common-O goes beyond just perception, inspired by cognitive tests for humans, to probe reasoning across scenes by asking ``what’s in common?''. We evaluate leading multimodal language models, including models specifically trained to reason. We find that perceiving objects in single images is easy for most models, yet reasoning across scenes is very challenging even for the best models, including reasoning models. Despite saturating many leaderboards focusing on perception, the best performing model only achieves 35\% on Common-O Bench---and on Common-O Complex, consisting of more complex scenes, the best model achieves only 1\%. Curiously, we find models are more prone to hallucinate when similar objects are present in the scene, suggesting models may be relying on object co-occurrence seen during training. Among the models we evaluated, we found scale can provide modest improvements while models explicitly trained with multi-image inputs show bigger improvements, suggesting scaled multi-image training may offer promise. We make our benchmark publicly available to spur research into the challenge of hallucination when reasoning across scenes. Candace Ross, Florian Bordes, Adina Williams, Polina Kirichenko, Mark Ibrahim |
NeurIPS | 1 |
| 2024 | Leveraging Diffusion Perturbations for Measuring Fairness in Computer VisionabstractComputer vision models have been known to encode harmful biases, leading to the potentially unfair treatment of historically marginalized groups, such as people of color. However, there remains a lack of datasets balanced along demographic traits that can be used to evaluate the downstream fairness of these models. In this work, we demonstrate that diffusion models can be leveraged to create such a dataset. We first use a diffusion model to generate a large set of images depicting various occupations. Subsequently, each image is edited using inpainting to generate multiple variants, where each variant refers to a different perceived race. Using this dataset, we benchmark several vision-language models on a multi-class occupation classification task. We find that images generated with non-Caucasian labels have a significantly higher occupation misclassification rate than images generated with Caucasian labels, and that several misclassifications are suggestive of racial biases. We measure a model’s downstream fairness by computing the standard deviation in the probability of predicting the true occupation label across the different identity groups. Using this fairness metric, we find significant disparities between the evaluated vision-and-language models. We hope that our work demonstrates the potential value of diffusion methods for fairness evaluations. Nicholas Lui, Bryan Chia, William Berrios, Candace Ross, Douwe Kiela |
AAAI | 4 |
| 2024 | Improving Geo-Diversity of Generated Images with Contextualized Vendi Score Guidance
Reyhane Askari Hemmat, Melissa Hall, Alicia Sun, Candace Ross, Michal Drozdzal, Adriana Romero-Soriano |
ECCV (87) | 4 |
| 2023 | FACET: Fairness in Computer Vision Evaluation BenchmarkabstractComputer vision models have known performance disparities across attributes such as gender and skin tone. This means during tasks such as classification and detection, model performance differs for certain classes based on the demographics of the people in the image. These disparities have been shown to exist, but until now there has not been a unified approach to measure these differences for common use-cases of computer vision models. We present a new benchmark named FACET (FAirness in Computer Vision EvaluaTion), a large, publicly available evaluation set of 32k images for some of the most common vision tasks - image classification, object detection and segmentation. For every image in FACET, we hired expert reviewers to manually annotate person-related attributes such as perceived skin tone and hair type, manually draw bounding boxes and label fine-grained person-related classes such as disk jockey or guitarist. In addition, we use FACET to benchmark state-of-the-art vision models and present a deeper understanding of potential performance disparities and challenges across sensitive demographic attributes. With the exhaustive annotations collected, we probe models using single demographics attributes as well as multiple attributes using an intersectional approach (e.g. hair color and perceived skin tone). Our results show that classification, detection, segmentation, and visual grounding models exhibit performance disparities across demographic attributes and intersections of attributes. These harms suggest that not all people represented in datasets receive fair and equitable treatment in these vision tasks. We hope current and future results using our benchmark will contribute to fairer, more robust vision models. FACET is available publicly at https://facet.metademolab.com. Laura Gustafson, Chloé Rolland, Nikhila Ravi, Quentin Duval, Aaron Adcock, Cheng-Yang Fu, Melissa Hall, Candace Ross |
ICCV | 8 |
| 2022 | Winoground: Probing Vision and Language Models for Visio-Linguistic CompositionalityabstractWe present a novel task and dataset for evaluating the ability of vision and language models to conduct visio-linguistic compositional reasoning, which we call Winoground. Given two images and two captions, the goal is to match them correctly-but crucially, both captions contain a completely identical set of words, only in a different order. The dataset was carefully hand-curated by expert annotators and is labeled with a rich set offine-grained tags to assist in analyzing model performance. We probe a diverse range of state-of-the-art vision and language models and find that, surprisingly, none of them do much better than chance. Evidently, these models are not as skilled at visio-linguistic compositional reasoning as we might have hoped. We perform an extensive analysis to obtain insights into how future work might try to mitigate these models' shortcomings. We aim for Winoground to serve as a useful evaluation set for advancing the state of the art and driving further progress in the field. The dataset is available at https://huggingface.co/datasets/facebook/winoground. Tristan Thrush, Ryan Jiang, Max Bartolo, Amanpreet Singh, Adina Williams, Douwe Kiela, Candace Ross |
CVPR | 7 |
| 2022 | Perturbation Augmentation for Fairer NLPabstractUnwanted and often harmful social biases are becoming ever more salient in NLP research, affecting both models and datasets.In this work, we ask whether training on demographically perturbed data leads to fairer language models.We collect a large dataset of human annotated text perturbations and train a neural perturbation model, which we show outperforms heuristic alternatives.We find that (i) language models (LMs) pre-trained on demographically perturbed corpora are typically more fair, and (ii) LMs finetuned on perturbed GLUE datasets exhibit less demographic bias on downstream tasks, and (iii) fairness improvements do not come at the expense of performance on downstream tasks.Lastly, we discuss outstanding questions about how best to evaluate the (un)fairness of large language models.We hope that this exploration of neural demographic perturbation will help drive more improvement towards fairer NLP. Rebecca Qian, Candace Ross, Jude Fernandes, Eric Michael Smith, Douwe Kiela, Adina Williams |
EMNLP | 2 |
| 2021 | Measuring Social Biases in Grounded Vision and Language EmbeddingsabstractWe generalize the notion of measuring social biases in word embeddings to visually grounded word embeddings.Biases are present in grounded embeddings, and indeed seem to be equally or more significant than for ungrounded embeddings.This is despite the fact that vision and language can suffer from different biases, which one might hope could attenuate the biases in both.Multiple ways exist to generalize metrics measuring bias in word embeddings to this new setting.We introduce the space of generalizations (Grounded-WEAT and Grounded-SEAT) and demonstrate that three generalizations answer different yet important questions about how biases, language, and vision interact.These metrics are used on a new dataset, the first for grounded bias, created by augmenting standard linguistic bias benchmarks with 10,228 images from COCO, Conceptual Captions, and Google Images.Dataset construction is challenging because vision datasets are themselves very biased.The presence of these biases in systems will begin to have real-world consequences as they are deployed, making carefully measuring bias and then mitigating it critical to building a fair society. Candace Ross, Boris Katz, Andrei Barbu |
NAACL-HLT | 1 |
| 2018 | Grounding language acquisition by training semantic parsers using captioned videosabstractWe develop a semantic parser that is trained in a grounded setting using pairs of videos captioned with sentences. This setting is both data-efficient, requiring little annotation, and similar to the experience of children where they observe their environment and listen to speakers. The semantic parser recovers the meaning of English sentences despite not having access to any annotated sentences. It does so despite the ambiguity inherent in vision where a sentence may refer to any combination of objects, object properties, relations or actions taken by any agent in a video. For this task, we collected a new dataset for grounded language acquisition. Learning a grounded semantic parser — turning sentences into logical forms using captioned videos — can significantly expand the range of data that parsers can be trained on, lower the effort of training a semantic parser, and ultimately lead to a better understanding of child language acquisition. Candace Ross, Andrei Barbu, Yevgeni Berzak, Battushig Myanganbayar, Boris Katz |
EMNLP | 1 |