VLDB 2026 Research / reviewers in the wild / expert
Katie Collins
dblp:284/4959 · also Katherine M. Collins, Katherine Maeve Collins
· DBLP profile ↗
23ranked-venue papers
8as first author
22since 2021 · last 2025
0000-0002-7032-716XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 7 first-author · 20 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 4 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 4 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Learning Personalized Decision Support PoliciesabstractIndividual human decision-makers may benefit from different forms of support to improve decision outcomes, but when will each form of support yield better outcomes? In this work, we posit that personalizing access to decision support tools can be an effective mechanism for instantiating the appropriate use of AI assistance. Specifically, we propose the general problem of learning a decision support policy that, for a given input, chooses which form of support to provide to decision-makers for whom we initially have no prior information. We develop Modiste, an interactive tool to learn personalized decision support policies. Modiste leverages stochastic contextual bandit techniques to personalize a decision support policy for each decision-maker. In our computational experiments, we characterize the expertise profiles of decision-makers for whom personalized policies will outperform offline policies, including population-wide baselines. Our experiments include realistic forms of support (e.g., expert consensus and predictions from a large language model) on vision and language tasks. Our human subject experiments add nuance to and bolster our computational experiments, demonstrating the practical utility of personalized policies when real users benefit from accessing support across tasks. Umang Bhatt, Valerie Chen, Katie Collins, Parameswaran Kamalaruban, Emma Kallina, Adrian Weller, Ameet Talwalkar |
AAAI | 3 |
| 2025 | Studying Mathematical Reasoning through the Gadget Game
Jonas Bayer, Jacob Loader, Katie Collins, Simon Frieder, Adrian Weller, Josh Tenenbaum, Timothy Gowers |
CogSci | 3 |
| 2025 | Empathy in Explanation
Katie Collins, Kartik Chandra, Jonathan Ragan-Kelley, Adrian Weller, Josh Tenenbaum |
CogSci | 1 |
| 2025 | Generation and Evaluation in the Human Invention Process through the Lens of Game Design
Katie Collins, Graham Todd, Cedegao E. Zhang, Adrian Weller, Julian Togelius, Junyi Chu, Lionel Wong, Thomas L. Griffiths 0001, Josh Tenenbaum |
CogSci | 1 |
| 2025 | Modeling Open-World Cognition as On-Demand Synthesis of Probabilistic Models
Lionel Wong, Katie Collins, Lance Ying, Cedegao E. Zhang, Adrian Weller, Tobias Gerstenberg, Timothy J. O'Donnell, Alexander K. Lew, Jacob Andreas, Tyler Brooke-Wilson, Josh Tenenbaum |
CogSci | 2 |
| 2025 | Meta-reasoning: Deciding which game to play, which problem to solve, and when to quit
Lionel Wong, Tracey Mills, Ionatan Kuperwajs, Katie Collins, Thomas L. Griffiths 0001 |
CogSci | 4 |
| 2025 | What's in the Box? Reasoning about Unseen Objects from Multimodal Cues
Lance Ying, Daniel Xu, Alicia Zhang, Katie Collins, Max H. Siegel, Josh Tenenbaum |
CogSci | 4 |
| 2025 | Can Large Language Models Understand Symbolic Graphics Programs?abstractAgainst the backdrop of enthusiasm for large language models (LLMs), there is a growing need to scientifically assess their capabilities and shortcomings. This is nontrivial in part because it is difficult to find tasks which the models have not encountered during training. Utilizing symbolic graphics programs, we propose a domain well-suited to test multiple spatial-semantic reasoning skills of LLMs. Popular in computer graphics, these programs procedurally generate visual data. While LLMs exhibit impressive skills in general program synthesis and analysis, symbolic graphics programs offer a new layer of evaluation: they allow us to test an LLM's ability to answer semantic questions about the images or 3D geometries without a vision encoder. To semantically understand the symbolic programs, LLMs would need to possess the ability to "imagine" and reason how the corresponding graphics content would look with only the symbolic description of the local curvatures and strokes. We use this task to evaluate LLMs by creating a large benchmark for the semantic visual understanding of symbolic graphics programs, built procedurally with minimal human effort. Particular emphasis is placed on transformations of images that leave the image level semantics invariant while introducing significant changes to the underlying program. We evaluate commercial and open-source LLMs on our benchmark to assess their ability to reason about visual output of programs, finding that LLMs considered stronger at reasoning generally perform better. Lastly, we introduce a novel method to improve this ability -- Symbolic Instruction Tuning (SIT), in which the LLM is finetuned with pre-collected instruction data on symbolic graphics programs. Interestingly, we find that SIT not only improves LLM's understanding on symbolic programs, but it also improves general reasoning ability on various other benchmarks. Zeju Qiu, Weiyang Liu, Haiwen Feng, Zhen Liu 0019, Tim Z. Xiao, Katie Collins, Josh Tenenbaum, Adrian Weller, Michael J. Black, Bernhard Schölkopf |
ICLR | 6 |
| 2024 | Beyond Thumbs Up/Down: Untangling Challenges of Fine-Grained Feedback for Text-to-Image GenerationabstractHuman feedback plays a critical role in learning and refining reward models for text-to-image generation, but the optimal form the feedback should take for learning an accurate reward function has not been conclusively established. This paper investigates the effectiveness of fine-grained feedback which captures nuanced distinctions in image quality and prompt-alignment, compared to traditional coarse-grained feedback (for example, thumbs up/down or ranking between a set of options). While fine-grained feedback holds promise, particularly for systems catering to diverse societal preferences, we show that demonstrating its superiority to coarse-grained feedback is not automatic. Through experiments on real and synthetic preference data, we surface the complexities of building effective models due to the interplay of model choice, feedback type, and the alignment between human judgment and computational interpretation. We identify key challenges in eliciting and utilizing fine-grained feedback, prompting a reassessment of its assumed benefits and practicality. Our findings -- e.g., that fine-grained feedback can lead to worse models for a fixed budget, in some settings; however, in controlled settings with known attributes, fine grained rewards can indeed be more helpful -- call for careful consideration of feedback attributes and potentially beckon novel modeling approaches to appropriately unlock the potential value of fine-grained feedback in-the-wild. Katie Collins, Najoung Kim, Yonatan Bitton, Verena Rieser, Shayegan Omidshafiei, Yushi Hu, Sherol Chen, Senjuti Dutta, Minsuk Chang, Kimin Lee, Youwei Liang, Georgina Evans, Sahil Singla 0005, Gang Li 0021, Adrian Weller, Junfeng He, Deepak Ramachandran, Krishnamurthy Dvijotham |
AIES (1) | 1 |
| 2024 | COGGRAPH: Building bridges between cognitive science and computer graphics
Kartik Chandra, Anne H. K. Harrington, Katie Collins, Christopher J. Kymn, Kushin Mukherjee, Sean P. Anderson, Arnav Verma, Judith E. Fan |
CogSci | 3 |
| 2024 | People use fast, goal-directed simulation to reason about novel games
Cedegao E. Zhang, Katie Collins, Lionel Wong, Adrian Weller, Josh Tenenbaum |
CogSci | 2 |
| 2024 | Rich Human Feedback for Text-to-Image GenerationabstractRecent Text-to-Image (T2I) generation models such as Stable Diffusion and Imagen have made significant progress in generating high-resolution images based on text descriptions. However, many generated images still suffer from issues such as artifacts/implausibility, misalignment with text descriptions, and low aesthetic quality. Inspired by the success of Reinforcement Learning with Human Feedback (RLHF) for large language models, prior works collected human-provided scores as feedback on generated images and trained a reward model to improve the T2I generation. In this paper, we enrich the feedback signal by (i) marking image regions that are implausible or misaligned with the text, and (ii) annotating which words in the text prompt are misrepresented or missing on the image. We collect such rich human feedback on 18K generated images (RichHF-18K) and train a multimodal transformer to predict the rich feedback automatically. We show that the predicted rich human feedback can be leveraged to improve image generation, for example, by selecting high-quality training data to finetune and improve the generative models, or by creating masks with predicted heatmaps to inpaint the problematic regions. Notably, the improvements generalize to models (Muse) beyond those used to generate the images on which human feedback data were collected (Stable Diffusion variants). The RichHF-18K data set will be released in our GitHub repository: https://github.com/google-research/google-research/tree/master/richhf_18k. Youwei Liang, Junfeng He, Gang Li 0021, Peizhao Li, Arseniy Klimovskiy, Nicholas Carolan, Jiao Sun, Jordi Pont-Tuset, Sarah Young, Feng Yang 0008, Junjie Ke, Krishnamurthy Dvijotham, Katie Collins, Yiwen Luo, Yang Li 0058, Kai Kohlhoff, Deepak Ramachandran, Vidhya Navalpakkam |
CVPR | 13 |
| 2024 | Large Language Models Must Be Taught to Know What They Don't KnowabstractWhen using large language models (LLMs) in high-stakes applications, we need to know when we can trust their predictions. Some works argue that prompting high-performance LLMs is sufficient to produce calibrated uncertainties, while others introduce sampling methods that can be prohibitively expensive. In this work, we first argue that prompting on its own is insufficient to achieve good calibration and then show that fine-tuning on a small dataset of correct and incorrect answers can create an uncertainty estimate with good generalization and small computational overhead. We show that a thousand graded examples are sufficient to outperform baseline methods and that training through the features of a model is necessary for good performance and tractable for large open-source models when using LoRA. We also investigate the mechanisms that enable reliable LLM uncertainty estimation, finding that many models can be used as general-purpose uncertainty estimators, applicable not just to their own uncertainties but also the uncertainty of other models. Lastly, we show that uncertainty estimates inform human use of LLMs in human-AI collaborative settings through a user study. Sanyam Kapoor, Nate Gruver, Manley Roberts, Katie Collins, Arka Pal, Umang Bhatt, Adrian Weller, Samuel Dooley, Micah Goldblum, Andrew Gordon Wilson |
NeurIPS | 4 |
| 2023 | Human Uncertainty in Concept-Based AI SystemsabstractPlacing a human in the loop may help abate the risks of deploying AI systems in safety-critical settings (e.g., a clinician working with a medical AI system). However, mitigating risks arising from human error and uncertainty within such human-AI interactions is an important and understudied issue. In this work, we study human uncertainty in the context of concept-based models, a family of AI systems that enable human feedback via concept interventions where an expert intervenes on human-interpretable concepts relevant to the task. Prior work in this space often assumes that humans are oracles who are always certain and correct. Yet, real-world decision-making by humans is prone to occasional mistakes and uncertainty. We study how existing concept-based models deal with uncertain interventions from humans using two novel datasets: UMNIST, a visual dataset with controlled simulated uncertainty based on the MNIST dataset, and CUB-S, a relabeling of the popular CUB concept dataset with rich, densely-annotated soft labels from humans. We show that training with uncertain concept labels may help mitigate weaknesses of concept-based systems when handling uncertain interventions. These results allow us to identify several open challenges, which we argue can be tackled through future multidisciplinary research on building interactive uncertainty-aware systems. To facilitate further research, we release a new elicitation platform, UElic, to collect uncertain feedback from humans in collaborative prediction tasks. Katie Collins, Matthew Barker, Mateo Espinosa Zarlenga, Naveen Raman 0001, Umang Bhatt, Mateja Jamnik, Ilia Sucholutsky, Adrian Weller, Krishnamurthy Dvijotham |
AIES | 1 |
| 2023 | Robotics for the Streets: Open-Source Robotics for AcademicsabstractOpen-source hardware are designs that are publicly available for anyone to modify, distribute, make, or sell. Open-source software is source code that anyone can access, inspect, modify, improve, and distribute. Open-source robotics builds upon the principles of open-source software and open-source hardware. The Open-Source Hardware Association (OSHWA) has a goal to encourage research and foster technical knowledge by making it more accessible and collaborative. The “Robotics for the Streets” project has a goal to expand open-source hardware in academia by documenting how to use it for service, teaching and research by making a library of resources for others to follow. There is a secondary goal of diversifying science, technology, engineering, and math (STEM) by using open-source robotics to increase access to and visibility of STEM technology for marginalized, minoritized, and under-resourced communities. This paper will describe the design, creation and dissemination of the open-source mobile robot platform, Flower∞Bots. Flower∞Bots consist of three robots at the novice, intermediate, and expert levels including Lily∞Bot, Daisy∞Bot and Rosie∞Bot, respectively. These robots have the flexibility and modularity for modification based upon the user's needs. The benefits of this platform include the ability to be appropriate for any level and used for a variety of use cases. This research project is a unique, novel, and innovative practice for engineering education and research because it encourages diverse perspectives and voices to contribute to the creation and improvement of technology. It was hypothesized that this educational robotics platform will serve as a model and pathway for teachers, professors, practitioners, and STEM enthusiasts, with limited resources to engage in robotics outreach, teaching, and research. The platform and guidebook will enable teachers and academics to meet their professional development goals at a low cost. Carlotta A. Berry, Alejandro Marcenido Larregola, Katie Collins, Josiah McGee |
FIE | 3 |
| 2023 | Learning to Receive Help: Intervention-Aware Concept Embedding ModelsabstractConcept Bottleneck Models (CBMs) tackle the opacity of neural architectures by constructing and explaining their predictions using a set of high-level concepts. A special property of these models is that they permit concept interventions, wherein users can correct mispredicted concepts and thus improve the model's performance. Recent work, however, has shown that intervention efficacy can be highly dependent on the order in which concepts are intervened on and on the model's architecture and training hyperparameters. We argue that this is rooted in a CBM's lack of train-time incentives for the model to be appropriately receptive to concept interventions. To address this, we propose Intervention-aware Concept Embedding models (IntCEMs), a novel CBM-based architecture and training paradigm that improves a model's receptiveness to test-time interventions. Our model learns a concept intervention policy in an end-to-end fashion from where it can sample meaningful intervention trajectories at train-time. This conditions IntCEMs to effectively select and receive concept interventions when deployed at test-time. Our experiments show that IntCEMs significantly outperform state-of-the-art concept-interpretable models when provided with test-time concept interventions, demonstrating the effectiveness of our approach. Mateo Espinosa Zarlenga, Katie Collins, Krishnamurthy Dvijotham, Adrian Weller, Zohreh Shams, Mateja Jamnik |
NeurIPS | 2 |
| 2023 | Human-in-the-Loop MixupabstractAligning model representations to humans has been found to improve robustness and generalization. However, such methods often focus on standard observational data. Synthetic data is proliferating and powering many advances in machine learning; yet, it is not always clear whether synthetic labels are perceptually aligned to humans – rendering it likely model representations are not human aligned. We focus on the synthetic data used in mixup: a powerful regularizer shown to improve model robustness, generalization, and calibration. We design a comprehensive series of elicitation interfaces, which we release as HILL MixE Suite, and recruit 159 participants to provide perceptual judgments along with their uncertainties, over mixup examples. We find that human perceptions do not consistently align with the labels traditionally used for synthetic points, and begin to demonstrate the applicability of these findings to potentially increase the reliability of downstream models, particularly when incorporating human uncertainty. We release all elicited judgments in a new data hub we call H-Mix. Katie Collins, Umang Bhatt, Weiyang Liu, Vihari Piratla, Ilia Sucholutsky, Bradley C. Love, Adrian Weller |
UAI | 1 |
| 2023 | On the informativeness of supervision signalsabstractSupervised learning typically focuses on learning transferable representations from training examples annotated by humans. While rich annotations (like soft labels) carry more information than sparse annotations (like hard labels), they are also more expensive to collect. For example, while hard labels only provide information about the closest class an object belongs to (e.g., “this is a dog”), soft labels provide information about the object’s relationship with multiple classes (e.g., “this is most likely a dog, but it could also be a wolf or a coyote”). We use information theory to compare how a number of commonly-used supervision signals contribute to representation-learning performance, as well as how their capacity is affected by factors such as the number of labels, classes, dimensions, and noise. Our framework provides theoretical justification for using hard labels in the big-data regime, but richer supervision signals for few-shot learning and out-of-distribution generalization. We validate these results empirically in a series of experiments with over 1 million crowdsourced image annotations and conduct a cost-benefit analysis to establish a tradeoff curve that enables users to optimize the cost of supervising representation learning on their own datasets. Ilia Sucholutsky, Ruairidh M. Battleday, Katie Collins, Raja Marjieh, Joshua C. Peterson, Pulkit Singh, Umang Bhatt, Nori Jacoby, Adrian Weller, Thomas L. Griffiths 0001 |
UAI | 3 |
| 2022 | Structured, flexible, and robust: benchmarking and improving large language models towards more human-like behavior in out-of-distribution reasoning tasks
Katie Collins, Catherine Wong, Jiahai Feng, Megan Wei, Josh Tenenbaum |
CogSci | 1 |
| 2022 | Eliciting and Learning with Soft Labels from Every AnnotatorabstractThe labels used to train machine learning (ML) models are of paramount importance. Typically for ML classification tasks, datasets contain hard labels, yet learning using soft labels has been shown to yield benefits for model generalization, robustness, and calibration. Earlier work found success in forming soft labels from multiple annotators' hard labels; however, this approach may not converge to the best labels and necessitates many annotators, which can be expensive and inefficient. We focus on efficiently eliciting soft labels from individual annotators. We collect and release a dataset of soft labels (which we call CIFAR-10S) over the CIFAR-10 test set via a crowdsourcing study (N=248). We demonstrate that learning with our labels achieves comparable model performance to prior approaches while requiring far fewer annotators -- albeit with significant temporal costs per elicitation. Our elicitation methodology therefore shows nuanced promise in enabling practitioners to enjoy the benefits of improved model performance and reliability with fewer annotators, and serves as a guide for future dataset curators on the benefits of leveraging richer information, such as categorical uncertainty, from individual annotators. Katie Collins, Umang Bhatt, Adrian Weller |
HCOMP | 1 |
| 2022 | Hybrid Memoised Wake-Sleep: Approximate Inference at the Discrete-Continuous Interface
Tuan Anh Le 0001, Katie Collins, Luke Hewitt, Kevin Ellis, N. Siddharth 0001, Samuel Gershman, Josh Tenenbaum |
ICLR | 2 |
| 2021 | Learning Signal-Agnostic Manifolds of Neural FieldsabstractDeep neural networks have been used widely to learn the latent structure of datasets, across modalities such as images, shapes, and audio signals. However, existing models are generally modality-dependent, requiring custom architectures and objectives to process different classes of signals. We leverage neural fields to capture the underlying structure in image, shape, audio and cross-modal audiovisual domains in a modality-independent manner. We cast our task as one of learning a manifold, where we aim to infer a low-dimensional, locally linear subspace in which our data resides. By enforcing coverage of the manifold, local linearity, and local isometry, our model -- dubbed GEM -- learns to capture the underlying structure of datasets across modalities. We can then travel along linear regions of our manifold to obtain perceptually consistent interpolations between samples, and can further use GEM to recover points on our manifold and glean not only diverse completions of input images, but cross-modal hallucinations of audio or image signals. Finally, we show that by walking across the underlying manifold of GEM, we may generate new samples in our signal domains. Yilun Du, Katie Collins, Josh Tenenbaum, Vincent Sitzmann |
NeurIPS | 2 |
| 2020 | Perceiving unseen objects
Katie Collins, Josh Tenenbaum, Kevin A. Smith 0001 |
CogSci | 1 |