Judith E. Fan

dblp:148/8845 · DBLP profile ↗
← Back
57ranked-venue papers
3as first author
52since 2021 · last 2026
0000-0002-0097-3254ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 51 · 2 first-author · 48 since 2021Applied, interdisciplinary, general and emerging computing · 46 · 2 first-author · 43 since 2021Human-computer interaction and ubiquitous computing · 6 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Presenting Large Language Models as Companions Affects What Mental Capacities People Attribute to Them
abstract
How might messages about large language models (LLMs) found in public discourse influence the way people think about and interact with these models? To explore this question, we randomly assigned participants (N = 470) to watch short informational videos presenting LLMs as either machines, tools, or companions -- or to watch no video. We then assessed how strongly they believed LLMs to possess various mental capacities, such as the ability to have intentions or remember things. We found that participants who watched video messages presenting LLMs as companions reported believing that LLMs more fully possessed these capacities than did participants in other groups. In a follow-up study (N = 604), we replicated these findings and found nuanced effects on how these videos also impact people's reliance on LLM-generated responses when seeking out factual information. Together, these studies suggest that messages about LLMs -- beyond technical advances -- may shape what people believe about these systems and how they rely on LLM-generated responses.
Allison Chen, Sunnie S. Y. Kim, Angel Nathaniel Franyutti-Cintron, Amaya Dharmasiri, Kushin Mukherjee, Olga Russakovsky, Judith E. Fan
CHI7
2026 Gesturing Toward Abstraction: Multimodal Convention Formation in Collaborative Physical Tasks
abstract
A quintessential feature of human intelligence is the ability to create ad hoc conventions over time to achieve shared goals efficiently. We investigate how communication strategies evolve through repeated collaboration as people coordinate on shared procedural abstractions. To this end, we conducted an online unimodal study (n = 98) using natural language to probe abstraction hierarchies. In a follow-up lab study (n = 40), we examined how multimodal communication (speech and gestures) changed during physical collaboration. Pairs used augmented reality to isolate their partner’s hand and voice; one participant viewed a 3D virtual tower and sent instructions to the other, who built the physical tower. Participants became faster and more accurate by establishing linguistic and gestural abstractions and using cross-modal redundancy to emphasize key changes from previous interactions. Based on these findings, we extend probabilistic models of convention formation to multimodal settings, capturing shifts in modality preferences. Our findings and model provide building blocks for designing convention-aware intelligent agents situated in the physical world.
Kiyosu Maeda, William P. McCarthy, Ching-Yi Tsai, Jeffrey Mu, Robert D. Hawkins, Judith E. Fan, Parastoo Abtahi
CHI7
2025 Consequences of prior experience on visual problem solving
Sean P. Anderson, Lionel Wong, Maddy Bowers, Judith E. Fan
CogSci4
2025 How do we get to know someone? Diagnostic questions for inferring personal traits
Erik Brockbank, Tobias Gerstenberg, Judith E. Fan, Robert D. Hawkins
CogSci3
2025 Portraying Large Language Models as Machines, Tools, or Companions Affects What Mental Capacities People Attribute to Them
Allison Chen, Sunnie S. Y. Kim, Amaya Dharmasiri, Olga Russakovsky, Judith E. Fan
CogSci5
2025 Minds in the Making: Cognitive Science and Design Thinking
Junyi Chu, Arnav Verma, Guy Davidson, Robbie Fraser, Judith E. Fan
CogSci5
2025 What makes people think a puzzle is fun to solve?
Junyi Chu, Kristine Zheng, Judith E. Fan
CogSci3
2025 Perception as a Foundation for Common-Sense Theories of the World
Abdul-Rahim Deeb, Kevin A. Smith 0001, Shari Liu, Judith E. Fan
CogSci4
2025 Minds at School: Advancing cognitive science by measuring and modeling human learning in situ
Judith E. Fan, Kristine Zheng, Benjamin Motz 0002, Shayan Doroudi, Ji Son, Candace Thille
CogSci1
2025 Using Gesture and Language to Establish Multimodal Conventions in Collaborative Physical Tasks
Kiyosu Maeda, Ching-Yi Tsai, Judith E. Fan, Parastoo Abtahi
CogSci3
2025 Measuring sustained attention across timescales to predict learning in real-world environments
Shawn T. Schwartz, Kristine Zheng, Judith E. Fan
CogSci3
2025 Exploring the mechanisms that enable multimodal reasoning about data visualizations in vision-language models
Alexa R. Tartaglini, Christopher Potts, Judith E. Fan
CogSci3
2025 Measuring and predicting variation in the difficulty of questions about data visualizations
Arnav Verma, Judith E. Fan
CogSci2
2025 Linking student psychological orientation, engagement, and learning in college-level introductory data science
Kristine Zheng, Erik Brockbank, Shawn T. Schwartz, David Yeager, Chris Bryan, Carol S. Dweck, Judith E. Fan
CogSci7
2025 Investigating children's performance on object- and picture-based vocabulary assessments in global contexts: Evidence from Kisumu, Kenya
Rebecca Zhu, Tabitha Nduku, Joab Ochieng Arieda, Arnav Verma, Judith E. Fan, Michael C. Frank
CogSci5
2025 SketchAgent: Language-Driven Sequential Sketch Generation
abstract
Sketching serves as a versatile tool for externalizing ideas, enabling rapid exploration and visual communication that spans various disciplines. While artificial systems have driven substantial advances in content creation and human-computer interaction, capturing the dynamic and abstract nature of human sketching remains challenging. In this work, we introduce SketchAgent, a language-driven, sequential sketch generation method that enables users to create, modify, and refine sketches through dynamic, conversational interactions. Our approach requires no training or fine-tuning. Instead, we leverage the sequential nature and rich prior knowledge of off-the-shelf multimodal large language models (LLMs). We present an intuitive sketching language, introduced to the model through in-context examples, enabling it to "draw" using string-based actions. These are processed into vector graphics and then rendered to create a sketch on a pixel canvas, which can be accessed again for further tasks. By drawing stroke by stroke, our agent captures the evolving, dynamic qualities intrinsic to sketching. We demonstrate that SketchAgent can generate sketches from diverse prompts, engage in dialogue-driven drawing, and collaborate meaningfully with human users.
Yael Vinker, Tamar Rott Shaham, Kristine Zheng, Alex Zhao, Judith E. Fan, Antonio Torralba 0001
CVPR5
2024 How do video content creation goals impact which concepts people prioritize for generating B-roll imagery?
abstract
B-roll is vital when producing high-quality videos, but finding the right images can be difficult and time-consuming. Moreover, what B-roll is most effective can depend on a video content creator’s intent—is the goal to entertain, to inform, or something else? While new text-to-image generation models provide promising avenues for streamlining B-roll production, it remains unclear how these tools can provide support for content creators with different goals. To close this gap, we aimed to understand how video content creator’s goals guide which visual concepts they prioritize for B-roll generation. Here we introduce a benchmark containing judgments from > 800 people as to which terms in 12 video transcripts should be assigned highest priority for B-roll imagery accompaniment. We verified that participants reliably prioritized different visual concepts depending on whether their goal was help produce informative or entertaining videos. We next explored how well several algorithms, including heuristic approaches and large language models (LLMs), could predict systematic patterns in human judgments. We found that none of these methods fully captured human judgments in either goal condition, with state-of-the-art LLMs (i.e., GPT-4) even underperforming a baseline that sampled only nouns or nouns and adjectives. Overall, our work identifies opportunities to develop improved algorithms to support video production workflows.
Holly Huey, Mackenzie Leake, Deepali Aneja, Matthew Fisher, Judith E. Fan
Creativity & Cognition5
2024 Communicating Design Intent Using Drawing and Text
abstract
Realizing a designer’s intent in software currently requires tedious manipulation of geometric primitives, such as points and curves. By contrast, designers routinely communicate more abstract design goals to one another using an efficient combination of natural language and drawings. What would it take to develop artificial systems that understand how humans naturally convey design intent, and thereby enable more seamless interactions between humans and machines throughout the design process? First, it is vital to establish benchmarks that showcase the full range of strategies that humans use to successfully communicate about design intent. Here we take initial steps towards that goal by conducting an online study in which pairs of human participants – a “Designer” and “Maker” – collaborated over multiple turns to recreate target designs. In each turn, Designers sent messages containing language, drawings, or both to the Maker, describing how to modify an existing design toward the target. We found a preference for communicating using drawings in early turns and observed several multimodal strategies for conveying design intent. By comparing how human Makers and GPT-4V carried out instructions, we identify a gap in human and machine understanding of multimodal instructions and suggest a path for bridging this gap.
William P. McCarthy, Justin Matejka, Karl D. D. Willis, Judith E. Fan, Yewen Pu
Creativity & Cognition4
2024 Without his cookies, he's just a monster: a counterfactual simulation model of social explanation
Erik Brockbank, Justin Yang, Mishika Govil, Judith E. Fan, Tobias Gerstenberg
CogSci4
2024 COGGRAPH: Building bridges between cognitive science and computer graphics
Kartik Chandra, Anne H. K. Harrington, Katie Collins, Christopher J. Kymn, Kushin Mukherjee, Sean P. Anderson, Arnav Verma, Judith E. Fan
CogSci8
2024 How does assembling an object affect memory for it?
William P. McCarthy, Sean P. Anderson, Judith E. Fan
CogSci3
2024 Evaluating human and machine understanding of data visualizations
Arnav Verma, Kushin Mukherjee, Christopher Potts, Elisa Kreiss, Judith E. Fan
CogSci5
2024 Probabilistic simulation supports generalizable intuitive physics
Khaled Jedoui, Rahul M. V., Felix J. Binder, Josh Tenenbaum, Judith E. Fan, Dan Yamins, Kevin A. Smith 0001
CogSci6
2024 Understanding Physical Dynamics with Counterfactual World Modeling
Rahul M. V., Kevin T. Feigelis, Daniel Bear, Khaled Jedoui, Klemen Kotar, Felix J. Binder, Wanhee Lee, Sherry Liu, Kevin A. Smith 0001, Judith E. Fan, Dan Yamins
ECCV (24)11
2023 Advancing Cognitive Science and AI with Cognitive-AI Benchmarking
Felix J. Binder, Logan Matthew Cross, Yoni Friedman, Robert D. Hawkins, Dan Yamins, Judith E. Fan
CogSci6
2023 Humans choose visual subgoals to reduce cognitive cost
Felix J. Binder, Marcelo G. Mattar, David Kirsh, Judith E. Fan
CogSci4
2023 How do communicative goals guide which data visualizations people think are effective?
Holly Huey, Lauren Oey, Hannah Lloyd, Judith E. Fan
CogSci4
2023 What is graph comprehension and how do you measure it?
Hannah Lloyd, Holly Huey, Erik Brockbank, Lace M. K. Padilla, Judith E. Fan
CogSci5
2023 Measuring and Modeling Physical Intrinsic Motivation
Julio Martinez, Felix J. Binder, Nick Haber, Judith E. Fan, Dan Yamins
CogSci5
2023 How does the mind discover useful abstractions?
Marcelo G. Mattar, Judith E. Fan, Wai Keen Vong, Lionel Wong
CogSci2
2023 Evaluating machine comprehension of sketch meaning at different levels of abstraction
Kushin Mukherjee, Xuanchen Lu, Holly Huey, Yael Vinker, Rio Aguina-Kang, Ariel Shamir, Judith E. Fan
CogSci7
2023 Marks and Meanings: new perspectives on the evolution of human symbolic behavior
Kristian Tylén, Mathias Sablé-Meyer, Judith E. Fan, Michelle C. Langley
CogSci3
2023 Learning Dense Correspondences between Photos and Sketches
abstract
Humans effortlessly grasp the connection between sketches and real-world objects, even when these sketches are far from realistic. Moreover, human sketch understanding goes beyond categorization -- critically, it also entails understanding how individual elements within a sketch correspond to parts of the physical world it represents. What are the computational ingredients needed to support this ability? Towards answering this question, we make two contributions: first, we introduce a new sketch-photo correspondence benchmark, PSC6k, containing 150K annotations of 6250 sketch-photo pairs across 125 object categories, augmenting the existing Sketchy dataset with fine-grained correspondence metadata. Second, we propose a self-supervised method for learning dense correspondences between sketch-photo pairs, building upon recent advances in correspondence learning for pairs of photos. Our model uses a spatial transformer network to estimate the warp flow between latent representations of a sketch and photo extracted by a contrastive learning-based ConvNet backbone. We found that this approach outperformed several strong baselines and produced predictions that were quantitatively consistent with other warp-based methods. However, our benchmark also revealed systematic differences between predictions of the suite of models we tested and those of humans. Taken together, our work suggests a promising path towards developing artificial systems that achieve more human-like understanding of visual images at different levels of abstraction. Project page: https://photo-sketch-correspondence.github.io
Xuanchen Lu, Xiaolong Wang 0004, Judith E. Fan
ICML3
2023 SEVA: Leveraging sketches to evaluate alignment between human and machine visual abstraction
abstract
Sketching is a powerful tool for creating abstract images that are sparse but meaningful. Sketch understanding poses fundamental challenges for general-purpose vision algorithms because it requires robustness to the sparsity of sketches relative to natural visual inputs and because it demands tolerance for semantic ambiguity, as sketches can reliably evoke multiple meanings. While current vision algorithms have achieved high performance on a variety of visual tasks, it remains unclear to what extent they understand sketches in a human-like way. Here we introduce $\texttt{SEVA}$, a new benchmark dataset containing approximately 90K human-generated sketches of 128 object concepts produced under different time constraints, and thus systematically varying in sparsity. We evaluated a suite of state-of-the-art vision algorithms on their ability to correctly identify the target concept depicted in these sketches and to generate responses that are strongly aligned with human response patterns on the same sketch recognition task. We found that vision algorithms that better predicted human sketch recognition performance also better approximated human uncertainty about sketch meaning, but there remains a sizable gap between model and human response patterns. To explore the potential of models that emulate human visual abstraction in generative tasks, we conducted further evaluations of a recently developed sketch generation algorithm (Vinker et al., 2022) capable of generating sketches that vary in sparsity. We hope that public release of this dataset and evaluation protocol will catalyze progress towards algorithms with enhanced capacities for human-like visual abstraction.
Kushin Mukherjee, Holly Huey, Xuanchen Lu, Yael Vinker, Rio Aguina-Kang, Ariel Shamir, Judith E. Fan
NeurIPS7
2023 Physion++: Evaluating Physical Scene Understanding that Requires Online Inference of Different Physical Properties
abstract
General physical scene understanding requires more than simply localizing and recognizing objects -- it requires knowledge that objects can have different latent properties (e.g., mass or elasticity), and that those properties affect the outcome of physical events. While there has been great progress in physical and video prediction models in recent years, benchmarks to test their performance typically do not require an understanding that objects have individual physical properties, or at best test only those properties that are directly observable (e.g., size or color). This work proposes a novel dataset and benchmark, termed Physion++, that rigorously evaluates visual physical prediction in artificial systems under circumstances where those predictions rely on accurate estimates of the latent physical properties of objects in the scene. Specifically, we test scenarios where accurate prediction relies on estimates of properties such as mass, friction, elasticity, and deformability, and where the values of those properties can only be inferred by observing how objects move and interact with other objects or fluids. We evaluate the performance of a number of state-of-the-art prediction models that span a variety of levels of learning vs. built-in knowledge, and compare that performance to a set of human predictions. We find that models that have been trained using standard regimes and datasets do not spontaneously learn to make inferences about latent properties, but also that models that encode objectness and physical states tend to make better predictions. However, there is still a huge gap between all models and human performance, and all models' predictions correlate poorly with those made by humans, suggesting that no state-of-the-art model is learning to make physical predictions in a human-like way. These results show that current deep learning models that succeed in some settings nevertheless fail to achieve human-level physical prediction in other cases, especially those where latent property inference is required. Project page: https://dingmyu.github.io/physion_v2/
Hsiao-Yu Fish Tung, Mingyu Ding, Zhenfang Chen, Daniel Bear, Chuang Gan 0001, Josh Tenenbaum, Dan Yamins, Judith E. Fan, Kevin A. Smith 0001
NeurIPS8
2022 How do people incorporate advice from artificial agents when making physical judgments?
Erik Brockbank, Justin Yang, Suvir Mirchandani, Erdem Biyik, Dorsa Sadigh, Judith E. Fan
CogSci7
2022 Developmental changes in the semantic part structure of drawn objects
Holly Huey, Bria Long, Justin Yang, Kaylee R. George, Judith E. Fan
CogSci5
2022 From Images to Symbols: Drawing as a Window into the Mind
Kushin Mukherjee, Holly Huey, Timothy T. Rogers, Judith E. Fan
CogSci4
2022 Decomposing objects into parts from vision and language
Maneesha Nagabandi, Justin Yang, Holly Huey, Judith E. Fan
CogSci4
2022 Generalizing physical prediction by composing forces and objects
Kelsey R. Allen, Ed Vul, Judith E. Fan
CogSci4
2022 Communicating understanding of physical dynamics in natural language
Jane Yang, Ronen Tamari, Judith E. Fan
CogSci4
2022 Identifying concept libraries from language about object structure
Catherine Wong, William P. McCarthy, Gabriel Grand, Yoni Friedman, Josh Tenenbaum, Jacob Andreas, Robert D. Hawkins, Judith E. Fan
CogSci8
2021 Visual scoping operations for physical assembly
Felix J. Binder, Marcelo G. Mattar, David Kirsh, Judith E. Fan
CogSci4
2021 Measuring and predicting variation in the interestingness of physical structures
Cameron Holdaway, Daniel Bear, Samaher Radwan, Michael C. Frank, Dan Yamins, Judith E. Fan
CogSci6
2021 Improvised Numerals Rely on 1-to-1 Correspondence
Sebastian Holt, David Barner, Judith E. Fan
CogSci3
2021 How do the semantic properties of visual explanations guide causal inference?
Holly Huey, Caren M. Walker, Judith E. Fan
CogSci3
2021 Predicting children's and adults' preferences in physical interactions via physics simulation
George Kachergis, Samaher Radwan, Bria Long, Judith E. Fan, Michael Lingelbach, Daniel Bear, Dan Yamins, Michael C. Frank
CogSci4
2021 Learning to communicate about shared procedural abstractions
William P. McCarthy, Robert D. Hawkins, Cameron Holdaway, Judith E. Fan
CogSci5
2021 Connecting perceptual and procedural abstractions in physical construction
William P. McCarthy, Marcelo G. Mattar, David Kirsh, Judith E. Fan
CogSci4
2021 Learning part-based abstractions for visual object concepts
Nadia Polikarpova, Judith E. Fan
CogSci3
2021 Theory Acquisition as Constraint-Based Program Synthesis
Ed Vul, Nadia Polikarpova, Judith E. Fan
CogSci4
2021 Visual communication of object concepts at different levels of abstraction
Justin Yang, Judith E. Fan
CogSci2
2020 Learning to build physical structures better over time
William P. McCarthy, David Kirsh, Judith E. Fan
CogSci3
2020 Schema and Metadata Guide the Collective Generation of Relevant and Diverse Work
abstract
While most crowd work seeks consistent answers, creative domains often seek more diverse input. The typical crowd mechanisms for controlling quality may stifle creativity, yet removing them altogether could just produce noise. Schemas and metadata provide two mechanisms for embedding existing knowledge into task environments. Schemas are expert-derived patterns designed to structure how people think through a problem. Metadata, on the other hand, illustrate a range of creative input that fits within the structure of a schema. To understand the relative effects of schemas and metadata, we conducted a study where crowd workers are asked to generate creative interpretations for a set of placemaking examples. Crowd workers were guided either by schema plus metadata, schema alone, or neither. We found that showing schema along with crowd-produced metadata helped workers contribute interpretations that are both more on-topic and diverse, compared to using the schema alone or no schema. We discuss the implications on how crowds can creatively build on insights shared by others.
Xiaotong (Tone) Xu, Judith E. Fan, Steven Dow
HCOMP2
2019 collabdraw: An Environment for Collaborative Sketching with an Artificial Agent
abstract
Sketching is one of the most accessible techniques for communicating our ideas quickly and for collaborating in real time. Here we present a web-based environment for collaborative sketching of everyday visual concepts. We explore the integration of an artificial agent, instantiated as a recurrent neural network, who is both cooperative and responsive to actions performed by its human collaborator. To evaluate the quality of the sketches produced in this environment, we conducted an experimental user study and found that sketches produced collaboratively carried as much semantically relevant information as those produced by humans on their own. Further control analyses suggest that the semantic information in these sketches were indeed the product of collaboration, rather than attributable to the contributions of the human or the artificial agent alone. Taken together, our findings attest to the potential of systems enabling real-time collaboration between humans and machines to create novel and meaningful content.
Judith E. Fan, Monica Dinculescu, David Ha
Creativity & Cognition1
2018 Drawings as a window into developmental changes in object representations
Bria Long, Judith E. Fan, Michael C. Frank
CogSci2
2015 Common object representations for visual recognition and production
Judith E. Fan, Dan Yamins, Nicholas B. Turk-Browne
CogSci1