Wai Keen Vong

dblp:176/1984 · DBLP profile ↗
← Back
14ranked-venue papers
6as first author
8since 2021 · last 2024
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 6 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 6 first-author · 6 since 2021
YearPublicationVenuePosition
2024 Beyond the Doors of Perception: Vision Transformers Represent Relations Between Objects
abstract
Though vision transformers (ViTs) have achieved state-of-the-art performance in a variety of settings, they exhibit surprising failures when performing tasks involving visual relations. This begs the question: how do ViTs attempt to perform tasks that require computing visual relations between objects? Prior efforts to interpret ViTs tend to focus on characterizing relevant low-level visual features. In contrast, we adopt methods from mechanistic interpretability to study the higher-level visual algorithms that ViTs use to perform abstract visual reasoning. We present a case study of a fundamental, yet surprisingly difficult, relational reasoning task: judging whether two visual entities are the same or different. We find that pretrained ViTs fine-tuned on this task often exhibit two qualitatively different stages of processing despite having no obvious inductive biases to do so: 1) a perceptual stage wherein local object features are extracted and stored in a disentangled representation, and 2) a relational stage wherein object representations are compared. In the second stage, we find evidence that ViTs can learn to represent somewhat abstract visual relations, a capability that has long been considered out of reach for artificial neural networks. Finally, we demonstrate that failures at either stage can prevent a model from learning a generalizable solution to our fairly simple tasks. By understanding ViTs in terms of discrete processing stages, one can more precisely diagnose and rectify shortcomings of existing and future models.
Michael A. Lepori, Alexa R. Tartaglini, Wai Keen Vong, Thomas Serre, Brenden M. Lake, Ellie Pavlick
NeurIPS3
2023 How does the mind discover useful abstractions?
Marcelo G. Mattar, Judith E. Fan, Wai Keen Vong, Lionel Wong
CogSci3
2022 Name that state: How language affects human reinforcement learning
Angela Radulescu, Wai Keen Vong, Todd M. Gureckis
CogSci2
2022 A Developmentally-Inspired Examination of Shape versus Texture Bias in Machines
Alexa R. Tartaglini, Wai Keen Vong, Brenden M. Lake
CogSci2
2022 Categorising images by generating natural language rules
Wai Keen Vong, Brenden M. Lake
CogSci1
2022 Abstract Visual Reasoning with Tangram Shapes
abstract
We introduce KILOGRAM, a resource for studying abstract visual reasoning in humans and machines.Drawing on the history of tangram puzzles as stimuli in cognitive science, we build a richly annotated dataset that, with >1k distinct stimuli, is orders of magnitude larger and more diverse than prior resources.It is both visually and linguistically richer, moving beyond whole shape descriptions to include segmentation maps and part labels.We use this resource to evaluate the abstract visual reasoning capacities of recent multi-modal models.We observe that pre-trained weights demonstrate limited abstract reasoning, which dramatically improves with fine-tuning.We also observe that explicitly describing parts aids abstract reasoning for both humans and models, especially when jointly encoding the linguistic and visual inputs.
Anya Ji, Noriyuki Kojima, Noah Rush, Alane Suhr, Wai Keen Vong, Robert D. Hawkins, Yoav Artzi
EMNLP5
2021 Fast and Flexible: Human program induction in abstract reasoning tasks
Aysja Johnson, Wai Keen Vong, Brenden M. Lake, Todd M. Gureckis
CogSci2
2021 Modeling artificial category learning from pixels: Revisiting Shepard, Hovland, and Jenkins (1961) with deep neural networks
Alexa R. Tartaglini, Wai Keen Vong, Brenden M. Lake
CogSci2
2020 Learning word-referent mappings and concepts from raw inputs
Wai Keen Vong, Brenden M. Lake
CogSci1
2018 Optimal Cooperative Inference
abstract
Cooperative transmission of data fosters rapid accumulation of knowledge by efficiently combining experiences across learners. Although well studied in human learning and increasingly in machine learning, we lack formal frameworks through which we may reason about the benefits and limitations of cooperative inference. We present such a framework. We introduce novel indices for measuring the effectiveness of probabilistic and cooperative information transmission. We relate our indices to the well-known Teaching Dimension in deterministic settings. We prove conditions under which optimal cooperative inference can be achieved, including a representation theorem that constrains the form of inductive biases for learners optimized for cooperative inference. We conclude by demonstrating how these principles may inform the design of machine learning algorithms and discuss implications for human and machine learning.
Scott Cheng-Hsin Yang, Arash Givchi, Wai Keen Vong, Patrick Shafto
AISTATS5
2018 Bayesian Teaching of Image Categories
Wai Keen Vong, Ravi B. Sojitra, Anderson Reyes, Scott Cheng-Hsin Yang, Patrick Shafto
CogSci1
2016 Do additional features help or harm during category learning? An exploration of the curse of dimensionality in human learners
Wai Keen Vong, Andrew Hendrickson, Andrew Perfors, Danielle J. Navarro
CogSci1
2014 The relevance of labels in semi-supervised learning depends on category structure
Wai Keen Vong, Andrew Perfors, Danielle J. Navarro
CogSci1
2013 The role of sampling assumptions in generalization with multiple categories
Wai Keen Vong, Andrew Hendrickson, Andrew Perfors, Danielle J. Navarro
CogSci1