Kevin A. Smith 0001

dblp:07/4241-3 · also Kevin Smith 0003 · DBLP profile ↗
← Back
42ranked-venue papers
11as first author
16since 2021 · last 2025
0000-0001-5009-0460ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 40 · 11 first-author · 16 since 2021Applied, interdisciplinary, general and emerging computing · 32 · 10 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Seeing through Occlusion: Uncertainty-aware Joint Physical Tracking and Prediction
Arijit Dasgupta, Andrew D. Bolton, Vikash Mansinghka 0001, Josh Tenenbaum, Kevin A. Smith 0001
CogSci5
2025 Perception as a Foundation for Common-Sense Theories of the World
Abdul-Rahim Deeb, Kevin A. Smith 0001, Shari Liu, Judith E. Fan
CogSci2
2025 Calculating probabilities from imagined possibilities: Limitations in 4-year-olds
Brian Leahy, Vicente Vivanco, Samuel J. Cheyette, Kevin A. Smith 0001, Lucy J. White, Roman Feiman, Laura Schulz, Josh Tenenbaum
CogSci4
2025 Ensemble Physics: Perceiving the Mass of Groups of Objects is More Than the Sum of Its Parts
Vicente Vivanco, Josh Tenenbaum, Vivian C. Paulun, Kevin A. Smith 0001
CogSci4
2025 From pixels to physics: an image-computable model of physical predictions
Nishad Gothoskar, Josh Tenenbaum, Kevin A. Smith 0001
CogSci4
2024 Probabilistic simulation supports generalizable intuitive physics
Khaled Jedoui, Rahul M. V., Felix J. Binder, Josh Tenenbaum, Judith E. Fan, Dan Yamins, Kevin A. Smith 0001
CogSci8
2024 Understanding Physical Dynamics with Counterfactual World Modeling
Rahul M. V., Kevin T. Feigelis, Daniel Bear, Khaled Jedoui, Klemen Kotar, Felix J. Binder, Wanhee Lee, Sherry Liu, Kevin A. Smith 0001, Judith E. Fan, Dan Yamins
ECCV (24)10
2024 Evaluating Large Vision-and-Language Models on Children's Mathematical Olympiads
abstract
Recent years have seen a significant progress in the general-purpose problem solving abilities of large vision and language models (LVLMs), such as ChatGPT, Gemini, etc.; some of these breakthroughs even seem to enable AI models to outperform human abilities in varied tasks that demand higher-order cognitive skills. Are the current large AI models indeed capable of generalized problem solving as humans do? A systematic analysis of AI capabilities for joint vision and text reasoning, however, is missing in the current scientific literature. In this paper, we make an effort towards filling this gap, by evaluating state-of-the-art LVLMs on their mathematical and algorithmic reasoning abilities using visuo-linguistic problems from children's Olympiads. Specifically, we consider problems from the Mathematical Kangaroo (MK) Olympiad, which is a popular international competition targeted at children from grades 1-12, that tests children's deeper mathematical abilities using puzzles that are appropriately gauged to their age and skills. Using the puzzles from MK, we created a dataset, dubbed SMART-840, consisting of 840 problems from years 2020-2024. With our dataset, we analyze LVLMs power on mathematical reasoning; their responses on our puzzles offer a direct way to compare against that of children. Our results show that modern LVLMs do demonstrate increasingly powerful reasoning skills in solving problems for higher grades, but lack the foundations to correctly answer problems designed for younger children. Further analysis shows that there is no significant correlation between the reasoning capabilities of AI models and that of young children, and their capabilities appear to be based on a different type of reasoning than the cumulative knowledge that underlies children's mathematical skills.
Anoop Cherian, Kuan-Chuan Peng, Suhas Lohit, Joanna Matthiesen, Kevin A. Smith 0001, Josh Tenenbaum
NeurIPS5
2023 "Just In Time" Representations for Mental Simulation in Intuitive Physics
Tony Chen 0003, Kelsey R. Allen, Samuel J. Cheyette, Josh Tenenbaum, Kevin A. Smith 0001
CogSci5
2023 Strategy choice for physical reasoning is (partially) sensitive to cognitive costs
Thomas Ngo, Samuel J. Cheyette, Josh Tenenbaum, Kevin A. Smith 0001
CogSci4
2023 Are Deep Neural Networks SMARTer Than Second Graders?
abstract
Recent times have witnessed an increasing number of applications of deep neural networks towards solving tasks that require superior cognitive abilities, e.g., playing Go, generating art, question answering (e.g., ChatGPT), etc. Such a dramatic progress raises the question: how generalizable are neural networks in solving problems that demand broad skills? To answer this question, we propose SMART: a Simple Multimodal Algorithmic Reasoning Task and the associated SMART-101 dataset11The SMART-101 dataset is publicly available at: https://doi.org/10.5281/zenodo.7761800, for evaluating the abstraction, deduction, and generalization abilities of neural networks in solving visuo-linguistic puzzles designed specifically for children in the 6–8 age group. Our dataset consists of 101 unique puzzles; each puzzle comprises a picture and a question, and their solution needs a mix of several elementary skills, including arithmetic, algebra, and spatial reasoning, among others. To scale our dataset towards training deep neural networks, we programmatically generate entirely new instances for each puzzle while retaining their solution algorithm. To benchmark the performance on the SMART-101 dataset, we propose a vision-and-language meta-learning model that can incorporate varied state-of-the-art neural backbones. Our experiments reveal that while powerful deep models offer reasonable performances on puzzles in a supervised setting, they are not better than random accuracy when analyzed for generalization –filling this gap may demand new multimodal learning approaches.
Anoop Cherian, Kuan-Chuan Peng, Suhas Lohit, Kevin A. Smith 0001, Josh Tenenbaum
CVPR4
2023 H-SAUR: Hypothesize, Simulate, Act, Update, and Repeat for Understanding Object Articulations from Interactions
abstract
The world is filled with articulated objects that are difficult to determine how to use from vision alone, e.g., a door might open inwards or outwards. Humans handle these objects with strategic trial-and-error: first pushing a door then pulling if that doesn't work. We enable these capabilities in autonomous agents by proposing “Hypothesize, Simulate, Act, Update, and Repeat” (H-SAUR), a probabilistic generative framework that simultaneously generates a distribution of hypotheses about how objects articulate given input observations, captures certainty over hypotheses over time, and infer plausible actions for exploration and goal-conditioned manipulation. We compare our model with existing work in manipulating objects after a handful of exploration actions, on the PartNet-Mobility dataset. We further propose a novel PuzzleBoxes benchmark that contains locked boxes that require multiple steps to solve. We show that the proposed model significantly outperforms the current state-of-the-art articulated object manipulation framework, despite using zero training data. We further improve the test-time efficiency of H-SAUR by integrating a learned prior from learning-based vision models.
Kei Ota, Hsiao-Yu Fish Tung, Kevin A. Smith 0001, Anoop Cherian, Tim K. Marks, Alan Sullivan, Asako Kanezaki, Josh Tenenbaum
ICRA3
2023 Physion++: Evaluating Physical Scene Understanding that Requires Online Inference of Different Physical Properties
abstract
General physical scene understanding requires more than simply localizing and recognizing objects -- it requires knowledge that objects can have different latent properties (e.g., mass or elasticity), and that those properties affect the outcome of physical events. While there has been great progress in physical and video prediction models in recent years, benchmarks to test their performance typically do not require an understanding that objects have individual physical properties, or at best test only those properties that are directly observable (e.g., size or color). This work proposes a novel dataset and benchmark, termed Physion++, that rigorously evaluates visual physical prediction in artificial systems under circumstances where those predictions rely on accurate estimates of the latent physical properties of objects in the scene. Specifically, we test scenarios where accurate prediction relies on estimates of properties such as mass, friction, elasticity, and deformability, and where the values of those properties can only be inferred by observing how objects move and interact with other objects or fluids. We evaluate the performance of a number of state-of-the-art prediction models that span a variety of levels of learning vs. built-in knowledge, and compare that performance to a set of human predictions. We find that models that have been trained using standard regimes and datasets do not spontaneously learn to make inferences about latent properties, but also that models that encode objectness and physical states tend to make better predictions. However, there is still a huge gap between all models and human performance, and all models' predictions correlate poorly with those made by humans, suggesting that no state-of-the-art model is learning to make physical predictions in a human-like way. These results show that current deep learning models that succeed in some settings nevertheless fail to achieve human-level physical prediction in other cases, especially those where latent property inference is required. Project page: https://dingmyu.github.io/physion_v2/
Hsiao-Yu Fish Tung, Mingyu Ding, Zhenfang Chen, Daniel Bear, Chuang Gan 0001, Josh Tenenbaum, Dan Yamins, Judith E. Fan, Kevin A. Smith 0001
NeurIPS9
2021 Meta-strategy learning in physical problem-solving: the effect of embodied experience
Kelsey R. Allen, Kevin A. Smith 0001, Laura-Ashleigh Bird, Josh Tenenbaum, Tamar R. Makin, Dorothy Cowie
CogSci2
2021 Unsupervised Discovery of 3D Physical Objects from Video
Yilun Du, Kevin A. Smith 0001, Tomer D. Ullman, Josh Tenenbaum, Jiajun Wu 0001
ICLR2
2021 AGENT: A Benchmark for Core Psychological Reasoning
abstract
For machine agents to successfully interact with humans in real-world settings, they will need to develop an understanding of human mental life. Intuitive psychology, the ability to reason about hidden mental variables that drive observable actions, comes naturally to people: even pre-verbal infants can tell agents from objects, expecting agents to act efficiently to achieve goals given constraints. Despite recent interest in machine agents that reason about other agents, it is not clear if such agents learn or hold the core psychology principles that drive human reasoning. Inspired by cognitive development studies on intuitive psychology, we present a benchmark consisting of a large dataset of procedurally generated 3D animations, AGENT (Action, Goal, Efficiency, coNstraint, uTility), structured around four scenarios (goal preferences, action efficiency, unobserved constraints, and cost-reward trade-offs) that probe key concepts of core intuitive psychology. We validate AGENT with human-ratings, propose an evaluation protocol emphasizing generalization, and compare two strong baselines built on Bayesian inverse planning and a Theory of Mind neural network. Our results suggest that to pass the designed tests of core intuitive psychology at human levels, a model must acquire or have built-in representations of how agents plan, combining utility computations and core knowledge of objects and physics.
Tianmin Shu, Abhishek Bhandwaldar, Chuang Gan 0001, Kevin A. Smith 0001, Shari Liu, Dan Gutfreund, Elizabeth S. Spelke, Josh Tenenbaum, Tomer D. Ullman
ICML4
2020 The Origins of Common Sense in Humans and Machines
Kevin A. Smith 0001, Eliza Kosoy, Alison Gopnik, Deepak Pathak, Alan Fern, Josh Tenenbaum, Tomer D. Ullman
CogSci1
2020 The fine structure of surprise in intuitive physics: when, why, and how much?
Kevin A. Smith 0001, Lingjie Mei, Shunyu Yao 0006, Jiajun Wu 0001, Elizabeth S. Spelke, Josh Tenenbaum, Tomer D. Ullman
CogSci1
2020 Abstract strategy learning underlies flexible transfer in physical problem solving
Kelsey R. Allen, Kevin A. Smith 0001, Ulyana Piterbarg, Josh Tenenbaum
CogSci2
2020 Perceiving unseen objects
Katie Collins, Josh Tenenbaum, Kevin A. Smith 0001
CogSci3
2019 Rapid Trial-and-Error Learning in Physical Problem Solving
Kelsey R. Allen, Kevin A. Smith 0001, Josh Tenenbaum
CogSci2
2019 Real-time inference of physical properties in dynamic scenes
Kevin A. Smith 0001, Mario Belledonne, Ilker Yildirim, Jiajun Wu 0001, Josh Tenenbaum
CogSci1
2019 The effects of object motion observations on physical prediction
Moyuru Yamada, Kevin A. Smith 0001, Josh Tenenbaum
CogSci2
2019 Differentiable Physics and Stable Modes for Tool-Use and Manipulation Planning - Extended Abtract
abstract
We propose to formulate physical reasoning and manipulation planning as an optimization problem that integrates first order logic, which we call Logic-Geometric Programming.
Marc Toussaint, Kelsey R. Allen, Kevin A. Smith 0001, Josh Tenenbaum
IJCAI3
2019 Modeling Expectation Violation in Intuitive Physics with Coarse Probabilistic Object Representations
abstract
From infancy, humans have expectations about how objects will move and interact. Even young children expect objects not to move through one another, teleport, or disappear. They are surprised by mismatches between physical expectations and perceptual observations, even in unfamiliar scenes with completely novel objects. A model that exhibits human-like understanding of physics should be similarly surprised, and adjust its beliefs accordingly. We propose ADEPT, a model that uses a coarse (approximate geometry) object-centric representation for dynamic 3D scene understanding. Inference integrates deep recognition networks, extended probabilistic physical simulation, and particle filtering for forming predictions and expectations across occlusion. We also present a new test set for measuring violations of physical expectations, using a range of scenarios derived from developmental psychology. We systematically compare ADEPT, baseline models, and human expectations on this test set. ADEPT outperforms standard network architectures in discriminating physically implausible scenes, and often performs this discrimination at the same level as people.
Kevin A. Smith 0001, Lingjie Mei, Shunyu Yao 0006, Jiajun Wu 0001, Elizabeth S. Spelke, Josh Tenenbaum, Tomer D. Ullman
NeurIPS1
2018 Learning to act by integrating mental simulations and physical experiments
Ishita Dasgupta 0001, Kevin A. Smith 0001, Eric Schulz, Josh Tenenbaum, Samuel Gershman
CogSci2
2018 Tiptoeing around it: Inference from absence in potentially offensive speech
Monica A. Gates, Tess L. Veuthey, Michael Henry Tessler, Kevin A. Smith 0001, Tobias Gerstenberg, Laurie Bayet, Josh Tenenbaum
CogSci4
2018 Strategies and representations in physical inference
Kevin A. Smith 0001, Josh Tenenbaum, Erin M. Anderson, Susan J. Hespos, Lance J. Rips, Chaz Firestone, Jessica B. Hamrick
CogSci1
2018 End-to-End Differentiable Physics for Learning and Control
abstract
We present a differentiable physics engine that can be integrated as a module in deep neural networks for end-to-end learning. As a result, structured physics knowledge can be embedded into larger systems, allowing them, for example, to match observations by performing precise simulations, while achieves high sample efficiency. Specifically, in this paper we demonstrate how to perform backpropagation analytically through a physical simulator defined via a linear complementarity problem. Unlike traditional finite difference methods, such gradients can be computed analytically, which allows for greater flexibility of the engine. Through experiments in diverse domains, we highlight the system's ability to learn physical parameters from data, efficiently match and simulate observed visual behavior, and readily enable control via gradient-based planning methods. Code for the engine and experiments is included with the paper.
Filipe de Avila Belbute-Peres, Kevin A. Smith 0001, Kelsey R. Allen, Josh Tenenbaum, J. Zico Kolter
NeurIPS2
2017 Simulation and heuristics in flexible tool use
Kelsey R. Allen, Kevin A. Smith 0001, Josh Tenenbaum
CogSci2
2017 Faulty Towers: A hypothetical simulation model of physical support
Tobias Gerstenberg, Kevin A. Smith 0001, Josh Tenenbaum
CogSci3
2017 Thinking inside the box: Motion prediction in contained spaces uses simulation
Kevin A. Smith 0001, Filipe Peres, Ed Vul, Joshua Tenebaum
CogSci1
2015 Think again? The amount of mental simulation tracks uncertainty in the outcome
Jessica B. Hamrick, Kevin A. Smith 0001, Thomas L. Griffiths 0001, Ed Vul
CogSci2
2015 Prospective uncertainty: The range of possible futures in physical prediction
Kevin A. Smith 0001, Ed Vul
CogSci1
2015 The 'Fundamental Attribution Error' is rational in an uncertain world
Drew Walker, Kevin A. Smith 0001, Ed Vul
CogSci2
2014 Empirical Evidence for Markov Chain Monte Carlo in Memory Search
David Bourgin, Joshua T. Abbott, Thomas L. Griffiths 0001, Kevin A. Smith 0001, Ed Vul
CogSci4
2014 Looking forwards and backwards: Similarities and differences in prediction and retrodiction
Kevin A. Smith 0001, Ed Vul
CogSci1
2013 Consistent physics underlying ballistic motion prediction
Kevin A. Smith 0001, Peter W. Battaglia, Ed Vul
CogSci1
2013 Physical predictions over time
Kevin A. Smith 0001, Eyal Dechter, Josh Tenenbaum, Ed Vul
CogSci1
2012 Sources of uncertainty in intuitive physics
Kevin A. Smith 0001, Ed Vul
CogSci1
2011 The Biomedical Resource Ontology (BRO) to enable resource discovery in clinical and translational research
abstract
The biomedical research community relies on a diverse set of resources, both within their own institutions and at other research centers. In addition, an increasing number of shared electronic resources have been developed. Without effective means to locate and query these resources, it is challenging, if not impossible, for investigators to be aware of the myriad resources available, or to effectively perform resource discovery when the need arises. In this paper, we describe the development and use of the Biomedical Resource Ontology (BRO) to enable semantic annotation and discovery of biomedical resources. We also describe the Resource Discovery System (RDS) which is a federated, inter-institutional pilot project that uses the BRO to facilitate resource discovery on the Internet. Through the RDS framework and its associated Biositemaps infrastructure, the BRO facilitates semantic search and discovery of biomedical resources, breaking down barriers and streamlining scientific research that will improve human health.
Jessica D. Tenenbaum, Patricia L. Whetzel, Kent Anderson, Charles D. Borromeo, Ivo D. Dinov, Davera Gabriel, Beth A. Kirschner, Barbara Mirel, Timothy D. Morris, Natasha F. Noy, Csongor Nyulas, David Rubenson, Paul R. Saxman, Nancy Whelan, Zachary C. Wright, Brian D. Athey, Michael J. Becich, Geoffrey S. Ginsburg, Mark A. Musen, Kevin A. Smith 0001, Alice F. Tarantal, Daniel L. Rubin, Peter Lyster
J. Biomed. Informatics21
2009 Application of Information Technology: The University of Michigan Honest Broker: A Web-based Service for Clinical and Translational Research and Practice
abstract
For the success of clinical and translational science, a seamless interoperation is required between clinical and research information technology. Addressing this need, the Michigan Clinical Research Collaboratory (MCRC) was created. The MCRC employed a standards-driven Web Services architecture to create the U-M Honest Broker, which enabled sharing of clinical and research data among medical disciplines and separate institutions. Design objectives were to facilitate sharing of data, maintain a master patient index (MPI), deidentification of data, and routing data to preauthorized destination systems for use in clinical care, research, or both. This article describes the architecture and design of the U-M HB system and the successful demonstration project. Seventy percent of eligible patients were recruited for a prospective study examining the correlation between interventional cardiac catheterizations and depression. The U-M Honest Broker delivered on the promise of using structured clinical knowledge shared among providers to help clinical and translational research.
Andrew D. Boyd, Paul R. Saxman, Dale A. Hunscher, Kevin A. Smith 0001, Timothy D. Morris, Michelle Kaston, Frederick Bayoff, Bruce Rogers, Pamela Hayes, Namrata Rajeev, Eva Kline-Rogers, Kim Eagle, Daniel J. Clauw, John F. Greden, Lee A. Green, Brian D. Athey
J. Am. Medical Informatics Assoc.4