Samuel Gershman

dblp:44/10432 · also Sam Gershman, Samuel J. Gershman, Samuel Joseph Gershman · DBLP profile ↗
← Back
59ranked-venue papers
15as first author
22since 2021 · last 2025
0000-0002-6546-3298ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 47 · 12 first-author · 16 since 2021Applied, interdisciplinary, general and emerging computing · 37 · 7 first-author · 15 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 Quantifying the cost of context sensitivity in decision making
Shuze Liu, Samuel Gershman, Bilal Bari
CogSci2
2025 Strategy selection in complex tasks through adaptive integration of learned and online metareasoning
Tracey Mills, Samuel Gershman, Josh Tenenbaum
CogSci2
2025 Language models assign responsibility based on actual rather than counterfactual contributions
Eric J. Bigelow, Tobias Gerstenberg, Tomer D. Ullman, Samuel Gershman
CogSci5
2025 Modeling intrinsic motivation as reflective planning
Samuel Gershman
CogSci2
2025 Adaptive Social Learning using Theory of Mind
Lance Ying, Ryan Truong, Josh Tenenbaum, Samuel Gershman
CogSci4
2025 Do Mice Grok? Glimpses of Hidden Progress in Sensory Cortex
abstract
Does learning of task-relevant representations stop when behavior stops changing? Motivated by recent work in machine learning and the intuitive observation that human experts continue to learn after mastery, we hypothesize that task-specific representation learning in cortex can continue, even when behavior saturates. In a novel reanalysis of recently published neural data, we find evidence for such learning in posterior piriform cortex of mice following continued training on a task, long after behavior saturates at near-ceiling performance ("overtraining"). We demonstrate that class representations in cortex continue to separate during overtraining, so that examples that were incorrectly classified at the beginning of overtraining can abruptly be correctly classified later on, despite no changes in behavior during that time. We hypothesize this hidden learning takes the form of approximate margin maximization; we validate this and other predictions in the neural data, as well as build and interpret a simple synthetic model that recapitulates these phenomena. We conclude by demonstrating how this model of late-time feature learning implies an explanation for the empirical puzzle of overtraining reversal in animal learning, where task-specific representations are more robust to particular task changes because the learned features can be reused.
Tanishq Kumar, Blake Bordelon, Cengiz Pehlevan, Venkatesh N. Murthy, Samuel Gershman
ICLR5
2025 Blending Complementary Memory Systems in Hybrid Quadratic-Linear Transformers
abstract
We develop hybrid memory architectures for general-purpose sequence processing neural networks, that combine key-value memory using softmax attention (KV-memory) with fast weight memory through dynamic synaptic modulation (FW-memory)---the core principles of quadratic and linear transformers, respectively. These two memory systems have complementary but individually limited properties: KV-memory offers precise retrieval but is constrained by quadratic complexity in sequence length, while FW-memory supports arbitrarily long sequences and enables more expressive computation but sacrifices precise recall. We propose and compare three methods to blend these two systems into a single memory system, differing in how and when input information is delivered to each system, to leverage the strengths of both. We conduct experiments on general language modeling and retrieval tasks by training 340M- and 1.3B-parameter models from scratch, as well as on synthetic algorithmic tasks designed to precisely illustrate the benefits of certain hybrid methods over others. We also evaluate our hybrid memory systems on reinforcement learning in partially observable environments. Overall, we demonstrate how a well-designed hybrid can overcome the limitations of its individual components, offering new insights into the design principle of neural memory systems.
Kazuki Irie, Morris Yau, Samuel Gershman
NeurIPS3
2025 Gradient Descent as Loss Landscape Navigation: a Normative Framework for Deriving Learning Rules
abstract
Learning rules—prescriptions for updating model parameters to improve performance—are typically assumed rather than derived. Why do some learning rules work better than others, and under what assumptions can a given rule be considered optimal? We propose a theoretical framework that casts learning rules as policies for navigating (partially observable) loss landscapes, and identifies optimal rules as solutions to an associated optimal control problem. A range of well-known rules emerge naturally within this framework under different assumptions: gradient descent from short-horizon optimization, momentum from longer-horizon planning, natural gradients from accounting for parameter space geometry, non-gradient rules from partial controllability, and adaptive optimizers like Adam from online Bayesian inference of loss landscape shape. We further show that continual learning strategies like weight resetting can be understood as optimal responses to task uncertainty. By unifying these phenomena under a single objective, our framework clarifies the computational structure of learning and offers a principled foundation for designing adaptive algorithms.
John J. Vastola, Samuel Gershman, Kanaka Rajan
NeurIPS2
2025 Action subsampling supports policy compression in large action spaces
abstract
Real-world decision-making often involves navigating large action spaces with state-dependent action values, taxing the limited cognitive resources at our disposal. While previous studies have explored cognitive constraints on generating action consideration sets or refining state-action mappings (policy complexity), their interplay remains underexplored. In this work, we present a resource-rational framework for policy compression that integrates both constraints, offering a unified perspective on decision-making under cognitive limitations. Through simulations, we characterize the suboptimality arising from reduced action consideration sets and reveal the complex interaction between policy complexity and action consideration set size in mitigating this suboptimality. We then use such normative insight to explain empirically observed phenomena in option generation, including the preferential sampling of generally valuable options and increased correlation in responses across contexts under cognitive load. We further validate the framework's predictions through a contextual multi-armed bandit experiment, showing how humans flexibly adapt their action consideration sets and policy complexity to maintain near-optimality in a task-dependent manner. Our study demonstrates the importance of accounting for fine-grained resource constraints in understanding human cognition, and highlights the presence of adaptive metacognitive strategies even in simple tasks.
Shuze Liu, Samuel Gershman
PLoS Comput. Biol.2
2025 Nucleus accumbens dopamine release reflects Bayesian inference during instrumental learning
abstract
Dopamine release in the nucleus accumbens has been hypothesized to signal the difference between observed and predicted reward, known as reward prediction error, suggesting a biological implementation for reinforcement learning. Rigorous tests of this hypothesis require assumptions about how the brain maps sensory signals to reward predictions, yet this mapping is still poorly understood. In particular, the mapping is non-trivial when sensory signals provide ambiguous information about the hidden state of the environment. Previous work using classical conditioning tasks has suggested that reward predictions are generated conditional on probabilistic beliefs about the hidden state, such that dopamine implicitly reflects these beliefs. Here we test this hypothesis in the context of an instrumental task (a two-armed bandit), where the hidden state switches stochastically. We measured choice behavior and recorded dLight signals that reflect dopamine release in the nucleus accumbens core. Model comparison among a wide set of cognitive models based on the behavioral data favored models that used Bayesian updating of probabilistic beliefs. These same models also quantitatively matched mesolimbic dLight measurements better than non-Bayesian alternatives. We conclude that probabilistic belief computation contributes to instrumental task performance in mice and is reflected in mesolimbic dopamine signaling.
Albert J. Qü, Lung-Hao Tai, Christopher D. Hall, Emilie M. Tu, Maria K. Eckstein, Karyna Mishchanchuk, Wan Chen Lin, Juliana Chase, Andrew F. Macaskill, Anne Gabrielle Eva Collins, Samuel Gershman, Linda Wilbrecht
PLoS Comput. Biol.11
2024 Grokking as the transition from lazy to rich training dynamics
abstract
We propose that the grokking phenomenon, where the train loss of a neural network decreases much earlier than its test loss, can arise due to a neural network transitioning from lazy training dynamics to a rich, feature learning regime. To illustrate this mechanism, we study the simple setting of vanilla gradient descent on a polynomial regression problem with a two layer neural network which exhibits grokking without regularization in a way that cannot be explained by existing theories. We identify sufficient statistics for the test loss of such a network, and tracking these over training reveals that grokking arises in this setting when the network first attempts to fit a kernel regression solution with its initial features, followed by late-time feature learning where a generalizing solution is identified after train loss is already low. We find that the key determinants of grokking are the rate of feature learning---which can be controlled precisely by parameters that scale the network output---and the alignment of the initial features with the target function $y(x)$. We argue this delayed generalization arises when (1) the top eigenvectors of the initial neural tangent kernel and the task labels $y(x)$ are misaligned, but (2) the dataset size is large enough so that it is possible for the network to generalize eventually, but not so large that train loss perfectly tracks test loss at all epochs, and (3) the network begins training in the lazy regime so does not learn features immediately. We conclude with evidence that this transition from lazy (linear model) to rich training (feature learning) can control grokking in more general settings, like on MNIST, one-layer Transformers, and student-teacher networks.
Tanishq Kumar, Blake Bordelon, Samuel Gershman, Cengiz Pehlevan
ICLR3
2024 Predictive Representations: Building Blocks of Intelligence
abstract
Adaptive behavior often requires predicting future events. The theory of reinforcement learning prescribes what kinds of predictive representations are useful and how to compute them. This review integrates these theoretical ideas with work on cognition and neuroscience. We pay special attention to the successor representation and its generalizations, which have been widely applied as both engineering tools and models of brain function. This convergence suggests that particular kinds of predictive representations may function as versatile building blocks of intelligence.
Wilka Carvalho, Momchil S. Tomov, William de Cothi, Caswell Barry, Samuel Gershman
Neural Comput.5
2024 Human decision making balances reward maximization and policy compression
abstract
Policy compression is a computational framework that describes how capacity-limited agents trade reward for simpler action policies to reduce cognitive cost. In this study, we present behavioral evidence that humans prefer simpler policies, as predicted by a capacity-limited reinforcement learning model. Across a set of tasks, we find that people exploit structure in the relationships between states, actions, and rewards to "compress" their policies. In particular, compressed policies are systematically biased towards actions with high marginal probability, thereby discarding some state information. This bias is greater when there is redundancy in the reward-maximizing action policy across states, and increases with memory load. These results could not be explained qualitatively or quantitatively by models that did not make use of policy compression under a capacity limit. We also confirmed the prediction that time pressure should further reduce policy complexity and increase action bias, based on the hypothesis that actions are selected via time-dependent decoding of a compressed code. These findings contribute to a deeper understanding of how humans adapt their decision-making strategies under cognitive resource constraints.
Lucy Lai, Samuel Gershman
PLoS Comput. Biol.2
2023 Learning About Scientists from Climate Consensus Messaging
Reed Orchinik, Rachit Dubey, Samuel Gershman, Derek Powell, Rahul Bhui
CogSci3
2023 Does a Curriculum Improve Perceptual Decision Making?
Younes Strittmatter, Markus Spitzer 0002, Miguel Ruiz-Garcia, Samuel Gershman, Sebastian Musslick
CogSci4
2023 Successor-Predecessor Intrinsic Exploration
abstract
Exploration is essential in reinforcement learning, particularly in environments where external rewards are sparse. Here we focus on exploration with intrinsic rewards, where the agent transiently augments the external rewards with self-generated intrinsic rewards. Although the study of intrinsic rewards has a long history, existing methods focus on composing the intrinsic reward based on measures of future prospects of states, ignoring the information contained in the retrospective structure of transition sequences. Here we argue that the agent can utilise retrospective information to generate explorative behaviour with structure-awareness, facilitating efficient exploration based on global instead of local information. We propose Successor-Predecessor Intrinsic Exploration (SPIE), an exploration algorithm based on a novel intrinsic reward combining prospective and retrospective information. We show that SPIE yields more efficient and ethologically plausible exploratory behaviour in environments with sparse rewards and bottleneck states than competing methods. We also implement SPIE in deep reinforcement learning agents, and show that the resulting agent achieves stronger empirical performance than existing methods on sparse-reward Atari games.
Changmin Yu, Neil Burgess, Maneesh Sahani, Samuel Gershman
NeurIPS4
2023 Emergence of belief-like representations through reinforcement learning
abstract
To behave adaptively, animals must learn to predict future reward, or value. To do this, animals are thought to learn reward predictions using reinforcement learning. However, in contrast to classical models, animals must learn to estimate value using only incomplete state information. Previous work suggests that animals estimate value in partially observable tasks by first forming "beliefs"-optimal Bayesian estimates of the hidden states in the task. Although this is one way to solve the problem of partial observability, it is not the only way, nor is it the most computationally scalable solution in complex, real-world environments. Here we show that a recurrent neural network (RNN) can learn to estimate value directly from observations, generating reward prediction errors that resemble those observed experimentally, without any explicit objective of estimating beliefs. We integrate statistical, functional, and dynamical systems perspectives on beliefs to show that the RNN's learned representation encodes belief information, but only when the RNN's capacity is sufficiently large. These results illustrate how animals can estimate value in tasks without explicitly estimating beliefs, yielding a representation useful for systems with limited capacity.
Jay A. Hennig, Sandra A. Romero Pinto, Takahiro Yamaguchi, Scott W. Linderman, Naoshige Uchida, Samuel Gershman
PLoS Comput. Biol.6
2022 Coding Strategies in Memory for 3D Objects: The Influence of Task Uncertainty
Christopher Bates, Samuel Gershman
CogSci2
2022 Combining mental simulation and abstract reasoning explains people's reaction time in an intuitive physics task
Felix Sosa, Samuel Gershman, Tomer D. Ullman
CogSci2
2022 Hybrid Memoised Wake-Sleep: Approximate Inference at the Discrete-Continuous Interface
Tuan Anh Le 0001, Katie Collins, Luke Hewitt, Kevin Ellis, N. Siddharth 0001, Samuel Gershman, Josh Tenenbaum
ICLR6
2021 Neural signatures of arbitration between Pavlovian and instrumental action selection
abstract
Pavlovian associations drive approach towards reward-predictive cues, and avoidance of punishment-predictive cues. These associations "misbehave" when they conflict with correct instrumental behavior. This raises the question of how Pavlovian and instrumental influences on behavior are arbitrated. We test a computational theory according to which Pavlovian influence will be stronger when inferred controllability of outcomes is low. Using a model-based analysis of a Go/NoGo task with human subjects, we show that theta-band oscillatory power in frontal cortex tracks inferred controllability, and that these inferences predict Pavlovian action biases. Functional MRI data revealed an inferior frontal gyrus correlate of action probability and a ventromedial prefrontal correlate of outcome valence, both of which were modulated by inferred controllability.
Samuel Gershman, Marc Guitart-Masip, James F. Cavanagh
PLoS Comput. Biol.1
2021 Rational inattention and tonic dopamine
abstract
Slow-timescale (tonic) changes in dopamine (DA) contribute to a wide variety of processes in reinforcement learning, interval timing, and other domains. Furthermore, changes in tonic DA exert distinct effects depending on when they occur (e.g., during learning vs. performance) and what task the subject is performing (e.g., operant vs. classical conditioning). Two influential theories of tonic DA-the average reward theory and the Bayesian theory in which DA controls precision-have each been successful at explaining a subset of empirical findings. But how the same DA signal performs two seemingly distinct functions without creating crosstalk is not well understood. Here we reconcile the two theories under the unifying framework of 'rational inattention,' which (1) conceptually links average reward and precision, (2) outlines how DA manipulations affect this relationship, and in so doing, (3) captures new empirical phenomena. In brief, rational inattention asserts that agents can increase their precision in a task (and thus improve their performance) by paying a cognitive cost. Crucially, whether this cost is worth paying depends on average reward availability, reported by DA. The monotonic relationship between average reward and precision means that the DA signal contains the information necessary to retrieve the precision. When this information is needed after the task is performed, as presumed by Bayesian inference, acute manipulations of DA will bias behavior in predictable ways. We show how this framework reconciles a remarkably large collection of experimental findings. In reinforcement learning, the rational inattention framework predicts that learning from positive and negative feedback should be enhanced in high and low DA states, respectively, and that DA should tip the exploration-exploitation balance toward exploitation. In interval timing, this framework predicts that DA should increase the speed of the internal clock and decrease the extent of interference by other temporal stimuli during temporal reproduction (the central tendency effect). Finally, rational inattention makes the new predictions that these effects should be critically dependent on the controllability of rewards, that post-reward delays in intertemporal choice tasks should be underestimated, and that average reward manipulations should affect the speed of the clock-thus capturing empirical findings that are unexplained by either theory alone. Our results suggest that a common computational repertoire may underlie the seemingly heterogeneous roles of DA.
John G. Mikhael, Lucy Lai, Samuel Gershman
PLoS Comput. Biol.3
2020 Downloading Culture.zip: Social learning by program induction
Max Kleiman-Weiner, Felix Sosa, Bill Thompson 0001, Sebastiaan van Opheusden, Thomas L. Griffiths 0001, Samuel Gershman, Fiery Cushman
CogSci6
2020 Discovery of hierarchical representations for efficient planning
abstract
We propose that humans spontaneously organize environments into clusters of states that support hierarchical planning, enabling them to tackle challenging problems by breaking them down into sub-problems at various levels of abstraction. People constantly rely on such hierarchical presentations to accomplish tasks big and small-from planning one's day, to organizing a wedding, to getting a PhD-often succeeding on the very first attempt. We formalize a Bayesian model of hierarchy discovery that explains how humans discover such useful abstractions. Building on principles developed in structure learning and robotics, the model predicts that hierarchy discovery should be sensitive to the topological structure, reward distribution, and distribution of tasks in the environment. In five simulations, we show that the model accounts for previously reported effects of environment structure on planning behavior, such as detection of bottleneck states and transitions. We then test the novel predictions of the model in eight behavioral experiments, demonstrating how the distribution of tasks and rewards can influence planning behavior via the discovered hierarchy, sometimes facilitating and sometimes hindering performance. We find evidence that the hierarchy discovery process unfolds incrementally across trials. Finally, we propose how hierarchy discovery and hierarchical planning might be implemented in the brain. Together, these findings present an important advance in our understanding of how the brain might use Bayesian inference to discover and exploit the hidden hierarchical structure of the environment.
Momchil S. Tomov, Samyukta Yagati, Agni Kumar, Wanqian Yang, Samuel Gershman
PLoS Comput. Biol.5
2019 Downloading Culture.zip: Social learning by program induction with execution traces
Max Kleiman-Weiner, Felix Sosa, Samuel Gershman, Fiery Cushman
CogSci3
2019 Hard choices: Children's understanding of the cost of action selection
Shari Liu, Fiery Cushman, Samuel Gershman, Wouter Kool 0002, Elizabeth S. Spelke
CogSci3
2019 Generalization as diffusion: human function learning on graphs
Charley M. Wu, Eric Schulz, Samuel Gershman
CogSci3
2019 Human Evaluation of Models Built for Interpretability
abstract
Recent years have seen a boom in interest in interpretable machine learning systems built on models that can be understood, at least to some degree, by domain experts. However, exactly what kinds of models are truly human-interpretable remains poorly understood. This work advances our understanding of precisely which factors make models interpretable in the context of decision sets, a specific class of logic-based model. We conduct carefully controlled human-subject experiments in two domains across three tasks based on human-simulatability through which we identify specific types of complexity that affect performance more heavily than others-trends that are consistent across tasks and domains. These results can inform the choice of regularizers during optimization to learn more interpretable models, and their consistency suggests that there may exist common design principles for interpretable machine learning systems.
Isaac Lage, Emily Chen, Jeffrey He, Menaka Narayanan, Been Kim, Samuel Gershman, Finale Doshi-Velez
HCOMP6
2019 Estimating Scale-Invariant Future in Continuous Time
abstract
Natural learners must compute an estimate of future outcomes that follow from a stimulus in continuous time. Widely used reinforcement learning algorithms discretize continuous time and estimate either transition functions from one step to the next (model-based algorithms) or a scalar value of exponentially discounted future reward using the Bellman equation (model-free algorithms). An important drawback of model-based algorithms is that computational cost grows linearly with the amount of time to be simulated. An important drawback of model-free algorithms is the need to select a timescale required for exponential discounting. We present a computational mechanism, developed based on work in psychology and neuroscience, for computing a scale-invariant timeline of future outcomes. This mechanism efficiently computes an estimate of inputs as a function of future time on a logarithmically compressed scale and can be used to generate a scale-invariant power-law-discounted estimate of expected future reward. The representation of future time retains information about what will happen when. The entire timeline can be constructed in a single parallel operation that generates concrete behavioral and neural predictions. This computational mechanism could be incorporated into future reinforcement learning algorithms.
Zoran Tiganj, Samuel Gershman, Per B. Sederberg, Marc W. Howard
Neural Comput.2
2018 Explaining Human Decision Making in Optimal Stopping Tasks
Christiane Baumann, Henrik Singmann, Vassilios E. Kaxiras, Samuel Gershman, Bettina von Helversen
CogSci4
2018 Evaluating Compositionality in Sentence Embeddings
Ishita Dasgupta 0001, Demi Guo, Andreas Stuhlmüller, Samuel Gershman, Noah D. Goodman
CogSci4
2018 Learning to act by integrating mental simulations and physical experiments
Ishita Dasgupta 0001, Kevin A. Smith 0001, Eric Schulz, Josh Tenenbaum, Samuel Gershman
CogSci5
2018 Moral Dynamics: A Computational Model of Moral Judgment
Felix Sosa, Tomer D. Ullman, Samuel Gershman, Josh Tenenbaum, Tobias Gerstenberg
CogSci3
2018 Human-in-the-Loop Interpretability Prior
abstract
We often desire our models to be interpretable as well as accurate. Prior work on optimizing models for interpretability has relied on easy-to-quantify proxies for interpretability, such as sparsity or the number of operations required. In this work, we optimize for interpretability by directly including humans in the optimization loop. We develop an algorithm that minimizes the number of user studies to find models that are both predictive and interpretable and demonstrate our approach on several data sets. Our human subjects results show trends towards different proxy notions of interpretability on different datasets, which suggests that different proxies are preferred on different tasks.
Isaac Lage, Andrew Slavin Ross, Samuel Gershman, Been Kim, Finale Doshi-Velez
NeurIPS3
2017 Amortized Hypothesis Generation
Ishita Dasgupta 0001, Eric Schulz, Noah D. Goodman, Samuel Gershman
CogSci4
2017 Variational Particle Approximations
abstract
Approximate inference in high-dimensional, discrete probabilistic models is a central problem in computational statistics and machine learning. This paper describes discrete particle variational inference (DPVI), a new approach that combines key strengths of Monte Carlo, variational and search- based techniques. DPVI is based on a novel family of particle- based variational approximations that can be fit using simple, fast, deterministic search techniques. Like Monte Carlo, DPVI can handle multiple modes, and yields exact results in a well- defined limit. Like unstructured mean-field, DPVI is based on optimizing a lower bound on the partition function; when this quantity is not of intrinsic interest, it facilitates convergence assessment and debugging. Like both Monte Carlo and combinatorial search, DPVI can take advantage of factorization, sequential structure, and custom search operators. This paper defines DPVI particle-based approximation family and partition function lower bounds, along with the sequential DPVI and local DPVI algorithm templates for optimizing them. DPVI is illustrated and evaluated via experiments on lattice Markov Random Fields, nonparametric Bayesian mixtures and block-models, and parametric as well as non-parametric hidden Markov models. Results include applications to real-world spike-sorting and relational modeling problems, and show that DPVI can offer appealing time/accuracy trade-offs as compared to multiple alternatives.
Ardavan Saeedi, Tejas D. Kulkarni, Vikash Mansinghka 0001, Samuel Gershman
J. Mach. Learn. Res.4
2017 Dopamine, Inference, and Uncertainty
abstract
The hypothesis that the phasic dopamine response reports a reward prediction error has become deeply entrenched. However, dopamine neurons exhibit several notable deviations from this hypothesis. A coherent explanation for these deviations can be obtained by analyzing the dopamine response in terms of Bayesian reinforcement learning. The key idea is that prediction errors are modulated by probabilistic beliefs about the relationship between cues and outcomes, updated through Bayesian inference. This account can explain dopamine responses to inferred value in sensory preconditioning, the effects of cue preexposure (latent inhibition), and adaptive coding of prediction errors when rewards vary across orders of magnitude. We further postulate that orbitofrontal cortex transforms the stimulus representation through recurrent dynamics, such that a simple error-driven learning rule operating on the transformed representation can implement the Bayesian reinforcement learning update.
Samuel Gershman
Neural Comput.1
2017 Predictive representations can link model-based reinforcement learning to model-free mechanisms
abstract
Humans and animals are capable of evaluating actions by considering their long-run future rewards through a process described using model-based reinforcement learning (RL) algorithms. The mechanisms by which neural circuits perform the computations prescribed by model-based RL remain largely unknown; however, multiple lines of evidence suggest that neural circuits supporting model-based behavior are structurally homologous to and overlapping with those thought to carry out model-free temporal difference (TD) learning. Here, we lay out a family of approaches by which model-based computation may be built upon a core of TD learning. The foundation of this framework is the successor representation, a predictive state representation that, when combined with TD learning of value predictions, can produce a subset of the behaviors associated with model-based learning, while requiring less decision-time computation than dynamic programming. Using simulations, we delineate the precise behavioral capabilities enabled by evaluating actions using this approach, and compare them to those demonstrated by biological organisms. We then introduce two new algorithms that build upon the successor representation while progressively mitigating its limitations. Because this framework can account for the full range of observed putatively model-based behaviors while still utilizing a core TD framework, we suggest that it represents a neurally plausible family of mechanisms for model-based evaluation.
Evan M. Russek, Ida Momennejad, Matt M. Botvinick, Samuel Gershman, Nathaniel D. Daw
PLoS Comput. Biol.4
2016 Coalescing the Vapors of Human Experience into a Viable and Meaningful Comprehension
Tomer D. Ullman, Max H. Siegel, Josh Tenenbaum, Samuel Gershman
CogSci4
2016 Probing the Compositionality of Intuitive Functions
abstract
How do people learn about complex functional structure? Taking inspiration from other areas of cognitive science, we propose that this is accomplished by harnessing compositionality: complex structure is decomposed into simpler building blocks. We formalize this idea within the framework of Bayesian regression using a grammar over Gaussian process kernels. We show that participants prefer compositional over non-compositional function extrapolations, that samples from the human prior over functions are best described by a compositional model, and that people perceive compositional functions as more predictable than their non-compositional but otherwise similar counterparts. We argue that the compositional nature of intuitive functions is consistent with broad principles of human cognition.
Eric Schulz, Josh Tenenbaum, David Duvenaud, Maarten Speekenbrink, Samuel Gershman
NIPS5
2016 When Does Model-Based Control Pay Off?
abstract
Many accounts of decision making and reinforcement learning posit the existence of two distinct systems that control choice: a fast, automatic system and a slow, deliberative system. Recent research formalizes this distinction by mapping these systems to "model-free" and "model-based" strategies in reinforcement learning. Model-free strategies are computationally cheap, but sometimes inaccurate, because action values can be accessed by inspecting a look-up table constructed through trial-and-error. In contrast, model-based strategies compute action values through planning in a causal model of the environment, which is more accurate but also more cognitively demanding. It is assumed that this trade-off between accuracy and computational demand plays an important role in the arbitration between the two strategies, but we show that the hallmark task for dissociating model-free and model-based strategies, as well as several related variants, do not embody such a trade-off. We describe five factors that reduce the effectiveness of the model-based strategy on these tasks by reducing its accuracy in estimating reward outcomes and decreasing the importance of its choices. Based on these observations, we describe a version of the task that formally and empirically obtains an accuracy-demand trade-off between model-free and model-based strategies. Moreover, we show that human participants spontaneously increase their reliance on model-based control on this task, compared to the original paradigm. Our novel task and our computational analyses may prove important in subsequent empirical investigations of how humans balance accuracy and demand.
Wouter Kool 0002, Fiery Cushman, Samuel Gershman
PLoS Comput. Biol.3
2015 Phrase similarity in humans and machines
Samuel Gershman, Josh Tenenbaum
CogSci1
2015 Assessing the Perceived Predictability of Functions
Eric Schulz, Josh Tenenbaum, David N. Reshef, Maarten Speekenbrink, Samuel Gershman
CogSci5
2015 Distance Dependent Infinite Latent Feature Models
abstract
Latent feature models are widely used to decompose data into a small number of components. Bayesian nonparametric variants of these models, which use the Indian buffet process (IBP) as a prior over latent features, allow the number of features to be determined from the data. We present a generalization of the IBP, the distance dependent Indian buffet process (dd-IBP), for modeling non-exchangeable data. It relies on distances defined between data points, biasing nearby data to share more features. The choice of distance measure allows for many kinds of dependencies, including temporal and spatial. Further, the original IBP is a special case of the dd-IBP. We develop the dd-IBP and theoretically characterize its feature-sharing properties. We derive a Markov chain Monte Carlo sampler for a linear Gaussian model with a dd-IBP prior and study its performance on real-world non-exchangeable data.
Samuel Gershman, Peter I. Frazier, David M. Blei
IEEE Trans. Pattern Anal. Mach. Intell.1
2015 A Unifying Probabilistic View of Associative Learning
abstract
Two important ideas about associative learning have emerged in recent decades: (1) Animals are Bayesian learners, tracking their uncertainty about associations; and (2) animals acquire long-term reward predictions through reinforcement learning. Both of these ideas are normative, in the sense that they are derived from rational design principles. They are also descriptive, capturing a wide range of empirical phenomena that troubled earlier theories. This article describes a unifying framework encompassing Bayesian and reinforcement learning theories of associative learning. Each perspective captures a different aspect of associative learning, and their synthesis offers insight into phenomena that neither perspective can explain on its own.
Samuel Gershman
PLoS Comput. Biol.1
2014 Amortized Inference in Probabilistic Reasoning
Samuel Gershman, Noah D. Goodman
CogSci1
2014 Information Selection in Noisy Environments with Large Action Spaces
Pedro Tsividis, Samuel Gershman, Josh Tenenbaum, Laura Schulz
CogSci2
2014 Design Principles of the Hippocampal Cognitive Map
Kimberly L. Stachenfeld, Matt M. Botvinick, Samuel Gershman
NIPS3
2014 Dopamine Ramps Are a Consequence of Reward Prediction Errors
abstract
Temporal difference learning models of dopamine assert that phasic levels of dopamine encode a reward prediction error. However, this hypothesis has been challenged by recent observations of gradually ramping stratal dopamine levels as a goal is approached. This note describes conditions under which temporal difference learning models predict dopamine ramping. The key idea is representational: a quadratic transformation of proximity to the goal implies approximately linear ramping, as observed experimentally.
Samuel Gershman
Neural Comput.1
2014 Statistical Computations Underlying the Dynamics of Memory Updating
abstract
Psychophysical and neurophysiological studies have suggested that memory is not simply a carbon copy of our experience: Memories are modified or new memories are formed depending on the dynamic structure of our experience, and specifically, on how gradually or abruptly the world changes. We present a statistical theory of memory formation in a dynamic environment, based on a nonparametric generalization of the switching Kalman filter. We show that this theory can qualitatively account for several psychophysical and neural phenomena, and present results of a new visual memory experiment aimed at testing the theory directly. Our experimental findings suggest that humans can use temporal discontinuities in the structure of the environment to determine when to form new memory traces. The statistical perspective we offer provides a coherent account of the conditions under which new experience is integrated into an old memory versus forming a new memory, and shows that memory formation depends on inferences about the underlying structure of our experience.
Samuel Gershman, Angela Radulescu, Kenneth A. Norman, Yael Niv
PLoS Comput. Biol.1
2013 Bayesian vector analysis and the perception of hierarchical motion
Samuel Gershman, Frank Jäkel, Josh Tenenbaum
CogSci1
2013 Structured cognitive representations and complex inference in neural systems
Samuel Gershman, Josh Tenenbaum, Alexandre Pouget, Matt M. Botvinick, Peter Dayan
CogSci1
2012 Nonparametric variational inference
Samuel Gershman, Matthew Hoffman 0001, David M. Blei
ICML1
2012 The Successor Representation and Temporal Context
abstract
The successor representation was introduced into reinforcement learning by Dayan ( 1993 ) as a means of facilitating generalization between states with similar successors. Although reinforcement learning in general has been used extensively as a model of psychological and neural processes, the psychological validity of the successor representation has yet to be explored. An interesting possibility is that the successor representation can be used not only for reinforcement learning but for episodic learning as well. Our main contribution is to show that a variant of the temporal context model (TCM; Howard & Kahana, 2002 ), an influential model of episodic memory, can be understood as directly estimating the successor representation using the temporal difference learning algorithm (Sutton & Barto, 1998 ). This insight leads to a generalization of TCM and new experimental predictions. In addition to casting a new normative light on TCM, this equivalence suggests a previously unexplored point of contact between different learning systems.
Samuel Gershman, Christopher D. Moore, Michael T. Todd, Kenneth A. Norman, Per B. Sederberg
Neural Comput.1
2012 Multistability and Perceptual Inference
abstract
Ambiguous images present a challenge to the visual system: How can uncertainty about the causes of visual inputs be represented when there are multiple equally plausible causes? A Bayesian ideal observer should represent uncertainty in the form of a posterior probability distribution over causes. However, in many real-world situations, computing this distribution is intractable and requires some form of approximation. We argue that the visual system approximates the posterior over underlying causes with a set of samples and that this approximation strategy produces perceptual multistability--stochastic alternation between percepts in consciousness. Under our analysis, multistability arises from a dynamic sample-generating process that explores the posterior through stochastic diffusion, implementing a rational form of approximate Bayesian inference known as Markov chain Monte Carlo (MCMC). We examine in detail the most extensively studied form of multistability, binocular rivalry, showing how a variety of experimental phenomena--gamma-like stochastic switching, patchy percepts, fusion, and traveling waves--can be understood in terms of MCMC sampling over simple graphical models of the underlying perceptual tasks. We conjecture that the stochastic nature of spiking neurons may lend itself to implementing sample-based posterior approximations in the brain.
Samuel Gershman, Ed Vul, Josh Tenenbaum
Neural Comput.1
2011 Computational, Neuroscientific, and Lifespan Perspectives on the Exploration-Exploitation Dilemma
A. Ross Otto, W. Bradley Knox, Bradley C. Love, Samuel Gershman, Yael Niv, Darrell A. Worthy, W. Todd Maddox, Jared M. Hotaling, Jerome R. Busemeyer, Richard M. Shiffrin
CogSci4
2010 The Neural Costs of Optimal Control
abstract
Optimal control entails combining probabilities and utilities. However, for most practical problems probability densities can be represented only approximately. Choosing an approximation requires balancing the benefits of an accurate approximation against the costs of computing it. We propose a variational framework for achieving this balance and apply it to the problem of how a population code should optimally represent a distribution under resource constraints. The essence of our analysis is the conjecture that population codes are organized to maximize a lower bound on the log expected utility. This theory can account for a plethora of experimental data, including the reward-modulation of sensory receptive fields.
Samuel Gershman
NIPS1
2009 Perceptual Multistability as Markov Chain Monte Carlo Inference
abstract
While many perceptual and cognitive phenomena are well described in terms of Bayesian inference, the necessary computations are intractable at the scale of real-world tasks, and it remains unclear how the human mind approximates Bayesian inference algorithmically. We explore the proposal that for some tasks, humans use a form of Markov Chain Monte Carlo to approximate the posterior distribution over hidden variables. As a case study, we show how several phenomena of perceptual multistability can be explained as MCMC inference in simple graphical models for low-level vision.
Samuel Gershman, Ed Vul, Josh Tenenbaum
NIPS1
2009 A Bayesian Analysis of Dynamics in Free Recall
abstract
We develop a probabilistic model of human memory performance in free recall experiments. In these experiments, a subject first studies a list of words and then tries to recall them. To model these data, we draw on both previous psychological research and statistical topic models of text documents. We assume that memories are formed by assimilating the semantic meaning of studied words (represented as a distribution over topics) into a slowly changing latent context (represented in the same space). During recall, this context is reinstated and used as a cue for retrieving studied words. By conceptualizing memory retrieval as a dynamic latent variable model, we are able to use Bayesian inference to represent uncertainty and reason about the cognitive processes underlying memory. We present a particle filter algorithm for performing approximate posterior inference, and evaluate our model on the prediction of recalled words in experimental data. By specifying the model hierarchically, we are also able to capture inter-subject variability.
Richard Socher, Samuel Gershman, Adler J. Perotte, Per B. Sederberg, David M. Blei, Kenneth A. Norman
NIPS2