Nathaniel D. Daw

dblp:38/929 · DBLP profile ↗
← Back
44ranked-venue papers
8as first author
14since 2021 · last 2025
0000-0001-5029-1430ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 32 · 7 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 26 · 10 since 2021Systems, architecture and hardware · 1 · 1 first-author
YearPublicationVenuePosition
2025 How goals affect information seeking
Gili Karni, Nathaniel D. Daw, Yael Niv
CogSci2
2025 Discovering Symbolic Cognitive Models from Human and Animal Behavior
abstract
Symbolic models play a key role in cognitive science, expressing computationally precise hypotheses about how the brain implements a cognitive process. Identifying an appropriate model typically requires a great deal of effort and ingenuity on the part of a human scientist. Here, we adapt FunSearch (Romera-Paredes et al. 2024), a recently developed tool that uses Large Language Models (LLMs) in an evolutionary algorithm, to automatically discover symbolic cognitive models that accurately capture human and animal behavior. We consider datasets from three species performing a classic reward-learning task that has been the focus of substantial modeling effort, and find that the discovered programs outperform state-of-the-art cognitive models for each. The discovered programs can readily be interpreted as hypotheses about human and animal cognition, instantiating interpretable symbolic learning and decision-making algorithms. Broadly, these results demonstrate the viability of using LLM-powered program synthesis to propose novel scientific hypotheses regarding mechanisms of human and animal cognition.
Pablo Samuel Castro, Nenad Tomasev, Ankit Anand, Navodita Sharma, Rishika Mohanta, Aparna Dev, Kuba Perlin, Siddhant Jain, Kyle Levin, Noémi Élteto, Will Dabney, Alexander Novikov 0001, Glenn C. Turner, Maria K. Eckstein, Nathaniel D. Daw, Kevin J. Miller, Kimberly L. Stachenfeld
ICML15
2025 Trial-by-trial learning of successor representations in human behavior
abstract
Decisions in humans and other organisms depend, in part, on learning and using models that capture the statistical structure of the world, including the long-run expected outcomes of our actions. One prominent approach to forecasting such long-run outcomes is the successor representation (SR), which predicts future states aggregated over multiple timesteps. Although much behavioral and neural evidence suggests that people and animals use such a representation, it remains unknown how they acquire it. It has frequently been assumed to be learned by temporal difference bootstrapping (SR-TD(0)), but this assumption has largely not been empirically tested or compared to alternatives including eligibility traces (SR-TD([Formula: see text])). Here we address this gap by leveraging trial-by-trial reaction times in graph sequence learning tasks, which are favorable for studying learning dynamics because the long horizons in these studies differentiate the transient update dynamics of different learning rules. We examined the behavior of SR-TD(λ) on a probabilistic graph learning task alongside a number of alternatives, and found that behavior was best explained by a hybrid model which learned via SR-TD(λ) alongside an additional predictive model of recency. The relatively large λ we estimate indicates a predominant role of eligibility trace mechanisms over the bootstrap-based chaining typically assumed. Our results provide insight into how humans learn predictive representations, and demonstrate that people simultaneously learn the SR alongside lower-order predictions.
Ari E. Kahn, Danielle S. Bassett, Nathaniel D. Daw
PLoS Comput. Biol.3
2024 Anxiety symptoms of major depression associated with increased willingness to exert cognitive, but not physical effort
Laura Bustamante, Deanna M. Barch, Johanne Solis, Temitope Oshinowo, Ivan Grahek, Anna B. Konova, Nathaniel D. Daw, Jonathan D. Cohen 0003
CogSci7
2024 Program-Based Strategy Induction for Reinforcement Learning
Carlos G. Correa, Thomas L. Griffiths 0001, Nathaniel D. Daw
CogSci3
2024 Temporally extended decision-making through episodic sampling
Corey Yishan Zhou, Deborah Talmi, Nathaniel D. Daw, Marcelo G. Mattar
CogSci3
2024 The role of training variability for model-based and model-free learning of an arbitrary visuomotor mapping
abstract
A fundamental feature of the human brain is its capacity to learn novel motor skills. This capacity requires the formation of vastly different visuomotor mappings. Using a grid navigation task, we investigated whether training variability would enhance the flexible use of a visuomotor mapping (key-to-direction rule), leading to better generalization performance. Experiments 1 and 2 show that participants trained to move between multiple start-target pairs exhibited greater generalization to both distal and proximal targets compared to participants trained to move between a single pair. This finding suggests that limited variability can impair decisions even in simple tasks without planning. In addition, during the training phase, participants exposed to higher variability were more inclined to choose options that, counterintuitively, moved the cursor away from the target while minimizing its actual distance under the constrained mapping, suggesting a greater engagement in model-based computations. In Experiments 3 and 4, we showed that the limited generalization performance in participants trained with a single pair can be enhanced by a short period of variability introduced early in learning or by incorporating stochasticity into the visuomotor mapping. Our computational modeling analyses revealed that a hybrid model between model-free and model-based computations with different mixing weights for the training and generalization phases, best described participants' data. Importantly, the differences in the model-based weights between our experimental groups, paralleled the behavioral findings during training and generalization. Taken together, our results suggest that training variability enables the flexible use of the visuomotor mapping, potentially by preventing the consolidation of habits due to the continuous demand to change responses.
Carlos A. Velázquez-Vargas, Nathaniel D. Daw, Jordan A. Taylor
PLoS Comput. Biol.2
2023 Predictive and Interpretable: Combining Artificial Neural Networks and Classic Cognitive Models to Understand Human Learning and Decision Making
Maria K. Eckstein, Christopher Summerfield, Nathaniel D. Daw, Kevin J. Miller
CogSci3
2023 Would I have gotten that reward? Long-term credit assignment by counterfactual contribution analysis
abstract
To make reinforcement learning more sample efficient, we need better credit assignment methods that measure an action’s influence on future rewards. Building upon Hindsight Credit Assignment (HCA), we introduce Counterfactual Contribution Analysis (COCOA), a new family of model-based credit assignment algorithms. Our algorithms achieve precise credit assignment by measuring the contribution of actions upon obtaining subsequent rewards, by quantifying a counterfactual query: ‘Would the agent still have reached this reward if it had taken another action?’. We show that measuring contributions w.r.t. rewarding _states_, as is done in HCA, results in spurious estimates of contributions, causing HCA to degrade towards the high-variance REINFORCE estimator in many relevant environments. Instead, we measure contributions w.r.t. rewards or learned representations of the rewarding objects, resulting in gradient estimates with lower variance. We run experiments on a suite of problems specifically designed to evaluate long-term credit assignment capabilities. By using dynamic programming, we measure ground-truth policy gradients and show that the improved performance of our new model-based credit assignment methods is due to lower bias and variance compared to HCA and common baselines. Our results demonstrate how modeling action contributions towards rewarding outcomes can be leveraged for credit assignment, opening a new path towards sample-efficient reinforcement learning.
Alexander Meulemans, Simon Schug, Seijin Kobayashi, Nathaniel D. Daw, Greg Wayne
NeurIPS4
2023 Humans decompose tasks by trading off utility and computational cost
abstract
Human behavior emerges from planning over elaborate decompositions of tasks into goals, subgoals, and low-level actions. How are these decompositions created and used? Here, we propose and evaluate a normative framework for task decomposition based on the simple idea that people decompose tasks to reduce the overall cost of planning while maintaining task performance. Analyzing 11,117 distinct graph-structured planning tasks, we find that our framework justifies several existing heuristics for task decomposition and makes predictions that can be distinguished from two alternative normative accounts. We report a behavioral study of task decomposition (N = 806) that uses 30 randomly sampled graphs, a larger and more diverse set than that of any previous behavioral study on this topic. We find that human responses are more consistent with our framework for task decomposition than alternative normative accounts and are most consistent with a heuristic-betweenness centrality-that is justified by our approach. Taken together, our results suggest the computational cost of planning is a key principle guiding the intelligent structuring of goal-directed behavior.
Carlos G. Correa, Mark K. Ho, Frederick Callaway, Nathaniel D. Daw, Thomas L. Griffiths 0001
PLoS Comput. Biol.4
2023 Disentangling Abstraction from Statistical Pattern Matching in Human and Machine Learning
abstract
The ability to acquire abstract knowledge is a hallmark of human intelligence and is believed by many to be one of the core differences between humans and neural network models. Agents can be endowed with an inductive bias towards abstraction through meta-learning, where they are trained on a distribution of tasks that share some abstract structure that can be learned and applied. However, because neural networks are hard to interpret, it can be difficult to tell whether agents have learned the underlying abstraction, or alternatively statistical patterns that are characteristic of that abstraction. In this work, we compare the performance of humans and agents in a meta-reinforcement learning paradigm in which tasks are generated from abstract rules. We define a novel methodology for building "task metamers" that closely match the statistics of the abstract tasks but use a different underlying generative process, and evaluate performance on both abstract and metamer tasks. We find that humans perform better at abstract tasks than metamer tasks whereas common neural network architectures typically perform worse on the abstract tasks than the matched metamers. This work provides a foundation for characterizing differences between humans and machine learning that can be used in future work towards developing machines with more human-like behavior.
Sreejan Kumar, Ishita Dasgupta 0001, Nathaniel D. Daw, Jonathan D. Cohen 0003, Thomas L. Griffiths 0001
PLoS Comput. Biol.3
2022 Leveraging psychometrics of rational inattention to estimate individual differences in the capacity for cognitive control
Ham Huang, Ivan Grahek, Laura Bustamante, Nathaniel D. Daw, Andrew Caplin, Sebastian Musslick
CogSci4
2022 Using natural language and program abstractions to instill human inductive biases in machines
abstract
Strong inductive biases give humans the ability to quickly learn to perform a variety of tasks. Although meta-learning is a method to endow neural networks with useful inductive biases, agents trained by meta-learning may sometimes acquire very different strategies from humans. We show that co-training these agents on predicting representations from natural language task descriptions and programs induced to generate such tasks guides them toward more human-like inductive biases. Human-generated language descriptions and program induction models that add new learned primitives both contain abstract concepts that can compress description length. Co-training on these representations result in more human-like behavior in downstream meta-reinforcement learning agents than less abstract controls (synthetic language descriptions, program induction without learned primitives), suggesting that the abstraction supported by these representations is key.
Sreejan Kumar, Carlos G. Correa, Ishita Dasgupta 0001, Raja Marjieh, Michael Y. Hu, Robert D. Hawkins, Jonathan D. Cohen 0003, Nathaniel D. Daw, Karthik Narasimhan, Thomas L. Griffiths 0001
NeurIPS8
2021 Meta-Learning of Structured Task Distributions in Humans and Machines
Sreejan Kumar, Ishita Dasgupta 0001, Jonathan D. Cohen 0003, Nathaniel D. Daw, Thomas L. Griffiths 0001
ICLR4
2020 A simple model for learning in volatile environments
abstract
Sound principles of statistical inference dictate that uncertainty shapes learning. In this work, we revisit the question of learning in volatile environments, in which both the first and second-order statistics of observations dynamically evolve over time. We propose a new model, the volatile Kalman filter (VKF), which is based on a tractable state-space model of uncertainty and extends the Kalman filter algorithm to volatile environments. The proposed model is algorithmically simple and encompasses the Kalman filter as a special case. Specifically, in addition to the error-correcting rule of Kalman filter for learning observations, the VKF learns volatility according to a second error-correcting rule. These dual updates echo and contextualize classical psychological models of learning, in particular hybrid accounts of Pearce-Hall and Rescorla-Wagner. At the computational level, compared with existing models, the VKF gives up some flexibility in the generative model to enable a more faithful approximation to exact inference. When fit to empirical data, the VKF is better behaved than alternatives and better captures human choice data in two independent datasets of probabilistic learning tasks. The proposed model provides a coherent account of learning in stable or volatile environments and has implications for decision neuroscience research.
Payam Piray, Nathaniel D. Daw
PLoS Comput. Biol.2
2019 Hierarchical Bayesian inference for concurrent model fitting and comparison for group studies
abstract
Computational modeling plays an important role in modern neuroscience research. Much previous research has relied on statistical methods, separately, to address two problems that are actually interdependent. First, given a particular computational model, Bayesian hierarchical techniques have been used to estimate individual variation in parameters over a population of subjects, leveraging their population-level distributions. Second, candidate models are themselves compared, and individual variation in the expressed model estimated, according to the fits of the models to each subject. The interdependence between these two problems arises because the relevant population for estimating parameters of a model depends on which other subjects express the model. Here, we propose a hierarchical Bayesian inference (HBI) framework for concurrent model comparison, parameter estimation and inference at the population level, combining previous approaches. We show that this framework has important advantages for both parameter estimation and model comparison theoretically and experimentally. The parameters estimated by the HBI show smaller errors compared to other methods. Model comparison by HBI is robust against outliers and is not biased towards overly simplistic models. Furthermore, the fully Bayesian approach of our theory enables researchers to make inference on group-level parameters by performing HBI t-test.
Payam Piray, Amir Dezfouli, Tom Heskes, Michael J. Frank, Nathaniel D. Daw
PLoS Comput. Biol.5
2018 Novel methods for measuring the cost of cognitive control in a patch foraging task and a demand selection task with Stroop
Laura Bustamante, Augustus Baker, Allison Burton, Amitai Shenhav, Chloe Hoeber, Nathaniel D. Daw, Jonathan D. Cohen 0003
CogSci6
2017 Pre-term infants exhibit impaired prediction and learning in Audio-Visual association paradigm
Sagi Jaffe-Dax, Alex M. Boldin, Nathaniel D. Daw, Lauren L. Emberson
CogSci3
2017 Mechanisms of overharvesting in patch foraging
Gary Kane, Aaron M. Bornstein, Amitai Shenhav, Nathaniel D. Daw, Jonathan D. Cohen 0003
CogSci5
2017 Intolerance to uncertainty is associated with diminished exploration
Jennifer Lenow, Nathaniel D. Daw, Elizabeth A. Phelps
CogSci2
2017 Suboptimal Criterion Learning in Static and Dynamic Environments
abstract
Humans often make decisions based on uncertain sensory information. Signal detection theory (SDT) describes detection and discrimination decisions as a comparison of stimulus "strength" to a fixed decision criterion. However, recent research suggests that current responses depend on the recent history of stimuli and previous responses, suggesting that the decision criterion is updated trial-by-trial. The mechanisms underpinning criterion setting remain unknown. Here, we examine how observers learn to set a decision criterion in an orientation-discrimination task under both static and dynamic conditions. To investigate mechanisms underlying trial-by-trial criterion placement, we introduce a novel task in which participants explicitly set the criterion, and compare it to a more traditional discrimination task, allowing us to model this explicit indication of criterion dynamics. In each task, stimuli were ellipses with principal orientations drawn from two categories: Gaussian distributions with different means and equal variance. In the covert-criterion task, observers categorized a displayed ellipse. In the overt-criterion task, observers adjusted the orientation of a line that served as the discrimination criterion for a subsequently presented ellipse. We compared performance to the ideal Bayesian learner and several suboptimal models that varied in both computational and memory demands. Under static and dynamic conditions, we found that, in both tasks, observers used suboptimal learning rules. In most conditions, a model in which the recent history of past samples determines a belief about category means fit the data best for most observers and on average. Our results reveal dynamic adjustment of discrimination criterion, even after prolonged training, and indicate how decision criteria are updated over time.
Elyse H. Norton, Stephen M. Fleming, Nathaniel D. Daw, Michael S. Landy
PLoS Comput. Biol.3
2017 Predictive representations can link model-based reinforcement learning to model-free mechanisms
abstract
Humans and animals are capable of evaluating actions by considering their long-run future rewards through a process described using model-based reinforcement learning (RL) algorithms. The mechanisms by which neural circuits perform the computations prescribed by model-based RL remain largely unknown; however, multiple lines of evidence suggest that neural circuits supporting model-based behavior are structurally homologous to and overlapping with those thought to carry out model-free temporal difference (TD) learning. Here, we lay out a family of approaches by which model-based computation may be built upon a core of TD learning. The foundation of this framework is the successor representation, a predictive state representation that, when combined with TD learning of value predictions, can produce a subset of the behaviors associated with model-based learning, while requiring less decision-time computation than dynamic programming. Using simulations, we delineate the precise behavioral capabilities enabled by evaluating actions using this approach, and compare them to those demonstrated by biological organisms. We then introduce two new algorithms that build upon the successor representation while progressively mitigating its limitations. Because this framework can account for the full range of observed putatively model-based behaviors while still utilizing a core TD framework, we suggest that it represents a neurally plausible family of mechanisms for model-based evaluation.
Evan M. Russek, Ida Momennejad, Matt M. Botvinick, Samuel Gershman, Nathaniel D. Daw
PLoS Comput. Biol.5
2016 Boredom, Information-Seeking and Exploration
Andra Geana, Nathaniel D. Daw, Jonathan D. Cohen 0003
CogSci3
2016 Information-Seeking, Learning and the Marginal Value Theorem: A Normative Approach to Adaptive Exploration
Andra Geana, Nathaniel D. Daw, Jonathan D. Cohen 0003
CogSci3
2014 Cognitive Control Mode Predicts Behavioral Expression of Model-Based Reinforcement-Learning
A. Ross Otto, Anya Skatova, Seth Madlon-Kay, Nathaniel D. Daw
CogSci4
2013 Cortical and Hippocampal Correlates of Deliberation during Model-Based Decisions for Rewards in Humans
abstract
How do we use our memories of the past to guide decisions we've never had to make before? Although extensive work describes how the brain learns to repeat rewarded actions, decisions can also be influenced by associations between stimuli or events not directly involving reward - such as when planning routes using a cognitive map or chess moves using predicted countermoves - and these sorts of associations are critical when deciding among novel options. This process is known as model-based decision making. While the learning of environmental relations that might support model-based decisions is well studied, and separately this sort of information has been inferred to impact decisions, there is little evidence concerning the full cycle by which such associations are acquired and drive choices. Of particular interest is whether decisions are directly supported by the same mnemonic systems characterized for relational learning more generally, or instead rely on other, specialized representations. Here, building on our previous work, which isolated dual representations underlying sequential predictive learning, we directly demonstrate that one such representation, encoded by the hippocampal memory system and adjacent cortical structures, supports goal-directed decisions. Using interleaved learning and decision tasks, we monitor predictive learning directly and also trace its influence on decisions for reward. We quantitatively compare the learning processes underlying multiple behavioral and fMRI observables using computational model fits. Across both tasks, a quantitatively consistent learning process explains reaction times, choices, and both expectation- and surprise-related neural activity. The same hippocampal and ventral stream regions engaged in anticipating stimuli during learning are also engaged in proportion to the difficulty of decisions. These results support a role for predictive associations learned by the hippocampal memory system to be recalled during choice formation.
Aaron M. Bornstein, Nathaniel D. Daw
PLoS Comput. Biol.2
2013 Testing Whether Humans Have an Accurate Model of Their Own Motor Uncertainty in a Speeded Reaching Task
abstract
In many motor tasks, optimal performance presupposes that human movement planning is based on an accurate internal model of the subject's own motor error. We developed a motor choice task that allowed us to test whether the internal model implicit in a subject's choices differed from the actual in isotropy (elongation) and variance. Subjects were first trained to hit a circular target on a touch screen within a time limit. After training, subjects were repeatedly shown pairs of targets differing in size and shape and asked to choose the target that was easier to hit. On each trial they simply chose a target - they did not attempt to hit the chosen target. For each subject, we tested whether the internal model implicit in her target choices was consistent with her true error distribution in isotropy and variance. For all subjects, movement end points were anisotropic, distributed as vertically elongated bivariate Gaussians. However, in choosing targets, almost all subjects effectively assumed an isotropic distribution rather than their actual anisotropic distribution. Roughly half of the subjects chose as though they correctly estimated their own variance and the other half effectively assumed a variance that was more than four times larger than the actual, essentially basing their choices merely on the areas of the targets. The task and analyses we developed allowed us to characterize the internal model of motor error implicit in how humans plan reaching movements. In this task, human movement planning - even after extensive training - is based on an internal model of human motor error that includes substantial and qualitative inaccuracies.
Hang Zhang 0028, Nathaniel D. Daw, Laurence T. Maloney
PLoS Comput. Biol.2
2012 Neural Computations Supporting Cognition: Rumelhart Prize Symposium in Honor of Peter Dayan
Kenji Doya, John P. O'Doherty, Alexandre Pouget, Peter Bossaerts, Nathaniel D. Daw, Yael Niv
CogSci5
2011 Moving Beyond Where and What to How: Using Models and fMRI to Understand Brain-Behavior Relations
Bradley C. Love, John P. Spencer, Nathaniel D. Daw, John P. O'Doherty
CogSci3
2011 Environmental statistics and the trade-off between model-based and TD learning in humans
abstract
There is much evidence that humans and other animals utilize a combination of model-based and model-free RL methods. Although it has been proposed that these systems may dominate according to their relative statistical efficiency in different circumstances, there is little specific evidence -- especially in humans -- as to the details of this trade-off. Accordingly, we examine the relative performance of different RL approaches under situations in which the statistics of reward are differentially noisy and volatile. Using theory and simulation, we show that model-free TD learning is relatively most disadvantaged in cases of high volatility and low noise. We present data from a decision-making experiment manipulating these parameters, showing that humans shift learning strategies in accord with these predictions. The statistical circumstances favoring model-based RL are also those that promote a high learning rate, which helps explain why, in psychology, the distinction between these strategies is traditionally conceived in terms of rule-based vs. incremental learning.
Dylan A. Simon, Nathaniel D. Daw
NIPS2
2011 Grid Cells, Place Cells, and Geodesic Generalization for Spatial Reinforcement Learning
abstract
Reinforcement learning (RL) provides an influential characterization of the brain's mechanisms for learning to make advantageous choices. An important problem, though, is how complex tasks can be represented in a way that enables efficient learning. We consider this problem through the lens of spatial navigation, examining how two of the brain's location representations--hippocampal place cells and entorhinal grid cells--are adapted to serve as basis functions for approximating value over space for RL. Although much previous work has focused on these systems' roles in combining upstream sensory cues to track location, revisiting these representations with a focus on how they support this downstream decision function offers complementary insights into their characteristics. Rather than localization, the key problem in learning is generalization between past and present situations, which may not match perfectly. Accordingly, although neural populations collectively offer a precise representation of position, our simulations of navigational tasks verify the suggestion that RL gains efficiency from the more diffuse tuning of individual neurons, which allows learning about rewards to generalize over longer distances given fewer training experiences. However, work on generalization in RL suggests the underlying representation should respect the environment's layout. In particular, although it is often assumed that neurons track location in Euclidean coordinates (that a place cell's activity declines "as the crow flies" away from its peak), the relevant metric for value is geodesic: the distance along a path, around any obstacles. We formalize this intuition and present simulations showing how Euclidean, but not geodesic, representations can interfere with RL by generalizing inappropriately across barriers. Our proposal that place and grid responses should be modulated by geodesic distances suggests novel predictions about how obstacles should affect spatial firing fields, which provides a new viewpoint on data concerning both spatial codes.
Nicholas J. Gustafson, Nathaniel D. Daw
PLoS Comput. Biol.2
2007 The rat as particle filter
abstract
Although theorists have interpreted classical conditioning as a laboratory model of Bayesian belief updating, a recent reanalysis showed that the key features that theoretical models capture about learning are artifacts of averaging over subjects. Rather than learning smoothly to asymptote (reflecting, according to Bayesian models, the gradual tradeoff from prior to posterior as data accumulate), subjects learn suddenly and their predictions fluctuate perpetually. We suggest that abrupt and unstable learning can be modeled by assuming subjects are conducting in- ference using sequential Monte Carlo sampling with a small number of samples — one, in our simulations. Ensemble behavior resembles exact Bayesian models since, as in particle filters, it averages over many samples. Further, the model is capable of exhibiting sophisticated behaviors like retrospective revaluation at the ensemble level, even given minimally sophisticated individuals that do not track uncertainty in their beliefs over trials.
Nathaniel D. Daw, Aaron C. Courville
NIPS1
2006 Representation and Timing in Theories of the Dopamine System
abstract
Although the responses of dopamine neurons in the primate midbrain are well characterized as carrying a temporal difference (TD) error signal for reward prediction, existing theories do not offer a credible account of how the brain keeps track of past sensory events that may be relevant to predicting future reward. Empirically, these shortcomings of previous theories are particularly evident in their account of experiments in which animals were exposed to variation in the timing of events. The original theories mispredicted the results of such experiments due to their use of a representational device called a tapped delay line. Here we propose that a richer understanding of history representation and a better account of these experiments can be given by considering TD algorithms for a formal setting that incorporates two features not originally considered in theories of the dopaminergic response: partial observability (a distinction between the animal's sensory experience and the true underlying state of the world) and semi-Markov dynamics (an explicit account of variation in the intervals between events). The new theory situates the dopaminergic system in a richer functional and anatomical context, since it assumes (in accord with recent computational theories of cortex) that problems of partial observability and stimulus history are solved in sensory cortex using statistical modeling and inference and that the TD system predicts reward using the results of this inference rather than raw sensory data. It also accounts for a range of experimental data, including the experiments involving programmed temporal variability and other previously unmodeled dopaminergic response phenomena, which we suggest are related to subjective noise in animals' interval timing. Finally, it offers new experimental predictions and a rich theoretical framework for designing future experiments.
Nathaniel D. Daw, Aaron C. Courville, David S. Touretzky
Neural Comput.1
2006 The misbehavior of value and the discipline of the will
Peter Dayan, Yael Niv, Ben Seymour, Nathaniel D. Daw
Neural Networks4
2005 How fast to work: Response vigor, motivation and tonic dopamine
abstract
Reinforcement learning models have long promised to unify computa- tional, psychological and neural accounts of appetitively conditioned be- havior. However, the bulk of data on animal conditioning comes from free-operant experiments measuring how fast animals will work for rein- forcement. Existing reinforcement learning (RL) models are silent about these tasks, because they lack any notion of vigor. They thus fail to ad- dress the simple observation that hungrier animals will work harder for food, as well as stranger facts such as their sometimes greater produc- tivity even when working for irrelevant outcomes such as water. Here, we develop an RL framework for free-operant behavior, suggesting that subjects choose how vigorously to perform selected actions by optimally balancing the costs and benefits of quick responding. Motivational states such as hunger shift these factors, skewing the tradeoff. This accounts normatively for the effects of motivation on response rates, as well as many other classic findings. Finally, we suggest that tonic levels of dopamine may be involved in the computation linking motivational state to optimal responding, thereby explaining the complex vigor-related ef- fects of pharmacological manipulation of dopamine.
Yael Niv, Nathaniel D. Daw, Peter Dayan
NIPS2
2004 Similarity and Discrimination in Classical Conditioning: A Latent Variable Account
abstract
We propose a probabilistic, generative account of configural learning phenomena in classical conditioning. Configural learning experiments probe how animals discriminate and generalize between patterns of si- multaneously presented stimuli (such as tones and lights) that are dif- ferentially predictive of reinforcement. Previous models of these issues have been successful more on a phenomenological than an explanatory level: they reproduce experimental findings but, lacking formal founda- tions, provide scant basis for understanding why animals behave as they do. We present a theory that clarifies seemingly arbitrary aspects of pre- vious models while also capturing a broader set of data. Key patterns of data, e.g. concerning animals' readiness to distinguish patterns with varying degrees of overlap, are shown to follow from statistical inference.
Aaron C. Courville, Nathaniel D. Daw, David S. Touretzky
NIPS2
2003 Model Uncertainty in Classical Conditioning
abstract
We develop a framework based on Bayesian model averaging to explain how animals cope with uncertainty about contingencies in classical con- ditioning experiments. Traditional accounts of conditioning fit parame- ters within a fixed generative model of reinforcer delivery; uncertainty over the model structure is not considered. We apply the theory to ex- plain the puzzling relationship between second-order conditioning and conditioned inhibition, two similar conditioning regimes that nonethe- less result in strongly divergent behavioral outcomes. According to the theory, second-order conditioning results when limited experience leads animals to prefer a simpler world model that produces spurious corre- lations; conditioned inhibition results when a more complex model is justified by additional experience.
Aaron C. Courville, Nathaniel D. Daw, Geoffrey J. Gordon, David S. Touretzky
NIPS2
2002 Timing and Partial Observability in the Dopamine System
abstract
According to a series of influential models, dopamine (DA) neurons sig- nal reward prediction error using a temporal-difference (TD) algorithm. We address a problem not convincingly solved in these accounts: how to maintain a representation of cues that predict delayed consequences. Our new model uses a TD rule grounded in partially observable semi-Markov processes, a formalism that captures two largely neglected features of DA experiments: hidden state and temporal variability. Previous models pre- dicted rewards using a tapped delay line representation of sensory inputs; we replace this with a more active process of inference about the under- lying state of the world. The DA system can then learn to map these inferred states to reward predictions using TD. The new model can ex- plain previously vexing data on the responses of DA neurons in the face of temporal variability. By combining statistical model-based learning with a physiologically grounded TD theory, it also brings into contact with physiology some insights about behavior that had previously been confined to more abstract psychological models.
Nathaniel D. Daw, Aaron C. Courville, David S. Touretzky
NIPS1
2002 Long-Term Reward Prediction in TD Models of the Dopamine System
abstract
This article addresses the relationship between long-term reward predictions and slow-timescale neural activity in temporal difference (TD) models of the dopamine system. Such models attempt to explain how the activity of dopamine (DA) neurons relates to errors in the prediction of future rewards. Previous models have been mostly restricted to short-term predictions of rewards expected during a single, somewhat artificially defined trial. Also, the models focused exclusively on the phasic pause-and-burst activity of primate DA neurons; the neurons' slower, tonic background activity was assumed to be constant. This has led to difficulty in explaining the results of neurochemical experiments that measure indications of DA release on a slow timescale, results that seem at first glance inconsistent with a reward prediction model. In this article, we investigate a TD model of DA activity modified so as to enable it to make longer-term predictions about rewards expected far in the future. We show that these predictions manifest themselves as slow changes in the baseline error signal, which we associate with tonic DA activity. Using this model, we make new predictions about the behavior of the DA system in a number of experimental situations. Some of these predictions suggest new computational explanations for previously puzzling data, such as indications from microdialysis studies of elevated DA activity triggered by aversive events.
Nathaniel D. Daw, David S. Touretzky
Neural Comput.1
2002 Local analysis of behaviour in the adjusting-delay task for assessing choice of delayed reinforcement
Rudolf N. Cardinal, Nathaniel D. Daw, Trevor W. Robbins, Barry J. Everitt
Neural Networks2
2002 Opponent interactions between serotonin and dopamine
Nathaniel D. Daw, Sham M. Kakade, Peter Dayan
Neural Networks1
2001 Operant behavior suggests attentional gating of dopamine system inputs
Nathaniel D. Daw, David S. Touretzky
Neurocomputing1
2000 Embedded Compilation for Multimedia Applications
abstract
Reconfigurable computing obtains its performance advantage over fixed processors by creating hardware configurations that are specialized for a particular application. In some cases, this advantage can be pushed even further, by creating hardware specialized to a particular instance of an application. For many problems where this approach is applicable, such as automatic target recognition, template matching and encryption, the problem parameters can change often, even within a single program execution, requiring periodic, and potentially expensive, hardware reconfigurations. To support these applications, we propose a method for on-chip configuration generation, or embedded compilation, for use with Carnegie Mellon University's PipeRench reconfigurable processor. We describe PipeRench's performance in detail for one problem, template matching, relative to the newest general-purpose processors, and show how embedded compilation can be used to support multiple problem instances for a second problem, IDEA encryption.
Nathaniel D. Daw, Seth Copen Goldstein, Dennis Strelow
FCCM1
2000 Behavioral considerations suggest an average reward TD model of the dopamine system
Nathaniel D. Daw, David S. Touretzky
Neurocomputing1