VLDB 2026 Research / reviewers in the wild / expert
Peter Dayan
dblp:22/522 · also Peter S. Dayan
· DBLP profile ↗
157ranked-venue papers
30as first author
34since 2021 · last 2026
0000-0003-3476-1839ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 115 · 29 first-author · 18 since 2021Applied, interdisciplinary, general and emerging computing · 57 · 1 first-author · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ℵ-IPOMDP: Mitigating Deception in a Cognitive Hierarchy with Off-Policy Counterfactual Anomaly DetectionabstractSocial agents with finitely nested opponent models are vulnerable to manipulation by agents with deeper recursive capabilities. This imbalance, rooted in logic and the theory of recursive modelling frameworks, cannot be solved directly. We propose a computational framework called ℵ-IPOMDP, which augments the Bayesian inference of model-based RL agents with an anomaly detection algorithm and an out-of-belief policy. Our mechanism allows agents to realize that they are being deceived, even if they cannot understand how, and to deter opponents via a credible threat. We test this framework in both a mixed-motive and a zero-sum game. Our results demonstrate the ℵ-mechanism’s effectiveness, leading to more equitable outcomes and less exploitation by more sophisticated agents. We discuss implications for AI safety, cybersecurity, cognitive science, and psychiatry. Nitay Alon, Joseph M. Barnby, Stefan Sarkadi, Lion Schulz, Jeffrey S. Rosenschein, Peter Dayan |
J. Artif. Intell. Res. | 6 |
| 2026 | Metacognitive efficiency in learned value-based choiceabstractMetacognition, the ability to assess the quality of our own decisions, is a critical form of higher-order information processing. Metacognitive efficiency is therefore an essential measure of cognitive capability. When making decisions is simple, metacognitive judgments are also straightforward; thus, assessing metacognitive efficiency requires normalizing for the quality of underlying task performance. This is duly common in measures of efficiency popular in perceptual decision-making such as the M-ratio. However, such normalization is hard in reinforcement learning problems, because task difficulty changes dynamically. We therefore repurposed the central idea underlying the M-ratio, using confidence judgments to fashion a notional decision-maker (which we call a Backward model), assessing metacognitive sensitivity according to the quality of its virtual decisions, and quantifying metacognitive efficiency by comparing virtual and modelled actual qualities in the original decision-making task. We used simulated and empirical data to show that our measure of metacognitive sensitivity, the Backward performance, has comparable properties to other measures such as quadratic scoring, and that our measure of efficiency, the MetaRL.Ratio, is independent of empirical performance and is preserved across levels of task difficulty. We suggest that the MetaRL.Ratio as a promising tool for assessing metacognitive efficiency in value-based learning/decision-making. Sara Ershadmanesh, Ali Gholamzadeh, Kobe Desender, Peter Dayan |
PLoS Comput. Biol. | 4 |
| 2026 | Homeostasis after injury: How intertwined inference and control underpin post-injury pain and behaviourabstractInjuries are an unfortunate but inevitable fact of life, leading to an evolutionary mandate for powerful homeostatic processes of recovery and recuperation. The physiological responses of the body and the immune system must be coordinated with behaviour to allow protected time for this to happen, and to prevent further damage to the affected bodily parts. Reacting appropriately requires an internal control system that represents the nature and state of the injury and specifies and withholds actions accordingly. We bring the formal uncertainties embodied in this system into the framework of a partially observable Markov decision process. We discuss nociceptive phenomena in light of this analysis, noting particularly the counter-intuitive behaviours associated with injury investigation, and the propensity for transitions from normative, tonic, to pathological, chronic pain states. Importantly, these simulation results provide a quantitative account and enable us to sketch a much needed roadmap for future theoretical and experimental studies on injury, tonic pain, and the transition to chronic pain. Pranav Mahajan, Peter Dayan, Ben Seymour |
PLoS Comput. Biol. | 2 |
| 2026 | Composing egocentric and allocentric maps for flexible navigationabstractEgocentric representations of the environment have historically been relegated to being used only for simple forms of spatial behaviour such as stimulus-response learning. However, in the many cases that critical aspects of policies are best defined relative to the self, egocentric representations can be advantageous. Furthermore, there is evidence that forms of egocentric representation might exist in the wider hippocampal formation. Nevertheless, egocentric representations have yet to be fully incorporated as a component of modern navigational methods. Here we investigate egocentric successor representations (SRs) and their combination with allocentric representations. We build a reinforcement learning agent that combines an egocentric SR with a conventional allocentric SR to navigate complex 2D environments. We demonstrate that the agent learns generalisable egocentric and allocentric value functions which, even when only additively composed, allow it to learn policies efficiently and to adapt to new environments quickly. Our work shows the benefit for egocentric relational structure to be captured, as well as allocentric. We offer a new perspective on how cognitive maps could usefully be composed from multiple simple maps representing associations between state features defined in different reference frames. Daniel Shani, Peter Dayan |
PLoS Comput. Biol. | 2 |
| 2025 | Complexity in Complexity: Understanding Visual Complexity Through Structure, Color, and Surprise
Karahan Saritas, Peter Dayan, Tingke Shen, Surabhi S. Nath |
CogSci | 2 |
| 2025 | Striking the Right Chord Between Reuse and Improvisation: Melody Learning as Resource-Rational Program Induction
Hanqi Zhou, David G. Nagy, Peter Dayan, Charley M. Wu |
CogSci | 3 |
| 2025 | Building, Reusing, and Generalizing Abstract Representations from Concrete SequencesabstractHumans excel at learning abstract patterns across different sequences, filtering out
irrelevant details, and transferring these generalized concepts to new sequences.
In contrast, many sequence learning models lack the ability to abstract, which
leads to memory inefficiency and poor transfer. We introduce a non-parametric
hierarchical variable learning model (HVM) that learns chunks from sequences
and abstracts contextually similar chunks as variables. HVM efficiently organizes
memory while uncovering abstractions, leading to compact sequence representations.
When learning on language datasets such as babyLM, HVM learns a more efficient
dictionary than standard compression algorithms such as Lempel-Ziv. In a sequence
recall task requiring the acquisition and transfer of variables embedded in sequences,
we demonstrate HVM’s sequence likelihood correlates with human recall times. In
contrast, large language models (LLMs) struggle to transfer abstract variables as
effectively as humans. From HVM’s adjustable layer of abstraction, we demonstrate
that the model realizes a precise trade-off between compression and generalization.
Our work offers a cognitive model that captures the learning and transfer of abstract
representations in human cognition and differentiates itself from LLMs. Shuchen Wu, Mirko Thalmann, Peter Dayan, Zeynep Akata, Eric Schulz |
ICLR | 3 |
| 2025 | Concept-Guided Interpretability via Neural ChunkingabstractNeural networks are often described as
black boxes, reflecting the significant challenge of understanding their internal workings and interactions. We propose a different perspective that challenges the prevailing view: rather than being inscrutable, neural networks exhibit patterns in their raw population activity that mirror regularities in the training data. We refer to this as the \textit{Reflection Hypothesis} and provide evidence for this phenomenon in both simple recurrent neural networks (RNNs) and complex large language models (LLMs).
Building on this insight, we propose to leverage cognitively-inspired methods of \textit{chunking} to segment high-dimensional neural population dynamics into interpretable units that reflect underlying concepts.
We propose three methods to extract these emerging entities, complementing each other based on label availability and neural data dimensionality. Discrete sequence chunking (DSC) creates a dictionary of entities in a lower-dimensional neural space; population averaging (PA) extracts recurring entities that correspond to known labels; and unsupervised chunk discovery (UCD) can be used when labels are absent.
We demonstrate the effectiveness of these methods in extracting entities across varying model sizes, ranging from inducing compositionality in RNNs to uncovering recurring neural population states in large language models with diverse architectures, and illustrate their advantage to other interpretability methods.
Throughout, we observe a robust correspondence between the extracted entities and concrete or abstract concepts in the sequence. Artificially inducing the extracted entities in neural populations effectively alters the network's generation of associated concepts.
Our work points to a new direction for interpretability, one that harnesses both cognitive principles and the structure of naturalistic data to reveal the hidden computations of complex learning systems, gradually transforming them from black boxes into systems we can begin to understand.
Implementation and code are publicly available at _https://github.com/swu32/Chunk-Interpretability_ Shuchen Wu, Stephan Alaniz, Shyamgopal Karthik, Peter Dayan, Eric Schulz, Zeynep Akata |
NeurIPS | 4 |
| 2025 | Noradrenergic and Dopaminergic modulation of meta-cognition and meta-controlabstractHumans and animals use multiple control systems for decision-making. This involvement is subject to meta-cognitive regulation - as a form of control over control or meta-control. However, the nature of this meta-control is unclear. For instance, Model-based (MB) control may be boosted when decision-makers generally lack confidence as it is more statistically efficient; or it may be suppressed, since the MB controller can correctly assess its own unreliability. Since control and metacontrol are themselves subject to the influence of neuromodulators, we examined the effects of perturbing the noradrenergic (NE) and dopaminergic (DA) systems with propranolol and L-DOPA, respectively. We first administered a simple perceptual task to examine the effects of the manipulations on meta-cognitive ability. Using Bayesian analyses, we found that 81% of group M-ratio samples were lower under propranolol relative to placebo, suggesting a decrease of meta-cognitive ability; and 60% of group M-ratio samples were higher under L-DOPA relative to placebo, considered as no effect of L-DOPA on meta-cognitive ability . We then asked subjects to provide choices and confidence ratings in a two-outcome decision-making task that has been used to dissociate Model-free (MF) and MB control. MB behavior was enhanced by propranolol, while MF behavior was not significantly affected by either drug. The interaction between confidence and MF/MB behavior was highly variable under propranolol, but under L-DOPA, the interaction was significantly lower/higher relative to placebo. Our results suggest a decrease in metacognitive ability under the influence of propranolol and an enhancement of MB behavior and meta-control under the influence of propranolol and L-DOPA, respectively. These findings shed light on the role of NE and DA in different aspects of control and meta-control and suggest potential avenues for mitigating dysfunction. Sara Ershadmanesh, Sahar Rajabi, Reza Rostami, Rani Moran, Peter Dayan |
PLoS Comput. Biol. | 5 |
| 2025 | Mechanisms of mistrust: A Bayesian account of misinformation learningabstractFrom the intimate realm of personal interactions to the sprawling arena of political discourse, discerning the trustworthy from the dubious is crucial. Here, we present a novel behavioral task and accompanying Bayesian models that allow us to study key aspects of this learning process in a tightly controlled setting. In our task, participants are confronted with several different types of (mis-)information sources, ranging from ones that lie to ones with biased reporting, and have to learn these attributes under varying degrees of feedback. We formalize inference in this setting as a doubly Bayesian learning process where agents simultaneously learn about the ground truth as well as the qualities of an information source reporting on this ground truth. Our model and detailed analyses reveal how participants can generally follow Bayesian learning dynamics, highlighting a basic human ability to learn about diverse information sources. This learning is also reflected in explicit trust reports about the sources. We additionally show how participants approached the inference problem with priors that held sources to be helpful. Finally, when outside feedback was noisier, participants still learned along Bayesian lines but struggled to pick up on biases in information. Our work pins down computationally the generally impressive human ability to learn the trustworthiness of information sources while revealing minor fault lines when it comes to noisier environments and news sources with a slant. Lion Schulz, Yannick Streicher, Eric Schulz, Rahul Bhui, Peter Dayan |
PLoS Comput. Biol. | 5 |
| 2024 | Optimal and sub-optimal temporal decisions can explain procrastination in a real-world task
Sahiti Chebolu, Peter Dayan |
CogSci | 2 |
| 2024 | State-Independent and State-Dependent Learning in a Motivational Go/NoGo task
Azadeh Nazemorroaya, Dan Bang, Peter Dayan |
CogSci | 3 |
| 2024 | Simplicity in Complexity: Explaining Visual Complexity using Deep Segmentation Models
Tingke Shen, Surabhi S. Nath, Aenne Brielmann, Peter Dayan |
CogSci | 4 |
| 2024 | Characterising the Creative Process in Humans and Large Language Models
Surabhi S. Nath, Peter Dayan, Claire E. Stevenson |
ICCC | 2 |
| 2024 | Simplifying Latent Dynamics with Softly State-Invariant World ModelsabstractTo solve control problems via model-based reasoning or planning, an agent needs to know how its actions affect the state of the world. The actions an agent has at its disposal often change the state of the environment in systematic ways. However, existing techniques for world modelling do not guarantee that the effect of actions are represented in such systematic ways. We introduce the Parsimonious Latent Space Model (PLSM), a world model that regularizes the latent dynamics to make the effect of the agent's actions more predictable. Our approach minimizes the mutual information between latent states and the change that an action produces in the agent's latent state, in turn minimizing the dependence the state has on the dynamics. This makes the world model softly state-invariant. We combine PLSM with different model classes used for i) future latent state prediction, ii) planning, and iii) model-free reinforcement learning. We find that our regularization improves accuracy, generalization, and performance in downstream tasks, highlighting the importance of systematic treatment of actions in world models. Tankred Saanum, Peter Dayan, Eric Schulz |
NeurIPS | 2 |
| 2024 | Wagers for work: Decomposing the costs of cognitive effortabstractSome aspects of cognition are more taxing than others. Accordingly, many people will avoid cognitively demanding tasks in favor of simpler alternatives. Which components of these tasks are costly, and how much, remains unknown. Here, we use a novel task design in which subjects request wages for completing cognitive tasks and a computational modeling procedure that decomposes their wages into the costs driving them. Using working memory as a test case, our approach revealed that gating new information into memory and protecting against interference are costly. Critically, other factors, like memory load, appeared less costly. Other key factors which may drive effort costs, such as error avoidance, had minimal influence on wage requests. Our approach is sensitive to individual differences, and could be used in psychiatric populations to understand the true underlying nature of apparent cognitive deficits. Sarah L. Master, Clayton E. Curtis, Peter Dayan |
PLoS Comput. Biol. | 3 |
| 2023 | Habits of Mind: Reusing Action Sequences for Efficient Planning
Noémi Élteto, Peter Dayan |
CogSci | 2 |
| 2023 | Compositionality under time pressure
Valerio Rubino, Mani Hamidi, Peter Dayan, Charley M. Wu |
CogSci | 3 |
| 2023 | Reinforcement Learning with Simple Sequence PriorsabstractIn reinforcement learning (RL), simplicity is typically quantified on an action-by-action basis -- but this timescale ignores temporal regularities, like repetitions, often present in sequential strategies. We therefore propose an RL algorithm that learns to solve tasks with sequences of actions that are compressible. We explore two possible sources of simple action sequences: Sequences that can be learned by autoregressive models, and sequences that are compressible with off-the-shelf data compression algorithms. Distilling these preferences into sequence priors, we derive a novel information-theoretic objective that incentivizes agents to learn policies that maximize rewards while conforming to these priors. We show that the resulting RL algorithm leads to faster learning, and attains higher returns than state-of-the-art model-free approaches in a series of continuous control tasks from the DeepMind Control Suite. These priors also produce a powerful information-regularized agent that is robust to noisy observations and can perform open-loop control. Tankred Saanum, Noémi Élteto, Peter Dayan, Marcel Binz, Eric Schulz |
NeurIPS | 3 |
| 2023 | Reframing dopamine: A controlled controller at the limbic-motor interfaceabstractPavlovian influences notoriously interfere with operant behaviour. Evidence suggests this interference sometimes coincides with the release of the neuromodulator dopamine in the nucleus accumbens. Suppressing such interference is one of the targets of cognitive control. Here, using the examples of active avoidance and omission behaviour, we examine the possibility that direct manipulation of the dopamine signal is an instrument of control itself. In particular, when instrumental and Pavlovian influences come into conflict, dopamine levels might be affected by the controlled deployment of a reframing mechanism that recasts the prospect of possible punishment as an opportunity to approach safety, and the prospect of future reward in terms of a possible loss of that reward. We operationalize this reframing mechanism and fit the resulting model to rodent behaviour from two paradigmatic experiments in which accumbens dopamine release was also measured. We show that in addition to matching animals' behaviour, the model predicts dopamine transients that capture some key features of observed dopamine release at the time of discriminative cues, supporting the idea that modulation of this neuromodulator is amongst the repertoire of cognitive control strategies. Kevin Lloyd, Peter Dayan |
PLoS Comput. Biol. | 2 |
| 2022 | Neural Network Poisson Models for Behavioural and Neural Spike Train DataabstractOne of the most important and challenging application areas for complex machine learning methods is to predict, characterize and model rich, multi-dimensional, neural data. Recent advances in neural recording techniques have made it possible to monitor the activity of a large number of neurons across different brain regions as animals perform behavioural tasks. This poses the critical challenge of establishing links between neural activity at a microscopic scale, which might for instance represent sensory input, and at a macroscopic scale, which then generates behaviour. Predominant modeling methods apply rather disjoint techniques to these scales; by contrast, we suggest an end-to-end model which exploits recent developments of flexible, but tractable, neural network point-process models to characterize dependencies between stimuli, actions, and neural data. We apply this model to a public dataset collected using Neuropixel probes in mice performing a visually-guided behavioural task as well as a synthetic dataset produced from a hierarchical network model with reciprocally connected sensory and integration circuits intended to characterize animal behaviour in a fixed-duration motion discrimination task. We show that our model outperforms previous approaches and contributes novel insights into the relationships between neural activity and behaviour. Moein Khajehnejad, Forough Habibollahi, Richard Nock, Ehsan Arabzadeh, Peter Dayan, Amir Dezfouli |
ICML | 5 |
| 2022 | Optimism and pessimism in optimised replayabstractThe replay of task-relevant trajectories is known to contribute to memory consolidation and improved task performance. A wide variety of experimental data show that the content of replayed sequences is highly specific and can be modulated by reward as well as other prominent task variables. However, the rules governing the choice of sequences to be replayed still remain poorly understood. One recent theoretical suggestion is that the prioritization of replay experiences in decision-making problems is based on their effect on the choice of action. We show that this implies that subjects should replay sub-optimal actions that they dysfunctionally choose rather than optimal ones, when, by being forgetful, they experience large amounts of uncertainty in their internal models of the world. We use this to account for recent experimental data demonstrating exactly pessimal replay, fitting model parameters to the individual subjects' choices. Georgy Antonov, Chris Gagne 0001, Eran Eldar, Peter Dayan |
PLoS Comput. Biol. | 4 |
| 2022 | Vigilance, arousal, and acetylcholine: Optimal control of attention in a simple detection taskabstractPaying attention to particular aspects of the world or being more vigilant in general can be interpreted as forms of 'internal' action. Such arousal-related choices come with the benefit of increasing the quality and situational appropriateness of information acquisition and processing, but incur potentially expensive energetic and opportunity costs. One implementational route for these choices is widespread ascending neuromodulation, including by acetylcholine (ACh). The key computational question that elective attention poses for sensory processing is when it is worthwhile paying these costs, and this includes consideration of whether sufficient information has yet been collected to justify the higher signal-to-noise ratio afforded by greater attention and, particularly if a change in attentional state is more expensive than its maintenance, when states of heightened attention ought to persist. We offer a partially observable Markov decision-process treatment of optional attention in a detection task, and use it to provide a qualitative model of the results of studies using modern techniques to measure and manipulate ACh in rodents performing a similar task. Sahiti Chebolu, Peter Dayan, Kevin Lloyd |
PLoS Comput. Biol. | 2 |
| 2022 | The pursuit of happiness: A reinforcement learning perspective on habituation and comparisonsabstractIn evaluating our choices, we often suffer from two tragic relativities. First, when our lives change for the better, we rapidly habituate to the higher standard of living. Second, we cannot escape comparing ourselves to various relative standards. Habituation and comparisons can be very disruptive to decision-making and happiness, and till date, it remains a puzzle why they have come to be a part of cognition in the first place. Here, we present computational evidence that suggests that these features might play an important role in promoting adaptive behavior. Using the framework of reinforcement learning, we explore the benefit of employing a reward function that, in addition to the reward provided by the underlying task, also depends on prior expectations and relative comparisons. We find that while agents equipped with this reward function are less happy, they learn faster and significantly outperform standard reward-based agents in a wide range of environments. Specifically, we find that relative comparisons speed up learning by providing an exploration incentive to the agents, and prior expectations serve as a useful aid to comparisons, especially in sparsely-rewarded and non-stationary environments. Our simulations also reveal potential drawbacks of this reward function and show that agents perform sub-optimally when comparisons are left unchecked and when there are too many similar options. Together, our results help explain why we are prone to becoming trapped in a cycle of never-ending wants and desires, and may shed light on psychopathologies such as depression, materialism, and overconsumption. Rachit Dubey, Thomas L. Griffiths 0001, Peter Dayan |
PLoS Comput. Biol. | 3 |
| 2022 | Tracking human skill learning with a hierarchical Bayesian sequence modelabstractHumans can implicitly learn complex perceptuo-motor skills over the course of large numbers of trials. This likely depends on our becoming better able to take advantage of ever richer and temporally deeper predictive relationships in the environment. Here, we offer a novel characterization of this process, fitting a non-parametric, hierarchical Bayesian sequence model to the reaction times of human participants' responses over ten sessions, each comprising thousands of trials, in a serial reaction time task involving higher-order dependencies. The model, adapted from the domain of language, forgetfully updates trial-by-trial, and seamlessly combines predictive information from shorter and longer windows onto past events, weighing the windows proportionally to their predictive power. As the model implies a posterior over window depths, we were able to determine how, and how many, previous sequence elements influenced individual participants' internal predictions, and how this changed with practice. Already in the first session, the model showed that participants had begun to rely on two previous elements (i.e., trigrams), thereby successfully adapting to the most prominent higher-order structure in the task. The extent to which local statistical fluctuations in trigram frequency influenced participants' responses waned over subsequent sessions, as participants forgot the trigrams less and evidenced skilled performance. By the eighth session, a subset of participants shifted their prior further to consider a context deeper than two previous elements. Finally, participants showed resistance to interference and slow forgetting of the old sequence when it was changed in the final sessions. Model parameters for individual participants covaried appropriately with independent measures of working memory and error characteristics. In sum, the model offers the first principled account of the adaptive complexity and nuanced dynamics of humans' internal sequence representations during long-term implicit skill learning. Noémi Élteto, Dezso Németh, Karolina Janacsek, Peter Dayan |
PLoS Comput. Biol. | 4 |
| 2022 | Biased belief priors versus biased belief updating: Differential correlates of depression and anxietyabstractIndividuals prone to anxiety and depression often report beliefs and make judgements about themselves that are more negative than those reported by others. We use computational modeling of a richly naturalistic task to disentangle the role of negative priors versus negatively biased belief updating and to investigate their association with different dimensions of Internalizing psychopathology. Undergraduate participants first provided profiles for a hypothetical tech internship. They then viewed pairs of other profiles and selected the individual they would prefer to work alongside out of each pair. In a subsequent phase of the experiment, participants made judgments about their relative popularity as hypothetical internship partners both before any feedback and after each of 20 items of feedback revealing whether or not they had been selected as the preferred teammate from a given pairing. Scores on latent factors of general negative affect, anxiety-specific affect and depression-specific affect were estimated using participants' self-report scores on standardized measures of anxiety and depression together with factor loadings from a bifactor analysis conducted previously. Higher scores on the depression-specific factor were linked to more negative prior beliefs but were not associated with differences in belief updating. In contrast, higher scores on the anxiety-specific factor were associated with a negative bias in belief updating but no difference in prior beliefs. These findings indicate that, to at least some extent, distinct processes may impact the formation of belief priors and in-the-moment belief updating and that these processes may be differentially disrupted in depression and anxiety. Future directions for enquiry include examination of the possibility that prior beliefs biases in depression might reflect generalization from prior experiences or global schema whereas belief updating biases in anxiety might be more situationally specific. Chris Gagne 0001, Sharon Agai, Christian Ramiro, Peter Dayan, Sonia J. Bishop |
PLoS Comput. Biol. | 4 |
| 2022 | Interactions between attributions and beliefs at trial-by-trial level: Evidence from a novel computer game taskabstractInferring causes of the good and bad events that we experience is part of the process of building models of our own capabilities and of the world around us. Making such inferences can be difficult because of complex reciprocal relationships between attributions of the causes of particular events, and beliefs about the capabilities and skills that influence our role in bringing them about. Abnormal causal attributions have long been studied in connection with psychiatric disorders, notably depression and paranoia; however, the mechanisms behind attributional inferences and the way they can go awry are not fully understood. We administered a novel, challenging, game of skill to a substantial population of healthy online participants, and collected trial-by-trial time series of both their beliefs about skill and attributions about the causes of the success and failure of real experienced outcomes. We found reciprocal relationships that provide empirical confirmation of the attribution-self representation cycle theory. This highlights the dynamic nature of the processes involved in attribution, and validates a framework for developing and testing computational accounts of attribution-belief interactions. Elena Zamfir, Peter Dayan |
PLoS Comput. Biol. | 2 |
| 2021 | Exploring learning trajectories with dynamic infinite hidden Markov models
Sebastian A. Bruijns, Peter Dayan |
CogSci | 3 |
| 2021 | Tracking the Unknown: Modeling Long-Term Implicit Skill Acquisition as Non-Parametric Bayesian Sequence Learning
Noémi Élteto, Dezso Németh, Karolina Janacsek, Peter Dayan |
CogSci | 4 |
| 2021 | Confidence in control: Metacognitive computations for information search
Lion Schulz, Stephen M. Fleming, Peter Dayan |
CogSci | 3 |
| 2021 | Correcting experience replay for multi-agent communication
Sanjeevan Ahilan, Peter Dayan |
ICLR | 2 |
| 2021 | Two steps to risk sensitivityabstractDistributional reinforcement learning (RL) – in which agents learn about all the possible long-term consequences of their actions, and not just the expected value – is of great recent interest. One of the most important affordances of a distributional view is facilitating a modern, measured, approach to risk when outcomes are not completely certain. By contrast, psychological and neuroscientific investigations into decision making under risk have utilized a variety of more venerable theoretical models such as prospect theory that lack axiomatically desirable properties such as coherence. Here, we consider a particularly relevant risk measure for modeling human and animal planning, called conditional value-at-risk (CVaR), which quantifies worst-case outcomes (e.g., vehicle accidents or predation). We first adopt a conventional distributional approach to CVaR in a sequential setting and reanalyze the choices of human decision-makers in the well-known two-step task, revealing substantial risk aversion that had been lurking under stickiness and perseveration. We then consider a further critical property of risk sensitivity, namely time consistency, showing alternatives to this form of CVaR that enjoy this desirable characteristic. We use simulations to examine settings in which the various forms differ in ways that have implications for human and animal planning and behavior. Chris Gagne 0001, Peter Dayan |
NeurIPS | 2 |
| 2021 | Internality and the internalisation of failure: Evidence from a novel taskabstractA critical facet of adjusting one's behaviour after succeeding or failing at a task is assigning responsibility for the ultimate outcome. Humans have trait- and state-like tendencies to implicate aspects of their own behaviour (called 'internal' ascriptions) or facets of the particular task or Lady Luck ('chance'). However, how these tendencies interact with actual performance is unclear. We designed a novel task in which subjects had to learn the likelihood of achieving their goals, and the extent to which this depended on their efforts. High internality (Levenson I-score) was associated with decision making patterns that are less vulnerable to failure. Our computational analyses suggested that this depended heavily on the adjustment in the perceived achievability of riskier goals following failure. We found beliefs about chance not to be explanatory of choice behaviour in our task. Beliefs about powerful others were strong predictors of behaviour, but only when subjects lacked substantial influence over the outcome. Our results provide an evidentiary basis for heuristics and learning differences that underlie the formation and maintenance of control expectations by the self. Federico Mancinelli, Jonathan P. Roiser, Peter Dayan |
PLoS Comput. Biol. | 3 |
| 2021 | Dissecting the links between reward and loss, decision-making, and self-reported affect using a computational approachabstractLinks between affective states and risk-taking are often characterised using summary statistics from serial decision-making tasks. However, our understanding of these links, and the utility of decision-making as a marker of affect, needs to accommodate the fact that ongoing (e.g., within-task) experience of rewarding and punishing decision outcomes may alter future decisions and affective states. To date, the interplay between affect, ongoing reward and punisher experience, and decision-making has received little detailed investigation. Here, we examined the relationships between reward and loss experience, affect, and decision-making in humans using a novel judgement bias task analysed with a novel computational model. We demonstrated the influence of within-task favourability on decision-making, with more risk-averse/'pessimistic' decisions following more positive previous outcomes and a greater current average earning rate. Additionally, individuals reporting more negative affect tended to exhibit greater risk-seeking decision-making, and, based on our model, estimated time more poorly. We also found that individuals reported more positive affective valence during periods of the task when prediction errors and offered decision outcomes were more positive. Our results thus provide new evidence that (short-term) within-task rewarding and punishing experiences determine both future decision-making and subjectively experienced affective states. Vikki Neville, Peter Dayan, Iain D. Gilchrist, Elizabeth S. Paul, Michael Mendl |
PLoS Comput. Biol. | 2 |
| 2020 | A Local Temporal Difference Code for Distributional Reinforcement LearningabstractRecent theoretical and experimental results suggest that the dopamine system implements distributional temporal difference backups, allowing learning of the entire distributions of the long-run values of states rather than just their expected values. However, the distributional codes explored so far rely on a complex imputation step which crucially relies on spatial non-locality: in order to compute reward prediction errors, units must know not only their own state but also the states of the other units. It is far from clear how these steps could be implemented in realistic neural circuits. Here, we introduce the Laplace code: a local temporal difference code for distributional reinforcement learning that is representationally powerful and computationally straightforward. The code decomposes value distributions and prediction errors across three separated dimensions: reward magnitude (related to distributional quantiles), temporal discounting (related to the Laplace transform of future rewards) and time horizon (related to eligibility traces). Besides lending itself to a local learning rule, the decomposition recovers the temporal evolution of the immediate reward distribution, indicating all possible rewards at all future times. This increases representational capacity and allows for temporally-flexible computations that immediately adjust to changing horizons or discount factors. Pablo Tano, Peter Dayan, Alexandre Pouget |
NeurIPS | 2 |
| 2020 | Static and Dynamic Values of Computation in MCTSabstractMonte-Carlo Tree Search (MCTS) is one of the most-widely used methodsfor planning, and has powered many recent advances in artificialintelligence. In MCTS, one typically performs computations(i.e., simulations) to collect statistics about the possible futureconsequences of actions, and then chooses accordingly. Manypopular MCTS methods such as UCT and its variants decide whichcomputations to perform by trading-off exploration and exploitation. Inthis work, we take a more direct approach, and explicitly quantify thevalue of a computation based on its expected impact on the quality ofthe action eventually chosen. Our approach goes beyond the \emph{myopic}limitations of existing computation-value-based methods in two senses:(I) we are able to account for the impact of non-immediate (ie, future)computations (II) on non-immediate actions. We show that policies thatgreedily optimize computation values are optimal under certainassumptions and obtain results that are competitive with thestate-of-the-art. Eren Sezener, Peter Dayan |
UAI | 2 |
| 2020 | Combined model-free and model-sensitive reinforcement learning in non-human primatesabstractContemporary reinforcement learning (RL) theory suggests that potential choices can be evaluated by strategies that may or may not be sensitive to the computational structure of tasks. A paradigmatic model-free (MF) strategy simply repeats actions that have been rewarded in the past; by contrast, model-sensitive (MS) strategies exploit richer information associated with knowledge of task dynamics. MF and MS strategies should typically be combined, because they have complementary statistical and computational strengths; however, this tradeoff between MF/MS RL has mostly only been demonstrated in humans, often with only modest numbers of trials. We trained rhesus monkeys to perform a two-stage decision task designed to elicit and discriminate the use of MF and MS methods. A descriptive analysis of choice behaviour revealed directly that the structure of the task (of MS importance) and the reward history (of MF and MS importance) significantly influenced both choice and response vigour. A detailed, trial-by-trial computational analysis confirmed that choices were made according to a combination of strategies, with a dominant influence of a particular form of model sensitivity that persisted over weeks of testing. The residuals from this model necessitated development of a new combined RL model which incorporates a particular credit assignment weighting procedure. Finally, response vigor exhibited a subtly different collection of MF and MS influences. These results provide new illumination onto RL behavioural processes in non-human primates. Bruno Miranda, W. M. Nishantha Malalasekera, Timothy Edward John Behrens, Peter Dayan, Steven W. Kennerley |
PLoS Comput. Biol. | 4 |
| 2019 | Disentangled behavioural representationsabstractIndividual characteristics in human decision-making are often quantified by fitting a parametric cognitive model to subjects' behavior and then studying differences between them in the associated parameter space. However, these models often fit behavior more poorly than recurrent neural networks (RNNs), which are more flexible and make fewer assumptions about the underlying decision-making processes. Unfortunately, the parameter and latent activity spaces of RNNs are generally high-dimensional and uninterpretable, making it hard to use them to study individual differences. Here, we show how to benefit from the flexibility of RNNs while representing individual differences in a low-dimensional and interpretable space. To achieve this, we propose a novel end-to-end learning framework in which an encoder is trained to map the behavior of subjects into a low-dimensional latent space. These low-dimensional representations are used to generate the parameters of individual RNNs corresponding to the decision-making process of each subject. We introduce terms into the loss function that ensure that the latent dimensions are informative and disentangled, i.e., encouraged to have distinct effects on behavior. This allows them to align with separate facets of individual differences. We illustrate the performance of our framework on synthetic data as well as a dataset including the behavior of patients with psychiatric disorders. Amir Dezfouli, Hassan Ashtiani, Omar Ghattas, Richard Nock, Peter Dayan, Cheng Soon Ong |
NeurIPS | 5 |
| 2019 | Learning to use past evidence in a sophisticated world modelabstractHumans and other animals are able to discover underlying statistical structure in their environments and exploit it to achieve efficient and effective performance. However, such structure is often difficult to learn and use because it is obscure, involving long-range temporal dependencies. Here, we analysed behavioural data from an extended experiment with rats, showing that the subjects learned the underlying statistical structure, albeit suffering at times from immediate inferential imperfections as to their current state within it. We accounted for their behaviour using a Hidden Markov Model, in which recent observations are integrated with evidence from the past. We found that over the course of training, subjects came to track their progress through the task more accurately, a change that our model largely attributed to improved integration of past evidence. This learning reflected the structure of the task, decreasing reliance on recent observations, which were potentially misleading. Sanjeevan Ahilan, Rebecca B. Solomon, Yannick-André Breton, Kent L. Conover, Ritwik K. Niyogi, Peter Shizgal, Peter Dayan |
PLoS Comput. Biol. | 7 |
| 2019 | Models that learn how humans learn: The case of decision-making and its disordersabstractPopular computational models of decision-making make specific assumptions about learning processes that may cause them to underfit observed behaviours. Here we suggest an alternative method using recurrent neural networks (RNNs) to generate a flexible family of models that have sufficient capacity to represent the complex learning and decision- making strategies used by humans. In this approach, an RNN is trained to predict the next action that a subject will take in a decision-making task and, in this way, learns to imitate the processes underlying subjects' choices and their learning abilities. We demonstrate the benefits of this approach using a new dataset drawn from patients with either unipolar (n = 34) or bipolar (n = 33) depression and matched healthy controls (n = 34) making decisions on a two-armed bandit task. The results indicate that this new approach is better than baseline reinforcement-learning methods in terms of overall performance and its capacity to predict subjects' choices. We show that the model can be interpreted using off-policy simulations and thereby provides a novel clustering of subjects' learning processes-something that often eludes traditional approaches to modelling and behavioural analysis. Amir Dezfouli, Kristi Griffiths, Fabio Ramos 0001, Peter Dayan, Bernard W. Balleine |
PLoS Comput. Biol. | 4 |
| 2019 | A computational account of threat-related attentional biasabstractVisual selective attention acts as a filter on perceptual information, facilitating learning and inference about important events in an agent's environment. A role for visual attention in reward-based decisions has previously been demonstrated, but it remains unclear how visual attention is recruited during aversive learning, particularly when learning about multiple stimuli concurrently. This question is of particular importance in psychopathology, where enhanced attention to threat is a putative feature of pathological anxiety. Using an aversive reversal learning task that required subjects to learn, and exploit, predictions about multiple stimuli, we show that the allocation of visual attention is influenced significantly by aversive value but not by uncertainty. Moreover, this relationship is bidirectional in that attention biases value updates for attended stimuli, resulting in heightened value estimates. Our findings have implications for understanding biased attention in psychopathology and support a role for learning in the expression of threat-related attentional biases in anxiety. Toby Wise, Jochen Michely, Peter Dayan, Ray Dolan |
PLoS Comput. Biol. | 3 |
| 2018 | Fast Parametric Learning with Activation MemorizationabstractNeural networks trained with backpropagation often struggle to identify classes that have been observed a small number of times. In applications where most class labels are rare, such as language modelling, this can become a performance bottleneck. One potential remedy is to augment the network with a fast-learning non-parametric model which stores recent activations and class labels into an external memory. We explore a simplified architecture where we treat a subset of the model parameters as fast memory stores. This can help retain information over longer time intervals than a traditional memory, and does not require additional space or compute. In the case of image classification, we display faster binding of novel classes on an Omniglot image curriculum task. We also show improved performance for word-based language models on news reports (GigaWord), books (Project Gutenberg) and Wikipedia articles (WikiText-103) - the latter achieving a state-of-the-art perplexity of 29.2. Jack W. Rae, Chris Dyer, Peter Dayan, Timothy P. Lillicrap |
ICML | 3 |
| 2018 | Integrated accounts of behavioral and neuroimaging data using flexible recurrent neural network modelsabstractNeuroscience studies of human decision-making abilities commonly involve subjects completing a decision-making task while BOLD signals are recorded using fMRI. Hypotheses are tested about which brain regions mediate the effect of past experience, such as rewards, on future actions. One standard approach to this is model-based fMRI data analysis, in which a model is fitted to the behavioral data, i.e., a subject's choices, and then the neural data are parsed to find brain regions whose BOLD signals are related to the model's internal signals. However, the internal mechanics of such purely behavioral models are not constrained by the neural data, and therefore might miss or mischaracterize aspects of the brain. To address this limitation, we introduce a new method using recurrent neural network models that are flexible enough to be jointly fitted to the behavioral and neural data. We trained a model so that its internal states were suitably related to neural activity during the task, while at the same time its output predicted the next action a subject would execute. We then used the fitted model to create a novel visualization of the relationship between the activity in brain regions at different times following a reward and the choices the subject subsequently made. Finally, we validated our method using a previously published dataset. We found that the model was able to recover the underlying neural substrates that were discovered by explicit model engineering in the previous work, and also derived new results regarding the temporal pattern of brain activity. Amir Dezfouli, Richard W. Morris, Fabio Ramos 0001, Peter Dayan, Bernard W. Balleine |
NeurIPS | 4 |
| 2018 | Control of neurite growth and guidance by an inhibitory cell-body signalabstractThe development of a functional nervous system requires tight control of neurite growth and guidance by extracellular chemical cues. Neurite growth is astonishingly sensitive to shallow concentration gradients, but a widely observed feature of both growth and guidance regulation, with important consequences for development and regeneration, is that both are only elicited over the same relatively narrow range of concentrations. Here we show that all these phenomena can be explained within one theoretical framework. We first test long-standing explanations for the suppression of the trophic effects of nerve growth factor at high concentrations, and find they are contradicted by experiment. Instead we propose a new hypothesis involving inhibitory signalling among the cell bodies, and then extend this hypothesis to show how both growth and guidance can be understood in terms of a common underlying signalling mechanism. This new model for the first time unifies several key features of neurite growth regulation, quantitatively explains many aspects of experimental data, and makes new predictions about unknown details of developmental signalling. Brendan A. Bicknell, Zac Pujic, Peter Dayan, Geoffrey J. Goodhill |
PLoS Comput. Biol. | 3 |
| 2018 | A model of risk and mental state shifts during social interactionabstractCooperation and competition between human players in repeated microeconomic games offer a window onto social phenomena such as the establishment, breakdown and repair of trust. However, although a suitable starting point for the quantitative analysis of such games exists, namely the Interactive Partially Observable Markov Decision Process (I-POMDP), computational considerations and structural limitations have limited its application, and left unmodelled critical features of behavior in a canonical trust task. Here, we provide the first analysis of two central phenomena: a form of social risk-aversion exhibited by the player who is in control of the interaction in the game; and irritation or anger, potentially exhibited by both players. Irritation arises when partners apparently defect, and it potentially causes a precipitate breakdown in cooperation. Failing to model one's partner's propensity for it leads to substantial economic inefficiency. We illustrate these behaviours using evidence drawn from the play of large cohorts of healthy volunteers and patients. We show that for both cohorts, a particular subtype of player is largely responsible for the breakdown of trust, a finding which sheds new light on borderline personality disorder. Andreas Hula, Iris Vilares, Terry Lohrenz, Peter Dayan, P. Read Montague |
PLoS Comput. Biol. | 4 |
| 2018 | Interrupting behaviour: Minimizing decision costs via temporal commitment and low-level interruptsabstractIdeal decision-makers should constantly assess all sources of information about opportunities and threats, and be able to redetermine their choices promptly in the face of change. However, perpetual monitoring and reassessment impose inordinate sensing and computational costs, making them impractical for animals and machines alike. The obvious alternative of committing for extended periods of time to limited sensory strategies associated with particular courses of action can be dangerous and wasteful. Here, we explore the intermediate possibility of making provisional temporal commitments whilst admitting interruption based on limited broader observation. We simulate foraging under threat of predation to elucidate the benefits of such a scheme. We relate our results to diseases of distractibility and roving attention, and consider mechanistic substrates such as noradrenergic neuromodulation. Kevin Lloyd, Peter Dayan |
PLoS Comput. Biol. | 2 |
| 2018 | Change, stability, and instability in the Pavlovian guidance of behaviour from adolescence to young adulthoodabstractPavlovian influences are important in guiding decision-making across health and psychopathology. There is an increasing interest in using concise computational tasks to parametrise such influences in large populations, and especially to track their evolution during development and changes in mental health. However, the developmental course of Pavlovian influences is uncertain, a problem compounded by the unclear psychometric properties of the relevant measurements. We assessed Pavlovian influences in a longitudinal sample using a well characterised and widely used Go-NoGo task. We hypothesized that the strength of Pavlovian influences and other 'psychomarkers' guiding decision-making would behave like traits. As reliance on Pavlovian influence is not as profitable as precise instrumental decision-making in this Go-NoGo task, we expected this influence to decrease with higher IQ and age. Additionally, we hypothesized it would correlate with expressions of psychopathology. We found that Pavlovian effects had weak temporal stability, while model-fit was more stable. In terms of external validity, Pavlovian effects decreased with increasing IQ and experience within the task, in line with normative expectations. However, Pavlovian effects were poorly correlated with age or psychopathology. Thus, although this computational construct did correlate with important aspects of development, it does not meet conventional requirements for tracking individual development. We suggest measures that might improve psychometric properties of task-derived Pavlovian measures for future studies. Michael Moutoussis, Edward T. Bullmore, Ian M. Goodyer, Peter Fonagy, Peter B. Jones, Ray Dolan, Peter Dayan |
PLoS Comput. Biol. | 7 |
| 2017 | A model of structure learning, inference, and generation for scene understanding
David Raposo, Peter Dayan, Demis Hassabis, Peter W. Battaglia |
CogSci | 2 |
| 2017 | Increased decision thresholds enhance information gathering performance in juvenile Obsessive-Compulsive Disorder (OCD)abstractPatients with obsessive-compulsive disorder (OCD) can be described as cautious and hesitant, manifesting an excessive indecisiveness that hinders efficient decision making.However, excess caution in decision making may also lead to better performance in specific situations where the cost of extended deliberation is small.We compared 16 juvenile OCD patients with 16 matched healthy controls whilst they performed a sequential information gathering task under different external cost conditions.We found that patients with OCD outperformed healthy controls, winning significantly more points.The groups also differed in the number of draws required prior to committing to a decision, but not in decision accuracy.A novel Bayesian computational model revealed that subjective sampling costs arose as a non-linear function of sampling, closely resembling an escalating urgency signal.Group difference in performance was best explained by a later emergence of these subjective costs in the OCD group, also evident in an increased decision threshold.Our findings present a novel computational model and suggest that enhanced information gathering in OCD can be accounted for by a higher decision threshold arising out of an altered perception of costs that, in some specific contexts, may be advantageous. Author summaryPatients with obsessive-compulsive disorder (OCD) report to suffer from indecisiveness and overly cautious decision making.Although many studies captured such a bias experimentally, little is known about the cognitive mechanisms driving such an indecisiveness.In this study, we investigated 16 juvenile OCD patients and compared their performance Tobias U. Hauser, Michael Moutoussis, Reto Iannaccone, Silvia Brem, Susanne Walitza, Renate Drechsler, Peter Dayan, Ray Dolan |
PLoS Comput. Biol. | 7 |
| 2016 | How People Use Social Information to Find out What to Want in the Paradigmatic Case of Inter-temporal PreferencesabstractThe weight with which a specific outcome feature contributes to preference quantifies a person's 'taste' for that feature. However, far from being fixed personality characteristics, tastes are plastic. They tend to align, for example, with those of others even if such conformity is not rewarded. We hypothesised that people can be uncertain about their tastes. Personal tastes are therefore uncertain beliefs. People can thus learn about them by considering evidence, such as the preferences of relevant others, and then performing Bayesian updating. If a person's choice variability reflects uncertainty, as in random-preference models, then a signature of Bayesian updating is that the degree of taste change should correlate with that person's choice variability. Temporal discounting coefficients are an important example of taste-for patience. These coefficients quantify impulsivity, have good psychometric properties and can change upon observing others' choices. We examined discounting preferences in a novel, large community study of 14-24 year olds. We assessed discounting behaviour, including decision variability, before and after participants observed another person's choices. We found good evidence for taste uncertainty and for Bayesian taste updating. First, participants displayed decision variability which was better accounted for by a random-taste than by a response-noise model. Second, apparent taste shifts were well described by a Bayesian model taking into account taste uncertainty and the relevance of social information. Our findings have important neuroscientific, clinical and developmental significance. Michael Moutoussis, Ray Dolan, Peter Dayan |
PLoS Comput. Biol. | 3 |
| 2015 | Staying afloat on Neurath's boat - Heuristics for sequential causal learning
Neil Bramley, Peter Dayan, David A. Lagnado |
CogSci | 2 |
| 2015 | Simple Plans or Sophisticated Habits? State, Transition and Learning Interactions in the Two-Step TaskabstractThe recently developed 'two-step' behavioural task promises to differentiate model-based from model-free reinforcement learning, while generating neurophysiologically-friendly decision datasets with parametric variation of decision variables. These desirable features have prompted its widespread adoption. Here, we analyse the interactions between a range of different strategies and the structure of transitions and outcomes in order to examine constraints on what can be learned from behavioural performance. The task involves a trade-off between the need for stochasticity, to allow strategies to be discriminated, and a need for determinism, so that it is worth subjects' investment of effort to exploit the contingencies optimally. We show through simulation that under certain conditions model-free strategies can masquerade as being model-based. We first show that seemingly innocuous modifications to the task structure can induce correlations between action values at the start of the trial and the subsequent trial events in such a way that analysis based on comparing successive trials can lead to erroneous conclusions. We confirm the power of a suggested correction to the analysis that can alleviate this problem. We then consider model-free reinforcement learning strategies that exploit correlations between where rewards are obtained and which actions have high expected value. These generate behaviour that appears model-based under these, and also more sophisticated, analyses. Exploiting the full potential of the two-step task as a tool for behavioural neuroscience requires an understanding of these issues. Thomas E. Akam, Rui Costa, Peter Dayan |
PLoS Comput. Biol. | 3 |
| 2015 | Monte Carlo Planning Method Estimates Planning Horizons during Interactive Social ExchangeabstractReciprocating interactions represent a central feature of all human exchanges. They have been the target of various recent experiments, with healthy participants and psychiatric populations engaging as dyads in multi-round exchanges such as a repeated trust task. Behaviour in such exchanges involves complexities related to each agent's preference for equity with their partner, beliefs about the partner's appetite for equity, beliefs about the partner's model of their partner, and so on. Agents may also plan different numbers of steps into the future. Providing a computationally precise account of the behaviour is an essential step towards understanding what underlies choices. A natural framework for this is that of an interactive partially observable Markov decision process (IPOMDP). However, the various complexities make IPOMDPs inordinately computationally challenging. Here, we show how to approximate the solution for the multi-round trust task using a variant of the Monte-Carlo tree search algorithm. We demonstrate that the algorithm is efficient and effective, and therefore can be used to invert observations of behavioural choices. We use generated behaviour to elucidate the richness and sophistication of interactive inference. Andreas Hula, P. Read Montague, Peter Dayan |
PLoS Comput. Biol. | 3 |
| 2015 | Tamping Ramping: Algorithmic, Implementational, and Computational Explanations of Phasic Dopamine Signals in the AccumbensabstractSubstantial evidence suggests that the phasic activity of dopamine neurons represents reinforcement learning's temporal difference prediction error. However, recent reports of ramp-like increases in dopamine concentration in the striatum when animals are about to act, or are about to reach rewards, appear to pose a challenge to established thinking. This is because the implied activity is persistently predictable by preceding stimuli, and so cannot arise as this sort of prediction error. Here, we explore three possible accounts of such ramping signals: (a) the resolution of uncertainty about the timing of action; (b) the direct influence of dopamine over mechanisms associated with making choices; and (c) a new model of discounted vigour. Collectively, these suggest that dopamine ramps may be explained, with only minor disturbance, by standard theoretical ideas, though urgent questions remain regarding their proximal cause. We suggest experimental approaches to disentangling which of the proposed mechanisms are responsible for dopamine ramps. Kevin Lloyd, Peter Dayan |
PLoS Comput. Biol. | 2 |
| 2015 | A Probabilistic Palimpsest Model of Visual Short-term MemoryabstractWorking memory plays a key role in cognition, and yet its mechanisms remain much debated. Human performance on memory tasks is severely limited; however, the two major classes of theory explaining the limits leave open questions about key issues such as how multiple simultaneously-represented items can be distinguished. We propose a palimpsest model, with the occurrent activity of a single population of neurons coding for several multi-featured items. Using a probabilistic approach to storage and recall, we show how this model can account for many qualitative aspects of existing experimental data. In our account, the underlying nature of a memory item depends entirely on the characteristics of the population representation, and we provide analytical and numerical insights into critical issues such as multiplicity and binding. We consider representations in which information about individual feature values is partially separate from the information about binding that creates single items out of multiple features. An appropriate balance between these two types of information is required to capture fully the different types of error seen in human experimental data. Our model provides the first principled account of misbinding errors. We also suggest a specific set of stimuli designed to elucidate the representations that subjects actually employ. Loïc Matthey, Paul M. Bays, Peter Dayan |
PLoS Comput. Biol. | 3 |
| 2015 | Anticipation and Choice Heuristics in the Dynamic Consumption of Pain ReliefabstractHumans frequently need to allocate resources across multiple time-steps. Economic theory proposes that subjects do so according to a stable set of intertemporal preferences, but the computational demands of such decisions encourage the use of formally less competent heuristics. Few empirical studies have examined dynamic resource allocation decisions systematically. Here we conducted an experiment involving the dynamic consumption over approximately 15 minutes of a limited budget of relief from moderately painful stimuli. We had previously elicited the participants' time preferences for the same painful stimuli in one-off choices, allowing us to assess self-consistency. Participants exhibited three characteristic behaviors: saving relief until the end, spreading relief across time, and early spending, of which the last was markedly less prominent. The likelihood that behavior was heuristic rather than normative is suggested by the weak correspondence between one-off and dynamic choices. We show that the consumption choices are consistent with a combination of simple heuristics involving early-spending, spreading or saving of relief until the end, with subjects predominantly exhibiting the last two. Giles W. Story, Ivo Vlaev, Peter Dayan, Ben Seymour, Ara Darzi, Ray Dolan |
PLoS Comput. Biol. | 3 |
| 2014 | Bayes-Adaptive Simulation-based Search with Value Function Approximation
Arthur Guez, Nicolas Heess, David Silver 0001, Peter Dayan |
NIPS | 4 |
| 2014 | Some Work and Some Play: Microscopic and Macroscopic Approaches to Labor and LeisureabstractGiven the option, humans and other animals elect to distribute their time between work and leisure, rather than choosing all of one and none of the other. Traditional accounts of partial allocation have characterised behavior on a macroscopic timescale, reporting and studying the mean times spent in work or leisure. However, averaging over the more microscopic processes that govern choices is known to pose tricky theoretical problems, and also eschews any possibility of direct contact with the neural computations involved. We develop a microscopic framework, formalized as a semi-Markov decision process with possibly stochastic choices, in which subjects approximately maximise their expected returns by making momentary commitments to one or other activity. We show macroscopic utilities that arise from microscopic ones, and demonstrate how facets such as imperfect substitutability can arise in a more straightforward microscopic manner. Ritwik K. Niyogi, Peter Shizgal, Peter Dayan |
PLoS Comput. Biol. | 3 |
| 2014 | Optimal Recall from Bounded Metaplastic Synapses: Predicting Functional Adaptations in Hippocampal Area CA3abstractA venerable history of classical work on autoassociative memory has significantly shaped our understanding of several features of the hippocampus, and most prominently of its CA3 area, in relation to memory storage and retrieval. However, existing theories of hippocampal memory processing ignore a key biological constraint affecting memory storage in neural circuits: the bounded dynamical range of synapses. Recent treatments based on the notion of metaplasticity provide a powerful model for individual bounded synapses; however, their implications for the ability of the hippocampus to retrieve memories well and the dynamics of neurons associated with that retrieval are both unknown. Here, we develop a theoretical framework for memory storage and recall with bounded synapses. We formulate the recall of a previously stored pattern from a noisy recall cue and limited-capacity (and therefore lossy) synapses as a probabilistic inference problem, and derive neural dynamics that implement approximate inference algorithms to solve this problem efficiently. In particular, for binary synapses with metaplastic states, we demonstrate for the first time that memories can be efficiently read out with biologically plausible network dynamics that are completely constrained by the synaptic plasticity rule, and the statistics of the stored patterns and of the recall cue. Our theory organises into a coherent framework a wide range of existing data about the regulation of excitability, feedback inhibition, and network oscillations in area CA3, and makes novel and directly testable predictions that can guide future experiments. Cristina Savin, Peter Dayan, Máté Lengyel |
PLoS Comput. Biol. | 2 |
| 2013 | Structured cognitive representations and complex inference in neural systems
Samuel Gershman, Josh Tenenbaum, Alexandre Pouget, Matt M. Botvinick, Peter Dayan |
CogSci | 5 |
| 2013 | Correlations strike back (again): the case of associative memory retrievalabstractIt has long been recognised that statistical dependencies in neuronal activity need to be taken into account when decoding stimuli encoded in a neural population. Less studied, though equally pernicious, is the need to take account of dependencies between synaptic weights when decoding patterns previously encoded in an auto-associative memory. We show that activity-dependent learning generically produces such correlations, and failing to take them into account in the dynamics of memory retrieval leads to catastrophically poor recall. We derive optimal network dynamics for recall in the face of synaptic correlations caused by a range of synaptic plasticity rules. These dynamics involve well-studied circuit motifs, such as forms of feedback inhibition and experimentally observed dendritic nonlinearities. We therefore show how addressing the problem of synaptic correlations leads to a novel functional account of key biophysical features of the neural substrate. Cristina Savin, Peter Dayan, Máté Lengyel |
NIPS | 2 |
| 2013 | Scalable and Efficient Bayes-Adaptive Reinforcement Learning Based on Monte-Carlo Tree SearchabstractBayesian planning is a formally elegant approach to learning optimal behaviour under model uncertainty, trading off exploration and exploitation in an ideal way. Unfortunately, planning optimally in the face of uncertainty is notoriously taxing, since the search space is enormous. In this paper we introduce a tractable, sample-based method for approximate Bayes-optimal planning which exploits Monte-Carlo tree search. Our approach avoids expensive applications of Bayes rule within the search tree by sampling models from current beliefs, and furthermore performs this sampling in a lazy manner. This enables it to outperform previous Bayesian model-based reinforcement learning algorithms by a significant margin on several well-known benchmark problems. As we show, our approach can even work in problems with an infinite state space that lie qualitatively out of reach of almost all previous work in Bayesian exploration. Arthur Guez, David Silver 0001, Peter Dayan |
J. Artif. Intell. Res. | 3 |
| 2013 | Informing the design of clinical decision support services for evaluation of children with minor blunt head trauma in the emergency department: A sociotechnical analysis
Barbara Sheehan, Lise E. Nigrovic, Peter Dayan, Nathan Kuppermann, Dustin W. Ballard, Evaline Alessandrini, Lalit Bajaj, Howard Goldberg, Jeffrey Hoffman, Steven R. Offerman, Dustin G. Mark, Marguerite Swietlik, Eric Tham, Leah Tzimenatos, David R. Vinson, Grant S. Jones, Suzanne Bakken |
J. Biomed. Informatics | 3 |
| 2013 | Sparse Coding Can Predict Primary Visual Cortex Receptive Field Changes Induced by Abnormal Visual InputabstractReceptive fields acquired through unsupervised learning of sparse representations of natural scenes have similar properties to primary visual cortex (V1) simple cell receptive fields. However, what drives in vivo development of receptive fields remains controversial. The strongest evidence for the importance of sensory experience in visual development comes from receptive field changes in animals reared with abnormal visual input. However, most sparse coding accounts have considered only normal visual input and the development of monocular receptive fields. Here, we applied three sparse coding models to binocular receptive field development across six abnormal rearing conditions. In every condition, the changes in receptive field properties previously observed experimentally were matched to a similar and highly faithful degree by all the models, suggesting that early sensory development can indeed be understood in terms of an impetus towards sparsity. As previously predicted in the literature, we found that asymmetries in inter-ocular correlation across orientations lead to orientation-specific binocular receptive fields. Finally we used our models to design a novel stimulus that, if present during rearing, is predicted by the sparsity principle to lead robustly to radically abnormal receptive fields. Jonathan J. Hunt, Peter Dayan, Geoffrey J. Goodhill |
PLoS Comput. Biol. | 2 |
| 2012 | Computational, Cognitive, and Neural Models of Decision-making Biases
Jonathan Malmaud, Josh Tenenbaum, Peter Dayan, Laurence T. Maloney, Ed Vul, Nick Chater |
CogSci | 3 |
| 2012 | Dynamic decision making: neuronal, computational, and cognitive underpinnings
Magda Osman, Maarten Speekenbrink, Peter Dayan, Masataka Watanabe, Nigel Harvey |
CogSci | 3 |
| 2012 | Efficient Bayes-Adaptive Reinforcement Learning using Sample-Based SearchabstractBayesian model-based reinforcement learning is a formally elegant approach to learning optimal behaviour under model uncertainty, trading off exploration and exploitation in an ideal way. Unfortunately, finding the resulting Bayes-optimal policies is notoriously taxing, since the search space becomes enormous. In this paper we introduce a tractable, sample-based method for approximate Bayes-optimal planning which exploits Monte-Carlo tree search. Our approach outperformed prior Bayesian model-based RL algorithms by a significant margin on several well-known benchmark problems -- because it avoids expensive applications of Bayes rule within the search tree by lazily sampling models from the current beliefs. We illustrate the advantages of our approach by showing it working in an infinite state space domain which is qualitatively out of reach of almost all previous work in Bayesian exploration. Arthur Guez, David Silver 0001, Peter Dayan |
NIPS | 3 |
| 2012 | Cortical Surround Interactions and Perceptual Salience via Natural Scene StatisticsabstractSpatial context in images induces perceptual phenomena associated with salience and modulates the responses of neurons in primary visual cortex (V1). However, the computational and ecological principles underlying contextual effects are incompletely understood. We introduce a model of natural images that includes grouping and segmentation of neighboring features based on their joint statistics, and we interpret the firing rates of V1 neurons as performing optimal recognition in this model. We show that this leads to a substantial generalization of divisive normalization, a computation that is ubiquitous in many neural areas and systems. A main novelty in our model is that the influence of the context on a target stimulus is determined by their degree of statistical dependence. We optimized the parameters of the model on natural image patches, and then simulated neural and perceptual responses on stimuli used in classical experiments. The model reproduces some rich and complex response patterns observed in V1, such as the contrast dependence, orientation tuning and spatial asymmetry of surround suppression, while also allowing for surround facilitation under conditions of weak stimulation. It also mimics the perceptual salience produced by simple displays, and leads to readily testable predictions. Our results provide a principled account of orientation-based contextual modulation in early vision and its sensitivity to the homogeneity and spatial arrangement of inputs, and lends statistical support to the theory that V1 computes visual salience. Ruben Coen Cagli, Peter Dayan, Odelia Schwartz |
PLoS Comput. Biol. | 2 |
| 2012 | Bonsai Trees in Your Head: How the Pavlovian System Sculpts Goal-Directed Choices by Pruning Decision TreesabstractWhen planning a series of actions, it is usually infeasible to consider all potential future sequences; instead, one must prune the decision tree. Provably optimal pruning is, however, still computationally ruinous and the specific approximations humans employ remain unknown. We designed a new sequential reinforcement-based task and showed that human subjects adopted a simple pruning strategy: during mental evaluation of a sequence of choices, they curtailed any further evaluation of a sequence as soon as they encountered a large loss. This pruning strategy was Pavlovian: it was reflexively evoked by large losses and persisted even when overwhelmingly counterproductive. It was also evident above and beyond loss aversion. We found that the tendency towards Pavlovian pruning was selectively predicted by the degree to which subjects exhibited sub-clinical mood disturbance, in accordance with theories that ascribe Pavlovian behavioural inhibition, via serotonin, a role in mood disorders. We conclude that Pavlovian behavioural inhibition shapes highly flexible, goal-directed choices in a manner that may be important for theories of decision-making in mood disorders. Quentin J. M. Huys, Neir Eshel, Elizabeth J. P. O'Nions, Luke Sheridan, Peter Dayan, Jonathan P. Roiser |
PLoS Comput. Biol. | 5 |
| 2012 | Computational Phenotyping of Two-Person Interactions Reveals Differential Neural Response to Depth-of-ThoughtabstractReciprocating exchange with other humans requires individuals to infer the intentions of their partners. Despite the importance of this ability in healthy cognition and its impact in disease, the dimensions employed and computations involved in such inferences are not clear. We used a computational theory-of-mind model to classify styles of interaction in 195 pairs of subjects playing a multi-round economic exchange game. This classification produces an estimate of a subject's depth-of-thought in the game (low, medium, high), a parameter that governs the richness of the models they build of their partner. Subjects in each category showed distinct neural correlates of learning signals associated with different depths-of-thought. The model also detected differences in depth-of-thought between two groups of healthy subjects: one playing patients with psychiatric disease and the other playing healthy controls. The neural response categories identified by this computational characterization of theory-of-mind may yield objective biomarkers useful in the identification and characterization of pathologies that perturb the capacity to model and interact with other humans. Ting Xiang, Debajyoti Ray, Terry Lohrenz, Peter Dayan, P. Read Montague |
PLoS Comput. Biol. | 4 |
| 2011 | Two is better than one: distinct roles for familiarity and recollection in retrieving palimpsest memoriesabstractStoring a new pattern in a palimpsest memory system comes at the cost of interfering with the memory traces of previously stored items. Knowing the age of a pattern thus becomes critical for recalling it faithfully. This implies that there should be a tight coupling between estimates of age, as a form of familiarity, and the neural dynamics of recollection, something which current theories omit. Using a normative model of autoassociative memory, we show that a dual memory system, consisting of two interacting modules for familiarity and recollection, has best performance for both recollection and recognition. This finding provides a new window onto actively contentious psychological and neural aspects of recognition memory. Cristina Savin, Peter Dayan, Máté Lengyel |
NIPS | 2 |
| 2011 | Bayes-Optimal ChemotaxisabstractChemotaxis plays a crucial role in many biological processes, including nervous system development. However, fundamental physical constraints limit the ability of a small sensing device such as a cell or growth cone to detect an external chemical gradient. One of these is the stochastic nature of receptor binding, leading to a constantly fluctuating binding pattern across the cell's array of receptors. This is analogous to the uncertainty in sensory information often encountered by the brain at the systems level. Here we derive analytically the Bayes-optimal strategy for combining information from a spatial array of receptors in both one and two dimensions to determine gradient direction. We also show how information from more than one receptor species can be optimally integrated, derive the gradient shapes that are optimal for guiding cells or growth cones over the longest possible distances, and illustrate that polarized cell behavior might arise as an adaptation to slowly varying environments. Together our results provide closed-form predictions for variations in chemotactic performance over a wide range of gradient conditions. Duncan Mortimer, Peter Dayan, Kevin Burrage, Geoffrey J. Goodhill |
Neural Comput. | 2 |
| 2011 | Disentangling the Roles of Approach, Activation and Valence in Instrumental and Pavlovian RespondingabstractHard-wired, Pavlovian, responses elicited by predictions of rewards and punishments exert significant benevolent and malevolent influences over instrumentally-appropriate actions. These influences come in two main groups, defined along anatomical, pharmacological, behavioural and functional lines. Investigations of the influences have so far concentrated on the groups as a whole; here we take the critical step of looking inside each group, using a detailed reinforcement learning model to distinguish effects to do with value, specific actions, and general activation or inhibition. We show a high degree of sophistication in Pavlovian influences, with appetitive Pavlovian stimuli specifically promoting approach and inhibiting withdrawal, and aversive Pavlovian stimuli promoting withdrawal and inhibiting approach. These influences account for differences in the instrumental performance of approach and withdrawal behaviours. Finally, although losses are as informative as gains, we find that subjects neglect losses in their instrumental learning. Our findings argue for a view of the Pavlovian system as a constraint or prior, facilitating learning by alleviating computational costs that come with increased flexibility. Quentin J. M. Huys, Roshan Cools, Martin Gölzer, Eva Friedel, Andreas Heinz, Ray Dolan, Peter Dayan |
PLoS Comput. Biol. | 7 |
| 2010 | Change-Based Inference in Attractor Nets: Linear AnalysisabstractOne standard interpretation of networks of cortical neurons is that they form dynamical attractors. Computations such as stimulus estimation are performed by mapping inputs to points on the networks' attractive manifolds. These points represent population codes for the stimulus values. However, this standard interpretation is hard to reconcile with the observation that the firing rates of such neurons constantly change following presentation of stimuli. We have recently suggested an alternative interpretation according to which computations are realized by systematic changes in the states of such networks over time. This way of performing computations is fast, accurate, readily learnable, and robust to various forms of noise. Here we analyze the computation of stimulus discrimination in this change-based setting, relating it directly to the computation of stimulus estimation in the conventional attractor-based view. We use a common linear approximation to compare the two methods and show that perfect performance at estimation implies chance performance at discrimination. Reza Moazzezi, Peter Dayan |
Neural Comput. | 2 |
| 2010 | Pavlovian-Instrumental Interaction in 'Observing Behavior'abstractSubjects typically choose to be presented with stimuli that predict the existence of future reinforcements. This so-called 'observing behavior' is evident in many species under various experimental conditions, including if the choice is expensive, or if there is nothing that subjects can do to improve their lot with the information gained. A recent study showed that the activities of putative midbrain dopamine neurons reflect this preference for observation in a way that appears to challenge the common prediction-error interpretation of these neurons. In this paper, we provide an alternative account according to which observing behavior arises from a small, possibly Pavlovian, bias associated with the operation of working memory. Ulrik R. Beierholm, Peter Dayan |
PLoS Comput. Biol. | 2 |
| 2009 | Statistical Models of Linear and Nonlinear Contextual Interactions in Early Visual ProcessingabstractA central hypothesis about early visual processing is that it represents inputs in a coordinate system matched to the statistics of natural scenes. Simple versions of this lead to Gabor-like receptive fields and divisive gain modulation from local surrounds; these have led to influential neural and psychological models of visual processing. However, these accounts are based on an incomplete view of the visual context surrounding each point. Here, we consider an approximate model of linear and non-linear correlations between the responses of spatially distributed Gabor-like receptive fields, which, when trained on an ensemble of natural scenes, unifies a range of spatial context effects. The full model accounts for neural surround data in primary visual cortex (V1), provides a statistical foundation for perceptual phenomena associated with Lis (2002) hypothesis that V1 builds a saliency map, and fits data on the tilt illusion. Ruben Coen Cagli, Peter Dayan, Odelia Schwartz |
NIPS | 2 |
| 2009 | Know Thy Neighbour: A Normative Theory of Synaptic DepressionabstractSynapses exhibit an extraordinary degree of short-term malleability, with release probabilities and effective synaptic strengths changing markedly over multiple timescales. From the perspective of a fixed computational operation in a network, this seems like a most unacceptable degree of added noise. We suggest an alternative theory according to which short term synaptic plasticity plays a normatively-justifiable role. This theory starts from the commonplace observation that the spiking of a neuron is an incomplete, digital, report of the analog quantity that contains all the critical information, namely its membrane potential. We suggest that one key task for a synapse is to solve the inverse problem of estimating the pre-synaptic membrane potential from the spikes it receives and prior expectations, as in a recursive filter. We show that short-term synaptic depression has canonical dynamics which closely resemble those required for optimal estimation, and that it indeed supports high quality estimation. Under this account, the local postsynaptic potential and the level of synaptic resources track the (scaled) mean and variance of the estimated presynaptic membrane potential. We make experimentally testable predictions for how the statistics of subthreshold membrane potential fluctuations and the form of spiking non-linearity should be related to the properties of short-term plasticity in any particular cell type. Jean-Pascal Pfister, Peter Dayan, Máté Lengyel |
NIPS | 2 |
| 2009 | Goal-directed control and its antipodes
Peter Dayan |
Neural Networks | 1 |
| 2008 | Load and Attentional BayesabstractSelective attention is a most intensively studied psychological phenomenon, rife with theoretical suggestions and schisms. A critical idea is that of limited capacity, the allocation of which has produced half a century's worth of conflict about such phenomena as early and late selection. An influential resolution of this debate is based on the notion of perceptual load (Lavie, 2005, TICS, 9: 75), which suggests that low-load, easy tasks, because they underuse the total capacity of attention, mandatorily lead to the processing of stimuli that are irrelevant to the current attentional set; whereas high-load, difficult tasks grab all resources for themselves, leaving distractors high and dry. We argue that this theory presents a challenge to Bayesian theories of attention, and suggest an alternative, statistical, account of key supporting data. Peter Dayan |
NIPS | 1 |
| 2008 | Psychiatry: Insights into depression through normative decision-making modelsabstractDecision making lies at the very heart of many psychiatric diseases. It is also a central theoretical concern in a wide variety of fields and has undergone detailed, in-depth, analyses. We take as an example Major Depressive Disorder (MDD), applying insights from a Bayesian reinforcement learning framework. We focus on anhedonia and helplessness. Helplessness—a core element in the conceptual- izations of MDD that has lead to major advances in its treatment, pharmacolog- ical and neurobiological understanding—is formalized as a simple prior over the outcome entropy of actions in uncertain environments. Anhedonia, which is an equally fundamental aspect of the disease, is related to the effective reward size. These formulations allow for the design of specific tasks to measure anhedonia and helplessness behaviorally. We show that these behavioral measures capture explicit, questionnaire-based cognitions. We also provide evidence that these tasks may allow classification of subjects into healthy and MDD groups based purely on a behavioural measure and avoiding any verbal reports. There are strong ties between decision making and psychiatry, with maladaptive decisions and be- haviors being very prominent in people with psychiatric disorders. Depression is classically seen as following life events such as divorces and job losses. Longitudinal studies, however, have revealed that a significant fraction of the stressors associated with depression do in fact follow MDD onset, and that they are likely due to maladaptive behaviors prominent in MDD (Kendler et al., 1999). Clinically effective ’talking’ therapies for MDD such as cognitive and dialectical behavior therapies (DeRubeis et al., 1999; Bortolotti et al., 2008; Gotlib and Hammen, 2002; Power, 2005) explicitly concentrate on altering patients’ maladaptive behaviors and decision making processes. Decision making is a promising avenue into psychiatry for at least two more reasons. First, it offers powerful analytical tools. Control problems related to decision making are prevalent in a huge diversity of fields, ranging from ecology to economics, computer science and engineering. These fields have produced well-founded and thoroughly characterized frameworks within which many issues in decision making can be framed. Here, we will focus on framing issues identified in psychiatric settings within a normative decision making framework. Its second major strength comes from its relationship to neurobiology, and particularly those neuro- modulatory systems which are powerfully affected by all major clinically effective pharmacothera- pies in psychiatry. The understanding of these systems has benefited significantly from theoretical accounts of optimal control such as reinforcement learning (Montague et al., 1996; Kapur and Rem- ington, 1996; Smith et al., 1999; Yu and Dayan, 2005; Dayan and Yu, 2006). Such accounts may be useful to identify in more specific terms the roles of the neuromodulators in psychiatry (Smith et al., 2004; Williams and Dayan, 2005; Moutoussis et al., 2008; Dayan and Huys, 2008). ∗[email protected], [email protected], [email protected]; www.gatsby.ucl.ac.uk/∼qhuys/pub.html Quentin J. M. Huys, Joshua T. Vogelstein, Peter Dayan |
NIPS | 3 |
| 2008 | Bayesian Model of Behaviour in Economic GamesabstractClassical Game Theoretic approaches that make strong rationality assumptions have difficulty modeling observed behaviour in Economic games of human subjects. We investigate the role of finite levels of iterated reasoning and non-selfish utility functions in a Partially Observable Markov Decision Process model that incorporates Game Theoretic notions of interactivity. Our generative model captures a broad class of characteristic behaviours in a multi-round Investment game. We invert the generative process for a recognition model that is used to classify 200 subjects playing an Investor-Trustee game against randomly matched opponents. Debajyoti Ray, Brooks King-Casas, P. Read Montague, Peter Dayan |
NIPS | 4 |
| 2008 | Encoding and Decoding Spikes for Dynamic StimuliabstractNaturally occurring sensory stimuli are dynamic. In this letter, we consider how spiking neural populations might transmit information about continuous dynamic stimulus variables. The combination of simple encoders and temporal stimulus correlations leads to a code in which information is not readily available to downstream neurons. Here, we explore a complex encoder that is paired with a simple decoder that allows representation and manipulation of the dynamic information in neural systems. The encoder we present takes the form of a biologically plausible recurrent spiking neural network where the output population recodes its inputs to produce spikes that are independently decodeable. We show that this network can be learned in a supervised manner by a simple local learning rule. Rama Natarajan, Quentin J. M. Huys, Peter Dayan, Richard S. Zemel |
Neural Comput. | 3 |
| 2008 | Serotonin, Inhibition, and Negative MoodabstractPavlovian predictions of future aversive outcomes lead to behavioral inhibition, suppression, and withdrawal. There is considerable evidence for the involvement of serotonin in both the learning of these predictions and the inhibitory consequences that ensue, although less for a causal relationship between the two. In the context of a highly simplified model of chains of affectively charged thoughts, we interpret the combined effects of serotonin in terms of pruning a tree of possible decisions, (i.e., eliminating those choices that have low or negative expected outcomes). We show how a drop in behavioral inhibition, putatively resulting from an experimentally or psychiatrically influenced drop in serotonin, could result in unexpectedly large negative prediction errors and a significant aversive shift in reinforcement statistics. We suggest an interpretation of this finding that helps dissolve the apparent contradiction between the fact that inhibition of serotonin reuptake is the first-line treatment of depression, although serotonin itself is most strongly linked with aversive rather than appetitive outcomes and predictions. Peter Dayan, Quentin J. M. Huys |
PLoS Comput. Biol. | 1 |
| 2007 | Hippocampal Contributions to Control: The Third WayabstractRecent experimental studies have focused on the specialization of different neural structures for different types of instrumental behavior. Recent theoretical work has provided normative accounts for why there should be more than one control system, and how the output of different controllers can be integrated. Two par- ticlar controllers have been identified, one associated with a forward model and the prefrontal cortex and a second associated with computationally simpler, habit- ual, actor-critic methods and part of the striatum. We argue here for the normative appropriateness of an additional, but so far marginalized control system, associ- ated with episodic memory, and involving the hippocampus and medial temporal cortices. We analyze in depth a class of simple environments to show that episodic control should be useful in a range of cases characterized by complexity and in- ferential noise, and most particularly at the very early stages of learning, long before habitization has set in. We interpret data on the transfer of control from the hippocampus to the striatum in the light of this hypothesis. Máté Lengyel, Peter Dayan |
NIPS | 2 |
| 2007 | Fast Population CodingabstractUncertainty coming from the noise in its neurons and the ill-posed nature of many tasks plagues neural computations. Maybe surprisingly, many studies show that the brain manipulates these forms of uncertainty in a probabilistically consistent and normative manner, and there is now a rich theoretical literature on the capabilities of populations of neurons to implement computations in the face of uncertainty. However, one major facet of uncertainty has received comparatively little attention: time. In a dynamic, rapidly changing world, data are only temporarily relevant. Here, we analyze the computational consequences of encoding stimulus trajectories in populations of neurons. For the most obvious, simple, instantaneous encoder, the correlations induced by natural, smooth stimuli engender a decoder that requires access to information that is nonlocal both in time and across neurons. This formally amounts to a ruinous representation. We show that there is an alternative encoder that is computationally and representationally powerful in which each spike contributes independent information; it is independently decodable, in other words. We suggest this as an appropriate foundation for understanding time-varying population codes. Furthermore, we show how adaptation to temporal stimulus statistics emerges directly from the demands of simple decoding. Quentin J. M. Huys, Richard S. Zemel, Rama Natarajan, Peter Dayan |
Neural Comput. | 4 |
| 2006 | Uncertainty, phase and oscillatory hippocampal recallabstractMany neural areas, notably, the hippocampus, show structured, dynamical, population behavior such as coordinated oscillations. It has long been observed that such oscillations provide a substrate for representing analog information in the firing phases of neurons relative to the underlying population rhythm. However, it has become increasingly clear that it is essential for neural populations to represent uncertainty about the information they capture, and the substantial recent work on neural codes for uncertainty has omitted any analysis of oscillatory systems. Here, we observe that, since neurons in an oscillatory network need not only fire once in each cycle (or even at all), uncertainty about the analog quantities each neuron represents by its firing phase might naturally be reported through the degree of concentration of the spikes that it fires. We apply this theory to memory in a model of oscillatory associative recall in hippocampal area CA3. Although it is not well treated in the literature, representing and manipulating uncertainty is fundamental to competent memory; our theory enables us to view CA3 as an effective uncertainty-aware, retrieval system. Máté Lengyel, Peter Dayan |
NIPS | 2 |
| 2006 | Images, Frames, and Connectionist HierarchiesabstractThe representation of hierarchically structured knowledge in systems using distributed patterns of activity is an abiding concern for the connectionist solution of cognitively rich problems. Here, we use statistical unsupervised learning to consider semantic aspects of structured knowledge representation. We meld unsupervised learning notions formulated for multilinear models with tensor product ideas for representing rich information. We apply the model to images of faces. Peter Dayan |
Neural Comput. | 1 |
| 2006 | Soft Mixer Assignment in a Hierarchical Generative Model of Natural Scene StatisticsabstractGaussian scale mixture models offer a top-down description of signal generation that captures key bottom-up statistical characteristics of filter responses to images. However, the pattern of dependence among the filters for this class of models is prespecified. We propose a novel extension to the gaussian scale mixture model that learns the pattern of dependence from observed inputs and thereby induces a hierarchical representation of these inputs. Specifically, we propose that inputs are generated by gaussian variables (modeling local filter structure), multiplied by a mixer variable that is assigned probabilistically to each input from a set of possible mixers. We demonstrate inference of both components of the generative model, for synthesized data and for different classes of natural images, such as a generic ensemble and faces. For natural images, the mixer variable assignments show invariances resembling those of complex cells in visual cortex; the statistics of the gaussian components of the model are in accord with the outputs of divisive normalization models. We also show how our model helps interrelate a wide range of models of image statistics and cortical processing. Odelia Schwartz, Terrence J. Sejnowski, Peter Dayan |
Neural Comput. | 3 |
| 2006 | The misbehavior of value and the discipline of the will
Peter Dayan, Yael Niv, Ben Seymour, Nathaniel D. Daw |
Neural Networks | 1 |
| 2006 | Pre-attentive visual selection
Zhaoping Li 0001, Peter Dayan |
Neural Networks | 2 |
| 2005 | Differential Priors for Elastic Nets
Miguel Á. Carreira-Perpiñán, Peter Dayan, Geoffrey J. Goodhill |
IDEAL | 2 |
| 2005 | Norepinephrine and Neural InterruptsabstractAngela J. Yu Center for Brain, Mind & Behavior Green Hall, Princeton University Princeton, NJ 08540, USA [email protected] Experimental data indicate that norepinephrine is critically involved in aspects of vigilance and attention. Previously, we considered the func- tion of this neuromodulatory system on a time scale of minutes and longer, and suggested that it signals global uncertainty arising from gross changes in environmental contingencies. However, norepinephrine is also known to be activated phasically by familiar stimuli in well- learned tasks. Here, we extend our uncertainty-based treatment of nore- pinephrine to this phasic mode, proposing that it is involved in the de- tection and reaction to state uncertainty within a task. This role of nore- pinephrine can be understood through the metaphor of neural interrupts. Peter Dayan, Angela J. Yu |
NIPS | 1 |
| 2005 | How fast to work: Response vigor, motivation and tonic dopamineabstractReinforcement learning models have long promised to unify computa- tional, psychological and neural accounts of appetitively conditioned be- havior. However, the bulk of data on animal conditioning comes from free-operant experiments measuring how fast animals will work for rein- forcement. Existing reinforcement learning (RL) models are silent about these tasks, because they lack any notion of vigor. They thus fail to ad- dress the simple observation that hungrier animals will work harder for food, as well as stranger facts such as their sometimes greater produc- tivity even when working for irrelevant outcomes such as water. Here, we develop an RL framework for free-operant behavior, suggesting that subjects choose how vigorously to perform selected actions by optimally balancing the costs and benefits of quick responding. Motivational states such as hunger shift these factors, skewing the tradeoff. This accounts normatively for the effects of motivation on response rates, as well as many other classic findings. Finally, we suggest that tonic levels of dopamine may be involved in the computation linking motivational state to optimal responding, thereby explaining the complex vigor-related ef- fects of pharmacological manipulation of dopamine. Yael Niv, Nathaniel D. Daw, Peter Dayan |
NIPS | 3 |
| 2005 | A Bayesian Framework for Tilt Perception and ConfidenceabstractThe misjudgement of tilt in images lies at the heart of entertaining visual illusions and rigorous perceptual psychophysics. A wealth of findings has attracted many mechanistic models, but few clear computational principles. We adopt a Bayesian approach to perceptual tilt estimation, showing how a smoothness prior offers a powerful way of addressing much confusing data. In particular, we faithfully model recent results showing that confidence in estimation can be systematically affected by the same aspects of images that affect bias. Confidence is central to Bayesian modeling approaches, and is applicable in many other perceptual domains. Perceptual anomalies and illusions, such as the misjudgements of motion and tilt evident in so many psychophysical experiments, have intrigued researchers for decades.13 A Bayesian view48 has been particularly influential in models of motion processing, treating such anomalies as the normative product of prior information (often statistically codifying Gestalt laws) with likelihood information from the actual scenes presented. Here, we expand the range of statistically normative accounts to tilt estimation, for which there are classes of results (on estimation confidence) that are so far not available for motion. The tilt illusion arises when the perceived tilt of a center target is misjudged (ie bias) in the presence of flankers. Another phenomenon, called Crowding, refers to a loss in the confidence (ie sensitivity) of perceived target tilt in the presence of flankers. Attempts have been made to formalize these phenomena quantitatively. Crowding has been modeled as compulsory feature pooling (ie averaging of orientations), ignoring spatial positions.9, 10 The tilt illusion has been explained by lateral interactions11, 12 in populations of orientationtuned units; and by calibration.13 However, most models of this form cannot explain a number of crucial aspects of the data. First, the geometry of the positional arrangement of the stimuli affects attraction versus repulsion in bias, as emphasized by Kapadia et al14 (figure 1A), and others.15, 16 Second, Solomon et al. recently measured bias and sensitivity simultaneously.11 The rich and surprising range of sensitivities, far from flat as a function of flanker angles (figure 1B), are outside the reach of standard models. Moreover, current explanations do not offer a computational account of tilt perception as the outcome of a normative inference process. Here, we demonstrate that a Bayesian framework for orientation estimation, with a prior favoring smoothness, can naturally explain a range of seemingly puzzling tilt data. We explicitly consider both the geometry of the stimuli, and the issue of confidence in the esti- (A) 6 5 4 3 2 1 0 -1 -2 Odelia Schwartz, Terrence J. Sejnowski, Peter Dayan |
NIPS | 3 |
| 2004 | Rate- and Phase-coded Autoassociative MemoryabstractAreas of the brain involved in various forms of memory exhibit patterns of neural activity quite unlike those in canonical computational models. We show how to use well-founded Bayesian probabilistic autoassociative recall to derive biologically reasonable neuronal dynamics in recurrently coupled models, together with appropriate values for parameters such as the membrane time constant and inhibition. We explicitly treat two cases. One arises from a standard Hebbian learning rule, and involves activity patterns that are coded by graded firing rates. The other arises from a spike timing dependent learning rule, and involves patterns coded by the phase of spike times relative to a coherent local field potential oscillation. Our model offers a new and more complete understanding of how neural dynamics may support autoassociation. Máté Lengyel, Peter Dayan |
NIPS | 2 |
| 2004 | Assignment of Multiplicative Mixtures in Natural ImagesabstractIn the analysis of natural images, Gaussian scale mixtures (GSM) have been used to account for the statistics of (cid:2)lter responses, and to inspire hi- erarchical cortical representational learning schemes. GSMs pose a crit- ical assignment problem, working out which (cid:2)lter responses were gen- erated by a common multiplicative factor. We present a new approach to solving this assignment problem through a probabilistic extension to the basic GSM, and show how to perform inference in the model using Gibbs sampling. We demonstrate the ef(cid:2)cacy of the approach on both synthetic and image data. Understanding the statistical structure of natural images is an important goal for visual neuroscience. Neural representations in early cortical areas decompose images (and likely other sensory inputs) in a way that is sensitive to sophisticated aspects of their probabilistic structure. This structure also plays a key role in methods for image processing and coding. A striking aspect of natural images that has re(cid:3)ections in both top-down and bottom-up modeling is coordination across nearby locations, scales, and orientations. From a top- down perspective, this structure has been modeled using what is known as a Gaussian Scale Mixture model (GSM).1(cid:150)3 GSMs involve a multi-dimensional Gaussian (each di- mension of which captures local structure as in a linear (cid:2)lter), multiplied by a spatialized collection of common hidden scale variables or mixer variables(cid:3) (which capture the coordi- nation). GSMs have wide implications in theories of cortical receptive (cid:2)eld development, eg the comprehensive bubbles framework of Hyv¤arinen.4 The mixer variables provide the top-down account of two bottom-up characteristics of natural image statistics, namely the ‘bowtie’ statistical dependency,5,6 and the fact that the marginal distributions of receptive (cid:2)eld-like (cid:2)lters have high kurtosis.7,8 In hindsight, these ideas also bear a close relation- ship with Ruderman and Bialek’s multiplicative bottom-up image analysis framework9 and statistical models for divisive gain control.6 Coordinated structure has also been addressed in other image work,10(cid:150)14 and in other domains such as speech15 and (cid:2)nance.16 Many approaches to the unsupervised speci(cid:2)cation of representations in early cortical areas rely on the coordinated structure.17(cid:150)21 The idea is to learn linear (cid:2)lters (eg modeling simple cells as in22,23), and then, based on the coordination, to (cid:2)nd combinations of these (perhaps non-linearly transformed) as a way of (cid:2)nding higher order (cid:2)lters (eg complex cells). One critical facet whose speci(cid:2)cation from data is not obvious is the neighborhood arrangement, ie which linear (cid:2)lters share which mixer variables. (cid:3)Mixer variables are also called mutlipliers, but are unrelated to the scales of a wavelet. Here, we suggest a method for (cid:2)nding the neighborhood based on Bayesian inference of the GSM random variables. In section 1, we consider estimating these components based on information from different-sized neighborhoods and show the modes of failure when inference is too local or too global. Based on these observations, in section 2 we propose an extension to the GSM generative model, in which the mixer variables can overlap prob- abilistically. We solve the neighborhood assignment problem using Gibbs sampling, and demonstrate the technique on synthetic data. In section 3, we apply the technique to image data. 1 GSM inference of Gaussian and mixer variables In a simple, n-dimensional, version of a GSM, (cid:2)lter responses l are synthesized y by mul- tiplying an n-dimensional Gaussian with values g = fg1 : : : gng, by a common mixer variable v. (1) We assume g are uncorrelated ((cid:27)2 along diagonal of the covariance matrix). For the ana- lytical calculations, we assume that v has a Rayleigh distribution: Odelia Schwartz, Terrence J. Sejnowski, Peter Dayan |
NIPS | 3 |
| 2004 | Inference, Attention, and Decision in a Bayesian Neural ArchitectureabstractWe study the synthesis of neural coding, selective attention and percep- tual decision making. A hierarchical neural architecture is proposed, which implements Bayesian integration of noisy sensory input and top- down attentional priors, leading to sound perceptual discrimination. The model offers an explicit explanation for the experimentally observed modulation that prior information in one stimulus feature (location) can have on an independent feature (orientation). The network's intermediate levels of representation instantiate known physiological properties of vi- sual cortical neurons. The model also illustrates a possible reconciliation of cortical and neuromodulatory representations of uncertainty. Angela J. Yu, Peter Dayan |
NIPS | 2 |
| 2004 | Probabilistic Computation in Spiking PopulationsabstractAs animals interact with their environments, they must constantly update estimates about their states. Bayesian models combine prior probabil- ities, a dynamical model and sensory evidence to update estimates op- timally. These models are consistent with the results of many diverse psychophysical studies. However, little is known about the neural rep- resentation and manipulation of such Bayesian information, particularly in populations of spiking neurons. We consider this issue, suggesting a model based on standard neural architecture and activations. We illus- trate the approach on a simple random walk example, and apply it to a sensorimotor integration task that provides a particularly compelling example of dynamic probabilistic computation. Bayesian models have been used to explain a gamut of experimental results in tasks which require estimates to be derived from multiple sensory cues. These include a wide range of psychophysical studies of perception;13 motor action;7 and decision-making.3, 5 Central to Bayesian inference is that computations are sensitive to uncertainties about afferent and efferent quantities, arising from ignorance, noise, or inherent ambiguity (e.g., the aperture problem), and that these uncertainties change over time as information accumulates and dissipates. Understanding how neurons represent and manipulate uncertain quantities is therefore key to understanding the neural instantiation of these Bayesian inferences. Most previous work on representing probabilistic inference in neural populations has fo- cused on the representation of static information.1, 12, 15 These encompass various strategies for encoding and decoding uncertain quantities, but do not readily generalize to real-world dynamic information processing tasks, particularly the most interesting cases with stim- uli changing over the same timescale as spiking itself.11 Notable exceptions are the re- cent, seminal, but, as we argue, representationally restricted, models proposed by Gold and Shadlen,5 Rao,10 and Deneve.4 In this paper, we first show how probabilistic information varying over time can be repre- sented in a spiking population code. Second, we present a method for producing spiking codes that facilitate further processing of the probabilistic information. Finally, we show the utility of this method by applying it to a temporal sensorimotor integration task. 1 TRAJECTORY ENCODING AND DECODING We assume that population spikes R(t) arise stochastically in relation to the trajectory X(t) of an underlying (but hidden) variable. We use RT and XT for the whole trajectory and spike trains respectively from time 0 to T . The spikes RT constitute the observations and are assumed to be probabilistically related to the signal by a tuning function f (X, i): P (R(i, T )|X(T )) f (X, i) (1) for the spike train of the ith neuron, with parameters i. Therefore, via standard Bayesian inference, RT determines a distribution over the hidden variable at time T , P (X(T )|RT ). We first consider a version of the dynamics and input coding that permits an analytical examination of the impact of spikes. Let X(t) follow a stationary Gaussian process such that the joint distribution P (X(t1), X(t2), . . . , X(tm)) is Gaussian for any finite collection of times, with a covariance matrix which depends on time differences: Ctt = c(|t - t |). Function c(|t|) controls the smoothness of the resulting random walks. Then, P (X(T )|RT ) p(X(T )) dX(T )P (R X(T ) T |X(T ))P (X(T )|X (T )) (2) where P (X(T )|X(T )) is the distribution over the whole trajectory X(T ) conditional on the value of X(T ) at its end point. If RT are a set of conditionally independent inhomoge- neous Poisson processes, we have P (RT |X(T )) f (X(t d f (X( ), i i ), i) exp - i i) , (3) where ti are the spike times of neuron i in RT . Let = [X(ti )] be the vector of stimulus positions at the times at which we observed a spike and = [(ti )] be the vector of spike positions. If the tuning functions are Gaussian f (X, i) exp(-(X - i)2/22) and sufficiently dense that d f (X, i i) is independent of X (a standard assumption in population coding), then P (RT |X(T )) exp(- - 2/22) and in Equation 2, we can marginalize out X(T ) except at the spike times ti : P (X(T )|RT ) p(X(T )) d exp -[, X(T )]T C-1 [, X(T )] - - 2 (4) 2 22 and C is the block covariance matrix between X(ti ), x(T ) at the spike times [tt ] and the final time T . This Gaussian integral has P (X(T )|RT ) N ((T ), (T )), with (T ) = CT t(Ctt + I2)-1 = k (T ) = CT T - kCtT (5) CT T is the T, T th element of the covariance matrix and CT t is similarly a row vector. The dependence in on past spike times is specified chiefly by the inverse covariance matrix, and acts as an effective kernel (k). This kernel is not stationary, since it depends on factors such as the local density of spiking in the spike train RT . For example, consider where X(t) evolves according to a diffusion process with drift: dX = -Xdt + dN (t) (6) where prevents it from wandering too far, N (t) is white Gaussian noise with mean zero and 2 variance. Figure 1A shows sample kernels for this process. Inspection of Figure 1A reveals some important traits. First, the monotonically decreasing kernel magnitude as the time span between the spike and the current time T grows matches the intuition that recent spikes play a more significant role in determining the posterior over X(T ). Second, the kernel is nearly exponential, with a time constant that depends on the time constant of the covariance function and the density of the spikes; two settings of these parameters produced the two groupings of kernels in the figure. Finally, the fully adaptive kernel k can be locally well approximated by a metronomic kernel k (shown in red in Figure 1A) that assumes regular spiking. This takes advantage of the general fact, indicated by the grouping of kernels, that the kernel depends weakly on the actual spike pattern, but strongly on the average rate. The merits of the metronomic kernel are that it is stationary and only depends on a single mean rate rather than the full spike train RT . It also justifies Kernels k and ks Variance ratio Full kernel A B D -0.5 -2 10 10 2 / 0 -4 2 5 Space 10 Kernel size (weight) 0 0.5 0 0.03 0.06 0.09 0.04 0.06 0.08 0.1 t-tspike Time C True stimulus and means Regular, stationary kernel E 0.5 -0.5 0 0 Space Space -0.5 0.5 0.03 0.04 0.05 0.06 0.07 0.08 0.09 0.1 0.03 0.04 0.05 0.06 0.07 0.08 0.09 0.1 Time Time Figure 1: Exact and approximate spike decoding with the Gaussian process prior. Spikes are shown in yellow, the true stimulus in green, and P (X(T )|RT ) in gray. Blue: exact inference with nonstationary and red: approximate inference with regular spiking. A Ker- nel samples for a diffusion process as defined by equations 5, 6. B, C: Mean and variance of the inference. D: Exact inference with full kernel k and E: approximation based on metronomic kernel k. (Equation 7). the form of decoder used for the network model in the next section.6 Figure 1D shows an example of how well Equation 5 specifies a distribution over X(t) through very few spikes. Finally, 1E shows a factorized approximation with the stationary kernel similar to that used by Hinton and Brown6 and in our recurrent network: ^ t P (X(t)|R(t)) f (X, kst j=0 j ij = exp(-E(X(t), R(t), t)), (7) i i) By design, the mean is captured very well, but not the variance, which in this example grows too rapidly for long interspike intervals (Figure 1B, C). Using a slower kernel im- proves performance on the variance, but at the expense of the mean. We thus turn to the net- work model with recurrent connections that are available to reinstate the spike-conditional characteristics of the full kernel. 2 NETWORK MODEL FORMULATION Above we considered how population spikes RT specify a distribution over X(T ). We now extend this to consider how interconnected populations of neurons can specify distributions over time-varying variables. We frame the problem and our approach in terms of a two-level network, connecting one population of neurons to another; this construction is intended to apply to any level of processing. The network maps input population spikes R(t) to output population spikes S(t), where input and output evolve over time. As with the input spikes, ST indicates the output spike trains from time 0 to T , and these output spikes are assumed to determine a distribution over a related hidden variable. For the recurrent and feedforward computation in the network, we start with the de- ceptively simple goal9 of producing output spikes in such a way that the distribution Q(X(T )|ST ) they imply over the same hidden variable X(T ) as the input, faithfully matches P (X(T )|RT ). This might seem a strange goal, since one could surely just lis- ten to the input spikes. However, in order for the output spikes to track the hidden variable, the dynamics of the interactions between the neurons must explicitly capture the dynamics of the process X(T ). Once this `identity mapping' problem has been solved, more general, complex computations can be performed with ease. We illustrate this on a multisensory integration task, tracking a hidden variable that depends on multiple sensory cues. The aim of the recurrent network is to take the spikes R(t) as inputs, and produce output spikes that capture the probabilistic dynamics. We proceed in two steps. We first consider the probabilistic decoding process which turns ST into Q(X(t)|ST ). Then we discuss the recurrent and feedforward processing that produce appropriate ST given this decoder. Note that this decoding process is not required for the network processing; it instead provides a computational objective for the spiking dynamics in the system. We use a simple log-linear decoder based on a spatiotemporal kernel:6 Q(X(T )|ST ) exp(-E(X(T ), ST , T )) , where (8) E(X, S T T , T ) = S(j, T - ) j =0 j (X, ) (9) is an energy function, and the spatiotemporal kernels are assumed separable: j(X, ) = gj(X)( ). The spatial kernel gj(X) is related to the receptive field f (X, j) of neuron j and the temporal kernel j(X, ) to k The dynamics of processing in the network follows a standard recurrent neural architecture for modeling cortical responses, in the case that network inputs R(t) and outputs S(t) are spikes. The effect of a spike on other neurons in the network is assumed to have some simple temporal dynamics, described here again by the temporal kernel ( ): ri(t) = T R(i, T - )( ) s S(j, T - )( ) =0 j (t) = T =0 where T is the extent of the kernel. The response of an output neuron is governed by a stochastic spiking rule, where the probability that neuron j spikes at time t is given by: P (S(j, t) = 1) = (uj(t)) = ( w v i ij ri(t) + k kj sk (t - 1)) (10) where () is the logistic function, and W and V are the feedforward and recurrent weights. If ( ) = exp(- ), then uj(T ) = (0)(Wj R(T ) + Vj S(T )) + (1)uj(T - 1); this corresponds to a discretization of the standard dynamics for the membrane potential of a leaky integrate-and-fire neuron: duj = -u dt j +WR+VS, where the leak is determined by the temporal kernel. The task of the network is to make Q(X(T )|ST ) of Equation 8 match P (X(T )|RT ) com- ing from one of the two models above (exact dynamic or approximate stationary kernel). We measure the discrepancy using the Kullback-Leibler (KL) divergence: J = KL [P (X(T )|R t T )||Q(X (T )|ST )] (11) and, as a proof of principle in the experiments below, find optimal W and V by minimizing the KL divergence J using back-propagation through time (BPTT). In or- der to implement this in the most straightforward way, we convert the stochastic spik- ing rule (Equation 10) to a deterministic rule via the mean-field assumption: Sj(t) = ( w v i ij ri(t) + k kj sk (t - 1)). The gradients are tedious, but can be neatly ex- pressed in a temporally recursive form. Note that our current focus in the system is on the representational capability of the system, rather than its learning. Our results establish that the system can faithfully represent the posterior distribution. We return to the issue of more plausible learning rules below. The resulting network can be seen as a dynamic spiking analogue of the recurrent network scheme of Pouget et al.:9 both methods formulate feedforward and recurrent connections so that a simple decoding of the output can match optimal but complex decoding applied to the inputs. A further advantage of the scheme proposed here is that it facilitates downstream processing of the probabilistic information, as the objective encourages the formation of distributions at the output that factorize across the units. 3 RELATED MODELS Ideas about the representation of probabilistic information in spiking neurons are in vogue. One treatment considers Poisson spiking in populations with regular tuning functions, as- suming that stimuli change slowly compared with the inter-spike intervals.8 This leads to a Kalman filter account with much formal similarity to the models of P (X(T )|RT ). However, because of the slow timescale, recurrent dynamics can be allowed to settle to an underlying attractor. In another approach, the spiking activity of either a single neuron4 or a pair of neurons5 is considered as reporting (logarithmic) probabilistic information about an underlying binary hypothesis. A third treatment proposes that a population of neurons directly represents the (logarithmic) probability over the state of a hidden Markov model.10 Our method is closely related to the latter two models. Like Deneve's4 we consider the transformation of input spikes to output spikes with a fixed assumed decoding scheme so that the dynamics of an underlying process is captured. Our decoding mechanism produces something like the predictive coding apparent in Deneve's scheme, except that here, a neu- ron may not need to spike not only if it itself has recently spiked and thereby conveyed the appropriate information; but also if one of its population neighbors has recently spiked. This is explicitly captured by the recurrent interactions among the population. Our scheme also resembles Rao's10 approach in that it involves population codes. Our representational scheme is more general, however, in that the spatiotemporal decoder defines the relation- ship between output spikes and Q(X(T )|ST ), whereas his method assumes a direct en- coding, with each output neuron's activity proportional to log Q(X(T )|ST ). Our decoder can produce such a direct encoding if the spatial and temporal kernels are delta functions, but other kernels permit coordination amongst the population to take into account temporal effects, and to produce higher fidelity in the output distribution. $('.dropdown-menu a.dropdown-toggle').on('click', function (e) { if (!$(this).next().hasClass('show')) { $(this).parents('.dropdown-menu').first().find('.show').removeClass("show"); } var $subMenu = $(this).next(".dropdown-menu"); $subMenu.toggleClass('show'); $(this).parents('li.nav-item.dropdown.show').on('hidden.bs.dropdown', function (e) { $('.dropdown-submenu .show').removeClass("show"); }); return false; }); Name Change Policy × Requests for name changes in the electronic proceedings will be accepted with no questions asked. However name changes may cause bibliographic tracking issues. Authors are asked to consider this carefully and discuss it with their co-authors prior to requesting a name change in the electronic proceedings. Use the "Report an Issue" link to request a name change. Report an Issue | Name Change Policy Do not remove: This comment is monitored to verify that the site is working properly Richard S. Zemel, Quentin J. M. Huys, Rama Natarajan, Peter Dayan |
NIPS | 4 |
| 2003 | Plasticity Kernels and Temporal StatisticsabstractComputational mysteries surround the kernels relating the magnitude and sign of changes in efficacy as a function of the time difference between pre- and post-synaptic activity at a synapse. One important idea34 is that kernels result from fil(cid:173) tering, ie an attempt by synapses to eliminate noise corrupting learning. This idea has hitherto been applied to trace learning rules; we apply it to experimentally-defined kernels, using it to reverse-engineer assumed signal statistics. We also extend it to consider the additional goal for filtering of weighting learning according to statistical surprise, as in the Z-score transform. This provides a fresh view of observed kernels and can lead to different, and more natural, signal statistics. Peter Dayan, Michael Häusser |
NIPS | 1 |
| 2003 | Dopamine Modulation in a Basal Ganglio-cortical Network Implements Saliency-based Gating of Working Memory
Aaron J. Gruber, Peter Dayan, Boris Gutkin, Sara A. Solla |
NIPS | 2 |
| 2003 | Doubly Distributional Population Codes: Simultaneous Representation of Uncertainty and MultiplicityabstractPerceptual inference fundamentally involves uncertainty, arising from noise in sensation and the ill-posed nature of many perceptual problems. Accurate perception requires that this uncertainty be correctly represented, manipulated, and learned about. The choices subjects make in various psychophysical experiments suggest that they do indeed take such uncertainty into account when making perceptual inferences, posing the question as to how uncertainty is represented in the activities of neuronal populations. Most theoretical investigations of population coding have ignored this issue altogether; the few existing proposals that address it do so in such a way that it is fatally conflated with another facet of perceptual problems that also needs correct handling: multiplicity (that is, the simultaneous presence of multiple distinct stimuli). We present and validate a more powerful proposal for the way that population activity may encode uncertainty, both distinctly from and simultaneously with multiplicity. Maneesh Sahani, Peter Dayan |
Neural Comput. | 2 |
| 2002 | Adaptation and Unsupervised LearningabstractAdaptation is a ubiquitous neural and psychological phenomenon, with a wealth of instantiations and implications. Although a basic form of plasticity, it has, bar some notable exceptions, attracted computational theory of only one main variety. In this paper, we study adaptation from the perspective of factor analysis, a paradigmatic technique of unsuper- vised learning. We use factor analysis to re-interpret a standard view of adaptation, and apply our new model to some recent data on adaptation in the domain of face discrimination. Peter Dayan, Maneesh Sahani, Gregoire Deback |
NIPS | 1 |
| 2002 | Replay, Repair and ConsolidationabstractA standard view of memory consolidation is that episodes are stored tem- porarily in the hippocampus, and are transferred to the neocortex through replay. Various recent experimental challenges to the idea of transfer, particularly for human memory, are forcing its re-evaluation. However, although there is independent neurophysiological evidence for replay, short of transfer, there are few theoretical ideas for what it might be doing. We suggest and demonstrate two important computational roles associated with neocortical indices. Szabolcs Káli, Peter Dayan |
NIPS | 2 |
| 2002 | Expected and Unexpected Uncertainty: ACh and NE in the NeocortexabstractInference and adaptation in noisy and changing, rich sensory environ- ments are rife with a variety of specific sorts of variability. Experimental and theoretical studies suggest that these different forms of variability play different behavioral, neural and computational roles, and may be reported by different (notably neuromodulatory) systems. Here, we re- fine our previous theory of acetylcholine’s role in cortical inference in the (oxymoronic) terms of expected uncertainty, and advocate a theory for norepinephrine in terms of unexpected uncertainty. We suggest that norepinephrine reports the radical divergence of bottom-up inputs from prevailing top-down interpretations, to influence inference and plasticity. We illustrate this proposal using an adaptive factor analysis model. Angela J. Yu, Peter Dayan |
NIPS | 2 |
| 2002 | Structure in the Space of Value Functions
David J. Foster, Peter Dayan |
Mach. Learn. | 2 |
| 2002 | Opponent interactions between serotonin and dopamine
Nathaniel D. Daw, Sham M. Kakade, Peter Dayan |
Neural Networks | 3 |
| 2002 | Introduction for 2002 Special Issue: Computational Models of Neuromodulation
Kenji Doya, Peter Dayan, Michael E. Hasselmo |
Neural Networks | 2 |
| 2002 | Dopamine: generalization and bonuses
Sham M. Kakade, Peter Dayan |
Neural Networks | 2 |
| 2002 | Acetylcholine in cortical inference
Angela J. Yu, Peter Dayan |
Neural Networks | 2 |
| 2001 | Motivated Reinforcement LearningabstractThe standard reinforcement learning view of the involvement of neuromodulatory systems in instrumental conditioning in(cid:173) cludes a rather straightforward conception of motivation as prediction of sum future reward. Competition between actions is based on the motivating characteristics of their consequent states in this sense. Substantial, careful, experiments reviewed in Dickinson & Balleine, 12,13 into the neurobiology and psychol(cid:173) ogy of motivation shows that this view is incomplete. In many cases, animals are faced with the choice not between many dif(cid:173) ferent actions at a given state, but rather whether a single re(cid:173) sponse is worth executing at all. Evidence suggests that the motivational process underlying this choice has different psy(cid:173) chological and neural properties from that underlying action choice. We describe and model these motivational systems, and consider the way they interact. Peter Dayan |
NIPS | 1 |
| 2001 | ACh, Uncertainty, and Cortical InferenceabstractAcetylcholine (ACh) has been implicated in a wide variety of tasks involving attentional processes and plasticity. Following extensive animal studies, it has previously been suggested that ACh reports on uncertainty and controls hippocampal, cortical and cortico-amygdalar plasticity. We extend this view and consider its effects on cortical representational inference, arguing that ACh controls the balance between bottom-up inference, in(cid:3)uenced by input stimuli, and top-down inference, in(cid:3)uenced by contextual information. We illustrate our proposal using a hierarchical hid- den Markov model. Peter Dayan, Angela J. Yu |
NIPS | 1 |
| 2001 | A familiarity-based learning procedure for the establishment of place fields in area CA3 of the rat hippocampus
Szabolcs Káli, Peter Dayan |
Neurocomputing | 2 |
| 2000 | Competition and Arbors in Ocular DominanceabstractHebbian and competitive Hebbian algorithms are almost ubiquitous in modeling pattern formation in cortical development. We analyse in the(cid:173) oretical detail a particular model (adapted from Piepenbrock & Ober(cid:173) mayer, 1999) for the development of Id stripe-like patterns, which places competitive and interactive cortical influences, and free and restricted ini(cid:173) tial arborisation onto a common footing. Peter Dayan |
NIPS | 1 |
| 2000 | Explaining Away in Weight SpaceabstractExplaining away has mostly been considered in terms of inference of states in belief networks. We show how it can also arise in a Bayesian context in inference about the weights governing relationships such as those between stimuli and reinforcers in conditioning experiments such as bacA,'Ward blocking. We show how explaining away in weight space can be accounted for using an extension of a Kalman filter model; pro(cid:173) vide a new approximate way of looking at the Kalman gain matrix as a whitener for the correlation matrix of the observation process; suggest a network implementation of this whitener using an architecture due to Goodall; and show that the resulting model exhibits backward blocking. Peter Dayan, Sham M. Kakade |
NIPS | 1 |
| 2000 | Dopamine BonusesabstractSubstantial data support a temporal difference (TO) model of dopamine (OA) neuron activity in which the cells provide a global error signal for reinforcement learning. However, in certain cir(cid:173) cumstances, OA activity seems anomalous under the TO model, responding to non-rewarding stimuli. We address these anoma(cid:173) lies by suggesting that OA cells multiplex information about re(cid:173) ward bonuses, including Sutton's exploration bonuses and Ng et al's non-distorting shaping bonuses. We interpret this additional role for OA in terms of the unconditional attentional and psy(cid:173) chomotor effects of dopamine, having the computational role of guiding exploration. Sham M. Kakade, Peter Dayan |
NIPS | 2 |
| 2000 | Hippocampally-Dependent Consolidation in a Hierarchical Model of NeocortexabstractIn memory consolidation, declarative memories which initially require the hippocampus for their recall, ultimately become independent of it. Consolidation has been the focus of numerous experimental and qualita(cid:173) tive modeling studies, but only little quantitative exploration. We present a consolidation model in which hierarchical connections in the cortex, that initially instantiate purely semantic information acquired through probabilistic unsupervised learning, come to instantiate episodic infor(cid:173) mation as well. The hippocampus is responsible for helping complete partial input patterns before consolidation is complete, while also train(cid:173) ing the cortex to perform appropriate completion by itself. Szabolcs Káli, Peter Dayan |
NIPS | 2 |
| 2000 | Position Variance, Recurrence and Perceptual LearningabstractStimulus arrays are inevitably presented at different positions on the retina in visual tasks, even those that nominally require fixation. In par(cid:173) ticular, this applies to many perceptual learning tasks. We show that per(cid:173) ceptual inference or discrimination in the face of positional variance has a structurally different quality from inference about fixed position stimuli, involving a particular, quadratic, non-linearity rather than a purely lin(cid:173) ear discrimination. We show the advantage taking this non-linearity into account has for discrimination, and suggest it as a role for recurrent con(cid:173) nections in area VI, by demonstrating the superior discrimination perfor(cid:173) mance of a recurrent network. We propose that learning the feedforward and recurrent neural connections for these tasks corresponds to the fast and slow components of learning observed in perceptual learning tasks. Zhaoping Li 0001, Peter Dayan |
NIPS | 2 |
| 1999 | Acquisition in Autoshaping
Sham M. Kakade, Peter Dayan |
NIPS | 2 |
| 1999 | The Effect of Correlated Variability on the Accuracy of a Population CodeabstractWe study the impact of correlated neuronal firing rate variability on the accuracy with which an encoded quantity can be extracted from a population of neurons. Contrary to widespread belief, correlations in the variabilities of neuronal firing rates do not, in general, limit the increase in coding accuracy provided by using large populations of encoding neurons. Furthermore, in some cases, but not all, correlations improve the accuracy of a population code. L. F. Abbott, Peter Dayan |
Neural Comput. | 2 |
| 1999 | Recurrent Sampling Models for the Helmholtz MachineabstractMany recent analysis-by-synthesis density estimation models of cortical learning and processing have made the crucial simplifying assumption that units within a single layer are mutually independent given the states of units in the layer below or the layer above. In this article, we suggest using either a Markov random field or an alternative stochastic sampling architecture to capture explicitly particular forms of dependence within each layer. We develop the architectures in the context of real and binary Helmholtz machines. Recurrent sampling can be used to capture correlations within layers in the generative or the recognition models, and we also show how these can be combined. Peter Dayan |
Neural Comput. | 1 |
| 1998 | Computational Differences between Asymmetrical and Symmetrical Networks
Zhaoping Li 0001, Peter Dayan |
NIPS | 2 |
| 1998 | Distributional Population Codes and Multiple Motion Models
Richard S. Zemel, Peter Dayan |
NIPS | 2 |
| 1998 | Analytical Mean Squared Error Curves for Temporal Difference Learning
Satinder Singh 0001, Peter Dayan |
Mach. Learn. | 2 |
| 1998 | A Hierarchical Model of Binocular RivalryabstractBinocular rivalry is the alternating percept that can result when the two eyes see different scenes. Recent psychophysical evidence supports the notion that some aspects of binocular rivalry bear functional similarities to other bistable percepts. We build a model based on the hypothesis (Logothetis & Schall, 1989; Leopold & Logothetis, 1996; Logothetis, Leopold & Sheinberg, 1996) that alternation can be generated by competition between top-down cortical explanations for the inputs, rather than by direct competition between the inputs. Recent neurophysiological evidence shows that some binocular neurons are modulated with the changing percept; others are not, even if they are selective between the stimuli presented to the eyes. We extend our model to a hierarchy to address these effects. Peter Dayan |
Neural Comput. | 1 |
| 1998 | Probabilistic Interpretation of Population CodesabstractWe present a general encoding-decoding framework for interpreting the activity of a population of units. A standard population code interpretation method, the Poisson model, starts from a description as to how a single value of an underlying quantity can generate the activities of each unit in the population. In casting it in the encoding-decoding framework, we find that this model is too restrictive to describe fully the activities of units in population codes in higher processing areas, such as the medial temporal area. Under a more powerful model, the population activity can convey information not only about a single value of some quantity but also about its whole distribution, including its variance, and perhaps even the certainty the system has in the actual presence in the world of the entity generating this quantity. We propose a novel method for forming such probabilistic interpretations of population codes and compare it to the existing method. Richard S. Zemel, Peter Dayan, Alexandre Pouget |
Neural Comput. | 2 |
| 1998 | Bayesian retrieval in associative memories with storage errorsabstractIt is well known that for finite-sized networks, onestep retrieval in the autoassociative Willshaw net is a suboptimal way to extract the information stored in the synapses. Iterative retrieval strategies are much better, but have hitherto only had heuristic justification. We show how they emerge naturally from considerations of probabilistic inference under conditions of noisy and partial input and a corrupted weight matrix. We start from the conditional probability distribution over possible patterns for retrieval. This contains all possible information that is available to an observer of the network and the initial input. Since this distribution is over exponentially many patterns, we use it to develop two approximate, but tractable, iterative retrieval methods. One performs maximum likelihood inference to find the single most likely pattern, using the (negative log of the) conditional probability as a Lyapunov function for retrieval. In physics terms, if storage errors are present, then the modified iterative update equations contain an additional antiferromagnetic interaction term and site dependent threshold values. The second method makes a mean field assumption to optimize a tractable estimate of the full conditional probability distribution. This leads to iterative mean field equations which can be interpreted in terms of a network of neurons with sigmoidal responses but with the same interactions and thresholds as in the maximum likelihood update equations. In the absence of storage errors, both models become very similiar to the Willshaw model, where standard retrieval is iterated using a particular form of linear threshold strategy. Friedrich T. Sommer, Peter Dayan |
IEEE Trans. Neural Networks | 2 |
| 1997 | Combining Probabilistic Population Codes
Richard S. Zemel, Peter Dayan |
IJCAI | 2 |
| 1997 | Statistical Models of Conditioning
Peter Dayan, Theresa Long |
NIPS | 1 |
| 1997 | Hippocampal Model of Rat Spatial Abilities Using Temporal Difference Learning
David J. Foster, Richard G. M. Morris, Peter Dayan |
NIPS | 3 |
| 1997 | Using Expectation-Maximization for Reinforcement LearningabstractWe discuss Hinton's (1989) relative payoff procedure (RPP), a static reinforcement learning algorithm whose foundation is not stochastic gradient ascent. We show circumstances under which applying the RPP is guaranteed to increase the mean return, even though it can make large changes in the values of the parameters. The proof is based on a mapping between the RPP and a form of the expectation-maximization procedure of Dempster, Laird, and Rubin (1977). Peter Dayan, Geoffrey E. Hinton |
Neural Comput. | 1 |
| 1997 | Factor Analysis Using Delta-Rule Wake-Sleep LearningabstractWe describe a linear network that models correlations between real-valued visible variables using one or more real-valued hidden variables-a factor analysis model. This model can be seen as a linear version of the Helmholtz machine, and its parameters can be learned using the wake-sleep method, in which learning of the primary generative model is assisted by a recognition model, whose role is to fill in the values of hidden variables based on the values of visible variables. The generative and recognition models are jointly learned in wake and sleep phases, using just the delta rule. This learning procedure is comparable in simplicity to Hebbian learning, which produces a somewhat different representation of correlations in terms of principal components. We argue that the simplicity of wake-sleep learning makes factor analysis a plausible alternative to Hebbian learning as a model of activity-dependent cortical plasticity. Radford M. Neal, Peter Dayan |
Neural Comput. | 2 |
| 1997 | Modeling the manifolds of images of handwritten digitsabstractThis paper describes two new methods for modeling the manifolds of digitized images of handwritten digits. The models allow a priori information about the structure of the manifolds to be combined with empirical data. Accurate modeling of the manifolds allows digits to be discriminated using the relative probability densities under the alternative models. One of the methods is grounded in principal components analysis, the other in factor analysis. Both methods are based on locally linear low-dimensional approximations to the underlying data manifold. Links with other methods that model the manifold are discussed. Geoffrey E. Hinton, Peter Dayan, Michael Revow |
IEEE Trans. Neural Networks | 2 |
| 1996 | A Hierarchical Model of Visual Rivalry
Peter Dayan |
NIPS | 1 |
| 1996 | Neural Models for Part-Whole Hierarchies
Maximilian Riesenhuber, Peter Dayan |
NIPS | 2 |
| 1996 | Analytical Mean Squared Error Curves in Temporal Difference Learning
Satinder Singh 0001, Peter Dayan |
NIPS | 2 |
| 1996 | Probabilistic Interpretation of Population Codes
Richard S. Zemel, Peter Dayan, Alexandre Pouget |
NIPS | 2 |
| 1996 | Exploration Bonuses and Dual Control
Peter Dayan, Terrence J. Sejnowski |
Mach. Learn. | 1 |
| 1996 | Varieties of Helmholtz Machine
Peter Dayan, Geoffrey E. Hinton |
Neural Networks | 1 |
| 1995 | Predictive Hebbian Learningabstractestablished from the perspective of psychological experiments, the neural mechanisms that underlie this pre- Terrence J. Sejnowski, Peter Dayan, P. Read Montague |
COLT | 2 |
| 1995 | Improving Policies without Measuring Merits
Peter Dayan, Satinder Singh 0001 |
NIPS | 1 |
| 1995 | Does the Wake-sleep Algorithm Produce Good Density Estimators?
Brendan J. Frey, Geoffrey E. Hinton, Peter Dayan |
NIPS | 3 |
| 1995 | The Helmholtz machineabstractDiscovering the structure inherent in a set of patterns is a fundamental aim of statistical inference or learning. One fruitful approach is to build a parameterized stochastic generative model, independent draws from which are likely to produce the patterns. For all but the simplest generative models, each pattern can be generated in exponentially many ways. It is thus intractable to adjust the parameters to maximize the probability of the observed patterns. We describe a way of finessing this combinatorial explosion by maximizing an easily computed lower bound on the probability of the observations. Our method can be viewed as a form of hierarchical self-supervised learning that may relate to the function of bottom-up and top-down cortical processing pathways. Peter Dayan, Geoffrey E. Hinton, Radford M. Neal, Richard S. Zemel |
Neural Comput. | 1 |
| 1995 | Competition and Multiple Cause ModelsabstractIf different causes can interact on any occasion to generate a set of patterns, then systems modeling the generation have to model the interaction too. We discuss a way of combining multiple causes that is based on the Integrated Segmentation and Recognition architecture of Keeler et al. (1991). It is more cooperative than the scheme embodied in the mixture of experts architecture, which insists that just one cause generate each output, and more competitive than the noisy-or combination function, which was recently suggested by Saund (1994a,b). Simulations confirm its efficacy. Peter Dayan, Richard S. Zemel |
Neural Comput. | 1 |
| 1994 | Recognizing Handwritten Digits Using Mixtures of Linear ModelsabstractWe construct a mixture of locally linear generative models of a col(cid:173) lection of pixel-based images of digits, and use them for recogni(cid:173) tion. Different models of a given digit are used to capture different styles of writing, and new images are classified by evaluating their log-likelihoods under each model. We use an EM-based algorithm in which the M-step is computationally straightforward principal components analysis (PCA). Incorporating tangent-plane informa(cid:173) tion [12] about expected local deformations only requires adding tangent vectors into the sample covariance matrices for the PCA, and it demonstrably improves performance. Geoffrey E. Hinton, Michael Revow, Peter Dayan |
NIPS | 3 |
| 1994 | TD(lambda) Converges with Probability 1
Peter Dayan, Terrence J. Sejnowski |
Mach. Learn. | 1 |
| 1993 | Foraging in an Uncertain Environment Using Predictive Hebbian Learning
P. Read Montague, Peter Dayan, Terrence J. Sejnowski |
NIPS | 2 |
| 1993 | Temporal Difference Learning of Position Evaluation in the Game of Go
Nicol N. Schraudolph, Peter Dayan, Terrence J. Sejnowski |
NIPS | 2 |
| 1993 | Arbitrary Elastic Topologies and Ocular DominanceabstractThe elastic net, which has been used to produce accounts of the formation of topology-preserving maps and ocular dominance stripes (OD), embodies a nearest neighbor topology. A Hebbian account of OD is not so restricted—and indeed makes the prediction that the width of the stripes depends on the nature of the (more general) neighborhood relations. Elastic and Hebbian accounts have recently been unified, raising a question mark about their different determiners of stripe widths. This paper considers this issue, and demonstrates theoretically that it is possible to use more general topologies in the elastic net, including those effectively adopted in the Hebbian model. Peter Dayan |
Neural Comput. | 1 |
| 1993 | Improving Generalization for Temporal Difference Learning: The Successor RepresentationabstractEstimation of returns over time, the focus of temporal difference (TD) algorithms, imposes particular constraints on good function approximators or representations. Appropriate generalization between states is determined by how similar their successors are, and representations should follow suit. This paper shows how TD machinery can be used to learn such representations, and illustrates, using a navigation task, the appropriately distributed nature of the result. Peter Dayan |
Neural Comput. | 1 |
| 1993 | The Variance of Covariance Rules for Associative Matrix Memories and Reinforcement LearningabstractHebbian synapses lie at the heart of most associative matrix memories (Kohonen 1987; Hinton and Anderson 1981) and are also biologically plausible (Brown et al. 1990; Baudry and Davis 1991). Their analytical and computational tractability make these memories the best understood form of distributed information storage. A variety of Hebbian algorithms for estimating the covariance between input and output patterns has been proposed. This note points out that one class of these involves stochastic estimation of the covariance, shows that the signal-to-noise ratios of the rules are governed by the variances of their estimates, and considers some parallels in reinforcement learning. Peter Dayan, Terrence J. Sejnowski |
Neural Comput. | 1 |
| 1992 | Feudal Reinforcement Learning
Peter Dayan, Geoffrey E. Hinton |
NIPS | 1 |
| 1992 | Using Aperiodic Reinforcement for Directed Self-Organization During Development
P. Read Montague, Peter Dayan, Steven J. Nowlan, Terrence J. Sejnowski |
NIPS | 2 |
| 1992 | The Convergence of TD(lambda) for General lambda
Peter Dayan |
Mach. Learn. | 1 |
| 1992 | Technical Note Q-Learning
Christopher J. C. H. Watkins, Peter Dayan |
Mach. Learn. | 2 |
| 1991 | Perturbing Hebbian Rules
Peter Dayan, Geoffrey J. Goodhill |
NIPS | 1 |
| 1990 | Navigating Through Temporal Difference
Peter Dayan |
NIPS | 1 |
| 1990 | Optimal Plasticity from Matrix Memories: What Goes Up Must Come DownabstractA recent article (Stanton and Sejnowski 1989) on long-term synaptic depression in the hippocampus has reopened the issue of the computational efficiency of particular synaptic learning rules (Hebb 1949; Palm 1988a; Morris and Willshaw 1989) — homosynaptic versus heterosynaptic and monotonic versus nonmonotonic changes in synaptic efficacy. We have addressed these questions by calculating and maximizing the signal-to-noise ratio, a measure of the potential fidelity of recall, in a class of associative matrix memories. Up to a multiplicative constant, there are three optimal rules, each providing for synaptic depression such that positive and negative changes in synaptic efficacy balance out. For one rule, which is found to be the Stent-Singer rule (Stent 1973; Rauschecker and Singer 1979), the depression is purely heterosynaptic; for another (Stanton and Sejnowski 1989), the depression is purely homosynaptic; for the third, which is a generalization of the first two, and has a higher signal-to-noise ratio, it is both heterosynaptic and homosynaptic. The third rule takes the form of a covariance rule (Sejnowski 1977a,b) and includes, as a special case, the prescription due to Hopfield (1982) and others (Willshaw 1971; Kohonen 1972). David J. Willshaw, Peter Dayan |
Neural Comput. | 2 |