Sander M. Bohté

dblp:15/5737 · DBLP profile ↗
← Back
50ranked-venue papers
17as first author
13since 2021 · last 2025
0000-0002-7866-278XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 37 · 15 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-authorTheory of computation · 2 · 2 first-authorComputer networks · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2025 Curriculum Design for Scalable Biologically Plausible Deep Reinforcement Learning
abstract
Humans have a remarkable capacity for learning, yet neuronal learning is constrained to locality in time and space and limited feedback. While neural learning rules have been designed that adhere to these principles and constraints, they exhibit difficulty in scaling to deep networks and complicated datasets. BrainProp is a biologically plausible learning rule, learning from trial-and-error feedback through reinforcement learning, that does generalise to deep networks and achieves good performance on traditional machine learning benchmarks. It does however falter on problems with a large number of output categories, such as the classical ImageNet vision benchmark: while standard BrainProp eventually succeeds, learning is not robust and highly sensitive to hyper-parameter optimisation and proper initialisation. Here, we leverage insights from behavioural science by developing a curriculum that structures how samples are presented to a network to optimise learning. The key features of the curriculum involve progressively introducing new classes to the dataset based on performance metrics, and using a recency bias to protect recently acquired classes. We demonstrate that our curriculum approach makes BrainProp-style learning robust and more rapid, while substantially improving classification accuracy. We also show the curriculum similarly improves performance for networks trained using error-backpropagation. We thus establish a new state-of-the-art performance for large-scale deep reinforcement learning. Our results show the potential of curriculum learning in local learning settings with limited feedback and further bridges the gap between biologically plausible learning rules and error-backpropagation.
Alexandra R. Van Den Berg, Pieter R. Roelfsema, Sander M. Bohté
IJCNN3
2025 Energy optimization induces predictive-coding properties in a multi-compartment spiking neural network model
abstract
Predictive coding is a prominent theoretical framework for understanding hierarchical sensory processing in the brain, yet how it could be implemented in networks of cortical neurons is still unclear. While most existing studies have taken a hand-wiring approach to creating microcircuits that match experimental results, recent work in rate-based artificial neural networks revealed that suitable cortical connectivity might result from self-organisation given some fundamental computational principle, such as energy efficiency. As no corresponding approach has studied this in more plausible networks of spiking neurons, we here investigate whether predictive coding properties in a multi-compartment spiking neural network can emerge from energy optimisation. We find that a model trained with an energy objective in addition to a task-relevant objective is able to reconstruct internal representations given top-down expectation signals alone. Additionally, neurons in the energy-optimised model show differential responses to expected versus unexpected stimuli, qualitatively similar to experimental evidence for predictive coding. These findings indicate that predictive-coding-like behaviour might be an emergent property of energy optimisation, providing a new perspective on how predictive coding could be achieved in the cortex.
Mingfang Zhang 0006, Raluca Chitic, Sander M. Bohté
PLoS Comput. Biol.3
2024 Masked Image Modeling as a Framework for Self-Supervised Learning Across Eye Movements
Robin Weiler, Matthias Brucklacher, Cyriel M. A. Pennartz, Sander M. Bohté
ICANN (4)4
2024 Balanced Resonate-and-Fire Neurons
abstract
The resonate-and-fire (RF) neuron, introduced over two decades ago, is a simple, efficient, yet biologically plausible spiking neuron model, which can extract frequency patterns within the time domain due to its resonating membrane dynamics. However, previous RF formulations suffer from intrinsic shortcomings that limit effective learning and prevent exploiting the principled advantage of RF neurons. Here, we introduce the balanced RF (BRF) neuron, which alleviates some of the intrinsic limitations of vanilla RF neurons and demonstrates its effectiveness within recurrent spiking neural networks (RSNNs) on various sequence learning tasks. We show that networks of BRF neurons achieve overall higher task performance, produce only a fraction of the spikes, and require significantly fewer parameters as compared to modern RSNNs. Moreover, BRF-RSNN consistently provide much faster and more stable training convergence, even when bridging many hundreds of time steps during backpropagation through time (BPTT). These results underscore that our BRF-RSNN is a strong candidate for future large-scale RSNN architectures, further lines of research in SNN methodology, and more efficient hardware implementations.
Saya Higuchi, Sebastian Kairat, Sander M. Bohté, Sebastian Otte
ICML3
2024 Recurrent neural networks that learn multi-step visual routines with reinforcement learning
abstract
Many cognitive problems can be decomposed into series of subproblems that are solved sequentially by the brain. When subproblems are solved, relevant intermediate results need to be stored by neurons and propagated to the next subproblem, until the overarching goal has been completed. We will here consider visual tasks, which can be decomposed into sequences of elemental visual operations. Experimental evidence suggests that intermediate results of the elemental operations are stored in working memory as an enhancement of neural activity in the visual cortex. The focus of enhanced activity is then available for subsequent operations to act upon. The main question at stake is how the elemental operations and their sequencing can emerge in neural networks that are trained with only rewards, in a reinforcement learning setting. We here propose a new recurrent neural network architecture that can learn composite visual tasks that require the application of successive elemental operations. Specifically, we selected three tasks for which electrophysiological recordings of monkeys' visual cortex are available. To train the networks, we used RELEARNN, a biologically plausible four-factor Hebbian learning rule, which is local both in time and space. We report that networks learn elemental operations, such as contour grouping and visual search, and execute sequences of operations, solely based on the characteristics of the visual stimuli and the reward structure of a task. After training was completed, the activity of the units of the neural network elicited by behaviorally relevant image items was stronger than that elicited by irrelevant ones, just as has been observed in the visual cortex of monkeys solving the same tasks. Relevant information that needed to be exchanged between subroutines was maintained as a focus of enhanced activity and passed on to the subsequent subroutines. Our results demonstrate how a biologically plausible learning rule can train a recurrent neural network on multistep visual tasks.
Sami Mollard, Catherine Wacongne, Sander M. Bohté, Pieter R. Roelfsema
PLoS Comput. Biol.3
2023 Efficient Uncertainty Estimation in Spiking Neural Networks via MC-dropout
Bojian Yin, Sander M. Bohté
ICANN (1)3
2023 Mechanisms of human dynamic object recognition revealed by sequential deep neural networks
abstract
Humans can quickly recognize objects in a dynamically changing world. This ability is showcased by the fact that observers succeed at recognizing objects in rapidly changing image sequences, at up to 13 ms/image. To date, the mechanisms that govern dynamic object recognition remain poorly understood. Here, we developed deep learning models for dynamic recognition and compared different computational mechanisms, contrasting feedforward and recurrent, single-image and sequential processing as well as different forms of adaptation. We found that only models that integrate images sequentially via lateral recurrence mirrored human performance (N = 36) and were predictive of trial-by-trial responses across image durations (13-80 ms/image). Importantly, models with sequential lateral-recurrent integration also captured how human performance changes as a function of image presentation durations, with models processing images for a few time steps capturing human object recognition at shorter presentation durations and models processing images for more time steps capturing human object recognition at longer presentation durations. Furthermore, augmenting such a recurrent model with adaptation markedly improved dynamic recognition performance and accelerated its representational dynamics, thereby predicting human trial-by-trial responses using fewer processing resources. Together, these findings provide new insights into the mechanisms rendering object recognition so fast and effective in a dynamic visual world.
Lynn K. A. Sörensen, Sander M. Bohté, Dorina De Jong, Heleen A. Slagter, H. Steven Scholte
PLoS Comput. Biol.2
2022 A Taxonomy of Recurrent Learning Rules
Guillermo Martín-Sánchez, Sander M. Bohté, Sebastian Otte
ICANN (1)2
2022 Real-time classification of LIDAR data using discrete-time Recurrent Spiking Neural Networks
abstract
With the advancement of Edge AI and autonomous systems, AI applications are increasingly subject to energy, latency and environmental constraints. Biological neural systems naturally adhere to these constraints and, as such are a source of inspiration. Spiking Neural Networks (SNNs) are a more detailed model of biological neural processing. Recent work shows that they perform well in object recognition and detection in general, and in Autonomous Driving tasks based on ranged LIDAR data in particular. However, these LIDAR-SNN approaches do not optimize for latency, as they require the entire frame to be scanned before processing. They also require large SNNs, limiting the energy efficiency achieved. To reach both low-latency and high energy efficiency in LIDAR object recognition, we develop a compact recurrent SNN. First, we propose and examine an open LIDAR labeled dataset by processing the point clouds from the KITTI Vision Benchmark. We then train our recurrent SNNs on this dataset and propose specific optimizations, including input encoding, sparse connectivity and truncation of error-backpropagation. With these optimizations, we show that compact recurrent SNNs can exceed the performance of classical RNNs like LSTMs and approach the performance of large non-spiking CNNs. Additionally, they significantly reduce latency by allowing early and online object classification before the end of the sequence.
Anca-Diana Vicol, Bojian Yin, Sander M. Bohté
IJCNN3
2022 Arousal state affects perceptual decision-making by modulating hierarchical sensory processing in a large-scale visual system model
abstract
Arousal levels strongly affect task performance. Yet, what arousal level is optimal for a task depends on its difficulty. Easy task performance peaks at higher arousal levels, whereas performance on difficult tasks displays an inverted U-shape relationship with arousal, peaking at medium arousal levels, an observation first made by Yerkes and Dodson in 1908. It is commonly proposed that the noradrenergic locus coeruleus system regulates these effects on performance through a widespread release of noradrenaline resulting in changes of cortical gain. This account, however, does not explain why performance decays with high arousal levels only in difficult, but not in simple tasks. Here, we present a mechanistic model that revisits the Yerkes-Dodson effect from a sensory perspective: a deep convolutional neural network augmented with a global gain mechanism reproduced the same interaction between arousal state and task difficulty in its performance. Investigating this model revealed that global gain states differentially modulated sensory information encoding across the processing hierarchy, which explained their differential effects on performance on simple versus difficult tasks. These findings offer a novel hierarchical sensory processing account of how, and why, arousal state affects task performance.
Lynn K. A. Sörensen, Sander M. Bohté, Heleen A. Slagter, H. Steven Scholte
PLoS Comput. Biol.2
2021 LocalNorm: Robust Image Classification Through Dynamically Regularized Normalization
Bojian Yin, H. Steven Scholte, Sander M. Bohté
ICANN (4)3
2021 Learning continuous-time working memory tasks with on-policy neural reinforcement learning
abstract
An animals’ ability to learn how to make decisions based on sensory evidence is often well described by Reinforcement Learning (RL) frameworks. These frameworks, however, typically apply to event-based representations and lack the explicit and fine-grained notion of time needed to study psychophysically relevant measures like reaction times and psychometric curves. Here, we develop and use a biologically plausible continuous-time RL scheme of CT-AuGMEnT (Continuous-Time Attention-Gated MEmory Tagging) to study these behavioural quantities. We show how CT-AuGMEnT implements on-policy SARSA learning as a biologically plausible form of reinforcement learning with working memory units using ‘attentional’ feedback. We show that the CT-AuGMEnT model efficiently learns tasks in continuous time and can learn to accumulate relevant evidence through time. This allows the model to link task difficulty to psychophysical measurements such as accuracy and reaction-times. We further show how the implementation of a separate accessory network for feedback allows the model to learn continuously, also in case of significant transmission delays between the network’s feedforward and feedback layers and even when the accessory network is randomly initialized. Our results demonstrate that CT-AuGMEnT represents a fully time-continuous biologically plausible end-to-end RL model for learning to integrate evidence and make decisions.
Davide Zambrano, Pieter R. Roelfsema, Sander M. Bohté
Neurocomputing3
2021 Flexible Working Memory Through Selective Gating and Attentional Tagging
abstract
Working memory is essential: it serves to guide intelligent behavior of humans and nonhuman primates when task-relevant stimuli are no longer present to the senses. Moreover, complex tasks often require that multiple working memory representations can be flexibly and independently maintained, prioritized, and updated according to changing task demands. Thus far, neural network models of working memory have been unable to offer an integrative account of how such control mechanisms can be acquired in a biologically plausible manner. Here, we present WorkMATe, a neural network architecture that models cognitive control over working memory content and learns the appropriate control operations needed to solve complex working memory tasks. Key components of the model include a gated memory circuit that is controlled by internal actions, encoding sensory information through untrained connections, and a neural circuit that matches sensory inputs to memory content. The network is trained by means of a biologically plausible reinforcement learning rule that relies on attentional feedback and reward prediction errors to guide synaptic updates. We demonstrate that the model successfully acquires policies to solve classical working memory tasks, such as delayed recognition and delayed pro-saccade/anti-saccade tasks. In addition, the model solves much more complex tasks, including the hierarchical 12-AX task or the ABAB ordered recognition task, both of which demand an agent to independently store and updated multiple items separately in memory. Furthermore, the control strategies that the model acquires for these tasks subsequently generalize to new task contexts with novel stimuli, thus bringing symbolic production rule qualities to a neural network architecture. As such, WorkMATe provides a new solution for the neural implementation of flexible memory control.
Wouter Kruijne, Sander M. Bohté, Pieter R. Roelfsema, Christian N. L. Olivers
Neural Comput.2
2020 Attention-Gated Brain Propagation: How the brain can implement reward-based error backpropagation
abstract
Much recent work has focused on biologically plausible variants of supervised learning algorithms. However, there is no teacher in the motor cortex that instructs the motor neurons and learning in the brain depends on reward and punishment. We demonstrate a biologically plausible reinforcement learning scheme for deep networks with an arbitrary number of layers. The network chooses an action by selecting a unit in the output layer and uses feedback connections to assign credit to the units in successively lower layers that are responsible for this action. After the choice, the network receives reinforcement and there is no teacher correcting the errors. We show how the new learning scheme – Attention-Gated Brain Propagation (BrainProp) – is mathematically equivalent to error backpropagation, for one output unit at a time. We demonstrate successful learning of deep fully connected, convolutional and locally connected networks on classical and hard image-classification benchmarks; MNIST, CIFAR10, CIFAR100 and Tiny ImageNet. BrainProp achieves an accuracy that is equivalent to that of standard error-backpropagation, and better than state-of-the-art biologically inspired learning schemes. The trial-and-error nature of learning is associated with limited additional training time so that BrainProp is a factor of 1-3.5 times slower. Our results thereby provide new insights into how deep learning may be implemented in the brain.
Isabella Pozzi, Sander M. Bohté, Pieter R. Roelfsema
NeurIPS2
2020 Depth in convolutional neural networks solves scene segmentation
abstract
Feed-forward deep convolutional neural networks (DCNNs) are, under specific conditions, matching and even surpassing human performance in object recognition in natural scenes. This performance suggests that the analysis of a loose collection of image features could support the recognition of natural object categories, without dedicated systems to solve specific visual subtasks. Research in humans however suggests that while feedforward activity may suffice for sparse scenes with isolated objects, additional visual operations ('routines') that aid the recognition process (e.g. segmentation or grouping) are needed for more complex scenes. Linking human visual processing to performance of DCNNs with increasing depth, we here explored if, how, and when object information is differentiated from the backgrounds they appear on. To this end, we controlled the information in both objects and backgrounds, as well as the relationship between them by adding noise, manipulating background congruence and systematically occluding parts of the image. Results indicate that with an increase in network depth, there is an increase in the distinction between object- and background information. For more shallow networks, results indicated a benefit of training on segmented objects. Overall, these results indicate that, de facto, scene segmentation can be performed by a network of sufficient depth. We conclude that the human brain could perform scene segmentation in the context of object identification without an explicit mechanism, by selecting or "binding" features that belong to the object and ignoring other features, in a manner similar to a very deep convolutional neural network.
Noor Seijdel, Nikos Tsakmakidis, Edward H. F. de Haan, Sander M. Bohté, H. Steven Scholte
PLoS Comput. Biol.4
2019 Using the structure of genome data in the design of deep neural networks for predicting amyotrophic lateral sclerosis from genotype
abstract
MOTIVATION: Amyotrophic lateral sclerosis (ALS) is a neurodegenerative disease caused by aberrations in the genome. While several disease-causing variants have been identified, a major part of heritability remains unexplained. ALS is believed to have a complex genetic basis where non-additive combinations of variants constitute disease, which cannot be picked up using the linear models employed in classical genotype-phenotype association studies. Deep learning on the other hand is highly promising for identifying such complex relations. We therefore developed a deep-learning based approach for the classification of ALS patients versus healthy individuals from the Dutch cohort of the Project MinE dataset. Based on recent insight that regulatory regions harbor the majority of disease-associated variants, we employ a two-step approach: first promoter regions that are likely associated to ALS are identified, and second individuals are classified based on their genotype in the selected genomic regions. Both steps employ a deep convolutional neural network. The network architecture accounts for the structure of genome data by applying convolution only to parts of the data where this makes sense from a genomics perspective. RESULTS: Our approach identifies potentially ALS-associated promoter regions, and generally outperforms other classification methods. Test results support the hypothesis that non-additive combinations of variants contribute to ALS. Architectures and protocols developed are tailored toward processing population-scale, whole-genome data. We consider this a relevant first step toward deep learning assisted genotype-phenotype association in whole genome-sized data. AVAILABILITY AND IMPLEMENTATION: Our code will be available on Github, together with a synthetic dataset (https://github.com/byin-cwi/ALS-Deeplearning). The data used in this study is available to bona-fide researchers upon request. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Bojian Yin, Marleen Balvert, Rick A. A. van der Spek, Bas E. Dutilh, Sander M. Bohté, Jan Veldink, Alexander Schönhuth
Bioinform.5
2018 A Deep Predictive Coding Network for Inferring Hierarchical Causes Underlying Sensory Inputs
Shirin Dora, Cyriel M. A. Pennartz, Sander M. Bohté
ICANN (3)3
2018 Continuous-Time Spike-Based Reinforcement Learning for Working Memory Tasks
Marios Karamanis, Davide Zambrano, Sander M. Bohté
ICANN (2)3
2018 Gating Sensory Noise in a Spiking Subtractive LSTM
Isabella Pozzi, Roeland Nusselder, Davide Zambrano, Sander M. Bohté
ICANN (1)4
2018 An image representation based convolutional network for DNA classification
Bojian Yin, Marleen Balvert, Davide Zambrano, Alexander Schönhuth, Sander M. Bohté
ICLR (Poster)5
2015 Continuous-time on-policy neural Reinforcement Learning of working memory tasks
abstract
As living organisms, one of our primary characteristics is the ability to rapidly process and react to unknown and unexpected events. To this end, we are able to recognize an event or a sequence of events and learn to respond properly. Despite advances in machine learning, current cognitive robotic systems are not able to rapidly and efficiently respond in the real world: the challenge is to learn to recognize both what is important, and also when to act. Reinforcement Learning (RL) is typically used to solve complex tasks: to learn the how. To respond quickly - to learn when - the environment has to be sampled often enough. For “enough”, a programmer has to decide on the step-size as a time-representation, choosing between a fine-grained representation of time (many state-transitions; difficult to learn with RL) or to a coarse temporal resolution (easier to learn with RL but lacking precise timing). Here, we derive a continuous-time version of on-policy SARSA-learning in a working-memory neural network model, AuGMEnT. Using a neural working memory network resolves the what problem, our when solution is built on the notion that in the real world, instantaneous actions of duration dt are actually impossible. We demonstrate how we can decouple action duration from the internal time-steps in the neural RL model using an action selection system. The resultant CT-AuGMEnT successfully learns to react to the events of a continuous-time task, without any pre-imposed specifications about the duration of the events or the delays between them.
Davide Zambrano, Pieter R. Roelfsema, Sander M. Bohté
IJCNN3
2015 How Attention Can Create Synaptic Tags for the Learning of Working Memories in Sequential Tasks
abstract
Intelligence is our ability to learn appropriate responses to new stimuli and situations. Neurons in association cortex are thought to be essential for this ability. During learning these neurons become tuned to relevant features and start to represent them with persistent activity during memory delays. This learning process is not well understood. Here we develop a biologically plausible learning scheme that explains how trial-and-error learning induces neuronal selectivity and working memory representations for task-relevant information. We propose that the response selection stage sends attentional feedback signals to earlier processing levels, forming synaptic tags at those connections responsible for the stimulus-response mapping. Globally released neuromodulators then interact with tagged synapses to determine their plasticity. The resulting learning rule endows neural networks with the capacity to create new working memory representations of task relevant information as persistent activity. It is remarkably generic: it explains how association neurons learn to store task-relevant information for linear as well as non-linear stimulus-response mappings, how they become tuned to category boundaries or analog variables, depending on the task demands, and how they learn to integrate probabilistic evidence for perceptual decisions.
Jaldert O. Rombouts, Sander M. Bohté, Pieter R. Roelfsema
PLoS Comput. Biol.2
2014 Spiking Neural Networks: Principles and Challenges
André Grüning, Sander M. Bohté
ESANN2
2014 Learning resets of neural working memory
Jaldert O. Rombouts, Pieter R. Roelfsema, Sander M. Bohté
ESANN3
2014 Spiking AGREL
Davide Zambrano, Jaldert O. Rombouts, Cecilia Laschi, Sander M. Bohté
ESANN4
2012 Biologically Plausible Multi-dimensional Reinforcement Learning in Neural Networks
Jaldert O. Rombouts, Arjen van Ooyen, Pieter R. Roelfsema, Sander M. Bohté
ICANN (1)4
2012 Efficient Spike-Coding with Multiplicative Adaptation in a Spike Response Model
abstract
Neural adaptation underlies the ability of neurons to maximize encoded information over a wide dynamic range of input stimuli. While adaptation is an intrinsic feature of neuronal models like the Hodgkin-Huxley model, the challenge is to integrate adaptation in models of neural computation. Recent computational models like the Adaptive Spike Response Model implement adaptation as spike-based addition of fixed-size fast spike-triggered threshold dynamics and slow spike-triggered currents. Such adaptation has been shown to accurately model neural spiking behavior over a limited dynamic range. Taking a cue from kinetic models of adaptation, we propose a multiplicative Adaptive Spike Response Model where the spike-triggered adaptation dynamics are scaled multiplicatively by the adaptation state at the time of spiking. We show that unlike the additive adaptation model, the firing rate in the multiplicative adaptation model saturates to a maximum spike-rate. When simulating variance switching experiments, the model also quantitatively fits the experimental data over a wide dynamic range. Furthermore, dynamic threshold models of adaptation suggest a straightforward interpretation of neural activity in terms of dynamic signal encoding with shifted and weighted exponential kernels. We show that when thus encoding rectified filtered stimulus signals, the multiplicative Adaptive Spike Response Model achieves a high coding efficiency and maintains this efficiency over changes in the dynamic signal range of several orders of magnitude, without changing model parameters.
Sander M. Bohté
NIPS1
2012 Neurally Plausible Reinforcement Learning of Working Memory Tasks
abstract
A key function of brains is undoubtedly the abstraction and maintenance of information from the environment for later use. Neurons in association cortex play an important role in this process: during learning these neurons become tuned to relevant features and represent the information that is required later as a persistent elevation of their activity. It is however not well known how these neurons acquire their task-relevant tuning. Here we introduce a biologically plausible learning scheme that explains how neurons become selective for relevant information when animals learn by trial and error. We propose that the action selection stage feeds back attentional signals to earlier processing levels. These feedback signals interact with feedforward signals to form synaptic tags at those connections that are responsible for the stimulus-response mapping. A globally released neuromodulatory signal interacts with these tagged synapses to determine the sign and strength of plasticity. The learning scheme is generic because it can train networks in different tasks, simply by varying inputs and rewards. It explains how neurons in association cortex learn to (1) temporarily store task-relevant information in non-linear stimulus-response mapping tasks and (2) learn to optimally integrate probabilistic evidence for perceptual decision making.
Jaldert O. Rombouts, Sander M. Bohté, Pieter R. Roelfsema
NIPS2
2011 Error-Backpropagation in Networks of Fractionally Predictive Spiking Neurons
Sander M. Bohté
ICANN (1)1
2010 Fractionally Predictive Spiking Neurons
abstract
Recent experimental work has suggested that the neural firing rate can be interpreted as a fractional derivative, at least when signal variation induces neural adaptation. Here, we show that the actual neural spike-train itself can be considered as the fractional derivative, provided that the neural signal is approximated by a sum of power-law kernels. A simple standard thresholding spiking neuron suffices to carry out such an approximation, given a suitable refractory response. Empirically, we find that the online approximation of signals with a sum of power-law kernels is beneficial for encoding signals with slowly varying components, like long-memory self-similar signals. For such signals, the online power-law kernel approximation typically required less than half the number of spikes for similar SNR as compared to sums of similar but exponentially decaying kernels. As power-law kernels can be accurately approximated using sums or cascades of weighted exponentials, we demonstrate that the corresponding decoding of spike-trains by a receiving neuron allows for natural and transparent temporal signal filtering by tuning the weights of the decoding kernel.
Sander M. Bohté, Jaldert O. Rombouts
NIPS1
2009 Optimization of Online Patient Scheduling with Urgencies and Preferences
Ivan B. Vermeulen, Sander M. Bohté, Peter A. N. Bosman, Sylvia G. Elkhuizen, Piet J. M. Bakker, Han La Poutré
AIME2
2009 Adaptive resource allocation for efficient patient scheduling
Ivan B. Vermeulen, Sander M. Bohté, Sylvia G. Elkhuizen, Han Lameris, Piet J. M. Bakker, Han La Poutré
Artif. Intell. Medicine2
2007 Adaptive Optimization of Hospital Resource Calendars
Ivan B. Vermeulen, Sander M. Bohté, Sylvia G. Elkhuizen, J. S. Lameris, Piet J. M. Bakker, Han La Poutré
AIME2
2007 Reducing the Variability of Neural Responses: A Computational Theory of Spike-Timing-Dependent Plasticity
abstract
Experimental studies have observed synaptic potentiation when a presynaptic neuron fires shortly before a postsynaptic neuron and synaptic depression when the presynaptic neuron fires shortly after. The dependence of synaptic modulation on the precise timing of the two action potentials is known as spike-timing dependent plasticity (STDP). We derive STDP from a simple computational principle: synapses adapt so as to minimize the postsynaptic neuron's response variability to a given presynaptic input, causing the neuron's output to become more reliable in the face of noise. Using an objective function that minimizes response variability and the biophysically realistic spike-response model of Gerstner (2001), we simulate neurophysiological experiments and obtain the characteristic STDP curve along with other phenomena, including the reduction in synaptic plasticity as synaptic efficacy increases. We compare our account to other efforts to derive STDP from computational principles and argue that our account provides the most comprehensive coverage of the phenomena. Thus, reliability of neural response in the face of noise may be a key goal of unsupervised cortical adaptation.
Sander M. Bohté, Michael C. Mozer
Neural Comput.1
2007 Multi-agent Pareto appointment exchanging in hospital patient scheduling
abstract
We present a dynamic and distributed approach to the hospital patient scheduling problem, in which patients can have multiple appointments that have to be scheduled to different resources. To efficiently solve this problem we develop a multi-agent Pareto-improvement appointment exchanging algorithm: MPAEX. It respects the decentralization of scheduling authorities and continuously improves patient schedules in response to the dynamic environment. We present models of the hospital patient scheduling problem in terms of the health care cycle where a doctor repeatedly orders sets of activities to diagnose and/or treat a patient. We introduce the Theil index to the health care domain to characterize different hospital patient scheduling problems in terms of the degree of relative workload inequality between required resources. In experiments that simulate a broad range of hospital patient scheduling problems, we extensively compare the performance of MPAEX to a set of scheduling benchmarks. The distributed and dynamic MPAEX performs almost as good as the best centralized and static scheduling heuristic, and is robust for variations in the model settings.
Ivan B. Vermeulen, Sander M. Bohté, Koye Somefun, Han La Poutré
Serv. Oriented Comput. Appl.2
2006 Strategic Foresighted Learning in Competitive Multi-Agent Games
Pieter Jan't Hoen, Sander M. Bohté, Han La Poutré
ECAI2
2005 Applications of spiking neural networks
Sander M. Bohté, Joost N. Kok
Inf. Process. Lett.1
2004 Nonparametric classification with polynomial MPMC cascades
Sander M. Bohté, Markus Breitenbach, Gregory Z. Grudic
ICML1
2004 Reducing Spike Train Variability: A Computational Theory Of Spike-Timing Dependent Plasticity
abstract
Experimental studies have observed synaptic potentiation when a presynaptic neuron fires shortly before a postsynaptic neuron, and synaptic depression when the presynaptic neuron fires shortly af- ter. The dependence of synaptic modulation on the precise tim- ing of the two action potentials is known as spike-timing depen- dent plasticity or STDP. We derive STDP from a simple compu- tational principle: synapses adapt so as to minimize the postsy- naptic neuron's variability to a given presynaptic input, causing the neuron's output to become more reliable in the face of noise. Using an entropy-minimization objective function and the biophys- ically realistic spike-response model of Gerstner (2001), we simu- late neurophysiological experiments and obtain the characteristic STDP curve along with other phenomena including the reduction in synaptic plasticity as synaptic efficacy increases. We compare our account to other efforts to derive STDP from computational princi- ples, and argue that our account provides the most comprehensive coverage of the phenomena. Thus, reliability of neural response in the face of noise may be a key goal of cortical adaptation. 1 Introduction Experimental studies have observed synaptic potentiation when a presynaptic neu- ron fires shortly before a postsynaptic neuron, and synaptic depression when the presynaptic neuron fires shortly after. The dependence of synaptic modulation on the precise timing of the two action potentials, known as spike-timing dependent plasticity or STDP, is depicted in Figure 1. Typically, plasticity is observed only when the presynaptic and postsynaptic spikes (hereafter, pre and post) occur within a 2030 ms time window, and the transition from potentiation to depression is very rapid. Another important observation is that synaptic plasticity decreases with in- creased synaptic efficacy. The effects are long lasting, and are therefore referred to as long-term potentiation (LTP) and depression (LTD). For detailed reviews of the evidence for STDP, see [1, 2]. Because these intriguing findings appear to describe a fundamental learning mech- anism in the brain, a flurry of models have been developed that focus on different aspects of STDP, from biochemical models that explain the underlying mechanisms giving rise to STDP [3], to models that explore the consequences of a STDP-like learning rules in an ensemble of spiking neurons [4, 5, 6, 7], to models that pro- pose fundamental computational justifications for STDP. Most commonly, STDP Figure 1: (a) Measuring STDP experimentally: pre-post spike pairs are repeatedly in- duced at a fixed interval tpre-post, and the resulting change to the strength of the synapse is assessed; (b) change in synaptic strength after repeated spike pairing as a function of the difference in time between the pre and post spikes (data from Zhang et al., 1998). We have superimposed an exponential fit of LTP and LTD. is viewed as a type of asymmetric Hebbian learning with a temporal dimension. However, this perspective is hardly a fundamental computational rationale, and one would hope that such an intuitively sensible learning rule would emerge from a first-principle computational justification. Several researchers have tried to derive a learning rule yielding STDP from first principles. Rao and Sejnowski [8] show that STDP emerges when a neuron attempts to predict its membrane potential at some time t from the potential at time t - t. However, STDP emerges only for a narrow range of t values, and the qualitative nature of the modeling makes it unclear whether a quantitative fit can be obtained. Dayan and H ausser [9] show that STDP can be viewed as an optimal noise-removal filter for certain noise distributions. However, even small variation from these noise distributions yield quite different learning rules, and the noise statistics of biological neurons are unknown. Eisele (private communication) has shown that an STDP-like learning rule can be derived from the goal of maintaining the relevant connections in a network. Chechik [10] is most closely related to the present work. He relates STDP to information theory via maximization of mutual information between input and output spike trains. This approach derives the LTP portion of STDP, but fails to yield the LTD portion. The computational approach of Chechik (as well as Dayan and H ausser) is premised on a rate-coding neuron model that disregards the relative timing of spikes. It seems quite odd to argue for STDP using rate codes: if spike timing is irrelevant to information transmission, then STDP is likely an artifact and is not central to understanding mechanisms of neural computation. Further, as noted in [9], because STDP is not quite additive in the case of multiple input or output spikes that are near in time [11], one should consider interpretations that are based on individual spikes, not aggregates over spike trains. Here, we present an alternative computational motivation for STDP. We conjecture that a fundamental objective of cortical computation is to achieve reliable neural re- sponses, that is, neurons should produce the identical response--both in the number and timing of spikes--given a fixed input spike train. Reliability is an issue if neu- rons are affected by noise influences, because noise leads to variability in a neuron's dynamics and therefore in its response. Minimizing this variability will reduce the effect of noise and will therefore increase the informativeness of the neuron's output signal. The source of the noise is not important; it could be intrinsic to a neuron (e.g., a noisy threshold) or it could originate in unmodeled external sources causing fluctuations in the membrane potential uncorrelated with a particular input. We are not suggesting that increasing neural reliability is the only learning objective. If it were, a neuron would do well to give no response regardless of the input. Rather, reliability is but one of many objectives that learning tries to achieve. This form of unsupervised learning must, of course, be complemented by supervised and reinforcement learning that allow an organism to achieve its goals and satisfy drives. We derive STDP from the following computational principle: synapses adapt so as to minimize the entropy of the postsynaptic neuron's output in response to a given presynaptic input. In our simulations, we follow the methodology of neurophysiolog- ical experiments. This approach leads to a detailed fit to key experimental results. We model not only the shape (sign and time course) of the STDP curve, but also the fact that potentiation of a synapse depends on the efficacy of the synapse--it decreases with increased efficacy. In addition to fitting these key STDP phenom- ena, the model allows us to make predictions regarding the relationship between properties of the neuron and the shape of the STDP curve. Before delving into the details of our approach, we attempt to give a basic intu- ition about the approach. Noise in spiking neuron dynamics leads to variability in the number and timing of spikes. Given a particular input, one spike train might be more likely than others, but the output is nondeterministic. By the entropy- minimization principle, adaptation should reduce the likelihood of these other pos- sibilities. To be concrete, consider a particular experimental paradigm. In [12], a pre neuron is identified with a weak synapse to a post neuron, such that the pre is unlikely to cause the post to fire. However, the post can be induced to fire via a second presynaptic connection. In a typical trial, the pre is induced to fire a single spike, and with a variable delay, the post is also induced to fire (typically) a single spike. To increase the likelihood of the observed post response, other response pos- sibilities must be suppressed. With presynaptic input preceding the postsynaptic spike, the most likely alternative response is no output spikes at all. Increasing the synaptic connection weight should then reduce the possibility of this alternative response. With presynaptic input following the postsynaptic spike, the most likely alternative response is a second output spike. Decreasing the synaptic connection weight should reduce the possibility of this alternative response. Because both of these alternatives become less likely as the lag between pre and post spikes is in- creased, one would expect that the magnitude of synaptic plasticity diminishes with the lag, as is observed in the STDP curve. Our approach to reducing response variability given a particular input pattern in- volves computing the gradient of synaptic weights with respect to a differentiable model of spiking neuron behavior. We use the Spike Response Model (SRM) of [13] with a stochastic threshold, where the stochastic threshold models fluctuations of the membrane potential or the threshold outside of experimental control. For the stochastic SRM, the response probability is differentiable with respect to the synap- tic weights, allowing us to calculate the entropy gradient with respect to the weights conditional on the presented input. Learning is presumed to take a gradient step to reduce this conditional entropy. In modeling neurophysiological experiments, we demonstrate that this learning rule yields the typical STDP curve. We can predict the relationship between the exact shape of the STDP curve and physiologically measurable parameters, and we show that our results are robust to the choice of the few free parameters of the model. Two papers in these proceedings are closely related to our work. They also find STDP-like curves when attempting to maximize an information-theoretic measure-- the mutual information between input and output--for a Spike Response Model [14, 15]. Bell & Parra [14] use a deterministic SRM model which does not model the LTD component of STDP properly. The derivation by Toyoizumi et al. [15] is valid only for an essentially constant membrane potential with small fluctuations. Neither of these approaches has succeeded in quantitatively modeling specific experimental data with neurobiologically-realistic timing parameters, and neither explains the saturation of LTD/LTP with increasing weights as we do. Nonetheless, these models make an interesting contrast to ours by suggesting a computational principle of optimization of information transmission, as contrasted with our principle of neural noise reduction. Perhaps experimental tests can be devised to distinguish between these competing theories. 2 The Stochastic Spike Response Model The Spike Response Model (SRM), defined by Gerstner [13], is a generic integrate- and-fire model of a spiking neuron that closely corresponds to the behavior of a biological spiking neuron and is characterized in terms of a small set of easily inter- pretable parameters [16]. The standard SRM formulation describes the temporal evolution of the membrane potential based on past neuronal events, specifically as a weighted sum of postsynaptic potentials (PSPs) modulated by reset and thresh- old effects of previous postsynaptic spiking events. Following [13], the membrane potential of cell i at time t, ui(t), is defined as: ui(t) = (t - ^ fi) + wij (t - ^ fi, t - fj), (1) ji fj F t j where i is the set of inputs connected to neuron i, Ft is the set of times prior to j t that neuron j has spiked, ^ fi is the time of the last spike of neuron i, wij is the synaptic weight from neuron j to neuron i, (t - ^ fi, t - fj) is the PSP in neuron i due to an input spike from neuron j at time fj, and (t - ^ fi) is the refractory response due to the postsynaptic spike at time ^ fi. Neuron i fires when the potential ui(t) exceeds a threshold () from below. The postsynaptic potential is modeled as the differential alpha function in [13], defined with respect to two variables: the time since the most recent postsynaptic spike, x, and the time since the presynaptic spike, s: 1 s s (x, s) = exp - - exp - H(s)H(x - s)+ (2) 1 - s m s m s - x x x +exp - exp - - exp - H(x)H(s - x) , s m s where s and m are the rise and decay time-constants of the PSP, and H is the Heaviside function. The refractory reset function is defined to be [13]: x + x (x) = u abs absH(abs - x)H(-x) + uabsexp - + usexp - , (3) r f s r r where uabs is a large negative contribution to the potential to model the absolute refractory period, with duration abs. We smooth this refractory response by a fast decaying exponential with time constant f . The third term in the sum represents r the slow decaying exponential recovery of an elevated threshold, us, with time r constant s. (Graphs of these and functions can be found in [13].) We made r a minor modification to the SRM described in [13] by relaxing the constraint that s = r m; smoothing the absolute refractory function is mentioned in [13] but not explicitly defined as we do here. In all simulations presented, abs = 2ms, s = 4 r m, and f = 0.1 r m. The SRM we just described is deterministic. Gerstner [13] introduces a stochas- tic variant of the SRM (sSRM) by incorporating the notion of a stochastic firing threshold: given membrane potential ui(t), the probability density of the neuron firing at time t is specified by (ui(t)). Herrmann & Gerstner [17] find that then for a realistic escape-rate noise model the firing probability density as a function of the potential is initially small and constant, transitioning to asymptotically linear increasing around threshold . In our simulations, we use such a function: (v) = (ln[1 + exp(( - v))] - ( - v)), (4) where is the firing threshold in the absence of noise, determines the abruptness of the constant-to-linear probability density transition around , and determines the slope of the increasing part. Experiments with sigmoidal and exponential density functions were found to not qualitatively affect the results. 3 Minimizing Conditional Entropy We now derive the rule for adjusting the weight from a presynaptic neuron j to a postsynaptic sSRM neuron i, so as to minimize the entropy of i's response given a particular spike sequence from j. A spike sequence is described by the set of all times at which spikes have occurred within some interval between 0 and T , denoted F T for neuron j. We assume the interval is wide enough that spikes outside the j interval do not influence the state of the neuron within the interval (e.g., through threshold reset effects). We can then treat intervals as independent of each other. Let the postsynaptic neuron i produce a response i, where i is the set of all possible responses given the input, FT , and g() is the probability density over i responses. The differential conditional entropy h(i) of neuron i's response is then defined as: h(i) = - g()log g() d. (5) i To minimize the differential conditional entropy by adjusting the neuron's weights, we compute the gradient of the conditional entropy with respect to the weights: h(i) log(g()) = - g() log(g()) + 1 d. (6) wij wij i For a differentiable neuron model, log(g())/wij can be expressed as follows when neuron i fires once at time ^ fi [18]: log(g()) T (u u (t - ^ fi) - (ui(t)) = i(t)) i(t) dt, (7) wij t=0 ui(t) wij (ui(t)) where (.) is the Dirac delta, and (ui(t)) is the firing probability-density of neuron i at time t. (See [18] for the generalization to multiple postsynaptic spikes.) With the sSRM we can compute the partial derivatives (ui(t))/ui(t) and ui(t)/wij. Given the density function (4), (ui(t)) u = , i(t) = (t - ^ f u i, t - fj ). i(t) 1 + exp(( - ui(t)) wij To perform gradient descent in the conditional entropy, we use the weight update h( w i) ij - (8) wij T (t - ^ fi, t - fj) (t - ^ fi) - (ui(t)) g() log(g()) + 1 dt d. (1 + exp(( - ui(t)))(ui(t)) i t=0 We can use numerical methods to evaluate Equation (8). However, it seems bio- logically unrealistic to suppose a neuron can integrate over all possible responses . This dilemma can be circumvented in two ways. First, the resulting learning rule might be cached in some form through evolution so that the full computation is not necessary (e.g., in an STDP curve). Second, the specific response produced by a neuron on a single trial might be considered to be a sample from the distribution g(), and the integration is performed by a sampling process over repeated trials; Figure 2: (a) Experimental setup of Zhang et al. and (b) their experimental STDP curve (small squares) vs. our model (solid line). Model parameters: s = 1.5ms, m = 12.25ms. each trial would produce a stochastic gradient step. 4 Simulation Methodology We model in detail the experiment of Zhang et al. [12] (Figure 2a). In this exper- iment, a post neuron is identified that has two neurons projecting to it, call them the pre and the driver. The pre is subthreshold: it produces depolarization but no spike. The driver is suprathreshold: it induces a spike in the post. Plasticity of the pre-post synapse is measured as a function of the timing between pre and post spikes (tpre-post) by varying the timing between induced spikes in the pre and the driver (tpre-driver). This measurement yields the well-known STDP curve (Figure 1b).1 The experiment imposes several constraints on a simulation: The driver alone causes spiking > 70% of the time, the pre alone causes spiking < 10% of the time, synchronous firing of driver and pre cause LTP if and only if the post fires, and the time constants of the EPSPs--s and m in the sSRM--are in the range of 13ms and 1015ms respectively. These constraints remove many free parameters from our simulation. We do not explicitly model the two input cells; instead, we model the EPSPs they produce. The magnitude of these EPSPs are picked to satisfy the experimental constraints: the driver EPSP alone causes a spike in the post on 77.4% of trials, and the pre EPSP alone causes a spike on fewer than 0.1% of trials. Free parameters of the simulation are and in the spike-probability function ( can be folded into ), and the magnitude (us, u , f , r abs) and reset time constants ( s r r abs). The dependent variable of the simulation is tpre-driver, and we measure the time of the post spike to determine tpre-post. We estimate the weight update for a given tpre-driver using Equation 8, approximating the integral by a summation over all time-discretized output responses consisting of 0, 1, or 2 spikes. Three or more spikes have a probability that is vanishingly small.
Sander M. Bohté, Michael C. Mozer
NIPS1
2004 The evidence for neural information processing with precise spike-times: A survey
Sander M. Bohté
Nat. Comput.1
2004 Introduction
Sander M. Bohté, Michiel C. van Wezel, Joost N. Kok
Nat. Comput.1
2004 Market-based recommendation: Agents that compete for consumer attention
abstract
The amount of attention space available for recommending suppliers to consumers on e-commerce sites is typically limited. We present a competitive distributed recommendation mechanism based on adaptive software agents for efficiently allocating the "consumer attention space," or banners. In the example of an electronic shopping mall, the task is delegated to the individual shops, each of which evaluates the information that is available about the consumer and his or her interests (e.g. keywords, product queries, and available parts of a profile). Shops make a monetary bid in an auction where a limited amount of "consumer attention space" for the arriving consumer is sold. Each shop is represented by a software agent that bids for each consumer. This allows shops to rapidly adapt their bidding strategy to focus on consumers interested in their offerings. For various basic and simple models for on-line consumers, shops, and profiles, we demonstrate the feasibility of our system by evolutionary simulations as in the field of agent-based computational economics (ACE). We also develop adaptive software agents that learn bidding-strategies, based on neural networks and strategy exploration heuristics. Furthermore, we address the commercial and technological advantages of this distributed market-based approach. The mechanism we describe is not limited to the example of the electronic shopping mall, but can easily be extended to other domains.
Sander M. Bohté, Enrico H. Gerding, Han La Poutré
ACM Trans. Internet Techn.1
2003 COllective INtelligence with Sequences of Actions - Coordinating Actions in Multi-agent Systems
Pieter Jan't Hoen, Sander M. Bohté
ECML2
2002 Modeling efficient conjunction detection with spiking neural networks
Sander M. Bohté, Joost N. Kok, Han La Poutré
ESANN1
2002 Error-backpropagation in temporally encoded networks of spiking neurons
Sander M. Bohté, Joost N. Kok, Han La Poutré
Neurocomputing1
2002 Unsupervised clustering with spiking neurons by sparse temporal coding and multilayer RBF networks
abstract
We demonstrate that spiking neural networks encoding information in the timing of single spikes are capable of computing and learning clusters from realistic data. We show how a spiking neural network based on spike-time coding and Hebbian learning can successfully perform unsupervised clustering on real-world data, and we demonstrate how temporal synchrony in a multilayer network can induce hierarchical clustering. We develop a temporal encoding of continuously valued data to obtain adjustable clustering capacity and precision with an efficient use of neurons: input variables are encoded in a population code by neurons with graded and overlapping sensitivity profiles. We also discuss methods for enhancing scale-sensitivity of the network and show how the induced synchronization of neurons within early RBF layers allows for the subsequent detection of complex clusters.
Sander M. Bohté, Han La Poutré, Joost N. Kok
IEEE Trans. Neural Networks1
2001 Competitive market-based allocation of consumer attention space
abstract
The amount of attention space available for recommending suppliers to consumers on e-commerce sites is typically limited. We present a competitive distributed recommendation mechanism based on adaptive software agents for efficiently allocating the "consumer attention space", or banners. In our approach, each agent bids in an auction for the momentary attention of each consumer. Successive auctions allow agents to rapidly adapt their bidding strategy to focus on consumers interested in their offerings. We demonstrate the feasibility of our system by an evolutionary simulation, and reflect on the advantages of this distributed market-based approach.
Sander M. Bohté, Enrico H. Gerding, Han La Poutré
EC1
2000 SpikeProp: backpropagation for networks of spiking neurons
Sander M. Bohté, Joost N. Kok, Han La Poutré
ESANN1
2000 Unsupervised Classification of Complex Clusters in Networks of Spiking Neurons
abstract
For unsupervised clustering in a network of spiking neurons we develop a temporal encoding of continuously valued data to obtain arbitrary clustering capacity and precision with an efficient use of neurons. Input variables are encoded independently in a population code by neurons with 1D graded and overlapping sensitivity profiles. Using a temporal Hebbian learning rule, the network architecture yields reliable clustering of high-dimensional multi-modal data. Additionally, multi-scale sensitivity to the input is achieved by using an appropriate choice of local activation functions. We present a multilayer version of the algorithm to perform a form of hierarchical clustering. We show how synchronous spiking of neurons can emerge with a local Hebbian learning rule and can be exploited by subsequent RBF layers employing the same local learning rule. Neuronal synchrony thus naturally enhances the clustering capabilities of artificial spiking neural networks, which has been widely suggested in neurobiology.
Sander M. Bohté, Han La Poutré, Joost N. Kok
IJCNN (3)1
2000 The Effects of Pair-wise and Higher-order Correlations on the Firing Rate of a Postsynaptic Neuron
abstract
Coincident firing of neurons projecting to a common target cell is likely to raise the probability of firing of this postsynaptic cell. Therefore, synchronized firing constitutes a significant event for postsynaptic neurons and is likely to play a role in neuronal information processing. Physiological data on synchronized firing in cortical networks are based primarily on paired recordings and cross-correlation analysis. However, pair-wise correlations among all inputs onto a postsynaptic neuron do not uniquely determine the distribution of simultaneous postsynaptic events. We develop a framework in order to calculate the amount of synchronous firing that, based on maximum entropy, should exist in a homogeneous neural network in which the neurons have known pair-wise correlations and higher-order structure is absent. According to the distribution of maximal entropy, synchronous events in which a large proportion of the neurons participates should exist even in the case of weak pair-wise correlations. Network simulations also exhibit these highly synchronous events in the case of weak pair-wise correlations. If such a group of neurons provides input to a common postsynaptic target, these network bursts may enhance the impact of this input, especially in the case of a high postsynaptic threshold. The proportion of neurons participating in synchronous bursts can be approximated by our method under restricted conditions. When these conditions are not fulfilled, the spike trains have less than maximal entropy, which is indicative of the presence of higher-order structure. In this situation, the degree of synchronicity cannot be derived from the pair-wise correlations.
Sander M. Bohté, Henk Spekreijse, Pieter R. Roelfsema
Neural Comput.1