Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Andrea Soltoggio

dblp:04/6283 · DBLP profile ↗
← Back
29ranked-venue papers
10as first author
8since 2021 · last 2026
0000-0002-9750-8358ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 27 · 10 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Reinforcement learning · 48% Transfer learning and domain adaptation · 24% Learning paradigms · 10%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Embedded and real-time systems · 40% Storage systems · 20% Cloud and datacenter computing · 20%

Topics — the 10 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
multi-agent reinforcement learning
1.012026
Policy Search, Retrieval, and Composition via Task Similarity in Collaborative Agentic Systems · AAAI 2026
Machine learning › Reinforcement learning › transfer learning in reinforcement learning
policy reuse
1.012026
Policy Search, Retrieval, and Composition via Task Similarity in Collaborative Agentic Systems · AAAI 2026
Machine learning › Transfer learning and domain adaptation
task similarity
1.012026
Policy Search, Retrieval, and Composition via Task Similarity in Collaborative Agentic Systems · AAAI 2026
Storage systems
deep reinforcement learning
0.712023
$\mathrm{R}^{3}$: On-Device Real-Time Deep Reinforcement Learning for Autonomous Robotics · RTSS 2023
Memory systems
memory management
0.712023
$\mathrm{R}^{3}$: On-Device Real-Time Deep Reinforcement Learning for Autonomous Robotics · RTSS 2023
Embedded and real-time systems › embedded machine learning
on-device training
0.712023
$\mathrm{R}^{3}$: On-Device Real-Time Deep Reinforcement Learning for Autonomous Robotics · RTSS 2023
Embedded and real-time systems
real-time scheduling
0.712023
$\mathrm{R}^{3}$: On-Device Real-Time Deep Reinforcement Learning for Autonomous Robotics · RTSS 2023
Cloud and datacenter computing
resource management
0.712023
$\mathrm{R}^{3}$: On-Device Real-Time Deep Reinforcement Learning for Autonomous Robotics · RTSS 2023
Machine learning › Learning paradigms
continual learning
0.412020
Sliced Cramer Synaptic Consolidation for Preserving Deeply Learned Representations · ICLR 2020
Knowledge, reasoning and agents › Multi-agent systems
agentic AI
0.312026
Policy Search, Retrieval, and Composition via Task Similarity in Collaborative Agentic Systems · AAAI 2026

Methods — techniques the papers use, named apart from their topics

wasserstein task embedding · 1.0policy masking · 1.0fine-tuning · 1.0cosine similarity · 1.0runtime profiling · 0.7dynamic batch sizing · 0.7deep reinforcement learning · 0.7sliced cramer distance · 0.4
YearPublicationVenuePosition
2026 Policy Search, Retrieval, and Composition via Task Similarity in Collaborative Agentic Systems
abstract
Agentic AI aims to create systems that set their own goals, adapt proactively to change, and refine behavior through continuous experience. Recent advances suggest that, when facing multiple and unforeseen tasks, agents could benefit from sharing machine-learned knowledge and reusing policies that have already been fully or partially learned by other agents. However, how to query, select, and retrieve policies from a pool of agents, and how to integrate such policies remains a largely unexplored area. This study explores how an agent decides what knowledge to select, from whom, and when and how to integrate it in its own policy in order to accelerate its own learning. The proposed algorithm, Modular Sharing and Composition in Collective Learning (MOSAIC), improves learning in agentic collectives by combining (1) knowledge selection using performance signals and cosine similarity on Wasserstein task embeddings, (2) modular and transferable neural representations via masks, and (3) policy integration, composition and fine-tuning. MOSAIC outperforms isolated learners and global sharing approaches in both learning speed and overall performance, and in some cases solves tasks that isolated agents cannot. The results also demonstrate that selective, goal-driven reuse leads to less susceptibility to task interference. We also observe the emergence of self-organization, where agents solving simpler tasks accelerate the learning of harder ones through shared knowledge.
Saptarshi Nath, Christos Peridis, Eseoghene Benjamin, Soheil Kolouri, Peter Kinnell, Zexin Li 0001, Cong Liu 0005, Shirin Dora, Andrea Soltoggio
AAAI10
2025 Design of electric powertrains to achieve NVH performance using autoencoders and a physical meaningful latent space
abstract
Abstract One of the fundamental differences in the perception of electric (e-) vehicles is how their radiated noise is perceived with respect to classic internal combustion engines. Even though e-vehicles are usually quieter, the tonal content of the radiated noise can be more annoying. This paper proposes a novel approach that starts from the assumed radiated noise spectrum profile as input to a neural network that can return powertrain design parameters that would lead to generation of that specific noise profile. The proposed network acts as an autoencoder where the latent space is forced to have a physical meaning. As diverse combinations of powertrain parameters can result in similar noise profiles, a variational autoencoder is used to learn a structured latent representation, ensuring continuity and smooth transitions between possible solutions. The network predictions are validated against results of a three-dimensional CAE e-powertrain model. Overall, the mean absolute error is around 5 dBA for this feasibility study, which aims to demonstrate the concept. This work takes an inverse approach to the optimisation problem by starting from the user-perceived noise to predict the parameters required to achieve that. Although this study focuses solely on gear teeth microgeometry changes and bearing preloads, additional powertrain parameters could be incorporated as needed.
Marcos Ricardo Souza, Günter Offner, Andrea Soltoggio, Mahdi Mohammadpour, Stephanos Theodossiades
Neural Comput. Appl.3
2025 Wasserstein task embedding for measuring task similarities
abstract
Measuring similarities between different tasks is critical in a broad spectrum of machine learning problems, including transfer, multi-task, continual, and meta-learning. Most current approaches to measuring task similarities are architecture-dependent: (1) relying on pre-trained models, or (2) training networks on tasks and using forward transfer as a proxy for task similarity. In this paper, we leverage the optimal transport theory and define a novel task embedding for supervised classification that is model-agnostic, training-free, and capable of handling (partially) disjoint label sets. In short, given a dataset with ground-truth labels, we perform a label embedding through multi-dimensional scaling and concatenate dataset samples with their corresponding label embeddings. Then, we define the distance between two datasets as the 2-Wasserstein distance between their updated samples. Lastly, we leverage the 2-Wasserstein embedding framework to embed tasks into a vector space in which the Euclidean distance between the embedded points approximates the proposed 2-Wasserstein distance between tasks. We show that the proposed embedding leads to a significantly faster comparison of tasks compared to related approaches like the Optimal Transport Dataset Distance (OTDD). Furthermore, we demonstrate the effectiveness of our embedding through various numerical experiments and show statistically significant correlations between our proposed distance and the forward and backward transfer among tasks on a wide variety of image recognition datasets.
Yikun Bai, Yuzhe Lu, Andrea Soltoggio, Soheil Kolouri
Neural Networks4
2024 SLoSH: Set Locality Sensitive Hashing via Sliced-Wasserstein Embeddings
abstract
Learning from set-structured data is an essential problem with many applications in machine learning and computer vision. This paper focuses on a non-parametric, data-independent, and efficient learning algorithm from setstructured data using optimal transport and approximate nearest neighbor (ANN) solutions, particularly localitysensitive hashing. We consider the problem of set retrieval from an input set query. This retrieval problem requires 1) an efficient mechanism to calculate the distances/dissimilarities between sets and 2) an appropriate data structure for a fast nearest-neighbor search. To that end, we propose to use Sliced-Wasserstein embedding as a computationally efficient "set-2-vector" operator that enables downstream ANN with theoretical guarantees. The set elements are treated as samples from an unknown underlying distribution, and the Sliced-Wasserstein distance is used to compare sets. We demonstrate the effectiveness of our algorithm, denoted as Set Locality Sensitive Hashing (SLoSH), on various set retrieval datasets and compare our proposed embedding with standard set embedding approaches, including Generalized Mean (GeM) embedding/pooling, Featurewise Sort Pooling (FSPool), Covariance Pooling, and Wasserstein embedding and show consistent improvement in retrieval results, both in terms of accuracy and computational efficiency.
Yuzhe Lu, Andrea Soltoggio, Soheil Kolouri
WACV3
2023 $\mathrm{R}^{3}$: On-Device Real-Time Deep Reinforcement Learning for Autonomous Robotics
abstract
Autonomous robotic systems, like autonomous vehicles and robotic search and rescue, require efficient on-device training for continuous adaptation of Deep Reinforcement Learning (DRL) models in dynamic environments. This research is fundamentally motivated by the need to understand and address the challenges of on-device real-time DRL, which involves balancing timing and algorithm performance under memory constraints, as exposed through our extensive empirical studies. This intricate balance requires co-optimizing two pivotal parameters of DRL training - batch size and replay buffer size. Configuring these parameters significantly affects timing and algorithm performance, while both (unfortunately) require substantial memory allocation to achieve near-optimal performance. This paper presents$\mathbf{R}^{3}$, a holistic solution for managing timing, memory, and algorithm performance in on-device real-time DRL training.$\mathbf{R}^{3}$employs (i) a deadline-driven feedback loop with dynamic batch sizing for optimizing timing, (ii) efficient memory management to reduce memory footprint and allow larger replay buffer sizes, and (iii) a runtime coordinator guided by heuristic analysis and a runtime profiler for dynamically adjusting memory resource reservations. These components collaboratively tackle the trade-offs in on-device DRL training, improving timing and algorithm performance while minimizing the risk of out-of-memory (OOM) errors. We implemented and evaluated$\mathbf{R}^{3}$extensively across various DRL frameworks and benchmarks on three hardware platforms commonly adopted by autonomous robotic systems. Additionally, we integrate$\mathbf{R}^{3}$with a popular realistic autonomous car simulator to demonstrate its real-world applicability. Evaluation results show that$\mathbf{R}^{3}$achieves efficacy across diverse platforms, ensuring consistent latency performance and timing predictability with minimal overhead. Moreover,$\mathbf{R}^{3}$showcases versatility by handling varied optimization goals and adapting to fluctuating systems scenarios.
Zexin Li 0001, Aritra Samanta, Yufei Li 0001, Andrea Soltoggio, Hyoseung Kim 0001, Cong Liu 0005
RTSS4
2023 A domain-agnostic approach for characterization of lifelong learning systems
Megan M. Baker, Alexander New, Mario Aguilar-Simon, Ziad Al-Halah, Sébastien M. R. Arnold, Eseoghene Benjamin, Andrew P. Brna, Ethan Brooks, Ryan C. Brown, Zachary A. Daniels, Anurag Reddy Daram, Fabien Delattre, Ryan Dellana, Eric Eaton, Haotian Fu, Kristen Grauman, Jesse Hostetler, Shariq Iqbal, Cassandra Kent, Nicholas Ketz, Soheil Kolouri, George Dimitri Konidaris, Dhireesha Kudithipudi, Erik G. Learned-Miller, Michael L. Littman, Sandeep Madireddy, Jorge A. Mendez, Eric Q. Nguyen, Christine D. Piatko, Praveen K. Pilly, Aswin Raghavan, Abrar Rahman, Santhosh K. Ramakrishnan, Neale Ratzlaff, Andrea Soltoggio, Peter Stone 0001, Indranil Sur, Zhipeng Tang, Saket Tiwari, Kyle Vedder, Felix Wang, Zifan Xu, Angel Yanguas-Gil, Harel Yedidsion, Shangqun Yu, Gautam K. Vallabha
Neural Networks36
2022 Context meta-reinforcement learning via neuromodulation
abstract
Meta-reinforcement learning (meta-RL) algorithms enable agents to adapt quickly to tasks from few samples in dynamic environments. Such a feat is achieved through dynamic representations in an agent's policy network (obtained via reasoning about task context, model parameter updates, or both). However, obtaining rich dynamic representations for fast adaptation beyond simple benchmark problems is challenging due to the burden placed on the policy network to accommodate different policies. This paper addresses the challenge by introducing neuromodulation as a modular component to augment a standard policy network that regulates neuronal activities in order to produce efficient dynamic representations for task adaptation. The proposed extension to the policy network is evaluated across multiple discrete and continuous control environments of increasing complexity. To prove the generality and benefits of the extension in meta-RL, the neuromodulated network was applied to two state-of-the-art meta-RL algorithms (CAVIA and PEARL). The result demonstrates that meta-RL augmented with neuromodulation produces significantly better result and richer dynamic representations in comparison to the baselines.
Eseoghene Benjamin, Jeffery Dick, Nicholas Ketz, Praveen K. Pilly, Andrea Soltoggio
Neural Networks5
2022 Deep Reinforcement Learning With Modulated Hebbian Plus Q-Network Architecture
abstract
In this article, we consider a subclass of partially observable Markov decision process (POMDP) problems which we termed confounding POMDPs. In these types of POMDPs, temporal difference (TD)-based reinforcement learning (RL) algorithms struggle, as TD error cannot be easily derived from observations. We solve these types of problems using a new bio-inspired neural architecture that combines a modulated Hebbian network (MOHN) with deep Q-network (DQN), which we call modulated Hebbian plus Q-network architecture (MOHQA). The key idea is to use a Hebbian network with rarely correlated bio-inspired neural traces to bridge temporal delays between actions and rewards when confounding observations and sparse rewards result in inaccurate TD errors. In MOHQA, DQN learns low-level features and control, while the MOHN contributes to high-level decisions by associating rewards with past states and actions. Thus, the proposed architecture combines two modules with significantly different learning algorithms, a Hebbian associative network and a classical DQN pipeline, exploiting the advantages of both. Simulations on a set of POMDPs and on the Malmo environment show that the proposed algorithm improved DQN's results and even outperformed control tests with advantage-actor critic (A2C), quantile regression DQN with long short-term memory (QRDQN + LSTM), Monte Carlo policy gradient (REINFORCE), and aggregated memory for reinforcement learning (AMRL) algorithms on most difficult POMDPs with confounding stimuli and sparse rewards.
Pawel Ladosz, Eseoghene Benjamin, Jeffery Dick, Nicholas Ketz, Soheil Kolouri, Jeffrey L. Krichmar, Praveen K. Pilly, Andrea Soltoggio
IEEE Trans. Neural Networks Learn. Syst.8
2020 Evolving inborn knowledge for fast adaptation in dynamic POMDP problems
abstract
Rapid online adaptation to changing tasks is an important problem in machine learning and, recently, a focus of meta-reinforcement learning. However, reinforcement learning (RL) algorithms struggle in POMDP environments because the state of the system, essential in a RL framework, is not always visible. Additionally, hand-designed meta-RL architectures may not include suitable computational structures for specific learning problems. The evolution of online learning mechanisms, on the contrary, has the ability to incorporate learning strategies into an agent that can (i) evolve memory when required and (ii) optimize adaptation speed to specific online learning problems. In this paper, we exploit the highly adaptive nature of neuromodulated neural networks to evolve a controller that uses the latent space of an autoencoder in a POMDP. The analysis of the evolved networks reveals the ability of the proposed algorithm to acquire inborn knowledge in a variety of aspects such as the detection of cues that reveal implicit rewards, and the ability to evolve location neurons that help with navigation. The integration of inborn knowledge and online plasticity enabled fast adaptation and better performance in comparison to some non-evolutionary meta-reinforcement learning algorithms. The algorithm proved also to succeed in the 3D gaming environment Malmo Minecraft.
Eseoghene Benjamin, Pawel Ladosz, Jeffery Dick, Wen-Hua Chen 0001, Praveen K. Pilly, Andrea Soltoggio
GECCO6
2020 Sliced Cramer Synaptic Consolidation for Preserving Deeply Learned Representations
Soheil Kolouri, Nicholas Ketz, Andrea Soltoggio, Praveen K. Pilly
ICLR3
2019 A fully convolutional two-stream fusion network for interactive image segmentation
Yang Hu 0004, Andrea Soltoggio, Russell Lock, Steve Carter
Neural Networks2
2018 Convolutional neural networks for automated targeted analysis of raw gas chromatography-mass spectrometry data
abstract
Through their breath, humans exhale hundreds of volatile organic compounds (VOCs) that can reveal pathologies, including many types of cancer at early stages. Gas chromatography-mass spectrometry (GC-MS) is an analytical method used to separate and detect compounds in the mixture contained in breath samples. The identification of VOCs is based on the recognition of their specific ion patterns in GC-MS data, which requires labour-intensive and time-consuming preprocessing and analysis by domain experts. This paper explores the original idea of applying supervised machine learning, and in particular convolutional neural networks (CNNs), to learn ion patterns directly from raw GC-MS data. The method adapts to machine specific characteristics, and once trained, can quickly analyse breath samples bypassing the time-consuming preprocessing phase. The CNN classification performance is compared to those of shallow neural networks and support vector machines. All considered machine learning tools achieved high accuracy in experiments with clinical data from participants. In particular, the CNN-based approach detected the lowest number of false positives. The results indicate that the proposed method is a promising tool to improve accuracy, specificity, and in particular speed in the detection of VOCs of interest in large-scale data analysis.
Angelika Skarysz, Yaser Alkhalifah, Kareen Darnley, Michael Eddleston, Yang Hu 0004, Duncan B. McLaren, William H. Nailon, Dahlia Salman, Martin Sykora, C. L. Paul Thomas, Andrea Soltoggio
IJCNN11
2018 Born to learn: The inspiration, progress, and future of evolved plastic artificial neural networks
Andrea Soltoggio, Kenneth O. Stanley, Sebastian Risi
Neural Networks1
2018 Distributed Task Rescheduling With Time Constraints for the Optimization of Total Task Allocations in a Multirobot System
abstract
This paper considers the problem of maximizing the number of task allocations in a distributed multirobot system under strict time constraints, where other optimization objectives need also be considered. It builds upon existing distributed task allocation algorithms, extending them with a novel method for maximizing the number of task assignments. The fundamental idea is that a task assignment to a robot has a high cost if its reassignment to another robot creates a feasible time slot for unallocated tasks. Multiple reassignments among networked robots may be required to create a feasible time slot and an upper limit to this number of reassignments can be adjusted according to performance requirements. A simulated rescue scenario with task deadlines and fuel limits is used to demonstrate the performance of the proposed method compared with existing methods, the consensus-based bundle algorithm and the performance impact (PI) algorithm. Starting from existing (PI-generated) solutions, results show up to a 20% increase in task allocations using the proposed method.
Joanna Turner, Qinggang Meng, Gerald Schaefer, Amanda Whitbrook, Andrea Soltoggio
IEEE Trans. Cybern.5
2017 Building Efficient Deep Hebbian Networks for Image Classification Tasks
Yanis Bahroun, Eugénie Hunsicker, Andrea Soltoggio
ICANN (1)3
2017 Online Representation Learning with Single and Multi-layer Hebbian Networks for Image Classification
Yanis Bahroun, Andrea Soltoggio
ICANN (1)2
2017 Neural Networks for Efficient Nonlinear Online Clustering
Yanis Bahroun, Eugénie Hunsicker, Andrea Soltoggio
ICONIP (1)3
2014 POET: An Evo-Devo Method to Optimize the Weights of Large Artificial Neural Networks
abstract
Large search spaces as those of artificial neural networks are difficult to search with machine learning techniques. The large amount of parameters is the main challenge for search techniques that do not exploit correlations expressed as patterns in the parameter space. Evolutionary computation with indirect genotype-phenotype mapping was proposed as a possible solution, but current methods often fail when the space is fractured and presents irregularities. This study employs an evolutionary indirect encoding inspired by developmental biology. Cellular proliferations and deletions of variable size allow for the definition of both regular large areas and small detailed areas in the parameter space. The method is tested on the search of the weights of a neural network for the classification of the MNIST dataset. The results demonstrate that even large networks such as those required for image classification can be effectively automatically designed by the proposed evolutionary developmental method. The combination of real-world problems like vision and classification, evolution and development, endows the proposed method with aspects of particular relevance to artificial life.
Alessandro Fontana, Andrea Soltoggio, Borys Wróbel
ALIFE2
2014 Real-time Hebbian Learning from Autoencoder Features for Control Tasks
abstract
Neural plasticity and in particular Hebbian learning play an important role in many research areas related to artficial life. By allowing artificial neural networks (ANNs) to adjust their weights in real time, Hebbian ANNs can adapt over their lifetime. However, even as researchers improve and extend Hebbian learning, a fundamental limitation of such systems is that they learn correlations between preexisting static fea-tures and network outputs. A Hebbian ANN could in principle achieve significantly more if it could accumulate new features over its lifetime from which to learn correlations. Interest-ingly, autoencoders, which have recently gained prominence in deep learning, are themselves in effect a kind of feature accumulator that extract meaningful features from their in-puts. The insight in this paper is that if an autoencoder is connected to a Hebbian learning layer, then the resulting Real-time Autoencoder-Augmented Hebbian Network (RAAHN) can actually learn new features (with the autoencoder) while si-multaneously learning control policies from those new features (with the Hebbian layer) in real time as an agent experiences its environment. In this paper, the RAAHN is shown in a simu-lated robot maze navigation experiment to enable a controller to learn the perfect navigation strategy significantly more of-ten than several Hebbian-based variant approaches that lack the autoencoder. In the long run, this approach opens up the intriguing possibility of real-time deep learning for control.
Justin K. Pugh, Andrea Soltoggio, Kenneth O. Stanley
ALIFE2
2013 Solving the Distal Reward Problem with Rare Correlations
abstract
In the course of trial-and-error learning, the results of actions, manifested as rewards or punishments, occur often seconds after the actions that caused them. How can a reward be associated with an earlier action when the neural activity that caused that action is no longer present in the network? This problem is referred to as the distal reward problem. A recent computational study proposes a solution using modulated plasticity with spiking neurons and argues that precise firing patterns in the millisecond range are essential for such a solution. In contrast, the study reported in this letter shows that it is the rarity of correlating neural activity, and not the spike timing, that allows the network to solve the distal reward problem. In this study, rare correlations are detected in a standard rate-based computational model by means of a threshold-augmented Hebbian rule. The novel modulated plasticity rule allows a randomly connected network to learn in classical and instrumental conditioning scenarios with delayed rewards. The rarity of correlations is shown to be a pivotal factor in the learning and in handling various delays of the reward. This study additionally suggests the hypothesis that short-term synaptic plasticity may implement eligibility traces and thereby serve as a selection mechanism in promoting candidate synapses for long-term storage.
Andrea Soltoggio, Jochen J. Steil
Neural Comput.1
2012 From modulated Hebbian plasticity to simple behavior learning through noise and weight saturation
Andrea Soltoggio, Kenneth O. Stanley
Neural Networks1
2011 Evolution of neural symmetry and its coupled alignment to body plan morphology
abstract
Body morphology is thought to have heavily influenced the evolution of neural architecture. However, the extent of this interaction and its underlying principles are largely unclear. To help us elucidate these principles, we examine the artificial evolution of a hypothetical nervous system embedded in a fish-inspired animat. The aim is to observe the evolution of neural structures in relation to both body morphology and required motor primitives. Our investigations reveal that increasing the pressure to evolve a wider range of movements also results in higher levels of neural symmetry. We further examine how different body shapes affect the evolution of neural structure; we find that, in order to achieve optimal movements, the neural structure integrates and compensates for asymmetrical body morphology. Our study clearly indicates that different parts of the animat - specifically, nervous system and body plan - evolve in concert with and become highly functional with respect to the other parts. The autonomous emergence of morphological and neural computation in this model contributes to unveiling the surprisingly strong coupling of such systems in nature.
Ben Jones, Andrea Soltoggio, Bernhard Sendhoff, Xin Yao 0001
GECCO2
2009 Novelty of behaviour as a basis for the neuro-evolution of operant reward learning
abstract
An agent that deviates from a usual or previous course of action can be said to display novel or varying behaviour. Novelty of behaviour can be seen as the result of real or apparent randomness in decision making, which prevents an agent from repeating exactly past choices. In this paper, novelty of behaviour is considered as an evolutionary precursor of the exploring skill in reward learning, and conservative behaviour as the precursor of exploitation. Novelty of behaviour in neural control is hypothesised to be an important factor in the neuro-evolution of operant reward learning. Agents capable of varying behaviour, as opposed to conservative, when exposed to reward stimuli appear to acquire on a faster evolutionary scale the meaning and use of such reward information. The hypothesis is validated by comparing the performance during evolution in two environments that either favour or are neutral to novelty. Following these findings, we suggest that neuro-evolution of operant reward learning is fostered by environments where behavioural novelty is intrinsically beneficial, i.e. where varying or exploring behaviour is associated with low risk.
Andrea Soltoggio, Ben Jones
GECCO1
2008 Evolutionary Advantages of Neuromodulated Plasticity in Dynamic, Reward-based Scenarios
Andrea Soltoggio, John A. Bullinaria, Claudio Mattiussi, Peter Dürr, Dario Floreano
ALIFE1
2008 Neural Plasticity and Minimal Topologies for Reward-Based Learning
abstract
Artificial neural networks for online learning problems are often implemented with synaptic plasticity to achieve adaptive behaviour. A common problem is that the overall learning dynamics are emergent properties strongly dependent on the correct combination of neural architectures, plasticity rules and environmental features. Which complexity in architectures and learning rules is required to match specific control and learning problems is not clear. Here a set of homosynaptic plasticity rules is applied to topologically unconstrained neural controllers while operating and evolving in dynamic reward-based scenarios. Performances are monitored on simulations of bee foraging problems and T-maze navigation. Varying reward locations compel the neural controllers to adapt their foraging strategies over time, fostering online reward-based learning. In contrast to previous studies, the results here indicate that reward-based learning in complex dynamic scenarios can be achieved with basic plasticity rules and minimal topologies.
Andrea Soltoggio
HIS1
2007 Evolving neuromodulatory topologies for reinforcement learning-like problems
abstract
Environments with varying reward contingencies constitute a challenge to many living creatures. In such conditions, animals capable of adaptation and learning derive an advantage. Recent studies suggest that neuromodulatory dynamics are a key factor in regulating learning and adaptivity when reward conditions are subject to variability. In biological neural networks, specific circuits generate modulatory signals, particularly in situations that involve learning cues such as a reward or novel stimuli. Modulatory signals are then broadcast and applied onto target synapses to activate or regulate synaptic plasticity. Artificial neural models that include modulatory dynamics could prove their potential in uncertain environments when online learning is required. However, a topology that synthesises and delivers modulatory signals to target synapses must be devised. So far, only handcrafted architectures of such kind have been attempted. Here we show that modulatory topologies can be designed autonomously by artificial evolution and achieve superior learning capabilities than traditional fixed-weight or Hebbian networks. In our experiments, we show that simulated bees autonomously evolved a modulatory network to maximise the reward in a reinforcement learning-like environment.
Andrea Soltoggio, Peter Dürr, Claudio Mattiussi, Dario Floreano
IEEE Congress on Evolutionary Computation1
2006 A simple line search operator for ridged landscapes
abstract
This paper describes a new simple operator for Evolutionary Algorithms (EA) to climb ridged landscapes.
Andrea Soltoggio
GECCO1
2005 An enhanced GA to improve the search process reliability in tuning of control systems
abstract
Evolutionary Algorithms (EAs) have been largely applied to optimisation and synthesis of controllers. In spite of several successful applications and competitive solutions, the stochastic nature of EAs and the uncertainty of the results have considerably hindered their use in industrial applications. In this paper we propose a Genetic Algorithm (GA) for tuning controllers for classical first and second order plants with actuator nonlinearities. To increase the robustness of the algorithm we introduce two features: 1) genetic operators that perform directional mutations, 2) selection tournaments organized by genome vicinity. The experiment results show that the proposed GA is able to guarantee high performance and low variance in the results from different runs. The increased reliability, compared to the results from a classical GA, seems to favour particularly the application of Evolutionary Computation (EC) in tuning of control systems, where, thanks to this approach, a large search space can be searched repeatedly with high consistency in the solutions.
Andrea Soltoggio
GECCO1
2004 A Comparison of Genetic Programming and Genetic Algorithms in the Design of a Robust, Saturated Control System
Andrea Soltoggio
GECCO (2)1