Ann Nowé

dblp:95/232 · DBLP profile ↗
← Back
137ranked-venue papers
4as first author
30since 2021 · last 2026
0000-0001-6346-4564ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 105 · 1 first-author · 25 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 6 · 2 first-authorHuman-computer interaction and ubiquitous computing · 5 · 1 first-authorComputer networks · 4 · 1 since 2021Systems, architecture and hardware · 3 · 1 first-authorTheory of computation · 3Software engineering, systems software and programming languages · 2
YearPublicationVenuePosition
2026 Generalized policy improvement for efficient and robust multi-objective reinforcement learning
abstract
Multi-objective reinforcement learning (MORL) algorithms tackle sequential decision problems where agents may have different preferences over (possibly conflicting) reward functions. These algorithms often learn a set of policies, each optimized for a particular agent preference, that are later reused when optimizing policies for different preferences. We introduce a novel algorithm that builds upon Generalized Policy Improvement (GPI) to construct principled, formally-derived prioritization schemes that improve sample efficiency. These correspond to active-learning strategies by which the agent can identify (i) the most promising preferences/objectives to train on at each moment; and (ii) the most relevant previous experiences to learn policies for new agent preferences through a novel Dyna-style MORL method. We prove our algorithm is guaranteed to always converge to an optimal solution in a finite number of steps, or an $$\epsilon $$ -optimal solution (for a bounded $$\epsilon $$ ) if the agent can only identify sub-optimal policies. Our method monotonically improves the quality of its partial solutions while learning. We also introduce a bound that characterizes the maximum utility loss (with respect to the optimal solution) incurred by intermediate policies identified by our method during learning. Finally, we propose a novel epistemic uncertainty-aware extension of GPI that exploits high-confidence lower bounds to mitigate the impact of unreliable action-value estimates in GPI policies, and prove that it provides tighter performance bounds than the current state of the art. We empirically show that our method outperforms state-of-the-art MORL algorithms in challenging multi-objective tasks.
Lucas Nunes Alegre, Ana L. C. Bazzan, Diederik M. Roijers, Ann Nowé, Bruno C. da Silva 0001
Auton. Agents Multi Agent Syst.4
2026 Multi-objective reinforcement learning for provably incentivising alignment with value systems
abstract
This paper addresses the problem of ensuring that autonomous learning agents align with multiple moral values. Specifically, we present the theoretical principles and algorithmic tools necessary for creating an environment where we ensure that the agent learns a behaviour aligned with multiple moral values while striving to achieve its individual objective. To address this value alignment problem, we adopt the Multi-Objective Reinforcement Learning framework and propose a novel algorithm that combines techniques from Multi-Objective Reinforcement Learning and Linear Programming. In addition, we illustrate our value alignment process with an example involving an autonomous vehicle. Here, we demonstrate that the agent learns to behave in alignment with the ethical values of safety, achievement, and comfort, with achievement representing the agent’s individual objective. Such ethical behaviour differs depending on the ordering between values. We also use a synthetic multi-objective environment to evaluate the computational costs of guaranteeing ethical learning as the number of values increases.
Manel Rodriguez-Soto, Roxana Radulescu, Filippo Bistaffa, Oriol Ricart, Arnau Mayoral-Macau, Maite López-Sánchez, Juan A. Rodríguez-Aguilar, Ann Nowé
Artif. Intell.8
2026 Benchmarking knowledge graph embedding models for the prediction of oligogenic combinations
abstract
Identifying the potential oligogenic causes of rare diseases remains a challenge, notwithstanding the advancements made in the last decade. While a variety of predictive and ranking approaches have been proposed, their precision remains limited, as only a small number of high-quality training cases are available and it remains difficult to know which features may be most relevant for the design of new predictors. We hypothesize here that structured biological information, which provides an integration of various relevant biological networks and ontologies in a single heterogeneous knowledge graph, can make a difference as it allows for learning a relevant genetic representation through KGE methods. An exhaustive benchmarking is performed here wherein we assess the performance of various state-of-the-art embedding models for the task of identifying potentially pathogenic gene pairs. The results obtained show that these KGE provide highly accurate predictions, leading to an Area Under the Precision-Recall Curve of up to $0.93$, representing also a significant advancement over previous approaches for predicting gene pairs involved in oligogenic diseases. We show nonetheless that care needs to be taken in the cross-validation when using embeddings, as data leakage between folds in embedding space will reveal overly optimistic results. The further evaluation of the methods on a holdout set as well as on a group of new male infertility cases show that three Translational Distance models (TransE, MurE, and RotatE) and two of the Semantic Matching models (DisMult and QuatE) provide the better results. The analysis is concluded by comparing all known gene combinations for these top-ranking models, examining their similarities and differences. Overall, KGE provide a predictive advancement but new steps will need to be taken generate explanations as to why the pairs are relevant for oligogenic diseases.
Inas Bosch, Barbara Gravel, Alexandre Renaux, Ann Nowé, Maris Laan, Tom Lenaerts
Briefings Bioinform.4
2025 Congestion-Aware Multi-Agent Path Planning for Pick-Up and Delivery Tasks
abstract
Mobile robotic systems play a pivotal role in logistics, particularly in warehouse operations, where efficient and collision-free navigation is essential for completing tasks. However, managing a large number of robots often leads to congestion, causing delays and adversely affecting system scalability. This paper proposes a novel online algorithm for solving the Multi-Agent Pickup and Delivery (MAPD) problem. The algorithm addresses local collision detection and global congestion avoidance by integrating a congestion prediction model to enhance process efficiency. A deep neural network is employed to approximate congestion predictions independently of the number of agents, reducing computational complexity. Simulation experiments demonstrate that the proposed approach significantly improves system throughput and scalability, with a notable average doubling of throughput in specific scenarios. The findings provide a foundation for advanced congestion management strategies in multi-agent systems, paving the way for efficient and scalable deployment in logistics and beyond.
Mehrdad Asadi, Ann Nowé, Javad Ghofrani
GECCO2
2025 Reinforcement Learning for Model-Free Control of a Cooling Network with Uncertain Future Demands
abstract
Optimal control of complex systems often requires access to a high-fidelity model, and information about the (future) external stimuli applied to the system (load, demand, ...).An example of such a system is a cooling network, in which one or more chillers provide cooled liquid to a set of users with a variable demand.In this paper, we propose a Reinforcement Learning (RL) method for such a system with 3 chillers.It does not assume any model, and does not observe the future cooling demand, nor approximations of it.Still, we show that, after a training phase in a simulator, the learned controller achieves a performance better than classical rule-based controllers, and similar to a model predictive controller that does rely on a model and demand predictions.We show that the RL algorithm has learned implicitly how to anticipate, without requiring explicit predictions.This demonstrates that RL can allow to produce high-quality controllers in challenging industrial contexts.
Jeroen Willems, Denis Steckelmacher, Wouter Scholte, Bruno Depraetere, Edward Kikken, Abdellatif Bey-Temsamani, Ann Nowé
ICINCO (1)7
2025 A JAX-Accelerated Simulation Framework for Multi-Agent Energy Management in Energy Communities
Hicham Azmani, Andries Rosseau, Marjon Blondeel, Ann Nowé
AAMAS4
2025 Composing Reinforcement Learning Policies, with Formal Guarantees
Florent Delgrange, Guy Avni, Anna Lukina, Christian Schilling 0001, Ann Nowé, Guillermo A. Pérez
AAMAS5
2025 Curiosity-Driven Partner Selection Accelerates Convention Emergence in Language Games
Chin-Wing Leung, Paolo Turrini, Ann Nowé
AAMAS3
2025 Divide and Conquer: Provably Unveiling the Pareto Front with Multi-Objective Reinforcement Learning
Willem Röpke, Mathieu Reymond, Patrick Mannion, Diederik M. Roijers, Ann Nowé, Roxana Radulescu
AAMAS5
2025 Collective Intelligence in Decision-Making with Non-Stationary Experts
abstract
When sufficient experience to make informed decisions is unavailable, expert advice can help us navigate uncertainty. As expertise evolves, driven by continuous learning in human experts or model updates in artificial experts, it is crucial to adopt adaptive approaches. Existing methods for exploiting non-stationary experts focus on competing with the single best expert. In contrast, this work harnesses the power of collective intelligence to facilitate better decision-making in the face of evolving expertise or dynamic environments. To achieve this, we propose the novel CORVAL approach which optimally combines the insights of multiple experts. By adapting to drifts in expertise, our novel approach can surpass the performance of the single best expert as well as previous approaches. Empirical evaluations on a diverse range of non-stationary problems, including active learning applications, showcase the improved performance of our approach in collective decision-making scenarios.
Axel Abels, Vito Trianni, Ann Nowé, Tom Lenaerts
J. Artif. Intell. Res.3
2025 A framework for flexibly guiding learning agents
Mahmoud Elbarbari, Florent Delgrange, Ivo Vervlimmeren, Kyriakos Efthymiadis, Bram Vanderborght, Ann Nowé
Neural Comput. Appl.6
2025 Preference communication in multi-objective normal-form games
Willem Röpke, Diederik M. Roijers, Ann Nowé, Roxana Radulescu
Neural Comput. Appl.3
2024 The Wasserstein Believer: Learning Belief Updates for Partially Observable Environments through Reliable Latent Space Models
abstract
Partially Observable Markov Decision Processes (POMDPs) are used to model environments where the state cannot be perceived, necessitating reasoning based on past observations and actions. However, remembering the full history is generally intractable due to the exponential growth in the history space. Maintaining a probability distribution that models the belief over the current state can be used as a sufficient statistic of the history, but its computation requires access to the model of the environment and is often intractable. While SOTA algorithms use Recurrent Neural Networks to compress the observation-action history aiming to learn a sufficient statistic, they lack guarantees of success and can lead to sub-optimal policies. To overcome this, we propose the Wasserstein Belief Updater, an RL algorithm that learns a latent model of the POMDP and an approximation of the belief update under the assumption that the state is observable during training. Our approach comes with theoretical guarantees on the quality of our approximation ensuring that our latent beliefs allow for learning the optimal value function.
Raphaël Avalos, Florent Delgrange, Ann Nowé, Guillermo A. Pérez, Diederik M. Roijers
ICLR3
2024 A Systematic Analysis of Deep Learning Algorithms in High-Dimensional Data Regimes of Limited Size
abstract
There is a substantial demand for deep learning methods that can work with limited, high-dimensional, and noisy datasets. Nonetheless, current research mostly neglects this area, especially in the absence of prior expert knowledge or knowledge transfer. In this work, we bridge this gap by studying the performance of deep learning methods on the true data distribution in a limited, high-dimensional, and noisy data setting. To this end, we conduct a systematic evaluation that reduces the available training data while retaining the challenging properties mentioned above. Furthermore, we extensively search the space of hyperparameters and compare state-of-the-art architectures and models built and trained from scratch to advocate for the use of multi-objective tuning strategies. Our experiments highlight the lack of performative deep learning models in current literature and investigate the impact of training hyperparameters. We analyze the complexity of the models and demonstrate the advantage of choosing models tuned under multi-objective criteria in lower data regimes to reduce the likelihood to overfit. Lastly, we demonstrate the importance of selecting a proper inductive bias given a limited-sized dataset. Given our results, we conclude that tuning models using a multi-objective criterion results in simpler yet competitive models when reducing the number of data points.
Simon Jaxy, Ann Nowé, Pieter Libin
ICTAI2
2024 Prioritization of oligogenic variant combinations in whole exomes
abstract
MOTIVATION: Whole exome sequencing (WES) has emerged as a powerful tool for genetic research, enabling the collection of a tremendous amount of data about human genetic variation. However, properly identifying which variants are causative of a genetic disease remains an important challenge, often due to the number of variants that need to be screened. Expanding the screening to combinations of variants in two or more genes, as would be required under the oligogenic inheritance model, simply blows this problem out of proportion. RESULTS: We present here the High-throughput oligogenic prioritizer (Hop), a novel prioritization method that uses direct oligogenic information at the variant, gene and gene pair level to detect digenic variant combinations in WES data. This method leverages information from a knowledge graph, together with specialized pathogenicity predictions in order to effectively rank variant combinations based on how likely they are to explain the patient's phenotype. The performance of Hop is evaluated in cross-validation on 36 120 synthetic exomes for training and 14 280 additional synthetic exomes for independent testing. Whereas the known pathogenic variant combinations are found in the top 20 in approximately 60% of the cross-validation exomes, 71% are found in the same ranking range when considering the independent set. These results provide a significant improvement over alternative approaches that depend simply on a monogenic assessment of pathogenicity, including early attempts for digenic ranking using monogenic pathogenicity scores. AVAILABILITY AND IMPLEMENTATION: Hop is available at https://github.com/oligogenic/HOP.
Barbara Gravel, Alexandre Renaux, Sofia Papadimitriou, Guillaume Smits, Ann Nowé, Tom Lenaerts
Bioinform.5
2024 Exploring the Pareto front of multi-objective COVID-19 mitigation policies using reinforcement learning
abstract
Infectious disease outbreaks can have a disruptive impact on public health and societal processes. As decision-making in the context of epidemic mitigation is multi-dimensional hence complex, reinforcement learning in combination with complex epidemic models provides a methodology to design refined prevention strategies. Current research focuses on optimizing policies with respect to a single objective, such as the pathogen’s attack rate. However, as the mitigation of epidemics involves distinct, and possibly conflicting, criteria (i.a., mortality, morbidity, economic cost, well-being), a multi-objective decision approach is warranted to obtain balanced policies. To enhance future decision-making, we propose a deep multi-objective reinforcement learning approach by building upon a state-of-the-art algorithm called Pareto Conditioned Networks (PCN) to obtain a set of solutions for distinct outcomes of the decision problem. We consider different deconfinement strategies after the first Belgian lockdown within the COVID-19 pandemic and aim to minimize both COVID-19 cases (i.e., infections and hospitalizations) and the societal burden induced by the mitigation measures. As such, we connected a multi-objective Markov decision process with a stochastic compartment model designed to approximate the Belgian COVID-19 waves and explore reactive strategies. As these social mitigation measures are implemented in a continuous action space that modulates the contact matrix of the age-structured epidemic model, we extend PCN to this setting. We evaluate the solution set that PCN returns, and observe that it explored the whole range of possible social restrictions, leading to high-quality trade-offs, as it captured the problem dynamics. In this work, we demonstrate that multi-objective reinforcement learning adds value to epidemiological modeling and provides essential insights to balance mitigation policies.
Mathieu Reymond, Conor F. Hayes, Lander Willem, Roxana Radulescu, Steven Abrams, Diederik M. Roijers, Enda Howley, Patrick Mannion, Niel Hens, Ann Nowé, Pieter Libin
Expert Syst. Appl.10
2024 Dynamic Size Message Scheduling for Multi-Agent Communication Under Limited Bandwidth
abstract
Communication plays a vital role in multi-agent systems, fostering collaboration and coordination. However, in real-world scenarios where communication is bandwidth-limited, existing multi-agent reinforcement learning (MARL) algorithms often provide agents with a binary choice: either transmitting a fixed amount of data or no information at all. This rigid communication strategy hinders the ability to effectively utilize bandwidth. To overcome this challenge, we present the Dynamic Size Message Scheduling (DSMS) method, which introduces finer-grained communication scheduling by considering the actual size of the information being exchanged. Our approach lies in adapting message sizes using Fourier transform-based compression techniques with clipping, enabling agents to tailor their messages to match the allocated bandwidth according to importance weights. This method realizes a balance between information loss and bandwidth utilization. Receiving agents reliably decompress the messages using the inverse Fourier transform. We evaluate DSMS in cooperative tasks where the agent has partial observability. Experimental results demonstrate that DSMS significantly improves performance by optimizing the utilization of bandwidth and effectively balancing information importance.
Qingshuang Sun, Denis Steckelmacher, Yuan Yao 0004, Ann Nowé, Raphaël Avalos
IEEE Trans. Mob. Comput.4
2023 Wasserstein Auto-encoded MDPs: Formal Verification of Efficiently Distilled RL Policies with Many-sided Guarantees
Florent Delgrange, Ann Nowé, Guillermo A. Pérez
ICLR2
2023 Expertise Trees Resolve Knowledge Limitations in Collective Decision-Making
abstract
Experts advising decision-makers are likely to display expertise which varies as a function of the problem instance. In practice, this may lead to sub-optimal or discriminatory decisions against minority cases. In this work, we model such changes in depth and breadth of knowledge as a partitioning of the problem space into regions of differing expertise. We provide here new algorithms that explicitly consider and adapt to the relationship between problem instances and experts’ knowledge. We first propose and highlight the drawbacks of a naive approach based on nearest neighbor queries. To address these drawbacks we then introduce a novel algorithm — expertise trees — that constructs decision trees enabling the learner to select appropriate models. We provide theoretical insights and empirically validate the improved performance of our novel approach on a range of problems for which existing methods proved to be inadequate.
Axel Abels, Tom Lenaerts, Vito Trianni, Ann Nowé
ICML4
2023 Distributional Multi-Objective Decision Making
abstract
For effective decision support in scenarios with conflicting objectives, sets of potentially optimal solutions can be presented to the decision maker. We explore both what policies these sets should contain and how such sets can be computed efficiently. With this in mind, we take a distributional approach and introduce a novel dominance criterion relating return distributions of policies directly. Based on this criterion, we present the distributional undominated set and show that it contains optimal policies otherwise ignored by the Pareto front. In addition, we propose the convex distributional undominated set and prove that it comprises all policies that maximise expected utility for multivariate risk-averse decision makers. We propose a novel algorithm to learn the distributional undominated set and further contribute pruning operators to reduce the set to the convex distributional undominated set. Through experiments, we demonstrate the feasibility and effectiveness of these methods, making this a valuable new approach for decision support in real-world problems.
Willem Röpke, Conor F. Hayes, Patrick Mannion, Enda Howley, Ann Nowé, Diederik M. Roijers
IJCAI5
2023 Multi-Step Generalized Policy Improvement by Leveraging Approximate Models
abstract
We introduce a principled method for performing zero-shot transfer in reinforcement learning (RL) by exploiting approximate models of the environment. Zero-shot transfer in RL has been investigated by leveraging methods rooted in generalized policy improvement (GPI) and successor features (SFs). Although computationally efficient, these methods are model-free: they analyze a library of policies---each solving a particular task---and identify which action the agent should take. We investigate the more general setting where, in addition to a library of policies, the agent has access to an approximate environment model. Even though model-based RL algorithms can identify near-optimal policies, they are typically computationally intensive. We introduce $h$-GPI, a multi-step extension of GPI that interpolates between these extremes---standard model-free GPI and fully model-based planning---as a function of a parameter, $h$, regulating the amount of time the agent has to reason. We prove that $h$-GPI's performance lower bound is strictly better than GPI's, and show that $h$-GPI generally outperforms GPI as $h$ increases. Furthermore, we prove that as $h$ increases, $h$-GPI's performance becomes arbitrarily less susceptible to sub-optimality in the agent's policy library. Finally, we introduce novel bounds characterizing the gains achievable by $h$-GPI as a function of approximation errors in both the agent's policy library and its (possibly learned) model. These bounds strictly generalize those known in the literature. We evaluate $h$-GPI on challenging tabular and continuous-state problems under value function approximation and show that it consistently outperforms GPI and state-of-the-art competing methods under various levels of approximation errors.
Lucas Nunes Alegre, Ana L. C. Bazzan, Ann Nowé, Bruno C. da Silva 0001
NeurIPS3
2023 A Toolkit for Reliable Benchmarking and Research in Multi-Objective Reinforcement Learning
abstract
Multi-objective reinforcement learning algorithms (MORL) extend standard reinforcement learning (RL) to scenarios where agents must optimize multiple---potentially conflicting---objectives, each represented by a distinct reward function. To facilitate and accelerate research and benchmarking in multi-objective RL problems, we introduce a comprehensive collection of software libraries that includes: (i) MO-Gymnasium, an easy-to-use and flexible API enabling the rapid construction of novel MORL environments. It also includes more than 20 environments under this API. This allows researchers to effortlessly evaluate any algorithms on any existing domains; (ii) MORL-Baselines, a collection of reliable and efficient implementations of state-of-the-art MORL algorithms, designed to provide a solid foundation for advancing research. Notably, all algorithms are inherently compatible with MO-Gymnasium; and(iii) a thorough and robust set of benchmark results and comparisons of MORL-Baselines algorithms, tested across various challenging MO-Gymnasium environments. These benchmarks were constructed to serve as guidelines for the research community, underscoring the properties, advantages, and limitations of each particular state-of-the-art method.
Florian Felten, Lucas Nunes Alegre, Ann Nowé, Ana L. C. Bazzan, El-Ghazali Talbi, Grégoire Danoy, Bruno C. da Silva 0001
NeurIPS3
2023 Actor-critic multi-objective reinforcement learning for non-linear utility functions
Mathieu Reymond, Conor F. Hayes, Denis Steckelmacher, Diederik M. Roijers, Ann Nowé
Auton. Agents Multi Agent Syst.5
2023 Dealing with expert bias in collective decision-making
Axel Abels, Tom Lenaerts, Vito Trianni, Ann Nowé
Artif. Intell.4
2023 A knowledge graph approach to predict and interpret disease-causing gene interactions
abstract
BACKGROUND: Understanding the impact of gene interactions on disease phenotypes is increasingly recognised as a crucial aspect of genetic disease research. This trend is reflected by the growing amount of clinical research on oligogenic diseases, where disease manifestations are influenced by combinations of variants on a few specific genes. Although statistical machine-learning methods have been developed to identify relevant genetic variant or gene combinations associated with oligogenic diseases, they rely on abstract features and black-box models, posing challenges to interpretability for medical experts and impeding their ability to comprehend and validate predictions. In this work, we present a novel, interpretable predictive approach based on a knowledge graph that not only provides accurate predictions of disease-causing gene interactions but also offers explanations for these results. RESULTS: We introduce BOCK, a knowledge graph constructed to explore disease-causing genetic interactions, integrating curated information on oligogenic diseases from clinical cases with relevant biomedical networks and ontologies. Using this graph, we developed a novel predictive framework based on heterogenous paths connecting gene pairs. This method trains an interpretable decision set model that not only accurately predicts pathogenic gene interactions, but also unveils the patterns associated with these diseases. A unique aspect of our approach is its ability to offer, along with each positive prediction, explanations in the form of subgraphs, revealing the specific entities and relationships that led to each pathogenic prediction. CONCLUSION: Our method, built with interpretability in mind, leverages heterogenous path information in knowledge graphs to predict pathogenic gene interactions and generate meaningful explanations. This not only broadens our understanding of the molecular mechanisms underlying oligogenic diseases, but also presents a novel application of knowledge graphs in creating more transparent and insightful predictors for genetic research.
Alexandre Renaux, Chloé Terwagne, Michael Cochez, Ilaria Tiddi, Ann Nowé, Tom Lenaerts
BMC Bioinform.5
2023 Faster and more accurate pathogenic combination predictions with VarCoPP2.0
abstract
BACKGROUND: The prediction of potentially pathogenic variant combinations in patients remains a key task in the field of medical genetics for the understanding and detection of oligogenic/multilocus diseases. Models tailored towards such cases can help shorten the gap of missing diagnoses and can aid researchers in dealing with the high complexity of the derived data. The predictor VarCoPP (Variant Combinations Pathogenicity Predictor) that was published in 2019 and identified potentially pathogenic variant combinations in gene pairs (bilocus variant combinations), was the first important step in this direction. Despite its usefulness and applicability, several issues still remained that hindered a better performance, such as its False Positive (FP) rate, the quality of its training set and its complex architecture. RESULTS: We present VarCoPP2.0: the successor of VarCoPP that is a simplified, faster and more accurate predictive model identifying potentially pathogenic bilocus variant combinations. Results from cross-validation and on independent data sets reveal that VarCoPP2.0 has improved in terms of both sensitivity (95% in cross-validation and 98% during testing) and specificity (5% FP rate). At the same time, its running time shows a significant 150-fold decrease due to the selection of a simpler Balanced Random Forest model. Its positive training set now consists of variant combinations that are more confidently linked with evidence of pathogenicity, based on the confidence scores present in OLIDA, the Oligogenic Diseases Database ( https://olida.ibsquare.be ). The improvement of its performance is also attributed to a more careful selection of up-to-date features identified via an original wrapper method. We show that the combination of different variant and gene pair features together is important for predictions, highlighting the usefulness of integrating biological information at different levels. CONCLUSIONS: Through its improved performance and faster execution time, VarCoPP2.0 enables a more accurate analysis of larger data sets linked to oligogenic diseases. Users can access the ORVAL platform ( https://orval.ibsquare.be ) to apply VarCoPP2.0 on their data.
Nassim Versbraegen, Barbara Gravel, Charlotte Nachtegael, Alexandre Renaux, Emma Verkinderen, Ann Nowé, Tom Lenaerts, Sofia Papadimitriou
BMC Bioinform.6
2022 Distillation of RL Policies with Formal Guarantees via Variational Abstraction of Markov Decision Processes
abstract
We consider the challenge of policy simplification and verification in the context of policies learned through reinforcement learning (RL) in continuous environments. In well-behaved settings, RL algorithms have convergence guarantees in the limit. While these guarantees are valuable, they are insufficient for safety-critical applications. Furthermore, they are lost when applying advanced techniques such as deep-RL. To recover guarantees when applying advanced RL algorithms to more complex environments with (i) reachability, (ii) safety-constrained reachability, or (iii) discounted-reward objectives, we build upon the DeepMDP framework to derive new bisimulation bounds between the unknown environment and a learned discrete latent model of it. Our bisimulation bounds enable the application of formal methods for Markov decision processes. Finally, we show how one can use a policy obtained via state-of-the-art RL to efficiently train a variational autoencoder that yields a discrete latent model with provably approximately correct bisimulation guarantees. Additionally, we obtain a distilled version of the policy for the latent model.
Florent Delgrange, Ann Nowé, Guillermo A. Pérez
AAAI2
2022 A practical guide to multi-objective reinforcement learning and planning
abstract
Abstract Real-world sequential decision-making tasks are generally complex, requiring trade-offs between multiple, often conflicting, objectives. Despite this, the majority of research in reinforcement learning and decision-theoretic planning either assumes only a single objective, or that multiple objectives can be adequately handled via a simple linear combination. Such approaches may oversimplify the underlying problem and hence produce suboptimal results. This paper serves as a guide to the application of multi-objective methods to difficult problems, and is aimed at researchers who are already familiar with single-objective reinforcement learning and planning methods who wish to adopt a multi-objective perspective on their research, as well as practitioners who encounter multi-objective decision problems in practice. It identifies the factors that may influence the nature of the desired solution, and illustrates by example how these influence the design of multi-objective decision-making systems for complex problems.
Conor F. Hayes, Roxana Radulescu, Eugenio Bargiacchi, Johan Källström, Matthew Macfarlane, Mathieu Reymond, Timothy Verstraeten, Luisa M. Zintgraf, Richard Dazeley, Fredrik Heintz, Enda Howley, Athirai Aravazhi Irissappane, Patrick Mannion, Ann Nowé, Gabriel de Oliveira Ramos, Marcello Restelli, Peter Vamplew 0001, Diederik M. Roijers
Auton. Agents Multi Agent Syst.14
2022 On nash equilibria in normal-form games with vectorial payoffs
Willem Röpke, Diederik M. Roijers, Ann Nowé, Roxana Radulescu
Auton. Agents Multi Agent Syst.3
2022 Opponent learning awareness and modelling in multi-objective normal form games
Roxana Radulescu, Timothy Verstraeten, Patrick Mannion, Diederik M. Roijers, Ann Nowé
Neural Comput. Appl.6
2020 Fleet Control Using Coregionalized Gaussian Process Policy Iteration
abstract
In many settings, as for example wind farms, multiple machines are instantiated to perform the same task, which is called a fleet. The recent advances with respect to the Internet of Things allow control devices and/or machines to connect through cloud-based architectures in order to share information about their status and environment. Such an infrastructure allows seamless data sharing between fleet members, which could greatly improve the sample-efficiency of reinforcement learning techniques. However in practice, these machines, while almost identical in design, have small discrepancies due to production errors or degradation, preventing control algorithms to simply aggregate and employ all fleet data. We propose a novel reinforcement learning method that learns to transfer knowledge between similar fleet members and creates member-specific dynamics models for control. Our algorithm uses Gaussian processes to establish cross-member covariances. This is significantly different from standard transfer learning methods, as the focus is not on sharing information over tasks, but rather over system specifications. We demonstrate our approach on two benchmarks and a realistic wind farm setting. Our method significantly outperforms two baseline approaches, namely individual learning and joint learning where all samples are aggregated, in terms of the median and variance of the results.
Timothy Verstraeten, Pieter Libin, Ann Nowé
ECAI3
2020 An Interpretable Semi-supervised Classifier using Rough Sets for Amended Self-labeling
abstract
Semi-supervised classifiers combine labeled and unlabeled data during the learning phase in order to increase classifier's generalization capability. However, most successful semi-supervised classifiers involve complex ensemble structures and iterative algorithms which make it difficult to explain the outcome, thus behaving like black boxes. Furthermore, during an iterative self-labeling process, mistakes can be propagated if no amending procedure is used. In this paper, we build upon an interpretable self-labeling grey-box classifier that uses a black box to estimate the missing class labels and a white box to make the final predictions. We propose a Rough Set based approach for amending the self-labeling process. We compare its performance to the vanilla version of our self-labeling grey-box and the use of a confidence-based amending. In addition, we introduce some measures to quantify the interpretability of our model. The experimental results suggest that the proposed amending improves accuracy and interpretability of the self-labeling grey-box, thus leading to superior results when compared to state-of-the-art semi-supervised classifiers.
Isel Grau, Dipankar Sengupta, María Matilde García Lorenzo, Ann Nowé
FUZZ-IEEE4
2020 Collective Decision-Making as a Contextual Multi-armed Bandit Problem
Axel Abels, Tom Lenaerts, Vito Trianni, Ann Nowé
ICCCI4
2020 How Expert Confidence Can Improve Collective Decision-Making in Contextual Multi-Armed Bandit Problems
Axel Abels, Tom Lenaerts, Vito Trianni, Ann Nowé
ICCCI4
2020 Interactive Multi-objective Reinforcement Learning in Multi-armed Bandits with Gaussian Process Utility Models
Diederik M. Roijers, Luisa M. Zintgraf, Pieter Libin, Mathieu Reymond, Eugenio Bargiacchi, Ann Nowé
ECML/PKDD (3)6
2020 Multi-objective multi-agent decision making: a utility-based analysis and survey
Roxana Radulescu, Patrick Mannion, Diederik M. Roijers, Ann Nowé
Auton. Agents Multi Agent Syst.4
2020 AI-Toolbox: A C++ library for Reinforcement Learning and Planning (with Python Bindings)
abstract
This paper describes AI-Toolbox, a C++ software library that contains reinforcement learning and planning algorithms, and supports both single and multi agent problems, as well as partial observability. It is designed for simplicity and clarity, and contains extensive documentation of its API and code. It supports Python to enable users not comfortable with C++ to take advantage of the library's speed and functionality. AI-Toolbox is free software, and is hosted online at https://github.com/Svalorzen/AI-Toolbox.
Eugenio Bargiacchi, Diederik M. Roijers, Ann Nowé
J. Mach. Learn. Res.3
2020 Nonparametric user activity modelling and prediction
Yannick De Bock, Andres Auquilla, Ann Nowé, Joost R. Duflou
User Model. User Adapt. Interact.3
2019 Deep hybrid approach for 3D plane segmentation
Felipe Gomez Marulanda, Pieter Libin, Timothy Verstraeten, Ann Nowé
ESANN4
2019 A Multi-objective Reinforcement Learning Algorithm for JSSP
Beatriz M. Méndez-Hernández, Erick Rodríguez Bazan, Yailen Martínez-Jiménez, Pieter Libin, Ann Nowé
ICANN (1)5
2019 Dynamic Weights in Multi-Objective Deep Reinforcement Learning
abstract
Many real-world decision problems are characterized by multiple conflicting objectives which must be balanced based on their relative importance. In the dynamic weights setting the relative importance changes over time and specialized algorithms that deal with such change, such as a tabular Reinforcement Learning (RL) algorithm by Natarajan and Tadepalli (2005), are required. However, this earlier work is not feasible for RL settings that necessitate the use of function approximators. We generalize across weight changes and high-dimensional inputs by proposing a multi-objective Q-network whose outputs are conditioned on the relative importance of objectives and we introduce Diverse Experience Replay (DER) to counter the inherent non-stationarity of the Dynamic Weights setting. We perform an extensive experimental evaluation and compare our methods to adapted algorithms from Deep Multi-Task/Multi-Objective Reinforcement Learning and show that our proposed network in combination with DER dominates these adapted algorithms across weight change scenarios and problem domains.
Axel Abels, Diederik M. Roijers, Tom Lenaerts, Ann Nowé, Denis Steckelmacher
ICML4
2019 Per-Decision Option Discounting
abstract
In order to solve complex problems an agent must be able to reason over a sufficiently long horizon. Temporal abstraction, commonly modeled through options, offers the ability to reason at many timescales, but the horizon length is still determined by the discount factor of the underlying Markov Decision Process. We propose a modification to the options framework that naturally scales the agent’s horizon with option length. We show that the proposed option-step discount controls a bias-variance trade-off, with larger discounts (counter-intuitively) leading to less estimation variance.
Anna Harutyunyan, Peter Vrancx, Philippe Hamel, Ann Nowé, Doina Precup
ICML4
2019 Bayesian Anytime m-top Exploration
abstract
We introduce Boundary Focused Thompson sampling (BFTS), a new Bayesian algorithm to solve the anytime m-top exploration problem, where the objective is to identify the m best arms in a multi-armed bandit. First, we consider a set of existing benchmark problems that consider sub-Gaussian reward distributions (i.e., Gaussian with fixed variance and categorical reward). Next, we introduce a new environment inspired by a real world decision problem concerning insect control for organic agriculture. This new environment encodes a Poisson rewards distribution. For all these benchmarks, we experimentally show that BFTS consistently outperforms AT-LUCB, the current state of the art algorithm.
Pieter Libin, Timothy Verstraeten, Diederik M. Roijers, Kristof Theys, Ann Nowé
ICTAI6
2019 Synaptic Learning of Long-Term Cognitive Networks with Inputs
abstract
In contrast with the extense variety of machine learning algorithms, to fully automate the reasoning process, only a few can take advantage of the expert knowledge. Fuzzy Cognitive Maps (FCMs) are neural networks that can naturally integrate this kind of knowledge in the inference process. Nevertheless, FCMs have serious drawbacks difficult to overcome from the absence of an intrinsically learning algorithm or limited prediction horizon of the activation space of the neurons. Recently, some variants of the FCMs like Short-Term Cognitive Networks (STCN) and Long Term Cognitive Networks (LTCN) have been proposed to solve these problems. In this paper, we propose a new neural network model as a variant of LTCNs called Long-Term Cognitive Networks with Inputs (LTCNIs). A new kind of input neuron which is not present in the traditional FCMs approach or the derived algorithms STCNs and LTCNs is introduced, in order to model inputs like energy or mass in physical systems. The performance of the method is discussed through the modeling of a passive circuit problem. As a second contribution, a new flexible reasoning strategy, which preserves the expert knowledge through synaptic learning is presented. A synaptic learning based on a gradient descent method is implemented limited by a set of restrictions that preserves the model semantics.
Richar Sosa, Alejandro Alfonso, Gonzalo Nápoles, Rafael Bello 0001, Koen Vanhoof, Ann Nowé
IJCNN6
2019 Sample-Efficient Model-Free Reinforcement Learning with Off-Policy Critics
Denis Steckelmacher, Hélène Plisnier, Diederik M. Roijers, Ann Nowé
ECML/PKDD (3)4
2019 GUARDIAML: Machine Learning-Assisted Dynamic Information Flow Control
abstract
Developing JavaScript and web applications with confidentiality and integrity guarantees is challenging. Information flow control enables the enforcement of such guarantees. However, the integration of this technique into software tools used by developers in their workflow is missing. In this paper we present GUARDIAML, a machine learning-assisted dynamic information flow control tool for JavaScript web applications. GUARDIAML enables developers to detect unwanted information flow from sensitive sources to public sinks. It can handle the DOM and interaction with internal and external libraries and services. Because the specification of sources and sinks can be tedious, GUARDIAML assists in this process by suggesting the tagging of sources and sinks via a machine learning component.
Angel Luis Scull Pupo, Jens Nicolay, Kyriakos Efthymiadis, Ann Nowé, Coen De Roover, Elisa Gonzalez Boix
SANER4
2018 Learning With Options That Terminate Off-Policy
abstract
A temporally abstract action, or an option, is specified by a policy and a termination condition: the policy guides the option behavior, and the termination condition roughly determines its length. Generally, learning with longer options (like learning with multi-step returns) is known to be more efficient. However, if the option set for the task is not ideal, and cannot express the primitive optimal policy well, shorter options offer more flexibility and can yield a better solution. Thus, the termination condition puts learning efficiency at odds with solution quality. We propose to resolve this dilemma by decoupling the behavior and target terminations, just like it is done with policies in off-policy learning. To this end, we give a new algorithm, Q(beta), that learns the solution with respect to any termination condition, regardless of how the options actually terminate. We derive Q(beta) by casting learning with options into a common framework with well-studied multi-step off policy learning. We validate our algorithm empirically, and show that it holds up to its motivating claims.
Anna Harutyunyan, Peter Vrancx, Pierre-Luc Bacon, Doina Precup, Ann Nowé
AAAI5
2018 Adapting to Concept Drift in Credit Card Transaction Data Streams Using Contextual Bandits and Decision Trees
abstract
Credit card transactions predicted to be fraudulent by automated detection systems are typically handed over to human experts for verification. To limit costs, it is standard practice to select only the most suspicious transactions for investigation. We claim that a trade-off between exploration and exploitation is imperative to enable adaptation to changes in behavior (concept drift). Exploration consists of the selection and investigation of transactions with the purpose of improving predictive models, and exploitation consists of investigating transactions detected to be suspicious. Modeling the detection of fraudulent transactions as rewarding, we use an incremental Regression Tree learner to create clusters of transactions with similar expected rewards. This enables the use of a Contextual Multi-Armed Bandit (CMAB) algorithm to provide the exploration/exploitation trade-off. We introduce a novel variant of a CMAB algorithm that makes use of the structure of this tree, and use Semi-Supervised Learning to grow the tree using unlabeled data. The approach is evaluated on a real dataset and data generated by a simulator that adds concept drift by adapting the behavior of fraudsters to avoid detection. It outperforms frequently used offline models in terms of cumulative rewards, in particular in the presence of concept drift.
Dennis J. N. J. Soemers, Tim Brys, Kurt Driessens, Mark H. M. Winands, Ann Nowé
AAAI5
2018 Reinforcement Learning in POMDPs With Memoryless Options and Option-Observation Initiation Sets
abstract
Many real-world reinforcement learning problems have a hierarchical nature, and often exhibit some degree of partial observability. While hierarchy and partial observability are usually tackled separately (for instance by combining recurrent neural networks and options), we show that addressing both problems simultaneously is simpler and more efficient in many cases. More specifically, we make the initiation set of options conditional on the previously-executed option, and show that options with such Option-Observation Initiation Sets (OOIs) are at least as expressive as Finite State Controllers (FSCs), a state-of-the-art approach for learning in POMDPs. OOIs are easy to design based on an intuitive description of the task, lead to explainable policies and keep the top-level and option policies memoryless. Our experiments show that OOIs allow agents to learn optimal policies in challenging POMDPs, while being much more sample-efficient than a recurrent neural network over options.
Denis Steckelmacher, Diederik M. Roijers, Anna Harutyunyan, Peter Vrancx, Hélène Plisnier, Ann Nowé
AAAI6
2018 Learning to Coordinate with Coordination Graphs in Repeated Single-Stage Multi-Agent Decision Problems
abstract
Learning to coordinate between multiple agents is an important problem in many reinforcement learning problems. Key to learning to coordinate is exploiting loose couplings, i.e., conditional independences between agents. In this paper we study learning in repeated fully cooperative games, multi-agent multi-armed bandits (MAMABs), in which the expected rewards can be expressed as a coordination graph. We propose multi-agent upper confidence exploration (MAUCE), a new algorithm for MAMABs that exploits loose couplings, which enables us to prove a regret bound that is logarithmic in the number of arm pulls and only linear in the number of agents. We empirically compare MAUCE to sparse cooperative Q-learning, and a state-of-the-art combinatorial bandit approach, and show that it performs much better on a variety of settings, including learning control policies for wind farms.
Eugenio Bargiacchi, Timothy Verstraeten, Diederik M. Roijers, Ann Nowé, Hado van Hasselt
ICML4
2018 IPC-Net: 3D Point-Cloud Segmentation Using Deep Inter-Point Convolutional Layers
abstract
Over the last decade, the demand for better segmentation and classification algorithms in 3D spaces has significantly grown due to the popularity of new 3D sensor technologies and advancements in the field of robotics. Point-clouds are one of the most popular representations to store a digital description of 3D shapes. However, point-clouds are stored in irregular and unordered structures, which limits the direct use of segmentation algorithms such as Convolutional Neural Networks. The objective of our work is twofold: First, we aim to provide a full analysis of the PointNet architecture to illustrate which features are being extracted from the point-clouds. Second, to propose a new network architecture called IPC-Net to improve the state-of-the-art point cloud architectures. We show that IPC-Net extracts a larger set of unique features allowing the model to produce more accurate segmentations compared to the PointNet architecture. In general, our approach outperforms PointNet on every family of 3D geometries on which the models were tested. A high generalisation improvement was observed on every 3D shape, especially on the rockets dataset. Our experiments demonstrate that our main contribution, inter-point activation on the network's layers, is essential to accurately segment 3D point-clouds.
Felipe Gomez Marulanda, Pieter Libin, Timothy Verstraeten, Ann Nowé
ICTAI4
2018 Bayesian Best-Arm Identification for Selecting Influenza Mitigation Strategies
Pieter Libin, Timothy Verstraeten, Diederik M. Roijers, Jelena Grujic, Kristof Theys, Philippe Lemey, Ann Nowé
ECML/PKDD (3)7
2018 Modelling incomplete information in Boolean games using possibilistic logic
Sofie De Clercq, Steven Schockaert, Ann Nowé, Martine De Cock
Int. J. Approx. Reason.3
2017 Coordinating Human and Agent Behavior in Collective-Risk Scenarios
abstract
Various social situations entail a collective risk. A well-known example is climate change, wherein the risk of a future environmental disaster clashes with the immediate economic interest of developed and developing countries. The collective-risk game operationalizes this kind of situations. The decision process of the participants is determined by how good they are in evaluating the probability of future risk as well as their ability to anticipate the actions of the opponents. Anticipatory behavior contrasts with the reactive theories often used to analyze social dilemmas. Our initial work can already show that anticipative agents are a better model to human behavior than reactive ones. All the agents we studied used a recurrent neural network, however, only the ones that used it to predict future outcomes (anticipative agents) were able to account for changes in the context of games, a behavior also observed in experiments with humans. This extended abstract aims to explain how we wish to investigate anticipation within the context of the collective-risk game and the relevance these results may have for the field of hybrid socio-technical systems.
Elias Fernández Domingos, Juan C. Burguillo, Ann Nowé, Tom Lenaerts
AAAI3
2017 Exact and heuristic methods for solving Boolean games
Sofie De Clercq, Kim Bauters, Steven Schockaert, Mihail Mihaylov, Ann Nowé, Martine De Cock
Auton. Agents Multi Agent Syst.5
2017 PhyloGeoTool: interactively exploring large phylogenies in an epidemiological context
abstract
MOTIVATION: Clinicians, health officials and researchers are interested in the epidemic spread of pathogens in both space and time to support the optimization of intervention measures and public health policies. Large sequence databases of virus sequences provide an interesting opportunity to study this spread through phylogenetic analysis. To infer knowledge from large phylogenetic trees, potentially encompassing tens of thousands of virus strains, an efficient method for data exploration is required. The clades that are visited during this exploration should be annotated with strain characteristics (e.g. transmission risk group, tropism, drug resistance profile) and their geographic context. RESULTS: PhyloGeoTool implements a visual method to explore large phylogenetic trees and to depict characteristics of strains and clades, including their geographic context, in an interactive way. PhyloGeoTool also provides the possibility to position new virus strains relative to the existing phylogenetic tree, allowing users to gain insight in the placement of such new strains without the need to perform a de novo reconstruction of the phylogeny. AVAILABILITY AND IMPLEMENTATION: https://github.com/rega-cev/phylogeotool (Freely available: open source software project). CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Pieter Libin, Ewout Vanden Eynden, Francesca Incardona, Ann Nowé, Antonia Bezenchek, Anders Sönnerborg, Anne-Mieke Vandamme, Kristof Theys, Guy Baele
Bioinform.4
2017 Discovering knowledge from data clustering using automatically-defined interval type-2 fuzzy predicates
Diego S. Comas, Gustavo J. Meschino, Ann Nowé, Virginia Laura Ballarin
Expert Syst. Appl.3
2017 Multi-objectivization and ensembles of shapings in reinforcement learning
Tim Brys, Anna Harutyunyan, Peter Vrancx, Ann Nowé, Matthew E. Taylor
Neurocomputing4
2016 Case study: An analysis of accidental complexity in a state-of-the-art hyper-heuristic for HyFlex
abstract
While simplicity is an important factor affecting algorithm re-usability, it is often overlooked in algorithm design, which has a tendency to produce overly complex methods. In this paper we demonstrate Accidental Complexity Analysis (ACA), a research practice targeted at detecting and eliminating accidental complexity, without loss of performance (c.f. refactoring in software engineering), using it to analyze the presence of accidental complexity in GIHH, a state-of-the-art selection hyper-heuristic for HyFlex. We identify various algorithmic sub-mechanisms contributing little to GIHH's overall performance, and validate many other. As an outcome we present Lean-GIHH, a simplified, re-implementation of GIHH.
Steven Adriaensen, Ann Nowé
CEC2
2016 Formalizing Commitment-Based Deals in Boolean Games
abstract
Boolean games (BGs) are a strategic framework in which agents' goals are described using propositional logic. Despite the popularity of BGs, the problem of how agents can coordinate with others to (at least partially) achieve their goals has hardly received any attention. However, negotiation protocols that have been developed outside the setting of BGs can be adopted for this purpose, provided that we can formalize (i) how agents can make commitments and (ii) how deals between coalitions of agents can be identified given a set of active commitments. In this paper, we focus on these two aims. First, we show how agents can formulate commitments that are in accordance with their goals, and what it means for the commitments of an agent to be consistent. Second, we formalize deals in terms of coalitions who can achieve their goals without help from others. We show that verifying the consistency of a set of commitments of one agent is ΠP2-complete while checking the existence of a deal in a set of mutual commitments is Σp2
Sofie De Clercq, Steven Schockaert, Ann Nowé, Martine De Cock
ECAI3
2016 Measuring Diversity of Socio-Cognitively Inspired ACO Search
Ewelina Swiderska, Jakub Lasisz, Aleksander Byrski, Tom Lenaerts, Dana Samson, Bipin Indurkhya, Ann Nowé, Marek Kisiel-Dorohinicki
EvoApplications (1)7
2016 Towards a White Box Approach to Automated Algorithm Design
Steven Adriaensen, Ann Nowé
IJCAI2
2016 Combining Occupancy User Profiles in a Multi-user Environment: An Academic Office Case Study
abstract
In a worldwide context, space heating is the largest energy consumer in commercial buildings, it accounts for 35% of the total energy consumed in the US. Energy efficient thermostats, that learn occupancy patterns and user preferences, haven been studied in literature. However, they are oriented to single-user environments, therefore, they are not applicable in offices where several users interact, i.e. multi-user environments. To expand the single-user techniques in order to cope with multi-user environments, two methods are proposed to derive the user's expected temperatures demands based on their occupancy profiles and individual preferences in terms of desired temperature and tolerance. This paper presents the implications of the implementation of such techniques by means of a case study of two users in an academic office. We observed that the proposed methods reduced the operational time up to 33% compared to a reference fixed schedule of 12 hours while maintaining user comfort. In conclusion, smart thermostats can also reduce energy consumption in multi-user environments while guaranteeing individual user expectations.
Andres Auquilla, Yannick De Bock, Ann Nowé, Joost R. Duflou
Intelligent Environments3
2016 An adaptive rule-based classifier for mining big biological data
Dewan Md. Farid 0001, M. Abdulla Al Mamun, Bernard Manderick, Ann Nowé
Expert Syst. Appl.4
2016 Solving stable matching problems using answer set programming
abstract
Abstract Since the introduction of the stable marriage problem (SMP) by Gale and Shapley (1962), several variants and extensions have been investigated. While this variety is useful to widen the application potential, each variant requires a new algorithm for finding the stable matchings. To address this issue, we propose an encoding of the SMP using answer set programming (ASP), which can straightforwardly be adapted and extended to suit the needs of specific applications. The use of ASP also means that we can take advantage of highly efficient off-the-shelf solvers. To illustrate the flexibility of our approach, we show how our ASP encoding naturally allows us to select optimal stable matchings, i.e. matchings that are optimal according to some user-specified criterion. To the best of our knowledge, our encoding offers the first exact implementation to find sex-equal, minimum regret, egalitarian or maximum cardinality stable matchings for SMP instances in which individuals may designate unacceptable partners and ties between preferences are allowed.
Sofie De Clercq, Steven Schockaert, Martine De Cock, Ann Nowé
Theory Pract. Log. Program.4
2015 Expressing Arbitrary Reward Functions as Potential-Based Advice
abstract
Effectively incorporating external advice is an important problem in reinforcement learning, especially as it moves into the real world. Potential-based reward shaping is a way to provide the agent with a specific form of additional reward, with the guarantee of policy invariance. In this work we give a novel way to incorporate an arbitrary reward function with the same guarantee, by implicitly translating it into the specific form of dynamic advice potentials, which are maintained as an auxiliary value function learnt at the same time. We show that advice provided in this way captures the input reward function in expectation, and demonstrate its efficacy empirically.
Anna Harutyunyan, Sam Devlin, Peter Vrancx, Ann Nowé
AAAI4
2015 A benchmark set extension and comparative study for the HyFlex framework
abstract
In this work we conduct a comparative study of several publicly available, state-of-the-art hyper-heuristics for HyFlex in order to assess their generality across domains. To this purpose we extend the HyFlex benchmark set with 3 new problem domains: The 0-1 Knap Sack, Quadratic Assignment and Max-Cut Problem. To our knowledge, this is the first public extension of the benchmark since the CHeSC 2011 competition. In addition, this is the first study testing the Fair-Share Iterated Local Search (FS-ILS) method, designed in prior research, using a semi-automated design approach, on new unseen problem domains. We show that, of the methods compared, Adap-HH (CHeSC 2011 winner) clearly perfoms the most consistently, overall. In addition, we identify a weakness of, as well as a way to further simplify the FS-ILS method. Finally, we found that, overall, the state-of-the-art methods compared, generalized much better than a naive baseline.
Steven Adriaensen, Gabriela Ochoa, Ann Nowé
CEC3
2015 Risk-sensitivity through multi-objective reinforcement learning
abstract
Usually in reinforcement learning, the goal of the agent is to maximize the expected return. However, in practical applications, algorithms that solely focus on maximizing the mean return could be inappropriate as they do not account for the variability of their solutions. Thereby, a variability measure could be included to accommodate for a risk-sensitive setting, i.e. where the system engineer can explicitly define the tolerated level of variance. Our approach is based on multi-objectivization where a standard single-objective environment is extended with one (or more) additional objectives. More precisely, we augment the standard feedback signal of an environment with an additional objective that defines the variance of the solution. We highlight that our algorithm, named risk-sensitive Pareto Q-learning, is (1) specifically tailored to learn a set of Pareto non-dominated policies that trade-off these two objectives. Additionally (2), the algorithm can also retrieve every policy that has been learned throughout the state-action space. This in contrast to standard risk-sensitive approaches where only a single trade-off between mean and variance is learned at a time.
Kristof Van Moffaert, Tim Brys, Ann Nowé
CEC3
2015 Online Distributed Voltage Control of an offshore MTdc network using reinforcement learning
abstract
This paper addresses one of the main challenges on the way to an offshore transnational multi-terminal dc (MTdc) network: its control and operation. The main objective is to demonstrate the feasibility of using reinforcement learning (RL) techniques to control, in real time, a multi-terminal dc network aimed at integrating offshore wind farms (OWFs). This method of controlling MTdc networks using RL techniques is called Online Distributed Voltage Control (ODVC). The ODVC strategy uses Continuous Action Reinforcement Learning Automata (CARLA) to optimize power flows in real time. To validate the effectiveness of the proposed control method, dynamic simulations are carried out using a MTdc grid model composed of six nodes, interconnecting three offshore wind farms to three European countries. The results obtained demonstrate the advantages of implementing an online distributed voltage control strategy to obtain feasible controlled power flows with low transmission losses. The results obtained demonstrate the feasibility of the proposed method to control, in real time, MTdc networks and that the RL techniques are well-suited for this problem due to their inherent advantages of coping with stochastic environments.
S. Rodrigues, Rodrigo Teixeira Pinto, Pavol Bauer, Tim Brys, Ann Nowé
CEC5
2015 Reinforcement Learning from Demonstration through Shaping
Tim Brys, Anna Harutyunyan, Halit Bener Suay, Sonia Chernova, Matthew E. Taylor, Ann Nowé
IJCAI6
2015 Multilateral Negotiation in Boolean Games with Incomplete Information Using Generalized Possibilistic Logic
Sofie De Clercq, Steven Schockaert, Ann Nowé, Martine De Cock
IJCAI3
2015 Schedule-based multi-channel communication in wireless sensor networks: A complete design and performance evaluation
Kieu-Ha Phung, Bart Lemmens, Marnix Goossens, Ann Nowé, Lan Tran, Kris Steenhaut
Ad Hoc Networks4
2015 A Reinforcement Learning Approach for Interdomain Routing with Link Prices
abstract
In today’s Internet, the commercial aspects of routing are gaining importance. Current technology allows Internet Service Providers (ISPs) to renegotiate contracts online to maximize profits. Changing link prices will influence interdomain routing policies that are now driven by monetary aspects as well as global resource and performance optimization. In this article, we consider an interdomain routing game in which the ISP’s action is to set the price for its transit links. Assuming a cheapest path routing scheme, the optimal action is the price setting that yields the highest utility (i.e., profit) and depends both on the network load and the actions of other ISPs. We adapt a continuous and a discrete action learning automaton (LA) to operate in this framework as a tool that can be used by ISP operators to learn optimal price setting. In our model, agents representing different ISPs learn only on the basis of local information and do not need any central coordination or sensitive information exchange. Simulation results show that a single ISP employing LAs is able to learn the optimal price in a stationary environment. By introducing a selective exploration rule, LAs are also able to operate in nonstationary environments. When two ISPs employ LAs, we show that they converge to stable and fair equilibrium strategies.
Peter Vrancx, Pasquale Gurzi, Abdel Rodríguez, Kris Steenhaut, Ann Nowé
ACM Trans. Auton. Adapt. Syst.5
2014 Reinforcement Learning on Multiple Correlated Signals
abstract
This extended abstract provides a brief overview of my PhD research on multi-objectivization and ensemble techniques in reinforcement learning.
Tim Brys, Ann Nowé
AAAI2
2014 Combining Multiple Correlated Reward and Shaping Signals by Measuring Confidence
abstract
Multi-objective problems with correlated objectives are a class of problems that deserve specific attention. In contrast to typical multi-objective problems, they do not require the identification of trade-offs between the objectives, as (near-) optimal solutions for any objective are (near-) optimal for every objective. Intelligently combining the feedback from these objectives, instead of only looking at a single one, can improve optimization. This class of problems is very relevant in reinforcement learning, as any single-objective reinforcement learning problem can be framed as such a multi-objective problem using multiple reward shaping functions. After discussing this problem class, we propose a solution technique for such reinforcement learning problems, called adaptive objective selection. This technique makes a temporal difference learner estimate the Q-function for each objective in parallel, and introduces a way of measuring confidence in these estimates. This confidence metric is then used to choose which objective's estimates to use for action selection. We show significant improvements in performance over other plausible techniques on two problem domains. Finally, we provide an intuitive analysis of the technique's decisions, yielding insights into the nature of the problems being solved.
Tim Brys, Ann Nowé, Daniel Kudenko, Matthew E. Taylor
AAAI2
2014 Pareto Upper Confidence Bounds algorithms: An empirical study
abstract
Many real-world stochastic environments are inherently multi-objective environments with conflicting objectives. The multi-objective multi-armed bandits (MOMAB) are extensions of the classical, i.e. single objective, multi-armed bandits to reward vectors and multi-objective optimisation techniques are often required to design mechanisms with an efficient exploration / exploitation trade-off. In this paper, we propose the improved Pareto Upper Confidence Bound (iPUCB) algorithm that straightforwardly extends the single objective improved UCB algorithm to reward vectors by deleting the suboptimal arms. The goal of the improved Pareto UCB algorithm, i.e. iPUCB, is to identify the set of best arms, or the Pareto front, in a fixed budget of arm pulls. We experimentally compare the performance of the proposed Pareto upper confidence bound algorithm with the Pareto UCB1 algorithm and the Hoeffding race on a bi-objective example coming from an industrial control applications, i.e. the engagement of wet clutches. We propose a new regret metric based on the Kullback-Leibler divergence to measure the performance of a multi-objective multi-armed bandit algorithm. We show that iPUCB outperforms the other two tested algorithms on the given multi-objective environment.
Madalina M. Drugan, Ann Nowé, Bernard Manderick
ADPRL2
2014 Designing reusable metaheuristic methods: A semi-automated approach
abstract
Many interesting optimization problems cannot be solved efficiently. Recently, a lot of work has been done on meta-heuristic optimization methods that quickly find approximate solutions to otherwise intractable problems. While successful, the field suffers from a notable lack of reuse of methods, both in practical applications as in research. In this paper, we describe a semi-automated approach to design more re-usable methods, based on key principles of re-usability such as simplicity, modularity and generality. We illustrate this methodology by designing general metaheuristics (using hyperheuristics) and show that the methods obtained are competitive with the contestants of the Cross-Domain Heuristic Search Competition (2011). In particular, we find a method performing better than the competition's winner, which can be considered the state-of-the-art in domain-independent metaheuristic search.
Steven Adriaensen, Tim Brys, Ann Nowé
IEEE Congress on Evolutionary Computation3
2014 Using Ensemble Techniques and Multi-Objectivization to Solve Reinforcement Learning Problems
abstract
Recent work on multi-objectivization has shown how a single-objective reinforcement learning problem can be turned into a multi-objective problem with correlated objectives, by providing multiple reward shaping functions. The information contained in these correlated objectives can be exploited to solve the base, single-objective problem faster and better, given techniques specifically aimed at handling such correlated objectives. In this paper, we identify ensemble techniques as a set of methods that is suitable to solve multi-objectivized reinforcement learning problems. We empirically demonstrate their use on the Pursuit domain.
Tim Brys, Matthew E. Taylor, Ann Nowé
ECAI3
2014 Off-Policy Shaping Ensembles in Reinforcement Learning
abstract
In this work we propose learning an ensemble of policies related through potential-based shaping rewards via the off-policy Horde framework.
Anna Harutyunyan, Tim Brys, Peter Vrancx, Ann Nowé
ECAI4
2014 Fair-share ILS: a simple state-of-the-art iterated local search hyperheuristic
abstract
In this work we present a simple state-of-the-art selection hyperheuristic called Fair-Share Iterated Local Search (FS-ILS). FS-ILS is an iterated local search method using a conservative restart condition. Each iteration, a perturbation heuristic is selected proportionally to the acceptance rate of its previously proposed candidate solutions (after iterative improvement) by a domain-independent variant of the Metropolis condition. FS-ILS was developed in prior work using a semi-automated design approach. That work focused on how the method was found, rather than the method itself. As a result, it lacked a detailed explanation and analysis of the method, which will be the main contribution of this work. In our experiments we analyze FS-ILS's parameter sensitivity, accidental complexity and compare it to the contestants of the CHeSC (2011) competition.
Steven Adriaensen, Tim Brys, Ann Nowé
GECCO3
2014 Decentralized Computation of Pareto Optimal Pure Nash Equilibria of Boolean Games with Privacy Concerns
abstract
In Boolean games, agents try to reach a goal formulated as a Boolean formula. These games are attractive because of their compact representations. However, few methods are available to compute the solutions and they are either limited or do not take privacy or communication concerns into account. In this paper we propose the use of an algorithm related to reinforcement learning to address this problem. Our method is decentralized in the sense that agents try to achieve their goals without knowledge of the other agents’ goals. We prove that this is a sound method to compute a Pareto optimal pure Nash equilibrium for an interesting class of Boolean games. Experimental results are used to investigate the performance of the algorithm.
Sofie De Clercq, Kim Bauters, Steven Schockaert, Mihail Mihaylov, Martine De Cock, Ann Nowé
ICAART (2)6
2014 Multi-objectivization of reinforcement learning problems by reward shaping
abstract
Multi-objectivization is the process of transforming a single objective problem into a multi-objective problem. Research in evolutionary optimization has demonstrated that the addition of objectives that are correlated with the original objective can make the resulting problem easier to solve compared to the original single-objective problem. In this paper we investigate the multi-objectivization of reinforcement learning problems. We propose a novel method for the multi-objectivization of Markov Decision problems through the use of multiple reward shaping functions. Reward shaping is a technique to speed up reinforcement learning by including additional heuristic knowledge in the reward signal. The resulting composite reward signal is expected to be more informative during learning, leading the learner to identify good actions more quickly. Good reward shaping functions are by definition correlated with the target value function for the base reward signal, and we show in this paper that adding several correlated signals can help to solve the basic single objective problem faster and better. We prove that the total ordering of solutions, and by consequence the optimality of solutions, is preserved in this process, and empirically demonstrate the usefulness of this approach on two reinforcement learning tasks: a pathfinding problem and the Mario domain.
Tim Brys, Anna Harutyunyan, Peter Vrancx, Matthew E. Taylor, Daniel Kudenko, Ann Nowé
IJCNN6
2014 Scalarization based Pareto optimal set of arms identification algorithms
abstract
Multi-objective multi-armed bandits (MOMAB) is an extension of the multi-objective multi-armed bandits framework that considers reward vectors instead of scalar reward values. Scalarization functions transform the reward vectors into reward values in order to use the standard multi-armed bandits (MAB) algorithms. However for many applications it is not obvious to come up with a good scalarization set and therefore there is needed to develop MAB that discover the whole Pareto set of arms. Our approach to this multi-objective MAB problem is two folded: i) identify the set of Pareto optimal arms and ii) identify the minimum subset of scalarization functions that optimize the set of Pareto optimal arms. We experimentally compare the proposed MOMAB algorithms on a multi-objective Bernoulli problem.
Madalina M. Drugan, Ann Nowé
IJCNN2
2014 A novel adaptive weight selection algorithm for multi-objective multi-agent reinforcement learning
abstract
To solve multi-objective problems, multiple reward signals are often scalarized into a single value and further processed using established single-objective problem solving techniques. While the field of multi-objective optimization has made many advances in applying scalarization techniques to obtain good solution trade-offs, the utility of applying these techniques in the multi-objective multi-agent learning domain has not yet been thoroughly investigated. Agents learn the value of their decisions by linearly scalarizing their reward signals at the local level, while acceptable system wide behaviour results. However, the non-linear relationship between weighting parameters of the scalarization function and the learned policy makes the discovery of system wide trade-offs time consuming. Our first contribution is a thorough analysis of well known scalarization schemes within the multi-objective multi-agent reinforcement learning setup. The analysed approaches intelligently explore the weight-space in order to find a wider range of system trade-offs. In our second contribution, we propose a novel adaptive weight algorithm which interacts with the underlying local multi-objective solvers and allows for a better coverage of the Pareto front. Our third contribution is the experimental validation of our approach by learning bi-objective policies in self-organising smart camera networks. We note that our algorithm (i) explores the objective space faster on many problem instances, (ii) obtained solutions that exhibit a larger hypervolume, while (iii) acquiring a greater spread in the objective space.
Kristof Van Moffaert, Tim Brys, Arjun Chandra, Lukas Esterle, Peter R. Lewis 0001, Ann Nowé
IJCNN6
2014 Multi-objective χ-Armed bandits
abstract
Many of the standard optimization algorithms focus on optimizing a single, scalar feedback signal. However, real-life optimization problems often require a simultaneous optimization of more than one objective. In this paper, we propose a multi-objective extension to the standard χ-armed bandit problem. As the feedback signal is now vector-valued, the goal of the agent is to sample actions in the Pareto dominating area of the objective space. Therefore, we propose the multi-objective Hierarchical Optimistic Optimization strategy that discretizes the continuous action space in relation to the Pareto optimal solutions obtained in the multi-objective objective space. We experimentally validate the approach on two well-known multi-objective test functions and a simulation of a real life application, the filling phase of a wet clutch. We demonstrate that the strategy allows to identify the Pareto front after just a few epochs and to sample accordingly. After learning, several multi-objective quality indicators indicate that the set of sampled solutions by the algorithm very closely approximates the Pareto front.
Kristof Van Moffaert, Kevin Van Vaerenbergh, Peter Vrancx, Ann Nowé
IJCNN4
2014 Possibilistic Boolean Games: Strategic Reasoning under Incomplete Information
Sofie De Clercq, Steven Schockaert, Martine De Cock, Ann Nowé
JELIA4
2014 Using Answer Set Programming for Solving Boolean Games
Sofie De Clercq, Kim Bauters, Steven Schockaert, Martine De Cock, Ann Nowé
KR5
2014 A decentralized approach for convention emergence in multi-agent systems
Mihail Mihaylov, Karl Tuyls, Ann Nowé
Auton. Agents Multi Agent Syst.3
2014 Multi-objective reinforcement learning using sets of pareto dominating policies
Kristof Van Moffaert, Ann Nowé
J. Mach. Learn. Res.2
2013 Scalarized multi-objective reinforcement learning: Novel design techniques
abstract
In multi-objective problems, it is key to find compromising solutions that balance different objectives. The linear scalarization function is often utilized to translate the multi-objective nature of a problem into a standard, single-objective problem. Generally, it is noted that such as linear combination can only find solutions in convex areas of the Pareto front, therefore making the method inapplicable in situations where the shape of the front is not known beforehand, as is often the case. We propose a non-linear scalarization function, called the Chebyshev scalarization function, as a basis for action selection strategies in multi-objective reinforcement learning. The Chebyshev scalarization method overcomes the flaws of the linear scalarization function as it can (i) discover Pareto optimal solutions regardless of the shape of the front, i.e. convex as well as non-convex , (ii) obtain a better spread amongst the set of Pareto optimal solutions and (iii) is not particularly dependent on the actual weights used.
Kristof Van Moffaert, Madalina M. Drugan, Ann Nowé
ADPRL3
2013 Meta-Evolutionary Algorithms and recombination operators for satisfiability solving in fuzzy logics
abstract
In this work, we develop a new paradigm, called Meta-Evolutionary Algorithms, motivated by the challenging, continuous problems encountered in the domain of satisfiability in fuzzy logics (SAT∞). In Meta-Evolutionary Algorithms, the individuals in a population are optimization algorithms themselves. Mutation at the meta-population level is handled by performing an optimization step in each optimization algorithm, and recombination at the meta-population level is handled by exchanging information between different algorithms. We analyse different recombination operators and empirically show that simple Meta-Evolutionary Algorithms are able to outperform CMA-ES on a set of SAT∞benchmark problems.
Tim Brys, Madalina M. Drugan, Ann Nowé
IEEE Congress on Evolutionary Computation3
2013 Hypervolume-Based Multi-Objective Reinforcement Learning
Kristof Van Moffaert, Madalina M. Drugan, Ann Nowé
EMO3
2013 Solving satisfiability in fuzzy logics by mixing CMA-ES
abstract
Satisfiability in propositional logic is well researched and many approaches to checking and solving exist. In infinite-valued or fuzzy logics, however, there have only recently been attempts at developing methods for solving satisfiability. In this paper, we propose new benchmark problems and analyse the function landscape of different problem classes, focussing our analysis on plateaus. Based on this study, we develop Mixing CMA-ES (M-CMA-ES), an extension to CMA-ES that is well suited to solving problems with many large plateaus. We empirically show the relation between certain function landscape properties and M-CMA-ES performance.
Tim Brys, Madalina M. Drugan, Peter A. N. Bosman, Martine De Cock, Ann Nowé
GECCO5
2013 Reinforcement Learning for Multi-purpose Schedules
Kristof Van Moffaert, Yann-Michaël De Hauwere, Peter Vrancx, Ann Nowé
ICAART (2)4
2013 On the Behaviour of Scalarization Methods for the Engagement of a Wet Clutch
abstract
Many industrial problems are inherently multi-objective, and require special attention to find different trade-off solutions. Typical multi-objective approaches calculate a scalarization of the different objectives and subsequently optimize the problem using a single-objective optimization method. Several scalarization techniques are known in the literature, and each has its own advantages and drawbacks. In this paper, we explore various of these scalarization techniques in the context of an industrial application, namely the engagement of a wet clutch using reinforcement learning. We analyse the approximate Pareto front obtainable by each technique, and discuss the causes of the differences observed. Finally, we show how a simple search algorithm can help explore the parameter space of the scalarization techniques, to efficiently identify possible trade-off solutions.
Tim Brys, Kristof Van Moffaert, Kevin Van Vaerenbergh, Ann Nowé
ICMLA (1)4
2013 Designing multi-objective multi-armed bandits algorithms: A study
abstract
We propose an algorithmic framework for multi-objective multi-armed bandits with multiple rewards. Different partial order relationships from multi-objective optimization can be considered for a set of reward vectors, such as scalarization functions and Pareto search. A scalarization function transforms the multi-objective environment into a single objective environment and are a popular choice in multi-objective reinforcement learning. Scalarization techniques can be straightforwardly implemented into the current multi-armed bandit framework, but the efficiency of these algorithms depends very much on their type, linear or non-linear (e.g. Chebyshev), and their parameters. Using Pareto dominance order relationship allows to explore the multi-objective environment directly, however this can result in large sets of Pareto optimal solutions. In this paper we propose and evaluate the performance of multi-objective MABs using three regret metric criteria. The standard UCB1 is extended to scalarized multi-objective UCB1 and we propose a Pareto UCB1 algorithm. Both algorithms are proven to have a logarithmic upper bound for their expected regret. We also introduce a variant of the scalarized multi-objective UCB1 that removes online inefficient scalarizations in order to improve the algorithm's efficiency. These algorithms are experimentally compared on multi-objective Bernoulli distributions, Pareto UCB1 being the algorithm with the best empirical performance.
Madalina M. Drugan, Ann Nowé
IJCNN2
2013 Networks as a tool to save energy while keeping up general user comfort in buildings
abstract
Devices, like Smartphones, tablets, notebooks, desktop computers, smart televisions are all connected to our home network. Analyzing the traffic on a network has already been widely studied in order to be able to detect anomalies in the network such as broken hardware, bad routing schemes or intrusions. We claim that network traffic analysis however can also bring forward valuable information about the users of that network. In this paper we demonstrate with a practical applications how we can save energy by analysing the usage on the network, as network activity informs us about the activity of the user.
Yann-Michaël De Hauwere, Kristof Van Moffaert, Paul-Armand Verhaegen, Ann Nowé
LANMAN4
2013 Batch effect removal methods for microarray gene expression data integration: a survey
abstract
Genomic data integration is a key goal to be achieved towards large-scale genomic data analysis. This process is very challenging due to the diverse sources of information resulting from genomics experiments. In this work, we review methods designed to combine genomic data recorded from microarray gene expression (MAGE) experiments. It has been acknowledged that the main source of variation between different MAGE datasets is due to the so-called 'batch effects'. The methods reviewed here perform data integration by removing (or more precisely attempting to remove) the unwanted variation associated with batch effects. They are presented in a unified framework together with a wide range of evaluation tools, which are mandatory in assessing the efficiency and the quality of the data integration process. We provide a systematic description of the MAGE data integration methodology together with some basic recommendation to help the users in choosing the appropriate tools to integrate MAGE data for large-scale analysis; and also how to evaluate them from different perspectives in order to quantify their efficiency. All genomic data used in this study for illustration purposes were retrieved from InSilicoDB http://insilico.ulb.ac.be.
Cosmin Lazar, Stijn Meganck, Jonatan Taminau, David Steenhoff, Alain Coletta, Colin Molter, David Y. Weiss Solís, Robin Duque, Hugues Bersini, Ann Nowé
Briefings Bioinform.10
2013 GENESHIFT: A Nonparametric Approach for Integrating Microarray Gene Expression Data Based on the Inner Product as a Distance Measure between the Distributions of Genes
abstract
The potential of microarray gene expression (MAGE) data is only partially explored due to the limited number of samples in individual studies. This limitation can be surmounted by merging or integrating data sets originating from independent MAGE experiments, which are designed to study the same biological problem. However, this process is hindered by batch effects that are study-dependent and result in random data distortion; therefore numerical transformations are needed to render the integration of different data sets accurate and meaningful. Our contribution in this paper is two-fold. First we propose GENESHIFT, a new nonparametric batch effect removal method based on two key elements from statistics: empirical density estimation and the inner product as a distance measure between two probability density functions; second we introduce a new validation index of batch effect removal methods based on the observation that samples from two independent studies drawn from a same population should exhibit similar probability density functions. We evaluated and compared the GENESHIFT method with four other state-of-the-art methods for batch effect removal: Batch-mean centering, empirical Bayes or COMBAT, distance-weighted discrimination, and cross-platform normalization. Several validation indices providing complementary information about the efficiency of batch effect removal methods have been employed in our validation framework. The results show that none of the methods clearly outperforms the others. More than that, most of the methods used for comparison perform very well with respect to some validation indices while performing very poor with respect to others. GENESHIFT exhibits robust performances and its average rank is the highest among the average ranks of all methods used for comparison.
Cosmin Lazar, Jonatan Taminau, Stijn Meganck, David Steenhoff, Alain Coletta, David Y. Weiss Solís, Colin Molter, Robin Duque, Hugues Bersini, Ann Nowé
IEEE ACM Trans. Comput. Biol. Bioinform.10
2012 Improving Convergence of CMA-ES Through Structure-Driven Discrete Recombination
abstract
Evolutionary Strategies (ES) are a class of continuous optimization algorithms that have proven to perform very well on hard optimization problems. Whereas in earlier literature, both intermediate and discrete recombination operators were used, we now see that most ES, e.g. CMA-ES, use only intermediate recombination. While CMA-ES is considered state-of-the-art in continuous optimization, we believe that reintroducing discrete recombination can improve the algorithms' ability to escape local optima. Specifically, we look at using information on the problem's structure to create building blocks for recombination.
Tim Brys, Ann Nowé
AAAI2
2012 Local linear spectral unmixing via cluster analysis and non-negative matrix factorization for hyperspectral (CHRIS/PROBA) imagery
abstract
We present a novel approach for spectral unmixing in hyperspectral imagery called local linear spectral unmixing (LLSU). Our new proposal relies on an existing strategy for non-linear data modelling where a general non-linear model is approximated via several piecewise linear models. The algorithm is the result of hybridizing two well-known strategies for exploratory high-dimensional data analysis: cluster analysis and non-negative matrix factorization. It has been proposed to answer the limitations of global linear unmixing models which are widely used for spectral unmixing. Our strategy is to first group similar pixels from the hyperspctral image via cluster analysis and then to consider a local linear unmixing model for each cluster in order to obtain the constituent endmembers. Subsequently, the resulted local endmembers from each cluster of mixed pixels are at their turn clustered based on their spectral similarity; the final solution is given by the clusters' centroids.
Cosmin Lazar, Luca Demarchi, David Steenhoff, Jonathan Cheung-Wai Chan, Ann Nowé, Hichem Sahli
IGARSS5
2012 Improving wet clutch engagement with reinforcement learning
abstract
A common approach when applying reinforcement learning to address control problems is that of first learning a policy based on an approximated model of the plant, whose behavior can be quickly and safely explored in simulation; and then implementing the obtained policy to control the actual plant. Here we follow this approach to learn to engage a transmission clutch, with the aim of obtaining a rapid and smooth engagement, with a small torque loss. Using an approximated model of a wet clutch, which simulates a portion of the whole engagement, we first learn an open loop control signal, which is then transferred on the actual wet clutch, and improved by further learning with a different reward function, based on the actual torque loss observed.
Kevin Van Vaerenbergh, Abdel Rodríguez, Matteo Gagliolo, Peter Vrancx, Ann Nowé, Julian Stoev, Stijn Goossens, Gregory Pinte, Wim Symens
IJCNN5
2012 Unlocking the potential of publicly available microarray data using inSilicoDb and inSilicoMerging R/Bioconductor packages
abstract
BACKGROUND: With an abundant amount of microarray gene expression data sets available through public repositories, new possibilities lie in combining multiple existing data sets. In this new context, analysis itself is no longer the problem, but retrieving and consistently integrating all this data before delivering it to the wide variety of existing analysis tools becomes the new bottleneck. RESULTS: We present the newly released inSilicoMerging R/Bioconductor package which, together with the earlier released inSilicoDb R/Bioconductor package, allows consistent retrieval, integration and analysis of publicly available microarray gene expression data sets. Inside the inSilicoMerging package a set of five visual and six quantitative validation measures are available as well. CONCLUSIONS: By providing (i) access to uniformly curated and preprocessed data, (ii) a collection of techniques to remove the batch effects between data sets from different sources, and (iii) several validation tools enabling the inspection of the integration process, these packages enable researchers to fully explore the potential of combining gene expression data for downstream analysis. The power of using both packages is demonstrated by programmatically retrieving and integrating gene expression studies from the InSilico DB repository [https://insilicodb.org/app/].
Jonatan Taminau, Stijn Meganck, Cosmin Lazar, David Steenhoff, Alain Coletta, Colin Molter, Robin Duque, Virginie de Schaetzen, David Y. Weiss Solís, Hugues Bersini, Ann Nowé
BMC Bioinform.11
2012 Optimization of common pool resource sharing in multidomain IP-over-WDM networks
Dimitri Staessens, Didier Colle, Mario Pickavet, Ann Nowé, Kris Steenhaut, Piet Demeester
Comput. Commun.4
2012 A Survey on Filter Techniques for Feature Selection in Gene Expression Microarray Analysis
abstract
A plenitude of feature selection (FS) methods is available in the literature, most of them rising as a need to analyze data of very high dimension, usually hundreds or thousands of variables. Such data sets are now available in various application areas like combinatorial chemistry, text mining, multivariate imaging, or bioinformatics. As a general accepted rule, these methods are grouped in filters, wrappers, and embedded methods. More recently, a new group of methods has been added in the general framework of FS: ensemble techniques. The focus in this survey is on filter feature selection methods for informative feature discovery in gene expression microarray (GEM) analysis, which is also known as differentially expressed genes (DEGs) discovery, gene prioritization, or biomarker discovery. We present them in a unified framework, using standardized notations in order to reveal their technical details and to highlight their common characteristics as well as their particularities.
Cosmin Lazar, Jonatan Taminau, Stijn Meganck, David Steenhoff, Alain Coletta, Colin Molter, Virginie de Schaetzen, Robin Duque, Hugues Bersini, Ann Nowé
IEEE ACM Trans. Comput. Biol. Bioinform.10
2011 Adaptive State Representations for Multi-agent Reinforcement Learning
Yann-Michaël De Hauwere, Peter Vrancx, Ann Nowé
ICAART (2)3
2011 Self-organizing Synchronicity and Desynchronicity using Reinforcement Learning
Mihail Mihaylov, Yann-Aël Le Borgne, Ann Nowé, Karl Tuyls
ICAART (2)3
2011 Continuous Action Reinforcement Learning Automata - Performance and Convergence
Abdel Rodríguez, Ricardo del Corazón Grau-Ábalo, Ann Nowé
ICAART (2)3
2011 Transfer Learning for Multi-agent Coordination
Peter Vrancx, Yann-Michaël De Hauwere, Ann Nowé
ICAART (2)3
2011 Learning a pricing strategy in multi-domain DWDM networks
abstract
In today's Internet the commercial aspect of routing is gaining more and more importance. Commercial agreements between ISPs (i.e. transit and peering agreements) influence the inter-domain routing policies which are now driven by monetary aspects as well as global resource and performance optimization. To allow scalability and protect business critical topology information, hierarchical routing and topology aggregation became a fundamental issue in modern inter-domain networks. In this paper, we introduce a pricing mechanism that takes into account the effects of load dependent internal costs on the domain income and allows ISPs to set link prices during the topology aggregation process. We adapt a Continuous Action Reinforcement Learning Automata (CARLA), to operate in this framework as a tool used by ISP operators to learn the best price according to the network state. The reinforcement signal is proportional only to the domain's utility (i.e. profit) and thus does not need any central authority or sensitive information exchange among domains. Simulation results show that one ISP using CARLA can significantly improve its utility compared to other ISPs that statically choose their link prices. When two ISPs employ the same CARLA they can reach an equilibrium strategy while still improving their utilities.
Pasquale Gurzi, Kris Steenhaut, Ann Nowé, Peter Vrancx
LANMAN3
2011 inSilicoDb: an R/Bioconductor package for accessing human Affymetrix expert-curated datasets from GEO
abstract
Abstract Microarray technology has become an integral part of biomedical research and increasing amounts of datasets become available through public repositories. However, re-use of these datasets is severely hindered by unstructured, missing or incorrect biological samples information; as well as the wide variety of preprocessing methods in use. The inSilicoDb R/Bioconductor package is a command-line front-end to the InSilico DB, a web-based database currently containing 86 104 expert-curated human Affymetrix expression profiles compiled from 1937 GEO repository series. The use of this package builds on the Bioconductor project's focus on reproducibility by enabling a clear workflow in which not only analysis, but also the retrieval of verified data is supported. Availability: inSilicoDb is available as part of the Bioconductor project. There is a companion web interface that can be used for browsing available datasets before importing them into R/Bioconductor (http://insilico.ulb.ac.be). Contact: [email protected]
Jonatan Taminau, David Steenhoff, Alain Coletta, Stijn Meganck, Cosmin Lazar, Virginie de Schaetzen, Robin Duque, Colin Molter, Hugues Bersini, Ann Nowé, David Y. Weiss Solís
Bioinform.10
2010 Demonstrating principal component aggregation for distributed spatial pattern recognition
abstract
The Principal Component Aggregation has recently been proposed as a versatile distributed information extraction technique for sensor networks [3]. This demonstration illustrates its use for a network-level pattern recognition task. Four different patterns, or events, may be sensed by light measurements of a network of 27 nodes. The sens measurements are fused on the fly along a routing tree up to the base station, where the monitored pattern is recognized by a prediction algorithm.
Yann-Aël Le Borgne, Ann Nowé, Kris Steenhaut, Gianluca Bontempi
IPSN2
2010 Lifetime optimization for wireless sensor networks with correlated data gathering
Nashat Abughalieh, Yann-Aël Le Borgne, Kris Steenhaut, Ann Nowé
WiOpt4
2010 Analyzing the dynamics of stigmergetic interactions through pheromone games
Peter Vrancx, Katja Verbeeck, Ann Nowé
Theor. Comput. Sci.3
2009 The coevolution of loyalty and cooperation
abstract
Humans are inclined to engage in long-lasting relationships whose stability does not only rely on cooperation, but often also on loyalty - our tendency to keep interacting with the same partners even when better alternatives exist. Yet, what is the evolutionary mechanism behind such irrational behavior? Furthermore, under which conditions are individuals tempted to abandon their loyalty, and how does this affect the overall level of cooperation? Here, we study a model in which individuals interact along the edges of a dynamical graph, being able to adjust both their behavior and their social ties. Their willingness to sever interactions is determined by an individual characteristic and subject to evolution. We show that defectors ultimately loose any commitment to their social contacts, a result of their inability to establish any social tie under mutual agreement. Ironically, defectors' constant search for new partners to exploit leads to heterogeneous networks in which cooperation survives more easily. Cooperators, on the other hand, develop much more stable and long-term relationships. Their loyalty to their partners only decreases when the competition with defectors becomes fierce. These results indicate how our innate commitment to partners is related to mutual agreement among cooperators and how this commitment is evolutionary disadvantageous in times of conflict, both from an individual and a group perspective.
Sven Van Segbroeck, Francisco C. Santos, Ann Nowé, Jorge M. Pacheco, Tom Lenaerts
IEEE Congress on Evolutionary Computation3
2009 Power Aware Fulfilment of Latency Requirements by Exploiting Heterogeneity in Wireless Sensor and Actuator Networks
abstract
One can imagine a lot of realistic scenarios, such as home or factory automation, where a wide variety of sensors and actuators form a wireless network in order to perform various tasks. There is still a lot of work to be done to ensure scalable and reliable functionality of such a heterogeneous network. We present a MAC layer algorithm that ensures end-to-end latency requirements while minimizing maintenance. We demonstrate that our algorithm scales down the duty cycles along constrained paths in such a way that nodes with more capabilities take on more responsibility in meeting the end-to-end application constraints.
Joris Borms, Kris Steenhaut, Bart Lemmens, Ann Nowé
DSD4
2009 Stochastic Simulation of the Chemoton
abstract
Gánti's chemoton model is an illustrious example of a minimal cell model. It is composed of three stoichiometrically coupled autocatalytic subsystems: a metabolism, a template replication process, and a membrane enclosing the other two. Earlier studies on chemoton dynamics yield inconsistent results. Furthermore, they all appealed to deterministic simulations, which do not take into account the stochastic effects induced by small population sizes. We present, for the first time, results of a chemoton simulation in which these stochastic effects have been taken into account. We investigate the dynamics of the system and analyze in depth the mechanisms responsible for the observed behavior. Our results suggest that, in contrast to the most recent study by Munteanu and Solé, the stochastic chemoton reaches a unique stable division time after a short transient phase. We confirm the existence of an optimal template length and show that this is a consequence of the monomer concentration, which depends on the template length and the initiation threshold. Since longer templates imply shorter division times, these results motivate the selective pressure toward longer templates observed in nature.
Sven Van Segbroeck, Ann Nowé, Tom Lenaerts
Artif. Life2
2008 An Efficient Distributed Self-Organizing Routing Algorithm for Wireless Sensor Networks
abstract
This paper presents a simple routing protocol based on topology control that improves the lifetime of a Wireless Sensor Network in the usual convergecast pattern, by allowing the nodes to choose between two predefined power-levels to forward data towards the sink. The proposed protocol takes advantage of non-homogeneous topologies, where the nodes are grouped in clouds. Nodes will only use the highest power, to establish a link, when necessary, like for bridging the distance between two clouds. Within the clouds, only low-power links are used. The underlying distributed algorithm is shown to converge and finds, for each node of the network, an efficient path to the sink, provided the network is potentially connected at the highest of the two available transmission powers. We also propose a simple routing information refreshing technique that adds robustness to the proposed algorithm.
Hugues Smeets, Kris Steenhaut, Ann Nowé
CISIS3
2008 Using Generalized Learning Automata for State Space Aggregation in MAS
Yann-Michaël De Hauwere, Peter Vrancx, Ann Nowé
KES (1)3
2008 A Learning Automata Approach to Multi-agent Policy Gradient Learning
Maarten Peeters, Ville Könönen, Katja Verbeeck, Ann Nowé
KES (2)4
2008 Coordinated Exploration in Conflicting Multi-stage Games
Maarten Peeters, Ville Könönen, Katja Verbeeck, Sven Van Segbroeck, Ann Nowé
KES (2)5
2008 Decentralized Learning in Markov Games
abstract
Learning automata (LA) were recently shown to be valuable tools for designing multiagent reinforcement learning algorithms. One of the principal contributions of the LA theory is that a set of decentralized independent LA is able to control a finite Markov chain with unknown transition probabilities and rewards. In this paper, we propose to extend this algorithm to Markov games--a straightforward extension of single-agent Markov decision problems to distributed multiagent decision problems. We show that under the same ergodic assumptions of the original theorem, the extended algorithm will converge to a pure equilibrium point between agent policies.
Peter Vrancx, Katja Verbeeck, Ann Nowé
IEEE Trans. Syst. Man Cybern. Part B3
2007 Two-Stage ACO to Solve the Job Shop Scheduling Problem
Amilkar Puris, Rafael Bello 0001, Yaima Trujillo, Ann Nowé, Yailen Martínez-Jiménez
CIARP4
2007 Two-Step Particle Swarm Optimization to Solve the Feature Selection Problem
abstract
In this paper we propose a new model of particle swarm optimization called two-step PSO. The basic idea is to split the heuristic search performed by particles into two stages. We have studied the performance of this new algorithm for the feature selection problem by using the reduct concept of the rough set theory. Experimental results obtained show that the two-step approach improves over the PSO model in calculating reducts, with the same computational cost.
Rafael Bello 0001, Yudel Gómez, María Matilde García Lorenzo, Ann Nowé
ISDA4
2007 QoS in GMPLS based IP/DWDM Metro Networks
abstract
Due to the unpredictability of IP traffic, MANs are migrating from SONET/SDH ring architectures to DWDM mesh infrastructures. Firstly, the fast advance in optical technologies has provided DWDM networks with extremely high bandwidth capability (i.e. OC-192 and OC-768). Secondly, Generalized Multi Protocol Label Switching (GMPLS) and Multilayer Traffic Engineering (MTE) paradigms have turned them into highly flexible, scalable and cost efficient infrastructures able to improve service provider's return on investment (ROI). With the GMPLS automated control plane, IP/DWDM networks can take advantage of a better integration between the electrical and optical domains and consequently of a more optimized resource usage. GMPLS enables network-state-depended dynamic routing and grooming and provides network operators with the tools to differentiate service types and increase the offered QoS. This paper discusses the benefits of GMPLS as an integrator paradigm in IP/DWDM metro infrastructures and proposes a QoS control framework which allows a network operator to accommodate MPLS/DiffServ traffic on different virtual topologies that can guarantee different QoS levels. The benefits of GMPLS/MTE and of the proposed scheme are discussed by means of simulation results.
Walter Colitti, Kris Steenhaut, Ann Nowé
LANMAN3
2007 Exploring selfish reinforcement learning in repeated games with stochastic rewards
Katja Verbeeck, Ann Nowé, Johan Parent, Karl Tuyls
Auton. Agents Multi Agent Syst.2
2006 Two Step Ant Colony System to Solve the Feature Selection Problem
Rafael Bello 0001, Amilkar Puris, Ann Nowé, Yailen Martínez-Jiménez, María Matilde García Lorenzo
CIARP3
2005 Linear genetic programming using a compressed genotype representation
abstract
This paper presents a modularization strategy for linear genetic programming (GP) based on a substring compression/substitution scheme. The purpose of this substitution scheme is to protect building blocks and is in other words a form of learning linkage. The compression of the genotype provides both a protection mechanism and a form of genetic code reuse. This paper presents results for synthetic genetic algorithm (GA) reference problems like SEQ and OneMax as well as several standard GP problems. These include a real world application of GP to data compression. Results show that despite the fact that the compression substrings assumes a tight linkage between alleles, this approach improves the search process.
Johan Parent, Ann Nowé, Kris Steenhaut, Anne Defaweux
Congress on Evolutionary Computation2
2005 A model based on ant colony system and rough set theory to feature selection
abstract
In this paper we propose a hybrid approach to feature selection based on Ant Colony System algorithm and Rough Set Theory. Rough Set Theory offers the heuristic function to measure the quality of a single subset. We have studied the influence of the setting of the parameters for this problem, in particular for finding reducts. Experimental results show this hybrid approach is a promising method for features selection.
Rafael Bello 0001, Ann Nowé, Yailé Caballero Mota, Yudel Gómez, Peter Vrancx
GECCO2
2003 Extended Replicator Dynamics as a Key to Reinforcement Learning in Multi-agent Systems
Karl Tuyls, Dries Heytens, Ann Nowé, Bernard Manderick
ECML3
2002 Evolving Compression Preprocessors With Genetic Programming
Johan Parent, Ann Nowé
GECCO2
2002 Colonies of learning automata
abstract
Originally, learning automata (LAs) were introduced to describe human behavior from both a biological and psychological point of view. In this paper, we show that a set of interconnected LAs is also able to describe the behavior of an ant colony, capable of finding the shortest path from their nest to food sources and back. The field of ant colony optimization (ACO) models ant colony behavior using artificial ant algorithms. These algorithms find applications in a whole range of optimization problems and have been experimentally proved to work very well. It turns out that a known model of interconnected LA, used to control Markovian decision problems (MDPs) in a decentralized fashion, matches perfectly with these ant algorithms. The field of LAs can thus both impart in the understanding of why ant algorithms work so well and may also become an important theoretical tool for learning in multiagent systems (MAS) in general. To illustrate this, we give an example of how LAs can be used directly in common Markov game problems.
Katja Verbeeck, Ann Nowé
IEEE Trans. Syst. Man Cybern. Part B2
2001 Social Agents Playing a Periodical Policy
Ann Nowé, Johan Parent, Katja Verbeeck
ECML1
1998 Q-learning for adaptive load based routing
abstract
The paper deals with the control problem of routing in packet switched networks. Using Q-learning an adaptive, distributed and autonomous routing strategy can be obtained. The objective of the Q-learner under study is to balance the load such that average packet delivery time is optimised. If pure Q-learning is applied to routing each source has to learn the expected cost for sending a message via each of its neighbours for all destinations. Since Q-learning is basically a trial and error method packets have to be sent along non-optimal paths, which artificially increases the load. To reduce this effect and to speed up the learning a variant of Q-learning has been developed. Exploration and exploitation are partially decoupled such that stabilising features can be included in the Q-learning algorithm to cope with instabilities and overhead that might be caused by the costly exploitation in search of alternative paths. In the paper the above statements are justified mathematically and supported by the results of experiments.
Ann Nowé, Kris Steenhaut, Mohamed Fakir, Katja Verbeeck
SMC1
1998 Sugeno, Mamdani, and fuzzy Mamdani controllers put in a uniform interpolation framework
abstract
In this paper we present a general framework for proving that fuzzy controllers are universal approximators. The framework results from the duality between fuzzy control and interpolation. An important aspect of this framework is its constructive nature, which allows us to prove the approximation capabilities of “crisp valued” as well as “fuzzy valued” fuzzy controllers. We also discuss a possible semantic interpretation of a fuzzy conclusion as a fuzzy restriction on the possible values a control action can take. Based on this semantic point of view, robust fuzzy controllers can be built. Where appropriate we refer to related work on the approximation capabilities of fuzzy controllers. © 1998 John Wiley & Sons, Inc.13: 243–256, 1998
Ann Nowé
Int. J. Intell. Syst.1
1992 A self-tuning robust fuzzy controller
Ann Nowé
Microprocess. Microprogramming1
1989 An environment for knowledge based transformational implementation
Viviane Jonckers, Ann Nowé
IEA/AIE (2)2