EDBT 2026 Demo / reviewers in the wild / expert
Kevin Mets
dblp:23/11531
· DBLP profile ↗
19ranked-venue papers
1as first author
15since 2021 · last 2026
0000-0002-4812-4841ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 10 since 2021Systems, architecture and hardware · 4 · 4 since 2021Computer networks · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Constraint-aware reinforcement learning for energy-efficient and regulation-compliant wastewater treatmentabstractThis paper addresses the sustainable and regulation-compliant control of wastewater treatment plants (WWTPs) characterized by nonlinear process dynamics, external disturbances, and multiple, often competing, effluent-quality constraints (e.g., biochemical oxygen demand, ammonia, nitrate, and phosphorus limits). We propose an Adaptive Lagrangian Soft Actor–Critic (AL-SAC) algorithm that incorporates a Lagrangian relaxation term into the maximum-entropy SAC objective in order to impose hard bounds on these key effluent variables. Lagrange multipliers are updated online based on instantaneous constraint violations, eliminating manual penalty tuning and enhancing convergence stability. AL-SAC algorithm is evaluated on the Benchmark Simulation Model No.1 (BSM1) and is found to provide up to a 26% reduction in energy consumption (i.e., aeration and pumping), improve the effluent quality index (EQI) by 3.94%, and significantly reduce the period over which total nitrogen (TN) and ammonia (SNH) concentrations exceed regulatory thresholds. These results demonstrate AL-SAC’s promise for energy-efficient and fully compliant WWTP operation. Omid Sobhani, Thomas Huybrechts, Hamid Toliati, Cristian Camilo Gomez Cortes, Kevin Mets, Siegfried Mercelis |
Expert Syst. Appl. | 5 |
| 2026 | Flexible and Efficient Feature-Level Fusion With Wireless Acoustic Sensors Using Graph Attention NetworksabstractWireless Acoustic Sensor Networks (WASNs), or the Internet of Audio Things (IoAuT), enable intelligent acoustic sensing in IoT applications such as smart homes. A key challenge in such deployments is achieving accurate results under limited bandwidth and energy constraints. To efficiently process audio, sensor fusion techniques can be used. They can aggregate the raw audio signals, features, or local decisions of all sensors to make the final decision for tasks such as acoustic event classification. In contrast to a wired setting, connections within WASNs may be unstable due to interference. Additionally, transmitting large amounts of data reduces the sensors’ battery lifespan. A data/signal-level fusion preserves the full information by transmitting and fusing the raw audio, but it imposes an impractical bandwidth burden for WASNs. Conversely, decision-level fusion is communication-efficient and supports a variable number of sensors, it may compromise accuracy. In contrast, existing feature-level fusion methods can transmit richer information to the fusion center, but incur a higher communication overhead and often necessitate a fixed sensor topology, rendering them less suitable for wireless IoT settings. In this work, we propose a new feature-level fusion framework based on Graph Attention Networks (GATs) for acoustic event classification tasks using a WASN. Our approach supports dynamic WASN topologies and introduces a message condensation layer that reduces the volume of transmitted data, lowering the communication cost and bandwidth usage. Empirical results of a domestic acoustic event classification task show that our framework outperforms decision-level fusion techniques while maintaining a similar communication cost. Moreover, our framework outperforms prior feature-level and data-level fusion methods with a notably reduced communication cost by reducing the size of transmitted messages, and the additional flexibility of supporting dynamic WASN topologies, thus being robust to sensor failures, or the addition or removal of sensors. Wei Wei 0058, Matthias Hutsebaut-Buysse, Thomas Avé, Tom De Schepper, Kevin Mets |
IEEE Internet Things J. | 5 |
| 2026 | An in-depth analysis of discretization methods for communication learning using backpropagation with multi-agent reinforcement learningabstractCommunication is crucial in multi-agent reinforcement learning when agents are not able to observe the full state of the environment. The most common approach to allow learned communication between agents is the use of a differentiable communication channel that allows gradients to flow between agents as a form of feedback. However, this is challenging when we want to use discrete messages to reduce the message size, since gradients cannot flow through a discrete communication channel. Previous work proposed methods to deal with this problem. However, these methods are tested in different communication learning architectures and environments, making it hard to compare them. In this paper, we compare several state-of-the-art discretization methods as well as a novel approach. We do this comparison in the context of communication learning using gradients from other agents and perform tests on several environments. In addition, we present COMA-DIAL, a communication learning approach based on DIAL and COMA extended with learning rate scaling and adapted exploration. COMA-DIAL uses COMA to learn the action policy while it uses the mechanism introduced in DIAL to learn the communication policy. Using COMA-DIAL allows us to perform experiments on more complex environments. Our results show that the novel ST-DRU method, proposed in this paper, achieves the best results out of all discretization methods across the different environments. It achieves the best or close to the best performance in each of the experiments and is the only method that does not fail on any of the tested environments. Astrid Vanneste, Simon Vanneste, Tom De Schepper, Siegfried Mercelis, Peter Hellinckx, Kevin Mets |
Neural Comput. Appl. | 6 |
| 2025 | Reducing the stability gap for continual learning at the edge with class balancingabstractContinual learning (CL) at the edge requires the model to learn from sequentially arriving small batches of data.A naive online learning strategy fails due to the catastrophic forgetting phenomenon.Previous literature introduced the 'latent replay' for CL at the edge, where the input is transformed into latent representations using a pre-trained feature extractor.These latent representations are used, in combination with the real inputs, to train the adaptive classification layers.This approach is prone to the stability gap problem, where the accuracies of learned classes drop when learning a new class, and they only recover during subsequent training iterations.We hypothesize that this is caused by the class imbalance between new class data from the new task, and the old class data in the replay memory.We validate this by applying two class balancing strategies in a latent replay-based CL method.Our empirical results demonstrate that class balancing strategies provide a notable accuracy improvement, and a reduction of the stability gap when using a latent replay-based CL method with a small replay memory size. Wei Wei 0058, Matthias Hutsebaut-Buysse, Tom De Schepper, Kevin Mets |
ESANN | 4 |
| 2025 | Advancing MOSFET Fault Type Detection Through Data-Driven Unsupervised LearningabstractThe Metal-Oxide-Semiconductor Field-Effect Transistor (MOSFET) is a fundamental component in modern electronics, playing a vital role in the amplification and switching of signals within digital circuits. It regulates current flow in response to applied voltage. MOSFETs consist of four terminals: Gate (IG), Drain (ID), Source (IS), and Background or Body (IB). Nevertheless, during the design and manufacturing processes, certain devices are classified as non-functional due to specific abnormalities in their terminals. Identifying the nature of these defects during production can significantly enhance quality control by enabling more accurate detection and correction of faults. This study investigates the use of Artificial Intelligence (AI) and data-driven approaches by applying unsupervised learning techniques to classify the characteristics of faulty MOSFETs. Unsupervised learning enables the analysis of large datasets to examine abnormal devices and distinguish between them based on their distinct properties. By integrating AI-driven diagnostics into the production pipeline, our objective is to establish a system that not only improves yield and operational efficiency but also enhances the reliability of MOSFETs in end-use applications. This approach aims to establish a new standard for precision in quality control semiconductor manufacturing. Abdallah Alfaham, Murat Kocak, Furkan Elmaz, Jérôme Mitard, Joris Vanderschrick, Kevin Mets, Siegfried Mercelis |
IECON | 6 |
| 2025 | Temporal distillation: compressing a policy in space and timeabstractAbstract Deploying deep reinforcement learning on resource-constrained devices remains a significant challenge due to the energy-intensive nature of the sequential decision-making process. Model compression can reduce the spatial (e.g. storage, memory) requirements of a policy network, but this does not always translate to a proportional increase in inference speed and computational efficiency. We introduce a novel temporal compression paradigm that improves the efficiency more directly, by reducing the number of predictions needed to complete a task. This method, based on policy distillation, allows a student model to learn when a change of action will be required by observing sequences of identical actions in the trajectories of an existing teacher model. At each decision, the student can then predict both an action and how many times to perform this action consecutively. This approach allows any existing policy for discrete action spaces to be optimized for energy efficiency through both spatial and temporal compression simultaneously. Experiments on devices ranging from a microcontroller and smartphone processor to a data centre GPU show how this method can decrease the average time it takes to predict an action by up to 13.5 times, compared to 4 times through spatial compression alone, while maintaining a similar average return as the original teacher. In practice, this allows complex models to be deployed on ultra-low-power devices, enabling them to conserve energy by remaining in sleep mode for longer periods, and still achieve high runtime and task performance. Thomas Avé, Matthias Hutsebaut-Buysse, Kevin Mets |
Mach. Learn. | 3 |
| 2025 | GPI-tree search: algorithms for decision-time planning with the general policy improvement theoremabstractIn Reinforcement Learning, Unsupervised Skill Discovery tackles the learning of several policies for downstream task transfer. Once these skills are learnt, the question of how best to use and combine them remains an open problem. The General Policy Improvement Theorem (GPI) creates a policy stronger than any individual skill by selecting the highest-valued policy, generally evaluated with Successor Features. However, the GPI policy is unable to mix and combine the skills at decision time to formulate stronger plans. In this paper, we propose to adopt a model-based setting in order to make such planning possible, and formally show that a forward search improves on the GPI policy and any shallower searches under some approximation term. We argue for decision-time planning, and design a family of algorithms, GPI-Tree Search Algorithms , to use Monte Carlo Tree Search (MCTS) with GPI. These algorithms foster the skills and Q-value priors of the GPI framework to guide and improve the search, which we back up with visual intuition for the different design choices. Our experiments show that the resulting policies are much stronger than the GPI policy alone, even under approximation; they can also improve beyond the linear constraint of Successor Features. Louis Bagot, Lynn D'eer, Steven Latré, Tom De Schepper, Kevin Mets |
Neural Comput. Appl. | 5 |
| 2025 | Scalable reinforcement learning-based neural architecture searchabstractAbstract We assess the feasibility of a reusable neural architecture search agent aimed at amortizing the initial time-investment in building a good search strategy. We do this through the use of Reinforcement Learning, where an agent learns to iteratively select the best way to modify a given neural network architecture. This is achieved using a transformer-based agent design trained using the Ape-X algorithm. We consider both the NAS-Bench-101 and NAS-Bench-301 settings, and compare against various known strong baselines, such as local search and random search. While achieving competitive performance on both benchmarks, the amount of training required for the much larger NAS-Bench-301 is only marginally greater than NAS-Bench-101, illustrating the strong scaling properties of our agent. Our agent is able to achieve strong performance, but the choice of values for certain parameters are crucial to ensuring the succesful training of the agent. We provide some guidance for the selection of appropriate values for hyperparameters through a detailed description of our experimental setup and several ablation studies. Amber Cassimon, Siegfried Mercelis, Kevin Mets |
Neural Comput. Appl. | 3 |
| 2025 | Learning to communicate using a communication critic and counterfactual reasoningabstractLearning to communicate in order to share state information is an active problem in the area of multi-agent reinforcement learning. The credit assignment problem, the non-stationarity of the communication environment and the problem of encouraging the agents to be influenced by incoming messages are major challenges within this research field which need to be overcome in order to learn a valid communication protocol. This paper introduces the novel multi-agent counterfactual communication learning (MACC) method which adapts counterfactual reasoning in order to overcome the credit assignment problem for communicating agents. Next, the non-stationarity of the communication environment, while learning the communication Q -function, is overcome by creating the communication Q -function using the action policy of the other agents and the Q -function of the action environment. As the exact method to create the communication Q -function can be computationally intensive for a large number of agents, two approximation methods are proposed. Additionally, a social loss function is introduced in order to create influenceable agents, which is required to learn a valid communication protocol. Our experiments show that MACC is able to outperform the state-of-the-art baselines in four different scenarios in the particle environment. Finally, we demonstrate the scalability of MACC in a matrix environment. Simon Vanneste, Astrid Vanneste, Kevin Mets, Tom De Schepper, Ali Anwar 0002, Siegfried Mercelis, Peter Hellinckx |
Neural Comput. Appl. | 3 |
| 2024 | Online Adaptation of Compressed Models by Pre-Training and Task-Relevant PruningabstractNeural networks are increasingly deployed on edge devices, where they must adapt to new data in dynamic environments.Here, model compression techniques like pruning are essential.This involves removing redundant neurons, increasing efficiency at the cost of accuracy, and creating a conflict between efficiency and adaptability.We propose a novel method for training and compressing models that maintains and extends their ability to generalize to new data, improving online adaptation without reducing compression rates.By pre-training the model on additional knowledge and identifying the parts of the deep neural network that actually encode task-relevant knowledge, we can effectively prune the model by 80% and achieve 16% higher accuracies when adapting to new domains. Thomas Avé, Matthias Hutsebaut-Buysse, Wei Wei 0058, Kevin Mets |
ESANN | 4 |
| 2024 | Policy Compression for Low-Power Intelligent Scaling in Software-Based Network ArchitecturesabstractModern networks, characterized by their complexity and heterogeneity, are transitioning from manual to automated and intelligent management, according to the vision of Autonomous Networks (ANs). Leveraging data-driven techniques, ANs aim to provide the "Zero-X" and "Self-X" experience, where intelligent and adaptable network operations are key cornerstones. Such is the case of intelligent resource scaling, where the goal is to optimize resource orchestration to maximize efficiency, reduce latency, and maintain high-quality service, even amid fluctuating network loads and changing service requirements. Unfortunately, current approaches for auto-scaling are computationally expensive to deploy in resource-constrained devices such as those found at the edge or beyond. This paper introduces an innovative four-phase approach to train a compact, data-driven Deep Reinforcement Learning (DRL) scaler that can be deployed on low-power devices. Our results demonstrate the scalability and efficiency of this model, achieving state-of-the-art scaling with up to 1003x fewer parameters, enhancing interpretability and computational efficiency, making it a robust solution for intelligent resource scaling in network environments. The resulting 1487x runtime speed improvement and 28.5x reduction in memory requirements allow the scaler to be deployed on low-power devices and still operate in real-time, which is essential for mission-critical and latency-sensitive applications. Thomas Avé, Paola Soto, Miguel Camelo, Tom De Schepper, Kevin Mets |
NOMS | 5 |
| 2023 | Directed Real-World Learned ExplorationabstractAutomated Guided Vehicles (AGV) are omnipresent, and are able to carry out various kind of preprogrammed tasks. Unfortunately, a lot of manual configuration is still required in order to make these systems operational, and configuration needs to be re-done when the environment or task is changed. As an alternative to current inflexible methods, we employ a learning based method in order to perform directed exploration of a previously unseen environment. Instead of relying on handcrafted heuristic representations, the agent learns its own environmental representation through its embodiment. Our method offers loose coupling between the Reinforcement Learning (RL) agent, which is trained in simulation, and a separate, on real-world images trained task module. The uncertainty of the task module is used to direct the exploration behavior. As an example, we use a warehouse inventory task, and we show how directed exploration can improve the task performance through active data collection. We also propose a novel environment representation to efficiently tackle the sim2real gap in both sensing and actuation. We empirically evaluate the approach both in simulated environments and a real-world warehouse. Matthias Hutsebaut-Buysse, Ferran Gebelli Guinjoan, Erwin Rademakers, Steven Latré, Abdellatif Bey-Temsamani, Kevin Mets, Erik Mannens, Tom De Schepper |
IROS | 6 |
| 2022 | Object Detection To Enable Autonomous Vessels On European Inland WaterwaysabstractTo enable autonomous vessels to operate on inland waterways, they need to detect, track and localize objects at close range to safely navigate. We deployed current deep learning techniques to detect and track these objects. As there are no large labeled datasets of European inland waterways, we used transfer learning to overcome the lack of data. By using preexisting similar datasets, we were able to significantly decrease the required amount of labeled data from the target distribution. Furthermore, we improved the mean Average Precision from 0.461 to 0.814 by using a limited number of labeled target data samples. We estimated the relative distance of the objects based on the generated bounding boxes. The information from the camera is then combined with LiDar data to generate a top-view map of the environment which is used as input for an object-avoidance control agent. All these methods can run in real-time on the vessel with an fps of 1.83 on a 2.7GHz vCPU. Mattias Billast, Robin Janssens, Astrid Vanneste, Simon Vanneste, Olivier Vasseur, Ali Anwar 0002, Kevin Mets, Tom De Schepper, José Oramas M., Steven Latré, Peter Hellinckx |
IECON | 7 |
| 2022 | Safety Aware Autonomous Path Planning Using Model Predictive Reinforcement Learning for Inland WaterwaysabstractIn recent years, interest in autonomous shipping in urban waterways has increased significantly due to the trend of keeping cars and trucks out of city centers. Classical approaches such as Frenet frame based planning and potential field navigation often require tuning of many configuration parameters and sometimes even require a different configuration depending on the situation. In this paper, we propose a novel path planning approach based on reinforcement learning called Model Predictive Reinforcement Learning (MPRL). MPRL calculates a series of waypoints for the vessel to follow. The environment is represented as an occupancy grid map, allowing us to deal with any shape of waterway and any number and shape of obstacles. We demonstrate our approach on two scenarios and compare the resulting path with path planning using a Frenet frame and path planning based on a proximal policy optimization (PPO) agent. Our results show that MPRL outperforms both baselines in both test scenarios. The PPO based approach was not able to reach the goal in either scenario while the Frenet frame approach failed in the scenario consisting of a corner with obstacles. MPRL was able to safely (collision free) navigate to the goal in both of the test scenarios. Astrid Vanneste, Simon Vanneste, Olivier Vasseur, Robin Janssens, Mattias Billast, Ali Anwar 0002, Kevin Mets, Tom De Schepper, Siegfried Mercelis, Peter Hellinckx |
IECON | 7 |
| 2021 | Disagreement Options: Task Adaptation Through Temporally Extended Actions
Matthias Hutsebaut-Buysse, Tom De Schepper, Kevin Mets, Steven Latré |
ECML/PKDD (1) | 3 |
| 2020 | Language Grounded Task-Adaptation in Reinforcement Learning
Matthias Hutsebaut-Buysse, Kevin Mets, Steven Latré |
ESANN | 2 |
| 2013 | Design of a management infrastructure for smart grid pilot data processing and analysis
Matthias Strobbe, Tom Verschueren, Stijn Melis, Dieter Verslype, Kevin Mets, Filip De Turck, Chris Develder |
IM | 5 |
| 2012 | Distributed multi-agent algorithm for residential energy management in smart gridsabstractDistributed renewable power generators, such as solar cells and wind turbines are difficult to predict, making the demand-supply problem more complex than in the traditional energy production scenario. They also introduce bidirectional energy flows in the low-voltage power grid, possibly causing voltage violations and grid instabilities. In this article we describe a distributed algorithm for residential energy management in smart power grids. This algorithm consists of a market-oriented multi-agent system using virtual energy prices, levels of renewable energy in the real-time production mix, and historical price information, to achieve a shifting of loads to periods with a high production of renewable energy. Evaluations in our smart grid simulator for three scenarios show that the designed algorithm is capable of improving the self consumption of renewable energy in a residential area and reducing the average and peak loads for externally supplied power. Kevin Mets, Matthias Strobbe, Tom Verschueren, Thomas Roelens, Filip De Turck, Chris Develder |
NOMS | 1 |
| 2012 | Design and evaluation of an architecture for future smart grid service provisioningabstractThe increase of distributed renewable electricity generators, such as solar cells and wind turbines, requires new energy management systems where real-time measurements and communication between end users, suppliers and utilities are vital. To address this need, we propose a common service architecture that allows houses with renewable energy generation and smart energy devices to plug into a distributed energy management system, integrated with the public power grid. The presented architecture facilitates end-users to optimize their energy consumption, enables power network operators to better balance supply and demand, and creates a platform where new market players (e.g. ESCOs) can easily provide new services. This service architecture has been implemented and is currently evaluated in a field trial with 21 users, of which we present the initial results. Matthias Strobbe, Tom Verschueren, Kevin Mets, Stijn Melis, Chris Develder, Filip De Turck, Thierry Pollet, Stijn Van de Veire |
NOMS | 3 |