EDBT 2026 Demo / reviewers in the wild / expert
Sebastien Gros
dblp:125/5655 · also Sébastien Gros
· DBLP profile ↗
17ranked-venue papers
0as first author
16since 2021 · last 2026
0000-0001-6054-2133ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 10 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021Systems, architecture and hardware · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | corl: Reinforcement Learning of MILP Policies Solved via Branch-and-Bound
Akhil S. Anand, Elias Aarekol, Martin Mziray Dalseg, Magnus Stålhane, Sebastien Gros |
CPAIOR | 5 |
| 2026 | Predicting What Matters: Training AI Models for Better DecisionsabstractArtificial intelligence (AI) models that predict the future behavior of real-world systems and processes (also known as predictive AI models) are central to intelligent decision-making. They are often employed in model-based decision-making frameworks to optimize decisions for real-world tasks based on their predictions. However, decisions optimized using such predictive AI models often result in suboptimal performance when applied in the real world. This is primarily because these models are typically constructed to best fit the behavior of the real-world system, and hence to predict the most likely future rather than to optimize the best possible decisions for a given task. Due to this objective mismatch, their predictions cannot be guaranteed to support optimal decision-making in theory or in practice. In fact, there is increasing empirical evidence and consensus that predictive models must be tailored to decision-making objectives to achieve optimal real-world performance. Supporting this observation, we establish formal (necessary and sufficient) conditions that a predictive model (AI-based or not) must satisfy for a decision-making policy derived using that model to achieve optimal performance in the real world. We then discuss their implications for building predictive AI models for optimal sequential decision-making. Akhil S. Anand, Shambhuraj Sawant, Dirk Reinhardt, Sebastien Gros |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2025 | Sky Savers: Leveraging Drone Technology for Victim Localization in Avalanche Rescue via Transceiver Signal Analysis
Robin Vetsch, Samuel Kranz, Tindaro Pittorino, Peter de Baets, Martial Châteauvieux, Christoph Würsch, Daniel Lenz, Sebastien Gros |
ICINCO (1) | 8 |
| 2025 | Offline Guarded Safe Reinforcement Learning for Medical Treatment Optimization StrategiesabstractWhen applying offline reinforcement learning (RL) in healthcare scenarios, the out-of-distribution (OOD) issues pose significant risks, as inappropriate generalization beyond clinical expertise can result in potentially harmful recommendations.
While existing methods like conservative Q-learning (CQL) attempt to address the OOD issue, their effectiveness is limited by only constraining action selection by suppressing uncertain actions. This action-only regularization imitates clinician actions that prioritize short-term rewards, but it fails to regulate downstream state trajectories, thereby limiting the discovery of improved long-term treatment strategies. To safely improve policy beyond clinician recommendations while ensuring that state-action trajectories remain in-distribution, we propose \textit{Offline Guarded Safe Reinforcement Learning} ($\mathsf{OGSRL}$), a theoretically grounded model-based offline RL framework. $\mathsf{OGSRL}$ introduces a novel dual constraint mechanism for improving policy with reliability and safety. First, the OOD guardian is established to specify clinically validated regions for safe policy exploration. By constraining optimization within these regions, it enables the reliable exploration of treatment strategies that outperform clinician behavior by leveraging the full patient state history, without drifting into unsupported state-action trajectories. Second, we introduce a safety cost constraint that encodes medical knowledge about physiological safety boundaries, providing domain-specific safeguards even in areas where training data might contain potentially unsafe interventions. Notably, we provide theoretical guarantees on safety and near-optimality: policies that satisfy these constraints remain in safe and reliable regions and achieve performance close to the best possible policy supported by the data. When evaluated on the MIMIC-III sepsis treatment dataset, $\mathsf{OGSRL}$ demonstrated significantly better OOD handling than baselines. $\mathsf{OGSRL}$ achieved a 78\% reduction in mortality estimates and a 51\% increase in reward compared to clinician decisions. Runze Yan, Xun Shen, Akifumi Wachi, Sebastien Gros, Anni Zhao, Xiao Hu 0002 |
NeurIPS | 4 |
| 2025 | Automatic selection of inducing points in sparse Gaussian process for approximations of finite element analysesabstractGaussian process regression (GPR) is a widely used regression model, but it has poor scalability. Sparse approximation methods improve scalability by using inducing points to approximate the GPR, but determining the optimal number and placement of these points is challenging. Increasing the number of inducing points generally improves the predictive accuracy, but it comes at a computational cost. This article presents a method to estimate the necessary number of inducing points for accurate predictions of finite element method (FEM) analyses using approximate GPR. The approach leverages the proper orthogonal decomposition (POD) technique, using its modes to determine the inducing points. Results demonstrate that the proposed method identifies a sufficient number of inducing points for approximate GPR to achieve predictive accuracy comparable to full GPR, but with half the training time. This approach ensures computational efficiency without significant loss in accuracy, making it a valuable tool for scalable regression in engineering applications . POD has previously been combined with GPR to provide computationally efficient predictions for the full solution field across unseen variable combinations, treating spatial components separately via reduced basis functions. However, this work treats the spatial component as a variable within the GPR approximation, allowing continuous spatial predictions. This ensures that the covariance in the spatial dimension is captured by a single GPR. The method is applied to simulations of a three-span, post-tensioned concrete girder bridge. Heine Havneraas Røstum, Sebastien Gros, Ketil Aas-Jakobsen, Joseph Morlier |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | Application of Soft Actor-Critic algorithms in optimizing wastewater treatment with time delays integrationabstractWastewater treatment plants face unique challenges for process control due to their complex dynamics, slow time constants, and stochastic delays in observations and actions. These characteristics make conventional control methods, such as Proportional-Integral-Derivative controllers, suboptimal for achieving efficient phosphorus removal, a critical component of wastewater treatment to ensure environmental sustainability. This study addresses these challenges using a novel deep reinforcement learning approach based on the Soft Actor-Critic algorithm, integrated with a custom simulator designed to model the delayed feedback inherent in wastewater treatment plants. The simulator incorporates Long Short-Term Memory networks for accurate multi-step state predictions, enabling realistic training scenarios. To account for the stochastic nature of delays, agents were trained under three delay scenarios: no delay, constant delay, and random delay. The results demonstrate that incorporating random delays into the reinforcement learning framework significantly improves phosphorus removal efficiency while reducing operational costs. Specifically, the delay-aware agent achieved 36 % reduction in phosphorus emissions, 55 % higher reward, 77 % lower target deviation from the regulatory limit, and 9 % lower total costs than traditional control methods in the simulated environment. These findings underscore the potential of reinforcement learning to overcome the limitations of conventional control strategies in wastewater treatment, providing an adaptive and cost-effective solution for phosphorus removal. • Novel SAC framework handles time delays in wastewater treatment optimization. • Delay-aware RL models improve phosphorus control efficiency by 36%. • SAC agents reduce target deviations by 77% and operational costs by 9%. • Custom LSTM-based simulator enables realistic training for delay scenarios. • Demonstrates RL’s superiority over PID controllers in dynamic industrial processes. Esmaeel Mohammadi, Daniel Ortiz Arroyo, Aviaja Anna Hansen, Mikkel Stokholm-Bjerregaard, Sebastien Gros, Akhil S. Anand, Petar Durdevic |
Expert Syst. Appl. | 5 |
| 2024 | Noise2Inverse for 3D Low-Dose Cone-Beam Computed TomographyabstractCone-beam computed tomography (CBCT) is a noninvasive x-ray imaging technique with multiple clinical applications (orthopedics, dentistry, radiation therapy, etc.). These systems can provide sub-millimeter resolution in images of high diagnostic quality with short scanning times. However, CBCT scans require subjecting the patient with a sufficient radiation dose to obtain a high-quality image. Although relatively low, some clinical applications such as image guidance in radiotherapy (IGRT) require obtaining CBCT daily. This repeated exposure to x-ray radiation outside the treatment volume can have detrimental effects and may increase the lifetime risk of a secondary malignancy in younger patients. Hence, there is a clinical need to reduce CBCT imaging dose. However, images produced with a reduced dose often have greater noise and artifacts which deteriorate CBCT image quality, reducing their clinical efficacy. Deep learning-based methods have become a popular choice showing impressive performance for noise reduction in low-dose CBCT. The success of these methods critically depends on the availability of paired images for training which is a major obstacle for most applications. Recently, the Noise2Inverse method was designed specifically for denoising computed tomography (CT) data without requiring paired noisy and clean images. While the Noise2Inverse method was shown to reduce noise in challenging real-world experimental datasets, it processes 3D-CT volumes as individual 2D images. In this work, we extend the Noise2Inverse method to work directly with 3D-CBCT volumes requiring the use of high-end HPC resources. We compare the results between a 2D and 3D UNet++ model utilizing the ICASSP-2024 3D Low Dose CBCT Grand Challenge dataset with CBCT data corresponding to two dose levels, clinical and low-dose. Experimental results show that both 2D and 3D models can significantly reduce the noise in CBCT increasing the SSIM up to 148.05% compared to the low-dose FDK reconstruction. Furthermore, the 3D model improves the SSIM by 0.26% and 1.46% for clinical- and low-dose, respectively, over the 2D model. Austin Yunker, Jason Luce, John C. Roeske, Rajkumar Kettimuthu, Hyejoo Kang, Sebastien Gros, Alec M. Block |
IEEE Big Data | 6 |
| 2024 | Optimal Power Management of Multi-energy Community Considering The Local Energy MarketabstractThis paper proposes a market mechanism that enables the advanced distribution management system (ADMS) for energy trading in the local energy market. Two primary functions of the ADMS are discussed: reducing operational costs and coordinating the energy community (EC). The integration of these two functionalities considers all the constraints for maintaining distribution system reliability. Two energy transaction actions are allowed for the ECs: 1) each prosumer in the EC trades energy locally to maximize the global social welfare and pay network usage fees to the distributed system operator (DSO); 2) the EC trades energy with the DSO to sell their exceed energy or purchase energy to meet their load demand. The payment among the ECs is subject to a clearing price. In the market, the Newton distributed algorithm has been used to achieve the equilibrium point, and the ADMS actively optimizes voltage, and reactive power (Volt-VAR) via controlling distributed energy resources (DER) and managing the iterative process. To confirm the functionality of the model, 33-bus networks have been used to model the ECs. Younes Zahraoui, Sebastien Gros, Irina Oleinikova |
IECON | 2 |
| 2024 | Flipping-based Policy for Chance-Constrained Markov Decision ProcessesabstractSafe reinforcement learning (RL) is a promising approach for many real-world decision-making problems where ensuring safety is a critical necessity. In safe RL research, while expected cumulative safety constraints (ECSCs) are typically the first choices, chance constraints are often more pragmatic for incorporating safety under uncertainties. This paper proposes a \textit{flipping-based policy} for Chance-Constrained Markov Decision Processes (CCMDPs). The flipping-based policy selects the next action by tossing a potentially distorted coin between two action candidates. The probability of the flip and the two action candidates vary depending on the state. We establish a Bellman equation for CCMDPs and further prove the existence of a flipping-based policy within the optimal solution sets. Since solving the problem with joint chance constraints is challenging in practice, we then prove that joint chance constraints can be approximated into Expected Cumulative Safety Constraints (ECSCs) and that there exists a flipping-based policy in the optimal solution sets for constrained MDPs with ECSCs. As a specific instance of practical implementations, we present a framework for adapting constrained policy optimization to train a flipping-based policy. This framework can be applied to other safe RL algorithms. We demonstrate that the flipping-based policy can improve the performance of the existing safe RL algorithms under the same limits of safety constraints on Safety Gym benchmarks. Xun Shen, Akifumi Wachi, Kazumune Hashimoto, Sebastien Gros |
NeurIPS | 5 |
| 2023 | Optimization of the model predictive control meta-parameters through reinforcement learningabstractModel predictive control (MPC) is increasingly being considered for control of fast systems and embedded applications. However, MPC has some significant challenges for such systems, such as its high computational complexity. Further, the MPC parameters must be tuned, which is largely a trial-and-error process that affects the control performance, the robustness, and the computational complexity of the controller to a high degree. This paper presents a multivariate optimization method based on reinforcement learning (RL) that automatically tunes the control algorithm’s parameters from data to achieve optimal closed-loop performance. The main contribution of our method is the inclusion of state-dependent optimization of the meta-parameters of MPC, i.e. parameters that are non-differentiable wrt. the MPC solution. Our control algorithm is based on an event-triggered MPC, where we learn when the MPC should be re-computed, and a dual-mode MPC and linear state feedback control law applied in between MPC computations. We formulate a novel mixture-distribution RL policy determining the meta-parameters of our control algorithm and show that with joint optimization we achieve improvements that do not present themselves with univariate optimization of the same parameters. We demonstrate our framework on the inverted pendulum control task, reducing the total computation time of the control system by 36% while also improving the control performance by 18.4%. Eivind Bøhn, Sebastien Gros, Signe Moe, Tor Arne Johansen |
Eng. Appl. Artif. Intell. | 2 |
| 2023 | Energy management in residential microgrid using model predictive control-based reinforcement learning and Shapley valueabstractThis paper presents an Energy Management (EM) strategy for residential microgrid systems using Model Predictive Control (MPC)-based Reinforcement Learning (RL) and Shapley value. We construct a typical residential microgrid system that considers fluctuating spot-market prices, highly uncertain user demand and renewable generation, and collective peak power penalties. To optimize the benefits for all residential prosumers, the EM problem is formulated as a Cooperative Coalition Game (CCG). The objective is to first find an energy trading policy that reduces the collective economic cost (including spot-market cost and peak-power cost) of the residential coalition, and then to distribute the profits obtained through cooperation to all residents. An MPC-based RL approach, which compensates for the shortcomings of MPC and RL and benefits from the advantages of both, is proposed to reduce the monthly collective cost despite the system uncertainties. To determine the amount of monthly electricity bill each resident should pay, we transfer the cost distribution problem into a profit distribution problem. Then, the Shapley value approach is applied to equitably distribute the profits (i.e., cost savings) gained through cooperation to all residents based on the weighted average of their respective marginal contributions. Finally, simulations are performed on a three-household microgrid system located in Oslo, Norway, to validate the proposed strategy, where a real-world dataset of April 2020 is used. Simulation results show that the proposed MPC-based RL approach could effectively reduce the long-term economic cost by about 17.5%, and the Shapley value method provides a solution for allocating the collective bills fairly. Wenqi Cai, Arash Bahari Kordabad, Sebastien Gros |
Eng. Appl. Artif. Intell. | 3 |
| 2023 | Numerical Strategies for Mixed-Integer Optimization of Power-Split and Gear Selection in Hybrid Electric VehiclesabstractThis paper presents numerical strategies for a computationally efficient energy management system that co-optimizes the power split and gear selection of a hybrid electric vehicle (HEV). We formulate a mixed-integer optimal control problem (MIOCP) that is transcribed using multiple-shooting into a mixed-integer nonlinear program (MINLP) and then solved by nonlinear model predictive control. We present two different numerical strategies, a Selective Relaxation Approach (SRA), which decomposes the MINLP into several subproblems, and a Round-n-Search Approach (RSA), which is an enhancement of the known ‘relax-n-round’ strategy. Subsequently, the resulting algorithmic performance and optimality of the solution of the proposed strategies are analyzed against two benchmark strategies; one using rule-based gear selection, which is typically used in production vehicles, and the other using dynamic programming (DP), which provides a global optimum of a quantized version of the MINLP. The results show that both SRA and RSA enable about 3.6% cost reduction compared to the rule-based strategy, while still being within 1% of the DP solution. Moreover, for the case studied RSA takes about 35% less mean computation time compared to SRA, while both SRA and RSA being about 99 times faster than DP. Furthermore, both SRA and RSA were able to overcome the infeasibilities encountered by a typical rounding strategy under different drive cycles. The results show the computational benefit of the proposed strategies, as well as the energy saving possibility of co-optimization strategies in which actuator dynamics are explicitly included. Anand Ganesan, Sebastien Gros, Nikolce Murgovski |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2023 | Distributed Eco-Driving Control of a Platoon of Electric Vehicles Through Riccati RecursionabstractThis paper presents a distributed optimization procedure for the cooperative eco-driving control problem of a platoon of electric vehicles subject to safety and travel time constraints. Individual optimal trajectories are generated for each platoon member to account for heterogeneous vehicles and for the road slope. By rearranging the problem variables, the Riccati recursion can be applied along the chain-like structure of the platoon and be used to solve the problem by repeatedly transmitting information up and down the platoon. Since each vehicle is only responsible for its own part of the computations, the proposed control strategy is privacy-preserving and could therefore be deployed by any group of vehicles to form a platoon spontaneously while driving. The energy efficiency of this control strategy is evaluated in numerical experiments for platoons of electric trucks with different masses and rated motor powers. Rémi Lacombe, Sebastien Gros, Nikolce Murgovski, Balázs Kulcsár |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2022 | Can AI Abuse Personal Information in an EV Fast-Charging Market?abstractIn order to alleviate the range anxiety of electric vehicle users (EVUs), several researches focus on facilitating the efficiency of fast-electric vehicle charging stations (fast-EVCSs) using artificial intelligence (AI). This paper first proposes a fast-EVCS revenue maximization pricing policy using an AI approach, and we argue that the AI algorithm can learn to abuse EVUs information for maximizing its revenue. In order to investigate the hypothesis, firstly, a simulation environment is developed using vehicle performance models and an EVU’s charging station selection game. Then, we formulate the charging station revenue maximization problem as a Markov decision process (MDP) and propose a personalized dynamic pricing policy using a model-free reinforcement learning (RL) algorithm. From numerical simulation results, it is found that if the RL approach focuses solely on increasing revenue of the fast-EVCSs, it can learn to misuse personal information without any human intervention. To prevent such abuse, we suggest intuitive guidelines for policymakers and urban planners via numerical experiments. Sangjun Bae, Sebastien Gros, Balázs Kulcsár |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2022 | A Game Approach for Charging Station Placement Based on User Preferences and CrowdednessabstractThe placement of electric vehicle charging stations (EVCSs), which encourages the rapid development of electric vehicles (EVs), should be considered from not only operational perspective such as minimizing installation costs, but also user perspective so that their strategic and competitive charging behaviors can be reflected. This paper proposes a methodological framework to consider crowdedness and individual preferences of electric vehicle users (EVUs) in the selection of locations for fast-charging stations. The electric vehicle charging station placement problem (EVCSPP) is solved via a decentralized game theoretical decision-making algorithm and$k$-means clustering algorithm. The proposed algorithm, referred to as$k$-GRAPE, determines the locations of charging stations to maximize the sum of utilities of EVUs. In particular, we analytically present that 50% of suboptimality of the solution can be at least guaranteed, which is about 17% better than the existing game theoretical based framework. We show a few variants to describe the utility functions that may capture the difference in preferences of EVUs. Finally, we demonstrate the viability of the decision framework via three real-world data-based experiments. The results of the experiments, including a comparison with a baseline method are then discussed. Sangjun Bae, Inmo Jang, Sebastien Gros, Balázs Kulcsár, Jonas Hellgren |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2022 | Bilevel Optimization for Bunching Mitigation and Eco-Driving of Electric Bus LinesabstractThe problems of bus bunching mitigation and the energy management of groups of vehicles have traditionally been treated separately in the literature and been formulated in two different frameworks. The present work bridges this gap by formulating the optimal control problem of the bus line eco-driving and regularity control as a smooth, multi-objectivenonlinear program. Since this nonlinear program has only a few coupling variables, it is shown how it can be solved in parallel aboard each bus, such that only a marginal amount of computations need to be carried out centrally. This procedure leverages the structure of the bus line by enabling parallel computations and reducing the communication loads between the buses, which makes the problem resolution scalable in terms of the number of buses. Closed-loop control is then achieved by embedding this procedure in amodel predictive control. Stochastic simulations based on real passengers and travel times data are realized for several scenarios with different levels of bunching for a line of electric buses. Our method achieves fast recoveries to regular headways as well as energy savings of up to 9.3% when compared with traditional holding or speed control baselines. Rémi Lacombe, Sebastien Gros, Nikolce Murgovski, Balázs Kulcsár |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2018 | Super-short Term Wind Speed Prediction based on Artificial Neural Networks for Wind Turbine Control ApplicationsabstractIn this paper, an Artificial Neural Network (ANN) methodology to cast super-short term (under 30 seconds) wind speed predictions is presented. The aim is to obtain computationally efficient super-short term predictions that will be used in Wind Turbine (WT) real-time control applications in the future. A combination of power measurements and meteorological data are used to obtain the estimated rotor effective wind speed. This signal is then used as an input to train the ANNs. Additionally, a polynomial fitting is proposed to enhance the ANN results at each prediction step. The proposed strategy is compared with a classic persistence approach in order to quantify the achieved improvement. Julio Luna, Sebastien Gros, Jens Geisler, Ole Falkenberg, Rafal Noga, Axel Schild |
IECON | 2 |