Yassine Chemingui

dblp:371/4860 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
4since 2021 · last 2026
0000-0001-6288-1729ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Reinforcement learning · 82% Optimization for machine learning · 18%
Theoretical computer science
2 papers
Mathematical optimization · 100%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational science and engineering · 100%

Topics — the 12 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning › safe reinforcement learning
offline safe reinforcement learning
1.722025
Online Optimization for Offline Safe Reinforcement Learning · NeurIPS 2025
Constraint-Adaptive Policy Switching for Offline Safe Reinforcement Learning · AAAI 2025
Machine learning › Reinforcement learning
safe reinforcement learning
1.722025
Online Optimization for Offline Safe Reinforcement Learning · NeurIPS 2025
Constraint-Adaptive Policy Switching for Offline Safe Reinforcement Learning · AAAI 2025
Machine learning › Optimization for machine learning › model-based optimization › bayesian optimization
surrogate model
1.012026
Nanoporous Materials Discovery via Search Bias-Guided Surrogate Modeling · AAAI 2026
Computational science and engineering › materials science
materials discovery
1.012026
Nanoporous Materials Discovery via Search Bias-Guided Surrogate Modeling · AAAI 2026
Machine learning › Reinforcement learning
constrained reinforcement learning
0.912025
Constraint-Adaptive Policy Switching for Offline Safe Reinforcement Learning · AAAI 2025
Machine learning › Reinforcement learning
offline reinforcement learning
0.812024
Offline Model-Based Optimization via Policy-Guided Gradient Search · AAAI 2024
Machine learning › Reinforcement learning › model-based reinforcement learning
policy-guided search
0.812024
Offline Model-Based Optimization via Policy-Guided Gradient Search · AAAI 2024
Mathematical optimization
black-box optimization
0.812024
Offline Model-Based Optimization via Policy-Guided Gradient Search · AAAI 2024
Mathematical optimization
continuous optimization
0.812024
Offline Model-Based Optimization via Policy-Guided Gradient Search · AAAI 2024
Mathematical optimization › black-box optimization
offline model-based optimization
0.812024
Offline Model-Based Optimization via Policy-Guided Gradient Search · AAAI 2024
Machine learning › Optimization for machine learning › model-based optimization
bayesian optimization
0.312026
Nanoporous Materials Discovery via Search Bias-Guided Surrogate Modeling · AAAI 2026
Mathematical optimization
online optimization
0.312025
Online Optimization for Offline Safe Reinforcement Learning · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

value matching loss · 2.0surrogate modeling · 2.0optimization bias regularization · 2.0offline RL oracle · 1.7no-regret online optimization · 1.7minimax objective · 1.7offline reinforcement learning · 1.5gradient search · 1.5shared representation learning · 0.9policy switching · 0.9surrogate model · 0.8
YearPublicationVenuePosition
2026 Nanoporous Materials Discovery via Search Bias-Guided Surrogate Modeling
abstract
Nanoporous materials (NPMs) are suitable for solving some of the society's biggest challenges including carbon capture and conversion, storing hydrogen and methane, and sensing gases. The key challenge in discovering high-performing NPMs for a target application is that making and evaluating candidate NPMs requires performing resource-expensive wet-lab experiments. We consider the problem of discovering NPMs using existing experimental data of NPM evaluations. The overall goal is to find better NPMs than the best NPMs from the past experimental data. A simple approach is to create a surrogate model to match the objective values on the given dataset and employ it to score candidate NPMs to discover optimized NPMs. However, this surrogate model will fail because it does not have the appropriate search bias for the goal of optimization. To address this challenge, we propose a novel surrogate modeling approach that combines value matching loss with an optimization bias as regularizer. The key idea is to algorithmically realize search bias is to mimic the search behavior of monotonically increasing sequences of NPMs from the given dataset. Experiments on multiple real-world NPM discovery tasks demonstrate that our proposed surrogate model discovers significantly better NPMs than baselines including value matching surrogate model and one-step Bayesian optimization.
Azza Fadhel, Yassine Chemingui, Aryan Deshwal, Trong Nghia Hoang, Janardhan Rao Doppa
AAAI2
2025 Constraint-Adaptive Policy Switching for Offline Safe Reinforcement Learning
abstract
Offline safe reinforcement learning (OSRL) involves learning a decision-making policy to maximize rewards from a fixed batch of training data to satisfy pre-defined safety constraints. However, adapting to varying safety constraints during deployment without retraining remains an under-explored challenge. To address this challenge, we introduce constraint-adaptive policy switching (CAPS), a wrapper framework around existing offline RL algorithms. During training, CAPS uses offline data to learn multiple policies with a shared representation that optimize different reward and cost trade-offs. During testing, CAPS switches between those policies by selecting at each state the policy that maximizes future rewards among those that satisfy the current cost constraint. Our experiments on 38 tasks from the DSRL benchmark demonstrate that CAPS consistently outperforms existing methods, establishing a strong wrapper-based baseline for OSRL.
Yassine Chemingui, Aryan Deshwal, Honghao Wei, Alan Fern, Janardhan Rao Doppa
AAAI1
2025 Online Optimization for Offline Safe Reinforcement Learning
abstract
We study the problem of Offline Safe Reinforcement Learning (OSRL), where the goal is to learn a reward-maximizing policy from fixed data under a cumulative cost constraint. We propose a novel OSRL approach that frames the problem as a minimax objective and solves it by combining offline RL with online optimization algorithms. We prove the approximate optimality of this approach when integrated with an approximate offline RL oracle and no-regret online optimization. We also present a practical approximation that can be combined with any offline RL algorithm, eliminating the need for offline policy evaluation. Empirical results on the DSRL benchmark demonstrate that our method reliably enforces safety constraints under stringent cost budgets, while achieving high rewards. The code is available at https://github.com/yassineCh/O3SRL.
Yassine Chemingui, Aryan Deshwal, Alan Fern, Thanh Nguyen-Tang, Janardhan Rao Doppa
NeurIPS1
2024 Offline Model-Based Optimization via Policy-Guided Gradient Search
abstract
Offline optimization is an emerging problem in many experimental engineering domains including protein, drug or aircraft design, where online experimentation to collect evaluation data is too expensive or dangerous. To avoid that, one has to optimize an unknown function given only its offline evaluation at a fixed set of inputs. A naive solution to this problem is to learn a surrogate model of the unknown function and optimize this surrogate instead. However, such a naive optimizer is prone to erroneous overestimation of the surrogate (possibly due to over-fitting on a biased sample of function evaluation) on inputs outside the offline dataset. Prior approaches addressing this challenge have primarily focused on learning robust surrogate models. However, their search strategies are derived from the surrogate model rather than the actual offline data. To fill this important gap, we introduce a new learning-to-search perspective for offline optimization by reformulating it as an offline reinforcement learning problem. Our proposed policy-guided gradient search approach explicitly learns the best policy for a given surrogate model created from the offline data. Our empirical results on multiple benchmarks demonstrate that the learned optimization policy can be combined with existing offline surrogates to significantly improve the optimization performance.
Yassine Chemingui, Aryan Deshwal, Trong Nghia Hoang, Janardhan Rao Doppa
AAAI1