EDBT 2026 Demo / reviewers in the wild / expert
Elad Sarafian
dblp:220/4267
· DBLP profile ↗
7ranked-venue papers
3as first author
5since 2021 · last 2026
0000-0002-3271-6308ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Reinforcement learning · 52% Optimization for machine learning · 32% Generative modeling · 9% | |
| Theoretical computer science
2 papers |
Mathematical optimization · 100% |
Topics — the 15 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Mathematical optimization
black-box optimization |
1.4 | 2 | 2026 | EvoGrad: Evolutionary-Weighted Gradient and Hessian Learning for Black-Box Optimization · AAAI 2026 Explicit Gradient Learning for Black-Box Optimization · ICML 2020 |
Machine learning › Optimization for machine learning
gradient estimation |
1.0 | 1 | 2026 | EvoGrad: Evolutionary-Weighted Gradient and Hessian Learning for Black-Box Optimization · AAAI 2026 |
Machine learning › Optimization for machine learning › second-order optimization
hessian approximation |
1.0 | 1 | 2026 | EvoGrad: Evolutionary-Weighted Gradient and Hessian Learning for Black-Box Optimization · AAAI 2026 |
Machine learning › Reinforcement learning
imitation learning |
0.7 | 1 | 2023 | A Coupled Flow Approach to Imitation Learning · ICML 2023 |
Machine learning › Generative modeling
normalizing flow |
0.7 | 1 | 2023 | A Coupled Flow Approach to Imitation Learning · ICML 2023 |
Machine learning › Reinforcement learning
actor-critic methods |
0.5 | 1 | 2021 | Recomposing the Reinforcement Learning Building Blocks with Hypernetworks · ICML 2021 |
Machine learning › Deep learning architectures and training
hypernetwork |
0.5 | 1 | 2021 | Recomposing the Reinforcement Learning Building Blocks with Hypernetworks · ICML 2021 |
Machine learning › Reinforcement learning
meta-reinforcement learning |
0.5 | 1 | 2021 | Recomposing the Reinforcement Learning Building Blocks with Hypernetworks · ICML 2021 |
Machine learning › Reinforcement learning › value function approximation
q-function approximation |
0.5 | 1 | 2021 | Recomposing the Reinforcement Learning Building Blocks with Hypernetworks · ICML 2021 |
Machine learning › Reinforcement learning › safe reinforcement learning
constrained policy optimization |
0.4 | 1 | 2020 | Constrained Policy Improvement for Efficient Reinforcement Learning · IJCAI 2020 |
Machine learning › Optimization for machine learning
gradient learning |
0.4 | 1 | 2020 | Explicit Gradient Learning for Black-Box Optimization · ICML 2020 |
Machine learning › Reinforcement learning
offline reinforcement learning |
0.4 | 1 | 2020 | Constrained Policy Improvement for Efficient Reinforcement Learning · IJCAI 2020 |
Machine learning › Reinforcement learning
off-policy reinforcement learning |
0.4 | 1 | 2020 | Constrained Policy Improvement for Efficient Reinforcement Learning · IJCAI 2020 |
Machine learning › Reinforcement learning › policy optimization
policy improvement |
0.4 | 1 | 2020 | Constrained Policy Improvement for Efficient Reinforcement Learning · IJCAI 2020 |
Mathematical optimization
gradient estimation |
0.4 | 1 | 2020 | Explicit Gradient Learning for Black-Box Optimization · ICML 2020 |
Methods — techniques the papers use, named apart from their topics
taylor approximation · 2.0hessian estimation · 2.0evolutionary algorithm · 2.0CMA-ES · 2.0convergence analysis · 0.9normalizing flow · 0.7donsker-varadhan representation · 0.7KL divergence · 0.7hypernetwork · 0.5gradient estimation · 0.5neural network · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EvoGrad: Evolutionary-Weighted Gradient and Hessian Learning for Black-Box OptimizationabstractBlack-box algorithms aim to optimize functions without access to their analytical structure or gradient information, making them essential when gradients are unavailable or computationally expensive to obtain. Traditional methods for black-box optimization (BBO) primarily utilize non-parametric models, but these approaches often struggle to scale effectively in large input spaces. Conversely, parametric approaches, which rely on neural estimators and gradient signals via backpropagation, frequently encounter substantial gradient estimation errors, limiting their reliability. Explicit Gradient Learning (EGL), a recent advancement, directly learns gradients using a first-order Taylor approximation and has demonstrated superior performance compared to both parametric and non-parametric methods. However, EGL inherently remains local and myopic, often faltering on highly non-convex optimization landscapes. In this work, we address this limitation by integrating global statistical insights from the evolutionary algorithm CMA-ES into the gradient learning framework, effectively biasing gradient estimates towards regions with higher optimization potential. Moreover, we enhance the gradient learning process by estimating the Hessian matrix, allowing us to correct the second-order residual of the Taylor series approximation. Our proposed algorithm, EvoGrad2 (Evolutionary Gradient Learning with second-order approximation), achieves state-of-the-art results on the synthetic COCO test suite, exhibiting significant advantages in high-dimensional optimization problems. We further demonstrate EvoGrad2's effectiveness on challenging real-world machine learning tasks, including adversarial training and code generation, highlighting its ability to produce more robust, high-quality solutions. Our results underscore EvoGrad2's potential as a powerful tool for researchers and practitioners facing complex, high-dimensional, and non-linear optimization problems. Yedidya Kfir, Elad Sarafian, Yoram Louzoun, Sarit Kraus |
AAAI | 2 |
| 2023 | A Coupled Flow Approach to Imitation LearningabstractIn reinforcement learning and imitation learning, an object of central importance is the state distribution induced by the policy. It plays a crucial role in the policy gradient theorem, and references to it–along with the related state-action distribution–can be found all across the literature. Despite its importance, the state distribution is mostly discussed indirectly and theoretically, rather than being modeled explicitly. The reason being an absence of appropriate density estimation tools. In this work, we investigate applications of a normalizing flow based model for the aforementioned distributions. In particular, we use a pair of flows coupled through the optimality point of the Donsker-Varadhan representation of the Kullback-Leibler (KL) divergence, for distribution matching based imitation learning. Our algorithm, Coupled Flow Imitation Learning (CFIL), achieves state-of-the-art performance on benchmark tasks with a single expert trajectory and extends naturally to a variety of other settings, including the subsampled and state-only regimes. Gideon Freund, Elad Sarafian, Sarit Kraus |
ICML | 2 |
| 2022 | Analyzing and Overcoming Degradation in Warm-Start Reinforcement LearningabstractReinforcement Learning (RL) for robotic applications can benefit from a warm-start where the agent is initialized with a pretrained behavioral policy. However, when transitioning to RL updates, degradation in performance can occur, which may compromise the robot's safety. This degradation, which constitutes an inability to properly utilize the pretrained policy, is attributed to extrapolation error in the value function, a result of high values being assigned to Out-Of-Distribution actions not present in the behavioral policy's data. We investigate why the magnitude of degradation varies across policies and why the policy fails to quickly return to behavioral performance. We present visual confirmation of our analysis and draw comparisons to the Offline RL setting which suffers from similar difficulties. We propose a novel method, Confidence Constrained Learning (CCL) for Warm-Start RL, that reduces degradation by balancing between the policy gradient and constrained learning according to a confidence measure of the Q-values. For the constrained learning component we propose a novel objective, Positive Q-value Distance (CCL-PQD). We investigate a variety of constraint-based methods that aim to overcome the degradation, and find they constitute solutions for a multi-objective optimization problem between maximimal performance and miniminal degradation. Our results demonstrate that hyperparameter tuning for CCL-PQD produces solutions on the Pareto Front of this multi-objective problem, allowing the user to balance between performance and tolerable compromises to the robot's safety. Benjamin Wexler, Elad Sarafian, Sarit Kraus |
IROS | 2 |
| 2021 | Recomposing the Reinforcement Learning Building Blocks with HypernetworksabstractThe Reinforcement Learning (RL) building blocks, i.e. $Q$-functions and policy networks, usually take elements from the cartesian product of two domains as input. In particular, the input of the $Q$-function is both the state and the action, and in multi-task problems (Meta-RL) the policy can take a state and a context. Standard architectures tend to ignore these variables’ underlying interpretations and simply concatenate their features into a single vector. In this work, we argue that this choice may lead to poor gradient estimation in actor-critic algorithms and high variance learning steps in Meta-RL algorithms. To consider the interaction between the input variables, we suggest using a Hypernetwork architecture where a primary network determines the weights of a conditional dynamic network. We show that this approach improves the gradient approximation and reduces the learning step variance, which both accelerates learning and improves the final performance. We demonstrate a consistent improvement across different locomotion tasks and different algorithms both in RL (TD3 and SAC) and in Meta-RL (MAML and PEARL). Elad Sarafian, Shai Keynan, Sarit Kraus |
ICML | 1 |
| 2021 | A Domain Adaptation Approach for Performance Estimation of Spatial PredictionsabstractSpatial predictions, like other supervised learning tasks, require some criterion for a predictor's quality. Typical data-splitting schemes, such as holdouts and k-fold cross-validation, ignore the fact that the training data are usually not available where predictions are being made. The common data-splitting schemes are thus biased estimates of a predictor's performance, which in turn may lead to choosing suboptimal predictors. In this contribution, we borrow ideas from the domain adaptation machine-learning literature, to suggest the importance-weighted source risk (IWSR). IWSR is a principled approach for weighting the prediction risk, which allows the practitioner to explicitly state the target locations for prediction. IWSR essentially consists of down-weighting training locations and up-weighting target locations. We show that, unlike the usual (unweighted) empirical risk, IWSR is an unbiased estimator of the prediction error. Equipped with this risk estimator, we use it to learn a model in the empirical risk minimization framework and to evaluate the existing predictors. We show the superiority of this weighted risk, using both simulated data and an empirical control: air-temperature prediction in France. Ron Sarafian, Itai Kloog, Elad Sarafian, Ian Hough, Jonathan D. Rosenblatt |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2020 | Explicit Gradient Learning for Black-Box OptimizationabstractBlack-Box Optimization (BBO) methods can find optimal policies for systems that interact with complex environments with no analytical representation. As such, they are of interest in many Artificial Intelligence (AI) domains. Yet classical BBO methods fall short in high-dimensional non-convex problems. They are thus often overlooked in real-world AI tasks. Here we present a BBO method, termed Explicit Gradient Learning (EGL), that is designed to optimize high-dimensional ill-behaved functions. We derive EGL by finding weak spots in methods that fit the objective function with a parametric Neural Network (NN) model and obtain the gradient signal by calculating the parametric gradient. Instead of fitting the function, EGL trains a NN to estimate the objective gradient directly. We prove the convergence of EGL to a stationary point and its robustness in the optimization of integrable functions. We evaluate EGL and achieve state-of-the-art results in two challenging problems: (1) the COCO test suite against an assortment of standard BBO methods; and (2) in a high-dimensional non-convex image generation task. Elad Sarafian, Mor Sinay, Yoram Louzoun, Noa Agmon, Sarit Kraus |
ICML | 1 |
| 2020 | Constrained Policy Improvement for Efficient Reinforcement LearningabstractWe propose a policy improvement algorithm for Reinforcement Learning (RL) termed Rerouted Behavior Improvement (RBI). RBI is designed to take into account the evaluation errors of the Q-function. Such errors are common in RL when learning the Q-value from finite experience data. Greedy policies or even constrained policy optimization algorithms that ignore these errors may suffer from an improvement penalty (i.e., a policy impairment). To reduce the penalty, the idea of RBI is to attenuate rapid policy changes to actions that were rarely sampled. This approach is shown to avoid catastrophic performance degradation and reduce regret when learning from a batch of transition samples. Through a two-armed bandit example, we show that it also increases data efficiency when the optimal action has a high variance. We evaluate RBI in two tasks in the Atari Learning Environment: (1) learning from observations of multiple behavior policies and (2) iterative RL. Our results demonstrate the advantage of RBI over greedy policies and other constrained policy optimization algorithms both in learning from observations and in RL tasks. Elad Sarafian, Aviv Tamar, Sarit Kraus |
IJCAI | 1 |