EDBT 2026 Demo / reviewers in the wild / expert
Yarden As
dblp:312/4578
· DBLP profile ↗
9ranked-venue papers
3as first author
9since 2021 · last 2025
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Reinforcement learning · 54% Optimization for machine learning · 14% Transfer learning and domain adaptation · 8% | |
| Theoretical computer science
1 paper |
Mathematical optimization · 75% Distributed computing theory · 25% |
Topics — the 21 heaviest of 23, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
safe reinforcement learning |
3.1 | 4 | 2025 | SPiDR: A Simple Approach for Zero-Shot Safety in Sim-to-Real Transfer · NeurIPS 2025 ActSafe: Active Exploration with Safety Constraints for Reinforcement Learning · ICLR 2025 Log Barriers for Safe Black-box Optimization with Application to Safe Reinforcement Learning · J. Mach. Learn. Res. 2024 |
Machine learning › Reinforcement learning
model-based reinforcement learning |
2.2 | 3 | 2025 | ActSafe: Active Exploration with Safety Constraints for Reinforcement Learning · ICLR 2025 When to Sense and Control? A Time-adaptive Approach for Continuous-Time RL · NeurIPS 2024 Constrained Policy Optimization via Bayesian World Models · ICLR 2022 |
Machine learning › Reinforcement learning › exploration › exploration strategies
constrained exploration |
0.9 | 1 | 2025 | ActSafe: Active Exploration with Safety Constraints for Reinforcement Learning · ICLR 2025 |
Machine learning › Trustworthy machine learning
interpretability |
0.9 | 1 | 2025 | Learning Safety Constraints for Large Language Models · ICML 2025 |
Natural language and speech › Language models and text generation
large language model safety |
0.9 | 1 | 2025 | Learning Safety Constraints for Large Language Models · ICML 2025 |
Machine learning › Reinforcement learning › exploration
optimistic exploration |
0.9 | 1 | 2025 | ActSafe: Active Exploration with Safety Constraints for Reinforcement Learning · ICLR 2025 |
Machine learning › Representation and self-supervised learning › representation analysis
representation space analysis |
0.9 | 1 | 2025 | Learning Safety Constraints for Large Language Models · ICML 2025 |
Machine learning › Transfer learning and domain adaptation
sim-to-real transfer |
0.9 | 1 | 2025 | SPiDR: A Simple Approach for Zero-Shot Safety in Sim-to-Real Transfer · NeurIPS 2025 |
Mathematical optimization › distributed optimization
communication-efficient optimization |
0.9 | 1 | 2025 | Safe-EF: Error Feedback for Non-smooth Constrained Optimization · ICML 2025 |
Mathematical optimization
constrained optimization |
0.9 | 1 | 2025 | Safe-EF: Error Feedback for Non-smooth Constrained Optimization · ICML 2025 |
Mathematical optimization
distributed optimization |
0.9 | 1 | 2025 | Safe-EF: Error Feedback for Non-smooth Constrained Optimization · ICML 2025 |
Distributed computing theory › distributed learning
federated learning |
0.9 | 1 | 2025 | Safe-EF: Error Feedback for Non-smooth Constrained Optimization · ICML 2025 |
Machine learning › Efficient and distributed learning
active learning |
0.8 | 1 | 2024 | Transductive Active Learning: Theory and Applications · NeurIPS 2024 |
Machine learning › Optimization for machine learning › model-based optimization
bayesian optimization |
0.8 | 1 | 2024 | Transductive Active Learning: Theory and Applications · NeurIPS 2024 |
Machine learning › Optimization for machine learning › model-based optimization › bayesian optimization
safe bayesian optimization |
0.8 | 1 | 2024 | Transductive Active Learning: Theory and Applications · NeurIPS 2024 |
Machine learning › Reinforcement learning › safe reinforcement learning
constrained policy optimization |
0.6 | 1 | 2022 | Constrained Policy Optimization via Bayesian World Models · ICLR 2022 |
Machine learning › Reinforcement learning › model-based reinforcement learning
world model |
0.6 | 1 | 2022 | Constrained Policy Optimization via Bayesian World Models · ICLR 2022 |
Natural language and speech › Language models and text generation
large language model |
0.3 | 1 | 2025 | Learning Safety Constraints for Large Language Models · ICML 2025 |
Machine learning › Trustworthy machine learning
robustness |
0.3 | 1 | 2025 | SPiDR: A Simple Approach for Zero-Shot Safety in Sim-to-Real Transfer · NeurIPS 2025 |
Machine learning › Transfer learning and domain adaptation › fine-tuning
active finetuning |
0.2 | 1 | 2024 | Transductive Active Learning: Theory and Applications · NeurIPS 2024 |
Machine learning › Transfer learning and domain adaptation
fine-tuning |
0.2 | 1 | 2024 | Transductive Active Learning: Theory and Applications · NeurIPS 2024 |
Methods — techniques the papers use, named apart from their topics
uncertainty quantification · 0.9safety polytope · 0.9representation engineering · 0.9model-based planning · 0.9lower complexity bounds · 0.9geometric steering · 0.9error feedback · 0.9domain randomization · 0.9contractive compression · 0.9constrained optimization · 0.9uncertainty minimization · 0.8extended MDP formulation · 0.8decision rules · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ActSafe: Active Exploration with Safety Constraints for Reinforcement LearningabstractReinforcement learning (RL) is ubiquitous in the development of modern AI systems. However, state-of-the-art RL agents require extensive, and potentially
unsafe, interactions with their environments to learn effectively. These limitations
confine RL agents to simulated environments, hindering their ability to learn
directly in real-world settings. In this work, we present ActSafe, a novel
model-based RL algorithm for safe and efficient exploration. ActSafe learns
a well-calibrated probabilistic model of the system and plans optimistically
w.r.t. the epistemic uncertainty about the unknown dynamics, while enforcing
pessimism w.r.t. the safety constraints. Under regularity assumptions on the
constraints and dynamics, we show that ActSafe guarantees safety during
learning while also obtaining a near-optimal policy in finite time. In addition, we
propose a practical variant of ActSafe that builds on latest model-based RL advancements and enables safe exploration even in high-dimensional settings such
as visual control. We empirically show that ActSafe obtains state-of-the-art
performance in difficult exploration tasks on standard safe deep RL benchmarks
while ensuring safety during learning. Yarden As, Bhavya Sukhija, Lenart Treven, Carmelo Sferrazza, Stelian Coros, Andreas Krause 0001 |
ICLR | 1 |
| 2025 | Learning Safety Constraints for Large Language ModelsabstractLarge language models (LLMs) have emerged as powerful tools but pose significant safety risks through harmful outputs and vulnerability to adversarial attacks. We propose SaP–short for Safety Polytope–a geometric approach to LLM safety, that learns and enforces multiple safety constraints directly in the model's representation space. We develop a framework that identifies safe and unsafe regions via the polytope's facets, enabling both detection and correction of unsafe outputs through geometric steering. Unlike existing approaches that modify model weights, SaP operates post-hoc in the representation space, preserving model capabilities while enforcing safety constraints. Experiments across multiple LLMs demonstrate that our method can effectively detect unethical inputs, reduce adversarial attack success rates while maintaining performance on standard tasks, thus highlighting the importance of having an explicit geometric model for safety. Analysis of the learned polytope facets reveals emergence of specialization in detecting different semantic notions of safety, providing interpretable insights into how safety is captured in LLMs' representation space. Yarden As, Andreas Krause 0001 |
ICML | 2 |
| 2025 | Safe-EF: Error Feedback for Non-smooth Constrained OptimizationabstractFederated learning faces severe communication bottlenecks due to the high dimensionality of model updates. Communication compression with contractive compressors (e.g., Top-$K$) is often preferable in practice but can degrade performance without proper handling. Error feedback (EF) mitigates such issues but has been largely restricted for smooth, unconstrained problems, limiting its real-world applicability where non-smooth objectives and safety constraints are critical. We advance our understanding of EF in the canonical non-smooth convex setting by establishing new lower complexity bounds for first-order algorithms with contractive compression. Next, we propose Safe-EF, a novel algorithm that matches our lower bound (up to a constant) while enforcing safety constraints essential for practical applications. Extending our approach to the stochastic setting, we bridge the gap between theory and practical implementation. Extensive experiments in a reinforcement learning setup, simulating distributed humanoid robot training, validate the effectiveness of Safe-EF in ensuring safety and reducing communication complexity. Rustem Islamov, Yarden As, Ilyas Fatkhullin |
ICML | 2 |
| 2025 | SPiDR: A Simple Approach for Zero-Shot Safety in Sim-to-Real TransferabstractDeploying reinforcement learning (RL) safely in the real world is challenging, as policies trained in simulators must face the inevitable *sim-to-real gap*. Robust safe RL techniques are provably safe, however difficult to scale, while domain randomization is more practical yet prone to unsafe behaviors. We address this gap by proposing SPiDR, short for Sim-to-real via Pessimistic Domain Randomization—a scalable algorithm with provable guarantees for safe sim-to-real transfer. SPiDR uses domain randomization to incorporate the uncertainty about the sim-to-real gap into the safety constraints, making it versatile and highly compatible with existing training pipelines. Through extensive experiments on sim-to-sim benchmarks and two distinct real-world robotic platforms, we demonstrate that SPiDR effectively ensures safety despite the sim-to-real gap while maintaining strong performance. Yarden As, Chengrui Qu, Benjamin Unger, Dongho Kang, Max van der Hart, Laixi Shi, Stelian Coros, Adam Wierman, Andreas Krause 0001 |
NeurIPS | 1 |
| 2025 | SafeRPlan: Safe deep reinforcement learning for intraoperative planning of pedicle screw placementabstractSpinal fusion surgery requires highly accurate implantation of pedicle screw implants, which must be conducted in critical proximity to vital structures with a limited view of the anatomy. Robotic surgery systems have been proposed to improve placement accuracy. Despite remarkable advances, current robotic systems still lack advanced mechanisms for continuous updating of surgical plans during procedures, which hinders attaining higher levels of robotic autonomy. These systems adhere to conventional rigid registration concepts, relying on the alignment of preoperative planning to the intraoperative anatomy. In this paper, we propose a safe deep reinforcement learning (DRL) planning approach (SafeRPlan) for robotic spine surgery that leverages intraoperative observation for continuous path planning of pedicle screw placement. The main contributions of our method are (1) the capability to ensure safe actions by introducing an uncertainty-aware distance-based safety filter; (2) the ability to compensate for incomplete intraoperative anatomical information, by encoding a-priori knowledge of anatomical structures with neural networks pre-trained on pre-operative images; and (3) the capability to generalize over unseen observation noise thanks to the novel domain randomization techniques. Planning quality was assessed by quantitative comparison with the baseline approaches, gold standard (GS) and qualitative evaluation by expert surgeons. In experiments with human model datasets, our approach was capable of achieving over 5% higher safety rates compared to baseline approaches, even under realistic observation noise. To the best of our knowledge, SafeRPlan is the first safety-aware DRL planning approach specifically designed for robotic spine surgery. Yunke Ao, Hooman Esfandiari, Fabio Carrillo, Christoph J. Laux, Yarden As, Ruixuan Li 0003, Kaat Van Assche, Ayoob Davoodi, Nicola Cavalcanti, Mazda Farshad, Benjamin F. Grewe, Emmanuel B. Vander Poorten, Andreas Krause 0001, Philipp Fürnstahl |
Medical Image Anal. | 5 |
| 2024 | Transductive Active Learning: Theory and ApplicationsabstractWe study a generalization of classical active learning to real-world settings with concrete prediction targets where sampling is restricted to an accessible region of the domain, while prediction targets may lie outside this region.
We analyze a family of decision rules that sample adaptively to minimize uncertainty about prediction targets.
We are the first to show, under general regularity assumptions, that such decision rules converge uniformly to the smallest possible uncertainty obtainable from the accessible data.
We demonstrate their strong sample efficiency in two key applications: active fine-tuning of large neural networks and safe Bayesian optimization, where they achieve state-of-the-art performance. Jonas Hübotter, Bhavya Sukhija, Lenart Treven, Yarden As, Andreas Krause 0001 |
NeurIPS | 4 |
| 2024 | When to Sense and Control? A Time-adaptive Approach for Continuous-Time RLabstractReinforcement learning (RL) excels in optimizing policies for discrete-time Markov decision processes (MDP). However, various systems are inherently continuous in time, making discrete-time MDPs an inexact modeling choice.
In many applications, such as greenhouse control or medical treatments, each interaction (measurement or switching of action) involves manual intervention and thus is inherently costly. Therefore,
we generally prefer a time-adaptive approach with fewer interactions with the system.
In this work, we formalize an RL framework,
**T**ime-**a**daptive **Co**ntrol \& **S**ensing (**TaCoS**), that tackles this challenge by optimizing over policies that besides control predict the duration of its application. Our formulation results in an extended MDP that any standard RL algorithm can solve.
We demonstrate that state-of-the-art RL algorithms trained on TaCoS drastically reduce the interaction amount over their discrete-time counterpart while retaining the same or improved performance, and exhibiting robustness over discretization frequency.
Finally, we propose OTaCoS, an efficient model-based algorithm for our setting. We show that OTaCoS enjoys sublinear regret for systems with sufficiently smooth dynamics and empirically results in further sample-efficiency gains. Lenart Treven, Bhavya Sukhija, Yarden As, Florian Dörfler, Andreas Krause 0001 |
NeurIPS | 3 |
| 2024 | Log Barriers for Safe Black-box Optimization with Application to Safe Reinforcement LearningabstractOptimizing noisy functions online, when evaluating the objective requires experiments on a deployed system, is a crucial task arising in manufacturing, robotics and various other domains. Often, constraints on safe inputs are unknown ahead of time, and we only obtain noisy information, indicating how close we are to violating the constraints. Yet, safety must be guaranteed at all times, not only for the final output of the algorithm. We introduce a general approach for seeking a stationary point in high dimensional non-linear stochastic optimization problems in which maintaining safety during learning is crucial. Our approach called LB-SGD, is based on applying stochastic gradient descent (SGD) with a carefully chosen adaptive step size to a logarithmic barrier approximation of the original problem. We provide a complete convergence analysis of non-convex, convex, and strongly-convex smooth constrained problems, with first-order and zeroth-order feedback. Our approach yields efficient updates and scales better with dimensionality compared to existing approaches. We empirically compare the sample complexity and the computational cost of our method with existing safe learning approaches. Beyond synthetic benchmarks, we demonstrate the effectiveness of our approach on minimizing constraint violation in policy search tasks in safe reinforcement learning (RL). Ilnura Usmanova, Yarden As, Maryam Kamgarpour, Andreas Krause 0001 |
J. Mach. Learn. Res. | 2 |
| 2022 | Constrained Policy Optimization via Bayesian World Models
Yarden As, Ilnura Usmanova, Sebastian Curi, Andreas Krause 0001 |
ICLR | 1 |