Omer Gottesman

dblp:217/4338 · DBLP profile ↗
← Back
17ranked-venue papers
4as first author
11since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 4 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Geometry of Neural Reinforcement Learning in Continuous State and Action Spaces
abstract
Advances in reinforcement learning (RL) have led to its successful application in complex tasks with continuous state and action spaces. Despite these advances in practice, most theoretical work pertains to finite state and action spaces. We propose building a theoretical understanding of continuous state and action spaces by employing a geometric lens to understand the locally attained set of states. The set of all parametrised policies learnt through a semi-gradient based approach induce a set of attainable states in RL. We show that training dynamics of a two layer neural policy induce a low dimensional manifold of attainable states embedded in the high-dimensional nominal state space trained using an actor-critic algorithm. We prove that, under certain conditions, the dimensionality of this manifold is of the order of the dimensionality of the action space. This is the first result of its kind, linking the geometry of the state space to the dimensionality of the action space. We empirically corroborate this upper bound for four MuJoCo environments and also demonstrate the results in a toy environment with varying dimensionality. We also show the applicability of this theoretical result by introducing a local manifold learning layer to the policy and value function networks to improve the performance in control environments with very high degrees of freedom by changing one layer of the neural network to learn sparse representations.
Saket Tiwari, Omer Gottesman, George Dimitri Konidaris
ICLR2
2024 Mitigating Partial Observability in Sequential Decision Processes via the Lambda Discrepancy
abstract
Reinforcement learning algorithms typically rely on the assumption that the environment dynamics and value function can be expressed in terms of a Markovian state representation. However, when state information is only partially observable, how can an agent learn such a state representation, and how can it detect when it has found one? We introduce a metric that can accomplish both objectives, without requiring access to---or knowledge of---an underlying, unobservable state space. Our metric, the λ-discrepancy, is the difference between two distinct temporal difference (TD) value estimates, each computed using TD(λ) with a different value of λ. Since TD(λ=0) makes an implicit Markov assumption and TD(λ=1) does not, a discrepancy between these estimates is a potential indicator of a non-Markovian state representation. Indeed, we prove that the λ-discrepancy is exactly zero for all Markov decision processes and almost always non-zero for a broad class of partially observable environments. We also demonstrate empirically that, once detected, minimizing the λ-discrepancy can help with learning a memory function to mitigate the corresponding partial observability. We then train a reinforcement learning agent that simultaneously constructs two recurrent value networks with different λ parameters and minimizes the difference between them as an auxiliary loss. The approach scales to challenging partially observable domains, where the resulting agent frequently performs significantly better (and never performs worse) than a baseline recurrent agent with only a single value network.
Cameron Allen, Aaron Kirtland, Ruo Yu Tao, Sam Lobel, Daniel Scott, Nicholas Petrocelli, Omer Gottesman, Ronald Parr, Michael L. Littman, George Dimitri Konidaris
NeurIPS7
2023 Coarse-Grained Smoothness for Reinforcement Learning in Metric Spaces
abstract
Principled decision-making in continuous state–action spaces is impossible without some assumptions. A common approach is to assume Lipschitz continuity of the Q-function. We show that, unfortunately, this property fails to hold in many typical domains. We propose a new coarse-grained smoothness definition that generalizes the notion of Lipschitz continuity, is more widely applicable, and allows us to compute significantly tighter bounds on Q-functions, leading to improved learning. We provide a theoretical analysis of our new smoothness definition, and discuss its implications and impact on control and exploration in continuous domains.
Omer Gottesman, Kavosh Asadi, Cameron Allen, Sam Lobel, George Dimitri Konidaris, Michael L. Littman
AISTATS1
2023 Performance Bounds for Model and Policy Transfer in Hidden-parameter MDPs
Haotian Fu, Jiayu Yao, Omer Gottesman, Finale Doshi-Velez, George Dimitri Konidaris
ICLR3
2023 TD Convergence: An Optimization Perspective
abstract
We study the convergence behavior of the celebrated temporal-difference (TD) learning algorithm. By looking at the algorithm through the lens of optimization, we first argue that TD can be viewed as an iterative optimization algorithm where the function to be minimized changes per iteration. By carefully investigating the divergence displayed by TD on a classical counter example, we identify two forces that determine the convergent or divergent behavior of the algorithm. We next formalize our discovery in the linear TD setting with quadratic loss and prove that convergence of TD hinges on the interplay between these two forces. We extend this optimization perspective to prove convergence of TD in a much broader setting than just linear approximation and squared loss. Our results provide a theoretical explanation for the successful application of TD in reinforcement learning.
Kavosh Asadi, Shoham Sabach, Yao Liu 0009, Omer Gottesman, Rasool Fakoor
NeurIPS4
2023 Effectively Learning Initiation Sets in Hierarchical Reinforcement Learning
abstract
An agent learning an option in hierarchical reinforcement learning must solve three problems: identify the option's subgoal (termination condition), learn a policy, and learn where that policy will succeed (initiation set). The termination condition is typically identified first, but the option policy and initiation set must be learned simultaneously, which is challenging because the initiation set depends on the option policy, which changes as the agent learns. Consequently, data obtained from option execution becomes invalid over time, leading to an inaccurate initiation set that subsequently harms downstream task performance. We highlight three issues---data non-stationarity, temporal credit assignment, and pessimism---specific to learning initiation sets, and propose to address them using tools from off-policy value estimation and classification. We show that our method learns higher-quality initiation sets faster than existing methods (in MiniGrid and Montezuma's Revenge), can automatically discover promising grasps for robot manipulation (in Robosuite), and improves the performance of a state-of-the-art option discovery method in a challenging maze navigation task in MuJoCo.
Akhil Bagaria, Ben Abbatematteo, Omer Gottesman, Matthew Corsaro, Sreehari Rammohan, George Dimitri Konidaris
NeurIPS3
2023 iCVS - Inferring Cardio-Vascular hidden States from physiological signals available at the bedside
abstract
Intensive care medicine is complex and resource-demanding. A critical and common challenge lies in inferring the underlying physiological state of a patient from partially observed data. Specifically for the cardiovascular system, clinicians use observables such as heart rate, arterial and venous blood pressures, as well as findings from the physical examination and ancillary tests to formulate a mental model and estimate hidden variables such as cardiac output, vascular resistance, filling pressures and volumes, and autonomic tone. Then, they use this mental model to derive the causes for instability and choose appropriate interventions. Not only this is a very hard problem due to the nature of the signals, but it also requires expertise and a clinician's ongoing presence at the bedside. Clinical decision support tools based on mechanistic dynamical models offer an appealing solution due to their inherent explainability, corollaries to the clinical mental process, and predictive power. With a translational motivation in mind, we developed iCVS: a simple, with high explanatory power, dynamical mechanistic model to infer hidden cardiovascular states. Full model estimation requires no prior assumptions on physiological parameters except age and weight, and the only inputs are arterial and venous pressure waveforms. iCVS also considers autonomic and non-autonomic modulations. To gain more information without increasing model complexity, both slow and fast timescales of the blood pressure traces are exploited, while the main inference and dynamic evolution are at the longer, clinically relevant, timescale of minutes. iCVS is designed to allow bedside deployment at pediatric and adult intensive care units and for retrospective investigation of cardiovascular mechanisms underlying instability. In this paper, we describe iCVS and inference system in detail, and using a dataset of critically-ill children, we provide initial indications to its ability to identify bleeding, distributive states, and cardiac dysfunction, in isolation and in combination.
Neta Ravid Tannenbaum, Omer Gottesman, Azadeh Assadi, Mjaye Mazwi, Uri Shalit, Danny Eytan
PLoS Comput. Biol.2
2022 Optimistic Initialization for Exploration in Continuous Control
abstract
Optimistic initialization underpins many theoretically sound exploration schemes in tabular domains; however, in the deep function approximation setting, optimism can quickly disappear if initialized naively. We propose a framework for more effectively incorporating optimistic initialization into reinforcement learning for continuous control. Our approach uses metric information about the state-action space to estimate which transitions are still unexplored, and explicitly maintains the initial Q-value optimism for the corresponding state-action pairs. We also develop methods for efficiently approximating these training objectives, and for incorporating domain knowledge into the optimistic envelope to improve sample efficiency. We empirically evaluate these approaches on a variety of hard exploration problems in continuous control, where our method outperforms existing exploration techniques.
Sam Lobel, Omer Gottesman, Cameron Allen, Akhil Bagaria, George Dimitri Konidaris
AAAI2
2022 Faster Deep Reinforcement Learning with Slower Online Network
abstract
Deep reinforcement learning algorithms often use two networks for value function optimization: an online network, and a target network that tracks the online network with some delay. Using two separate networks enables the agent to hedge against issues that arise when performing bootstrapping. In this paper we endow two popular deep reinforcement learning algorithms, namely DQN and Rainbow, with updates that incentivize the online network to remain in the proximity of the target network. This improves the robustness of deep reinforcement learning in presence of noisy updates. The resultant agents, called DQN Pro and Rainbow Pro, exhibit significant performance improvements over their original counterparts on the Atari benchmark demonstrating the effectiveness of this simple idea in deep reinforcement learning. The code for our paper is available here: Github.com/amazon-research/fast-rl-with-slow-updates.
Kavosh Asadi, Rasool Fakoor, Omer Gottesman, Taesup Kim, Michael L. Littman, Alexander J. Smola
NeurIPS3
2021 State Relevance for Off-Policy Evaluation
abstract
Importance sampling-based estimators for off-policy evaluation (OPE) are valued for their simplicity, unbiasedness, and reliance on relatively few assumptions. However, the variance of these estimators is often high, especially when trajectories are of different lengths. In this work, we introduce Omitting-States-Irrelevant-to-Return Importance Sampling (OSIRIS), an estimator which reduces variance by strategically omitting likelihood ratios associated with certain states. We formalize the conditions under which OSIRIS is unbiased and has lower variance than ordinary importance sampling, and we demonstrate these properties empirically.
Simon P. Shen, Yecheng Jason Ma 0001, Omer Gottesman, Finale Doshi-Velez
ICML3
2021 Learning Markov State Abstractions for Deep Reinforcement Learning
abstract
A fundamental assumption of reinforcement learning in Markov decision processes (MDPs) is that the relevant decision process is, in fact, Markov. However, when MDPs have rich observations, agents typically learn by way of an abstract state representation, and such representations are not guaranteed to preserve the Markov property. We introduce a novel set of conditions and prove that they are sufficient for learning a Markov abstract state representation. We then describe a practical training procedure that combines inverse model estimation and temporal contrastive learning to learn an abstraction that approximately satisfies these conditions. Our novel training objective is compatible with both online and offline training: it does not require a reward signal, but agents can capitalize on reward information when available. We empirically evaluate our approach on a visual gridworld domain and a set of continuous control benchmarks. Our approach learns representations that capture the underlying structure of the domain and lead to improved sample efficiency over state-of-the-art deep reinforcement learning with visual features---often matching or exceeding the performance achieved with hand-designed compact state information.
Cameron Allen, Neev Parikh, Omer Gottesman, George Dimitri Konidaris
NeurIPS3
2020 Interpretable Off-Policy Evaluation in Reinforcement Learning by Highlighting Influential Transitions
abstract
Off-policy evaluation in reinforcement learning offers the chance of using observational data to improve future outcomes in domains such as healthcare and education, but safe deployment in high stakes settings requires ways of assessing its validity. Traditional measures such as confidence intervals may be insufficient due to noise, limited data and confounding. In this paper we develop a method that could serve as a hybrid human-AI system, to enable human experts to analyze the validity of policy evaluation estimates. This is accomplished by highlighting observations in the data whose removal will have a large effect on the OPE estimate, and formulating a set of rules for choosing which ones to present to domain experts for validation. We develop methods to compute exactly the influence functions for fitted Q-evaluation with two different function classes: kernel-based and linear least squares, as well as importance sampling methods. Experiments on medical simulations and real-world intensive care unit data demonstrate that our method can be used to identify limitations in the evaluation process and make evaluation more robust.
Omer Gottesman, Joseph Futoma, Yao Liu 0009, Sonali Parbhoo, Leo A. Celi, Emma Brunskill, Finale Doshi-Velez
ICML1
2020 Learning to search efficiently for causally near-optimal treatments
abstract
Finding an effective medical treatment often requires a search by trial and error. Making this search more efficient by minimizing the number of unnecessary trials could lower both costs and patient suffering. We formalize this problem as learning a policy for finding a near-optimal treatment in a minimum number of trials using a causal inference framework. We give a model-based dynamic programming algorithm which learns from observational data while being robust to unmeasured confounding. To reduce time complexity, we suggest a greedy algorithm which bounds the near-optimality constraint. The methods are evaluated on synthetic and real-world healthcare data and compared to model-free reinforcement learning. We find that our methods compare favorably to the model-free baseline while offering a more transparent trade-off between search time and treatment efficacy.
Samuel Håkansson, Viktor Lindblom, Omer Gottesman, Fredrik D. Johansson
NeurIPS3
2019 Combining parametric and nonparametric models for off-policy evaluation
abstract
We consider a model-based approach to perform batch off-policy evaluation in reinforcement learning. Our method takes a mixture-of-experts approach to combine parametric and non-parametric models of the environment such that the final value estimate has the least expected error. We do so by first estimating the local accuracy of each model and then using a planner to select which model to use at every time step as to minimize the return error estimate along entire trajectories. Across a variety of domains, our mixture-based approach outperforms the individual models alone as well as state-of-the-art importance sampling-based estimators.
Omer Gottesman, Yao Liu 0009, Scott Sussex, Emma Brunskill, Finale Doshi-Velez
ICML1
2018 Weighted Tensor Decomposition for Learning Latent Variables with Partial Data
abstract
Tensor decomposition methods are popular tools for learning latent variables given only lowerorder moments of the data. However, the standard assumption is that we have sufficient data to estimate these moments to high accuracy. In this work, we consider the case in which certain dimensions of the data are not always observed–common in applied settings, where not all measurements may be taken for all observations–resulting in moment estimates of varying quality. We derive a weighted tensor decomposition approach that is computationally as efficient as the non-weighted approach, and demonstrate that it outperforms methods that do not appropriately leverage these less-observed dimensions.
Omer Gottesman, Finale Doshi-Velez
AISTATS1
2018 Improving Sepsis Treatment Strategies by Combining Deep and Kernel-Based Reinforcement Learning
Xuefeng Peng, David Wihl, Omer Gottesman, Matthieu Komorowski, Li-Wei H. Lehman, Andrew Slavin Ross, A. Aldo Faisal, Finale Doshi-Velez
AMIA4
2018 Representation Balancing MDPs for Off-policy Policy Evaluation
abstract
We study the problem of off-policy policy evaluation (OPPE) in RL. In contrast to prior work, we consider how to estimate both the individual policy value and average policy value accurately. We draw inspiration from recent work in causal reasoning, and propose a new finite sample generalization error bound for value estimates from MDP models. Using this upper bound as an objective, we develop a learning algorithm of an MDP model with a balanced representation, and show that our approach can yield substantially lower MSE in common synthetic benchmarks and a HIV treatment simulation domain.
Yao Liu 0009, Omer Gottesman, Aniruddh Raghu, Matthieu Komorowski, A. Aldo Faisal, Finale Doshi-Velez, Emma Brunskill
NeurIPS2