Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Scott M. Jordan

dblp:222/1982 · also Scott Michael Jordan · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
5since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Reinforcement learning · 76% Learning theory · 8% Planning, search and constraint satisfaction · 6%

Topics — the 20 heaviest of 22, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning › temporal difference learning
eligibility traces
0.812024
From Past to Future: Rethinking Eligibility Traces · AAAI 2024
Machine learning › Reinforcement learning
hierarchical reinforcement learning
0.812024
Goal-Space Planning with Subgoal Models · J. Mach. Learn. Res. 2024
Machine learning › Reinforcement learning
model-based reinforcement learning
0.812024
Goal-Space Planning with Subgoal Models · J. Mach. Learn. Res. 2024
Machine learning › Reinforcement learning
policy evaluation
0.812024
From Past to Future: Rethinking Eligibility Traces · AAAI 2024
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › hierarchical planning
subgoal planning
0.812024
Goal-Space Planning with Subgoal Models · J. Mach. Learn. Res. 2024
Natural language and speech › Language models and text generation › alignment
behavioral alignment
0.712023
Behavior Alignment via Reward Function Optimization · NeurIPS 2023
Machine learning › Reinforcement learning
reward design
0.712023
Behavior Alignment via Reward Function Optimization · NeurIPS 2023
Machine learning › Reinforcement learning
reward maximization
0.712023
Behavior Alignment via Reward Function Optimization · NeurIPS 2023
Machine learning › Reinforcement learning › reward design
reward shaping
0.712023
Behavior Alignment via Reward Function Optimization · NeurIPS 2023
Machine learning › Learning theory
generalization bounds
0.512021
High Confidence Generalization for Reinforcement Learning · ICML 2021
Machine learning › Reinforcement learning
safe reinforcement learning
0.512021
High Confidence Generalization for Reinforcement Learning · ICML 2021
Machine learning › Reinforcement learning › markov decision process
non-stationary MDP
0.412020
Towards Safe Policy Improvement for Non-Stationary MDPs · NeurIPS 2020
Machine learning › Reinforcement learning
off-policy evaluation
0.412020
Towards Safe Policy Improvement for Non-Stationary MDPs · NeurIPS 2020
Machine learning › Reinforcement learning › safe reinforcement learning
safe policy improvement
0.412020
Towards Safe Policy Improvement for Non-Stationary MDPs · NeurIPS 2020
Machine learning › Learning theory › hypothesis testing
sequential testing
0.412020
Towards Safe Policy Improvement for Non-Stationary MDPs · NeurIPS 2020
Computer vision › Video understanding and tracking › video representation learning
action representation learning
0.412019
Learning Action Representations for Reinforcement Learning · ICML 2019
Machine learning › Reinforcement learning
generalization in reinforcement learning
0.412019
Learning Action Representations for Reinforcement Learning · ICML 2019
Machine learning › Reinforcement learning
value function
0.212024
From Past to Future: Rethinking Eligibility Traces · AAAI 2024
Machine learning › Reinforcement learning
policy optimization
0.212023
Behavior Alignment via Reward Function Optimization · NeurIPS 2023
Machine learning › Representation and self-supervised learning › representation learning › latent representation learning
state representation learning
0.112019
Learning Action Representations for Reinforcement Learning · ICML 2019

Methods — techniques the papers use, named apart from their topics

performance evaluation · 1.2benchmarking · 1.2temporal difference learning · 0.8temporal abstraction · 0.8subgoal-conditioned model · 0.8dyna · 0.8potential-based reward shaping · 0.7bi-level optimization · 0.7probabilistic safety guarantees · 0.5MDP generalization · 0.5
YearPublicationVenuePosition
2024 From Past to Future: Rethinking Eligibility Traces
abstract
In this paper, we introduce a fresh perspective on the challenges of credit assignment and policy evaluation. First, we delve into the nuances of eligibility traces and explore instances where their updates may result in unexpected credit assignment to preceding states. From this investigation emerges the concept of a novel value function, which we refer to as the ????????????? ????? ????????. Unlike traditional state value functions, bidirectional value functions account for both future expected returns (rewards anticipated from the current state onward) and past expected returns (cumulative rewards from the episode's start to the present). We derive principled update equations to learn this value function and, through experimentation, demonstrate its efficacy in enhancing the process of policy evaluation. In particular, our results indicate that the proposed learning approach can, in certain challenging contexts, perform policy evaluation more rapidly than TD(λ)–a method that learns forward value functions, v^π, ????????. Overall, our findings present a new perspective on eligibility traces and potential advantages associated with the novel value function it inspires, especially for policy evaluation.
Dhawal Gupta, Scott M. Jordan, Shreyas Chaudhari, Bo Liu 0006, Philip S. Thomas, Bruno C. da Silva 0001
AAAI2
2024 Position: Benchmarking is Limited in Reinforcement Learning Research
abstract
Novel reinforcement learning algorithms, or improvements on existing ones, are commonly justified by evaluating their performance on benchmark environments and are compared to an ever-changing set of standard algorithms. However, despite numerous calls for improvements, experimental practices continue to produce misleading or unsupported claims. One reason for the ongoing substandard practices is that conducting rigorous benchmarking experiments requires substantial computational time. This work investigates the sources of increased computation costs in rigorous experiment designs. We show that conducting rigorous performance benchmarks will likely have computational costs that are often prohibitive. As a result, we argue for using an additional experimentation paradigm to overcome the limitations of benchmarking.
Scott M. Jordan, Adam White 0001, Bruno C. da Silva 0001, Martha White, Philip S. Thomas
ICML1
2024 Goal-Space Planning with Subgoal Models
abstract
This paper investigates a new approach to model-based reinforcement learning using background planning: mixing (approximate) dynamic programming updates and model-free updates, similar to the Dyna architecture. Background planning with learned models is often worse than model-free alternatives, such as Double DQN, even though the former uses significantly more memory and computation. The fundamental problem is that learned models can be inaccurate and often generate invalid states, especially when iterated many steps. In this paper, we avoid this limitation by constraining background planning to a given set of (abstract) subgoals and learning only local, subgoal-conditioned models. This goal-space planning (GSP) approach is more computationally efficient, naturally incorporates temporal abstraction for faster long-horizon planning, and avoids learning the transition dynamics entirely. We show that our GSP algorithm can propagate value from an abstract space in a manner that helps a variety of base learners learn significantly faster in different domains.
Chunlok Lo, Kevin Roice, Parham Mohammad Panahi, Scott M. Jordan, Adam White 0001, Gabor Mihucz, Farzane Aminmansour, Martha White
J. Mach. Learn. Res.4
2023 Behavior Alignment via Reward Function Optimization
abstract
Designing reward functions for efficiently guiding reinforcement learning (RL) agents toward specific behaviors is a complex task. This is challenging since it requires the identification of reward structures that are not sparse and that avoid inadvertently inducing undesirable behaviors. Naively modifying the reward structure to offer denser and more frequent feedback can lead to unintended outcomes and promote behaviors that are not aligned with the designer's intended goal. Although potential-based reward shaping is often suggested as a remedy, we systematically investigate settings where deploying it often significantly impairs performance. To address these issues, we introduce a new framework that uses a bi-level objective to learn \emph{behavior alignment reward functions}. These functions integrate auxiliary rewards reflecting a designer's heuristics and domain knowledge with the environment's primary rewards. Our approach automatically determines the most effective way to blend these types of feedback, thereby enhancing robustness against heuristic reward misspecification. Remarkably, it can also adapt an agent's policy optimization process to mitigate suboptimalities resulting from limitations and biases inherent in the underlying RL algorithms. We evaluate our method's efficacy on a diverse set of tasks, from small-scale experiments to high-dimensional control challenges. We investigate heuristic auxiliary rewards of varying quality---some of which are beneficial and others detrimental to the learning process. Our results show that our framework offers a robust and principled way to integrate designer-specified heuristics. It not only addresses key shortcomings of existing approaches but also consistently leads to high-performing solutions, even when given misaligned or poorly-specified auxiliary reward functions.
Dhawal Gupta, Yash Chandak, Scott M. Jordan, Philip S. Thomas, Bruno C. da Silva 0001
NeurIPS3
2021 High Confidence Generalization for Reinforcement Learning
abstract
We present several classes of reinforcement learning algorithms that safely generalize to Markov decision processes (MDPs) not seen during training. Specifically, we study the setting in which some set of MDPs is accessible for training. The goal is to generalize safely to MDPs that are sampled from the same distribution, but which may not be in the set accessible for training. For various definitions of safety, our algorithms give probabilistic guarantees that agents can safely generalize to MDPs that are sampled from the same distribution but are not necessarily in the training set. These algorithms are a type of Seldonian algorithm (Thomas et al., 2019), which is a class of machine learning algorithms that return models with probabilistic safety guarantees for user-specified definitions of safety.
James E. Kostas, Yash Chandak, Scott M. Jordan, Georgios Theocharous, Philip S. Thomas
ICML3
2020 Evaluating the Performance of Reinforcement Learning Algorithms
abstract
Performance evaluations are critical for quantifying algorithmic advances in reinforcement learning. Recent reproducibility analyses have shown that reported performance results are often inconsistent and difficult to replicate. In this work, we argue that the inconsistency of performance stems from the use of flawed evaluation metrics. Taking a step towards ensuring that reported results are consistent, we propose a new comprehensive evaluation methodology for reinforcement learning algorithms that produces reliable measurements of performance both on a single environment and when aggregated across environments. We demonstrate this method by evaluating a broad class of reinforcement learning algorithms on standard benchmark tasks.
Scott M. Jordan, Yash Chandak, Mengxue Zhang, Philip S. Thomas
ICML1
2020 Towards Safe Policy Improvement for Non-Stationary MDPs
abstract
Many real-world sequential decision-making problems involve critical systems with financial risks and human-life risks. While several works in the past have proposed methods that are safe for deployment, they assume that the underlying problem is stationary. However, many real-world problems of interest exhibit non-stationarity, and when stakes are high, the cost associated with a false stationarity assumption may be unacceptable. We take the first steps towards ensuring safety, with high confidence, for smoothly-varying non-stationary decision problems. Our proposed method extends a type of safe algorithm, called a Seldonian algorithm, through a synthesis of model-free reinforcement learning with time-series analysis. Safety is ensured using sequential hypothesis testing of a policy’s forecasted performance, and confidence intervals are obtained using wild bootstrap.
Yash Chandak, Scott M. Jordan, Georgios Theocharous, Martha White, Philip S. Thomas
NeurIPS2
2019 Learning Action Representations for Reinforcement Learning
abstract
Most model-free reinforcement learning methods leverage state representations (embeddings) for generalization, but either ignore structure in the space of actions or assume the structure is provided a priori. We show how a policy can be decomposed into a component that acts in a low-dimensional space of action representations and a component that transforms these representations into actual actions. These representations improve generalization over large, finite action sets by allowing the agent to infer the outcomes of actions similar to actions already taken. We provide an algorithm to both learn and use action representations and provide conditions for its convergence. The efficacy of the proposed method is demonstrated on large-scale real-world problems.
Yash Chandak, Georgios Theocharous, James E. Kostas, Scott M. Jordan, Philip S. Thomas
ICML4