Ray Jiang

dblp:217/3543 · DBLP profile ↗
← Back
8ranked-venue papers
5as first author
3since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 5 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Reinforcement learning · 79% Trustworthy machine learning · 8% Generative modeling · 7%
Databases, data mining, and information retrieval
2 papers
Recommender systems · 100%
Theoretical computer science
1 paper
Mathematical optimization · 100%

Topics — the 12 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
off-policy reinforcement learning
1.122022
Learning Expected Emphatic Traces for Deep RL · AAAI 2022
Emphatic Algorithms for Deep Reinforcement Learning · ICML 2021
Machine learning › Reinforcement learning
temporal difference learning
1.122022
Learning Expected Emphatic Traces for Deep RL · AAAI 2022
Emphatic Algorithms for Deep Reinforcement Learning · ICML 2021
Machine learning › Reinforcement learning
deep reinforcement learning
0.822023
Human-level Atari 200x faster · ICLR 2023
Emphatic Algorithms for Deep Reinforcement Learning · ICML 2021
Machine learning › Reinforcement learning › sample efficiency
sample-efficient reinforcement learning
0.712023
Human-level Atari 200x faster · ICLR 2023
Machine learning › Reinforcement learning › reinforcement learning theory
deadly triad
0.512021
Emphatic Algorithms for Deep Reinforcement Learning · ICML 2021
Machine learning › Trustworthy machine learning
fairness
0.412020
A General Approach to Fairness with Optimal Transport · AAAI 2020
Mathematical optimization
optimal transport
0.412020
A General Approach to Fairness with Optimal Transport · AAAI 2020
Machine learning › Generative modeling › variational autoencoder
conditional variational autoencoder
0.412019
Beyond Greedy Ranking: Slate Optimization via List-CVAE · ICLR (Poster) 2019
Machine learning › Learning theory
online learning
0.412019
Learning from Delayed Outcomes via Proxies with Applications to Recommender Systems · ICML 2019
Recommender systems › interactive recommendation
slate recommendation
0.412019
Beyond Greedy Ranking: Slate Optimization via List-CVAE · ICLR (Poster) 2019
Machine learning › Reinforcement learning › off-policy reinforcement learning
experience replay
0.212022
Learning Expected Emphatic Traces for Deep RL · AAAI 2022
Machine learning › Reinforcement learning › deep reinforcement learning
atari game playing
0.112021
Emphatic Algorithms for Deep Reinforcement Learning · ICML 2021

Methods — techniques the papers use, named apart from their topics

reinforcement learning · 1.3optimal transport · 0.9distribution transport · 0.9list-CVAE · 0.8factorization · 0.8time-reversed n-step TD learning · 0.6emphatic weighting · 0.6multi-step returns · 0.5emphatic temporal difference · 0.5ETD(lambda) · 0.5neural network · 0.4
YearPublicationVenuePosition
2023 Human-level Atari 200x faster
Steven Kapturowski, Victor Campos 0001, Ray Jiang, Nemanja Rakicevic, Hado van Hasselt, Charles Blundell, Adrià Puigdomènech Badia
ICLR3
2022 Learning Expected Emphatic Traces for Deep RL
abstract
Off-policy sampling and experience replay are key for improving sample efficiency and scaling model-free temporal difference learning methods. When combined with function approximation, such as neural networks, this combination is known as the deadly triad and is potentially unstable. Recently, it has been shown that stability and good performance at scale can be achieved by combining emphatic weightings and multi-step updates. This approach, however, is generally limited to sampling complete trajectories in order, to compute the required emphatic weighting. In this paper we investigate how to combine emphatic weightings with non-sequential, off-line data sampled from a replay buffer. We develop a multi-step emphatic weighting that can be combined with replay, and a time-reversed n-step TD learning algorithm to learn the required emphatic weighting. We show that these state weightings reduce variance compared with prior approaches, while providing convergence guarantees. We tested the approach at scale on Atari 2600 video games, and observed that the new X-ETD(n) agent improved over baseline agents, highlighting both the scalability and broad applicability of our approach.
Ray Jiang, Shangtong Zhang, Veronica Chelu, Adam White 0001, Hado van Hasselt
AAAI1
2021 Emphatic Algorithms for Deep Reinforcement Learning
abstract
Off-policy learning allows us to learn about possible policies of behavior from experience generated by a different behavior policy. Temporal difference (TD) learning algorithms can become unstable when combined with function approximation and off-policy sampling—this is known as the “deadly triad”. Emphatic temporal difference (ETD($\lambda$)) algorithm ensures convergence in the linear case by appropriately weighting the TD($\lambda$) updates. In this paper, we extend the use of emphatic methods to deep reinforcement learning agents. We show that naively adapting ETD($\lambda$) to popular deep reinforcement learning algorithms, which use forward view multi-step returns, results in poor performance. We then derive new emphatic algorithms for use in the context of such algorithms, and we demonstrate that they provide noticeable benefits in small problems designed to highlight the instability of TD methods. Finally, we observed improved performance when applying these algorithms at scale on classic Atari games from the Arcade Learning Environment.
Ray Jiang, Tom Zahavy, Zhongwen Xu, Adam White 0001, Matteo Hessel, Charles Blundell, Hado van Hasselt
ICML1
2020 A General Approach to Fairness with Optimal Transport
abstract
We propose a general approach to fairness based on transporting distributions corresponding to different sensitive attributes to a common distribution. We use optimal transport theory to derive target distributions and methods that allow us to achieve fairness with minimal changes to the unfair model. Our approach is applicable to both classification and regression problems, can enforce different notions of fairness, and enable us to achieve a Pareto-optimal trade-off between accuracy and fairness. We demonstrate that it outperforms previous approaches in several benchmark fairness datasets.
Silvia Chiappa, Ray Jiang, Thomas S. Stepleton, Aldo Pacchiano, Heinrich Jiang, John Aslanides
AAAI2
2019 Degenerate Feedback Loops in Recommender Systems
abstract
Machine learning is used extensively in recommender systems deployed in products. The decisions made by these systems can influence user beliefs and preferences which in turn affect the feedback the learning system receives - thus creating a feedback loop. This phenomenon can give rise to the so-called "echo chambers" or "filter bubbles" that have user and societal implications. In this paper, we provide a novel theoretical analysis that examines both the role of user dynamics and the behavior of recommender systems, disentangling the echo chamber from the filter bubble effect. In addition, we offer practical solutions to slow down system degeneracy. Our study contributes toward understanding and developing solutions to commonly cited issues in the complex temporal scenario, an area that is still largely unexplored.
Ray Jiang, Silvia Chiappa, Tor Lattimore, András György 0001, Pushmeet Kohli
AIES1
2019 Beyond Greedy Ranking: Slate Optimization via List-CVAE
Ray Jiang, Sven Gowal, Yuqiu Qian, Timothy A. Mann, Danilo Jimenez Rezende
ICLR (Poster)1
2019 Learning from Delayed Outcomes via Proxies with Applications to Recommender Systems
abstract
Predicting delayed outcomes is an important problem in recommender systems (e.g., if customers will finish reading an ebook). We formalize the problem as an adversarial, delayed online learning problem and consider how a proxy for the delayed outcome (e.g., if customers read a third of the book in 24 hours) can help minimize regret, even though the proxy is not available when making a prediction. Motivated by our regret analysis, we propose two neural network architectures: Factored Forecaster (FF) which is ideal if the proxy is informative of the outcome in hindsight, and Residual Factored Forecaster (RFF) that is robust to a non-informative proxy. Experiments on two real-world datasets for predicting human behavior show that RFF outperforms both FF and a direct forecaster that does not make use of the proxy. Our results suggest that exploiting proxies by factorization is a promising way to mitigate the impact of long delays in human-behavior prediction tasks.
Timothy A. Mann, Sven Gowal, András György 0001, Huiyi Hu, Ray Jiang, Balaji Lakshminarayanan, Prav Srinivasan
ICML5
2019 Wasserstein Fair Classification
Ray Jiang, Aldo Pacchiano, Thomas S. Stepleton, Heinrich Jiang, Silvia Chiappa
UAI1