EDBT 2026 Demo / reviewers in the wild / expert
Ray Jiang
dblp:217/3543
· DBLP profile ↗
8ranked-venue papers
5as first author
3since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 5 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Reinforcement learning · 79% Trustworthy machine learning · 8% Generative modeling · 7% | |
| Databases, data mining, and information retrieval
2 papers |
Recommender systems · 100% | |
| Theoretical computer science
1 paper |
Mathematical optimization · 100% |
Topics — the 12 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
off-policy reinforcement learning |
1.1 | 2 | 2022 | Learning Expected Emphatic Traces for Deep RL · AAAI 2022 Emphatic Algorithms for Deep Reinforcement Learning · ICML 2021 |
Machine learning › Reinforcement learning
temporal difference learning |
1.1 | 2 | 2022 | Learning Expected Emphatic Traces for Deep RL · AAAI 2022 Emphatic Algorithms for Deep Reinforcement Learning · ICML 2021 |
Machine learning › Reinforcement learning
deep reinforcement learning |
0.8 | 2 | 2023 | Human-level Atari 200x faster · ICLR 2023 Emphatic Algorithms for Deep Reinforcement Learning · ICML 2021 |
Machine learning › Reinforcement learning › sample efficiency
sample-efficient reinforcement learning |
0.7 | 1 | 2023 | Human-level Atari 200x faster · ICLR 2023 |
Machine learning › Reinforcement learning › reinforcement learning theory
deadly triad |
0.5 | 1 | 2021 | Emphatic Algorithms for Deep Reinforcement Learning · ICML 2021 |
Machine learning › Trustworthy machine learning
fairness |
0.4 | 1 | 2020 | A General Approach to Fairness with Optimal Transport · AAAI 2020 |
Mathematical optimization
optimal transport |
0.4 | 1 | 2020 | A General Approach to Fairness with Optimal Transport · AAAI 2020 |
Machine learning › Generative modeling › variational autoencoder
conditional variational autoencoder |
0.4 | 1 | 2019 | Beyond Greedy Ranking: Slate Optimization via List-CVAE · ICLR (Poster) 2019 |
Machine learning › Learning theory
online learning |
0.4 | 1 | 2019 | Learning from Delayed Outcomes via Proxies with Applications to Recommender Systems · ICML 2019 |
Recommender systems › interactive recommendation
slate recommendation |
0.4 | 1 | 2019 | Beyond Greedy Ranking: Slate Optimization via List-CVAE · ICLR (Poster) 2019 |
Machine learning › Reinforcement learning › off-policy reinforcement learning
experience replay |
0.2 | 1 | 2022 | Learning Expected Emphatic Traces for Deep RL · AAAI 2022 |
Machine learning › Reinforcement learning › deep reinforcement learning
atari game playing |
0.1 | 1 | 2021 | Emphatic Algorithms for Deep Reinforcement Learning · ICML 2021 |
Methods — techniques the papers use, named apart from their topics
reinforcement learning · 1.3optimal transport · 0.9distribution transport · 0.9list-CVAE · 0.8factorization · 0.8time-reversed n-step TD learning · 0.6emphatic weighting · 0.6multi-step returns · 0.5emphatic temporal difference · 0.5ETD(lambda) · 0.5neural network · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Human-level Atari 200x faster
Steven Kapturowski, Victor Campos 0001, Ray Jiang, Nemanja Rakicevic, Hado van Hasselt, Charles Blundell, Adrià Puigdomènech Badia |
ICLR | 3 |
| 2022 | Learning Expected Emphatic Traces for Deep RLabstractOff-policy sampling and experience replay are key for improving sample efficiency and scaling model-free temporal difference learning methods. When combined with function approximation, such as neural networks, this combination is known as the deadly triad and is potentially unstable. Recently, it has been shown that stability and good performance at scale can be achieved by combining emphatic weightings and multi-step updates. This approach, however, is generally limited to sampling complete trajectories in order, to compute the required emphatic weighting. In this paper we investigate how to combine emphatic weightings with non-sequential, off-line data sampled from a replay buffer. We develop a multi-step emphatic weighting that can be combined with replay, and a time-reversed n-step TD learning algorithm to learn the required emphatic weighting. We show that these state weightings reduce variance compared with prior approaches, while providing convergence guarantees. We tested the approach at scale on Atari 2600 video games, and observed that the new X-ETD(n) agent improved over baseline agents, highlighting both the scalability and broad applicability of our approach. Ray Jiang, Shangtong Zhang, Veronica Chelu, Adam White 0001, Hado van Hasselt |
AAAI | 1 |
| 2021 | Emphatic Algorithms for Deep Reinforcement LearningabstractOff-policy learning allows us to learn about possible policies of behavior from experience generated by a different behavior policy. Temporal difference (TD) learning algorithms can become unstable when combined with function approximation and off-policy sampling—this is known as the “deadly triad”. Emphatic temporal difference (ETD($\lambda$)) algorithm ensures convergence in the linear case by appropriately weighting the TD($\lambda$) updates. In this paper, we extend the use of emphatic methods to deep reinforcement learning agents. We show that naively adapting ETD($\lambda$) to popular deep reinforcement learning algorithms, which use forward view multi-step returns, results in poor performance. We then derive new emphatic algorithms for use in the context of such algorithms, and we demonstrate that they provide noticeable benefits in small problems designed to highlight the instability of TD methods. Finally, we observed improved performance when applying these algorithms at scale on classic Atari games from the Arcade Learning Environment. Ray Jiang, Tom Zahavy, Zhongwen Xu, Adam White 0001, Matteo Hessel, Charles Blundell, Hado van Hasselt |
ICML | 1 |
| 2020 | A General Approach to Fairness with Optimal TransportabstractWe propose a general approach to fairness based on transporting distributions corresponding to different sensitive attributes to a common distribution. We use optimal transport theory to derive target distributions and methods that allow us to achieve fairness with minimal changes to the unfair model. Our approach is applicable to both classification and regression problems, can enforce different notions of fairness, and enable us to achieve a Pareto-optimal trade-off between accuracy and fairness. We demonstrate that it outperforms previous approaches in several benchmark fairness datasets. Silvia Chiappa, Ray Jiang, Thomas S. Stepleton, Aldo Pacchiano, Heinrich Jiang, John Aslanides |
AAAI | 2 |
| 2019 | Degenerate Feedback Loops in Recommender SystemsabstractMachine learning is used extensively in recommender systems deployed in products. The decisions made by these systems can influence user beliefs and preferences which in turn affect the feedback the learning system receives - thus creating a feedback loop. This phenomenon can give rise to the so-called "echo chambers" or "filter bubbles" that have user and societal implications. In this paper, we provide a novel theoretical analysis that examines both the role of user dynamics and the behavior of recommender systems, disentangling the echo chamber from the filter bubble effect. In addition, we offer practical solutions to slow down system degeneracy. Our study contributes toward understanding and developing solutions to commonly cited issues in the complex temporal scenario, an area that is still largely unexplored. Ray Jiang, Silvia Chiappa, Tor Lattimore, András György 0001, Pushmeet Kohli |
AIES | 1 |
| 2019 | Beyond Greedy Ranking: Slate Optimization via List-CVAE
Ray Jiang, Sven Gowal, Yuqiu Qian, Timothy A. Mann, Danilo Jimenez Rezende |
ICLR (Poster) | 1 |
| 2019 | Learning from Delayed Outcomes via Proxies with Applications to Recommender SystemsabstractPredicting delayed outcomes is an important problem in recommender systems (e.g., if customers will finish reading an ebook). We formalize the problem as an adversarial, delayed online learning problem and consider how a proxy for the delayed outcome (e.g., if customers read a third of the book in 24 hours) can help minimize regret, even though the proxy is not available when making a prediction. Motivated by our regret analysis, we propose two neural network architectures: Factored Forecaster (FF) which is ideal if the proxy is informative of the outcome in hindsight, and Residual Factored Forecaster (RFF) that is robust to a non-informative proxy. Experiments on two real-world datasets for predicting human behavior show that RFF outperforms both FF and a direct forecaster that does not make use of the proxy. Our results suggest that exploiting proxies by factorization is a promising way to mitigate the impact of long delays in human-behavior prediction tasks. Timothy A. Mann, Sven Gowal, András György 0001, Huiyi Hu, Ray Jiang, Balaji Lakshminarayanan, Prav Srinivasan |
ICML | 5 |
| 2019 | Wasserstein Fair Classification
Ray Jiang, Aldo Pacchiano, Thomas S. Stepleton, Heinrich Jiang, Silvia Chiappa |
UAI | 1 |