EDBT 2026 Demo / reviewers in the wild / expert
Chongyi Zheng
dblp:250/9267
· DBLP profile ↗
6ranked-venue papers
3as first author
5since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 3 first-author · 5 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Reinforcement learning · 62% Representation and self-supervised learning · 31% Transfer learning and domain adaptation · 4% |
Topics — the 12 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
goal-conditioned reinforcement learning |
2.0 | 3 | 2024 | Contrastive Difference Predictive Coding · ICLR 2024 Stabilizing Contrastive RL: Techniques for Robotic Goal Reaching from Offline Data · ICLR 2024 Learning Domain Invariant Representations in Goal-conditioned Block MDPs · NeurIPS 2021 |
Machine learning › Representation and self-supervised learning
contrastive learning |
1.6 | 2 | 2025 | Can a MISL Fly? Analysis and Ingredients for Mutual Information Skill Learning · ICLR 2025 Learning Temporal Distances: Contrastive Successor Features Can Provide a Metric Structure for Decision-Making · ICML 2024 |
Machine learning › Reinforcement learning › function approximation › representation learning for reinforcement learning
successor features |
1.6 | 2 | 2025 | Can a MISL Fly? Analysis and Ingredients for Mutual Information Skill Learning · ICLR 2025 Learning Temporal Distances: Contrastive Successor Features Can Provide a Metric Structure for Decision-Making · ICML 2024 |
Machine learning › Reinforcement learning › hierarchical reinforcement learning › skill learning
skill discovery |
0.9 | 1 | 2025 | Can a MISL Fly? Analysis and Ingredients for Mutual Information Skill Learning · ICLR 2025 |
Machine learning › Representation and self-supervised learning › contrastive learning › temporal contrastive learning
contrastive predictive coding |
0.8 | 1 | 2024 | Contrastive Difference Predictive Coding · ICLR 2024 |
Machine learning › Reinforcement learning › function approximation › representation learning for reinforcement learning
contrastive reinforcement learning |
0.8 | 1 | 2024 | Stabilizing Contrastive RL: Techniques for Robotic Goal Reaching from Offline Data · ICLR 2024 |
Machine learning › Reinforcement learning
offline reinforcement learning |
0.8 | 1 | 2024 | Stabilizing Contrastive RL: Techniques for Robotic Goal Reaching from Offline Data · ICLR 2024 |
Machine learning › Transfer learning and domain adaptation
domain generalization |
0.5 | 1 | 2021 | Learning Domain Invariant Representations in Goal-conditioned Block MDPs · NeurIPS 2021 |
Machine learning › Representation and self-supervised learning › representation learning › invariant representation learning
domain-invariant representation |
0.5 | 1 | 2021 | Learning Domain Invariant Representations in Goal-conditioned Block MDPs · NeurIPS 2021 |
Machine learning › Reinforcement learning › multi-agent reinforcement learning › value-based multi-agent reinforcement learning
value decomposition |
0.4 | 1 | 2020 | Learning Nearly Decomposable Value Functions Via Communication Minimization · ICLR 2020 |
Machine learning › Reinforcement learning › markov decision process › partially observable MDP
Block MDP |
0.1 | 1 | 2021 | Learning Domain Invariant Representations in Goal-conditioned Block MDPs · NeurIPS 2021 |
Machine learning › Efficient and distributed learning › distributed training › communication-efficient training
communication optimization |
0.1 | 1 | 2020 | Learning Nearly Decomposable Value Functions Via Communication Minimization · ICLR 2020 |
Methods — techniques the papers use, named apart from their topics
contrastive learning · 2.3successor features · 1.5wasserstein distance · 0.9mutual information · 0.9self-supervised learning · 0.8quasi-metric · 0.8theoretical framework · 0.5PA-SkewFit · 0.5value function factorization · 0.4information bottleneck · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Can a MISL Fly? Analysis and Ingredients for Mutual Information Skill LearningabstractSelf-supervised learning has the potential of lifting several of the key challenges in reinforcement learning today, such as exploration, representation learning, and reward design. Recent work (METRA) has effectively argued that moving away from mutual information and instead optimizing a certain Wasserstein distance is important for good performance. In this paper, we argue that the benefits seen in that paper can largely be explained within the existing framework of mutual information skill learning (MISL).
Our analysis suggests a new MISL method (contrastive successor features) that retains the excellent performance of METRA with fewer moving parts, and highlights connections between skill learning, contrastive representation learning, and successor features. Finally, through careful ablation studies, we provide further insight into some of the key ingredients for both our method and METRA. Chongyi Zheng, Jens Tuyls, Joanne Peng, Benjamin Eysenbach |
ICLR | 1 |
| 2024 | Stabilizing Contrastive RL: Techniques for Robotic Goal Reaching from Offline DataabstractRobotic systems that rely primarily on self-supervised learning have the potential to decrease the amount of human annotation and engineering effort required to learn control strategies. In the same way that prior robotic systems have leveraged self-supervised techniques from computer vision (CV) and natural language processing (NLP), our work builds on prior work showing that the reinforcement learning (RL) itself can be cast as a self-supervised problem: learning to reach any goal without human-specified rewards or labels. Despite the seeming appeal, little (if any) prior work has demonstrated how self-supervised RL methods can be practically deployed on robotic systems. By first studying a challenging simulated version of this task, we discover design decisions about architectures and hyperparameters that increase the success rate by $2 \times$. These findings lay the groundwork for our main result: we demonstrate that a self-supervised RL algorithm based on contrastive learning can solve real-world, image-based robotic manipulation tasks, with tasks being specified by a single goal image provided after training. Chongyi Zheng, Benjamin Eysenbach, Homer Walke, Patrick Yin, Kuan Fang, Ruslan Salakhutdinov, Sergey Levine |
ICLR | 1 |
| 2024 | Contrastive Difference Predictive CodingabstractPredicting and reasoning about the future lie at the heart of many time-series questions. For example, goal-conditioned reinforcement learning can be viewed as learning representations to predict which states are likely to be visited in the future. While prior methods have used contrastive predictive coding to model time series data, learning representations that encode long-term dependencies usually requires large amounts of data. In this paper, we introduce a temporal difference version of contrastive predictive coding that stitches together pieces of different time series data to decrease the amount of data required to learn predictions of future events. We apply this representation learning method to derive an off-policy algorithm for goal-conditioned RL. Experiments demonstrate that, compared with prior RL methods, ours achieves $2 \times$ median improvement in success rates and can better cope with stochastic environments. In tabular settings, we show that our method is about $20\times$ more sample efficient than the successor representation and $1500 \times$ more sample efficient than the standard (Monte Carlo) version of contrastive predictive coding. Chongyi Zheng, Ruslan Salakhutdinov, Benjamin Eysenbach |
ICLR | 1 |
| 2024 | Learning Temporal Distances: Contrastive Successor Features Can Provide a Metric Structure for Decision-MakingabstractTemporal distances lie at the heart of many algorithms for planning, control, and reinforcement learning that involve reaching goals, allowing one to estimate the transit time between two states. However, prior attempts to define such temporal distances in stochastic settings have been stymied by an important limitation: these prior approaches do not satisfy the triangle inequality. This is not merely a definitional concern, but translates to an inability to generalize and find shortest paths. In this paper, we build on prior work in contrastive learning and quasimetrics to show how successor features learned by contrastive learning (after a change of variables) form a temporal distance that does satisfy the triangle inequality, even in stochastic settings. Importantly, this temporal distance is computationally efficient to estimate, even in high-dimensional and stochastic settings. Experiments in controlled settings and benchmark suites demonstrate that an RL algorithm based on these new temporal distances exhibits combinatorial generalization (i.e., "stitching") and can sometimes learn more quickly than prior methods, including those based on quasimetrics. Vivek Myers, Chongyi Zheng, Anca D. Dragan, Sergey Levine, Benjamin Eysenbach |
ICML | 2 |
| 2021 | Learning Domain Invariant Representations in Goal-conditioned Block MDPsabstractDeep Reinforcement Learning (RL) is successful in solving many complex Markov Decision Processes (MDPs) problems. However, agents often face unanticipated environmental changes after deployment in the real world. These changes are often spurious and unrelated to the underlying problem, such as background shifts for visual input agents. Unfortunately, deep RL policies are usually sensitive to these changes and fail to act robustly against them. This resembles the problem of domain generalization in supervised learning. In this work, we study this problem for goal-conditioned RL agents. We propose a theoretical framework in the Block MDP setting that characterizes the generalizability of goal-conditioned policies to new environments. Under this framework, we develop a practical method PA-SkewFit that enhances domain generalization. The empirical evaluation shows that our goal-conditioned RL agent can perform well in various unseen test environments, improving by 50\% over baselines. Beining Han, Chongyi Zheng, Harris Chan, Keiran Paster, Michael R. Zhang, Jimmy Ba |
NeurIPS | 2 |
| 2020 | Learning Nearly Decomposable Value Functions Via Communication Minimization
Tonghan Wang 0001, Chongyi Zheng, Chongjie Zhang |
ICLR | 3 |