Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Yucen Wang

dblp:349/7802 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2025
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Reinforcement learning · 88% Robot navigation and mapping · 9% Deep learning architectures and training · 3%

Topics — the 14 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning › model-based reinforcement learning
world model
1.622025
FOUNDER: Grounding Foundation Models in World Models for Open-Ended Embodied Decision Making · ICML 2025
AD3: Implicit Action is the Key for World Models to Distinguish the Diverse Visual Distractors · ICML 2024
Robotics › Robot navigation and mapping › embodied navigation
embodied decision-making
0.912025
FOUNDER: Grounding Foundation Models in World Models for Open-Ended Embodied Decision Making · ICML 2025
Machine learning › Reinforcement learning
goal-conditioned reinforcement learning
0.912025
FOUNDER: Grounding Foundation Models in World Models for Open-Ended Embodied Decision Making · ICML 2025
Machine learning › Reinforcement learning › unsupervised reinforcement learning
reward-free reinforcement learning
0.912025
FOUNDER: Grounding Foundation Models in World Models for Open-Ended Embodied Decision Making · ICML 2025
Machine learning › Reinforcement learning › reward learning › reward modeling
reward model evaluation
0.912025
Reward Models in Deep Reinforcement Learning: A Survey · IJCAI 2025
Machine learning › Reinforcement learning › reward learning
reward modeling
0.912025
Reward Models in Deep Reinforcement Learning: A Survey · IJCAI 2025
Machine learning › Reinforcement learning
model-based reinforcement learning
0.812024
AD3: Implicit Action is the Key for World Models to Distinguish the Diverse Visual Distractors · ICML 2024
Machine learning › Reinforcement learning
imitation learning
0.712023
SeMAIL: Eliminating Distractors in Visual Imitation via Separated Models · ICML 2023
Machine learning › Reinforcement learning › imitation learning
model-based imitation learning
0.712023
SeMAIL: Eliminating Distractors in Visual Imitation via Separated Models · ICML 2023
Machine learning › Reinforcement learning › imitation learning › learning from observation
visual imitation learning
0.712023
SeMAIL: Eliminating Distractors in Visual Imitation via Separated Models · ICML 2023
Machine learning › Deep learning architectures and training
foundation model
0.312025
FOUNDER: Grounding Foundation Models in World Models for Open-Ended Embodied Decision Making · ICML 2025
Machine learning › Reinforcement learning
policy optimization
0.312025
Reward Models in Deep Reinforcement Learning: A Survey · IJCAI 2025
Machine learning › Reinforcement learning › deep reinforcement learning › visual reinforcement learning
visual control
0.212024
AD3: Implicit Action is the Key for World Models to Distinguish the Diverse Visual Distractors · ICML 2024
Machine learning › Reinforcement learning
sample efficiency
0.212023
SeMAIL: Eliminating Distractors in Visual Imitation via Separated Models · ICML 2023

Methods — techniques the papers use, named apart from their topics

world model · 0.9reward learning · 0.9goal-conditioned policy learning · 0.9foundation model · 0.9separated world models · 0.8implicit action generator · 0.8separated model · 0.7dynamics decoupling · 0.7adversarial imitation learning · 0.7
YearPublicationVenuePosition
2025 FOUNDER: Grounding Foundation Models in World Models for Open-Ended Embodied Decision Making
abstract
Foundation Models (FMs) and World Models (WMs) offer complementary strengths in task generalization at different levels. In this work, we propose FOUNDER, a framework that integrates the generalizable knowledge embedded in FMs with the dynamic modeling capabilities of WMs to enable open-ended task solving in embodied environments in a reward-free manner. We learn a mapping function that grounds FM representations in the WM state space, effectively inferring the agent’s physical states in the world simulator from external observations. This mapping enables the learning of a goal-conditioned policy through imagination during behavior learning, with the mapped task serving as the goal state. Our method leverages the predicted temporal distance to the goal state as an informative reward signal. FOUNDER demonstrates superior performance on various multi-task offline visual control benchmarks, excelling in capturing the deep-level semantics of tasks specified by text or videos, particularly in scenarios involving complex observations or domain gaps where prior methods struggle. The consistency of our learned reward function with the ground-truth reward is also empirically validated. Our project website is https://sites.google.com/view/founder-rl.
Yucen Wang, Shenghua Wan, Le Gan, De-Chuan Zhan
ICML1
2025 Reward Models in Deep Reinforcement Learning: A Survey
abstract
In reinforcement learning (RL), agents continually interact with the environment and use the feedback to refine their behavior. To guide policy optimization, reward models are introduced as proxies of the desired objectives, such that when the agent maximizes the accumulated reward, it also fulfills the task designer's intentions. Recently, significant attention from both academic and industrial researchers has focused on developing reward models that not only align closely with the true objectives but also facilitate policy optimization. In this survey, we provide a comprehensive review of reward modeling techniques within the RL literature. We begin by outlining the background and preliminaries in reward modeling. Next, we present an overview of recent reward modeling approaches, categorizing them based on the source, the mechanism, and the reward learning paradigm. Building on this understanding, we discuss various applications of these reward modeling techniques and review methods for evaluating reward models. Finally, we conclude by highlighting promising research directions in reward modeling. Altogether, this survey includes both established and emerging methods, filling the vacancy of a systematic review of reward models in current literature.
Shenghua Wan, Yucen Wang, Chenxiao Gao, Le Gan, Zongzhang Zhang, De-Chuan Zhan
IJCAI3
2024 AD3: Implicit Action is the Key for World Models to Distinguish the Diverse Visual Distractors
abstract
Model-based methods have significantly contributed to distinguishing task-irrelevant distractors for visual control. However, prior research has primarily focused on heterogeneous distractors like noisy background videos, leaving homogeneous distractors that closely resemble controllable agents largely unexplored, which poses significant challenges to existing methods. To tackle this problem, we propose Implicit Action Generator (IAG) to learn the implicit actions of visual distractors, and present a new algorithm named implicit Action-informed Diverse visual Distractors Distinguisher (AD3), that leverages the action inferred by IAG to train separated world models. Implicit actions effectively capture the behavior of background distractors, aiding in distinguishing the task-irrelevant components, and the agent can optimize the policy within the task-relevant state space. Our method achieves superior performance on various visual control tasks featuring both heterogeneous and homogeneous distractors. The indispensable role of implicit actions learned by IAG is also empirically validated.
Yucen Wang, Shenghua Wan, Le Gan, De-Chuan Zhan
ICML1
2023 SeMAIL: Eliminating Distractors in Visual Imitation via Separated Models
abstract
Model-based imitation learning (MBIL) is a popular reinforcement learning method that improves sample efficiency on high-dimension input sources, such as images and videos. Following the convention of MBIL research, existing algorithms are highly deceptive by task-irrelevant information, especially moving distractors in videos. To tackle this problem, we propose a new algorithm - named Separated Model-based Adversarial Imitation Learning (SeMAIL) - decoupling the environment dynamics into two parts by task-relevant dependency, which is determined by agent actions, and training separately. In this way, the agent can imagine its trajectories and imitate the expert behavior efficiently in task-relevant state space. Our method achieves near-expert performance on various visual control tasks with complex observations and the more challenging tasks with different backgrounds from expert observations.
Shenghua Wan, Yucen Wang, Minghao Shao, Ruying Chen, De-Chuan Zhan
ICML2