Wancong Zhang

dblp:270/8288 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2025
0009-0009-5461-2874ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Reinforcement learning · 67% Representation and self-supervised learning · 33%

Topics — the 3 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Representation and self-supervised learning › representation learning › latent representation learning › state representation learning
latent dynamics model
0.912025
Learning from Reward-Free Offline Data: A Case for Planning with Latent Dynamics Models · NeurIPS 2025
Machine learning › Reinforcement learning
offline reinforcement learning
0.912025
Learning from Reward-Free Offline Data: A Case for Planning with Latent Dynamics Models · NeurIPS 2025
Machine learning › Reinforcement learning › model-based reinforcement learning › model-based planning
planning with learned models
0.912025
Learning from Reward-Free Offline Data: A Case for Planning with Latent Dynamics Models · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

model-based planning · 0.9joint embedding predictive architecture · 0.9goal-conditioned RL · 0.9
YearPublicationVenuePosition
2025 Learning from Reward-Free Offline Data: A Case for Planning with Latent Dynamics Models
abstract
A long-standing goal in AI is to develop agents capable of solving diverse tasks across a range of environments, including those never seen during training. Two dominant paradigms address this challenge: (i) reinforcement learning (RL), which learns policies via trial and error, and (ii) optimal control, which plans actions using a known or learned dynamics model. However, their comparative strengths in the offline setting—where agents must learn from reward-free trajectories—remain underexplored. In this work, we systematically evaluate RL and control-based methods on a suite of navigation tasks, using offline datasets of varying quality. On the RL side, we consider goal-conditioned and zero-shot methods. On the control side, we train a latent dynamics model using the Joint Embedding Predictive Architecture (JEPA) and employ it for planning. We investigate how factors such as data diversity, trajectory quality, and environment variability influence the performance of these approaches. Our results show that model-free RL benefits most from large amounts of high-quality data, whereas model-based planning generalizes better to unseen layouts and is more data-efficient, while achieving trajectory stitching performance comparable to leading model-free methods. Notably, planning with a latent dynamics model proves to be a strong approach for handling suboptimal offline data and adapting to diverse environments.
Uladzislau Sobal, Wancong Zhang, Kyunghyun Cho, Randall Balestriero, Tim G. J. Rudner, Yann LeCun
NeurIPS2
2023 AttnOD: An Attention-Based OD Prediction Model with Adaptive Graph Convolution
Wancong Zhang, Gang Wang 0058, Tongyu Zhu
ICONIP (12)1