Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Zhengbang Zhu

dblp:277/0869 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
10since 2021 · last 2026
0009-0005-9310-3598ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 2 first-author · 8 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Reinforcement learning · 57% Generative modeling · 24% Multi-agent systems · 12%
Databases, data mining, and information retrieval
1 paper
Recommender systems · 100%

Topics — the 15 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
2.432025
ContraDiff: Planning Towards High Return States via Contrastive Learning · ICLR 2025
MADiff: Offline Multi-agent Learning with Diffusion Models · NeurIPS 2024
DiffStitch: Boosting Offline Reinforcement Learning with Diffusion-based Trajectory Stitching · ICML 2024
Machine learning › Reinforcement learning
offline reinforcement learning
2.432025
ContraDiff: Planning Towards High Return States via Contrastive Learning · ICLR 2025
MADiff: Offline Multi-agent Learning with Diffusion Models · NeurIPS 2024
DiffStitch: Boosting Offline Reinforcement Learning with Diffusion-based Trajectory Stitching · ICML 2024
Machine learning › Reinforcement learning
multi-agent reinforcement learning
0.912025
Reconstruction-Guided Policy: Enhancing Decision-Making through Agent-Wise State Consistency · ICLR 2025
Machine learning › Reinforcement learning
partial observability
0.912025
Reconstruction-Guided Policy: Enhancing Decision-Making through Agent-Wise State Consistency · ICLR 2025
Machine learning › Representation and self-supervised learning › representation learning › latent representation learning
state representation learning
0.912025
Reconstruction-Guided Policy: Enhancing Decision-Making through Agent-Wise State Consistency · ICLR 2025
Machine learning › Generative modeling › diffusion model
diffusion-based data augmentation
0.812024
DiffStitch: Boosting Offline Reinforcement Learning with Diffusion-based Trajectory Stitching · ICML 2024
Knowledge, reasoning and agents › Multi-agent systems
multi-agent coordination
0.812024
MADiff: Offline Multi-agent Learning with Diffusion Models · NeurIPS 2024
Machine learning › Reinforcement learning › multi-agent reinforcement learning
offline multi-agent reinforcement learning
0.812024
MADiff: Offline Multi-agent Learning with Diffusion Models · NeurIPS 2024
Knowledge, reasoning and agents › Multi-agent systems › human-agent interaction
teammate modeling
0.812024
MADiff: Offline Multi-agent Learning with Diffusion Models · NeurIPS 2024
Machine learning › Reinforcement learning › offline reinforcement learning
trajectory stitching
0.812024
DiffStitch: Boosting Offline Reinforcement Learning with Diffusion-based Trajectory Stitching · ICML 2024
Recommender systems
recommender system evaluation
0.812024
Understanding or Manipulation: Rethinking Online Performance Gains of Modern Recommender Systems · ACM Trans. Inf. Syst. 2024
Recommender systems › user modeling
user preference modeling
0.812024
Understanding or Manipulation: Rethinking Online Performance Gains of Modern Recommender Systems · ACM Trans. Inf. Syst. 2024
Machine learning › Reinforcement learning
imitation learning
0.612022
Plan Your Target and Learn Your Skills: Transferable State-Only Imitation Learning via Decoupled Policy Optimization · ICML 2022
Machine learning › Reinforcement learning › imitation learning
learning from observation
0.612022
Plan Your Target and Learn Your Skills: Transferable State-Only Imitation Learning via Decoupled Policy Optimization · ICML 2022
Machine learning › Reinforcement learning
policy optimization
0.612022
Plan Your Target and Learn Your Skills: Transferable State-Only Imitation Learning via Decoupled Policy Optimization · ICML 2022

Methods — techniques the papers use, named apart from their topics

diffusion model · 1.6state reconstruction · 0.9contrastive learning · 0.9agent-wise attention · 0.9preference shift metric · 0.8manipulation score · 0.8data augmentation · 0.8centralized controller · 0.8attention-based diffusion model · 0.8policy gradient · 0.6inverse dynamics model · 0.6generative adversarial training · 0.6
YearPublicationVenuePosition
2026 SocialDriveGen: generating diverse traffic scenarios with controllable social interactions
Jiaguo Tian, Zhengbang Zhu, Shenyu Zhang 0001, Weijie Peng, Shizeng Yao, Weinan Zhang 0001
Frontiers Comput. Sci.2
2025 Reconstruction-Guided Policy: Enhancing Decision-Making through Agent-Wise State Consistency
abstract
An important challenge in multi-agent reinforcement learning is partial observability, where agents cannot access the global state of the environment during execution and can only receive observations within their field of view. To address this issue, previous works typically use the dimensional-wise state, which is obtained by applying MLP or dimensional-based attention on the global state, for decision-making during training and relying on a reconstructed dimensional-wise state during execution. However, dimensional-wise states tend to divert agent attention to specific features, neglecting potential dependencies between agents, making it difficult to make optimal decisions. Moreover, the inconsistency between the states used in training and execution further increases additional errors. To resolve these issues, we propose a method called Reconstruction-Guided Policy (RGP) to reconstruct the agent-wise state, which represents the information of inter-agent relationships, as input for decision-making during both training and execution. This not only preserves the potential dependencies between agents but also ensures consistency between the states used in training and execution. We conducted extensive experiments on both discrete and continuous action environments to evaluate RGP, and the results demonstrates its superior effectiveness. Our code is public in https://anonymous.4open.science/r/RGP-9F79
Qifan Liang, Yixiang Shan, Zhengbang Zhu, Ting Long, Weinan Zhang 0001, Yuan Tian 0016
ICLR4
2025 ContraDiff: Planning Towards High Return States via Contrastive Learning
abstract
The performance of offline reinforcement learning (RL) is sensitive to the proportion of high-return trajectories in the offline dataset. However, in many simulation environments and real-world scenarios, there are large ratios of low-return trajectories rather than high-return trajectories, which makes learning an efficient policy challenging. In this paper, we propose a method called Contrastive Diffuser (ContraDiff) to make full use of low-return trajectories and improve the performance of offline RL algorithms. Specifically, ContraDiff groups the states of trajectories in the offline dataset into high-return states and low-return states and treats them as positive and negative samples correspondingly. Then, it designs a contrastive mechanism to pull the planned trajectory of an agent toward high-return states and push them away from low-return states. Through the contrast mechanism, trajectories with low returns can serve as negative examples for policy learning, guiding the agent to avoid areas associated with low returns and achieve better performance. Through the contrast mechanism, trajectories with low returns provide a ``counteracting force'' guides the agent to avoid areas associated with low returns and achieve better performance. Experiments on 27 sub-optimal datasets demonstrate the effectiveness of our proposed method. Our code is publicly available at https://github.com/Looomo/contradiff.
Yixiang Shan, Zhengbang Zhu, Ting Long, Qifan Liang, Yi Chang 0001, Weinan Zhang 0001
ICLR2
2025 DriveGen: Towards Infinite Diverse Traffic Scenarios with Large Models
abstract
Microscopic traffic simulation has become an important tool for autonomous driving training and testing. Although recent data-driven approaches advance realistic behavior generation, their learning still relies primarily on a single real-world dataset, which limits their diversity and thereby hinders downstream algorithm optimization. In this paper, we propose DriveGen, a novel traffic simulation framework with large models for more diverse traffic generation that supports further customized designs. DriveGen consists of two internal stages: the initialization stage uses a large language model and retrieval technique to generate map and vehicle assets; the rollout stage outputs trajectories with selected waypoint goals from a visual language model and a specifically designed diffusion planner. Through this two-staged process, DriveGen fully utilizes large models’ high-level cognition and reasoning of driving behavior, obtaining greater diversity beyond datasets while maintaining high realism. To support effective downstream optimization, we additionally develop DriveGen-CS, an automatic corner case generation pipeline that uses failures of the driving algorithm as additional prompt knowledge for large models without the need for retraining or fine-tuning. Experiments show that our generated scenarios and corner cases have superior performance compared to state-of-the-art baselines. Downstream experiments further verify that the synthesized traffic of DriveGen provides better optimization of the performance of typical driving algorithms, demonstrating the effectiveness of our framework.
Shenyu Zhang 0001, Jiaguo Tian, Zhengbang Zhu, Weinan Zhang 0001
IROS3
2024 Multi-Agent Trajectory Prediction with Scalable Diffusion Transformer
abstract
Accurate prediction of multi-agent spatiotemporal systems is critical to various real-world applications, such as autonomous driving, sports, and multiplayer games.Unfortunately, modeling multiagent trajectories is challenging due to its complicated, interactive, and multi-modal nature.Recently, diffusion models have achieved great success in modeling multi-modal distribution and trajectory generation, showing promising ability in resolving this problem.Motivated by this, in this paper, we propose a novel multi-agent trajectory prediction framework, dubbed Scalable Diffusion Transformer (SDT), which is naturally designed to learn the complicated distribution and implicit interactions among agents.We evaluate SDT on a set of real-world benchmark datasets and compare it with representative baseline methods, which demonstrates the better multi-agent trajectory prediction ability of SDT in terms of accuracy and diversity.
Shenyu Zhang 0001, Shixiong Kai, Chang Chen 0015, Yuzheng Zhuang, Zhengbang Zhu, Minghuan Liu, Weinan Zhang 0001
DAI5
2024 DiffStitch: Boosting Offline Reinforcement Learning with Diffusion-based Trajectory Stitching
abstract
In offline reinforcement learning (RL), the performance of the learned policy highly depends on the quality of offline datasets. However, the offline dataset contains very limited optimal trajectories in many cases. This poses a challenge for offline RL algorithms, as agents must acquire the ability to transit to high-reward regions. To address this issue, we introduce Diffusionbased Trajectory Stitching (DiffStitch), a novel diffusion-based data augmentation pipeline that systematically generates stitching transitions between trajectories. DiffStitch effectively connects low-reward trajectories with high-reward trajectories, forming globally optimal trajectories and thereby mitigating the challenges faced by offline RL algorithms in learning trajectory stitching. Empirical experiments conducted on D4RL datasets demonstrate the effectiveness of our pipeline across RL methodologies. Notably, DiffStitch demonstrates substantial enhancements in the performance of one-step methods(IQL), imitation learning methods(TD3+BC) and trajectory optimization methods(DT). Our code is publicly available at https://github.com/guangheli12/DiffStitch
Guanghe Li, Yixiang Shan, Zhengbang Zhu, Ting Long, Weinan Zhang 0001
ICML3
2024 MADiff: Offline Multi-agent Learning with Diffusion Models
abstract
Offline reinforcement learning (RL) aims to learn policies from pre-existing datasets without further interactions, making it a challenging task. Q-learning algorithms struggle with extrapolation errors in offline settings, while supervised learning methods are constrained by model expressiveness. Recently, diffusion models (DMs) have shown promise in overcoming these limitations in single-agent learning, but their application in multi-agent scenarios remains unclear. Generating trajectories for each agent with independent DMs may impede coordination, while concatenating all agents’ information can lead to low sample efficiency. Accordingly, we propose MADiff, which is realized with an attention-based diffusion model to model the complex coordination among behaviors of multiple agents. To our knowledge, MADiff is the first diffusion-based multi-agent learning framework, functioning as both a decentralized policy and a centralized controller. During decentralized executions, MADiff simultaneously performs teammate modeling, and the centralized controller can also be applied in multi-agent trajectory predictions. Our experiments demonstrate that MADiff outperforms baseline algorithms across various multi-agent learning tasks, highlighting its effectiveness in modeling complex multi-agent interactions.
Zhengbang Zhu, Minghuan Liu, Liyuan Mao, Bingyi Kang, Minkai Xu, Yong Yu 0001, Stefano Ermon, Weinan Zhang 0001
NeurIPS1
2024 Understanding or Manipulation: Rethinking Online Performance Gains of Modern Recommender Systems
abstract
Recommender systems are expected to be assistants that help human users find relevant information automatically without explicit queries. As recommender systems evolve, increasingly sophisticated learning techniques are applied and have achieved better performance in terms of user engagement metrics such as clicks and browsing time. The increase in the measured performance, however, can have two possible attributions: a better understanding of user preferences, and a more proactive ability to utilize human bounded rationality to seduce user over-consumption. A natural following question is whether current recommendation algorithms are manipulating user preferences. If so, can we measure the manipulation level? In this article, we present a general framework for benchmarking the degree of manipulations of recommendation algorithms, in both slate recommendation and sequential recommendation scenarios. The framework consists of four stages, initial preference calculation, training data collection, algorithm training and interaction, and metrics calculation that involves two proposed metrics, Manipulation Score and Preference Shift. We benchmark some representative recommendation algorithms in both synthetic and real-world datasets under the proposed framework. We have observed that a high online click-through rate does not necessarily mean a better understanding of user initial preference, but ends in prompting users to choose more documents they initially did not favor. Moreover, we find that the training data have notable impacts on the manipulation degrees, and algorithms with more powerful modeling abilities are more sensitive to such impacts. The experiments also verified the usefulness of the proposed metrics for measuring the degree of manipulations. We advocate that future recommendation algorithm studies should be treated as an optimization problem with constrained user preference manipulations.
Zhengbang Zhu, Rongjun Qin, Xinyi Dai, Yang Yu 0001, Yong Yu 0001, Weinan Zhang 0001
ACM Trans. Inf. Syst.1
2023 RITA: Boost Driving Simulators with Realistic Interactive Traffic Flow
abstract
High-quality traffic flow generation is the core module in building simulators for autonomous driving. However, the majority of available simulators are incapable of replicating traffic patterns that accurately reflect the various features of real-world data while also simulating human-like reactive responses to the tested autopilot driving strategies. Taking one step forward to addressing such a problem, we propose Realistic Interactive TrAffic flow (RITA) as an integrated component of existing driving simulators to provide high-quality traffic flow for the evaluation and optimization of the tested driving strategies. RITA is developed with consideration of three key features, i.e., fidelity, diversity, and controllability, and consists of two core modules called RITABackend and RITAKit. RITABackend is built to support vehicle-wise control and provide traffic generation models from real-world datasets, while RITAKit is developed with easy-to-use interfaces for controllable traffic generation via RITABackend. We demonstrate RITA’s capacity to create diversified and high-fidelity traffic simulations in several highly interactive highway scenarios. The experimental findings demonstrate that our produced RITA traffic flows exhibit all three key features, hence enhancing the completeness of driving strategy evaluation. Moreover, we showcase the possibility for further improvement of baseline strategies through online fine-tuning with RITA traffic flows.
Zhengbang Zhu, Shenyu Zhang 0001, Yuzheng Zhuang, Yuecheng Liu, Minghuan Liu, Ziqing Gong, Shixiong Kai, Qiang Gu, Bin Wang 0034, Siyuan Cheng 0012, Xinyu Wang 0001, Jianye Hao, Yong Yu 0001
DAI1
2022 Plan Your Target and Learn Your Skills: Transferable State-Only Imitation Learning via Decoupled Policy Optimization
abstract
Recent progress in state-only imitation learning extends the scope of applicability of imitation learning to real-world settings by relieving the need for observing expert actions. However, existing solutions only learn to extract a state-to-action mapping policy from the data, without considering how the expert plans to the target. This hinders the ability to leverage demonstrations and limits the flexibility of the policy. In this paper, we introduce Decoupled Policy Optimization (DePO), which explicitly decouples the policy as a high-level state planner and an inverse dynamics model. With embedded decoupled policy gradient and generative adversarial training, DePO enables knowledge transfer to different action spaces or state transition dynamics, and can generalize the planner to out-of-demonstration state regions. Our in-depth experimental analysis shows the effectiveness of DePO on learning a generalized target state planner while achieving the best imitation performance. We demonstrate the appealing usage of DePO for transferring across different tasks by pre-training, and the potential for co-training agents with various skills.
Minghuan Liu, Zhengbang Zhu, Yuzheng Zhuang, Weinan Zhang 0001, Jianye Hao, Yong Yu 0001, Jun Wang 0012
ICML2
2020 Weakly-Supervised Reconstruction of 3D Objects with Large Shape Variation from Single In-the-Wild Images
Shichen Sun, Zhengbang Zhu, Xiaowei Dai, Qijun Zhao, Jing Li 0060
ACCV (1)2