EDBT 2026 Demo / reviewers in the wild / expert
Zhengbang Zhu
dblp:277/0869
· DBLP profile ↗
11ranked-venue papers
3as first author
10since 2021 · last 2026
0009-0005-9310-3598ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 2 first-author · 8 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Reinforcement learning · 57% Generative modeling · 24% Multi-agent systems · 12% | |
| Databases, data mining, and information retrieval
1 paper |
Recommender systems · 100% |
Topics — the 15 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
diffusion model |
2.4 | 3 | 2025 | ContraDiff: Planning Towards High Return States via Contrastive Learning · ICLR 2025 MADiff: Offline Multi-agent Learning with Diffusion Models · NeurIPS 2024 DiffStitch: Boosting Offline Reinforcement Learning with Diffusion-based Trajectory Stitching · ICML 2024 |
Machine learning › Reinforcement learning
offline reinforcement learning |
2.4 | 3 | 2025 | ContraDiff: Planning Towards High Return States via Contrastive Learning · ICLR 2025 MADiff: Offline Multi-agent Learning with Diffusion Models · NeurIPS 2024 DiffStitch: Boosting Offline Reinforcement Learning with Diffusion-based Trajectory Stitching · ICML 2024 |
Machine learning › Reinforcement learning
multi-agent reinforcement learning |
0.9 | 1 | 2025 | Reconstruction-Guided Policy: Enhancing Decision-Making through Agent-Wise State Consistency · ICLR 2025 |
Machine learning › Reinforcement learning
partial observability |
0.9 | 1 | 2025 | Reconstruction-Guided Policy: Enhancing Decision-Making through Agent-Wise State Consistency · ICLR 2025 |
Machine learning › Representation and self-supervised learning › representation learning › latent representation learning
state representation learning |
0.9 | 1 | 2025 | Reconstruction-Guided Policy: Enhancing Decision-Making through Agent-Wise State Consistency · ICLR 2025 |
Machine learning › Generative modeling › diffusion model
diffusion-based data augmentation |
0.8 | 1 | 2024 | DiffStitch: Boosting Offline Reinforcement Learning with Diffusion-based Trajectory Stitching · ICML 2024 |
Knowledge, reasoning and agents › Multi-agent systems
multi-agent coordination |
0.8 | 1 | 2024 | MADiff: Offline Multi-agent Learning with Diffusion Models · NeurIPS 2024 |
Machine learning › Reinforcement learning › multi-agent reinforcement learning
offline multi-agent reinforcement learning |
0.8 | 1 | 2024 | MADiff: Offline Multi-agent Learning with Diffusion Models · NeurIPS 2024 |
Knowledge, reasoning and agents › Multi-agent systems › human-agent interaction
teammate modeling |
0.8 | 1 | 2024 | MADiff: Offline Multi-agent Learning with Diffusion Models · NeurIPS 2024 |
Machine learning › Reinforcement learning › offline reinforcement learning
trajectory stitching |
0.8 | 1 | 2024 | DiffStitch: Boosting Offline Reinforcement Learning with Diffusion-based Trajectory Stitching · ICML 2024 |
Recommender systems
recommender system evaluation |
0.8 | 1 | 2024 | Understanding or Manipulation: Rethinking Online Performance Gains of Modern Recommender Systems · ACM Trans. Inf. Syst. 2024 |
Recommender systems › user modeling
user preference modeling |
0.8 | 1 | 2024 | Understanding or Manipulation: Rethinking Online Performance Gains of Modern Recommender Systems · ACM Trans. Inf. Syst. 2024 |
Machine learning › Reinforcement learning
imitation learning |
0.6 | 1 | 2022 | Plan Your Target and Learn Your Skills: Transferable State-Only Imitation Learning via Decoupled Policy Optimization · ICML 2022 |
Machine learning › Reinforcement learning › imitation learning
learning from observation |
0.6 | 1 | 2022 | Plan Your Target and Learn Your Skills: Transferable State-Only Imitation Learning via Decoupled Policy Optimization · ICML 2022 |
Machine learning › Reinforcement learning
policy optimization |
0.6 | 1 | 2022 | Plan Your Target and Learn Your Skills: Transferable State-Only Imitation Learning via Decoupled Policy Optimization · ICML 2022 |
Methods — techniques the papers use, named apart from their topics
diffusion model · 1.6state reconstruction · 0.9contrastive learning · 0.9agent-wise attention · 0.9preference shift metric · 0.8manipulation score · 0.8data augmentation · 0.8centralized controller · 0.8attention-based diffusion model · 0.8policy gradient · 0.6inverse dynamics model · 0.6generative adversarial training · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SocialDriveGen: generating diverse traffic scenarios with controllable social interactions
Jiaguo Tian, Zhengbang Zhu, Shenyu Zhang 0001, Weijie Peng, Shizeng Yao, Weinan Zhang 0001 |
Frontiers Comput. Sci. | 2 |
| 2025 | Reconstruction-Guided Policy: Enhancing Decision-Making through Agent-Wise State ConsistencyabstractAn important challenge in multi-agent reinforcement learning is partial observability, where agents cannot access the global state of the environment during execution and can only receive observations within their field of view. To address this issue, previous works typically use the dimensional-wise state, which is obtained by applying MLP or dimensional-based attention on the global state, for decision-making during training and relying on a reconstructed dimensional-wise state during execution. However, dimensional-wise states tend to divert agent attention to specific features, neglecting potential dependencies between agents, making it difficult to make optimal decisions. Moreover, the inconsistency between the states used in training and execution further increases additional errors. To resolve these issues, we propose a method called Reconstruction-Guided Policy (RGP) to reconstruct the agent-wise state, which represents the information of inter-agent relationships, as input for decision-making during both training and execution. This not only preserves the potential dependencies between agents but also ensures consistency between the states used in training and execution. We conducted extensive experiments on both discrete and continuous action environments to evaluate RGP, and the results demonstrates its superior effectiveness. Our code is public in https://anonymous.4open.science/r/RGP-9F79 Qifan Liang, Yixiang Shan, Zhengbang Zhu, Ting Long, Weinan Zhang 0001, Yuan Tian 0016 |
ICLR | 4 |
| 2025 | ContraDiff: Planning Towards High Return States via Contrastive LearningabstractThe performance of offline reinforcement learning (RL) is sensitive to the proportion of high-return trajectories in the offline dataset. However, in many simulation environments and real-world scenarios, there are large ratios of low-return trajectories rather than high-return trajectories, which makes learning an efficient policy challenging. In this paper, we propose a method called Contrastive Diffuser (ContraDiff) to make full use of low-return trajectories and improve the performance of offline RL algorithms. Specifically, ContraDiff groups the states of trajectories in the offline dataset into high-return states and low-return states and treats them as positive and negative samples correspondingly. Then, it designs a contrastive mechanism to pull the planned trajectory of an agent toward high-return states and push them away from low-return states. Through the contrast mechanism, trajectories with low returns can serve as negative examples for policy learning, guiding the agent to avoid areas associated with low returns and achieve better performance. Through the contrast mechanism, trajectories with low returns provide a ``counteracting force'' guides the agent to avoid areas associated with low returns and achieve better performance.
Experiments on 27 sub-optimal datasets demonstrate the effectiveness of our proposed method. Our code is publicly available at https://github.com/Looomo/contradiff. Yixiang Shan, Zhengbang Zhu, Ting Long, Qifan Liang, Yi Chang 0001, Weinan Zhang 0001 |
ICLR | 2 |
| 2025 | DriveGen: Towards Infinite Diverse Traffic Scenarios with Large ModelsabstractMicroscopic traffic simulation has become an important tool for autonomous driving training and testing. Although recent data-driven approaches advance realistic behavior generation, their learning still relies primarily on a single real-world dataset, which limits their diversity and thereby hinders downstream algorithm optimization. In this paper, we propose DriveGen, a novel traffic simulation framework with large models for more diverse traffic generation that supports further customized designs. DriveGen consists of two internal stages: the initialization stage uses a large language model and retrieval technique to generate map and vehicle assets; the rollout stage outputs trajectories with selected waypoint goals from a visual language model and a specifically designed diffusion planner. Through this two-staged process, DriveGen fully utilizes large models’ high-level cognition and reasoning of driving behavior, obtaining greater diversity beyond datasets while maintaining high realism. To support effective downstream optimization, we additionally develop DriveGen-CS, an automatic corner case generation pipeline that uses failures of the driving algorithm as additional prompt knowledge for large models without the need for retraining or fine-tuning. Experiments show that our generated scenarios and corner cases have superior performance compared to state-of-the-art baselines. Downstream experiments further verify that the synthesized traffic of DriveGen provides better optimization of the performance of typical driving algorithms, demonstrating the effectiveness of our framework. Shenyu Zhang 0001, Jiaguo Tian, Zhengbang Zhu, Weinan Zhang 0001 |
IROS | 3 |
| 2024 | Multi-Agent Trajectory Prediction with Scalable Diffusion TransformerabstractAccurate prediction of multi-agent spatiotemporal systems is critical to various real-world applications, such as autonomous driving, sports, and multiplayer games.Unfortunately, modeling multiagent trajectories is challenging due to its complicated, interactive, and multi-modal nature.Recently, diffusion models have achieved great success in modeling multi-modal distribution and trajectory generation, showing promising ability in resolving this problem.Motivated by this, in this paper, we propose a novel multi-agent trajectory prediction framework, dubbed Scalable Diffusion Transformer (SDT), which is naturally designed to learn the complicated distribution and implicit interactions among agents.We evaluate SDT on a set of real-world benchmark datasets and compare it with representative baseline methods, which demonstrates the better multi-agent trajectory prediction ability of SDT in terms of accuracy and diversity. Shenyu Zhang 0001, Shixiong Kai, Chang Chen 0015, Yuzheng Zhuang, Zhengbang Zhu, Minghuan Liu, Weinan Zhang 0001 |
DAI | 5 |
| 2024 | DiffStitch: Boosting Offline Reinforcement Learning with Diffusion-based Trajectory StitchingabstractIn offline reinforcement learning (RL), the performance of the learned policy highly depends on the quality of offline datasets. However, the offline dataset contains very limited optimal trajectories in many cases. This poses a challenge for offline RL algorithms, as agents must acquire the ability to transit to high-reward regions. To address this issue, we introduce Diffusionbased Trajectory Stitching (DiffStitch), a novel diffusion-based data augmentation pipeline that systematically generates stitching transitions between trajectories. DiffStitch effectively connects low-reward trajectories with high-reward trajectories, forming globally optimal trajectories and thereby mitigating the challenges faced by offline RL algorithms in learning trajectory stitching. Empirical experiments conducted on D4RL datasets demonstrate the effectiveness of our pipeline across RL methodologies. Notably, DiffStitch demonstrates substantial enhancements in the performance of one-step methods(IQL), imitation learning methods(TD3+BC) and trajectory optimization methods(DT). Our code is publicly available at https://github.com/guangheli12/DiffStitch Guanghe Li, Yixiang Shan, Zhengbang Zhu, Ting Long, Weinan Zhang 0001 |
ICML | 3 |
| 2024 | MADiff: Offline Multi-agent Learning with Diffusion ModelsabstractOffline reinforcement learning (RL) aims to learn policies from pre-existing datasets without further interactions, making it a challenging task. Q-learning algorithms struggle with extrapolation errors in offline settings, while supervised learning methods are constrained by model expressiveness. Recently, diffusion models (DMs) have shown promise in overcoming these limitations in single-agent learning, but their application in multi-agent scenarios remains unclear. Generating trajectories for each agent with independent DMs may impede coordination, while concatenating all agents’ information can lead to low sample efficiency. Accordingly, we propose MADiff, which is realized with an attention-based diffusion model to model the complex coordination among behaviors of multiple agents. To our knowledge, MADiff is the first diffusion-based multi-agent learning framework, functioning as both a decentralized policy and a centralized controller. During decentralized executions, MADiff simultaneously performs teammate modeling, and the centralized controller can also be applied in multi-agent trajectory predictions. Our experiments demonstrate that MADiff outperforms baseline algorithms across various multi-agent learning tasks, highlighting its effectiveness in modeling complex multi-agent interactions. Zhengbang Zhu, Minghuan Liu, Liyuan Mao, Bingyi Kang, Minkai Xu, Yong Yu 0001, Stefano Ermon, Weinan Zhang 0001 |
NeurIPS | 1 |
| 2024 | Understanding or Manipulation: Rethinking Online Performance Gains of Modern Recommender SystemsabstractRecommender systems are expected to be assistants that help human users find relevant information automatically without explicit queries. As recommender systems evolve, increasingly sophisticated learning techniques are applied and have achieved better performance in terms of user engagement metrics such as clicks and browsing time. The increase in the measured performance, however, can have two possible attributions: a better understanding of user preferences, and a more proactive ability to utilize human bounded rationality to seduce user over-consumption. A natural following question is whether current recommendation algorithms are manipulating user preferences. If so, can we measure the manipulation level? In this article, we present a general framework for benchmarking the degree of manipulations of recommendation algorithms, in both slate recommendation and sequential recommendation scenarios. The framework consists of four stages, initial preference calculation, training data collection, algorithm training and interaction, and metrics calculation that involves two proposed metrics, Manipulation Score and Preference Shift. We benchmark some representative recommendation algorithms in both synthetic and real-world datasets under the proposed framework. We have observed that a high online click-through rate does not necessarily mean a better understanding of user initial preference, but ends in prompting users to choose more documents they initially did not favor. Moreover, we find that the training data have notable impacts on the manipulation degrees, and algorithms with more powerful modeling abilities are more sensitive to such impacts. The experiments also verified the usefulness of the proposed metrics for measuring the degree of manipulations. We advocate that future recommendation algorithm studies should be treated as an optimization problem with constrained user preference manipulations. Zhengbang Zhu, Rongjun Qin, Xinyi Dai, Yang Yu 0001, Yong Yu 0001, Weinan Zhang 0001 |
ACM Trans. Inf. Syst. | 1 |
| 2023 | RITA: Boost Driving Simulators with Realistic Interactive Traffic FlowabstractHigh-quality traffic flow generation is the core module in building simulators for autonomous driving. However, the majority of available simulators are incapable of replicating traffic patterns that accurately reflect the various features of real-world data while also simulating human-like reactive responses to the tested autopilot driving strategies. Taking one step forward to addressing such a problem, we propose Realistic Interactive TrAffic flow (RITA) as an integrated component of existing driving simulators to provide high-quality traffic flow for the evaluation and optimization of the tested driving strategies. RITA is developed with consideration of three key features, i.e., fidelity, diversity, and controllability, and consists of two core modules called RITABackend and RITAKit. RITABackend is built to support vehicle-wise control and provide traffic generation models from real-world datasets, while RITAKit is developed with easy-to-use interfaces for controllable traffic generation via RITABackend. We demonstrate RITA’s capacity to create diversified and high-fidelity traffic simulations in several highly interactive highway scenarios. The experimental findings demonstrate that our produced RITA traffic flows exhibit all three key features, hence enhancing the completeness of driving strategy evaluation. Moreover, we showcase the possibility for further improvement of baseline strategies through online fine-tuning with RITA traffic flows. Zhengbang Zhu, Shenyu Zhang 0001, Yuzheng Zhuang, Yuecheng Liu, Minghuan Liu, Ziqing Gong, Shixiong Kai, Qiang Gu, Bin Wang 0034, Siyuan Cheng 0012, Xinyu Wang 0001, Jianye Hao, Yong Yu 0001 |
DAI | 1 |
| 2022 | Plan Your Target and Learn Your Skills: Transferable State-Only Imitation Learning via Decoupled Policy OptimizationabstractRecent progress in state-only imitation learning extends the scope of applicability of imitation learning to real-world settings by relieving the need for observing expert actions. However, existing solutions only learn to extract a state-to-action mapping policy from the data, without considering how the expert plans to the target. This hinders the ability to leverage demonstrations and limits the flexibility of the policy. In this paper, we introduce Decoupled Policy Optimization (DePO), which explicitly decouples the policy as a high-level state planner and an inverse dynamics model. With embedded decoupled policy gradient and generative adversarial training, DePO enables knowledge transfer to different action spaces or state transition dynamics, and can generalize the planner to out-of-demonstration state regions. Our in-depth experimental analysis shows the effectiveness of DePO on learning a generalized target state planner while achieving the best imitation performance. We demonstrate the appealing usage of DePO for transferring across different tasks by pre-training, and the potential for co-training agents with various skills. Minghuan Liu, Zhengbang Zhu, Yuzheng Zhuang, Weinan Zhang 0001, Jianye Hao, Yong Yu 0001, Jun Wang 0012 |
ICML | 2 |
| 2020 | Weakly-Supervised Reconstruction of 3D Objects with Large Shape Variation from Single In-the-Wild Images
Shichen Sun, Zhengbang Zhu, Xiaowei Dai, Qijun Zhao, Jing Li 0060 |
ACCV (1) | 2 |