EDBT 2026 Demo / reviewers in the wild / expert
Guowei Zou
dblp:414/8367
· DBLP profile ↗
1ranked-venue papers
1as first author
1since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Motion planning and robot control · 30% Robot manipulation · 30% Reinforcement learning · 30% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Robotics › Robot manipulation
diffusion policy |
1.0 | 1 | 2026 | D²PPO: Diffusion Policy Policy Optimization with Dispersive Loss · AAAI 2026 |
Machine learning › Reinforcement learning
policy optimization |
1.0 | 1 | 2026 | D²PPO: Diffusion Policy Policy Optimization with Dispersive Loss · AAAI 2026 |
Robotics › Motion planning and robot control
robot learning |
1.0 | 1 | 2026 | D²PPO: Diffusion Policy Policy Optimization with Dispersive Loss · AAAI 2026 |
Machine learning › Representation and self-supervised learning › representation analysis
representation collapse |
0.3 | 1 | 2026 | D²PPO: Diffusion Policy Policy Optimization with Dispersive Loss · AAAI 2026 |
Methods — techniques the papers use, named apart from their topics
policy optimization · 1.0dispersive loss regularization · 1.0diffusion policy · 1.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | D²PPO: Diffusion Policy Policy Optimization with Dispersive LossabstractDiffusion policies excel at robotic manipulation by naturally modeling multimodal action distributions in high-dimensional spaces. Nevertheless, diffusion policies suffer from diffusion representation collapse: semantically similar observations are mapped to indistinguishable features, ultimately impairing their ability to handle subtle but critical variations required for complex robotic manipulation. To address this problem, we propose D²PPO (Diffusion Policy Policy Optimization with Dispersive Loss). D²PPO introduces dispersive loss regularization that combats representation collapse by treating all hidden representations within each batch as negative pairs. D²PPO compels the network to learn discriminative representations of similar observations, thereby enabling the policy to identify subtle yet crucial differences necessary for precise manipulation. In evaluation, we find that early-layer regularization benefits simple tasks, while late-layer regularization sharply enhances performance on complex manipulation tasks. On RoboMimic benchmarks, D²PPO achieves an average improvement of 22.7% in pre-training and 26.1% after fine-tuning, setting new SOTA results. In comparison with SOTA, the results of real-world experiments on a Franka Emika Panda robot show the excitingly high success rate of our method. The superiority of our method is especially evident in complex tasks. Guowei Zou, Weibing Li, Hejun Wu, Yukun Qian |
AAAI | 1 |