EDBT 2026 Demo / reviewers in the wild / expert
Per Mattsson
dblp:125/5370
· DBLP profile ↗
2ranked-venue papers
0as first author
2since 2021 · last 2025
0000-0002-2678-1330ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Reinforcement learning · 61% Robot manipulation · 30% Generative modeling · 9% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Robotics › Robot manipulation
diffusion policy |
0.8 | 1 | 2024 | Entropy-regularized Diffusion Policy with Q-Ensembles for Offline Reinforcement Learning · NeurIPS 2024 |
Machine learning › Reinforcement learning
offline reinforcement learning |
0.8 | 1 | 2024 | Entropy-regularized Diffusion Policy with Q-Ensembles for Offline Reinforcement Learning · NeurIPS 2024 |
Machine learning › Reinforcement learning › value-based reinforcement learning
q-value estimation |
0.8 | 1 | 2024 | Entropy-regularized Diffusion Policy with Q-Ensembles for Offline Reinforcement Learning · NeurIPS 2024 |
Machine learning › Generative modeling
diffusion model |
0.2 | 1 | 2024 | Entropy-regularized Diffusion Policy with Q-Ensembles for Offline Reinforcement Learning · NeurIPS 2024 |
Methods — techniques the papers use, named apart from their topics
stochastic differential equation · 0.8q-ensembles · 0.8entropy regularization · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Real-Time Diffusion Policies for Games: Enhancing Consistency Policies with Q-EnsemblesabstractDiffusion models have shown impressive performance in capturing complex and multi-modal action distributions for game agents, but their slow inference speed prevents practical deployment in real-time game environments. While consistency models offer a promising approach for one-step generation, they often suffer from training instability and performance degradation when applied to policy learning. In this paper, we present CPQE (Consistency Policy with Q-Ensembles), which combines consistency models with Q-ensembles to address these challenges. CPQE leverages uncertainty estimation through Q-ensembles to provide more reliable value function approximations, resulting in better training stability and improved performance compared to classic double Q-network methods. Our extensive experiments across multiple game scenarios demonstrate that CPQE achieves inference speeds of up to 60 Hz - a significant improvement over state-of-the-art diffusion policies that operate at only 20 Hz - while maintaining comparable performance to multi-step diffusion approaches. CPQE consistently outperforms state-of-theart consistency model approaches, showing both higher rewards and enhanced training stability throughout the learning process. These results indicate that CPQE offers a practical solution for deploying diffusion-based policies in games and other real-time applications where both multi-modal behavior modeling and rapid inference are critical requirements. Ruoqi Zhang, Ziwei Luo 0002, Jens Sjölund, Per Mattsson, Linus Gisslén, Alessandro Sestini |
CoG | 4 |
| 2024 | Entropy-regularized Diffusion Policy with Q-Ensembles for Offline Reinforcement LearningabstractDiffusion policy has shown a strong ability to express complex action distributions in offline reinforcement learning (RL). However, it suffers from overestimating Q-value functions on out-of-distribution (OOD) data points due to the offline dataset limitation. To address it, this paper proposes a novel entropy-regularized diffusion policy and takes into account the confidence of the Q-value prediction with Q-ensembles. At the core of our diffusion policy is a mean-reverting stochastic differential equation (SDE) that transfers the action distribution into a standard Gaussian form and then samples actions conditioned on the environment state with a corresponding reverse-time process. We show that the entropy of such a policy is tractable and that can be used to increase the exploration of OOD samples in offline RL training. Moreover, we propose using the lower confidence bound of Q-ensembles for pessimistic Q-value function estimation. The proposed approach demonstrates state-of-the-art performance across a range of tasks in the D4RL benchmarks, significantly improving upon existing diffusion-based policies. The code is available at https://github.com/ruoqizzz/entropy-offlineRL. Ruoqi Zhang, Ziwei Luo 0002, Jens Sjölund, Thomas B. Schön, Per Mattsson |
NeurIPS | 5 |