Per Mattsson

dblp:125/5370 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2025
0000-0002-2678-1330ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Reinforcement learning · 61% Robot manipulation · 30% Generative modeling · 9%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Robotics › Robot manipulation
diffusion policy
0.812024
Entropy-regularized Diffusion Policy with Q-Ensembles for Offline Reinforcement Learning · NeurIPS 2024
Machine learning › Reinforcement learning
offline reinforcement learning
0.812024
Entropy-regularized Diffusion Policy with Q-Ensembles for Offline Reinforcement Learning · NeurIPS 2024
Machine learning › Reinforcement learning › value-based reinforcement learning
q-value estimation
0.812024
Entropy-regularized Diffusion Policy with Q-Ensembles for Offline Reinforcement Learning · NeurIPS 2024
Machine learning › Generative modeling
diffusion model
0.212024
Entropy-regularized Diffusion Policy with Q-Ensembles for Offline Reinforcement Learning · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

stochastic differential equation · 0.8q-ensembles · 0.8entropy regularization · 0.8
YearPublicationVenuePosition
2025 Real-Time Diffusion Policies for Games: Enhancing Consistency Policies with Q-Ensembles
abstract
Diffusion models have shown impressive performance in capturing complex and multi-modal action distributions for game agents, but their slow inference speed prevents practical deployment in real-time game environments. While consistency models offer a promising approach for one-step generation, they often suffer from training instability and performance degradation when applied to policy learning. In this paper, we present CPQE (Consistency Policy with Q-Ensembles), which combines consistency models with Q-ensembles to address these challenges. CPQE leverages uncertainty estimation through Q-ensembles to provide more reliable value function approximations, resulting in better training stability and improved performance compared to classic double Q-network methods. Our extensive experiments across multiple game scenarios demonstrate that CPQE achieves inference speeds of up to 60 Hz - a significant improvement over state-of-the-art diffusion policies that operate at only 20 Hz - while maintaining comparable performance to multi-step diffusion approaches. CPQE consistently outperforms state-of-theart consistency model approaches, showing both higher rewards and enhanced training stability throughout the learning process. These results indicate that CPQE offers a practical solution for deploying diffusion-based policies in games and other real-time applications where both multi-modal behavior modeling and rapid inference are critical requirements.
Ruoqi Zhang, Ziwei Luo 0002, Jens Sjölund, Per Mattsson, Linus Gisslén, Alessandro Sestini
CoG4
2024 Entropy-regularized Diffusion Policy with Q-Ensembles for Offline Reinforcement Learning
abstract
Diffusion policy has shown a strong ability to express complex action distributions in offline reinforcement learning (RL). However, it suffers from overestimating Q-value functions on out-of-distribution (OOD) data points due to the offline dataset limitation. To address it, this paper proposes a novel entropy-regularized diffusion policy and takes into account the confidence of the Q-value prediction with Q-ensembles. At the core of our diffusion policy is a mean-reverting stochastic differential equation (SDE) that transfers the action distribution into a standard Gaussian form and then samples actions conditioned on the environment state with a corresponding reverse-time process. We show that the entropy of such a policy is tractable and that can be used to increase the exploration of OOD samples in offline RL training. Moreover, we propose using the lower confidence bound of Q-ensembles for pessimistic Q-value function estimation. The proposed approach demonstrates state-of-the-art performance across a range of tasks in the D4RL benchmarks, significantly improving upon existing diffusion-based policies. The code is available at https://github.com/ruoqizzz/entropy-offlineRL.
Ruoqi Zhang, Ziwei Luo 0002, Jens Sjölund, Thomas B. Schön, Per Mattsson
NeurIPS5