Konpat Preechakul

dblp:255/6268 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
5since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Generative modeling · 51% Segmentation and scene understanding · 16% 3D vision · 14%
Computer graphics and multimedia
3 papers
Computer animation and physical simulation · 70% Image and video processing · 30%

Topics — the 11 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
1.532024
Guided Motion Diffusion for Controllable Human Motion Synthesis · ICCV 2023
Diffusion Autoencoders: Toward a Meaningful and Decodable Representation · CVPR 2022
Optimizing Diffusion Noise Can Serve As Universal Motion Priors · CVPR 2024
Computer animation and physical simulation › motion synthesis
human motion synthesis
1.422024
Optimizing Diffusion Noise Can Serve As Universal Motion Priors · CVPR 2024
Guided Motion Diffusion for Controllable Human Motion Synthesis · ICCV 2023
Computer vision › Segmentation and scene understanding
scene understanding
0.912025
Visual Jenga: Discovering Object Dependencies via Counterfactual Inpainting · NeurIPS 2025
Image and video processing › image restoration
image inpainting
0.912025
Visual Jenga: Discovering Object Dependencies via Counterfactual Inpainting · NeurIPS 2025
Machine learning › Generative modeling › diffusion model
motion diffusion
0.712023
Guided Motion Diffusion for Controllable Human Motion Synthesis · ICCV 2023
Computer animation and physical simulation › motion synthesis
controllable motion generation
0.712023
Guided Motion Diffusion for Controllable Human Motion Synthesis · ICCV 2023
Machine learning › Generative modeling › diffusion model › diffusion-based representation learning
diffusion autoencoder
0.612022
Diffusion Autoencoders: Toward a Meaningful and Decodable Representation · CVPR 2022
Machine learning › Representation and self-supervised learning › representation matching › feature alignment › embedding alignment
latent space alignment
0.512021
Set Prediction in the Latent Space · NeurIPS 2021
Computer vision › Image recognition and object detection
object detection
0.512021
Set Prediction in the Latent Space · NeurIPS 2021
Computer vision › 3D vision › geometric deep learning › set learning
set prediction
0.512021
Set Prediction in the Latent Space · NeurIPS 2021
Computer vision › 3D vision › 3d scene understanding
physical scene understanding
0.312025
Visual Jenga: Discovering Object Dependencies via Counterfactual Inpainting · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

large inpainting model · 1.7counterfactual inpainting · 1.7gradient-based latent optimization · 1.5diffusion noise optimization · 1.5feature projection · 1.3diffusion model · 1.3dense guidance · 1.3diffusion probabilistic model · 0.6autoencoding · 0.6encoding networks · 0.5
YearPublicationVenuePosition
2025 Visual Jenga: Discovering Object Dependencies via Counterfactual Inpainting
abstract
This paper proposes a novel scene understanding task called Visual Jenga. Drawing inspiration from the game Jenga, the proposed task involves progressively removing objects from a single image until only the background remains. Just as Jenga players must understand structural dependencies to maintain tower stability, our task reveals the intrinsic relationships between scene elements by systematically exploring which objects can be removed while preserving scene coherence in both physical and geometric sense. As a starting point for tackling the Visual Jenga task, we propose a simple, data-driven, training-free approach that is surprisingly effective on a range of real-world images. The principle behind our approach is to utilize the asymmetry in the pairwise relationships between objects within a scene and employ a large inpainting model to generate a set of counterfactuals to quantify the asymmetry.
Anand Bhattad, Konpat Preechakul, Alexei A. Efros
NeurIPS2
2024 Optimizing Diffusion Noise Can Serve As Universal Motion Priors
abstract
We propose Diffusion Noise Optimization (DNO), a new method that effectively leverages existing motion diffusion models as motion priors for a wide range of motion-related tasks. Instead of training a task-specific diffusion model for each new task, DNO operates by optimizing the diffusion latent noise of an existing pre-trained text-to-motion model. Given the corresponding latent noise of a human motion, it propagates the gradient from the target criteria defined on the motion space through the whole denoising process to update the diffusion latent noise. As a result, DNO supports any use cases where criteria can be defined as a function of motion. In particular, we show that, for motion editing and control, DNO outperforms existing meth-ods in both achieving the objective and preserving the motion content. DNO accommodates a diverse range of editing modes, including changing trajectory, pose, joint lo-cations, or avoiding newly added obstacles. In addition, DNO is effective in motion denoising and completion, pro-ducing smooth and realistic motion from noisy and partial inputs. DNO achieves these results at inference time with-out the need for model retraining, offering great versatility for any defined reward or loss function on the motion rep-resentation.
Korrawe Karunratanakul, Konpat Preechakul, Emre Aksan, Thabo Beeler, Supasorn Suwajanakorn, Siyu Tang 0001
CVPR2
2023 Guided Motion Diffusion for Controllable Human Motion Synthesis
abstract
Denoising diffusion models have shown great promise in human motion synthesis conditioned on natural language descriptions. However, integrating spatial constraints, such as pre-defined motion trajectories and obstacles, remains a challenge despite being essential for bridging the gap between isolated human motion and its surrounding environment. To address this issue, we propose Guided Motion Diffusion (GMD), a method that incorporates spatial constraints into the motion generation process. Specifically, we propose an effective feature projection scheme that manipulates motion representation to enhance the coherency between spatial information and local poses. Together with a new imputation formulation, the generated motion can reliably conform to spatial constraints such as global motion trajectories. Furthermore, given sparse spatial constraints (e.g. sparse keyframes), we introduce a new dense guidance approach to turn a sparse signal, which is susceptible to being ignored during the reverse steps, into denser signals to guide the generated motion to the given constraints. Our extensive experiments justify the development of GMD, which achieves a significant improvement over state-of-the-art methods in text-based motion generation while allowing control of the synthesized motions with spatial constraints.
Korrawe Karunratanakul, Konpat Preechakul, Supasorn Suwajanakorn, Siyu Tang 0001
ICCV2
2022 Diffusion Autoencoders: Toward a Meaningful and Decodable Representation
abstract
Diffusion probabilistic models (DPMs) have achieved remarkable quality in image generation that rivals GANs'. But unlike GANs, DPMs use a set of latent variables that lack semantic meaning and cannot serve as a useful representation for other tasks. This paper explores the possibility of using DPMs for representation learning and seeks to extract a meaningful and decodable representation of an input image via autoencoding. Our key idea is to use a learnable encoder for discovering the high-level semantics, and a DPM as the decoder for modeling the remaining stochastic variations. Our method can encode any image into a two-part latent code where the first part is semantically meaningful and linear, and the second part captures stochastic details, allowing near-exact reconstruction. This capability enables challenging applications that currently foil GAN-based methods, such as attribute manipulation on real images. We also show that this two-level encoding improves denoising efficiency and naturally facilitates various downstream tasks including few-shot conditional sampling. Please visit our page: https://Diff-AE.github.io/
Konpat Preechakul, Nattanat Chatthee, Suttisak Wisadwongsa, Supasorn Suwajanakorn
CVPR1
2021 Set Prediction in the Latent Space
abstract
Set prediction tasks require the matching between predicted set and ground truth set in order to propagate the gradient signal. Recent works have performed this matching in the original feature space thus requiring predefined distance functions. We propose a method for learning the distance function by performing the matching in the latent space learned from encoding networks. This method enables the use of teacher forcing which was not possible previously since matching in the feature space must be computed after the entire output sequence is generated. Nonetheless, a naive implementation of latent set prediction might not converge due to permutation instability. To address this problem, we provide sufficient conditions for permutation stability which begets an algorithm to improve the overall model convergence. Experiments on several set prediction tasks, including image captioning and object detection, demonstrate the effectiveness of our method.
Konpat Preechakul, Chawan Piansaddhayanon, Burin Naowarat, Tirasan Khandhawit, Sira Sriswasdi, Ekapol Chuangsuwanich
NeurIPS1