Hongkun Dou

dblp:285/8223 · DBLP profile ↗
← Back
14ranked-venue papers
5as first author
13since 2021 · last 2026
0000-0001-6185-5369ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 4 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Constrained Particle Seeking: Solving Diffusion Inverse Problems with Just Forward Passes
abstract
Diffusion models have gained prominence as powerful generative tools for solving inverse problems due to their ability to model complex data distributions. However, existing methods typically rely on complete knowledge of the forward observation process to compute gradients for guided sampling, limiting their applicability in scenarios where such information is unavailable. In this work, we introduce *Constrained Particle Seeking (CPS)*, a novel gradient-free approach that leverages all candidate particle information to actively search for the optimal particle while incorporating constraints aligned with high-density regions of the unconditional prior. Unlike previous methods that passively select promising candidates, CPS reformulates the inverse problem as a constrained optimization task, enabling more flexible and efficient particle seeking. We demonstrate that CPS can effectively solve both image and scientific inverse problems, achieving results comparable to gradient-based methods while significantly outperforming gradient-free alternatives.
Hongkun Dou, Zike Chen, Hongjue Li, Yue Deng 0001
AAAI1
2026 You Only Look One Step: Accelerating Backpropagation in Diffusion Sampling With Gradient Shortcuts
abstract
Diffusion models (DMs) have recently demonstrated remarkable success in modeling large-scale data distributions. However, many downstream tasks require guiding the generated content based on specific differentiable metrics, typically necessitating backpropagation during the generation process. This approach is computationally expensive, as generating with DMs often demands tens to hundreds of recursive network calls, resulting in high memory usage and significant time consumption. In this paper, we propose a more efficient alternative that approaches the problem from the perspective of parallel denoising. We show that full backpropagation throughout the entire generation process is unnecessary. The downstream metrics can be optimized by retaining the computational graph of only one step during generation, thus providing a shortcut for gradient propagation. The resulting method, which we call Shortcut Diffusion Optimization (SDO), is generic, high-performance, and computationally lightweight, capable of optimizing all parameter types in diffusion sampling. We demonstrate the effectiveness of SDO on several real-world tasks, including controlling generation by optimizing latent and aligning the DMs by fine-tuning network parameters. Compared to full backpropagation, our approach reduces computational costs by $\sim\! 90\%$∼90% while maintaining superior performance. Code is available at https://github.com/deng-ai-lab/SDO.
Hongkun Dou, Xingyu Jiang 0003, Hongjue Li, Wen Yao 0001, Yue Deng 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2026 Global Modeling Matters: A Fast, Lightweight, and Effective Baseline for Efficient Image Restoration
abstract
Natural image quality is often degraded by adverse weather conditions, significantly impairing the performance of downstream tasks. Image restoration has emerged as a core solution to this challenge and has been widely discussed in the literature. Although recent transformer-based approaches have made remarkable progress in image restoration, their increasing system complexity poses significant challenges for real-time processing, particularly in real-world deployment scenarios. To this end, most existing methods attempt to simplify the self-attention mechanism, such as by channel self-attention or state space model. However, these methods primarily focus on network architecture while neglecting the inherent characteristics of image restoration itself. In this context, we explore a pyramid Wavelet-Fourier iterative pipeline to demonstrate the potential of Wavelet-Fourier processing for image restoration. Inspired by the above findings, we propose a novel and efficient restoration baseline, named Pyramid Wavelet-Fourier Network (PW-FNet). Specifically, PW-FNet features two key design principles: 1) at the inter-block level, integrates a pyramid wavelet-based multi-input multi-output structure to achieve multi-scale and multi-frequency bands decomposition; and 2) at the intra-block level, incorporates Fourier transforms as an efficient alternative to self-attention mechanisms, effectively reducing computational complexity while preserving global modeling capability. Extensive experiments on tasks such as image deraining, raindrop removal, image super-resolution, motion deblurring, image dehazing, image desnowing and underwater/low-light enhancement demonstrate that PW-FNet not only surpasses state-of-the-art methods in restoration quality but also achieves superior efficiency, with significantly reduced parameter size, computational cost and inference time. The code is available at: https://github.com/deng-ai-lab/PW-FNet.
Xingyu Jiang 0003, Ning Gao 0004, Hongkun Dou, Xiuhui Zhang, Xiaoqing Zhong, Yue Deng 0001, Hongjue Li
IEEE Trans. Image Process.3
2025 DPoser-X: Diffusion Model as Robust 3D Whole-Body Human Pose Prior
abstract
We present DPoser-X, a diffusion-based prior model for 3D whole-body human poses. Building a versatile and robust full-body human pose prior remains challenging due to the inherent complexity of articulated human poses and the scarcity of high-quality whole-body pose datasets. To address these limitations, we introduce a Diffusion model as body Pose prior (DPoser) and extend it to DPoser-X for expressive whole-body human pose modeling. Our approach unifies various pose-centric tasks as inverse problems, solving them through variational diffusion sampling. To enhance performance on downstream applications, we introduce a novel truncated timestep scheduling method specifically designed for pose data characteristics. We also propose a masked training mechanism that effectively combines whole-body and part-specific datasets, enabling our model to capture interdependencies between body parts while avoiding overfitting to specific actions. Extensive experiments demonstrate DPoser-X's robustness and versatility across multiple benchmarks for body, hand, face, and full-body pose modeling. Our model consistently outperforms state-of-the-art alternatives, establishing a new benchmark for whole-body human pose prior modeling.
Junzhe Lu 0001, Hongkun Dou, Ailing Zeng, Yue Deng 0001, Zhongang Cai, Lei Yang 0059, Yulun Zhang 0001, Haoqian Wang, Ziwei Liu 0002
ICCV3
2025 Hybrid Regularization Improves Diffusion-based Inverse Problem Solving
abstract
Diffusion models, recognized for their effectiveness as generative priors, have become essential tools for addressing a wide range of visual challenges. Recently, there has been a surge of interest in leveraging Denoising processes for Regularization (DR) to solve inverse problems. However, existing methods often face issues such as mode collapse, which results in excessive smoothing and diminished diversity. In this study, we perform a comprehensive analysis to pinpoint the root causes of gradient inaccuracies inherent in DR. Drawing on insights from diffusion model distillation, we propose a novel approach called Consistency Regularization (CR), which provides stabilized gradients without the need for ODE simulations. Building on this, we introduce Hybrid Regularization (HR), a unified framework that combines the strengths of both DR and CR, harnessing their synergistic potential. Our approach proves to be effective across a broad spectrum of inverse problems, encompassing both linear and nonlinear scenarios, as well as various measurement noise statistics. Experimental evaluations on benchmark datasets, including FFHQ and ImageNet, demonstrate that our proposed framework not only achieves highly competitive results compared to state-of-the-art methods but also offers significant reductions in wall-clock time and memory consumption.
Hongkun Dou, Jinyang Du, Wen Yao 0001, Yue Deng 0001
ICLR1
2025 Value-aligned Behavior Cloning for Offline Reinforcement Learning via Bi-level Optimization
abstract
Offline reinforcement learning (RL) aims to optimize policies under pre-collected data, without requiring any further interactions with the environment. Derived from imitation learning, Behavior cloning (BC) is extensively utilized in offline RL for its simplicity and effectiveness. Although BC inherently avoids out-of-distribution deviations, it lacks the ability to discern between high and low-quality data, potentially leading to sub-optimal performance when facing with poor-quality data. Current offline RL algorithms attempt to enhance BC by incorporating value estimation, yet often struggle to effectively balance these two critical components, specifically the alignment between the behavior policy and the pre-trained value estimations under in-sample offline data. To address this challenge, we propose the Value-aligned Behavior Cloning via Bi-level Optimization (VACO), a novel bi-level framework that seamlessly integrates an inner loop for weighted supervised behavior cloning (BC) with an outer loop dedicated to value alignment. In this framework, the inner loop employs a meta-scoring network to evaluate and appropriately weight each training sample, while the outer loop maximizes value estimation for alignment with controlled noise to facilitate limited exploration. This bi-level structure allows VACO to identify the optimal weighted BC policy, ultimately maximizing the expected estimated return conditioned on the learned value function. We conduct a comprehensive evaluation of VACO across a variety of continuous control benchmarks in offline RL, where it consistently achieves superior performance compared to existing state-of-the-art methods.
Xingyu Jiang 0003, Ning Gao 0004, Xiuhui Zhang, Hongkun Dou, Yue Deng 0001
ICLR4
2025 Physics-aligned field reconstruction with diffusion bridge
abstract
The reconstruction of physical fields from sparse measurements is pivotal in both scientific research and engineering applications. Traditional methods are increasingly supplemented by deep learning models due to their efficacy in extracting features from data. However, except for the low accuracy on complex physical systems, these models often fail to comply with essential physical constraints, such as governing equations and boundary conditions. To overcome this limitation, we introduce a novel data-driven field reconstruction framework, termed the Physics-aligned Schr\"{o}dinger Bridge (PalSB). This framework leverages a diffusion bridge mechanism that is specifically tailored to align with physical constraints. The PalSB approach incorporates a dual-stage training process designed to address both local reconstruction mapping and global physical principles. Additionally, a boundary-aware sampling technique is implemented to ensure adherence to physical boundary conditions. We demonstrate the effectiveness of PalSB through its application to three complex nonlinear systems: cylinder flow from Particle Image Velocimetry experiments, two-dimensional turbulence, and a reaction-diffusion system. The results reveal that PalSB not only achieves higher accuracy but also exhibits enhanced compliance with physical constraints compared to existing methods. This highlights PalSB's capability to generate high-quality representations of intricate physical interactions, showcasing its potential for advancing field reconstruction techniques. The source code can be found at https://github.com/lzy12301/PalSB.
Hongkun Dou, Shen Fang, Wang Han, Yue Deng 0001
ICLR2
2025 Spatiotemporal Degradation-Aware 3D Gaussian Splatting for Realistic Underwater Scene Reconstruction
abstract
Reconstructing realistic underwater scenes from underwater video remains a meaningful yet challenging task in the multimedia domain. The inherent spatiotemporal degradations in underwater imaging, including caustics, flickering, attenuation, and backscattering, frequently result in inaccurate geometry and appearance in existing 3D reconstruction methods. While a few recent works have explored underwater degradation-aware reconstruction, they often address either spatial or temporal degradation alone, falling short in more real-world underwater scenarios where both types of degradation occur. We propose MarineSTD-GS, a novel 3D Gaussian Splatting-based framework that explicitly models both temporal and spatial degradations for realistic underwater scene reconstruction. Specifically, we introduce two paired Gaussian primitives: Intrinsic Gaussians represent the true scene, while Degraded Gaussians render the degraded observations. The color of each Degraded Gaussian is physically derived from its paired Intrinsic Gaussian via a Spatiotemporal Degradation Modeling (SDM) module, enabling self-supervised disentanglement of realistic appearance from degraded images. To ensure stable training and accurate geometry, we further propose a Depth-Guided Geometry Loss and a Multi-Stage Optimization strategy. We also construct a simulated benchmark with diverse spatial and temporal degradations and ground-truth appearances for comprehensive evaluation. Experiments on both simulated and real-world datasets show that MarineSTD-GS robustly handles spatiotemporal degradations and outperforms existing methods in novel view synthesis with realistic, water-free scene appearances.
Shaohua Liu 0003, Ning Gao 0004, Zuoya Gu, Hongkun Dou, Yue Deng 0001, Hongjue Li
ACM Multimedia4
2025 Image-to-Image Bayesian Flow Networks With Structurally Informative Priors
abstract
Generative models represented by diffusion models have recently shown great potential in image generation. They usually use a reverse iteration process to map noise into the data. However, for many real-world applications such as image restoration and translation, the model input comes from a distribution that is not random noise, making it difficult for these models to adapt directly to these tasks. In this paper, we introduce Image-to-Image Bayesian Flow Networks (I2I-BFNs), a novel framework for general-purpose image-to-image translation (I2I) that operates within the parameter space of distributions. This method upholds Gaussian distributions over pixel intensities, refining distribution parameters through closed-form Bayesian inference, steered by the network's predictions for the target image. An essential aspect of our approach is the utilization of the conditional image as a robust prior parameter, initializing the translation process from a deterministic, clean image to reduce variance and produce interpretable generation. Additionally, we introduce a skip sampling technique that enhances the efficiency of I2I-BFNs, facilitating rapid translation in diverse image restoration and general I2I tasks. Our experimental evaluations showcase the model's competitive edge in various settings, underscoring its efficacy and adaptability. This work contributes new insights and opportunities for the large-scale development of efficient conditional generation systems.
Hongkun Dou, Jinyang Du, Xingyu Jiang 0003, Hongjue Li, Wen Yao 0001, Yue Deng 0001
IEEE Trans. Image Process.1
2025 SEGSID: A Semantic-Guided Framework for Sonar Image Despeckling
abstract
Sonar imagery is substantially degraded by speckle noise, making the task of despeckling crucial for improving image quality. Self-supervised despeckling methods, represented by blind-spot networks (BSNs), have shown promise in this regard. However, these methods consistently face significant challenges due to the spatial correlation of speckle noise and the inherent information loss within BSNs. In this paper, we introduce SEGSID, a BSN-based, semantic-guided sonar despeckling framework designed to address these challenges. Specifically, the SEGSID framework primarily comprises a Receptive Field Augmentation (RFA) module and a Global Semantic Enhancement (GSE) module. To address the noise spatial correlation, the RFA module is crafted to strategically extract valuable local information while avoiding the exploitation of noise-correlated pixels. Concurrently, the GSE module extracts the global semantic information from entire images and injects it into the extracted local features. This enhances BSNs' ability to harness more comprehensive image information and compensates for their inherent information loss. Furthermore, to bolster efficiency, we employ knowledge distillation techniques to transfer the expertise from the trained SEGSID into a more streamlined network suitable for broader practical applications. Extensive experiments on three distinct sonar datasets demonstrate that SEGSID outperforms both traditional despeckling methods and state-of-the-art self-supervised despeckling techniques. The implementation is publicly accessible at https://github.com/deng-ai-lab/SEGSID.
Shaohua Liu 0003, Junzhe Lu 0001, Hongkun Dou, Yue Deng 0001
IEEE Trans. Image Process.3
2025 Score-Based Neural Processes
abstract
Neural processes (NPs) have recently emerged as a powerful meta-learning framework capable of making predictions based on an arbitrary number of context points. However, the learning of NPs and their variants is hindered by the need for explicit reliance on the log-likelihood of predictive distributions, which complicates the training process. To tackle this problem, we introduce score-based NP (SNP) models, drawing inspiration from recently developed score-based generative models (SGMs) that restore data from noise by reversing a perturbation process. With denoising score matching (DSM) techniques, the SNPs bypass the intractable log-likelihood calculations, learning parameterized score functions instead. We also demonstrate that score functions possess excellent attributes that enable us to represent a wide family of conditional distributions naturally. Moreover, as data points are inherently unordered, it is crucial to incorporate appropriate inductive biases into SNPs. To this end, we propose building blocks for parameterizing permutation equivariant score functions, which induce the SNPs with the desired properties. Through extensive experimentation on both synthetic and real-world datasets, our SNPs exhibit remarkable performance and outperform existing state-of-the-art NP approaches.
Hongkun Dou, Junzhe Lu 0001, Xiaoqing Zhong, Wen Yao 0001, Yue Deng 0001
IEEE Trans. Neural Networks Learn. Syst.1
2023 Task-aware world model learning with meta weighting via bi-level optimization
abstract
Aligning the world model with the environment for the agent’s specific task is crucial in model-based reinforcement learning. While value-equivalent models may achieve better task awareness than maximum-likelihood models, they sacrifice a large amount of semantic information and face implementation issues. To combine the benefits of both types of models, we propose Task-aware Environment Modeling Pipeline with bi-level Optimization (TEMPO), a bi-level model learning framework that introduces an additional level of optimization on top of a maximum-likelihood model by incorporating a meta weighter network that weights each training sample. The meta weighter in the upper level learns to generate novel sample weights by minimizing a proposed task-aware model loss. The model in the lower level focuses on important samples while maintaining rich semantic information in state representations. We evaluate TEMPO on a variety of continuous and discrete control tasks from the DeepMind Control Suite and Atari video games. Our results demonstrate that TEMPO achieves state-of-the-art performance regarding asymptotic performance, training stability, and convergence speed.
Huining Yuan 0002, Hongkun Dou, Xingyu Jiang 0003, Yue Deng 0001
NeurIPS2
2022 Boosting Supervised Dehazing Methods via Bi-level Patch Reweighting
Xingyu Jiang 0003, Hongkun Dou, Chengwei Fu, Bingquan Dai, Tianrun Xu, Yue Deng 0001
ECCV (18)2
2020 Multi-Scale Remote Sensing Targets Detection with Rotated Feature Pyramid
abstract
For solving the difficult problem of multi-scale and multi-class target detection in complex environments of remote sensing, a target detection network is proposed based on rotated feature pyramid (RFP) and multi-scale context. Proposed method can overcome the interference caused by widely dispersed range in scale and terrain background. By extracting rotated anchors in four feature layers, the RFP module gains ample direction information to enhance plying-up target's contour. Through rotating anchors with a certain angle, RFP can decrease feature information of non-target area and avoid big scale anchor regression. Furthermore, we construct an anchor optimization method using multi-scale context which adjusts the anchor size proportion between different scales to improve the anchor selection accuracy. Experimental results on DIOR dataset demonstrate that the proposed network outperforms six state-of-the-art methods with 4.2% average precision higher. Beyond applicable to different backbones, our network has better performance for multi-class remote sensing targets.
Yinan Mao, Hongkun Dou, Danpei Zhao
IGARSS3