Tao Hu 0013

dblp:41/5865-13 · DBLP profile ↗
← Back
12ranked-venue papers
3as first author
12since 2021 · last 2026
0000-0002-4123-2962ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 11 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 HVI-CIDNet+: Beyond Extreme Darkness for Low-Light Image Enhancement
Kangbiao Shi, Yixu Feng, Tao Hu 0013, Peng Wu 0015, Guansong Pang, Qingsen Yan
IEEE Trans. Circuits Syst. Video Technol.4
2026 High Dynamic Range Imaging via Spatial-Frequency Interaction
abstract
publicly available.High Dynamic Range (HDR) imaging aims to reconstruct scenes with a wide range of luminance by fusing multi-exposure Low Dynamic Range (LDR) images. In dynamic scenes with pronounced foreground motion or camera jitter, especially under challenging conditions including extremely low or high luminance, widespread saturation, and substantial motion, existing approaches often encounter ghosting artifacts, spatial misalignment, and degradation of fine structural details. Traditional techniques based on handcrafted priors struggle to generalize to complex motion patterns, while most deep learning-based methods operate exclusively in the spatial domain, limiting their ability to capture global contextual cues and restore high-frequency structures that are better represented in the frequency domain. To address these challenges, we introduce a Dual-Domain Parallel Fusion Network with Prompt Refinement (DDPF-PR), which jointly leverages spatial and frequency-domain features for enhanced HDR reconstruction. Specifically, the proposed framework consists of a Bi-Domain Interaction Module(BDIM), which integrates spatial features for local detail and frequency features for global structure to suppress ghosting artifacts caused by motion. In addition, a Prompt Refinement Module(PRM) is designed to recover fine details in degraded regions such as saturated or misaligned areas by adaptively generating structural cues. Extensive experiments demonstrate that DDPF-PR consistently outperforms state-of-the-art methods across multiple benchmarks in both qualitative and quantitative evaluations. The code will be made publicly available.
Weiyu Zhou, Yongqing Yang, Tao Hu 0013, Pu Hui, Yu Cao 0016, Qingsen Yan, Yanning Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.3
2026 Boosting HDR Image Reconstruction via Semantic Knowledge Transfer
abstract
Recovering High Dynamic Range (HDR) images from multiple Standard Dynamic Range (SDR) images becomes challenging when the SDR images exhibit noticeable degradation and missing content. Leveraging scene-specific semantic priors offers a promising solution for restoring heavily degraded regions. However, these priors are typically extracted from sRGB SDR images, the domain/format gap poses a significant challenge when applying it to HDR imaging. To address this issue, we propose a general framework that transfers semantic knowledge derived from SDR domain via self-distillation to boost existing HDR reconstruction. Specifically, the proposed framework first introduces the Semantic Priors Guided Reconstruction Model (SPGRM), which leverages SDR image semantic knowledge to address ill-posed problems in the initial HDR reconstruction results. Subsequently, we leverage a self-distillation mechanism that constrains the color and content information with semantic knowledge, aligning the external outputs between the baseline and SPGRM. Furthermore, to transfer the semantic knowledge of the internal features, we utilize a Semantic Knowledge Alignment Module (SKAM) to fill the missing semantic contents with the complementary masks. Extensive experiments demonstrate that our framework significantly boosts HDR imaging quality for existing methods without altering the network architecture.
Tao Hu 0013, Longyao Wu, Wei Dong 0010, Peng Wu 0015, Jinqiu Sun, Xiaogang Xu 0002, Qingsen Yan, Yanning Zhang 0001
IEEE Trans. Image Process.1
2026 Ghost-Free HDR Imaging via Latent Low-Frequency Priors and Deformable Attention Alignment
abstract
Recovering ghost-free High Dynamic Range (HDR) images from multiple Low Dynamic Range (LDR) images becomes challenging when the LDR images exhibit saturation and significant motion. Recent Diffusion Models (DMs) have been introduced in HDR imaging field, showing promising performance, particularly in achieving visually perceptible better results compared to previous DNN-based methods. However, DMs require extensive iterations with large models to estimate entire images, resulting in inefficiency that hinders their practical application. To address this challenge, we propose the Low-Frequency aware Diffusion (LF-Diff) model for ghost-free HDR imaging. The key idea of LF-Diff is implementing the DMs in a highly compacted latent space and integrating it into a regression-based model to enhance the details of reconstructed images. Specifically, as low-frequency information is closely related to human visual perception we propose to utilize DMs to create compact low-frequency priors for the reconstruction process. These priors are integrated into a carefully designed Dynamic HDR Reconstruction Network (DHRNet), which employs a regression-based approach to produce high-quality HDR images. Furthermore, we introduce the Attention-guided Deformable Alignment Module (ADAM) that utilizes correlation-driven feature matching to learn deformable receptive fields for self-attention, enabling efficient pre-alignment of LDR images by focusing on salient regions. Extensive experiments on synthetic and real-world benchmark datasets demonstrate that our LF-Diff performs favorably against several state-of-the-art methods and is $10\times $ faster than previous DM-based methods.
Tao Hu 0013, Qingsen Yan, Wei Dong 0010, Peng Wu 0015, Yuankai Qi, Weisi Lin, Yanning Zhang 0001
IEEE Trans. Image Process.1
2025 Efficient Image Enhancement With a Diffusion-Based Frequency Prior
abstract
Due to the lack of appropriate priors, generating the content of dark regions remains a challenge in low-light image enhancement tasks. Currently, diffusion models employ robust image generation capabilities for enhancing low-light images. However, diffusion models require multiple iterations at the image feature level to generate details and content, which limits the speed. Moreover, the diffusion-based methods tend to generate unexpected artifacts in the degraded regions. To address these issues, we propose a Frequency Priors-guided Image Enhancement (FPIE) network, including a frequency prior generation network and an image restoration network. FPIE significantly accelerates inference by learning abstract prior with frequency domain constraints. Concretely, to learn compacted priors at the frequency domain, we introduce a joint training approach for the prior generation and restoration models to constrain the distribution of priors. Furthermore, to better utilize frequency-domain features for enhancing the network’s generation capabilities, a wavelet-based transformer block is introduced to produce intricate details and avoid the artifacts of the output. Extensive experimental results on the commonly used benchmarks demonstrate that our approach achieves state-of-the-art performances and well generalization to real-world images.
Qingsen Yan, Tao Hu 0013, Peng Wu 0015, Duwei Dai, Shuhang Gu, Wei Dong 0010, Yanning Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2025 From Dynamic to Static: Stepwisely Generate HDR Image for Ghost Removal
abstract
Generating high-quality high dynamic range (HDR) images in dynamic scenes is particularly challenging due to the influence of large motion. Despite the effectiveness of existing deep learning methods, they still suffer from ghosting artifacts when saturation and motion coexist. Inspired by fusion on static scenes, we propose an inpainting and fusion strategy to enhance the quality of the generated HDR images. The proposed method consists of pseudo-static LDR generation and detail-guided HDR generation, which creates pseudo-static images and then generates ghost-free HDR images. Specifically, the pseudo-static LDR generation network utilizes semantic information to identify the motion regions, and employs a diffusion model-based inpainting approach to produce pseudo-static LDR images that closely resemble real scenes. In the detail-guided HDR generation network, we employ a detail enhancement module to refine diverse high-frequency features with detailed information extracted from pseudo-static LDR images, which effectively enhances the visual quality. Extensive experiments on four public datasets demonstrate the superiority of the proposed method, both quantitatively and qualitatively.
Qingsen Yan, Kangzhen Yang, Tao Hu 0013, Genggeng Chen, Kexin Dai, Peng Wu 0015, Wenqi Ren, Yanning Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.3
2024 Generating Content for HDR Deghosting from Frequency View
abstract
Recovering ghost-free High Dynamic Range (HDR) images from multiple Low Dynamic Range (LDR) images becomes challenging when the LDR images exhibit saturation and significant motion. Recent Diffusion Models (DMs) have been introduced in HDR imaging field, demonstrating promising performance, particularly in achieving visually perceptible results compared to previous DNN-based methods. However, DMs require extensive iterations with large models to estimate entire images, resulting in inefficiency that hinders their practical application. To address this challenge, we propose the Low-Frequency aware Diffusion (LF-Diff) model for ghost-free HDR imaging. The key idea of LF-Diff is implementing the DMs in a highly compacted latent space and integrating it into a regression-based model to enhance the details of reconstructed images. Specifically, as low-frequency information is closely related to human visual perception we propose to utilize DMs to create compact low-frequency priors for the reconstruction process. In addition, to take full advantage of the above low-frequency priors, the Dynamic HDR Reconstruction Network (DHRNet) is carried out in a regression-based manner to obtain final HDR images. Extensive experiments conducted on synthetic and real-world benchmark datasets demonstrate that our LF-Diff performs favorably against several state-of-the-art methods and is 10x faster than previous DM-based methods.
Tao Hu 0013, Qingsen Yan, Yuankai Qi, Yanning Zhang 0001
CVPR1
2024 Efficient Content Reconstruction for High Dynamic Range Imaging
abstract
High Dynamic Range (HDR) images can be reconstructed from multiple Low Dynamic Range (LDR) images using existing deep neural network (DNN) techniques. Despite notable advancements, DNN-based methods still exhibit ghosting artifacts when handling LDR images with saturation and significant motion. Recent Diffusion models (DMs) have been introduced in HDR imaging, showcasing promising performance, especially in achieving visually perceptible results. However, DMs typically require numerous inference iterations to recover the clean image from Gaussian noise, demanding substantial computational resources. Additionally, DM only learns a probability distribution of the added noise in each step but neglects image space constraints on HDR images, limiting distortion-based metrics. To tackle these challenges, we propose an efficient network that integrates DM modules into existing regression-based models, providing reliable content reconstruction for HDR while avoiding limitations in distortion-based metrics.
Tao Hu 0013, Jiashuang He, Qingsen Yan
ICASSP2
2024 EiffHDR: An Efficient Network for Multi-Exposure High Dynamic Range Imaging
abstract
While recent progress in Multi-exposure HDR imaging is promising, the growing complexity of state-of-the-art (SOTA) methods poses challenges for their analysis and comparison. In this paper, we analyze the motivations and approaches behind previous SOTA works and introduce EiffHDR, an efficient Multi-exposure HDR imaging technique. In contrast to prior methods employing multiple branches spatial attention mechanisms, EiffHDR adopts a streamlined gating mechanism for information flow control at both spatial and channel levels, enabling implicit alignment. Subsequently, we process these features through proposed Efficient Merging Network, facilitating long-range correlations and multi-scale information perception, ultimately producing high-quality HDR images. Our experiments demonstrate that EiffHDR not only achieves outstanding performance but also significantly reduces computational complexity, making it a valuable contribution to the field.
Tao Hu 0013, Qingsen Yan
ICASSP3
2024 HL-HDR: Multi-Exposure High Dynamic Range Reconstruction with High-Low Frequency Decomposition
abstract
Generating high-quality High Dynamic Range (HDR) images in dynamic scenes is particularly challenging. Recent Transformer have been introduced in HDR imaging, demonstrating promising performance, particularly in scenarios involving large-scale motion compared to previous CNN-based methods. However, Transformer-based methods face hurdles capturing local details and come with high computational complexity, hindering further progress. In this paper, inspired by the distinct characteristics of high and low-frequency in image patterns, we propose a Frequency Decomposition Processing Block (FDPB) for ghost-free HDR imaging. In the image reconstruction process, FDPB decouples features into resolution-invariant high-frequency features and resolution-reduced low-frequency features to separately address local and global information. Specifically, considering the characteristics of different frequencies, for the high-frequency components, we design a Local Feature Extractor (LFE) based on CNN to extract local feature maps. Meanwhile, for the low-frequency components, we propose a Global Feature Extractor (GFE) that learns long-range dependencies through carefully designed Transformer modules. Importantly, the downscaled low-frequency features exploit Transformer’s remote learning capabilities while substantially reducing self-attention computational costs. By incorporating the FDPB as basic components, we further build a Low/High-Frequency Aware Network (HL-HDR), a hierarchical network to reconstruct high-quality ghost-free HDR images. Extensive experiments on four public datasets confirm the superior performance of the proposed method, both in terms of quantitative and qualitative evaluations.
Genggeng Chen, Tao Hu 0013, Kangzhen Yang, Qingsen Yan
IJCNN3
2024 Toward High-Quality HDR Deghosting With Conditional Diffusion Models
abstract
High Dynamic Range (HDR) images can be recovered from several Low Dynamic Range (LDR) images by existing Deep Neural Networks (DNNs) techniques. Despite the remarkable progress, DNN-based methods still generate ghosting artifacts when LDR images have saturation and large motion, which hinders potential applications in real-world scenarios. To address this challenge, we formulate the HDR deghosting problem as an image generation that leverages LDR features as the diffusion model’s condition, consisting of the feature condition generator and the noise predictor. Feature condition generator employs attention and Domain Feature Alignment (DFA) layer to transform the intermediate features to avoid ghosting artifacts. With the learned features as conditions, the noise predictor leverages a stochastic iterative denoising process for diffusion models to generate an HDR image by steering the sampling process. Furthermore, to mitigate semantic confusion caused by the saturation problem of LDR images, we design a sliding window noise estimator to sample smooth noise in a patch-based manner. In addition, an image space loss is proposed to avoid the color distortion of the estimated HDR results. We empirically evaluate our model on benchmark datasets for HDR imaging. The results demonstrate that our approach achieves state-of-the-art performances and well generalization to real-world images.
Qingsen Yan, Tao Hu 0013, Hao Tang 0005, Yu Zhu 0004, Wei Dong 0010, Luc Van Gool, Yanning Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2022 GaTector: A Unified Framework for Gaze Object Prediction
abstract
Gaze object prediction is a newly proposed task that aims to discover the objects being stared at by humans. It is of great application significance but still lacks a unified solution framework. An intuitive solution is to incorporate an object detection branch into an existing gaze prediction method. However, previous gaze prediction methods usually use two different networks to extract features from scene image and head image, which would lead to heavy network architecture and prevent each branch from joint optimization. In this paper, we build a novel framework named GaTector to tackle the gaze object prediction problem in a unified way. Particularly, a specific-general-specific (SGS) feature extractor is firstly proposed to utilize a shared backbone to extract general features for both scene and head images. To better consider the specificity of inputs and tasks, SGS introduces two input-specific blocks before the shared backbone and three task-specific blocks after the shared backbone. Specifically, a novel Defocus layer is designed to generate object-specific features for the object detection task without losing information or requiring extra computations. Moreover, the energy aggregation loss is introduced to guide the gaze heatmap to concentrate on the stared box. In the end, we propose a novel wUoC metric that can reveal the difference between boxes even when they share no overlapping area. Extensive experiments on the GOO dataset verify the superiority of our method in all three tracks, i.e. object detection, gaze estimation, and gaze object prediction.
Binglu Wang, Tao Hu 0013, Baoshan Li
CVPR2