Detian Huang

dblp:214/4255 · DBLP profile ↗
← Back
16ranked-venue papers
7as first author
16since 2021 · last 2026
0000-0002-8542-3728ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 4 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Industrial visual defect detection oriented context-modulated and cross-layer interaction network for image super-resolution
Xiancheng Zhu, Detian Huang, Jianqing Zhu, Huanqiang Zeng
Expert Syst. Appl.3
2026 Lightweight image super-resolution network with adaptive token selection and feature enhancement
Detian Huang, Mingxin Lin, Xinwei Gan, Luanyuan Dai, Huanqiang Zeng
Knowl. Based Syst.1
2026 Progressive local self-attention for content-aligned super-resolution
Detian Huang, Xiancheng Zhu, Fei Shen 0004, Taotao Lai, Huanqiang Zeng, Junhui Hou
Pattern Recognit.1
2026 Selection, Aggregation, and Enhancement: Trajectory Consistent Diffusion Model for Image Super-Resolution
abstract
Diffusion models have shown strong promise for image super-resolution (ISR). However, current approaches often underuse pretrained diffusion backbones and lack constraints on the sampling trajectory, which degrades structural consistency and fine details. For that, we introduce the trajectory consistent diffusion model (TCDM) for super-resolution, which jointly optimizes the sampling process through lightweight components and inference-time strategies while keeping the diffusion backbone frozen, yielding high-fidelity, detail-rich reconstructions. First, we propose a dynamic semantic selection (DSS) mechanism that records early intermediates, matches them to upsampled low-resolution features, and reconditions sampling with the best match to reduce the mismatch between conditioning and noise scale. Next, we design a cross-step aggregation guidance (CAG) strategy that aggregates features from the current state with the selected intermediate to enforce trajectory-level consistency in noise prediction. Finally, we present a plug-and-play frequency enhancement adapter (FE-Adapter) that injects different frequency-domain cues into the encoder during training, strengthening high-frequency perception while preserving global structures. Extensive experiments on multiple ISR benchmarks show that TCDM achieves strong structural fidelity and competitive no-reference perceptual quality, offering a favorable fidelity-perception trade-off.
Detian Huang, Yaohui Guo, Luanyuan Dai, Fei Shen 0004, Huanqiang Zeng
IEEE Trans. Image Process.1
2026 Image Super-Resolution Using Hierarchical Cross-Scale Self-Similarity
abstract
Previous studies have revealed that extending the spatial range of informative pixels offers positive performance gains for image Super-Resolution (SR). To activate more informative pixels, considerable efforts have been devoted to exploring various variants of non-local attention mechanisms for capturing image self-similarity. However, even the state-of-the-art non-local attention mechanisms ignore an inherent property of images, namely hierarchical cross-scale self-similarity. In this paper, we propose the first Hierarchical Cross-Scale Attention (HCSA). Specifically, we first extend the search space to multiple feature maps from a single feature map, and then model cross-scale feature correspondences among different layers. This allows HCSA to activate more informative pixels for image SR by adaptively rescaling and aggregating input pixels and large-scale patches within different feature maps. To ensure accurate cross scale feature matching, we propose to replace plain down sampling operations (e.g., interpolation, pooling) with Haar Wavelet Transform (HWT) encoding, which transfers spatial information of feature maps into the channel dimension, effectively avoiding important information loss. Considering that softmax normal ization in the standard non-local attention often leads to homogeneous feature aggregation due to the amplification of small similarity weights, we propose a simple yet effective Adaptive Selection (AS) operator. This operator generates a learnable sparse mask to remove redundant features, enabling HCSA to perform discriminative feature aggregation. As a generic building block, the proposed HCSA can be flexibly integrated into existing CNN- or Transformer-based SR models, significantly strengthening cross-layer information interaction and cross-scale feature representation. Quantitative and qualitative results demonstrate that our HCSA facilitates existing SR models to achieve superior accuracy and visual quality.
Xiancheng Zhu, Detian Huang, Taiheng Zeng, Xiaoqian Huang, Zhenzhen Hu 0004, Huanqiang Zeng
IEEE Trans. Multim.2
2025 EPDiff: Enhancing Prior-guided Diffusion model for Real-world Image Super-Resolution
abstract
Diffusion Models (DMs) have achieved promising success in Real-world Image Super-Resolution (Real-ISR), where they reconstruct High-Resolution (HR) images from available Low-Resolution (LR) counterparts with unknown degradation by leveraging pre-trained Text-to-Image (T2I) diffusion models. However, due to the randomness nature of DMs and the severe degradation commonly presented in LR images, most DMs-based Real-ISR methods neglect the structure-level and semantic information, which results in reconstructed HR images suffering not only from important edge missing, but also from undesired regional information confusion. To tackle these challenges, we propose an Enhancing Prior-guided Diffusion model (EPDiff) for Real-ISR, which leverages high-frequency priors and semantic guidance to generate reconstructed images with realistic details. Firstly, we design a Guide Adapter (GA) module that extracts latent texture and edge features from LR images to provide high-frequency priors. Subsequently, we introduce a Semantic Prompt Extractor (SPE) that generates high-quality semantic prompts to enhance image understanding. Additionally, we build a Feature Rectify ControlNet (FRControlNet) to refine feature modulation, enabling realistic detail generation. Extensive experiments demonstrate that the proposed EPDiff outperforms state-of-the-art methods on both synthetic and real-world datasets.
Detian Huang, Miaohua Ruan, Yaohui Guo, Zhenzhen Hu 0004, Huanqiang Zeng
Comput. Vis. Image Underst.1
2025 One-step diffusion for real-world image super-resolution via degradation removal and text prompts
Yaohui Guo, Luanyuan Dai, Xinwei Gan, Miaohua Ruan, Detian Huang
Image Vis. Comput.6
2025 CMASR: Lightweight image super-resolution with cluster and match attention
Detian Huang, Mingxin Lin, Huanqiang Zeng
Image Vis. Comput.1
2025 Multi-Modal Prior-Guided Diffusion Model for Blind Image Super-Resolution
abstract
Recently, diffusion models have achieved remarkable success in blind image super-resolution. However, most existing methods rely solely on uni-modal degraded low-resolution images to guide diffusion models for restoring high-fidelity images, resulting in inferior realism. In this letter, we propose a Multi-modal Prior-Guided diffusion model for blind image Super-Resolution (MPGSR), which fine-tunes Stable Diffusion (SD) by utilizing the superior visual-and-textual guidance for restoring realistic high-resolution images. Specifically, our MPGSR involves two stages, i.e., multi-modal guidance extraction and adaptive guidance injection. For the former, we propose a composited transformer and further incorporate it with GPT-CLIP to extract the representative visual-and-textual guidance. For the latter, we design a feature calibration ControlNet to inject the visual guidance and employ the cross-attention layer provided by the frozen SD to inject the textual guidance, thus effectively activating the powerful text-to-image generation potential. Extensive experiments show that our MPGSR outperforms state-of-the-art methods in restoration quality and convergence time.
Detian Huang, Jiaxun Song, Xiaoqian Huang, Zhenzhen Hu 0004, Huanqiang Zeng
IEEE Signal Process. Lett.1
2025 Torch-Advent-Civilization-Evolution: Accelerating Diffusion Model for Image Restoration
abstract
Recently, diffusion models as a hot paradigm have shown considerable superiority in image restoration with an unsupervised manner. However, they require iterative refinement from isotropic Gaussian through thousands of steps to produce a sample with exceptional quality. Most existing methods are devoted to designing fast solvers for the reverse stochastic differential equation (SDE) to accelerate sampling, while neglecting the potential of forward SDE. To better stimulate this potential, we propose the Torch-Advent-Civilization-Evolution (TACE), a novel diffusion model-based zero-shot framework for image restoration. Specifically, we propose the “Torch”, a latent vector that explicitly contains content information from the measurement image. By utilizing the Torch instead of isotropic Gaussian as initialization, our TACE significantly accelerates image restoration with better consistency and realness. To acquire the Torch from the forward process, we propose Prometheus SDEs, a cluster of equivalent SDEs. Furthermore, we construct a Conditional Guidance Projection (CGP) for the reverse SDE to strengthen the consistency of restored images. Finally, we design a Civilization Shuttle Strategy (CSS) for the generation process to enhance the realness of restored images. Extensive experiments validate that our TACE achieves state-of-the-art performance with fewer sampling steps in various typical tasks, such as super-resolution, deblurring, and colorization.
Jiaxun Song, Detian Huang, Xiaoqian Huang, Miaohua Ruan, Huanqiang Zeng
IEEE Trans. Circuits Syst. Video Technol.2
2024 Dtsr: detail-enhanced transformer for image super-resolution
Xiaoqian Huang, Detian Huang, Caixia Huang, Zhengjun Xu
Vis. Comput.2
2023 DDFormer: Dual-domain and Dual-aggregation Transformer for Multi-contrast MRI Super-Resolution
abstract
Multi-contrast high-resolution (HR) magnetic resonance (MR) images enrich available information for diagnosis and analysis. However, the resolution of MR images is typically low due to the limitations of hardware conditions and scanning time. Although convolutional neural network (CNN)-based and Transformer-based super-resolution (SR) methods are effective in improving the resolution and sharpness of images, most SR methods for multi-contrast magnetic resonance imaging (MRI) still have the following shortcomings. Firstly, most CNN-based methods are deficient in capturing global information, which is essential for regions with complicated anatomical structures. Secondly, although numerous Transformer-based methods capture long-range dependencies in the spatial dimension, they neglect the self-attention in the channel dimension, which is also important for low-level vision tasks. To address the above issues, we propose a novel dual-domain and dual-aggregation Transformer (DDFormer) for multi-contrast MRI SR. The dual-domain detail-enhanced Transformer (DDT) generates global features of the target domain; the dual-aggregation Transformer (DAT) effectively captures long-range dependencies in both spatial and channel dimensions. Specifically, we design DDT to model long-range dependencies in both reference and target images, so that the proposed DDFormer restores sharp structures and natural textures. Moreover, we propose a Channel-wise-Spatial Locally-enhanced Self-Attention layer to construct DAT to capture local features within a patch as well as global contextual features between patches in a single-channel feature and deeply aggregate semantic information of multiple protocols for MR images. Extensive experiments verify the effectiveness of our DDFormer, which outperforms state-of-the-art methods on benchmarks quantitatively and visually.
Xiaoqian Huang, Jiaxun Song, Caixia Huang, Detian Huang
BIBM6
2023 CLSR: Cross-Layer Interaction Pyramid Super-Resolution Network
abstract
Convolutional Neural Network (CNN) achieves impressive success in image super-resolution (SR), where global context interaction is critical for reconstructing reliable edge and texture details. However, most CNN-based SR models focus on modeling the global contextual information within a single feature map by using attention mechanisms, and ignore the dependencies among hierarchical features, resulting in blurred or even distorted detail restoration, especially for SR tasks with large scaling factors (i.e.,$\times 4$,$\times 8$). To tackle the above issue, we propose a Cross-Layer interaction pyramid Super-Resolution (CLSR) network that reconstructs the desired SR images progressively in a coarse-to-fine fashion. Specifically, we propose a novel Cross-Layer Non-Local attention (CLNL) for accurate detail restoration. Through explicitly modeling the long-range feature-wise similarities within and between layers, the proposed CLNL is able to discriminatively explore complementary patches from hierarchical features to reconstruct the target LR patches. Then, to further strengthen the information interaction among hierarchical features at different scales, we propose a novel Gradient Consistency-Aware learning framework (GCA) by constructing a closed loop (LR$\rightarrow $HR$\rightarrow $LR) on the gradient space. The proposed GCA is able to effectively capture the interdependence between LR and HR gradient maps to guide our CLSR for reliable detail restoration. Extensive experiments validate that our CLSR outperforms the state-of-the-art methods in terms of both reconstruction accuracy and visual quality.
Detian Huang, Xiancheng Zhu, Huanqiang Zeng
IEEE Trans. Circuits Syst. Video Technol.1
2022 Accurate visual tracking via reliable patch
Mengwei Yang, Yanming Lin, Detian Huang, Lingke Kong
Vis. Comput.3
2021 Breaking the Dilemma of Medical Image-to-image Translation
abstract
Supervised Pix2Pix and unsupervised Cycle-consistency are two modes that dominate the field of medical image-to-image translation. However, neither modes are ideal. The Pix2Pix mode has excellent performance. But it requires paired and well pixel-wise aligned images, which may not always be achievable due to respiratory motion or anatomy change between times that paired images are acquired. The Cycle-consistency mode is less stringent with training data and works well on unpaired or misaligned images. But its performance may not be optimal. In order to break the dilemma of the existing modes, we propose a new unsupervised mode called RegGAN for medical image-to-image translation. It is based on the theory of "loss-correction". In RegGAN, the misaligned target images are considered as noisy labels and the generator is trained with an additional registration network to fit the misaligned noise distribution adaptively. The goal is to search for the common optimal solution to both image-to-image translation and registration tasks. We incorporated RegGAN into a few state-of-the-art image-to-image translation methods and demonstrated that RegGAN could be easily combined with these methods to improve their performances. Such as a simple CycleGAN in our mode surpasses latest NICEGAN even though using less network parameters. Based on our results, RegGAN outperformed both Pix2Pix on aligned data and Cycle-consistency on misaligned or unpaired data. RegGAN is insensitive to noises which makes it a better choice for a wide range of scenarios, especially for medical image-to-image translation tasks in which well pixel-wise aligned data are not available. Code and dataset are available at https://github.com/Kid-Liet/Reg-GAN.
Lingke Kong, Chenyu Lian, Detian Huang, Yanle Hu, Qichao Zhou
NeurIPS3
2021 Multiple improved residual networks for medical image super-resolution
Defu Qiu, Lixin Zheng, Jianqing Zhu, Detian Huang
Future Gener. Comput. Syst.4