VLDB 2026 Research / reviewers in the wild / expert
Detian Huang
dblp:214/4255
· DBLP profile ↗
16ranked-venue papers
7as first author
16since 2021 · last 2026
0000-0002-8542-3728ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 4 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Industrial visual defect detection oriented context-modulated and cross-layer interaction network for image super-resolution
Xiancheng Zhu, Detian Huang, Jianqing Zhu, Huanqiang Zeng |
Expert Syst. Appl. | 3 |
| 2026 | Lightweight image super-resolution network with adaptive token selection and feature enhancement
Detian Huang, Mingxin Lin, Xinwei Gan, Luanyuan Dai, Huanqiang Zeng |
Knowl. Based Syst. | 1 |
| 2026 | Progressive local self-attention for content-aligned super-resolution
Detian Huang, Xiancheng Zhu, Fei Shen 0004, Taotao Lai, Huanqiang Zeng, Junhui Hou |
Pattern Recognit. | 1 |
| 2026 | Selection, Aggregation, and Enhancement: Trajectory Consistent Diffusion Model for Image Super-ResolutionabstractDiffusion models have shown strong promise for image super-resolution (ISR). However, current approaches often underuse pretrained diffusion backbones and lack constraints on the sampling trajectory, which degrades structural consistency and fine details. For that, we introduce the trajectory consistent diffusion model (TCDM) for super-resolution, which jointly optimizes the sampling process through lightweight components and inference-time strategies while keeping the diffusion backbone frozen, yielding high-fidelity, detail-rich reconstructions. First, we propose a dynamic semantic selection (DSS) mechanism that records early intermediates, matches them to upsampled low-resolution features, and reconditions sampling with the best match to reduce the mismatch between conditioning and noise scale. Next, we design a cross-step aggregation guidance (CAG) strategy that aggregates features from the current state with the selected intermediate to enforce trajectory-level consistency in noise prediction. Finally, we present a plug-and-play frequency enhancement adapter (FE-Adapter) that injects different frequency-domain cues into the encoder during training, strengthening high-frequency perception while preserving global structures. Extensive experiments on multiple ISR benchmarks show that TCDM achieves strong structural fidelity and competitive no-reference perceptual quality, offering a favorable fidelity-perception trade-off. Detian Huang, Yaohui Guo, Luanyuan Dai, Fei Shen 0004, Huanqiang Zeng |
IEEE Trans. Image Process. | 1 |
| 2026 | Image Super-Resolution Using Hierarchical Cross-Scale Self-SimilarityabstractPrevious studies have revealed that extending the spatial range of informative pixels offers positive performance gains for image Super-Resolution (SR). To activate more informative pixels, considerable efforts have been devoted to exploring various variants of non-local attention mechanisms for capturing image self-similarity. However, even the state-of-the-art non-local attention mechanisms ignore an inherent property of images, namely hierarchical cross-scale self-similarity. In this paper, we propose the first Hierarchical Cross-Scale Attention (HCSA). Specifically, we first extend the search space to multiple feature maps from a single feature map, and then model cross-scale feature correspondences among different layers. This allows HCSA to activate more informative pixels for image SR by adaptively rescaling and aggregating input pixels and large-scale patches within different feature maps. To ensure accurate cross scale feature matching, we propose to replace plain down sampling operations (e.g., interpolation, pooling) with Haar Wavelet Transform (HWT) encoding, which transfers spatial information of feature maps into the channel dimension, effectively avoiding important information loss. Considering that softmax normal ization in the standard non-local attention often leads to homogeneous feature aggregation due to the amplification of small similarity weights, we propose a simple yet effective Adaptive Selection (AS) operator. This operator generates a learnable sparse mask to remove redundant features, enabling HCSA to perform discriminative feature aggregation. As a generic building block, the proposed HCSA can be flexibly integrated into existing CNN- or Transformer-based SR models, significantly strengthening cross-layer information interaction and cross-scale feature representation. Quantitative and qualitative results demonstrate that our HCSA facilitates existing SR models to achieve superior accuracy and visual quality. Xiancheng Zhu, Detian Huang, Taiheng Zeng, Xiaoqian Huang, Zhenzhen Hu 0004, Huanqiang Zeng |
IEEE Trans. Multim. | 2 |
| 2025 | EPDiff: Enhancing Prior-guided Diffusion model for Real-world Image Super-ResolutionabstractDiffusion Models (DMs) have achieved promising success in Real-world Image Super-Resolution (Real-ISR), where they reconstruct High-Resolution (HR) images from available Low-Resolution (LR) counterparts with unknown degradation by leveraging pre-trained Text-to-Image (T2I) diffusion models. However, due to the randomness nature of DMs and the severe degradation commonly presented in LR images, most DMs-based Real-ISR methods neglect the structure-level and semantic information, which results in reconstructed HR images suffering not only from important edge missing, but also from undesired regional information confusion. To tackle these challenges, we propose an Enhancing Prior-guided Diffusion model (EPDiff) for Real-ISR, which leverages high-frequency priors and semantic guidance to generate reconstructed images with realistic details. Firstly, we design a Guide Adapter (GA) module that extracts latent texture and edge features from LR images to provide high-frequency priors. Subsequently, we introduce a Semantic Prompt Extractor (SPE) that generates high-quality semantic prompts to enhance image understanding. Additionally, we build a Feature Rectify ControlNet (FRControlNet) to refine feature modulation, enabling realistic detail generation. Extensive experiments demonstrate that the proposed EPDiff outperforms state-of-the-art methods on both synthetic and real-world datasets. Detian Huang, Miaohua Ruan, Yaohui Guo, Zhenzhen Hu 0004, Huanqiang Zeng |
Comput. Vis. Image Underst. | 1 |
| 2025 | One-step diffusion for real-world image super-resolution via degradation removal and text prompts
Yaohui Guo, Luanyuan Dai, Xinwei Gan, Miaohua Ruan, Detian Huang |
Image Vis. Comput. | 6 |
| 2025 | CMASR: Lightweight image super-resolution with cluster and match attention
Detian Huang, Mingxin Lin, Huanqiang Zeng |
Image Vis. Comput. | 1 |
| 2025 | Multi-Modal Prior-Guided Diffusion Model for Blind Image Super-ResolutionabstractRecently, diffusion models have achieved remarkable success in blind image super-resolution. However, most existing methods rely solely on uni-modal degraded low-resolution images to guide diffusion models for restoring high-fidelity images, resulting in inferior realism. In this letter, we propose a Multi-modal Prior-Guided diffusion model for blind image Super-Resolution (MPGSR), which fine-tunes Stable Diffusion (SD) by utilizing the superior visual-and-textual guidance for restoring realistic high-resolution images. Specifically, our MPGSR involves two stages, i.e., multi-modal guidance extraction and adaptive guidance injection. For the former, we propose a composited transformer and further incorporate it with GPT-CLIP to extract the representative visual-and-textual guidance. For the latter, we design a feature calibration ControlNet to inject the visual guidance and employ the cross-attention layer provided by the frozen SD to inject the textual guidance, thus effectively activating the powerful text-to-image generation potential. Extensive experiments show that our MPGSR outperforms state-of-the-art methods in restoration quality and convergence time. Detian Huang, Jiaxun Song, Xiaoqian Huang, Zhenzhen Hu 0004, Huanqiang Zeng |
IEEE Signal Process. Lett. | 1 |
| 2025 | Torch-Advent-Civilization-Evolution: Accelerating Diffusion Model for Image RestorationabstractRecently, diffusion models as a hot paradigm have shown considerable superiority in image restoration with an unsupervised manner. However, they require iterative refinement from isotropic Gaussian through thousands of steps to produce a sample with exceptional quality. Most existing methods are devoted to designing fast solvers for the reverse stochastic differential equation (SDE) to accelerate sampling, while neglecting the potential of forward SDE. To better stimulate this potential, we propose the Torch-Advent-Civilization-Evolution (TACE), a novel diffusion model-based zero-shot framework for image restoration. Specifically, we propose the “Torch”, a latent vector that explicitly contains content information from the measurement image. By utilizing the Torch instead of isotropic Gaussian as initialization, our TACE significantly accelerates image restoration with better consistency and realness. To acquire the Torch from the forward process, we propose Prometheus SDEs, a cluster of equivalent SDEs. Furthermore, we construct a Conditional Guidance Projection (CGP) for the reverse SDE to strengthen the consistency of restored images. Finally, we design a Civilization Shuttle Strategy (CSS) for the generation process to enhance the realness of restored images. Extensive experiments validate that our TACE achieves state-of-the-art performance with fewer sampling steps in various typical tasks, such as super-resolution, deblurring, and colorization. Jiaxun Song, Detian Huang, Xiaoqian Huang, Miaohua Ruan, Huanqiang Zeng |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Dtsr: detail-enhanced transformer for image super-resolution
Xiaoqian Huang, Detian Huang, Caixia Huang, Zhengjun Xu |
Vis. Comput. | 2 |
| 2023 | DDFormer: Dual-domain and Dual-aggregation Transformer for Multi-contrast MRI Super-ResolutionabstractMulti-contrast high-resolution (HR) magnetic resonance (MR) images enrich available information for diagnosis and analysis. However, the resolution of MR images is typically low due to the limitations of hardware conditions and scanning time. Although convolutional neural network (CNN)-based and Transformer-based super-resolution (SR) methods are effective in improving the resolution and sharpness of images, most SR methods for multi-contrast magnetic resonance imaging (MRI) still have the following shortcomings. Firstly, most CNN-based methods are deficient in capturing global information, which is essential for regions with complicated anatomical structures. Secondly, although numerous Transformer-based methods capture long-range dependencies in the spatial dimension, they neglect the self-attention in the channel dimension, which is also important for low-level vision tasks. To address the above issues, we propose a novel dual-domain and dual-aggregation Transformer (DDFormer) for multi-contrast MRI SR. The dual-domain detail-enhanced Transformer (DDT) generates global features of the target domain; the dual-aggregation Transformer (DAT) effectively captures long-range dependencies in both spatial and channel dimensions. Specifically, we design DDT to model long-range dependencies in both reference and target images, so that the proposed DDFormer restores sharp structures and natural textures. Moreover, we propose a Channel-wise-Spatial Locally-enhanced Self-Attention layer to construct DAT to capture local features within a patch as well as global contextual features between patches in a single-channel feature and deeply aggregate semantic information of multiple protocols for MR images. Extensive experiments verify the effectiveness of our DDFormer, which outperforms state-of-the-art methods on benchmarks quantitatively and visually. Xiaoqian Huang, Jiaxun Song, Caixia Huang, Detian Huang |
BIBM | 6 |
| 2023 | CLSR: Cross-Layer Interaction Pyramid Super-Resolution NetworkabstractConvolutional Neural Network (CNN) achieves impressive success in image super-resolution (SR), where global context interaction is critical for reconstructing reliable edge and texture details. However, most CNN-based SR models focus on modeling the global contextual information within a single feature map by using attention mechanisms, and ignore the dependencies among hierarchical features, resulting in blurred or even distorted detail restoration, especially for SR tasks with large scaling factors (i.e.,$\times 4$,$\times 8$). To tackle the above issue, we propose a Cross-Layer interaction pyramid Super-Resolution (CLSR) network that reconstructs the desired SR images progressively in a coarse-to-fine fashion. Specifically, we propose a novel Cross-Layer Non-Local attention (CLNL) for accurate detail restoration. Through explicitly modeling the long-range feature-wise similarities within and between layers, the proposed CLNL is able to discriminatively explore complementary patches from hierarchical features to reconstruct the target LR patches. Then, to further strengthen the information interaction among hierarchical features at different scales, we propose a novel Gradient Consistency-Aware learning framework (GCA) by constructing a closed loop (LR$\rightarrow $HR$\rightarrow $LR) on the gradient space. The proposed GCA is able to effectively capture the interdependence between LR and HR gradient maps to guide our CLSR for reliable detail restoration. Extensive experiments validate that our CLSR outperforms the state-of-the-art methods in terms of both reconstruction accuracy and visual quality. Detian Huang, Xiancheng Zhu, Huanqiang Zeng |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Accurate visual tracking via reliable patch
Mengwei Yang, Yanming Lin, Detian Huang, Lingke Kong |
Vis. Comput. | 3 |
| 2021 | Breaking the Dilemma of Medical Image-to-image TranslationabstractSupervised Pix2Pix and unsupervised Cycle-consistency are two modes that dominate the field of medical image-to-image translation. However, neither modes are ideal. The Pix2Pix mode has excellent performance. But it requires paired and well pixel-wise aligned images, which may not always be achievable due to respiratory motion or anatomy change between times that paired images are acquired. The Cycle-consistency mode is less stringent with training data and works well on unpaired or misaligned images. But its performance may not be optimal. In order to break the dilemma of the existing modes, we propose a new unsupervised mode called RegGAN for medical image-to-image translation. It is based on the theory of "loss-correction". In RegGAN, the misaligned target images are considered as noisy labels and the generator is trained with an additional registration network to fit the misaligned noise distribution adaptively. The goal is to search for the common optimal solution to both image-to-image translation and registration tasks. We incorporated RegGAN into a few state-of-the-art image-to-image translation methods and demonstrated that RegGAN could be easily combined with these methods to improve their performances. Such as a simple CycleGAN in our mode surpasses latest NICEGAN even though using less network parameters. Based on our results, RegGAN outperformed both Pix2Pix on aligned data and Cycle-consistency on misaligned or unpaired data. RegGAN is insensitive to noises which makes it a better choice for a wide range of scenarios, especially for medical image-to-image translation tasks in which well pixel-wise aligned data are not available. Code and dataset are available at https://github.com/Kid-Liet/Reg-GAN. Lingke Kong, Chenyu Lian, Detian Huang, Yanle Hu, Qichao Zhou |
NeurIPS | 3 |
| 2021 | Multiple improved residual networks for medical image super-resolution
Defu Qiu, Lixin Zheng, Jianqing Zhu, Detian Huang |
Future Gener. Comput. Syst. | 4 |