Deqiang Cheng 0001

dblp:27/208-1 · DBLP profile ↗
← Back
32ranked-venue papers
7as first author
28since 2021 · last 2026
0000-0001-8831-1994ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 19 · 4 first-author · 16 since 2021Artificial intelligence and machine learning · 10 · 3 first-author · 10 since 2021Systems, architecture and hardware · 1Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Lightweight image super resolution method inspired by memory consolidation mechanism
Peng Wang 0221, Yuze Wang 0009, Deqiang Cheng 0001, Qiqi Kou
Expert Syst. Appl.5
2026 Tree-Shaped Recursive Network: Rethinking Directional Features in Lightweight Single Image Super Resolution
abstract
Single Image Super Resolution (SISR) is a post-processing technology for Internet of Things (IoT) devices that reduces bandwidth usage and improves communication efficiency. While most SISR models perform similarly on images with simple textures, their differences become more pronounced when handling complex directional textures, which are challenging to reconstruct. In view of this, we propose a Tree Recursive Network (TRN) for SISR that utilizes a tree-shaped topology and a non-local coordinate attention mechanism to capture global directional textures. First, a binary tree residual block is introduced to preserve the attention and residual features without increasing the number of parameters. Second, a non-local coordinate attention mechanism is applied to integrate non-local and directional features to effectively capture global information while reducing computational complexity. Third, a directional attention fusion technique that integrates attention and residual features is presented. Extensive experimental results show that TRN outperforms other models in terms of evaluation metrics such as PSNR and SSIM. Notably, on the ×2 Urban100 dataset, the PSNR of TRN is 0.24 dB higher than that of the popular model CFIN, which highlights its potential and usefulness.
Fudi Yi, Mujtaba Asad, Deqiang Cheng 0001
IEEE Internet Things J.7
2026 LLMDNet: An Aautonomous mining truck object detection network in low-light conditions
Feixiang Xu, Deqiang Cheng 0001, Jiansheng Qian, Fengqian Sun, Lige Xue
Knowl. Based Syst.5
2026 Dual-consistency framework for cross-domain zero-shot image retrieval via prompt reconstruction and semantic mining
Deqiang Cheng 0001, Qiqi Kou
Multim. Syst.4
2026 Few-label blind image quality assessment via samples chosen from new and existing scenes
Deqiang Cheng 0001, Tianshu Song, Qiqi Kou, Leida Li
Pattern Recognit.1
2026 HMSR: Hypercomplex-guided mamba for fine-texture coupling in single image super-resolution
Fengqian Sun, Qiqi Kou, Deqiang Cheng 0001, Guangtao Zhai, Wenjun Zhang 0001
Pattern Recognit.6
2026 TGCADNet: Text-Guided Context-Aware Detection via CLIP for Small Objects in UAV Scenes
abstract
Recent studies have highlighted the importance of contextual information for small object detection. However, existing methods rely solely on visual features and lack additional semantic guidance, which limits their ability to model key scene-level context in semantically rich, globally complex environments and to suppress irrelevant local context in densely cluttered scenes. These limitations hinder their effectiveness in Unmanned Aerial Vehicle (UAV) and similar complex scenes. To address these challenges, we propose TGCADNet (Text-Guided Context-Aware Detection Network). TGCADNet is a small object detection framework that leverages the CLIP (Contrastive Language– Image Pretraining) model’s global semantic understanding and image-text alignment capabilities for enhancing context-aware detection. TGCADNet mainly consists of Text-Guided Scene-level Context-Aware (TG-SCA) and Text-Guided Local-Context Filtering (TG-LCF). Specifically, TG-SCA uses CLIP-generated text features to guide the model in accurately extracting key scene-level context from globally complex environments. Meanwhile, TG-LCF performs interactive computation between text and image features to filter high-quality local context, thereby reducing the impact of dense and cluttered local regions in UAV scenes. We validate the effectiveness of TGCADNet on the VisDrone, UAVDT, and AI-TOD-v2 datasets. Compared to the baseline, TGCADNet achieves an improvement of 1.8 in mAP@50 and 1.3 in mAP@50:95 on the VisDrone dataset. On the UAVDT and AI-TOD-v2 datasets, TGCADNet observes improvements of 2.5 and 2.3 in mAP@50, respectively. Furthermore, TGCADNet surpasses recent SOTA methods in both accuracy and efficiency, demonstrating its effectiveness in detecting small objects in UAV and similar remote sensing scenes.
Fengqian Sun, Deqiang Cheng 0001, Tianshu Song, Qiqi Kou
IEEE Trans. Circuits Syst. Video Technol.2
2026 Triple-Way Visual Modulation for Zero-Shot Sketch-Based Image Retrieval
abstract
Zero-Shot Sketch-Based Image Retrieval (ZS-SBIR) seeks to correlate unseen hand-drawn sketches with unseen real images by leveraging trained models on visible categories. Recent CLIP-based models, which primarily focus on visual-textual interaction, have demonstrated strong competitiveness in ZS-SBIR. However, they still fall short in the exploration of cross-modal visual representations, especially in terms of cross-modal visual shared and specific information. Differing from the aforementioned researches, we start with class-level, prompt-level, and patch-level visual information, committed to unlocking the potential of visual feature representation. On this foundation, we introduce the vision-centric Triple-way visual modulATion (TAT) framework to enhance the model’s perception of visual shared and specific information. Specifically, we establish unified multi-modal perception by integrating visual-level modality prompter into the CLIP architecture. We then conduct triple-way modulation modeling on prompt, token, and patch levels to effectively mine shared and specific features. Lastly, we develop an enhanced calibration strategy incorporating prompt-aware, token-aware, and logit-aware alignment modules to amplify the model’s proficiency in probing shared-specific features. We thoroughly test our approach to confirm its excellence and the efficacy of individual components. The comparison results on the popular datasets Sketchy, Sketchyv2, Tuberlin, and QuickDraw show that the developed algorithm significantly surpasses the current state-of-the-art technologies.
Qiqi Kou, Tianshu Song, Deqiang Cheng 0001
IEEE Trans. Circuits Syst. Video Technol.6
2026 FGDepth: Fine-Grained Boundary Perception Enhancement in Self-Supervised Indoor Depth Estimation
abstract
Self-supervised depth estimation has been widely ap plied in indoor environments. However, the presence of numerous objects and complex structural boundaries often leads existing methods to generate blurred or imprecise depth edges. To address this challenge, we propose FGDepth, a framework designed to enhance depth estimation through fine-grained boundary perception. Firstly, we introduce an SR (Super Resolution) auxiliary training branch that shares feature layers with the encoder of the depth estimation network. By utilizing the powerful detail recovery capabilities of the SR task, we improve the depth network's sensitivity to indoor object boundaries. Notably, the SR branch is used only during training, ensuring no added computational cost during inference. As far as we know, we are the first to employ the SR task as an auxiliary method for indoor self-supervised depth estimation. Second, we observe that outdoor scenes display significant depth variations due to their broader depth range, whereas indoor scenes typically lack clear boundary distinctions in background areas. This is because distant objects in indoor settings often appear as continuous surfaces with similar depths. This characteristic necessitates careful mitigation of background depth uniformity interference when enhancing depth boundaries in indoor scenes. Therefore, we design a depth-adaptive object mask to provide target object boundary information and use a triplet loss to align these differences with the depth map. Experimental results on three benchmark datasets show that our method outperforms existing approaches. We also conduct ablation studies to validate the contributions of each component.
Chenggong Han, Chen Lv 0002, Qiqi Kou, Deqiang Cheng 0001, Stefano Mattoccia
IEEE Trans. Multim.5
2025 DMNet: Image dehazing via Dual-Domain Modulation
Qiqi Kou, Jiapeng Chen, Tianshu Song, Deqiang Cheng 0001
Image Vis. Comput.6
2025 DCL-depth: monocular depth estimation network based on iam and depth consistency loss
Chenggong Han, Chen Lv 0002, Qiqi Kou, Deqiang Cheng 0001
Multim. Tools Appl.5
2024 Indicative Vision Transformer for end-to-end zero-shot sketch-based image retrieval
Deqiang Cheng 0001, Qiqi Kou, Mujtaba Asad
Adv. Eng. Informatics2
2024 MMAIndoor: Patched MLP and multi-dimensional cross attention based self-supervised indoor depth estimation
Chen Lv 0002, Chenggong Han, Tianshu Song, Qiqi Kou, Jiansheng Qian, Deqiang Cheng 0001
Neurocomputing7
2024 GDM-depth: Leveraging global dependency modelling for self-supervised indoor depth estimation
Chen Lv 0002, Chenggong Han, Jochen Lang 0001, Deqiang Cheng 0001, Jiansheng Qian
Image Vis. Comput.5
2024 Using full-scale feature fusion for self-supervised indoor depth estimation
Deqiang Cheng 0001, Junhui Chen, Chen Lv 0002, Chenggong Han
Multim. Tools Appl.1
2024 Intermediate-term memory mechanism inspired lightweight single image super resolution
Deqiang Cheng 0001, Yuze Wang 0009, Qiqi Kou
Multim. Tools Appl.1
2024 Single image detail enhancement via metropolis theorem
Mujtaba Asad, Deqiang Cheng 0001
Multim. Tools Appl.5
2024 Task-like training paradigm in CLIP for zero-shot sketch-based image retrieval
Deqiang Cheng 0001, Qiqi Kou
Multim. Tools Appl.2
2024 Quality-aware blind image motion deblurring
Tianshu Song, Leida Li, Jinjian Wu, Weisheng Dong, Deqiang Cheng 0001
Pattern Recognit.5
2024 Active Learning-Based Sample Selection for Label-Efficient Blind Image Quality Assessment
abstract
Despite the considerable effort devoted to high-generalizable blind image quality assessment (BIQA), the generalization performance of the state-of-the-art metrics remains limited when facing new visual scenes. A straightforward way to address the dilemma is labeling a great number of images from the new scene and subsequently training a new model, which is quite labor-intensive and cost-expensive. Hence, there is an urgent need to mitigate the dependency on labeled samples by designing a data-efficient BIQA algorithm. Motivated by the above facts, this paper presents an Active Learning-based IQA (AL-IQA) framework, which reduces the requirement for training samples by selecting representative images from two perspectives, including distortion and content. Specifically, in terms of distortion, we design distortion prompts and adopt Contrastive Language-Image Pre-Training (CLIP) to predict image distortion in a zero-shot manner. Then, we employ curriculum learning-inspired strategy to select samples with gradually increasing difficulty (measured by prediction uncertainty of CLIP), in order to facilitate model training. Meantime, in terms of content, we adopt distribution matching-based dataset distillation to distill unlabeled images into several high-density informative synthetic images. Then, feature distances between unlabeled images and distilled images are compared to identify images with the most representative content. Finally, Borda count is adopted to capture a consensus of both distortion and content through weighted counting, and prompt tuning is utilized for adapting the model to the IQA task. Extensive experiments are conducted on five IQA datasets, and the results demonstrate that the proposed AL-IQA not only effectively reduces the number of training samples but also achieves state-of-the-art prediction accuracy and generalization performance. The source code is available athttps://github.com/esnthere/AL-IQA.
Tianshu Song, Leida Li, Deqiang Cheng 0001, Pengfei Chen 0003, Jinjian Wu
IEEE Trans. Circuits Syst. Video Technol.3
2023 Ontology-Aware Network for Zero-Shot Sketch-Based Image Retrieval
abstract
Zero-Shot Sketch-Based Image Retrieval (ZSSBIR) is an emerging task. The pioneering work focused on the modal gap but ignored inter-class information. Although recent work has begun to consider the triplet-based or contrast-based loss to mine inter-class information, positive and negative samples need to be carefully selected, or the model is prone to lose modality-specific information. To respond to these issues, an Ontology-Aware Network (OAN) is proposed. Specifically, the smooth inter-class independence learning mechanism is put forward to maintain inter-class peculiarity. Meanwhile, distillation-based consistency preservation is utilized to keep modality-specific information. Extensive experiments have demonstrated the superior performance of our algorithm on two challenging Sketchy and Tu-Berlin datasets.
Deqiang Cheng 0001
ICASSP4
2023 Single image super-resolution based on sparse representation using edge-preserving regularization and a low-rank constraint
abstract
Abstract Sparse representation‐based non‐local self‐similarity approaches have demonstrated promising performance in single image super‐resolution reconstruction. This type of method, however, cannot always effectively preserve the key details of an image, resulting in edge artifacts and local structure blurring. To better preserve the image edge information, in this paper, a sparse representation super‐resolution method based on non‐local self‐similarity is proposed. First, we impose slide window gradient domain guided filtering on both the low‐resolution input image and the degraded restored image. Then, we utilize their difference as an edge‐preserving regularization term and incorporate this regularization term into the non‐local self‐similarity‐based sparse representation model to build a sparse coding model, which can enhance the restored high‐resolution image patches’ detail information. Finally, the iterative threshold algorithm is used to calculate the sparse representation coefficients so that a high‐resolution image can be estimated. Furthermore, to explore the potential structures of the subspaces spanned by similar patches, we enforce a low‐rank matrix recovery technique on the generated super‐resolution image, which can further refine the reconstruction quality. Experimental results prove that the new approach preserves the critical edge structures while suppressing noise and exceeding some popular methods both quantitatively and qualitatively.
Deqiang Cheng 0001, Qiqi Kou
IET Image Process.2
2023 Self-supervised monocular Depth estimation with multi-scale structure similarity loss
Chenggong Han, Deqiang Cheng 0001, Qiqi Kou, Jiamin Zhao
Multim. Tools Appl.2
2022 H-net: Unsupervised domain adaptation person re-identification network based on hierarchy
Deqiang Cheng 0001, Jiahan Li, Qiqi Kou, Ruihang Liu
Image Vis. Comput.1
2022 Structure-Preserving and Color-Restoring Up-Sampling for Single Low-Light Image
abstract
Single image super-resolution (SR), as a basic computer vision task, has been widely studied. However, existing single image SR methods, applied in images with normal light, perform poorly on low-light images. To address this limitation, a method of structure-preserving and color-restoring up-sampling for single low-light image is proposed. Theoretical and experimental analysis indicates that existing methods suffer negative effects due to the suppressed color and the weakened texture in low-light image. Therefore, combined with the Retinex theory, the single low-light image up-sampling model is established for the first time, which avoids the conflict between color and texture by distinguishing reflectance and illumination. Further, we develop a structure-preserving and color-restoring up-sampling network for single low-light image SR. In the network, the reflectance and illumination components are obtained by decomposing the observed image, and then up-sampling of reflectance and enhancement of illumination are performed to complete the primary SR. In addition, the gradient information is reconstructed and fused into the up-sampling process to further enrich the SR texture. Experiments demonstrate that our method obtains competitive qualitative and quantitative evaluation on the produced dataset and real-world images.
Deqiang Cheng 0001, Qiqi Kou
IEEE Trans. Circuits Syst. Video Technol.3
2022 Light-Guided and Cross-Fusion U-Net for Anti-Illumination Image Super-Resolution
abstract
The learning-based methods for single image super- resolution (SISR) can reconstruct realistic details, but they suffer severe performance degradation for low-light images because of their ignorance of negative effects of illumination, and even produce overexposure for unevenly illuminated images. In this paper, we pioneer an anti-illumination approach toward SISR named Light-guided and Cross-fusion U-Net (LCUN), which can simultaneously improve the texture details and lighting of low-resolution images. In our design, we develop a U-Net for SISR (SRU) to reconstruct super- resolution (SR) images from coarse to fine, effectively suppressing noise and absorbing illuminance information. In particular, the proposed Intensity Estimation Unit (IEU) generates the light intensity map and innovatively guides SRU to adaptively brighten inconsistent illumination. Further, aiming at efficiently utilizing key features and avoiding light interference, an Advanced Fusion Block (AFB) is developed to cross-fuse low-resolution features, reconstructed features and illuminance features in pairs. Moreover, SRU introduces a gate mechanism to dynamically adjust its composition, overcoming the limitations of fixed-scale SR. LCUN is compared with the retrained SISR methods and the combined SISR methods on low-light and uneven-light images. Extensive experiments demonstrate that LCUN advances the state-of-the-arts SISR methods in terms of objective metrics and visual effects, and it can reconstruct relatively clear textures and cope with complex lighting.
Deqiang Cheng 0001, Chen Lv 0002, Qiqi Kou
IEEE Trans. Circuits Syst. Video Technol.1
2021 Activity guided multi-scales collaboration based on scaled-CNN for saliency prediction
Deqiang Cheng 0001, Ruihang Liu, Jiahan Li, Song Liang, Qiqi Kou
Image Vis. Comput.1
2021 Unsupervised Person Re-Identification Based on Measurement Axis
abstract
The main focus of unsupervised person re-identification is the clustering of unlabeled samples in the target domain. However, most existing studies neglected to mine the deep semantic information of the target domain and did not consider a better combination of the source domain and the target domain. In this letter, we not only consider the changes of the target domain within its own domain but also mine the deep semantic information of the images by designing a measurement axis component. Then, the deep semantic information mined by the axis is used as the judgment basis of hard negative samples. Moreover, a new loss function is designed in this work to improve the migration ability of the network. Experimental results on two person re-identification domains show that our technology accuracy outperforms the state of the art by a large margin.
Jiahan Li, Deqiang Cheng 0001, Ruihang Liu, Qiqi Kou
IEEE Signal Process. Lett.2
2020 Improved Low-Power Cost-Effective DCT Implementation Based on Markov Random Field and Stochastic Logic
abstract
Discrete Cosine Transform (DCT) is a commonly used building block for image and video compression. In this article, we present a Markov Random Field (MRF)-based design for DCT implementation because MRF logic gates outperform standard non-MRF units by achieving high noise immunity for applications to logic-based computing systems in deep sub-micron condition. Furthermore, it is found that stochastic logic, a low-cost form of number representation, can also efficiently simplify computations. By combining these two techniques, we present an improved DCT hardware circuit. The example eight-point one-dimensional DCT (1D DCT) system is simulated using 65 nm CMOS technology. Simulation results show that the proposed MRF design can achieve 13% higher noise immunity and 47% area saving, compared with the typical stochastic 1D DCT using classical Master-and-Slave architecture. While achieving the same error rate of 0.21, power consumption is reduced by 52%.
Yufeng Li 0003, I-Chyn Wey, Deqiang Cheng 0001, Fan Yang 0001, Xuan Zeng 0001, Jie Chen 0002
IEEE Trans. Circuits Syst. Video Technol.4
2019 Cross-Complementary Local Binary Pattern for Robust Texture Classification
abstract
The performance of local binary pattern (LBP) and many LBP-based variants is usually limited by rotation, illumination, scale, viewpoint, and the number of training samples. In view of this, this letter presents a robust image descriptor named crosscomplementary LBP (CCLBP) for texture classification. Based on the continuous rotation invariance and highly discriminative characteristic of principal curvatures, significant local geometrical information, which is complementary to LBP is obtained. Then, the resulting information is quantized and encoded into a binary pattern. To enhance the robustness to scale, viewpoint, and the number of training samples, a multiscale and multiresolution analysis is explored by diversifying two parameters accordantly. Subsequently, a cross-scale joint feature representation is conducted on the generated complementary binary responses, resulting in the proposed CCLBP, which captures a highly discriminative information but with low dimensionality. Experimental results on three standard texture databases demonstrate that the proposed CCLBP achieves competitive performance or outperforms state-of-the-art texture descriptors while enjoying a succinct feature representation. Impressively, under the premise of maintaining complete rotation invariance, the performance of the CCLBP approach against illumination, viewpoint, and scale changes has been improved, especially when the number of training samples is limited.
Qiqi Kou, Deqiang Cheng 0001, Huandong Zhuang
IEEE Signal Process. Lett.2
2018 Low Power Area-Efficient DCT Implementation Based on Markov Random Field-Stochastic Logic
abstract
Markov Random Field (MRF) has been adopted to achieve high noise immunity for computing systems in deep sub-micron condition. However, complete MRF designs consume large area overhead, limiting its direct hardware implementation for one-dimensional discrete cosine transform. As a low-cost number representation, stochastic logic can efficiently simplify computing circuits. By combining the two techniques, we present an MRF-based gate group design in order to achieve area and power saving with high noise immunity for stochastic adders used in discrete cosine transform. To validate the performance of our design, we implement an 8-point one-dimensional discrete cosine transform (1D-DCT) system applied the proposed design in 65 nm CMOS technology. Simulation results show that the proposed design can achieve 7% higher noise-immunity with 31% area-saving for stochastic adders and 52% power-saving, compared with the area-saving Master-and-slave stochastic 1D-DCT. The proposed design benefits outdoor sensors and biological portable devices dealing with image compression.
Yufeng Li 0003, Deqiang Cheng 0001, Jie Chen 0002
ISCAS3
2007 A Fuzzy Multicast Island Partitioning Model for Overlay Multicast Network
abstract
In order to solve the problem of partitioning nodes into MIs (Multicast Islands) in overlay multicast network, a MI partitioning model is established. The model uses RTT (Round Trip Time) of links and the network-accessing frequencies of nodes as properties to partition nodes to the nearest MSNs (Multicast Service Nodes), namely, the centres of MIs. Also, the similarity degree between nodes and relevant MSNs and the objective function of MI are designed under fuzzy condition. Given MSNs, an optimal fuzzy recognition matrix is calculated and analyzed including its algorithm. Results from the simulation verify the feasibility and validity of the model.
Deqiang Cheng 0001
ICME1