EDBT 2026 Demo / reviewers in the wild / expert
Qiqi Kou
dblp:221/9451
· DBLP profile ↗
22ranked-venue papers
2as first author
21since 2021 · last 2026
0000-0003-2873-2636ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 1 first-author · 12 since 2021Artificial intelligence and machine learning · 8 · 1 first-author · 8 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Environment-encoder guided adaptive optimization for real-time detection and tracking in underground coal mines
Song Liang, Ruihang Liu, Jiansheng Qian, Qiqi Kou |
Eng. Appl. Artif. Intell. | 5 |
| 2026 | Lightweight image super resolution method inspired by memory consolidation mechanism
Peng Wang 0221, Yuze Wang 0009, Deqiang Cheng 0001, Qiqi Kou |
Expert Syst. Appl. | 6 |
| 2026 | Dual-consistency framework for cross-domain zero-shot image retrieval via prompt reconstruction and semantic mining
Deqiang Cheng 0001, Qiqi Kou |
Multim. Syst. | 5 |
| 2026 | Few-label blind image quality assessment via samples chosen from new and existing scenes
Deqiang Cheng 0001, Tianshu Song, Qiqi Kou, Leida Li |
Pattern Recognit. | 4 |
| 2026 | HMSR: Hypercomplex-guided mamba for fine-texture coupling in single image super-resolution
Fengqian Sun, Qiqi Kou, Deqiang Cheng 0001, Guangtao Zhai, Wenjun Zhang 0001 |
Pattern Recognit. | 5 |
| 2026 | TGCADNet: Text-Guided Context-Aware Detection via CLIP for Small Objects in UAV ScenesabstractRecent studies have highlighted the importance of contextual information for small object detection. However, existing methods rely solely on visual features and lack additional semantic guidance, which limits their ability to model key scene-level context in semantically rich, globally complex environments and to suppress irrelevant local context in densely cluttered scenes. These limitations hinder their effectiveness in Unmanned Aerial Vehicle (UAV) and similar complex scenes. To address these challenges, we propose TGCADNet (Text-Guided Context-Aware Detection Network). TGCADNet is a small object detection framework that leverages the CLIP (Contrastive Language– Image Pretraining) model’s global semantic understanding and image-text alignment capabilities for enhancing context-aware detection. TGCADNet mainly consists of Text-Guided Scene-level Context-Aware (TG-SCA) and Text-Guided Local-Context Filtering (TG-LCF). Specifically, TG-SCA uses CLIP-generated text features to guide the model in accurately extracting key scene-level context from globally complex environments. Meanwhile, TG-LCF performs interactive computation between text and image features to filter high-quality local context, thereby reducing the impact of dense and cluttered local regions in UAV scenes. We validate the effectiveness of TGCADNet on the VisDrone, UAVDT, and AI-TOD-v2 datasets. Compared to the baseline, TGCADNet achieves an improvement of 1.8 in mAP@50 and 1.3 in mAP@50:95 on the VisDrone dataset. On the UAVDT and AI-TOD-v2 datasets, TGCADNet observes improvements of 2.5 and 2.3 in mAP@50, respectively. Furthermore, TGCADNet surpasses recent SOTA methods in both accuracy and efficiency, demonstrating its effectiveness in detecting small objects in UAV and similar remote sensing scenes. Fengqian Sun, Deqiang Cheng 0001, Tianshu Song, Qiqi Kou |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2026 | Triple-Way Visual Modulation for Zero-Shot Sketch-Based Image RetrievalabstractZero-Shot Sketch-Based Image Retrieval (ZS-SBIR) seeks to correlate unseen hand-drawn sketches with unseen real images by leveraging trained models on visible categories. Recent CLIP-based models, which primarily focus on visual-textual interaction, have demonstrated strong competitiveness in ZS-SBIR. However, they still fall short in the exploration of cross-modal visual representations, especially in terms of cross-modal visual shared and specific information. Differing from the aforementioned researches, we start with class-level, prompt-level, and patch-level visual information, committed to unlocking the potential of visual feature representation. On this foundation, we introduce the vision-centric Triple-way visual modulATion (TAT) framework to enhance the model’s perception of visual shared and specific information. Specifically, we establish unified multi-modal perception by integrating visual-level modality prompter into the CLIP architecture. We then conduct triple-way modulation modeling on prompt, token, and patch levels to effectively mine shared and specific features. Lastly, we develop an enhanced calibration strategy incorporating prompt-aware, token-aware, and logit-aware alignment modules to amplify the model’s proficiency in probing shared-specific features. We thoroughly test our approach to confirm its excellence and the efficacy of individual components. The comparison results on the popular datasets Sketchy, Sketchyv2, Tuberlin, and QuickDraw show that the developed algorithm significantly surpasses the current state-of-the-art technologies. Qiqi Kou, Tianshu Song, Deqiang Cheng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | FGDepth: Fine-Grained Boundary Perception Enhancement in Self-Supervised Indoor Depth EstimationabstractSelf-supervised depth estimation has been widely ap plied in indoor environments. However, the presence of numerous objects and complex structural boundaries often leads existing methods to generate blurred or imprecise depth edges. To address this challenge, we propose FGDepth, a framework designed to enhance depth estimation through fine-grained boundary perception. Firstly, we introduce an SR (Super Resolution) auxiliary training branch that shares feature layers with the encoder of the depth estimation network. By utilizing the powerful detail recovery capabilities of the SR task, we improve the depth network's sensitivity to indoor object boundaries. Notably, the SR branch is used only during training, ensuring no added computational cost during inference. As far as we know, we are the first to employ the SR task as an auxiliary method for indoor self-supervised depth estimation. Second, we observe that outdoor scenes display significant depth variations due to their broader depth range, whereas indoor scenes typically lack clear boundary distinctions in background areas. This is because distant objects in indoor settings often appear as continuous surfaces with similar depths. This characteristic necessitates careful mitigation of background depth uniformity interference when enhancing depth boundaries in indoor scenes. Therefore, we design a depth-adaptive object mask to provide target object boundary information and use a triplet loss to align these differences with the depth map. Experimental results on three benchmark datasets show that our method outperforms existing approaches. We also conduct ablation studies to validate the contributions of each component. Chenggong Han, Chen Lv 0002, Qiqi Kou, Deqiang Cheng 0001, Stefano Mattoccia |
IEEE Trans. Multim. | 4 |
| 2025 | DMNet: Image dehazing via Dual-Domain Modulation
Qiqi Kou, Jiapeng Chen, Tianshu Song, Deqiang Cheng 0001 |
Image Vis. Comput. | 1 |
| 2025 | DCL-depth: monocular depth estimation network based on iam and depth consistency loss
Chenggong Han, Chen Lv 0002, Qiqi Kou, Deqiang Cheng 0001 |
Multim. Tools Appl. | 3 |
| 2024 | Indicative Vision Transformer for end-to-end zero-shot sketch-based image retrieval
Deqiang Cheng 0001, Qiqi Kou, Mujtaba Asad |
Adv. Eng. Informatics | 3 |
| 2024 | MMAIndoor: Patched MLP and multi-dimensional cross attention based self-supervised indoor depth estimation
Chen Lv 0002, Chenggong Han, Tianshu Song, Qiqi Kou, Jiansheng Qian, Deqiang Cheng 0001 |
Neurocomputing | 5 |
| 2024 | Intermediate-term memory mechanism inspired lightweight single image super resolution
Deqiang Cheng 0001, Yuze Wang 0009, Qiqi Kou |
Multim. Tools Appl. | 5 |
| 2024 | Task-like training paradigm in CLIP for zero-shot sketch-based image retrieval
Deqiang Cheng 0001, Qiqi Kou |
Multim. Tools Appl. | 5 |
| 2023 | Single image super-resolution based on sparse representation using edge-preserving regularization and a low-rank constraintabstractAbstract Sparse representation‐based non‐local self‐similarity approaches have demonstrated promising performance in single image super‐resolution reconstruction. This type of method, however, cannot always effectively preserve the key details of an image, resulting in edge artifacts and local structure blurring. To better preserve the image edge information, in this paper, a sparse representation super‐resolution method based on non‐local self‐similarity is proposed. First, we impose slide window gradient domain guided filtering on both the low‐resolution input image and the degraded restored image. Then, we utilize their difference as an edge‐preserving regularization term and incorporate this regularization term into the non‐local self‐similarity‐based sparse representation model to build a sparse coding model, which can enhance the restored high‐resolution image patches’ detail information. Finally, the iterative threshold algorithm is used to calculate the sparse representation coefficients so that a high‐resolution image can be estimated. Furthermore, to explore the potential structures of the subspaces spanned by similar patches, we enforce a low‐rank matrix recovery technique on the generated super‐resolution image, which can further refine the reconstruction quality. Experimental results prove that the new approach preserves the critical edge structures while suppressing noise and exceeding some popular methods both quantitatively and qualitatively. Deqiang Cheng 0001, Qiqi Kou |
IET Image Process. | 3 |
| 2023 | Self-supervised monocular Depth estimation with multi-scale structure similarity loss
Chenggong Han, Deqiang Cheng 0001, Qiqi Kou, Jiamin Zhao |
Multim. Tools Appl. | 3 |
| 2022 | H-net: Unsupervised domain adaptation person re-identification network based on hierarchy
Deqiang Cheng 0001, Jiahan Li, Qiqi Kou, Ruihang Liu |
Image Vis. Comput. | 3 |
| 2022 | Structure-Preserving and Color-Restoring Up-Sampling for Single Low-Light ImageabstractSingle image super-resolution (SR), as a basic computer vision task, has been widely studied. However, existing single image SR methods, applied in images with normal light, perform poorly on low-light images. To address this limitation, a method of structure-preserving and color-restoring up-sampling for single low-light image is proposed. Theoretical and experimental analysis indicates that existing methods suffer negative effects due to the suppressed color and the weakened texture in low-light image. Therefore, combined with the Retinex theory, the single low-light image up-sampling model is established for the first time, which avoids the conflict between color and texture by distinguishing reflectance and illumination. Further, we develop a structure-preserving and color-restoring up-sampling network for single low-light image SR. In the network, the reflectance and illumination components are obtained by decomposing the observed image, and then up-sampling of reflectance and enhancement of illumination are performed to complete the primary SR. In addition, the gradient information is reconstructed and fused into the up-sampling process to further enrich the SR texture. Experiments demonstrate that our method obtains competitive qualitative and quantitative evaluation on the produced dataset and real-world images. Deqiang Cheng 0001, Qiqi Kou |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Light-Guided and Cross-Fusion U-Net for Anti-Illumination Image Super-ResolutionabstractThe learning-based methods for single image super- resolution (SISR) can reconstruct realistic details, but they suffer severe performance degradation for low-light images because of their ignorance of negative effects of illumination, and even produce overexposure for unevenly illuminated images. In this paper, we pioneer an anti-illumination approach toward SISR named Light-guided and Cross-fusion U-Net (LCUN), which can simultaneously improve the texture details and lighting of low-resolution images. In our design, we develop a U-Net for SISR (SRU) to reconstruct super- resolution (SR) images from coarse to fine, effectively suppressing noise and absorbing illuminance information. In particular, the proposed Intensity Estimation Unit (IEU) generates the light intensity map and innovatively guides SRU to adaptively brighten inconsistent illumination. Further, aiming at efficiently utilizing key features and avoiding light interference, an Advanced Fusion Block (AFB) is developed to cross-fuse low-resolution features, reconstructed features and illuminance features in pairs. Moreover, SRU introduces a gate mechanism to dynamically adjust its composition, overcoming the limitations of fixed-scale SR. LCUN is compared with the retrained SISR methods and the combined SISR methods on low-light and uneven-light images. Extensive experiments demonstrate that LCUN advances the state-of-the-arts SISR methods in terms of objective metrics and visual effects, and it can reconstruct relatively clear textures and cope with complex lighting. Deqiang Cheng 0001, Chen Lv 0002, Qiqi Kou |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2021 | Activity guided multi-scales collaboration based on scaled-CNN for saliency prediction
Deqiang Cheng 0001, Ruihang Liu, Jiahan Li, Song Liang, Qiqi Kou |
Image Vis. Comput. | 5 |
| 2021 | Unsupervised Person Re-Identification Based on Measurement AxisabstractThe main focus of unsupervised person re-identification is the clustering of unlabeled samples in the target domain. However, most existing studies neglected to mine the deep semantic information of the target domain and did not consider a better combination of the source domain and the target domain. In this letter, we not only consider the changes of the target domain within its own domain but also mine the deep semantic information of the images by designing a measurement axis component. Then, the deep semantic information mined by the axis is used as the judgment basis of hard negative samples. Moreover, a new loss function is designed in this work to improve the migration ability of the network. Experimental results on two person re-identification domains show that our technology accuracy outperforms the state of the art by a large margin. Jiahan Li, Deqiang Cheng 0001, Ruihang Liu, Qiqi Kou |
IEEE Signal Process. Lett. | 4 |
| 2019 | Cross-Complementary Local Binary Pattern for Robust Texture ClassificationabstractThe performance of local binary pattern (LBP) and many LBP-based variants is usually limited by rotation, illumination, scale, viewpoint, and the number of training samples. In view of this, this letter presents a robust image descriptor named crosscomplementary LBP (CCLBP) for texture classification. Based on the continuous rotation invariance and highly discriminative characteristic of principal curvatures, significant local geometrical information, which is complementary to LBP is obtained. Then, the resulting information is quantized and encoded into a binary pattern. To enhance the robustness to scale, viewpoint, and the number of training samples, a multiscale and multiresolution analysis is explored by diversifying two parameters accordantly. Subsequently, a cross-scale joint feature representation is conducted on the generated complementary binary responses, resulting in the proposed CCLBP, which captures a highly discriminative information but with low dimensionality. Experimental results on three standard texture databases demonstrate that the proposed CCLBP achieves competitive performance or outperforms state-of-the-art texture descriptors while enjoying a succinct feature representation. Impressively, under the premise of maintaining complete rotation invariance, the performance of the CCLBP approach against illumination, viewpoint, and scale changes has been improved, especially when the number of training samples is limited. Qiqi Kou, Deqiang Cheng 0001, Huandong Zhuang |
IEEE Signal Process. Lett. | 1 |