Qimin Cheng

dblp:118/5313 · DBLP profile ↗
← Back
17ranked-venue papers
5as first author
7since 2021 · last 2026
0000-0001-6713-7544ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 12 · 4 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 SHNet: Spectral Bias Guidance and Hierarchical Dependency Modeling Network for Camouflaged Object Detection
Qimin Cheng, Yingjie Du
MMM (1)2
2026 ChangeViT: Unleashing plain vision transformers for change detection in remote sensing images
Duowang Zhu, Xiaohu Huang, Qimin Cheng
Pattern Recognit.4
2025 Efficient Yet Effective: A Dynamic Self-Distillation Framework for Remote Sensing Image-Text Retrieval
abstract
Remote Sensing Image-Text Retrieval (RSITR) is essential for bridging the gap between heterogeneous data modalities. Recently, Vision-Language Models such as CLIP have gained popularity and become a dominant paradigm in this field. However, the high semantic similarity in remote sensing scenes introduces substantial noise into the binary supervision of contrastive learning, making domain adaptation difficult. Existing methods mitigate this by introducing large-scale additional data, but at the expense of significant training cost. To address this, we propose DSD-RSITR, a Dynamic Self-Distillation Framework that improves learning efficiency and reduces reliance on extensive data. It maintains teacher encoders for both modalities to retain prior knowledge and predict semantic relations before alignment. A dynamic EMA strategy and adaptive distillation loss support rapid teacher refinement in early training stages and stronger supervision as training progresses. Experiments show that DSD-RSITR surpasses RemoteCLIP by 1.55% and 0.29% mR on RSICD and RSITMD, respectively, despite using only 4.7% and 2.1% of the training data. Code is available at: https://github.com/zzl0107/DSD-RSITR.
Qimin Cheng, Zilei Zhou, Linfeng Yuan, Yingjie Du
IEEE Geosci. Remote. Sens. Lett.1
2025 Enhancing Road Maintenance Through Cyber-Physical Integration: The LEE-YOLO Model for Drone-Assisted Pavement Crack Detection
abstract
Drones for road crack detection provide a real-time Cyber-Physical System (CPS) solution, improving road surface assessment for efficient highway maintenance. CPS integrates computational algorithms with physical components, offering advantages over traditional manual inspections, which are slow and prone to false positives, especially in complex environments. To address these challenges, this paper presents the LEE-YOLO model, a novel and lightweight solution that leverages drones within a CPS framework to enable efficient road crack detection. The model incorporates a streamlined network built upon the YOLOv8n model and introduces an innovative lightweight fusion structure, SGC2f, which uses grouped convolutions and linear operations to reduce model weight and complexity markedly. Additionally, the proposed Efficient Bidirectional Feature Pyramid Network (EBiFPN) optimizes output channel utilization within the feature network, enhancing the model’s capacity to detect targets at multiple scales. Furthermore, the Efficient Multi-Scale Attention (EMA) module in the model’s backbone layer is designed to reduce redundant data and improve the extraction of relevant features. Experimental results demonstrate that the LEE-YOLO model outperforms YOLOv8n, achieving a 4.3% higher precision and a 4.0% improvement in mean average precision (mAP). Importantly, it achieves a 43.5% reduction in weight and a 39.5% decrease in computational demand, effectively balancing performance and efficiency. These advancements significantly enhance the potential for drone applications in highway maintenance and detection within the Cyber-Physical Systems domain.
Yingjie Du, Qimin Cheng, Xiaofeng Liu 0003, Yuwei Yi
IEEE Trans. Intell. Transp. Syst.2
2024 Zigzag Attention: A Structural Aware Module For Lane Detection
abstract
Lane detection presents a formidable challenge in the realm of autonomous driving, given the real-time processing demands, diverse acquisition conditions, and the unique elongated and angular characteristics of lane lines. While a multitude of network design strategies have been proposed to tackle this challenge, few effectively address the distinct morphology of lane lines. Approaches that leverage global relationships across all positions, e.g. attention mechanisms, are often hard to meet real-time processing requirements due to their computational complexity. To address these challenges and unique attributes of lane lines, including issues like occlusion and dashed lines, we present a specialized and plug-and-play attention module. It employs zigzag transformations to cohesively assemble spatially disparate, lane-relevant regions, thereby transforming the challenge into one of localized feature learning, which can be easily enhanced via lightweight convolutions and fully connected layers. Additionally, we harness the symmetry inherent in lane lines to bolster the learning process and enhance accuracy. Comprehensive experimentation validates the efficacy of our proposed module across a range of algorithms, demonstrating superior performance metrics, including parameters, computational complexity, and runtime, when compared to other attention approaches.
Jiajun Ling, Qimin Cheng, Xiao Huang 0003
ICASSP3
2024 MFTrans: Modality-Masked Fusion Transformer for Incomplete Multi-Modality Brain Tumor Segmentation
abstract
Brain tumor segmentation is a fundamental task and existing approaches usually rely on multi-modality magnetic resonance imaging (MRI) images for accurate segmentation. However, the common problem of missing/incomplete modalities in clinical practice would severely degrade their segmentation performance, and existing fusion strategies for incomplete multi-modality brain tumor segmentation are far from ideal. In this work, we propose a novel framework named M$^{2}$FTrans to explore and fuse cross-modality features through modality-masked fusion transformers under various incomplete multi-modality settings. Considering vanilla self-attention is sensitive to missing tokens/inputs, both learnable fusion tokens and masked self-attention are introduced to stably build long-range dependency across modalities while being more flexible to learn from incomplete modalities. In addition, to avoid being biased toward certain dominant modalities, modality-specific features are further re-weighted through spatial weight attention and channel-wise fusion transformers for feature redundancy reduction and modality re-balancing. In this way, the fusion strategy in M$^{2}$FTrans is more robust to missing modalities. Experimental results on the widely-used BraTS2018, BraTS2020, and BraTS2021 datasets demonstrate the effectiveness of M$^{2}$FTrans, outperforming the state-of-the-art approaches with large margins under various incomplete modalities for brain tumor segmentation.
Li Yu 0003, Qimin Cheng, Xin Yang 0008, Kwang-Ting Cheng, Zengqiang Yan
IEEE J. Biomed. Health Informatics3
2022 NWPU-Captions Dataset and MLCA-Net for Remote Sensing Image Captioning
abstract
Recently, the burgeoning demands for captioning-related applications have inspired great endeavors in the remote sensing community. However, current benchmark datasets are deficient in data volume, category variety, and description richness, which hinders the advancement of new remote sensing image captioning approaches, especially those based on deep learning. To overcome this limitation, we present a larger and more challenging benchmark dataset, termed NWPU-Captions. NWPU-Captions contains 157,500 sentences, with all 31,500 images annotated manually by 7 experienced volunteers. The superiority of NWPU-Captions over current publicly available benchmark datasets not only lies in its much larger scale but also in its wider coverage of complex scenes and the richness and variety of describing vocabularies. Further, a novel encoder-decoder architecture, multi-level and contextual attention network (MLCA-Net), is proposed. MLCA-Net employs a multi-level attention module to adaptively aggregate image features of specific spatial regions and scales and introduces a contextual attention module to explore the latent context hidden in remote sensing images. MLCA-Net improves the flexibility and diversity of the generated captions while keeping their accuracy and conciseness by exploring the properties of scale variations and semantic ambiguity. Finally, the effectiveness, robustness, and generalization of MLCA-Net are proved through extensive experiments on existing datasets and NWPU-Captions.
Qimin Cheng, Yuzhuo Zhou, Huanying Li, Zhongyuan Wang 0001
IEEE Trans. Geosci. Remote. Sens.1
2020 Self-learning based medical image representation for rigid real-time and multimodal slice-to-volume registration
Qiuling Gui, Qimin Cheng, Wenguang Hou, Mingyue Ding
Inf. Sci.5
2018 A survey and analysis on automatic image annotation
Qimin Cheng, Peng Fu 0004, Conghuan Tu
Pattern Recognit.1
2004 Research on embedded GIS based on wireless networks
abstract
With the rapid development of GIS, embedded technology and wireless communication technology, embedded GIS based on wireless networks becomes an active research area in the field of GIS. On the one hand, mobile embedded GIS enables users to enjoy spatial information service by wireless networks anytime and anywhere. On the other hand users can customize information at will so as to meet the requirements in different professional fields. The purpose of the paper is to propose a solution of embedded GIS based on embedded devices and wireless networks. The main contributions include a novel new system architecture and detailed function of each component. Besides, the method of data organization, the visualization of digital maps, and spatial data transmitting by wireless networks are presented. Based on our strategy, a practical system is developed to prove the validity of our strategy'. Experimental results demonstrate our strategy can meet the actual need of the visualization, querying, editing and analyzing of spatial information satisfyingly.
Feixiang Chen, Chongjun Yang, Wenyang Yu, Qimin Cheng
IGARSS4
2004 Application of M-band wavelet theory to texture analysis in content-based aerial image retrieval
abstract
With the rapid development of 3S technologies, content-based retrieval from remote sensing images provides an important means of information acquisition and sharing and has become a key technology in digital earth construction. This work puts the emphasis on using texture features for retrieval of aerial image database. M-band wavelet theory is introduced and M-band wavelet histogram technology is applied to texture extraction. Some experimental results are given to evaluate the retrieval performance of our method and to compare with that of method based on traditional two-band wavelet transform, followed by conclusions.
Qimin Cheng, Chongjun Yang, Feixiang Chen
IGARSS1
2004 An approach to terrain geometry modeling and visualization in virtual city
abstract
The technique of terrain geometry modeling is one of the most important research fields in the simulation of virtual terrain scene and virtual city. Concerned problems include 3D model, 3D data management, 3DGIS, visualization, and so on. This work firstly discusses the advantages and disadvantages of the existing methods for geometry modeling. Based on the requirements of 3D models, general 3D data structure and simplified spatial model, a new way of constructing terrain geometry models for implementing minor data storage and fast visualization is proposed. The integrated algorithm of terrain mode and 3D object is also developed in order to keep the consistency of topological relationship among geometric models. Thus the virtual city structure (terrain + 3D objects) is easily developed to store and manage the geometric and attribute data of digital city, which is the foundation of data management, visualization and spatial analysis.
Wenyang Yu, Chongjun Yang, Feixiang Chen, Qimin Cheng, Xiaoqiu Le
IGARSS4
2003 OpenGIS WMS-based prototype system of Spatial Information Search Engine
abstract
Spatial Information Search Engine (SISE), aiming to search the World Wide Web for online geographic information to meet the end users needs, is a totally new research area in the field of spatial information sharing. In this paper we introduce an OpenGIS WMS specification based SISE prototype system. This SISE prototype is capable of discovering available WMS servers dynamically, choosing qualified WMS servers according to the client request, and fetching the matched geographical information from them transparently. The system architecture, working principles, detailed function of each component, implementation strategies and system performance test results of this SISE prototype are introduced.
Chongjun Yang, Lingling Guo, Qimin Cheng
IGARSS4
2003 A prototype system of content-based retrieval of remote sensing images
abstract
The problem of content-based retrieval of remote sensing images presents a major challenge not only because of the surprisingly increasing volume of images acquired from a wide range of sensors but also because of the complexity of images themselves. In this paper, a prototype software system for content-based retrieval of remote sensing images, namely CBRRSI, is introduced. The main contribution of our research work is the novel wavelet-based feature representation approach and the flexible progressive retrieval strategies based on multi-descriptors. Some experimental results are given to prove the validity of the CBRRSI.
Qimin Cheng, Chongjun Yang
IGARSS1
2003 An effective buffer generation method in GIS
abstract
Buffer Analysis is one of the most important functions of spatial analysis in GIS. This paper uses rotation transform point formula and recursion method to further improve on the vector buffer generation algorithm of double parallel lines and circular arcs, simplifies the process of the generation of parallel lines and the circular correction of sharp angles, and finds a better solution to intersection problem of borderlines of buffer zone.
Chongjun Yang, Xiaoping Rui, Liqiang Zhang 0001, Qimin Cheng
IGARSS5
2003 Coal mine WebGIS developing with Java
abstract
Discusses the development of the SLYWGIS system, the steps on how to use Java to develop a miniature coal mine WebGIS system. Data collection and pre-treatment, system framework design and the methods of embedding the Java class into HTML language are introduced in detail.
Xiaoping Rui, Chongjun Yang, Qimin Cheng
IGARSS4
2003 A topological 3D reconstruction of complicated buildings and crossroads
abstract
The generation of 3D model for buildings and crossroads presents a challenge due to the complexity of manmade objects and lack of image understanding algorithms. In this paper, a topological strategy for semi-automatic 3D reconstruction of complicated buildings and crossroads from aerial image pair is presented based on a topology-based 3D data model in which data collection and data modeling are merged into an organic whole. Finally, a prototype software system is developed to prove the validity of the presented approach.
DeRen Li, Qimin Cheng
IGARSS3