VLDB 2026 Research / reviewers in the wild / expert
Yi Li 0054
dblp:59/871-54
· DBLP profile ↗
10ranked-venue papers
1as first author
10since 2021 · last 2025
0000-0003-4226-6635ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | FASTCC: A lightweight human pose detection method leveraging SimCCabstractAs a focal area within computer vision algorithms, human pose estimation algorithms find applications in diverse fields such as security and virtual reality. Achieving a balance between speed and accuracy is imperative for practical applications. Existing methods often present a trade-off between high accuracy and real-time performance. In response, this paper introduces the Fast Coordinate Classification (FastCC) detection head. It employs a shared fully connected Transformer for global self-attention operations on feature layers from the backbone network. The spatial attention coordinate encoder then outputs the coordinates of the horizontal and vertical axes of the keypoints, which are subsequently combined to derive the actual keypoint positions. Experimental evaluations conducted on the COCO and MPII datasets demonstrate that our detection head enhances the accuracy of human pose estimation algorithms while maintaining a lightweight design, outperforming the traditional heatmap method. Yi Li 0054, Yongtao Wang, Dou Quan, Yabo Yan, Qinghai Yang |
Neurocomputing | 1 |
| 2024 | LM-Net: A Lightweight Matching Network for Remote Sensing Image Matching and RegistrationabstractDeep feature learning methods have shown significant advantages over handcrafted feature-based methods in remote sensing image matching and registration. Existing deep learning methods usually introduce complex modules into the deep convolutional network for more robust feature learning. However, they usually require high computation and memory resources for the computing device and have expensive time costs for image registration. As a basic image-processing task, it is crucial to build a lightweight matching network (LM-Net) for fast and accurate image matching and registration. Unfortunately, the image-matching performance will decrease significantly when we directly compress the deep model to a lightweight one. This article proposes an LM-Net based on the knowledge distillation (KD) learning framework for remote sensing image matching and registration. We first build an LM-Net with three convolutional layers. Then, this article proposes an effective KD approach for network optimization, which transfers the effective knowledge from the deep matching network to LM-Net to improve image-matching performances. Specifically, this article considers the useful information in the instance samples and the relation information between samples. It designs the feature and feature relation distillation learning for LM-Net training. Extensive experimental results and analysis have shown the effectiveness and advantages of the proposed LM-Net. LM-Net can reduce the number of parameters and computational complexity of the matching network. Meanwhile, LM-Net can significantly decrease the time cost and achieve results comparable to those of the deep model. It reduces the average image registration time by 42% on remote sensing image matching and registration. Additionally, LM-Net generalizes well on other multimodal remote sensing images. Dou Quan, Chonghua Lv, Shuang Wang 0001, Yi Li 0054, Bo Ren 0001, Jocelyn Chanussot, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | A Unified Deep Learning Network for Remote Sensing Image Registration and Change DetectionabstractImage registration and change detection are crucial for multitemporal remote sensing image analysis. The images should be registered before the change information detection. Existing deep learning methods have shown significant advantages in image registration and change detection tasks. They usually design two independent task-specific deep networks for image registration and change detection, respectively. These independent deep networks will learn from scratch and rely on many task-specific labeled training datasets. This article finds that image registration and change detection have similar learning mechanisms, which focus on extracting discriminative features. Inspired by this, we propose a Unified image Registration and Change detection Network (URCNet) that can perform image alignment and change information detection through a single network. Additionally, this article proposes various deep collaborative learning methods for URCNet optimization, which enforce that the URCNet can effectively support remote sensing image registration and change detection simultaneously. Extensive experiments demonstrate the effectiveness of the proposed URCNet for image registration and change detection, which can achieve comparable and better results with task-specific and more complex deep networks. The proposed URCNet can support multitasks based on the same scene images, different scene images, and even multimodal images. Moreover, URCNet shows significant advantages over other deep networks in change detection under limited labeled datasets. Rufan Zhou, Dou Quan, Shuang Wang 0001, Chonghua Lv, Xianwei Cao, Jocelyn Chanussot, Yi Li 0054, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2024 | A Concurrent Multiscale Detector for End-to-End Image MatchingabstractThis article focuses on end-to-end image matching through joint key-point detection and descriptor extraction. To find repeatable and high discrimination key points, we improve the deep matching network from the perspectives of network structure and network optimization. First, we propose a concurrent multiscale detector (CS-det) network, which consists of several parallel convolutional networks to extract multiscale features and multilevel discriminative information for key-point detection. Moreover, we introduce an attention module to fuse the response maps of various features adaptively. Importantly, we propose two novel rank consistent losses (RC-losses) for network optimization, significantly improving image matching performances. On the one hand, we propose a score rank consistent loss (RC-S-loss) to ensure that the key points have high repeatability. Different from the score difference loss merely focusing on the absolute score of an individual key point, our proposed RC-S-loss pays more attention to the relative score of key points in the image. On the other hand, we propose a score-discrimination RC-loss to ensure that the key point has high discrimination, which can reduce the confusion from other key points in subsequent matching and then further enhance the accuracy of image matching. Extensive experimental results demonstrate that the proposed CS-det improves the mean matching result of deep detector by 1.4%-2.1%, and the proposed RC-losses can boost the matching performances by 2.7%-3.4% than score difference loss. Our source codes are available at https://github.com/iquandou/CS-Net. Dou Quan, Shuang Wang 0001, Ning Huyan, Yi Li 0054, Ruiqi Lei, Jocelyn Chanussot, Biao Hou, Licheng Jiao |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2022 | Deep Modality Independent Descriptor Learning for Optical and SAR Image Patch MatchingabstractDue to the complementary information between multi-modal images, they are widely used in various applications. However, there are significant differences in appearance caused by different imaging mechanisms, which bring great challenges to multi-modal image patch matching. To solve this problem, this paper proposes a deep modality independent descriptor learning network (DMID-Net) for multi-modal image patch matching. DMID-Net computes the self-similarity of deep features as the structure descriptor for image patch matching, which is independent of image modality and shared between multi-modal images. Thus, the acquired deep modality independent descriptor(DMID) can reduce the influence of significant differences between multi-modal images, further improving the matching performances. Experimental results on a large number of optical and SAR image-pairs demonstrate the effectiveness of DMID-Net on multi-modal image patch matching. Huiyuan Wei, Dou Quan, Ruiqi Lei, Baorui Duan, Shuang Wang 0001, Yi Li 0054, Biao Hou, Licheng Jiao |
IGARSS | 6 |
| 2022 | Self-Distillation Feature Learning Network for Optical and SAR Image RegistrationabstractOptical and SAR image registration is important for multi-modal remote sensing image information fusion. Recently, deep matching networks have shown better performances than traditional methods on image matching. However, due to significant differences between optical and SAR images, the performances of existing deep learning methods still need to be further improved. This paper proposes a self-distillation feature learning network (SDNet) for optical and SAR image registration, improving performance from network structure and network optimization. Firstly, we explore the impact of different weight-sharing strategies on optical and SAR image matching. Then, we design a partially unshared feature learning network for multi-modal image feature learning. It has fewer parameters than the fully unshared network and has more flexibility than the fully shared network. Additionally, the limited binary supervised information (matching or non-matching) is insufficient to train the deep matching networks for optical-SAR image registration. Thus, we propose a self-distillation feature learning method to exploit more similarity information for deep network optimization enhancing, such as the similarity ordering between a series of non-matching patch-pairs. The exploited rich similarity information will significantly enhance network training and improve matching accuracy. Finally, considering that existing deep learning methods brute-force constrain the features of the matching optical and SAR image patches are similar, which will be lost many discriminative information, degenerating matching performances. Thus, we build an auxiliary task reconstruction learning to optimize the feature learning network to keep more discriminative information. Extensive experiments demonstrate the effectiveness of our proposed method on multi-modal image registration. Dou Quan, Huiyuan Wei, Shuang Wang 0001, Ruiqi Lei, Baorui Duan, Yi Li 0054, Biao Hou, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2022 | Cluster Alignment With Target Knowledge Mining for Unsupervised Domain Adaptation Semantic SegmentationabstractUnsupervised domain adaptation (UDA) carries out knowledge transfer from the labeled source domain to the unlabeled target domain. Existing feature alignment methods in UDA semantic segmentation achieve this goal by aligning the feature distribution between domains. However, these feature alignment methods ignore the domain-specific knowledge of the target domain. In consequence, 1) the correlation among pixels of the target domain is not explored; and 2) the classifier is not explicitly designed for the target domain distribution. To conquer these obstacles, we propose a novel cluster alignment framework, which mines the domain-specific knowledge when performing the alignment. Specifically, we design a multi-prototype clustering strategy to make the pixel features within the same class tightly distributed for the target domain. Subsequently, a contrastive strategy is developed to align the distributions between domains, with the clustered structure maintained. After that, a novel affinity-based normalized cut loss is devised to learn task-specific decision boundaries. Our method enhances the model's adaptability in the target domain, and can be used as a pre-adaptation for self-training to boost its performance. Sufficient experiments prove the effectiveness of our method against existing state-of-the-art methods on representative UDA benchmarks. Shuang Wang 0001, Dong Zhao 0007, Yuwei Guo 0001, Qi Zang, Yu Gu 0015, Yi Li 0054, Licheng Jiao |
IEEE Trans. Image Process. | 7 |
| 2021 | A Feature Decomposition Framework for Multi-Modal Image Patch MatchingabstractMulti-modal remote sensing images have complementary information which is conducive to enhancing the performance of various applications. Image patch matching plays a crucial role in the combination of multi-modal images. However, there are great differences in appearance and texture of multi-modal images, which brings great difficulties to image patching matching. To solve this problem, we propose a novel feature decomposition framework for multi-modal image patch matching. It aims to eliminate the hinder caused by the significant difference in multi-modal images. Specifically, this paper proposes to decompose the feature of images into common feature and modal private feature. Then, only the common feature is used for image patch matching, so as to improve the matching accuracy. Experimental results on optical and SAR images demonstrate that our proposed feature decomposition framework can significantly improve the performance of multi-modal image patch matching. Baorui Duan, Dou Quan, Yi Li 0054, Ruiqi Lei, Shuang Wang 0001, Biao Hou, Licheng Jiao |
IGARSS | 3 |
| 2021 | Deep Global Feature-Based Template Matching for Fast Multi-Modal Image RegistrationabstractDue to the different imaging mechanisms, there is a significant non-line difference between multi-modal images, which brings difficulties to multi-modal image registration. The traditional methods based on grayscale and handcraft features are difficult with obtain common features between different source images. The performances of deep local features matching methods rely on the quality and quantity of the detected keypoints, which can be quite time-consuming to register images. To achieve fast and accurate multi-modal image registration, we propose a deep global feature-based template matching method (GFTM) which uses a deep convolutional network to extract common global deep features from multi-modal images. Then, fast template matching is performed on global deep features to search the position with maximal similarity. Additionally, we build a similarity label map and design three losses to optimize our network, including contrast loss, error loss and peak loss. Extensive experimental results on optical and SAR images demonstrated that our proposed method is effective on multi-modal image registration. Ruiqi Lei, Bowu Yang, Dou Quan, Yi Li 0054, Baorui Duan, Shuang Wang 0001, Huarong Jia, Biao Hou, Licheng Jiao |
IGARSS | 4 |
| 2021 | Multi-Relation Attention Network for Image Patch MatchingabstractDeep convolutional neural networks attract increasing attention in image patch matching. However, most of them rely on a single similarity learning model, such as feature distance and the correlation of concatenated features. Their performances will degenerate due to the complex relation between matching patches caused by various imagery changes. To tackle this challenge, we propose a multi-relation attention learning network (MRAN) for image patch matching. Specifically, we propose to fuse multiple feature relations (MR) for matching, which can benefit from the complementary advantages between different feature relations and achieve significant improvements on matching tasks. Furthermore, we propose a relation attention learning module to learn the fused relation adaptively. With this module, meaningful feature relations are emphasized and the others are suppressed. Extensive experiments show that our MRAN achieves best matching performances, and has good generalization on multi-modal image patch matching, multi-modal remote sensing image patch matching and image retrieval tasks. Dou Quan, Shuang Wang 0001, Yi Li 0054, Bowu Yang, Ning Huyan, Jocelyn Chanussot, Biao Hou, Licheng Jiao |
IEEE Trans. Image Process. | 3 |