Rui Yang 0038

dblp:92/1942-38 · DBLP profile ↗
← Back
13ranked-venue papers
6as first author
13since 2021 · last 2025
0000-0002-3209-0456ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 7 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2025 Predicting Spectral Information for Self-Supervised Signal Classification
abstract
Deep learning methods have demonstrated remarkable performance across various communication signal processing tasks. However, most signal classification methods require a substantial amount of labeled samples for training, posing significant challenges in the field of communication signals, as labeling necessitates expert knowledge. This paper proposes a novel self-supervised signal classification method called Spectral-Guided Self-Supervised Signal Classification (SGSSC). Specifically, to leverage frequency-domain information with modulation semantics as prior knowledge for the model, we design a previously unexplored pretext task tailored to the format of signal data. This task involves predicting spectral information from masked time-domain signals, enabling the model to learn implicit signal features through cross-domain pattern transformation. Furthermore, the pretext task in the SGSSC method is relevant to the downstream classification task, and using traditional fine-tuning strategies on the downstream task may lead to the loss of certain features associated with the pretext task. Therefore, we propose an attention mechanism-based fine-tuning strategy that adaptively integrates pre-trained features from different levels. Extensive experimental results validate the superiority of the SGSSC method. For instance, when the proportion of labeled samples is only 0.5%, our method achieves an average improvement of 2.3% in downstream classification tasks compared to the best-performing self-supervised training strategies.
Shuang Wang 0001, Hantong Xing, Chenxu Wang 0001, Dou Quan, Rui Yang 0038, Dong Zhao 0007, Luyang Mei
IJCAI6
2024 TfNet: Building Detection in Remote Sensing Images Using Multi-Scale Feature Fusion
abstract
Building detection in remote sensing images is significant to urban land planning, battlefield environment perception, and illegal building monitoring. Existing methods excel in detecting small-scale building areas but struggle with remote sensing images with diverse and multi-scale characteristics. To solve this problem, this paper proposes Two-stage Feature fusion Network(TFNet). Specifically, we propose dense short connection module to enable internal interaction between features from different layers, achieving intra-stage feature fusion. Moreover, the holistically-nested edge detection network is used to integrate the effective information between different layers to achieve inter-stage feature fusion. In addition, the adaptive fusion weight is introduced to make the model adaptively select the weights of varying levels of features. Experimental results demonstrate the effectiveness of TFNet in improving the detection performance of multi-scale buildings in remote sensing images.
Chunlei Han, Luyang Mei, ZhongQian Jin, Shuang Wang 0001, Siyu Cao, Rui Yang 0038
IGARSS8
2024 MfrNet: A New Multi-Scale Feature Refining Method for Remote Sensing Image Change Captioning
abstract
Remote Sensing Image Change Captioning (RSICC) is an emerging multimodal field with promising prospects. This paper introduces a remote sensing image change caption model based on multi-scale and refined features. First, it extracts multi-scale features from dual-temporal images and then feeds them into the JointAtt and Dence Fusion (JADF) module for attention mutual guidance and feature refinement to eliminate noise. Next, the features are input into a transformer-based sentence generator for change statement generation. We conducted experiments on the Levir-CC dataset comparing our approach with existing methods, the results indicate that our MFRNet outperforms state-of-the-art methods in all metrics.
Kaiqi Xu, Yingping Han, Rui Yang 0038, Xiutiao Ye, Yanhe Guo, Hantong Xing, Shuang Wang 0001
IGARSS3
2024 Selection and Reconstruction of Key Locals: A Novel Specific Domain Image-Text Retrieval Method
abstract
In recent years, Vision-Language Pre-training (VLP) models have demonstrated rich prior knowledge for multimodal alignment, prompting investigations into their application in Specific Domain Image-Text Retrieval(SDITR) such as Text-Image Person Re-identification (TIReID) and Remote Sensing Image-Text Retrieval (RSITR). Due to the unique data characteristics in specific scenarios, the primary challenge is to leverage discriminative fine-grained local information for improved mapping of images and text into a shared space. Current approaches interact with all multimodal local features for alignment, implicitly focusing on discriminative local information to distinguish data differences, which may bring noise and uncertainty. Furthermore, their VLP feature extractors like CLIP often focus on instance-level representations, potentially reducing the discriminability of fine-grained local features. To alleviate these issues, we propose an Explicit Key Local information Selection and Reconstruction Framework (EKLSR), which explicitly selects key local information to enhance feature representation. Specifically, we introduce a Key Local information Selection and Fusion (KLSF) that utilizes hidden knowledge from the VLP model to select interpretably and fuse key local information. Secondly, we employ Key Local segment Reconstruction (KLR) based on multimodal interaction to reconstruct the key local segments of images (text), significantly enriching their discriminative information and enhancing both inter-modal and intra-modal interaction alignment. To demonstrate the effectiveness of our approach, we conducted experiments on five datasets across TIReID and RSITR. Notably, our EKLSR model achieves state-of-the-art performance on two RSITR datasets.
Yu Liao, Rui Yang 0038, Jianwei Tao, Bai Liu 0002, Zhipeng Hu, Shuang Wang 0001, Zeng Zhao
ACM Multimedia3
2024 Accurate and Lightweight Learning for Specific Domain Image-Text Retrieval
abstract
Recent advances in vision-language pre-trained models like CLIP have greatly enhanced general domain image-text retrieval performance. This success has led scholars to develop methods for applying CLIP to Specific Domain Image-Text Retrieval (SDITR) tasks such as Remote Sensing Image-Text Retrieval (RSITR) and Text-Image Person Re-identification (TIReID). However, these methods for SDITR often neglect two critical aspects: the enhancement of modal-level distribution consistency within the retrieval space and the reduction of CLIP's computational cost during inference. To address these issues, this paper presents a novel framework, Accurate and lightweight learning for specific domain Image-text Retrieval (AIR), based on the CLIP. AIR incorporates a Modal-Level distribution Consistency Enhancement regularization (MLCE) loss and a Self-Pruning Distillation Strategy (SPDS) to improve retrieval precision and computational efficiency. The MLCE loss harmonizes the sample distance distributions within image and text modalities, fostering a retrieval space closer to the ideal state. SPDS employs a strategic knowledge distillation process to transfer deep multimodal insights from CLIP to a shallower level, maintaining only the essential layers for inference, thus achieving model light-weighting. Comprehensive experiments across various datasets in RSITR and TIReID reveal that MLCE loss secures optimal retrieval, while SPDS achieves a favorable balance between accuracy and computational demand during testing.
Rui Yang 0038, Shuang Wang 0001, Jianwei Tao, Yingping Han, Qiaoling Lin, Yanhe Guo, Biao Hou, Licheng Jiao
ACM Multimedia1
2024 Continual learning for cross-modal image-text retrieval based on domain-selective attention
Rui Yang 0038, Shuang Wang 0001, Yu Gu 0015, Jihui Wang, Yingzhi Sun, Yu Liao, Licheng Jiao
Pattern Recognit.1
2024 Transcending Fusion: A Multiscale Alignment Method for Remote Sensing Image-Text Retrieval
abstract
Remote sensing image-text retrieval (RSITR) is pivotal for knowledge services and data mining in the remote sensing (RS) domain. Considering the multiscale representations in image content and text vocabulary can enable the models to learn richer representations and enhance retrieval. Current multiscale RSITR approaches typically align multiscale fused image features with text features but overlook aligning image-text pairs at distinct scales separately. This oversight restricts their ability to learn joint representations suitable for effective retrieval. We introduce a novel multiscale alignment (MSA) method to overcome this limitation. Our method comprises three key innovations: 1) a multiscale cross-modal alignment transformer (MSCMAT), which computes cross-attention between single-scale image features and localized text features, integrating global textual context to derive a matching score matrix within a mini-batch; 2) a multiscale cross-modal semantic alignment loss (MSCMA loss) that enforces semantic alignment across scales; and 3) a cross-scale multimodal semantic consistency loss (CSMMC loss) that uses the matching matrix from the largest scale to guide alignment at smaller scales. We evaluated our method across multiple datasets, demonstrating its efficacy with various visual backbones and establishing its superiority over existing state-of-the-art methods. The GitHub URL for our project ishttps://github.com/yr666666/MSA.
Rui Yang 0038, Shuang Wang 0001, Yingping Han, Yuanheng Li, Dong Zhao 0007, Dou Quan, Yanhe Guo, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.1
2023 Learning Pseudo-Relations for Cross-domain Semantic Segmentation
abstract
Domain adaptive semantic segmentation aims to adapt a model trained on labeled source domain to unlabeled target domain. Self-training shows competitive potential in this field. Existing methods along this stream mainly focus on selecting reliable predictions on target data as pseudo-labels for category learning, while ignoring the useful relations between pixels for relation learning. In this paper, we propose a pseudo-relation learning framework, Relation Teacher (RTea), which can exploitable pixel relations to efficiently use unreliable pixels and learn generalized representations. In this framework, we build reasonable pseudo-relations on local grids and fuse them with low-level relations in the image space, which are motivated by the reliable local relations prior and available low-level relations prior. Then, we design a pseudo-relation learning strategy and optimize the class probability to meet the relation consistency by finding the optimal sub-graph division. In this way, the model’s certainty and consistency of prediction are enhanced on the target domain, and the cross-domain inadaptation is further eliminated. Extensive experiments on three datasets demonstrate the effectiveness of the proposed method. The code will be available at https://github.com/DZhaoXd/RTea.
Dong Zhao 0007, Shuang Wang 0001, Qi Zang, Dou Quan, Xiutiao Ye, Rui Yang 0038, Licheng Jiao
ICCV6
2023 A Fast and Accurate Method for Remote Sensing Image-Text Retrieval Based On Large Model Knowledge Distillation
abstract
With the increasing development of remote sensing (RS) technology, remote sensing cross-modal image-text retrieval (RSCMITR) task has gradually attracted wide attention. At present, the large-scale pre-training model is brilliant in the field of natural images cross-modal retrieval, but the current RSCMITR models do not focus on it, resulting in less retrieval performance improvement. This paper proposes a lightweight network structure based on large-scale pre-training model and knowledge distillation, designing a lightweight model based on separable convolution and text convolution. Knowledge distillation technology is used to make the Light model learn the hidden knowledge of large-scale model CLIP-RS, which realizes fast and accurate retrieval. The proposed method achieves state-of-the-art performance on four commonly used RSCMITR datasets.
Yu Liao, Rui Yang 0038, Hantong Xing, Dou Quan, Shuang Wang 0001, Biao Hou
IGARSS2
2023 A Texture and Saliency Enhanced Image Learning Method For Cross-Modal Remote Sensing Image-Text Retrieval
abstract
Cross-modal remote sensing image-text retrieval (CMRSITR) can retrieve images of interest from a vast amount of remote sensing images and has received significant attention in recent years. However, existing methods do not consider saliency and texture information, which are essential for remote sensing images when extracting image features. Therefore, this paper proposes a novel texture and saliency enhanced image learning method for CMRSITR. We constructed a multi-task image feature extractor in this new method. A texture map and a saliency map are created by extracting texture and detecting the saliency of each RS image. Both maps are set as supervised information during training to make the extracted saliency and texture features gradually reconstructed to a saliency map and a texture map, respectively. At the same time, the retrieval features of each RS image are obtained from the retrieval feature branch of the image. Experiments conducted on two commonly used CMRSITR datasets, RSICD and UCM, showed that the proposed method is effective in improving retrieval performance and achieved state-of-the-art retrieval performance compared to existing methods.
Rui Yang 0038, Yanhe Guo, Shuang Wang 0001
IGARSS1
2023 Knowledge Decomposition and Replay: A Novel Cross-modal Image-Text Retrieval Continual Learning Method
abstract
To enable machines to mimic human cognitive abilities and alleviate the catastrophic forgetting problem in cross-modal image-text retrieval (CMITR), this paper proposes a novel continual learning method, Knowledge Decomposition and Replay (KDR), which emulates the process of knowledge decomposition and replay exhibited by humans in complex and changing environments. KDR has two components: a feature Decomposition-based CMITR Model (DCM) and a cross-task Generic Knowledge Replay strategy (GKR). DCM decomposes text and image features into task-specific and generic knowledge features, mimicking the human cognitive process of knowledge decomposition. Specifically, it employs a generic knowledge features extraction module for all tasks and a task-specific module for each task with a few trainable fully connected layers. Similarly, GKR emulates the human behavior of knowledge replay by utilizing the image-text similarity matrix output from the old task model with inputting the previous samples to induce the learning of the image-text similarity matrix output from the current task model with inputting the previous samples, using knowledge distillation technology. To demonstrate the effect of KDR, we adapted a continual learning dataset Seq-COCO from MSCOCO. Extensive experiments on Seq-COCO showed that KDR reduces catastrophic forgetting and consolidates general knowledge, improving the model's learning ability in CMITR.
Rui Yang 0038, Shuang Wang 0001, Yanhe Guo, Xiutiao Ye, Biao Hou, Licheng Jiao
ACM Multimedia1
2022 A Transformer-Based Cross-Modal Image-Text Retrieval Method using Feature Decoupling and Reconstruction
abstract
With the increasing application of remote sensing technology, the task of cross-modal retrieval of remote sensing images (CMRRS) has gradually attracted widespread attention. Ex-isting methods often completely map the features of different modalities to a shared space and do not decouple between the modal-invariant information and modal-heterogeneous in-formation, which leads to redundant information in feature mapping and usually gets sub-optimal retrieval performance. This paper proposes a Transformer-based CMRRS method using feature decoupling and reconstruction (TBFDR) to solve this problem. TBFDR achieves state-of-the-art performance in remote sensing image-text retrieval task on Sydney-Captions dataset.
Yingzhi Sun, Yu Liao, Rui Yang 0038, Shuang Wang 0001, Biao Hou, Licheng Jiao
IGARSS5
2021 Cross-Modal Feature Fusion Retrieval for Remote Sensing Image-Voice Retrieval
abstract
With the increasing popularity of remote sensing technology applications, some emergency scenarios require rapid retrieval of remote sensing images, such as earthquake rescue, etc. Due to the high efficiency of voice input, researchers have focused on cross-modal remote sensing image-voice retrieval methods. However, these methods have two major drawbacks: speech input lacks discrimination and the intra-modal semantic information is under used. To address these drawbacks, we propose a novel cross-modal feature fusion retrieval model. Our model provides a more optimized cross-modal common feature space than previous models and thus optimizes the retrieval performance. First, our model adds the extra textual keyword information to the audio feature for remote sensing image retrieval. Second, it introduces inter-modality adversarial learning and intra-modality semantic discrimination into the remote sensing image-voice retrieval task. We conducted experiments on two datasets modified from the UCM-Captions dataset and the Remote Sensing Image Caption Dataset. The experimental results show that our model outperforms state-of-the-art models in this task.
Rui Yang 0038, Yu Gu 0015, Yu Liao, Yingzhi Sun, Shuang Wang 0001, Biao Hou, Licheng Jiao
IGARSS1