Zhenghang Yuan

dblp:229/6328 · DBLP profile ↗
← Back
9ranked-venue papers
8as first author
7since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 9 · 8 first-author · 7 since 2021
YearPublicationVenuePosition
2024 Referring Image Segmentation for Remote Sensing Data
abstract
In this paper, we present a new task: referring image segmentation for remote sensing data, which targets segmenting out specific objects referred to by natural language. Due to the absence of a dataset for this task, we construct a dataset based on the SkyScapes dataset. Our dataset is designed with linguistically structured expressions that focus on object categories, attributes, and spatial relationships, enabling the generation of binary masks from semantic segmentation maps. To benchmark this task, we evaluate and compare the performance of three different convolutional neural network (CNN)-based methods and a Transformer-based method. Experimental results provide valuable insights into the adaptability of these methods to remote sensing data, highlighting the potential of our dataset as a resource for the remote sensing community to further explore vision-language tasks.
Zhenghang Yuan, Lichao Mou, Yuansheng Hua, Xiao Xiang Zhu 0001
IGARSS1
2024 RRSIS: Referring Remote Sensing Image Segmentation
abstract
Localizing desired objects from remote sensing images is of great use in practical applications. Referring image segmentation, which aims at segmenting out the objects to which a given expression refers, has been extensively studied in natural images. However, almost no research attention is given to this task of remote sensing imagery. Considering its potential for real-world applications, in this paper, we introduce referring remote sensing image segmentation (RRSIS) to fill in this gap and make some insightful explorations. Specifically, we create a new dataset, called RefSegRS, for this task, enabling us to evaluate different methods. Afterward, we benchmark referring image segmentation methods of natural images on the RefSegRS dataset and find that these models show limited efficacy in detecting small and scattered objects. To alleviate this issue, we propose a language-guided cross-scale enhancement (LGCE) module that utilizes linguistic features to adaptively enhance multi-scale visual features by integrating both deep and shallow features. The proposed dataset, benchmarking results, and the designed LGCE module provide insights into the design of a better RRSIS model. The dataset and code will be available at https://gitlab.lrz.de/ai4eo/reasoning/rrsis.
Zhenghang Yuan, Lichao Mou, Yuansheng Hua, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.1
2023 Overcoming Language Bias in Remote Sensing Visual Question Answering Via Adversarial Training
abstract
The Visual Question Answering (VQA) system offers a user-friendly interface and enables human-computer interaction. However, VQA models commonly face the challenge of language bias, resulting from the learned superficial correlation between questions and answers. To address this issue, in this study, we present a novel framework to reduce the language bias of the VQA for remote sensing data (RSVQA). Specifically, we add an adversarial branch to the original VQA framework. Based on the adversarial branch, we introduce two regularizers to constrain the training process against language bias. Furthermore, to evaluate the performance in terms of language bias, we propose a new metric that combines standard accuracy with the performance drop when incorporating question and random image information. Experimental results demonstrate the effectiveness of our method. We believe that our method can shed light on future work for reducing language bias on the RSVQA task.
Zhenghang Yuan, Lichao Mou, Xiao Xiang Zhu 0001
IGARSS1
2022 Change-Aware Visual Question Answering
abstract
Change detection has been a hot research topic in the field of remote sensing, and it can provide information on observing changes of Earth's surface. However, segmentation-based change results are not very friendly to end users. Thus, in order to improve user experience and offer them high-level semantic information on change detection, we introduce a new task: change-aware visual question answering (VQA) on multi-temporal aerial images. Specifically, given a pair of multi-temporal aerial images and questions, this task aims to automatically provide natural language answers. By doing so, end users have better access to easy-to-understand change information through natural language. Besides, we also create a dataset made of multi-temporal image-question-answer triplets and a baseline method for this task. Experimental results offer valuable insights for the further research on this task.
Zhenghang Yuan, Lichao Mou, Xiao Xiang Zhu 0001
IGARSS1
2022 From Easy to Hard: Learning Language-Guided Curriculum for Visual Question Answering on Remote Sensing Data
abstract
Visual question answering (VQA) for remote sensing scene has great potential in intelligent human-computer interaction system. Although VQA in computer vision has been widely researched, VQA for remote sensing data (RSVQA) is still in its infancy. There are two characteristics that need to be specially considered for the RSVQA task. 1) No object annotations are available in RSVQA datasets, which makes it difficult for models to exploit informative region representation; 2) There are questions with clearly different difficulty levels for each image in the RSVQA task. Directly training a model with questions in a random order may confuse the model and limit the performance. To address these two problems, in this paper, a multi-level visual feature learning method is proposed to jointly extract language-guided holistic and regional image features. Besides, a self-paced curriculum learning (SPCL)-based VQA model is developed to train networks with samples in an easy-to-hard way. To be more specific, a language-guided SPCL method with a soft weighting strategy is explored in this work. The proposed model is evaluated on three public datasets, and extensive experimental results show that the proposed RSVQA framework can achieve promising performance. Code will be available at https://gitlab.lrz.de/ai4eo/reasoning/VQA-easy2hard.
Zhenghang Yuan, Lichao Mou, Qi Wang 0009, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.1
2022 Change Detection Meets Visual Question Answering
abstract
The Earth’s surface is continually changing, and identifying changes plays an important role in urban planning and sustainability. Although change detection techniques have been successfully developed for many years, these techniques are still limited to experts and facilitators in related fields. In order to provide every user with flexible access to change information and help them better understand land-cover changes, we introduce a novel task: change detection-based visual question answering (CDVQA) on multi-temporal aerial images. In particular, multi-temporal images can be queried to obtain high level change-based information according to content changes between two input images. We first build a CDVQA dataset including multi-temporal image-question-answer triplets using an automatic question-answer generation method. Then, a baseline CDVQA framework is devised in this work, and it contains four parts: multi-temporal feature encoding, multi-temporal fusion, multi-modal fusion, and answer prediction. In addition, we also introduce a change enhancing module to multi-temporal feature encoding, aiming at incorporating more change-related information. Finally, effects of different backbones and multi-temporal fusion strategies are studied on the performance of CDVQA task. The experimental results provide useful insights for developing better CDVQA models, which are important for future research on this task. The dataset will be available at https://github.com/YZHJessica/CDVQA.
Zhenghang Yuan, Lichao Mou, Zhitong Xiong, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.1
2021 Self-Paced Curriculum Learning for Visual Question Answering on Remote Sensing Data
abstract
Answering questions with natural language by extracting information from image has great potential in various applications. Although visual question answering (VQA) for natural image has been broadly studied, VQA for remote sensing data is still in the early research stage. For the same remote sensing image, there exist questions with dramatically different difficulty-levels. Treating these questions equally may mislead the model and limit the VQA model performance. Considering this problem, in this work, we propose a self-paced curriculum learning (SPCL) based VQA model with hard and soft weighting strategies for remote sensing data. Like human learning process, the model is trained from easy to hard question samples gradually. Extensive experimental results on two datasets demonstrate that the proposed training method can achieve promising performance.
Zhenghang Yuan, Lichao Mou, Xiao Xiang Zhu 0001
IGARSS1
2019 GETNET: A General End-to-End 2-D CNN Framework for Hyperspectral Image Change Detection
abstract
Change detection (CD) is an important application of remote sensing, which provides timely change information about large-scale Earth surface. With the emergence of hyperspectral imagery, CD technology has been greatly promoted, as hyperspectral data with high spectral resolution are capable of detecting finer changes than using the traditional multispectral imagery. Nevertheless, the high dimension of the hyperspectral data makes it difficult to implement traditional CD algorithms. Besides, endmember abundance information at subpixel level is often not fully utilized. In order to better handle high-dimension problem and explore abundance information, this paper presents a general end-to-end 2-D convolutional neural network (CNN) framework for hyperspectral image CD (HSI-CD). The main contributions of this paper are threefold: 1) mixed-affinity matrix that integrates subpixel representation is introduced to mine more cross-channel gradient features and fuse multisource information; 2) 2-D CNN is designed to learn the discriminative features effectively from the multisource data at a higher level and enhance the generalization ability of the proposed CD algorithm; and 3) the new HSI-CD data set is designed for objective comparison of different methods. Experimental results on real hyperspectral data sets demonstrate that the proposed method outperforms most of the state of the arts.
Qi Wang 0009, Zhenghang Yuan, Qian Du 0001, Xuelong Li 0001
IEEE Trans. Geosci. Remote. Sens.2
2018 ROBUST PCANet for Hyperspectral Image Change Detection
abstract
Deep learning is an effective tool for handling high-dimensional data and modeling nonlinearity, which can tackle the hyperspectral data well. Usually deep learning methods need a large number of training samples. However, there is no labeled data for training in change detection (CD). Considering these, this paper develops an unsupervised Robust PCA network (RPCANet) for hyperspectral image CD task. The main contributions of this work are twofold: 1) An unsupervised convolutional neural networks named RPCANet is proposed to handle the hyperspectral image CD; 2) An effective CD framework using the RPCANet and change vector analysis (CVA) is designed to achieve better CD performance with more powerful features. Experimental results on real hyperspectral data sets demonstrate the effectiveness of the proposed method.
Zhenghang Yuan, Qi Wang 0009, Xuelong Li 0001
IGARSS1