Kaixuan Jiang

dblp:326/5156 · DBLP profile ↗
← Back
7ranked-venue papers
6as first author
7since 2021 · last 2026
0000-0002-7223-1452ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 From segmentation to change: Releasing segment anything model for remote sensing change detection
Kaixuan Jiang, Chen Wu 0003, Zhenghui Zhao, Bo Du 0001, Liangpei Zhang
Pattern Recognit.1
2025 Beyond the Destination: A Novel Benchmark for Exploration-Aware Embodied Question Answering
abstract
Embodied Question Answering (EQA) is a challenging task in embodied intelligence that requires agents to dynamically explore 3D environments, actively gather visual information, and perform multi-step reasoning to answer questions. However, current EQA approaches suffer from critical limitations in exploration efficiency, dataset design, and evaluation metrics. Moreover, existing datasets often introduce biases or prior knowledge, leading to disembodied reasoning, while frontier-based exploration strategies struggle in cluttered environments and fail to ensure fine-grained exploration of task-relevant areas. To address these challenges, we construct the EXPloration-awaRe Embodied queStion anSwering Benchmark (EXPRESS-Bench), the largest dataset designed specifically to evaluate both exploration and reasoning capabilities. EXPRESS-Bench consists of 777 exploration trajectories and 2,044 question-trajectory pairs. To improve exploration efficiency, we propose Fine-EQA, a hybrid exploration model that integrates frontier-based and goal-oriented navigation to guide agents toward task-relevant regions more effectively. Additionally, we introduce a novel evaluation metric, Exploration-Answer Consistency (EAC), which ensures faithful assessment by measuring the alignment between answer grounding and exploration reliability. Extensive experimental comparisons with state-of-the-art EQA models demonstrate the effectiveness of our EXPRESS-Bench in advancing embodied exploration and question reasoning.
Kaixuan Jiang, Yang Liu 0119, Jingzhou Luo, Ziliang Chen 0001, Ling Pan, Guanbin Li, Liang Lin 0004
ICCV1
2025 LGCANet: Local-Global and Change-Aware Network via Segment Anything Model for Remote Sensing Images Change Detection
abstract
Change detection (CD) is a very fundamental and challenging task in remote sensing. Many deep learning-based CD methods generally utilize Siamese networks to extract image features. However, the semantic features extracted by these methods are still not fine-grained. In addition, these CD methods ignore the object scale diverse in remote sensing images and the interaction information between bi-temporal images, which leads to the problem that the network is unable to capture more efficient feature embeddings, with ambiguous or erroneous detection results. To alleviate the above issues, we propose Local-Global and Change Aware Network via Fast Segment Anything Model (LGCANet). The Segment Everything Model (SAM) can accurately segment objects in various scene images. In this work, we intend to utilize the powerful recognition capabilities of SAM to refine the CD task. Therefore, LGCANet employs more efficient FastSAM and ResNet as encoders to extract potential feature representations in remote sensing images. FastSAM can effectively extract global contextual information, combined with ResNet’s powerful deep feature extraction capability, which enables the network to comprehensively model features. LGCANet contains three modules: content aware attention module (CAAM), fore-background aware module (FAM), and edge-reinforce hybrid-selection module (EHM). CAAM delivers feature extraction from local to global perception, realizing dynamic attention to various scales of objects. FAM can effectively learn foreground and background representations through feature interaction, which significantly enhances the model’s capability of recognizing changed regions. EHM can utilize direction-awareness to extract edge information and generate fine-grained detection maps by adaptively selecting discriminative features through designed attention mechanisms. Experiments on publicly available CD datasets show that LGCANet achieves superior detection performance compared to other state-of-the-art methods. The code is available at https://github.com/Jscript10/LGCANet.
Kaixuan Jiang, Chen Wu 0003
IEEE Trans. Geosci. Remote. Sens.1
2023 MANet: An Efficient Multidimensional Attention-Aggregated Network for Remote Sensing Image Change Detection
abstract
Deep learning has significantly advanced the change detection in remote sensing image with its excellent performance. For change detection tasks, there are two critical issues. First, with scale variance of different objects in remote sensing images, effectively aggregating multi-scale features helps to generate fine-grained change objects. Second, it is critical but challenging to fully exploit the variance information between bi-temporal images to avoid pseudo-variation and region blurring. To alleviate the above issues, this paper proposes an efficient multi-dimensional attention-aggregation network (MANet), which keeps better feature aggregation while maintaining excellent differential attention ability. This paper carries three main contributions. First, we propose a multiscale asymmetric convolutional attention (MACA) module. Due to the asymmetric convolution’s ability to focus on feature contours effectively, the MACA can not only aggregate multi-scale features effectively, but also refine the edge information of features. Second, we propose a dual-dimensional attention (DDA) module for adaptively fusing shallow and deep features, which is used to generate rich feature representations. Third, the difference guidance (DG) module is exploited for enhancing the attention of changed regions to mitigate the influence of uncorrelated changes on the change detection result. Experiments on four popular change detection datasets show that our network can accomplish higher detection accuracy than the state-of-the-art networks.
Kaixuan Jiang, Jia Liu 0020, Fang Liu 0034, Liang Xiao 0001
IEEE Trans. Geosci. Remote. Sens.1
2022 Dual Unet: A Novel Siamese Network for Change Detection with Cascade Differential Fusion
abstract
Change detection (CD) of remote sensing images is to detect the change region by analyzing the difference between two bitemporal images. It is extensively used in land resource planning, natural hazards monitoring and other fields. In our study, we propose a novel Siamese neural network for change detection task, namely Dual-UNet. In contrast to previous individually encoded the bitemporal images, we design an encoder differential-attention module to focus on the spatial difference relationships of pixels. In order to improve the generalization of networks, it computes the attention weights between any pixels between bitemporal images and uses them to engender more discriminating features. In order to improve the feature fusion and avoid gradient vanishing, multi-scale weighted variance map fusion strategy is proposed in the decoding stage. Experiments demonstrate that the proposed approach consistently outperforms the most advanced methods on popular seasonal change detection datasets.
Kaixuan Jiang, Jia Liu 0020, Fang Liu 0034, Yangguang Liu, Jiao Shi
IGARSS1
2022 Spatial-Adaptive and Feature-Enhanced Siamese Network for Change Detection
abstract
Change detection (CD) plays an increasingly important role in earth observation and reveals surface changes according to multi-temporal images. Although deep learning-based CD methods work well for their excellent modeling ability, objects in different size and shape are generally processed by the same filter kernels in feature extraction, which leads to spatial blurring and degrades the CD performance. In this paper, a spatial adaptive and feature enhanced (SAFE) siamese network is proposed to tackle this problem, where the SAFE consists of a spatial-adaptive (SA) part and a feature-enhanced (FE) part. Specifically, pixel belonging to different objects possesses its own spatial knowledge, which is captured by a soft fusion of multi-scale difference images (DIs) called SA part. Changed and unchanged areas are strengthened or weakened by the FE, which combines object features with each DI accordingly. Moreover, since there are more unchanged pixels than changed pixels, a weight-pair is introduced to balance changed and unchanged objects in the training process. The experimental results verify that compared with four representative CD algorithms, our proposed method performs best on the Change Detection Dataset (CDD).
Yangguang Liu, Fang Liu 0034, Jia Liu 0020, Xu Tang 0004, Kaixuan Jiang, Liang Xiao 0001
IGARSS5
2022 Joint Variation Learning of Fusion and Difference Features for Change Detection in Remote Sensing Images
abstract
Remote sensing (RS) image change detection (CD) is an earth observation technique for detecting surface changes in the same area during a period. With the rapid development of deep learning, various deep neural networks especially Siamese ones have been widely used in the field of CD. However, they have the deficiency of insufficient contextual information aggregation, resulting in false and missed detections, and it is difficult to refine the detection of change edges. To alleviate these problems and obtain more accurate results, we propose an efficient self-weighted spatial-temporal attention network (SSANet). In contrast to the Siamese structure, our network is a novel joint learning framework composed of fusion sub-network, difference sub-network, and decoder. Fusion sub-network is used to extract multiscale object features where we propose a multi-core channel-aligning attention (MCA) module to capture the long-range semantic information for multi-scale context aggregation. Difference sub-network is used to extract the difference variation features, where we propose a feature differential reconfiguration (FDR) module to learn the temporal change information. FDR can effectively filter change information and reconstruct features to improve the perception of changed regions. To better balance the MCA and FDR modules, an asymmetric weighting (AW) module is proposed in the decoder to self-weight the multi-scale features and generate the change map. Experiments demonstrate the efficiency of proposed sub-networks and modules, and the state-of-the-art performance of SSANet.
Kaixuan Jiang, Jia Liu 0020, Fang Liu 0034, Liang Xiao 0001
IEEE Trans. Geosci. Remote. Sens.1