EDBT 2026 Demo / reviewers in the wild / expert
Jiawei Jiang 0002
dblp:185/1521-2
· DBLP profile ↗
22ranked-venue papers
6as first author
22since 2021 · last 2026
0000-0002-9200-9189ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 18 · 5 first-author · 18 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Computer networks · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Information-coupled MRI acceleration via multi-modal mapping and progressive masking
Jiawei Jiang 0002, Honghui Xu 0002, Jianwei Zheng 0001 |
Pattern Recognit. | 1 |
| 2026 | Arbitrary-Scale Fusion Operator for High-Resolution Hyperspectral ImagingabstractFor high-resolution hyperspectral (HrHs) imaging, spatial-spectral fusion offers a promising alternative to expensive equipment. However, retraining multiple models for varied scaling factors is currently unavoidable, costing extra computational resource and human labor. To address this issue, we propose Arbitrary-scale Fusion Operator (AFO), a lightweight solution for HrHs fusion given arbitrary scalings, turning the retraining strategy into “training-free” ones. Specifically, AFO treats low-resolution hyperspectral (LrHs) images and high resolution multispectral (HrMs) images as light-wise degraded functions within the spectrum, which are initially embedded into a high-dimensional space to simulate the original light signals, tapping the potential of enriched prior learning. Then, a flow of kernel integration (KI) is meticulously crafted, followed by a rival step of dimension reduction for HrHs generation. For a well-behaved KI computation, an Attention-Driven Convolution Integration (ADCI) is engineered to restore the broken discretization invariance derived by convolutions, yet with the locally inductive bias preserved. In addition, we propose an Implicit Neural Functional Integration (INFI) to achieve cross domain interaction of spatial degradation functions, followed by the use of Galerkin-type Integration (GI) as a decoder to handle high-frequency information. Finally, the bonded activation functions are improved for the principle of continuous-discrete equivalence. Extensive experiments validate the superiority of our approach over cutting-edge methods. Notably, our proposal holds significantly better generalization on arbitrary scaling factors, yet requires only 0.07M parameters. Honghui Xu 0002, Wei Li 0034, Jiawei Jiang 0002, Zhi Liu 0009, Jianwei Zheng 0001 |
IEEE Trans. Multim. | 4 |
| 2025 | Breaking Information Isolation: Accelerating MRI via Inter-sequence Mapping and Progressive MaskingabstractDeep unfolding network (DUN) has shed new light on multi-sequence MRI reconstruction, providing both high interpretability and acceptable performance. However, current approaches still suffer from the plight of information isolation, i.e., learning features of multi-suquences individually and leaving the mask departed from model updating. In this work, we propose a new unfolding solution, namely Information-coupled MRI Acceleration (IMA), to address the isolation issue. Concretely, two specific mechanisms are presented. On the one hand, the latent connections across different sequences are explicitly molded via two auxiliary matrices. While the first matrix is meticulously engineered to assemble the spatial details, the second one hammers at capturing the depth information conditioned on the enriched channels. On the other hand, following a deep analysis on the non-uniform distribution in low- and high-frequency components of the given mask, we elaborate a new unfolding flow using a progressive masking scheme, featuring a dilation-contraction mechanism during forward propagation of successive stages. Massive experiments are conducted under various sampling patterns and acceleration rates, whose results demonstrate that, without any sophisticated architectures, our IMA outperforms the current cutting-edge methods both visually and numerically. Jianwei Zheng 0001, Xiaomin Yao, Guojiang Shen, Wei Li 0034, Jiawei Jiang 0002 |
AAAI | 5 |
| 2025 | LMO: Linear Mamba Operator for MRI ReconstructionabstractInterpretability and consistency have long been crucial factors in MRI reconstruction. While interpretability has been significantly innovated with the emerging deep unfolding networks, current solutions still suffer from inconsistency issues and produce inferior anatomical structures. Especially in out-of-distribution cases, e.g., when the acceleration rate (AR) varies, the generalization performance is often catastrophic. To counteract the dilemma, we propose an innovative Linear Mamba Operator (LMO) to ensure consistency and generalization, while still enjoying desirable interpretability. Theoretically, we argue that mapping between function spaces, rather than between signal instances, provides a solid foundation of high generalization. Technically, LMO achieves a good balance between global integration facilitated by a state space model that scans the whole function domain, and local integration engaged with an appealing property of continuous-discrete equivalence. On that basis, learning holistic features can be guaranteed, tapping the potential of maximizing data consistency. Quantitative and qualitative results demonstrate that LMO significantly outperforms other state-of-the-arts. More importantly, LMO is the unique model that, with AR changed, achieves retraining performance without retraining steps. Codes are available at https://github.com/ZhengJianwei2/LMO. Wei Li 0034, Jiawei Jiang 0002, Kaihao Yu, Jianwei Zheng 0001 |
CVPR | 2 |
| 2025 | Controllable Face Inpainting via Pseudo-Style EmbeddingabstractImage inpainting, a critical facet of computer vision, is in full bloom accompanied by the rapid innovation of convolution neural networks and transformers, revolutionizing the practical management of abnormity disposal, image editing, etc. Of these applications, face inpainting is more challenging due to the higher demand for semantic accuracy in key regions such as eyes and nose. Classical face inpainting methods are celebrated for their fast generation speed and refined texture details. However, they often lack the level of controllability required for complex tasks. In contrast, existing multi-modal controllable inpainting techniques offer enhanced guidance through image-text integration but tend to be time-consuming and produce suboptimal texture refinement. To address these limitations, we propose the Multi-modal Pseudo-style Embedded Transformer (MPET), a novel and efficient multi-modal inpainting algorithm that seamlessly integrates the strengths of both approaches, achieving state-of-the-art performance. Specifically, edge completion facilitates a cost-efficient and simple bridging of the contour continuity. Multi-modal pseudo-style generation amalgamates the image-text modalities, successfully embedding text features within the visual vectors, thereby culminating in the formation of pseudo-style diagrams rich in diverse attributes. On that basis, a controllable style-embedded siamese network is elaborated, effectively orchestrating the interaction among style attributes while ensuring high-precision pixel infusion. Extensive experiments on public datasets demonstrate the superiority of our approach through both quantitative and qualitative evaluations, highlighting its potential to advance the field of face inpainting. Jiawei Jiang 0002, Yueqian Quan, Honghui Xu 0002, Jianwei Zheng 0001 |
ECAI | 2 |
| 2025 | Spatial-Spectral Fusion Neural OperatorabstractWith the rapid development of deep learning, spatial-spectral fusion (SSF) has emerged as an ideal alternative to traditional, costly hyperspectral image (HSI) acquisition methods. However, current solutions necessitate training and storing multiple models for different scaling factors. Besides, a meticulously designed network architecture to meet desirable performance often suffers from a severe computational burden. To counteract the dilemma, we propose SFNO, a lightweight spatial-spectral fusion neural operator for arbitrary-scale SSF. SFNO leverages approximation theory by embedding features from two degraded functions into a high-dimensional latent space, enabling efficient learning of basis functions. Kernel integration mechanisms are then used to approximate certain priors, followed by dimensionality reduction to generate high-resolution HSIs. Moreover, with the aid of discrete invariance property, we propose a new mechanism of progressive resampling (PR), which allows for the shrinkage of necessary spatial domain without any performance degradation. Extensive experiments on CAVE and Harvard datasets show that SFNO and its variant significantly improve performance, especially in out-of-domain fusion, requiring only 0.098M parameters and 0.966G FLOPs. Wei Li 0034, Jiawei Jiang 0002, Ni Xu, Yan Li 0083, Jianwei Zheng 0001 |
ICME | 2 |
| 2025 | SpecSolver: Solving Spatial-Spectral Fusion via Semantic TransformerabstractBy clustering pixels with locally similar values, superpixel-based approaches have shown great potential in processing hyperspectral images (HSI), thereby reducing the computational burden associated with large spatial dimensions. However, specific for spatial-spectral fusion (SSF), superpixel segmentation is inherently non-differentiable and irreversible; hence it is inapplicable. To address the issues, we propose a semantic transformer-based solver, namely SpecSolver, which is basically inspired by the benefits of superpixel-based approaches, yet with the inner mechanism completely improved. The core idea lies in learning the intrinsic semantic states of HSIs hidden behind discretized pixel representations. Specifically, we propose a new Semantic-Attention to adaptively split the image domain into a series of learnable slices of flexible shapes, where image pixels under similar semantic states will be ascribed to the same slice. By calculating attention to the Semantic-Superpixel tokens encoded from slices, SpecSolver can effectively capture intricate semantic correlations from the vast number of pixels, which also empowers the solver with an endogenous capacity for modeling different magnification scales and allows for efficient computation in linear complexity. On that basis, we elaborate a SpatialNet module, which extracts multiscale local spectral information, and a FreqNet module, which supplements global information, capturing subtle details and variations across different spectra. Experiments on two benchmark SSF datasets verify the state-of-the-art (SOTA) performance of the proposed method, both visually and quantitatively. Also, ablation studies validate the mentioned contributions. Wei Li 0034, Honghui Xu 0002, Jiawei Jiang 0002, Jianwei Zheng 0001 |
ACM Multimedia | 4 |
| 2025 | Arbitrary-scale Fusion Neural OperatorabstractSpatial-spectral fusion offers a promising alternative to expensive equipment in high-resolution hyperspectral (HrHs) imaging. However, training separate models for different scaling factors remains costly. To address this, we propose the Arbitrary-scale Fusion Neural Operator (AFNO), a lightweight solution for HrHs fusion across arbitrary scalings. Instead of entities, AFNO treats low-resolution hyperspectral (LrHs) and high-resolution multispectral (HrMs) images as functions and performs meticulously designed integrations as the mapping operator. The key components include Attention-Driven Convolution Integration (ADCI) to restore discretization invariance disrupted by convolutions, Implicit Neural Functional Integration (INFI) for cross-domain interaction of spatial degradations, and Galerkin-type Integration as a decoder for high-frequency details. Additionally, the bonded activation opeartor are improved for the principle of continuous-discrete equivalence. Extensive experiments validate the superiority of our approach over cutting-edge methods. Notably, AFNO holds significantly better generalization on arbitrary scaling factors, yet requiring only 0.07M parameters. Wei Li 0034, Honghui Xu 0002, Jiawei Jiang 0002, Zhi Liu 0009, Jianwei Zheng 0001 |
ACM Multimedia | 4 |
| 2025 | Semantic-Spatial Attention for Refined Object Placement in Text-to-Image SynthesisabstractSolely based on given prompts, text-guided diffusion models have enjoyed a unique capability in generating diverse and creative images. Nevertheless, the conveyance of image information through text presents a series of challenges, particularly in controlling the positioning of objects in synthesized images. Despite attempts of recent efforts in exploring alternative conditions, such as bounding box/mask-image pairs, the requirement of a substantial amount of paired data and time-consuming fine-tuning emerge as new issues. Given the observations that not only prompt-related cross-attention maps reveal the spatial arrangement and centroid positions of the objects, but also out-of-prompt markers enjoy rich semantic information, we thus engineer a weighted optimization loss. Specifically, three spatial sub-losses, namely inner box reinforcement loss, outer box attenuation loss, and centroid loss, are devised and seamlessly integrated into the sampling step of current vanilla diffusion models. Without any annotations of layout data required, the final approach runs in a training-free fashion. Extensive experiments with new performance scores demonstrate that our proposal not only successfully addresses the issue of object positioning but also boosts the capabilities of most current models, such as Stable Diffusion and GLIGEN, in high-quality synthesis and coverage of various concepts. Moreover, the proposed mechanism plays a plug-and-play role. Jianwei Zheng 0001, Ni Xu, Wei Li 0034, Jiawei Jiang 0002, Xiaoqin Zhang 0002 |
IEEE Trans. Multim. | 4 |
| 2025 | Axial-shunted Spatial-temporal Conversation for Change DetectionabstractBenefitting from the maturing of intelligence techniques and advanced sensors, recent years have witnessed the full flourishing of change detection (CD) on multi-temporal remote sensing images. However, extraneous interference caused by normal temporal evolution and the extreme sparsity of spatial changes still plague the detection accuracy. To counteract this dilemma, a lightweight axial-shunted spatial-temporal conversation network (ASCNet) is proposed, which models the intrinsic representations in dually augmented images with a parallel treatment of convolutions and attentions. Specifically, for the features of weakly augmented bi-temporal image pairs from Siamese CNN, a roundtable attention-based and intra-scale axial-shunted interaction, with linear complexity, is presented. By splitting horizontally or vertically into multiple chunks and then performing axial-squeeze operation, axial-shunted scheme can achieve fine-grained attention while maintaining linear complexity. Moreover, roundtable attention pursues efficient bi-temporal modeling by incorporating both self-attention and cross-attention in a single attentional computation, while imposing change guiding and difference gating for focusing on changes. Simultaneously, a video transformer is introduced for the modeling of strongly augmented sequences, followed by an inter-scale spatial-temporal alignment to recalibrate the feature responses. ASCNet demonstrates state-of-the-art performance on four publicly available CD datasets while maintaining superior computational efficiency. The source code is available at https://github.com/fengyuchao97/ASCNet . Yuchao Feng, Jiawei Jiang 0002, Jintao Lai, Jianwei Zheng 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | Null Space Matters: Range-Null Decomposition for Consistent Multi-Contrast MRI ReconstructionabstractConsistency and interpretability have long been the critical issues in MRI reconstruction. While interpretability has been dramatically improved with the employment of deep unfolding networks (DUNs), current methods still suffer from inconsistencies and generate inferior anatomical structure. Especially in multi-contrast scenes, different imaging protocols often exacerbate the concerned issue. In this paper, we propose a range-null decomposition-assisted DUN architecture to ensure consistency while still providing desirable interpretability. Given the input decomposed, we argue that the inconsistency could be analytically relieved by feeding solely the null-space component into proximal mapping, while leaving the range-space counterpart fixed. More importantly, a correlation decoupling scheme is further proposed to narrow the information gap for multi-contrast fusion, which dynamically borrows isotropic features from the opponent while maintaining the modality-specific ones. Specifically, the two features are attached to different frequencies and learned individually by the newly designed isotropy encoder and anisotropy encoder. The former strives for the contrast-shared information, while the latter serves to capture the contrast-specific features. The quantitative and qualitative results show that our proposal outperforms most cutting-edge methods by a large margin. Codes will be released on https://github.com/chenjiachengzzz/RNU. Jiawei Jiang 0002, Fei Wu 0026, Jianwei Zheng 0001 |
AAAI | 2 |
| 2024 | Memory-Augmented Dual-Domain Unfolding Network for MRI ReconstructionabstractThe compressed sensing MRI aims to recover high-fidelity images from undersampled k-space data, which enables MRI acceleration and meanwhile mitigates problems caused by prolonged acquisition time, such as physiological motion artifacts, patient discomfort, and delayed medical care. In this regard, the deep unfolding network (DUN) has emerged as the predominant solution due to the benefits of better interpretability and model capacity. However, existing algorithms remain inadequate for two principal reasons. First, directly unrolling a typical optimization algorithm is ill-considered for the structure information and domain knowledge. Second, the incorporation of the MRI-oriented imaging mechanism is inadequate. To tackle these two issues, we propose a Memoryaugmented Dual-domain Unfolding Network (MDUNet). Particularly, the per-iteratively learned memory is held to facilitate a better efficacy of feature representation. Besides, with the scheme of memory augmentation alternatively employed in the k-space and image domain, both the regional structure and global information can be complementarily integrated in a spiral manner. Comprehensive experiments conducted on diverse datasets, sampling rates, and sampling patterns demonstrate that our method, while maintaining a relatively small number of parameters, surpasses the latest methods. Codes will be available on the GitHub homepage of the corresponding author. Jiawei Jiang 0002, Yueqian Quan, Jianwei Zheng 0001 |
ICASSP | 1 |
| 2024 | Multi-dimensional visual data completion via weighted hybrid graph-Laplacian
Jiawei Jiang 0002, Yile Xu, Honghui Xu 0002, Guojiang Shen, Jianwei Zheng 0001 |
Signal Process. | 1 |
| 2024 | Cascading Blend Network for Image InpaintingabstractImage inpainting refers to filling in unknown regions with known knowledge, which is in full flourish accompanied by the popularity and prosperity of deep convolutional networks. Current inpainting methods have excelled in completing small-sized corruption or specifically masked images. However, for large-proportion corrupted images, most attention-based and structure-based approaches, though reported with state-of-the-art performance, fail to reconstruct high-quality results due to the short consideration of semantic relevance. To relieve the above problem, in this paper, we propose a novel image inpainting approach, namely cascading blend network (CBNet), to strengthen the capacity of feature representation. As a whole, we introduce an adjacent transfer attention (ATA) module in the decoder, which preserves contour structure reasonably from the deep layer and blends structure-texture information from the shadow layer. In a coarse to delicate manner, a multi-scale contextual blend (MCB) block is further designed to felicitously assemble the multi-stage feature information. In addition, to ensure a high qualified hybrid of the feature information, extra deep supervision is applied to the intermediate features through a cascaded loss. Qualitative and quantitative experiments on the Paris StreetView, CelebA, and Places2 datasets demonstrate the superior performance of our approach compared with most state-of-the-art algorithms. Yiting Jin, Wanliang Wang, Yidong Yan, Jiawei Jiang 0002, Jianwei Zheng 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2023 | Building Change Detection Using Cross-Temporal Feature Interaction NetworkabstractBuilding change detection of remote sensing images is in full flourishing accompanied by the prosperity of convolutional neural networks. For spatial-temporal context modeling, existing solutions disregard the inter-image interactions, albeit their positive contribution to the acquisition of differences. To fill the gap, we propose a cross-temporal feature interaction network to effectively derive the change representations. Specifically, we propose a linearized cross-attention, which motivates each counterpart to glimpse the representation of another image while preserving its own features. In addition, to circumvent the misalignment caused by step-down sampling in the backbone, we introduce multi-level feature alignment using learnable affine transformation and stepwise aggregation. Based on a naive backbone (ResNet18) without sophisticated structures, our model outperforms other state-of-the-art methods on three datasets in terms of both efficiency and effectiveness. Yuchao Feng, Jiawei Jiang 0002, Honghui Xu 0002, Jianwei Zheng 0001 |
ICASSP | 2 |
| 2023 | Low-Dose CT Reconstruction Via Optimization-Inspired GANabstractMost research on Low-dose Computed Tomography (LDCT) reconstruction is designed as a black box, lacking controllability and interpretability. In this paper, a Proximal Linear ADMM framework-based Generative Adversarial Network (PLA-GAN) is proposed. Specifically, without loss of interpretability, channel attention blocks and NonLocal Sparse Attention (NLSA) modules are embedded into two regularizers respectively and iterated alternately, driving the network to cope with real and complex CT image degradation through a multi-scale and adaptive way. To further promote the visual quality, a discriminator containing NLSA module is also introduced. The comparisons with state-of-the-arts on the Mayo dataset validate the superiority of our proposed algorithm both numerically and visually. The advantages of generalizability and interpretability are also evident. Jiawei Jiang 0002, Yuchao Feng, Honghui Xu 0002, Jianwei Zheng 0001 |
ICASSP | 1 |
| 2023 | Compact Intertemporal Coupling Network for Remote Sensing Change DetectionabstractChange detection of multi-temporal remote sensing images is in full flourishing accompanied by the popularity and prosperity of deep learning. The prominent challenge lies in the crude distribution of the newly constructed and demolished changes, the interferences of massive irrelevant objects, and the spatial-temporal changes from the passage of time. For context modeling, existing solutions waste massive attention on task-irrelevant features, spotlighting insufficiently on the genuinely changed regions. To fill the gap, we propose a compact intertemporal coupling network (CICNet) to derive the change representations. Specifically, to underpin the interaction of spatial-temporal differences in a global perspective, we detach and innovate the solo-head self-attention into a lightweight intertemporal-attention, favorably bridging the intra-level features. In addition, to circumvent the misalignment imposed by spatial sampling, lightweight global channel- and spatial- attentions are globally incorporated for stepwise calibration between localization seduced low-level information and semantics abundant high-level features. Based on a naive backbone (ResNet18/34) without sophisticated structures, our model outperforms other state-of-the-art methods on four datasets in terms of both efficiency and effectiveness. Yuchao Feng, Honghui Xu 0002, Jiawei Jiang 0002, Jianwei Zheng 0001 |
ICME | 3 |
| 2023 | GA-HQS: MRI reconstruction via a generically accelerated unfolding approachabstractDeep unfolding networks (DUNs) are the foremost methods in the realm of compressed sensing MRI, as they can employ learnable networks to facilitate interpretable forward-inference operators. However, several daunting issues still exist, including the heavy dependency on first-order optimization algorithms, the insufficient information fusion mechanisms, and the limitation of capturing long-range relationships. To address the issues, we propose a Generically Accelerated Half-Quadratic Splitting (GA-HQS) algorithm that incorporates second-order gradient information and pyramid attention modules for the delicate fusion of inputs at the pixel level. Moreover, a multi-scale split transformer is also designed to enhance the global feature representation. Comprehensive experiments demonstrate that our method surpasses previous ones on single-coil MRI acceleration tasks. Jiawei Jiang 0002, Honghui Xu 0002, Yuchao Feng, Jianwei Zheng 0001 |
ICME | 1 |
| 2023 | Latent-space Unfolding for MRI ReconstructionabstractTo circumvent the problems caused by prolonged acquisition periods, compressed sensing MRI enjoys a high usage profile to accelerate the recovery of high-quality images from under-sampled k-space data. Most current solutions dedicate to solving this issue with the pursuit of certain prior properties, yet the treatments are all enforced in the original space, resulting in limited feature information. To achieve a performance promotion yet with the guarantee of running efficiency, in this work, we propose a latent-space unfolding network (LsUNet). Specifically, by an elaborately designed reversible network, the inputs are first mapped to a channel-lifted latent space, which taps the potential of capturing spatial-invariant features sufficiently. Within the latent space, we then unfold an accelerated optimization algorithm to iterate an efficient and feasible solution, in which a parallelly dual-domain update is equipped for better feature fusion. Finally, an inverse embedding transformation of the recovered high-dimensional representation is applied to achieve the expected estimation. LsUNet enjoys high interpretability due to the physically induced modules, which not only facilitates an intuitive understanding of the internal operating mechanism but also endows it with high generalization ability. Comprehensive experiments on different datasets and various sampling rates/patterns demonstrate the advantages of our proposal over the latest methods both visually and numerically. Jiawei Jiang 0002, Yuchao Feng, Dongyan Guo, Jianwei Zheng 0001 |
ACM Multimedia | 1 |
| 2023 | Tensor completion via hybrid shallow-and-deep priors
Honghui Xu 0002, Jiawei Jiang 0002, Yuchao Feng, Yiting Jin, Jianwei Zheng 0001 |
Appl. Intell. | 2 |
| 2023 | Change Detection on Remote Sensing Images Using Dual-Branch Multilevel Intertemporal NetworkabstractChange detection (CD) of remote sensing (RS) images is mushrooming up accompanied by the on-going innovation of convolutional neural networks (CNNs). Yet with the high-speed technology upgrade, the obstacle that identifies unbalanced variations in foreground–background categories still lies on the table, especially in cases with limited samples and massive interference such as seasonal turnover, illumination intensity, and building reformation. Moreover, to date, neither of the off-the-shelf methods probes the feasibility of direct interaction between bitemporal images before accessing difference features. In this article, we propose a dual-branch multilevel intertemporal network (DMINet) to efficiently and effectively derive the change representations. Specifically, by unifying self-attention (SelfAtt) and cross-attention (CrossAtt) in a single module, we present an intertemporal joint-attention (JointAtt) block to steer the global feature distribution of each input, motivating information coupling between intralevel representations and meanwhile suppressing the task-irrelevant interferences. In addition, centering more on the detection of difference features, a reliable architecture is designed by spotlighting two concerns, i.e., the difference acquisition using subtraction and concatenation as well as the multilevel difference aggregation using incremental feature alignment. Based on a naive backbone without sophisticated structures, i.e., ResNet18, our model outperforms other state-of-the-art (SOTA) methods on four CD datasets, especially in cases with rarely samples. Moreover, the achievement is attained with light overheads. Yuchao Feng, Jiawei Jiang 0002, Honghui Xu 0002, Jianwei Zheng 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | ICIF-Net: Intra-Scale Cross-Interaction and Inter-Scale Feature Fusion Network for Bitemporal Remote Sensing Images Change DetectionabstractChange detection (CD) of remote sensing (RS) images has enjoyed remarkable success by virtue of convolutional neural networks (CNNs) with promising discriminative capabilities. However, CNNs lack the capability of modeling long-range dependencies in bitemporal image pairs, resulting in inferior identifiability against the same semantic targets yet with varying features. The recently thriving Transformer, on the contrary, is warranted, for practice, with global receptive fields. To jointly harvest the local-global features and circumvent the misalignment issues caused by step-by-step downsampling operations in traditional backbone networks, we propose an intra-scale cross-interaction and inter-scale feature fusion network (ICIF-Net), explicitly tapping the potential of integrating CNN and Transformer. In particular, the local features and global features, respectively, extracted by CNN and Transformer, are interactively communicated at the same spatial resolution using a linearized Conv Attention module, which motivates the counterpart to glimpse the representation of another branch while preserving its own features. In addition, with the introduction of two attention-based inter-scale fusion schemes, including mask-based aggregation and spatial alignment (SA), information integration is enforced at different resolutions. Finally, the integrated features are fed into a conventional change prediction head to generate the output. Extensive experiments conducted on four CD datasets of bitemporal (RS) images demonstrate that our ICIF-Net surpasses the other state-of-the-art (SOTA) approaches. Yuchao Feng, Honghui Xu 0002, Jiawei Jiang 0002, Jianwei Zheng 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |