EDBT 2026 Demo / reviewers in the wild / expert
Wei Zhang 0243
dblp:10/4661-243
· DBLP profile ↗
25ranked-venue papers
0as first author
25since 2021 · last 2026
0000-0002-4424-079XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 18 · 18 since 2021Artificial intelligence and machine learning · 8 · 8 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning Personalised Human Internal Cognition from External Expressive Behaviours for Real Personality RecognitionabstractAutomatic real personality recognition (RPR) aims to evaluate human real personality traits from their expressive behaviours. However, most existing solutions generally act as external observers to infer observers' personality impressions based on target individuals' expressive behaviours, which significantly deviate from their real personalities and consistently lead to inferior recognition performance. Inspired by the association between real personality and human internal cognition underlying the generation of expressive behaviours, we propose a novel RPR approach that efficiently simulates personalised internal cognition from external short audio-visual behaviours expressed by target individual. The simulated personalised cognition, represented as a set of network weights that enforce the personalised network to reproduce the individual-specific facial reactions, is further encoded as a graph containing two-dimensional node and edge feature matrices, with a novel 2D Graph Neural Network (2D-GNN) proposed for inferring real personality traits from it. To simulate real personality-related cognition, an end-to-end (E2E) strategy is designed to jointly train our cognition simulation, 2D graph construction, and personality recognition modules. Experiments show our approach’s effectiveness in capturing real personality traits with superior computational efficiency. Xiangyu Kong 0001, Hengde Zhu, Haoqin Sun, Jiayan Gu, Xinyi Ni, Wei Zhang 0243, Shizhe Liu, Siyang Song |
AAAI | 7 |
| 2026 | MAUGen: A Unified Diffusion Approach for Multi-Identity Facial Expression and AU Label GenerationabstractThe lack of large-scale, demographically diverse face images with precise Action Unit (AU) occurrence and intensity annotations has long been recognized as a fundamental bottleneck in developing generalizable facial AU recognition systems. In this paper, we propose MAUGen, a diffusion-based multi-modal framework that jointly generates a large collection of photorealistic facial expressions and anatomically consistent AU labels, including both occurrence and intensity, conditioned on a single descriptive text prompt. Our MAUGen involves two key modules: (1) a Multi-modal Representation Learning (MRL) module that captures the relationships among the paired facial textual description, facial identity, facial expression image, and AU activations within a unified latent space; and (2) a Diffusion-based Image-label Generator (DIG) that decodes the obtained joint representation into aligned facial image-label pairs across diverse identities. Under this framework, we introduce the Multi-Identity Facial Action (MIFA), a large-scale multi-modal (i.e., text descriptions, face images with labels) synthetic dataset that features comprehensive AU annotations and identity variations. Extensive experiments demonstrate that MAUGen outperforms existing methods in synthesizing photorealistic, demographically diverse facial images, along with semantically aligned AU labels. Ye Lou, Ao Gao, Wei Zhang 0243, Siyang Song |
AAAI | 4 |
| 2025 | In2NeCT: Inter-class and Intra-class Neural Collapse Tuning for Semantic Segmentation of Imbalanced Remote Sensing ImagesabstractRemote sensing images (RSIs) are frequently characterized by multi-scale inter-class objects and inconsistently distributed objects due to scene limitations, which would cause a significant data imbalance challenging the corresponding semantic segmentation. Recent methods have leveraged various deep learning techniques to capture high-quality representations for RSI semantic segmentation, but are hardly capable of addressing the afore-mentioned challenge given their limited explorations towards the mechanisms behind the representations. The recently discovered Neural Collapse (NC) phenomenon in computer vision models suggests the simplex equiangular tight frame (ETF) as the optimal representation structure, which has motivated us to observe that the optimal structure of last-layer representations is disrupted and inter-class representations for minor classes tend to become closer to each other beacuse of data imbalance. To address these issues, we propose Inter-class and Intra-class Neural Collapse Tuning (In2NeCT) to optimize the representations that satisfy the simplex ETF, which facilitates the discrimination of inter-class representations and the coherence of intra-class representations. Extensive experiments on three datasets demonstrate that our In2NeCT consistently leads to significant improvements in performance and outperforms the state-of-the-art methods. Junao Shen, Qiyun Hu, Tian Feng 0001, Xinyu Wang 0036, Hui Cui 0002, Sensen Wu, Wei Zhang 0243 |
AAAI | 7 |
| 2025 | PerReactor: Offline Personalised Multiple Appropriate Facial Reaction GenerationabstractIn dyadic human-human interactions, individuals may express multiple different facial reactions in response to the same/similar behaviours expressed by their conversational partners depending on their personalised behaviour patterns. As a result, frequently-employed reconstruction loss-based strategies lead the training of previous automatic facial reaction generation (FRG) models to not only suffer from the 'one-to-many mapping' problem, but also fail to comprehensively consider the quality of the generated facial reactions. Besides, none of them considered such personalised behaviour patterns in generating facial reactions. In this paper, we propose the first adversarial FRG model training strategy which jointly learns appropriateness and realism discriminators to provide comprehensive task-specific supervision for training the target facial reaction generators, and reformulates the 'one-to-many (facial reactions) mapping' training problem as a 'one-to-one (distribution) mapping' training task, i.e., the FRG model is trained to output a distribution representing multiple appropriate/plausible facial reaction from each input human behaviour. In addition, our approach also serves as the first offline FRG approach that considers personalised behaviour patterns in generating of target individuals' facial reactions. Experiments show that our PerReactor not only largely outperformed all existing offline solutions for generating more appropriate, diverse and realistic facial reactions, but also is the first approach that can effectively generate personalised appropriate facial reactions. Hengde Zhu, Xiangyu Kong 0001, Weicheng Xie 0001, Xilin He, Lu Liu 0001, LinLin Shen, Wei Zhang 0243, Hatice Gunes, Siyang Song |
AAAI | 8 |
| 2025 | SSFMamba: Spatial-Spectral Fusion State Space Model for PansharpeningabstractPansharpening aims to fuse the panchromatic (PAN) and low-resolution multispectral (LR-MS) images, finally generating the high-resolution multispectral (HR-MS) images by reconstructing the spatial-spectral properties. Recently, VMamba-based methods built upon the visual state space (VSS) have shown great potential in pansharpening, but they are unable to explicitly characterize the spatial-spectral properties from the given LR-MS and PAN images. Thus, we propose spatial-spectral fusion state space model (SSFMamba), which consists of multi-scale spatial-wise visual state space (MSpa-VSS) block, bi-directional spectral-wise visual state space (BSpe-VSS) block, and gated spatial-spectral fusion (GSSF) block. Specifically, the MSpa-VSS block introduce multi-scale convolution opeartion in the VSS, thus simultaneously modelling spatial dependencies and capturing multi-scale spatial property; the BSpe-VSS block bi-directionally scans the spectral channels for learning continuous spectral property; moreover, we design the GSSF block for fusing spatial-spectral properties in the adaptive manner. Extensive experimental results on several datasets demonstrate the superiority performance of the proposed SSFMamba. Mengting Ma, Mengjiao Zhao, Yizhen Jiang, Wei Zhang 0243 |
ICASSP | 5 |
| 2025 | SLGN: Spatiotemporal Language-Guided Graph Network for Referring Video SegmentationabstractReferring video segmentation uses text descriptions to identify and segment objects. This requires the model to effectively perform spatiotemporal modeling of videos under the guidance of linguistic information. However, previous works have not explicitly considered the differences among objects in videos, nor have they adequately aligned language information with object-level features, leading to misclassifications in complex scenes. In this paper, we rethink the relationship between objects and language in videos, introducing a novel graph-based network (SLGN) for referring video segmentation to address the problem. Specifically, we design a Spatiotemporal Graph Perception (SGP) module that uses temporal, semantics, and positional priors to construct multidimensional edges between objects and employs graph convolution to model their spatiotemporal relationships. Meanwhile, we design a Clue Graph Perception (CGP) module that leverages text descriptions and potential objects to construct an object-word graph, achieving modality alignment at the object level. Experimental results demonstrate that our method outperforms recent representative methods in performance. Rongrong Lian, Zhenkai Wu, Mengting Ma, Wei Zhang 0243 |
ICME | 5 |
| 2025 | HetSSNet: Spatial-Spectral Heterogeneous Graph Learning Network for Panchromatic and Multispectral Images FusionabstractRemote sensing pansharpening aims to reconstruct spatial-spectral properties during the fusion of panchromatic (PAN) images and lowresolution multi-spectral (LR-MS) images, finally generating the high-resolution multi-spectral (HRMS) images. In the mainstream modeling strategies, i.e., CNN and Transformer, the input images are treated as the equal-sized grid of pixels in the Euclidean space. They have limitations in facing remote sensing images with irregular ground objects. Graph is the more flexible structure, however, there are two major challenges when modeling spatial-spectral properties with graph: 1) constructing the customized graph structure for spatial-spectral relationship priors; 2) learning the unified spatial-spectral representation through the graph. To address these challenges, we propose the spatial-spectral heterogeneous graph learning network, named HetSSNet. Specifically, HetSSNet initially constructs the heterogeneous graph structure for pansharpening, which explicitly describes pansharpening-specific relationships. Subsequently, the basic relationship pattern generation module is designed to extract the multiple relationship patterns from the heterogeneous graph. Finally, relationship pattern aggregation module is exploited to collaboratively learn unified spatial-spectral representation across different relationships among nodes with adaptive importance learning from local and global perspectives. Extensive experiments demonstrate the significant superiority and generalization of HetSSNet. Mengting Ma, Yizhen Jiang, Mengjiao Zhao, Wei Zhang 0243 |
ICML | 5 |
| 2025 | High-precision short-term industrial energy consumption forecasting via parallel-NN with Adaptive Universal Decomposition
Fan Yang 0100, Shuning Ge, Jian Liu 0046, Ke Yan 0001, Ao Gao, Yijie Dong, Wei Zhang 0243 |
Expert Syst. Appl. | 8 |
| 2025 | CDxLSTM: Boosting Remote Sensing Change Detection With Extended Long Short-Term MemoryabstractIn complex scenes and varied conditions, effectively integrating spatial-temporal context is crucial for accurately identifying changes. However, current RS-CD methods lack a balanced consideration of performance and efficiency. CNNs lack global context, Transformers are computationally expensive, and Mambas face CUDA dependence and local correlation loss. In this paper, we propose CDXLSTM, with a core component that is a powerful XLSTM-based feature enhancement layer, integrating the advantages of linear computational complexity, global context perception, and strong interpret-ability. Specifically, we introduce a scale-specific Feature Enhancer layer, incorporating a Cross-Temporal Global Perceptron customized for semantic-accurate deep features, and a Cross-Temporal Spatial Refiner customized for detail-rich shallow features. Additionally, we propose a Cross-Scale Interactive Fusion module to progressively interact global change representations with spatial responses. Extensive experimental results demonstrate that CDXLSTM achieves state-of-the-art performance across three benchmark datasets, offering a compelling balance between efficiency and accuracy. Code is available at https://github.com/xwmaxwma/rschange. Zhenkai Wu, Rongrong Lian, Wei Zhang 0243 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2025 | Unleashing Fourier-Domain Potential: Spatial-Spectral Reconstruction Framework for Remote Sensing PansharpeningabstractPansharpening aims to generate high-resolution multispectral (HR-MS) images by fusing the corresponding low-resolution multispectral (LR-MS) and high-resolution panchromatic (PAN) images. While the spatial-spectral properties modeling plays crucial roles in generating high-quality HR-MS images, existing approaches suffer from: 1) directly modeling the entangled spatial-spectral properties; and 2) lacking task-specific priors for spatial and spectral properties modeling. This paper proposes a novel Fourier domain-based approach (Fourier-SSR) to address these problems, where phase and amplitude components of PAN and LR-MS are considered to individually model spatial and spectral properties for the target HR-MS image. Our Fourier-SSR is motivated by the crucial findings in Fourier domain: (i) manifestation of spatial-spectral properties, i.e., the spatial and spectral properties of remote sensing images can be individually manifested in their phase and amplitude components in Fourier domain; (ii) spatial property-related prior, i.e., only reconstructing the phase component of PAN image in Fourier domain, could generate spatial property required for target HR-MS image; and (iii) spectral property-related prior, i.e., jointly models the amplitude components of PAN and LR-MS images could generate the required spectral property for target HR-MS image. Based on the aforementioned findings, we design aFourier-guided spatial Mixerand aFourier-guided spectral Mixer, which innovatively employ complex feature interaction strategies to individually reconstruct the phase and amplitude components for target HR-MS. Experiments show that our methods unleash Fourier domain potential in individually modeling spatial and spectral properties for the target HR-MS image, leading to superior performance over previous state-of-the-art. Our code is provided in Supplementary Material. Mengting Ma, Yizhen Jiang, Mengjiao Zhao, Wei Zhang 0243, Siyang Song |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | LOGCAN++: Adaptive Local-Global Class-Aware Network for Semantic Segmentation of Remote Sensing ImagesabstractRemote sensing images are usually characterized by complex backgrounds, scale and orientation variations, and large intraclass variance. General semantic segmentation methods usually fail to fully investigate the above issues, and thus their performances on remote sensing image segmentation are limited. In this article, we propose our LOGCAN++, a semantic segmentation model customized for remote sensing images, which is made up of a global class-aware (GCA) module and several local class-aware (LCA) modules. The GCA module captures global representations for class-level context modeling to reduce the interference of background noise. The LCA module generates local class representations as intermediate perceptual elements to indirectly associate pixels with the global class representations, targeting dealing with the large intraclass variance problem. In particular, we introduce affine transformations in the LCA module for adaptive extraction of local class representations to effectively tolerate scale and orientation variations in remote sensing images. Extensive experiments on three benchmark datasets show that our LOGCAN++ outperforms current mainstream general and remote sensing semantic segmentation methods and achieves a better trade-off between speed and accuracy. Rongrong Lian, Zhenkai Wu, Fan Yang 0100, Mengting Ma, Sensen Wu, Zhenhong Du, Wei Zhang 0243, Siyang Song |
IEEE Trans. Geosci. Remote. Sens. | 9 |
| 2025 | Supervised Detail-Guided Multiscale State-Space Model for Pan-SharpeningabstractPan-sharpening reconstructs the high-resolution multispectral (HR-MS) image from its corresponding panchromatic (PAN) image and low-resolution multispectral (LR-MS) image. However, existing deep learning (DL)-based pan-sharpening methods typically suffer from three challenges: 1) the vanilla LR-MS image upsampling employed by them fails to consider domain knowledge, thereby disregarding crucial information; 2) while remote sensing images exhibit multiscale complex land features, existing methods fail to fully exploit crucial multiscale spatial information, that is, scale transformation layers in their models are not effective; and 3) existing convolutional neural network (CNN) and transformer-based pan-sharpening backbones are constrained by inherent local receptive fields or quadratic computational complexity, making them difficult to balance their effectiveness and efficiency. To address these issues, we propose a novel supervised detail-guided multiscale state-space model for pan-sharpening, namely SDMSPan. Our SDMSPan consists of three residual state-space modules (Res-SSMs) that are responsible for handling image information at three spatial scales, where each Res-SSM aims to model both local and long-range dependencies between PAN and LR-MS images at a specific spatial scale with lower computational cost. Between each pair of Res-SSMs, a novel detail-guided upsampling block (DGUB) is proposed to apply spatial details of the PAN image to guide effective and task-aware intermediate feature upsampling, where a novel multiscale intermediate spatial-spectral supervision strategy is also proposed to supervise the training of every DGUB. Experimental results demonstrate that our proposed approach significantly outperforms other state-of-the-art methods in performance. Our code is provided athttps://github.com/zhaomengjiao123/SDMSPan. Mengjiao Zhao, Mengting Ma, Wei Zhang 0243, Siyang Song |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | CGMGM: A Cross-Gaussian Mixture Generative Model for Few-Shot Semantic SegmentationabstractFew-shot semantic segmentation (FSS) aims to segment unseen objects in a query image using a few pixel-wise annotated support images, thus expanding the capabilities of semantic segmentation. The main challenge lies in extracting sufficient information from the limited support images to guide the segmentation process. Conventional methods typically address this problem by generating single or multiple prototypes from the support images and calculating their cosine similarity to the query image. However, these methods often fail to capture meaningful information for modeling the de facto joint distribution of pixel and category. Consequently, they result in incomplete segmentation of foreground objects and mis-segmentation of the complex background. To overcome this issue, we propose the Cross Gaussian Mixture Generative Model (CGMGM), a novel Gaussian Mixture Models~(GMMs)-based FSS method, which establishes the joint distribution of pixel and category in both the support and query images. Specifically, our method initially matches the feature representations of the query image with those of the support images to generate and refine an initial segmentation mask. It then employs GMMs to accurately model the joint distribution of foreground and background using the support masks and the initial segmentation mask. Subsequently, a parametric decoder utilizes the posterior probability of pixels in the query image, by applying the Bayesian theorem, to the joint distribution, to generate the final segmentation mask. Experimental results on PASCAL-5i and COCO-20i datasets demonstrate our CGMGM's effectiveness and superior performance compared to the state-of-the-art methods. Junao Shen, Kun Kuang 0001, Xinyu Wang 0036, Tian Feng 0001, Wei Zhang 0243 |
AAAI | 6 |
| 2024 | Deep Unfolding Network with Spatial-spectral Perception Enhanced for Pan-sharpening
Mengjiao Zhao, Mengting Ma, Ao Gao, Siyang Song, Wei Zhang 0243 |
BMVC | 6 |
| 2024 | CROCFUN: Cross-Modal Conditional Fusion Network for PansharpeningabstractPansharpening aims to reconstruct a high-fidelity multispectral (HR-MS) image by fusing a multispectral (MS) image and a panchromatic (PAN) image. However, conventional pansharpening methods often struggle to address the modal gap between PAN and MS images. In this paper, we propose a novel cross-modal conditional fusion network (CroCFuN) for pansharpening, which builds upon recent advancements in optimal transport (OT) theory. Specifically, we formulate the modal alignment in pansharpening as an OT problem and thus design a feature alignment module (FAM) to adjust the MS image features to align with the PAN image features. Meanwhile, we propose a feature-activation normalization fusion module (FANFM), which adopts the multi-stage conditional fusion strategy to generate high-quality fusion features. Experimental results demonstrate that our CroCFuN outperforms recent representative methods for pan-sharpening regarding both visual and quantitative qualities. Code will be available at https://github.com/Florina2333. Mengting Ma, Chenlu Hu, Huanting Zhang, Tian Feng 0001, Wei Zhang 0243 |
ICASSP | 6 |
| 2024 | Frequency-Spatial Domain Information Fusion Network for Pan-SharpeningabstractPan-sharpening aims to fuse panchromatic (PAN) images with low-resolution multi-spectral (LR-MS) images to generate high-resolution multi-spectral (HR-MS) images. Despite the impressive performance of existing learning-based methods, they are constrained by coarse fusion strategies in frequency or spatial domain. In this paper, we discover that PAN images can provide all the spatial textures required for HR-MS images, while spectral information must be provided jointly by PAN and LR-MS images. Inspired by this, we propose a noval frequency-spatial domain information fusion network for pan-sharpening, called FSDNet. Specifically, we design a Dual-Domain Information Processing Module (DDPM) to construct FSDNet. It consists of a Frequency Domain Feature Processing Block (FDB), a Spatial Domain Information Processing Block (SDB), and an Information Fusion Block (IFB). The FDB in the frequency domain uses the Adaptive Amplitude Fusion Block (AAFB) and convolution layers to finely modulate amplitude and phase components, exploring global information. The SDB uses cascaded residual blocks to capture and enhance local information in the spatial domain. The IFB based on invertible neural networks (INNs) introduces Multi-Scale Self-Attention Block (MSAB), achieves effective information fusion and reduce information loss. Extensive experiments on the QuickBird and GaoFen-2 datasets demonstrate the effectiveness and superiority of our method. Mengjiao Zhao, Mengting Ma, Ao Gao, Wei Zhang 0243 |
ICIP | 4 |
| 2024 | SSETPAN: Spatial-Spectral Enhanced Transformer based network for pansharpeningabstractPansharpening aims for effective spatial-spectral fusion of low-resolution multispectral (LR-MS) and panchromatic (PAN) images, yielding high-resolution multispectral (HR-MS) images. PAN images contain rich spatial details and LR-MS images contain abundant spectral features. However, most of the learning-based methods ignore their distinct attributes, and employ weaker fusion strategies. Besides, Transformer has recently gained considerable popularity in target feature extraction. Therefore, our paper develops a novel Transformer-based network for pansharpening, dubbed Spatial-Spectral Enhanced Transformer based network (SSETPAN), proficient in fine-grained spatial-spectral feature extraction and interaction. SSET-PAN comprises three main modules: channel-wise transformer (CTM), spatial-wise transformer (STM), and adaptive spatial-spectral feature fusion (ASSFM). CTM extracts LR-MS explicit spectral features, while STM obtains high-quality spatial features. ASSFM achieves adaptive spatial-spectral feature fusion via kernel-varied convolution combination. Extensive experiments on GaoFen-2 and WorldView-3 datasets demonstrate that SSETPAN achieves favorable performance against existing pansharpening methods. Huanting Zhang, Mengting Ma, Xinyu Wang 0036, Wei Zhang 0243 |
ICME | 6 |
| 2024 | DuCoFPan: Dual-Condition Flow-based Network for Pan-sharpeningabstractPan-sharpening aims to reconstruct high-resolution multi-spectral (HR-MS) images from panchromatic (PAN) images and low-resolution multi-spectral (LR-MS) images. Despite demonstrated performance, existing learning-based methods struggle to address the ill-posed problem from spectral and spatial perspectives. Generally, mitigating the ill-posed issue involves obtaining the probability distribution of HR-MS images. In this paper, we propose DuCoFPan, a novel dual-condition flow-based network for pan-sharpening, which learns the distributions of HR-MS images guided by spectral and spatial conditions, respectively. Specifically, we design a Dual-Condition Flow Module (DCFM) that adopts spectral and spatial conditions through reversible affine transformations. For condition injection, we present a Spectral-based Condition Injection Block (SPEB) capturing fine-grained spectral features in the Fourier domain and a Spatial-based Condition Injection Block (SPAB) extracting spatial features. Additionally, we devise a Feature Interaction Block (FIB) to promote information flow. Extensive experiments on QuickBird and GaoFen-2 datasets demonstrate the effectiveness and superiority of our method. Mengjiao Zhao, Mengting Ma, Xinyu Wang 0036, Ao Gao, Wei Zhang 0243 |
ICME | 7 |
| 2024 | DOCNet: Dual-Domain Optimized Class-Aware Network for Remote Sensing Image SegmentationabstractThe spatial attention mechanism has been frequently employed for the semantic segmentation of remote sensing images, given its renowned capability to model long-range dependencies. As remote sensing images often exhibit intricate backgrounds, significant intraclass variability, and a foreground-background imbalance, spatial attention mechanism-based methods somehow tend to introduce an extensive amount of background context through intensive affinity operations, causing unsatisfactory segmentation outcomes. While several class-aware methods attempt to attenuate the interference of background context by generating class representations as representative features, they still encounter challenges related to independent correlation calculation and single-confidence scale class representations. We introduce a dual-domain optimized class-aware network designed to address these challenges. In the semantic domain, we use category confidence as a scaling criterion to derive class representations at multiple confidence scales, effectively modeling pixel-class relationships. In the spatial domain, we leverage pixel-class relationships and their consensus to enhance relevant correlations while suppressing erroneous ones. Experimental results on three datasets demonstrate that the proposed method surpasses previous state-of-the-art ones for remote sensing image segmentation. Code is available athttps://github.com/xwmaxwma/rssegmentation. Rui Che, Xinyu Wang 0036, Mengting Ma, Sensen Wu, Tian Feng 0001, Wei Zhang 0243 |
IEEE Geosci. Remote. Sens. Lett. | 7 |
| 2023 | Log-Can: Local-Global Class-Aware Network For Semantic Segmentation of Remote Sensing ImagesabstractRemote sensing images are known of having complex backgrounds, high intra-class variance and large variation of scales, which bring challenge to semantic segmentation. We present LoG-CAN, a multi-scale semantic segmentation network with a global class-aware (GCA) module and local class-aware (LCA) modules to remote sensing images. Specifically, the GCA module captures the global representations of class-wise context modeling to circumvent back-ground interference; the LCA modules generate local class representations as intermediate aware elements, indirectly associating pixels with global class representations to reduce variance within a class; and a multi-scale architecture with GCA and LCA modules yields effective segmentation of objects at different scales via cascaded refinement and fusion of features. Through the evaluation on the ISPRS Vaihingen dataset and the ISPRS Potsdam dataset, experimental results indicate that LoG-CAN outperforms the state-of-the-art methods for general semantic segmentation, while significantly reducing network parameters and computation. Code is available at https://github.com/xwmaxwma/rssegmentation. Mengting Ma, Chenlu Hu, Zhiyuan Song, Tian Feng 0001, Wei Zhang 0243 |
ICASSP | 7 |
| 2023 | STANet: Spatiotemporal Adaptive Network for Remote Sensing ImagesabstractSpatiotemporal fusion aims to generate remote sensing images with high spatial and temporal resolutions. Conventional spatiotemporal fusion methods usually use convolution operations for feature extraction, which limits their capability of capturing long-range dependencies. Meanwhile, the significant difference of spatial resolutions of images brings great difficulty to the reconstruction of detailed textures. To address these issues, we propose a GAN-based multi-stage spatiotemporal adaptive network (STANet) for remote sensing images using temporal feature refinement and spatial texture transfer. In particular, we design a temporal interaction module (TIM) to extract useful information on surface changes over time, using a cross-temporal gating mechanism that emphasizes feature changes throughout the task. We employ the adaptive instance normalization (AdaIN) layers to learn the global spatial correlation via texture transfer from the fine image to the coarse image. Experiments on two datasets show that the proposed method outperforms other state-of-the-art methods in several metrics. Chenlu Hu, Mengting Ma, Huanting Zhang, Dun Wu, Wei Zhang 0243 |
ICIP | 7 |
| 2023 | SACANet: scene-aware class attention network for semantic segmentation of remote sensing imagesabstractSpatial attention mechanism has been widely used in semantic segmentation of remote sensing images given its capability to model long-range dependencies. Many methods adopting spatial attention mechanism aggregate contextual information using direct relationships between pixels within an image, while ignoring the scene awareness of pixels (i.e., being aware of the global context of the scene where the pixels are located and perceiving their relative positions). Given the observation that scene awareness benefits context modeling with spatial correlations of ground objects, we design a scene-aware attention module based on a refined spatial attention mechanism embedding scene awareness. Besides, we present a local-global class attention mechanism to address the problem that general attention mechanism introduces excessive background noises while hardly considering the large intra-class variance in remote sensing images. In this paper, we integrate both scene-aware and class attentions to propose a scene-aware class attention network (SACANet) for semantic segmentation of remote sensing images. Experimental results on three datasets show that SACANet outperforms other state-of-the-art methods and validate its effectiveness. Code is available at https://github.com/xwmaxwma/rssegmentation. Rui Che, Tingfeng Hong, Mengting Ma, Tian Feng 0001, Wei Zhang 0243 |
ICME | 7 |
| 2023 | STNet: Spatial and Temporal feature fusion network for change detection in remote sensing imagesabstractAs an important task in remote sensing image analysis, remote sensing change detection (RSCD) aims to identify changes of interest in a region from spatially co-registered multi-temporal remote sensing images, so as to monitor the local development. Existing RSCD methods usually formulate RSCD as a binary classification task, representing changes of interest by merely feature concatenation or feature subtraction and recovering the spatial details via densely connected change representations, whose performances need further improvement. In this paper, we propose STNet, a RSCD network based on spatial and temporal feature fusions. Specifically, we design a temporal feature fusion (TFF) module to combine bitemporal features using a cross-temporal gating mechanism for emphasizing changes of interest; a spatial feature fusion module is deployed to capture fine-grained information using a cross-scale attention mechanism for recovering the spatial details of change representations. Experimental results on three benchmark datasets for RSCD demonstrate that the proposed method achieves the state-of-the-art performance. Code is available at https://github.com/xwmaxwma/rschange. Tingfeng Hong, Mengting Ma, Tian Feng 0001, Wei Zhang 0243 |
ICME | 7 |
| 2023 | DBDAN: Dual-Branch Dynamic Attention Network for Semantic Segmentation of Remote Sensing Images
Rui Che, Tingfeng Hong, Xinyu Wang 0036, Tian Feng 0001, Wei Zhang 0243 |
PRCV (4) | 6 |
| 2023 | MAPMaN: Multi-Stage U-Shaped Adaptive Pattern Matching Network for Semantic Segmentation of Remote Sensing ImagesabstractAbstract Remote sensing images (RSIs) often possess obvious background noises, exhibit a multi‐scale phenomenon, and are characterized by complex scenes with ground objects in diversely spatial distribution pattern, bringing challenges to the corresponding semantic segmentation. CNN‐based methods can hardly address the diverse spatial distributions of ground objects, especially their compositional relationships, while Vision Transformers (ViTs) introduce background noises and have a quadratic time complexity due to dense global matrix multiplications. In this paper, we introduce Adaptive Pattern Matching (APM), a lightweight method for long‐range adaptive weight aggregation. Our APM obtains a set of pixels belonging to the same spatial distribution pattern of each pixel, and calculates the adaptive weights according to their compositional relationships. In addition, we design a tiny U‐shaped network using the APM as a module to address the large variance of scales of ground objects in RSIs. This network is embedded after each stage in a backbone network to establish a Multi‐stage U‐shaped Adaptive Pattern Matching Network (MAPMaN), for nested multi‐scale modeling of ground objects towards semantic segmentation of RSIs. Experiments on three datasets demonstrate that our MAPMaN can outperform the state‐of‐the‐art methods in common metrics. The code can be available at https://github.com/INiid/MAPMaN . Tingfeng Hong, Xinyu Wang 0036, Rui Che, Chenlu Hu, Tian Feng 0001, Wei Zhang 0243 |
Comput. Graph. Forum | 7 |