EDBT 2026 Demo / reviewers in the wild / expert
Yaxiong Chen
dblp:10/9546
· DBLP profile ↗
73ranked-venue papers
26as first author
66since 2021 · last 2026
0000-0002-2903-6723ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 37 · 19 first-author · 36 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 5 first-author · 22 since 2021Artificial intelligence and machine learning · 21 · 5 first-author · 17 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ProPL: Universal Semi-Supervised Ultrasound Image Segmentation via Prompt-Guided Pseudo-LabelingabstractExisting approaches for the problem of ultrasound image segmentation, whether supervised or semi-supervised, are typically specialized for specific anatomical structures or tasks, limiting their practical utility in clinical settings. In this paper, we pioneer the task of universal semi-supervised ultrasound image segmentation and propose ProPL, a framework that can handle multiple organs and segmentation tasks while leveraging both labeled and unlabeled data. At its core, ProPL employs a shared vision encoder coupled with prompt-guided dual decoders, enabling flexible task adaptation through a prompting-upon-decoding mechanism and reliable self-training via an uncertainty-driven pseudo-label calibration (UPLC) module. To facilitate research in this direction, we introduce a comprehensive ultrasound dataset spanning 5 organs and 8 segmentation tasks. Extensive experiments demonstrate that ProPL outperforms state-of-the-art methods across various metrics, establishing a new benchmark for universal ultrasound image segmentation. Yaxiong Chen, Qicong Wang, Jingliang Hu, Yilei Shi, Shengwu Xiong 0001, Xiao Xiang Zhu 0001, Lichao Mou |
AAAI | 1 |
| 2026 | Exploring a double task learning framework for makeup transfer
Zhaoyang Sun, Shengwu Xiong 0001, Yaxiong Chen |
Eng. Appl. Artif. Intell. | 3 |
| 2026 | HLF-LKNet: A self-supervised denoising network with high-low frequency fusion and large-kernel high-frequency enhancement
Yaxiong Chen, Yutong Yang, Yongqing Yan, Sai Zhong, Shili Xiong |
Neurocomputing | 1 |
| 2026 | Medical video segmentation model based on text reference
Zichan Li, Qiang Zhang 0031, Junjian Hu, Yaxiong Chen |
Knowl. Based Syst. | 4 |
| 2026 | WCEDNet: A Weighted Cascaded Encoder-Decoder Network for Hyperspectral Change Detection Based on Spatial-Spectral Difference FeaturesabstractThe core of hyperspectral change detection lies in accurately capturing spectral feature differences across different temporal phases to determine whether surface objects have changed. Since spectral variations of different ground objects often manifest more prominently in specific wavelength bands, we design a Weighted Cascaded Encoder-Decoder Network based on spatial-spectral difference features for hyperspectral change detection. Firstly, unlike conventional change detection frameworks based on siamese networks, our proposed single-branch approach focuses more intensively on extracting spatial-spectral difference features. Secondly, the weighted cascaded structure introduced in the encoder stage enables differential attention to different bands, enhancing focus on spectral bands with high responsiveness. Furthermore, we have developed a spatial-spectral cross-attention module to model intra-feature correlations within spatial and spectral domains. Our method was evaluated on three challenging hyperspectral change detection datasets, and experimental results demonstrate its superior performance compared to competitive models. The detailed code has been open-sourced at https://github.com/WUTCM-Lab/WCEDNet. Bo Zhang 0069, Yaxiong Chen, Ruilin Yao, Shengwu Xiong 0001 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2026 | ArtGlyphDiffuser: Text-driven artistic glyph generation via Style-to-CLIP Projection and Multi-Level Controlled diffusion
Xiongbo Lu, Yaxiong Chen, Shengwu Xiong 0001 |
Pattern Recognit. | 2 |
| 2026 | ReCoTR: Reducing Semantic Cognitive Shift via Dual-Consensus Token Compression for Remote Sensing Image-Text RetrievalabstractWith the rapid advancement of vision-language models (VLMs) in general-purpose settings, their application to cross-modal retrieval and semantic understanding of large-scale multimodal remote sensing (RS) data is emerging as a key enabler for urban governance, environmental monitoring, and disaster response. However, the pervasive issue of semantic shift in RS image poses a significant challenge to the transferability of pre-trained VLMs. To address this limitation, we propose ReCoTR, an enhanced CLIP-based cross-modal retrieval framework tailored for remote sensing applications. ReCoTR tackles region-level granularity bias and contextual semantic drift through a Dual Consensus Token Evaluation (DCTE) module, which leverages a mixture-of-experts strategy to fuse inter-modal semantic consensus with intra-modal structural consistency, enabling fine-grained estimation of semantic confidence for visual tokens. Moreover, to mitigate representational contamination caused by background noise, we introduce the Semantic Confidence Token Compression (SCTC) module. This module selectively filters and aggregates tokens with high semantic relevance, thus reducing redundancy and alleviating the noise amplification inherent in CLIP's average pooling. Experimental results on three benchmark RS cross-modal retrieval datasets demonstrate that ReCoTR consistently outperforms existing methods on bidirectional image-text retrieval tasks, validating its effectiveness and robustness in remote sensing semantic alignment scenarios. Our source codes are available at: https://github.com/Jerry710/ReCoTR.git. Jirui Huang, Yaxiong Chen, Chuang Du, Shengwu Xiong 0001, Xiaoqiang Lu |
IEEE Trans. Image Process. | 2 |
| 2025 | RealisID: Scale-Robust and Fine-Controllable Identity Customization via Local and Global ComplementationabstractRecently, the success of text-to-image synthesis has greatly advanced the development of identity customization techniques, whose main goal is to produce realistic identity-specific photographs based on text prompts and reference face images. However, it is difficult for existing identity customization methods to simultaneously meet the various requirements of different real-world applications, including the identity fidelity of small face, the control of face location, pose and expression, as well as the customization of multiple persons. To this end, we propose a scale-robust and fine-controllable method, namely RealisID, which learns different control capabilities through the cooperation between a pair of local and global branches. Specifically, by using cropping and up-sampling operations to filter out face-irrelevant information, the local branch concentrates the fine control of facial details and the scale-robust identity fidelity within the face region. Meanwhile, the global branch manages the overall harmony of the entire image. It also controls the face location by taking the location guidance as input. As a result, RealisID can benefit from the complementarity of these two branches. Finally, by implementing our branches with two different variants of ControlNet, our method can be easily extended to handle multi-person customization, even only trained on single-person datasets. Extensive experiments and ablation studies indicate the effectiveness of RealisID and verify its ability in fulfilling all the requirements mentioned above. Zhaoyang Sun, Yaxiong Chen, Shengwu Xiong 0001 |
AAAI | 5 |
| 2025 | SwinV2-SF: A Parallel Polyp Segmentation Architecture Leveraging Swin Transformer V2 and Spatial-Frequency Dual-Domain AttentionabstractPolyps are abnormally growing tissues on the intestinal mucosal surface, often serving as precursors to cellular carcinogenesis. Accurate polyp segmentation is crucial for the diagnosis and treatment of early-stage colorectal cancer. Despite significant advancements in this field enabled by deep learning techniques, complex backgrounds in the gastrointestinal tract-such as luminal folds, air bubbles, and gastric acid secretions-still pose major challenges for existing models. These complexities hinder the models' ability to accurately perceive polyp boundaries and effectively extract contextual information. Due to the prevalent neglect of frequency-domain features in existing segmentation methods, this paper proposes SwinV2-SF-a dual-branch hybrid model integrating the spatial-frequency dualdomain network SFUNet with Swin Transformer V2. It extracts frequency-domain features through Residual Fast Fourier Transform blocks (RFFT), combines spatial and frequency features via a dynamically weighted Spatial-Frequency Attention Assembly (SFAA), and introduces the AutoDFLoss function, to resolve collectively boundary ambiguity and class imbalance. We evaluated the segmentation performance of SwinV2-SF on benchmark datasets including Kvasir-SEG and ETIS-Larib. Experimental results demonstrate that our model achieves competitive performance in terms of mDice and mIoU metrics, notably outperforming state-of-the-art models on the Kvasir-SEG dataset. The code is located at https://github.com/LeedJohn/SwinV2-SF. Zhengqin Wang, Mingming Guo, Zihan Lu, Yaxiong Chen, Qiang Zhang 0031 |
BIBM | 5 |
| 2025 | LAMM-ViT: AI Face Detection via Layer-Aware Modulation of Region-Guided AttentionabstractDetecting AI-synthetic faces presents a critical challenge: it is hard to capture consistent structural relationships between facial regions across diverse generation techniques. Current methods, which focus on specific artifacts rather than fundamental inconsistencies, often fail when confronted with novel generative models. To address this limitation, we introduce Layer-aware Mask Modulation Vision Transformer (LAMM-ViT), a Vision Transformer designed for robust facial forgery detection. This model integrates distinct Region-Guided Multi-Head Attention (RG-MHA) and Layer-aware Mask Modulation (LAMM) components within each layer. RG-MHA utilizes facial landmarks to create regional attention masks, guiding the model to scrutinize architectural inconsistencies across different facial areas. Crucially, the separate LAMM module dynamically generates layer-specific parameters, including mask weights and gating values, based on network context. These parameters then modulate the behavior of RG-MHA, enabling adaptive adjustment of regional focus across network depths. This architecture facilitates the capture of subtle, hierarchical forgery cues ubiquitous among diverse generation techniques, such as GANs and Diffusion Models. In cross-model generalization tests, LAMM-ViT demonstrates superior performance, achieving 94.09% mean ACC (a +5.45% improvement over SoTA) and 98.62% mean AP (a +3.09% improvement). These results demonstrate LAMM-ViT’s exceptional ability to generalize and its potential for reliable deployment against evolving synthetic media threats.The code is available at https://github.com/WHUT-ZJL/LAMM-ViT. Jiangling Zhang, Jirui Huang, Yaxiong Chen |
ECAI | 4 |
| 2025 | ROME: Radar Sparsity Improvement and Omnimodal Enhancement for 3D Object Detection in Bird's Eye ViewsabstractCombining omnimodal feature interaction using LiDAR, surround-view camera, and Radar to form a network has a great guarantee for the safety of autonomous driving, but most of the current omnimodal fusion methods focus on the interaction enhancement of LiDAR and surround-view camera, ignoring the focus on Radar. Enhancing the contextual representation of Radar can ensure better all-weather capability of the perceptual network. To this end, we design the ROME method based on Radar sparsity improvement to better enhance the performance and robustness of the model in terms of alleviating Radar sparsity shortcomings. Firstly, we design the Autocorrelation Point Enhancement (APE) module to improve Radar sparsity leveraging the point-to-point autocorrelation of Radar. Moreover, for omnimodal Bird’s Eye View (BEV) features, an Omnimodal Adaptive Fusion (OAF) module is designed to improve the robustness of BEV features. With the improved Radar modality, the performance of BEV features for the whole driving scene is further improved. Comprehensive experiments on the nuScenes dataset and comparisons with state-of-the-art methods demonstrate the advantages of our proposed method. Yilong Guo, Junyin Wang, Chenghu Du, Shengwu Xiong 0001, Yaxiong Chen |
ICASSP | 5 |
| 2025 | Exploring Flexibility in Incremental Few-Shot Object DetectionabstractIncremental few-shot object detection (iFSD) is critical for real-world applications, enabling rapid adaptation to novel categories with minimal data while mitigating catastrophic forgetting. However, existing methods lack flexibility, particularly in feature representation. The pursuit of a flexible approach to iFSD presents a substantial challenge. To address this, we propose an Attention-Based Feature Aggregation (AFA) that dynamically refines feature representations guided by limited support samples, and a Conditional Classifier (CC) that dynamically refines the generated class prototypes based on the existing knowledge, while conditioning on the limited support images, enhancing flexibility and adaptability. We conducted comprehensive experiments on the MS COCO and LVIS datasets to validate the superiority of our approach. Dongdong Gong, Tengfei Gong, Yaxiong Chen, Jinglin Yuan, Shengwu Xiong 0001 |
ICME | 3 |
| 2025 | AnyArtisticGlyph: Multilingual Controllable Artistic Glyph GenerationabstractArtistic Glyph Image Generation (AGIG) differs from current creativity-focused generation models by offering finely controllable deterministic generation. It transfers the style of a reference image to a source while preserving its content. Although advanced and promising, current methods may reveal flaws when scrutinizing synthesized image details, often producing blurred or incorrect textures, posing a significant challenge. Hence, we introduce AnyArtisticGlyph, a diffusion-based, multilingual controllable artistic glyph generation model. It includes a font fusion and embedding module, which generates latent features for detailed structure creation, and a vision-text fusion and embedding module that uses the CLIP model to encode references and blends them with transformation caption embeddings for seamless global image generation. Moreover, we incorporate a coarse-grained feature-level loss to enhance generation accuracy. Experiments show that it produces natural, detailed artistic glyph images with state-of-the-art performance. Our project will be open-sourced on https://github.com/jiean001/AnyArtisticGlyph to advance text generation technology. Xiongbo Lu, Yaxiong Chen, Shengwu Xiong 0001 |
ICME | 2 |
| 2025 | Location-Oriented Sound Event Localization and Detection with Spatial Mapping and Regression LocalizationabstractSound Event Localization and Detection (SELD) combines the Sound Event Detection (SED) with the corresponding Direction Of Arrival (DOA). Recently, adopted event-oriented multi-track methods affect the generality in polyphonic environments due to the limitation of the number of tracks. To enhance the generality in polyphonic environments, we propose Spatial Mapping and Regression Localization for SELD (SMRL-SELD). SMRL-SELD segments the 3D spatial space, mapping it to a 2D plane, and a new regression localization loss is proposed to help the results converge toward the location of the corresponding event. SMRL-SELD is location-oriented, allowing the model to learn event features based on orientation. Thus, the method enables the model to process polyphonic sounds regardless of the number of overlapping events. We conducted experiments on STARSS23 and STARSS22 datasets and our proposed SMRL-SELD outperforms the existing SELD methods in overall evaluation and polyphony environments. Xueping Zhang, Yaxiong Chen, Ruilin Yao, Yunfei Zi, Shengwu Xiong 0001 |
ICME | 2 |
| 2025 | Multimodal Co-aware Scale-Spatial Network for Medical Image Segmentation
Enming Huang, Teng Fei Ggong, Yaxiong Chen, Shengwu Xiong 0001 |
PRCV (14) | 3 |
| 2025 | Progressive language-aware encoding and decoding for referring expression comprehension
Yichen Zhao, Yaxiong Chen, Shengwu Xiong 0001 |
Sci. China Inf. Sci. | 2 |
| 2025 | Cross-Domain Density Map-Generated Ship Counting Network for Remote Sensing ImageabstractIn recent years, with the continuous development of remote sensing technology, maritime ship monitoring has become an important research area. Accurately counting the number of ships in remote sensing images is crucial for maritime traffic safety, fisheries management, and marine environmental protection. Existing methods typically use Gaussian kernel functions to generate density maps; however, due to the varied shapes of ships that do not conform to the Gaussian kernel, the resulting density maps fail to accurately reflect the true forms of ships, thereby affecting counting performance. To overcome these limitations, we introduce the cross-domain density map-generated ship counting network (CDDMNet). This network innovatively incorporates a cross-domain feature fusion module (CDFFM), which effectively adapts to ships of varying sizes and shapes. In addition, we have introduced the feature correlation regularization constraint (FCRC) and the integrated loss function, which effectively overcome the disturbances that may arise from variations in ship sizes and enhance the model’s adaptability to changes in ship types and environmental conditions. Experimental results show that the CDDMNet has achieved excellent performance across multiple remote sensing image datasets. Finally, on the RSOC dataset, the mean absolute error (MAE) reached 52.80 and the root mean squared error (RMSE) reached 69.77. Yaxiong Chen, Qijian Li, Kai Yan 0001, Shengwu Xiong 0001 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2025 | Temporal-Aware Spatial Interaction Transformer for Crop Yield Prediction Based on Multisensor Satellite Image Time SeriesabstractCrop yield prediction is crucial for agricultural decision making. Satellite Image Time Series (SITS) data, which provide continuous temporal observations of vegetation changes, have become a standard for accurate prediction. Recent studies have shown that the integration of multimodal data from different satellite sensors significantly enhances the performance of SITS-based crop yield predictions. However, existing methods often rely on simplistic combinations of multimodal data. Temporal inconsistency between different modals is not considered. In addition, the influence of spatial interaction on crop growth is ignored. To address this issue, we propose the TASI-Transformer (Temporal-Aware Spatial Interaction Transformer) for crop yield prediction using multisensor satellite image time series, which incorporates two innovative modules: a Temporal Enhanced Position Encoding (TEPE) module that incorporates crop growth dates to extract unique temporal information from different satellite time series. A Spatial Enhanced Multimodal Interaction (SEMI) module that learns the impact of spatial relationships between different regions and the interaction between multiple modals. Experimental results on the SICKLE and CROPNET datasets demonstrate that the proposed method achieves state-of-the-art performance, with a Mean Absolute Percentage Error (MAPE) as low as 26.99% in SICKLE using actual season data, and a Root Mean Squared Error (RMSE) of 6.85, Coefficient of Determination(R²) of 0.51, and Pearson Correlation (CORR) of 0.71 in CROPNET, outperforming existing methods. Tengfei Gong, Xinchao Zhu, Yaxiong Chen, Shengwu Xiong 0001 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2025 | Spatial Invariant Hash Based on Self-Attention Mechanism for Remote Sensing Ship Image RetrievalabstractIn the task of remote sensing ship image retrieval, due to the significant Angle changes and multi-scale characteristics of ships in the image, it is difficult to extract advanced features and the efficiency of feature descriptors is low. Therefore, this paper proposes a spatial invariant hash based on self-attention mechanism for remote sensing ship image retrieval (SIHS). The algorithm consists of two core modules: First, a module based on spatial invariance is designed, which uses deep convolutional neural network to extract continuous real-valued descriptors, and introduces the spatial transformation attention mechanism, and enhances the adaptation ability and learning efficiency of the model to the spatial invariance features through self-learning affine transformation and attention calculation; Secondly, a module based on self-attention hashing is proposed, which improves the efficiency of image representation by multi-scale image embedding, optimizes the attention regularization in the visual encoder, and effectively solves the problem of quantization loss in hash mapping. The experimental results show that the retrieval performance of SIHS algorithm on GGWS, DSCR, FGSC-23 and FGSCR-42 datasets is superior to the existing methods based on deep features. Fuwei Huang, Yaxiong Chen, Kai Yan 0001, Yin Ye, Xuehu Liu, Shengwu Xiong 0001 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2025 | Bow Direction Detection Based on Angular Coding With Heading Intersection Over Union LossabstractAccurate bow direction detection is essential for ship trajectory prediction and port monitoring. Existing ship detection networks typically output angles within 180°, while extending to 360° introduces cyclic issues affecting rotation intersection over union (RIoU) accuracy. This study proposes a novel bow direction detection algorithm that extends network output to 360° and integrates a heading intersection over union (HIoU) loss to enhance detection accuracy and robustness. Additionally, an HIoU loss function is designed to improve bow direction identification and reduce quantization errors in hash codes. The algorithm is evaluated on three datasets: FGSD, OHD-SJTU-S, and OHD-SJTU-L. On FGSD, it achieves mean average precision (mAP) of 91.14%. On OHD-SJTU-S, it attains an$\text {mAP}_{50:95}$of 63.3% and a bow direction prediction accuracy of 90.7%. On OHD-SJTU-L, the$\text {mAP}_{50:95}$is 29.2%, with an accuracy of 80.2%. Yaxiong Chen, Qiangqiang Huang, Hao Sun 0014, Shengwu Xiong 0001, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2025 | Bilinear Parallel Fourier Transformer for Multimodal Remote Sensing ClassificationabstractVision Transformers (ViTs) have shown promise in multimodal fusion image classification, yet face performance challenges in complex remote sensing scenarios. Single fusion frameworks often fail to fully utilize multimodal diversity, and the uneven distribution of image categories complicates the accurate construction of spatial structures by Transformers. Additionally, traditional cross-entropy tends to favor majority classes, neglecting minority classes, resulting in suboptimal predictions and reduced overall accuracy (OA). To solve these challenges, we propose a novel deep neural network, a bilinear parallel Fourier Transformer (BPFT). We propose a novel dual-fusion feature interaction (DFFI) module that utilizes two distinct types of fused features for learning, namely the spatial-spectral fusion feature and the global fusion feature. Besides, we introduce a dual-feature interaction (DFI) module to improve the utilization of fused feature information. To enable the Transformer to better establish spatial structural relationships, we employ the Fourier transform in place of the self-attention mechanism. To address the focus on minority class labels, we propose an exponential label smoothing cross-entropy loss function. This loss function comprises two components: exponential cross-entropy and label smoothing. The exponential cross-entropy component applies a strong penalty to misclassified samples, thereby increasing attention on minority class labels. To validate the efficacy of our approach, extensive experiments are conducted across two multimodal remote sensing datasets: Augsburg and Berlin, encompassing hyperspectral imaging (HSI) data and synthetic aperture radar (SAR) data. The results of these experiments affirm the superior performance of our proposed BPFT model compared to existing state-of-the-art models in multimodal remote sensing image classification tasks. Yaxiong Chen, Qicong Wang, Yichen Zhao, Shengwu Xiong 0001, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2025 | Global-Local Fusion With Semantic Information Guidance for Accurate Small Object Detection in UAV Aerial ImagesabstractIn recent years, the rapid development of the unmanned aerial vehicle (UAV) technology has generated a large number of aerial photography images captured by UAV. Consequently, the object detection in UAV aerial images has emerged as a recent research focus. However, due to the flexible flight heights and diverse shooting angles of UAV, two significant challenges have arisen in UAV aerial images: extreme variation in target scale and the presence of numerous small targets. To address these challenges, this article introduces a semantic information-guided fusion module specifically tailored for small targets. This module utilizes high-level semantic information to guide and align the underlying texture information, thereby enhancing the semantic representation of small targets at the feature level and subsequently improving the model’s ability to detect them. In addition, this article introduces a novel global–local fusion detection strategy to strengthen the detection of small targets. We have redesigned the foreground region assembly method to address the drawbacks of previous methods that involved multiple inferences. Extensive experiments conducted on the VisDrone and UAVDT datasets demonstrate that our two self-designed modules can significantly enhance the detection capability of small targets compared with the YOLOX-M model. Our code is publicly available at:https://github.com/LearnYZZ/GLSDet. Yaxiong Chen, Zhengze Ye, Haokai Sun 0001, Tengfei Gong, Shengwu Xiong 0001, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2025 | VGRSS: Datasets and Models for Visual Grounding in Remote Sensing Ship ImagesabstractThis paper introduces a task named Visual Grounding of Remote Sensing Ship Images (VGRSS). The goal of VGRSS is to locate ship objects in remote sensing images guided by natural language. Extensive research has been conducted on multimodal processing of remote sensing images and text to retrieve rich information from remote sensing images using natural language. However, due to the unique characteristics of remote sensing ship images, ship localization using natural language remains a challenge. Therefore, in this work, we construct datasets for the VGRSS task and explore deep learning models. Specifically, our contributions can be summarized as follows: First, we construct two remote sensing ship datasets for visual grounding. One is based on the optical remote sensing dataset, named RSSVG, while the other is based on the synthetic aperture radar (SAR) dataset, named SARVG. Second, we propose a Language-Guided Visual Feature Enhancement (LVFE) module. This module enhances visual features through language guidance before Visual-Linguistic Fusion. Third, we propose a Visual-Linguistic Fusion (VLF) module based on multimodal feature stacking. This module inputs the stacked language and visual features, and then performs feature fusion using a Transformer, enabling effective cross-modal interaction and integration. Fourth, we introduce a novel loss calculation method by incorporating Enhanced Intersection over Union (EIoU) into the loss function. Finally, we benchmark extensive state-of-the-art (SOTA) natural image visual grounding methods on the constructed RSSVG and SARVG datasets, then provide insightful analysis based on the results. This work offers valuable insights for developing better VGRSS models. Yaxiong Chen, Liwen Zhan, Yichen Zhao, Shengwu Xiong 0001, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2025 | A Dual-Stage Wavelet and Linear Attention Enhancement Network for Agricultural Hyperspectral Image ClassificationabstractHyperspectral image (HSI) classification faces unique challenges in agricultural scenario due to spectral-spatial feature similarity caused by complex planting structures and high spectral similarity. Existing spatial-spectral joint feature extraction methods fail to fully exploit the advantages of spatial and spectral information, thus have certain limitations and cannot effectively distinguish similar crops in agricultural scenarios. To address these limitations, we proposed a dual-stage wavelet and linear attention enhancement network (DSW-LAN) for agricultural HSI classification, addressing the challenges of complex spatial-spectral information and high redundancy. We integrates a direction factorized deformable 3D Convolution (DFDWConv3D) module to capture multi-scale spatial-spectral features through adaptive kernel adjustments, while wavelet transform decomposes spatial features into low-frequency (structural) and high-frequency (textural) components for targeted enhancement. Additionally, a spectral probe-guided linear attention mechanism efficiently models long-range spectral dependencies with reduced computational complexity by prioritizing discriminative bands. Experimental results demonstrate superior performance on three challenging agricultural HSI datasets, achieving enhanced classification accuracy with reduced computational complexity. Yaxiong Chen, Bo Zhang 0069, Shili Xiong, Shengwu Xiong 0001, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2025 | Discover the Unknown Ones in Fine-Grained Ship DetectionabstractRemote sensing image-based ship identification technology has great applications in areas such as national defense and fishery management. However, existing remote sensing ship studies mainly focus on a closed environment and overlook actual sea conditions, while new military ships will be encountered. These unknown categories of ships will be ignored or misclassified by existing models, dramatically affecting the accurate assessment of the maritime situation. Furthermore, existing unknown detection methods for natural images fail to tackle the remote sensing ship detection problem for the property of high similarity in overall appearance. To cope with this problem, this paper proposes a fine-grained unknown ship detection network. Firstly, we explore a class-balanced proposal sampler to avoid inefficient information learning. Secondly, we propose a finegrained memory bank-based contrastive learning strategy to separate different categories. Finally, to further separate unknown classes, we adopt an uncertainty-aware unknown learner with logit to reduce the uncertainty of fine-grained predictions. Experiments conducted in three public ship detection datasets ShipRSImageNet, DOSR, and HRSC2016 show that the method not only achieves good detection on unknown class ships, but also improves the detection accuracy on known classes. The code is available at https://github.com/FoRGEU/DUONet. Tengfei Gong, Yaxiong Chen, Shengwu Xiong 0001, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Multibranch Fusion-Based Feature Enhance for Remote-Sensing Scene ClassificationabstractRemote-sensing (RS) scene classification is a fundamental and significant task in RS image interpretation, involving the annotation of semantic content. RS scene images are characterized by complex backgrounds, rich content, and multiscale targets, exhibiting both intraclass separation and interclass convergence. Therefore, extracting features that effectively express the intrinsic attributes of images and possess high discriminative is crucial for RS scene classification. Existing global-based methods often lack the ability to capture significant detailed information in similar scenes. Conversely, methods based on local discriminative features tend to overlook the interrelationships of objects within the same scene. To address these issues, this article proposes a unified framework named MBFNet to align and fuse features of different scales and levels for accurate RS scene classification. We utilize a multibranch feature-extracting network structure with parallel convolution and Transformer modules. Simultaneously, a kernel-selected multiscale aggregation (KSMSA) module is designed to efficiently process the diverse scale features emanating from these parallel branches. By selecting different convolution kernels, a dynamic receptive field is established to adaptively process features of different scales, reducing semantic differences to achieve effective aggregation of multiscale features. Moreover, a learnable multilevel aggregation (LMLA) module is designed to integrate shallow features, such as shape information, into deep features for more comprehensive feature fusion. Benefiting from KSMSA and LMLA, the proposed MBFNet improves the discriminability of features, thereby enhancing classification performance. Comprehensive experiments on three benchmark datasets demonstrate that the proposed method outperforms state-of-the-art RS scene classification methods in terms of performance. Xiongbo Lu, Meng Yang 0034, Yaxiong Chen, Shengwu Xiong 0001, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Class-Aware Consistency Learning for Open-Set Semi-Supervised Hyperspectral Image ClassificationabstractSemi-supervised hyperspectral image (HSI) classification methods focus on exploring the spectral and spatial information of unlabeled samples. However, existing methods generally follow the closed-set setting, assuming that unlabeled samples do not contain novel classes, which is hard to hold in practical applications. This paper aims to study semi-supervised HSI classification in the open-set setting, i.e., unlabeled samples fall into novel classes, and proposes a class-aware consistency learning (CACL) method. First, to explore discriminative spectral-spatial features, a position-aware transformer is developed, which effectively models spatial position priors between the center pixel and its neighboring pixels via a symmetric position-aware encoding. Then, to reduce the interference from novel class samples on the model’s discrimination, a prototype-driven consistency learning is proposed, which accurately selects unlabeled samples belonging to known classes via a known class sampler, and efficiently utilizes their spectral-spatial information by modeling consistent predictions across different views. Finally, to further improve the distinguishability between known classes, a prototype contrastive optimization is proposed to decrease the distance between samples from the same class and increase the distances between those from different classes in the feature domain. Furthermore, an adaptive segmentation threshold is designed to accurately predict known classes and reject novel classes. Extensive experiments verify that our CACL outperforms the state-of-the-art methods, achieving the overall accuracy of 81.65%, 88.61%, and 92.88% on the Indian Pines, Salinas, and Pavia University datasets, with 10 labeled samples in each known class. The code is available at https://github.com/rock-in/CACL-main. Hao Sun 0014, Renyi Chen, Huaxiong Yao, Yaxiong Chen, Wei Xie 0008, Guirong Feng, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | SSPNet: Spatial-Spectral Perception Network for Mineral Hyperspectral Image ClassificationabstractUnlike general scenes, mineral hyperspectral images often exhibit similar spatial and spectral characteristics across different mines, making traditional classification methods less effective due to compromised robustness. To address this, we propose a Spatial-Spectral Perception Network for mineral hyperspectral image classification. This approach divides spatial-spectral feature extraction into two stages. In the spatial feature perception stage, we introduce a Spatial Frequency Perceptron that maps three-dimensional spatial features into low-frequency and high-frequency domains. We then apply Triple-Cross-Attention to each frequency domain to better differentiate spatial features of similar mines. In the spectral perception stage, we design a Spectral Linear Perceptron using Absolute Linear Attention, which captures fine-grained spectral differences by establishing internal relationships between spectral features through Absolute Positional Weighting. This enables effective separation of similar spectra for final classification. Extensive experiments on three publicly available mineral hyperspectral image datasets and one agricultural hyperspectral dataset show that our method outperforms popular alternatives in both effectiveness and robustness. The open-source code can be accessed at https://github.com/WUTCM-Lab/SSPNet. Bo Zhang 0069, Yaxiong Chen, Ruilin Yao, Shili Xiong, Shengwu Xiong 0001, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Hyperspectral Image Classification via Cascaded Spatial Cross-Attention NetworkabstractIn hyperspectral images (HSIs), different land cover (LC) classes have distinct reflective characteristics at various wavelengths. Therefore, relying on only a few bands to distinguish all LC classes often leads to information loss, resulting in poor average accuracy. To address this problem, we propose a method called Cascaded Spatial Cross-Attention Network (CSCANet) for HSI classification. We design a cascaded spatial cross-attention module, which first performs cross-attention on local and global features in the spatial context, then uses a group cascade structure to sequentially propagate important spatial regions within the different channels, and finally obtains joint attention features to improve the robustness of the network. Moreover, we also design a two-branch feature separation structure based on spatial-spectral features to separate different LC Tokens as much as possible, thereby improving the distinguishability of different LC classes. Extensive experiments demonstrate that our method achieves excellent performance in enhancing classification accuracy and robustness. The source code can be obtained from https://github.com/WUTCM-Lab/CSCANet. Bo Zhang 0069, Yaxiong Chen, Shengwu Xiong 0001, Xiaoqiang Lu |
IEEE Trans. Image Process. | 2 |
| 2025 | SSAT++: A Semantic-Aware and Versatile Makeup Transfer Network With Local Color Consistency ConstraintabstractThe purpose of makeup transfer (MT) is to transfer makeup from a reference image to a target face while preserving the target's content. Existing methods have made remarkable progress in generating realistic results but do not perform well in terms of semantic correspondence and color fidelity. In addition, the straightforward extension of processing videos frame by frame tends to produce flickering results in most methods. These limitations restrict the applicability of previous methods in real-world scenarios. To address these issues, we propose a symmetric semantic-aware transfer network (SSAT++) to improve makeup similarity and video temporal consistency. For MT, the feature fusion (FF) module first integrates the content and semantic features of the input images, producing multiscale fusion features. Then, the semantic correspondence from the reference to the target is obtained by measuring the correlation of fusion features at each position. According to semantic correspondence, the symmetric mask semantic transfer (SMST) module aligns the reference makeup features with the target content features to generate MT results. Meanwhile, the semantic correspondence from the target to the reference is obtained by transposing the correlation matrix and applied to the makeup removal task. To enhance color fidelity, we propose a novel local color loss that forces the transferred results to have the same color histogram distribution as the reference. Furthermore, a morphing simulation is designed to ensure temporal consistency for video MT without requiring additional video frame input and optical flow estimation. To evaluate the effectiveness of our SSAT++, extensive experiments have been conducted on the MT dataset which has a variety of makeup styles, and on the MT-Wild dataset which contains images with diverse poses and expressions. The experiments show that SSAT++ outperforms existing MT methods through qualitative and quantitative evaluation and provides more flexible makeup control. Code and trained model will be available at https://gitee.com/sunzhaoyang0304/ssat-msp and https://github.com/Snowfallingplum/SSAT. Zhaoyang Sun, Yaxiong Chen, Shengwu Xiong 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | MakeupDiffuse: a double image-controlled diffusion model for exquisite makeup transfer
Xiongbo Lu, Yaxiong Chen, Shengwu Xiong 0001 |
Vis. Comput. | 4 |
| 2024 | Content-Style Decoupling for Unsupervised Makeup Transfer without Generating Pseudo Ground TruthabstractThe absence of real targets to guide the model training is one of the main problems with the makeup transfer task. Most existing methods tackle this problem by synthesizing pseudo ground truths (PGTs). However, the generated PGTs are often sub-optimal and their imprecision will eventually lead to performance degradation. To alleviate this issue, in this paper, we propose a novel Content-Style Decoupled Makeup Transfer (CSD-MT) method, which works in a purely unsupervised manner and thus eliminates the negative effects of generating PGTs. Specifically, based on the frequency characteristics analysis, we assume that the low-frequency (LF) component of a face image is more associated with its makeup style information, while the high-frequency (HF) component is more related to its content details. This assumption allows CSD-MT to decouple the content and makeup style information in each face image through the frequency decomposition. After that, CSD-MT realizes makeup transfer by maximizing the consistency of these two types of information between the transferred result and input images, respectively. Two newly designed loss functions are also introduced to further improve the transfer performance. Extensive quantitative and qualitative analyses show the effectiveness of our CSD-MT method. Our code is available at https://github.com/Snowfallingplum/CSD-MT. Zhaoyang Sun, Shengwu Xiong 0001, Yaxiong Chen |
CVPR | 3 |
| 2024 | Adaptive Learning via a Negative Selection Strategy for Few-Shot Bioacoustic Event DetectionabstractAlthough the Prototypical Network (ProtoNet) has demonstrated effectiveness in few-shot biological event detection, two persistent issues remain. Firstly, there is difficulty in constructing a representative negative prototype due to the absence of explicitly annotated negative samples. Secondly, the durations of the target biological vocalisations vary across tasks, making it challenging for the model to consistently yield optimal results across all tasks. To address these issues, we propose a novel adaptive learning framework with an adaptive learning loss to guide classifier updates. Additionally, we propose a negative selection strategy to construct a more representative negative prototype for ProtoNet. All experiments ware performed on the DCASE 2023 TASK5 few-shot bioacoustic event detection dataset. The results show that our proposed method achieves an F-measure of 0.703, an improvement of 12.84%. Yaxiong Chen, Xueping Zhang, Yunfei Zi, Shengwu Xiong 0001 |
ICME | 1 |
| 2024 | CausalCLIPSeg: Unlocking CLIP's Potential in Referring Medical Image Segmentation with Causal Intervention
Yaxiong Chen, Minghong Wei, Zixuan Zheng, Jingliang Hu, Yilei Shi, Shengwu Xiong 0001, Xiao Xiang Zhu 0001, Lichao Mou |
MICCAI (3) | 1 |
| 2024 | Striving for Simplicity: Simple Yet Effective Prior-Aware Pseudo-labeling for Semi-supervised Ultrasound Image Segmentation
Yaxiong Chen, Zixuan Zheng, Jingliang Hu, Yilei Shi, Shengwu Xiong 0001, Xiao Xiang Zhu 0001, Lichao Mou |
MICCAI (9) | 1 |
| 2024 | SHMT: Self-supervised Hierarchical Makeup Transfer via Latent Diffusion ModelsabstractThis paper studies the challenging task of makeup transfer, which aims to apply diverse makeup styles precisely and naturally to a given facial image. Due to the absence of paired data, current methods typically synthesize sub-optimal pseudo ground truths to guide the model training, resulting in low makeup fidelity. Additionally, different makeup styles generally have varying effects on the person face, but existing methods struggle to deal with this diversity. To address these issues, we propose a novel Self-supervised Hierarchical Makeup Transfer (SHMT) method via latent diffusion models. Following a "decoupling-and-reconstruction" paradigm, SHMT works in a self-supervised manner, freeing itself from the misguidance of imprecise pseudo-paired data. Furthermore, to accommodate a variety of makeup styles, hierarchical texture details are decomposed via a Laplacian pyramid and selectively introduced to the content representation. Finally, we design a novel Iterative Dual Alignment (IDA) module that dynamically adjusts the injection condition of the diffusion model, allowing the alignment errors caused by the domain gap between content and makeup representations to be corrected. Extensive quantitative and qualitative analyses demonstrate the effectiveness of our method. Our code is available at https://github.com/Snowfallingplum/SHMT. Zhaoyang Sun, Shengwu Xiong 0001, Yaxiong Chen |
NeurIPS | 3 |
| 2024 | A Fine Rendering High-Resolution Makeup Transfer network via inversion-editing strategy
Zhaoyang Sun, Shengwu Xiong 0001, Yaxiong Chen |
Eng. Appl. Artif. Intell. | 3 |
| 2024 | ResCount: A Residual Feature Fusion Network for Ship Counting in Remote Sensing ImagesabstractShip counting is used to count the number of ships in an image. It has a wide range of research backgrounds in areas such as port management and maritime security. In specific areas such as ports, due to their large number of ships, the ships captured by remote sensing images often have problems of uneven distribution and large differences in ship sizes, which will affect the performance of ship counting. To address the above problems, this letter proposes a residual feature fusion network for ship counting (ResCount). The model first uses a feature extraction network to extract the feature map of the image, and then uses a dual-branch structure to further enhance the feature map. One branch uses a visual encoder module to learn the connection between different regions in the image to improve the problem of decreased counting accuracy in scenes with uneven distribution of ships. However, the visual encoder will lose information such as the outline and texture of the ship. Therefore, the other branch uses a regional context feature fusion module (CAF) proposed in this letter to extract local features of different scales and context features of ships to improve the counting accuracy in scenes with large differences in ship size. In addition, this letter proposes a residual feature fusion (RFF) module to enhance the model’s attention to sparse areas and finally regress to obtain a density map. In addition, we conducted a large number of experiments to verify the method. Finally, on the remote sensing object counting dataset (RSOC), the mean absolute error (MAE) index reached 60.08 and the root mean squared error (RMSE) index reached 79.62. Kai Yan 0001, Yaxiong Chen, Shengwu Xiong 0001 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2024 | Scale-Aware Adaptive Refinement and Cross-Interaction for Remote Sensing Audio-Visual Cross-Modal Retrieval
Yaxiong Chen, Chuang Du, Yunfei Zi, Shengwu Xiong 0001, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Multiscale Salient Alignment Learning for Remote-Sensing Image-Text RetrievalabstractRemote-sensing image–text (RSIT) retrieval involves the use of either textual descriptions or remote-sensing images (RSI) as queries to retrieve relevant RSIs or corresponding text descriptions. Many traditional cross-modal RSIT retrieval methods tend to overlook the importance of capturing salient information and establishing the prior similarity between RSIs and texts, leading to a decline in cross-modal retrieval performance. In this article, we address these challenges by introducing a novel approach known as multiscale salient image-guided text alignment (MSITA). This approach is designed to learn salient information by aligning text with images for effective cross-modal RSIT retrieval. The MSITA approach first incorporates a multiscale fusion module and a salient learning module to facilitate the extraction of salient information. In addition, it introduces an image-guided text alignment (IGTA) mechanism that uses image information to guide the alignment of texts, enabling the effective capture of fine-grained correspondences between RSI regions and textual descriptions. In addition to these components, a novel loss function is devised to enhance the similarity across different modalities and reinforce the prior similarity between RSIs and texts. Extensive experiments conducted on four widely adopted RSIT datasets affirm that the MSITA approach significantly enhances cross-modal RSIT retrieval performance in comparison to other state-of-the-art methods. Yaxiong Chen, Shengwu Xiong 0001, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Thread the Needle: Cues-Driven Multiassociation for Remote Sensing Cross-Modal RetrievalabstractRapid advances in Earth observation technologies have yielded numerous remotely sensed images and corresponding text data, enabling cross-modal image–text retrieval to extract valuable clues. However, current methods often focus on learning global semantic information from text and remote sensing (RS) images, while neglecting fine-grained semantic alignment and correlation. In addition, contrastive learning between modalities is often insufficient. To address these issues, we propose an innovative cues-driven multiassociation feature matching network (CDMAN) for cross-modal RS image retrieval. The proposed method primarily involves two key steps: 1) aligning positive samples and enhancing fusion for negative samples based on modal cues. To achieve precise alignment between RS images and text and facilitate the learning process for negative samples in contrastive learning, we have developed a novel fine-grained cues injection module that aligns and guides modalities using fine-grained cues; and 2) establishing multigranularity associative learning. To address the issue of insufficient association between RS images and text, we have implemented multigranularity collaborative associative learning, focusing on general and fine-grained modal associations. By fully leveraging modal cues, our method maintains both detailed associations and overall consistency in global associations. Experiments demonstrate that, compared to baseline methods, this approach achieves more accurate cross-modal retrieval (MCR) by combining fine-grained alignment and multigranularity associations. Yaxiong Chen, Jirui Huang, Zhaoyang Sun, Shengwu Xiong 0001, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Integrating Multisubspace Joint Learning With Multilevel Guidance for Cross-Modal Retrieval of Remote Sensing ImagesabstractIn recent years, with the continuous advancement of remote sensing technology and text processing techniques, there has been a growing abundance of remote sensing images and associated textual data. Combining remote sensing images with their corresponding textual data allows for integrated analysis and retrieval, which holds significant practical implications across multiple application domains, including geographic information systems (GIS), environmental monitoring, and agricultural management. Remote sensing images have the characteristics of multi-targets and multi-scales, and the textual descriptions of these targets are not fully utilized, leading to a decrease in retrieval accuracy. Previous methods have struggled to balance inter-modality information interaction and intra-modality feature fusion, and they have paid little attention to the consistency of distribution within modalities. In light of this, this paper proposes a symmetric multi-level guidance network (SMLGN) for cross-modal retrieval in remote sensing. SMLGN first introduces fusion guidance between local and global within modalities and fine-grained bidirectional guidance between modalities, allowing for the learning of a common semantic space. Furthermore, to address the distribution differences of different modalities within the common semantic space, we design an adversarial joint learning framework and a multi-objective loss function to optimize the SMLGN method and achieve consistency in data distribution. The experimental results demonstrate that the SMLGN method performs well in the task of cross-modal retrieval between remote sensing images and textual data. It effectively integrates the information from both modalities, improving the accuracy and reliability of the retrieval process. Yaxiong Chen, Jirui Huang, Shengwu Xiong 0001, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Integrating Detailed Features and Global Contexts for Semantic Segmentation in Ultrahigh-Resolution Remote Sensing ImagesabstractSemantic segmentation of ultrahigh-resolution (UHR) remote sensing images is a fundamental task for many downstream applications. Achieving precise pixel-level classification is paramount for obtaining exceptional segmentation results. This challenge becomes even more complex due to the need to address intricate segmentation boundaries and accurately delineate small objects within the remote sensing imagery. To meet these demands effectively, it is critical to integrate two crucial components: global contextual information and spatial detail feature information. In response to this imperative, the multilevel context-aware segmentation network (MCSNet) emerges as a promising solution. MCSNet is engineered to not only model the overarching global context but also extract intricate spatial detail features, thereby optimizing segmentation outcomes. The strength of MCSNet lies in its two pivotal modules, the spatial detail feature extraction (SDFE) module and the refined multiscale feature fusion (RMFF) module. Moreover, to further harness the potential of MCSNet, a multitask learning approach is employed. This approach integrates boundary detection and semantic segmentation, ensuring that the network is well-rounded in its segmentation capabilities. The efficacy of MCSNet is rigorously demonstrated through comprehensive experiments conducted on two established international society for photogrammetry and remote sensing (ISPRS) 2-D semantic labeling datasets: Potsdam and Vaihingen. These experiments unequivocally establish MCSNet stands as a pioneering solution, that delivers state-of-the-art performance, as evidenced by its outstanding mean intersection over union (mIoU) and mean$F1$-score (mF1) metrics. The code is available at:https://github.com/WUTCM-Lab/MCSNet. Yaxiong Chen, Shengwu Xiong 0001, Xiaoqiang Lu, Xiao Xiang Zhu 0001, Lichao Mou |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | A Joint Saliency Temporal-Spatial-Spectral Information Network for Hyperspectral Image Change DetectionabstractHyperspectral image change detection (HSI-CD) is a fundamental task in the field of remote sensing (RS) observation, which utilizes the rich spectral and spatial information in bitemporal HSIs to detect subtle changes on the Earth’s surface. However, modern deep learning (DL)-based HSI-CD methods mostly rely on patch-based methods, which leads to spectral band redundancy and spatial information noise in limited receiving domains, thus ignoring the extraction and utilization of saliency information and limiting the improvement of CD performance. To address these issues, this article proposes a joint saliency temporal–spatial–spectral information network (STSS-Net) for HSI-CD. The principal contributions of this article can be summarized: 1) we have designed a spatial saliency information extraction (SSIE) module for denoising based on distance from center pixels and spectral similarity of the substance, which increases the attention to spatial differences between similar spectral substances and different spectral substances; 2) we have designed a compact high-level spectral information tokenizer (CHLSIT) for spectral saliency information, where the high-level conceptual information of changes in spectral interest can be represented by nonlinear combinations of spectral bands, and redundancy can be removed by extracting high-level spectral conceptual features; and 3) utilizing the advantages of CNN and transformer architectures to combine temporal–spatial–spectral information. The experimental results on three real HSI-CD datasets show that STSS-Net can improve the accuracy of CD and has a certain improvement in the detection of edge information and complex information. Yaxiong Chen, Shengwu Xiong 0001, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Learning Starts From Optimizing the Composition of Temporal Information for Hyperspectral Change DetectionabstractHyperspectral image change detection (HSI-CD) is a task that utilizes both spectral and spatial features to more effectively detect changes. The features of HSI captured at different times are often influenced by external factors. The annoying variability is not useful for CD. Current methods extract effective information through complex learning together with it. They did not focus on whether different categories (changed and unchanged) of sample pairs have the same information composition. They lack a fundamental analysis and optimization based on the composition of information pairs to address current issues such as insufficient recognition of changed samples and poor feature fusion. To address these issues, we rethink the information composition of bitemporal sample pairs in different categories for HSI-CD, and design a frequency-domain information exchange and generation network with Siamese U-shaped structure (FDIEG-UNet) for HSI-CD. Our contribution can be summarized as follows: 1) by designing the FDGJLS module for learning and separating global-joint spectral-spatial features, to provide better basic frequency-domain information for subsequent feature-domain optimization; 2) based on the information from FDGJLS, we use the ULDT and SSTDFE modules to remove the influence of temporal-correlated invalid style information in feature fusion. Two modules work together to improve the equality of feature description between changed and unchanged samples in CD; and 3) the experimental results on three real HSI-CD datasets demonstrate the effectiveness of our proposed method in solving the above problems. The code is available athttps://github.com/WUTCM-Lab/FDIEG-UNet. Yaxiong Chen, Jirui Huang, Shengwu Xiong 0001, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Cross-Modal Remote Sensing Image-Audio Retrieval With Adaptive Learning for Aligning CorrelationabstractAn important challenge that existing work has yet to address is the relatively small differences in audio representations compared with the rich content provided by remote sensing (RS) images, making it easy to overlook certain details in the images. This imbalance in information between modalities poses a challenge in maintaining consistent representations. In response to this challenge, we propose a novel cross-modal RS image-audio (RSIA) retrieval method called adaptive learning for aligning correlation (ALAC). ALAC integrates region-level learning into image annotation through a region-enhanced learning attention (RELA) module. By collaboratively suppressing features at different region levels, ALAC is able to provide a more comprehensive visual feature representation. In addition, a novel adaptive knowledge transfer (AKT) strategy has been proposed, which guides the learning process of the frontend network using aligned feature vectors. This approach allows the model to adaptively acquire alignment information during the learning process, thereby facilitating better alignment between the two modalities. Finally, to better use mutual information between different modalities, we introduce a plug-and-play result rerank module. This module optimizes the similarity matrix using retrieval mutual information between modalities as weights, significantly improving retrieval accuracy. Experimental results on four RSIA datasets demonstrate that ALAC outperforms other methods in retrieval performance. Compared with state-of-the-art methods, improvements of 1.49%, 2.25%, 4.24%, and 1.33% were, respectively, achieved by ALAC. The codes are accessible athttps://github.com/huangjh98/ALAC. Yaxiong Chen, Shengwu Xiong 0001, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Visual Contextual Semantic Reasoning for Cross-Modal Drone Image-Text RetrievalabstractThe cross-modal drone image-text (DIT) retrieval task involves using either text or drone images as queries to retrieve relevant drone images or corresponding text. The primary challenge stems from the diverse and intricate nature of drone images, making effective alignment between image and text challenging. In response, we propose an innovative approach called visual contextual semantic reasoning (VCSR), aimed at precisely aligning information across different modalities. VCSR employs textual cues to guide rich semantic reasoning within the visual context, reducing redundancy in visual information. Furthermore, the method captures drone image information relevant to the text, revealing subtle correspondences between drone image regions and textual content. To enhance visual semantic learning, context region learning (CRL) term and consistency semantic alignment (CSA) terms are introduced for stronger guidance, further intensifying the cross-modal interaction between textual and visual data, resulting in more robust feature representation. Extensive experiments conducted on two self-constructed DIT datasets demonstrate that VCSR outperforms alternative methods in terms of DIT retrieval performance. The codes are accessible athttps://github.com/huangjh98/VCSR. Yaxiong Chen, Shengwu Xiong 0001, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Oriented Object Detector With Gaussian Distribution Cost Label Assignment and Task-Decoupled HeadabstractRecently, oriented object detection in remote sensing images has garnered significant attention due to its broad range of applications. Early oriented object detection adhered to the established general object detection frameworks, utilizing the label assignment strategy based on the horizontal bounding box annotations or rotation-agnostic cost function. Such strategy may not reflect the large aspect ratio and rotation of arbitrary-oriented objects in remota sensing images and require high parameter-tuning efforts in training process, which will eventually harm the detector performance. Furthermore, the localization quality of oriented object depends on precise rotation angle prediction, exacerbating the inconsistency between classification and regression tasks in oriented object detection. To address these issues, we propose the Gaussian Distribution Cost Optimal Transport Assignment (GCOTA) and Decoupled Layer Attention Angle Head (DLAAH). Specifically, GCOTA utilize Gaussian distribution based cost function for the optimal transport label assignment in training process, alleviating the impact of rotation angle and large aspect ratio in remote sensing images. DLAAH predicts rotation angle independently and incorporates layer attention to obtain the task-specific features based on the shared FPN features, enhancing the angle prediction and improving consistency across different tasks. Based on these proposed components, we present an anchor-free oriented detector, namely Gaussian Distribution and Task-Decoupled head oriented Detector(GTDet) and a a multi-class ship detection dataset in real scenarios (CGWX), which provides a benchmark for fine-grained object recognition in remote sensing images. Comprehensive experiments are conducted on CGWX and several public challenging datasets, including DOTAv1.0, HRSC2016, to demonstrate that our method achieves superior performance on oriented object detection task. The code is available at https://github.com/WUTCM-Lab/GTDet. Qiangqiang Huang, Ruilin Yao, Xiaoqiang Lu, Jishuai Zhu, Shengwu Xiong 0001, Yaxiong Chen |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | SANet: A Self-Attention Network for Agricultural Hyperspectral Image ClassificationabstractUnlike conventional hyperspectral image (HSI) classification in general scenes, agricultural HSI classification poses greater challenges due to the increased occurrence of “same spectrum different object” and “different spectrum same object” phenomena caused by class similarities. Furthermore, the dense spatial distribution of land cover categories in agricultural scenes and the mixing of spatial–spectral features at crop boundaries add to the complexity of agricultural HSIs. To tackle these issues, we propose SANet, a network designed to enhance crop classification. SANet integrates spectral and contextual information while emphasizing self-correlation within the HSIs. It combines the spatial–spectral nonlocal block structure and the multiscale spectral self-attention (SSA) structure, allocating more attention resources to spatial and spectral dimensions and modeling the existing correlations within the spectral–spatial domain. Additionally, we introduce a two-branch spatial–spectral semantic extraction and fusion structure that can adaptively learn results from both branches. Experimental results demonstrate the promising performance of SANet in agricultural HSI classification by effectively utilizing spectral data, contextual information, and self-attention mechanisms. Bo Zhang 0069, Yaxiong Chen, Shengwu Xiong 0001, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Global-Group Attention Network With Focal Attention Loss for Aerial Scene ClassificationabstractAerial scene classification, aiming at assigning a specific semantic class to each aerial image, is a fundamental task in the remote sensing community. Aerial scene images have more diverse and complex geological features. While some statistics of images can be well fit using convolution, it limits such models to capturing the global context hidden in aerial scenes. Furthermore, to optimize the feature space, many methods add class information to the feature embedding space. However, they seldom combine model structure with class information to obtain more separable feature representations. In this article, we propose to address these limitations in a unified framework (i.e., CGFNet) from two aspects: focusing on the key information of input images and optimizing the feature space. Specifically, we propose a global-group attention module (GGAM) to adaptively learn and selectively focus on important information from input images. GGAM consists of two parallel branches: the adaptive global attention branch (AGAB) and the region-aware attention branch (RAAB). AGAB utilizes an adaptive pooling operation to better model the global context in aerial scenes. As a supplement to AGAB, RAAB combines grouping features with spatial attention to spatially enhance the semantic distribution of features (i.e., selectively focus on effective regions of features and ignore irrelevant semantic regions). In parallel, a focal attention loss (FA-Loss) is exploited to introduce class information into attention vector space, which can improve intraclass consistency and interclass separability. Experimental results on four publicly available and challenging datasets demonstrate the effectiveness of our method. The source code will be released at:https://github.com/zoecheno/CGFNet. Yichen Zhao, Yaxiong Chen, Shengwu Xiong 0001, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Co-Enhanced Global-Part Integration for Remote-Sensing Scene ClassificationabstractRemote sensing (RS) scene classification aims to classify remote sensing images with similar scene characteristics into one category. Plenty of RS images are complex in background, rich in content, and multi-scale in target, exhibiting the characteristics of both intra-class separation and inter-class convergence. Therefore, discriminative feature representations designed to highlight the differences between classes are the key to RS scene classification. Existing methods represent scene images by extracting either global context or discriminative part features from RS images. However, global-based methods often lack salient details in similar RS scenes, while part-based methods tend to ignore the relationships between local ground objects, thus weakening the discriminative feature representation. In this paper, we propose to combine global context and part-level discriminative features within a unified framework called CGINet for accurate RS scene classification. To be specific, we develop a light context-aware attention block (LCAB) to explicitly model the global context to obtain larger receptive fields and contextual information. A co-enhanced loss module (CELM) is also devised to encourage the model to actively locate discriminative parts for feature enhancement. In particular, CELM is only used during training and not activated during inference, which introduces less computational cost. Benefiting from LCAB and CELM, our proposed CGINet improves the discriminability of features, thereby improving classification performance. Comprehensive experiments over four benchmark datasets show that the proposed method achieves consistent performance gains over state-of-the-art RS scene classification methods. Yichen Zhao, Yaxiong Chen, Shengwu Xiong 0001, Xiaoqiang Lu, Xiao Xiang Zhu 0001, Lichao Mou |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | ESPT: A Self-Supervised Episodic Spatial Pretext Task for Improving Few-Shot LearningabstractSelf-supervised learning (SSL) techniques have recently been integrated into the few-shot learning (FSL) framework and have shown promising results in improving the few-shot image classification performance. However, existing SSL approaches used in FSL typically seek the supervision signals from the global embedding of every single image. Therefore, during the episodic training of FSL, these methods cannot capture and fully utilize the local visual information in image samples and the data structure information of the whole episode, which are beneficial to FSL. To this end, we propose to augment the few-shot learning objective with a novel self-supervised Episodic Spatial Pretext Task (ESPT). Specifically, for each few-shot episode, we generate its corresponding transformed episode by applying a random geometric transformation to all the images in it. Based on these, our ESPT objective is defined as maximizing the local spatial relationship consistency between the original episode and the transformed one. With this definition, the ESPT-augmented FSL objective promotes learning more transferable feature representations that capture the local spatial features of different images and their inter-relational structural information in each input episode, thus enabling the model to generalize better to new categories with only a few samples. Extensive experiments indicate that our ESPT method achieves new state-of-the-art performance for few-shot image classification on three mainstay benchmark datasets. The source code will be available at: https://github.com/Whut-YiRong/ESPT. Xiongbo Lu, Zhaoyang Sun, Yaxiong Chen, Shengwu Xiong 0001 |
AAAI | 4 |
| 2023 | Information bottleneck disentanglement based sparse representation for fair classification
Xiongbo Lu, Yaxiong Chen, Shengwu Xiong 0001 |
Pattern Recognit. Lett. | 3 |
| 2023 | Aerial image recognition in discriminative bi-transformer
Yichen Zhao, Yaxiong Chen, Xiongbo Lu, Lei Zhou 0008, Shengwu Xiong 0001 |
Signal Process. | 2 |
| 2023 | Deep Saliency Smoothing Hashing for Drone Image RetrievalabstractDeep hashing algorithms are widely exploited in retrieval tasks due to its low storage and retrieval efficiency. Most of which focus on global feature learning, whilst neglecting local fine-grained features and saliency information for drone images. In this paper, we tackle these dilemmas with a novelDeep Saliency Smoothing Hashing(DSSH) algorithm, which can leverage saliency capture mechanism, distribution smoothing term, global features and local fine-grained features to learn effective hash codes for drone image retrieval. The DSSH algorithm first designs information extraction module to capture global features and local fine-grained features for drone images. Meanwhile, a saliency capture module is proposed to perform information interaction attention and visual enhancement attention, which can capture the saliency area of drone images effectively. On top of the two paths, a novel objective function is designed to preserve the similarity of hash codes, smooth the distribution of drone image datasets and reduce the quantization errors between hash codes and hash-like codes concurrently. Extensive experiments on the Drone Action Dataset and ERA Drone Dataset demonstrate that the DSSH algorithm can further improve the retrieval performance compared to other deep hashing algorithms. Yaxiong Chen, Lichao Mou, Pu Jin, Shengwu Xiong 0001, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Fine Aligned Discriminative Hashing for Remote Sensing Image-Audio RetrievalabstractFor cross-modal remote sensing image-audio retrieval task, hashing technology has attracted much attention in recent works. Most of them focus on mappingRemote Sensing(RS) images and audios into a Hamming space, whilst neglecting discriminative information of RS images and fine alignment for RS images and audios. In this paper, we tackle these dilemmas with a novelFine Aligned Discriminative Hashing(FADH) approach, which can learn hash codes to capture discriminative information of RS images and learn the corresponding detailed information between RS images and audios simultaneously. We first develop a new discriminative information learning module to learn discriminative information of RS images. Meanwhile, a fine alignment module is proposed to unearth the fine correspondence for RS image regions and audios, which can effectively improve the retrieval performance. On top of the two paths, we design a new objective function, which can maintain the similarity of hash codes, preserve the semantic information of RS image features and audio features and eliminate cross-modal differences. The reliability and significance of the designed framework are effectively demonstrated by diverse experiments on three remote sensing image-audio datasets. Yaxiong Chen, Shengwu Xiong 0001, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Self-Supervision Interactive Alignment for Remote Sensing Image-Audio RetrievalabstractCross-modal remote sensing image-audio retrieval aims to use audio or remote sensing images as queries to retrieve relevant remote sensing images or corresponding audios. Although many approaches leverage labeled samples to achieve good performance, the performance cost of labeled samples is high, because cross-modal remote sensing labeled samples usually requires huge labor resources. Therefore, unsupervised cross-modal learning is very important in real-world applications. In this paper, we propose a novel unsupervised cross-modal remote sensing image-audio retrieval approach, namedSelf-Supervision Interactive Alignment(SSIA), which can take advantage of large amounts of unlabeled samples to learn the salient information, cross-modal alignment and the similarity between remote sensing images and audios. Since self-supervised learning lacks the supervision of label information, we leverage the similarity between the input remote sensing image information and audio information as the supervision information. Besides, to perform cross-modal alignment, a novel interactive alignment module is designed to explore fine correspondence relation for remote sensing images and audios. Moreover, we design an audio guided image de-redundant module to reduce the redundant information of visual information, which can capture salient information of remote sensing images. Extensive experiments on four widely-used remote sensing image-audio datasets testify that the SSIA perform gain better remote sensing image-audio retrieval performance than other compared approaches. Yaxiong Chen, Shengwu Xiong 0001, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | MATNet: A Combining Multi-Attention and Transformer Network for Hyperspectral Image ClassificationabstractHyperspectral image (HSI) has rich spatial-spectral information, high spectral correlation and large redundancy between information. Due to the sparse background distribution of HSI, existing methods generally perform poorly for the classification of class pixels located in the boundary areas of land cover categories. This is largely because the network is vulnerable to surrounding redundant information during the training stage, leading to inaccurate feature extraction and thus poor generalization ability of the model. Based on previous work, we propose a HSI classification network called MATNet which combines multi-attention and Transformer. The network first uses spatial attention and channel attention to pay more attention to the more significant information parts, then uses tokenizer module to make a semantic level representation of different categories of ground objects, and then performs deep semantic feature extraction using the transformer encoder module. Finally, we design a loss function called Lpoly, which adds a polynomial to the label smoothing loss to tune the original first polynomial to accommodate different datasets and tasks. We perform experiments in several well-known HSI datasets as well as for visualization. The results show that our proposed MATNet performs well in extracting spatial-spectral features of HSIs as well as understanding semantic degrees of semantic degrees. Bo Zhang 0069, Yaxiong Chen, Shengwu Xiong 0001, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | SSAT: A Symmetric Semantic-Aware Transformer Network for Makeup Transfer and RemovalabstractMakeup transfer is not only to extract the makeup style of the reference image, but also to render the makeup style to the semantic corresponding position of the target image. However, most existing methods focus on the former and ignore the latter, resulting in a failure to achieve desired results. To solve the above problems, we propose a unified Symmetric Semantic-Aware Transformer (SSAT) network, which incorporates semantic correspondence learning to realize makeup transfer and removal simultaneously. In SSAT, a novel Symmetric Semantic Corresponding Feature Transfer (SSCFT) module and a weakly supervised semantic loss are proposed to model and facilitate the establishment of accurate semantic correspondence. In the generation process, the extracted makeup features are spatially distorted by SSCFT to achieve semantic alignment with the target image, then the distorted makeup features are combined with unmodified makeup irrelevant features to produce the final result. Experiments show that our method obtains more visually accurate makeup transfer results, and user study in comparison with other state-of-the-art makeup transfer methods reflects the superiority of our method. Besides, we verify the robustness of the proposed method in the difference of expression and pose, object occlusion scenes, and extend it to video makeup transfer. Zhaoyang Sun, Yaxiong Chen, Shengwu Xiong 0001 |
AAAI | 2 |
| 2022 | Feature Space Disentangling Based on Spatial Attention for Makeup TransferabstractMakeup transfer aims at rendering the makeup style from a given reference image to a source image. Most existing works have achieved promising progress by disentangled representation. However, these methods do not consider the spatial distribution of makeup style, which inevitably change the makeup-irrelevant regions. To solve the problem, we introduce a novel feature space disentangling framework based on spatial attention mechanism for makeup transfer. In particular, we first utilize a single encoder to extract all the features of the image. Then we propose a learnable spatial semantic classifier to classify the extracted features into makeup-specific and makeup-irrelevant features. Finally, we complete makeup transfer by swapping the classified features. Experiments demonstrate that the makeup-specific features precisely signify the spatial distribution of makeup style. The superiority of our approach is well demonstrated by the experiment that it produces promising visual results and keeps those makeup-irrelevant regions unchanged. Jinli Zhou, Yaxiong Chen, Zhaoyang Sun, Chang Zhan, Shengwu Xiong 0001 |
ICIP | 2 |
| 2022 | Fusing Acoustic and Text Emotional Features for Expressive Speech SynthesisabstractProminent methods based on Tacotron2 and advanced models have improved the quality of synthesized speech. However, most data-driven Text- To-Speech (TTS) synthesis methods only aim to achieve reasonable neutral prosody, so the synthesized speech is less expressive. In this paper, a method was proposed which fuses acoustic and text emotional features to produce more vivid and realistic speech. Specifically, to obtain acoustic features, two acoustic encoders are leveraged to extract utterance-level and phoneme-level vectors from the target speech, respectively. To obtain the objective sentiment features of the text, the sentiment analysis model is exploited to extract the sentiment vector from the text and expand it. The expanded vector is feature- fused with the output vector of the acoustic model. The experimental results on the LJSpeech dataset show that the naturalness and expressiveness of the MOS score are 3.63 and 3.45, respectively, and the similarity of the SMOS score is 4.14. Pengfei Duan 0005, Yunfei Zi, Yaxiong Chen, Shengwu Xiong 0001 |
ICME | 4 |
| 2022 | Deep Semantic Ranking Hashing Based on Self-Attention for Medical Image RetrievalabstractWith the rapid progress of medical image technology, medical image retrieval has attracted wide attention in medical data processing fields. Deep hashing methods have been proven effective for massive medical image retrieval. However, existing medical image retrieval methods ignore lesion context and category-level semantics, so it is difficult to correctly correspond to the context information and category of the lesion, resulting in poor performance. This paper addresses this dilemma with a novel Deep Semantic Ranking Hashing Based on Self-Attention (DSHA) approach. We first divide the medical triplet into smaller patches and send them to the multi-head self-attention module, which can more effectively encode the context information of the lesion and the interaction among patches. Meanwhile, the weight-sharing triplet networks are used to learn hash codes, which are forced to be semantically aligned. The proposed semantic enhancement loss effectively enhances the category-level semantics of hash codes. In addition, we added a semantic ranking penalty loss to optimize the retrieval accuracy. Extensive experiments on diverse medical image datasets prove that our DSHA method achieves remarkable results compared with the state-of-the-art medical image retrieval methods. Yibo Tang, Yaxiong Chen, Shengwu Xiong 0001 |
ICPR | 2 |
| 2022 | Mutual information maximizing GAN inversion for real face with identity preservation
Chengde Lin, Shengwu Xiong 0001, Yaxiong Chen |
J. Vis. Commun. Image Represent. | 3 |
| 2022 | Deep Quadruple-Based Hashing for Remote Sensing Image-Sound RetrievalabstractWith the rapid progress of earth observation technology, cross-modal remote sensing (RS) image-sound retrieval has attracted much attention from the field of RS data processing. Existing approaches usually learn the pairwise similarity relations between RS images and sounds. However, these approaches ignore relative semantic similarity relationships, which leads to poor performance of cross-modal RS image-sound retrieval. In this article, we address this dilemma with a noveldeep quadruple-based hashing(DQH) approach. We first devise a novel quadruple-based hashing network to learn relative semantic similarity relationships of hash codes. Meanwhile, we propose a quadruple construction hard module, which randomly selects two triplet hard units to directly learn relative semantic similarity relationships. On top of the two paths, we develop a new objective function to perform effective hash codes learning. The new objective function not only captures the relative semantic correlation of hash codes across different modalities and learns the relative semantic correlation of deep features but also enhances category-level semantics of hash codes and reduces the quantization error between hash-like codes and hash codes. The reasonableness and effectiveness of the proposed architecture are well illustrated by comprehensive experiments on diverse RS image-sound datasets. Yaxiong Chen, Shengwu Xiong 0001, Lichao Mou, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2021 | A bi-level distribution mixture framework for unsupervised driving performance evaluation from naturalistic truck driving data
Lin Lu 0002, Shengwu Xiong 0001, Yaxiong Chen |
Eng. Appl. Artif. Intell. | 3 |
| 2021 | Deep Category-Level and Regularized Hashing With Global Semantic Similarity LearningabstractThe hashing technique has been extensively used in large-scale image retrieval applications due to its low storage and fast computing speed. Most existing deep hashing approaches cannot fully consider the global semantic similarity and category-level semantic information, which result in the insufficient utilization of the global semantic similarity for hash codes learning and the semantic information loss of hash codes. To tackle these issues, we propose a novel deep hashing approach with triplet labels, namely, deep category-level and regularized hashing (DCRH), to leverage the global semantic similarity of deep feature and category-level semantic information to enhance the semantic similarity of hash codes. There are four contributions in this article. First, we design a novel global semantic similarity constraint about the deep feature to make the anchor deep feature more similar to the positive deep feature than to the negative deep feature. Second, we leverage label information to enhance category-level semantics of hash codes for hash codes learning. Third, we develop a new triplet construction module to select good image triplets for effective hash functions learning. Finally, we propose a new triplet regularized loss (Reg-L) term, which can force binary-like codes to approximate binary codes and eventually minimize the information loss between binary-like codes and binary codes. Extensive experimental results in three image retrieval benchmark datasets show that the proposed DCRH approach achieves superior performance over other state-of-the-art hashing approaches. Yaxiong Chen, Xiaoqiang Lu |
IEEE Trans. Cybern. | 1 |
| 2020 | Deep discrete hashing with pairwise correlation learning
Yaxiong Chen, Xiaoqiang Lu |
Neurocomputing | 1 |
| 2020 | Supervised deep hashing with a joint deep network
Yaxiong Chen, Xiaoqiang Lu, Xuelong Li 0001 |
Pattern Recognit. | 1 |
| 2020 | Deep Cross-Modal Image-Voice Retrieval in Remote SensingabstractWith the rapid progress of satellite and aircraft technologies, cross-modal remote sensing image-voice retrieval has been studied in geography recently. However, there still exist some bottlenecks: how to consider the characteristics of remote sensing data adequately and how to reduce the memory and improve the retrieval efficiency in large-scale remote sensing data. In this article, we propose a novel deep cross-modal remote sensing image-voice retrieval approach, namely, deep image-voice retrieval (DIVR), to capture more information of remote sensing data to generate hash codes with low memory and fast retrieval properties. Especially, the DIVR approach proposes inception dilated convolution module to capture multiscale contextual information of remote sensing images and voices. Moreover, in order to enhance cross-modal similarity, the deep features' similarity term is designed to make paired similar deep features as close as possible and paired dissimilar deep features as mutually far as possible. In addition, the quantization error term is designed to drive hash-like codes to approximate hash codes, which can effectively reduce the quantization error for hash codes' learning. Extensive experimental results on three remote sensing image-voice data sets show that the proposed DIVR approach can outperform other cross-modal retrieval approaches. Yaxiong Chen, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2020 | Discrete Deep Hashing With Ranking Optimization for Image RetrievalabstractFor large-scale image retrieval task, a hashing technique has attracted extensive attention due to its efficient computing and applying. By using the hashing technique in image retrieval, it is crucial to generate discrete hash codes and preserve the neighborhood ranking information simultaneously. However, both related steps are treated independently in most of the existing deep hashing methods, which lead to the loss of key category-level information in the discretization process and the decrease in discriminative ranking relationship. In order to generate discrete hash codes with notable discriminative information, we integrate the discretization process and the ranking process into one architecture. Motivated by this idea, a novel ranking optimization discrete hashing (RODH) method is proposed, which directly generates discrete hash codes (e.g., +1/-1) from raw images by balancing the effective category-level information of discretization and the discrimination of ranking information. The proposed method integrates convolutional neural network, discrete hash function learning, and ranking function optimizing into a unified framework. Meanwhile, a novel loss function based on label information and mean average precision (MAP) is proposed to preserve the label consistency and optimize the ranking information of hash codes simultaneously. Experimental results on four benchmark data sets demonstrate that RODH can achieve superior performance over the state-of-the-art hashing methods. Xiaoqiang Lu, Yaxiong Chen, Xuelong Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2020 | Siamese Dilated Inception Hashing With Intra-Group Correlation Enhancement for Image RetrievalabstractFor large-scale image retrieval, hashing has been extensively explored in approximate nearest neighbor search methods due to its low storage and high computational efficiency. With the development of deep learning, deep hashing methods have made great progress in image retrieval. Most existing deep hashing methods cannot fully consider the intra-group correlation of hash codes, which leads to the correlation decrease problem of similar hash codes and ultimately affects the retrieval results. In this article, we propose an end-to-end siamese dilated inception hashing (SDIH) method that takes full advantage of multi-scale contextual information and category-level semantics to enhance the intra-group correlation of hash codes for hash codes learning. First, a novel siamese inception dilated network architecture is presented to generate hash codes with the intra-group correlation enhancement by exploiting multi-scale contextual information and category-level semantics simultaneously. Second, we propose a new regularized term, which can force the continuous values to approximate discrete values in hash codes learning and eventually reduces the discrepancy between the Hamming distance and the Euclidean distance. Finally, experimental results in five public data sets demonstrate that SDIH can outperform other state-of-the-art hashing algorithms. Xiaoqiang Lu, Yaxiong Chen, Xuelong Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2019 | Deep Voice-Visual Cross-Modal Retrieval with Deep Feature Similarity Learning
Yaxiong Chen, Xiaoqiang Lu, Yachuang Feng |
PRCV (3) | 1 |
| 2018 | Hierarchical Recurrent Neural Hashing for Image Retrieval With Hierarchical Convolutional FeaturesabstractHashing has been an important and effective technology in image retrieval due to its computational efficiency and fast search speed. The traditional hashing methods usually learn hash functions to obtain binary codes by exploiting hand-crafted features, which cannot optimally represent the information of the sample. Recently, deep learning methods can achieve better performance, since deep learning architectures can learn more effective image representation features. However, these methods only use semantic features to generate hash codes by shallow projection but ignore texture details. In this paper, we proposed a novel hashing method, namely hierarchical recurrent neural hashing (HRNH), to exploit hierarchical recurrent neural network to generate effective hash codes. There are three contributions of this paper. First, a deep hashing method is proposed to extensively exploit both spatial details and semantic information, in which, we leverage hierarchical convolutional features to construct image pyramid representation. Second, our proposed deep network can exploit directly convolutional feature maps as input to preserve the spatial structure of convolutional feature maps. Finally, we propose a new loss function that considers the quantization error of binarizing the continuous embeddings into the discrete binary codes, and simultaneously maintains the semantic similarity and balanceable property of hash codes. Experimental results on four widely used data sets demonstrate that the proposed HRNH can achieve superior performance over other state-of-the-art hashing methods.Hashing has been an important and effective technology in image retrieval due to its computational efficiency and fast search speed. The traditional hashing methods usually learn hash functions to obtain binary codes by exploiting hand-crafted features, which cannot optimally represent the information of the sample. Recently, deep learning methods can achieve better performance, since deep learning architectures can learn more effective image representation features. However, these methods only use semantic features to generate hash codes by shallow projection but ignore texture details. In this paper, we proposed a novel hashing method, namely hierarchical recurrent neural hashing (HRNH), to exploit hierarchical recurrent neural network to generate effective hash codes. There are three contributions of this paper. First, a deep hashing method is proposed to extensively exploit both spatial details and semantic information, in which, we leverage hierarchical convolutional features to construct image pyramid representation. Second, our proposed deep network can exploit directly convolutional feature maps as input to preserve the spatial structure of convolutional feature maps. Finally, we propose a new loss function that considers the quantization error of binarizing the continuous embeddings into the discrete binary codes, and simultaneously maintains the semantic similarity and balanceable property of hash codes. Experimental results on four widely used data sets demonstrate that the proposed HRNH can achieve superior performance over other state-of-the-art hashing methods. Xiaoqiang Lu, Yaxiong Chen, Xuelong Li 0001 |
IEEE Trans. Image Process. | 2 |