VLDB 2026 Research / reviewers in the wild / expert
Shuli Cheng
dblp:243/9226
· DBLP profile ↗
40ranked-venue papers
6as first author
39since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 1 first-author · 22 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 2 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 6 since 2021Computer networks · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Adaptive structural-semantic fusion and optimization for hashing image retrieval
Shuli Cheng, Anyu Du, Tingjie Liu |
Expert Syst. Appl. | 2 |
| 2026 | Confidence aware Mamba interaction hashing for cross-modal retrieval
Jiapeng Tian, Shuli Cheng, Anyu Du |
Expert Syst. Appl. | 2 |
| 2026 | Multiple Local-Guided Modality-Shared Feature Learning Transformer for Visible-Infrared Person Re-IdentificationabstractVisible-Infrared Person Re-Identification (VI-ReID) aims to identify the same person in heterogeneous images captured by different spectral cameras. Due to the different imaging mechanisms of heterogeneous cameras, there are significant modality differences between cross-modal images. Directly extracting modality-shared features from heterogeneous images can make it difficult to align local discriminative information between different modalities, leading to local semantic inconsistency. To address this, we propose a novel Multiple Local-Guided Modality-Shared Feature Learning Transformer (MLMT) model to mine modality-shared local semantic similarity information, thereby facilitating the effective alignment of cross-modal local semantics. Firstly, we design a parameter-free Random Modality Gap Bridging (RMGB) module, which employs a low-complexity strategy involving random local region grayscale transformation and inverse grayscale transformation. This approach reduces the model’s sensitivity to modality changes from a cross-modal local perspective, thus smoothly and efficiently bridging the inherent differences between the two modalities. Secondly, to better learn the shared local discriminative information, we design a Multiple Local-Guided Feature Learning (ML) strategy. This strategy leverages learnable global and local tokens and employs orthogonal constraint learning to focus on distinct discriminative information in overlapping local regions, thereby exploring shared local similarities between different modalities. Furthermore, we introduce a Similarity Inference Reinforcement (SIR) module, which utilizes high similarity information between person images to optimize the distance matrix, thereby improving matching performance. Extensive experiments on two benchmark datasets for the VI-ReID task demonstrate the effectiveness of the MLMT method. Shuli Cheng |
IEEE Internet Things J. | 3 |
| 2026 | Spatial-frequency cross-shift learning perceptive transformer features in decoupled hashing
Shuli Cheng |
Pattern Recognit. | 2 |
| 2025 | PromptHash: Affinity-Prompted Collaborative Cross-Modal Learning for Adaptive Hashing RetrievalabstractCross-modal hashing is a promising approach for efficient data retrieval and storage optimization. However, contemporary methods exhibit significant limitations in semantic preservation, contextual integrity, and information redundancy, which constrains retrieval efficacy. We present PromptHash, an innovative framework leveraging affinity prompt-aware collaborative learning for adaptive cross-modal hashing. We propose an end-to-end framework for affinity-prompted collaborative hashing, with the following fundamental technical contributions: (i) a text affinity prompt learning mechanism that preserves contextual information while maintaining parameter efficiency, (ii) an adaptive gated selection fusion architecture that synthesizes State Space Model with Transformer network for precise cross-modal feature integration, and (iii) a prompt affinity alignment strategy that bridges modal heterogeneity through hierarchical contrastive learning. To the best of our knowledge, this study presents the first investigation into affinity prompt awareness within collaborative cross-modal adaptive hash learning, establishing a paradigm for enhanced semantic consistency across modalities. Through comprehensive evaluation on three benchmark multi-label datasets, PromptHash demonstrates substantial performance improvements over existing approaches. Notably, on the NUS-WIDE dataset, our method achieves significant gains of 18.22% and 18.65% in image-to-text and text-to-image retrieval tasks, respectively. The code is publicly available at https://github.com/ShiShuMo/PromptHash. Qiang Zou 0002, Shuli Cheng |
CVPR | 2 |
| 2025 | Entity-Level Alignment with Prompt-Guided Adapter for Remote Sensing Image-Text Retrieval
Shuoshuo Li, Shuli Cheng |
ACM Multimedia | 2 |
| 2025 | Covariance Attention Guidance Mamba Hashing for cross-modal retrieval
Shuli Cheng, Anyu Du, Qiang Zou 0002 |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | Spatial and learnable frequency dynamic collaborative visual perception network for remote sensing images semantic segmentationabstractRemote sensing semantic segmentation , as a current research hotspot, is widely applied in scenarios such as agricultural planning and urban construction planning . Deep learning models based on spatial and frequency domain co-perception have rapidly developed due to their comprehensive perception capabilities. However, current networks primarily apply frequency domain perception within various attention mechanisms , lacking complementary learning with the spatial domain. In addition, whether the encoder–decoder pairing in existing U-shaped architectures can fully exploit the advantages of dual-domain perception requires further discussion. To address these issues, this paper proposes the Spatial and Learnable Frequency Dynamic Collaborative Visual Perception Network (SFCVPNet). The encoder uses a cascade of Spatial and Frequency Domain-Aware Transformer (SFFormer) blocks and Convolutional Neural Networks (CNNs) blocks to perceive features, while the decoder employs simple skip connections, forming a novel integration method of CNNs and Transformer blocks in the U-shaped architecture. SFFormer includes the Spatial and Frequency Domain Collaborative Perception Module (SFCPM) and the Spatial and Frequency Guided Multi-Layer Perceptron (SF-MLP). Both SFCPM and SF-MLP utilize a dual-domain perception approach to extract and integrate features. Additionally, we adopt learnable operations for frequency domain feature perception, ensuring that frequency domain features meet the learning requirements of the network model. On three public datasets: ISPRS Vaihingen, ISPRS Potsdam, and LoveDA, respectively, we achieved mIoU scores of 85.90%, 88.19%, and 55.6%. The code will be released at https://github.com/cslxju/SFCVPNet . Shuli Cheng, Anyu Du |
Expert Syst. Appl. | 1 |
| 2025 | Content-adaptive perception mixer visual transformer for hashing image retrieval
Tingjie Liu, Shuli Cheng, Anyu Du |
Expert Syst. Appl. | 2 |
| 2025 | Label-guided diversified learning model for occluded person re-identification
Shuli Cheng |
Expert Syst. Appl. | 2 |
| 2025 | Embedded Separate Deep Localization Feature Information Vision Transformer for Hash Image Retrieval
Shuli Cheng |
Expert Syst. Appl. | 2 |
| 2025 | Occlusion Simulation and Token-Constrained Feature Coupling Network for Occluded Person ReidentificationabstractOccluded person reidentification (Re-ID) aims to learn the features of pedestrians with different identities under occlusion, and is a pivotal technology for intelligent security surveillance systems for the Internet of Things (IoT). In real scenarios, occlusion may occur at any spatial location and exhibits irregularity, occluded people Re-ID remains a challenging task. Most existing methods distinguish visible body parts by leveraging location-based segmentation or external cues, but they are often inefficient or increase network complexity. To address these problems, we propose an occlusion simulation and token-constrained feature coupling network (OSTCNet). Specifically, the occlusion simulation based on block mixing (OSBBM) strategy segments pedestrian images with different labels into blocks during the training phase and proportionally blends the image blocks based on coordinates, thereby generating occluded samples with diverse occlusion patterns. Additionally, we propose the local–global feature coupling (LGFC) module to perform multiscale coupling of local features and global representations, enhancing classification accuracy and capturing more comprehensive feature information. Given the lack of explicit constraints in ViT to differentiate the similarity between class token and patch tokens, which can cause accumulation of similar information in the feature embedding space, we introduce the token orthogonal embedding (TOE) module to enforce constraints on class token, ensuring representational differences from patch tokens. Experimental results on occluded, partial, and holistic ReID datasets validate the effectiveness of our method. Specifically, on the Occluded-DukeMTMC dataset, OSTCNet achieves Rank-1 accuracy of 75.2% and mean-average precision (mAP) score of 65.7%. Shuli Cheng, Anyu Du |
IEEE Internet Things J. | 2 |
| 2025 | Frequency Decoupling Enhancement and Mamba Depth Extraction-Based Feature Fusion in Transformer Hashing Image Retrieval
Shuli Cheng, Qiang Zou 0002 |
Knowl. Based Syst. | 2 |
| 2025 | Dual-stream feature extraction and semantic similarity for deep hashing image retrieval
Shuli Cheng, Anyu Du, Tingjie Liu |
Knowl. Based Syst. | 2 |
| 2024 | Enhanced transformer encoder and hybrid cascaded upsampler for medical image segmentation
Shuli Cheng |
Expert Syst. Appl. | 3 |
| 2024 | Multi-view similarity aggregation and multi-level gap optimization for unsupervised person re-identification
Shuli Cheng, Anyu Du |
Expert Syst. Appl. | 2 |
| 2024 | Pooling-based Visual Transformer with low complexity attention hashing for image retrieval
Shuli Cheng |
Expert Syst. Appl. | 3 |
| 2024 | Solving flexible job shop scheduling problems via deep reinforcement learning
Erdong Yuan, Shuli Cheng, Shiji Song |
Expert Syst. Appl. | 3 |
| 2024 | Deep hashing image retrieval based on hybrid neural network and optimized metric learning
Xingming Xiao, Shu Cao, Shuli Cheng, Erdong Yuan |
Knowl. Based Syst. | 4 |
| 2024 | ER-Swin: Feature Enhancement and Refinement Network Based on Swin Transformer for Semantic Segmentation of Remote Sensing ImagesabstractAs the field of remote sensing images processing continues to advance, semantic segmentation has become a focal point in this domain. The emergence of Swin Transformer has greatly alleviated the computational complexities associated with Transformers, leading to its widespread application in the field of semantic segmentation. However, most current network models lack a feature enhancement process internally, and the model’s tail lacks refinement modules to prevent category misjudgments caused by feature redundancy. To address this issue, we propose ER-Swin to explore the potential of utilizing Swin Transformer as the backbone network for semantic segmentation in remote sensing images. Addressing the need for feature enhancement in the backbone network, we propose the Interactive Feature Enhancement Attention (IFEA), which leverages diagonal information interaction to augment features. Additionally, we design the Semantic Selective Refinement Module (SSRM) to refine the rich features at the tail end of the network, thereby enhancing segmentation outcomes. We evaluate our model on the Vaihingen, Potsdam and LoveDA datasets, and achieved accuracies of 84.89%, 87.20%, and 55.1% on the mIoU metric. Through comparative experiments, we demonstrate the superior segmentation performance of our model, affirming its competitivenes. Shuli Cheng, Anyu Du |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2024 | Medical image segmentation based on dynamic positioning and region-aware attention
Zhongmiao Huang, Shuli Cheng |
Pattern Recognit. | 2 |
| 2024 | Multi-Stage Auxiliary Learning for Visible-Infrared Person Re-IdentificationabstractVisible-infrared person re-identification (VI-ReID) is a challenging study aimed at retrieving the same person across cameras, time, and modalities. Existing methods usually employ dual-stream networks with integrated constraints, or compensate for modality information to reduce the significant modality discrepancies among heterogeneous images. However, the effectiveness of designed constraints is often limited due to substantial cross-modality differences, and methods that compensate for modality information may introduce noise and additional computational cost. In this paper, we propose a novel Multi-Stage Auxiliary Learning strategy called MSALNet. Specifically, in our approach, the training process is bifurcated into two stages: 1) training with auxiliary modality pairs obtained from grayscale histogram equalization, and 2) training with visible and infrared image pairs to gradually extract more discriminative modality-shared features. We propose the Heterogeneous Feature Compensation Learning (HFCL) module for information compensation and fusion between visible and infrared features, generating auxiliary branches to learn more cross-modality-related information. Additionally, we propose the Modality Similarity Reinforcement (MSR) module to improve the consistency of cross-modality feature representation by suppressing interference information and leveraging pixel similarity probability distribution as supervisory information. Lastly, we design the Distance Center Alignment (DCA) loss to reduce intra-class variations within and between modalities, enhancing the distinguishability among different identities. Experimental results demonstrate MSALNet’s superior performance over most existing methods on two mainstream VI-ReID datasets and effectively saves computational cost. Shuli Cheng, Anyu Du |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | CACFTNet: A Hybrid Cov-Attention and Cross-Layer Fusion Transformer Network for Hyperspectral Image ClassificationabstractHyperspectral(HS) image classification has become an important research area. Although previous work on HS image classification has achieved impressive results, finding a proper balance between extracting spatial-spectral information and capturing band similarities remains a challenge. To address this problem, we designed some efficient modules and constructed a hybrid covariance attention and cross-layer fusion Transformer network (CACFTNet), which efficiently models the extraction of spatial-spectral information and band similarity. First, our approach combines local and global perspectives to process the desired feature information. We introduce the Dual Branch Feature Processing (DBFP) module, which can model the desired spatial-spectral information and band similarity. Secondly, we design the Dual Branch Feature Fusion (DBFF) module, which combines a Convolutional Neural Network (CNN) and a transformer to catch spatial-spectral information at different scales and fuse them effectively. To further process the features, we construct the Unparameterized Covariance Attention (UPCA) module, which utilizes the covariance matrix to capture the correlation between different spectral bands. This allows the network to concentrate on bands that are more useful for classification task. Additionally, we designed a hybrid activation function (HAF) that maps the channel values to specific ranges and emphasizes the extent to which the correlation varies between different bands. Finally, in order to incorporate important information from different layers, we propose the Cross-Layer Adaptive Attention Fusion (CAAF) module, which fully fuses information between layers and enriches the overall information representation. We evaluate our proposed model on three well-known public datasets and demonstrate its superiority over existing approaches. Shuli Cheng, Runze Chan, Anyu Du |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | MS2I2Former: Multiscale Spatial-Spectral Information Interactive Transformer for Hyperspectral Image ClassificationabstractTransformer models are increasingly used in hyperspectral image (HSI) classification, thanks to their excellent global feature extraction capabilities. However, these networks still need to be improved in recognizing locally complex feature shapes at different scales and handling linear and nonlinear complex correlations between spectral channels. To this end, we propose an innovative multiscale spatial–spectral information interaction transformer (MS2I2Former) architecture. The architecture skillfully integrates lightweight convolution and Transformer, effectively integrates local and global multiscale spatial features and spectral information, and realizes effective interaction between different scales. We design a multiscale spatial–spectral information interaction (MS2I2) module, which efficiently captures multiscale spatial–spectral features by combining deep convolution of convolution kernels of different sizes and orientations with the frequency domain. Based on this, we propose a distance mean cross-covariance representation (DMC2R) based on distance covariance, which aims to deeply explore the linear and nonlinear relationships between different spectral channels. Considering the convolutional kernel parameters and the comprehensive extraction of joint spectral–space features, we developed the hybrid convolution (HC) module, which combines multiple lightweight convolutions to extract deeper spectral-space features. To model complex remote feature relationships, we innovatively propose the multiscale double cross-symmetric transformer (MDCST) module. This module feeds the rich feature representations after multiscale mapping into double cross-symmetric attention (DCSA), which enhances the internal interactions and fusions among features to capture a wider range of feature dependencies. Experimental results show that on four public datasets, MS2I2Former achieves excellent classification results with fewer training samples compared to existing methods. The source code link is available athttps://github.com/cslxju/MS2I2Former. Shuli Cheng, Runze Chan, Anyu Du |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | HCGNet: A Hybrid Change Detection Network Based on CNN and GNNabstractImage features occur at different scales, e.g., short-term and long-term ones, and both of them are significant in the change detection (CD) of remote sensing images. To the best of our knowledge, however, it is still a challenge on how to effectively combine them together for a full-scale CD. The development of deep learning techniques brings the light on this issue. In this work, we propose a hybrid initiative called HCGNet, combining convolutional neural network (CNN) and vision graph neural network (ViG) for capturing the local and global features, respectively, in which we conduct two main adaptions for high accuracy: 1) a shift graph convolution module to establish the association between a node and its surrounding nodes for enhancing the local feature extraction capability and 2) a dual-branch decoder structure that efficiently utilizes the multiscale features acquired from the encoder to enhance the accuracy of the change map. The results show that: 1) our proposal outperforms the state-of-the-art works and 2) the individual functions of each component are obvious in the ablation experiments. Shuli Cheng |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | A Multiscale Cascaded Cross-Attention Hierarchical Network for Change Detection on Bitemporal Remote Sensing ImagesabstractRemote sensing image change detection (RSCD) is an important task in remote sensing image interpretation. Some recent RSCD works focus on the extraction and interaction of global and local information. However, the current work underutilizes hierarchical features and may introduce noise from shallow encoders. In this paper, we propose a multi-scale cascaded cross-attention hierarchical network (MSCCA-Net). This network utilizes a large kernel convolution formed by stacking small kernel convolutions combined with Efficient Transformer as the backbone network to achieve local and global feature extraction and fusion. We proposed for the first time the idea of bottom-up level-by-level fusion of hierarchical features, based on which we designed the multi- scale cascade cross-attention (MSCCA) cross-fusion hierarchical features level by level from the bottom upwards, realizing the redistribution of spatial and semantic information, and thus enhancing the gainful effect of the skip connection mechanism in the field of RSCD. Our experiments on three public datasets show that MSCCA is able to efficiently perform the reorganization of hierarchical features thus avoiding misdetection and omission of small targets. Meanwhile, MSCCA-Net has more excellent comprehensive performance compared with other state-of-the-art methods. Shuli Cheng |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | H2MWSTNet: A Hybrid Hierarchical Multigranularity Window Shift Transformer Network for Hyperspectral Image ClassificationabstractCurrently, hybrid network models combining convolutional neural networks (CNNs) and Transformers have garnered significant interest among researchers in hyperspectral image (HSI) classification. However, some hybrid models only consider the combination of a single-branch CNN and Transformer, neglecting the shallow multiscale spatial-spectral information and the deeper local spatial information. To address these issues, we propose a hybrid hierarchical multigranularity window shift transformer network (H2MWSTNet), which can efficiently extract multiscale spatial-spectral information and capture local spatial information in sequences. Specifically, we construct a new dual-branch spatial spectral convolution (DBSSC) block. The block effectively realizes the fusion of spatial-spectral information at different scales by aggregating the outputs of both spatial and spectral branches. Furthermore, we have designed a global feature enhancement (GFE) module that effectively enhances the features of the input token sequences. Along with GFE, we propose a multigranularity window shift attention (MWSA) module, which dynamically shifts the key matrix, enhancing the distinguishability of deeply extracted local spatial features. The GFE and MWSA modules provide a novel perspective for HSI feature extraction, constituting a feature mixer based on GFE and multigranularity window shift attention (GMSA) module. Building on the generic architecture and GMSA, we further developed a multigranularity window shift transformer (MWSFormer) encoder block. Finally, our CNN-Transformer hybrid model is designed in a hierarchical fashion, effectively reducing dimensionality and markedly improving classification accuracy. We conducted extensive experiments on six public HSI datasets, and the results demonstrate that H2MWSTNet outperforms the most advanced classification algorithms. Linshan Zhong, Shuli Cheng |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Combined query image retrieval based on hybrid coding of CNN and Mix-Transformer
Zhiwei Zhang 0008, Shuli Cheng |
Expert Syst. Appl. | 2 |
| 2023 | MLP-based classification of COVID-19 and skin diseases
Ruize Zhang 0002, Shuli Cheng, Shiji Song |
Expert Syst. Appl. | 3 |
| 2023 | Residual dense collaborative network for salient object detectionabstractAbstract Owing to the renaissance of deep convolutional neural networks (CNN), salient object detection based on fully convolutional neural networks (FCNs) has attracted widespread attention. However, the scale variation of prominent objects, complex background features and fuzzy edges have historically been a great challenge to us. All these are closely associated with the utilization of multi‐level and multi‐scale features. At the same time, deep learning methods meet the challenges of computation and memory consumption in practice. To address these problems, the authors propose a different salient object detection method based on residuals learning and dense fusion learning framework. The proposed network is named Residual Dense Collaborative Network (RDCNet). First of all, the authors design a multi‐layer residual learning (MRL) module to extract salient object features in more detail, getting the utmost out of the object's multi‐scale and multi‐level information. Then, on the basis of the vigoroso stage‐wise convolution feature, the authors put forward the dilated convolution module (DCM) to acquire a rough global saliency map. Finally, the final accurate saliency detection map is obtained through dense cooperation learning (DCL), and the remaining learning is also used to improve gradually, so as to achieve high compactness and high‐efficiency results. Experimental results show that this method is the most advanced method for five widely used datasets (DUTS‐TE, HKU‐IS, PASCAL‐S, ECSSD, DUT‐OMRON) without any pre‐processing and post‐processing. Especially on the ECSSD dataset, the F‐measure of RDCNet achieves 95.2%. Yibo Han, Shuli Cheng, Anyu Du |
IET Image Process. | 3 |
| 2023 | Deep internally connected transformer hashing for image retrievalabstractTransformer based on self-attention mechanism has made remarkable achievements in natural language processing , which inspired the application research of Transformer in computer vision . The current deep hashing algorithms extract image features through the convolutional neural network (CNN). CNN concentrates on local information , and features lack global dependency information, which has an impact on image retrieval accuracy . To remedy the above defects, this paper proposes deep internally connected Transformer hashing for image retrieval (DICTH). DICTH has designed an improved Transformer block: internally connected Transformer block (ICT). ICT performs an embedded transformation on the feature maps, splices the generated Keys and Queries, to explore the rich context information between Query–Key pairs, and then dynamically encodes through multi-layer convolution to learn the context multi-head self-attention matrix. By combining ICT and ResNet18 to achieve self-attention injection, a long-distance dependency is established in the feature space to make up for the shortcomings of pure CNN in the feature extraction process and guide the algorithm to learn more accurate hash codes. At the same time, in the face of complex label information in big data sets, this paper uses an improved cross-entropy loss function: T-cross-entropy loss, to promote network learning of hash codes with more ability to distinguish between classes. In this paper, a lot of experiments have been conducted on CIFAR10, NUS-WIDE and MS-COCO datasets to verify the performance of DICTH. Zijian Chao, Shuli Cheng |
Knowl. Based Syst. | 2 |
| 2023 | Lightweight Remote-Sensing Image Super-Resolution via Attention-Based Multilevel Feature Fusion NetworkabstractIn recent years, advancements in remote-sensing image super-resolution have achieved remarkable performance. However, many methods demand significant computational resources. This is problematic for edge devices with limited computational capabilities. To alleviate this problem, we propose an attention-based multi-level feature fusion network (AMFFN) to enhance the resolution of remote-sensing images. This proposed network integrates three efficient design strategies to provide a lightweight solution. Initially, we design the partial shallow residual block (PSRB) to replace the redundant convolution operation. The PSRB optimizes feature extraction via partial convolution and capitalizes on information across channels using pointwise convolution. Subsequently, integrating the PSRB, our dynamic feature distillation block (DFDB) leverages an information distillation mechanism to distill and capture only the crucial features securing a robust feature depiction. Conclusively, for superior feature fusion, we conceptualized an attention-based multi-level feature fusion (AMFF) mechanism. The attention intrinsic to AMFF weighs the significance of features from varied branches, assuring that the resulting output is comprehensive and discerning. We conduct thorough experimental validation on two datasets of remote-sensing images and measure network complexity by evaluating network parameters and multi-adds operations. The results show that our method effectively balances computational complexity and performance. In addition, we have expanded the application of AMFFN to the field of natural image super-resolution. Experimental results on five benchmark test datasets further confirm the effectiveness of our method. Shuli Cheng, Anyu Du |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Asymmetric Cross-Attention Hierarchical Network Based on CNN and Transformer for Bitemporal Remote Sensing Images Change DetectionabstractAs an important task in the field of remote sensing (RS) image processing, RS image change detection (CD) has made significant advances through the use of convolutional neural networks (CNNs). The transformer has recently been introduced into the field of CD due to its excellent global perception capabilities. Some works have attempted to combine CNN and transformer to jointly harvest local-global features; however, these works have not paid much attention to the interaction between the features extracted by both. Also, the use of the transformer has resulted in significant resource consumption. In this article, we propose the Asymmetric Cross-attention Hierarchical Network (ACAHNet) by combining CNN and transformer in a series-parallel manner. The proposed Asymmetric Multiheaded Cross Attention (AMCA) module reduces the quadratic computational complexity of the transformer to linear, and the module enhances the interaction between features extracted from the CNN and the transformer. Different from the early and late fusion strategies employed in previous work, the effectiveness of the mid-term fusion strategy employed by ACAHNet shows a new choice of timing for feature fusion in the CD task. Our experiments on the proposed method on three public datasets show that our network has a better performance in terms of effectiveness and computational resource consumption compared to other comparative methods. Shuli Cheng, Haojin Li 0002 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Dynamic dual attention iterative network for image super-resolution
Shuli Cheng, Anyu Du |
Appl. Intell. | 3 |
| 2022 | Low-Rank Semantic Feature Reconstruction Hashing for Remote Sensing RetrievalabstractRemote sensing image retrieval (RSIR) is the main technology for automatic analysis and understanding of remote sensing big data, which has been widely concerned in recent years. The mainstream attention mechanism based on high order tensor plays an important role in computer vision, but the model complexity is high. Low-rank feature reconstruction can reconstruct high-order context semantics based on low-rank tensor, and the feature reconstruction layer can realize the fusion of high-order context semantics. In order to reconstruct high-order remote sensing semantics, we propose a novel low-rank semantic feature reconstruction hashing (LRSFRH) using a lightweight dual-attention mechanism and semantic reservation loss to capture remote contextual semantic information of remote sensing scenes for remote sensing retrieval. Its main contributions are as follows: 1) lightweight dual-attention mechanism is proposed based on effective channel attention (ECA) and high-order tensor reconstruction (HTR). Among them, HTR can explore high-order contextual semantic information of remote sensing with low-order constraints, and ECA’s cross-channel interaction can significantly reduce model complexity while maintaining performance; 2) in the remote sensing feature hashing space, we use second-order global covariance pooling (GCP) to accelerate model convergence and enrich remote sensing semantic representation; and 3) in metric learning, we propose a new multiple semantic reconstruction loss (MSRL) to optimize network parameters. Experimental results show that LRSFRH outperforms most existing hash algorithms on two public benchmark datasets (AID and UC Merced), and the proposed algorithm achieves state-of-the-art (SOTA) performance in RSIR tasks. Anyu Du, Shuli Cheng |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | SwinSUNet: Pure Transformer Network for Remote Sensing Image Change DetectionabstractConvolutional neural network (CNN) can extract effective semantic features, so it was widely used for remote sensing image change detection (CD) in the latest years. CNN has acquired great achievements in the field of CD, but due to the intrinsic locality of convolution operation, it could not capture global information in space-time. The transformer was proposed in recent years and it can effectively extract global information, so it was used to solve computer vision (CV) tasks and achieved amazing success. In this article, we design a pure transformer network with Siamese U-shaped structure to solve CD problems and name it SwinSUNet. SwinSUNet contains encoder, fusion, and decoder, and all of them use Swin transformer blocks as basic units. Encoder has a Siamese structure based on hierarchical Swin transformer, so encoder can process bitemporal images in parallel and extract their multiscale features. Fusion is mainly responsible for the merge operation of the bitemporal features generated by the encoder. Like encoder, the decoder is also based on hierarchical Swin transformer. Different from the encoder, the decoder uses upsampling and merging (UM) block and Swin transformer blocks to recover the details of the change information. The encoder uses patch merging and Swin transformer blocks to generate effective semantic features. After the sequential process of these three modules, SwinSUNet will output the change maps. We did expensive experiments on four CD datasets, and in these experiments, SwinSUNet achieved better results than other related methods. Shuli Cheng |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2021 | Bidirectional Focused Semantic Alignment Attention Network for Cross-Modal RetrievalabstractCross-modal retrieval is a very challenging and significant task in intelligent understanding. Researchers have tried to capture modal semantic information through a weighted attention mechanism. Still, they cannot eliminate irrelevant semantic information's negative effects and cannot capture fine-grained modal semantic information. In order to further accurately capture the multi-modal semantic information, a bidirectional focused semantic alignment attention network (BFSAAN) is proposed to handle cross-modal retrieval tasks. Core ideas of BFSAAN are as follows: 1) Bidirectional focused attention mechanism is adopted to share modal semantic information, further eliminating the negative influence of irrelevant semantic information. 2) Strip pooling is applied to image and text modalities, a lightweight spatial attention mechanism to capture modal spatial semantic information. 3) Second-order covariance pooling is explored to obtain multi-modal semantic representation, capturing modal channel semantic information and achieving semantic alignment between image-text modalities. The experiment is executed in two standard cross-modal retrieval datasets (Flickr30K and MS COCO). The experimental design includes four aspects: performance comparison, ablation analysis, algorithm convergence, and visual analysis. Experimental results show that BFSAAN has better crossmodal retrieval performance. Shuli Cheng, Anyu Du |
ICASSP | 1 |
| 2021 | Fusion layer attention for image-text matching
Depeng Wang, Shiji Song, Gao Huang 0001, Shuli Cheng, Naixiang Ao, Anyu Du |
Neurocomputing | 6 |
| 2021 | A privacy-preserving image retrieval scheme based secure kNN, DNA coding and deep hashing
Shuli Cheng, Gao Huang 0001, Anyu Du |
Multim. Tools Appl. | 1 |
| 2019 | A novel deep hashing method for fast image retrieval
Shuli Cheng, Huicheng Lai, Jiwei Qin |
Vis. Comput. | 1 |