VLDB 2026 Research / reviewers in the wild / expert
Anyu Du
dblp:243/8997
· DBLP profile ↗
20ranked-venue papers
1as first author
20since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 10 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Adaptive structural-semantic fusion and optimization for hashing image retrieval
Shuli Cheng, Anyu Du, Tingjie Liu |
Expert Syst. Appl. | 3 |
| 2026 | Confidence aware Mamba interaction hashing for cross-modal retrieval
Jiapeng Tian, Shuli Cheng, Anyu Du |
Expert Syst. Appl. | 3 |
| 2025 | Covariance Attention Guidance Mamba Hashing for cross-modal retrieval
Shuli Cheng, Anyu Du, Qiang Zou 0002 |
Eng. Appl. Artif. Intell. | 3 |
| 2025 | Spatial and learnable frequency dynamic collaborative visual perception network for remote sensing images semantic segmentationabstractRemote sensing semantic segmentation , as a current research hotspot, is widely applied in scenarios such as agricultural planning and urban construction planning . Deep learning models based on spatial and frequency domain co-perception have rapidly developed due to their comprehensive perception capabilities. However, current networks primarily apply frequency domain perception within various attention mechanisms , lacking complementary learning with the spatial domain. In addition, whether the encoder–decoder pairing in existing U-shaped architectures can fully exploit the advantages of dual-domain perception requires further discussion. To address these issues, this paper proposes the Spatial and Learnable Frequency Dynamic Collaborative Visual Perception Network (SFCVPNet). The encoder uses a cascade of Spatial and Frequency Domain-Aware Transformer (SFFormer) blocks and Convolutional Neural Networks (CNNs) blocks to perceive features, while the decoder employs simple skip connections, forming a novel integration method of CNNs and Transformer blocks in the U-shaped architecture. SFFormer includes the Spatial and Frequency Domain Collaborative Perception Module (SFCPM) and the Spatial and Frequency Guided Multi-Layer Perceptron (SF-MLP). Both SFCPM and SF-MLP utilize a dual-domain perception approach to extract and integrate features. Additionally, we adopt learnable operations for frequency domain feature perception, ensuring that frequency domain features meet the learning requirements of the network model. On three public datasets: ISPRS Vaihingen, ISPRS Potsdam, and LoveDA, respectively, we achieved mIoU scores of 85.90%, 88.19%, and 55.6%. The code will be released at https://github.com/cslxju/SFCVPNet . Shuli Cheng, Anyu Du |
Expert Syst. Appl. | 3 |
| 2025 | Content-adaptive perception mixer visual transformer for hashing image retrieval
Tingjie Liu, Shuli Cheng, Anyu Du |
Expert Syst. Appl. | 3 |
| 2025 | Occlusion Simulation and Token-Constrained Feature Coupling Network for Occluded Person ReidentificationabstractOccluded person reidentification (Re-ID) aims to learn the features of pedestrians with different identities under occlusion, and is a pivotal technology for intelligent security surveillance systems for the Internet of Things (IoT). In real scenarios, occlusion may occur at any spatial location and exhibits irregularity, occluded people Re-ID remains a challenging task. Most existing methods distinguish visible body parts by leveraging location-based segmentation or external cues, but they are often inefficient or increase network complexity. To address these problems, we propose an occlusion simulation and token-constrained feature coupling network (OSTCNet). Specifically, the occlusion simulation based on block mixing (OSBBM) strategy segments pedestrian images with different labels into blocks during the training phase and proportionally blends the image blocks based on coordinates, thereby generating occluded samples with diverse occlusion patterns. Additionally, we propose the local–global feature coupling (LGFC) module to perform multiscale coupling of local features and global representations, enhancing classification accuracy and capturing more comprehensive feature information. Given the lack of explicit constraints in ViT to differentiate the similarity between class token and patch tokens, which can cause accumulation of similar information in the feature embedding space, we introduce the token orthogonal embedding (TOE) module to enforce constraints on class token, ensuring representational differences from patch tokens. Experimental results on occluded, partial, and holistic ReID datasets validate the effectiveness of our method. Specifically, on the Occluded-DukeMTMC dataset, OSTCNet achieves Rank-1 accuracy of 75.2% and mean-average precision (mAP) score of 65.7%. Shuli Cheng, Anyu Du |
IEEE Internet Things J. | 3 |
| 2025 | Dual-stream feature extraction and semantic similarity for deep hashing image retrieval
Shuli Cheng, Anyu Du, Tingjie Liu |
Knowl. Based Syst. | 3 |
| 2024 | Multi-view similarity aggregation and multi-level gap optimization for unsupervised person re-identification
Shuli Cheng, Anyu Du |
Expert Syst. Appl. | 3 |
| 2024 | ER-Swin: Feature Enhancement and Refinement Network Based on Swin Transformer for Semantic Segmentation of Remote Sensing ImagesabstractAs the field of remote sensing images processing continues to advance, semantic segmentation has become a focal point in this domain. The emergence of Swin Transformer has greatly alleviated the computational complexities associated with Transformers, leading to its widespread application in the field of semantic segmentation. However, most current network models lack a feature enhancement process internally, and the model’s tail lacks refinement modules to prevent category misjudgments caused by feature redundancy. To address this issue, we propose ER-Swin to explore the potential of utilizing Swin Transformer as the backbone network for semantic segmentation in remote sensing images. Addressing the need for feature enhancement in the backbone network, we propose the Interactive Feature Enhancement Attention (IFEA), which leverages diagonal information interaction to augment features. Additionally, we design the Semantic Selective Refinement Module (SSRM) to refine the rich features at the tail end of the network, thereby enhancing segmentation outcomes. We evaluate our model on the Vaihingen, Potsdam and LoveDA datasets, and achieved accuracies of 84.89%, 87.20%, and 55.1% on the mIoU metric. Through comparative experiments, we demonstrate the superior segmentation performance of our model, affirming its competitivenes. Shuli Cheng, Anyu Du |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2024 | Multi-Stage Auxiliary Learning for Visible-Infrared Person Re-IdentificationabstractVisible-infrared person re-identification (VI-ReID) is a challenging study aimed at retrieving the same person across cameras, time, and modalities. Existing methods usually employ dual-stream networks with integrated constraints, or compensate for modality information to reduce the significant modality discrepancies among heterogeneous images. However, the effectiveness of designed constraints is often limited due to substantial cross-modality differences, and methods that compensate for modality information may introduce noise and additional computational cost. In this paper, we propose a novel Multi-Stage Auxiliary Learning strategy called MSALNet. Specifically, in our approach, the training process is bifurcated into two stages: 1) training with auxiliary modality pairs obtained from grayscale histogram equalization, and 2) training with visible and infrared image pairs to gradually extract more discriminative modality-shared features. We propose the Heterogeneous Feature Compensation Learning (HFCL) module for information compensation and fusion between visible and infrared features, generating auxiliary branches to learn more cross-modality-related information. Additionally, we propose the Modality Similarity Reinforcement (MSR) module to improve the consistency of cross-modality feature representation by suppressing interference information and leveraging pixel similarity probability distribution as supervisory information. Lastly, we design the Distance Center Alignment (DCA) loss to reduce intra-class variations within and between modalities, enhancing the distinguishability among different identities. Experimental results demonstrate MSALNet’s superior performance over most existing methods on two mainstream VI-ReID datasets and effectively saves computational cost. Shuli Cheng, Anyu Du |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | CACFTNet: A Hybrid Cov-Attention and Cross-Layer Fusion Transformer Network for Hyperspectral Image ClassificationabstractHyperspectral(HS) image classification has become an important research area. Although previous work on HS image classification has achieved impressive results, finding a proper balance between extracting spatial-spectral information and capturing band similarities remains a challenge. To address this problem, we designed some efficient modules and constructed a hybrid covariance attention and cross-layer fusion Transformer network (CACFTNet), which efficiently models the extraction of spatial-spectral information and band similarity. First, our approach combines local and global perspectives to process the desired feature information. We introduce the Dual Branch Feature Processing (DBFP) module, which can model the desired spatial-spectral information and band similarity. Secondly, we design the Dual Branch Feature Fusion (DBFF) module, which combines a Convolutional Neural Network (CNN) and a transformer to catch spatial-spectral information at different scales and fuse them effectively. To further process the features, we construct the Unparameterized Covariance Attention (UPCA) module, which utilizes the covariance matrix to capture the correlation between different spectral bands. This allows the network to concentrate on bands that are more useful for classification task. Additionally, we designed a hybrid activation function (HAF) that maps the channel values to specific ranges and emphasizes the extent to which the correlation varies between different bands. Finally, in order to incorporate important information from different layers, we propose the Cross-Layer Adaptive Attention Fusion (CAAF) module, which fully fuses information between layers and enriches the overall information representation. We evaluate our proposed model on three well-known public datasets and demonstrate its superiority over existing approaches. Shuli Cheng, Runze Chan, Anyu Du |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | MS2I2Former: Multiscale Spatial-Spectral Information Interactive Transformer for Hyperspectral Image ClassificationabstractTransformer models are increasingly used in hyperspectral image (HSI) classification, thanks to their excellent global feature extraction capabilities. However, these networks still need to be improved in recognizing locally complex feature shapes at different scales and handling linear and nonlinear complex correlations between spectral channels. To this end, we propose an innovative multiscale spatial–spectral information interaction transformer (MS2I2Former) architecture. The architecture skillfully integrates lightweight convolution and Transformer, effectively integrates local and global multiscale spatial features and spectral information, and realizes effective interaction between different scales. We design a multiscale spatial–spectral information interaction (MS2I2) module, which efficiently captures multiscale spatial–spectral features by combining deep convolution of convolution kernels of different sizes and orientations with the frequency domain. Based on this, we propose a distance mean cross-covariance representation (DMC2R) based on distance covariance, which aims to deeply explore the linear and nonlinear relationships between different spectral channels. Considering the convolutional kernel parameters and the comprehensive extraction of joint spectral–space features, we developed the hybrid convolution (HC) module, which combines multiple lightweight convolutions to extract deeper spectral-space features. To model complex remote feature relationships, we innovatively propose the multiscale double cross-symmetric transformer (MDCST) module. This module feeds the rich feature representations after multiscale mapping into double cross-symmetric attention (DCSA), which enhances the internal interactions and fusions among features to capture a wider range of feature dependencies. Experimental results show that on four public datasets, MS2I2Former achieves excellent classification results with fewer training samples compared to existing methods. The source code link is available athttps://github.com/cslxju/MS2I2Former. Shuli Cheng, Runze Chan, Anyu Du |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Residual dense collaborative network for salient object detectionabstractAbstract Owing to the renaissance of deep convolutional neural networks (CNN), salient object detection based on fully convolutional neural networks (FCNs) has attracted widespread attention. However, the scale variation of prominent objects, complex background features and fuzzy edges have historically been a great challenge to us. All these are closely associated with the utilization of multi‐level and multi‐scale features. At the same time, deep learning methods meet the challenges of computation and memory consumption in practice. To address these problems, the authors propose a different salient object detection method based on residuals learning and dense fusion learning framework. The proposed network is named Residual Dense Collaborative Network (RDCNet). First of all, the authors design a multi‐layer residual learning (MRL) module to extract salient object features in more detail, getting the utmost out of the object's multi‐scale and multi‐level information. Then, on the basis of the vigoroso stage‐wise convolution feature, the authors put forward the dilated convolution module (DCM) to acquire a rough global saliency map. Finally, the final accurate saliency detection map is obtained through dense cooperation learning (DCL), and the remaining learning is also used to improve gradually, so as to achieve high compactness and high‐efficiency results. Experimental results show that this method is the most advanced method for five widely used datasets (DUTS‐TE, HKU‐IS, PASCAL‐S, ECSSD, DUT‐OMRON) without any pre‐processing and post‐processing. Especially on the ECSSD dataset, the F‐measure of RDCNet achieves 95.2%. Yibo Han, Shuli Cheng, Anyu Du |
IET Image Process. | 5 |
| 2023 | Lightweight Remote-Sensing Image Super-Resolution via Attention-Based Multilevel Feature Fusion NetworkabstractIn recent years, advancements in remote-sensing image super-resolution have achieved remarkable performance. However, many methods demand significant computational resources. This is problematic for edge devices with limited computational capabilities. To alleviate this problem, we propose an attention-based multi-level feature fusion network (AMFFN) to enhance the resolution of remote-sensing images. This proposed network integrates three efficient design strategies to provide a lightweight solution. Initially, we design the partial shallow residual block (PSRB) to replace the redundant convolution operation. The PSRB optimizes feature extraction via partial convolution and capitalizes on information across channels using pointwise convolution. Subsequently, integrating the PSRB, our dynamic feature distillation block (DFDB) leverages an information distillation mechanism to distill and capture only the crucial features securing a robust feature depiction. Conclusively, for superior feature fusion, we conceptualized an attention-based multi-level feature fusion (AMFF) mechanism. The attention intrinsic to AMFF weighs the significance of features from varied branches, assuring that the resulting output is comprehensive and discerning. We conduct thorough experimental validation on two datasets of remote-sensing images and measure network complexity by evaluating network parameters and multi-adds operations. The results show that our method effectively balances computational complexity and performance. In addition, we have expanded the application of AMFFN to the field of natural image super-resolution. Experimental results on five benchmark test datasets further confirm the effectiveness of our method. Shuli Cheng, Anyu Du |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Dynamic dual attention iterative network for image super-resolution
Shuli Cheng, Anyu Du |
Appl. Intell. | 4 |
| 2022 | LKASR: Large kernel attention for lightweight image super-resolution
Anyu Du |
Knowl. Based Syst. | 4 |
| 2022 | Low-Rank Semantic Feature Reconstruction Hashing for Remote Sensing RetrievalabstractRemote sensing image retrieval (RSIR) is the main technology for automatic analysis and understanding of remote sensing big data, which has been widely concerned in recent years. The mainstream attention mechanism based on high order tensor plays an important role in computer vision, but the model complexity is high. Low-rank feature reconstruction can reconstruct high-order context semantics based on low-rank tensor, and the feature reconstruction layer can realize the fusion of high-order context semantics. In order to reconstruct high-order remote sensing semantics, we propose a novel low-rank semantic feature reconstruction hashing (LRSFRH) using a lightweight dual-attention mechanism and semantic reservation loss to capture remote contextual semantic information of remote sensing scenes for remote sensing retrieval. Its main contributions are as follows: 1) lightweight dual-attention mechanism is proposed based on effective channel attention (ECA) and high-order tensor reconstruction (HTR). Among them, HTR can explore high-order contextual semantic information of remote sensing with low-order constraints, and ECA’s cross-channel interaction can significantly reduce model complexity while maintaining performance; 2) in the remote sensing feature hashing space, we use second-order global covariance pooling (GCP) to accelerate model convergence and enrich remote sensing semantic representation; and 3) in metric learning, we propose a new multiple semantic reconstruction loss (MSRL) to optimize network parameters. Experimental results show that LRSFRH outperforms most existing hash algorithms on two public benchmark datasets (AID and UC Merced), and the proposed algorithm achieves state-of-the-art (SOTA) performance in RSIR tasks. Anyu Du, Shuli Cheng |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2021 | Bidirectional Focused Semantic Alignment Attention Network for Cross-Modal RetrievalabstractCross-modal retrieval is a very challenging and significant task in intelligent understanding. Researchers have tried to capture modal semantic information through a weighted attention mechanism. Still, they cannot eliminate irrelevant semantic information's negative effects and cannot capture fine-grained modal semantic information. In order to further accurately capture the multi-modal semantic information, a bidirectional focused semantic alignment attention network (BFSAAN) is proposed to handle cross-modal retrieval tasks. Core ideas of BFSAAN are as follows: 1) Bidirectional focused attention mechanism is adopted to share modal semantic information, further eliminating the negative influence of irrelevant semantic information. 2) Strip pooling is applied to image and text modalities, a lightweight spatial attention mechanism to capture modal spatial semantic information. 3) Second-order covariance pooling is explored to obtain multi-modal semantic representation, capturing modal channel semantic information and achieving semantic alignment between image-text modalities. The experiment is executed in two standard cross-modal retrieval datasets (Flickr30K and MS COCO). The experimental design includes four aspects: performance comparison, ablation analysis, algorithm convergence, and visual analysis. Experimental results show that BFSAAN has better crossmodal retrieval performance. Shuli Cheng, Anyu Du |
ICASSP | 3 |
| 2021 | Fusion layer attention for image-text matching
Depeng Wang, Shiji Song, Gao Huang 0001, Shuli Cheng, Naixiang Ao, Anyu Du |
Neurocomputing | 8 |
| 2021 | A privacy-preserving image retrieval scheme based secure kNN, DNA coding and deep hashing
Shuli Cheng, Gao Huang 0001, Anyu Du |
Multim. Tools Appl. | 4 |