Shiyong Lan

dblp:154/7047 · DBLP profile ↗
← Back
30ranked-venue papers
2as first author
30since 2021 · last 2026
0000-0002-7109-9170ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 1 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 7 since 2021Human-computer interaction and ubiquitous computing · 5 · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Dual-branch normalizing flow for anomaly detection and localization from images
Yao Li 0019, Shiyong Lan, Wenwu Wang 0001, Yixin Qiao, Guonan Deng
Neurocomputing4
2026 Small object detection using multi-scale detail enhancement and decoupled detection head
Yixin Qiao, Xinyuan Zhou, Shiyong Lan, Wenwu Wang 0001, Yao Li 0019, Guonan Deng
Neurocomputing3
2026 STGFMamba: Spatio-temporal graph Fourier-enhanced Mamba for traffic prediction
Xinyuan Zhou, Ruiyi Lu, Zhiang Hou, Yao Ren, Wenwu Wang 0001, Shiyong Lan
Inf. Sci.7
2026 DSTFGCN: A dynamic spatial-temporal fusion graph convolution network for traffic flow forecasting
abstract
Traffic flow prediction is one of the core technology of Intelligent Transportation System. Its fundamental challenge is to effectively model the complex spatial-temporal dependencies. Although extensive research has been conducted in this field, the limitations of current methods restrict their effectiveness in accurate predictions. For temporal dependence, existing methods based on recurrent neural networks only focus on local dependencies and ignore global dependencies. For spatial dependencies, existing methods use predefined or adaptive adjacency matrices that cannot accurately reflect the relationships between real traffic flow. To overcome these limitations, we propose a dynamic spatial-temporal fusion graph convolution network (DSTFGCN). In the temporal aspect, we introduce gated dilated causal convolution to capture the local dependencies and node-independent temporal graph convolution to capture the global dependencies specific to each node. In the spatial aspect, we propose a dynamic graph convolution block. It can construct dynamic graphs based on the characteristics of the input data and aggregate both local and global spatial dependencies. Experiments on six real-world datasets have shown that DSTFGCN outperforms current mainstream methods. The codes are available at https://github.com/SYLan2019/DSTFGCN.
Tianyi Pan, Xinyuan Zhou, Shiyong Lan, Wenwu Wang 0001, Hongyu Yang 0002, Zhiang Hou, Yao Ren
Neural Networks3
2026 DMAGaze : Gaze estimation using feature disentanglement and multi-scale attention
Haohan Chen, Hongjia Liu, Shiyong Lan, Wenwu Wang 0001, Yixin Qiao, Yao Li 0019, Guonan Deng
Pattern Recognit. Lett.3
2025 SAR Ship Detector Using Cross-stage Feature Fusion and Decoupled Head with Mutual Guidance
abstract
Deep learning-based SAR ship detection methods enhance resilience to noise, distortion, and interference in ocean environments, establishing them as the foremost approach for ship detection nowadays. Nonetheless, substantial difficulties persist in separating ships from the complex backgrounds found in SAR images: 1) Many ships are highly similar to the sea surface clutter noise, making them susceptible to false alarms; 2) Ship targets exhibit a wide range of variations in size and shape. In this paper, we propose a novel network to address above problems. Firstly, the deformable convolution is incorporated into the backbone network to adapt to the wide-range of ship shapes. Secondly, the cross-stage feature fusion module (CSFFM) is introduced to realize local self-supervised interaction between two adjacent layers, thereby reducing the impact of receptive field differences between different feature layers and mitigating the influence of complex background noise. Finally, the mutually guided decoupled-head (MGDH) is designed to achieve mutual guidance between classification and regression, thus further enhancing the significant regions of the feature maps. Through extensive experiments, it has been verified that our proposed method has achieved the most promising performance compared to well-known baselines. The codes will be available at https://github.com/SYLan2019/CSFF-MGDH.
Yixin Qiao, Xiaoxiao Yin, Xinyuan Zhou, Shiyong Lan, Guonan Deng
ICASSP4
2025 MRGNN: Mamba-Register-Based Graph Neural Network for Unsupervised Anomaly Detection in Multivariate Time Series
Danling Meng, Xinyuan Zhou, Shiyong Lan, Wenwu Wang 0001, Weihong Yuan, Ruiyi Lu
ICIC (12)3
2025 EMGPose: An Efficient Multi-Granularity Representation for Human Pose Estimation
abstract
Current Transformer-based methods typically unwisely represent the entire image at a single granularity. A high-resolution representation of the regions of interest can significantly improve the accuracy of human pose estimation while causing unnecessary computational costs for other regions. To overcome this limitation, we propose an efficient two-stage framework using adaptive Multi-Granularity representation for different important image regions for Human pose estimation (EMGPose). In the first stage, the image is split into coarse-grained patches for simple inference. If without sufficient accuracy, important patches will be resplit into multiple finer-grained patches for the second stage of inference. Furthermore, we propose a new token-merge strategy based on token importance and similarity in Transformer, effectively reducing the computational load from low-information background patches. Extensive experiments demonstrate the excellent performance of the proposed method. Specifically, our model EMGPose-Base achieves 76.3 AP (+0.5 AP) and 62.2 AP (+2.6 AP) and higher efficiency than baseline ViTPose-Base on the COCO validation set and OCHuman test set, respectively.
Guonan Deng, Shiyong Lan, Wenwu Wang 0001, Yixin Qiao, Yao Li 0019, Haohan Chen, Hongyu Yang 0002
ICME2
2025 DCGNet: Detail and Context Guided Small Object Detection Network with Decoupled Detection Head
abstract
Small object detection plays a significant role in many fields, yet it still faces numerous challenges. These include the lack of key features due to its tiny size, and its sensitivity to environmental noise. To address these issues, we introduce a novel network termed Detail and Context Guided Small Object Detection Network (DCGNet). Specifically, 1) we design a Detail and Context Enhanced Feature Pyramid Network (DCE-FPN), which transforms the feature from the temporal domain into the frequency domain, followed by noise reduction on high-frequency components to accentuate the detailed information of small objects and leverages multiple branches with different receptive fields to capture multi-scale contextual information to further differentiate small objects from the background. 2) We propose a Enhanced Decoupled Interaction RCNN (EDI-RCNN) that independently enhances the specific features for each task and subsequently facilitates comprehensive feature interaction. This helps avoid ambiguity and sparsity of the features of the small objects extracted with coupled detection head in existing methods. The results on the well-known small object detection datasets (VisDrone2019 and AI-TODv2) show that our proposed method achieves SOTA in both Average Precision (AP) and Average Precision for small objects (APs) metrics.
Yixin Qiao, Shiyong Lan, Wenwu Wang 0001, Haohan Chen, Yao Li 0019, Guonan Deng
ICME2
2025 Diff-DTF: Dynamic Temporal Feature Extraction and Refinement with Diffusion Model in Time Series Anomaly Detection
Ruiyi Lu, Weihong Yuan, Xinyuan Zhou, Shiyong Lan
PRICAI (5)4
2025 Spatiotemporal Trend Fusion Feature Graph Convolution Network for Spatial Interpolation in Traffic Scenes⋆
abstract
Sensors are always sparsely distributed in traffic networks due to high deployment costs, sensor damage, etc. Insufficient data may affect our perception of traffic scenarios, resulting in the inability of intelligent transportation systems (ITS) to efficiently perform traffic monitoring and scenario decisions. Spatial interpolation methods are used to infer the status of locations where no sensors are deployed. However, existing methods still have the following limitations: (1) Mainly interpolated unobserved nodes by extracting spatiotemporal dependencies between nodes, but ignored complex interaction patterns in feature dimensions. (2) Recent methods generally used standard TCN to extract temporal correlations, which is affected by abnormal data. (3) When extracting spatial correlations, the use of deep GCN layer can lead to over-smoothing problem. To mitigate these limitations, we propose a novel spatial interpolation model, namely, Spatiotemporal Trend Fusion Feature Graph Convolution Network (STFGCN). Specifically, a novel feature graph convolution network is used to capture complex interaction patterns. Secondly, the trend capture branch is used to alleviate the impact of abnormal data on TCN. Finally, a dual-stage spatial module is used to solve the degradation of detailed feature representations in deep GCN layer. Experimental results on six real traffic datasets demonstrate that our method outperforms state-of-the-art baseline models.
Zhiang Hou, Shiyong Lan, Xinyuan Zhou, Wujiang Zhu, Yao Ren
SMC2
2025 FE-HGAT: Frequency-Enhanced Hybrid Graph Attention Network For Traffic Prediction
abstract
Traffic flow prediction is crucial for urban traffic management and planning. However, although existing research has achieved promising results, most methods primarily focus on time-domain processing, with insufficient exploration of frequency-domain signal characteristics. Moreover, existing approaches often fail to effectively distinguish and simultaneously model the spatial dependencies between nearby and distant nodes. To address these issues, this paper proposes a Frequency-Enhanced Hybrid Graph Attention Network (FE-HGAT) for traffic flow prediction. Our approach employs a dynamic filter in the temporal dimension, which utilizes Fast Fourier Transform (FFT) to enhance key frequency-domain features, thereby better characterizing the temporal dependencies in traffic data. Besides, to capture spatial dependencies between nodes at varying distances, we design a dynamic threshold module to distinguish between nearby and distant nodes, employing external attention (EA) and a mixture-of-experts-enhanced graph attention network (MOE-GAT) to model local dependencies and long-distance semantic similarities, respectively. Experiments demonstrate that FE-HGAT outperforms existing baseline models on several public transportation datasets, validating its effectiveness in traffic forecasting. The code is available at https://github.com/ry123scuer/FE-HGAT.
Yao Ren, Wujiang Zhu, Shiyong Lan, Xinyuan Zhou, Hongyu Yang 0002, Zhiang Hou
SMC3
2025 Spatial Interpolation Based on Causal Spatiotemporal Modeling
abstract
The problem of spatial interpolation is a common challenge in fields such as traffic flow analysis. However, most existing methods directly utilize information from neighboring nodes to infer status of the unobserved location, without considering whether this information contains confounding factors or whether there is a true causal relationship, despite the fact that some unknown confounding factors are inevitably included in the data collection process. To address this, this paper proposes a Causal Attention Spatial Temporal Interpolation (CASI), which leverages causal relationships between nodes for spatial interpolation. The proposed CASI employs dilated convolutions and gating mechanisms to capture temporal dependencies, and introduces the Causal Spatiotemporal Attention (CSTA) mechanism, to uncover spatial causal dependencies between nodes. Subsequently, a novel Graph-Guided Feature Enhance Module (GFEM) is designed, which leverages causal probabilities from CSTA’s Gumbel-Softmax to weight the adjacency matrix in traditional GCN, composing GS-GCN, then adopts self-attention on temporal and spatial dimension respectively to further enhance the features from the improved GCN. We evaluate CASI on three real-world datasets, where it outperforms the optimal baselines across MAE, RMSE and MAPE by up to 3.7%, 1.7%, and 5.4% on PEMS04, 8.5%, 6.0%, and 5.9% on PEMS08, respectively. Ablation studies further demonstrated the effectiveness of the proposed modules in causal spatiotemporal modeling.
Shiyong Lan, Yao Ren, Weihong Yuan, Xinyuan Zhou, Zhiang Hou
SMC2
2025 MADFlow: Multimodal difference compensation flow for multimodal anomaly detection
Yao Li 0019, Xinyuan Zhou, Shiyong Lan, Wenwu Wang 0001, Yixin Qiao
Neurocomputing3
2025 RemoteDPL: A Semi-Supervised Object Detector With Dense Pseudo-Labels for Remote Sensing
abstract
Deep learning-based object detection has seen substantial advancements, however, its practical deployment is often constrained by the need for large-scale labeled datasets. This limitation becomes even more critical in remote sensing imagery, where objects are densely distributed and exhibit significant scale variations. To address these challenges, we introduce RemoteDPL, a novel semi-supervised object detection (SSOD) framework that leverages dense pseudo-labeling (DPL) and multi-scale learning. RemoteDPL offers three key contributions. First, a fusion module is designed to dynamically integrate spatial and channel features across scales, improving detection across varied object sizes. Second, an instance density prediction branch is introduced to support pseudo-label mining, enhancing detection performance in densely populated regions. Lastly, we propose a two-stage pseudo-label filtering strategy that first selects "pending" class predictions and then refines them using a joint confidence score based on both classification and density information. Extensive experiments on the DOTA-v1.0 and NWPU datasets confirm the effectiveness of RemoteDPL, demonstrating its clear advantage over existing state-of-the-art (SOTA) semi-supervised object detection methods. On the NWPU dataset, RemoteDPL outperforms the SOTA baseline by +3.44%, +1.10%, and +1.62% under the settings of data labelled with 30%, 40%, and 50%, respectively, highlighting its strong capability in low-label remote sensing scenarios.
Yongjie Ma, Xinyuan Zhou, Shiyong Lan, Wenwu Wang 0001, Zicheng Sun, Yixin Qiao
IEEE Trans. Geosci. Remote. Sens.3
2024 Dense Pseudo-Labels based Semi-supervised Object Detection for Remote Sensing⋆
abstract
Deep-learning-based object detection has recently played an increasingly important role in analyzing geographic spatial information. However, object detection performance is strongly correlated with the quality and quantity of manually labeled data. Furthermore, object detection in Remote Sensing differs from natural scenes and faces two challenges: 1) Dense instance distribution and 2) Significant scale variations. These issues also contribute to the difficulty of manual annotation. To this end, based on a dense pseudo-labeling framework and multi-scale learning, this paper proposes a novel semi-supervised object detection (SSOD) framework for remote sensing, namely RemoteDPL. Firstly, a fusion module is proposed to adaptively integrate spatial and channel features of images at different scales, improving the detection of objects at various scales. Second, a specialized branch predicts instance density and aids in pseudo-label mining to enhance detection in dense scenarios. Finally, a two-stage filtering strategy is devised for pseudo-label mining, which first filters to obtain the "pending" class prediction boxes and then further filters this portion of prediction boxes to obtain the pseudo-labels according to the joint confidence based on the classification and density scores. Extensive experiments on DOTA-v1.0 have demonstrated that our proposed RemoteDPL surpasses the current state-of-the-art SSOD methods in various semi-supervised settings.
Yongjie Ma, Shiyong Lan, Xiaoxiao Yin
IJCNN2
2024 Visual Object Tracker Based on Video Temporal-Spatial Features and Long-term Memory
abstract
High-performance video object tracking is pivotal for video comprehension and analysis. There exists evident temporal correlation information among consecutive video frames. Nevertheless, current methods fail to effectively leverage this temporal information, leading to inaccurate feature representation of visual targets and heightened risks of tracking failure. To tackle this issue, we introduce a novel transformer-based tracker, dubbed the Video Temporal-Spatial Features and Long-term Memory (TSFLM) tracker. Firstly, the Encoder sequentially integrates multiple self-attention modules to extract spatial and temporal features, respectively. Secondly, we design a novel continuous template update module capable of preserving the long-term memory of the target template. Thirdly, we employ the long-term memory template to further augment the feature representation of the input image (search frame). Finally, the tracking results are derived through the decoder. Extensive comparative experiments against baselines on multiple challenging benchmarks demonstrate that our tracker achieves state-of-the-art performance. The source codes will be accessible at https://github.com/SYLan2019/TSFLM-Tracker.
Yongyang Gao, Shiyong Lan, Piaoyang Li
SMC2
2024 Dynamic Spatial Feature Enhancement for Local Climate Zone Classification in SAR and Multi-Spectral Data
abstract
Local Climate Zone (LCZ) classification from remote sensing images plays a crucial role in quantifying the urban heat island effect. However, the performance of LCZ classification has not been satisfactory so far, especially for built-up area categories. To alleviate this issue, we introduce a novel network architecture, DS-LCZNET, which incorporates a Dynamic Spatial Feature Enhancement (DSFE) module for capturing complex spatial information and a SAR-MS Fusion (SMF) module to improve feature integration from SAR and MS data. Extensive experiments demonstrate that DS-LCZNET significantly enhances classification performance, achieving a 3.55% increase in overall accuracy, a 1.18% improvement in average accuracy (AA), and a 3.88% rise in the kappa (x100) coefficient compared to the current leading baseline, MsF- LCZ-Net. The codes will be publicly available at: https://github.com/zhyilin97/DSLCZNET.
Yilin Zheng, Shiyong Lan, Guonan Deng
SMC2
2024 Semantic-aware normalizing flow with feature fusion for image anomaly detection
Yao Li 0019, Shiyong Lan, Wenwu Wang 0001, Weikang Huang, Wujiang Zhu
Neurocomputing3
2023 Siamese Network Based on MLP and Multi-head Cross Attention for Visual Object Tracking
Piaoyang Li, Shiyong Lan, Shipeng Sun, Wenwu Wang 0001, Yongyang Gao, Yongyu Yang, Guangyu Yu
ICANN (10)2
2023 GanNeXt: A New Convolutional GAN for Anomaly Detection
Bowei Pu, Shiyong Lan, Wenwu Wang 0001, Caiying Yang, Hongyu Yang 0002
ICANN (3)2
2023 Visual-Haptic-Kinesthetic Object Recognition with Multimodal Transformer
Xinyuan Zhou, Shiyong Lan, Wenwu Wang 0001, Hongyu Yang 0002
ICANN (7)2
2023 Flow-Based One-Class Anomaly Detection with Multi-Frequency Feature Fusion
abstract
Anomaly detection in computer vision seeks to identify samples outside of a predefined distribution, including texture defect detection and semantic anomaly detection. However, existing methods are difficult to simultaneously achieve high performance for both types of anomaly detection. To address this issue, we propose a new flow-based anomaly detection method. Firstly, we use semantic features extracted from a pre-trained backbone to learn the distribution of normal data from a semantic perspective. Secondly, we introduce a multi-frequency feature fusion module to aggregate semantic and texture information, which substantially improves performance for both types of anomaly detection at the same time. Extensive experiments on multiple well-known datasets demonstrate that our proposed method performs well in both types of anomaly detection, specially, achieves state-of-the-art performance in one-class anomaly detection. The codes will be available at https://github.com/SYLan2019/FOAD-MFFF.
Shiyong Lan, Weikang Huang, Yitong Ma, Hongyu Yang 0002, Yilin Zheng
ICIP2
2023 DLAHSD: Dynamic Label Adopted In Auxiliary Head for SAR Detection
abstract
Ship detection in synthetic aperture radar (SAR) images is a major issue in maritime surveillance and port management. Existing challenges are mainly as follows: (1) Tiny ships are mixed with scattered noise spots on the sea. (2) Ships are present in extreme aspect-ratios and various scales. (3) The land background blurs the outline of coastal ships. To address these problems, we propose an efficient detection neural network (DLAHSD) that integrates the Multi-scale Feature Location Fusion (MFLF) module and the Auxiliary Detection Head (ADH) based CenterNet. In addition, we designed a Dynamic Elliptic Gaussian (DEG) module to label the heatmap of ships. Experimental results on the challenging SSDD dataset show that our model offers improved performance over the baseline methods. The codes will be available at https://github.com/SYLan2019/DLAHSD.
Xiaoxiao Yin, Shiyong Lan, Weikang Huang, Yitong Ma, Wenwu Wang 0001, Hongyu Yang 0002, Yilin Zheng
ICIP2
2023 A Semantics-Aware Normalizing Flow Model for Anomaly Detection
abstract
Anomaly detection in computer vision aims to detect outliers from input image data. Examples include texture defect detection and semantic discrepancy detection. However, existing methods are limited in detecting both types of anomalies, especially for the latter. In this work, we propose a novel semantics-aware normalizing flow model to address the above challenges. First, we employ the semantic features extracted from a backbone network as the initial input of the normalizing flow model, which learns the mapping from the normal data to a normal distribution according to semantic attributes, thus enhances the discrimination of semantic anomaly detection. Second, we design a new feature fusion module in the normalizing flow model to integrate texture features and semantic features, which can substantially improve the fitting of the distribution function with input data, thus achieving improved performance for the detection of both types of anomalies. Extensive experiments on five well-known datasets for semantic anomaly detection show that the proposed method outperforms the state-of-the-art baselines. The codes will be available at https://github.com/SYLan2019/SANF-AD.
Shiyong Lan, Weikang Huang, Wenwu Wang 0001, Hongyu Yang 0002, Yitong Ma, Yongjie Ma
ICME2
2022 Face Super-Resolution with Spatial Attention Guided by Multiscale Receptive-Field Features
Weikang Huang, Shiyong Lan, Wenwu Wang 0001, Xuedong Yuan, Hongyu Yang 0002, Piaoyang Li
ICANN (1)2
2022 A Transformer-Based GAN for Anomaly Detection
Caiyin Yang, Shiyong Lan, Weikang Huang, Wenwu Wang 0001, Hongyu Yang 0002, Piaoyang Li
ICANN (2)2
2022 DSTAGNN: Dynamic Spatial-Temporal Aware Graph Neural Network for Traffic Flow Forecasting
abstract
As a typical problem in time series analysis, traffic flow prediction is one of the most important application fields of machine learning. However, achieving highly accurate traffic flow prediction is a challenging task, due to the presence of complex dynamic spatial-temporal dependencies within a road network. This paper proposes a novel Dynamic Spatial-Temporal Aware Graph Neural Network (DSTAGNN) to model the complex spatial-temporal interaction in road network. First, considering the fact that historical data carries intrinsic dynamic information about the spatial structure of road networks, we propose a new dynamic spatial-temporal aware graph based on a data-driven strategy to replace the pre-defined static graph usually used in traditional graph convolution. Second, we design a novel graph neural network architecture, which can not only represent dynamic spatial relevance among nodes with an improved multi-head attention mechanism, but also acquire the wide range of dynamic temporal dependency from multi-receptive field features via multi-scale gated convolution. Extensive experiments on real-world data sets demonstrate that our proposed method significantly outperforms the state-of-the-art methods.
Shiyong Lan, Yitong Ma, Weikang Huang, Wenwu Wang 0001, Hongyu Yang 0002, Pyang Li
ICML1
2021 Robust Visual Object Tracking with Spatiotemporal Regularisation and Discriminative Occlusion Deformation
abstract
Spatiotemporal regularized Discriminative Correlation Filters (DCF) have been proposed recently for visual tracking, achieving state-of-the-art performance. However, the tracking performance of the online learning model used in this kind methods is highly dependent on the quality of the appearance feature of the target, and the target feature appearance could be heavily deformed due to the occlusion by other objects or the variations in their dynamic self-appearance. In this paper, we propose a new approach to mitigate these two kinds of appearance deformation. Firstly, we embed the occlusion perception block into the model update stage, then we adaptively adjust the model update according to the situation of occlusion. Secondly, we use the relatively stable colour statistics to deal with the appearance shape changes in large targets, and compute the histogram response scores as a complementary part of final correlation response. Extensive experiments are performed on four well-known datasets, i.e. OTB100, VOT-2018, UAV123, and TC128. The results show that the proposed approach outperforms the baseline DCF method, especially, on the TC128/UAV123 datasets, with a gain of over 4.05 %/2.43% in mean overlap precision. We will release our code at https://github.com/SYLan2019/STD0D.
Shiyong Lan, Shipeng Sun, Wenwu Wang 0001
ICIP1
2021 SAGAN: Skip-Attention GAN For Anomaly Detection
abstract
Generative Adversarial Networks (GANs) have been used recently for anomaly detection from images, where the anomaly scores are obtained by comparing the global difference between the input and generated image. However, the anomalies often appear in local areas of an image scene, and ignoring such information can lead to unreliable detection of anomalies. In this paper, we propose an efficient anomaly detection network Skip-Attention GAN (SAGAN), which adds attention modules to capture local information to improve the accuracy of latent representation of images, and uses depth-wise separable convolutions to reduce the number of parameters in the model. We evaluate the proposed method on the CIFAR-10 dataset and the LBOT dataset (built by ourselves), and show that the performance of our method in terms of area under curve (AUC) on both datasets is improved by more than 10% on average, as compared with three recent baseline methods.
Shiyong Lan, Weikang Huang, Wenwu Wang 0001
ICIP2