Shiren Li

dblp:217/7163 · DBLP profile ↗
← Back
14ranked-venue papers
1as first author
11since 2021 · last 2026
0000-0002-8746-1787ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 4 · 3 since 2021
YearPublicationVenuePosition
2026 SFIF -Net: Spatial-Frequency Interactive Feature Learning for Medical Image Segmentation
abstract
ABSTRACT Accurate lesion segmentation in medical image analysis is critical for diagnosis and treatment planning. However, traditional U‐shaped architectures often struggle with large variation in lesion size and blurred boundaries. To address these challenges, a model called SFIF‐Net has been proposed in this study. In particular, SFIF‐Net strengthens feature interaction through four key components: a Hierarchical Feature Aggregation (HFA) module to enable cross‐layer feature fusion guidance; a Layer‐wise Feature Aggregation (LWFA) module in skip connections for dynamic multiscale fusion; an Interactive Feature Fusion (IFF) module equipped with a Spectral Feature Migration (SFM) component in the decoder to restore fine boundaries via spatial‐frequency fusion and a Multiscale Feature Enhancement (MFE) module applied across stages to improve multilevel feature learning. Experiments have been conducted on four public datasets, which are ISIC2018, BUSI, GlaS and CVC‐ClinicDB, and four metrics (Dice, mIoU, HD95 and Specificity) are used for evaluating the model performance. Experimental results show that SFIF‐Net outperforms other popular models. It achieves the highest average Dice score across all four datasets, outperforming the second‐best models by 1.02%, 1.22%, 0.06% and 0.06%, respectively. The source code is available at https://github.com/shen123shen/SFIF‐Net‐main .
Haozhou Shen, Shiren Li, Novalee Sayaxang, Maksim Davydov, Latifah Kamarudin, Guangguang Yang
Expert Syst. J. Knowl. Eng.2
2026 Enhancing medical image segmentation with the modification of U-shaped network
Shiren Li, Maksim Davydov, Serestina Viriri, Irsa Talib, Zhihao Yuan, Guangguang Yang
Vis. Comput.2
2026 Enhancing medical image segmentation with adaptive convolution and dynamic high-frequency feature enhancement
Wenguang Xu, Shiren Li, Kamoliddin Shukurov, Maksim Davydov, Jawad Hussain, Guangguang Yang
Vis. Comput.3
2025 PC-UNet: a pure convolutional UNet with channel shuffle average for medical image segmentation
Shiren Li, Yongliang Xiong, Guangguang Yang
Appl. Intell.3
2025 Cfseg-Net: context feature extraction network for medical image segmentation
Shiren Li, Yaoxue Lin, Sihua Tang, Wenguang Xu, Kangxian Chen, Guangguang Yang
Vis. Comput.2
2024 ESAformer: Enhanced Self-Attention for Automatic Speech Recognition
abstract
In this paper, an Enhanced Self-Attention (ESA) module has been put forward for feature extraction. The proposed ESA is integrated with the recursive gated convolution and self-attention mechanism. In particular, the former is used to capture multi-order feature interaction and the latter is for global feature extraction. In addition, the location of interest that is suitable for inserting the ESA is also worth being explored. In this paper, the ESA is embedded into the encoder layer of the Transformer network for automatic speech recognition (ASR) tasks, and this newly proposed model is named ESAformer. The effectiveness of the ESAformer has been validated using three datasets, that are Aishell-1, HKUST and WSJ. Experimental results show that, compared with the Transformer network, 0.8% CER, 1.2% CER and 0.7%/0.4% WER, improvement for these three mentioned datasets, respectively, can be achieved.
Zhikui Duan, Shiren Li, Xinmei Yu, Guangguang Yang
IEEE Signal Process. Lett.3
2023 LFEformer: Local Feature Enhancement Using Sliding Window With Deformability for Automatic Speech Recognition
abstract
A module using sliding window with deformablity, abbreviated as SWD, has been proposed for local feature enhancement. In particular, the proposed SWD module adopts windows with variable size based on the depth of the embedded network layers. Moreover, the proposed SWD module is inserted into the Transformer network, referred as LFEformer, for automatic speech recognition. Such network is particularly good at capturing both local and global features, and this is beneficial for model improvement. It is worth mentioning that the local and global features are extracted by SWD module and the attention mechanism in Transformer network, respectively. The effectiveness of the LFEformer has been validated on three widely used datasets, which are Aishell-1, HKUST and WSJ (dev93/eval92). The experimental results demonstrate that 0.5% CER, 0.8% CER and 0.7%/0.3% WER improvement can be obtained in the correspondent datasets.
Guangyong Wei, Zhikui Duan, Shiren Li, Xinmei Yu, Guangguang Yang
IEEE Signal Process. Lett.3
2023 2D arcsine and sine combined logistic map for image encryption
Zhikui Duan, Shiren Li
Vis. Comput.3
2022 Source-free unsupervised multi-source domain adaptation via proxy task for person re-identification
Zhikui Duan, Shiren Li
Vis. Comput.3
2021 Learning to locate for fine-grained image recognition
Jianguo Hu, Shiren Li
Comput. Vis. Image Underst.3
2021 Revisiting Hard Example for Action Recognition
abstract
Video-based action recognition, which needs to handle temporal motion and spatial cues simultaneously, remains a challenging task. In this paper, our motivation is to address this issue by fully utilizing temporal information. Specially, a novel light-weight Voting-based Temporal Correlation (VTC) module is proposed to enhance temporal information. Multiple branches with different temporal sampling intervals are included in this module and they are regarded as voters. The final classification result is “voted” by these branches together. VTC module integrates sparse temporal sampling strategy into feature sequences, so it mitigates the effect of redundant information and focuses more on temporal modeling. Additionally, we propose a simple and intuitive Similarity Loss (SL) to guide the training procedure of the VTC module and the backbone network. When we introduce confusion in the predicted vector intentionally, SL eases intra-class variation by discovering class-specific common motion patterns rather than sample-specific discriminative information. SL neither needs excessive parameter tuning during training nor adds significant computation overhead during test time. By combining VTC module and SL with complementary advances in the field, we clearly outperform state-of-the-art results and achieve 83.0, 98.4, 49.6 and 77.8 accuracy on HMDB51, UCF101, something-something-v1, and Kinetics respectively.
Jianguo Hu, Shiren Li, Zhihao Yuan
IEEE Trans. Circuits Syst. Video Technol.3
2020 Rethinking Temporal-Related Sample for Human Action Recognition
abstract
Temporal-related samples always have huge intra-class appearance variation, on which lots of existing action recognition algorithms have poor performance. In this paper, our motivation is to address this issue by utilizing temporal information more effectively. A novel light-weight Voting-based Temporal Correlation module (VTC) is proposed to enhance temporal cues. VTC integrates sparse temporal sampling strategy into feature sequences, so it mitigates the effect of redundant information and focuses more on temporal modeling. Furthermore, we propose a simple and intuitive Similarity Loss (SL) to guide the training procedure for VTC. Introducing confusion in the predicted vector intentionally, SL eases intra-class variation by discovering class-specific common motion pattern rather than sample-specific discriminative information. Combining VTC and SL with complementary advances in this field, we clearly outperform state-of-the-art results on HMDB51, UCF101, and Something-something-v1 dataset. The code has been made publicly available on https://github.com/FingerRec/TRS.
Shiren Li, Zhikui Duan, Zhihao Yuan
ICASSP2
2020 Combination of temporal-channels correlation information and bilinear feature for action recognition
abstract
In this study, the authors focus on improving the spatio–temporal representation ability of three‐dimensional (3D) convolutional neural networks (CNNs) in the video domain. They observe two unfavourable issues: (i) the convolutional filters only dedicate to learning local representation along input channels. Also they treat channel‐wise features equally, without emphasising the important features; (ii) traditional global average pooling layer only captures first‐order statistics, ignoring finer detail features useful for classification. To mitigate these problems, they proposed two modules to boost 3D CNNs’ performance, which are temporal‐channel correlation (TCC) and bilinear pooling module. The TCC module can capture the information of inter‐channel correlations over the temporal domain. Moreover, the TCC module generates channel‐wise dependencies, which can adaptively re‐weight the channel‐wise features. Therefore, the network can focus on learning important features. With regards to the bilinear pooling module, it can capture more complex second‐order statistics in deep features and generate a second‐order classification vector. We can get more accurate classification results by combining the first‐order and second‐order classification vector. Extensive experiments show that adding our proposed modules to I3D network could consistently improve the performance and outperform the state‐of‐the‐art methods. The code and models are available at https://github.com/caijh33/I3D_TCC_Bilinear .
Jiahui Cai, Jianguo Hu, Shiren Li, Jialing Lin
IET Comput. Vis.3
2018 Fast detection method of quick response code based on run-length coding
abstract
Quick response (QR) code, one of the two‐dimensional barcodes, is now being widely used in all fields. The effectiveness of decoding, however, needs to be improved in real‐time application. In most cases, the decoding procedure is time consuming, in which the detection of QR code plays an essential part. Therefore, this study proposes a fast detection method of QR code based on run‐length coding: firstly, a novel approach is proposed to detect the minimum region containing position detection pattern (PDP) in QR code. Second, coordinates of central PDP in QR code are calculated by using run‐length coding. The highlight in this step is the calculation, which utilises modified Knuth–Morris–Pratt algorithm. By this means, the computational complexity can be reduced tremendously. Finally, QR code can be detected successfully with the coordinates. The experimental results show that the proposed method is time saving and suitable for real‐time application.
Shiren Li, Jiayu Shang 0001, Zhikui Duan, Junwei Huang
IET Image Process.1