Shujie Ding

dblp:370/9723 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
6since 2021 · last 2025
0000-0001-6160-5765ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2025 LRDTN: Spectral-Spatial Convolutional Fusion Long-Range Dependence Transformer Network for Hyperspectral Image Classification
abstract
Recently, deep learning has achieved remarkable breakthroughs in hyperspectral image (HSI) classification tasks, particularly with methods based on convolutional neural networks (CNNs) and transformers. However, these methods have several limitations: 1) the limited receptive field inherent in the convolutional layer greatly hampers capturing feature contextual information on a large scale and 2) transformers cannot establish strong local relationships, making it challenging to characterize complex dependencies between distant pixels and different bands in HSIs. Moreover, as the network complexity increases, so does the number of network parameters. We propose a novel network called the spectral-spatial convolutional fusion long-range dependence transformer network (LRDTN) for HSI classification to address these challenges. LRDTN comprises three key components: dynamic-dependent convolutional (DDC) module, the multiscale enhanced fusion (MsEF) module, and the local–global perception transformer (LGPT). Specifically, the DDC dynamically models local features, while the MsEF integrates information from different scales to capture contextual relationships in HSI features effectively. Additionally, the ability to mine and utilize HSI local-global features and complex long-range dependencies is enhanced by the proposed transformer variant, LGPT. Ultimately, through the ingeniously designed structure of the LRDTN, the model effectively maintains its performance while reducing the number of network parameters. Extensive experiments conducted on four typical HSI datasets, including urban areas, agricultural areas, and swamps, demonstrate the superiority of LRDTN over other state-of-the-art networks. The code is available athttps://github.com/ybyangjing/LRDTN.
Shujie Ding, Xiaoli Ruan, Jing Yang 0017, Chengjiang Li, Jie Sun 0033, Xianghong Tang, Zhidong Su
IEEE Trans. Geosci. Remote. Sens.1
2024 Covid-IRLNet: A COVID-19 Diagnostic Model For Extracting CT Image Features and CT Sequence Features
abstract
At the end of 2019, the COVID-19 outbreak emerged abruptly. Chinese health authorities highlighted the role of CT scans, X-rays, and other computerized lung imaging in aiding COVID-19 diagnosis. This study aims to develop a computer-based system to assist healthcare professionals in diagnosing COVID-19 infections based on computerized imaging analysis. This approach aims to alleviate the workload of COVID-19 specialists, improving diagnostic and treatment efficiency and allowing specialists to focus on devising appropriate patient care plans promptly. The proposed method focuses on analyzing COVID-19 lesion characteristics within individual CT slices and their serial characteristics across CT sequences. This approach mirrors the diagnostic process of radiologists closely. To validate our model, we compiled a dataset from real medical diagnostic settings, minimizing the impact of lesion-like artifacts. We conducted a series of comparative and ablation experiments to evaluate the model's performance. Results indicate that our model outperforms the classic classification models and other commonly used models for COVID-19 diagnosis on our constructed dataset.
Jingxiang Xu, Jianqiang Li 0002, Linna Zhao, Shujie Ding
COMPSAC5
2024 A Progressive Soft Erase and Multi-Scale Feature Fusion Method for Medical Education Management of Breast Cancer
abstract
Breast ultrasound (BUS) image lesion segmentation is an important step in computer-aided diagnosis (CAD) systems for breast cancer screening. Due to acquiring pixel-level labels for fully supervised BUS lesion segmentation being extremely expensive and time-consuming, many studies have adopted weakly supervised methods to mitigate the reliance of models on pixel-level labels. However, in weakly supervised methods, the class activation map (CAM) often excessively focuses on the most distinctive regions of the target while overlooking other lesion parts. This limitation often leads to insufficient CAM activation and compromises the accuracy of lesion segmentation. To address this problem, we propose a weakly supervised framework for lesion segmentation on BUS images using image-level labels. Specifically, the method consists of two stages: classification and segmentation. During the classification stage, the network based on the progressive soft erase (PSE) module reduces the contribution of the most discriminative feature region in CAM, making the network focus on the features of non-dominant regions and generating more high-quality pseudo-masks. In the segmentation stage, we propose a dual-branch feature fusion (DBFF) network, which can not only capture more contextual information by fusing from different scales but also generate a more accurate segmentation mask. Extensive experiments on the publicly available dataset BUS I show the effectiveness of our method. Furthermore, this model significantly supports medical education management by providing a more efficient and cost-effective approach to breast lesion segmentation.
Jianqiang Li 0002, Shujie Ding, Linna Zhao, Zhaolei Liu
SMC5
2024 MGS-Net: Fusing Global and Local Feature Enhancements for Healthcare Education Management of Myasthenia Gravis Using Speech Data
abstract
Myasthenia gravis (MG) is a neurological disease that is difficult to diagnose and requires long-term management. The progression of this disease is reflected to some extent in changes in speech, such as hoarseness and articulation disorders. However, it is difficult for general neurologists to grasp the diagnostic patterns of such rare diseases, especially in underdeveloped regions. As an emerging field, speech-based intelligent diagnostic assistance provides a safe, non-invasive, and convenient solution for healthcare education management. To this end, we firstly constructed a novel Chinese speech dataset of myasthenia gravis patients (MGCS). Then we proposed a network named Myasthenia Gravis Speech Net (MGS-Net) for the classification of myasthenia gravis pathological speech, which is mainly composed of two blocks: the Local Feature Enhancement (LFE) block and the Feedforward Dense (FFD) block. The LFE block extracts temporal local features using a sliding window approach, while the FFD block captures the global representation of the data. Compared to existing methods, our pipeline achieves an accuracy of 98.75% and a recall rate of 99.17%. We validated the effectiveness of existing acoustic feature sets in pathological speech classification of MG, which will provide an important tool for health education management of neurological diseases.
Jianqiang Li 0002, Jingchen Zou, Yuning Huang, Shujie Ding, Linna Zhao
SMC5
2024 Sign Language Recognition and Translation Methods Promote Sign Language Education: A Review
abstract
Sign language recognition and translation (SLRT) aims to convert sign language into textual representation, which holds significant importance for the deaf community. Sign language possesses complex and diverse grammatical structures, with each sign language having distinct motion trajectories and gesture variations, making SLRT a complex research domain. In recent years, numerous researchers have proposed differ-ent modeling approaches, achieving significant advancements through the utilization of large language models. In this survey, we systematically review the developmental trajectory of SLRT, encompassing an introduction to key technical approaches at each stage and the latest research progress. Through a comprehensive examination of these methods, valuable insights are provided for future research and practical applications. Lastly, we identify the existing limitations of current methods and propose potential avenues for future research.
Jingchen Zou, Jianqiang Li 0002, Yuning Huang, Shujie Ding
SMC5
2024 LCTCS: Low-Cost and Two-Channel Sparse Network for Hyperspectral Image Classification
abstract
Abstract Using convolutional neural networks (CNNs) in classifying hyperspectral images (HSIs) has achieved quite good results in recent years. It is widely used in agricultural remote sensing, geological exploration, environmental monitoring, and marine remote sensing. Unfortunately, the complexity of network structures used for hyperspectral image classification challenges the efficient delivery of HSI data extremely, and existing methods suffer from a large amount of redundancy in the network weight parameters during training, as they either require huge computational resources or make inefficient use of storage space when designing the network structure, and many of the parameters that waste computational resources contribute less to the rich spectral and spatial information transfer in HSI. So we introduce LCTCS, a better low-memory and less-parametric network approach. LCTCS aims to improve the efficiency of computational resource utilization with advanced classification performance and lower levels of computational resources. Unlike the conventional 2D and 3D convolution used previously, we use simple and efficient 3D grouped convolution as a vehicle to convey the semantic features of HSIs. More specifically, we design a novel two-channel sparse network to classify HSIs since grouped 3D convolution conveys the properties of hyperspectral data well in the time and space domains.We have compared LCTCS with eight widely used network methods on four publicly available hyperspectral datasets for learning HSI information. A series of experiments shows that the model architecture designed has $$65.89 \%$$ 65.89 % less storage space than the DBDA method, consumes $$67.36 \%$$ 67.36 % fewer computational resources than the SSRN method on the IP dataset, and accomplishes a highly accurate classification task with the number of parameters accounting for only $$1.99 \%$$ 1.99 % that of the DBMA method.
Jie Sun 0033, Jing Yang 0017, Shujie Ding, Shaobo Li 0001, Jianjun Hu
Neural Process. Lett.4