Shuanglin Li

dblp:254/9190 · DBLP profile ↗
← Back
12ranked-venue papers
6as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Graph learning · 77% Face, body and person analysis · 23%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Medical and health informatics · 100%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Graph learning
graph neural network
0.912025
DepMGNN: Matrixial Graph Neural Network for Video-based Automatic Depression Assessment · AAAI 2025
Medical and health informatics › mental health informatics
depression assessment
0.912025
DepMGNN: Matrixial Graph Neural Network for Video-based Automatic Depression Assessment · AAAI 2025
Computer vision › Face, body and person analysis
facial behavior analysis
0.312025
DepMGNN: Matrixial Graph Neural Network for Video-based Automatic Depression Assessment · AAAI 2025

Methods — techniques the papers use, named apart from their topics

matrix-style edge features · 1.7graph-style data structure · 1.7
YearPublicationVenuePosition
2026 A preplanned-online rolling-horizon re-optimization method for post-disaster large-scale search-and-rescue under uncertain chance encounters
Shuanglin Li
Expert Syst. Appl.1
2026 Joint economic lot sizing shipment policy in multi-level integrated network supply chains of bottleneck material
Yang Sheng, Shuanglin Li
Expert Syst. Appl.2
2025 DepMGNN: Matrixial Graph Neural Network for Video-based Automatic Depression Assessment
abstract
Depression can be reflected by long-term human spatio-temporal facial behaviours. While human face videos recorded in real-world usually have long and variable lengths, existing video-based depression assessment approaches frequently re-sample/down-sample such videos to short and equal-length videos, or split each video into several equal-length segments, where segment-level spatio-temporal facial behaviours are suppressed as a vector-style representations for RNN-based long-term (video-level) modelling. Both strategies lead to crucial information loss and distortion. In this paper, we propose a novel graph-style data structure called Matrixial Graph and an effective Matrixial Graph Neural Network (MGNN) for face video-based depression assessment, which can directly and end-to-end model long-term depression-specific spatio-temporal facial cues from variable-length videos without resampling/splitting videos or suppressing video segments to vectors. Importantly, the nodes in our matrixial graph are capable of including matrices of different shapes, and thus nodes of a matrix graph can directly represent all frame-level 2D facial feature maps (or images themselves) of an entire video regardless of its length. Then, our MGNN is the first GNN that can jointly process matrixial graphs containing varying numbers of nodes, which further learns matrix-style edge features, thereby facilitating to explicit model video-level multi-scale spatio-temporal facial behaviours among matrixial graph nodes for depression assessment. Experiments show that the explicit spatio-temporal modeling on 2D facial feature maps, facilitated by our matrixial graph/MGNN, provided significant benefits, leading our approach to achieve new state-of-the-art performances on AVEC2013 and AVEC2014 datasets with large advantages.
Leijing Zhou, Shuanglin Li, Changzeng Fu, Jun Lu 0006, Jing Han 0009, Yi Zhang 0036, Siyang Song
AAAI3
2025 A Frequency-aware Augmentation Network for Mental Disorders Assessment from Audio
abstract
Depression and Attention Deficit Hyperactivity Disorder (ADHD) stand out as the common mental health challenges today. In affective computing, speech signals serve as effective biomarkers for mental disorder assessment. Current research, relying on labor-intensive hand-crafted features or simplistic time-frequency representations, often overlooks critical details by not accounting for the differential impacts of various frequency bands and temporal fluctuations. Therefore, we propose a frequency-aware augmentation network with dynamic convolution for depression and ADHD assessment. In the proposed method, the spectrogram is used as the input feature and adopts a multi-scale convolution to help the network focus on discriminative frequency bands related to mental disorders. A dynamic convolution is also designed to aggregate multiple convolution kernels dynamically based upon their attentions which are input-independent to capture dynamic information. Finally, a feature augmentation block is proposed to enhance the feature representation ability and make full use of the captured information. Experimental results on AVEC 2014 and self-recorded ADHD dataset prove the robustness of our method, an RMSE of 9.23 was attained for estimating depression severity, along with an accuracy of 89.8% in detecting ADHD.
Shuanglin Li, Siyang Song, Rajesh Nair, Syed M. Naqvi
ICASSP1
2025 Efficient Long Speech Sequence Modelling for Time-Domain Depression Level Estimation
abstract
Depression significantly affects emotions, thoughts, and daily activities. Recent research indicates that speech signals contain vital cues about depression, sparking interest in audiobased deep-learning methods for estimating its severity. However, most methods rely on time-frequency representations of speech which have recently been criticized for their limitations due to the loss of information when performing time-frequency projections, e.g. Fourier transform, and Mel-scale transformation. Furthermore, segmenting real-world speech into brief intervals risks losing critical interconnections between recordings. Additionally, such an approach may not adequately reflect real-world scenarios, as individuals with depression often pause and slow down in their conversations and interactions. Building on these observations, we present an efficient method for depression level estimation using long speech signals in the time domain. The proposed method leverages a state space model coupled with the dual-path structure-based long sequence modelling module and temporal external attention module to reconstruct and enhance the detection of depression-related cues hidden in the raw audio waveforms. Experimental results on the AVEC2013 and AVEC2014 datasets show promising results in capturing consequential long-sequence depression cues and demonstrate outstanding performance over the state-of-the-art.
Shuanglin Li, Zhijie Xie, Syed M. Naqvi
ICASSP1
2025 Analysis of Distributed Element-Level Resource Allocation Efficiency
abstract
The problem of efficient collaboration of cross-domain fragmented resources under communication constraints is analyzed. Firstly, the concept of element-level resource allocation under communication constraints is proposed. Secondly, the distributed resource allocation framework, algorithm and models are put forward. Finally, the efficiency of traditional centralized and distributed resource allocation under different communication conditions is quantitatively analyzed through simulation. The simulation results show that with the increase of communication miss rate, the resource allocation efficiency of distributed element-level is significantly improved compared with that of centralized mode, with an increase of about$\mathbf{1 4 \%}$in typical scenarios.
Chunlei Han, Shuanglin Li, Shihui Ji
ICPADS4
2025 A Multi-Scale Feature Refinement and Dual-Attention Enhanced Dynamic Convolutional Network for Speech-Based Depression and ADHD Assessment
abstract
In the area of affective computing, speech has been identified as a promising biomarker for assessing depression and attention deficit hyperactivity disorder (ADHD). These disorders manifest as abnormalities in speech across various frequency bands and exhibit temporal variations. Most existing work on speech features relies on the magnitude spectrogram, which discards phase information and also does not consider the impact of different frequency bands on depression and ADHD detection. Inspired by these, we propose a novel multi-scale complex feature refinement and dynamic convolution attention-aware network to enhance speech-based assessment of depression and ADHD. Our approach incorporates three key components: multi-scale complex feature refinement (MSFR), dynamic convolutional neural network (Dy-CNN), and dual-attention feature enhancement (DAFE) module. The MSFR module utilizes depth-wise convolutional networks to process both magnitude and phase input, selectively emphasizing frequency bands associated with depression and ADHD. Importantly, the Dy-CNN module employs an attention mechanism to autonomously generate multiple convolution kernels that adapt to input features and capture relevant temporal dynamics linked to depression and ADHD. Additionally, the DAFE module enhances feature representation and detection performance by incorporating channel shuffle attention (CSA) and spatial axial attention (SAA) mechanisms, which leverage both inter- and intra-channel relationships and examine time-frequency characteristics of the feature map. Extensive experiments conducted on four publicly available datasets, i.e., AVEC2013, AVEC2014, E-DAIC, and a self-collected authentic ADHD dataset demonstrated that the proposed method outperforms previous approaches and exhibits superior generalization capabilities across different language settings (i.e., English, German) for speech-based depression and ADHD assessment.
Shuanglin Li, Siyang Song, Syed M. Naqvi
IEEE Trans. Affect. Comput.1
2024 A Novel Audio-Visual Information Fusion System for Mental Disorders Detection
abstract
Mental disorders are among the foremost contributors to the global healthcare challenge. Research indicates that timely diagnosis and intervention are vital in treating various mental disorders. However, the early somatization symptoms of certain mental disorders may not be immediately evident, often resulting in their oversight and misdiagnosis. Additionally, the traditional diagnosis methods incur high time and cost. Deep learning methods based on fMRI and EEG have improved the efficiency of the mental disorder detection process. However, the cost of the equipment and trained staff are generally huge. Moreover, most systems are only trained for a specific mental disorder and are not general-purpose. Recently, physiological studies have shown that there are some speech and facial-related symptoms in a few mental disorders (e.g., depression and ADHD). In this paper, we focus on the emotional expression features of mental disorders and introduce a multimodal mental disorder diagnosis system based on audio-visual information input. Our proposed system is based on spatial-temporal attention networks and innovative uses a less computationally intensive pre-train audio recognition network to fine-tune the video recognition module for better results. We also apply the unified system for multiple mental disorders (ADHD and depression) for the first time. The proposed system achieves over 80% accuracy on the real multimodal ADHD dataset and achieves state-of-the-art results on the depression dataset AVEC 2014.
Yichun Li, Shuanglin Li, Syed M. Naqvi
FUSION2
2024 Ensemble multi-objective optimization approach for heterogeneous drone delivery problem
Xupeng Wen, Guohua Wu 0001, Shuanglin Li, Ling Wang 0001
Expert Syst. Appl.3
2023 Enhancing ADHD Detection Using Diva Interview-Based Audio Signals and A Two-Stream Network
abstract
Attention deficit hyperactivity disorder (ADHD) is a neurodevelopmental condition that results in altered behaviour in social development and communication patterns. However, due to the dearth of medical psychiatrists globally, the diagnosis of ADHD is frequently delayed. With the burgeoning development of artificial intelligence, it is rational to introduce deep learning to facilitate the ADHD diagnosis. Previous deep learning methods mainly use functional magnetic resonance imaging (fMRI) or Electroencephalography (EEG) signals to detect ADHD, where the data are expensive to acquire, i.e. equipment cost and specialised staff for data collection. Over the past years, speech signals have gained increasing attention owing to their cost-effectiveness in data collection and non-intrusive characteristics. In this work, based on the Diagnostic Interview for ADHD in adults (DIVA), we design a questionnaire and collected the audio data of ADHD patients and normal controls in collaboration with the Cumbria, Northumberland, Tyne and Wear NHS Foundation Trust. Besides, we propose a two-stream model (TSM) to exploit local and global features to assist ADHD detection. Applying the TSM to the collected real ADHD audio data, the performance of the proposed method is promising with an average accuracy of 84.9%.
Shuanglin Li, Yang Sun 0003, Rajesh Nair, Syed M. Naqvi
IPCCC1
2023 A novel min-max robust model for post-disaster relief kit assembly and distribution
Yarui Zhang, Shuanglin Li, Shuangyan Li
Expert Syst. Appl.3
2022 Post-Disaster Distribution System Restoration With Logistics Support and Geographical Characteristics
abstract
Repair scheduling and routing and logistics support are interdependent and critical for post-disaster distribution system restoration (PDSR), which is also influenced by the geographical characteristics of outage area. Hence, we develop a co-optimization model for the PDSR with logistics support and geographical characteristics. A hybrid improved bacterial colony chemotaxis algorithm is proposed to solve the model, in which A* algorithm is employed to route repair crews and material delivery in the transportation network considering geographical characteristics, and an improved bacterial colony chemotaxis algorithm is proposed to determine the repair scheduling and material allocation in the distribution system. Different scale of distribution system instances with different damage levels and different geographical characteristics are used to demonstrate the effectiveness of the proposed methodology.
Shuanglin Li, Zujun Ma, Tsan-Ming Choi
IEEE Trans. Intell. Transp. Syst.1