Jian Li 0063

dblp:33/5448-63 · DBLP profile ↗
← Back
13ranked-venue papers
9as first author
13since 2021 · last 2026
0000-0001-5880-2565ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 6 first-author · 10 since 2021Computer networks · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 A robust and interpretable framework for sports activity recognition based on wearable sensor signals and image representations
Jian Li 0063, Yibo Fan, Junhui Gong, Ruoyu Chen 0001, Yuliang Zhao
Eng. Appl. Artif. Intell.1
2025 Adaptive Deformable Convolutional Neural Network Framework for depression-related behavioral analysis in mice
Jian Li 0063, Xiaoyong Lyu, Yuliang Zhao
Eng. Appl. Artif. Intell.1
2025 CIR-DFENet: Incorporating cross-modal image representation and dual-stream feature enhanced network for activity recognition
Yuliang Zhao, Jin-Liang Shao, Xiru Lin, Tianang Sun, Jian Li 0063, Chao Lian, Xiaoyong Lyu, Binqiang Si, Zhikun Zhan
Expert Syst. Appl.5
2025 Disease and personality information enhanced depression detection based on the TransGCL framework
Yuliang Zhao, Jian Li 0063, Chao Lian, Kaixuan Tian, Changzeng Fu
Neurocomputing5
2025 An Intelligent Badminton Handle With Multinode MEMS Sensors for Explainable Motion Recognition
abstract
Intelligent sensing technologies are transforming sports training by enabling precise motion analysis, critical for skill development and performance optimization. This study introduces a badminton racket handle embedded with a lightweight, multi-node MEMS-based sensing system designed for real-time motion recognition. To capture distributed grip forces, swing trajectories, and impact mechanics at the player-equipment interface, the system employs an ergonomic design ensuring natural gameplay. A hybrid feature extraction approach, integrating time-and frequency-domain features with a 1D-CNN, achieves a classification accuracy of 97.89% across ten badminton actions. To enhance interpretability and provide actionable insights, explainable AI using SMDL-attribution identifies key motion features, revealing biomechanical inefficiencies in grip strength, swing consistency, and wrist motion. Seamlessly integrated with Virtual Reality (VR) platforms, the system delivers immersive, real-time feedback, transforming training into an interactive and data-driven experience. By combining advanced sensing, machine learning, and explainable AI, this system establishes a new benchmark for intelligent sports monitoring, with broad applications in sports training, rehabilitation, and human-computer interaction.
Jian Li 0063, Yibo Fan, Ruoyu Chen 0001, Siyuan Liang 0004, Yuliang Zhao
IEEE Internet Things J.1
2025 Dynamic Multisensor Fusion Framework With Adaptive Spatiotemporal Optimization for IoT-Based Motion Recognition
abstract
Motion recognition in IoT-based sensor systems is crucial for applications such as healthcare and human-computer interaction. However, one key challenge—data structure inconsistencies—complicates the performance of existing systems, particularly in dynamic real-world environments. Traditional fusion approaches lack the adaptability required to address sensor inconsistencies and fail to fully leverage the potential of multisensor data. To overcome this challenge, we propose a dynamic multisensor fusion framework (DMSFF) with adaptive spatiotemporal optimization. This framework introduces a dynamic sensor weighting mechanism that prioritizes reliable data while suppressing noise, ensuring robustnesss. A transformer-based fusion architecture captures spatiotemporal features, modeling complex intersensor relationships and long-term dependencies. Additionally, a motion kernel matching module aligns the data with canonical motion patterns, improving feature extraction and enhancing the recognition of subtle activities. The framework is validated on benchmark datasets, including those with real-world noise and structural inconsistencies, achieving an accuracy of 99.48%. This work establishes a new benchmark for multisensor motion recognition, providing scalable and robust solutions for smart healthcare and human-computer interaction.
Jian Li 0063, Yibo Fan, Xiaoyong Lyu, Yuliang Zhao
IEEE Internet Things J.1
2025 Multi-temporal image fusion empowered convolutional neural networks for recognition of 9 common mice actions
Jian Li 0063, Yuliang Zhao
Knowl. Based Syst.1
2025 MPRNet: A Temporal-Aware Cross-Modal Encoding Framework for Personality Recognition
abstract
Recent advances in personality recognition have improved trait inference from multimodal data, yet many existing methods rely on short-term video segments or static images, limiting the modeling of temporal dynamics due to short video durations, sparse frame-level annotations, and inconsistent modality coverage across audio, text, and visual channels. These limitations make it difficult to model how personality traits manifest over time and across modalities in naturalistic settings. To address these challenges, we introduce the Northeast University Personality Recognition (NEUPR) dataset, comprising 654 self-reported and discussion-based videos collected through MBTI assessments. NEUPR offers naturally expressed multimodal dataincluding audio, facial expressions, eye movements, and speech transcripts-captured across diverse participants and real-world settings. Building on this dataset, we propose MPRNet, a unified framework for dynamic personality recognition featuring two core innovations: (1) a multimodal encoder that leverages LSTM to capture temporal dependencies across longer sequences and integrates latent personality embeddings extracted from BERT representations of text to enrich semantic context, fused through adaptive weighting and enhanced by Gram encoding to preserve local feature patterns; and (2) a feature enhancement module that incorporates learnable positional encoding and channel attention to address modality imbalance and improve sensitivity to spatially salient features across modalities. Experimental results demonstrate that MPRNet outperforms state-of-the-art methods across multiple datasets, while ablation studies confirm the effectiveness of its components. By explicitly modeling temporal variation and enhancing cross-modal fusion, MPRNet enables more robust personality inference. This work establishes both a benchmark dataset and an adaptive modeling framework for multimodal personality analysis, advancing dynamic trait recognition.
Jian Li 0063, Junhui Gong, Shifeng Wang, Yuliang Zhao
IEEE Trans. Affect. Comput.1
2025 Sparse Emotion Dictionary and CWT Spectrogram Fusion With Multi-Head Self-Attention for Depression Recognition in Parkinson's Disease Patients
abstract
Depression is prevalent in patients with Parkinson's disease (PD), due to the dramatic negative impact that behavioral disorders have on daily life. Regrettably, most researchers in the past ignored the study of depression in PD patients, especially when depressive symptoms and PD symptoms are coupled together, it is difficult for researchers to recognize depression from the macro physiological signs of PD patients. Researchers are increasingly turning their attention to the subtle phenomena of emotional expression in conversation, using the textual and spectral features extracted from the audio of interviews as the primary support for understanding emotional states. However, there is still a lack of effective technical means to fuse these two features to recognize depression in PD patients. In this study, we proposed an innovative image fusion approach, fusing a sparse emotion dictionary with textual features and a Continuous Wavelet Transform (CWT) spectrogram with spectral features for the precise recognition of depression in PD patients. The fusion process integrates low-dimensional emotion-related textual cues, contributing to a more comprehensive extraction of emotionally relevant information. Subsequently, we introduce a High and Low Frequency Feature Fusion Multi-headed Self-Attention (HL-MSA) mechanism within a high and low frequency feature fusion network to amalgamate information across different frequency features within the images. The results underscore the efficacy of this novel fusion approach in effectively extracting depressive features in PD patients, attaining advanced recognition performance. Notably, this endeavor represents a pioneering stride in seamlessly fusing a sparse emotion dictionary and CWT spectrogram, exemplifying a promising and effective initiative for recognizing depression in PD patients.
Jian Li 0063, Yuliang Zhao, Yinghao Liu, Yuanyi Wu, Wanyue Wang
IEEE Trans. Affect. Comput.1
2025 Image Encoding and Fusion of Multi-Modal Data Enhance Depression Diagnosis in Parkinson's Disease Patients
abstract
The diagnosis of depression in individuals with Parkinson's Disease (PD) through the utilization of multimodal fusion techniques represents a significant domain. The primary challenge involves the creation of a robust fusion framework to address the heterogeneity among different modalities effectively. However, previous studies primarily focused on interactions between heterogeneous data, neglecting the structural similarities among isomorphic data, resulting in a substantial loss of feature information when merging heterogeneous data. In this study, we introduced a multi-modal data image encoding and fusion approach for diagnosing depression in PD patients. Additionally, we proposed a multi-modal dataset encompassing motion, facial expression, and audio data. First, we designed an RGB and sparse coding method to encode the multi-modal data, achieving the isomorphic transformation of multi-modal information and extracting feature information from lower-dimensional spaces. Furthermore, we introduced a Spatial-Temporal Network (STN) to fuse the three types of encoded images. We incorporated the Relation Global Attention (RGA) to enhance feature extraction and leverage all encoded image location feature nodes for balanced decision attention. Finally, recognizing the limitations of traditional machine learning algorithms in handling multi-tasks in medical diagnosis, we established a multi-task weighted loss function to achieve depression identification and severity prediction through Multi-Task learning (MTL).
Jian Li 0063, Yuliang Zhao, Wayne Jason Li, Changzeng Fu, Chao Lian
IEEE Trans. Affect. Comput.1
2025 Multimodal Depression Assessment Framework Integrating Personality and Gait for Older Adults With Medical Conditions
abstract
Elderly individuals often suffer from underlying medical conditions, resulting in a significant decline in quality of life and a heightened susceptibility to depression. Presently, AI screening tools based on behavioral indicators offer an objective and effective approach to diagnosing depression. However, current AI depression screening tools are primarily tailored to adolescents and adults, exhibiting shortcomings in their applicability and accuracy for elderly individuals with underlying medical conditions. To address the above issues, first, this paper constructs a depression dataset for elderly people with underlying diseases by using semi-structured interviews. Second, based on cognitive science insights, it is recognized that personality factors significantly influence behavioral expressions and also determine the attitudes of elderly individuals toward current life circumstances/health issues. Therefore, besides annotating depression severity, the Big Five-10 personality scale was utilized to annotate participant personalities. Finally, a late fusion-based multi-task learning framework was proposed, and the effects of introducing gait information and personality annotation on the performance of depression assessment were investigated. The experimental findings affirm the importance of integrating gait information and personality assessment in improving depression detection effectiveness. This study provides valuable foundational resources, as well as beneficial references and insights, for the research on depression in the elderly.
Yuliang Zhao, Jian Li 0063, Siyang Song, Chao Lian, Yinghao Liu, Changzeng Fu
IEEE Trans. Affect. Comput.3
2024 Global joint information extraction convolution neural network for Parkinson's disease diagnosis
Yuliang Zhao, Yinghao Liu, Jian Li 0063, Xiaoai Wang, Ruige Yang, Chao Lian, Zhikun Zhan, Changzeng Fu
Expert Syst. Appl.3
2023 RSMNet: A Robust Stacked Multiscale Feature Fusion Network for Visible RS Images
abstract
The aircrafts and ships in visible remote sensing (RS) images are of different scales. They are difficult to detect as they may be easily obscured by complex weather conditions such as snow and cloud. Therefore, it is important to eliminate the interference of complex weather conditions in order to detect these multi-scale objects accurately. This letter proposes an improved robust stacked multi-scale feature fusion network RSMNet to address this problem from two aspects. First, a stacked dilated convolution is used to enlarge the receptive fields of high-resolution images and improve the ability to extract multi-scale information. Second, the maps of extracted features are resized and integrated to refine the connection among different layers. Compared to the original Faster R-CNN model, RSMNet provides a 2% and 3.9% higher AP in the detection of aircrafts and ships, respectively. RSMNet also shows much more robust performance than the original model in detection under cloudy and snowy conditions.
Jian Li 0063, Ruige Yang, Yuliang Zhao, Xiaoai Wang, Lianjiang Li, Qiang Fu 0017
IEEE Geosci. Remote. Sens. Lett.1