Yingying Jiao

dblp:150/4189 · DBLP profile ↗
← Back
24ranked-venue papers
10as first author
20since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 5 first-author · 13 since 2021Artificial intelligence and machine learning · 11 · 4 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 first-author · 3 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Security and privacy · 2 · 2 since 2021
YearPublicationVenuePosition
2026 DiffusionPose: Markov-Optimized Diffusion Model for Human Pose Estimation
abstract
Video-based human pose estimation has long been a nontrivial task due to its dynamic nature and challenging detection scenarios such as occlusion and defocus. Inspired by the success of diffusion models, researchers have applied them to video pose estimation, outperforming traditional joint detection methods. However, existing diffusion model-based methods still face challenges like slow convergence and unstable pose generation. To tackle these issues, we propose DiffusionPose, a novel framework for video pose estimation that integrates diffusion models with optimization strategies: (1) We combine the emerging Mamba with Transformers to balance global and local spatio-temporal modeling. (2) We integrate Markov Random Fields into the reverse diffusion process to enhance the denoising of pose heatmaps, particularly addressing the issue of confused generation of occluded joints. (3) We mathematically formulate a Markov objective to supervise the heatmap denoising process, enabling the model to generate anatomically plausible skeletons. Our method achieves state-of-the-art performance on three large-scale benchmark datasets. Interestingly, it shows surprising robustness in challenging video scenarios, improving the accuracy of the most difficult ankle joint by 16.9% compared to the previous best diffusion model-based method on the Challenging-PoseTrack dataset.
Zhenguang Liu, Shaojing Fan, Sifan Wu 0001, Yingying Jiao
AAAI5
2026 Dual Coding Theory in Action: Language-Assisted Human Pose Estimation in Videos
abstract
Video-based human pose estimation aims to localize keypoints across frames, enabling robust analysis of human motion in applications such as sports, surveillance, and healthcare. However, existing methods rely solely on visual cues, limiting their robustness in complex scenes involving occlusion, motion blur, or poor lighting. In contrast, dual coding theory from psychology suggests that human cognition is inherently multimodal: we learn by integrating visual perception with linguistic context to form structured, semantic understandings of the world. Visual input provides concrete spatiotemporal grounding, while language offers symbolic abstraction that enhances reasoning and generalization. Motivated by this cognitive principle, we present the first framework that explicitly incorporates language as an auxiliary modality to enhance video-based pose estimation. To address the lack of paired video-text datasets, we first employ a Multimodal Large Language Model (MLLM) to generate textual descriptions of human interactions from videos. We then propose a novel coarse-to-fine multimodal alignment pipeline: a cross-modal semantic interaction module establishes initial grounding between spatiotemporal visual features and textual embeddings, while an optimal transport-based feature matching mechanism enforces fine-grained, geometry-aware alignment. This cognitively inspired design enables more accurate and robust pose estimation, especially in visually challenging scenes like occlusion and motion blur. Extensive experiments on three benchmarks confirm that our method consistently outperforms state-of-the-art approaches.
Sifan Wu 0001, Haipeng Chen 0002, Yingda Lyu, Shaojing Fan, Zhenguang Liu, Yingying Jiao
AAAI7
2026 Attentive Keypoint Identification: Progressive Spatiotemporal Refinement for Video-based Human Pose Estimation
abstract
Video-based human pose estimation has vast applications such as action recognition, sports analytics, and crime detection. However, this task is challenging as it involves interpreting both spatial context and temporal dynamics to accurately localize human anatomical keypoints in video sequences. Current approaches, often based on attention mechanisms, perform well but struggle in challenging scenarios like rapid motion and pose occlusion. We attribute these failures to two fundamental limitations: spatial uniformity, where models indiscriminately assign attention to both joint-relevant features and background clutter, thereby introducing spatial noise; and temporal rigidity, an inability to adapt to large joint displacements, resulting in severe feature misalignment during rapid motion. To overcome these challenges, we introduce PSTPose, a novel progressive spatiotemporal refinement framework. Specifically, to address the spatial uniformity problem, we propose a Discriminative Feature Enhancement (DFE) module that emphasizes joint-relevant features and a Feature Cluster Grouping (FCG) module that forms compact, semantically meaningful regions. For the temporal rigidity problem, we introduce a Deformable Spatiotemporal Fusion (DSF) module that adaptively aligns features across consecutive frames via deformation-aware sampling. This design ensures robust keypoint localization, particularly in cluttered and dynamic scenes. Extensive experiments on three large-scale benchmarks, PoseTrack2017, PoseTrack2018, PoseTrack21, demonstrate that PSTPose establishes a new state-of-the-art.
Sifan Wu 0001, Haipeng Chen 0002, Yingda Lyu, Shaojing Fan, Zhenguang Liu, Yingying Jiao
AAAI7
2026 Unveiling Deepfakes: A Frequency-Aware Triple Branch Network for Deepfake Detection
abstract
Advanced deepfake technologies are blurring the lines between real and fake, presenting both revolutionary opportunities and alarming threats. While it unlocks novel applications in fields like entertainment and education, its malicious use has sparked urgent ethical and societal concerns ranging fromidentity theftto thedissemination of misinformation. To tackle these challenges, feature analysis using frequency features has emerged as a promising direction for deepfake detection. However, one aspect that has been overlooked so far is that existing methods tend to concentrate on one or a few specific frequency domains, which risks overfitting to particular artifacts and significantly undermines their robustness when facing diverse forgery patterns. Another underexplored aspect we observe is that different features often attend to the same forged region, resulting in redundant feature representations and limiting the diversity of the extracted clues. This may undermine the ability of a model to capture complementary information across different facets, thereby compromising its generalization capability to diverse manipulations. In this paper, we seek to tackle these challenges from two aspects: (1) we propose a triple-branch network that jointly captures spatial and frequency features by learning from both original image and image reconstructed by different frequency channels, and (2) we mathematically derive feature decoupling and fusion losses grounded in the mutual information theory, which enhances the model to focus on task-relevant features across the original image and the image reconstructed by different frequency channels. Extensive experiments onsixlarge-scale benchmark datasets demonstrate that our method consistently achieves state-of-the-art performance. Our code is released athttps://github.com/injooker/Unveiling_Deepfake.
Qihao Shen, Jiaxing Xuan, Zhenguang Liu, Sifan Wu 0001, Yutong Xie 0019, Zhaoyan Ming, Yingying Jiao, Kui Ren 0001
IEEE Trans. Dependable Secur. Comput.7
2026 Rethinking Skeleton-Based Action Recognition From Action-Class Prediction Distribution Perspective
abstract
Action recognition has long been a fundamental and compelling problem in the field of computer vision. However, one aspect that has been overlooked so far is that current action recognition approaches often produce an unfavourable multi-peaked distribution when identifying the action class of a given motion sequence, which is ambiguous and hard to learn for neural networks. Moreover, current methods heavily rely on neural networks to extract action features for differentiating actions, lacking theoretical constraints ensuring that action-specific features are selectively extracted and ambiguous features common to multiple actions are effectively reduced. These shortcomings culminate in inadequate action recognition accuracy. Motivated by this, in this paper we seek to tackle the problem from three aspects: 1) We try to eliminate ambiguity by enforcing a smooth single-peaked distribution instead of a multi-peaked one for action-class prediction. 2) We theoretically analyze the lower bound of the label prediction log-likelihood and derive a training objective, which focuses on the extraction of action-specific features and the reduction of ambiguous features. 3) We further advocate feeding the model with richer information, including positive information like body-part structures and negative information like masked inputs. Empirically, our approach sets the new state-of-the-art performance on five large-scale benchmarks. Our code is released at https://github.com/ActionR-Group/DPM to facilitate future research.
Yingying Jiao, Haipeng Chen 0002, Yingda Lyu, Shuang Wu 0002, Zhenguang Liu
IEEE Trans. Image Process.1
2026 Detecting Driver Sleepiness From Physiological Indicators Using a CNN-LSTM Self-Attention Model
abstract
Sleepiness at the wheel is an important factor contributing to road traffic accidents. Based on the characteristic changes in Electroencephalography (EEG) and Electrooculography (EOG) signals, a dozing state is refined into three sub-states: the onset, duration, and end state. Each state is characterized by different physiological indicators such as the EEG alpha waves, the rising edge, and falling edge waveforms in EOG signals. To enable real-time detection of these physiological indicators, we propose a framework integrating three Convolutional Neural Network-Long Short-Term Memory-Self-Attention (CLSA) models, which combine CNN-based local feature extraction with self-attention mechanism for global context capture. The framework is evaluated for performance on continuous test data from 12 subjects. Our results demonstrate that by detecting alpha waves and the rising edge waveform, the alpha wave epoch (AWE) at the onset of the dozing state can be identified with high accuracy and precision. Thus, the onset sub-state is calculated as the period from the start time of the rising edge waveform to the time when the AWE is valid. Subsequently, the duration sub-state corresponds to the sustained presence of alpha waves. Furthermore, the falling edge waveform is detected with high accuracy, enabling the classification of the end state into two distinct phenomena: alpha blocking phenomenon or alpha wave attenuation-disappearance phenomenon, representing the sleepiness level-relaxed wakefulness or sleep onset, respectively. Utilizing three-channel signal processing, this framework provides a promising approach for real-time sleepiness detection in real-world driving scenarios.
Yingying Jiao
IEEE J. Biomed. Health Informatics1
2025 Causal-Inspired Multitask Learning for Video-Based Human Pose Estimation
abstract
Video-based human pose estimation has long been a fundamental yet challenging problem in computer vision. Previous studies focus on spatio-temporal modeling through the enhancement of architecture design and optimization strategies. However, they overlook the causal relationships in the joints, leading to models that may be overly tailored and thus estimate poorly to challenging scenes. Therefore, adequate causal reasoning capability, coupled with good interpretability of model, are both indispensable and prerequisite for achieving reliable results. In this paper, we pioneer a causal perspective on pose estimation and introduce a causal-inspired multitask learning framework, consisting of two stages. In the first stage, we try to endow the model with causal spatio-temporal modeling ability by introducing two self-supervision auxiliary tasks. Specifically, these auxiliary tasks enable the network to infer challenging keypoints based on observed keypoint information, thereby imbuing causal reasoning capabilities into the model and making it robust to challenging scenes. In the second stage, we argue that not all feature tokens contribute equally to pose estimation. Prioritizing causal (keypoint-relevant) tokens is crucial to achieve reliable results, which could improve the interpretability of the model. To this end, we propose a Token Causal Importance Selection module to identify the causal tokens and non-causal tokens (e.g., background and objects). Additionally, non-causal tokens could provide potentially beneficial cues but may be redundant. We further introduce a non-causal tokens clustering module to merge the similar non-causal tokens. Extensive experiments show that our method outperforms state-of-the-art methods on three large-scale benchmark datasets.
Haipeng Chen 0002, Sifan Wu 0001, Yifang Yin, Yingying Jiao, Yingda Lyu, Zhenguang Liu
AAAI5
2025 Optimizing Human Pose Estimation Through Focused Human and Joint Regions
abstract
Human pose estimation has given rise to a broad spectrum of novel and compelling applications, including action recognition, sports analysis, as well as surveillance. However, accurate video pose estimation remains an open challenge. One aspect that has been overlooked so far is that existing methods learn motion clues from all pixels rather than focusing on the target human body, making them easily misled and disrupted by unimportant information such as background changes or movements of other people. Additionally, while the current Transformer-based pose estimation methods has demonstrated impressive performance with global modeling, they struggle with local context perception and precise positional identification. In this paper, we try to tackle these challenges from three aspects: (1) We propose a bilayer Human-Keypoint Mask module that performs coarse-to-fine visual token refinement, which gradually zooms in on the target human body and keypoints while masking out unimportant figure regions. (2) We further introduce a novel deformable cross attention mechanism and a bidirectional separation strategy to adaptively aggregate spatial and temporal motion clues from constrained surrounding contexts. (3) We mathematically formulate the deformable cross attention, constraining that the model focuses solely on the regions centered at the target person body. Empirically, our method achieves state-of-the-art performance on three large-scale benchmark datasets. A remarkable highlight is that our method achieves an 84.8 mean Average Precision (mAP) on the challenging wrist joint, which significantly outperforms the 81.5 mAP achieved by the current state-of-the-art method on the PoseTrack2017 dataset.
Yingying Jiao, Zhenguang Liu, Shaojing Fan, Sifan Wu 0001, Zheqi Wu, Zhuoyue Xu
AAAI1
2025 SpatioTemporal Learning for Human Pose Estimation in Sparsely-Labeled Videos
abstract
Human pose estimation in videos remains a challenge, largely due to the reliance on extensive manual annotation of large datasets, which is expensive and labor-intensive. Furthermore, existing approaches often struggle to capture long-range temporal dependencies and overlook the complementary relationship between temporal pose heatmaps and visual features. To address these limitations, we introduce STDPose, a novel framework that enhances human pose estimation by learning spatiotemporal dynamics in sparsely-labeled videos. STDPose incorporates two key innovations: 1) A novel Dynamic-Aware Mask to capture long-range motion context, allowing for a nuanced understanding of pose changes. 2) A system for encoding and aggregating spatiotemporal representations and motion dynamics to effectively model spatiotemporal relationships, improving the accuracy and robustness of pose estimation. STDPose establishes a new performance benchmark for both video pose propagation (i.e., propagating pose annotations from labeled frames to unlabeled frames) and pose estimation tasks, across three large-scale evaluation datasets. Additionally, utilizing pseudo-labels generated by pose propagation, STDPose achieves competitive performance with only 26.7% labeled data.
Yingying Jiao, Sifan Wu 0001, Shaojing Fan, Zhenguang Liu, Zhuoyue Xu, Zheqi Wu
AAAI1
2025 HVIS: A Human-like Vision and Inference System for Human Motion Prediction
abstract
Grasping the intricacies of human motion, which involve perceiving spatio-temporal dependence and multi-scale effects, is essential for predicting human motion. While humans inherently possess the requisite skills to navigate this issue, it proves to be markedly more challenging for machines to emulate. To bridge the gap, we propose the Human-like Vision and Inference System (HVIS) for human motion prediction, which is designed to emulate human observation and forecast future movements. HVIS comprises two components: the human-like vision encode (HVE) module and the human-like motion inference (HMI) module. The HVE module mimics and refines the human visual process, incorporating a retina-analog component that captures spatiotemporal information separately to avoid unnecessary crosstalk. Additionally, a visual cortex-analogy component is designed to hierarchically extract and treat complex motion features, focusing on both global and local features of human poses. The HMI is employed to simulate the multi-stage learning model of the human brain. The spontaneous learning network simulates the neuronal fracture generation process for the adversarial generation of future motions. Subsequently, the deliberate learning network is optimized for hard-to-train joints to prevent misleading learning. Experimental results demonstrate that our method achieves new state-of-the-art performance, significantly outperforming existing methods by 19.8 % on Human3.6M, 15.7 % on CMU Mocap, and 11.1 % on G3D.
Kedi Lyu, Haipeng Chen 0003, Zhenguang Liu, Yifang Yin, Yukang Lin, Yingying Jiao
AAAI6
2025 CNN-Transformer Based Real-Time Detection of Driver Sleep Onset Using EEG and EOG Signals
abstract
Falling asleep at the wheel is a significant contributing factor to road traffic accidents. In simulated driving experiments, subjects exhibited frequent eye-closure episodes, which were accompanied by alpha blocking or alpha wave attenuation in electroencephalographic (EEG) signals. Based on the observation and analysis of 10 datasets, we defined two physiological markers that can indicate driver sleep onset: (1) the rising-edge waveform in vertical electrooculographic (VEOG) signals caused by eyelid closure behavior, and (2) the subsequent emergence of EEG alpha waves. A novel Convolutional Neural Network (CNN)-Transformer framework was proposed for realtime sleep onset detection. This framwork transforms the detection of the rising-edge waveform and EEG alpha waves into separate binary classification tasks, each solved by a dedicated CNN-Transformer model. The experimental results demonstrate that the sleep onset detection based on these two physiological markers achieves an average precision of 99.5 % and a sensitivity of$\mathbf{9 6. 2} \%$across all subjects, confirming the superior performance of the proposed framework. This framework leverages onechannel EEG and two-channel VEOG signals to enable real-time sleep onset detection in real-world driving scenarios, offering a promising practical solution.
Yingying Jiao
BIBM1
2025 Multi-Grained Feature Pruning for Video-Based Human Pose Estimation
abstract
Human pose estimation, with its broad applications in action recognition and motion capture, has experienced significant advancements. However, current Transformer-based methods for video pose estimation often face challenges in managing redundant temporal information and achieving fine-grained perception because they only focus on processing low-resolution features. To address these challenges, we propose a novel multi-scale resolution framework that encodes spatiotemporal representations at varying granularities and executes fine-grained perception compensation. Furthermore, we employ a density peaks clustering method to dynamically identify and prioritize tokens that offer important semantic information. This strategy effectively prunes redundant feature tokens, especially those arising from multi-frame features, thereby optimizing computational efficiency without sacrificing semantic richness. Empirically, it sets new benchmarks for both performance and efficiency on three large-scale datasets. Our method achieves a 93.8% improvement in inference speed compared to the baseline, while also enhancing pose estimation accuracy, reaching 87.4 mAP on the PoseTrack2017 dataset.
Shaojing Fan, Zhenguang Liu, Zheqi Wu, Sifan Wu 0001, Yingying Jiao
ICASSP6
2025 Enhancing Human Pose Estimation in Internet of Things via Diffusion Generative Models
abstract
With the ongoing development of public video surveillance technology, accurate human pose estimation is becoming increasingly important in urban administration and law enforcement. However, existing methods rely on large-scale dense annotations, which are labor-intensive and time-consuming. To tackle this, we propose SparsePose which leverages training videos with sparse annotations (labeled every k frames) to learn to propagate temporal poses that help to estimate the poses in unlabeled frames. Technically, we engage in a novel dual-branch architecture that combines 1) pose forecasting of the consecutive neighboring frames with 2) visual clues of the current frame and the nearest labeled frames. We theoretically derive the intrabranch and interbranch mutual information loss to supervise that maximized pose-relevant features are extracted from the current frame and different branches complement each other to approach precise pose estimation. Additionally, we propose a diffusion generative enhancement, which improves the robustness of the model to challenging scenes from the perspective of diversity. Empirical results show that our method significantly outperforms the state-of-the-art methods in sparsely labeled pose estimation on three benchmark datasets.
Sifan Wu 0001, Hongzhe Zhang, Zhenguang Liu, Haipeng Chen 0002, Yingying Jiao
IEEE Internet Things J.5
2025 Dual Space Representation Learning for Skeleton-Based Action Recognition
abstract
Skeleton-based action recognition is crucial for machine intelligence. Current methods generally learn from 3D articulated motion sequences in the straightforward Euclidean space. Yet, thevanillaEuclidean space may not be the optimal choice for modeling the intricate correlations among human body joints. This challenge arises from the non-Euclidean nature of human anatomy, where joint correlations often vary non-linearly during movement. To address this, we propose a dual space representation learning method. Specifically, we represent the motion sequences in Hyperbolic space, leveraging its intrinsic properties to capture the non-Euclidean latent anatomy of human motions. We then incorporate the motion features from both Hyperbolic and Euclidean spaces, allowing us to precisely model the non-linear joint correlations while effectively sketching human poses. The proposed method empirically achieves state-of-the-art performance on the NTU RGB+D 60, NTURGB+D 120, and NW-UCLA datasets.
Haipeng Chen 0002, Zhenguang Liu, Sihao Hu, Yingying Jiao
IEEE Signal Process. Lett.5
2025 I Know Who Clones Your Code: Interpretable Smart Contract Similarity Detection
abstract
Widespread reuse of open-source code in smart contract development boosts programming efficiency but significantly amplifies bug propagation across contracts, while dedicated methods for detecting similar smart contract functions remain very limited. Conventional abstract-syntax-tree (AST) based methods for smart contract similarity detection face challenges in handling intricate tree structures, which impedes detailed semantic comparison of code. Recent deep-learning based approaches tend to overlook code syntax and detection interpretability, resulting in suboptimal performance. To fill this research gap, we introduceSmartDetector, a novel approach for computing similarity between smart contract functions, explainable at the fine-grained statement level. Technically,SmartDetectordecomposes the AST of a smart contract function into a series of smaller statement trees, each reflecting a structural element of the source code. Then,SmartDetectoruses a classifier to compute the similarity score of two functions by comparing each pair of their statement trees. To address the infinite hyperparameter space of the classifier, we mathematically derive a cosine-wise diffusion process to efficiently search optimal hyperparameters. Extensive experiments conducted on three large real-world datasets demonstrate thatSmartDetectoroutperforms current state-of-the-art methods by an average improvement of 14.01% in F1-score, achieving an overall average F1-score of 95.88%.
Zhenguang Liu, Lixun Ma, Zhongzheng Mu, Chengkun Wei, Yingying Jiao, Kui Ren 0001
IEEE Trans. Dependable Secur. Comput.6
2024 Joint-Motion Mutual Learning for Pose Estimation in Video
abstract
Human pose estimation in videos has long been a compelling yet challenging task within the realm of computer vision. Nevertheless, this task remains difficult because of the complex video scenes, such as video defocus and self-occlusion. Recent methods strive to integrate multi-frame visual features generated by a backbone network for pose estimation. However, they often ignore the useful joint information encoded in the initial heatmap, which is a by-product of the backbone generation. Comparatively, methods that attempt to refine the initial heatmap fail to consider any spatio-temporal motion features. As a result, the performance of existing methods for pose estimation falls short due to the lack of ability to leverage both local joint (heatmap) information and global motion (feature) dynamics.
Sifan Wu 0001, Haipeng Chen 0002, Yifang Yin, Sihao Hu, Runyang Feng, Yingying Jiao, Zhenguang Liu
ACM Multimedia6
2024 A column-shared histogramming TDC with pixel-to-pixel coincidence detection and compact analog counters for Flash LiDAR sensor
Kaiming Nie, Qinglong Lin, Yingying Jiao
Sci. China Inf. Sci.5
2022 Spatial-Temporal Correlation Modeling for Motion Prediction
abstract
Human motion prediction is fundamental for many applications in computer vision. Current methods typically handle motion prediction with seqential models, which ignore the fact that joint movement is driven by forces. In this paper, we provide a novel mechanical view to decompose force into magnitude and direction, which contributes to modeling the temporal evolution of joints. Moreover, existing graph convolution-based methods merely utilize the deep-level features, which is difficult to capture the complex spatial dependencies contexts. We introduce a novel spatial connections encoding model to capture the multi-level spatial dependencies between joints. Finally, to encode abundant temporal dependencies, we present a multi-head temporal encoding module. Comprehensive experiments show that our model sets the state-of-the-art performance on the largest human motion benchmark datasets.
Yingying Jiao, Haipeng Chen 0002, Chang Yao 0001, Pengxiang Su, Chong Fu 0002, Xiang Wang 0010
ICME1
2022 Multimodal medical image segmentation using multi-scale context-aware network
Zhanshan Li, Yingying Jiao
Neurocomputing4
2022 GLPose: Global-Local Representation Learning for Human Pose Estimation
abstract
Multi-frame human pose estimation is at the core of many computer vision tasks. Although state-of-the-art approaches have demonstrated remarkable results for human pose estimation on static images, their performances inevitably come short when being applied to videos. A central issue lies in the visual degeneration of video frames induced by rapid motion and pose occlusion in dynamic environments. This problem, by nature, is insurmountable for a single frame. Therefore, incorporating complementary visual cues from other video frames becomes an intuitive paradigm. Current state-of-the-art methods usually leverage information from adjacent frames, which unfortunately place excessive focus on only the temporally nearby frames. In this paper, we argue that combining global semantically similar information and local temporal visual context will deliver more comprehensive and more robust representations for human pose estimation. Towards this end, we present an effective framework, namely global-local enhanced pose estimation ( GLPose ) network. Our framework consists of a feature processing module that conditionally incorporates global semantic information and local visual context to generate a robust human representation and a feature enhancement module that excavates complementary information from this aggregated representation to enhance keyframe features for precise estimation. We empirically find that the proposed GLpose outperforms existing methods by a large margin and achieves new state-of-the-art results on large benchmark datasets.
Yingying Jiao, Haipeng Chen 0002, Runyang Feng, Haoming Chen, Sifan Wu 0001, Yifang Yin, Zhenguang Liu
ACM Trans. Multim. Comput. Commun. Appl.1
2020 Driver sleepiness detection from EEG and EOG signals using GAN and LSTM networks
Yingying Jiao, Yini Deng, Bao-Liang Lu
Neurocomputing1
2018 Driver Sleepiness Detection Using LSTM Neural Network
Yini Deng, Yingying Jiao, Bao-Liang Lu
ICONIP (4)2
2017 Detecting driver sleepiness from EEG alpha wave during daytime driving
abstract
Drowsy driving is the main reason for sleep-related crashes. We have observed that an alpha wave attenuation-disappearance phenomenon and a typical alpha blocking phenomenon commonly exist in the eye closure events during daytime simulated driving experiments. These two alpha-related phenomena prove to respectively represent two different sleepiness levels: the sleep onset and the relaxed wakefulness. Therefore, we propose a novel algorithm for tracking the alpha wave change and detecting the two alpha-related phenomena in real-time for recognizing driver sleepiness. Our proposed algorithm adopts continuous wavelet transform for charactering the signal change and support vector machine for classification. The experimental results indicate that the algorithm is able to detect the start and end points of alpha waves during eye-closed period and distinguish the two types of end points of alpha waves in two alpha-related phenomena with high sensitivity and precision. The eye-closed period detected by our algorithm with alpha waves has a high overlapping rate with that marked by human experts. The main contributions of the proposed algorithm are twofold: to detect alpha waves during the eye-closed period in real-time and serve as an indicator for judging the current sleepiness level as the sleep onset or relaxed wakefulness at the end points of alpha waves.
Yingying Jiao, Bao-Liang Lu
BIBM1
2014 Recognizing slow eye movement for driver fatigue detection with machine learning approach
abstract
Slow eye movement (SEM) regarded as a sign of onset of sleep is very significant for detecting driver fatigue, but its characteristics and detection algorithm have been rarely involved in the study of driver fatigue detection. In this study, some new features were extracted based on wavelet singularity analysis and statistics to detect SEMs. Six subjects participated in this simulated driving experiment, and for each subject, a more than 2 hours electro-oculogram (EOG) session was recorded. Each session was divided into SEM epochs and non-SEM epochs according to the common judgments made by the two of three experts by the visual recognition criteria of SEMs. Regarding the problem of detecting SEMs as an imbalance classification problem, and through the under-sampling and over-sampling methods a 2s horizontal electro-oculogram (HEO) signal could finally be recognized as the category of SEMs or non-SEMs with the classifiers SVM, GELM, and KNN respectively. Results prove that the proposed features was a little better than the wavelet energy features, and through the combination of the wavelet energy features and the new features based on wavelet singularity analysis and statistics, the classification results were improved obviously.
Yingying Jiao, Yong Peng 0001, Bao-Liang Lu, Shanguang Chen
IJCNN1