VLDB 2026 Research / reviewers in the wild / expert
Shasha Yuan
dblp:146/1908
· DBLP profile ↗
55ranked-venue papers
6as first author
38since 2021 · last 2026
0000-0003-4792-9880ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 37 · 3 first-author · 28 since 2021Artificial intelligence and machine learning · 16 · 3 first-author · 10 since 2021Systems, architecture and hardware · 1Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Improved Cross-Branch Fusion Method and Contrastive Learning for Seizure Type Recognition
Zaiwang Li, Manman Yuan, Chenchen Jiang, Longfei Qi, Shasha Yuan |
ICIC (29) | 5 |
| 2026 | Integrating modularity maximization and contrastive learning for identifying spatial domain from spatial transcriptomics
Shasha Yuan, Shengjun Li |
Expert Syst. Appl. | 2 |
| 2026 | Autoencoder-aided graph convolutional networks integrating multi-view and multi-scale for improving spatial domain identification
Juan Wang 0003, Xuena Liang, Shasha Yuan, Jin-Xing Liu 0001, Junliang Shang |
Knowl. Based Syst. | 3 |
| 2026 | TVFNet: text and visual attention feature fusion network for multi-lesion segmentation of diabetic retinopathy
Yanfei Guo, Yuanke Zhang, Fei Ma 0004, Jing Meng 0001, Shasha Yuan, Jindong Sun |
Neural Comput. Appl. | 5 |
| 2026 | MLRR-ATV: A Robust Manifold Nonnegative Low-Rank Representation With Adaptive Total-Variation Regularization for scRNA-seq Data ClusteringabstractSince genomics was proposed, the exploration of genes has been the focus of research. The emergence of single-cell RNA sequencing (scRNA-seq) technology makes it possible to explore gene expression at the single-cell level. Due to the limitations of sequencing technology, the data contains a lot of noise. At the same time, it also has the characteristics of high-dimensional and sparse. Clustering is a common method of analyzing scRNA-seq data. This paper proposes a novel single-cell clustering method called Robust Manifold Nonnegative Low-Rank Representation with Adaptive Total-Variation Regularization (MLRR-ATV). The Adaptive Total-Variation (ATV) regularization is introduced into Low-Rank Representation (LRR) model to reduce the influence of noise through gradient learning. Then, the linear and nonlinear manifold structures in the data are learned through Euclidean distance and cosine similarity, and more valuable information is retained. Because the model is non-convex, we use the Alternating Direction Method of Multipliers (ADMM) to optimize the model. We tested the performance of the MLRR-ATV model on eight real scRNA-seq datasets and selected nine state-of-the-art methods as comparison methods. The experimental results show that the performance of the MLRR-ATV model is better than the other nine methods. Gao-Fei Wang, Juan Wang 0003, Shasha Yuan, Chun-Hou Zheng 0001, Jin-Xing Liu 0001 |
IEEE Trans. Comput. Biol. Bioinform. | 3 |
| 2026 | Epileptic Seizure Prediction Using Multi-Strategy Data Augmentation and Hierarchical Contrastive LearningabstractAccurate early prediction of epileptic seizures is crucial for improving patients' quality of life. However, existing seizure prediction methods often rely on large-scale labeled datasets and face challenges in generalization and real-time performance. To address these issues, this study proposes an efficient seizure prediction framework that achieves high performance even with limited labeled data, significantly reducing dependence on extensive annotations. To better distinguish preictal states, contrastive learning is employed to enhance feature separation between interictal and preictal periods, leading to improved sensitivity in detecting early seizure patterns. First, a data augmentation strategy is designed, incorporating wavelet-based frequency mixing, temporal masking, and window-based masking to enhance model robustness and generalization. Second, a hierarchical contrastive loss function is introduced, integrating instance-level and temporal contrastive learning to improve the model's ability to capture preictal patterns. Finally, a lightweight SE-EEGNet is developed and optimized as a feature extractor, strengthening critical feature extraction and enabling real-time seizure prediction. On the CHB-MIT dataset, the proposed method achieves 94.51% accuracy, 95.05% sensitivity, a 0.024/h false positive rate (FPR), and a 20.12-minute prediction time using only 30% labeled data. On the Siena dataset, it achieves 93.14% accuracy, 92.77% sensitivity, and a 0.030/h FPR. Moreover, performance improves further as the amount of labeled data increases, validating the effectiveness and practical applicability of the proposed approach in seizure prediction. Longfei Qi, Feng Li 0033, Junliang Shang, Shihan Wang 0009, Shasha Yuan |
IEEE J. Biomed. Health Informatics | 6 |
| 2026 | Cluster-Guided Contrastive Learning With Masked Autoencoder for Spatial Domain Identification Based on Spatial TranscriptomicsabstractRecent advancements in spatial transcriptomics technology have enabled the capture of gene expression profiles while maintaining spatial information. Accurately identifying spatial clustering plays a pivotal role in analyzing spatial transcriptomics data and understanding tissue microenvironments. However, current spatial domain identification methods cannot explore the complex relationship of gene expression profiles and spatial topology. To alleviate this issue, we propose STMCCL, a novel self-supervised learning framework that jointly trains a masked autoencoder and cluster-guided contrastive learning. This framework extracts informative latent representations from gene expression profiles and spatial information. Specifically, we first use data augmentation strategies to build augmented views and employ a masked encoder to generate a feature view. Then, encoders are applied to learn view-unique embeddings of each view. Furthermore, we introduce a multiple cluster-perspectives module that considers both geometric and structural relationships between clusters to produce more reliable cluster assignments. Finally, to derive more discriminative positives and negatives, the cluster-guided contrastive module calculates the confidence of each sample based on the initial cluster. Comprehensive experiments on 7 public datasets demonstrate that STMCCL outperforms the state-of-the-art baselines with finer-scale spatial domain identification. Juan Wang 0003, Shasha Yuan, Junliang Shang |
IEEE J. Biomed. Health Informatics | 3 |
| 2025 | TAGCL-DDI: Two-Stage Augmentation-Driven Graph Contrastive Learning for Drug-Drug Interaction PredictionabstractDrug-drug interactions (DDIs) occur when the simultaneous administration of multiple drugs alters their pharmacological effects or causes adverse reactions. Existing DDIs prediction methods typically extract features of drug pairs from a DDI event graph and utilize a multilayer perceptron classifier to predict interactions. However, relying solely on the single DDI event graph limits the encoder's ability to capture comprehensive drug features. To overcome this limitation, we propose TAGCLDDI, a novel Two-stage Augmentation-driven Graph Contrastive Learning framework enhance drug feature representation. In the first stage, TAGCL-DDI employs the multi-head graph attention network to adaptively learn edge weights within the DDI event graph, and then generates multiple augmented graph views by mapping the graph into diverse subspaces through different encoders, enabling the capture of richer interaction information. In the second stage, node-level contrastive learning further refines feature representations across subspaces. Experimental results on two benchmark datasets show that TAGCL-DDI achieves state-of-the-art performance in DDIs prediction. Jiquan Zhao, Feng Li 0033, Yuzhuo Yuan, Wentian Xin, Shasha Yuan |
BIBM | 6 |
| 2025 | Epileptic Seizure Detection Using ECA-EEGNet with Earth Mover's Distance-Based Metric LearningabstractAccurate and efficient detection of epileptic seizures from electroencephalogram (EEG) signals is of great significance for clinical diagnosis and real-time monitoring. However, traditional EEG analysis methods face limitations in feature extraction and classification accuracy, primarily due to their heavy reliance on handcrafted features and rigid decision boundaries. Moreover, in clinical settings, the scarcity of seizure EEG signals and the difficulty in obtaining labeled data often lead to overfitting, especially when training data is insufficient. To address these challenges, this paper proposes a novel end-to-end seizure detection framework that integrates an attention-guided lightweight neural network with an advanced metric learning strategy. Based on the baseline EEGNet architecture, the proposed framework incorporates an Efficient Channel Attention (ECA) module to enhance the extraction of discriminative features from multichannel EEG signals. Furthermore, to improve the separability between seizure and non-seizure interictal states, we introduce a triplet loss-based metric learning method using the Earth Mover's Distance. By employing Earth Mover's Distance as the distance metric in the feature space and introducing a triplet loss function to constrain the relative distance relationships between samples, the proposed method ensures more compact embeddings for intra-class samples and better separation between inter-class feature distributions, thereby effectively mitigating overfitting risks under limited data conditions. Experimental evaluations on the CHB-MIT dataset demonstrate the superior performance of the proposed method, achieving an average accuracy of 97.51 %, sensitivity of 95.56 %, and specificity of 97.93 %. These results indicate that the proposed framework provides a promising and computationally efficient solution for automatic seizure detection in practical EEG analysis. Shihan Wang 0009, Junliang Shang, Juan Wang 0003, Longfei Qi, Shasha Yuan |
BIBM | 5 |
| 2025 | An Adaptive Single-Cell Sequencing Data Cluster Method Under Weight Fusion ConstraintabstractSingle-cell RNA sequencing (scRNA-seq) provides the transcriptome of a single cell, allowing researchers to study cellular phenomena at a higher resolution level. Nevertheless, noise generated by technical limitations and other results seriously interferes with the downstream analysis of sequencing data such as clustering. How to minimize the impact of noise on the accuracy of clustering methods has become a focus of current research. In this case, we propose a novel cell clustering algorithm called low-rank representation constrained clustering based on noise weight fusion (LRBNW). First, we mitigate the noise interference by introducing a noise weight matrix and assigning different weights to the noise through a reliability assessment strategy. By assigning larger weights to smaller reconstruction errors, we then highlight useful features with small errors, which clean features more representative in data analysis. Finally, we impose a k-block diagonal constraint on the affinity matrix through a block strategy, grouping related features into the same block, to eliminate redundancy among them and avoid over-reliance on related features. Extensive experiments demonstrate that LRBNW achieves higher accuracy results than existing state-of-the-art clustering methods on 10 real scRNA-seq datasets. In addition, downstream analysis experiments also indicated that LRBNW can identify biologically significant groups and reduce noise interference in scRNA-seq data. The result proves LRBNW is a powerful cell type identification tool, and has potential in predicting new cell types. Zhenchang Wang, Shasha Yuan, Feng Li 0033, Juan Wang 0003 |
BIBM | 2 |
| 2025 | Semi-Supervised Gaussian Mixture Variational Autoencoder with Graph Representation for Epileptic Seizure DetectionabstractAccurate electroencephalogram(EEG) annotation is essential for seizure detection but costly and error-prone, which can affect subsequent tasks. Moreover, the brain is a non-Euclidean topological structure, which contains spatial information for seizure detection. Based on these, this paper proposes a semi- supervised model based on Gaussian mixture variational autoencoder with graph representation, named GGMVAE. Firstly, for unlabeled EEG signals, we construct an adjacency matrix via Pearson correlation between channels. Then, matrix and EEG features are fed into a Gaussian Mixture VAE to learn the temporal and spatiall features through unsupervised training. Finally, the pre-trained encoder extracts low-dimensional features from partially labeled data for classification. The method is evaluated on the epilepsy dataset at the University of Helsinki, achieving the accuracy of 97.27 %, precision of 96.17 %, recall of 98.41 %, and F1-score of 97.28 %. The results indicate that this semi-supervised model can effectively learn the temporal and spatial features of EEG signals, improving the performance of seizure detection.. Shasha Yuan, Chenchen Jiang, Manman Yuan, Qianqian Ren, Yanfei Guo |
BIBM | 1 |
| 2025 | Low-Rank Multiple Kernel Model Based on Local Structures Learning and Adaptive Similarity Preserving for scRNA-seq Data Clustering
Juan Wang 0003, Tian-Jing Qiao, Zhenduo Zhang, Chun-Hou Zheng 0001, Shasha Yuan |
ICIC (25) | 5 |
| 2025 | Dynamic momentum contrastive learning network for diabetic retinopathy grading
Yanfei Guo, Chenglong Yang, Hangli Du, Yuanke Zhang, Fei Ma 0004, Shasha Yuan |
Eng. Appl. Artif. Intell. | 6 |
| 2025 | A Modified Transformer Network for Seizure Detection Using EEG SignalsabstractSeizures have a serious impact on the physical function and daily life of epileptic patients. The automated detection of seizures can assist clinicians in taking preventive measures for patients during the diagnosis process. The combination of deep learning (DL) model with convolutional neural network (CNN) and transformer network can effectively extract both local and global features, resulting in improved seizure detection performance. In this study, an enhanced transformer network named Inresformer is proposed for seizure detection, which is combined with Inception and Residual network extracting different scale features of electroencephalography (EEG) signals to enrich the feature representation. In addition, the improved transformer network replaces the existing Feedforward layers with two half-step Feedforward layers to enhance the nonlinear representation of the model. The proposed architecture utilizes discrete wavelet transform (DWT) to decompose the original EEG signals, and the three sub-bands are selected for signal reconstruction. Then, the Co-MixUp method is adopted to solve the problem of data imbalance, and the processed signals are sent to the Inresformer network for seizure information capture and recognition. Finally, discriminant fusion is performed on the results of three-scale EEG sub-signals to achieve final seizure recognition. The proposed network achieves the best accuracy of 100% on Bonn dataset and the average accuracy of 98.03%, sensitivity of 95.65%, and specificity of 98.57% on the long-term CHB-MIT dataset. Compared to the existing DL networks, the proposed method holds significant potential for clinical research and diagnosis applications with competitive performance. Wenrong Hu, Juan Wang 0003, Feng Li 0033, Qingwei Jia, Shasha Yuan |
Int. J. Neural Syst. | 7 |
| 2025 | A Contrastive Learning-Enhanced Residual Network for Predicting Epileptic Seizures Using EEG SignalsabstractThe models used to predict epileptic seizures based on electroencephalogram (EEG) signals often encounter substantial challenges due to the requirement for large, labeled datasets and the inherent complexity of EEG data, which hinders their robustness and generalization capability. This study proposes CLResNet, a framework for predicting epileptic seizures, which combines contrastive self-supervised learning with a modified deep residual neural network to address the above challenges. In contrast to traditional models, CLResNet uses unlabeled EEG data for pre-training to extract robust feature representations. It is then fine-tuned on a smaller labeled dataset to significantly reduce its reliance on labeled data while improving its efficiency and predictive accuracy. The contrastive learning (CL) framework enhances the ability of the model to distinguish between preictal and interictal states, thus improving its robustness and generalizability. The architecture of CLResNet contains residual connections that enable it to learn deep features of the data and ensure an efficient gradient flow. The results of the evaluation of the model on the CHB-MIT dataset showed that it outperformed prevalent methods in the field, with an accuracy of 92.97%, sensitivity of 94.18%, and false-positive rate of 0.043/h. On the Siena dataset, the model also achieved competitive performance, with an accuracy of 92.79%, a sensitivity of 91.47%, and a false-positive rate of 0.041/h. These results confirm the effectiveness of CLResNet in addressing variations in EEG data, and show that contrastive self-supervised learning is a robust and accurate approach for predicting seizures. Longfei Qi, Shasha Yuan, Feng Li 0033, Junliang Shang, Juan Wang 0003, Shihan Wang 0009 |
Int. J. Neural Syst. | 2 |
| 2024 | Improve spatial domain identification for spatial transcriptomics using high-order neighbor feature hybrid graph convolutional networksabstractRecent developments in spatial transcriptomics (ST) technologies have afforded us a profound understanding of gene expression patterns in the tissue microenvironment. Recently, several prominent spatial domain identification methods have been introduced to employ both spatial and expression information for precisely deciphering tissue structures. However, existing methods only focus on information from immediate neighbors, failing to capture the mixed relationships of neighbors at various scales and learn a general mixed feature from neighbors at different distances. To this end, we propose ST-HNHG, which fuses gene expression profiles, spatial information, and morphological images for deciphering spatial domains. Specifically, the high-order neighbor feature hybrid graph convolutional network (HNHGCN) is designed to capture feature representations between neighbors at different distances and learn their linear mixing. A data augmentation module is also proposed to enhance data diversity and model robustness. The attention mechanism is also introduced to integrate the embeddings learned from morphological and expression information, obtaining the latent representation for spatial domain identification. We test ST-HNHG on two ST datasets. The results indicate that ST-HNHG outperforms most existing methods, and considering the linear mixing between neighbors at various scales is beneficial for improving the accuracy of recognizing spatial domains. Xuena Liang, Shasha Yuan, Shengjun Li, Juan Wang 0003 |
BIBM | 2 |
| 2024 | Seizure Types Classification Based on Multi-branch Hybrid Deep Learning Network
Qingwei Jia, Jin-Xing Liu 0001, Junling Shang, Ling-Yun Dai, Wenrong Hu, Shasha Yuan |
ICIC (4) | 7 |
| 2024 | Epileptic Seizure Detection with an End-to-End Temporal Convolutional Network and Bidirectional Long Short-Term Memory ModelabstractAutomatic seizure detection plays a key role in assisting clinicians for rapid diagnosis and treatment of epilepsy. In view of the parallelism of temporal convolutional network (TCN) and the capability of bidirectional long short-term memory (BiLSTM) in mining the long-range dependency of multi-channel time-series, we propose an automatic seizure detection method with a novel end-to-end TCN-BiLSTM model in this work. First, raw EEG is filtered with a 0.5-45 Hz band-pass filter, and the filtered data are input into the proposed TCN-BiLSTM network for feature extraction and classification. Post-processing process including moving average filtering, thresholding and collar technique is then employed to further improve the detection performance. The method was evaluated on two EEG database. On the CHB-MIT scalp EEG database, our method achieved a segment-based sensitivity of 94.31%, specificity of 97.13%, and accuracy of 97.09%. Meanwhile, an event-based sensitivity of 96.48% and an average false detection rate (FDR) of 0.38/h were obtained. On the SH-SDU database we collected, the segment-based sensitivity of 94.99%, specificity of 93.25%, and accuracy of 93.27% were achieved. In addition, an event-based sensitivity of 99.35% and a false detection rate of 0.54/h were yielded. The total detection time consumed for 1[Formula: see text]h EEG data was 5.65[Formula: see text]s. These results demonstrate the superiority and promising potential of the proposed method in real-time monitoring of epileptic seizures. Xingchen Dong, Yiming Wen, Dezan Ji, Shasha Yuan, Zhen Liu 0028 |
Int. J. Neural Syst. | 4 |
| 2024 | Combining EEG Features and Convolutional Autoencoder for Neonatal Seizure DetectionabstractNeonatal epilepsy is a common emergency phenomenon in neonatal intensive care units (NICUs), which requires timely attention, early identification, and treatment. Traditional detection methods mostly use supervised learning with enormous labeled data. Hence, this study offers a semi-supervised hybrid architecture for detecting seizures, which combines the extracted electroencephalogram (EEG) feature dataset and convolutional autoencoder, called Fd-CAE. First, various features in the time domain and entropy domain are extracted to characterize the EEG signal, which helps distinguish epileptic seizures subsequently. Then, the unlabeled EEG features are fed into the convolutional autoencoder (CAE) for training, which effectively represents EEG features by optimizing the loss between the input and output features. This unsupervised feature learning process can better combine and optimize EEG features from unlabeled data. After that, the pre-trained encoder part of the model is used for further feature learning of labeled data to obtain its low-dimensional feature representation and achieve classification. This model is performed on the neonatal EEG dataset collected at the University of Helsinki Hospital, which has a high discriminative ability to detect seizures, with an accuracy of 92.34%, precision of 93.61%, recall rate of 98.74%, and F1-score of 95.77%, respectively. The results show that unsupervised learning by CAE is beneficial to the characterization of EEG signals, and the proposed Fd-CAE method significantly improves classification performance. Shasha Yuan, Jin-Xing Liu 0001, Wenrong Hu, Qingwei Jia, Fangzhou Xu |
Int. J. Neural Syst. | 2 |
| 2024 | EEG-based epileptic seizure detection using deep learning techniques: A survey
Jie Xu 0059, Kuiting Yan, Zengqian Deng, Yankai Yang, Jin-Xing Liu 0001, Juan Wang 0003, Shasha Yuan |
Neurocomputing | 7 |
| 2024 | A Clustering Method for Single-Cell RNA-Seq Data Based on Automatic Weighting Penalty and Low-Rank RepresentationabstractAdvances in high-throughput single-cell RNA sequencing (scRNA-seq) technology have provided more comprehensive biological information on cell expression. Clustering analysis is a critical step in scRNA-seq research and provides clear knowledge of the cell identity. Unfortunately, the characteristics of scRNA-seq data and the limitations of existing technologies make clustering encounter a considerable challenge. Meanwhile, some existing methods treat different features equally and ignore differences in feature contributions, which leads to a loss of information. To overcome limitations, we introduce a weighted distance constraint into the construction of the similarity graph and combine the similarity constraint. We propose the Joint Automatic Weighting Similarity Graph and Low-rank Representation (JAGLRR) clustering method. Evaluating the contributions of each feature and assigning various weight values can increase the significance of valuable features while decreasing the interference of redundant features. The similarity constraint allows the model to generate a more symmetric affinity matrix. Benefitting from that affinity matrix, JAGLRR recovers the original linear relationship of the data more accurately and obtains more discriminative information. The results on simulated datasets and 8 real datasets show that JAGLRR outperforms 11 existing comparison methods in clustering experiments, with higher clustering accuracy and stability. Juan Wang 0003, Zhen-Chang Wang, Shasha Yuan, Chun-Hou Zheng 0001, Jin-Xing Liu 0001, Junliang Shang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2024 | Spatiotemporal Network Based on GCN and BiGRU for Seizure DetectionabstractAs an important tool for detecting and diagnosing epilepsy, multi-channel EEG records the neuronal activities of different brain regions. Visual identification of abnormal EEG signals poses challenges, making the use of artificial intelligence techniques for automated seizure detection an inevitable trend. However, existing seizure detection methods often overlook the spatial relationship between EEG channels, which can't take full advantage of brain network structure. In this paper, we design an end-to-end spatiotemporal architecture for seizure detection based on Graph Convolutional Networks (GCN) and Bidirectional Gated Recurrent Units (BiGRU) to efficiently model the spatial dependence and temporal dynamics of EEG. Firstly, the original EEG signals are preprocessed by applying wavelet transform for temporal-frequency analysis. The Pearson correlation matrix is computed for specific frequency bands and GCN is utilized to extract spatial features between EEG channels. Then, these features are sent into the BiGRU network to capture temporal relationships. Finally, the detection decisions are achieved using fully connected layers and the multi-level decision rules are implemented to provide the final results. The proposed method is validated on CHB-MIT EEG dataset, achieving 98.85% sensitivity, 95.83% specificity, 97.35% accuracy, 97.4% F1-score, and 97.33% AUC. This network fusions multiple EEG characteristics in the spatial-temporal-frequency domains to improve the detection performance and the promising result demonstrates that the performance of this model is superior to or on par with existing methods. Jie Xu 0059, Shasha Yuan, Junliang Shang, Juan Wang 0003, Kuiting Yan, Yankai Yang |
IEEE J. Biomed. Health Informatics | 2 |
| 2023 | Epileptic EEG Signal Detection based on Uncorrelated Multilinear Principal Component Analysis and Metric LearningabstractEpilepsy is a chronically occurring neurological disorder which is characterized by uninterrupted repetitive seizures that can occur spontaneously, making it one of the most prevalent brain disorders. A novel method for detecting seizures is proposed in this paper, which utilizes uncorrelated multilinear principal component analysis (UMPCA) and metric learning based on doublet support vector machine (doublet-SVM). The system first segmented the EEG signal and performed modified Stockwell transform (MST) to obtain a 2-dimensional time-frequency spectrum, and the third-order tensor for multichannel EEG signals was constructed based on time, frequency and spatial domains. Then, the uncorrelated features were extracted using the UMPCA, which could distinguish seizure and non-seizure characteristics in the massive third-order EEG tensor. After that, the distance metric learning is approached by employing the doublet-SVM algorithm, which transforms it into a kernel classifier problem for efficient EEG classification. The performance of this epilepsy detection model was tested and evaluated on the Freiburg EEG database of 21 patients, and the average sensitivity, specificity and accuracy obtained were 98.74%, 98.11% and 98.12%, respectively. The results demonstrate the significant ability of this algorithm in detecting seizures. Yankai Yang, Kuiting Yan, Shasha Yuan |
BIBM | 5 |
| 2023 | Epileptic Seizure Detection Based on Feature Extraction and CNN-BiGRU Network with Attention Mechanism
Jie Xu 0059, Juan Wang 0003, Jin-Xing Liu 0001, Junliang Shang, Ling-Yun Dai, Kuiting Yan, Shasha Yuan |
ICIC (2) | 7 |
| 2023 | Seizure Prediction Based on Hybrid Deep Learning Model Using Scalp Electroencephalogram
Kuiting Yan, Junliang Shang, Juan Wang 0003, Jie Xu 0059, Shasha Yuan |
ICIC (2) | 5 |
| 2023 | scGASI: A Graph Autoencoder-Based Single-Cell Integration Clustering Method
Tian-Jing Qiao, Feng Li 0033, Shasha Yuan, Ling-Yun Dai, Juan Wang 0003 |
ISBRA | 3 |
| 2023 | BioSTD: A New Tensor Multi-View Framework via Combining Tensor Decomposition and Strong Complementarity Constraint for Analyzing Cancer Omics DataabstractAdvances in omics technology have enriched the understanding of the biological mechanisms of diseases, which has provided a new approach for cancer research. Multi-omics data contain different levels of cancer information, and comprehensive analysis of them has attracted wide attention. However, limited by the dimensionality of matrix models, traditional methods cannot fully use the key high-dimensional global structure of multi-omics data. Moreover, besides global information, local features within each omics are also critical. It is necessary to consider the potential local information together with the high-dimensional global information, ensuring that the shared and complementary features of the omics data are comprehensively observed. In view of the above, this article proposes a new tensor integrative framework called the strong complementarity tensor decomposition model (BioSTD) for cancer multi-omics data. It is used to identify cancer subtype specific genes and cluster subtype samples. Different from the matrix framework, BioSTD utilizes multi-view tensors to coordinate each omics to maximize high-dimensional spatial relationships, which jointly considers the different characteristics of different omics data. Meanwhile, we propose the concept of strong complementarity constraint applicable to omics data and introduce it into BioSTD. Strong complementarity is used to explore the potential local information, which can enhance the separability of different subtypes, allowing consistency and complementarity in the omics data to be fully represented. Experimental results on real cancer datasets show that our model outperforms other advanced models, which confirms its validity. Ying-Lian Gao, Juan Wang 0003, Shasha Yuan, Jin-Xing Liu 0001 |
IEEE J. Biomed. Health Informatics | 4 |
| 2023 | A Personalized Low-Rank Subspace Clustering Method Based on Locality and Similarity Constraints for scRNA-seq Data AnalysisabstractSingle-cell RNA sequencing (scRNA-seq) technology can provide expression profile of single cells, which propels biological research into a new chapter. Clustering individual cells based on their transcriptome is a critical objective of scRNA-seq data analysis. However, the high-dimensional, sparse and noisy nature of scRNA-seq data pose a challenge to single-cell clustering. Therefore, it is urgent to develop a clustering method targeting scRNA-seq data characteristics. Due to its powerful subspace learning capability and robustness to noise, the subspace segmentation method based on low-rank representation (LRR) is broadly used in clustering researches and achieves satisfactory results. In view of this, we propose a personalized low-rank subspace clustering method, namely PLRLS, to learn more accurate subspace structures from both global and local perspectives. Specifically, we first introduce the local structure constraint to capture the local structure information of the data, while helping our method to obtain better inter-cluster separability and intra-cluster compactness. Then, in order to retain the important similarity information that is ignored by the LRR model, we utilize the fractional function to extract similarity information between cells, and introduce this information as the similarity constraint into the LRR framework. The fractional function is an efficient similarity measure designed for scRNA-seq data, which has theoretical and practical implications. In the end, based on the LRR matrix learned from PLRLS, we perform downstream analyses on real scRNA-seq datasets, including spectral clustering, visualization and marker gene identification. Comparative experiments show that the proposed method achieves superior clustering accuracy and robustness. Tian-Jing Qiao, Jin-Xing Liu 0001, Junliang Shang, Shasha Yuan, Chun-Hou Zheng 0001, Juan Wang 0003 |
IEEE J. Biomed. Health Informatics | 4 |
| 2023 | NLRRC: A Novel Clustering Method of Jointing Non-Negative LRR and Random Walk Graph Regularized NMF for Single-Cell Type IdentificationabstractThe development of single-cell RNA sequencing (scRNA-seq) technology has opened up a new perspective for us to study disease mechanisms at the single cell level. Cell clustering reveals the natural grouping of cells, which is a vital step in scRNA-seq data analysis. However, the high noise and dropout of single-cell data pose numerous challenges to cell clustering. In this study, we propose a novel matrix factorization method named NLRRC for single-cell type identification. NLRRC joins non-negative low-rank representation (LRR) and random walk graph regularized NMF (RWNMFC) to accurately reveal the natural grouping of cells. Specifically, we find the lowest rank representation of single-cell samples by non-negative LRR to reduce the difficulty of analyzing high-dimensional samples and capture the global information of the samples. Meanwhile, by using random walk graph regularization (RWGR) and NMF, RWNMFC captures manifold structure and cluster information before generating a cluster allocation matrix. The cluster assignment matrix contains cluster labels, which can be used directly to get the clustering results. The performance of NLRRC is validated on simulated and real single-cell datasets. The results of the experiments illustrate that NLRRC has a significant advantage in single-cell type identification. Juan Wang 0003, Linping Wang, Shasha Yuan, Feng Li 0033, Jin-Xing Liu 0001, Junliang Shang |
IEEE J. Biomed. Health Informatics | 3 |
| 2023 | Automatic Seizure Detection Using Logarithmic Euclidean-Gaussian Mixture Models (LE-GMMs) and Improved Deep Forest LearningabstractAutomatic seizure detection could facilitate early detection, improve treatment planning, and reduce medical workload. This study describes a novel Logarithmic Euclidean-Gaussian Mixture Models (LE-GMMs) and an improved Deep Forest learning algorithm for epileptic seizure detection. The LE-GMMs could map the Riemannian manifold structure of Gaussian models to linear Euclidean space, which fully exploits the ability of GMMs to distinguish non-seizure and seizure EEG signals. The Multi-Pooling and error Screening Forest (MPSForest) learning method based on Deep Forest uses multi-pooling and out-of-bagging (OOB) error screening to reduce memory load and random tree construction. Firstly, variational modal decomposition (VMD) is applied to decompose electroencephalogram (EEG) signals into five layers, and the first three layers are chosen to construct EEG time-frequency distribution. Then Gaussian Mixture Models are estimated, and the LE-GMMs are constructed to extract valid EEG features. These features are input into the MPSForest model to classify seizure and non-seizure samples. After that, the outputs are subjected to post-processing to get the final seizure detection results, including moving average filtering and the adaptive collar technique. The proposed method achieves average sensitivity of 98.22% and specificity of 98.99% on the UPenn and Mayo Clinic dataset, and for the long-term Freiburg EEG dataset with 21 patients, the sensitivity of 98.47% and specificity of 98.57% are yielded respectively with the false detection rate of 0.24/h. The experimental results show that this proposed method has excellent accuracy in distinguishing non-seizure and seizure EEG signals and holds great potential for clinical research and diagnostics. Shasha Yuan, Junliang Shang, Jin-Xing Liu 0001, Juan Wang 0003 |
IEEE J. Biomed. Health Informatics | 1 |
| 2022 | An integrated Extreme learning machine based on kernel risk-sensitive loss of q-Gaussian and voting mechanism for sample classificationabstractEnsemble learning is to train and combine multiple learners to complete the corresponding learning tasks. It can improve the stability of the overall model, and a good ensemble method can further improve the accuracy of the model. At the same time, as one of the outstanding representatives of machine learning, Extreme Learning Machine has attracted the continuous attention of experts and scholars. to get a better representation of the feature space, we extend the Gaussian kernel in the kernel risk-sensitive loss and propose a Kernel Risk-Sensitive Loss of q-Gaussian kernel and Hyper-graph Regularized Extreme Learning Machine method. Since the contingency in the ELM training process cannot be completely avoided, the stability of most ELM methods is affected to some extent. What’s more, we introduce the voting mechanism and a new ELM classification model named Kernel Risk-Sensitive Loss of q-Gaussian kernel and Hyper-graph Regularized Integrated Extreme Learning Machine based on Voting Mechanism is proposed. It improves the stability of the model through the idea of ensemble learning. We apply the new model on six real data sets, and through observation and analysis of experimental results, we find that the new model has certain competitiveness, especially in classification accuracy and stability. Ying-Lian Gao, Zhen-Xin Niu, Shasha Yuan, Chun-Hou Zheng 0001, Jin-Xing Liu 0001 |
BIBM | 4 |
| 2022 | A Multi-Graph Laplacian Regularized Low-Rank Representation method for cancer sample clustering with integrated TCGA dataabstractRecently, cancer sample clustering research based on gene expression data has been completely developed. Moreover, studies discover that other genomic data in TCGA besides gene expression data also contain features that can be utilized to cluster. Thus, by integrating these genomic data, new cancer clustering feature source can be formed. As a powerful subspace clustering method, Low-Rank Representation (LRR) has delivered an important breakthrough in clustering cancer samples. However, most methods based on LRR are only employed to analyze gene expression data, and cannot make full use of the characteristic information of other genomic data. Based on the LRR method, this paper proposes a novel Multi-Graph Laplacian regularized Low-Rank Representation (MGLLRR) method for cancer sample clustering using multi-omics datasets. To preserve the local geometry in genomic data, multi-graph regularization is led into MGLLRR method. The multi-graph Laplacian can fully preserve the hidden non-linear manifold structure in the data to make sure the smoothness of the integrated data along the estimated manifold. Considering the noise effect of different genomic data, we also introduce the idea of block constraint. We set each genome data as a data block and impose different constraint on it. Therefore, it can avoid the influence of different noise in multiple genomic data and improve the reliability of tumor clustering. The clustering experimental results indicate the effectiveness of MGLLRR on cancer sample clustering. And MGLLRR is a practical and effective analysis method of multiple genomic data. Juan Wang 0003, Li-Hong Wang, Tian-Jing Qiao, Shasha Yuan |
BIBM | 4 |
| 2022 | A Tensor Robust Model Based on Enhanced Tensor Nuclear Norm and Low-Rank Constraint for Multi-view Cancer Genomics Data
Shasha Yuan, Junliang Shang, Jin-Xing Liu 0001 |
ISBRA | 2 |
| 2022 | Multi-view manifold regularized compact low-rank representation for cancer samples clustering on multi-omics dataabstractBACKGROUND: The identification of cancer types is of great significance for early diagnosis and clinical treatment of cancer. Clustering cancer samples is an important means to identify cancer types, which has been paid much attention in the field of bioinformatics. The purpose of cancer clustering is to find expression patterns of different cancer types, so that the samples with similar expression patterns can be gathered into the same type. In order to improve the accuracy and reliability of cancer clustering, many clustering methods begin to focus on the integration analysis of cancer multi-omics data. Obviously, the methods based on multi-omics data have more advantages than those using single omics data. However, the high heterogeneity and noise of cancer multi-omics data pose a great challenge to the multi-omics analysis method. RESULTS: In this study, in order to extract more complementary information from cancer multi-omics data for cancer clustering, we propose a low-rank subspace clustering method called multi-view manifold regularized compact low-rank representation (MmCLRR). In MmCLRR, each omics data are regarded as a view, and it learns a consistent subspace representation by imposing a consistence constraint on the low-rank affinity matrix of each view to balance the agreement between different views. Moreover, the manifold regularization and concept factorization are introduced into our method. Relying on the concept factorization, the dictionary can be updated in the learning, which greatly improves the subspace learning ability of low-rank representation. We adopt linearized alternating direction method with adaptive penalty to solve the optimization problem of MmCLRR method. CONCLUSIONS: Finally, we apply MmCLRR into the clustering of cancer samples based on multi-omics data, and the clustering results show that our method outperforms the existing multi-view methods. Juan Wang 0003, Cong-Hai Lu, Ling-Yun Dai, Shasha Yuan |
BMC Bioinform. | 5 |
| 2021 | Robust Tensor Method Based on Correntropy and Tensor Singular Value Decomposition for Cancer Genomics DataabstractThe analysis of biological sequencing data can provide significant support for researchers to unravel the mysteries of life further. This paper proposes a robust tensor data analysis method based on correntropy and tensor singular value decomposition (t-SVD) (CoTD) to analyze high-dimensional and multi-way cancer genomics data. CoTD uses the maximum correntropy criterion to increase the sparsity of the sparse tensor and fully exploits the vital information of the tensor data. It can effectively suppress outliers in the process of recovering low-rank and separating sparse data. In addition, through t-SVD, the internal spatial structure of the original tensor data can be well preserved. In this way, essential information can be retained in the low-rank part, which increases the clustering effect. The CoTD model is optimized by the half-quadratic technique and alternating direction method of multipliers (ADMM). Sample clustering and differentially expressed gene (DEG) extraction experiments are carried out on cancer genomics datasets. CoTD model is compared with four similar methods, which proves that the CoTD model has good performance. Ying-Lian Gao, Shasha Yuan, Jin-Xing Liu 0001 |
BIBM | 3 |
| 2021 | Joint CC and Bimax: A Biclustering Method for Single-Cell RNA-Seq Data Analysis
He-Ming Chu, Jin-Xing Liu 0001, Juan Wang 0003, Shasha Yuan, Ling-Yun Dai |
ISBRA | 5 |
| 2021 | The Automatic Detection of Seizure Based on Tensor Distance And Bayesian Linear Discriminant AnalysisabstractElectroencephalogram (EEG) plays an important role in recording brain activity to diagnose epilepsy. However, it is not only laborious, but also not very cost effective for medical experts to manually identify the features on EEG. Therefore, automatic seizure detection in accordance with the EEG recordings is significant for the diagnosis and treatment of epilepsy. Here, a new method for detecting seizures using tensor distance (TD) is proposed. First, the time-frequency characteristics of EEG signals are obtained by wavelet transformation, and the tensor representation of EEG signals is then obtained. Tucker decomposition is used to obtain the principal components of the EEG tensor. After, the distances between different categories of EEG tensors are calculated as the EEG features. Finally, the TD features are classified through the Bayesian Linear Discriminant Analysis (Bayesian LDA) classifier. The performance of this method is measured by the sensitivity, specificity, and recognition accuracy. Results indicate 95.12% sensitivity, 97.60% specificity, 97.60% recognition accuracy, and a false detection rate of 0.76 per hour in the invasive EEG dataset, which included 566.57[Formula: see text]h of EEG recording data from 21 patients. Taken together, the results show that TD has a good detection effect for seizure classification and that this method has high computational speed and great potential for real-time diagnosis. Delu Ma, Shasha Yuan, Junliang Shang, Jin-Xing Liu 0001, Ling-Yun Dai, Fangzhou Xu |
Int. J. Neural Syst. | 2 |
| 2021 | Logistic Weighted Profile-Based Bi-Random Walk for Exploring MiRNA-Disease Associations
Ling-Yun Dai, Jin-Xing Liu 0001, Juan Wang 0003, Shasha Yuan |
J. Comput. Sci. Technol. | 5 |
| 2020 | Dual Graph regularized PCA based on Different Norm Constraints for Bi-clustering Analysis on Single-cell RNA-seq DataabstractIn recent years, single-cell RNA sequencing (scRNA-seq) technology has made significant progress in many fields and become an important means to study cell dynamics. How to effectively mine valuable biological information from these sequencing data is a topic worthy of researching. In this paper, two new methods based on traditional principal component analysis (PCA) are proposed and used to scRNA-seq data. The first method named dual graph regularized PCA (DGPPCA) is based on Frobenius-norm and L2,p-norm constraints, and the method named the dual graph-regularization PCA (DG2PPCA) is based on the nonconvex proximal Lp-norm ( 02,p-norm constraints. We apply these two new methods to five scRNA-seq datasets, and perform bi-clustering on genes and samples at the same time. Extensive experiments are conducted to explore the influence of the combination of different norm constraints in the two optimization models. Jin-Xing Liu 0001, Juan Wang 0003, Shasha Yuan, Ling-Yun Dai |
BIBM | 5 |
| 2020 | Automatic Seizure Prediction based on Modified Stockwell Transform and Tensor DecompositionabstractReliable epileptic seizure prediction is significantly important in improving the life of patients and enhancing the therapy effect. In this paper, a novel seizure prediction algorithm is proposed employing the tensor decomposition on long-term intracranial EEG recordings. The modified Stockwell transform (MST) is conducted on the segmented EEG signals to transform into two-dimensional instantaneous power spectra. Then, the third-order tensor representation of the multi-channel EEG signals are structured with the models of time, frequency and space. Tucker decomposition, one valid tensor decomposition method, is applied to obtain the principal components of the EEG tensors and the smaller core tensors after decomposition are extracted as features of interictal EEG and preictal EEG. After that, the classification of preictal and interictal data is achieved by feeding the features into Bayesian Linear Discriminant Analysis (BLDA) classifier. The evaluation of the proposed algorithm is carried out on the Freiburg EEG database and a sensitivity of 88.49% for the seizure occurrence period of 30 min, meanwhile, a sensitivity of 97.62% for the seizure occurrence period of 50 min are yielded with a false alarm rate of 0. 25/h. The results show that this algorithm based on tensor analysis has notable performance for seizure prediction. Shasha Yuan, Jin-Xing Liu 0001, Junliang Shang, Fangzhou Xu, Ling-Yun Dai |
BIBM | 1 |
| 2020 | Tensor Robust Principal Component Analysis with Low-Rank Weight Constraints for Sample ClusteringabstractWith the rapid development of the next-generation sequencing technology, a large amount of genomics information has been obtained. The scale of biological sequencing data is particularly large and complex. The tensor robust principal component analysis (TRPCA) method can effectively preserve the spatial structure of tensor data, so it has received extensive attention. However, the low-rank tensor obtained by TRPCA may be damaged to a certain extent. To solve this problem, this paper proposes a model for weighting low-rank data based on the method of TRPCA. This model has an additional constraint penalty term that can repair corrupted low-rank data and the effective information in it can be fully utilized. In addition, the norm is used to constrain the sparse tensor to make the sparse effect better. In the experimental part, TRPCA model clusters samples by low-rank tensor. The experimental results on cancer omics data show that our method is superior to other methods. Yu-Ying Zhao, Maoli Wang, Juan Wang 0003, Shasha Yuan, Jin-Xing Liu 0001, Xiang-Zhen Kong |
BIBM | 4 |
| 2020 | IDSSIM: an lncRNA functional similarity calculation model based on an improved disease semantic similarity methodabstractBACKGROUND: It has been widely accepted that long non-coding RNAs (lncRNAs) play important roles in the development and progression of human diseases. Many association prediction models have been proposed for predicting lncRNA functions and identifying potential lncRNA-disease associations. Nevertheless, among them, little effort has been attempted to measure lncRNA functional similarity, which is an essential part of association prediction models. RESULTS: In this study, we presented an lncRNA functional similarity calculation model, IDSSIM for short, based on an improved disease semantic similarity method, highlight of which is the introduction of information content contribution factor into the semantic value calculation to take into account both the hierarchical structures of disease directed acyclic graphs and the disease specificities. IDSSIM and three state-of-the-art models, i.e., LNCSIM1, LNCSIM2, and ILNCSIM, were evaluated by applying their disease semantic similarity matrices and the lncRNA functional similarity matrices, as well as corresponding matrices of human lncRNA-disease associations coming from either lncRNADisease database or MNDR database, into an association prediction method WKNKN for lncRNA-disease association prediction. In addition, case studies of breast cancer and adenocarcinoma were also performed to validate the effectiveness of IDSSIM. CONCLUSIONS: Results demonstrated that in terms of ROC curves and AUC values, IDSSIM is superior to compared models, and can improve accuracy of disease semantic similarity effectively, leading to increase the association prediction ability of the IDSSIM-WKNKN model; in terms of case studies, most of potential disease-associated lncRNAs predicted by IDSSIM can be confirmed by databases and literatures, implying that IDSSIM can serve as a promising tool for predicting lncRNA functions, identifying potential lncRNA-disease associations, and pre-screening candidate lncRNAs to perform biological experiments. The IDSSIM code, all experimental data and prediction results are available online at https://github.com/CDMB-lab/IDSSIM . Wenwen Fan, Junliang Shang, Feng Li 0033, Shasha Yuan, Jin-Xing Liu 0001 |
BMC Bioinform. | 5 |
| 2020 | LncRNA-Disease Associations Prediction Using Bipartite Local Model With Nearest Profile-Based Association InferringabstractThere is much evidence that long non-coding RNA (lncRNA) is associated with many diseases. However, it is time-consuming and expensive to identify meaningful lncRNA-disease associations (LDAs) through medical or biological experiments. Therefore, investigating how to identify more meaningful LDAs is necessary, and at the same time it is conducive to the prevention, diagnosis and treatment of complex diseases. Considering the limitations of some current prediction models, a novel model based on bipartite local model with nearest profile-based association inferring, BLM-NPAI, is developed for predicting LDAs. This model predicts novel LDAs from the lncRNA side and the disease side, respectively. More importantly, for some lncRNAs and diseases without any association, the model can also be predicted by their nearest neighbors. Leave-one-out cross validation (LOOCV) and 5-fold cross validation are implemented for BLM-NPAI to evaluate the performance of this model. Our model is superior to current advanced methods in most cases. In addition, to verify the validity and reliability of BLM-NPAI, three disease cases and three lncRNA cases are analyzed to further evaluate BLM-NPAI. Finally, these predicted novel LDAs are confirmed by using the LncRNA-disease database. Jin-Xing Liu 0001, Ying-Lian Gao, Shasha Yuan |
IEEE J. Biomed. Health Informatics | 5 |
| 2020 | Epileptic seizure prediction based on local mean decomposition and deep convolutional neural network
Zuyi Yu, Weiwei Nie, Fangzhou Xu, Shasha Yuan, Yan Leng |
J. Supercomput. | 5 |
| 2019 | L2, 1-GRMF: an improved graph regularized matrix factorization method to predict drug-target interactionsabstractBACKGROUND: Predicting drug-target interactions is time-consuming and expensive. It is important to present the accuracy of the calculation method. There are many algorithms to predict global interactions, some of which use drug-target networks for prediction (ie, a bipartite graph of bound drug pairs and targets known to interact). Although these algorithms can predict some drug-target interactions to some extent, there is little effect for some new drugs or targets that have no known interaction. RESULTS: Since the datasets are usually located at or near low-dimensional nonlinear manifolds, we propose an improved GRMF (graph regularized matrix factorization) method to learn these flow patterns in combination with the previous matrix-decomposition method. In addition, we use one of the pre-processing steps previously proposed to improve the accuracy of the prediction. CONCLUSIONS: Cross-validation is used to evaluate our method, and simulation experiments are used to predict new interactions. In most cases, our method is superior to other methods. Finally, some examples of new drugs and new targets are predicted by performing simulation experiments. And the improved GRMF method can better predict the remaining drug-target interactions. Ying-Lian Gao, Jin-Xing Liu 0001, Ling-Yun Dai, Shasha Yuan |
BMC Bioinform. | 5 |
| 2018 | Sparse Orthogonal Nonnegative Matrix Factorization for Identifying Differentially Expressed Genes and Clustering Tumor Samples
Ling-Yun Dai, Jin-Xing Liu 0001, Mi-Xiao Hou, Shasha Yuan |
BIBM | 6 |
| 2018 | A Fast Quantum Clustering Approach for Cancer Gene Clustering
Guangshun Li, Jin-Xing Liu 0001, Ling-Yun Dai, Shasha Yuan, Ying Guo 0002 |
BIBM | 5 |
| 2018 | Identifying Characteristic Genes and Clustering via an Lp-Norm Robust Feature Selection Method for Integrated Data
Shasha Wu, Mi-Xiao Hou, Jin-Xing Liu 0001, Juan Wang 0003, Shasha Yuan |
ICIC (2) | 5 |
| 2018 | Epileptic Seizure Prediction Using Diffusion Distance and Bayesian Linear Discriminate Analysis on Intracranial EEGabstractEpilepsy is a chronic neurological disorder characterized by sudden and apparently unpredictable seizures. A system capable of forecasting the occurrence of seizures is crucial and could open new therapeutic possibilities for human health. This paper addresses an algorithm for seizure prediction using a novel feature - diffusion distance (DD) in intracranial Electroencephalograph (iEEG) recordings. Wavelet decomposition is conducted on segmented electroencephalograph (EEG) epochs and subband signals at scales 3, 4 and 5 are utilized to extract the diffusion distance. The features of all channels composing a feature vector are then fed into a Bayesian Linear Discriminant Analysis (BLDA) classifier. Finally, postprocessing procedure is applied to reduce false prediction alarms. The prediction method is evaluated on the public intracranial EEG dataset, which consists of 577.67[Formula: see text]h of intracranial EEG recordings from 21 patients with 87 seizures. We achieved a sensitivity of 85.11% for a seizure occurrence period of 30[Formula: see text]min and a sensitivity of 93.62% for a seizure occurrence period of 50[Formula: see text]min, both with the seizure prediction horizon of 10[Formula: see text]s. Our false prediction rate was 0.08/h. The proposed method yields a high sensitivity as well as a low false prediction rate, which demonstrates its potential for real-time prediction of seizures. Shasha Yuan |
Int. J. Neural Syst. | 1 |
| 2016 | An Improved Sparse Representation over Learned Dictionary Method for Seizure DetectionabstractAutomatic seizure detection has played an important role in the monitoring, diagnosis and treatment of epilepsy. In this paper, a patient specific method is proposed for seizure detection in the long-term intracranial electroencephalogram (EEG) recordings. This seizure detection method is based on sparse representation with online dictionary learning and elastic net constraint. The online learned dictionary could sparsely represent the testing samples more accurately, and the elastic net constraint which combines the 11-norm and 12-norm not only makes the coefficients sparse but also avoids over-fitting problem. First, the EEG signals are preprocessed using wavelet filtering and differential filtering, and the kernel function is applied to make the samples closer to linearly separable. Then the dictionaries of seizure and nonseizure are respectively learned from original ictal and interictal training samples with online dictionary optimization algorithm to compose the training dictionary. After that, the test samples are sparsely coded over the learned dictionary and the residuals associated with ictal and interictal sub-dictionary are calculated, respectively. Eventually, the test samples are classified as two distinct categories, seizure or nonseizure, by comparing the reconstructed residuals. The average segment-based sensitivity of 95.45%, specificity of 99.08%, and event-based sensitivity of 94.44% with false detection rate of 0.23/h and average latency of -5.14 s have been achieved with our proposed method. Shasha Yuan |
Int. J. Neural Syst. | 3 |
| 2016 | Epileptic Seizure Detection with Log-Euclidean Gaussian Kernel-Based Sparse RepresentationabstractEpileptic seizure detection plays an important role in the diagnosis of epilepsy and reducing the massive workload of reviewing electroencephalography (EEG) recordings. In this work, a novel algorithm is developed to detect seizures employing log-Euclidean Gaussian kernel-based sparse representation (SR) in long-term EEG recordings. Unlike the traditional SR for vector data in Euclidean space, the log-Euclidean Gaussian kernel-based SR framework is proposed for seizure detection in the space of the symmetric positive definite (SPD) matrices, which form a Riemannian manifold. Since the Riemannian manifold is nonlinear, the log-Euclidean Gaussian kernel function is applied to embed it into a reproducing kernel Hilbert space (RKHS) for performing SR. The EEG signals of all channels are divided into epochs and the SPD matrices representing EEG epochs are generated by covariance descriptors. Then, the testing samples are sparsely coded over the dictionary composed by training samples utilizing log-Euclidean Gaussian kernel-based SR. The classification of testing samples is achieved by computing the minimal reconstructed residuals. The proposed method is evaluated on the Freiburg EEG dataset of 21 patients and shows its notable performance on both epoch-based and event-based assessments. Moreover, this method handles multiple channels of EEG recordings synchronously which is more speedy and efficient than traditional seizure detection methods. Shasha Yuan |
Int. J. Neural Syst. | 1 |
| 2015 | Kernel Collaborative Representation-Based Automatic Seizure Detection in Intracranial EEGabstractAutomatic seizure detection is of great significance in the monitoring and diagnosis of epilepsy. In this study, a novel method is proposed for automatic seizure detection in intracranial electroencephalogram (iEEG) recordings based on kernel collaborative representation (KCR). Firstly, the EEG recordings are divided into 4s epochs, and then wavelet decomposition with five scales is performed. After that, detail signals at scales 3, 4 and 5 are selected to be sparsely coded over the training sets using KCR. In KCR, l2-minimization replaces l1-minimization and the sparse coefficients are computed with regularized least square (RLS), and a kernel function is utilized to improve the separability between seizure and nonseizure signals. The reconstructed residuals of each EEG epoch associated with seizure and nonseizure training samples are compared and EEG epochs are categorized as the class that minimizes the reconstructed residual. At last, a multi-decision rule is applied to obtain the final detection decision. In total, 595 h of iEEG recordings from 21 patients with 87 seizures are employed to evaluate the system. The average sensitivity of 94.41%, specificity of 96.97%, and false detection rate of 0.26/h are achieved. The seizure detection system based on KCR yields both a high sensitivity and a low false detection rate for long-term EEG. Shasha Yuan, Xueli Li, Xiuhe Zhao, Jiwen Wang |
Int. J. Neural Syst. | 1 |
| 2015 | Multifractal Analysis and Relevance Vector Machine-Based Automatic Seizure Detection in Intracranial EEGabstractAutomatic seizure detection technology is of great significance for long-term electroencephalogram (EEG) monitoring of epilepsy patients. The aim of this work is to develop a seizure detection system with high accuracy. The proposed system was mainly based on multifractal analysis, which describes the local singular behavior of fractal objects and characterizes the multifractal structure using a continuous spectrum. Compared with computing the single fractal dimension, multifractal analysis can provide a better description on the transient behavior of EEG fractal time series during the evolvement from interictal stage to seizures. Thus both interictal EEG and ictal EEG were analyzed by multifractal formalism and their differences in the multifractal features were used to distinguish the two class of EEG and detect seizures. In the proposed detection system, eight features (α0, α(min), α(max), Δα, f(α(min)), f(α(max)), Δf and R) were extracted from the multifractal spectrums of the preprocessed EEG to construct feature vectors. Subsequently, relevance vector machine (RVM) was applied for EEG patterns classification, and a series of post-processing operations were used to increase the accuracy and reduce false detections. Both epoch-based and event-based evaluation methods were performed to appraise the system's performance on the EEG recordings of 21 patients in the Freiburg database. The epoch-based sensitivity of 92.94% and specificity of 97.47% were achieved, and the proposed system obtained a sensitivity of 92.06% with a false detection rate of 0.34/h in event-based performance assessment. Shasha Yuan |
Int. J. Neural Syst. | 3 |
| 2015 | Iris recognition based on a novel variation of local binary pattern
Shasha Yuan |
Vis. Comput. | 3 |
| 2014 | Epileptic EEG Classification Based on Kernel Sparse RepresentationabstractThe automatic identification of epileptic EEG signals is significant in both relieving heavy workload of visual inspection of EEG recordings and treatment of epilepsy. This paper presents a novel method based on the theory of sparse representation to identify epileptic EEGs. At first, the raw EEG epochs are preprocessed via Gaussian low pass filtering and differential operation. Then, in the scheme of sparse representation based classification (SRC), a test EEG sample is sparsely represented on the training set by solving l1-minimization problem, and the represented residuals associated with ictal and interictal training samples are computed. The test EEG sample is categorized as the class that yields the minimum represented residual. So unlike the conventional EEG classification methods, the choice and calculation of EEG features are avoided in the proposed framework. Moreover, the kernel trick is employed to generate a kernel version of the SRC method for improving the separability between ictal and interictal classes. The satisfactory recognition accuracy of 98.63% for ictal and interictal EEG classification and for ictal and normal EEG classification has been achieved by the kernel SRC. In addition, the fast speed makes the kernel SRC suit for the real-time seizure monitoring application in the near future. Shasha Yuan, Xueli Li, Jiwen Wang, Guijuan Jia |
Int. J. Neural Syst. | 3 |