Liying Yang 0001

dblp:74/3584-1 · DBLP profile ↗
← Back
30ranked-venue papers
5as first author
27since 2021 · last 2025
0000-0002-4336-9014ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 28 · 5 first-author · 25 since 2021Artificial intelligence and machine learning · 7 · 7 since 2021
YearPublicationVenuePosition
2025 Target-Aware Cross-Subject EEG Emotion Recognition with Emotion Pattern Matching and Adversarial Adaptive Clustering
abstract
Cross-subject EEG emotion recognition is crucial yet challenged by limited fine-grained alignment and label scarcity in Semi-Supervised Domain Adaptation (SSDA). We propose EPM-AAC-TSLA, a framework addressing these issues. The Emotion Pattern Matching (EPM) module uses a bidirectional similarity and diversity-based greedy strategy to aggregate source subjects, effectively mitigating negative transfer. The Adversarial Adaptive Clustering (AAC) module performs robust, cluster-level cross-domain feature alignment. Finally, the Target-aware Source Label Adjustment (TSLA) module refines source knowledge by leveraging scarce target labels. Evaluated in a 3-shot setting on SEED and SEED-IV datasets, EPM-AAC-TSLA achieved peak average accuracies of 94.46% and 85.55%, respectively. These results confirm the framework's effectiveness and superior data efficiency in the few-shot setting for cross-subject EEG-based emotion recognition.
Jiulin Fu, Liying Yang 0001, Yubin Sun, Jingtao Du
BIBM2
2025 Dynamic Log-Determinant Gating for Cross-Subject EEG Domain Adaptation
abstract
Electroencephalography (EEG) has seen rapidly growing applications in healthcare and brain-computer interfaces, yet cross-subject generalization remains a critical challenge. Unsupervised Domain Adaptation (UDA) techniques are widely employed to mitigate domain shift, but the strong constraints introduced during adaptation often lead to feature space collapse and loss of information. Moreover, relying on empirical risk minimization with cross-entropy (CE) loss fails to effectively prevent such collapse. To address this issue, we propose a novel regularizer, Dynamic Log-Determinant Gating Loss (Dyn-LogDet). This strategy introduces a history-adaptive gating mechanism based on the log-determinant of the feature Gram matrix, such that the penalty is applied only when the current feature diversity drops below a historical baseline. This ultimately contributing to better generalization performance. We validate Dyn-LogDet on two representative categories of UDA approaches: (i) explicit distribution alignment, represented by the maximum mean discrepancy (MMD)-based regularization method, and (ii) adversarial distribution alignment, represented by adversarial learning approaches(ADV), and conduct experiments on the SEED and SEED-IV datasets. In both cases, Dyn-LogDet yields consistent and improvements in performance, demonstrating its efficacy in enhancing cross-subject generalization.
Yubin Sun, Liying Yang 0001, Jiulin Fu, Jingtao Du
BIBM2
2025 Dual-Maxup: A Dual-Maximization Data Augmentation Approach for Cross-Subject EEG Emotion Recognition via Meta-Transfer Learning
abstract
Electroencephalography (EEG) emotion recognition is a vital task in affective computing, but its progress is hampered by the scarcity of labeled data and the significant intersubject variability of EEG signals. While data augmentation and meta-transfer learning have been independently explored to address these issues, their effective integration remains a challenge. Traditional augmentation methods, which are applied uniformly, often fail to generalize in cross-subject scenarios and can disrupt the intrinsic task structure of meta-learning. To overcome this, we introduce Dual-Maxup, a novel dualmaximization data augmentation strategy specifically designed for the meta-transfer learning method. Our approach begins by constructing a comprehensive augmentation pool by combining five standard EEG augmentation methods with various metaaugmentation modes. The core of Dual-Maxup is an adversarial dual-stage selection mechanism. In the inner loop, it adaptively chooses the augmented support samples that produce the highest classification loss for the base learner, forcing it to adapt to more challenging variations. In the outer loop, it selects the augmented query samples that incur the highest loss for the meta learner, ensuring that the model's meta-parameters are updated to handle the most difficult evaluation scenarios. This adversarial selection process continuously exposes the model to hard-to-classify samples, thereby maximizing its robustness and generalization ability. We validate our approach through comprehensive experiments on the benchmark DEAP dataset, demonstrating that Dual-Maxup consistently outperforms state-of-the-art methods in cross-subject emotion recognition. The source code is available at https://github.com/qwangwl/DualMaxUp.
Liying Yang 0001, Qian Zhang 0074
BIBM2
2025 TPC-AMDA: Augmented Multi-Source Domain Adaptation Based on TreePurgeCluster for Cross-Subject EEG Emotion Recognition
abstract
Due to substantial individual variability in EEG signals, cross-subject EEG emotion recognition often suffers from poor generalizability. Although domain adaptation is widely used, single-source domain adaptation neglects the heterogeneity among distinct source subjects, while multi-source domain adaptation suffers from domain conflicts and negative transfer effects when managing multiple heterogeneous source domains simultaneously. To address these issues, we propose an augmented multi-source domain adaptation model based on TreePurgeCluster (TPC-AMDA), which incorporates TreePurgeCluster for source domain clustering and combines augmented multisource domain adversarial learning to mitigate domain conflicts, reduce computational overhead, and improve model stability. Specifically, we first employ EEG-Mixup data augmentation to generate more diverse feature samples. Next, we propose the TreePurgeCluster method to cluster different source subjects into multiple source domains, preserving the core characteristics of each domain while reducing excessive inter-domain differences. Finally, we perform multi-source domain adaptation, aligning each source domain with the target domain to minimize domain divergence and enhance the generalization ability of the model. Experiments on three benchmark datasets validate the superior performance of our approach achieving an accuracy of 93.70 % on SEED, 80.80% on SEED-IV, and 65.01% and 70.36% on the valence and arousal dimensions of DEAP, respectively.
Yumeng Ye, Liying Yang 0001, Qinyu Hai
BIBM2
2025 DDSPR: Dynamic Domain Selection and Pseudo-label Refinement for Cross-Subject EEG-based Emotion Recognition
Qinyu Hai, Liying Yang 0001, Yumeng Ye, Jingtao Du, Huanyu He
CogSci2
2025 AC-CDCN: A Cross-Subject EEG Emotion Recognition Model with Anti-Collapse Domain Generalization
Yubin Sun, Liying Yang 0001, Huanyu He, Jingtao Du
CogSci2
2025 JMS2A: Joint Multi-source Domain and Two-step Alignment Strategy for Cross-subject EEG Emotion Recognition
Liying Yang 0001, Jingtao Du, Huanyu He
CogSci2
2025 Fine-grained label propagation via density-based prototype matching for cross-subject EEG emotion recognition
Liying Yang 0001, Qian Zhang 0074, Jingtao Du, Yumeng Ye
Knowl. Based Syst.2
2024 Cross-Subject Emotion Classification with Residual Pseudo-Label Distance-aware Dual-Classifier
abstract
Emotion plays a crucial role in information exchange and decision-making processes. Emotion recognition based on electroencephalography (EEG) captures and analyzes brain electrical activity, effectively reflecting emotional characteristics and providing unique advantages for constructing intelligent emotional systems. However, due to significant differences in feature distribution between the source and target domains, traditional models struggle with generalization in cross-subject tasks. To address this issue, this paper proposes a new model for EEG-based emotion analysis that employs the Pseudo-Label Distance-aware Dual-Classifier (PL-DDC) strategy, which relies on a Residual Temporal-Frequency Analysis (RTFA) structure. We name this model Residual Pseudo-Label Distance-aware Dual-Classifier (RPL-DDC). The RTFA module effectively captures mixed dependencies in complex time-frequency signals, while the PL-DDC strategy introduces a distance loss between the source and target domains. By combining classifier output discrepancies and a pseudo-label mechanism, it progressively reduces the feature distribution gap between the two domains, thereby enhancing the model’s classification performance. Experimental results on the SEED and SEED-IV datasets show that the proposed model achieves accuracy of 89.03% ± 5.03 on the SEED dataset and 75.49% ± 10.44 on the SEED-IV dataset. Compared to traditional baseline models, these results validate the effectiveness of the proposed model in cross-subject emotion recognition.
Jingtao Du, Liying Yang 0001, Huanyu He, Jiulin Fu
BIBM2
2024 MDAC: EEG Emotion Recognition with Multi-Scale Dual Attention Capsule Network
abstract
In recent years, deep learning has exhibited significant prowess in the field of EEG-based affective recognition. The attention mechanism has always been a focal point of interest within the domain of deep learning. However, existing EEG analysis techniques still face challenges in accurately pinpointing emotion-related signals across both spatial and temporal scales. We propose a novel model, named MDAC, which is based on a multi-scale dual attention and a capsule network tuned for EEG data, to address the aforementioned issue. Initially, the MDA module applies varying scales of perceptual fields to the data and integrates them to simultaneously obtain attention weights at the pixel level and the sampling rate level, achieving precise weighting of EEG signals in both spatial and temporal resolutions. Furthermore, we have increased the number of convolutional channels and the dimensionality of the primary capsules in CapsNet to better align with the characteristics of EEG data, proposing a structural configuration more apt for EEG-based affective recognition. We conducted subject-dependent and subject-independent experiments on the DEAP dataset to validate our model. In the subject-dependent experiments, the accuracy rates for both the valence and arousal dimensions were 99.58%. In the subject-independent experiments, the accuracy rates for the valence and arousal dimensions were 98.15% and 98.04%, respectively. The experimental results corroborate the efficacy of the method we proposed in this paper for the task of emotion recognition.
Huanyu He, Liying Yang 0001, Jingtao Du, Qinyu Hai, Jiulin Fu
BIBM2
2024 ST-GCN: EEG Emotion Recognition via Spectral Graph and Temporal Analysis with Graph Convolutional Networks
abstract
Emotion recognition from electroencephalogram (EEG) signals is a key application in brain-computer interfaces (BCIs), but the high dimensionality and noise in EEG data pose significant challenges. Many existing approaches fail to adequately filter irrelevant information or fully capture complex inter-channel relationships and temporal dynamics, leading to suboptimal emotional representation.To address these challenges, we propose ST-GCN, a novel model that integrates spectral and temporal domain features using graph convolution for robust EEG emotion recognition. ST-GCN employs a channel information reconstruction layer, channel aggregation, and temporal feature extraction to learn discriminative representations across EEG channels and time. Evaluated on the DEAP dataset with a cross-validation setup, ST-GCN achieves state-of-the-art performance, with 98.43% accuracy for valence and 98.69% for arousal, demonstrating its effectiveness in EEG-based emotion recognition.
Chengchuang Tang, Liying Yang 0001, Jingtao Du, Qian Zhang 0074
BIBM2
2024 Cross-Subject Emotion Classification based on Dual-Attention Mechanism and Meta-Transfer Learning
Qian Zhang 0074, Liying Yang 0001
CogSci2
2024 An AER-based spiking convolution neural network system for image classification with low latency and high energy efficiency
Lichen Feng, Hongwei Shan, Liying Yang 0001, Zhangming Zhu
Neurocomputing4
2023 BiCCT: A Compact Convolutional Transformer for EEG Emotion Recognition
abstract
Emotion is a manifestation of human’s internal psychological and physiological reactions. Understanding and recognizing emotions is one of the important ways to understand human behavior and human-computer interaction. However, with the widespread application of deep learning in the field of EEG emotion recognition, the number of parameters and model size have increased accordingly. In this paper, we combined the Bi-hemisphere asymmetry theory and Compact Convolutional Transformer to propose a model named BiCCT to recognize emotions, which has fewer training parameters and can achieve higher recognition performance. We first constructed three different matrices of the recorded EEG information according to the international 10-20 system to preserve the temporal information and spatial information of the EEG signals. Next, we applied an improved Transformer architecture, which achieves fewer model parameters and a lightweight structure through the token pooling module and the Convolutional Tokenizer module. We conducted a set of subject-dependent and a set of subject-dependent shuffle experiments on the DEAP dataset. The first set of experiments used a subject-known movie to predict a completely unknown movie. In the second set of experiments, all the data of the subject were randomly divided into training set and test set. We obtained 67.42% for valence and 67.81% for arousal in the first set of experiments. We achieved 94.41% accuracy in the valence dimension and 95.15% accuracy in the arousal dimension in the second set of experiments. At the same time, our model parameters are only 0.17M, which is far lower than other models. It means that our model is lighter and faster in training speed, and has the ability to be deployed in some scenarios with limited computing resources potential.
Liying Yang 0001, Chengchuang Tang, Qian Zhang 0074, Huanyu He
BIBM2
2023 LUR: An Online Learning Model for EEG Emotion Recognition
abstract
Emotion recognition based on EEG (Electroen-cephalogram) has been widely used in may scenarios, such as Brain-computer interface, medical health and entertainment, etc. However, large differences exist in subjects due to individual characteristics, and EEG data of different time periods usually distribute inconsistently, which hinder the further development of EEG-based research. In this paper, we propose a LUR model to partially address these challenges. In the first stage, we extracted candidate features based on neuroscience research and learned them using SVM or Naive Bayes algorithm. If the performance during the validation phase is satisfactory, we keep the features. Otherwise we replace the candidate features and use an online learning method named LUR(Learn, Unlearn, and Relearn), which continuously prunes the irrelevant connections of the current data and retains important connections to constructing a model more suitable for the subject. Because the proposed model is based on streaming data, it can continuously relearn and solve the problem that the input data is not i.i.d. to a certain extent. Experiments were conducted on two public datasets, DEAP and DREAMER, and competitive results were obtained.
Liying Yang 0001, Qian Zhang 0074, Jianing Xi, Chengchuang Tang
BIBM2
2023 EEG-MLP: An all-MLP Architecture for EEG Emotion Recognition
abstract
Emotion recognition based on EEG has attracted widespread research interest in the field of brain-computer interfaces. To extract EEG intra- and inter-channel features and find discriminative representations for EEG emotion recognition, we propose EEG-Multilayer Perceptron (EEG-MLP) architecture. EEG-MLP is completely composed of MLPs and mainly consists of two modules, one is a temporal mixer that captures intra-channel (temporal) information, and the other is a channel mixer that captures inter-channel information. The two modules learn knowledge in a parallel manner, and then their outputs are fused to extract global information and classify EEG emotions. We conduct extensive experiments on DEAP dataset. EEG-MLP is first compared with five inter-channel interaction models (related to CNN or GCN) to verify its effectiveness. Then, five other models with similar architecture to EEG-MLP were also contrasted. Experimental results show that EEG-MLP achieves the best performance among the above methods, with accuracies of 94.87% and 95.32% in the valence and arousal dimensions, respectively. In addition, it has a strong discrimination ability for complex categories, and has low requirements for storage resources.
Dunhui Liu, Liying Yang 0001, Pei Ni, Haoxuan Sun, Qian Zhang 0074, Chengchuang Tang
BIBM2
2023 MEEG-Transformer: Transformer Network based on Multi-domain EEG for Emotion Recognition
abstract
Emotion recognition is a trending topic for research in the area of the brain computer interface (BCI). As an effective signal source, EEG(Electroencephalogram) is widely used in emotion recognition tasks, from which multiple features can be extracted in different domains, such as time domain and frequency domain. However, how to make full use of multiple domain features has become a challenge. In this study, we propose a transformer network for emotion recognition based on Multi-domain EEG features, named MEEG-Transformer. MEEG-Transformer can effectively capture the spatial information with the convolution layer, mine unique information within each domain, and explore the complementary information between features from different domains using self-attention mechanism. Specifically, we extract the features of time domain, frequency domain and wavelet domain respectively, construct the two-dimensional feature matrix of three domains based on the 10-20 system, and merge the three matrices into multi-domain EEG features. Using the DEAP dataset to perform experiments, the proposed model achieves 96.8% and 96.0% recognition accuracy in the arousal and valence dimensions respectively. It is indicated that the proposed method has a strong inspiration for emotion recognition tasks.
Haoxuan Sun, Liying Yang 0001, Dunhui Liu, Pei Ni
BIBM2
2023 A two-stream channel reconstruction and feature attention network for EEG emotion recognition
abstract
Research on human emotions based on EEG during multimedia stimuli is an emerging field that has made significant progress in EEG-based emotion classification. However, current studies often neglect the extraction of dynamic information from EEG signals and lack the exploration of local information. Moreover, many existing models are overly complex, demanding an excessive investment in training resources and time. In this paper, we propose a novel, simple two-stream channel reconstruction and feature attention network, named CRFAE-motionNet, for EEG emotion recognition. The main advantage of CRFAEmotionNet is its ability to simultaneously integrate static and dynamic information from EEG signals within a unified network. Additionally, it can extract continuous EEG temporal information through channel reconstruction and utilize feature attention to further explore local information. The proposed network was evaluated using the publicly available DEAP dataset. The experimental results indicate that the proposed CRFAEmotionNet outperforms the state-of-the-art baselines, achieving the accuracy of 98.7% for valence and 98.6% for arousal.
Liying Yang 0001, Haoxuan Sun, Qian Zhang 0074, Chengchuang Tang
BIBM2
2023 User-independent Emotion Classification based on Domain Adversarial Transfer Learning
Pei Ni, Liying Yang 0001, Dunhui Liu, Si Chao, Haoxuan Sun
CogSci2
2023 A Method for Predicting DNA Motif Length Based On Deep Learning
abstract
A DNA motif is a sequence pattern shared by the DNA sequence segments that bind to a specific protein. Discovering motifs in a given DNA sequence dataset plays a vital role in studying gene expression regulation. As an important attribute of the DNA motif, the motif length directly affects the quality of the discovered motifs. How to determine the motif length more accurately remains a difficult challenge to be solved. We propose a new motif length prediction scheme named MotifLen by using supervised machine learning. First, a method of constructing sample data for predicting the motif length is proposed. Secondly, a deep learning model for motif length prediction is constructed based on the convolutional neural network. Then, the methods of applying the proposed prediction model based on a motif found by an existing motif discovery algorithm are given. The experimental results show that i) the prediction accuracy of MotifLen is more than 90% on the validation set and is significantly higher than that of the compared methods on real datasets, ii) MotifLen can successfully optimize the motifs found by the existing motif discovery algorithms, and iii) it can effectively improve the time performance of some existing motif discovery algorithms.
Qiang Yu 0003, Yana Hu, Shengpin Chen, Liying Yang 0001
IEEE ACM Trans. Comput. Biol. Bioinform.5
2022 Electroencephalogram Emotion Recognition Based on Individual Frontal Asymmetry Hypothesis
abstract
The use of Electroencephalogram(EEG) for emotion recognition has tremendous potential across psychology and biomedicine. However, how the brain generates emotions remains unclear. Inspired by neuroscience and psychology, this paper puts forward the individual frontal asymmetry hypothesis and three methods of Electroencephalogram(EEG) emotion recognition based on this potential hypothesis are introduced, which recognizes and classifies the individual’s emotion effectively with signals from only four channels out of the total 32 channels. First, all EEG signals are filtered according to the EEG frequency band. Then, taking the filtered left and right frontal lobe signal differences as the input, three different models are used for classification with leave-one-out cross-validation. For each subject, one film is used for testing and the remaining films are used for training. We verify our idea on the public database DEAP, and recognition accuracy reaches 75.39% in the valence dimension and 68.13% in the arousal dimension, respectively. Since only four EEG channels were used, it greatly improves the operation efficiency and saves the running time. This work might be a demonstration that emotion recognition using individual frontal asymmetry hypothesis is effective, and it provides a potential direction for emotion recognition using portable EEG acquisition devices.
Liying Yang 0001, Pei Ni
BIBM2
2022 EEG emotion recognition via Identity based Multi-gate Mixture-of-Experts network
abstract
Empowering computer systems to automatically recognize human emotions has become an urgent need in the field of human-computer interaction (HCI). Two-dimensional emotion (Valence-Arousal) models are commonly used to represent emotions. Up to now, the correlation between emotion dimensions has rarely been investigated, and subject-independent EEG emotion recognition is still a challenging task. For this purpose, we introduce multi-task learning (MTL) into EEG emotion recognition. MTL learns different emotion dimensions simultaneously and extracts correlation information between dimensions in task-sharing space to coordinate the optimization of multiple emotion dimensions. We further propose Identity based Multi-gate Mixture-of-Experts (IDMMOE), which allocates part of model subspace for each subject in a customized manner according to the subject’s identity. Extensive experiments were conducted on DEAP dataset. Three MTL models were implemented: Shared-Bottom, Multi-gate Mixture-of-Experts, and Customized Gate Control respectively. They were compared with a single-task learning model trained separately on valence and arousal. Experimental results demonstrate that two emotion dimensions are intrinsically related, and MTL acquires such correlation information and improves prediction accuracy in both emotion dimensions. In addition, IDMMOE achieves average accuracies of 89.5% and 89.7% for valence and arousal respectively and it is effective for subject-independent experiment.
Liying Yang 0001, Dunhui Liu, Si Chao, Pei Ni, Haoxuan Sun
BIBM1
2022 Greedy-mRMR: An emotion recognition algorithm based on EEG using greedy algorithm
abstract
Electroencephalography (EEG)-based emotion recognition methods mostly adopt features as many as possible to achieve high performance. However, this is not feasible and flexible for portable or wearable helmet devices because of the limited computing resource and requirement of real-time performance. Aiming at improving EEG emotion recognition performance with features as few as possible, this paper firstly designed a feature called Relative Intensity Ratio Entropy (RIRE), then proposed a novel feature selection method based on greedy algorithm and Max-Relevance and Min-Redundancy(mRMR), termed as Greedy-mRMR. Greedy algorithm is adopted to take top features in the ranking list, and mRMR algorithm is used to select features that are most relevant and minimum redundant. Greedy-mRMR algorithm also uses dynamic termination threshold to guarantee recognition accuracy in each iteration. Experiments were carried on DEAP dataset. SVM classifier achieved average classification accuracy of 91.16% in valence and 91.70% in arousal, by extracting RIRE feature and selecting with Greedy-mRMR algorithm. Experimental results show that RIRE is suitable for EEG emotion recognition and Greedy-mRMR outperforms state-of-the-art methods.
Liying Yang 0001, Si Chao, Dunhui Liu, Xiguo Yuan
BIBM1
2021 Intelligent Feature Selection for EEG Emotion Classification
abstract
Emotion classification plays a critical role in the development of human-computer interaction. EEG (Electroencephalogram) signal is an important information source, from which different features have been extracted for the study of emotion states. However, there is a large amount of redundant and irrelevant information in EEG signals, which directly interferes with emotion classification. Aiming at selecting EEG emotional features precisely and efficiently, this paper proposed a feature selection framework, termed MIGA (Mutual Information and Genetic Algorithm). It combined mutual information and genetic algorithm to improve the quality of feature subset based on three perspectives, that is correlation, contribution and synergistic effect between features. A new emotional feature IQI (Intensity Quantity of Information) was also designed in this paper. IQI is able to mine the intensity information of EEG signals, and enlarge the samples’ distance as a result. Experiments were carried on DEAP (A Database for Emotion Analysis using Physiological Signals). Results show that the classification accuracy of IQI is 5% higher than that of the traditional frequency domain feature, and MIGA reduces feature amount by 2/3 while ensuring classification accuracy. For DEAP, classification accuracy of MIGA in valence and arousal reached 88.55% and 88.14% respectively. It is indicated that, compared with existing methods, MIGA improves emotion recognition with much fewer features from EEG.
Liying Yang 0001, Si Chao
BIBM1
2021 A Grouped Dynamic EEG Channel Selection Method for Emotion Recognition
abstract
EEG signals directly reflect the active state of the brain, so they are widely used for emotion recognition. At present, many researchers have achieved noteworthy results by using multi-channel EEG signals. However, too many EEG channels will cause slow transmission, high experimental costs, and low efficiency. This paper proposed a grouped dynamic EEG channel selection method based on ReliefF and random forest (RF), termed GDCSBR. We divided all channels into four groups, which provided more choices and flexibility to subsequent dynamic channel selection. GDCSBR selected channels iteratively. With the iteration increased, the number of alternative channels decreased. At each iteration, we adopted the strategy with a minor loss of recognition accuracy. Finally, the results on subject-independent data were taken as the final choice, since there were relatively large differences between subjects. Experiments were carried out on the DEAP dataset. The recognition accuracy for valence reaches 81.27% while 10 channels are selected. As for arousal scale, 11 channels can obtain 82.36% of classification accuracy. In addition, we found that high-frequency bands play a crucial part in emotion recognition, and the selected channels were mostly located in the frontal and parietal lobes. These findings are coincident with previous work. Experimental results demonstrate the effectiveness of the proposed method.
Liying Yang 0001, Si Chao, Pei Ni, Dunhui Liu
BIBM1
2021 STIC: Predicting Single Nucleotide Variants and Tumor Purity in Cancer Genome
abstract
Single nucleotide variant (SNV) plays an important role in cellular proliferation and tumorigenesis in various types of human cancer. Next-generation sequencing (NGS) has provided high-throughput data at an unprecedented resolution to predict SNVs. Currently, there exist many computational methods for either germline or somatic SNV discovery from NGS data, but very few of them are versatile enough to adapt to any situations. In the absence of matched normal samples, the prediction of somatic SNVs from single-tumor samples becomes considerably challenging, especially when the tumor purity is unknown. Here, we propose a new approach, STIC, to predict somatic SNVs and estimate tumor purity from NGS data without matched normal samples. The main features of STIC include: (1) extracting a set of SNV-relevant features on each site and training the BP neural network algorithm on the features to predict SNVs; (2) creating an iterative process to distinguish somatic SNVs from germline ones by disturbing allele frequency; and (3) establishing a reasonable relationship between tumor purity and allele frequencies of somatic SNVs to accurately estimate the purity. We quantitatively evaluate the performance of STIC on both simulation and real sequencing datasets, the results of which indicate that STIC outperforms competing methods.
Xiguo Yuan, Haiyong Zhao, Liying Yang 0001, Shuzhen Wang, Jianing Xi
IEEE ACM Trans. Comput. Biol. Bioinform.4
2021 CNV_IFTV: An Isolation Forest and Total Variation-Based Detection of CNVs from Short-Read Sequencing Data
abstract
Accurate detection of copy number variations (CNVs) from short-read sequencing data is challenging due to the uneven distribution of reads and the unbalanced amplitudes of gains and losses. The direct use of read depths to measure CNVs tends to limit performance. Thus, robust computational approaches equipped with appropriate statistics are required to detect CNV regions and boundaries. This study proposes a new method called CNV_IFTV to address this need. CNV_IFTV assigns an anomaly score to each genome bin through a collection of isolation trees. The trees are trained based on isolation forest algorithm through conducting subsampling from measured read depths. With the anomaly scores, CNV_IFTV uses a total variation model to smooth adjacent bins, leading to a denoised score profile. Finally, a statistical model is established to test the denoised scores for calling CNVs. CNV_IFTV is tested on both simulated and real data in comparison to several peer methods. The results indicate that the proposed method outperforms the peer methods. CNV_IFTV is a reliable tool for detecting CNVs from short-read sequencing data even for low-level coverage and tumor purity. The detection results on tumor samples can aid to evaluate known cancer genes and to predict target drugs for disease diagnosis.
Xiguo Yuan, Jianing Xi, Liying Yang 0001, Junliang Shang, Junbo Duan
IEEE ACM Trans. Comput. Biol. Bioinform.4
2020 CONDEL: Detecting Copy Number Variation and Genotyping Deletion Zygosity from Single Tumor Samples Using Sequence Data
abstract
Characterizing copy number variations (CNVs) from sequenced genomes is a both feasible and cost-effective way to search for driver genes in cancer diagnosis. A number of existing algorithms for CNV detection only explored part of the features underlying sequence data and copy number structures, resulting in limited performance. Here, we describe CONDEL, a method for detecting CNVs from single tumor samples using high-throughput sequence data. CONDEL utilizes a novel statistic in combination with a peel-off scheme to assess the statistical significance of genome bins, and adopts a Bayesian approach to infer copy number gains, losses, and deletion zygosity based on statistical mixture models. We compare CONDEL to six peer methods on a large number of simulation datasets, showing improved performance in terms of true positive and false positive rates, and further validate CONDEL on three real datasets derived from the 1000 Genomes Project and the EGA archive. CONDEL obtained higher consistent results in comparison with other three single sample-based methods, and exclusively identified a number of CNVs that were previously associated with cancers. We conclude that CONDEL is a powerful tool for detecting copy number variations on single tumor samples even if these are sequenced at low-coverage.
Xiguo Yuan, Liying Yang 0001, Junbo Duan, Meihong Gao
IEEE ACM Trans. Comput. Biol. Bioinform.4
2017 Analysis of breast cancer subtypes by AP-ISA biclustering
abstract
BACKGROUND: Gene expression profiling has led to the definition of breast cancer molecular subtypes: Basal-like, HER2-enriched, LuminalA, LuminalB and Normal-like. Different subtypes exhibit diverse responses to treatment. In the past years, several traditional clustering algorithms have been applied to analyze gene expression profiling. However, accurate identification of breast cancer subtypes, especially within highly variable LuminalA subtype, remains a challenge. Furthermore, the relationship between DNA methylation and expression level in different breast cancer subtypes is not clear. RESULTS: In this study, a modified ISA biclustering algorithm, termed AP-ISA, was proposed to identify breast cancer subtypes. Comparing with ISA, AP-ISA provides the optimized strategy to select seeds and thresholds in the circumstance that prior knowledge is absent. Experimental results on 574 breast cancer samples were evaluated using clinical ER/PR information, PAM50 subtypes and the results of five peer to peer methods. One remarkable point in the experiment is that, AP-ISA divided the expression profiles of the luminal samples into four distinct classes. Enrichment analysis and methylation analysis showed obvious distinction among the four subgroups. Tumor variability within the Luminal subtype is observed in the experiments, which could contribute to the development of novel directed therapies. CONCLUSIONS: Aiming at breast cancer subtype classification, a novel biclustering algorithm AP-ISA is proposed in this paper. AP-ISA classifies breast cancer into seven subtypes and we argue that there are four subtypes in luminal samples. Comparison with other methods validates the effectiveness of AP-ISA. New genes that would be useful for targeted treatment of breast cancer were also obtained in this study.
Liying Yang 0001, Yunyan Shen, Xiguo Yuan, Jianhua Wei
BMC Bioinform.1
2012 Geometric Linear Regression and Geometric Relation
Kaijun Wang, Liying Yang 0001
ICIC (1)2