Guan Ning Lin

dblp:22/7237 · DBLP profile ↗
← Back
16ranked-venue papers
3as first author
8since 2021 · last 2026
0000-0001-9496-0149ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 16 · 3 first-author · 8 since 2021
YearPublicationVenuePosition
2026 GT-Mamba: a Topology-Aware Graph-State space model for robust and interpretable epigenetic age prediction
abstract
MOTIVATION: Current epigenetic clocks face a trade-off between predictive accuracy and biological interpretability, often relying on dataset-specific correction to generalize across cohorts. We propose GT-Mamba, a novel architecture that integrates a Structure-Aware Graph Transformer with the Mamba state space model. This design captures CpG topological correlations and genome-wide long-range dependencies. RESULTS: GT-Mamba demonstrates strong out-of-the-box robustness across heterogeneous independent validation cohorts, achieving a weighted average MAE of 4.43 years. Notably, it effectively generalizes to EPIC 850k arrays despite partial feature missingness, and maintains consistent performance across homologous age distribution shifts (MAE 2.94 years in a young cohort). Ablation studies confirm that graph topology contributes to improved robustness against noise. Mechanistic analysis suggests that the model captures methylation patterns associated with both developmental and functional processes. AVAILABILITY: Source code and pre-trained models are freely available at https://github.com/NENUBioCompute/GT-Mamba and archived on Zenodo (DOI: 10.5281/zenodo.19703155).
Yanting Tong, Qu Jing, Guan Ning Lin
Bioinform.6
2026 EpiSight: an intelligent EEG automated analysis platform for clinical epilepsy diagnosis
Liujinxiang Zhu, Guan Ning Lin
Frontiers Comput. Sci.4
2025 Mutation-drug sensitivity data resource (MDSDR): a comprehensive resource for studying and addressing drug resistance
Yihang Bao, Shunying Yu, Huafang Li, Guan Ning Lin
Frontiers Comput. Sci.6
2024 EpilepsyNet: A Self-Calibrating Deep Learning Network for Accurate and Robust Seizure Detection
abstract
Epilepsy affects millions of people worldwide and poses significant challenges for diagnosis and treatment. While traditional EEG analysis remains essential for seizure detection, it is often time-intensive and requires specialized expertise. To overcome these limitations, we proposed EpilepsyNet, an innovative deep learning architecture for automatic seizure detection. EpilepsyNet incorporates a self-calibration mechanism that enhances feature extraction by expanding the receptive field and enabling dynamic interactions across multiple EEG channels. Using neonatal EEG data, we validated EpilepsyNet’s effectiveness through five-fold cross-validation, demonstrating high classification accuracy and robustness across various performance metrics. Ablation experiments further underscored the important role of the self-calibrated convolution, as removing this component led to a notable decline in accuracy. These findings suggest that EpilepsyNet provides a scalable and reliable solution for seizure detection, with the potential to significantly improve clinical interventions by automating and optimizing the analysis of EEG data.
Shi Chang, Zhenhong Ye, Yihang Bao, Jingtong Zhao, Guan Ning Lin
BIBM7
2024 PSGAnalyzer: The Intelligent Sleep Analysis Platform Empowering Sleep Disorder Diagnosis and Treatment
abstract
The growing prevalence of sleep disorders and the need for efficient diagnostic tools have heightened interest in automated sleep analysis solutions. Traditional polysomnography (PSG) analysis requires intensive manual interpretation, limiting its scalability in the face of increasing diagnostic demand. PSGAnalyzer addresses this challenge by providing a comprehensive and automated platform that integrates multiple analysis modules. It includes (a) a Sleep Stage Classification module, which accurately segments sleep stages following American Academy of Sleep Medicine (AASM) standards, supported by advanced artifact-removal techniques for high-quality data; (b) a Blood Oxygen Saturation Analysis module that tracks SpO₂ levels and highlights potential hypoxic events to aid in respiratory disorder assessment; and (c) a Heart Rate Monitoring module, which analyzes heart rate variability across sleep stages to provide insights into cardiovascular health and autonomic nervous system activity during sleep. By offering a detailed, efficient, and fully automated PSG analysis, PSGAnalyzer is poised to support clinicians in diagnosing and monitoring sleep disorders more effectively, meeting the growing clinical demand for high-throughput, data-driven sleep health assessments.
Dongbin Lyu, Ruiyi Qian, Liujinxiang Zhu, Guan Ning Lin
BIBM7
2024 MEMO-stab: Sequence-based Annotation of Mutation Effect on Transmembrane Protein Stability with Protein Language Model-Driven Machine Learning
abstract
Transmembrane proteins are pivotal drug targets. Accurately predicting the impact of mutations on their stability can serve as an early indicator of alterations in protein function. Existing prediction methods based on protein crystal structures lack the capability for large-scale annotation of mutation effects on transmembrane proteins. In this study, we present a sequence-based protein language model-driven machine learning framework, MEMO-stab, to predict the effect of mutations on the stability of transmembrane proteins. MEMO-stab is an end-to-end binary classifier capable of rapidly and extensively annotating whether a mutation is destabilizing. We evaluated the integrative capability of MEMO-stab with eight diverse protein language models and conducted a performance comparison against alternative existing methods. We also designed case studies to further illustrate the advantage of our methods. MEMO-stab is the first sequence-based transmembrane protein-specific prediction tool capable of achieving reasonable results for destabilizing mutation prediction.
Yihang Bao, Weidi Wang, Guan Ning Lin
BIBM6
2024 Predicting Psychosis Progression in Clinical High-Risk Individuals Using Peripheral Transcriptomic and Epigenomic Profiles: A Machine Learning Approach
abstract
Advancing the prediction of psychosis in individuals identified as clinically high risk for psychosis (CHRP), a state indicating an elevated risk of developing psychosis, requires the identification of reliable biomarkers. Here, we employ a comprehensive approach using transcriptomic and epigenomic profiling from peripheral blood samples to predict the conversion from high-risk to psychosis in a longitudinal cohort of 56 individuals. We employed ensemble tree-based machine learning techniques, specifically XGBoost and LightGBM models, to analyze and select biomarkers capable of forecasting the conversion to psychosis at one-year and two-year follow-up intervals. The predictive accuracy of these models was robust, with the one-year model achieving an area under the curve (AUC) of 0.990 and the two-year model an AUC of 0.934. Our analysis identified eight gene expression markers (such AC026691.1, CA8, STXBP6 and TMEM89) and seven DNA methylation markers (such as cg14258341 and cg10663442) that were associated with a one-year conversion to psychosis. Additionally, two gene expression markers were found to predict a two-year conversion. Notably, long non-coding RNAs were validated as contributors to the predictive models. Our study also contributes to the advancement of CHR-P prediction by further distinguishing between patients who maintain symptoms and those who recover. These results affirm the potential of non-invasive blood biomarkers to stratify CHR-P individuals by their risk of conversion, address the broader spectrum of clinical outcomes, thus facilitating early and targeted intervention strategies in psychosis.
Weidi Wang, Mingxia Zhai, An Gu, Shunying Yu, Guan Ning Lin
BIBM6
2023 Probing Transmembrane Proteins Binding Domain via Multi-level Molecule Learning
abstract
The study of transmembrane proteins (TMPs) and their binding activities holds significant importance in the pharmaceutical industry. Due to their physicochemical properties, known binding information regarding TMPs remains comparatively scarce and prevented researchers to dig information from known samples. However, research into general binding structure basis can circumvent this barrier and provide more mechanism insights, which we previously demonstrated its existence and named it as TMPs binding Domain. In this study, we try to discover the TMPs binding domain more precisely. Through atomic-level heterogenous graph convolutions, we significantly improved the classification performance of binding domains. This lays the algorithmic groundwork for utilizing binding domains in the study of TMPs binding activities and further boost the drug target research or new drug development.
Yihang Bao, Yuanzhao Guo, Guan Ning Lin, Zehua Sun, Han Wang 0028
BIBM4
2019 Varying Mutational Classes Illuminate Differential Genetic Patterns Between Schizophrenia and Bipolar Disorder
abstract
Schizophrenia (SCZ) and bipolar disorder (BPD) are two severe psychiatric disorders that share considerable comorbidities in both clinicals and genetics. In both diseases, the genetic susceptibility often arises from a combination of different types of variants, suggesting that they may contribute their effects to various biological pathways. Here, we compared susceptible risk factors in SCZ and BPD impacted by different types of variants, such as de novo mutations (DNM), single nucleotide polymorphisms (SNP), and copy number variants (CNV), to explore the shared and distinct biological pathways underlying the complex structure of disease genetics in SCZ and BPD. We found that genes carrying DNMs in SCZ and genes impacted by CNVs in BPD explain disease-relevant functional enrichment patterns more than other types of variants. In addition, we evaluated the haploinsufficiency, mutation burden, and network degree of variants impacted genes. The results suggested that DNM affected genes contributed more to affected pathways than genes hit by CNVs in SCZ, not in BPD. As a whole, we investigated the impacts of various types of variants on biological pathways and provided insights on the convergent and divergent genetic patterns in SCZ and BPD.
Weidi Wang, Wu Huang, Guan Ning Lin
BIBM6
2019 Label Propagation Based Semi-supervised Feature Selection to Decode Clinical Phenotype of Huntington's Disease
Weidi Wang, Weichen Song, Guan Ning Lin
ICIC (1)5
2019 Integrative Enrichment Analysis of Intra- and Inter- Tissues' Differentially Expressed Genes Based on Perceptron
Weihao Pan, Weidi Wang, Weichen Song, Guan Ning Lin
ICIC (2)6
2018 Genome-Wide miRNA Expression Alterations in Nucleus Accumbens Provide Insights into Chronic Stress and Treatment in Depression
Weichen Song, Guan Ning Lin, Sufang Peng, Yanhua Zhang, Yifeng Shen, Huafang Li, Shunying Yu
BIBM2
2017 When loss-of-function is loss of function: assessing mutational signatures and impact of loss-of-function genetic variants
abstract
MOTIVATION: Loss-of-function genetic variants are frequently associated with severe clinical phenotypes, yet many are present in the genomes of healthy individuals. The available methods to assess the impact of these variants rely primarily upon evolutionary conservation with little to no consideration of the structural and functional implications for the protein. They further do not provide information to the user regarding specific molecular alterations potentially causative of disease. RESULTS: To address this, we investigate protein features underlying loss-of-function genetic variation and develop a machine learning method, MutPred-LOF, for the discrimination of pathogenic and tolerated variants that can also generate hypotheses on specific molecular events disrupted by the variant. We investigate a large set of human variants derived from the Human Gene Mutation Database, ClinVar and the Exome Aggregation Consortium. Our prediction method shows an area under the Receiver Operating Characteristic curve of 0.85 for all loss-of-function variants and 0.75 for proteins in which both pathogenic and neutral variants have been observed. We applied MutPred-LOF to a set of 1142 de novo vari3ants from neurodevelopmental disorders and find enrichment of pathogenic variants in affected individuals. Overall, our results highlight the potential of computational tools to elucidate causal mechanisms underlying loss of protein function in loss-of-function variants. AVAILABILITY AND IMPLEMENTATION: http://mutpred.mutdb.org. CONTACT: [email protected].
Kymberleigh A. Pagel, Vikas Pejaver, Guan Ning Lin, Hyun-Jun Nam, Matthew E. Mort, David N. Cooper, Jonathan Sebat, Lilia M. Iakoucheva, Sean D. Mooney, Predrag Radivojac
Bioinform.3
2010 SeqRate: sequence-based protein folding type classification and rates prediction
abstract
BACKGROUND: Protein folding rate is an important property of a protein. Predicting protein folding rate is useful for understanding protein folding process and guiding protein design. Most previous methods of predicting protein folding rate require the tertiary structure of a protein as an input. And most methods do not distinguish the different kinetic nature (two-state folding or multi-state folding) of the proteins. Here we developed a method, SeqRate, to predict both protein folding kinetic type (two-state versus multi-state) and real-value folding rate using sequence length, amino acid composition, contact order, contact number, and secondary structure information predicted from only protein sequence with support vector machines. RESULTS: We systematically studied the contributions of individual features to folding rate prediction. On a standard benchmark dataset, the accuracy of folding kinetic type classification is 80%. The Pearson correlation coefficient and the mean absolute difference between predicted and experimental folding rates (sec-1) in the base-10 logarithmic scale are 0.81 and 0.79 for two-state protein folders, and 0.80 and 0.68 for three-state protein folders. SeqRate is the first sequence-based method for protein folding type classification and its accuracy of fold rate prediction is improved over previous sequence-based methods. Its performance can be further enhanced with additional information, such as structure-based geometric contacts, as inputs. CONCLUSIONS: Both the web server and software of predicting folding rate are publicly available at http://casp.rnet.missouri.edu/fold_rate/index.html.
Guan Ning Lin, Zheng Wang 0025, Dong Xu 0002, Jianlin Cheng
BMC Bioinform.1
2009 Sequence-Based Prediction of Protein Folding Rates Using Contacts, Secondary Structures and Support Vector Machines
abstract
Predicting protein folding rate is useful for understanding protein folding process and guiding protein design. Most previous methods of predicting folding rate require the tertiary structure of a protein as an input. And most methods do not distinguish the different kinetic natures (two-state folding and multi-state folding) of the proteins. Here we developed a method, SeqRate, to predict both protein folding kinetic type (two-state versus multi-state) and real-value folding rate using features extracted from only protein sequence with support vector machines. On a standard benchmark dataset, the accuracy of folding kinetic type classification is 80%. The Pearson correlation coefficient and the mean absolute difference between predicted and experimental folding rates (sec-1) in the base-10 logarithmic scale are 0.81 and 0.79 for two-state protein folders, and 0.80 and 0.68 for three-state protein folders. SeqRate is the first sequence-based method for protein folding type classification and its accuracy of fold rate prediction is improved over previous sequence-based methods. Both the Web server and software of predicting folding rate are publicly available at http://casp.rnet.missouri.edu/fold_rate/index.html.
Guan Ning Lin, Zheng Wang 0025, Dong Xu 0002, Jianlin Cheng
BIBM1
2009 ComPhy: prokaryotic composite distance phylogenies inferred from whole-genome gene sets
abstract
BACKGROUND: With the increasing availability of whole genome sequences, it is becoming more and more important to use complete genome sequences for inferring species phylogenies. We developed a new tool ComPhy, 'Composite Distance Phylogeny', based on a composite distance matrix calculated from the comparison of complete gene sets between genome pairs to produce a prokaryotic phylogeny. RESULTS: The composite distance between two genomes is defined by three components: Gene Dispersion Distance (GDD), Genome Breakpoint Distance (GBD) and Gene Content Distance (GCD). GDD quantifies the dispersion of orthologous genes along the genomic coordinates from one genome to another; GBD measures the shared breakpoints between two genomes; GCD measures the level of shared orthologs between two genomes. The phylogenetic tree is constructed from the composite distance matrix using a neighbor joining method. We tested our method on 9 datasets from 398 completely sequenced prokaryotic genomes. We have achieved above 90% agreement in quartet topologies between the tree created by our method and the tree from the Bergey's taxonomy. In comparison to several other phylogenetic analysis methods, our method showed consistently better performance. CONCLUSION: ComPhy is a fast and robust tool for genome-wide inference of evolutionary relationship among genomes. It can be downloaded from http://digbio.missouri.edu/ComPhy.
Guan Ning Lin, Zhipeng Cai 0001, Guohui Lin, Sounak Chakraborty, Dong Xu 0002
BMC Bioinform.1