VLDB 2026 Research / reviewers in the wild / expert
Peilin Jia
dblp:82/7391
· DBLP profile ↗
29ranked-venue papers
11as first author
8since 2021 · last 2026
0000-0003-4523-4153ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 26 · 10 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Subsystem-Aware Stackelberg Game for Optimal Anti-Jamming Power Control in Periodically Switched Cyber-Physical SystemsabstractThis paper presents a subsystem-aware Stackelberg game framework for optimal power control in periodically switched cyber-physical systems (PSCPS) operating over wireless networks under jamming attacks. To account for subsystem switching dynamics, a switching Kalman filter is employed for mode-dependent state estimation, with corresponding error covariance matrices derived for each mode. The impact of switching is quantified using a subsystem-importance-based weighting factor, obtained by analyzing trace variations in the covariance matrices. This factor is integrated into the defender’s utility function to enable adaptive power control. First, a subsystem-aware Stackelberg game model without power constraints is formulated, and the equilibrium strategies of the defender (leader) and the attacker (follower) are obtained via the convex optimization. The framework is then extended to the power-constrained scenario, where feasibility issues arise. By analyzing the boundary and extreme points of the solution space, closed-form equilibrium strategies are derived. Simulation results show that the proposed approach enhances the resilience of PSCPS against jamming attacks in wireless communication environments. Peilin Jia, Jiyun Tian, Jie Lian 0001, Susanto Rahardja |
IEEE Internet Things J. | 1 |
| 2026 | Almost Sure Convergence of Nonhomogeneous Switching Markov Chains With Absorbing States: A Graph-Based ApproachabstractThis article investigates the problem of almost sure convergence for nonhomogeneous switching Markov chains with absorbing states. Two types of graphs are used to describe the switching walks of state components in the Markov chain under deterministic transfer-restricted switching and arbitrary switching. The transfer-restricted switching among nonhomogeneous Markov chains is characterized by a switching directed graph, while the stochastic transitions within each Markov chain are depicted by state transition component graphs. The reachability relationship between nonabsorbing and absorbing state components is then analyzed by using the Lyapunov function constructed from the state space of Markov chains. Building on the two types of graphs, cycle-dependent switching strategies for almost sure convergence to the target absorbing state are established. Furthermore, necessary and sufficient conditions for almost sure convergence to the common absorbing state under arbitrary switching are derived, which utilize the properties of absorbing Markov chains and state transition component graphs with self-loops. Finally, the effectiveness of the proposed method is validated through two examples. Jie Lian 0001, Wanchen Wang, Peilin Jia |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2025 | Finite-Time Prescribed Performance Adaptive Control for a Class of Switched Uncertain Nonlinear SystemsabstractThis paper investigates a finite-time prescribed performance adaptive control (FTPPAC) scheme for a class of full-state dependent switched uncertain nonlinear systems in the non-strict feedback form. Firstly, a model transformation is introduced to handle unknown control coefficients and unknown nonlinear functions, after which a state observer is designed to estimate the unknown state of the transformed system. Subsequently, the backstepping method and adaptive control are employed to ensure that the tracking error satisfies the pre-designed performance indices. Compared with the traditional design of prescribed performance control (PPC) for switched systems, a novel switching-dependent prescribed performance bound (SDPPB) is presented, which is specifically tailored to satisfy the requirements and characteristics of different subsystems. This design approach allows for the adjustment of the initial value of SDPPB at the switching instants based on the tracking error and subsystem change. Finally, by applying the average dwell time switching and multiple Lyapunov functions method, sufficient conditions for finite-time stability of the closed-loop system are derived. The effectiveness of the proposed control scheme is demonstrated through two simulations. Note to Practitioners—Finite-time stability analysis and quantitative evaluation of control performance are hot topics in the field of control, which play a crucial role in engineering practice. This paper proposes the finite-time prescribed performance adaptive control scheme for a class of switched uncertain nonlinear systems, which can be applied to motion tracking control of servo mechanisms, command tracking of aero-engine and attitude control of spacecraft. It is worth noting that unknown nonlinear dynamics, unknown control coefficients and system state performance constraint problems often exist in practical systems. Therefore, the model transformation is presented to address unknown control coefficients and unknown nonlinear functions. The state observer is used to estimate the unknown state of the transformed system. By designing switching-dependent prescribed performance bound, whose initial value can vary with the subsystem and tracking error, and by utilizing adaptive control method that remains continuous at the switching instant, we ensure that the tracking error satisfies the pre-designed performance indices and that all signals of the closed-loop system are bounded in finite time. In view of the growing demand for transient and steady-state performance in modern industrial production, this control scheme holds promise for practical problem-solving. Wanchen Wang, Jie Lian 0001, Peilin Jia |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2025 | A Cross-Layer Game-Theoretic Approach to Resilient Control of Networked Switched Systems Against DoS AttacksabstractThis article investigates the resilient control strategies of networked switched systems (NSSs) against denial-of-service (DoS) attacks and external disturbance. In the network layer, both the defender and the attacker allocate energy over multiple channels. Considering the impact of switching characteristic in the physical layer on the network layer, a dynamic regulating factor is proposed to adjust the total energy of the defender. To optimize the signal-to-interference-noise ratio and energy consumption simultaneously at each player's side, a multiobjective game problem is formulated. Furthermore, a nondominated sorting genetic algorithm framework is employed, incorporating the knee point selection mechanism to attain the Pareto-Nash equilibrium, based on which the optimal defense strategy can be derived to achieve resilience against DoS attacks. In the physical-layer, taking the dynamic packet loss caused by DoS attacks and external disturbance into account, an minimax controller containing control inputs and the switching signal is designed to guarantee the optimal performance for NSSs through the dynamic game-theoretic approach. Finally, the networked continuous stirred tank reactor system is provided to verify the effectiveness of the proposed method. Jie Lian 0001, Peilin Jia, Feiyue Wu |
IEEE Trans. Cybern. | 2 |
| 2024 | Comprehensive assessment of long-read sequencing platforms and calling algorithms for detection of copy number variationabstractCopy number variations (CNVs) play pivotal roles in disease susceptibility and have been intensively investigated in human disease studies. Long-read sequencing technologies offer opportunities for comprehensive structural variation (SV) detection, and numerous methodologies have been developed recently. Consequently, there is a pressing need to assess these methods and aid researchers in selecting appropriate techniques for CNV detection using long-read sequencing. Hence, we conducted an evaluation of eight CNV calling methods across 22 datasets from nine publicly available samples and 15 simulated datasets, covering multiple sequencing platforms. The overall performance of CNV callers varied substantially and was influenced by the input dataset type, sequencing depth, and CNV type, among others. Specifically, the PacBio CCS sequencing platform outperformed PacBio CLR and Nanopore platforms regarding CNV detection recall rates. A sequencing depth of 10x demonstrated the capability to identify 85% of the CNVs detected in a 50x dataset. Moreover, deletions were more generally detectable than duplications. Among the eight benchmarked methods, cuteSV, Delly, pbsv, and Sniffles2 demonstrated superior accuracy, while SVIM exhibited high recall rates. Peilin Jia |
Briefings Bioinform. | 2 |
| 2023 | Benchmark of embedding-based methods for accurate and transferable prediction of drug responseabstractPrediction of therapy response has been a major challenge in cancer precision medicine due to the extensive tumor heterogeneity. Recently, several deep learning methods have been developed to predict drug response by utilizing various omics data. Most of them train models by using the drug-response screening data generated from cell lines and then use these models to predict response in cancer patient data. In this study, we focus on and evaluate deep learning methods using transcriptome data for the long-standing question of personalized drug-response prediction. We developed an embedding-based approach for drug-response prediction and benchmarked similar methods for their performance. For all methods, we used pretreatment transcriptome data to train models and then conducted a comprehensive evaluation and comparison of the models using cross-panels, cross-datasets and target genes. We further validated the methods using three independent datasets assessing multiple compounds for their predictive capability of drug response, survival outcome and cell line status. As a result, the methods building on gene embeddings had an overall competitive performance with reduced overfitting when we applied evaluation parameters for model fitting as well as the correlation with clinical outcomes in the validation data. We further developed an ensemble model to combine the results from the three most competitive methods for an overall prediction. Finally, we developed DrVAEN (https://bioinfo.uth.edu/drvaen), a user-friendly and easy-accessible web-server that hosts all these methods for drug-response prediction and model comparison for broad use in cancer research, method evaluation and drug development. Peilin Jia, Ruifeng Hu 0002, Zhongming Zhao |
Briefings Bioinform. | 1 |
| 2021 | Distinct effect of prenatal and postnatal brain expression across 20 brain disorders and anthropometric social traits: a systematic study of spatiotemporal modularityabstractDifferent spatiotemporal abnormalities have been implicated in different neuropsychiatric disorders and anthropometric social traits, yet an investigation in the temporal network modularity with brain tissue transcriptomics has been lacking. We developed a supervised network approach to investigate the genome-wide association study (GWAS) results in the spatial and temporal contexts and demonstrated it in 20 brain disorders and anthropometric social traits. BrainSpan transcriptome profiles were used to discover significant modules enriched with trait susceptibility genes in a developmental stage-stratified manner. We investigated whether, and in which developmental stages, GWAS-implicated genes are coordinately expressed in brain transcriptome. We identified significant network modules for each disorder and trait at different developmental stages, providing a systematic view of network modularity at specific developmental stages for a myriad of brain disorders and traits. Specifically, we observed a strong pattern of the fetal origin for most psychiatric disorders and traits [such as schizophrenia (SCZ), bipolar disorder, obsessive-compulsive disorder and neuroticism], whereas increased co-expression activities of genes were more strongly associated with neurological diseases [such as Alzheimer's disease (AD) and amyotrophic lateral sclerosis] and anthropometric traits (such as college completion, education and subjective well-being) in postnatal brains. Further analyses revealed enriched cell types and functional features that were supported and corroborated prior knowledge in specific brain disorders, such as clathrin-mediated endocytosis in AD, myelin sheath in multiple sclerosis and regulation of synaptic plasticity in both college completion and education. Our study provides a landscape view of the spatiotemporal features in a myriad of brain-related disorders and traits. Peilin Jia, Astrid Marilyn Manuel, Brisa S. Fernandes, Yulin Dai, Zhongming Zhao |
Briefings Bioinform. | 1 |
| 2021 | Deep4mC: systematic assessment and computational prediction for DNA N4-methylcytosine sites by deep learningabstractDNA N4-methylcytosine (4mC) modification represents a novel epigenetic regulation. It involves in various cellular processes, including DNA replication, cell cycle and gene expression, among others. In addition to experimental identification of 4mC sites, in silico prediction of 4mC sites in the genome has emerged as an alternative and promising approach. In this study, we first reviewed the current progress in the computational prediction of 4mC sites and systematically evaluated the predictive capacity of eight conventional machine learning algorithms as well as 12 feature types commonly used in previous studies in six species. Using a representative benchmark dataset, we investigated the contribution of feature selection and stacking approach to the model construction, and found that feature optimization and proper reinforcement learning could improve the performance. We next recollected newly added 4mC sites in the six species' genomes and developed a novel deep learning-based 4mC site predictor, namely Deep4mC. Deep4mC applies convolutional neural networks with four representative features. For species with small numbers of samples, we extended our deep learning framework with a bootstrapping method. Our evaluation indicated that Deep4mC could obtain high accuracy and robust performance with the average area under curve (AUC) values greater than 0.9 in all species (range: 0.9005-0.9722). In comparison, Deep4mC achieved an AUC value improvement from 10.14 to 46.21% when compared to previous tools in these six species. A user-friendly web server (https://bioinfo.uth.edu/Deep4mC) was built for predicting putative 4mC sites in a genome. Hao-Dong Xu, Peilin Jia, Zhongming Zhao |
Briefings Bioinform. | 2 |
| 2020 | Critical microRNAs and regulatory motifs in cleft palate identified by a conserved miRNA-TF-gene network approach in humans and miceabstractCleft palate (CP) is the second most common congenital birth defect. The etiology of CP is complicated, with involvement of various genetic and environmental factors. To investigate the gene regulatory mechanisms, we designed a powerful regulatory analytical approach to identify the conserved regulatory networks in humans and mice, from which we identified critical microRNAs (miRNAs), target genes and regulatory motifs (miRNA-TF-gene) related to CP. Using our manually curated genes and miRNAs with evidence in CP in humans and mice, we constructed miRNA and transcription factor (TF) co-regulation networks for both humans and mice. A consensus regulatory loop (miR17/miR20a-FOXE1-PDGFRA) and eight miRNAs (miR-140, miR-17, miR-18a, miR-19a, miR-19b, miR-20a, miR-451a and miR-92a) were discovered in both humans and mice. The role of miR-140, which had the strongest association with CP, was investigated in both human and mouse palate cells. The overexpression of miR-140-5p, but not miR-140-3p, significantly inhibited cell proliferation. We further examined whether miR-140 overexpression could suppress the expression of its predicted target genes (BMP2, FGF9, PAX9 and PDGFRA). Our results indicated that miR-140-5p overexpression suppressed the expression of BMP2 and FGF9 in cultured human palate cells and Fgf9 and Pdgfra in cultured mouse palate cells. In summary, our conserved miRNA-TF-gene regulatory network approach is effective in detecting consensus miRNAs, motifs, and regulatory mechanisms in human and mouse CP. Peilin Jia, Saurav Mallik, Rong Fei, Hiroki Yoshioka, Akiko Suzuki, Junichi Iwata, Zhongming Zhao |
Briefings Bioinform. | 2 |
| 2020 | 6mA-Finder: a novel online tool for predicting DNA N6-methyladenine sites in genomesabstractMOTIVATION: DNA N6-methyladenine (6 mA) has recently been found as an essential epigenetic modification, playing its roles in a variety of cellular processes. The abnormal status of DNA 6 mA modification has been reported in cancer and other disease. The annotation of 6 mA marks in genome is the first crucial step to explore the underlying molecular mechanisms including its regulatory roles. RESULTS: We present a novel online DNA 6 mA site tool, 6 mA-Finder, by incorporating seven sequence-derived information and three physicochemical-based features through recursive feature elimination strategy. Our multiple cross-validations indicate the promising accuracy and robustness of our model. 6 mA-Finder outperforms its peer tools in general and species-specific 6 mA site prediction, suggesting it can provide a useful resource for further experimental investigation of DNA 6 mA modification. AVAILABILITY AND IMPLEMENTATION: https://bioinfo.uth.edu/6mA_Finder. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Hao-Dong Xu, Ruifeng Hu 0002, Peilin Jia, Zhongming Zhao |
Bioinform. | 3 |
| 2019 | Distinct telomere length and molecular signatures in seminoma and non-seminoma of testicular germ cell tumorabstractTesticular germ cell tumors (TGCTs) are classified into two main subtypes, seminoma (SE) and non-seminoma (NSE), but their molecular distinctions remain largely unexplored. Here, we used expression data for mRNAs and microRNAs (miRNAs) from The Cancer Genome Atlas (TCGA) to perform a systematic investigation to explain the different telomere length (TL) features between NSE (n = 48) and SE (n = 55). We found that TL elongation was dominant in NSE, whereas TL shortening prevailed in SE. We further showed that both mRNA and miRNA expression profiles could clearly distinguish these two subtypes. Notably, four telomere-related genes (TelGenes) showed significantly higher expression and positively correlated with telomere elongation in NSE than SE: three telomerase activity-related genes (TERT, WRAP53 and MYC) and an independent telomerase activity gene (ZSCAN4). We also found that the expression of genes encoding Yamanaka factors was positively correlated with telomere lengthening in NSE. Among them, SOX2 and MYC were highly expressed in NSE versus SE, while POU5F1 and KLF4 had the opposite patterns. These results suggested that enhanced expression of both TelGenes (TERT, WRAP53, MYC and ZSCAN4) and Yamanaka factors might induce telomere elongation in NSE. Conversely, the relative lack of telomerase activation and low expression of independent telomerase activity pathway during cell division may be contributed to telomere shortening in SE. Taken together, our results revealed the potential molecular profiles and regulatory roles involving the TL difference between NSE and SE, and provided a better molecular understanding of this complex disease. Hua Sun 0003, Pora Kim, Peilin Jia, Aekyung Park, Zhongming Zhao |
Briefings Bioinform. | 3 |
| 2019 | Translational bioinformatics in mental health: open access data sources and computational biomarker discoveryabstractMental illness is increasingly recognized as both a significant cost to society and a significant area of opportunity for biological breakthrough. As -omics and imaging technologies enable researchers to probe molecular and physiological underpinnings of multiple diseases, opportunities arise to explore the biological basis for behavioral health and disease. From individual investigators to large international consortia, researchers have generated rich data sets in the area of mental health, including genomic, transcriptomic, metabolomic, proteomic, clinical and imaging resources. General data repositories such as the Gene Expression Omnibus (GEO) and Database of Genotypes and Phenotypes (dbGaP) and mental health (MH)-specific initiatives, such as the Psychiatric Genomics Consortium, MH Research Network and PsychENCODE represent a wealth of information yet to be gleaned. At the same time, novel approaches to integrate and analyze data sets are enabling important discoveries in the area of mental and behavioral health. This review will discuss and catalog into an organizing framework the increasingly diverse set of MH data resources available, using schizophrenia as a focus area, and will describe novel and integrative approaches to molecular biomarker discovery that make use of mental health data. Jessica D. Tenenbaum, Krithika Bhuvaneshwar, Jane P. Gagliardi, Kate Fultz Hollis, Peilin Jia, Radhakrishnan Nagarajan, Gopalkumar Rakesh, Vignesh Subbian, Shyam Visweswaran, Zhongming Zhao, Leon Rozenblit |
Briefings Bioinform. | 5 |
| 2019 | CNet: a multi-omics approach to detecting clinically associated, combinatory genomic signaturesabstractMOTIVATION: Genome-wide multi-omics profiling of complex diseases provides valuable resources and opportunities to discover associations between various measures of genes and diseases. Currently, a pressing challenge is how to effectively detect functional genes associated with or causing phenotypic outcomes. We developed CNet to identify groups of genomic signatures whose combinatory effect is significantly associated with clinical and phenotypical outcomes. RESULTS: CNet builds on a generalized sequential feedforward method, augmented by a down-sampling bootstrap strategy to reduce random hitchhiking signatures. It further applies a dynamic trimming procedure to remove relatively less informative signatures at every step. CNet can manage heterogeneous genomic signature profiles simultaneously and select the best signature to represent a specific gene. To deal with various forms of clinical and phenotypical measurements, we introduced four models to deal with continuous, categorical and censored data. We tested CNet using drug-response data, multidimensional cancer genomics data and genome-wide association study data for multiple traits. Our results demonstrated that in various scenarios, CNet could effectively identify signatures that are associated with the outcomes. In addition, we applied CNet to identify likely disease-causing chains involving somatic mutations, pathway activities and patient outcomes. With appropriate setting, CNet can be applied in many biological conditions. AVAILABILITY AND IMPLEMENTATION: CNet can be downloaded at https://github.com/bsml320/CNet. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Peilin Jia, Guangsheng Pei, Zhongming Zhao |
Bioinform. | 1 |
| 2019 | deTS: tissue-specific enrichment analysis to decode tissue specificityabstractMOTIVATION: Diseases and traits are under dynamic tissue-specific regulation. However, heterogeneous tissues are often collected in biomedical studies, which reduce the power in the identification of disease-associated variants and gene expression profiles. RESULTS: We present deTS, an R package, to conduct tissue-specific enrichment analysis with two built-in reference panels. Statistical methods are developed and implemented for detecting tissue-specific genes and for enrichment test of different forms of query data. Our applications using multi-trait genome-wide association studies data and cancer expression data showed that deTS could effectively identify the most relevant tissues for each query trait or sample, providing insights for future studies. AVAILABILITY AND IMPLEMENTATION: https://github.com/bsml320/deTS and CRAN https://cran.r-project.org/web/packages/deTS/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Guangsheng Pei, Yulin Dai, Zhongming Zhao, Peilin Jia |
Bioinform. | 4 |
| 2018 | Kinase impact assessment in the landscape of fusion genes that retain kinase domains: a pan-cancer studyabstractAssessing the impact of kinase in gene fusion is essential for both identifying driver fusion genes (FGs) and developing molecular targeted therapies. Kinase domain retention is a crucial factor in kinase fusion genes (KFGs), but such a systematic investigation has not been done yet. To this end, we analyzed kinase domain retention (KDR) status in chimeric protein sequences of 914 KFGs covering 312 kinases across 13 major cancer types. Based on 171 kinase domain-retained KFGs including 101 kinases, we studied their recurrence, kinase groups, fusion partners, exon-based expression depth, short DNA motifs around the break points and networks. Our results, such as more KDR than 5'-kinase fusion genes, combinatorial effects between 3'-KDR kinases and their 5'-partners and a signal transduction-specific DNA sequence motif in the break point intronic sequences, supported positive selection on 3'-kinase fusion genes in cancer. We introduced a degree-of-frequency (DoF) score to measure the possible number of KFGs of a kinase. Interestingly, kinases with high DoF scores tended to undergo strong gene expression alteration at the break points. Furthermore, our KDR gene fusion network analysis revealed six of the seven kinases with the highest DoF scores (ALK, BRAF, MET, NTRK1, NTRK3 and RET) were all observed in thyroid carcinoma. Finally, we summarized common features of 'effective' (highly recurrent) kinases in gene fusions such as expression alteration at break point, redundant usage in multiple cancer types and 3'-location tendency. Collectively, our findings are useful for prioritizing driver kinases and FGs and provided insights into KFGs' clinical implications. Pora Kim, Peilin Jia, Zhongming Zhao |
Briefings Bioinform. | 2 |
| 2017 | Impacts of somatic mutations on gene expression: an association perspectiveabstractAssessing the functional impacts of somatic mutations in cancer genomes is critical for both identifying driver mutations and developing molecular targeted therapies. Currently, it remains a fundamental challenge to distinguish the patterns through which mutations execute their biological effects and to infer biological mechanisms underlying these patterns. To this end, we systematically studied the association between somatic mutations in protein-coding regions and expression profiles, which represents an indirect measurement of impacts. We defined mutation features (mutation type, cluster and status) and built linear regression models to assess mutation associations with mRNA expression and protein expression. Our results presented a comprehensive landscape of the associations between mutation features and expression profile in multiple cancer types, including 62 genes showing mutation type associated expression changes, 21 genes showing mutation cluster associations and 51 genes showing mutation status associations. We revealed four characteristics of the patterns that mutations impact on expression. First, we showed that mutation type (truncation versus amino acid-altering mutations) was the most important determinant of expression levels. Second, we detected mutation clusters in well-studied oncogenes that were associated with gene expression. Third, we found both similarities and differences in association patterns existed within and across cancer types. Fourth, although many of the observed associations stay stable at both mRNA and protein expression levels, there are also novel associations uniquely observed at the protein level, which warrant future investigation. Taken together, our findings provided implications for cancer driver gene prioritization and insights into the functional consequences of somatic mutations. Peilin Jia, Zhongming Zhao |
Briefings Bioinform. | 1 |
| 2015 | EW_dmGWAS: edge-weighted dense module search for genome-wide association studies and gene expression profilesabstractAbstract Summary: We previously developed dmGWAS to search for dense modules in a human protein–protein interaction (PPI) network; it has since become a popular tool for network-assisted analysis of genome-wide association studies (GWAS). dmGWAS weights nodes by using GWAS signals. Here, we introduce an upgraded algorithm, EW_dmGWAS, to boost GWAS signals in a node- and edge-weighted PPI network. In EW_dmGWAS, we utilize condition-specific gene expression profiles for edge weights. Specifically, differential gene co-expression is used to infer the edge weights. We applied EW_dmGWAS to two diseases and compared it with other relevant methods. The results suggest that EW_dmGWAS is more powerful in detecting disease-associated signals. Availability and implementation: The algorithm of EW_dmGWAS is implemented in the R package dmGWAS_3.0 and is available at http://bioinfo.mc.vanderbilt.edu/dmGWAS. Contact: [email protected] or [email protected] Supplementary information: Supplementary materials are available at Bioinformatics online. Quan Wang 0004, Zhongming Zhao, Peilin Jia |
Bioinform. | 4 |
| 2015 | A Gene Gravity Model for the Evolution of Cancer Genomes: A Study of 3, 000 Cancer Genomes across 9 Cancer TypesabstractCancer development and progression result from somatic evolution by an accumulation of genomic alterations. The effects of those alterations on the fitness of somatic cells lead to evolutionary adaptations such as increased cell proliferation, angiogenesis, and altered anticancer drug responses. However, there are few general mathematical models to quantitatively examine how perturbations of a single gene shape subsequent evolution of the cancer genome. In this study, we proposed the gene gravity model to study the evolution of cancer genomes by incorporating the genome-wide transcription and somatic mutation profiles of ~3,000 tumors across 9 cancer types from The Cancer Genome Atlas into a broad gene network. We found that somatic mutations of a cancer driver gene may drive cancer genome evolution by inducing mutations in other genes. This functional consequence is often generated by the combined effect of genetic and epigenetic (e.g., chromatin regulation) alterations. By quantifying cancer genome evolution using the gene gravity model, we identified six putative cancer genes (AHNAK, COL11A1, DDX3X, FAT4, STAG2, and SYNE1). The tumor genomes harboring the nonsynonymous somatic mutations in these genes had a higher mutation density at the genome level compared to the wild-type groups. Furthermore, we provided statistical evidence that hypermutation of cancer driver genes on inactive X chromosomes is a general feature in female cancer genomes. In summary, this study sheds light on the functional consequences and evolutionary characteristics of somatic mutations during tumorigenesis by propelling adaptive cancer genome evolution, which would provide new perspectives for cancer research and therapeutics. Feixiong Cheng, Chen-Ching Lin, Junfei Zhao, Peilin Jia, Wen-Hsiung Li, Zhongming Zhao |
PLoS Comput. Biol. | 5 |
| 2015 | Deciphering Signaling Pathway Networks to Understand the Molecular Mechanisms of Metformin ActionabstractA drug exerts its effects typically through a signal transduction cascade, which is non-linear and involves intertwined networks of multiple signaling pathways. Construction of such a signaling pathway network (SPNetwork) can enable identification of novel drug targets and deep understanding of drug action. However, it is challenging to synopsize critical components of these interwoven pathways into one network. To tackle this issue, we developed a novel computational framework, the Drug-specific Signaling Pathway Network (DSPathNet). The DSPathNet amalgamates the prior drug knowledge and drug-induced gene expression via random walk algorithms. Using the drug metformin, we illustrated this framework and obtained one metformin-specific SPNetwork containing 477 nodes and 1,366 edges. To evaluate this network, we performed the gene set enrichment analysis using the disease genes of type 2 diabetes (T2D) and cancer, one T2D genome-wide association study (GWAS) dataset, three cancer GWAS datasets, and one GWAS dataset of cancer patients with T2D on metformin. The results showed that the metformin network was significantly enriched with disease genes for both T2D and cancer, and that the network also included genes that may be associated with metformin-associated cancer survival. Furthermore, from the metformin SPNetwork and common genes to T2D and cancer, we generated a subnetwork to highlight the molecule crosstalk between T2D and cancer. The follow-up network analyses and literature mining revealed that seven genes (CDKN1A, ESR1, MAX, MYC, PPARGC1A, SP1, and STK11) and one novel MYC-centered pathway with CDKN1A, SP1, and STK11 might play important roles in metformin's antidiabetic and anticancer effects. Some results are supported by previous studies. In summary, our study 1) develops a novel framework to construct drug-specific signal transduction networks; 2) provides insights into the molecular mode of metformin; 3) serves a model for exploring signaling pathways to facilitate understanding of drug action, disease pathogenesis, and identification of drug targets. Jingchun Sun, Min Zhao 0006, Peilin Jia, Lily Wang 0001, Yonghui Wu 0001, Carissa Iverson, Yubo Zhou, Erica A. Bowton, Dan M. Roden, Joshua C. Denny, Melinda Aldrich, Hua Xu 0001, Zhongming Zhao |
PLoS Comput. Biol. | 3 |
| 2014 | VarWalker: Personalized Mutation Network Analysis of Putative Cancer Genes from Next-Generation Sequencing DataabstractA major challenge in interpreting the large volume of mutation data identified by next-generation sequencing (NGS) is to distinguish driver mutations from neutral passenger mutations to facilitate the identification of targetable genes and new drugs. Current approaches are primarily based on mutation frequencies of single-genes, which lack the power to detect infrequently mutated driver genes and ignore functional interconnection and regulation among cancer genes. We propose a novel mutation network method, VarWalker, to prioritize driver genes in large scale cancer mutation data. VarWalker fits generalized additive models for each sample based on sample-specific mutation profiles and builds on the joint frequency of both mutation genes and their close interactors. These interactors are selected and optimized using the Random Walk with Restart algorithm in a protein-protein interaction network. We applied the method in >300 tumor genomes in two large-scale NGS benchmark datasets: 183 lung adenocarcinoma samples and 121 melanoma samples. In each cancer, we derived a consensus mutation subnetwork containing significantly enriched consensus cancer genes and cancer-related functional pathways. These cancer-specific mutation networks were then validated using independent datasets for each cancer. Importantly, VarWalker prioritizes well-known, infrequently mutated genes, which are shown to interact with highly recurrently mutated genes yet have been ignored by conventional single-gene-based approaches. Utilizing VarWalker, we demonstrated that network-assisted approaches can be effectively adapted to facilitate the detection of cancer driver genes in NGS data. Peilin Jia, Zhongming Zhao |
PLoS Comput. Biol. | 1 |
| 2013 | Network-based mutation analysis of putative cancer genes from next-generation sequencing dataabstractNext-generation sequencing (NGS) has enabled fast detection of somatic mutations in cancer genomes. A major challenge in interpreting the large volume of mutation data is to distinguish driver mutations from neutral passenger mutations. Current approaches are primarily single-gene based prioritization according to mutation frequencies, which harbors both high false positive and false negative discoveries. We propose a novel network-based method of mutation data for driver gene prioritization from large scale mutation data for cancer. Our method takes into consideration of the mutation profile of each patient by fitting sample-specific generalized additive models. It builds on joint frequency of both mutation genes and their close interactors, which are optimized by the algorithm Random Walk with Restart in a protein-protein interaction network. We demonstrated our method in two large-scale NGS datasets: a lung adenocarcinoma (LUAD) dataset including 183 patients and a melanoma dataset including 121 samples. In each cancer, we derived a consensus mutation subnetwork with significantly enriched consensus cancer genes and cancer-related functional pathways. The LUAD subnetwork recruited 70 genes of the Cancer Gene Census (CGC) collection (p-value <; 2.2×10-16, Fisher's Exact Test) and the melanoma subnetwork included 65 CGC genes (p-value <; 2.2x10-16). In addition, our results indicate that some well-known, infrequently mutated genes, which have been ignored by conventional single-gene based approaches, are also prioritized and are shown to interact with those highly recurrently mutated genes. In sum, our method is effective in prioritizing candidate driver genes from more than ten thousand mutation genes and provides biological interpretations for future work. Peilin Jia, Zhongming Zhao |
BIBM | 1 |
| 2013 | Application of next generation sequencing to human gene fusion detection: computational tools, features and perspectivesabstractGene fusions are important genomic events in human cancer because their fusion gene products can drive the development of cancer and thus are potential prognostic tools or therapeutic targets in anti-cancer treatment. Major advancements have been made in computational approaches for fusion gene discovery over the past 3 years due to improvements and widespread applications of high-throughput next generation sequencing (NGS) technologies. To identify fusions from NGS data, existing methods typically leverage the strengths of both sequencing technologies and computational strategies. In this article, we review the NGS and computational features of existing methods for fusion gene detection and suggest directions for future development. Qingguo Wang, Junfeng Xia, Peilin Jia, William Pao, Zhongming Zhao |
Briefings Bioinform. | 3 |
| 2013 | Computational tools for copy number variation (CNV) detection using next-generation sequencing data: features and perspectivesabstractCopy number variation (CNV) is a prevalent form of critical genetic variation that leads to an abnormal number of copies of large genomic regions in a cell. Microarray-based comparative genome hybridization (arrayCGH) or genotyping arrays have been standard technologies to detect large regions subject to copy number changes in genomes until most recently high-resolution sequence data can be analyzed by next-generation sequencing (NGS). During the last several years, NGS-based analysis has been widely applied to identify CNVs in both healthy and diseased individuals. Correspondingly, the strong demand for NGS-based CNV analyses has fuelled development of numerous computational methods and tools for CNV detection. In this article, we review the recent advances in computational methods pertaining to CNV detection using whole genome and whole exome sequencing data. Additionally, we discuss their strengths and weaknesses and suggest directions for future development. Min Zhao 0006, Qingguo Wang, Quan Wang 0004, Peilin Jia, Zhongming Zhao |
BMC Bioinform. | 4 |
| 2012 | Network-Assisted Investigation of Combined Causal Signals from Genome-Wide Association Studies in SchizophreniaabstractWith the recent success of genome-wide association studies (GWAS), a wealth of association data has been accomplished for more than 200 complex diseases/traits, proposing a strong demand for data integration and interpretation. A combinatory analysis of multiple GWAS datasets, or an integrative analysis of GWAS data and other high-throughput data, has been particularly promising. In this study, we proposed an integrative analysis framework of multiple GWAS datasets by overlaying association signals onto the protein-protein interaction network, and demonstrated it using schizophrenia datasets. Building on a dense module search algorithm, we first searched for significantly enriched subnetworks for schizophrenia in each single GWAS dataset and then implemented a discovery-evaluation strategy to identify module genes with consistent association signals. We validated the module genes in an independent dataset, and also examined them through meta-analysis of the related SNPs using multiple GWAS datasets. As a result, we identified 205 module genes with a joint effect significantly associated with schizophrenia; these module genes included a number of well-studied candidate genes such as DISC1, GNA12, GNA13, GNAI1, GPR17, and GRIN2B. Further functional analysis suggested these genes are involved in neuronal related processes. Additionally, meta-analysis found that 18 SNPs in 9 module genes had P(meta)<1 × 10⁻⁴, including the gene HLA-DQA1 located in the MHC region on chromosome 6, which was reported in previous studies using the largest cohort of schizophrenia patients to date. These results demonstrated our bi-directional network-based strategy is efficient for identifying disease-associated genes with modest signals in GWAS datasets. This approach can be applied to any other complex diseases/traits where multiple GWAS datasets are available. Peilin Jia, Lily Wang 0001, Ayman H. Fanous, Carlos N. Pato, Todd L. Edwards, Zhongming Zhao |
PLoS Comput. Biol. | 1 |
| 2011 | dmGWAS: dense module searching for genome-wide association studies in protein-protein interaction networksabstractMOTIVATION: An important question that has emerged from the recent success of genome-wide association studies (GWAS) is how to detect genetic signals beyond single markers/genes in order to explore their combined effects on mediating complex diseases and traits. Integrative testing of GWAS association data with that from prior-knowledge databases and proteome studies has recently gained attention. These methodologies may hold promise for comprehensively examining the interactions between genes underlying the pathogenesis of complex diseases. METHODS: Here, we present a dense module searching (DMS) method to identify candidate subnetworks or genes for complex diseases by integrating the association signal from GWAS datasets into the human protein-protein interaction (PPI) network. The DMS method extensively searches for subnetworks enriched with low P-value genes in GWAS datasets. Compared with pathway-based approaches, this method introduces flexibility in defining a gene set and can effectively utilize local PPI information. RESULTS: We implemented the DMS method in an R package, which can also evaluate and graphically represent the results. We demonstrated DMS in two GWAS datasets for complex diseases, i.e. breast cancer and pancreatic cancer. For each disease, the DMS method successfully identified a set of significant modules and candidate genes, including some well-studied genes not detected in the single-marker analysis of GWA studies. Functional enrichment analysis and comparison with previously published methods showed that the genes we identified by DMS have higher association signal. AVAILABILITY: dmGWAS package and documents are available at http://bioinfo.mc.vanderbilt.edu/dmGWAS.html. Peilin Jia, Siyuan Zheng, Jirong Long, Zhongming Zhao |
Bioinform. | 1 |
| 2011 | An efficient hierarchical generalized linear mixed model for pathway analysis of genome-wide association studiesabstractMOTIVATION: In genome-wide association studies (GWAS) of complex diseases, genetic variants having real but weak associations often fail to be detected at the stringent genome-wide significance level. Pathway analysis, which tests disease association with combined association signals from a group of variants in the same pathway, has become increasingly popular. However, because of the complexities in genetic data and the large sample sizes in typical GWAS, pathway analysis remains to be challenging. We propose a new statistical model for pathway analysis of GWAS. This model includes a fixed effects component that models mean disease association for a group of genes, and a random effects component that models how each gene's association with disease varies about the gene group mean, thus belongs to the class of mixed effects models. RESULTS: The proposed model is computationally efficient and uses only summary statistics. In addition, it corrects for the presence of overlapping genes and linkage disequilibrium (LD). Via simulated and real GWAS data, we showed our model improved power over currently available pathway analysis methods while preserving type I error rate. Furthermore, using the WTCCC Type 1 Diabetes (T1D) dataset, we demonstrated mixed model analysis identified meaningful biological processes that agreed well with previous reports on T1D. Therefore, the proposed methodology provides an efficient statistical modeling framework for systems analysis of GWAS. AVAILABILITY: The software code for mixed models analysis is freely available at http://biostat.mc.vanderbilt.edu/LilyWang. Lily Wang 0001, Peilin Jia, Russell D. Wolfinger, Britney L. Grayson, Thomas M. Aune, Zhongming Zhao |
Bioinform. | 2 |
| 2010 | Pathway- and network-based analysis of GWAS data revealed susceptibility gene sets to schizophreniaabstractMaterials and methods In this study, we uniquely examined GWAS data at the the gene set level (i.e., pathways and protein-protein interaction (PPI) subnetworks) rather than the SNP level. We first collected a comprehensive list of pathways from the KEGG and BioCarta databases. Additionally, we applied a network module searching approach to search for informative subnetworks in a nodeweighted human PPI network by GWAS markers’ P values. Using this module searching method, we identified ~100 candidate genes from top ranked schizophrenia-specific modules and used these genes as a gene set for follow up functional enrichment test. We applied two statistical methods (ge ne set enrichment analysis (GSEA) and hypergeometric test) and also combined them by Fisher’s combined method to analyze the gene sets defined by canonical pathways or our module genes. Results were further validated by permutation analysis. Results and conclusion In GSEA analysis, we found the gene set consisting of the network module genes was the most significant among all gene sets, indicating our module searching being efficient in finding the real effect of multiple genes in am ore flexible way than classically defined pathways. We also identified a few pathways that are consistently associated with schizophrenia by multiple methods; they included glutamate metabolism pathway, TNFR1 pathway, and TGF beta signaling pathway. These results not only improved our understanding of the underlying pathogenesis of schizophrenia, but also suggested that the gene set based approach is powerful to detect common variants conferring risk to complex diseases. Peilin Jia, Zhongming Zhao |
BMC Bioinform. | 1 |
| 2009 | A multi-dimensional evidence-based candidate gene prioritization approach for complex diseases-schizophrenia as a caseabstractMOTIVATION: During the past decade, we have seen an exponential growth of vast amounts of genetic data generated for complex disease studies. Currently, across a variety of complex biological problems, there is a strong trend towards the integration of data from multiple sources. So far, candidate gene prioritization approaches have been designed for specific purposes, by utilizing only some of the available sources of genetic studies, or by using a simple weight scheme. Specifically to psychiatric disorders, there has been no prioritization approach that fully utilizes all major sources of experimental data. RESULTS: Here we present a multi-dimensional evidence-based candidate gene prioritization approach for complex diseases and demonstrate it in schizophrenia. In this approach, we first collect and curate genetic studies for schizophrenia from four major categories: association studies, linkage analyses, gene expression and literature search. Genes in these data sets are initially scored by category-specific scoring methods. Then, an optimal weight matrix is searched by a two-step procedure (core genes and unbiased P-values in independent genome-wide association studies). Finally, genes are prioritized by their combined scores using the optimal weight matrix. Our evaluation suggests this approach generates prioritized candidate genes that are promising for further analysis or replication. The approach can be applied to other complex diseases. AVAILABILITY: The collected data, prioritized candidate genes, and gene prioritization tools are freely available at http://bioinfo.mc.vanderbilt.edu/SZGR/. Jingchun Sun, Peilin Jia, Ayman H. Fanous, Bradley Todd Webb, Edwin J. C. G. van den Oord, Xiangning Chen, József Bukszár, Kenneth S. Kendler, Zhongming Zhao |
Bioinform. | 2 |
| 2006 | Demonstration of two novel methods for predicting functional siRNA efficiencyabstractBACKGROUND: siRNAs are small RNAs that serve as sequence determinants during the gene silencing process called RNA interference (RNAi). It is well know that siRNA efficiency is crucial in the RNAi pathway, and the siRNA efficiency for targeting different sites of a specific gene varies greatly. Therefore, there is high demand for reliable siRNAs prediction tools and for the design methods able to pick up high silencing potential siRNAs. RESULTS: In this paper, two systems have been established for the prediction of functional siRNAs: (1) a statistical model based on sequence information and (2) a machine learning model based on three features of siRNA sequences, namely binary description, thermodynamic profile and nucleotide composition. Both of the two methods show high performance on the two datasets we have constructed for training the model. CONCLUSION: Both of the two methods studied in this paper emphasize the importance of sequence information for the prediction of functional siRNAs. The way of denoting a bio-sequence by binary system in mathematical language might be helpful in other analysis work associated with fixed-length bio-sequence. Peilin Jia, Tieliu Shi, Yu-Dong Cai 0001 |
BMC Bioinform. | 1 |