Tzu-Pin Lu

dblp:75/9077 · DBLP profile ↗
← Back
19ranked-venue papers
2as first author
10since 2021 · last 2026
0000-0003-3697-0386ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 19 · 2 first-author · 10 since 2021
YearPublicationVenuePosition
2026 Development of a novel multimodal deep learning approach to improve diagnostic precision in ovarian cancer
abstract
BACKGROUND: Ovarian cancer represents the primary cause of mortality from gynecological malignancies among women. Treatment strategies for benign versus malignant ovarian tumors differ significantly, making accurate preoperative diagnosis essential for clinical decision-making. Traditional ultrasound diagnosis is highly operator-dependent, introducing subjectivity and variability. To improve diagnostic precision in ovarian tumor classification, we developed a multimodal deep learning system that combines ultrasound images with corresponding clinical text reports. METHODS: We retrospectively analyzed 1342 ultrasound images from 1062 patients who received surgical treatment for ovarian tumors at National Taiwan University Hospital from 2011 to 2021. Patients were classified into benign (n = 612) and malignant (including borderline, n = 450) groups based on pathology. A multimodal deep learning architecture was developed, incorporating DenseNet-121 and Swin Transformer for image feature extraction and Bio-Clinical BERT for processing clinical text reports. The dataset was split using subject-level stratification with five-fold cross-validation and a 15% independent test set. Furthermore, an external validation cohort of 268 effective cases from 3 independent medical centers was utilized to evaluate the model's generalizability. RESULTS: The multimodal model achieved superior performance at the subject level with 81.77% (95% CI: 75.89%, 86.48%) accuracy, 79.59% (95% CI: 70.57%, 86.38%) sensitivity, 83.81% (95% CI: 75.59%, 89.64%) specificity, and an area under the curve (AUC) of 0.88 (95% CI: 0.83, 0.93). In the external validation, the model maintained robust performance with an accuracy of 88.81%, sensitivity of 92.59%, and specificity of 84.96%, outperforming the International Ovarian Tumor Analysis Simple Rules (accuracy 86.4%). Integration of clinical text information significantly improved diagnostic performance compared to image-only models. Backward selection analysis revealed that both uterine findings and ovarian tumor descriptions contributed synergistically to the final diagnosis. CONCLUSIONS: This study successfully developed a multimodal deep learning model with diagnostic performance superior to traditional operator-dependent approaches. The model shows promise as a diagnostic tool for ovarian tumor classification, offering clinicians a way to improve preoperative diagnostic accuracy and enhance patient care quality.
Po-Chun Chiu, Chia-Yi Lee, Heng-Cheng Hsu, Yi-Jou Tai, Ying-Cheng Chiang, Tzu-Pin Lu
Briefings Bioinform.6
2024 Cross-population enhancement of PrediXcan predictions with a gnomAD-based east Asian reference framework
abstract
Over the past decade, genome-wide association studies have identified thousands of variants significantly associated with complex traits. For each locus, gene expression levels are needed to further explore its biological functions. To address this, the PrediXcan algorithm leverages large-scale reference data to impute the gene expression level from single nucleotide polymorphisms, and thus the gene-trait associations can be tested to identify the candidate causal genes. However, a challenge arises due to the fact that most reference data are from subjects of European ancestry, and the accuracy and robustness of predicted gene expression in subjects of East Asian (EAS) ancestry remains unclear. Here, we first simulated a variety of scenarios to explore the impact of the level of population diversity on gene expression. Population differentiated variants were estimated by using the allele frequency information from The Genome Aggregation Database. We found that the weights of a variants was the main factor that affected the gene expression predictions, and that ~70% of variants were significantly population differentiated based on proportion tests. To provide insights into this population effect on gene expression levels, we utilized the allele frequency information to develop a gene expression reference panel, Predict Asian-Population (PredictAP), for EAS ancestry. PredictAP can be viewed as an auxiliary tool for PrediXcan when using genotype data from EAS subjects.
Han-Ching Chan, Amrita Chattopadhyay, Tzu-Pin Lu
Briefings Bioinform.3
2023 Twnbiome: a public database of the healthy Taiwanese gut microbiome
abstract
With new advances in next generation sequencing (NGS) technology at reduced costs, research on bacterial genomes in the environment has become affordable. Compared to traditional methods, NGS provides high-throughput sequencing reads and the ability to identify many species in the microbiome that were previously unknown. Numerous bioinformatics tools and algorithms have been developed to conduct such analyses. However, in order to obtain biologically meaningful results, the researcher must select the proper tools and combine them to construct an efficient pipeline. This complex procedure may include tens of tools, each of which require correct parameter settings. Furthermore, an NGS data analysis involves multiple series of command-line tools and requires extensive computational resources, which imposes a high barrier for biologists and clinicians to conduct NGS analysis and even interpret their own data. Therefore, we established a public gut microbiome database, which we call Twnbiome, created using healthy subjects from Taiwan, with the goal of enabling microbiota research for the Taiwanese population. Twnbiome provides users with a baseline gut microbiome panel from a healthy Taiwanese cohort, which can be utilized as a reference for conducting case-control studies for a variety of diseases. It is an interactive, informative, and user-friendly database. Twnbiome additionally offers an analysis pipeline, where users can upload their data and download analyzed results. Twnbiome offers an online database which non-bioinformatics users such as clinicians and doctors can not only utilize to access a control set of data, but also analyze raw data with a few easy clicks. All results are customizable with ready-made plots and easily downloadable tables. Database URL: http://twnbiome.cgm.ntu.edu.tw/ .
Amrita Chattopadhyay, Chien-Yueh Lee, Ya-Chin Lee, Chiang-Lin Liu, Hsin-Kuang Chen, Yung-Hua Li, Liang-Chuan Lai, Mong-Hsun Tsai, Yen-Hsuan Ni, Han-Mo Chiu, Tzu-Pin Lu, Eric Y. Chuang
BMC Bioinform.11
2023 Multi-ethnic Imputation System (MI-System): A genotype imputation server for high-dimensional data
abstract
OBJECTIVE: Genotype imputation is a commonly used technique that infers un-typed variants into a study's genotype data, allowing better identification of causal variants in disease studies. However, due to overrepresentation of Caucasian studies, there's a lack of understanding of genetic basis of health-outcomes in other ethnic populations. Therefore, facilitating imputation of missing key-predictor-variants that can potentially improve a risk health-outcome prediction model, specifically for Asian ancestry, is of utmost relevance. METHODS: We aimed to construct an imputation and analysis web-platform, that primarily facilitates, but is not limited to genotype imputation on East-Asians. The goal is to provide a collaborative imputation platform for researchers in the public domain towards rapidly and efficiently conducting accurate genotype imputation. RESULTS: We present an online genotype imputation platform, Multi-ethnic Imputation System (MI-System) (https://misystem.cgm.ntu.edu.tw/), that offers users 3 established pipelines, SHAPEIT2-IMPUTE2, SHAPEIT4-IMPUTE5, and Beagle5.1 for conducting imputation analyses. In addition to 1000 Genomes and Hapmap3, a new customized Taiwan Biobank (TWB) reference panel, specifically created for Taiwanese-Chinese ancestry is provided. MI-System further offers functions to create customized reference panels to be used for imputation, conduct quality control, split whole genome data into chromosomes, and convert genome builds. CONCLUSION: Users can upload their genotype data and perform imputation with minimum effort and resources. The utility functions further can be utilized to preprocess user uploaded data with easy clicks. MI-System potentially contributes to Asian-population genetics research, while eliminating the requirement for high performing computational resources and bioinformatics expertise. It will enable an increased pace of research and provide a knowledge-base for genetic carriers of complex diseases, therefore greatly enhancing patient-driven research. STATEMENT OF SIGNIFICANCE: Multi-ethnic Imputation System (MI-System), primarily facilitates, but is not limited to, imputation on East-Asians, through 3 established prephasing-imputation pipelines, SHAPEIT2-IMPUTE2, SHAPEIT4-IMPUTE5, and Beagle5.1, where users can upload their genotype data and perform imputation and other utility functions with minimum effort and resources. A new customized Taiwan Biobank (TWB) reference panel, specifically created for Taiwanese-Chinese ancestry is provided. Utility functions include (a) create customized reference panels, (b) conduct quality control, (c) split whole genome data into chromosomes, and (d) convert genome builds. Users can also combine 2 reference panels using the system and use combined panels as reference to conduct imputation using MI-System.
Amrita Chattopadhyay, Chien-Yueh Lee, Ying-Cheng Shen, Kuan-Chen Lu, Tzu-Hung Hsiao, Ching-Heng Lin, Liang-Chuan Lai, Mong-Hsun Tsai, Tzu-Pin Lu, Eric Y. Chuang
J. Biomed. Informatics9
2022 CLIN_SKAT: an R package to conduct association analysis using functionally relevant variants
abstract
BACKGROUND: Availability of next generation sequencing data, allows low-frequency and rare variants to be studied through strategies other than the commonly used genome-wide association studies (GWAS). Rare variants are important keys towards explaining the heritability for complex diseases that remains to be explained by common variants due to their low effect sizes. However, analysis strategies struggle to keep up with the huge amount of data at disposal therefore creating a bottleneck. This study describes CLIN_SKAT, an R package, that provides users with an easily implemented analysis pipeline with the goal of (i) extracting clinically relevant variants (both rare and common), followed by (ii) gene-based association analysis by grouping the selected variants. RESULTS: CLIN_SKAT offers four simple functions that can be used to obtain clinically relevant variants, map them to genes or gene sets, calculate weights from global healthy populations and conduct weighted case-control analysis. CLIN_SKAT introduces improvements by adding certain pre-analysis steps and customizable features to make the SKAT results clinically more meaningful. Moreover, it offers several plot functions that can be availed towards obtaining visualizations for interpretation of the analyses results. CLIN_SKAT is available on Windows/Linux/MacOS and is operative for R version 4.0.4 or later. It can be freely downloaded from https://github.com/ShihChingYu/CLIN_SKAT , installed through devtools::install_github("ShihChingYu/CLIN_SKAT", force=T) and executed by loading the package into R using library(CLIN_SKAT). All outputs (tabular and graphical) can be downloaded in simple, publishable formats. CONCLUSIONS: Statistical association analysis is often underpowered due to low sample sizes and high numbers of variants to be tested, limiting detection of causal ones. Therefore, retaining a subset of variants that are biologically meaningful seems to be a more effective strategy for identifying explainable associations while reducing the degrees of freedom. CLIN_SKAT offers users a one-stop R package that identifies disease risk variants with improved power via a series of tailor-made procedures that allows dimension reduction, by retaining functionally relevant variants, and incorporating ethnicity based priors. Furthermore, it also eliminates the requirement for high computational resources and bioinformatics expertise.
Amrita Chattopadhyay, Ching-Yu Shih, Yu-Chen Hsu, Jyh-Ming Jimmy Juang, Eric Y. Chuang, Tzu-Pin Lu
BMC Bioinform.6
2022 ceRNAR: An R package for identification and analysis of ceRNA-miRNA triplets
abstract
Competitive endogenous RNA (ceRNA) represents a novel mechanism of gene regulation that controls several biological and pathological processes. Recently, an increasing number of in silico methods have been developed to accelerate the identification of such regulatory events. However, there is still a need for a tool supporting the hypothesis that ceRNA regulatory events only occur at specific miRNA expression levels. To this end, we present an R package, ceRNAR, which allows identification and analysis of ceRNA-miRNA triplets via integration of miRNA and RNA expression data. The ceRNAR package integrates three main steps: (i) identification of ceRNA pairs based on a rank-based correlation between pairs that considers the impact of miRNA and a running sum correlation statistic, (ii) sample clustering based on gene-gene correlation by circular binary segmentation, and (iii) peak merging to identify the most relevant sample patterns. In addition, ceRNAR also provides downstream analyses of identified ceRNA-miRNA triplets, including network analysis, functional annotation, survival analysis, external validation, and integration of different tools. The performance of our proposed approach was validated through simulation studies of different scenarios. Compared with several published tools, ceRNAR was able to identify true ceRNA triplets with high sensitivity, low false-positive rates, and acceptable running time. In real data applications, the ceRNAs common to two lung cancer datasets were identified in both datasets. The bridging miRNA for one of these, the ceRNA for MAP4K3, was identified by ceRNAR as hsa-let-7c-5p. Since similar cancer subtypes do share some biological patterns, these results demonstrated that our proposed algorithm was able to identify potential ceRNA targets in real patients. In summary, ceRNAR offers a novel algorithm and a comprehensive pipeline to identify and analyze ceRNA regulation. The package is implemented in R and is available on GitHub (https://github.com/ywhsiao/ceRNAR).
Yi-Wen Hsiao, Lin Wang 0043, Tzu-Pin Lu
PLoS Comput. Biol.3
2021 High-performance deep learning pipeline predicts individuals in mixtures of DNA using sequencing data
abstract
In this study, we proposed a deep learning (DL) model for classifying individuals from mixtures of DNA samples using 27 short tandem repeats and 94 single nucleotide polymorphisms obtained through massively parallel sequencing protocol. The model was trained/tested/validated with sequenced data from 6 individuals and then evaluated using mixtures from forensic DNA samples. The model successfully identified both the major and the minor contributors with 100% accuracy for 90 DNA mixtures, that were manually prepared by mixing sequence reads of 3 individuals at different ratios. Furthermore, the model identified 100% of the major contributors and 50-80% of the minor contributors in 20 two-sample external-mixed-samples at ratios of 1:39 and 1:9, respectively. To further demonstrate the versatility and applicability of the pipeline, we tested it on whole exome sequence data to classify subtypes of 20 breast cancer patients and achieved an area under curve of 0.85. Overall, we present, for the first time, a complete pipeline, including sequencing data processing steps and DL steps, that is applicable across different NGS platforms. We also introduced a sliding window approach, to overcome the sequence length variation problem of sequencing data, and demonstrate that it improves the model performance dramatically.
Nam Nhut Phan, Amrita Chattopadhyay, Tsui-Ting Lee, Hsiang-I Yin, Tzu-Pin Lu, Liang-Chuan Lai, Hsiao-Lin Hwa, Mong-Hsun Tsai, Eric Y. Chuang
Briefings Bioinform.5
2021 Gene-set integrative analysis of multi-omics data using tensor-based association test
abstract
MOTIVATION: Facilitated by technological advances and the decrease in costs, it is feasible to gather subject data from several omics platforms. Each platform assesses different molecular events, and the challenge lies in efficiently analyzing these data to discover novel disease genes or mechanisms. A common strategy is to regress the outcomes on all omics variables in a gene set. However, this approach suffers from problems associated with high-dimensional inference. RESULTS: We introduce a tensor-based framework for variable-wise inference in multi-omics analysis. By accounting for the matrix structure of an individual's multi-omics data, the proposed tensor methods incorporate the relationship among omics effects, reduce the number of parameters, and boost the modeling efficiency. We derive the variable-specific tensor test and enhance computational efficiency of tensor modeling. Using simulations and data applications on the Cancer Cell Line Encyclopedia (CCLE), we demonstrate our method performs favorably over baseline methods and will be useful for gaining biological insights in multi-omics analysis. AVAILABILITY AND IMPLEMENTATION: R function and instruction are available from the authors' website: https://www4.stat.ncsu.edu/~jytzeng/Software/TR.omics/TRinstruction.pdf. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Sheng-Mao Chang, Wenbin Lu, Yu-Jyun Huang, Yueyang Huang, Hung Hung, Jeffrey C. Miecznikowski, Tzu-Pin Lu, Jung-Ying Tzeng
Bioinform.8
2021 Transcript annotation tool (TransAT): an R package for retrieving annotations for transcript-specific genetic variants
abstract
BACKGROUND: An individual's genetics play a role in how RNA transcripts are generated from DNA and consequently in their translation into protein. Transcriptional and translational profiling of patients furnishes the information that a specific marker is present; however, it fails to provide evidence whether the marker correlates with response to a therapeutic agent. A comparative analysis of the frequency of genetic variants, such as single nucleotide polymorphisms (SNPs), in diseased and general populations can identify pathogenic variants in individual patients. This is in part because SNPs have considerable effects on protein function and gene expression when they occur in coding regions and regulatory sequences, respectively. Therefore, a tool that can help users to obtain the allele frequency for a corresponding transcript is the need of the day. Several annotation tools such as SNPnexus and VariED are publicly available; however, none of them can use transcript IDs as input and provide the corresponding genomic positions of variants. RESULTS: In this study, we developed an R package, called transcript annotation tool (TransAT), that provides (i) SNP ID and genomic position for a user-provided transcript ID from patients, and (ii) allele frequencies for the SNPs from publicly available global populations. All data elements are extracted, collected, and displayed in an easily downloadable format in two simple command lines. TransAT is available on Windows/Linux/MacOS and is operative for R version 4.0.4 or later. It is available at https://github.com/ShihChingYu/TransAT and can be downloaded and installed using devtools::install_github("ShihChingYu/TransAT", force=T) on the R execution page. Thereafter, all functions can be executed by loading the package into R with library(TransAT). CONCLUSIONS: TransAT is a novel tool that seamlessly provides genetic annotations for queried transcripts. Such easily obtainable information would be greatly advantageous for physicians, assisting them to make individualized decisions about specific drug treatments. Moreover, allele frequencies from user-chosen global ethnic populations will highlight the importance of ethnicity and its effect on patient pathogenicity.
Ching-Yu Shih, Amrita Chattopadhyay, Chien-Hui Wu, Yu-Wen Tien, Tzu-Pin Lu
BMC Bioinform.5
2021 RNASeqR: An R Package for Automated Two-Group RNA-Seq Analysis Workflow
abstract
RNA-Seq analysis has revolutionized researchers' understanding of the transcriptome in biological research. Assessing the differences in transcriptomic profiles between tissue samples or patient groups enables researchers to explore the underlying biological impact of transcription. RNA-Seq analysis requires multiple processing steps and huge computational capabilities. There are many well-developed R packages for individual steps; however, there are few R/Bioconductor packages that integrate existing software tools into a comprehensive RNA-Seq analysis and provide fundamental end-to-end results in pure R environment so that researchers can quickly and easily get fundamental information in big sequencing data. To address this need, we have developed the open source R/Bioconductor package, RNASeqR. It allows users to run an automated RNA-Seq analysis with only six steps, producing essential tabular and graphical results for further biological interpretation. The features of RNASeqR include: six-step analysis, comprehensive visualization, background execution version, and the integration of both R and command-line software. RNASeqR provides fast, light-weight, and easy-to-run RNA-Seq analysis pipeline in pure R environment. It allows users to efficiently utilize popular software tools, including both R/Bioconductor and command-line tools, without predefining the resources or environments. RNASeqR is freely available for Linux and macOS operating systems from Bioconductor (https://bioconductor.org/packages/release/bioc/html/RNASeqR.html).
Kuan-Hao Chao, Yi-Wen Hsiao, Yi-Fang Lee, Chien-Yueh Lee, Liang-Chuan Lai, Mong-Hsun Tsai, Tzu-Pin Lu, Eric Y. Chuang
IEEE ACM Trans. Comput. Biol. Bioinform.7
2020 Genome-Wide Association Study (GWAS) on Metabolic Syndrome in Subjects with Abdominal Obesity in a Taiwanese Population
abstract
Granted that the association of SNPs with metabolic syndrome (MetS) has been reported in diverse populations including Taiwanese, the relationship between abdominal obesity and metabolic syndrome still be vague. In this GWAS study, we assessed SNPs that are associated with MetS in subjects with abdominal obesity. A total of 10,285 Taiwanese subjects were evaluated in this study. We found the association of MetS with three significant SNPs located on different chromosomes at the genome-wide significance level (P-8). The rs662799 SNP in the apolipoprotein A5 (APOA5) gene, the rs325 SNP in lipoprotein lipase (LPL) gene, and the rs247617 SNP in Cholesteryl Ester Transfer Protein (CETP) gene. Also, we surprisingly discovered that the rs325-C and rs247617-A are protective factors and the rs662799-G is a risk factor for subjects with abdominal obesity. Our results supported the APOA5, LPL, and CETP genes that may contribute to the risk of metabolic syndrome in a Taiwanese population.
Kuan-Hung Yeh, Ching-Heng Lin, Tzu-Hung Hsiao, Tzu-Pin Lu
BIBM4
2020 Association test using Copy Number Profile Curves (CONCUR) enhances power in rare copy number variant analysis
abstract
Copy number variants (CNVs) are the gain or loss of DNA segments in the genome that can vary in dosage and length. CNVs comprise a large proportion of variation in human genomes and impact health conditions. To detect rare CNV associations, kernel-based methods have been shown to be a powerful tool due to their flexibility in modeling the aggregate CNV effects, their ability to capture effects from different CNV features, and their accommodation of effect heterogeneity. To perform a kernel association test, a CNV locus needs to be defined so that locus-specific effects can be retained during aggregation. However, CNV loci are arbitrarily defined and different locus definitions can lead to different performance depending on the underlying effect patterns. In this work, we develop a new kernel-based test called CONCUR (i.e., copy number profile curve-based association test) that is free from a definition of locus and evaluates CNV-phenotype associations by comparing individuals' copy number profiles across the genomic regions. CONCUR is built on the proposed concepts of "copy number profile curves" to describe the CNV profile of an individual, and the "common area under the curve (cAUC) kernel" to model the multi-feature CNV effects. The proposed method captures the effects of CNV dosage and length, accounts for the numerical nature of copy numbers, and accommodates between- and within-locus etiological heterogeneity without the need to define artificial CNV loci as required in current kernel methods. In a variety of simulation settings, CONCUR shows comparable or improved power over existing approaches. Real data analyses suggest that CONCUR is well powered to detect CNV effects in the Swedish Schizophrenia Study and the Taiwan Biobank.
Amanda Brucker, Wenbin Lu, Rachel Marceau West, Qi-You Yu, Chuhsing Kate Hsiao, Tzu-Hung Hsiao, Ching-Heng Lin, Patrik K. E. Magnusson, Patrick F. Sullivan, Jin P. Szatkiewicz, Tzu-Pin Lu, Jung-Ying Tzeng
PLoS Comput. Biol.11
2019 anamiR: integrated analysis of MicroRNA and gene expression profiling
abstract
BACKGROUND: With advancements in high-throughput technologies, the cost of obtaining expression profiles of both mRNA and microRNA in the same individual has substantially decreased. Integrated analysis of these profiles can help to elucidate the functional effects of RNA expression in complex diseases, such as cancer. However, fundamental discrepancies are observed in the results from microRNA-mRNA target gene prediction algorithms, and few packages can be used to analyze microRNA and mRNA expression levels simultaneously. RESULTS: To address these issues, an R package, anamiR, was developed. A total of 10 experimental/prediction databases were integrated. Two analytical functions are provided in anamiR, including the single marker test and functional gene set enrichment analysis, and several parameters can be changed by users. Here we demonstrate the potential application of the anamiR package to 2 publicly available microarray datasets. CONCLUSION: The anamiR package is effective for an integrated analysis of both RNA and microRNA profiles. By characterizing biological functions and signaling pathways, this package helps identify dysregulated genes/miRNAs from biological and medical experiments. The source code and manual of the anamiR package are freely available at https://bioconductor.org/packages/release/bioc/html/anamiR.html .
Ti-Tai Wang, Chien-Yueh Lee, Liang-Chuan Lai, Mong-Hsun Tsai, Tzu-Pin Lu, Eric Y. Chuang
BMC Bioinform.5
2018 Probabilistic prioritization of candidate pathway association with pathway score
abstract
BACKGROUND: Current methods for gene-set or pathway analysis are usually designed to test the enrichment of a single gene-set. Once the analysis is carried out for each of the sets under study, a list of significant sets can be obtained. However, if one wishes to further prioritize the importance or strength of association of these sets, no such quantitative measure is available. Using the magnitude of p-value to rank the pathways may not be appropriate because p-value is not a measure for strength of significance. In addition, when testing each pathway, these analyses are often implicitly affected by the number of differentially expressed genes included in the set and/or affected by the dependence among genes. RESULTS: Here we propose a two-stage procedure to prioritize the pathways/gene-sets. In the first stage we develop a pathway-level measure with three properties. First, it contains all genes (differentially expressed or not) in the same set, and summarizes the collective effect of all genes per sample. Second, this pathway score accounts for the correlation between genes by synchronizing their correlation directions. Third, the score includes a rank transformation to enhance the variation among samples as well as to avoid the influence of extreme heterogeneity among genes. In the second stage, all scores are included simultaneously in a Bayesian logistic regression model which can evaluate the strength of association for each set and rank the sets based on posterior probabilities. Simulations from Gaussian distributions and human microarray data, and a breast cancer study with RNA-Seq are considered for demonstration and comparison with other existing methods. CONCLUSIONS: The proposed summary pathway score provides for each sample an overall evaluation of gene expression in a gene-set. It demonstrates the advantages of including all genes in the set and the synchronization of correlation direction. The simultaneous utilization of all pathway-level scores in a Bayesian model not only offers a probabilistic evaluation and ranking of the pathway association but also presents good accuracy in identifying the top-ranking pathways. The resulting recommendation list of ranked pathways can be a reference for potential target therapy or for future allocation of research resources.
Shu-Ju Lin, Tzu-Pin Lu, Qi-You Yu, Chuhsing Kate Hsiao
BMC Bioinform.2
2017 Differential correlation analysis of glioblastoma reveals immune ceRNA interactions predictive of patient survival
abstract
BACKGROUND: Recent studies illuminated a novel role of microRNA (miRNA) in the competing endogenous RNA (ceRNA) interaction: two genes (ceRNAs) can achieve coexpression by competing for a pool of common targeting miRNAs. Individual biological investigations implied ceRNA interaction performs crucial oncogenic/tumor suppressive functions in glioblastoma multiforme (GBM). Yet, a systematic analysis has not been conducted to explore the functional landscape and prognostic significance of ceRNA interaction. RESULTS: Incorporating the knowledge that ceRNA interaction is highly condition-specific and modulated by the expressional abundance of miRNAs, we devised a ceRNA inference by differential correlation analysis to identify the miRNA-modulated ceRNA pairs. Analyzing sample-paired miRNA and gene expression profiles of GBM, our data showed that this alternative layer of gene interaction is essential in global information flow. Functional annotation analysis revealed its involvement in activated processes in brain, such as synaptic transmission, as well as critical tumor-associated functions. Notably, a systematic survival analysis suggested the strength of ceRNA-ceRNA interactions, rather than expressional abundance of individual ceRNAs, among three immune response genes (CCL22, IL2RB, and IRF4) is predictive of patient survival. The prognostic value was validated in two independent cohorts. CONCLUSIONS: This work addresses the lack of a comprehensive exploration into the functional and prognostic relevance of ceRNA interaction in GBM. The proposed efficient and reliable method revealed its significance in GBM-related functions and prognosis. The highlighted roles of ceRNA interaction provide a basis for further biological and clinical investigations.
Yu-Chiao Chiu, Li-Ju Wang, Tzu-Pin Lu, Tzu-Hung Hsiao, Eric Y. Chuang, Yidong Chen 0002
BMC Bioinform.3
2017 iGC - an integrated analysis package of gene expression and copy number alteration
abstract
BACKGROUND: With the advancement in high-throughput technologies, researchers can simultaneously investigate gene expression and copy number alteration (CNA) data from individual patients at a lower cost. Traditional analysis methods analyze each type of data individually and integrate their results using Venn diagrams. Challenges arise, however, when the results are irreproducible and inconsistent across multiple platforms. To address these issues, one possible approach is to concurrently analyze both gene expression profiling and CNAs in the same individual. RESULTS: We have developed an open-source R/Bioconductor package (iGC). Multiple input formats are supported and users can define their own criteria for identifying differentially expressed genes driven by CNAs. The analysis of two real microarray datasets demonstrated that the CNA-driven genes identified by the iGC package showed significantly higher Pearson correlation coefficients with their gene expression levels and copy numbers than those genes located in a genomic region with CNA. Compared with the Venn diagram approach, the iGC package showed better performance. CONCLUSION: The iGC package is effective and useful for identifying CNA-driven genes. By simultaneously considering both comparative genomic and transcriptomic data, it can provide better understanding of biological and medical questions. The iGC package's source code and manual are freely available at https://www.bioconductor.org/packages/release/bioc/html/iGC.html .
Yi-Pin Lai, Liang-Bo Wang, Wei-An Wang, Liang-Chuan Lai, Mong-Hsun Tsai, Tzu-Pin Lu, Eric Y. Chuang
BMC Bioinform.6
2013 Identification of reproducible gene expression signatures in lung adenocarcinoma
abstract
BACKGROUND: Lung cancer is the leading cause of cancer-related death worldwide. Tremendous research efforts have been devoted to improving treatment procedures, but the average five-year overall survival rates are still less than 20%. Many biomarkers have been identified for predicting survival; challenges arise, however, in translating the findings into clinical practice due to their inconsistency and irreproducibility. In this study, we proposed an approach by identifying predictive genes through pathways. RESULTS: The microarrays from Shedden et al. were used as the training set, and the log-rank test was performed to select potential signature genes. We focused on 24 cancer-related pathways from 4 biological databases. A scoring scheme was developed by the Cox hazard regression model, and patients were divided into two groups based on the medians. Subsequently, their predictability and generalizability were evaluated by the 2-fold cross-validation and a resampling test in 4 independent datasets, respectively. A set of 16 genes related to apoptosis execution was demonstrated to have good predictability as well as generalizability in more than 700 lung adenocarcinoma patients and was reproducible in 4 independent datasets. This signature set was shown to have superior performances compared to 6 other published signatures. Furthermore, the corresponding risk scores derived from the set were found to associate with the efficacy of the anti-cancer drug ZD-6474 targeting EGFR. CONCLUSIONS: In summary, we presented a new approach to identify reproducible survival predictors for lung adenocarcinoma, and the identified genes may serve as both prognostic and predictive biomarkers in the future.
Tzu-Pin Lu, Eric Y. Chuang, James J. Chen
BMC Bioinform.1
2010 Concurrent analysis of copy number variation and gene expression: Application in paired non-smoking female lung cancer patients
abstract
This study developed a method to identify disease-correlated pathways by integrating copy numbers (CN) and gene expression (GE). To evaluate the correlation between CN and GE, a suitable window size was assessed by simulation. Gene Set Enrichment Analysis (GSEA) was utilized to identify the possible pathways by CN, GE, and their correlations, respectively. Each of those enriched pathways was further assigned a score to incorporate the information from CN, GE, and their correlations. A dataset of 44 female non-smoking lung cancer patients with both normal and tumor tissues was used to evaluate the performance of this method. To further appraise the predicting abilities of those pathways, patients were classified by support vector machines using the pathways identified by only copy number, only gene expression and incorporating CN, GE, and their correlations. The results showed that the proposed method earned higher accuracy, sensitivity and specificity than traditional methods.
Jung-Chih Chang, Tzu-Pin Lu, Eric Y. Chuang, Liang-Chuan Lai, Mong-Hsun Tsai, Chuhsing Kate Hsiao, Pei-Chun Chen
BIBM2
2010 Concurrent analysis of copy number variations and expression profiles to identify genes associated with tumorigenesis and survival outcome in lung adenocarcinoma
abstract
Lung cancer has been one of the major causes of cancer-related death worldwide. To predict survival outcomes of lung cancer patients, many prognosis gene sets were identified by using gene expression microarrays. However, these gene sets were often inconsistent across independent cohorts. To identify genes with more consistency, we combined gene expression and copy number variations (CNVs). Affymetrix SNP 6.0 and u133plus2.0 microarrays were performed on 42 pairs of lung adenocarcinoma patients. The copy number varied regions (CNVR) existed in more than 30% samples were identified and 475 differentially expressed genes with concordant changes were selected for pathway analysis. Thirteen pathways were significantly enriched among the 475 CNV-associated genes, and survival analyses showed these pathways had generally consistent and significant prediction probabilities across three independent microarray studies. Therefore, integration between gene expression and copy number may help to lower false discovery rate and identify genes used to predict survival outcomes.
Tzu-Pin Lu, Liang-Chuan Lai, Chuhsing Kate Hsiao, Pei-Chun Chen, Mong-Hsun Tsai, Eric Y. Chuang
BIBM1