EDBT 2026 Demo / reviewers in the wild / expert
Eric Y. Chuang
dblp:97/879
· DBLP profile ↗
21ranked-venue papers
0as first author
6since 2021 · last 2024
0000-0003-2530-0096ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 21 · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | NURECON: A Novel Online System for Determining Nutrition Requirements Based on Microbial CompositionabstractDietary habits have been proven to have an impact on the microbial composition and health of the human gut. Over the past decade, researchers have discovered that gut microbiota can use nutrients to produce metabolites that have major implications for human physiology. However, there is no comprehensive system that specifically focuses on identifying nutrient deficiencies based on gut microbiota, making it difficult to interpret and compare gut microbiome data in the literature. This study proposes an analytical platform, NURECON, that can predict nutrient deficiency information in individuals by comparing their metagenomic information to a reference baseline. NURECON integrates a next-generation bacterial 16S rRNA analytical pipeline (QIIME2), metabolic pathway prediction tools (PICRUSt2 and KEGG), and a food compound database (FooDB) to enable the identification of missing nutrients and provide personalized dietary suggestions. Metagenomic information from total number of 287 healthy subjects was used to establish baseline microbial composition and metabolic profiles. The uploaded data is analyzed and compared to the baseline for nutrient deficiency assessment. Visualization results include gut microbial composition, related enzymes, pathways, and nutrient abundance. NURECON is a user-friendly online platform that provides nutritional advice to support dietitians' research or menu design. Zhao-Qi Hu, Yuan-Mao Hung, Li-Han Chen, Liang-Chuan Lai, Min-Hsiung Pan, Eric Y. Chuang, Mong-Hsun Tsai |
IEEE ACM Trans. Comput. Biol. Bioinform. | 6 |
| 2023 | Twnbiome: a public database of the healthy Taiwanese gut microbiomeabstractWith new advances in next generation sequencing (NGS) technology at reduced costs, research on bacterial genomes in the environment has become affordable. Compared to traditional methods, NGS provides high-throughput sequencing reads and the ability to identify many species in the microbiome that were previously unknown. Numerous bioinformatics tools and algorithms have been developed to conduct such analyses. However, in order to obtain biologically meaningful results, the researcher must select the proper tools and combine them to construct an efficient pipeline. This complex procedure may include tens of tools, each of which require correct parameter settings. Furthermore, an NGS data analysis involves multiple series of command-line tools and requires extensive computational resources, which imposes a high barrier for biologists and clinicians to conduct NGS analysis and even interpret their own data. Therefore, we established a public gut microbiome database, which we call Twnbiome, created using healthy subjects from Taiwan, with the goal of enabling microbiota research for the Taiwanese population. Twnbiome provides users with a baseline gut microbiome panel from a healthy Taiwanese cohort, which can be utilized as a reference for conducting case-control studies for a variety of diseases. It is an interactive, informative, and user-friendly database. Twnbiome additionally offers an analysis pipeline, where users can upload their data and download analyzed results. Twnbiome offers an online database which non-bioinformatics users such as clinicians and doctors can not only utilize to access a control set of data, but also analyze raw data with a few easy clicks. All results are customizable with ready-made plots and easily downloadable tables. Database URL: http://twnbiome.cgm.ntu.edu.tw/ . Amrita Chattopadhyay, Chien-Yueh Lee, Ya-Chin Lee, Chiang-Lin Liu, Hsin-Kuang Chen, Yung-Hua Li, Liang-Chuan Lai, Mong-Hsun Tsai, Yen-Hsuan Ni, Han-Mo Chiu, Tzu-Pin Lu, Eric Y. Chuang |
BMC Bioinform. | 12 |
| 2023 | Multi-ethnic Imputation System (MI-System): A genotype imputation server for high-dimensional dataabstractOBJECTIVE: Genotype imputation is a commonly used technique that infers un-typed variants into a study's genotype data, allowing better identification of causal variants in disease studies. However, due to overrepresentation of Caucasian studies, there's a lack of understanding of genetic basis of health-outcomes in other ethnic populations. Therefore, facilitating imputation of missing key-predictor-variants that can potentially improve a risk health-outcome prediction model, specifically for Asian ancestry, is of utmost relevance. METHODS: We aimed to construct an imputation and analysis web-platform, that primarily facilitates, but is not limited to genotype imputation on East-Asians. The goal is to provide a collaborative imputation platform for researchers in the public domain towards rapidly and efficiently conducting accurate genotype imputation. RESULTS: We present an online genotype imputation platform, Multi-ethnic Imputation System (MI-System) (https://misystem.cgm.ntu.edu.tw/), that offers users 3 established pipelines, SHAPEIT2-IMPUTE2, SHAPEIT4-IMPUTE5, and Beagle5.1 for conducting imputation analyses. In addition to 1000 Genomes and Hapmap3, a new customized Taiwan Biobank (TWB) reference panel, specifically created for Taiwanese-Chinese ancestry is provided. MI-System further offers functions to create customized reference panels to be used for imputation, conduct quality control, split whole genome data into chromosomes, and convert genome builds. CONCLUSION: Users can upload their genotype data and perform imputation with minimum effort and resources. The utility functions further can be utilized to preprocess user uploaded data with easy clicks. MI-System potentially contributes to Asian-population genetics research, while eliminating the requirement for high performing computational resources and bioinformatics expertise. It will enable an increased pace of research and provide a knowledge-base for genetic carriers of complex diseases, therefore greatly enhancing patient-driven research. STATEMENT OF SIGNIFICANCE: Multi-ethnic Imputation System (MI-System), primarily facilitates, but is not limited to, imputation on East-Asians, through 3 established prephasing-imputation pipelines, SHAPEIT2-IMPUTE2, SHAPEIT4-IMPUTE5, and Beagle5.1, where users can upload their genotype data and perform imputation and other utility functions with minimum effort and resources. A new customized Taiwan Biobank (TWB) reference panel, specifically created for Taiwanese-Chinese ancestry is provided. Utility functions include (a) create customized reference panels, (b) conduct quality control, (c) split whole genome data into chromosomes, and (d) convert genome builds. Users can also combine 2 reference panels using the system and use combined panels as reference to conduct imputation using MI-System. Amrita Chattopadhyay, Chien-Yueh Lee, Ying-Cheng Shen, Kuan-Chen Lu, Tzu-Hung Hsiao, Ching-Heng Lin, Liang-Chuan Lai, Mong-Hsun Tsai, Tzu-Pin Lu, Eric Y. Chuang |
J. Biomed. Informatics | 10 |
| 2022 | CLIN_SKAT: an R package to conduct association analysis using functionally relevant variantsabstractBACKGROUND: Availability of next generation sequencing data, allows low-frequency and rare variants to be studied through strategies other than the commonly used genome-wide association studies (GWAS). Rare variants are important keys towards explaining the heritability for complex diseases that remains to be explained by common variants due to their low effect sizes. However, analysis strategies struggle to keep up with the huge amount of data at disposal therefore creating a bottleneck. This study describes CLIN_SKAT, an R package, that provides users with an easily implemented analysis pipeline with the goal of (i) extracting clinically relevant variants (both rare and common), followed by (ii) gene-based association analysis by grouping the selected variants. RESULTS: CLIN_SKAT offers four simple functions that can be used to obtain clinically relevant variants, map them to genes or gene sets, calculate weights from global healthy populations and conduct weighted case-control analysis. CLIN_SKAT introduces improvements by adding certain pre-analysis steps and customizable features to make the SKAT results clinically more meaningful. Moreover, it offers several plot functions that can be availed towards obtaining visualizations for interpretation of the analyses results. CLIN_SKAT is available on Windows/Linux/MacOS and is operative for R version 4.0.4 or later. It can be freely downloaded from https://github.com/ShihChingYu/CLIN_SKAT , installed through devtools::install_github("ShihChingYu/CLIN_SKAT", force=T) and executed by loading the package into R using library(CLIN_SKAT). All outputs (tabular and graphical) can be downloaded in simple, publishable formats. CONCLUSIONS: Statistical association analysis is often underpowered due to low sample sizes and high numbers of variants to be tested, limiting detection of causal ones. Therefore, retaining a subset of variants that are biologically meaningful seems to be a more effective strategy for identifying explainable associations while reducing the degrees of freedom. CLIN_SKAT offers users a one-stop R package that identifies disease risk variants with improved power via a series of tailor-made procedures that allows dimension reduction, by retaining functionally relevant variants, and incorporating ethnicity based priors. Furthermore, it also eliminates the requirement for high computational resources and bioinformatics expertise. Amrita Chattopadhyay, Ching-Yu Shih, Yu-Chen Hsu, Jyh-Ming Jimmy Juang, Eric Y. Chuang, Tzu-Pin Lu |
BMC Bioinform. | 5 |
| 2021 | High-performance deep learning pipeline predicts individuals in mixtures of DNA using sequencing dataabstractIn this study, we proposed a deep learning (DL) model for classifying individuals from mixtures of DNA samples using 27 short tandem repeats and 94 single nucleotide polymorphisms obtained through massively parallel sequencing protocol. The model was trained/tested/validated with sequenced data from 6 individuals and then evaluated using mixtures from forensic DNA samples. The model successfully identified both the major and the minor contributors with 100% accuracy for 90 DNA mixtures, that were manually prepared by mixing sequence reads of 3 individuals at different ratios. Furthermore, the model identified 100% of the major contributors and 50-80% of the minor contributors in 20 two-sample external-mixed-samples at ratios of 1:39 and 1:9, respectively. To further demonstrate the versatility and applicability of the pipeline, we tested it on whole exome sequence data to classify subtypes of 20 breast cancer patients and achieved an area under curve of 0.85. Overall, we present, for the first time, a complete pipeline, including sequencing data processing steps and DL steps, that is applicable across different NGS platforms. We also introduced a sliding window approach, to overcome the sequence length variation problem of sequencing data, and demonstrate that it improves the model performance dramatically. Nam Nhut Phan, Amrita Chattopadhyay, Tsui-Ting Lee, Hsiang-I Yin, Tzu-Pin Lu, Liang-Chuan Lai, Hsiao-Lin Hwa, Mong-Hsun Tsai, Eric Y. Chuang |
Briefings Bioinform. | 9 |
| 2021 | RNASeqR: An R Package for Automated Two-Group RNA-Seq Analysis WorkflowabstractRNA-Seq analysis has revolutionized researchers' understanding of the transcriptome in biological research. Assessing the differences in transcriptomic profiles between tissue samples or patient groups enables researchers to explore the underlying biological impact of transcription. RNA-Seq analysis requires multiple processing steps and huge computational capabilities. There are many well-developed R packages for individual steps; however, there are few R/Bioconductor packages that integrate existing software tools into a comprehensive RNA-Seq analysis and provide fundamental end-to-end results in pure R environment so that researchers can quickly and easily get fundamental information in big sequencing data. To address this need, we have developed the open source R/Bioconductor package, RNASeqR. It allows users to run an automated RNA-Seq analysis with only six steps, producing essential tabular and graphical results for further biological interpretation. The features of RNASeqR include: six-step analysis, comprehensive visualization, background execution version, and the integration of both R and command-line software. RNASeqR provides fast, light-weight, and easy-to-run RNA-Seq analysis pipeline in pure R environment. It allows users to efficiently utilize popular software tools, including both R/Bioconductor and command-line tools, without predefining the resources or environments. RNASeqR is freely available for Linux and macOS operating systems from Bioconductor (https://bioconductor.org/packages/release/bioc/html/RNASeqR.html). Kuan-Hao Chao, Yi-Wen Hsiao, Yi-Fang Lee, Chien-Yueh Lee, Liang-Chuan Lai, Mong-Hsun Tsai, Tzu-Pin Lu, Eric Y. Chuang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 8 |
| 2019 | anamiR: integrated analysis of MicroRNA and gene expression profilingabstractBACKGROUND: With advancements in high-throughput technologies, the cost of obtaining expression profiles of both mRNA and microRNA in the same individual has substantially decreased. Integrated analysis of these profiles can help to elucidate the functional effects of RNA expression in complex diseases, such as cancer. However, fundamental discrepancies are observed in the results from microRNA-mRNA target gene prediction algorithms, and few packages can be used to analyze microRNA and mRNA expression levels simultaneously. RESULTS: To address these issues, an R package, anamiR, was developed. A total of 10 experimental/prediction databases were integrated. Two analytical functions are provided in anamiR, including the single marker test and functional gene set enrichment analysis, and several parameters can be changed by users. Here we demonstrate the potential application of the anamiR package to 2 publicly available microarray datasets. CONCLUSION: The anamiR package is effective for an integrated analysis of both RNA and microRNA profiles. By characterizing biological functions and signaling pathways, this package helps identify dysregulated genes/miRNAs from biological and medical experiments. The source code and manual of the anamiR package are freely available at https://bioconductor.org/packages/release/bioc/html/anamiR.html . Ti-Tai Wang, Chien-Yueh Lee, Liang-Chuan Lai, Mong-Hsun Tsai, Tzu-Pin Lu, Eric Y. Chuang |
BMC Bioinform. | 6 |
| 2018 | Analyzing Differential Regulatory Networks Modulated by Continuous-State Genomic Features in Glioblastoma MultiformeabstractGene regulatory networks are a global representation of complex interactions between molecules that dictate cellular behavior. Study of a regulatory network modulated by single or multiple modulators' expression levels, including microRNAs (miRNAs) and transcription factors (TFs), in different conditions can further reveal the modulators' roles in diseases such as cancers. Existing computational methods for identifying such modulated regulatory networks are typically carried out by comparing groups of samples dichotomized with respect to the modulator status, ignoring the fact that most biological features are intrinsically continuous variables. Here, we devised a sliding window-based regression scheme and proposed the Regression-based Inference of Modulation (RIM) algorithm to infer the dynamic gene regulation modulated by continuous-state modulators. We demonstrated the improvement in performance as well as computation efficiency achieved by RIM. Applying RIM to genome-wide expression profiles of 520 glioblastoma multiforme (GBM) tumors, we investigated miRNA- and TF-modulated gene regulatory networks and showed their association with dynamic cellular processes and brain-related functions in GBM. Overall, the proposed algorithm provides an efficient and robust scheme for comprehensively studying modulated gene regulatory networks. Yu-Chiao Chiu, Tzu-Hung Hsiao, Li-Ju Wang, Yidong Chen 0002, Eric Y. Chuang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2017 | Differential correlation analysis of glioblastoma reveals immune ceRNA interactions predictive of patient survivalabstractBACKGROUND: Recent studies illuminated a novel role of microRNA (miRNA) in the competing endogenous RNA (ceRNA) interaction: two genes (ceRNAs) can achieve coexpression by competing for a pool of common targeting miRNAs. Individual biological investigations implied ceRNA interaction performs crucial oncogenic/tumor suppressive functions in glioblastoma multiforme (GBM). Yet, a systematic analysis has not been conducted to explore the functional landscape and prognostic significance of ceRNA interaction. RESULTS: Incorporating the knowledge that ceRNA interaction is highly condition-specific and modulated by the expressional abundance of miRNAs, we devised a ceRNA inference by differential correlation analysis to identify the miRNA-modulated ceRNA pairs. Analyzing sample-paired miRNA and gene expression profiles of GBM, our data showed that this alternative layer of gene interaction is essential in global information flow. Functional annotation analysis revealed its involvement in activated processes in brain, such as synaptic transmission, as well as critical tumor-associated functions. Notably, a systematic survival analysis suggested the strength of ceRNA-ceRNA interactions, rather than expressional abundance of individual ceRNAs, among three immune response genes (CCL22, IL2RB, and IRF4) is predictive of patient survival. The prognostic value was validated in two independent cohorts. CONCLUSIONS: This work addresses the lack of a comprehensive exploration into the functional and prognostic relevance of ceRNA interaction in GBM. The proposed efficient and reliable method revealed its significance in GBM-related functions and prognosis. The highlighted roles of ceRNA interaction provide a basis for further biological and clinical investigations. Yu-Chiao Chiu, Li-Ju Wang, Tzu-Pin Lu, Tzu-Hung Hsiao, Eric Y. Chuang, Yidong Chen 0002 |
BMC Bioinform. | 5 |
| 2017 | iGC - an integrated analysis package of gene expression and copy number alterationabstractBACKGROUND: With the advancement in high-throughput technologies, researchers can simultaneously investigate gene expression and copy number alteration (CNA) data from individual patients at a lower cost. Traditional analysis methods analyze each type of data individually and integrate their results using Venn diagrams. Challenges arise, however, when the results are irreproducible and inconsistent across multiple platforms. To address these issues, one possible approach is to concurrently analyze both gene expression profiling and CNAs in the same individual. RESULTS: We have developed an open-source R/Bioconductor package (iGC). Multiple input formats are supported and users can define their own criteria for identifying differentially expressed genes driven by CNAs. The analysis of two real microarray datasets demonstrated that the CNA-driven genes identified by the iGC package showed significantly higher Pearson correlation coefficients with their gene expression levels and copy numbers than those genes located in a genomic region with CNA. Compared with the Venn diagram approach, the iGC package showed better performance. CONCLUSION: The iGC package is effective and useful for identifying CNA-driven genes. By simultaneously considering both comparative genomic and transcriptomic data, it can provide better understanding of biological and medical questions. The iGC package's source code and manual are freely available at https://www.bioconductor.org/packages/release/bioc/html/iGC.html . Yi-Pin Lai, Liang-Bo Wang, Wei-An Wang, Liang-Chuan Lai, Mong-Hsun Tsai, Tzu-Pin Lu, Eric Y. Chuang |
BMC Bioinform. | 7 |
| 2015 | Analyzing differential regulatory networks modulated by continuous-state genomic features in glioblastoma multiformeabstractGene regulatory networks are a global representation of complex interactions between molecules that dictate cellular behavior. Study of a regulatory network modulated by single or multiple modulators' expression levels, including microRNAs (miRNAs) and transcription factors (TFs), in different conditions can further reveal the modulators' roles in diseases such as cancers. Existing computational methods for identifying such modulated regulatory networks are typically carried out by comparing groups of samples dichotomized with respect to the modulator status, ignoring the fact that most biological features are intrinsically continuous variables. Here we devised a sliding window-based regression scheme and proposed the Regression-based Inference of Modulation (RIM) algorithm to infer the dynamic gene regulation modulated by continuous-state modulators. We demonstrated the improvement in performance as well as computation efficiency achieved by RIM. Applying RIM to genome-wide expression profiles of 520 glioblastoma multiforme (GBM) tumors, we investigated miRNA- and TF-modulated gene regulatory networks and showed their association with dynamic cellular processes and brain-related functions in GBM. Overall, the proposed algorithm provides an efficient and robust scheme for comprehensively studying modulated gene regulatory networks. Yu-Chiao Chiu, Kai-Wen Liang, Tzu-Hung Hsiao, Yidong Chen 0002, Eric Y. Chuang |
BIBM | 5 |
| 2015 | Patho-finder - A fast and accurate program for pathogen identification through RNA-seqabstractTechnology of next generation sequencing to detect pathogens of sample can impact human health by revealing pathogens which cause disease. Several workflow has developed in purposed to detect pathogens in next generation sequencing data. However, the requirement of computation power of these workflows limited the application. The time consuming problem make the workflow difficult to detect datasets with large sample size. Here we presented Patho-finder, a fast and accurate workflow designed for detecting pathogen in RNA sequencing data. We have evaluated performance of Patho-finder by three aspects. First, we evaluate performance by alter the data features, to see how Patho-finder work under different simulation conditions. Next, we compare the time consuming and accuracy between Patho-finder and existing workflow. At last, we used Patho-finder on the RNA-seq of cell lines with known virus-infected. The validation result demonstrated our approach could finish the task in real datasets. Chin-Ting Wu, Tzu-Hung Hsiao, Yu-Chiao Chiu, Yu-Ching Hsu, Eric Y. Chuang, Yidong Chen 0002 |
BIBM | 5 |
| 2013 | Modeling competing endogenous RNA regulatory networks in glioblastoma multiformeabstractRecent studies postulated that genes harboring identical microRNA (miRNA) binding sites can crosstalk by competing for a limited pool of the binding miRNAs (the miRNA program), named as the regulation of competing endogenous RNAs (ceRNAs). Incorporating recent biological evidence that ceRNA regulation depends on miRNA program expression levels, we developed, in the present study, a mathematical model for systematically inferring ceRNA regulation that is dependent on expression levels of the miRNA programs from sample-paired mRNA and miRNA expression datasets. Applying the method to analyze glioblastoma datasets, a compact ceRNA regulatory network was constructed. Our data further demonstrated that ceRNA regulation plays an essential role in transient cellular responses to dynamic inter-cellular signals. The findings illuminate mechanism of ceRNA regulation and further provide biological insights into the complex human interactome. Yu-Chiao Chiu, Eric Y. Chuang, Tzu-Hung Hsiao, Yidong Chen 0002 |
BIBM | 2 |
| 2013 | Identification of reproducible gene expression signatures in lung adenocarcinomaabstractBACKGROUND: Lung cancer is the leading cause of cancer-related death worldwide. Tremendous research efforts have been devoted to improving treatment procedures, but the average five-year overall survival rates are still less than 20%. Many biomarkers have been identified for predicting survival; challenges arise, however, in translating the findings into clinical practice due to their inconsistency and irreproducibility. In this study, we proposed an approach by identifying predictive genes through pathways. RESULTS: The microarrays from Shedden et al. were used as the training set, and the log-rank test was performed to select potential signature genes. We focused on 24 cancer-related pathways from 4 biological databases. A scoring scheme was developed by the Cox hazard regression model, and patients were divided into two groups based on the medians. Subsequently, their predictability and generalizability were evaluated by the 2-fold cross-validation and a resampling test in 4 independent datasets, respectively. A set of 16 genes related to apoptosis execution was demonstrated to have good predictability as well as generalizability in more than 700 lung adenocarcinoma patients and was reproducible in 4 independent datasets. This signature set was shown to have superior performances compared to 6 other published signatures. Furthermore, the corresponding risk scores derived from the set were found to associate with the efficacy of the anti-cancer drug ZD-6474 targeting EGFR. CONCLUSIONS: In summary, we presented a new approach to identify reproducible survival predictors for lung adenocarcinoma, and the identified genes may serve as both prognostic and predictive biomarkers in the future. Tzu-Pin Lu, Eric Y. Chuang, James J. Chen |
BMC Bioinform. | 2 |
| 2012 | Using gene sets to identify putative drugs for breast cancerabstractThe number of current anti-cancer drugs was limited and the response rates were also not high. To "reposition" known drugs as anti-cancer drugs to increase the therapeutic efficiency, we presented a novel analysis framework to identify putative drugs for cancer. Using breast cancer as example, a "cancer - gene sets - drugs" network was constructed through two procedures. First, the "gene sets - drugs" network was built by applying the expression pattern of drugs for gene set enrichment analysis. Secondly, the breast cancer progression associated gene sets were identified by survival analysis of patient cohorts. By integrating the two results, 25 tumor progression associated gene sets and 360 putative anti-cancer drugs were identified. Our method has the ability to identify the "reposition" drugs and the potential affected mechanisms of tumor progression concurrently. It will be useful to speed up the development of anti-cancer drugs from bench to clinical application. Tzu-Hung Hsiao, Hung-I Harry Chen, Yidong Chen 0002, Yu-Heng Chen, Eric Y. Chuang |
BIBM | 5 |
| 2010 | Concurrent analysis of copy number variation and gene expression: Application in paired non-smoking female lung cancer patientsabstractThis study developed a method to identify disease-correlated pathways by integrating copy numbers (CN) and gene expression (GE). To evaluate the correlation between CN and GE, a suitable window size was assessed by simulation. Gene Set Enrichment Analysis (GSEA) was utilized to identify the possible pathways by CN, GE, and their correlations, respectively. Each of those enriched pathways was further assigned a score to incorporate the information from CN, GE, and their correlations. A dataset of 44 female non-smoking lung cancer patients with both normal and tumor tissues was used to evaluate the performance of this method. To further appraise the predicting abilities of those pathways, patients were classified by support vector machines using the pathways identified by only copy number, only gene expression and incorporating CN, GE, and their correlations. The results showed that the proposed method earned higher accuracy, sensitivity and specificity than traditional methods. Jung-Chih Chang, Tzu-Pin Lu, Eric Y. Chuang, Liang-Chuan Lai, Mong-Hsun Tsai, Chuhsing Kate Hsiao, Pei-Chun Chen |
BIBM | 3 |
| 2010 | Utilizing Cox regression model to assess the relations between predefined gene sets and the survival outcome of lung adenocarcinomaabstractThe risks of relapse for lung adenocarcinoma patients were still higher than 30%, even after complete surgical resections in early stages. Although lots of prognosis studies using genome-wide profiling had been published, biological meaning and interactions among the prognostic genes were poorly understood. Therefore, we developed a novel method integrating gene set enrichment analysis and Cox-hazard regression model to investigate the relations between predefined gene sets and the survival outcome in lung cancer. The method was able to select gene sets associated with the survival outcome, clustering of the prognostic genes sets, and selection of a representative gene set from each cluster. Furthermore, kernel matrix was used to visualize the similarities between those representative gene sets. In addition to survival outcome, our method can also use other continuous variables to explore other biological interpretation concealed in the predefined gene sets. Jo-Yang Lu, Eric Y. Chuang, Chuhsing Kate Hsiao, Mong-Hsun Tsai, Liang-Chuan Lai, Pei-Chun Chen |
BIBM | 2 |
| 2010 | Concurrent analysis of copy number variations and expression profiles to identify genes associated with tumorigenesis and survival outcome in lung adenocarcinomaabstractLung cancer has been one of the major causes of cancer-related death worldwide. To predict survival outcomes of lung cancer patients, many prognosis gene sets were identified by using gene expression microarrays. However, these gene sets were often inconsistent across independent cohorts. To identify genes with more consistency, we combined gene expression and copy number variations (CNVs). Affymetrix SNP 6.0 and u133plus2.0 microarrays were performed on 42 pairs of lung adenocarcinoma patients. The copy number varied regions (CNVR) existed in more than 30% samples were identified and 475 differentially expressed genes with concordant changes were selected for pathway analysis. Thirteen pathways were significantly enriched among the 475 CNV-associated genes, and survival analyses showed these pathways had generally consistent and significant prediction probabilities across three independent microarray studies. Therefore, integration between gene expression and copy number may help to lower false discovery rate and identify genes used to predict survival outcomes. Tzu-Pin Lu, Liang-Chuan Lai, Chuhsing Kate Hsiao, Pei-Chun Chen, Mong-Hsun Tsai, Eric Y. Chuang |
BIBM | 6 |
| 2008 | A probe-density-based analysis method for array CGH data: simulation, normalization and centralizationabstractMOTIVATION: Genomic instability is one of the fundamental factors in tumorigenesis and tumor progression. Many studies have shown that copy-number abnormalities at the DNA level are important in the pathogenesis of cancer. Array comparative genomic hybridization (aCGH), developed based on expression microarray technology, can reveal the chromosomal aberrations in segmental copies at a high resolution. However, due to the nature of aCGH, many standard expression data processing tools, such as data normalization, often fail to yield satisfactory results. RESULTS: We demonstrated a novel aCGH normalization algorithm, which provides an accurate aCGH data normalization by utilizing the dependency of neighboring probe measurements in aCGH experiments. To facilitate the study, we have developed a hidden Markov model (HMM) to simulate a series of aCGH experiments with random DNA copy number alterations that are used to validate the performance of our normalization. In addition, we applied the proposed normalization algorithm to an aCGH study of lung cancer cell lines. By using the proposed algorithm, data quality and the reliability of experimental results are significantly improved, and the distinct patterns of DNA copy number alternations are observed among those lung cancer cell lines. SUPPLEMENTARY INFORMATION: Source codes and.gures may be found at http://ntumaps.cgm.ntu.edu.tw/aCGH_supplementary. Hung-I Harry Chen, Fang-Han Hsu, Mong-Hsun Tsai, Pan-Chyr Yang, Paul S. Meltzer, Eric Y. Chuang, Yidong Chen 0002 |
Bioinform. | 7 |
| 2008 | A probe-density-based analysis method for array CGH data: simulation, normalization and centralizationabstractBioinformatics 2008; Vol. 24 no. 16: 1749–1756. The publishers regret that there was an error in the copyright line which should have been Open Access. This article is now freely available online. The copyright line of the article should have read: © 2008 The Author(s) This is an Open Access article distributed under the terms of the Creative Commons Attribution Non-Commercial License (http://creativecommons.org/licenses/by-nc/2.0/uk/) which permits unrestricted non-commercial use, distribution, and reproduction in any medium, provided the original work is properly cited. Hung-I Harry Chen, Fang-Han Hsu, Mong-Hsun Tsai, Pan-Chyr Yang, Paul S. Meltzer, Eric Y. Chuang, Yidong Chen 0002 |
Bioinform. | 7 |
| 2004 | Correcting log ratios for signal saturation in cDNA microarraysabstractMOTIVATION: Pixel saturation occurs when the pixel intensity exceeds a threshold and the recorded pixel intensity is truncated. Microarray experiments are commonly afflicted with saturated pixels. As a result, estimators of gene expression are biased, with the amount of bias increasing as a function of the proportion of pixels saturated. Saturation is directly related to the photomultiplier tube (PMT) voltage settings and RNA abundance and is not necessarily associated with poor array or poor spot quality. When choosing PMT settings, higher PMT settings are desired because of improved signal-to-noise ratios of low-intensity spots. This improved signal is somewhat offset by saturation of high-intensity spots. In practice, spots with saturated pixels are discarded or the biased value is used. Neither of these approaches is appealing, particularly the former approach when a highly expressed gene is discarded because of saturation. RESULTS: We present a method to correct for saturation using pixel-level data. The method is based on a censored regression model. Evaluations on several arrays indicate that the method performs well. Simulation studies suggest that the method is robust under certain model violations. Lori E. Dodd, Edward L. Korn, Lisa M. McShane, G. V. R. Chandramouli, Eric Y. Chuang |
Bioinform. | 5 |