Mong-Hsun Tsai

dblp:52/6004 · DBLP profile ↗
← Back
13ranked-venue papers
0as first author
5since 2021 · last 2024
0000-0001-8777-5818ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 13 · 5 since 2021
YearPublicationVenuePosition
2024 NURECON: A Novel Online System for Determining Nutrition Requirements Based on Microbial Composition
abstract
Dietary habits have been proven to have an impact on the microbial composition and health of the human gut. Over the past decade, researchers have discovered that gut microbiota can use nutrients to produce metabolites that have major implications for human physiology. However, there is no comprehensive system that specifically focuses on identifying nutrient deficiencies based on gut microbiota, making it difficult to interpret and compare gut microbiome data in the literature. This study proposes an analytical platform, NURECON, that can predict nutrient deficiency information in individuals by comparing their metagenomic information to a reference baseline. NURECON integrates a next-generation bacterial 16S rRNA analytical pipeline (QIIME2), metabolic pathway prediction tools (PICRUSt2 and KEGG), and a food compound database (FooDB) to enable the identification of missing nutrients and provide personalized dietary suggestions. Metagenomic information from total number of 287 healthy subjects was used to establish baseline microbial composition and metabolic profiles. The uploaded data is analyzed and compared to the baseline for nutrient deficiency assessment. Visualization results include gut microbial composition, related enzymes, pathways, and nutrient abundance. NURECON is a user-friendly online platform that provides nutritional advice to support dietitians' research or menu design.
Zhao-Qi Hu, Yuan-Mao Hung, Li-Han Chen, Liang-Chuan Lai, Min-Hsiung Pan, Eric Y. Chuang, Mong-Hsun Tsai
IEEE ACM Trans. Comput. Biol. Bioinform.7
2023 Twnbiome: a public database of the healthy Taiwanese gut microbiome
abstract
With new advances in next generation sequencing (NGS) technology at reduced costs, research on bacterial genomes in the environment has become affordable. Compared to traditional methods, NGS provides high-throughput sequencing reads and the ability to identify many species in the microbiome that were previously unknown. Numerous bioinformatics tools and algorithms have been developed to conduct such analyses. However, in order to obtain biologically meaningful results, the researcher must select the proper tools and combine them to construct an efficient pipeline. This complex procedure may include tens of tools, each of which require correct parameter settings. Furthermore, an NGS data analysis involves multiple series of command-line tools and requires extensive computational resources, which imposes a high barrier for biologists and clinicians to conduct NGS analysis and even interpret their own data. Therefore, we established a public gut microbiome database, which we call Twnbiome, created using healthy subjects from Taiwan, with the goal of enabling microbiota research for the Taiwanese population. Twnbiome provides users with a baseline gut microbiome panel from a healthy Taiwanese cohort, which can be utilized as a reference for conducting case-control studies for a variety of diseases. It is an interactive, informative, and user-friendly database. Twnbiome additionally offers an analysis pipeline, where users can upload their data and download analyzed results. Twnbiome offers an online database which non-bioinformatics users such as clinicians and doctors can not only utilize to access a control set of data, but also analyze raw data with a few easy clicks. All results are customizable with ready-made plots and easily downloadable tables. Database URL: http://twnbiome.cgm.ntu.edu.tw/ .
Amrita Chattopadhyay, Chien-Yueh Lee, Ya-Chin Lee, Chiang-Lin Liu, Hsin-Kuang Chen, Yung-Hua Li, Liang-Chuan Lai, Mong-Hsun Tsai, Yen-Hsuan Ni, Han-Mo Chiu, Tzu-Pin Lu, Eric Y. Chuang
BMC Bioinform.8
2023 Multi-ethnic Imputation System (MI-System): A genotype imputation server for high-dimensional data
abstract
OBJECTIVE: Genotype imputation is a commonly used technique that infers un-typed variants into a study's genotype data, allowing better identification of causal variants in disease studies. However, due to overrepresentation of Caucasian studies, there's a lack of understanding of genetic basis of health-outcomes in other ethnic populations. Therefore, facilitating imputation of missing key-predictor-variants that can potentially improve a risk health-outcome prediction model, specifically for Asian ancestry, is of utmost relevance. METHODS: We aimed to construct an imputation and analysis web-platform, that primarily facilitates, but is not limited to genotype imputation on East-Asians. The goal is to provide a collaborative imputation platform for researchers in the public domain towards rapidly and efficiently conducting accurate genotype imputation. RESULTS: We present an online genotype imputation platform, Multi-ethnic Imputation System (MI-System) (https://misystem.cgm.ntu.edu.tw/), that offers users 3 established pipelines, SHAPEIT2-IMPUTE2, SHAPEIT4-IMPUTE5, and Beagle5.1 for conducting imputation analyses. In addition to 1000 Genomes and Hapmap3, a new customized Taiwan Biobank (TWB) reference panel, specifically created for Taiwanese-Chinese ancestry is provided. MI-System further offers functions to create customized reference panels to be used for imputation, conduct quality control, split whole genome data into chromosomes, and convert genome builds. CONCLUSION: Users can upload their genotype data and perform imputation with minimum effort and resources. The utility functions further can be utilized to preprocess user uploaded data with easy clicks. MI-System potentially contributes to Asian-population genetics research, while eliminating the requirement for high performing computational resources and bioinformatics expertise. It will enable an increased pace of research and provide a knowledge-base for genetic carriers of complex diseases, therefore greatly enhancing patient-driven research. STATEMENT OF SIGNIFICANCE: Multi-ethnic Imputation System (MI-System), primarily facilitates, but is not limited to, imputation on East-Asians, through 3 established prephasing-imputation pipelines, SHAPEIT2-IMPUTE2, SHAPEIT4-IMPUTE5, and Beagle5.1, where users can upload their genotype data and perform imputation and other utility functions with minimum effort and resources. A new customized Taiwan Biobank (TWB) reference panel, specifically created for Taiwanese-Chinese ancestry is provided. Utility functions include (a) create customized reference panels, (b) conduct quality control, (c) split whole genome data into chromosomes, and (d) convert genome builds. Users can also combine 2 reference panels using the system and use combined panels as reference to conduct imputation using MI-System.
Amrita Chattopadhyay, Chien-Yueh Lee, Ying-Cheng Shen, Kuan-Chen Lu, Tzu-Hung Hsiao, Ching-Heng Lin, Liang-Chuan Lai, Mong-Hsun Tsai, Tzu-Pin Lu, Eric Y. Chuang
J. Biomed. Informatics8
2021 High-performance deep learning pipeline predicts individuals in mixtures of DNA using sequencing data
abstract
In this study, we proposed a deep learning (DL) model for classifying individuals from mixtures of DNA samples using 27 short tandem repeats and 94 single nucleotide polymorphisms obtained through massively parallel sequencing protocol. The model was trained/tested/validated with sequenced data from 6 individuals and then evaluated using mixtures from forensic DNA samples. The model successfully identified both the major and the minor contributors with 100% accuracy for 90 DNA mixtures, that were manually prepared by mixing sequence reads of 3 individuals at different ratios. Furthermore, the model identified 100% of the major contributors and 50-80% of the minor contributors in 20 two-sample external-mixed-samples at ratios of 1:39 and 1:9, respectively. To further demonstrate the versatility and applicability of the pipeline, we tested it on whole exome sequence data to classify subtypes of 20 breast cancer patients and achieved an area under curve of 0.85. Overall, we present, for the first time, a complete pipeline, including sequencing data processing steps and DL steps, that is applicable across different NGS platforms. We also introduced a sliding window approach, to overcome the sequence length variation problem of sequencing data, and demonstrate that it improves the model performance dramatically.
Nam Nhut Phan, Amrita Chattopadhyay, Tsui-Ting Lee, Hsiang-I Yin, Tzu-Pin Lu, Liang-Chuan Lai, Hsiao-Lin Hwa, Mong-Hsun Tsai, Eric Y. Chuang
Briefings Bioinform.8
2021 RNASeqR: An R Package for Automated Two-Group RNA-Seq Analysis Workflow
abstract
RNA-Seq analysis has revolutionized researchers' understanding of the transcriptome in biological research. Assessing the differences in transcriptomic profiles between tissue samples or patient groups enables researchers to explore the underlying biological impact of transcription. RNA-Seq analysis requires multiple processing steps and huge computational capabilities. There are many well-developed R packages for individual steps; however, there are few R/Bioconductor packages that integrate existing software tools into a comprehensive RNA-Seq analysis and provide fundamental end-to-end results in pure R environment so that researchers can quickly and easily get fundamental information in big sequencing data. To address this need, we have developed the open source R/Bioconductor package, RNASeqR. It allows users to run an automated RNA-Seq analysis with only six steps, producing essential tabular and graphical results for further biological interpretation. The features of RNASeqR include: six-step analysis, comprehensive visualization, background execution version, and the integration of both R and command-line software. RNASeqR provides fast, light-weight, and easy-to-run RNA-Seq analysis pipeline in pure R environment. It allows users to efficiently utilize popular software tools, including both R/Bioconductor and command-line tools, without predefining the resources or environments. RNASeqR is freely available for Linux and macOS operating systems from Bioconductor (https://bioconductor.org/packages/release/bioc/html/RNASeqR.html).
Kuan-Hao Chao, Yi-Wen Hsiao, Yi-Fang Lee, Chien-Yueh Lee, Liang-Chuan Lai, Mong-Hsun Tsai, Tzu-Pin Lu, Eric Y. Chuang
IEEE ACM Trans. Comput. Biol. Bioinform.6
2019 anamiR: integrated analysis of MicroRNA and gene expression profiling
abstract
BACKGROUND: With advancements in high-throughput technologies, the cost of obtaining expression profiles of both mRNA and microRNA in the same individual has substantially decreased. Integrated analysis of these profiles can help to elucidate the functional effects of RNA expression in complex diseases, such as cancer. However, fundamental discrepancies are observed in the results from microRNA-mRNA target gene prediction algorithms, and few packages can be used to analyze microRNA and mRNA expression levels simultaneously. RESULTS: To address these issues, an R package, anamiR, was developed. A total of 10 experimental/prediction databases were integrated. Two analytical functions are provided in anamiR, including the single marker test and functional gene set enrichment analysis, and several parameters can be changed by users. Here we demonstrate the potential application of the anamiR package to 2 publicly available microarray datasets. CONCLUSION: The anamiR package is effective for an integrated analysis of both RNA and microRNA profiles. By characterizing biological functions and signaling pathways, this package helps identify dysregulated genes/miRNAs from biological and medical experiments. The source code and manual of the anamiR package are freely available at https://bioconductor.org/packages/release/bioc/html/anamiR.html .
Ti-Tai Wang, Chien-Yueh Lee, Liang-Chuan Lai, Mong-Hsun Tsai, Tzu-Pin Lu, Eric Y. Chuang
BMC Bioinform.4
2017 iGC - an integrated analysis package of gene expression and copy number alteration
abstract
BACKGROUND: With the advancement in high-throughput technologies, researchers can simultaneously investigate gene expression and copy number alteration (CNA) data from individual patients at a lower cost. Traditional analysis methods analyze each type of data individually and integrate their results using Venn diagrams. Challenges arise, however, when the results are irreproducible and inconsistent across multiple platforms. To address these issues, one possible approach is to concurrently analyze both gene expression profiling and CNAs in the same individual. RESULTS: We have developed an open-source R/Bioconductor package (iGC). Multiple input formats are supported and users can define their own criteria for identifying differentially expressed genes driven by CNAs. The analysis of two real microarray datasets demonstrated that the CNA-driven genes identified by the iGC package showed significantly higher Pearson correlation coefficients with their gene expression levels and copy numbers than those genes located in a genomic region with CNA. Compared with the Venn diagram approach, the iGC package showed better performance. CONCLUSION: The iGC package is effective and useful for identifying CNA-driven genes. By simultaneously considering both comparative genomic and transcriptomic data, it can provide better understanding of biological and medical questions. The iGC package's source code and manual are freely available at https://www.bioconductor.org/packages/release/bioc/html/iGC.html .
Yi-Pin Lai, Liang-Bo Wang, Wei-An Wang, Liang-Chuan Lai, Mong-Hsun Tsai, Tzu-Pin Lu, Eric Y. Chuang
BMC Bioinform.5
2013 EBARDenovo: highly accurate de novo assembly of RNA-Seq with efficient chimera-detection
abstract
MOTIVATION: High-accuracy de novo assembly of the short sequencing reads from RNA-Seq technology is very challenging. We introduce a de novo assembly algorithm, EBARDenovo, which stands for Extension, Bridging And Repeat-sensing Denovo. This algorithm uses an efficient chimera-detection function to abrogate the effect of aberrant chimeric reads in RNA-Seq data. RESULTS: EBARDenovo resolves the complications of RNA-Seq assembly arising from sequencing errors, repetitive sequences and aberrant chimeric amplicons. In a series of assembly experiments, our algorithm is the most accurate among the examined programs, including de Bruijn graph assemblers, Trinity and Oases. AVAILABILITY AND IMPLEMENTATION: EBARDenovo is available at http://ebardenovo.sourceforge.net/. This software package (with patent pending) is free of charge for academic use only. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Hsueh-Ting Chu, William W. L. Hsiao, Jen-Chih Chen, Tze-Jung Yeh, Mong-Hsun Tsai, Yen-Wenn Liu, Sheng-An Lee, Chaur-Chin Chen, Theresa Tsao, Cheng-Yan Kao
Bioinform.5
2010 Concurrent analysis of copy number variation and gene expression: Application in paired non-smoking female lung cancer patients
abstract
This study developed a method to identify disease-correlated pathways by integrating copy numbers (CN) and gene expression (GE). To evaluate the correlation between CN and GE, a suitable window size was assessed by simulation. Gene Set Enrichment Analysis (GSEA) was utilized to identify the possible pathways by CN, GE, and their correlations, respectively. Each of those enriched pathways was further assigned a score to incorporate the information from CN, GE, and their correlations. A dataset of 44 female non-smoking lung cancer patients with both normal and tumor tissues was used to evaluate the performance of this method. To further appraise the predicting abilities of those pathways, patients were classified by support vector machines using the pathways identified by only copy number, only gene expression and incorporating CN, GE, and their correlations. The results showed that the proposed method earned higher accuracy, sensitivity and specificity than traditional methods.
Jung-Chih Chang, Tzu-Pin Lu, Eric Y. Chuang, Liang-Chuan Lai, Mong-Hsun Tsai, Chuhsing Kate Hsiao, Pei-Chun Chen
BIBM5
2010 Utilizing Cox regression model to assess the relations between predefined gene sets and the survival outcome of lung adenocarcinoma
abstract
The risks of relapse for lung adenocarcinoma patients were still higher than 30%, even after complete surgical resections in early stages. Although lots of prognosis studies using genome-wide profiling had been published, biological meaning and interactions among the prognostic genes were poorly understood. Therefore, we developed a novel method integrating gene set enrichment analysis and Cox-hazard regression model to investigate the relations between predefined gene sets and the survival outcome in lung cancer. The method was able to select gene sets associated with the survival outcome, clustering of the prognostic genes sets, and selection of a representative gene set from each cluster. Furthermore, kernel matrix was used to visualize the similarities between those representative gene sets. In addition to survival outcome, our method can also use other continuous variables to explore other biological interpretation concealed in the predefined gene sets.
Jo-Yang Lu, Eric Y. Chuang, Chuhsing Kate Hsiao, Mong-Hsun Tsai, Liang-Chuan Lai, Pei-Chun Chen
BIBM4
2010 Concurrent analysis of copy number variations and expression profiles to identify genes associated with tumorigenesis and survival outcome in lung adenocarcinoma
abstract
Lung cancer has been one of the major causes of cancer-related death worldwide. To predict survival outcomes of lung cancer patients, many prognosis gene sets were identified by using gene expression microarrays. However, these gene sets were often inconsistent across independent cohorts. To identify genes with more consistency, we combined gene expression and copy number variations (CNVs). Affymetrix SNP 6.0 and u133plus2.0 microarrays were performed on 42 pairs of lung adenocarcinoma patients. The copy number varied regions (CNVR) existed in more than 30% samples were identified and 475 differentially expressed genes with concordant changes were selected for pathway analysis. Thirteen pathways were significantly enriched among the 475 CNV-associated genes, and survival analyses showed these pathways had generally consistent and significant prediction probabilities across three independent microarray studies. Therefore, integration between gene expression and copy number may help to lower false discovery rate and identify genes used to predict survival outcomes.
Tzu-Pin Lu, Liang-Chuan Lai, Chuhsing Kate Hsiao, Pei-Chun Chen, Mong-Hsun Tsai, Eric Y. Chuang
BIBM5
2008 A probe-density-based analysis method for array CGH data: simulation, normalization and centralization
abstract
MOTIVATION: Genomic instability is one of the fundamental factors in tumorigenesis and tumor progression. Many studies have shown that copy-number abnormalities at the DNA level are important in the pathogenesis of cancer. Array comparative genomic hybridization (aCGH), developed based on expression microarray technology, can reveal the chromosomal aberrations in segmental copies at a high resolution. However, due to the nature of aCGH, many standard expression data processing tools, such as data normalization, often fail to yield satisfactory results. RESULTS: We demonstrated a novel aCGH normalization algorithm, which provides an accurate aCGH data normalization by utilizing the dependency of neighboring probe measurements in aCGH experiments. To facilitate the study, we have developed a hidden Markov model (HMM) to simulate a series of aCGH experiments with random DNA copy number alterations that are used to validate the performance of our normalization. In addition, we applied the proposed normalization algorithm to an aCGH study of lung cancer cell lines. By using the proposed algorithm, data quality and the reliability of experimental results are significantly improved, and the distinct patterns of DNA copy number alternations are observed among those lung cancer cell lines. SUPPLEMENTARY INFORMATION: Source codes and.gures may be found at http://ntumaps.cgm.ntu.edu.tw/aCGH_supplementary.
Hung-I Harry Chen, Fang-Han Hsu, Mong-Hsun Tsai, Pan-Chyr Yang, Paul S. Meltzer, Eric Y. Chuang, Yidong Chen 0002
Bioinform.4
2008 A probe-density-based analysis method for array CGH data: simulation, normalization and centralization
abstract
Bioinformatics 2008; Vol. 24 no. 16: 1749–1756. The publishers regret that there was an error in the copyright line which should have been Open Access. This article is now freely available online. The copyright line of the article should have read: © 2008 The Author(s) This is an Open Access article distributed under the terms of the Creative Commons Attribution Non-Commercial License (http://creativecommons.org/licenses/by-nc/2.0/uk/) which permits unrestricted non-commercial use, distribution, and reproduction in any medium, provided the original work is properly cited.
Hung-I Harry Chen, Fang-Han Hsu, Mong-Hsun Tsai, Pan-Chyr Yang, Paul S. Meltzer, Eric Y. Chuang, Yidong Chen 0002
Bioinform.4