Tin Chi Nguyen

dblp:94/10786 · also Tin Nguyen 0001 · DBLP profile ↗
← Back
12ranked-venue papers
2as first author
6since 2021 · last 2026
0000-0001-8001-9470ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 11 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 1
YearPublicationVenuePosition
2026 CSIE: cancer subtyping via inference and ensemble
abstract
While multi-omics integration is the gold standard for precision oncology, its clinical utility is severely hampered by the incomplete data problem, where cost and technical barriers often leave researchers with only single-omics profiles. Our manuscript introduces CSIE (cancer subtyping via inference and ensemble), a framework that bridges this gap by using a novel transformer-based inference module which incorporates systems-level knowledge to accurately infer missing omics layers from gene expression data. Furthermore, CSIE employs an ensemble clustering module that simultaneously integrates multi-omics data via different similarity metrics and clustering algorithms to capture molecular patterns of cancer subtypes. The robustness of CSIE is validated through extensive benchmarking against 12 state-of-the-art methods across 66 cancer datasets with over 15 000 patients and 22 diverse data modalities/platforms. Our results demonstrate that CSIE significantly outperforms existing tools, particularly in scenarios with incomplete data. This work shifts the paradigm from requiring exhaustive data collection to leveraging biological intelligence for data completion, offering a scalable solution for high-resolution cancer subtyping in real-world clinical settings. All source code of CSIE and scripts for regenerating results reported in this article are available at https://github.com/tinnlab/CSIE.
Dao Tran, Yen Thi-Hai Pham, Hung N. Luu, Juli Petereit, Manuel A Andrade-Rodriguez, Phi Bya, Tin Chi Nguyen
Briefings Bioinform.7
2024 CCPA: cloud-based, self-learning modules for consensus pathway analysis using GO, KEGG and Reactome
abstract
This manuscript describes the development of a resource module that is part of a learning platform named 'NIGMS Sandbox for Cloud-based Learning' (https://github.com/NIGMS/NIGMS-Sandbox). The module delivers learning materials on Cloud-based Consensus Pathway Analysis in an interactive format that uses appropriate cloud resources for data access and analyses. Pathway analysis is important because it allows us to gain insights into biological mechanisms underlying conditions. But the availability of many pathway analysis methods, the requirement of coding skills, and the focus of current tools on only a few species all make it very difficult for biomedical researchers to self-learn and perform pathway analysis efficiently. Furthermore, there is a lack of tools that allow researchers to compare analysis results obtained from different experiments and different analysis methods to find consensus results. To address these challenges, we have designed a cloud-based, self-learning module that provides consensus results among established, state-of-the-art pathway analysis techniques to provide students and researchers with necessary training and example materials. The training module consists of five Jupyter Notebooks that provide complete tutorials for the following tasks: (i) process expression data, (ii) perform differential analysis, visualize and compare the results obtained from four differential analysis methods (limma, t-test, edgeR, DESeq2), (iii) process three pathway databases (GO, KEGG and Reactome), (iv) perform pathway analysis using eight methods (ORA, CAMERA, KS test, Wilcoxon test, FGSEA, GSA, SAFE and PADOG) and (v) combine results of multiple analyses. We also provide examples, source code, explanations and instructional videos for trainees to complete each Jupyter Notebook. The module supports the analysis for many model (e.g. human, mouse, fruit fly, zebra fish) and non-model species. The module is publicly available at https://github.com/NIGMS/Consensus-Pathway-Analysis-in-the-Cloud. This manuscript describes the development of a resource module that is part of a learning platform named ``NIGMS Sandbox for Cloud-based Learning'' https://github.com/NIGMS/NIGMS-Sandbox. The overall genesis of the Sandbox is described in the editorial NIGMS Sandbox [1] at the beginning of this Supplement. This module delivers learning materials on the analysis of bulk and single-cell ATAC-seq data in an interactive format that uses appropriate cloud resources for data access and analyses.
Van-Dung Pham, Hung Nguyen 0005, Bang Tran, Juli Petereit, Tin Chi Nguyen
Briefings Bioinform.6
2023 A robust and accurate single-cell data trajectory inference method using ensemble pseudotime
abstract
BACKGROUND: The advance in single-cell RNA sequencing technology has enhanced the analysis of cell development by profiling heterogeneous cells in individual cell resolution. In recent years, many trajectory inference methods have been developed. They have focused on using the graph method to infer the trajectory using single-cell data, and then calculate the geodesic distance as the pseudotime. However, these methods are vulnerable to errors caused by the inferred trajectory. Therefore, the calculated pseudotime suffers from such errors. RESULTS: We proposed a novel framework for trajectory inference called the single-cell data Trajectory inference method using Ensemble Pseudotime inference (scTEP). scTEP utilizes multiple clustering results to infer robust pseudotime and then uses the pseudotime to fine-tune the learned trajectory. We evaluated the scTEP using 41 real scRNA-seq data sets, all of which had the ground truth development trajectory. We compared the scTEP with state-of-the-art methods using the aforementioned data sets. Experiments on real linear and non-linear data sets demonstrate that our scTEP performed superior on more data sets than any other method. The scTEP also achieved a higher average and lower variance on most metrics than other state-of-the-art methods. In terms of trajectory inference capacity, the scTEP outperforms those methods. In addition, the scTEP is more robust to the unavoidable errors resulting from clustering and dimension reduction. CONCLUSION: The scTEP demonstrates that utilizing multiple clustering results for the pseudotime inference procedure enhances its robustness. Furthermore, robust pseudotime strengthens the accuracy of trajectory inference, which is the most crucial component in the pipeline. scTEP is available at https://cran.r-project.org/package=scTEP .
Yifan Zhang 0034, Tin Chi Nguyen, Sergiu M. Dascalu, Frederick C. Harris Jr.
BMC Bioinform.3
2022 A comprehensive survey of the approaches for pathway analysis using multi-omics data integration
abstract
Pathway analysis has been widely used to detect pathways and functions associated with complex disease phenotypes. The proliferation of this approach is due to better interpretability of its results and its higher statistical power compared with the gene-level statistics. A plethora of pathway analysis methods that utilize multi-omics setup, rather than just transcriptomics or proteomics, have recently been developed to discover novel pathways and biomarkers. Since multi-omics gives multiple views into the same problem, different approaches are employed in aggregating these views into a comprehensive biological context. As a result, a variety of novel hypotheses regarding disease ideation and treatment targets can be formulated. In this article, we review 32 such pathway analysis methods developed for multi-omics and multi-cohort data. We discuss their availability and implementation, assumptions, supported omics types and databases, pathway analysis techniques and integration strategies. A comprehensive assessment of each method's practicality, and a thorough discussion of the strengths and drawbacks of each technique will be provided. The main objective of this survey is to provide a thorough examination of existing methods to assist potential users and researchers in selecting suitable tools for their data and analysis purposes, while highlighting outstanding challenges in the field that remain to be addressed for future development.
Zeynab Maghsoudi, Alireza Tavakkoli, Tin Chi Nguyen
Briefings Bioinform.4
2022 DrGA: cancer driver gene analysis in a simpler manner
abstract
BACKGROUND: To date, cancer still is one of the leading causes of death worldwide, in which the cumulative of genes carrying mutations was said to be held accountable for the establishment and development of this disease mainly. From that, identification and analysis of driver genes were vital. Our previous study indicated disagreement on a unifying pipeline for these tasks and then introduced a complete one. However, this pipeline gradually manifested its weaknesses as being unfamiliar to non-technical users, time-consuming, and inconvenient. RESULTS: This study presented an R package named DrGA, developed based on our previous pipeline, to tackle the mentioned problems above. It wholly automated four widely used downstream analyses for predicted driver genes and offered additional improvements. We described the usage of the DrGA on driver genes of human breast cancer. Besides, we also gave the users another potential application of DrGA in analyzing genomic biomarkers of a complex disease in another organism. CONCLUSIONS: DrGA facilitated the users with limited IT backgrounds and rapidly created consistent and reproducible results. DrGA and its applications, along with example data, were freely provided at https://github.com/huynguyen250896/DrGA .
Quang-Huy Nguyen 0001, Tin Chi Nguyen, Duc-Hau Le
BMC Bioinform.2
2021 A comprehensive survey of regulatory network inference methods using single cell RNA sequencing data
abstract
Gene regulatory network is a complicated set of interactions between genetic materials, which dictates how cells develop in living organisms and react to their surrounding environment. Robust comprehension of these interactions would help explain how cells function as well as predict their reactions to external factors. This knowledge can benefit both developmental biology and clinical research such as drug development or epidemiology research. Recently, the rapid advance of single-cell sequencing technologies, which pushed the limit of transcriptomic profiling to the individual cell level, opens up an entirely new area for regulatory network research. To exploit this new abundant source of data and take advantage of data in single-cell resolution, a number of computational methods have been proposed to uncover the interactions hidden by the averaging process in standard bulk sequencing. In this article, we review 15 such network inference methods developed for single-cell data. We discuss their underlying assumptions, inference techniques, usability, and pros and cons. In an extensive analysis using simulation, we also assess the methods' performance, sensitivity to dropout and time complexity. The main objective of this survey is to assist not only life scientists in selecting suitable methods for their data and analysis purposes but also computational scientists in developing new methods by highlighting outstanding challenges in the field that remain to be addressed in the future development.
Hung Nguyen 0005, Bang Tran, Bahadir Pehlivan, Tin Chi Nguyen
Briefings Bioinform.5
2020 GSMA: an approach to identify robust global and test Gene Signatures using Meta-Analysis
abstract
MOTIVATION: Recent advances in biomedical research have made massive amount of transcriptomic data available in public repositories from different sources. Due to the heterogeneity present in the individual experiments, identifying reproducible biomarkers for a given disease from multiple independent studies has become a major challenge. The widely used meta-analysis approaches, such as Fisher's method, Stouffer's method, minP and maxP, have at least two major limitations: (i) they are sensitive to outliers, and (ii) they perform only one statistical test for each individual study, and hence do not fully utilize the potential sample size to gain statistical power. RESULTS: Here, we propose a gene-level meta-analysis framework that overcomes these limitations and identifies a gene signature that is reliable and reproducible across multiple independent studies of a given disease. The approach provides a comprehensive global signature that can be used to understand the underlying biological phenomena, and a smaller test signature that can be used to classify future samples of a given disease. We demonstrate the utility of the framework by constructing disease signatures for influenza and Alzheimer's disease using nine datasets including 1108 individuals. These signatures are then validated on 12 independent datasets including 912 individuals. The results indicate that the proposed approach performs better than the majority of the existing meta-analysis approaches in terms of both sensitivity as well as specificity. The proposed signatures could be further used in diagnosis, prognosis and identification of therapeutic targets. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Adib Shafi, Tin Chi Nguyen, Azam Peyvandi-Pour, Sorin Draghici
Bioinform.2
2019 MGKA: A genetic algorithm-based clustering technique for genomic data
abstract
Advances in high-throughput technologies have generated enormous amounts of high-throughput genomic data. Cluster analysis is often the first step to gain insights into genomic data. K-means, the most widely used clustering algorithm, is known to produce sub-optimal clusters depending on the choice of initialized centers. In this paper, we propose a genetic algorithm-based unsupervised clustering method that searches for the optimal centers of clusters based on the concept of k-means. The genetic algorithm reduces k-means sensitivity to randomly initialized centers and reduces the probability of converging to local minima. Two clustering validity indexes are introduced to the selection process to automatically determine the appropriate number of clusters. The proposed algorithm is applied to 16 disease datasets and four single-cell datasets to demonstrate its performance. Results show that our approach outperforms the current state of the art algorithms on a majority of the datasets.
Hung Nguyen 0005, Sushil J. Louis, Tin Chi Nguyen
CEC3
2019 PINSPlus: a tool for tumor subtype discovery in integrated genomic data
abstract
SUMMARY: Since cancer is a heterogeneous disease, tumor subtyping is crucial for improved treatment and prognosis. We have developed a subtype discovery tool, called PINSPlus, that is: (i) robust against noise and unstable quantitative assays, (ii) able to integrate multiple types of omics data in a single analysis and (iii) dramatically superior to established approaches in identifying known subtypes and novel subgroups with significant survival differences. Our validation on 12,158 samples from 44 datasets shows that PINSPlus vastly outperforms other approaches. The software is easy-to-use and can partition hundreds of patients in a few minutes on a personal computer. AVAILABILITY AND IMPLEMENTATION: The package is available at https://cran.r-project.org/package=PINSPlus. Data and R script used in this manuscript are available at https://bioinformatics.cse.unr.edu/software/PINSPlus/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Hung Nguyen 0005, Sangam Shrestha, Sorin Draghici, Tin Chi Nguyen
Bioinform.4
2018 A survey of the approaches for identifying differential methylation using bisulfite sequencing data
abstract
DNA methylation is an important epigenetic mechanism that plays a crucial role in cellular regulatory systems. Recent advancements in sequencing technologies now enable us to generate high-throughput methylation data and to measure methylation up to single-base resolution. This wealth of data does not come without challenges, and one of the key challenges in DNA methylation studies is to identify the significant differences in the methylation levels of the base pairs across distinct biological conditions. Several computational methods have been developed to identify differential methylation using bisulfite sequencing data; however, there is no clear consensus among existing approaches. A comprehensive survey of these approaches would be of great benefit to potential users and researchers to get a complete picture of the available resources. In this article, we present a detailed survey of 22 such approaches focusing on their underlying statistical models, primary features, key advantages and major limitations. Importantly, the intrinsic drawbacks of the approaches pointed out in this survey could potentially be addressed by future research.
Adib Shafi, Cristina Mitrea, Tin Chi Nguyen, Sorin Draghici
Briefings Bioinform.3
2017 DANUBE: Data-Driven Meta-ANalysis Using UnBiased Empirical Distributions - Applied to Biological Pathway Analysis
abstract
Identifying the pathways and mechanisms that are significantly impacted in a given phenotype is challenging. Issues include patient heterogeneity and noise. Many experiments do not have a large enough sample size to achieve the statistical power necessary to identify significantly impacted pathways. Meta-analysis based on combining p-values from individual experiments has been used to improve power. However, all classical meta-analysis approaches work under the assumption that the p-values produced by experiment-level statistical tests follow a uniform distribution under the null hypothesis. Here we show that this assumption does not hold for three mainstream pathway analysis methods, and significant bias is likely to affect many, if not all such meta-analysis studies. We introduce DANUBE, a novel and unbiased approach to combine statistics computed from individual studies. Our framework uses control samples to construct empirical null distributions, from which empirical p-values of individual studies are calculated and combined using either a Central Limit Theorem approach or the additive method. We assess the performance of DANUBE using four different pathway analysis methods. DANUBE is compared with five meta-analysis approaches, as well as with a pathway analysis approach that employs multiple datasets (MetaPath). The 25 approaches have been tested on 16 different datasets related to two human diseases, Alzheimer's disease (7 datasets) and acute myeloid leukemia (9 datasets). We demonstrate that DANUBE overcomes bias in order to consistently identify relevant pathways. We also show how the framework improves results in more general cases, compared to classical meta-analysis performed with common experiment-level statistical tests such as Wilcoxon and t-test.
Tin Chi Nguyen, Cristina Mitrea, Rebecca Tagett, Sorin Draghici
Proc. IEEE1
2016 A novel bi-level meta-analysis approach: applied to biological pathway analysis
abstract
MOTIVATION: The accumulation of high-throughput data in public repositories creates a pressing need for integrative analysis of multiple datasets from independent experiments. However, study heterogeneity, study bias, outliers and the lack of power of available methods present real challenge in integrating genomic data. One practical drawback of many P-value-based meta-analysis methods, including Fisher's, Stouffer's, minP and maxP, is that they are sensitive to outliers. Another drawback is that, because they perform just one statistical test for each individual experiment, they may not fully exploit the potentially large number of samples within each study. RESULTS: We propose a novel bi-level meta-analysis approach that employs the additive method and the Central Limit Theorem within each individual experiment and also across multiple experiments. We prove that the bi-level framework is robust against bias, less sensitive to outliers than other methods, and more sensitive to small changes in signal. For comparative analysis, we demonstrate that the intra-experiment analysis has more power than the equivalent statistical test performed on a single large experiment. For pathway analysis, we compare the proposed framework versus classical meta-analysis approaches (Fisher's, Stouffer's and the additive method) as well as against a dedicated pathway meta-analysis package (MetaPath), using 1252 samples from 21 datasets related to three human diseases, acute myeloid leukemia (9 datasets), type II diabetes (5 datasets) and Alzheimer's disease (7 datasets). Our framework outperforms its competitors to correctly identify pathways relevant to the phenotypes. The framework is sufficiently general to be applied to any type of statistical meta-analysis. AVAILABILITY AND IMPLEMENTATION: The R scripts are available on demand from the authors. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Tin Chi Nguyen, Rebecca Tagett, Michele Donato, Cristina Mitrea, Sorin Draghici
Bioinform.1