Tianwei Yu

dblp:55/3282 · DBLP profile ↗
← Back
33ranked-venue papers
14as first author
12since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 32 · 13 first-author · 11 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 ADM: adaptive graph diffusion for meta-dimension reduction
abstract
Dimension reduction is essential for analyzing high-dimensional data, with various techniques developed to address diverse data characteristics. However, individual methods often struggle to capture all intricate patterns and complex structures simultaneously. To overcome this limitation, we introduce ADM (Adaptive graph Diffusion for Meta-dimension reduction), a novel meta-dimension reduction method grounded in graph diffusion theory. ADM integrates results from multiple dimension reduction techniques, leveraging their individual strengths while mitigating their specific weaknesses.ADM utilizes dynamic Markov processes to transform Euclidean space results into an information space, revealing intrinsic nonlinear manifold structures that are hard to capture by conventional methods. A critical advancement in ADM is its adaptive diffusion mechanism, which dynamically selects optimal diffusion time scales for each sample, enabling effective representation of multi-scale structures. This approach generates robust, high-quality low-dimensional representations that capture both local and global data structures while reducing noise and technique-specific distortions. We demonstrate ADM's efficacy on simulated and real-world datasets, including various omics data types. Results show that ADM provides clearer separation between biological groups and reveals more meaningful patterns compared to existing methods, advancing the analysis and visualization of complex biological data.
Junning Feng 0001, Yong Liang 0001, Tianwei Yu
Briefings Bioinform.3
2025 Dissecting genetic regulation of metabolic coordination
abstract
Understanding genetic regulation of metabolism is critical for gaining insights into the causes of metabolic diseases. Traditional metabolome-based genome-wide association studies (mGWAS) focus on static associations between single nucleotide polymorphisms (SNPs) and metabolite levels, overlooking the changing relationships caused by genotypes within the metabolic network. Notably, some metabolites exhibit changes in correlation patterns with other metabolites under certain physiological conditions while maintaining their overall abundance level. In this manuscript, we develop Metabolic Differential-coordination GWAS (mdGWAS), an innovative framework that detects SNPs associated with the changing correlation patterns between metabolites and metabolic pathways. This approach transcends and complements conventional mean-based analyses by identifying latent regulatory factors that govern the system-level metabolic coordination. Through comprehensive simulation studies, mdGWAS demonstrated robust performance in detecting SNP-metabolite-metabolite associations. Applying mdGWAS to genotyping and mass spectrometry (MS)-based metabolomics data of the METabolic Syndrome In Men (METSIM) Study revealed novel SNPs and genes potentially involved in the regulation of the coordination between metabolic pathways.
Emily C. Hector, Daiwei Zhang, Leqi Tian, Junning Feng 0001, Xianyong Yin, Markku Laakso, Jiashun Xiao, Jian Kang 0003, Tianwei Yu
Briefings Bioinform.11
2025 Nonlinear embedding and integration of omics data: a fast and tuning-free approach
abstract
The rapid progress of single-cell technology has facilitated cost-effective acquisition of diverse omics data, allowing biologists to unravel the complexities of cell populations, disease states, and more. Additionally, single-cell multi-omics technologies have opened new avenues for studying biological interactions. However, the high dimensionality and sparsity of omics data present significant analytical challenges. Dimension reduction (DR) techniques are hence essential for analyzing such complex data, yet many existing methods have inherent limitations. Linear methods like principal component analysis (PCA) struggle to capture intricate associations within data. In response, nonlinear techniques have emerged, but they may face scalability issues, be restricted to single-omics data, or prioritize visualization over generating informative embeddings. Here, we introduce dissimilarity based on conditional ordered list (DCOL) correlation, a novel measure for quantifying nonlinear relationships between variables. Based on this measure, we propose DCOL-PCA and DCOL-Canonical Correlation Analysis for dimension reduction and integration of single- and multi-omics data. In simulations, our methods outperformed nine DR methods and four joint dimension reduction methods, demonstrating stable performance across various settings. We also validated these methods on real datasets, with our method demonstrating its ability to detect intricate signals within and between omics data and generate lower dimensional embeddings that preserve the essential information and latent structures.
Tianwei Yu
Briefings Bioinform.2
2025 A Spatial-Aware Temporal Modeling Network for Imitation Learning-Based Drone Navigation
abstract
Imitation learning-based drone autonomous navigation has attracted significant attention due to the ability of leveraging deep neural networks to learn the control policy from human pilot demonstrations. However, most current studies generate the control command using only a single image, overlooking the semantic information embedded in the sequential input images. While some reinforcement learning-based methods have explored the temporal modeling of sequential input images, they often overlook the spatial relations between frames and vectorize 2D information of each image into a 1D feature. In this paper, we propose a novel imitation learning-based method, termed the spatial-aware temporal modeling network (SATMN), for autonomous drone navigation using sequential images as input. Specifically, we introduce a spatial-temporal-separated modeling mechanism to extract low-resolution spatial features from original images and then perceive spatial-temporal relations among these 2D features. SATMN preserves the spatial information of each 2D image feature during temporal modeling and enables real-time onboard computing on a drone. To validate the effectiveness of the proposed method, we design a compact quadrotor platform capable of autonomous navigation using SATMN, entirely powered by onboard computing devices. Comprehensive and reproducible experiments on public datasets demonstrate the superior performance of our method compared to existing approaches.
Tianwei Yu, Yuanjie Dang, Peng Chen 0008, Ronghua Liang
IEEE Trans. Intell. Transp. Syst.1
2024 Bayesian functional analysis for untargeted metabolomics data with matching uncertainty and small sample sizes
abstract
Untargeted metabolomics based on liquid chromatography-mass spectrometry technology is quickly gaining widespread application, given its ability to depict the global metabolic pattern in biological samples. However, the data are noisy and plagued by the lack of clear identity of data features measured from samples. Multiple potential matchings exist between data features and known metabolites, while the truth can only be one-to-one matches. Some existing methods attempt to reduce the matching uncertainty, but are far from being able to remove the uncertainty for most features. The existence of the uncertainty causes major difficulty in downstream functional analysis. To address these issues, we develop a novel approach for Bayesian Analysis of Untargeted Metabolomics data (BAUM) to integrate previously separate tasks into a single framework, including matching uncertainty inference, metabolite selection and functional analysis. By incorporating the knowledge graph between variables and using relatively simple assumptions, BAUM can analyze datasets with small sample sizes. By allowing different confidence levels of feature-metabolite matching, the method is applicable to datasets in which feature identities are partially known. Simulation studies demonstrate that, compared with other existing methods, BAUM achieves better accuracy in selecting important metabolites that tend to be functionally consistent and assigning confidence scores to feature-metabolite matches. We analyze a COVID-19 metabolomics dataset and a mouse brain metabolomics dataset using BAUM. Even with a very small sample size of 16 mice per group, BAUM is robust and stable. It finds pathways that conform to existing knowledge, as well as novel pathways that are biologically plausible.
Guoxuan Ma, Jian Kang 0003, Tianwei Yu
Briefings Bioinform.3
2024 A robust statistical approach for finding informative spatially associated pathways
abstract
Spatial transcriptomics offers deep insights into cellular functional localization and communication by mapping gene expression to spatial locations. Traditional approaches that focus on selecting spatially variable genes often overlook the complexity of biological pathways and the interactions among genes. Here, we introduce a novel framework that shifts the focus towards directly identifying functional pathways associated with spatial variability by adapting the Brownian distance covariance test in an innovative manner to explore the heterogeneity of biological functions over space. Unlike most other methods, this statistical testing approach is free of gene selection and parameter selection and allows nonlinear and complex dependencies. It allows for a deeper understanding of how cells coordinate their activities across different spatial domains through biological pathways. By analyzing real human and mouse datasets, the method found significant pathways that were associated with spatial variation, as well as different pathway patterns among inner- and edge-cancer regions. This innovative framework offers a new perspective on analyzing spatial transcriptomic data, contributing to our understanding of tissue architecture and disease pathology. The implementation is publicly available at https://github.com/tianlq-prog/STpathway.
Leqi Tian, Jiashun Xiao, Tianwei Yu
Briefings Bioinform.3
2023 Multi-Speed Global Contextual Subspace Matching for Few-Shot Action Recognition
abstract
Few-shot action recognition (FSAR) aims to classify unseen query actions into categories represented by a few labeled support videos. Most current FSAR methods adopt the frame-level matching mechanism that requires continuous actions to be represented by a fixed number of frame features. However, this could compromise the completeness of the contextual video information and make it difficult to handle video features of varying frame sampling speeds. In this paper, we propose a multi-speed global contextual subspace matching (MGCSM) method that generates global contextual action subspace representations from videos containing different numbers of frames to preserve contextual semantic information. Specifically, we propose to obtain the scale-agnostic information of embedding video features using a global contextual aggregation (GCA) module and then generate the discriminative action subspace representation with an action subspace generation (ASG) module. Furthermore, we introduce a multi-speed subspace matching (MSM) mechanism that generates a multi-speed classification score by integrating the similarities between query videos and support subspaces of varying sampling speeds. The proposed method is embedding-agnostic and can be combined with most mainstream embedding networks without model re-designs. Comprehensive and reproducible experiments on standard datasets demonstrate our method's superior performance compared to existing state-of-the-art methods.
Tianwei Yu, Peng Chen 0008, Yuanjie Dang, Ruohong Huan, Ronghua Liang
ACM Multimedia1
2023 An integrated deep learning framework for the interpretation of untargeted metabolomics data
abstract
Untargeted metabolomics is gaining widespread applications. The key aspects of the data analysis include modeling complex activities of the metabolic network, selecting metabolites associated with clinical outcome and finding critical metabolic pathways to reveal biological mechanisms. One of the key roadblocks in data analysis is not well-addressed, which is the problem of matching uncertainty between data features and known metabolites. Given the limitations of the experimental technology, the identities of data features cannot be directly revealed in the data. The predominant approach for mapping features to metabolites is to match the mass-to-charge ratio (m/z) of data features to those derived from theoretical values of known metabolites. The relationship between features and metabolites is not one-to-one since some metabolites share molecular composition, and various adduct ions can be derived from the same metabolite. This matching uncertainty causes unreliable metabolite selection and functional analysis results. Here we introduce an integrated deep learning framework for metabolomics data that take matching uncertainty into consideration. The model is devised with a gradual sparsification neural network based on the known metabolic network and the annotation relationship between features and metabolites. This architecture characterizes metabolomics data and reflects the modular structure of biological system. Three goals can be achieved simultaneously without requiring much complex inference and additional assumptions: (1) evaluate metabolite importance, (2) infer feature-metabolite matching likelihood and (3) select disease sub-networks. When applied to a COVID metabolomics dataset and an aging mouse brain dataset, our method found metabolic sub-networks that were easily interpretable.
Leqi Tian, Tianwei Yu
Briefings Bioinform.2
2022 BraceNet: Graph-Embedded Neural Network For Brain Network Analysis
abstract
Multimodal brain networks extracted from functional magnetic resonance imaging (fMRI) characterize complex connectivities among brain regions from both structural and functional, showing great potential for mental health analysis. Deep neural network models have led a tremendous success in various downstream tasks. However, common property in brain network data, the small number of samples compared with the huge amount of features, hinders the application of deep learning techniques for brain network analysis. This work presents a graph-embedding method by leveraging the unique characteristics of brain networks to unleash the power of deep neural networks and achieve outstanding prediction performance with proper explainability. Experiment results show clear advancements in our proposed braceNet on both real and synthetic datasets.
Xuan Kan, Yunchuan Kong, Tianwei Yu, Ying Guo 0003
IEEE Big Data3
2022 Metapone: a Bioconductor package for joint pathway testing for untargeted metabolomics data
abstract
MOTIVATION: Testing for pathway enrichment is an important aspect in the analysis of untargeted metabolomics data. Due to the unique characteristics of untargeted metabolomics data, some key issues have not been fully addressed in existing pathway testing algorithms: (i) matching uncertainty between data features and metabolites; (ii) lacking of method to analyze positive mode and negative mode liquid chromatography-mass spectrometry (LC/MS) data simultaneously on the same set of subjects; (iii) the incompleteness of pathways in individual software packages. RESULTS: We developed an innovative R/Bioconductor package: metabolic pathway testing with positive and negative mode data (metapone), which can perform two novel statistical tests that take matching uncertainty into consideration-(i) a weighted gene set enrichment analysis-type test and (ii) a permutation-based weighted hypergeometric test. The package is capable of combining positive- and negative-ion mode results in a single testing scheme. For comprehensiveness, the built-in pathways were manually curated from three sources: Kyoto Encyclopedia of Genes and Genomes, Mummichog and The Small Molecule Pathway Database. AVAILABILITY AND IMPLEMENTATION: The package is available at https://bioconductor.org/packages/devel/bioc/html/metapone.html. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Leqi Tian, Guoxuan Ma, Ziyin Tang, Siheng Wang, Jian Kang 0003, Donghai Liang, Tianwei Yu
Bioinform.9
2022 AIME: Autoencoder-based integrative multi-omics data embedding that allows for confounder adjustments
abstract
In the integrative analyses of omics data, it is often of interest to extract data representation from one data type that best reflect its relations with another data type. This task is traditionally fulfilled by linear methods such as canonical correlation analysis (CCA) and partial least squares (PLS). However, information contained in one data type pertaining to the other data type may be complex and in nonlinear form. Deep learning provides a convenient alternative to extract low-dimensional nonlinear data embedding. In addition, the deep learning setup can naturally incorporate the effects of clinical confounding factors into the integrative analysis. Here we report a deep learning setup, named Autoencoder-based Integrative Multi-omics data Embedding (AIME), to extract data representation for omics data integrative analysis. The method can adjust for confounder variables, achieve informative data embedding, rank features in terms of their contributions, and find pairs of features from the two data types that are related to each other through the data embedding. In simulation studies, the method was highly effective in the extraction of major contributing features between data types. Using two real microRNA-gene expression datasets, one with confounder variables and one without, we show that AIME excluded the influence of confounders, and extracted biologically plausible novel information. The R package based on Keras and the TensorFlow backend is available at https://github.com/tianwei-yu/AIME.
Tianwei Yu
PLoS Comput. Biol.1
2021 Accurate feature selection improves single-cell RNA-seq cell clustering
abstract
Cell clustering is one of the most important and commonly performed tasks in single-cell RNA sequencing (scRNA-seq) data analysis. An important step in cell clustering is to select a subset of genes (referred to as 'features'), whose expression patterns will then be used for downstream clustering. A good set of features should include the ones that distinguish different cell types, and the quality of such set could have a significant impact on the clustering accuracy. All existing scRNA-seq clustering tools include a feature selection step relying on some simple unsupervised feature selection methods, mostly based on the statistical moments of gene-wise expression distributions. In this work, we carefully evaluate the impact of feature selection on cell clustering accuracy. In addition, we develop a feature selection algorithm named FEAture SelecTion (FEAST), which provides more representative features. We apply the method on 12 public scRNA-seq datasets and demonstrate that using features selected by FEAST with existing clustering tools significantly improve the clustering accuracy.
Kenong Su, Tianwei Yu, Hao Wu 0003
Briefings Bioinform.2
2020 scBatch: batch-effect correction of RNA-seq data through sample distance matrix adjustment
abstract
MOTIVATION: Batch effect is a frequent challenge in deep sequencing data analysis that can lead to misleading conclusions. Existing methods do not correct batch effects satisfactorily, especially with single-cell RNA sequencing (RNA-seq) data. RESULTS: We present scBatch, a numerical algorithm for batch-effect correction on bulk and single-cell RNA-seq data with emphasis on improving both clustering and gene differential expression analysis. scBatch is not restricted by assumptions on the mechanism of batch-effect generation. As shown in simulations and real data analyses, scBatch outperforms benchmark batch-effect correction methods. AVAILABILITY AND IMPLEMENTATION: The R package is available at github.com/tengfei-emory/scBatch. The code to generate results and figures in this article is available at github.com/tengfei-emory/scBatch-paper-scripts. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Tianwei Yu
Bioinform.2
2020 forgeNet: a graph deep neural network model using tree-based ensemble classifiers for feature graph construction
abstract
MOTIVATION: A unique challenge in predictive model building for omics data has been the small number of samples (n) versus the large amount of features (p). This 'n≪p' property brings difficulties for disease outcome classification using deep learning techniques. Sparse learning by incorporating known functional relationships between the biological units, such as the graph-embedded deep feedforward network (GEDFN) model, has been a solution to this issue. However, such methods require an existing feature graph, and potential mis-specification of the feature graph can be harmful on classification and feature selection. RESULTS: To address this limitation and develop a robust classification model without relying on external knowledge, we propose a forest graph-embedded deep feedforward network (forgeNet) model, to integrate the GEDFN architecture with a forest feature graph extractor, so that the feature graph can be learned in a supervised manner and specifically constructed for a given prediction task. To validate the method's capability, we experimented the forgeNet model with both synthetic and real datasets. The resulting high classification accuracy suggests that the method is a valuable addition to sparse deep learning models for omics data. AVAILABILITY AND IMPLEMENTATION: The method is available at https://github.com/yunchuankong/forgeNet. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yunchuan Kong, Tianwei Yu
Bioinform.2
2019 DNLC: differential network local consistency analysis
abstract
BACKGROUND: The biological network is highly dynamic. Functional relations between genes can be activated or deactivated depending on the biological conditions. On the genome-scale network, subnetworks that gain or lose local expression consistency may shed light on the regulatory mechanisms related to the changing biological conditions, such as disease status or tissue developmental stages. RESULTS: In this study, we develop a new method to select genes and modules on the existing biological network, in which local expression consistency changes significantly between clinical conditions. The method is called DNLC: Differential Network Local Consistency. In simulations, our algorithm detected artificially created local consistency changes effectively. We applied the method on two publicly available datasets, and the method detected novel genes and network modules that were biologically plausible. CONCLUSIONS: The new method is effective in finding modules in which the gene expression consistency change between clinical conditions. It is a useful tool that complements traditional differential expression analyses to make discoveries from gene expression data. The R package is available at https://cran.r-project.org/web/packages/DNLC.
Yusheng Ding, Qingyang Xiao, Linqing Liu, Qingpo Cai, Yunchuan Kong, Tianwei Yu
BMC Bioinform.9
2018 Mitigating the adverse impact of batch effects in sample pattern detection
abstract
Motivation: It is well known that batch effects exist in RNA-seq data and other profiling data. Although some methods do a good job adjusting for batch effects by modifying the data matrices, it is still difficult to remove the batch effects entirely. The remaining batch effect can cause artifacts in the detection of patterns in the data. Results: In this study, we consider the batch effect issue in the pattern detection among the samples, such as clustering, dimension reduction and construction of networks between subjects. Instead of adjusting the original data matrices, we design an adaptive method to directly adjust the dissimilarity matrix between samples. In simulation studies, the method achieved better results recovering true underlying clusters, compared to the leading batch effect adjustment method ComBat. In real data analysis, the method effectively corrected distance matrices and improved the performance of clustering algorithms. Availability and implementation: The R package is available at: https://github.com/tengfei-emory/QuantNorm. Supplementary information: Supplementary data are available at Bioinformatics online.
Tengjiao Zhang, Weiyang Shi, Tianwei Yu
Bioinform.4
2018 Missing value imputation for LC-MS metabolomics data by incorporating metabolic network and adduct ion relations
abstract
Motivation: Metabolomics data generated from liquid chromatography-mass spectrometry platforms often contain missing values. Existing imputation methods do not consider underlying feature relations and the metabolic network information. As a result, the imputation results may not be optimal. Results: We proposed an imputation algorithm that incorporates the existing metabolic network, adduct ion relations even for unknown compounds, as well as linear and nonlinear associations between feature intensities to build a feature-level network. The algorithm uses support vector regression for missing value imputation based on features in the neighborhood on the network. We compared our proposed method with methods being widely used. As judged by the normalized root mean squared error in real data-based simulations, our proposed methods can achieve better accuracy. Availability and implementation: The R package is available at http://web1.sph.emory.edu/users/tyu8/MINMA. Contact: [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Zhuxuan Jin, Jian Kang 0003, Tianwei Yu
Bioinform.3
2018 A graph-embedded deep feedforward network for disease outcome classification and feature selection using gene expression data
abstract
Motivation: Gene expression data represents a unique challenge in predictive model building, because of the small number of samples (n) compared with the huge amount of features (p). This 'n≪p' property has hampered application of deep learning techniques for disease outcome classification. Sparse learning by incorporating external gene network information could be a potential solution to this issue. Still, the problem is very challenging because (i) there are tens of thousands of features and only hundreds of training samples, (ii) the scale-free structure of the gene network is unfriendly to the setup of convolutional neural networks. Results: To address these issues and build a robust classification model, we propose the Graph-Embedded Deep Feedforward Networks (GEDFN), to integrate external relational information of features into the deep neural network architecture. The method is able to achieve sparse connection between network layers to prevent overfitting. To validate the method's capability, we conducted both simulation experiments and real data analysis using a breast invasive carcinoma RNA-seq dataset and a kidney renal clear cell carcinoma RNA-seq dataset from The Cancer Genome Atlas. The resulting high classification accuracy and easily interpretable feature selection results suggest the method is a useful addition to the current graph-guided classification models and feature selection procedures. Availability and implementation: The method is available at https://github.com/yunchuankong/GEDFN. Supplementary information: Supplementary data are available at Bioinformatics online.
Yunchuan Kong, Tianwei Yu
Bioinform.2
2018 A new dynamic correlation algorithm reveals novel functional aspects in single cell and bulk RNA-seq data
abstract
Dynamic correlations are pervasive in high-throughput data. Large numbers of gene pairs can change their correlation patterns in response to observed/unobserved changes in physiological states. Finding changes in correlation patterns can reveal important regulatory mechanisms. Currently there is no method that can effectively detect global dynamic correlation patterns in a dataset. Given the challenging nature of the problem, the currently available methods use genes as surrogate measurements of physiological states, which cannot faithfully represent true underlying biological signals. In this study we develop a new method that directly identifies strong latent dynamic correlation signals from the data matrix, named DCA: Dynamic Correlation Analysis. At the center of the method is a new metric for the identification of pairs of variables that are highly likely to be dynamically correlated, without knowing the underlying physiological states that govern the dynamic correlation. We validate the performance of the method with extensive simulations. We applied the method to three real datasets: a single cell RNA-seq dataset, a bulk RNA-seq dataset, and a microarray gene expression dataset. In all three datasets, the method reveals novel latent factors with clear biological meaning, bringing new insights into the data.
Tianwei Yu
PLoS Comput. Biol.1
2017 Detecting subnetwork-level dynamic correlations
abstract
MOTIVATION: The biological regulatory system is highly dynamic. The correlations between many functionally related genes change over different biological conditions. Finding dynamic relations on the existing biological network may reveal important regulatory mechanisms. Currently no method is available to detect subnetwork-level dynamic correlations systematically on the genome-scale network. Two major issues hampered the development. The first is gene expression profiling data usually do not contain time course measurements to facilitate the analysis of dynamic relations, which can be partially addressed by using certain genes as indicators of biological conditions. Secondly, it is unclear how to effectively delineate subnetworks, and define dynamic relations between them. RESULTS: Here we propose a new method named LANDD (Liquid Association for Network Dynamics Detection) to find subnetworks that show substantial dynamic correlations, as defined by subnetwork A is concentrated with Liquid Association scouting genes for subnetwork B. The method produces easily interpretable results because of its focus on subnetworks that tend to comprise functionally related genes. Also, the collective behaviour of genes in a subnetwork is a much more reliable indicator of underlying biological conditions compared to using single genes as indicators. We conducted extensive simulations to validate the method's ability to detect subnetwork-level dynamic correlations. Using a real gene expression dataset and the human protein-protein interaction network, we demonstrate the method links subnetworks of distinct biological processes, with both confirmed relations and plausible new functional implications. We also found signal transduction pathways tend to show extensive dynamic relations with other functional groups. AVAILABILITY AND IMPLEMENTATION: The R package is available at https://cran.r-project.org/web/packages/LANDD CONTACTS: [email protected], [email protected] or [email protected] information: Supplementary data are available at Bioinformatics online.
Shangzhao Qiu, Zhuxuan Jin, Sihong Gong, Tianwei Yu
Bioinform.7
2016 Bayesian network feature finder (BANFF): an R package for gene network feature selection
abstract
MOTIVATION: Network marker selection on genome-scale networks plays an important role in the understanding of biological mechanisms and disease pathologies. Recently, a Bayesian nonparametric mixture model has been developed and successfully applied for selecting genes and gene sub-networks. Hence, extending this method to a unified approach for network-based feature selection on general large-scale networks and creating an easy-to-use software package is on demand. RESULTS: We extended the method and developed an R package, the Bayesian network feature finder (BANFF), providing a package of posterior inference, model comparison and graphical illustration of model fitting. The model was extended to a more general form, and a parallel computing algorithm for the Markov chain Monte Carlo -based posterior inference and an expectation maximization-based algorithm for posterior approximation were added. Based on simulation studies, we demonstrate the use of BANFF on analyzing gene expression on a protein-protein interaction network. AVAILABILITY: https://cran.r-project.org/web/packages/BANFF/index.html CONTACT: [email protected], [email protected] information: Supplementary data are available at Bioinformatics online.
Zhou Lan, Yize Zhao, Jian Kang 0003, Tianwei Yu
Bioinform.4
2014 Improving peak detection in high-resolution LC/MS metabolomics data using preexisting knowledge and machine learning approach
abstract
MOTIVATION: Peak detection is a key step in the preprocessing of untargeted metabolomics data generated from high-resolution liquid chromatography-mass spectrometry (LC/MS). The common practice is to use filters with predetermined parameters to select peaks in the LC/MS profile. This rigid approach can cause suboptimal performance when the choice of peak model and parameters do not suit the data characteristics. RESULTS: Here we present a method that learns directly from various data features of the extracted ion chromatograms (EICs) to differentiate between true peak regions from noise regions in the LC/MS profile. It utilizes the knowledge of known metabolites, as well as robust machine learning approaches. Unlike currently available methods, this new approach does not assume a parametric peak shape model and allows maximum flexibility. We demonstrate the superiority of the new approach using real data. Because matching to known metabolites entails uncertainties and cannot be considered a gold standard, we also developed a probabilistic receiver-operating characteristic (pROC) approach that can incorporate uncertainties. AVAILABILITY AND IMPLEMENTATION: The new peak detection approach is implemented as part of the apLCMS package available at http://web1.sph.emory.edu/apLCMS/ CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Tianwei Yu, Dean P. Jones
Bioinform.1
2014 Network-based modular latent structure analysis
abstract
BACKGROUND: High-throughput expression data, such as gene expression and metabolomics data, exhibit modular structures. Groups of features in each module follow a latent factor model, while between modules, the latent factors are quasi-independent. Recovering the latent factors can shed light on the hidden regulation patterns of the expression. The difficulty in detecting such modules and recovering the latent factors lies in the high dimensionality of the data, and the lack of knowledge in module membership. METHODS: Here we describe a method based on community detection in the co-expression network. It consists of inference-based network construction, module detection, and interacting latent factor detection from modules. RESULTS: In simulations, the method outperformed projection-based modular latent factor discovery when the input signals were not Gaussian. We also demonstrate the method's value in real data analysis. CONCLUSIONS: The new method nMLSA (network-based modular latent structure analysis) is effective in detecting latent structures, and is easy to extend to non-linear cases. The method is available as R code at http://web1.sph.emory.edu/users/tyu8/nMLSA/.
Tianwei Yu
BMC Bioinform.1
2013 xMSanalyzer: automated pipeline for improved feature detection and downstream analysis of large-scale, non-targeted metabolomics data
abstract
BACKGROUND: Detection of low abundance metabolites is important for de novo mapping of metabolic pathways related to diet, microbiome or environmental exposures. Multiple algorithms are available to extract m/z features from liquid chromatography-mass spectral data in a conservative manner, which tends to preclude detection of low abundance chemicals and chemicals found in small subsets of samples. The present study provides software to enhance such algorithms for feature detection, quality assessment, and annotation. RESULTS: xMSanalyzer is a set of utilities for automated processing of metabolomics data. The utilites can be classified into four main modules to: 1) improve feature detection for replicate analyses by systematic re-extraction with multiple parameter settings and data merger to optimize the balance between sensitivity and reliability, 2) evaluate sample quality and feature consistency, 3) detect feature overlap between datasets, and 4) characterize high-resolution m/z matches to small molecule metabolites and biological pathways using multiple chemical databases. The package was tested with plasma samples and shown to more than double the number of features extracted while improving quantitative reliability of detection. MS/MS analysis of a random subset of peaks that were exclusively detected using xMSanalyzer confirmed that the optimization scheme improves detection of real metabolites. CONCLUSIONS: xMSanalyzer is a package of utilities for data extraction, quality control assessment, detection of overlapping and unique metabolites in multiple datasets, and batch annotation of metabolites. The program was designed to integrate with existing packages such as apLCMS and XCMS, but the framework can also be used to enhance data extraction for other LC/MS data software.
Karan Uppal, Quinlyn A. Soltow, Frederick H. Strobel, William Stephen Pittard, Kim M. Gernert, Tianwei Yu, Dean P. Jones
BMC Bioinform.6
2013 Hierarchical Clustering of High- Throughput Expression Data Based on General Dependences
abstract
High-throughput expression technologies, including gene expression array and liquid chromatography--mass spectrometry (LC-MS) and so on, measure thousands of features, i.e., genes or metabolites, on a continuous scale. In such data, both linear and nonlinear relations exist between features. Nonlinear relations can reflect critical regulation patterns in the biological system. However, they are not identified and utilized by traditional clustering methods based on linear associations. Clustering based on general dependences, i.e., both linear and nonlinear relations, is hampered by the high dimensionality and high noise level of the data. We developed a sensitive nonparametric measure of general dependence between (groups of) random variables in high dimensions. Based on this dependence measure, we developed a hierarchical clustering method. In simulation studies, the method outperformed correlation- and mutual information (MI)-based hierarchical clustering methods in clustering features with nonlinear dependences. We applied the method to a microarray data set measuring the gene expression in cell-cycle time series to show it generates biologically relevant results. The R code is available at http://userwww.service.emory.edu/~tyu8/GDHC.
Tianwei Yu, Hesen Peng
IEEE ACM Trans. Comput. Biol. Bioinform.1
2011 Incorporating Nonlinear Relationships in Microarray Missing Value Imputation
abstract
Microarray gene expression data often contain missing values. Accurate estimation of the missing values is important for downstream data analyses that require complete data. Nonlinear relationships between gene expression levels have not been well-utilized in missing value imputation. We propose an imputation scheme based on nonlinear dependencies between genes. By simulations based on real microarray data, we show that incorporating nonlinear relationships could improve the accuracy of missing value imputation, both in terms of normalized root-mean-squared error and in terms of the preservation of the list of significant genes in statistical testing. In addition, we studied the impact of artificial dependencies introduced by data normalization on the simulation results. Our results suggest that methods relying on global correlation structures may yield overly optimistic simulation results when the data have been subjected to row (gene)-wise mean removal.
Tianwei Yu, Hesen Peng, Wei Sun 0006
IEEE ACM Trans. Comput. Biol. Bioinform.1
2010 An exploratory data analysis method to reveal modular latent structures in high-throughput data
abstract
BACKGROUND: Modular structures are ubiquitous across various types of biological networks. The study of network modularity can help reveal regulatory mechanisms in systems biology, evolutionary biology and developmental biology. Identifying putative modular latent structures from high-throughput data using exploratory analysis can help better interpret the data and generate new hypotheses. Unsupervised learning methods designed for global dimension reduction or clustering fall short of identifying modules with factors acting in linear combinations. RESULTS: We present an exploratory data analysis method named MLSA (Modular Latent Structure Analysis) to estimate modular latent structures, which can find co-regulative modules that involve non-coexpressive genes. CONCLUSIONS: Through simulations and real-data analyses, we show that the method can recover modular latent structures effectively. In addition, the method also performed very well on data generated from sparse global latent factor models. The R code is available at http://userwww.service.emory.edu/~tyu8/MLSA/.
Tianwei Yu
BMC Bioinform.1
2010 Quantification and deconvolution of asymmetric LC-MS peaks using the bi-Gaussian mixture model and statistical model selection
abstract
BACKGROUND: Liquid chromatography-mass spectrometry (LC-MS) is one of the major techniques for the quantification of metabolites in complex biological samples. Peak modeling is one of the key components in LC-MS data pre-processing. RESULTS: To quantify asymmetric peaks with high noise level, we developed an estimation procedure using the bi-Gaussian function. In addition, to accurately quantify partially overlapping peaks, we developed a deconvolution method using the bi-Gaussian mixture model combined with statistical model selection. CONCLUSIONS: Using extensive simulations and real data, we demonstrated the advantage of the bi-Gaussian mixture model over the Gaussian mixture model and the method of kernel smoothing combined with signal summation in peak quantification and deconvolution. The method is implemented in the R package apLCMS: http://www.sph.emory.edu/apLCMS/.
Tianwei Yu, Hesen Peng
BMC Bioinform.1
2009 apLCMS - adaptive processing of high-resolution LC/MS data
abstract
MOTIVATION: Liquid chromatography-mass spectrometry (LC/MS) profiling is a promising approach for the quantification of metabolites from complex biological samples. Significant challenges exist in the analysis of LC/MS data, including noise reduction, feature identification/ quantification, feature alignment and computation efficiency. RESULT: Here we present a set of algorithms for the processing of high-resolution LC/MS data. The major technical improvements include the adaptive tolerance level searching rather than hard cutoff or binning, the use of non-parametric methods to fine-tune intensity grouping, the use of run filter to better preserve weak signals and the model-based estimation of peak intensities for absolute quantification. The algorithms are implemented in an R package apLCMS, which can efficiently process large LC/ MS datasets. AVAILABILITY: The R package apLCMS is available at www.sph.emory.edu/apLCMS. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Tianwei Yu, Youngja Park, Jennifer M. Johnson, Dean P. Jones
Bioinform.1
2007 Detection of eQTL modules mediated by activity levels of transcription factors
abstract
MOTIVATION: Studies of gene expression quantitative trait loci (eQTL) in different organisms have shown the existence of eQTL hot spots: each being a small segment of DNA sequence that harbors the eQTL of a large number of genes. Two questions of great interest about eQTL hot spots arise: (1) which gene within the hot spot is responsible for the linkages, i.e. which gene is the quantitative trait gene (QTG)? (2) How does a QTG affect the expression levels of many genes linked to it? Answers to the first question can be offered by available biological evidence or by statistical methods. The second question is harder to address. One simple situation is that the QTG encodes a transcription factor (TF), which regulates the expression of genes linked to it. However, previous results have shown that TFs are not overrepresented in the eQTL hot spots. In this article, we consider the scenario that the propagation of genetic perturbation from a QTG to other linked genes is mediated by the TF activity. We develop a procedure to detect the eQTL modules (eQTL hot spots together with linked genes) that are compatible with this scenario. RESULTS: We first detect 27 eQTL modules from a yeast eQTL data, and estimate TF activity profiles using the method of Yu and Li (2005). Then likelihood ratio tests (LRTs) are conducted to find 760 relationships supporting the scenario of TF activity mediation: (DNA polymorphism --> cis-linked gene --> TF activity --> downstream linked gene). They are organized into 4 eQTL modules: an amino acid synthesis module featuring a cis-linked gene LEU2 and the mediating TF Leu3; a pheromone response module featuring a cis-linked gene GPA1 and the mediating TF Ste12; an energy-source control module featuring two cis-linked genes, GSY2 and HAP1, and the mediating TF Hap1; a mitotic exit module featuring four cis-linked genes, AMN1, CSH1, DEM1 and TOS1, and the mediating TF complex Ace2/Swi5. Gene Ontology is utilized to reveal interesting functional groups of the downstream genes in each module. AVAILABILITY: Our methods are implemented in an R package: eqtl.TF, which includes source codes and relevant data. It can be freely downloaded at http://www.stat.ucla.edu/~sunwei/software.htm. SUPPLEMENTARY INFORMATION: http://www.stat.ucla.edu/~sunwei/yeast_eQTL_TF/supplementary.pdf.
Wei Sun 0006, Tianwei Yu, Ker-Chau Li
Bioinform.2
2007 A forward-backward fragment assembling algorithm for the identification of genomic amplification and deletion breakpoints using high-density single nucleotide polymorphism (SNP) array
abstract
BACKGROUND: DNA copy number aberration (CNA) is one of the key characteristics of cancer cells. Recent studies demonstrated the feasibility of utilizing high density single nucleotide polymorphism (SNP) genotyping arrays to detect CNA. Compared with the two-color array-based comparative genomic hybridization (array-CGH), the SNP arrays offer much higher probe density and lower signal-to-noise ratio at the single SNP level. To accurately identify small segments of CNA from SNP array data, segmentation methods that are sensitive to CNA while resistant to noise are required. RESULTS: We have developed a highly sensitive algorithm for the edge detection of copy number data which is especially suitable for the SNP array-based copy number data. The method consists of an over-sensitive edge-detection step and a test-based forward-backward edge selection step. CONCLUSION: Using simulations constructed from real experimental data, the method shows high sensitivity and specificity in detecting small copy number changes in focused regions. The method is implemented in an R package FASeg, which includes data processing and visualization utilities, as well as libraries for processing Affymetrix SNP array data.
Tianwei Yu, Wei Sun 0006, Ker-Chau Li, Zugen Chen, Sharoni Jacobs, Dione K. Bailey, David T. Wong
BMC Bioinform.1
2005 Inference of transcriptional regulatory network by two-stage constrained space factor analysis
abstract
MOTIVATION: Microarray gene expression and cross-linking chromatin immunoprecipitation data contain voluminous information that can help the identification of transcriptional regulatory networks at the full genome scale. Such high-throughput data are noisy however. In contrast, from the biomedical literature, we can find many evidenced transcription factor (TF)-target gene binding relationships that have been elucidated at the molecular level. But such sporadically generated knowledge only offers glimpses on limited patches of the network. How to incorporate this valuable knowledge resource to build more reliable network models remains a question. RESULTS: We present a modified factor analysis approach. Our algorithm starts with the evidenced TF-gene linkages. It iterates between the network configuration estimation step and the connection strength estimation step, using the high-throughput data, till convergence. We report two comprehensive regulatory networks obtained for Saccharomyces cerevisiae, one under the normal growth condition and the other under the environmental stress condition. SUPPLEMENTARY INFORMATION: http://kiefer.stat.ucla.edu/lap2/download/bti656_supplement.pdf.
Tianwei Yu, Ker-Chau Li
Bioinform.1
2005 Study of coordinative gene expression at the biological process level
abstract
MOTIVATION: Cellular processes are not isolated groups of events. Nevertheless, in most microarray analyses, they tend to be treated as standalone units. To shed light on how various parts of the interlocked biological processes are coordinated at the transcription level, there is a need to study the between-unit expressional relationship directly. RESULTS: We approach this issue by constructing an index of correlation function to convey the global pattern of coexpression between genes from one process and genes from the entire genome. Processes with similar signatures are then identified and projected to a process-to-process association graph. This top-down method allows for detailed gene-level analysis between linked processes to follow up. Using the cell-cycle gene-expression profiles for Saccharomyces cerevisiae, we report well-organized networks of biological processes that would be difficult to find otherwise. Using another dataset, we report a sharply different network structure featuring cellular responses under environmental stress. SUPPLEMENTARY INFORMATION: http://kiefer.stat.ucla.edu/lap2/download/KL_supplement.pdf.
Tianwei Yu, Wei Sun 0006, Shinsheng Yuan, Ker-Chau Li
Bioinform.1