Veerabhadran Baladandayuthapani

dblp:76/9023 · DBLP profile ↗
← Back
15ranked-venue papers
0as first author
3since 2021 · last 2023
0000-0001-9107-3157ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 9 · 2 since 2021Artificial intelligence and machine learning · 4 · 1 since 2021Databases, data management, data science and information retrieval · 4

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Probabilistic and Bayesian machine learning · 100%
Interdisciplinary, comprehensive, and emerging computing
6 papers
Bioinformatics and computational biology · 100%
Databases, data mining, and information retrieval
2 papers
Distributed and cloud data management · 34% Machine learning and data management · 34% Database system architecture and tuning · 22%

Topics — the 19 heaviest of 23, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
bayesian inference
1.022022
Bayesian Covariate-Dependent Gaussian Graphical Models with Varying Structure · J. Mach. Learn. Res. 2022
Quantile Graphical Models: a Bayesian Approach · J. Mach. Learn. Res. 2020
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models
gaussian graphical model
1.022022
Bayesian Covariate-Dependent Gaussian Graphical Models with Varying Structure · J. Mach. Learn. Res. 2022
Quantile Graphical Models: a Bayesian Approach · J. Mach. Learn. Res. 2020
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models
bayesian graphical model
0.612022
Bayesian Covariate-Dependent Gaussian Graphical Models with Varying Structure · J. Mach. Learn. Res. 2022
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods
markov chain monte carlo
0.612022
Bayesian Covariate-Dependent Gaussian Graphical Models with Varying Structure · J. Mach. Learn. Res. 2022
Bioinformatics and computational biology › network bioinformatics › biological network analysis
gene co-expression network
0.612022
SpaceX: gene co-expression network estimation for spatial transcriptomics · Bioinform. 2022
Bioinformatics and computational biology › transcriptomics
spatial transcriptomics
0.612022
SpaceX: gene co-expression network estimation for spatial transcriptomics · Bioinform. 2022
Bioinformatics and computational biology › network bioinformatics › biological network analysis
differential network analysis
0.522018
iDINGO - integrative differential network analysis in genomics with Shiny application · Bioinform. 2018
DINGO: differential network analysis in genomics · Bioinform. 2015
Bioinformatics and computational biology
genomics
0.522018
iDINGO - integrative differential network analysis in genomics with Shiny application · Bioinform. 2018
iBAG: integrative Bayesian analysis of high-dimensional multiplatform genomics data · Bioinform. 2013
Machine learning › Probabilistic and Bayesian machine learning › structured models
graphical models
0.412020
Quantile Graphical Models: a Bayesian Approach · J. Mach. Learn. Res. 2020
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference
0.412020
Quantile Graphical Models: a Bayesian Approach · J. Mach. Learn. Res. 2020
Bioinformatics and computational biology › network bioinformatics › biological network analysis
biological network inference
0.412020
NExUS: Bayesian simultaneous network estimation across unequal sample sizes · Bioinform. 2020
Bioinformatics and computational biology
cancer genomics
0.222020
NExUS: Bayesian simultaneous network estimation across unequal sample sizes · Bioinform. 2020
DINGO: differential network analysis in genomics · Bioinform. 2015
Bioinformatics and computational biology
biomarker discovery
0.212013
iBAG: integrative Bayesian analysis of high-dimensional multiplatform genomics data · Bioinform. 2013
Bioinformatics and computational biology › data integration
cross-platform data integration
0.212013
iBAG: integrative Bayesian analysis of high-dimensional multiplatform genomics data · Bioinform. 2013
Bioinformatics and computational biology
gene expression analysis
0.112011
Bayesian ensemble methods for survival prediction in gene expression data · Bioinform. 2011
Bioinformatics and computational biology › survival analysis
survival prediction
0.112011
Bayesian ensemble methods for survival prediction in gene expression data · Bioinform. 2011
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › bayesian model selection
bayesian variable selection
0.112010
On the Computation of Stochastic Search Variable Selection in Linear Regression with UDFs · ICDM 2010
Database system architecture and tuning › analytical database system
in-database analytics
0.112010
On the Computation of Stochastic Search Variable Selection in Linear Regression with UDFs · ICDM 2010
Bioinformatics and computational biology
multi-omics data integration
0.112018
iDINGO - integrative differential network analysis in genomics with Shiny application · Bioinform. 2018

Methods — techniques the papers use, named apart from their topics

spatial poisson model · 0.6mixture prior · 0.6markov chain monte carlo · 0.6bayesian factor model · 0.6variational bayes · 0.4quantile regression · 0.4bayesian inference · 0.4bayesian hierarchical modeling · 0.4sufficient statistics · 0.4shiny · 0.3group-specific dependency estimation · 0.3user-defined functions · 0.2stochastic search variable selection · 0.2pathway-based network inference · 0.2joint graphical model estimation · 0.2sampling · 0.2job scheduling · 0.2hierarchical bayesian modeling · 0.2
YearPublicationVenuePosition
2023 Tumor radiogenomics in gliomas with Bayesian layered variable selection
abstract
We propose a statistical framework to analyze radiological magnetic resonance imaging (MRI) and genomic data to identify the underlying radiogenomic associations in lower grade gliomas (LGG). We devise a novel imaging phenotype by dividing the tumor region into concentric spherical layers that mimics the tumor evolution process. MRI data within each layer is represented by voxel-intensity-based probability density functions which capture the complete information about tumor heterogeneity. Under a Riemannian-geometric framework these densities are mapped to a vector of principal component scores which act as imaging phenotypes. Subsequently, we build Bayesian variable selection models for each layer with the imaging phenotypes as the response and the genomic markers as predictors. Our novel hierarchical prior formulation incorporates the interior-to-exterior structure of the layers, and the correlation between the genomic markers. We employ a computationally-efficient Expectation-Maximization-based strategy for estimation. Simulation studies demonstrate the superior performance of our approach compared to other approaches. With a focus on the cancer driver genes in LGG, we discuss some biologically relevant findings. Genes implicated with survival and oncogenesis are identified as being associated with the spherical layers, which could potentially serve as early-stage diagnostic markers for disease monitoring, prior to routine invasive approaches. We provide a R package that can be used to deploy our framework to identify radiogenomic associations.
Shariq Mohammed, Sebastian Kurtek, Karthik Bharath, Arvind Rao, Veerabhadran Baladandayuthapani
Medical Image Anal.5
2022 SpaceX: gene co-expression network estimation for spatial transcriptomics
abstract
MOTIVATION: The analysis of spatially resolved transcriptome enables the understanding of the spatial interactions between the cellular environment and transcriptional regulation. In particular, the characterization of the gene-gene co-expression at distinct spatial locations or cell types in the tissue enables delineation of spatial co-regulatory patterns as opposed to standard differential single gene analyses. To enhance the ability and potential of spatial transcriptomics technologies to drive biological discovery, we develop a statistical framework to detect gene co-expression patterns in a spatially structured tissue consisting of different clusters in the form of cell classes or tissue domains. RESULTS: We develop SpaceX (spatially dependent gene co-expression network), a Bayesian methodology to identify both shared and cluster-specific co-expression network across genes. SpaceX uses an over-dispersed spatial Poisson model coupled with a high-dimensional factor model which is based on a dimension reduction technique for computational efficiency. We show via simulations, accuracy gains in co-expression network estimation and structure by accounting for (increasing) spatial correlation and appropriate noise distributions. In-depth analysis of two spatial transcriptomics datasets in mouse hypothalamus and human breast cancer using SpaceX, detected multiple hub genes which are related to cognitive abilities for the hypothalamus data and multiple cancer genes (e.g. collagen family) from the tumor region for the breast cancer data. AVAILABILITY AND IMPLEMENTATION: The SpaceX R-package is available at github.com/bayesrx/SpaceX. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Satwik Acharyya, Veerabhadran Baladandayuthapani
Bioinform.3
2022 Bayesian Covariate-Dependent Gaussian Graphical Models with Varying Structure
abstract
We introduce Bayesian Gaussian graphical models with covariates (GGMx), a class of multivariate Gaussian distributions with covariate-dependent sparse precision matrix. We propose a general construction of a functional mapping from the covariate space to the cone of sparse positive definite matrices, which encompasses many existing graphical models for heterogeneous settings. Our methodology is based on a novel mixture prior for precision matrices with a non-local component that admits attractive theoretical and empirical properties. The flexible formulation of GGMx allows both the strength and the sparsity pattern of the precision matrix (hence the graph structure) change with the covariates. Posterior inference is carried out with a carefully designed Markov chain Monte Carlo algorithm, which ensures the positive definiteness of sparse precision matrices at any given covariates' values. Extensive simulations and a case study in cancer genomics demonstrate the utility of the proposed model.
Francesco C. Stingo, Veerabhadran Baladandayuthapani
J. Mach. Learn. Res.3
2020 NExUS: Bayesian simultaneous network estimation across unequal sample sizes
abstract
MOTIVATION: Network-based analyses of high-throughput genomics data provide a holistic, systems-level understanding of various biological mechanisms for a common population. However, when estimating multiple networks across heterogeneous sub-populations, varying sample sizes pose a challenge in the estimation and inference, as network differences may be driven by differences in power. We are particularly interested in addressing this challenge in the context of proteomic networks for related cancers, as the number of subjects available for rare cancer (sub-)types is often limited. RESULTS: We develop NExUS (Network Estimation across Unequal Sample sizes), a Bayesian method that enables joint learning of multiple networks while avoiding artefactual relationship between sample size and network sparsity. We demonstrate through simulations that NExUS outperforms existing network estimation methods in this context, and apply it to learn network similarity and shared pathway activity for groups of cancers with related origins represented in The Cancer Genome Atlas (TCGA) proteomic data. AVAILABILITY AND IMPLEMENTATION: The NExUS source code is freely available for download at https://github.com/priyamdas2/NExUS. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Priyam Das, Christine B. Peterson, Kim-Anh Do, Rehan Akbani, Veerabhadran Baladandayuthapani
Bioinform.5
2020 Quantile Graphical Models: a Bayesian Approach
abstract
Graphical models are ubiquitous tools to describe the interdependence between variables measured simultaneously such as large-scale gene or protein expression data. Gaussian graphical models (GGMs) are well-established tools for probabilistic exploration of dependence structures using precision matrices and they are generated under a multivariate normal joint distribution. However, they suffer from several shortcomings since they are based on Gaussian distribution assumptions. In this article, we propose a Bayesian quantile based approach for sparse estimation of graphs. We demonstrate that the resulting graph estimation is robust to outliers and applicable under general distributional assumptions. Furthermore, we develop efficient variational Bayes approximations to scale the methods for large data sets. Our methods are applied to a novel cancer proteomics data dataset where-in multiple proteomic antibodies are simultaneously assessed on tumor samples using reverse-phase protein arrays (RPPA) technology.
Nilabja Guha, Veerabhadran Baladandayuthapani, Bani K. Mallick
J. Mach. Learn. Res.2
2018 iDINGO - integrative differential network analysis in genomics with Shiny application
abstract
Motivation: Differential network analysis is an important way to understand network rewiring involved in disease progression and development. Building differential networks from multiple 'omics data provides insight into the holistic differences of the interactive system under different patient-specific groups. DINGO was developed to infer group-specific dependencies and build differential networks. However, DINGO and other existing tools are limited to analyze data arising from a single platform, and modeling each of the multiple 'omics data independently does not account for the hierarchical structure of the data. Results: We developed the iDINGO R package to estimate group-specific dependencies and make inferences on the integrative differential networks, considering the biological hierarchy among the platforms. A Shiny application has also been developed to facilitate easier analysis and visualization of results, including integrative differential networks and hub gene identification across platforms. Availability and implementation: R package is available on CRAN (https://cran.r-project.org/web/packages/iDINGO) and Shiny application at https://github.com/MinJinHa/iDINGO. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Caleb A. Class, Min Jin Ha, Veerabhadran Baladandayuthapani, Kim-Anh Do
Bioinform.3
2015 DINGO: differential network analysis in genomics
abstract
MOTIVATION: Cancer progression and development are initiated by aberrations in various molecular networks through coordinated changes across multiple genes and pathways. It is important to understand how these networks change under different stress conditions and/or patient-specific groups to infer differential patterns of activation and inhibition. Existing methods are limited to correlation networks that are independently estimated from separate group-specific data and without due consideration of relationships that are conserved across multiple groups. METHOD: We propose a pathway-based differential network analysis in genomics (DINGO) model for estimating group-specific networks and making inference on the differential networks. DINGO jointly estimates the group-specific conditional dependencies by decomposing them into global and group-specific components. The delineation of these components allows for a more refined picture of the major driver and passenger events in the elucidation of cancer progression and development. RESULTS: Simulation studies demonstrate that DINGO provides more accurate group-specific conditional dependencies than achieved by using separate estimation approaches. We apply DINGO to key signaling pathways in glioblastoma to build differential networks for long-term survivors and short-term survivors in The Cancer Genome Atlas. The hub genes found by mRNA expression, DNA copy number, methylation and microRNA expression reveal several important roles in glioblastoma progression. AVAILABILITY AND IMPLEMENTATION: R Package at: odin.mdacc.tmc.edu/∼vbaladan. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Min Jin Ha, Veerabhadran Baladandayuthapani, Kim-Anh Do
Bioinform.2
2014 Latent Feature Decompositions for Integrative Analysis of Multi-Platform Genomic Data
abstract
Increased availability of multi-platform genomics data on matched samples has sparked research efforts to discover how diverse molecular features interact both within and between platforms. In addition, simultaneous measurements of genetic and epigenetic characteristics illuminate the roles their complex relationships play in disease progression and outcomes. However, integrative methods for diverse genomics data are faced with the challenges of ultra-high dimensionality and the existence of complex interactions both within and between platforms. We propose a novel modeling framework for integrative analysis based on decompositions of the large number of platform-specific features into a smaller number of latent features. Subsequently we build a predictive model for clinical outcomes accounting for both within- and between-platform interactions based on Bayesian model averaging procedures. Principal components, partial least squares and non-negative matrix factorization as well as sparse counterparts of each are used to define the latent features, and the performance of these decompositions is compared both on real and simulated data. The latent feature interactions are shown to preserve interactions between the original features and not only aid prediction but also allow explicit selection of outcome-related features. The methods are motivated by and applied to a glioblastoma multiforme data set from The Cancer Genome Atlas to predict patient survival times integrating gene expression, microRNA, copy number and methylation data. For the glioblastoma data, we find a high concordance between our selected prognostic genes and genes with known associations with glioblastoma. In addition, our model discovers several relevant cross-platform interactions such as copy number variation associated gene dosing and epigenetic regulation through promoter methylation. On simulated data, we show that our proposed method successfully incorporates interactions within and between genomic platforms to aid accurate prediction and variable selection. Our methods perform best when principal components are used to define the latent features.
Karl B. Gregory, Amin A. Momin, Kevin R. Coombes, Veerabhadran Baladandayuthapani
IEEE ACM Trans. Comput. Biol. Bioinform.4
2014 Bayesian Variable Selection in Linear Regression in One Pass for Large Datasets
abstract
Bayesian models are generally computed with Markov Chain Monte Carlo (MCMC) methods. The main disadvantage about MCMC methods is the large number of iterations they need to sample the posterior distributions of model parameters, especially for large data sets. On the other hand, variable selection remains a challenging problem due to its combinatorial search space, where Bayesian models are a promising solution. In this work, we study how to accelerate Bayesian model computation for variable selection in linear regression. We propose a fast Gibbs sampler algorithm, a widely used MCMC method, that incorporates several optimizations. We use non-informative and conjugate prior distributions on several model parameters, which enable data set summarization in one pass exploiting an augmented set of sufficient statistics. Thereafter the algorithm can iterate in main memory. Sufficient statistics are indexed with a sparse binary vector to efficiently compute matrix projections based on selected variables. Discovered variable subsets probabilities, selecting and discarding each variable, are stored on a hash table for fast retrieval in future iterations. We study how to integrate our algorithm into a database management system (DBMS), exploiting aggregate User-Defined Functions for parallel data summarization and stored procedures to manipulate matrices with arrays. An experimental evaluation with real data sets evaluates accuracy and time performance, comparing our DBMS-based algorithm, with the R package. Our algorithm is shown to produce accurate results, scale linearly on data set size and run orders of magnitude faster than the R package.
Carlos Ordonez 0001, Carlos Garcia-Alvarado, Veerabhadran Baladandayuthapani
ACM Trans. Knowl. Discov. Data3
2013 A fast convergence clustering algorithm merging MCMC and EM methods
abstract
Clustering is a fundamental problem in statistics and machine learning, whose solution is commonly computed by the Expectation-Maximization (EM) method, which finds a locally optimal solution for an objective function called log-likelihood. Since the surface of the log-likelihood function is non convex, a stochastic search with Markov Chain Monte Carlo (MCMC) methods can help escaping locally optimal solutions. In this article, we tackle two fundamental conflicting goals: Finding higher quality solutions and achieving faster convergence. With that motivation in mind, we introduce an efficient algorithm that combines elements of the EM and MCMC methods to find clustering solutions that are qualitatively better than those found by the standard EM method. Moreover, our hybrid algorithm allows tuning model parameters and understanding the uncertainty in their estimation. The main issue with MCMC methods is that they generally require a very large number of iterations to explore the posterior of each model parameter. Convergence is accelerated by several algorithmic improvements which include sufficient statistics, simplified model parameter priors, fixing covariance matrices and iterative sampling from small blocks of the data set. A brief experimental evaluation shows promising results.
David Sergio Matusevich, Carlos Ordonez 0001, Veerabhadran Baladandayuthapani
CIKM3
2013 Data mining algorithms as a service in the cloud exploiting relational database systems
abstract
We present a novel cloud system based on DBMS technology, where data mining algorithms are offered as a service. A local DBMS connects to the cloud and the cloud system returns computed data mining models as small relational tables that are archived and which can be easily transferred, queried and integrated with the client database. Unlike other analytic systems, our solution is not based on MapReduce. Our system avoids exporting large tables outside the local DBMS and thus it avoids transmitting large volumes of data to the cloud. The system offers three processing modes: local, cloud and hybrid, where a linear cost model is used to choose processing mode. In hybrid mode processing is split between the local DBMS and the cloud DBMS. Our system has a job scheduler with FIFO, SJF and RR policies to enhance response time and get partial results early. The cloud DBMS performs dynamic job scheduling, model computation and model archive management. Our system incorporates several optimizations: local data set summarization with sufficient statistics, sampling, caching matrices in RAM and selectively transmitting small matrices, back and forth. We show that in general the most efficient computing mechanism is hybrid processing: summarizing or sampling the data set in the local DBMS, transferring small matrices back and forth, leaving mathematically complex methods as a task for the cloud DBMS.
Carlos Ordonez 0001, Javier García-García 0001, Carlos Garcia-Alvarado, Wellington Cabrera, Veerabhadran Baladandayuthapani, Mohammed S. Quraishi
SIGMOD Conference5
2013 iBAG: integrative Bayesian analysis of high-dimensional multiplatform genomics data
abstract
MOTIVATION: Analyzing data from multi-platform genomics experiments combined with patients' clinical outcomes helps us understand the complex biological processes that characterize a disease, as well as how these processes relate to the development of the disease. Current data integration approaches are limited in that they do not consider the fundamental biological relationships that exist among the data obtained from different platforms. Statistical Model: We propose an integrative Bayesian analysis of genomics data (iBAG) framework for identifying important genes/biomarkers that are associated with clinical outcome. This framework uses hierarchical modeling to combine the data obtained from multiple platforms into one model. RESULTS: We assess the performance of our methods using several synthetic and real examples. Simulations show our integrative methods to have higher power to detect disease-related genes than non-integrative methods. Using the Cancer Genome Atlas glioblastoma dataset, we apply the iBAG model to integrate gene expression and methylation data to study their associations with patient survival. Our proposed method discovers multiple methylation-regulated genes that are related to patient survival, most of which have important biological functions in other diseases but have not been previously studied in glioblastoma. AVAILABILITY: http://odin.mdacc.tmc.edu/∼vbaladan/. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Veerabhadran Baladandayuthapani, Jeffrey S. Morris, Bradley M. Broom, Ganiraju Manyam, Kim-Anh Do
Bioinform.2
2013 Integrative network-based Bayesian analysis of diverse genomics data
abstract
BACKGROUND: In order to better understand cancer as a complex disease with multiple genetic and epigenetic factors, it is vital to model the fundamental biological relationships among these alterations as well as their relationships with important clinical outcomes. METHODS: We develop an integrative network-based Bayesian analysis (iNET) approach that allows us to jointly analyze multi-platform high-dimensional genomic data in a computationally efficient manner. The iNET approach is formulated as an objective Bayesian model selection problem for Gaussian graphical models to model joint dependencies among platform-specific features using known biological mechanisms. Using both simulated datasets and a glioblastoma (GBM) study from The Cancer Genome Atlas (TCGA), we illustrate the iNET approach via integrating three data types, microRNA, gene expression (mRNA), and patient survival time. RESULTS: We show that the iNET approach has greater power in identifying cancer-related microRNAs than non-integrative approaches based on realistic simulated datasets. In the TCGA GBM study, we found many mRNA-microRNA pairs and microRNAs that are associated with patient survival time, with some of these associations identified in previous studies. CONCLUSIONS: The iNET discovers relationships consistent with the underlying biological mechanisms among these variables, as well as identifying important biomarkers that are potentially relevant to patient survival. In addition, we identified some microRNAs that can potentially affect patient survival which are missed by non-integrative approaches.
Veerabhadran Baladandayuthapani, Christopher C. Holmes, Kim-Anh Do
BMC Bioinform.2
2011 Bayesian ensemble methods for survival prediction in gene expression data
abstract
MOTIVATION: We propose a Bayesian ensemble method for survival prediction in high-dimensional gene expression data. We specify a fully Bayesian hierarchical approach based on an ensemble 'sum-of-trees' model and illustrate our method using three popular survival models. Our non-parametric method incorporates both additive and interaction effects between genes, which results in high predictive accuracy compared with other methods. In addition, our method provides model-free variable selection of important prognostic markers based on controlling the false discovery rates; thus providing a unified procedure to select relevant genes and predict survivor functions. RESULTS: We assess the performance of our method several simulated and real microarray datasets. We show that our method selects genes potentially related to the development of the disease as well as yields predictive performance that is very competitive to many other existing methods. AVAILABILITY: http://works.bepress.com/veera/1/.
Vinícius Bonato, Veerabhadran Baladandayuthapani, Bradley M. Broom, Erik P. Sulman, Kenneth D. Aldape, Kim-Anh Do
Bioinform.2
2010 On the Computation of Stochastic Search Variable Selection in Linear Regression with UDFs
abstract
Computing Bayesian statistics with traditional techniques is extremely slow, specially when large data has to be exported from a relational DBMS. We propose algorithms for large scale processing of stochastic search variable selection (SSVS) for linear regression that can work entirely inside a DBMS. The traditional SSVS algorithm requires multiple scans of the input data in order to compute a regression model. Due to our optimizations, SSVS can be done in either one scan over the input table for large number of records with sufficient statistics, or one scan per iteration for high-dimensional data. We consider storage layouts which efficiently exploit DBMS parallel processing of aggregate functions. Experimental results demonstrate correctness, convergence and performance of our algorithms. Finally, the algorithms show good scalability for data with a very large number of records, or a very high number of dimensions.
Mario Navas, Carlos Ordonez 0001, Veerabhadran Baladandayuthapani
ICDM3