EDBT 2026 Demo / reviewers in the wild / expert
Thao Vu
dblp:326/9589
· DBLP profile ↗
6ranked-venue papers
3as first author
6since 2021 · last 2025
0000-0001-5252-0006ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | cytoKernel: robust kernel embeddings for assessing differential expression of single-cell dataabstractMOTIVATION: High-throughput sequencing of single-cell data can be used to rigorously evaluate cell specification and enable intricate variations between groups or conditions to be identified. Many popular existing methods for differential expression target differences in aggregate measurement (mean, median, sum) and limit their approaches to detect only global differential changes. RESULTS: We present a robust method for differential expression of single-cell data using a kernel-based score test, cytoKernel. CytoKernel is specifically designed to assess the differential expression of single-cell RNA sequencing and high-dimensional flow or mass cytometry data using the full probability distribution pattern. cytoKernel is based on kernel embeddings which employs the probability distributions of the single-cell data, by calculating the pairwise divergence/distance between distributions of subjects. It can detect both patterns involving changes in the aggregate, as well as more elusive variations that are often overlooked due to the multimodal characteristics of single-cell data. We performed extensive benchmarks across both simulated and real data sets from mass cytometry data and single-cell RNA sequencing. The cytoKernel procedure effectively controls the false discovery rate and shows favorable performance compared to existing methods. The method is able to identify more differential patterns than existing approaches. We apply cytoKernel to assess gene expression and protein marker expression differences from cell subpopulations in various publicly available single-cell RNAseq and mass cytometry datasets. AVAILABILITY AND IMPLEMENTATION: The methods described in this paper are implemented in the open-source R package cytoKernel, which is freely available from Bioconductor at http://bioconductor.org/packages/cytoKernel. Tusharkanti Ghosh, Ryan M. Baxter, Souvik Seal, Victor G. Lui, Pratyaydipta Rudra, Thao Vu, Elena W. Y. Hsieh, Debashis Ghosh |
Bioinform. | 6 |
| 2024 | Smccnet 2.0: a comprehensive tool for multi-omics network inference with shiny visualizationabstractSparse multiple canonical correlation network analysis (SmCCNet) is a machine learning technique for integrating omics data along with a variable of interest (e.g., phenotype of complex disease), and reconstructing multi-omics networks that are specific to this variable. We present the second-generation SmCCNet (SmCCNet 2.0) that adeptly integrates single or multiple omics data types along with a quantitative or binary phenotype of interest. In addition, this new package offers a streamlined setup process that can be configured manually or automatically, ensuring a flexible and user-friendly experience. AVAILABILITY : This package is available in both CRAN: https://cran.r-project.org/web/packages/SmCCNet/index.html and Github: https://github.com/KechrisLab/SmCCNet under the MIT license. The network visualization tool is available at https://smccnet.shinyapps.io/smccnetnetwork/ . Weixuan Liu, Thao Vu, Iain R. Konigsberg, Katherine A. Pratte, Yonghua Zhuang, Katerina J. Kechris |
BMC Bioinform. | 2 |
| 2023 | NetSHy: network summarization via a hybrid approach leveraging topological propertiesabstractMOTIVATION: Biological networks can provide a system-level understanding of underlying processes. In many contexts, networks have a high degree of modularity, i.e. they consist of subsets of nodes, often known as subnetworks or modules, which are highly interconnected and may perform separate functions. In order to perform subsequent analyses to investigate the association between the identified module and a variable of interest, a module summarization, that best explains the module's information and reduces dimensionality is often needed. Conventional approaches for obtaining network representation typically rely only on the profiles of the nodes within the network while disregarding the inherent network topological information. RESULTS: In this article, we propose NetSHy, a hybrid approach which is capable of reducing the dimension of a network while incorporating topological properties to aid the interpretation of the downstream analyses. In particular, NetSHy applies principal component analysis (PCA) on a combination of the node profiles and the well-known Laplacian matrix derived directly from the network similarity matrix to extract a summarization at a subject level. Simulation scenarios based on random and empirical networks at varying network sizes and sparsity levels show that NetSHy outperforms the conventional PCA approach applied directly on node profiles, in terms of recovering the true correlation with a phenotype of interest and maintaining a higher amount of explained variation in the data when networks are relatively sparse. The robustness of NetSHy is also demonstrated by a more consistent correlation with the observed phenotype as the sample size decreases. Lastly, a genome-wide association study is performed as an application of a downstream analysis, where NetSHy summarization scores on the biological networks identify more significant single nucleotide polymorphisms than the conventional network representation. AVAILABILITY AND IMPLEMENTATION: R code implementation of NetSHy is available at https://github.com/thaovu1/NetSHy. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Thao Vu, Elizabeth Litkowski, Weixuan Liu, Katherine A. Pratte, Leslie Lange, Russell Bowler, Farnoush Banaei Kashani, Katerina J. Kechris |
Bioinform. | 1 |
| 2023 | FunSpace: A functional and spatial analytic approach to cell imaging data using entropy measuresabstractSpatial heterogeneity in the tumor microenvironment (TME) plays a critical role in gaining insights into tumor development and progression. Conventional metrics typically capture the spatial differential between TME cellular patterns by either exploring the cell distributions in a pairwise fashion or aggregating the heterogeneity across multiple cell distributions without considering the spatial contribution. As such, none of the existing approaches has fully accounted for the simultaneous heterogeneity caused by both cellular diversity and spatial configurations of multiple cell categories. In this article, we propose an approach to leverage spatial entropy measures at multiple distance ranges to account for the spatial heterogeneity across different cellular organizations. Functional principal component analysis (FPCA) is applied to estimate FPC scores which are then served as predictors in a Cox regression model to investigate the impact of spatial heterogeneity in the TME on survival outcome, potentially adjusting for other confounders. Using a non-small cell lung cancer dataset (n = 153) as a case study, we found that the spatial heterogeneity in the TME cellular composition of CD14+ cells, CD19+ B cells, CD4+ and CD8+ T cells, and CK+ tumor cells, had a significant non-zero effect on the overall survival (p = 0.027). Furthermore, using a publicly available multiplexed ion beam imaging (MIBI) triple-negative breast cancer dataset (n = 33), our proposed method identified a significant impact of cellular interactions between tumor and immune cells on the overall survival (p = 0.046). In simulation studies under different spatial configurations, the proposed method demonstrated a high predictive power by accounting for both clinical effect and the impact of spatial heterogeneity. Thao Vu, Souvik Seal, Tusharkanti Ghosh, Mansooreh Ahmadian, Julia Wrobel, Debashis Ghosh |
PLoS Comput. Biol. | 1 |
| 2022 | Effective Subject Representation based on Multi-omics Disease Networks using Graph EmbeddingabstractThe study of complex behavior of biological systems has become increasingly dependent on evolutionary network modeling. In particular, multi-omics networks capture interactions between biomolecules such as proteins and metabolites, providing a basis for predicting relationships between such biomolecules and various phenotypic traits of complex diseases. In this paper, we introduce an integrative framework that given a multi-omics network representing a cohort of subjects, learns expressive representations for network nodes, and combines the learned nodes representations with the biological profiles of individual subjects for enriched representation of the subjects. With extensive empirical evaluation using real-world multi-omics networks, we show that our proposed framework significantly outperforms existing and baseline methods in terms of subject representation accuracy, particularly when the multi-omics network representing the cohort is sparse and structured and therefore, more informative. Sundous Hussein, Thao Vu, Leslie Lange, Russell Bowler, Katerina J. Kechris, Farnoush Banaei Kashani |
BIBM | 2 |
| 2022 | SPF: A spatial and functional data analytic approach to cell imaging dataabstractThe tumor microenvironment (TME), which characterizes the tumor and its surroundings, plays a critical role in understanding cancer development and progression. Recent advances in imaging techniques enable researchers to study spatial structure of the TME at a single-cell level. Investigating spatial patterns and interactions of cell subtypes within the TME provides useful insights into how cells with different biological purposes behave, which may consequentially impact a subject's clinical outcomes. We utilize a class of well-known spatial summary statistics, the K-function and its variants, to explore inter-cell dependence as a function of distances between cells. Using techniques from functional data analysis, we introduce an approach to model the association between these summary spatial functions and subject-level outcomes, while controlling for other clinical scalar predictors such as age and disease stage. In particular, we leverage the additive functional Cox regression model (AFCM) to study the nonlinear impact of spatial interaction between tumor and stromal cells on overall survival in patients with non-small cell lung cancer, using multiplex immunohistochemistry (mIHC) data. The applicability of our approach is further validated using a publicly available multiplexed ion beam imaging (MIBI) triple-negative breast cancer dataset. Thao Vu, Julia Wrobel, Benjamin G. Bitler, Erin L. Schenk, Kimberly R. Jordan, Debashis Ghosh |
PLoS Comput. Biol. | 1 |