EDBT 2026 Demo / reviewers in the wild / expert
Peter J. Castaldi
dblp:72/9576
· DBLP profile ↗
15ranked-venue papers
1as first author
4since 2021 · last 2025
0000-0001-9920-4713ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Trustworthy machine learning · 61% Representation and self-supervised learning · 24% Probabilistic and Bayesian machine learning · 15% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Bioinformatics and computational biology · 80% Medical and health informatics · 20% | |
| Databases, data mining, and information retrieval
2 papers |
Data mining · 100% |
Topics — the 14 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
interpretability |
1.0 | 2 | 2022 | Explanations of Black-Box Models based on Directional Feature Interactions · ICLR 2022 Instance-wise Feature Grouping · NeurIPS 2020 |
Machine learning › Trustworthy machine learning › interpretability › explainable AI
interactive explanation |
0.6 | 1 | 2022 | Explanations of Black-Box Models based on Directional Feature Interactions · ICLR 2022 |
Machine learning › Trustworthy machine learning › interpretability › post-hoc explanation
model-agnostic explanation |
0.6 | 1 | 2022 | Explanations of Black-Box Models based on Directional Feature Interactions · ICLR 2022 |
Data mining
clustering |
0.5 | 2 | 2017 | Multiple Clustering Views from Multiple Uncertain Experts · ICML 2017 Interpretable Clustering via Discriminative Rectangle Mixture Model · ICDM 2016 |
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction › feature selection
feature grouping |
0.4 | 1 | 2020 | Instance-wise Feature Grouping · NeurIPS 2020 |
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
feature selection |
0.4 | 1 | 2020 | Instance-wise Feature Grouping · NeurIPS 2020 |
Bioinformatics and computational biology
multi-omics data integration |
0.4 | 1 | 2019 | Unsupervised discovery of phenotype-specific multi-omics networks · Bioinform. 2019 |
Bioinformatics and computational biology › biological network › network biology
network inference |
0.4 | 1 | 2019 | Unsupervised discovery of phenotype-specific multi-omics networks · Bioinform. 2019 |
Data mining › clustering
interpretable clustering |
0.2 | 1 | 2016 | Interpretable Clustering via Discriminative Rectangle Mixture Model · ICDM 2016 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference
bayesian nonparametric model |
0.2 | 1 | 2014 | Dual beta process priors for latent cluster discovery in chronic obstructive pulmonary disease · KDD 2014 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › bayesian nonparametric model
beta process |
0.2 | 1 | 2014 | Dual beta process priors for latent cluster discovery in chronic obstructive pulmonary disease · KDD 2014 |
Medical and health informatics › clinical data analysis › phenotyping
disease subtyping |
0.2 | 1 | 2014 | Dual beta process priors for latent cluster discovery in chronic obstructive pulmonary disease · KDD 2014 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference |
0.1 | 1 | 2017 | Multiple Clustering Views from Multiple Uncertain Experts · ICML 2017 |
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process |
0.1 | 1 | 2014 | Dual beta process priors for latent cluster discovery in chronic obstructive pulmonary disease · KDD 2014 |
Methods — techniques the papers use, named apart from their topics
variational inference · 1.2directional feature interaction · 0.6bayesian probabilistic model · 0.6variational lower bound · 0.4information theory · 0.4gumbel-softmax · 0.4sparse multiple canonical correlation analysis · 0.4gaussian process · 0.4beta process · 0.4mixture model · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A generalized higher-order correlation analysis framework for multi-omics network inferenceabstractMultiple -omics (genomics, proteomics, etc.) profiles are commonly generated to gain insight into a disease or physiological system. Constructing multi-omics networks with respect to the trait(s) of interest provides an opportunity to understand relationships between molecular features but integration is challenging due to multiple data sets with high dimensionality. One approach is to use canonical correlation to integrate one or two omics types and a single trait of interest. However, these types of methods may be limited due to (1) not accounting for higher-order correlations existing among features, (2) computational inefficiency when extending to more than two omics data when using a penalty term-based sparsity method, and (3) lack of flexibility for focusing on specific correlations (e.g., omics-to-phenotype correlation versus omics-to-omics correlations). In this work, we have developed a novel multi-omics network analysis pipeline called Sparse Generalized Tensor Canonical Correlation Analysis Network Inference (SGTCCA-Net) that can effectively overcome these limitations. We also introduce an implementation to improve the summarization of networks for downstream analyses. Simulation and real-data experiments demonstrate the effectiveness of our novel method for inferring omics networks and features of interest. Weixuan Liu, Katherine A. Pratte, Peter J. Castaldi, Craig P. Hersh, Russell Bowler, Farnoush Banaei Kashani, Katerina J. Kechris |
PLoS Comput. Biol. | 3 |
| 2024 | Partial correlation network analysis identifies coordinated gene expression within a regional cluster of COPD genome-wide association signalsabstractChronic obstructive pulmonary disease (COPD) is a complex disease influenced by well-established environmental exposures (most notably, cigarette smoking) and incompletely defined genetic factors. The chromosome 4q region harbors multiple genetic risk loci for COPD, including signals near HHIP, FAM13A, GSTCD, TET2, and BTC. Leveraging RNA-Seq data from lung tissue in COPD cases and controls, we estimated the co-expression network for genes in the 4q region bounded by HHIP and BTC (~70MB), through partial correlations informed by protein-protein interactions. We identified several co-expressed gene pairs based on partial correlations, including NPNT-HHIP, BTC-NPNT and FAM13A-TET2, which were replicated in independent lung tissue cohorts. Upon clustering the co-expression network, we observed that four genes previously associated to COPD: BTC, HHIP, NPNT and PPM1K appeared in the same network community. Finally, we discovered a sub-network of genes differentially co-expressed between COPD vs controls (including FAM13A, PPA2, PPM1K and TET2). Many of these genes were previously implicated in cell-based knock-out experiments, including the knocking out of SPP1 which belongs to the same genomic region and could be a potential local key regulatory gene. These analyses identify chromosome 4q as a region enriched for COPD genetic susceptibility and differential co-expression. Michele Gentili, Kimberly Glass, Enrico Maiorino, Brian D. Hobbs, Zhonghui Xu, Peter J. Castaldi, Michael H. Cho, Craig P. Hersh, Dandi Qiao, Jarrett D. Morrow, Vincent Carey, John Platig, Edwin K. Silverman |
PLoS Comput. Biol. | 6 |
| 2022 | Explanations of Black-Box Models based on Directional Feature Interactions
Aria Masoomi, Davin Hill, Zhonghui Xu, Craig P. Hersh, Edwin K. Silverman, Peter J. Castaldi, Stratis Ioannidis, Jennifer G. Dy |
ICLR | 6 |
| 2021 | Improved prediction of smoking status via isoform-aware RNA-seq deep learning modelsabstractMost predictive models based on gene expression data do not leverage information related to gene splicing, despite the fact that splicing is a fundamental feature of eukaryotic gene expression. Cigarette smoking is an important environmental risk factor for many diseases, and it has profound effects on gene expression. Using smoking status as a prediction target, we developed deep neural network predictive models using gene, exon, and isoform level quantifications from RNA sequencing data in 2,557 subjects in the COPDGene Study. We observed that models using exon and isoform quantifications clearly outperformed gene-level models when using data from 5 genes from a previously published prediction model. Whereas the test set performance of the previously published model was 0.82 in the original publication, our exon-based models including an exon-to-isoform mapping layer achieved a test set AUC (area under the receiver operating characteristic) of 0.88, which improved to an AUC of 0.94 using exon quantifications from a larger set of genes. Isoform variability is an important source of latent information in RNA-seq data that can be used to improve clinical prediction models. Zifeng Wang 0002, Aria Masoomi, Zhonghui Xu, Adel Boueiz, Sool Lee, Russell Bowler, Michael H. Cho, Edwin K. Silverman, Craig P. Hersh, Jennifer G. Dy, Peter J. Castaldi |
PLoS Comput. Biol. | 12 |
| 2020 | Instance-wise Feature GroupingabstractIn many learning problems, the domain scientist is often interested in discovering the groups of features that are redundant and are important for classification. Moreover, the features that belong to each group, and the important feature groups may vary per sample. But what do we mean by feature redundancy? In this paper, we formally define two types of redundancies using information theory: \textit{Representation} and \textit{Relevant redundancies}. We leverage these redundancies to design a formulation for instance-wise feature group discovery and reveal a theoretical guideline to help discover the appropriate number of groups. We approximate mutual information via a variational lower bound and learn the feature group and selector indicators with Gumbel-Softmax in optimizing our formulation. Experiments on synthetic data validate our theoretical claims. Experiments on MNIST, Fashion MNIST, and gene expression datasets show that our method discovers feature groups with high classification accuracies. Aria Masoomi, Chieh Wu, Zifeng Wang 0002, Peter J. Castaldi, Jennifer G. Dy |
NeurIPS | 5 |
| 2019 | Unsupervised discovery of phenotype-specific multi-omics networksabstractMOTIVATION: Complex diseases often involve a wide spectrum of phenotypic traits. Better understanding of the biological mechanisms relevant to each trait promotes understanding of the etiology of the disease and the potential for targeted and effective treatment plans. There have been many efforts towards omics data integration and network reconstruction, but limited work has examined the incorporation of relevant (quantitative) phenotypic traits. RESULTS: We propose a novel technique, sparse multiple canonical correlation network analysis (SmCCNet), for integrating multiple omics data types along with a quantitative phenotype of interest, and for constructing multi-omics networks that are specific to the phenotype. As a case study, we focus on miRNA-mRNA networks. Through simulations, we demonstrate that SmCCNet has better overall prediction performance compared to popular gene expression network construction and integration approaches under realistic settings. Applying SmCCNet to studies on chronic obstructive pulmonary disease (COPD) and breast cancer, we found enrichment of known relevant pathways (e.g. the Cadherin pathway for COPD and the interferon-gamma signaling pathway for breast cancer) as well as less known omics features that may be important to the diseases. Although those applications focus on miRNA-mRNA co-expression networks, SmCCNet is applicable to a variety of omics and other data types. It can also be easily generalized to incorporate multiple quantitative phenotype simultaneously. The versatility of SmCCNet suggests great potential of the approach in many areas. AVAILABILITY AND IMPLEMENTATION: The SmCCNet algorithm is written in R, and is freely available on the web at https://cran.r-project.org/web/packages/SmCCNet/index.html. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. W. Jenny Shi, Yonghua Zhuang, Pamela H. Russell, Brian D. Hobbs, Margaret M. Parker, Peter J. Castaldi, Pratyaydipta Rudra, Brian Vestal, Craig P. Hersh, Laura M. Saba, Katerina J. Kechris |
Bioinform. | 6 |
| 2018 | Crowdclustering with Partition LabelsabstractCrowdclustering is a practical way to incorporate domain knowledge into clustering, by combining opinions from multiple domain experts. Existing crowdclustering methods analyze binary pairwise similarity labels. However, in some applications, experts might provide partition labels. If we convert partition labels into pairwise similarity, then it would be difficult to understand the relationships between clustering solutions from different experts. In this paper, we propose a crowdclustering model that directly analyzes partition labels. The proposed model adopts a novel approach based on a modified multinomial logistic regression model, which simultaneously learns the number of clusters and determines hyper-planes that partition samples into clusters. The proposed model also learns a mapping between the latent clusters and expert labels, revealing the agreements and disagreements between experts. Experiments on benchmark data demonstrate that the proposed model simultaneously learns the number of clusters and discovers the clustering structure. An experiment on disease subtyping problem illustrates that the proposed model helps us understand the agreement and disagreement between experts. Junxiang Chen, Yale Chang, Peter J. Castaldi, Michael H. Cho, Brian D. Hobbs, Jennifer G. Dy |
AISTATS | 3 |
| 2017 | Clustering from Multiple Uncertain ExpertsabstractUtilizing expert input often improves clustering performance. However in a knowledge discovery problem, ground truth is unknown even to an expert. Thus, instead of one expert, we solicit the opinion from multiple experts. The key question motivating this work is: which experts should be assigned higher weights when there is disagreement on whether to put a pair of samples in the same group? To model the uncertainty in constraints from different experts, we build a probabilistic model for pairwise constraints through jointly modeling each expert’s accuracy and the mapping from features to latent cluster assignments. After learning our probabilistic discriminative clustering model and accuracies of different experts, 1) samples that were not annotated by any expert can be clustered using the discriminative clustering model; and 2) experts with higher accuracies are automatically assigned higher weights in determining the latent cluster assignments. Experimental results on UCI benchmark datasets and a real-world disease subtyping dataset demonstrate that our proposed approach outperforms competing alternatives, including semi-crowdsourced clustering, semi-supervised clustering with constraints from majority voting, and consensus clustering. Yale Chang, Junxiang Chen, Michael H. Cho, Peter J. Castaldi, Edwin K. Silverman, Jennifer G. Dy |
AISTATS | 4 |
| 2017 | Multiple Clustering Views from Multiple Uncertain ExpertsabstractExpert input can improve clustering performance. In today’s collaborative environment, the availability of crowdsourced multiple expert input is becoming common. Given multiple experts’ inputs, most existing approaches can only discover one clustering structure. However, data is multi-faced by nature and can be clustered in different ways (also known as views). In an exploratory analysis problem where ground truth is not known, different experts may have diverse views on how to cluster data. In this paper, we address the problem on how to automatically discover multiple ways to cluster data given potentially diverse inputs from multiple uncertain experts. We propose a novel Bayesian probabilistic model that automatically learns the multiple expert views and the clustering structure associated with each view. The benefits of learning the experts’ views include 1) enabling the discovery of multiple diverse clustering structures, and 2) improving the quality of clustering solution in each view by assigning higher weights to experts with higher confidence. In our approach, the expert views, multiple clustering structures and expert confidences are jointly learned via variational inference. Experimental results on synthetic datasets, benchmark datasets and a real-world disease subtyping problem show that our proposed approach outperforms competing baselines, including meta clustering, semi-supervised clustering, semi-crowdsourced clustering and consensus clustering. Yale Chang, Junxiang Chen, Michael H. Cho, Peter J. Castaldi, Edwin K. Silverman, Jennifer G. Dy |
ICML | 4 |
| 2017 | Clustering with Domain-Specific Usefulness ScoresabstractClustering is a challenging problem because given the same data set, it can be grouped in multiple different ways. Which of these clustering solutions is interesting depends on its domain application. Thus, incorporating domain expert input often improves clustering performance. However, most existing semi-supervised clustering techniques can only incorporate instance-level constraints (a few labels or must-link/cannot-link constraints), which domain experts may not be comfortable providing in knowledge discovery problems because categories are not known. Fortunately, domain experts often have an idea regarding properties that clustering solutions should have in order to be useful in domain application based on domain relevant scores. In this paper, we provide a framework for jointly optimizing the usefulness and quality of a clustering solution. Experiments on a synthetic data, a benchmark data, and a real-world disease subtyping problem demonstrate the usefulness of our proposed approach. Yale Chang, Junxiang Chen, Michael H. Cho, Peter J. Castaldi, Edwin K. Silverman, Jennifer G. Dy |
SDM | 4 |
| 2017 | A Bayesian Nonparametric Model for Disease Subtyping: Application to Emphysema PhenotypesabstractWe introduce a novel Bayesian nonparametric model that uses the concept of disease trajectories for disease subtype identification. Although our model is general, we demonstrate that by treating fractions of tissue patterns derived from medical images as compositional data, our model can be applied to study distinct progression trends between population subgroups. Specifically, we apply our algorithm to quantitative emphysema measurements obtained from chest CT scans in the COPDGene Study and show several distinct progression patterns. As emphysema is one of the major components of chronic obstructive pulmonary disease (COPD), the third leading cause of death in the United States [1], an improved definition of emphysema and COPD subtypes is of great interest. We investigate several models with our algorithm, and show that one with age , pack years (a measure of cigarette exposure), and smoking status as predictors gives the best compromise between estimated predictive performance and model complexity. This model identified nine subtypes which showed significant associations to seven single nucleotide polymorphisms (SNPs) known to associate with COPD. Additionally, this model gives better predictive accuracy than multiple, multivariate ordinary least squares regression as demonstrated in a five-fold cross validation analysis. We view our subtyping algorithm as a contribution that can be applied to bridge the gap between CT-level assessment of tissue composition to population-level analysis of compositional trends that vary between disease subtypes. James C. Ross, Peter J. Castaldi, Michael H. Cho, Junxiang Chen, Yale Chang, Jennifer G. Dy, Edwin K. Silverman, George R. Washko, Raúl San José Estépar |
IEEE Trans. Medical Imaging | 2 |
| 2016 | Interpretable Clustering via Discriminative Rectangle Mixture ModelabstractClustering is a technique that is usually applied as a tool for exploratory data analysis. Because of the exploratory nature of this task, it would be beneficial if a clustering method generates interpretable results, and allows incorporating domain knowledge. This motivates us to develop a probabilistic discriminative model that learns a rectangular decision rule for each cluster, we call Discriminative Rectangle Mixture (DReaM) model. DReaM gives interpretable clustering results, because the rectangular decision rules discovered explicitly illustrate how one cluster is defined and differs from other clusters. It also facilitates us to take advantage of existing rules because we can choose informative prior distributions for the rectangular rules. Moreover, DReaM allows that the features for generating rules do not have to be the same as the features for discovering cluster structure. We approximate the distribution for the rules discovered via variational inference. Experimental results demonstrate that DReaM gives more interpretable clustering results, and yet its performance is comparable to existing clustering methods when solving traditional clustering. Furthermore, in real applications, DReaM is able to effectively take advantage of domain knowledge, and to generate reasonable clustering results. Junxiang Chen, Yale Chang, Brian D. Hobbs, Peter J. Castaldi, Michael H. Cho, Edwin K. Silverman, Jennifer G. Dy |
ICDM | 4 |
| 2016 | Bipartite Community Structure of eQTLsabstractGenome Wide Association Studies (GWAS) and expression quantitative trait locus (eQTL) analyses have identified genetic associations with a wide range of human phenotypes. However, many of these variants have weak effects and understanding their combined effect remains a challenge. One hypothesis is that multiple SNPs interact in complex networks to influence functional processes that ultimately lead to complex phenotypes, including disease states. Here we present CONDOR, a method that represents both cis- and trans-acting SNPs and the genes with which they are associated as a bipartite graph and then uses the modular structure of that graph to place SNPs into a functional context. In applying CONDOR to eQTLs in chronic obstructive pulmonary disease (COPD), we found the global network "hub" SNPs were devoid of disease associations through GWAS. However, the network was organized into 52 communities of SNPs and genes, many of which were enriched for genes in specific functional classes. We identified local hubs within each community ("core SNPs") and these were enriched for GWAS SNPs for COPD and many other diseases. These results speak to our intuition: rather than single SNPs influencing single genes, we see groups of SNPs associated with the expression of families of functionally related genes and that disease SNPs are associated with the perturbation of those functions. These methods are not limited in their application to COPD and can be used in the analysis of a wide variety of disease processes and other phenotypic traits. John Platig, Peter J. Castaldi, Dawn L. DeMeo, John Quackenbush |
PLoS Comput. Biol. | 2 |
| 2014 | Dual beta process priors for latent cluster discovery in chronic obstructive pulmonary diseaseabstractChronic obstructive pulmonary disease (COPD) is a lung disease characterized by airflow limitation usually associated with an inflammatory response to noxious particles, such as cigarette smoke. COPD is currently the third leading cause of death in the United States and is the only leading cause of death that is increasing in prevalence. It also represents an enormous financial burden to society, costing tens of billions of dollars annually in the U.S. It is widely accepted by the medical community that COPD is a heterogeneous disease, with substantial evidence indicating that genetic variation contributes to varying levels of disease susceptibility. This heterogeneity makes it difficult to predict health decline and develop targeted treatments for better patient care. Although researchers have made several attempts to discover disease subtypes, results have been inconclusive, in part because standard clustering methods have not properly dealt with disease manifestations that may worsen with increased exposure. In this paper we introduce a transformative way of looking at the COPD subtyping task. Specifically, we model the relationship between risk factors (such as age and smoke exposure) and manifestations of disease severity using Gaussian Processes, which allow us to represent so-called "disease trajectories". We also posit that individuals can be associated with multiple disease types (latent clusters), which we assume are influenced by genetics. Furthermore, we predict that only subsets of the numerous disease-related quantitative features are useful for describing each latent subtype. We model these associations using two separate beta process priors, and we describe a variational inference approach to discover the most probable latent cluster assignments. Results are validated with associations to genetic markers. James C. Ross, Peter J. Castaldi, Michael H. Cho, Jennifer G. Dy |
KDD | 2 |
| 2011 | An empirical assessment of validation practices for molecular classifiersabstractProposed molecular classifiers may be overfit to idiosyncrasies of noisy genomic and proteomic data. Cross-validation methods are often used to obtain estimates of classification accuracy, but both simulations and case studies suggest that, when inappropriate methods are used, bias may ensue. Bias can be bypassed and generalizability can be tested by external (independent) validation. We evaluated 35 studies that have reported on external validation of a molecular classifier. We extracted information on study design and methodological features, and compared the performance of molecular classifiers in internal cross-validation versus external validation for 28 studies where both had been performed. We demonstrate that the majority of studies pursued cross-validation practices that are likely to overestimate classifier performance. Most studies were markedly underpowered to detect a 20% decrease in sensitivity or specificity between internal cross-validation and external validation [median power was 36% (IQR, 21-61%) and 29% (IQR, 15-65%), respectively]. The median reported classification performance for sensitivity and specificity was 94% and 98%, respectively, in cross-validation and 88% and 81% for independent validation. The relative diagnostic odds ratio was 3.26 (95% CI 2.04-5.21) for cross-validation versus independent validation. Finally, we reviewed all studies (n = 758) which cited those in our study sample, and identified only one instance of additional subsequent independent validation of these classifiers. In conclusion, these results document that many cross-validation practices employed in the literature are potentially biased and genuine progress in this field will require adoption of routine external validation of molecular classifiers, preferably in much larger studies than in current practice. Peter J. Castaldi, Issa J. Dahabreh, John P. A. Ioannidis |
Briefings Bioinform. | 1 |