VLDB 2026 Research / reviewers in the wild / expert
Ziv Bar-Joseph
dblp:66/6963
· DBLP profile ↗
71ranked-venue papers
9as first author
19since 2021 · last 2025
0000-0003-3430-6051ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 58 · 4 first-author · 18 since 2021Artificial intelligence and machine learning · 3Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-authorSystems, architecture and hardware · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 1Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Active Learning for Protein Structure Prediction
Zexin Xue, Michael Bailey, Ruijiang Li, Alejandro Corrochano-Navarro, Sizhen Li, Lorenzo Kogler-Anele, Qui Yu, Ziv Bar-Joseph, Sven Jager |
RECOMB | 9 |
| 2025 | Recovering time-varying networks from single-cell dataabstractMOTIVATION: Gene regulation is a dynamic process that underlies all aspects of human development, disease response, and other biological processes. The reconstruction of temporal gene regulatory networks has conventionally relied on regression analysis, graphical models, or other types of relevance networks. With the large increase in time series single-cell data, new approaches are needed to address the unique scale and nature of these data for reconstructing such networks. RESULTS: Here, we develop a deep neural network, Marlene, to infer dynamic graphs from time series single-cell gene expression data. Marlene constructs directed gene networks using a self-attention mechanism where the weights evolve over time using recurrent units. By employing meta learning, the model is able to recover accurate temporal networks even for rare cell types. In addition, Marlene can identify gene interactions relevant to specific biological responses, including COVID-19 immune response, fibrosis, and aging, paving the way for potential treatments. AVAILABILITY AND IMPLEMENTATION: The code used to train Marlene is available at https://github.com/euxhenh/Marlene. Euxhen Hasanaj, Barnabás Póczos, Ziv Bar-Joseph |
Bioinform. | 3 |
| 2025 | ENACT: End-to-End Analysis of Visium High Definition (HD) DataabstractMOTIVATION: Spatial transcriptomics (ST) enables the study of gene expression within its spatial context in histopathology samples. To date, a limiting factor has been the resolution of sequencing based ST products. The introduction of the Visium High Definition (HD) technology opens the door to cell resolution ST studies. However, challenges remain in the ability to accurately map transcripts to cells and in assigning cell types based on the transcript data. RESULTS: We developed ENACT, a self-contained pipeline that integrates advanced cell segmentation with Visium HD transcriptomics data to infer cell types across whole tissue sections. Our pipeline incorporates novel bin-to-cell assignment methods, enhancing the accuracy of single-cell transcript estimates. Validated on diverse synthetic and real datasets, our approach is both scalable to samples with hundreds of thousands of cells and effective, offering a robust solution for spatially resolved transcriptomics analysis. AVAILABILITY AND IMPLEMENTATION: ENACT source code is available at https://github.com/Sanofi-Public/enact-pipeline. Experimental data are available at https://zenodo.org/records/14748859. Mena Soliman Asaad Kamel, Yiwen Song, Ana Solbas, Sergio Villordo, Amrut Sarangi, Pavel Senin, Sunaal Mathew, Luis Cano Ayestas, Clément Levin, Seqian Wang, Marion Classe, Ziv Bar-Joseph, Albert Pla |
Bioinform. | 12 |
| 2025 | PyEvoCell: an LLM-augmented single-cell trajectory analysis dashboardabstractMOTIVATION: Several methods have been developed for trajectory inference in single-cell studies. However, identifying relevant lineages among several cell types and interpreting the results of downstream analysis remains a challenging task that requires deep understanding of various cell type transitions and progression patterns. Therefore, there is a need for methods that can aid researchers in the analysis and interpretation of such trajectories. RESULTS: We developed PyEvoCell, a dashboard for trajectory interpretation and analysis that is augmented by large language model (LLM) capabilities. PyEvoCell applies the LLM to the outputs of trajectory inference methods such as Monocle3, to suggest biologically relevant lineages. Once a lineage is defined, users can conduct differential expression and functional analyses which are also interpreted by the LLM. Finally, any hypothesis or claim derived from the analysis can be validated using the veracity filter, a feature enabled by the LLM, to confirm or reject claims by providing relevant PubMed citations. AVAILABILITY AND IMPLEMENTATION: The software is available at https://github.com/Sanofi-Public/PyEvoCell. It contains installation instructions, user manual, demo datasets, as well as license conditions. https://doi.org/10.5281/zenodo.15114803. Sachin Mathur, Mathieu Beauvais, Arnau Giribet, Nicolas Aragon Barrero, Chaorui Zhang, Towsif Rahman, Seqian Wang, Jeremy Huang, Nima Nouri, Andre H. Kurlovs, Ziv Bar-Joseph, Peyman Passban |
Bioinform. | 11 |
| 2024 | Integrating patients in time series clinical transcriptomics dataabstractMOTIVATION: Analysis of time series transcriptomics data from clinical trials is challenging. Such studies usually profile very few time points from several individuals with varying response patterns and dynamics. Current methods for these datasets are mainly based on linear, global orderings using visit times which do not account for the varying response rates and subgroups within a patient cohort. RESULTS: We developed a new method that utilizes multi-commodity flow algorithms for trajectory inference in large scale clinical studies. Recovered trajectories satisfy individual-based timing restrictions while integrating data from multiple patients. Testing the method on multiple drug datasets demonstrated an improved performance compared to prior approaches suggested for this task, while identifying novel disease subtypes that correspond to heterogeneous patient response patterns. AVAILABILITY AND IMPLEMENTATION: The source code and instructions to download the data have been deposited on GitHub at https://github.com/euxhenh/Truffle. Euxhen Hasanaj, Sachin Mathur, Ziv Bar-Joseph |
Bioinform. | 3 |
| 2024 | SpatialOne: end-to-end analysis of visium data at scaleabstractMOTIVATION: Spatial transcriptomics allow to quantify mRNA expression within the spatial context. Nonetheless, in-depth analysis of spatial transcriptomics data remains challenging and difficult to scale due to the number of methods and libraries required for that purpose. RESULTS: Here we present SpatialOne, an end-to-end pipeline designed to simplify the analysis of 10x Visium data by combining multiple state-of-the-art computational methods to segment, deconvolve, and quantify spatial information; this approach streamlines the analysis of reproducible spatial-data at scale. AVAILABILITY AND IMPLEMENTATION: SpatialOne source code and execution examples are available at https://github.com/Sanofi-Public/spatialone-pipeline, experimental data is available at https://zenodo.org/records/12605154. SpatialOne is distributed as a docker container image. Mena Soliman Asaad Kamel, Amrut Sarangi, Pavel Senin, Sergio Villordo, Sunaal Mathew, Het Barot, Seqian Wang, Ana Solbas, Luis Cano Ayestas, Marion Classe, Ziv Bar-Joseph, Albert Pla |
Bioinform. | 11 |
| 2024 | Representations of lipid nanoparticles using large language models for transfection efficiency predictionabstractMOTIVATION: Lipid nanoparticles (LNPs) are the most widely used vehicles for mRNA vaccine delivery. The structure of the lipids composing the LNPs can have a major impact on the effectiveness of the mRNA payload. Several properties should be optimized to improve delivery and expression including biodegradability, synthetic accessibility, and transfection efficiency. RESULTS: To optimize LNPs, we developed and tested models that enable the virtual screening of LNPs with high transfection efficiency. Our best method uses the lipid Simplified Molecular-Input Line-Entry System (SMILES) as inputs to a large language model. Large language model-generated embeddings are then used by a downstream gradient-boosting classifier. As we show, our method can more accurately predict lipid properties, which could lead to higher efficiency and reduced experimental time and costs. AVAILABILITY AND IMPLEMENTATION: Code and data links available at: https://github.com/Sanofi-Public/LipoBART. Saeed Moayedpour, Jonathan Broadbent, Saleh Riahi, Michael Bailey, Hoa V. Thu, Dimitar A. Dobchev, Akshay Balsubramani, Ricardo N. D. Santos, Lorenzo Kogler-Anele, Alejandro Corrochano-Navarro, Sizhen Li, Fernando U. Montoya, Vikram Agarwal, Ziv Bar-Joseph, Sven Jager |
Bioinform. | 14 |
| 2024 | Proxy endpoints - bridging clinical trials and real world dataabstractOBJECTIVE: Disease severity scores, or endpoints, are routinely measured during Randomized Controlled Trials (RCTs) to closely monitor the effect of treatment. In real-world clinical practice, although a larger set of patients is observed, the specific RCT endpoints are often not captured, which makes it hard to utilize real-world data (RWD) to evaluate drug efficacy in larger populations. METHODS: To overcome this challenge, we developed an ensemble technique which learns proxy models of disease endpoints in RWD. Using a multi-stage learning framework applied to RCT data, we first identify features considered significant drivers of disease available within RWD. To create endpoint proxy models, we use Explainable Boosting Machines (EBMs) which allow for both end-user interpretability and modeling of non-linear relationships. RESULTS: We demonstrate our approach on two diseases, rheumatoid arthritis (RA) and atopic dermatitis (AD). As we show, our combined feature selection and prediction method achieves good results for both disease areas, improving upon prior methods proposed for predictive disease severity scoring. CONCLUSION: Having disease severity over time for a patient is important to further disease understanding and management. Our results open the door to more use cases in the space of RA and AD such as treatment effect estimates or prognostic scoring on RWD. Our framework may be extended beyond RA and AD to other diseases where the severity score is not well measured in electronic health records. Maxim Kryukov, Kathleen P. Moriarty, Macarena Villamea, Ingrid O'dwyer, Ohn Chow, Flavio Dormont, Ramon Hernandez, Ziv Bar-Joseph, Brandon Rufino |
J. Biomed. Informatics | 8 |
| 2024 | Predicting lung aging using scRNA-Seq dataabstractAge prediction based on single cell RNA-Sequencing data (scRNA-Seq) can provide information for patients' susceptibility to various diseases and conditions. In addition, such analysis can be used to identify aging related genes and pathways. To enable age prediction based on scRNA-Seq data, we developed PolyEN, a new regression model which learns continuous representation for expression over time. These representations are then used by PolyEN to integrate genes to predict an age. Existing and new lung aging data we profiled demonstrated PolyEN's improved performance over existing methods for age prediction. Our results identified lung epithelial cells as the most significant predictors for non-smokers while lung endothelial cells led to the best chronological age prediction results for smokers. Alex Singh, John E. McDonough, Taylor Sterling Adams, Robin Vos, Ruben De Man, Greg Myers, Laurens J. Ceulemans, Bart M. Vanaudenaerde, Wim A. Wuyts, Xiting Yan, Jonas Schupp, James S. Hagood, Naftali Kaminski, Ziv Bar-Joseph |
PLoS Comput. Biol. | 15 |
| 2024 | Constrained Pseudo-Time Ordering for Clinical Transcriptomics DataabstractTime series RNASeq studies can enable understanding of the dynamics of disease progression and treatment response in patients. They also provide information on biomarkers, activated and repressed pathways, and more. While useful, data from multiple patients is challenging to integrate due to the heterogeneity in treatment response among patients, and the small number of timepoints that are usually profiled. Due to the heterogeneity among patients, relying on the sampled time points to integrate data across individuals is challenging and does not lead to correct reconstruction of the response patterns. To address these challenges, we developed a new constrained based pseudo-time ordering method for analyzing transcriptomics data in clinical and response studies. Our method allows the assignment of samples to their correct placement on the response curve while respecting the individual patient order. We use polynomials to represent gene expression over the duration of the study and an EM algorithm to determine parameters and locations. Application to four treatment response datasets shows that our method improves on prior methods and leads to accurate orderings that provide new biological insight on the disease and response. Sachin Mathur, Hamid Mattoo, Ziv Bar-Joseph |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2023 | Deriving spatial features from in situ proteomics imaging to enhance cancer survival analysisabstractMOTIVATION: Spatial proteomics data have been used to map cell states and improve our understanding of tissue organization. More recently, these methods have been extended to study the impact of such organization on disease progression and patient survival. However, to date, the majority of supervised learning methods utilizing these data types did not take full advantage of the spatial information, impacting their performance and utilization. RESULTS: Taking inspiration from ecology and epidemiology, we developed novel spatial feature extraction methods for use with spatial proteomics data. We used these features to learn prediction models for cancer patient survival. As we show, using the spatial features led to consistent improvement over prior methods that used the spatial proteomics data for the same task. In addition, feature importance analysis revealed new insights about the cell interactions that contribute to patient survival. AVAILABILITY AND IMPLEMENTATION: The code for this work can be found at gitlab.com/enable-medicine-public/spatsurv. Monica T. Dayao, Alexandro Trevino, Honesty Kim, Matthew Ruffalo, H. Blaize D'angio, Ryan Preska, Umamaheswar Duvvuri, Aaron T. Mayer, Ziv Bar-Joseph |
Bioinform. | 9 |
| 2022 | Unsupervised Cell Functional Annotation for Single-Cell RNA-Seq
Dongshunyi Li, Ziv Bar-Joseph |
RECOMB | 3 |
| 2022 | Clustering spatial transcriptomics dataabstractMOTIVATION: Recent advancements in fluorescence in situ hybridization (FISH) techniques enable them to concurrently obtain information on the location and gene expression of single cells. A key question in the initial analysis of such spatial transcriptomics data is the assignment of cell types. To date, most studies used methods that only rely on the expression levels of the genes in each cell for such assignments. To fully utilize the data and to improve the ability to identify novel sub-types, we developed a new method, FICT, which combines both expression and neighborhood information when assigning cell types. RESULTS: FICT optimizes a probabilistic function that we formalize and for which we provide learning and inference algorithms. We used FICT to analyze both simulated and several real spatial transcriptomics data. As we show, FICT can accurately identify cell types and sub-types, improving on expression only methods and other methods proposed for clustering spatial transcriptomics data. Some of the spatial sub-types identified by FICT provide novel hypotheses about the new functions for excitatory and inhibitory neurons. AVAILABILITY AND IMPLEMENTATION: FICT is available at: https://github.com/haotianteng/FICT. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Haotian Teng, Ziv Bar-Joseph |
Bioinform. | 3 |
| 2022 | Integrating longitudinal clinical and microbiome data to predict growth faltering in preterm infants
Jose Lugo-Martinez, Siwei Xu, Justine Levesque, Daniel Gallagher, Leslie A. Parker, Josef Neu, Christopher J. Stewart, Janet E. Berrington, Nicholas D. Embleton, Gregory Young, Katherine E. Gregory, Misty Good, Arti Tandon, David Genetti, Tracy Warren, Ziv Bar-Joseph |
J. Biomed. Informatics | 16 |
| 2022 | CINS: Cell Interaction Network inference from Single cell expression dataabstractStudies comparing single cell RNA-Seq (scRNA-Seq) data between conditions mainly focus on differences in the proportion of cell types or on differentially expressed genes. In many cases these differences are driven by changes in cell interactions which are challenging to infer without spatial information. To determine cell-cell interactions that differ between conditions we developed the Cell Interaction Network Inference (CINS) pipeline. CINS combines Bayesian network analysis with regression-based modeling to identify differential cell type interactions and the proteins that underlie them. We tested CINS on a disease case control and on an aging mouse dataset. In both cases CINS correctly identifies cell type interactions and the ligands involved in these interactions improving on prior methods suggested for cell interaction predictions. We performed additional mouse aging scRNA-Seq experiments which further support the interactions identified by CINS. Carlos Cosme, Taylor Sterling Adams, Jonas Schupp, Koji Sakamoto, Nikos Xylourgidis, Matthew Ruffalo, Naftali Kaminski, Ziv Bar-Joseph |
PLoS Comput. Biol. | 10 |
| 2021 | The Epigenetic Consensus Problem
Sabrina Rashid, Gadi Taubenfeld, Ziv Bar-Joseph |
SIROCCO | 3 |
| 2021 | Deep learning of gene relationships from single cell time-course expression dataabstractTime-course gene-expression data have been widely used to infer regulatory and signaling relationships between genes. Most of the widely used methods for such analysis were developed for bulk expression data. Single cell RNA-Seq (scRNA-Seq) data offer several advantages including the large number of expression profiles available and the ability to focus on individual cells rather than averages. However, the data also raise new computational challenges. Using a novel encoding for scRNA-Seq expression data, we develop deep learning methods for interaction prediction from time-course data. Our methods use a supervised framework which represents the data as 3D tensor and train convolutional and recurrent neural networks for predicting interactions. We tested our time-course deep learning (TDL) models on five different time-series scRNA-Seq datasets. As we show, TDL can accurately identify causal and regulatory gene-gene interactions and can also be used to assign new function to genes. TDL improves on prior methods for the above tasks and can be generally applied to new time-series scRNA-Seq data. Ziv Bar-Joseph |
Briefings Bioinform. | 2 |
| 2021 | Identifying signaling genes in spatial single-cell expression dataabstractMOTIVATION: Recent technological advances enable the profiling of spatial single-cell expression data. Such data present a unique opportunity to study cell-cell interactions and the signaling genes that mediate them. However, most current methods for the analysis of these data focus on unsupervised descriptive modeling, making it hard to identify key signaling genes and quantitatively assess their impact. RESULTS: We developed a Mixture of Experts for Spatial Signaling genes Identification (MESSI) method to identify active signaling genes within and between cells. The mixture of experts strategy enables MESSI to subdivide cells into subtypes. MESSI relies on multi-task learning using information from neighboring cells to improve the prediction of response genes within a cell. Applying the methods to three spatial single-cell expression datasets, we show that MESSI accurately predicts the levels of response genes, improving upon prior methods and provides useful biological insights about key signaling genes and subtypes of excitatory neuron cells. AVAILABILITY AND IMPLEMENTATION: MESSI is available at: https://github.com/doraadong/MESSI. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Dongshunyi Li, Ziv Bar-Joseph |
Bioinform. | 3 |
| 2021 | Dhaka: variational autoencoder for unmasking tumor heterogeneity from single cell genomic dataabstractMOTIVATION: Intra-tumor heterogeneity is one of the key confounding factors in deciphering tumor evolution. Malignant cells exhibit variations in their gene expression, copy numbers and mutation even when originating from a single progenitor cell. Single cell sequencing of tumor cells has recently emerged as a viable option for unmasking the underlying tumor heterogeneity. However, extracting features from single cell genomic data in order to infer their evolutionary trajectory remains computationally challenging due to the extremely noisy and sparse nature of the data. RESULTS: Here we describe 'Dhaka', a variational autoencoder method which transforms single cell genomic data to a reduced dimension feature space that is more efficient in differentiating between (hidden) tumor subpopulations. Our method is general and can be applied to several different types of genomic data including copy number variation from scDNA-Seq and gene expression from scRNA-Seq experiments. We tested the method on synthetic and six single cell cancer datasets where the number of cells ranges from 250 to 6000 for each sample. Analysis of the resulting feature space revealed subpopulations of cells and their marker genes. The features are also able to infer the lineage and/or differentiation trajectory between cells greatly improving upon prior methods suggested for feature extraction and dimensionality reduction of such data. AVAILABILITY AND IMPLEMENTATION: All the datasets used in the paper are publicly available and developed software package and supporting info is available on Github https://github.com/MicrosoftGenomics/Dhaka. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Sabrina Rashid, Sohrab Shah, Ziv Bar-Joseph, Ravi Pandya |
Bioinform. | 3 |
| 2020 | Supervised Adversarial Alignment of Single-Cell RNA-seq Data
Songwei Ge, Haohan Wang, Amir Alavi, Eric P. Xing, Ziv Bar-Joseph |
RECOMB | 5 |
| 2020 | Iterative point set registration for aligning scRNA-seq dataabstractSeveral studies profile similar single cell RNA-Seq (scRNA-Seq) data using different technologies and platforms. A number of alignment methods have been developed to enable the integration and comparison of scRNA-Seq data from such studies. While each performs well on some of the datasets, to date no method was able to both perform the alignment using the original expression space and generalize to new data. To enable such analysis we developed Single Cell Iterative Point set Registration (SCIPR) which extends methods that were successfully applied to align image data to scRNA-Seq. We discuss the required changes needed, the resulting optimization function, and algorithms for learning a transformation function for aligning data. We tested SCIPR on several scRNA-Seq datasets. As we show it successfully aligns data from several different cell types, improving upon prior methods proposed for this task. In addition, we show the parameters learned by SCIPR can be used to align data not used in the training and to identify key cell type-specific genes. Amir Alavi, Ziv Bar-Joseph |
PLoS Comput. Biol. | 2 |
| 2020 | Inferring TF activation order in time series scRNA-Seq studiesabstractMethods for the analysis of time series single cell expression data (scRNA-Seq) either do not utilize information about transcription factors (TFs) and their targets or only study these as a post-processing step. Using such information can both, improve the accuracy of the reconstructed model and cell assignments, while at the same time provide information on how and when the process is regulated. We developed the Continuous-State Hidden Markov Models TF (CSHMM-TF) method which integrates probabilistic modeling of scRNA-Seq data with the ability to assign TFs to specific activation points in the model. TFs are assumed to influence the emission probabilities for cells assigned to later time points allowing us to identify not just the TFs controlling each path but also their order of activation. We tested CSHMM-TF on several mouse and human datasets. As we show, the method was able to identify known and novel TFs for all processes, assigned time of activation agrees with both expression information and prior knowledge and combinatorial predictions are supported by known interactions. We also show that CSHMM-TF improves upon prior methods that do not utilize TF-gene interaction. Chieh Lin, Ziv Bar-Joseph |
PLoS Comput. Biol. | 3 |
| 2019 | Continuous-state HMMs for modeling time-series single-cell RNA-Seq dataabstractMOTIVATION: Methods for reconstructing developmental trajectories from time-series single-cell RNA-Seq (scRNA-Seq) data can be largely divided into two categories. The first, often referred to as pseudotime ordering methods are deterministic and rely on dimensionality reduction followed by an ordering step. The second learns a probabilistic branching model to represent the developmental process. While both types have been successful, each suffers from shortcomings that can impact their accuracy. RESULTS: We developed a new method based on continuous-state HMMs (CSHMMs) for representing and modeling time-series scRNA-Seq data. We define the CSHMM model and provide efficient learning and inference algorithms which allow the method to determine both the structure of the branching process and the assignment of cells to these branches. Analyzing several developmental single-cell datasets, we show that the CSHMM method accurately infers branching topology and correctly and continuously assign cells to paths, improving upon prior methods proposed for this task. Analysis of genes based on the continuous cell assignment identifies known and novel markers for different cell types. AVAILABILITY AND IMPLEMENTATION: Software and Supporting website: www.andrew.cmu.edu/user/chiehl1/CSHMM/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Chieh Lin, Ziv Bar-Joseph |
Bioinform. | 2 |
| 2019 | Network-guided prediction of aromatase inhibitor response in breast cancerabstractPrediction of response to specific cancer treatments is complicated by significant heterogeneity between tumors in terms of mutational profiles, gene expression, and clinical measures. Here we focus on the response of Estrogen Receptor (ER)+ post-menopausal breast cancer tumors to aromatase inhibitors (AI). We use a network smoothing algorithm to learn novel features that integrate several types of high throughput data and new cell line experiments. These features greatly improve the ability to predict response to AI when compared to prior methods. For a subset of the patients, for which we obtained more detailed clinical information, we can further predict response to a specific AI drug. Matthew Ruffalo, Roby Thomas, Adrian V. Lee, Steffi Oesterreich, Ziv Bar-Joseph |
PLoS Comput. Biol. | 6 |
| 2018 | iDREM: Interactive visualization of dynamic regulatory networksabstractThe Dynamic Regulatory Events Miner (DREM) software reconstructs dynamic regulatory networks by integrating static protein-DNA interaction data with time series gene expression data. In recent years, several additional types of high-throughput time series data have been profiled when studying biological processes including time series miRNA expression, proteomics, epigenomics and single cell RNA-Seq. Combining all available time series and static datasets in a unified model remains an important challenge and goal. To address this challenge we have developed a new version of DREM termed interactive DREM (iDREM). iDREM provides support for all data types mentioned above and combines them with existing interaction data to reconstruct networks that can lead to novel hypotheses on the function and timing of regulators. Users can interactively visualize and query the resulting model. We showcase the functionality of the new tool by applying it to microglia developmental data from multiple labs. James S. Hagood, Namasivayam Ambalavanan, Naftali Kaminski, Ziv Bar-Joseph |
PLoS Comput. Biol. | 5 |
| 2018 | Predicting protein targets for drug-like compounds using transcriptomicsabstractAn expanded chemical space is essential for improved identification of small molecules for emerging therapeutic targets. However, the identification of targets for novel compounds is biased towards the synthesis of known scaffolds that bind familiar protein families, limiting the exploration of chemical space. To change this paradigm, we validated a new pipeline that identifies small molecule-protein interactions and works even for compounds lacking similarity to known drugs. Based on differential mRNA profiles in multiple cell types exposed to drugs and in which gene knockdowns (KD) were conducted, we showed that drugs induce gene regulatory networks that correlate with those produced after silencing protein-coding genes. Next, we applied supervised machine learning to exploit drug-KD signature correlations and enriched our predictions using an orthogonal structure-based screen. As a proof-of-principle for this regimen, top-10/top-100 target prediction accuracies of 26% and 41%, respectively, were achieved on a validation of set 152 FDA-approved drugs and 3104 potential targets. We then predicted targets for 1680 compounds and validated chemical interactors with four targets that have proven difficult to chemically modulate, including non-covalent inhibitors of HRAS and KRAS. Importantly, drug-target interactions manifest as gene expression correlations between drug treatment and both target gene KD and KD of genes that act up- or down-stream of the target, even for relatively weak binders. These correlations provide new insights on the cellular response of disrupting protein interactions and highlight the complex genetic phenotypes of drug treatment. With further refinement, our pipeline may accelerate the identification and development of novel chemical classes by screening compound-target interactions. Nicolas A. Pabon, Samuel K. Estabrooks, Zhaofeng Ye, Amanda K. Herbrand, Evelyn Süß, Ricardo M. Biondi, Victoria A. Assimon, Jason E. Gestwicki, Jeffrey L. Brodsky, Carlos J. Camacho, Ziv Bar-Joseph |
PLoS Comput. Biol. | 12 |
| 2017 | MethRaFo: MeDIP-seq methylation estimate using a Random Forest RegressorabstractMOTIVATION: Profiling of genome wide DNA methylation is now routinely performed when studying development, cancer and several other biological processes. Although Whole genome Bisulfite Sequencing provides high-quality methylation measurements at the resolution of nucleotides, it is relatively costly and so several studies have used alternative methods for such profiling. One of the most widely used low cost alternatives is MeDIP-Seq. However, MeDIP-Seq is biased for CpG enriched regions and thus its results need to be corrected in order to determine accurate methylation levels. RESULTS: Here we present a method for correcting MeDIP-Seq results based on Random Forest regression. Applying the method to real data from several different tissues (brain, cortex, penis) we show that it achieves almost 4 fold decrease in run time while increasing accuracy by as much as 20% over prior methods developed for this task. AVAILABILITY AND IMPLEMENTATION: MethRaFo is freely available as a python package (with a R wrapper) at https://github.com/phoenixding/methrafo. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ziv Bar-Joseph |
Bioinform. | 2 |
| 2017 | TASIC: determining branching models from time series single cell dataabstractMOTIVATION: Single cell RNA-Seq analysis holds great promise for elucidating the networks and pathways controlling cellular differentiation and disease. However, the analysis of time series single cell RNA-Seq data raises several new computational challenges. Cells at each time point are often sampled from a mixture of cell types, each of which may be a progenitor of one, or several, specific fates making it hard to determine which cells should be used to reconstruct temporal trajectories. In addition, cells, even from the same time point, may be unsynchronized making it hard to rely on the measured time for determining these trajectories. RESULTS: We present TASIC a new method for determining temporal trajectories, branching and cell assignments in single cell time series experiments. Unlike prior approaches TASIC uses on a probabilistic graphical model to integrate expression and time information making it more robust to noise and stochastic variations. Applying TASIC to in vitro myoblast differentiation and in-vivo lung development data we show that it accurately reconstructs developmental trajectories from single cell experiments. The reconstructed models enabled us to identify key genes involved in cell fate determination and to obtain new insights about a specific type of lung cells and its role in development. AVAILABILITY AND IMPLEMENTATION: The TASIC software package is posted in the supporting website. The datasets used in the paper are publicly available. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Sabrina Rashid, Darrell N. Kotton, Ziv Bar-Joseph |
Bioinform. | 3 |
| 2016 | Shall We Dense? Comparing Design Strategies for Time Series Expression Experiments
Emre Sefer, Ziv Bar-Joseph |
RECOMB | 2 |
| 2016 | Distributed Gradient Descent in Bacterial Food Search
Shashank Singh 0005, Sabrina Rashid, Saket Navlakha, Ziv Bar-Joseph |
RECOMB | 4 |
| 2016 | Reconstructing the temporal progression of HIV-1 immune response pathwaysabstractMOTIVATION: Most methods for reconstructing response networks from high throughput data generate static models which cannot distinguish between early and late response stages. RESULTS: We present TimePath, a new method that integrates time series and static datasets to reconstruct dynamic models of host response to stimulus. TimePath uses an Integer Programming formulation to select a subset of pathways that, together, explain the observed dynamic responses. Applying TimePath to study human response to HIV-1 led to accurate reconstruction of several known regulatory and signaling pathways and to novel mechanistic insights. We experimentally validated several of TimePaths' predictions highlighting the usefulness of temporal models. AVAILABILITY AND IMPLEMENTATION: Data, Supplementary text and the TimePath software are available from http://sb.cs.cmu.edu/timepath CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Siddhartha Jain 0001, Joel Arrais, Narasimhan J. Venkatachari, Velpandi Ayyavoo, Ziv Bar-Joseph |
Bioinform. | 5 |
| 2016 | Genome wide predictions of miRNA regulation by transcription factorsabstractMOTIVATION: Reconstructing regulatory networks from expression and interaction data is a major goal of systems biology. While much work has focused on trying to experimentally and computationally determine the set of transcription-factors (TFs) and microRNAs (miRNAs) that regulate genes in these networks, relatively little work has focused on inferring the regulation of miRNAs by TFs. Such regulation can play an important role in several biological processes including development and disease. The main challenge for predicting such interactions is the very small positive training set currently available. Another challenge is the fact that a large fraction of miRNAs are encoded within genes making it hard to determine the specific way in which they are regulated. RESULTS: To enable genome wide predictions of TF-miRNA interactions, we extended semi-supervised machine-learning approaches to integrate a large set of different types of data including sequence, expression, ChIP-seq and epigenetic data. As we show, the methods we develop achieve good performance on both a labeled test set, and when analyzing general co-expression networks. We next analyze mRNA and miRNA cancer expression data, demonstrating the advantage of using the predicted set of interactions for identifying more coherent and relevant modules, genes, and miRNAs. The complete set of predictions is available on the supporting website and can be used by any method that combines miRNAs, genes, and TFs. AVAILABILITY AND IMPLEMENTATION: Code and full set of predictions are available from the supporting website: http://cs.cmu.edu/~mruffalo/tf-mirna/ CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Matthew Ruffalo, Ziv Bar-Joseph |
Bioinform. | 2 |
| 2015 | SMARTS: reconstructing disease response networks from multiple individuals using time series gene expression dataabstractMOTIVATION: Current methods for reconstructing dynamic regulatory networks are focused on modeling a single response network using model organisms or cell lines. Unlike these models or cell lines, humans differ in their background expression profiles due to age, genetics and life factors. In addition, there are often differences in start and end times for time series human data and in the rate of progress based on the specific individual. Thus, new methods are required to integrate time series data from multiple individuals when modeling and constructing disease response networks. RESULTS: We developed Scalable Models for the Analysis of Regulation from Time Series (SMARTS), a method integrating static and time series data from multiple individuals to reconstruct condition-specific response networks in an unsupervised way. Using probabilistic graphical models, SMARTS iterates between reconstructing different regulatory networks and assigning individuals to these networks, taking into account varying individual start times and response rates. These models can be used to group different sets of patients and to identify transcription factors that differentiate the observed responses between these groups. We applied SMARTS to analyze human response to influenza and mouse brain development. In both cases, it was able to greatly improve baseline groupings while identifying key relevant TFs that differ between the groups. Several of these groupings and TFs are known to regulate the relevant processes while others represent novel hypotheses regarding immune response and development. AVAILABILITY AND IMPLEMENTATION: Software and Supplementary information are available at http://sb.cs.cmu.edu/smarts/. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Aaron Wise, Ziv Bar-Joseph |
Bioinform. | 2 |
| 2015 | MassExodus: modeling evolving networks in harsh environments
Saket Navlakha, Christos Faloutsos, Ziv Bar-Joseph |
Data Min. Knowl. Discov. | 3 |
| 2015 | Decreasing-Rate Pruning Optimizes the Construction of Efficient and Robust Distributed NetworksabstractRobust, efficient, and low-cost networks are advantageous in both biological and engineered systems. During neural network development in the brain, synapses are massively over-produced and then pruned-back over time. This strategy is not commonly used when designing engineered networks, since adding connections that will soon be removed is considered wasteful. Here, we show that for large distributed routing networks, network function is markedly enhanced by hyper-connectivity followed by aggressive pruning and that the global rate of pruning, a developmental parameter not previously studied by experimentalists, plays a critical role in optimizing network structure. We first used high-throughput image analysis techniques to quantify the rate of pruning in the mammalian neocortex across a broad developmental time window and found that the rate is decreasing over time. Based on these results, we analyzed a model of computational routing networks and show using both theoretical analysis and simulations that decreasing rates lead to more robust and efficient networks compared to other rates. We also present an application of this strategy to improve the distributed design of airline networks. Thus, inspiration from neural network formation suggests effective ways to design distributed networks across several domains. Saket Navlakha, Alison L. Barth, Ziv Bar-Joseph |
PLoS Comput. Biol. | 3 |
| 2014 | Multitask Learning of Signaling and Regulatory Networks with Application to Studying Human Response to FluabstractReconstructing regulatory and signaling response networks is one of the major goals of systems biology. While several successful methods have been suggested for this task, some integrating large and diverse datasets, these methods have so far been applied to reconstruct a single response network at a time, even when studying and modeling related conditions. To improve network reconstruction we developed MT-SDREM, a multi-task learning method which jointly models networks for several related conditions. In MT-SDREM, parameters are jointly constrained across the networks while still allowing for condition-specific pathways and regulation. We formulate the multi-task learning problem and discuss methods for optimizing the joint target function. We applied MT-SDREM to reconstruct dynamic human response networks for three flu strains: H1N1, H5N1 and H3N2. Our multi-task learning method was able to identify known and novel factors and genes, improving upon prior methods that model each condition independently. The MT-SDREM networks were also better at identifying proteins whose removal affects viral load indicating that joint learning can still lead to accurate, condition-specific, networks. Supporting website with MT-SDREM implementation: http://sb.cs.cmu.edu/mtsdrem. Siddhartha Jain 0001, Anthony Gitter, Ziv Bar-Joseph |
PLoS Comput. Biol. | 3 |
| 2013 | Identifying proteins controlling key disease signaling pathwaysabstractMOTIVATION: Several types of studies, including genome-wide association studies and RNA interference screens, strive to link genes to diseases. Although these approaches have had some success, genetic variants are often only present in a small subset of the population, and screens are noisy with low overlap between experiments in different labs. Neither provides a mechanistic model explaining how identified genes impact the disease of interest or the dynamics of the pathways those genes regulate. Such mechanistic models could be used to accurately predict downstream effects of knocking down pathway members and allow comprehensive exploration of the effects of targeting pairs or higher-order combinations of genes. RESULTS: We developed methods to model the activation of signaling and dynamic regulatory networks involved in disease progression. Our model, SDREM, integrates static and time series data to link proteins and the pathways they regulate in these networks. SDREM uses prior information about proteins' likelihood of involvement in a disease (e.g. from screens) to improve the quality of the predicted signaling pathways. We used our algorithms to study the human immune response to H1N1 influenza infection. The resulting networks correctly identified many of the known pathways and transcriptional regulators of this disease. Furthermore, they accurately predict RNA interference effects and can be used to infer genetic interactions, greatly improving over other methods suggested for this task. Applying our method to the more pathogenic H5N1 influenza allowed us to identify several strain-specific targets of this infection. AVAILABILITY: SDREM is available from http://sb.cs.cmu.edu/sdrem. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Anthony Gitter, Ziv Bar-Joseph |
Bioinform. | 2 |
| 2013 | Integrating sequence, expression and interaction data to determine condition-specific miRNA regulationabstractMOTIVATION: MicroRNAs (miRNAs) are small non-coding RNAs that regulate gene expression post-transcriptionally. MiRNAs were shown to play an important role in development and disease, and accurately determining the networks regulated by these miRNAs in a specific condition is of great interest. Early work on miRNA target prediction has focused on using static sequence information. More recently, researchers have combined sequence and expression data to identify such targets in various conditions. RESULTS: We developed the Protein Interaction-based MicroRNA Modules (PIMiM), a regression-based probabilistic method that integrates sequence, expression and interaction data to identify modules of mRNAs controlled by small sets of miRNAs. We formulate an optimization problem and develop a learning framework to determine the module regulation and membership. Applying PIMiM to cancer data, we show that by adding protein interaction data and modeling cooperative regulation of mRNAs by a small number of miRNAs, PIMiM can accurately identify both miRNA and their targets improving on previous methods. We next used PIMiM to jointly analyze a number of different types of cancers and identified both common and cancer-type-specific miRNA regulators. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Hai-Son Le, Ziv Bar-Joseph |
Bioinform. | 2 |
| 2013 | A high-throughput framework to detect synapses in electron microscopy imagesabstractMOTIVATION: Synaptic connections underlie learning and memory in the brain and are dynamically formed and eliminated during development and in response to stimuli. Quantifying changes in overall density and strength of synapses is an important pre-requisite for studying connectivity and plasticity in these cases or in diseased conditions. Unfortunately, most techniques to detect such changes are either low-throughput (e.g. electrophysiology), prone to error and difficult to automate (e.g. standard electron microscopy) or too coarse (e.g. magnetic resonance imaging) to provide accurate and large-scale measurements. RESULTS: To facilitate high-throughput analyses, we used a 50-year-old experimental technique to selectively stain for synapses in electron microscopy images, and we developed a machine-learning framework to automatically detect synapses in these images. To validate our method, we experimentally imaged brain tissue of the somatosensory cortex in six mice. We detected thousands of synapses in these images and demonstrate the accuracy of our approach using cross-validation with manually labeled data and by comparing against existing algorithms and against tools that process standard electron microscopy images. We also used a semi-supervised algorithm that leverages unlabeled data to overcome sample heterogeneity and improve performance. Our algorithms are highly efficient and scalable and are freely available for others to use. AVAILABILITY: Code is available at http://www.cs.cmu.edu/∼saketn/detect_synapses/ Saket Navlakha, Joseph Suhan, Alison L. Barth, Ziv Bar-Joseph |
Bioinform. | 4 |
| 2013 | Beeping a maximal independent set
Yehuda Afek, Noga Alon, Ziv Bar-Joseph, Alejandro Cornejo, Bernhard Haeupler, Fabian Kuhn |
Distributed Comput. | 3 |
| 2012 | Matching experiments across species using expression values and textual informationabstractMOTIVATION: With the vast increase in the number of gene expression datasets deposited in public databases, novel techniques are required to analyze and mine this wealth of data. Similar to the way BLAST enables cross-species comparison of sequence data, tools that enable cross-species expression comparison will allow us to better utilize these datasets: cross-species expression comparison enables us to address questions in evolution and development, and further allows the identification of disease-related genes and pathways that play similar roles in humans and model organisms. Unlike sequence, which is static, expression data changes over time and under different conditions. Thus, a prerequisite for performing cross-species analysis is the ability to match experiments across species. RESULTS: To enable better cross-species comparisons, we developed methods for automatically identifying pairs of similar expression datasets across species. Our method uses a co-training algorithm to combine a model of expression similarity with a model of the text which accompanies the expression experiments. The co-training method outperforms previous methods based on expression similarity alone. Using expert analysis, we show that the new matches identified by our method indeed capture biological similarities across species. We then use the matched expression pairs between human and mouse to recover known and novel cycling genes as well as to identify genes with possible involvement in diabetes. By providing the ability to identify novel candidate genes in model organisms, our method opens the door to new models for studying diseases. AVAILABILITY: Source code and supplementary information is available at: www.andrew.cmu.edu/user/aaronwis/cotrain12. Aaron Wise, Zoltán N. Oltvai, Ziv Bar-Joseph |
Bioinform. | 3 |
| 2012 | A Network-based Approach for Predicting Missing Pathway InteractionsabstractEmbedded within large-scale protein interaction networks are signaling pathways that encode response cascades in the cell. Unfortunately, even for well-studied species like S. cerevisiae, only a fraction of all true protein interactions are known, which makes it difficult to reason about the exact flow of signals and the corresponding causal relations in the network. To help address this problem, we introduce a framework for predicting new interactions that aid connectivity between upstream proteins (sources) and downstream transcription factors (targets) of a particular pathway. Our algorithms attempt to globally minimize the distance between sources and targets by finding a small set of shortcut edges to add to the network. Unlike existing algorithms for predicting general protein interactions, by focusing on proteins involved in specific responses our approach homes-in on pathway-consistent interactions. We applied our method to extend pathways in osmotic stress response in yeast and identified several missing interactions, some of which are supported by published reports. We also performed experiments that support a novel interaction not previously reported. Our framework is general and may be applicable to edge prediction problems in other domains. Saket Navlakha, Anthony Gitter, Ziv Bar-Joseph |
PLoS Comput. Biol. | 3 |
| 2011 | Inferring Interaction Networks using the IBP applied to microRNA Target PredictionabstractDetermining interactions between entities and the overall organization and clustering of nodes in networks is a major challenge when analyzing biological and social network data. Here we extend the Indian Buffet Process (IBP), a nonparametric Bayesian model, to integrate noisy interaction scores with properties of individual entities for inferring interaction networks and clustering nodes within these networks. We present an application of this method to study how microRNAs regulate mRNAs in cells. Analysis of synthetic and real data indicates that the method improves upon prior methods, correctly recovers interactions and clusters, and provides accurate biological predictions. Hai-Son Le, Ziv Bar-Joseph |
NIPS | 2 |
| 2011 | Learning Cellular Sorting Pathways Using Protein Interactions and Sequence Motifs
Tien-ho Lin, Ziv Bar-Joseph, Robert F. Murphy |
RECOMB | 2 |
| 2011 | Beeping a Maximal Independent Set
Yehuda Afek, Noga Alon, Ziv Bar-Joseph, Alejandro Cornejo, Bernhard Haeupler, Fabian Kuhn |
DISC | 3 |
| 2011 | DECOD: fast and accurate discriminative DNA motif findingabstractMOTIVATION: Motif discovery is now routinely used in high-throughput studies including large-scale sequencing and proteomics. These datasets present new challenges. The first is speed. Many motif discovery methods do not scale well to large datasets. Another issue is identifying discriminative rather than generative motifs. Such discriminative motifs are important for identifying co-factors and for explaining changes in behavior between different conditions. RESULTS: To address these issues we developed a method for DECOnvolved Discriminative motif discovery (DECOD). DECOD uses a k-mer count table and so its running time is independent of the size of the input set. By deconvolving the k-mers DECOD considers context information without using the sequences directly. DECOD outperforms previous methods both in speed and in accuracy when using simulated and real biological benchmark data. We performed new binding experiments for p53 mutants and used DECOD to identify p53 co-factors, suggesting new mechanisms for p53 activation. AVAILABILITY: The source code and binaries for DECOD are available at http://www.sb.cs.cmu.edu/DECOD CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Peter Huggins, Idit Shiff, Rachel Beckerman, Oleg Laptenko, Carol Prives, Marcel H. Schulz, Itamar Simon, Ziv Bar-Joseph |
Bioinform. | 9 |
| 2011 | Discriminative Motif Finding for Predicting Protein Subcellular LocalizationabstractMany methods have been described to predict the subcellular location of proteins from sequence information. However, most of these methods either rely on global sequence properties or use a set of known protein targeting motifs to predict protein localization. Here, we develop and test a novel method that identifies potential targeting motifs using a discriminative approach based on hidden Markov models (discriminative HMMs). These models search for motifs that are present in a compartment but absent in other, nearby, compartments by utilizing an hierarchical structure that mimics the protein sorting mechanism. We show that both discriminative motif finding and the hierarchical structure improve localization prediction on a benchmark data set of yeast proteins. The motifs identified can be mapped to known targeting motifs and they are more conserved than the average protein sequence. Using our motif-based predictions, we can identify potential annotation errors in public databases for the location of some of the proteins. A software implementation and the data set described in this paper are available from http://murphylab.web.cmu.edu/software/2009_TCBB_motif/. Tien-ho Lin, Robert F. Murphy, Ziv Bar-Joseph |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2010 | Cross Species Expression Analysis using a Dirichlet Process Mixture Model with Latent MatchingsabstractRecent studies compare gene expression data across species to identify core and species specific genes in biological systems. To perform such comparisons researchers need to match genes across species. This is a challenging task since the correct matches (orthologs) are not known for most genes. Previous work in this area used deterministic matchings or reduced multidimensional expression data to binary representation. Here we develop a new method that can utilize soft matches (given as priors) to infer both, unique and similar expression patterns across species and a matching for the genes in both species. Our method uses a Dirichlet process mixture model which includes a latent data matching variable. We present learning and inference algorithms based on variational methods for this model. Applying our method to immune response data we show that it can accurately identify common and unique response patterns by improving the matchings between human and mouse genes. Hai-Son Le, Ziv Bar-Joseph |
NIPS | 2 |
| 2010 | Cross-species queries of large gene expression databasesabstractMOTIVATION: Expression databases, including the Gene Expression Omnibus and ArrayExpress, have experienced significant growth over the past decade and now hold hundreds of thousands of arrays from multiple species. Since most drugs are initially tested on model organisms, the ability to compare expression experiments across species may help identify pathways that are activated in a similar way in humans and other organisms. However, while several methods exist for finding co-expressed genes in the same species as a query gene, looking at co-expression of homologs or arbitrary genes in other species is challenging. Unlike sequence, which is static, expression is dynamic and changes between tissues, conditions and time. Thus, to carry out cross-species analysis using these databases, we need methods that can match experiments in one species with experiments in another species. RESULTS: To facilitate queries in large databases, we developed a new method for comparing expression experiments from different species. We define a distance metric between the ranking of orthologous genes in the two species. We show how to solve an optimization problem for learning the parameters of this function using a training dataset of known similar expression experiments pairs. The function we learn outperforms previous methods and simpler rank comparison methods that have been used in the past for single species analysis. We used our method to compare millions of array pairs from mouse and human expression experiments. The resulting matches can be used to find functionally related genes, to hypothesize about biological response mechanisms and to highlight conditions and diseases that are activating similar pathways in both species. AVAILABILITY: Supporting methods, results and a Matlab implementation are available from http://sb.cs.cmu.edu/ExpQ/. Hai-Son Le, Zoltán N. Oltvai, Ziv Bar-Joseph |
Bioinform. | 3 |
| 2009 | Cross Species Expression Analysis of Innate Immune Response
Ronald Rosenfeld, Gerard J. Nau, Ziv Bar-Joseph |
RECOMB | 4 |
| 2009 | Cross species analysis of microarray expression dataabstractMOTIVATION: Many biological systems operate in a similar manner across a large number of species or conditions. Cross-species analysis of sequence and interaction data is often applied to determine the function of new genes. In contrast to these static measurements, microarrays measure the dynamic, condition-specific response of complex biological systems. The recent exponential growth in microarray expression datasets allows researchers to combine expression experiments from multiple species to identify genes that are not only conserved in sequence but also operated in a similar way in the different species studied. RESULTS: In this review we discuss the computational and technical challenges associated with these studies, the approaches that have been developed to address these challenges and the advantages of cross-species analysis of microarray data. We show how successful application of these methods lead to insights that cannot be obtained when analyzing data from a single species. We also highlight current open problems and discuss possible ways to address them. Peter Huggins, Ziv Bar-Joseph |
Bioinform. | 3 |
| 2009 | Modeling spatial and temporal variation in motion dataabstractWe present a novel method to model and synthesize variation in motion data. Given a few examples of a particular type of motion as input, we learn a generative model that is able to synthesize a family of spatial and temporal variants that are statistically similar to the input examples. The new variants retain the features of the original examples, but are not exact copies of them. We learn a Dynamic Bayesian Network model from the input examples that enables us to capture properties of conditional independence in the data, and model it using a multivariate probability distribution. We present results for a variety of human motion, and 2D handwritten characters. We perform a user study to show that our new variants are less repetitive than typical game and crowd simulation approaches of re-playing a small number of existing motion clips. Our technique can synthesize new variants efficiently and has a small memory requirement. Manfred Lau, Ziv Bar-Joseph, James J. Kuffner |
ACM Trans. Graph. | 2 |
| 2008 | Alignment and classification of time series gene expression in clinical studiesabstractMOTIVATION: Classification of tissues using static gene-expression data has received considerable attention. Recently, a growing number of expression datasets are measured as a time series. Methods that are specifically designed for this temporal data can both utilize its unique features (temporal evolution of profiles) and address its unique challenges (different response rates of patients in the same class). RESULTS: We present a method that utilizes hidden Markov models (HMMs) for the classification task. We use HMMs with less states than time points leading to an alignment of the different patient response rates. To focus on the differences between the two classes we develop a discriminative HMM classifier. Unlike the traditional generative HMM, discriminative HMM can use examples from both classes when learning the model for a specific class. We have tested our method on both simulated and real time series expression data. As we show, our method improves upon prior methods and can suggest markers for specific disease and response stages that are not found when using traditional classifiers. AVAILABILITY: Matlab implementation is available from http://www.cs.cmu.edu/~thlin/tram/. Tien-ho Lin, Naftali Kaminski, Ziv Bar-Joseph |
ISMB | 3 |
| 2008 | Protein complex identification by supervised graph local clusteringabstractMOTIVATION: Protein complexes integrate multiple gene products to coordinate many biological functions. Given a graph representing pairwise protein interaction data one can search for subgraphs representing protein complexes. Previous methods for performing such search relied on the assumption that complexes form a clique in that graph. While this assumption is true for some complexes, it does not hold for many others. New algorithms are required in order to recover complexes with other types of topological structure. RESULTS: We present an algorithm for inferring protein complexes from weighted interaction graphs. By using graph topological patterns and biological properties as features, we model each complex subgraph by a probabilistic Bayesian network (BN). We use a training set of known complexes to learn the parameters of this BN model. The log-likelihood ratio derived from the BN is then used to score subgraphs in the protein interaction graph and identify new complexes. We applied our method to protein interaction data in yeast. As we show our algorithm achieved a considerable improvement over clique based algorithms in terms of its ability to recover known complexes. We discuss some of the new complexes predicted by our algorithm and determine that they likely represent true complexes. AVAILABILITY: Matlab implementation is available on the supporting website: www.cs.cmu.edu/~qyj/SuperComplex. Yanjun Qi, Fernanda Balem, Christos Faloutsos, Judith Klein-Seetharaman, Ziv Bar-Joseph |
ISMB | 5 |
| 2008 | A Combined Expression-Interaction Model for Inferring the Temporal Activity of Transcription Factors
Yanxin Shi, Itamar Simon, Tom M. Mitchell, Ziv Bar-Joseph |
RECOMB | 4 |
| 2008 | A Semi-Supervised Method for Predicting Transcription Factor-Gene Interactions in Escherichia coliabstractWhile Escherichia coli has one of the most comprehensive datasets of experimentally verified transcriptional regulatory interactions of any organism, it is still far from complete. This presents a problem when trying to combine gene expression and regulatory interactions to model transcriptional regulatory networks. Using the available regulatory interactions to predict new interactions may lead to better coverage and more accurate models. Here, we develop SEREND (SEmi-supervised REgulatory Network Discoverer), a semi-supervised learning method that uses a curated database of verified transcriptional factor-gene interactions, DNA sequence binding motifs, and a compendium of gene expression data in order to make thousands of new predictions about transcription factor-gene interactions, including whether the transcription factor activates or represses the gene. Using genome-wide binding datasets for several transcription factors, we demonstrate that our semi-supervised classification strategy improves the prediction of targets for a given transcription factor. To further demonstrate the utility of our inferred interactions, we generated a new microarray gene expression dataset for the aerobic to anaerobic shift response in E. coli. We used our inferred interactions with the verified interactions to reconstruct a dynamic regulatory network for this response. The network reconstructed when using our inferred interactions was better able to correctly identify known regulators and suggested additional activators and repressors as having important roles during the aerobic-anaerobic shift interface. Jason Ernst, Qasim K. Beg, Krin A. Kay, Gábor Balázsi, Zoltán N. Oltvai, Ziv Bar-Joseph |
PLoS Comput. Biol. | 6 |
| 2008 | Extracting Dynamics from Static Cancer Expression DataabstractStatic expression experiments analyze samples from many individuals. These samples are often snapshots of the progression of a certain disease such as cancer. This raises an intriguing question: Can we determine a temporal order for these samples? Such an ordering can lead to better understanding of the dynamics of the disease and to the identification of genes associated with its progression. In this paper we formally prove, for the first time, that under a model for the dynamics of the expression levels of a single gene, it is indeed possible to recover the correct ordering of the static expression datasets by solving an instance of the traveling salesman problem (TSP). In addition, we devise an algorithm that combines a TSP heuristic and probabilistic modeling for inferring the underlying temporal order of the microarray experiments. This algorithm constructs probabilistic continuous curves to represent expression profiles leading to accurate temporal reconstruction for human data. Applying our method to cancer expression data we show that the ordering derived agrees well with survival duration. A classifier that utilizes this ordering improves upon other classifiers suggested for this task. The set of genes displaying consistent behavior for the determined ordering are enriched for genes associated with cancer progression. Anupam Gupta 0001, Ziv Bar-Joseph |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2007 | Inferring pairwise regulatory relationships from multiple time series datasetsabstractMOTIVATION: Time series expression experiments have emerged as a popular method for studying a wide range of biological systems under a variety of conditions. One advantage of such data is the ability to infer regulatory relationships using time lag analysis. However, such analysis in a single experiment may result in many false positives due to the small number of time points and the large number of genes. Extending these methods to simultaneously analyze several time series datasets is challenging since under different experimental conditions biological systems may behave faster or slower making it hard to rely on the actual duration of the experiment. RESULTS: We present a new computational model and an associated algorithm to address the problem of inferring time-lagged regulatory relationships from multiple time series expression experiments with varying (unknown) time-scales. Our proposed algorithm uses a set of known interacting pairs to compute a temporal transformation between every two datasets. Using this temporal transformation we search for new interacting pairs. As we show, our method achieves a much lower false-positive rate compared to previous methods that use time series expression data for pairwise regulatory relationship discovery. Some of the new predictions made by our method can be verified using other high throughput data sources and functional annotation databases. AVAILABILITY: Matlab implementation is available from the supporting website: http://www.cs.cmu.edu/~yanxins/regulation_inference/index.html. Yanxin Shi, Tom M. Mitchell, Ziv Bar-Joseph |
Bioinform. | 3 |
| 2007 | A mixture of feature experts approach for protein-protein interaction predictionabstractBACKGROUND: High-throughput methods can directly detect the set of interacting proteins in model species but the results are often incomplete and exhibit high false positive and false negative rates. A number of researchers have recently presented methods for integrating direct and indirect data for predicting interactions. These methods utilize a common classifier for all pairs. However, due to missing data and high redundancy among the features used, different protein pairs may benefit from different features based on the set of attributes available. In addition, in many cases it is hard to directly determine which of the data sources contributed to a prediction. This information is important for biologists using these predications in the design of new experiments. RESULTS: To address these challenges we propose a Mixture-of-Feature-Experts method for protein-protein interaction prediction. We split the features into roughly homogeneous sets of feature experts. The individual experts use logistic regression and their scores are combined using another logistic regression. When combining the scores the weighting of each expert depends on the set of input attributes available for that pair. Thus, different experts will have different influence on the prediction depending on the available features. CONCLUSION: We applied our method to predict the set of interacting proteins in yeast and human cells. Our method improved upon the best previous methods for this task. In addition, the weighting of the experts provides means to evaluate the prediction based on the high scoring features. Yanjun Qi, Judith Klein-Seetharaman, Ziv Bar-Joseph |
BMC Bioinform. | 3 |
| 2006 | A Patient-Gene Model for Temporal Expression Profiles in Clinical Studies
Naftali Kaminski, Ziv Bar-Joseph |
RECOMB | 2 |
| 2006 | STEM: a tool for the analysis of short time series gene expression dataabstractBACKGROUND: Time series microarray experiments are widely used to study dynamical biological processes. Due to the cost of microarray experiments, and also in some cases the limited availability of biological material, about 80% of microarray time series experiments are short (3-8 time points). Previously short time series gene expression data has been mainly analyzed using more general gene expression analysis tools not designed for the unique challenges and opportunities inherent in short time series gene expression data. RESULTS: We introduce the Short Time-series Expression Miner (STEM) the first software program specifically designed for the analysis of short time series microarray gene expression data. STEM implements unique methods to cluster, compare, and visualize such data. STEM also supports efficient and statistically rigorous biological interpretations of short time series data through its integration with the Gene Ontology. CONCLUSION: The unique algorithms STEM implements to cluster and compare short time series gene expression data combined with its visualization capabilities and integration with the Gene Ontology should make STEM useful in the analysis of data from a significant portion of all microarray studies. STEM is available for download for free to academic and non-profit users at http://www.cs.cmu.edu/~jernst/stem. Jason Ernst, Ziv Bar-Joseph |
BMC Bioinform. | 2 |
| 2005 | Active learning for sampling in time-series experiments with application to gene expression analysisabstractMany time-series experiments seek to estimate some signal as a continuous function of time. In this paper, we address the sampling problem for such experiments: determining which time-points ought to be sampled in order to minimize the cost of data collection. We restrict our attention to a growing class of experiments which measure multiple signals at each time-point and where raw materials/observations are archived initially, and selectively analyzed later, this analysis being the more expensive step. We present an active learning algorithm for iteratively choosing time-points to sample, using the uncertainty in the quality of the currently estimated time-dependent curve as the objective function. Using simulated data as well as gene expression data, we show that our algorithm performs well, and can significantly reduce experimental cost without loss of information. Rohit Singh 0001, Nathan P. Palmer, David K. Gifford, Bonnie Berger, Ziv Bar-Joseph |
ICML | 5 |
| 2004 | Analyzing time series gene expression dataabstractMOTIVATION: Time series expression experiments are an increasingly popular method for studying a wide range of biological systems. However, when analyzing these experiments researchers face many new computational challenges. Algorithms that are specifically designed for time series experiments are required so that we can take advantage of their unique features (such as the ability to infer causality from the temporal response pattern) and address the unique problems they raise (e.g. handling the different non-uniform sampling rates). RESULTS: We present a comprehensive review of the current research in time series expression data analysis. We divide the computational challenges into four analysis levels: experimental design, data analysis, pattern recognition and networks. For each of these levels, we discuss computational and biological problems at that level and point out some of the methods that have been proposed to deal with these issues. Many open problems in all these levels are discussed. This review is intended to serve as both, a point of reference for experimental biologists looking for practical solutions for analyzing their data, and a starting point for computer scientists interested in working on the computational problems related to time series expression analysis. Ziv Bar-Joseph |
Bioinform. | 1 |
| 2003 | K-ary Clustering with Optimal Leaf Ordering for Gene Expression DataabstractMOTIVATION: A major challenge in gene expression analysis is effective data organization and visualization. One of the most popular tools for this task is hierarchical clustering. Hierarchical clustering allows a user to view relationships in scales ranging from single genes to large sets of genes, while at the same time providing a global view of the expression data. However, hierarchical clustering is very sensitive to noise, it usually lacks of a method to actually identify distinct clusters, and produces a large number of possible leaf orderings of the hierarchical clustering tree. In this paper we propose a new hierarchical clustering algorithm which reduces susceptibility to noise, permits up to k siblings to be directly related, and provides a single optimal order for the resulting tree. RESULTS: We present an algorithm that efficiently constructs a k-ary tree, where each node can have up to k children, and then optimally orders the leaves of that tree. By combining k clusters at each step our algorithm becomes more robust against noise and missing values. By optimally ordering the leaves of the resulting tree we maintain the pairwise relationships that appear in the original method, without sacrificing the robustness. Our k-ary construction algorithm runs in O(n(3)) regardless of k and our ordering algorithm runs in O(4(k)n(3)). We present several examples that show that our k-ary clustering algorithm achieves results that are superior to the binary tree results in both global presentation and cluster identification. AVAILABILITY: We have implemented the above algorithms in C++ on the Linux operating system. Ziv Bar-Joseph, Erik D. Demaine, David K. Gifford, Nathan Srebro, Angèle M. Foley, Tommi S. Jaakkola |
Bioinform. | 1 |
| 2003 | Hierarchical Context-based Pixel OrderingabstractAbstract We present a context‐based scanning algorithm which reorders the input image using a hierarchical representationof the image. Our algorithm optimally orders (permutes) the leaves corresponding to the pixels, by minimizing thesum of distances between neighboring pixels. The reordering results in an improved autocorrelation betweennearby pixels which leads to a smoother image. This allows us, for the first time, to improve image compressionrates using context‐based scans. The results presented in this paper greatly improve upon previous work in bothcompression rate and running time. Categories and Subject Descriptors (according to ACM CCS): I.3.5 [Computer Graphics]: Computational Geometryand Object Modeling I.3.6 [Computer Graphics]: Methodology and Techniques Ziv Bar-Joseph, Daniel Cohen-Or |
Comput. Graph. Forum | 1 |
| 2002 | A new approach to analyzing gene expression time series dataabstractWe present algorithms for time-series gene expression analysis that permit the principled estimation of unobserved time-points, clustering, and dataset alignment. Each expression profile is modeled as a cubic spline (piecewise polynomial) that is estimated from the observed data and every time point influences the overall smooth expression curve. We constrain the spline coefficients of genes in the same class to have similar expression patterns, while also allowing for gene specific parameters. We show that unobserved time-points can be reconstructed using our method with 10-15% less error when compared to previous best methods. Our clustering algorithm operates directly on the continuous representations of gene expression profiles, and we demonstrate that this is particularly effective when applied to non-uniformly sampled data. Our continuous alignment algorithm also avoids difficulties encountered by discrete approaches. In particular, our method allows for control of the number of degrees of freedom of the warp through the specification of parameterized functions, which helps to avoid overfitting. We demonstrate that our algorithm produces stable low-error alignments on real expression data and further show a specific application to yeast knockout data that produces biologically meaningful results. Ziv Bar-Joseph, Georg K. Gerber, David K. Gifford, Tommi S. Jaakkola, Itamar Simon |
RECOMB | 1 |
| 2002 | K-ary Clustering with Optimal Leaf Ordering for Gene Expression Data
Ziv Bar-Joseph, Erik D. Demaine, David K. Gifford, Angèle M. Foley, Tommi S. Jaakkola, Nathan Srebro |
WABI | 1 |
| 2002 | Early-Delivery Dynamic Atomic Broadcast
Ziv Bar-Joseph, Idit Keidar, Nancy A. Lynch |
DISC | 1 |
| 2001 | Texture Mixing and Texture Movie Synthesis Using Statistical LearningabstractWe present an algorithm based on statistical learning for synthesizing static and time-varying textures matching the appearance of an input texture. Our algorithm is general and automatic and it works well on various types of textures, including 1D sound textures, 2D texture images, and 3D texture movies. The same method is also used to generate 2D texture mixtures that simultaneously capture the appearance of a number of different input textures. In our approach, input textures are treated as sample signals generated by a stochastic process. We first construct a tree representing a hierarchical multiscale transform of the signal using wavelets. From this tree, new random trees are generated by learning and sampling the conditional probabilities of the paths in the original tree. Transformation of these random trees back into signals results in new random textures. In the case of 2D texture synthesis, our algorithm produces results that are generally as good as or better than those produced by previously described methods in this field. For texture mixtures, our results are better and more general than those produced by earlier methods. For texture movies, we present the first algorithm that is able to automatically generate movie clips of dynamic phenomena such as waterfalls, fire flames, a school of jellyfish, a crowd of people, etc. Our results indicate that the proposed technique is effective and robust. Ziv Bar-Joseph, Ran El-Yaniv, Dani Lischinski, Michael Werman |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2000 | Totally Ordered Multicast with Bounded Delays and Variable Rates
Ziv Bar-Joseph, Idit Keidar, Tal Anker, Nancy A. Lynch |
OPODIS | 1 |
| 1998 | A Tight Lower Bound for Randomized Synchronous ConsensusabstractWe prove tight upper and lower bounds of @(t/J-) on the expected number of rounds needed for randomized synchronous consensus protocols for a fail-stop, full information, dynamic adversary.In particular this proves that some restrictions are needed on the power of the adversary to allow randomized constant expected number of rounds protocols. Ziv Bar-Joseph, Michael Ben-Or |
PODC | 1 |