EDBT 2026 Demo / reviewers in the wild / expert
Jing Zhang 0062
dblp:05/3499-62
· DBLP profile ↗
24ranked-venue papers
2as first author
18since 2021 · last 2025
0000-0002-5970-0509ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 14 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 6 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MUSE: A Multi-slice Joint Analysis Method for Spatial Transcriptomics ExperimentsabstractRecent advances in spatial transcriptomics (ST) and cost reductions have enabled large-scale multi-slice ST data generation, enhancing the statistical power to detect subtle biological signals. However, cross-slice inconsistencies and data quality variability present significant analytical challenges. To overcome these limitations, we developed MUSE, a computational framework designed for multislice joint embedding, spatial domain identification, and gene expression imputation. Specifically, MUSE integrates a two-module architecture to ensure robust cross-slice alignment and data harmonization. The alignment module models each slice as a graph and employs optimal transport to align cells across slices while preserving spatial continuity. The optimization module further refines integration by incorporating an alignment loss, allowing lower-quality data to leverage structural information from higher-quality slices. Additionally, MUSE generates virtual neighbors from aligned cells, enriching contextual information and mitigating data sparsity. These design principles enable seamless integration with existing single-slice methods, extending their applicability to multi-slice ST analysis. To comprehensively evaluate its performance, we applied MUSE to 12 real and 48 simulated datasets spanning a range of data qualities. Across all metrics, MUSE consistently outperformed existing methods in cross-slice consistency, spatial domain identification, and gene expression imputation. To promote accessibility and adoption, we provide MUSE as an open-source software package. As multi-slice ST datasets become increasingly prevalent, MUSE provides a robust and extensible framework designed to effectively integrate growing numbers of slices, thereby advancing the analysis of tissue architectures and spatial gene expression in complex biological systems. Ziheng Duan, Zhiqing Xiao, Rex Ying, Jing Zhang 0062 |
CIKM | 5 |
| 2025 | DISCO: A Diffusion Model For Spatial Transcriptomics Data CompletionabstractSpatial transcriptomics enables the study of gene expression within the spatial context of tissues, offering valuable insights into tissue organization and function. However, technical limitations can result in large missing regions of data, which hinder accurate downstream analyses and biological interpretation. To address these challenges, we propose DISCO (DIffusion model for Spatial transcriptomics data COmpletion), a framework with three key features. First, DISCO employs a graph neural network-based region encoder to integrate spatial and gene expression information from observed regions, generating latent representations that guide the prediction of missing regions. Second, it uses two diffusion modules: a position diffusion module to predict the spatial layout of missing regions, and a gene expression diffusion module to generate gene expression profiles conditioned on the predicted coordinates. Third, DISCO incorporates neighboring region information during inference to guide the denoising process, ensuring smooth transitions and biologically coherent results. We validate DISCO across multiple sequencing platforms, species, and datasets, demonstrating its effectiveness in reconstructing large missing regions. DISCO is implemented as open-source software, providing researchers with a powerful tool to enhance data completeness and advance spatial transcriptomics research. Ziheng Duan, Zhuoyang Zhang, James Song, Jing Zhang 0062 |
ICIP | 5 |
| 2025 | Identifying Combinatorial Regulatory Genes for Cell Fate Decision via Reparameterizable Subset ExplanationsabstractCell fate decisions are highly coordinated processes governed by complex interactions among numerous regulatory genes, while disruptions in these mechanisms can lead to developmental abnormalities and disease. Traditional methods often fail to capture such combinatorial interactions, limiting their ability to fully model cell fate dynamics. Here, we introduce MetaVelo, a global feature explanation framework for identifying key regulatory gene sets influencing cell fate transitions. MetaVelo models these transitions as a black-box function and employs a differentiable neural ordinary differential equation (ODE) surrogate to enable efficient optimization. By reparameterizing the problem as a controllable data generation process, MetaVelo overcomes the challenges posed by the non-differentiable nature of cell fate dynamics. Benchmarking across diverse stand-alone and longitudinal single-cell RNA-seq datasets and three black-box cell fate models demonstrates its superiority over 12 baseline methods in predicting developmental trajectories and identifying combinatorial regulatory gene sets. MetaVelo further distinguishes independent from synergistic regulatory genes, offering novel insights into the gene interactions governing cell fate. With the growing availability of high-resolution single-cell data, MetaVelo provides a scalable and effective framework for advancing developmental biology and therapeutic applications. Junhao Liu 0001, Martin Renqiang Min, Jing Zhang 0062 |
KDD (2) | 4 |
| 2024 | Understanding Transcriptional Regulatory Redundancy by Learnable Global Subset Perturbations
Junhao Liu 0001, Siwei Xu, Dylan Riffle, Ziheng Duan, Martin Renqiang Min, Jing Zhang 0062 |
ACML | 6 |
| 2024 | iHAST: Integrating Hybrid Attention for Super-Resolution in Spatial Transcriptomics
Jing Zhang 0062, Ziheng Duan, Siwei Xu |
BMVC | 2 |
| 2024 | iMIRACLE: An Iterative Multi-View Graph Neural Network to Model Intercellular Gene Regulation From Spatial Transcriptomic DataabstractSpatial transcriptomics has transformed genomic research by measuring spatially resolved gene expressions, allowing us to investigate how cells adapt to their microenvironment via modulating their expressed genes. This essential process usually starts from cell-cell communication (CCC) via ligand-receptor (LR) interaction, leading to regulatory changes within the receiver cell. However, few methods were developed to connect them to provide biological insights into intercellular regulation. To fill this gap, we propose iMiracle, an iterative multi-view graph neural network that models each cell's intercellular regulation with three key features. Firstly, iMiracle integrates inter- and intra-cellular networks to jointly estimate cell-type- and micro-environment-driven gene expressions. Optionally, it allows prior knowledge of intra-cellular networks as pre-structured masks to maintain biological relevance. Secondly, iMiracle employs iterative learning to overcome the sparsity of spatial transcriptomic data and gradually fill in the missing edges in the CCC network. Thirdly, iMiracle infers a cell-specific ligand-gene regulatory score based on the contributions of different LR pairs to interpret inter-cellular regulation. We applied iMiracle to nine simulated and eight real datasets from three sequencing platforms and demonstrated that iMiracle consistently outperformed ten methods in gene expression imputation and four methods in regulatory score inference. Lastly, we developed iMiracle as an open-source software and anticipate that it can be a powerful tool in decoding the complexities of inter-cellular transcriptional regulation. Ziheng Duan, Siwei Xu, Cheyu Lee, Dylan Riffle, Jing Zhang 0062 |
CIKM | 5 |
| 2024 | scACT: Accurate Cross-modality Translation via Cycle-consistent Training from Unpaired Single-cell DataabstractSingle-cell sequencing technologies have revolutionized genomics by enabling the simultaneous profiling of various molecular modalities within individual cells. Their integration, especially cross-modality translation, offers deep insights into cellular regulatory mechanisms. Many methods have been developed for cross-modality translation, but their reliance on scarce high-quality co-assay data limits their applicability. Addressing this, we introduce scACT, a deep generative model designed to extract cross-modality biological insights from unpaired single-cell data. scACT tackles three major challenges: aligning unpaired multi-modal data via adversarial training, facilitating cross-modality translation without prior knowledge via cycle-consistent training, and enabling interpretable regulatory interconnections explorations via in-silico perturbations. To test its performance, we applied scACT on diverse single-cell datasets and found it outperformed existing methods in all three tasks. Finally, we have developed scACT as an individual open-source software package to advance single-cell omics data processing and analysis within the research community. Siwei Xu, Junhao Liu 0001, Jing Zhang 0062 |
CIKM | 3 |
| 2024 | scENCORE: leveraging single-cell epigenetic data to predict chromatin conformation using graph embeddingabstractDynamic compartmentalization of eukaryotic DNA into active and repressed states enables diverse transcriptional programs to arise from a single genetic blueprint, whereas its dysregulation can be strongly linked to a broad spectrum of diseases. While single-cell Hi-C experiments allow for chromosome conformation profiling across many cells, they are still expensive and not widely available for most labs. Here, we propose an alternate approach, scENCORE, to computationally reconstruct chromatin compartments from the more affordable and widely accessible single-cell epigenetic data. First, scENCORE constructs a long-range epigenetic correlation graph to mimic chromatin interaction frequencies, where nodes and edges represent genome bins and their correlations. Then, it learns the node embeddings to cluster genome regions into A/B compartments and aligns different graphs to quantify chromatin conformation changes across conditions. Benchmarking using cell-type-matched Hi-C experiments demonstrates that scENCORE can robustly reconstruct A/B compartments in a cell-type-specific manner. Furthermore, our chromatin confirmation switching studies highlight substantial compartment-switching events that may introduce substantial regulatory and transcriptional changes in psychiatric disease. In summary, scENCORE allows accurate and cost-effective A/B compartment reconstruction to delineate higher-order chromatin structure heterogeneity in complex tissues. Ziheng Duan, Siwei Xu, Shushrruth Sai Srinivasan, Ahyeon Hwang, Cheyu Lee, Mark Gerstein, Yu Luan, Matthew J. Girgenti, Jing Zhang 0062 |
Briefings Bioinform. | 10 |
| 2024 | Impeller: a path-based heterogeneous graph learning method for spatial transcriptomic data imputationabstractMOTIVATION: Recent advances in spatial transcriptomics allow spatially resolved gene expression measurements with cellular or even sub-cellular resolution, directly characterizing the complex spatiotemporal gene expression landscape and cell-to-cell interactions in their native microenvironments. Due to technology limitations, most spatial transcriptomic technologies still yield incomplete expression measurements with excessive missing values. Therefore, gene imputation is critical to filling in missing data, enhancing resolution, and improving overall interpretability. However, existing methods either require additional matched single-cell RNA-seq data, which is rarely available, or ignore spatial proximity or expression similarity information. RESULTS: To address these issues, we introduce Impeller, a path-based heterogeneous graph learning method for spatial transcriptomic data imputation. Impeller has two unique characteristics distinct from existing approaches. First, it builds a heterogeneous graph with two types of edges representing spatial proximity and expression similarity. Therefore, Impeller can simultaneously model smooth gene expression changes across spatial dimensions and capture similar gene expression signatures of faraway cells from the same type. Moreover, Impeller incorporates both short- and long-range cell-to-cell interactions (e.g. via paracrine and endocrine) by stacking multiple GNN layers. We use a learnable path operator in Impeller to avoid the over-smoothing issue of the traditional Laplacian matrices. Extensive experiments on diverse datasets from three popular platforms and two species demonstrate the superiority of Impeller over various state-of-the-art imputation methods. AVAILABILITY AND IMPLEMENTATION: The code and preprocessed data used in this study are available at https://github.com/aicb-ZhangLabs/Impeller and https://zenodo.org/records/11212604. Ziheng Duan, Dylan Riffle, Junhao Liu 0001, Martin Renqiang Min, Jing Zhang 0062 |
Bioinform. | 6 |
| 2023 | iHerd: an integrative hierarchical graph representation learning framework to quantify network changes and prioritize risk genes in diseaseabstractDifferent genes form complex networks within cells to carry out critical cellular functions, while network alterations in this process can potentially introduce downstream transcriptome perturbations and phenotypic variations. Therefore, developing efficient and interpretable methods to quantify network changes and pinpoint driver genes across conditions is crucial. We propose a hierarchical graph representation learning method, called iHerd. Given a set of networks, iHerd first hierarchically generates a series of coarsened sub-graphs in a data-driven manner, representing network modules at different resolutions (e.g., the level of signaling pathways). Then, it sequentially learns low-dimensional node representations at all hierarchical levels via efficient graph embedding. Lastly, iHerd projects separate gene embeddings onto the same latent space in its graph alignment module to calculate a rewiring index for driver gene prioritization. To demonstrate its effectiveness, we applied iHerd on a tumor-to-normal GRN rewiring analysis and cell-type-specific GCN analysis using single-cell multiome data of the brain. We showed that iHerd can effectively pinpoint novel and well-known risk genes in different diseases. Distinct from existing models, iHerd's graph coarsening for hierarchical learning allows us to successfully classify network driver genes into early and late divergent genes (EDGs and LDGs), emphasizing genes with extensive network changes across and within signaling pathway levels. This unique approach for driver gene classification can provide us with deeper molecular insights. The code is freely available at https://github.com/aicb-ZhangLabs/iHerd. All other relevant data are within the manuscript and supporting information files. Ziheng Duan, Ahyeon Hwang, Cheyu Lee, Kaichi Xie, Chutong Xiao, Min Xu 0009, Matthew J. Girgenti, Jing Zhang 0062 |
PLoS Comput. Biol. | 9 |
| 2022 | Deep Active Learning for Cryo-Electron Tomography ClassificationabstractCryo-Electron Tomography (cryo-ET) is an emerging 3D imaging technique which shows great potentials in structural biology research. One of the main challenges is to perform classification of macromolecules captured by cryo-ET. Re-cent efforts exploit deep learning to address this challenge. However, training reliable deep models usually requires a huge amount of labeled data in supervised fashion. Annotating cryo-ET data is arguably very expensive. Deep Active Learning (DAL) can be used to reduce labeling cost while not sacrificing the task performance too much. Nevertheless, most existing methods resort to auxiliary models or complex fashions (e.g. adversarial learning) for uncertainty estimation, the core of DAL. These models need to be highly customized for cryo-ET tasks which require 3D networks, and extra efforts are also indispensable for tuning these models, rendering a difficulty of deployment on cryo-ET tasks. To address these challenges, we propose a novel metric for data selection in DAL, which can also be leveraged as a regularizer of the empirical loss, further boosting the task model. We demonstrate the superiority of our method via extensive experiments on both simulated and real cryo-ET datasets. Our source Code and Appendix can be found at this URL. Tianyang Wang 0004, Bo Li 0013, Jing Zhang 0062, Mostofa Rafid Uddin, Min Xu 0009 |
ICIP | 3 |
| 2022 | Unsupervised Multi-Task Learning for 3D Subtomogram Image Alignment, Clustering and Segmentationabstract3D subtomogram image alignment, clustering, and segmentation are vital to macromolecular structure recognition in cryo-electron tomography (cryo-ET). However, acquiring ground-truth labels to train a unified deep learning model that can simultaneously deal with these tasks is unaffordable. To this end, we propose an end-to-end unified multi-task learning framework to simultaneously complete the three tasks, where models are trained in an unsupervised manner without using any labels. In particular, we have three parallel branches. In the alignment branch, we adopt a two-stage training scheme, i.e., self-supervised pretraining and constrained unsupervised training using our proposed skip correlation attention layer and constrained loss. Synchronously, in the clustering branch, the learned deep cluster features are utilized to iteratively cluster subtomograms into groups using pseudo-labels from an image-wise Gaussian Mixture Model (GMM). Meanwhile, in the segmentation branch, we use rough pseudo-labels generated from a voxel-wise GMM as supervision signals, and prior knowledge from humans is utilized to jointly learn how to correct these labels as well as predict reliable segmentation results. Benefiting from the end-to-end unified network architecture, our method achieves overall state-of-the-art performance on both simulated and real subtomogram processing benchmarks. Haoyi Zhu, Chuting Wang, Yuanxin Wang 0001, Zhaoxin Fan, Mostofa Rafid Uddin, Xin Gao 0001, Jing Zhang 0062, Min Xu 0009 |
ICIP | 7 |
| 2022 | Venus: An efficient virus infection detection and fusion site discovery method using single-cell and bulk RNA-seq dataabstractEarly and accurate detection of viruses in clinical and environmental samples is essential for effective public healthcare, treatment, and therapeutics. While PCR detects potential pathogens with high sensitivity, it is difficult to scale and requires knowledge of the exact sequence of the pathogen. With the advent of next-gen single-cell sequencing, it is now possible to scrutinize viral transcriptomics at the finest possible resolution-cells. This newfound ability to investigate individual cells opens new avenues to understand viral pathophysiology with unprecedented resolution. To leverage this ability, we propose an efficient and accurate computational pipeline, named Venus, for virus detection and integration site discovery in both single-cell and bulk-tissue RNA-seq data. Specifically, Venus addresses two main questions: whether a tissue/cell type is infected by viruses or a virus of interest? And if infected, whether and where has the virus inserted itself into the human genome? Our analysis can be broken into two parts-validation and discovery. Firstly, for validation, we applied Venus on well-studied viral datasets, such as HBV- hepatocellular carcinoma and HIV-infection treated with antiretroviral therapy. Secondly, for discovery, we analyzed datasets such as HIV-infected neurological patients and deeply sequenced T-cells. We detected viral transcripts in the novel target of the brain and high-confidence integration sites in immune cells. In conclusion, here we describe Venus, a publicly available software which we believe will be a valuable virus investigation tool for the scientific community at large. Cheyu Lee, Ziheng Duan, Min Xu 0009, Matthew J. Girgenti, Ke Xu 0010, Mark Gerstein, Jing Zhang 0062 |
PLoS Comput. Biol. | 8 |
| 2021 | Unsupervised Domain Alignment Based Open Set Structural Recognition of Macromolecules Captured By Cryo-Electron TomographyabstractCellular cryo-Electron Tomography (cryo-ET) provides three-dimensional views of structural and spatial information of various macromolecules in cells in a near-native state. Subtomogram classification is a key step for recognizing and differentiating these macromolecular structures. In recent years, deep learning methods have been developed for high-throughput subtomogram classification tasks; however, conventional supervised deep learning methods cannot recognize macromolecular structural classes that do not exist in the training data. This imposes a major weakness since most native macromolecular structures in cells are unknown and consequently, cannot be included in the training data. Therefore, open set learning which can recognize unknown macromolecular structures is necessary for boosting the power of automatic subtomogram classification. In this paper, we propose a method called Margin-based Loss for Unsupervised Domain Alignment (MLUDA) for open set recognition problems where only a few categories of interest are shared between cross-domain data. Through extensive experiments, we demonstrate that MLUDA performs well at cross-domain open-set classification on both public datasets and medical imaging datasets. So our method is of practical importance. Gregory Howe, Kai Yi, Jing Zhang 0062, Yi-Wei Chang, Min Xu 0009 |
ICIP | 5 |
| 2021 | SCAN-ATAC-Sim: a scalable and efficient method for simulating single-cell ATAC-seq data from bulk-tissue experimentsabstractSUMMARY: scATAC-seq is a powerful approach for characterizing cell-type-specific regulatory landscapes. However, it is difficult to benchmark the performance of various scATAC-seq analysis techniques (such as clustering and deconvolution) without having a priori a known set of gold-standard cell types. To simulate scATAC-seq experiments with known cell-type labels, we introduce an efficient and scalable scATAC-seq simulation method (SCAN-ATAC-Sim) that down-samples bulk ATAC-seq data (e.g. from representative cell lines or tissues). Our protocol uses a consistent but tunable signal-to-noise ratio across cell types in a scATAC-seq simulation for integrating bulk experiments with different levels of background noise, and it independently samples twice without replacement to account for the diploid genome. Because it uses an efficient weighted reservoir sampling algorithm and is highly parallelizable with OpenMP, our implementation in C++ allows millions of cells to be simulated in less than an hour on a laptop computer. AVAILABILITY AND IMPLEMENTATION: SCAN-ATAC-Sim is available at scan-atac-sim.gersteinlab.org. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Zhanlin Chen, Jing Zhang 0062, Jason Liu 0003, Jiangqi Zhu, Donghoon Lee 0006, Min Xu 0009, Mark Gerstein |
Bioinform. | 2 |
| 2021 | DECODE: a Deep-learning framework for Condensing enhancers and refining boundaries with large-scale functional assaysabstractMOTIVATION: Mapping distal regulatory elements, such as enhancers, is a cornerstone for elucidating how genetic variations may influence diseases. Previous enhancer-prediction methods have used either unsupervised approaches or supervised methods with limited training data. Moreover, past approaches have implemented enhancer discovery as a binary classification problem without accurate boundary detection, producing low-resolution annotations with superfluous regions and reducing the statistical power for downstream analyses (e.g. causal variant mapping and functional validations). Here, we addressed these challenges via a two-step model called Deep-learning framework for Condensing enhancers and refining boundaries with large-scale functional assays (DECODE). First, we employed direct enhancer-activity readouts from novel functional characterization assays, such as STARR-seq, to train a deep neural network for accurate cell-type-specific enhancer prediction. Second, to improve the annotation resolution, we implemented a weakly supervised object detection framework for enhancer localization with precise boundary detection (to a 10 bp resolution) using Gradient-weighted Class Activation Mapping. RESULTS: Our DECODE binary classifier outperformed a state-of-the-art enhancer prediction method by 24% in transgenic mouse validation. Furthermore, the object detection framework can condense enhancer annotations to only 13% of their original size, and these compact annotations have significantly higher conservation scores and genome-wide association study variant enrichments than the original predictions. Overall, DECODE is an effective tool for enhancer classification and precise localization. AVAILABILITY AND IMPLEMENTATION: DECODE source code and pre-processing scripts are available at decode.gersteinlab.org. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Zhanlin Chen, Jing Zhang 0062, Jason Liu 0003, Donghoon Lee 0006, Martin Renqiang Min, Min Xu 0009, Mark Gerstein |
Bioinform. | 2 |
| 2021 | Active learning to classify macromolecular structures in situ for less supervision in cryo-electron tomographyabstractMOTIVATION: Cryo-Electron Tomography (cryo-ET) is a 3D bioimaging tool that visualizes the structural and spatial organization of macromolecules at a near-native state in single cells, which has broad applications in life science. However, the systematic structural recognition and recovery of macromolecules captured by cryo-ET are difficult due to high structural complexity and imaging limits. Deep learning-based subtomogram classification has played critical roles for such tasks. As supervised approaches, however, their performance relies on sufficient and laborious annotation on a large training dataset. RESULTS: To alleviate this major labeling burden, we proposed a Hybrid Active Learning (HAL) framework for querying subtomograms for labeling from a large unlabeled subtomogram pool. Firstly, HAL adopts uncertainty sampling to select the subtomograms that have the most uncertain predictions. This strategy enforces the model to be aware of the inductive bias during classification and subtomogram selection, which satisfies the discriminativeness principle in AL literature. Moreover, to mitigate the sampling bias caused by such strategy, a discriminator is introduced to judge if a certain subtomogram is labeled or unlabeled and subsequently the model queries the subtomogram that have higher probabilities to be unlabeled. Such query strategy encourages to match the data distribution between the labeled and unlabeled subtomogram samples, which essentially encodes the representativeness criterion into the subtomogram selection process. Additionally, HAL introduces a subset sampling strategy to improve the diversity of the query set, so that the information overlap is decreased between the queried batches and the algorithmic efficiency is improved. Our experiments on subtomogram classification tasks using both simulated and real data demonstrate that we can achieve comparable testing performance (on average only 3% accuracy drop) by using less than 30% of the labeled subtomograms, which shows a very promising result for subtomogram classification task with limited labeling resources. AVAILABILITY AND IMPLEMENTATION: https://github.com/xulabs/aitom. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Xuefeng Du, Haohan Wang, Zhenxi Zhu, Yi-Wei Chang, Jing Zhang 0062, Eric P. Xing, Min Xu 0009 |
Bioinform. | 6 |
| 2021 | Bayesian structural time series for biomedical sensor data: A flexible modeling framework for evaluating interventionsabstractThe development of mobile-health technology has the potential to revolutionize personalized medicine. Biomedical sensors (e.g., wearables) can assist with determining treatment plans for individuals, provide quantitative information to healthcare providers, and give objective measurements of health, leading to the goal of precise phenotypic correlates for genotypes. Even though treatments and interventions are becoming more specific and datasets more abundant, measuring the causal impact of health interventions requires careful considerations of complex covariate structures, as well as knowledge of the temporal and spatial properties of the data. Thus, interpreting biomedical sensor data needs to make use of specialized statistical models. Here, we show how the Bayesian structural time series framework, widely used in economics, can be applied to these data. This framework corrects for covariates to provide accurate assessments of the significance of interventions. Furthermore, it allows for a time-dependent confidence interval of impact, which is useful for considering individualized assessments of intervention efficacy. We provide a customized biomedical adaptor tool, MhealthCI, around a specific implementation of the Bayesian structural time series framework that uniformly processes, prepares, and registers diverse biomedical data. We apply the software implementation of MhealthCI to a structured set of examples in biomedicine to showcase the ability of the framework to evaluate interventions with varying levels of data richness and covariate complexity and also compare the performance to other models. Specifically, we show how the framework is able to evaluate an exercise intervention's effect on stabilizing blood glucose in a diabetes dataset. We also provide a future-anticipating illustration from a behavioral dataset showcasing how the framework integrates complex spatial covariates. Overall, we show the robustness of the Bayesian structural time series framework when applied to biomedical sensor data, highlighting its increasing value for current and future datasets. Jason Liu 0003, Daniel J. Spakowicz, Garrett I. Ash, Rebecca Hoyd, Rohan Ahluwalia, Shaoke Lou, Donghoon Lee 0006, Jing Zhang 0062, Carolyn Presley, Ann Greene, Matthew Stults-Kolehmainen, Laura M. Nally, Julien S. Baker, Lisa M. Fucito, Stuart A. Weinzimer, Andrew V. Papachristos, Mark Gerstein |
PLoS Comput. Biol. | 9 |
| 2020 | TopicNet: a framework for measuring transcriptional regulatory network changeabstractMOTIVATION: Recently, many chromatin immunoprecipitation sequencing experiments have been carried out for a diverse group of transcription factors (TFs) in many different types of human cells. These experiments manifest large-scale and dynamic changes in regulatory network connectivity (i.e. network 'rewiring'), highlighting the different regulatory programs operating in disparate cellular states. However, due to the dense and noisy nature of current regulatory networks, directly comparing the gains and losses of targets of key TFs across cell states is often not informative. Thus, here, we seek an abstracted, low-dimensional representation to understand the main features of network change. RESULTS: We propose a method called TopicNet that applies latent Dirichlet allocation to extract functional topics for a collection of genes regulated by a given TF. We then define a rewiring score to quantify regulatory-network changes in terms of the topic changes for this TF. Using this framework, we can pinpoint particular TFs that change greatly in network connectivity between different cellular states (such as observed in oncogenesis). Also, incorporating gene expression data, we define a topic activity score that measures the degree to which a given topic is active in a particular cellular state. And we show how activity differences can indicate differential survival in various cancers. AVAILABILITY AND IMPLEMENTATION: The TopicNet framework and related analysis were implemented using R and all codes are available at https://github.com/gersteinlab/topicnet. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Shaoke Lou, Tianxiao Li 0001, Xiangmeng Kong, Jing Zhang 0062, Jason Liu 0003, Donghoon Lee 0006, Mark Gerstein |
Bioinform. | 4 |
| 2020 | DiNeR: a Differential graphical model for analysis of co-regulation Network RewiringabstractBACKGROUND: During transcription, numerous transcription factors (TFs) bind to targets in a highly coordinated manner to control the gene expression. Alterations in groups of TF-binding profiles (i.e. "co-binding changes") can affect the co-regulating associations between TFs (i.e. "rewiring the co-regulator network"). This, in turn, can potentially drive downstream expression changes, phenotypic variation, and even disease. However, quantification of co-regulatory network rewiring has not been comprehensively studied. RESULTS: To address this, we propose DiNeR, a computational method to directly construct a differential TF co-regulation network from paired disease-to-normal ChIP-seq data. Specifically, DiNeR uses a graphical model to capture the gained and lost edges in the co-regulation network. Then, it adopts a stability-based, sparsity-tuning criterion -- by sub-sampling the complete binding profiles to remove spurious edges -- to report only significant co-regulation alterations. Finally, DiNeR highlights hubs in the resultant differential network as key TFs associated with disease. We assembled genome-wide binding profiles of 104 TFs in the K562 and GM12878 cell lines, which loosely model the transition between normal and cancerous states in chronic myeloid leukemia (CML). In total, we identified 351 significantly altered TF co-regulation pairs. In particular, we found that the co-binding of the tumor suppressor BRCA1 and RNA polymerase II, a well-known transcriptional pair in healthy cells, was disrupted in tumors. Thus, DiNeR successfully extracted hub regulators and discovered well-known risk genes. CONCLUSIONS: Our method DiNeR makes it possible to quantify changes in co-regulatory networks and identify alterations to TF co-binding patterns, highlighting key disease regulators. Our method DiNeR makes it possible to quantify changes in co-regulatory networks and identify alterations to TF co-binding patterns, highlighting key disease regulators. Jing Zhang 0062, Jason Liu 0003, Donghoon Lee 0006, Shaoke Lou, Zhanlin Chen, Gamze Gürsoy, Mark Gerstein |
BMC Bioinform. | 1 |
| 2020 | NIMBus: a negative binomial regression based Integrative Method for mutation Burden AnalysisabstractBACKGROUND: Identifying frequently mutated regions is a key approach to discover DNA elements influencing cancer progression. However, it is challenging to identify these burdened regions due to mutation rate heterogeneity across the genome and across different individuals. Moreover, it is known that this heterogeneity partially stems from genomic confounding factors, such as replication timing and chromatin organization. The increasing availability of cancer whole genome sequences and functional genomics data from the Encyclopedia of DNA Elements (ENCODE) may help address these issues. RESULTS: We developed a negative binomial regression-based Integrative Method for mutation Burden analysiS (NIMBus). Our approach addresses the over-dispersion of mutation count statistics by (1) using a Gamma-Poisson mixture model to capture the mutation-rate heterogeneity across different individuals and (2) estimating regional background mutation rates by regressing the varying local mutation counts against genomic features extracted from ENCODE. We applied NIMBus to whole-genome cancer sequences from the PanCancer Analysis of Whole Genomes project (PCAWG) and other cohorts. It successfully identified well-known coding and noncoding drivers, such as TP53 and the TERT promoter. To further characterize the burdening of non-coding regions, we used NIMBus to screen transcription factor binding sites in promoter regions that intersect DNase I hypersensitive sites (DHSs). This analysis identified mutational hotspots that potentially disrupt gene regulatory networks in cancer. We also compare this method to other mutation burden analysis methods. CONCLUSION: NIMBus is a powerful tool to identify mutational hotspots. The NIMBus software and results are available as an online resource at github.gersteinlab.org/nimbus. Jing Zhang 0062, Jason Liu 0003, Patrick D. McGillivray, Caroline Yi, Lucas Lochovsky, Donghoon Lee 0006, Mark Gerstein |
BMC Bioinform. | 1 |
| 2020 | Epigenome-based splicing prediction using a recurrent neural networkabstractAlternative RNA splicing provides an important means to expand metazoan transcriptome diversity. Contrary to what was accepted previously, splicing is now thought to predominantly take place during transcription. Motivated by emerging data showing the physical proximity of the spliceosome to Pol II, we surveyed the effect of epigenetic context on co-transcriptional splicing. In particular, we observed that splicing factors were not necessarily enriched at exon junctions and that most epigenetic signatures had a distinctly asymmetric profile around known splice sites. Given this, we tried to build an interpretable model that mimics the physical layout of splicing regulation where the chromatin context progressively changes as the Pol II moves along the guide DNA. We used a recurrent-neural-network architecture to predict the inclusion of a spliced exon based on adjacent epigenetic signals, and we showed that distinct spatio-temporal features of these signals were key determinants of model outcome, in addition to the actual nucleotide sequence of the guide DNA strand. After the model had been trained and tested (with >80% precision-recall curve metric), we explored the derived weights of the latent factors, finding they highlight the importance of the asymmetric time-direction of chromatin context during transcription. Donghoon Lee 0006, Jing Zhang 0062, Jason Liu 0003, Mark Gerstein |
PLoS Comput. Biol. | 2 |
| 2020 | Few-shot learning for classification of novel macromolecular structures in cryo-electron tomogramsabstractCryo-electron tomography (cryo-ET) provides 3D visualization of subcellular components in the near-native state and at sub-molecular resolutions in single cells, demonstrating an increasingly important role in structural biology in situ. However, systematic recognition and recovery of macromolecular structures in cryo-ET data remain challenging as a result of low signal-to-noise ratio (SNR), small sizes of macromolecules, and high complexity of the cellular environment. Subtomogram structural classification is an essential step for such task. Although acquisition of large amounts of subtomograms is no longer an obstacle due to advances in automation of data collection, obtaining the same number of structural labels is both computation and labor intensive. On the other hand, existing deep learning based supervised classification approaches are highly demanding on labeled data and have limited ability to learn about new structures rapidly from data containing very few labels of such new structures. In this work, we propose a novel approach for subtomogram classification based on few-shot learning. With our approach, classification of unseen structures in the training data can be conducted given few labeled samples in test data through instance embedding. Experiments were performed on both simulated and real datasets. Our experimental results show that we can make inference on new structures given only five labeled samples for each class with a competitive accuracy (> 0.86 on the simulated dataset with SNR = 0.1), or even one sample with an accuracy of 0.7644. The results on real datasets are also promising with accuracy > 0.9 on both conditions and even up to 1 on one of the real datasets. Our approach achieves significant improvement compared with the baseline method and has strong capabilities of generalizing to other cellular components. Liangyong Yu, Bo Zhou 0009, Jing Zhang 0062, Xin Gao 0001, Rui Jiang 0001, Min Xu 0009 |
PLoS Comput. Biol. | 7 |
| 2018 | MOAT: efficient detection of highly mutated regions with the Mutations Overburdening Annotations ToolabstractSummary: Identifying genomic regions with higher than expected mutation count is useful for cancer driver detection. Previous parametric approaches require numerous cell-type-matched covariates for accurate background mutation rate (BMR) estimation, which is not practical for many situations. Non-parametric, permutation-based approaches avoid this issue but usually suffer from considerable compute-time cost. Hence, we introduce Mutations Overburdening Annotations Tool (MOAT), a non-parametric scheme that makes no assumptions about mutation process except requiring that the BMR changes smoothly with genomic features. MOAT randomly permutes single-nucleotide variants, or target regions, on a relatively large scale to provide robust burden analysis. Furthermore, we show how we can do permutations in an efficient manner using graphics processing unit acceleration, speeding up the calculation by a factor of ∼250. Availability and implementation: MOAT is available at moat.gersteinlab.org. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Lucas Lochovsky, Jing Zhang 0062, Mark Gerstein |
Bioinform. | 2 |