Zheng Wang 0049

dblp:181/2834-49 · DBLP profile ↗
← Back
16ranked-venue papers
0as first author
8since 2021 · last 2026
0000-0003-3892-6151ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 16 · 8 since 2021
YearPublicationVenuePosition
2026 Inferring the qualities of protein-RNA models with graph transformers
abstract
MOTIVATION: Breakthrough advancements in protein tertiary and quaternary structure prediction have accelerated structural bioinformatics research activity and drug development processes. However, many biological mechanisms involve more complicated interactions, such as those between amino and nucleic acids. Predicting the structure of protein-RNA complexes is highly relevant and challenging due to data scarcity and experimental difficulties. Understanding and interpreting these interactions can yield crucial insights into various human diseases and biological phenomena. Thus, quality assessment methods that specifically evaluate protein-RNA complex models can provide significant utility in this emerging area of protein-RNA structural bioinformatics research. RESULTS: We propose a novel graph transformer-based approach named complex quality assessment of RNA and protein (CARP) to infer multiple quality perspectives of protein-RNA complex models. For a single protein-RNA complex model, in one shot, CARP simultaneously predicts multiple overall fold, overall interface, and per-protein-RNA interface quality estimates. When evaluated against a non-redundant protein-RNA docking benchmark, our methods demonstrated obvious improved performance compared to almost all of the existing scoring tools, particularly when ordering and selecting the highest quality decoys. Furthermore, CARP consistently selected higher quality models relative to other predictors when tested on CASP16 targets. Specifically, CARP-predicted global interface and global protein-RNA interface qualities were ranked first and second, respectively, based on the selected top-3 models over all ten CASP16 protein-RNA complex targets. CARP also showed a strong ability, compared to both existing tools and AlphaFold3 self-estimates, in selecting high quality AlphaFold3 models. AVAILABILITY AND IMPLEMENTATION: CARP is freely available at github.com/zwang-bioinformatics/CARP/.
Andrew Jordan Siciliano, Yifan Bao, Bishal Shrestha, Zheng Wang 0049
Bioinform.4
2026 SCW: building the whole-genome 3D structures based on extremely sparse single-cell Hi-C data
abstract
BACKGROUND: The study of three-dimensional (3D) genome structures at the single-cell level is crucial for understanding cell-to-cell variability. However, it is challenging to reconstruct the 3D structures of the whole genome based on single-cell Hi-C data because of the sparseness of the single-cell Hi-C data and the complexity of the problem. RESULTS: To address this, we developed a new computational method, named SCW (single-cell whole-genome), to build the high-resolution 3D genome structures of the whole genome based on extremely sparse single-cell Hi-C data (either zeros or ones in the Hi-C matrix). We evaluated our reconstructed 3D genome structures on various types of cells, checked the fitness of the reconstructed 3D structures to the single-cell Hi-C data, cross-validated the reconstructed structures with FISH data, bulk Hi-C data, and gene expression data, and then compared SCW with the state-of-the-art tools Nuc_dynamics, Hickit, and Tensor-FLAMINGO. SCW achieved better robustness as Nuc_dynamics failed on extremely sparse Hi-C matrices that contained only zeros or ones. The Pearson Correlation between our reconstructed 3D structure and the FISH data can reach 0.63, which is higher than the structures built by Hickit. SCW can also build better structures compared to Tensor-FLAMINGO based on multiple evaluations, particularly with the 20 Kbp resolution. Both the intra- and inter-chromosomal contact patterns are maintained in our reconstructed 3D structures, which also match the findings from single-cell gene expression data. CONCLUSIONS: SCW enables high-precision whole-genome 3D reconstruction from extremely sparse single-cell Hi-C data. SCW outperforms existing tools in structural accuracy and robustly maintains intra/inter-chromosomal contacts. Its versatility is validated across diverse cell types.
Tong Liu 0028, Bishal Shrestha, Zheng Wang 0049
BMC Bioinform.4
2025 HiC4D-SPOT: a spatiotemporal outlier detection tool for Hi-C data
abstract
The 3D organization of chromatin is essential for the functioning of cellular processes, including transcriptional regulation, genome integrity, chromatin accessibility, and higher order nuclear architecture. However, detecting anomalous chromatin interactions in spatiotemporal Hi-C data remains a significant challenge. We present HiC4D-SPOT, an unsupervised deep-learning framework that models chromatin dynamics using a ConvLSTM-based autoencoder to identify structural anomalies. Benchmarking results demonstrate high reconstruction fidelity, with Pearson Correlation Coefficient and Spearman Correlation Coefficient values of 0.9, while accurately detecting deviations linked to temporal inconsistencies, topologically associating domain (TAD) and loop perturbations, and significant chromatin remodeling events. HiC4D-SPOT successfully identifies swapped time points in a time-swap experiment, captures simulated TAD and loop disruptions with high confidence scores and statistical significance of 0.01, and detects HERV-H boundary weakening during cardiomyocyte differentiation, as well as cohesin-mediated loop loss and recovery-aligning with experimentally observed chromatin remodeling events. These findings establish HiC4D-SPOT as an efficient tool for analyzing 3D chromatin dynamics, enabling the detection of biologically significant structural anomalies in spatiotemporal Hi-C data.
Bishal Shrestha, Zheng Wang 0049
Briefings Bioinform.2
2025 Generating three-dimensional genome structures with a variational quantum algorithm
abstract
Chromosome conformation capture experiments have revealed the underlying spatial interactions that govern three-dimensional (3D) genome organization and topology. Detecting 3D contacts between genomic loci considerably enhances our understanding of fundamental regulatory processes. Modeling 3D structures from experimental contact matrices can further contextualize the relationship between 3D genome organization and regulation. While classical algorithms have been successful in reconstructing genomic conformations, we investigate the prospect of quantum computation to aid in modeling the conformational space. In this context, we propose a novel variational quantum algorithm (VQA) to model the distribution of 3D genomic structures from experimental contact data. Through rigorous evaluations, we demonstrate the capability of our algorithm to sample ensembles of viable 3D conformations that agree well with experimental and simulated contact data. Furthermore, we extend our methodology to model the conformational space of a single cell or a population of cells. In the advent of sufficient quantum utility, the insights gained from this study can serve as a foundation for investigating high-resolution, large-scale ensembles of genomic conformations through generative VQAs.
Andrew Jordan Siciliano, Zheng Wang 0049
Briefings Bioinform.2
2024 Learning Micro-C from Hi-C with diffusion models
abstract
In the last few years, Micro-C has shown itself as an improved alternative to Hi-C. It replaced the restriction enzymes in Hi-C assays with micrococcal nuclease (MNase), resulting in capturing nucleosome resolution chromatin interactions. The signal-to-noise improvement of Micro-C allows it to detect more chromatin loops than high-resolution Hi-C. However, compared with massive Hi-C datasets available in the literature, there are only a limited number of Micro-C datasets. To take full advantage of these Hi-C datasets, we present HiC2MicroC, a computational method learning and then predicting Micro-C from Hi-C based on the denoising diffusion probabilistic models (DDPM). We trained our DDPM and other regression models in human foreskin fibroblast (HFFc6) cell line and evaluated these methods in six different cell types at 5-kb and 1-kb resolution. Our evaluations demonstrate that both HiC2MicroC and regression methods can markedly improve Hi-C towards Micro-C, and our DDPM-based HiC2MicroC outperforms regression in various terms. First, HiC2MicroC successfully recovers most of the Micro-C loops even those not detected in Hi-C maps. Second, a majority of the HiC2MicroC-recovered loops anchor CTCF binding sites in a convergent orientation. Third, HiC2MicroC loops share genomic and epigenetic properties with Micro-C loops, including linking promoters and enhancers, and their anchors are enriched for structural proteins (CTCF and cohesin) and histone modifications. Lastly, we find our recovered loops are also consistent with the loops identified from promoter capture Micro-C (PCMicro-C) and Chromatin Interaction Analysis by Paired-End Tag Sequencing (ChIA-PET). Overall, HiC2MicroC is an effective tool for further studying Hi-C data with Micro-C as a template. HiC2MicroC is publicly available at https://github.com/zwang-bioinformatics/HiC2MicroC/.
Tong Liu 0028, Zheng Wang 0049
PLoS Comput. Biol.3
2023 scHiMe: predicting single-cell DNA methylation levels based on single-cell Hi-C data
abstract
Recently a biochemistry experiment named methyl-3C was developed to simultaneously capture the chromosomal conformations and DNA methylation levels on individual single cells. However, the number of data sets generated from this experiment is still small in the scientific community compared with the greater amount of single-cell Hi-C data generated from separate single cells. Therefore, a computational tool to predict single-cell methylation levels based on single-cell Hi-C data on the same individual cells is needed. We developed a graph transformer named scHiMe to accurately predict the base-pair-specific (bp-specific) methylation levels based on both single-cell Hi-C data and DNA nucleotide sequences. We benchmarked scHiMe for predicting the bp-specific methylation levels on all of the promoters of the human genome, all of the promoter regions together with the corresponding first exon and intron regions, and random regions on the whole genome. Our evaluation showed a high consistency between the predicted and methyl-3C-detected methylation levels. Moreover, the predicted DNA methylation levels resulted in accurate classifications of cells into different cell types, which indicated that our algorithm successfully captured the cell-to-cell variability in the single-cell Hi-C data. scHiMe is freely available at http://dna.cs.miami.edu/scHiMe/.
Tong Liu 0028, Zheng Wang 0049
Briefings Bioinform.3
2023 DeepChIA-PET: Accurately predicting ChIA-PET from Hi-C and ChIP-seq with deep dilated networks
abstract
Chromatin interaction analysis by paired-end tag sequencing (ChIA-PET) can capture genome-wide chromatin interactions mediated by a specific DNA-associated protein. The ChIA-PET experiments have been applied to explore the key roles of different protein factors in chromatin folding and transcription regulation. However, compared with widely available Hi-C and ChIP-seq data, there are not many ChIA-PET datasets available in the literature. A computational method for accurately predicting ChIA-PET interactions from Hi-C and ChIP-seq data is needed that can save the efforts of performing wet-lab experiments. Here we present DeepChIA-PET, a supervised deep learning approach that can accurately predict ChIA-PET interactions by learning the latent relationships between ChIA-PET and two widely used data types: Hi-C and ChIP-seq. We trained our deep models with CTCF-mediated ChIA-PET of GM12878 as ground truth, and the deep network contains 40 dilated residual convolutional blocks. We first showed that DeepChIA-PET with only Hi-C as input significantly outperforms Peakachu, another computational method for predicting ChIA-PET from Hi-C but using random forests. We next proved that adding ChIP-seq as one extra input does improve the classification performance of DeepChIA-PET, but Hi-C plays a more prominent role in DeepChIA-PET than ChIP-seq. Our evaluation results indicate that our learned models can accurately predict not only CTCF-mediated ChIA-ET in GM12878 and HeLa but also non-CTCF ChIA-PET interactions, including RNA polymerase II (RNAPII) ChIA-PET of GM12878, RAD21 ChIA-PET of GM12878, and RAD21 ChIA-PET of K562. In total, DeepChIA-PET is an accurate tool for predicting the ChIA-PET interactions mediated by various chromatin-associated proteins from different cell types.
Tong Liu 0028, Zheng Wang 0049
PLoS Comput. Biol.2
2022 ST-ChIP: Accurate prediction of spatiotemporal ChIP-seq data with recurrent neural networks
abstract
Chromatin immunoprecipitation followed by sequencing (ChIP-seq) is a powerful method for locating protein-DNA binding sites. Spatiotemporal ChIP-seq data greatly contribute to the studies of dynamic biological processes as they contain information from both spatial and temporal dimensions. However, we can hardly find a computational method for forecasting spatiotemporal ChIP-seq data in the literature. Here we present ST-ChIP, a supervised method using Long Short-Term Memory (LSTM) for predicting coverage or peaks of spatiotemporal ChIP-seq data. We benchmarked three recurrent neural networks and found that two of them achieved higher predictive performances on recovering coverage or peaks of the forecasting time steps. Our results demonstrate that enhancer regions are enriched with our predicted H3K4me1 coverage, and promoter regions are enriched with our predicted H3K4me3 peaks, which match the findings from other studies. In total, ST-ChIP is an effective method for accurately predicting spatiotemporal ChIP-seq data. ST-ChIP is publicly available at http://dna.cs.miami.edu/ST-ChIP/.
Tong Liu 0028, Zheng Wang 0049
BIBM2
2020 MASS: predict the global qualities of individual protein models using random forests and novel statistical potentials
abstract
BACKGROUND: Protein model quality assessment (QA) is an essential procedure in protein structure prediction. QA methods can predict the qualities of protein models and identify good models from decoys. Clustering-based methods need a certain number of models as input. However, if a pool of models are not available, methods that only need a single model as input are indispensable. RESULTS: We developed MASS, a QA method to predict the global qualities of individual protein models using random forests and various novel energy functions. We designed six novel energy functions or statistical potentials that can capture the structural characteristics of a protein model, which can also be used in other protein-related bioinformatics research. MASS potentials demonstrated higher importance than the energy functions of RWplus, GOAP, DFIRE and Rosetta when the scores they generated are used as machine learning features. MASS outperforms almost all of the four CASP11 top-performing single-model methods for global quality assessment in terms of all of the four evaluation criteria officially used by CASP, which measure the abilities to assign relative and absolute scores, identify the best model from decoys, and distinguish between good and bad models. MASS has also achieved comparable performances with the leading QA methods in CASP12 and CASP13. CONCLUSIONS: MASS and the source code for all MASS potentials are publicly available at http://dna.cs.miami.edu/MASS/ .
Tong Liu 0028, Zheng Wang 0049
BMC Bioinform.2
2019 HiCNN: a very deep convolutional neural network to better enhance the resolution of Hi-C data
abstract
MOTIVATION: High-resolution Hi-C data are indispensable for the studies of three-dimensional (3D) genome organization at kilobase level. However, generating high-resolution Hi-C data (e.g. 5 kb) by conducting Hi-C experiments needs millions of mammalian cells, which may eventually generate billions of paired-end reads with a high sequencing cost. Therefore, it will be important and helpful if we can enhance the resolutions of Hi-C data by computational methods. RESULTS: We developed a new computational method named HiCNN that used a 54-layer very deep convolutional neural network to enhance the resolutions of Hi-C data. The network contains both global and local residual learning with multiple speedup techniques included resulting in fast convergence. We used mean squared errors and Pearson's correlation coefficients between real high-resolution and computationally predicted high-resolution Hi-C data to evaluate the method. The evaluation results show that HiCNN consistently outperforms HiCPlus, the only existing tool in the literature, when training and testing data are extracted from the same cell type (i.e. GM12878) and from two different cell types in the same or different species (i.e. GM12878 as training with K562 as testing, and GM12878 as training with CH12-LX as testing). We further found that the HiCNN-enhanced high-resolution Hi-C data are more consistent with real experimental high-resolution Hi-C data than HiCPlus-enhanced data in terms of indicating statistically significant interactions. Moreover, HiCNN can efficiently enhance low-resolution Hi-C data, which eventually helps recover two chromatin loops that were confirmed by 3D-FISH. AVAILABILITY AND IMPLEMENTATION: HiCNN is freely available at http://dna.cs.miami.edu/HiCNN/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Tong Liu 0028, Zheng Wang 0049
Bioinform.2
2019 SCL: a lattice-based approach to infer 3D chromosome structures from single-cell Hi-C data
abstract
MOTIVATION: In contrast to population-based Hi-C data, single-cell Hi-C data are zero-inflated and do not indicate the frequency of proximate DNA segments. There are a limited number of computational tools that can model the 3D structures of chromosomes based on single-cell Hi-C data. RESULTS: We developed single-cell lattice (SCL), a computational method to reconstruct 3D structures of chromosomes based on single-cell Hi-C data. We designed a loss function and a 2 D Gaussian function specifically for the characteristics of single-cell Hi-C data. A chromosome is represented as beads-on-a-string and stored in a 3 D cubic lattice. Metropolis-Hastings simulation and simulated annealing are used to simulate the structure and minimize the loss function. We evaluated the SCL-inferred 3 D structures (at both 500 and 50 kb resolutions) using multiple criteria and compared them with the ones generated by another modeling software program. The results indicate that the 3 D structures generated by SCL closely fit single-cell Hi-C data. We also found similar patterns of trans-chromosomal contact beads, Lamin-B1 enriched topologically associating domains (TADs), and H3K4me3 enriched TADs by mapping data from previous studies onto the SCL-inferred 3 D structures. AVAILABILITY AND IMPLEMENTATION: The C++ source code of SCL is freely available at http://dna.cs.miami.edu/SCL/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Zheng Wang 0049
Bioinform.2
2019 Exploring the 2D and 3D structural properties of topologically associating domains
abstract
BACKGROUND: Topologically associating domains (TADs) are genomic regions with varying lengths. The interactions within TADs are more frequent than those between different TADs. TADs or sub-TADs are considered the structural and functional units of the mammalian genomes. Although TADs are important for understanding how genomes function, we have limited knowledge about their 3D structural properties. RESULTS: In this study, we designed and benchmarked three metrics for capturing the three-dimensional and two-dimensional structural signatures of TADs, which can help better understand TADs' structural properties and the relationships between structural properties and genetic and epigenetic features. The first metric for capturing 3D structural properties is radius of gyration, which in this study is used to measure the spatial compactness of TADs. The mass value of each DNA bead in a 3D structure is novelly defined as one or more genetic or epigenetic feature(s). The second metric is folding degree. The last metric is exponent parameter, which is used to capture the 2D structural properties based on TADs' Hi-C contact matrices. In general, we observed significant correlations between the three metrics and the genetic and epigenetic features. We made the same observations when using H3K4me3, transcription start sites, and RNA polymerase II to represent the mass value in the modified radius-of-gyration metric. Moreover, we have found that the TADs in the clusters of depleted chromatin states apparently correspond to smaller exponent parameters and larger radius of gyrations. In addition, a new objective function of multidimensional scaling for modelling chromatin or TADs 3D structures was designed and benchmarked, which can handle the DNA bead-pairs with zero Hi-C contact values. CONCLUSIONS: The web server for reconstructing chromatin 3D structures using multiple different objective functions and the related source code are publicly available at http://dna.cs.miami.edu/3DChrom/.
Tong Liu 0028, Zheng Wang 0049
BMC Bioinform.2
2019 Predicting protein residue-residue contacts using random forests and deep networks
abstract
BACKGROUND: The ability to predict which pairs of amino acid residues in a protein are in contact with each other offers many advantages for various areas of research that focus on proteins. For example, contact prediction can be used to reduce the computational complexity of predicting the structure of proteins and even to help identify functionally important regions of proteins. These predictions are becoming especially important given the relatively low number of experimentally determined protein structures compared to the amount of available protein sequence data. RESULTS: Here we have developed and benchmarked a set of machine learning methods for performing residue-residue contact prediction, including random forests, direct-coupling analysis, support vector machines, and deep networks (stacked denoising autoencoders). These methods are able to predict contacting residue pairs given only the amino acid sequence of a protein. According to our own evaluations performed at a resolution of +/- two residues, the predictors we trained with the random forest algorithm were our top performing methods with average top 10 prediction accuracy scores of 85.13% (short range), 74.49% (medium range), and 54.49% (long range). Our ensemble models (stacked denoising autoencoders combined with support vector machines) were our best performing deep network predictors and achieved top 10 prediction accuracy scores of 75.51% (short range), 60.26% (medium range), and 43.85% (long range) using the same evaluation. These tests were blindly performed on targets from the CASP11 dataset; and the results suggested that our models achieved comparable performance to contact predictors developed by groups that participated in CASP11. CONCLUSIONS: Due to the challenging nature of contact prediction, it is beneficial to develop and benchmark a variety of different prediction methods. Our work has produced useful tools with a simple interface that can provide contact predictions to users without requiring a lengthy installation process. In addition to this, we have released our C++ implementation of the direct-coupling analysis method as a standalone software package. Both this tool and our RFcon web server are freely available to the public at http://dna.cs.miami.edu/RFcon /.
Joseph Luttrell IV, Tong Liu 0028, Zheng Wang 0049
BMC Bioinform.4
2018 Measuring the three-dimensional structural properties of topologically associating domains
Tong Liu 0028, Zheng Wang 0049
BIBM2
2018 scHiCNorm: a software package to eliminate systematic biases in single-cell Hi-C data
abstract
Summary: We build a software package scHiCNorm that uses zero-inflated and hurdle models to remove biases from single-cell Hi-C data. Our evaluations prove that our models can effectively eliminate systematic biases for single-cell Hi-C data, which better reveal cell-to-cell variances in terms of chromosomal structures. Availability and implementation: scHiCNorm is available at http://dna.cs.miami.edu/scHiCNorm/. Perl scripts are provided that can generate bias features. Pre-built bias features for human (hg19 and hg38) and mouse (mm9 and mm10) are available to download. R scripts can be downloaded to remove biases. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Tong Liu 0028, Zheng Wang 0049
Bioinform.2
2018 Reconstructing high-resolution chromosome three-dimensional structures by Hi-C complex networks
abstract
BACKGROUND: Hi-C data have been widely used to reconstruct chromosomal three-dimensional (3D) structures. One of the key limitations of Hi-C is the unclear relationship between spatial distance and the number of Hi-C contacts. Many methods used a fixed parameter when converting the number of Hi-C contacts to wish distances. However, a single parameter cannot properly explain the relationship between wish distances and genomic distances or the locations of topologically associating domains (TADs). RESULTS: We have addressed one of the key issues of using Hi-C data, that is, the unclear relationship between spatial distances and the number of Hi-C contacts, which is crucial to understand significant biological functions, such as the enhancer-promoter interactions. Specifically, we developed a new method to infer this converting parameter and pairwise Euclidean distances based on the topology of the Hi-C complex network (HiCNet). The inferred distances were modeled by clustering coefficient and multiple other types of constraints. We found that our inferred distances between bead-pairs within the same TAD were apparently smaller than those distances between bead-pairs from different TADs. Our inferred distances had a higher correlation with fluorescence in situ hybridization (FISH) data, fitted the localization patterns of Xist transcripts on DNA, and better matched 156 pairs of protein-enabled long-range chromatin interactions detected by ChIA-PET. Using the inferred distances and another round of optimization, we further reconstructed 40 kb high-resolution 3D chromosomal structures of mouse male ES cells. The high-resolution structures successfully illustrate TADs and DNA loops (peaks in Hi-C contact heatmaps) that usually indicate enhancer-promoter interactions. CONCLUSIONS: We developed a novel method to infer the wish distances between DNA bead-pairs from Hi-C contacts. High-resolution 3D structures of chromosomes were built based on the newly-inferred wish distances. This whole process has been implemented as a tool named HiCNet, which is publicly available at http://dna.cs.miami.edu/HiCNet/ .
Tong Liu 0028, Zheng Wang 0049
BMC Bioinform.2