Yong Zhang 0006

dblp:66/4615-6 · DBLP profile ↗
← Back
14ranked-venue papers
2as first author
4since 2021 · last 2026
0000-0001-6316-2734ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 14 · 2 first-author · 4 since 2021
YearPublicationVenuePosition
2026 ChromBERT-tools: a versatile toolkit for context-specific regulatory representations of transcription regulators across different cell types
abstract
SUMMARY: Representations that encode the genome-wide regulatory behavior of transcription regulators provide a foundation for flexible transcription modeling and in silico regulatory analysis. Existing regulator representations are commonly derived from gene co-expression, motif annotations, or static protein features, which capture useful but limited aspects of regulator identity but do not directly model how regulators participate in region-specific regulatory programs across the genome. ChromBERT addresses this gap by learning context-aware regulatory representations from large-scale ChIP-seq data. However, routine bioinformatics applications require lightweight, accessible, and modular tools for generating, adapting, and interpreting these representations in user-defined biological contexts. Here, we present ChromBERT-tools, a user-oriented toolkit built upon ChromBERT that converts its regulatory representation framework into practical workflows for customizable analysis across cellular contexts. ChromBERT-tools provides command-line interfaces and Python APIs organized into three functional layers: representation generation, predictive modeling, and regulatory interpretation. The representation generation layer produces representations of genomic regions and transcription regulators. The predictive modeling layer fine-tunes ChromBERT for genome-wide regulatory activity prediction through classification or regression tasks, with optimized implementation to reduce running time and computational resource requirements. The regulatory interpretation layer supports inference of the context-specific roles of cis-regulatory elements and transcription regulators. These modules can be used independently or integrated into end-to-end workflows, enabling flexible analyses across diverse datasets. ChromBERT-tools lowers the barrier to applying context-specific regulatory representations in routine genomic analyses. AVAILABILITY AND IMPLEMENTATION: ChromBERT-tools is freely available at https://github.com/TongjiZhanglab/ChromBERT-tools, with documentation at https://chrombert-tools.readthedocs.io/en/latest/. A frozen archival snapshot is available on Zenodo under DOI: 10.5281/zenodo.20094206.
Zhanhao Li, Zhaowei Yu, Yong Zhang 0006
Bioinform.4
2021 CStreet: a computed Cell State trajectory inference method for time-series single-cell RNA sequencing data
abstract
MOTIVATION: The increasing amount of time-series single-cell RNA sequencing (scRNA-seq) data raises the key issue of connecting cell states (i.e. cell clusters or cell types) to obtain the continuous temporal dynamics of transcription, which can highlight the unified biological mechanisms involved in cell state transitions. However, most existing trajectory methods are specifically designed for individual cells, so they can hardly meet the needs of accurately inferring the trajectory topology of the cell state, which usually contains cells assigned to different branches. RESULTS: Here, we present CStreet, a computed Cell State trajectory inference method for time-series scRNA-seq data. It uses time-series information to construct the k-nearest neighbor connections between cells within each time point and between adjacent time points. Then, CStreet estimates the connection probabilities of the cell states and visualizes the trajectory, which may include multiple starting points and paths, using a force-directed graph. By comparing the performance of CStreet with that of six commonly used cell state trajectory reconstruction methods on simulated data and real data, we demonstrate the high accuracy and high tolerance of CStreet. AVAILABILITY AND IMPLEMENTATION: CStreet is written in Python and freely available on the web at https://github.com/TongjiZhanglab/CStreet and https://doi.org/10.5281/zenodo.4483205. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Chengchen Zhao, Wenchao Xiu, Yuwei Hua, Naiqian Zhang, Yong Zhang 0006
Bioinform.5
2021 NUCOME: A comprehensive database of nucleosome organization referenced landscapes in mammalian genomes
abstract
BACKGROUND: Nucleosome organization is involved in many regulatory activities in various organisms. However, studies integrating nucleosome organization in mammalian genomes are very limited mainly due to the lack of comprehensive data quality control (QC) assessment and uneven data quality of public data sets. RESULTS: The NUCOME is a database focused on filtering qualified nucleosome organization referenced landscapes covering various cell types in human and mouse based on QC metrics. The filtering strategy guarantees the quality of nucleosome organization referenced landscapes and exempts users from redundant data set selection and processing. The NUCOME database provides standardized, qualified data source and informative nucleosome organization features at a whole-genome scale and on the level of individual loci. CONCLUSIONS: The NUCOME provides valuable data resources for integrative analyses focus on nucleosome organization. The NUCOME is freely available at http://compbio-zhanglab.org/NUCOME .
Xiaolan Chen, Guifen Liu, Yong Zhang 0006
BMC Bioinform.4
2021 GLEANER: a web server for GermLine cycle Expression ANalysis and Epigenetic Roadmap visualization
abstract
BACKGROUND: Germline cells are important carriers of genetic and epigenetic information transmitted across generations in mammals. During the mammalian germline cell development cycle (i.e., the germline cycle), cell potency changes cyclically, accompanied by dynamic transcriptional changes and epigenetic reprogramming. Recently, to understand these dynamic and regulatory mechanisms, multiomic analyses, including transcriptomic and epigenomic analyses of DNA methylation, chromatin accessibility and histone modifications of germline cells, have been performed for different stages in human and mouse germline cycles. However, the long time span of the germline cycle and material scarcity of germline cells have largely limited the understanding of these dynamic characteristic changes. A tool that integrates the existing multiomics data and visualizes the overall continuous dynamic trends in the germline cycle can partially overcome such limitations. RESULTS: Here, we present GLEANER, a web server for GermLine cycle Expression ANalysis and Epigenetics Roadmap visualization. GLEANER provides a comprehensive collection of the transcriptome, DNA methylome, chromatin accessibility, and H3K4me3, H3K27me3, and H3K9me3 histone modification characteristics in human and mouse germline cycles. For each input gene, GLEANER shows the integrative analysis results of its transcriptional and epigenetic features, the genes with correlated transcriptional changes, and the overall continuous dynamic trends in the germline cycle. We further used two case studies to demonstrate the detailed functionality of GLEANER and highlighted that it can provide valuable clues to the epigenetic regulation mechanisms in the genetic and epigenetic information transmitted during the germline cycle. CONCLUSIONS: To the best of our knowledge, GLEANER is the first web server dedicated to the analysis and visualization of multiomics data related to the mammalian germline cycle. GLEANER is freely available at http://compbio-zhanglab.org/GLEANER .
Shiyang Zeng, Yuwei Hua, Yong Zhang 0006, Guifen Liu, Chengchen Zhao
BMC Bioinform.3
2016 Dr.seq: a quality control and analysis pipeline for droplet sequencing
abstract
MOTIVATION: Drop-seq has recently emerged as a powerful technology to analyze gene expression from thousands of individual cells simultaneously. Currently, Drop-seq technology requires refinement and quality control (QC) steps are critical for such data analysis. There is a strong need for a convenient and comprehensive approach to obtain dedicated QC and to determine the relationships between cells for ultra-high-dimensional datasets. RESULTS: We developed Dr.seq, a QC and analysis pipeline for Drop-seq data. By applying this pipeline, Dr.seq provides four groups of QC measurements for given Drop-seq data, including reads level, bulk-cell level, individual-cell level and cell-clustering level QC. We assessed Dr.seq on simulated and published Drop-seq data. Both assessments exhibit reliable results. Overall, Dr.seq is a comprehensive QC and analysis pipeline designed for Drop-seq data that is easily extended to other droplet-based data types. AVAILABILITY AND IMPLEMENTATION: Dr.seq is freely available at: http://www.tongji.edu.cn/∼zhanglab/drseq and https://bitbucket.org/tarela/drseq CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Xiao Huo, Sheng'en Hu, Chengchen Zhao, Yong Zhang 0006
Bioinform.4
2016 ChiLin: a comprehensive ChIP-seq and DNase-seq quality control and analysis pipeline
abstract
BACKGROUND: Transcription factor binding, histone modification, and chromatin accessibility studies are important approaches to understanding the biology of gene regulation. ChIP-seq and DNase-seq have become the standard techniques for studying protein-DNA interactions and chromatin accessibility respectively, and comprehensive quality control (QC) and analysis tools are critical to extracting the most value from these assay types. Although many analysis and QC tools have been reported, few combine ChIP-seq and DNase-seq data analysis and quality control in a unified framework with a comprehensive and unbiased reference of data quality metrics. RESULTS: ChiLin is a computational pipeline that automates the quality control and data analyses of ChIP-seq and DNase-seq data. It is developed using a flexible and modular software framework that can be easily extended and modified. ChiLin is ideal for batch processing of many datasets and is well suited for large collaborative projects involving ChIP-seq and DNase-seq from different designs. ChiLin generates comprehensive quality control reports that include comparisons with historical data derived from over 23,677 public ChIP-seq and DNase-seq samples (11,265 datasets) from eight literature-based classified categories. To the best of our knowledge, this atlas represents the most comprehensive ChIP-seq and DNase-seq related quality metric resource currently available. These historical metrics provide useful heuristic quality references for experiment across all commonly used assay types. Using representative datasets, we demonstrate the versatility of the pipeline by applying it to different assay types of ChIP-seq data. The pipeline software is available open source at https://github.com/cfce/chilin . CONCLUSION: ChiLin is a scalable and powerful tool to process large batches of ChIP-seq and DNase-seq datasets. The analysis output and quality metrics have been structured into user-friendly directories and reports. We have successfully compiled 23,677 profiles into a comprehensive quality atlas with fine classification for users.
Shenglin Mei, Qiu Wu, Hanfei Sun, Lewyn Li, Len Taing, Sujun Chen, Fugen Li, Tao Liu 0022, Chongzhi Zang, Clifford A. Meyer, Yong Zhang 0006, Myles Brown, Henry Long
BMC Bioinform.14
2013 BSeQC: quality control of bisulfite sequencing experiments
abstract
MOTIVATION: Bisulfite sequencing (BS-seq) has emerged as the gold standard to study genome-wide DNA methylation at single-nucleotide resolution. Quality control (QC) is a critical step in the analysis pipeline to ensure that BS-seq data are of high quality and suitable for subsequent analysis. Although several QC tools are available for next-generation sequencing data, most of them were not designed to handle QC issues specific to BS-seq protocols. Therefore, there is a strong need for a dedicated QC tool to evaluate and remove potential technical biases in BS-seq experiments. RESULTS: We developed a package named BSeQC to comprehensively evaluate the quality of BS-seq experiments and automatically trim nucleotides with potential technical biases that may result in inaccurate methylation estimation. BSeQC takes standard SAM/BAM files as input and generates bias-free SAM/BAM files for downstream analysis. Evaluation based on real BS-seq data indicates that the use of the bias-free SAM/BAM file substantially improves the quantification of methylation level. AVAILABILITY AND IMPLEMENTATION: BSeQC is freely available at: http://code.google.com/p/bseqc/.
Xueqiu Lin, Deqiang Sun, Benjamin Rodriguez, Hanfei Sun, Yong Zhang 0006, Wei Li 0036
Bioinform.6
2013 CistromeFinder for ChIP-seq and DNase-seq data reuse
abstract
SUMMARY: Chromatin immunoprecipitation and DNase I hypersensitivity assays with high-throughput sequencing have greatly accelerated the understanding of transcriptional and epigenetic regulation, although data reuse for the community of experimental biologists has been challenging. We created a data portal CistromeFinder that can help query, evaluate and visualize publicly available Chromatin immunoprecipitation and DNase I hypersensitivity assays with high-throughput sequencing data in human and mouse. The database currently contains 6378 samples over 4391 datasets, 313 factors and 102 cell lines or cell populations. Each dataset has gone through a consistent analysis and quality control pipeline; therefore, users could evaluate the overall quality of each dataset before examining binding sites near their genes of interest. CistromeFinder is integrated with UCSC genome browser for visualization, Primer3Plus for ChIP-qPCR primer design and CistromeMap for submitting newly available datasets. It also allows users to leave comments to facilitate data evaluation and update. AVAILABILITY: http://cistrome.org/finder. CONTACT: [email protected] or [email protected].
Hanfei Sun, Tao Liu 0022, Juan Wang 0025, Xueqiu Lin, Len Taing, Prakash K. Rao, Myles Brown, Yong Zhang 0006, Henry Long, Xiaole Shirley Liu
Bioinform.12
2012 GFOLD: a generalized fold change for ranking differentially expressed genes from RNA-seq data
abstract
MOTIVATION: RNA-seq has been widely used in transcriptome analysis to effectively measure gene expression levels. Although sequencing costs are rapidly decreasing, almost 70% of all the human RNA-seq samples in the gene expression omnibus do not have biological replicates and more unreplicated RNA-seq data were published than replicated RNA-seq data in 2011. Despite the large amount of single replicate studies, there is currently no satisfactory method for detecting differentially expressed genes when only a single biological replicate is available. RESULTS: We present the GFOLD (generalized fold change) algorithm to produce biologically meaningful rankings of differentially expressed genes from RNA-seq data. GFOLD assigns reliable statistics for expression changes based on the posterior distribution of log fold change. In this way, GFOLD overcomes the shortcomings of P-value and fold change calculated by existing RNA-seq analysis methods and gives more stable and biological meaningful gene rankings when only a single biological replicate is available. AVAILABILITY: The open source C/C++ program is available at http://www.tongji.edu.cn/∼zhanglab/GFOLD/index.html
Jianxing Feng, Clifford A. Meyer, Jun S. Liu, Xiaole Shirley Liu, Yong Zhang 0006
Bioinform.6
2012 DiNuP: a systematic approach to identify regions of differential nucleosome positioning
abstract
MOTIVATION: With the rapid development of high-throughput sequencing technologies, the genome-wide profiling of nucleosome positioning has become increasingly affordable. Many future studies will investigate the dynamic behaviour of nucleosome positioning in cells that have different states or that are exposed to different conditions. However, a robust method to effectively identify the regions of differential nucleosome positioning (RDNPs) has not been previously available. RESULTS: We describe a novel computational approach, DiNuP, that compares nucleosome profiles generated by high-throughput sequencing under various conditions. DiNuP provides a statistical P-value for each identified RDNP based on the difference of read distributions. DiNuP also empirically estimates the false discovery rate as a cutoff when two samples have different sequencing depths and differentiate reliable RDNPs from the background noise. Evaluation of DiNuP showed it to be both sensitive and specific for the detection of changes in nucleosome location, occupancy and fuzziness. RDNPs that were identified using publicly available datasets revealed that nucleosome positioning dynamics are closely related to the epigenetic regulation of transcription. AVAILABILITY AND IMPLEMENTATION: DiNuP is implemented in Python and is freely available at http://www.tongji.edu.cn/~zhanglab/DiNuP. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Qianzi Tang, Jianxing Feng, Xiaole Shirley Liu, Yong Zhang 0006
Bioinform.5
2012 CistromeMap: a knowledgebase and web server for ChIP-Seq and DNase-Seq studies in mouse and human
abstract
Abstract Summary: Transcription and chromatin regulators, and histone modifications play essential roles in gene expression regulation. We have created CistromeMap as a web server to provide a comprehensive knowledgebase of all of the publicly available ChIP-Seq and DNase-Seq data in mouse and human. We have also manually curated metadata to ensure annotation consistency, and developed a user-friendly display matrix for quick navigation and retrieval of data for specific factors, cells and papers. Finally, we provide users with summary statistics of ChIP-Seq and DNase-Seq studies. Availability: Freely available on the web at http://cistrome.dfci.harvard.edu/pc/ Contact: [email protected]; [email protected]
Ying Ge, Len Taing, Tao Liu 0022, Junsheng Chen, Lingling Shen, Xikun Duan, Sheng'en Hu, Wei Li 0036, Henry Long, Yong Zhang 0006, Xiaole Shirley Liu
Bioinform.14
2006 Phylophenetic properties of metabolic pathway topologies as revealed by global analysis
abstract
BACKGROUND: As phenotypic features derived from heritable characters, the topologies of metabolic pathways contain both phylogenetic and phenetic components. In the post-genomic era, it is possible to measure the "phylophenetic" contents of different pathways topologies from a global perspective. RESULTS: We reconstructed phylophenetic trees for all available metabolic pathways based on topological similarities, and compared them to the corresponding 16S rRNA-based trees. Similarity values for each pair of trees ranged from 0.044 to 0.297. Using the quartet method, single pathways trees were merged into a comprehensive tree containing information from a large part of the entire metabolic networks. This tree showed considerably higher similarity (0.386) to the corresponding 16S rRNA-based tree than any tree based on a single pathway, but was, on the other hand, sufficiently distinct to preserve unique phylogenetic information not reflected by the 16S rRNA tree. CONCLUSION: We observed that the topology of different metabolic pathways provided different phylogenetic and phenetic information, depicting the compromise between phylogenetic information and varying evolutionary pressures forming metabolic pathway topologies in different organisms. The phylogenetic information content of the comprehensive tree is substantially higher than that of any tree based on a single pathway, which also gave clues to constraints working on the topology of the global metabolic networks, information that is only partly reflected by the topologies of individual metabolic pathways.
Yong Zhang 0006, Shaojuan Li, Geir Skogerbø, Shiwei Sun, Hongchao Lu, Baochen Shi, Runsheng Chen
BMC Bioinform.1
2006 Dynamic Changes in Subgraph Preference Profiles of Crucial Transcription Factors
abstract
Transcription factors with a large number of target genes--transcription hub(s), or THub(s)--are usually crucial components of the regulatory system of a cell, and the different patterns through which they transfer the transcriptional signal to downstream cascades are of great interest. By profiling normalized abundances (A(N)) of basic regulatory patterns of individual THubs in the yeast Saccharomyces cerevisiae transcriptional regulation network under five different cellular states and environmental conditions, we have investigated their preferences for different basic regulatory patterns. Subgraph-normalized abundances downstream of individual THubs often differ significantly from that of the network as a whole, and conversely, certain over-represented subgraphs are not preferred by any THub. The THub preferences changed substantially when the cellular or environmental conditions changed. This switching of regulatory pattern preferences suggests that a change in conditions does not only elicit a change in response by the regulatory network, but also a change in the mechanisms by which the response is mediated. The THub subgraph preference profile thus provides a novel tool for description of the structure and organization between the large-scale exponents and local regulatory patterns.
Changning Liu, Geir Skogerbø, Hongchao Lu, Baochen Shi, Yong Zhang 0006, Tao Wu 0002, Runsheng Chen
PLoS Comput. Biol.8
2004 Conservation analysis of small RNA genes in Escherichia coli
abstract
MOTIVATION: Small RNA (sRNA) genes in Escherichia coli have been in focus recently, as 44 out of 55 experimentally confirmed sRNA genes have been precisely located in the genome. The object of this study is to analyze quantitatively the conservation of these sRNA genes and compare it with the conservation of protein-encoding genes, function-unknown regions and tRNA genes. RESULTS: The results show that within an evolutionary distance of 0.26, both sRNA genes and protein-encoding genes display a similar tendency in their degrees of conservation at the nucleotide level. In addition, the conservation of sRNA genes is much stronger than function-unknown regions, but much weaker than tRNA genes. Based on the conservation of studied sRNA genes, we also give clues to estimate the total number of sRNA genes in E.coli. SUPPLEMENTARY INFORMATION: Supplementary information is available at http://www.bioinfo.org.cn/SM/sRNAconservation.htm
Yong Zhang 0006, Lunjiang Ling, Baochen Shi, Runsheng Chen
Bioinform.1