EDBT 2026 Demo / reviewers in the wild / expert
Bing Zhang 0003
dblp:74/2272-3
· DBLP profile ↗
16ranked-venue papers
2as first author
2since 2021 · last 2022
0000-0001-8676-2425ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 15 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
6 papers |
Bioinformatics and computational biology · 94% Medical and health informatics · 6% |
Topics — the 11 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology
omics data analysis |
0.6 | 1 | 2022 | OmicsEV: a tool for comprehensive quality evaluation of omics data tables · Bioinform. 2022 |
Bioinformatics and computational biology
cancer genomics |
0.2 | 1 | 2015 | Empowering biologists with multi-omics data: colorectal cancer as a paradigm · Bioinform. 2015 |
Bioinformatics and computational biology
multi-omics data integration |
0.2 | 1 | 2015 | Empowering biologists with multi-omics data: colorectal cancer as a paradigm · Bioinform. 2015 |
Bioinformatics and computational biology › protein sequence analysis
protein database search |
0.2 | 1 | 2013 | customProDB: an R package to generate customized protein databases from RNA-Seq data for proteomics search · Bioinform. 2013 |
Bioinformatics and computational biology
proteomics |
0.2 | 1 | 2013 | customProDB: an R package to generate customized protein databases from RNA-Seq data for proteomics search · Bioinform. 2013 |
Medical and health informatics
cancer prognosis |
0.1 | 1 | 2011 | Semi-supervised learning improves gene expression-based prediction of cancer recurrence · Bioinform. 2011 |
Bioinformatics and computational biology
gene expression analysis |
0.1 | 1 | 2011 | Semi-supervised learning improves gene expression-based prediction of cancer recurrence · Bioinform. 2011 |
Bioinformatics and computational biology › functional genomics › functional enrichment analysis
gene set enrichment analysis |
0.1 | 1 | 2008 | Supervised principal component analysis for gene set enrichment of microarray data with continuous or survival outcomes · Bioinform. 2008 |
Bioinformatics and computational biology › biological network › network biology
protein complex identification |
0.1 | 1 | 2008 | From pull-down data to protein interaction networks and complexes with biological relevance · Bioinform. 2008 |
Bioinformatics and computational biology › protein analysis › protein-protein interaction › protein-protein interaction network analysis
protein-protein interaction network inference |
0.1 | 1 | 2008 | From pull-down data to protein interaction networks and complexes with biological relevance · Bioinform. 2008 |
Bioinformatics and computational biology › proteomics
mass spectrometry proteomics |
0.0 | 1 | 2008 | From pull-down data to protein interaction networks and complexes with biological relevance · Bioinform. 2008 |
Methods — techniques the papers use, named apart from their topics
multi-omics concordance · 0.6data normalization assessment · 0.6batch effect detection · 0.6network visualization · 0.2data integration · 0.2support vector machine · 0.1semi-supervised learning · 0.1principal component analysis · 0.1mixture distribution · 0.1gumbel extreme value distribution · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Weakly Unpaired Image Translation from Hematoxylin and Eosin Staining Image to Immunohistochemistry Staining ImageabstractHistopathological imaging is one of the most critical disease diagnosis tools. Hematoxylin and eosin (H&E) staining images are routinely acquired and show all cells in the tissue. In contrast, immunohistochemistry (IHC) staining of cell type markers reveals spatial information of specific cell types of interest, and their use is limited by more complex laboratory processing and high cost. As an example, CD3 IHC images provide critical spatial information on the tumor-infiltrating lymphocytes (TILs), but for most tumor specimens, only H&E images are available. It is of high interest to predict the localization of CD 3 + cells based on H&E images without performing CD3 IHC, which may be achieved through image translation. Deep learning methods have been developed to perform image translation from exactly matched image pairs, but it is hard to acquire matched pairs of H&E and CD3 IHC images because they are typically from close but different slices of one tissue. In this paper, we propose a Registration Generative Adversarial Network (R-GAN) to translate H&E images into CD3 IHC images using weakly unpaired training samples. A registration module and a novel patch-level L1loss function are incorporated into Pix2PixGAN to address the challenges arising from weakly unpaired training samples. We conduct benchmarking experiments on a dataset constructed with tissue microarrays (TMAs) of liver cancer tissues, which contains H&E images and the corresponding CD3 IHC images for 1,073 tissue cores. The proposed R-GAN outperforms two commonly used translation methods in terms of both image quality and style relevance metrics. Kuan Huang, Bing Zhang 0003 |
BIBM | 4 |
| 2022 | OmicsEV: a tool for comprehensive quality evaluation of omics data tablesabstractSUMMARY: RNA-Seq and mass spectrometry-based studies generate omics data tables with measurements for tens of thousands of genes across all samples in a study. The success of a study relies on the quality of these data tables, which is determined by both experimental data generation and computational methods used to process raw experimental data into quantitative data tables. We present OmicsEV, an R package for the quality evaluation of omics data tables. For each data table, OmicsEV uses a series of methods to evaluate data depth, data normalization, batch effect, biological signal, platform reproducibility and multi-omics concordance, producing comprehensive visual and quantitative evaluation results that help assess the data quality of individual data tables and facilitate the identification of the optimal data processing method and parameters for the omics study under investigation. AVAILABILITY AND IMPLEMENTATION: The source code and the user manual of OmicsEV are available at https://github.com/bzhanglab/OmicsEV, and the source code is released under the GPL-3 license. Eric J. Jaehnig, Bing Zhang 0003 |
Bioinform. | 3 |
| 2018 | DLAD4U: deriving and prioritizing disease lists from PubMed literatureabstractBACKGROUND: Due to recent technology advancements, disease related knowledge is growing rapidly. It becomes nontrivial to go through all published literature to identify associations between human diseases and genetic, environmental, and life style factors, disease symptoms, and treatment strategies. Here we report DLAD4U (Disease List Automatically Derived For You), an efficient, accurate and easy-to-use disease search engine based on PubMed literature. RESULTS: DLAD4U uses the eSearch and eFetch APIs from the National Center for Biotechnology Information (NCBI) to find publications related to a query and to identify diseases from the retrieved publications. The hypergeometric test was used to prioritize identified diseases for displaying to users. DLAD4U accepts any valid queries for PubMed, and the output results include a ranked disease list, information associated with each disease, chronologically-ordered supporting publications, a summary of the run, and links for file export. DLAD4U outperformed other disease search engines in our comparative evaluation using selected genes and drugs as query terms and manually curated data as "gold standard". For 100 genes that are associated with only one disease in the gold standard, the Mean Average Precision (MAP) measure from DLAD4U was 0.77, which clearly outperformed other tools. For 10 genes that are associated with multiple diseases in the gold standard, the mean precision, recall and F-measure scores from DLAD4U were always higher than those from other tools. The superior performance of DLAD4U was further confirmed using 100 drugs as queries, with an MAP of 0.90. CONCLUSIONS: DLAD4U is a new, intuitive disease search engine that takes advantage of existing resources at NCBI to provide computational efficiency and uses statistical analyses to ensure accuracy. DLAD4U is publicly available at http://dlad4u.zhang-lab.org . Junhui Shen, Suhas V. Vasaikar, Bing Zhang 0003 |
BMC Bioinform. | 3 |
| 2016 | PGA: an R/Bioconductor package for identification of novel peptides using a customized database derived from RNA-SeqabstractBACKGROUND: Peptide identification based upon mass spectrometry (MS) is generally achieved by comparison of the experimental mass spectra with the theoretically digested peptides derived from a reference protein database. Obviously, this strategy could not identify peptide and protein sequences that are absent from a reference database. A customized protein database on the basis of RNA-Seq data is thus proposed to assist with and improve the identification of novel peptides. Correspondingly, development of a comprehensive pipeline, which provides an end-to-end solution for novel peptide detection with the customized protein database, is necessary. RESULTS: A pipeline with an R package, assigned as a PGA utility, was developed that enables automated treatment to the tandem mass spectrometry (MS/MS) data acquired from different MS platforms and construction of customized protein databases based on RNA-Seq data with or without a reference genome guide. Hence, PGA can identify novel peptides and generate an HTML-based report with a visualized interface. On the basis of a published dataset, PGA was employed to identify peptides, resulting in 636 novel peptides, including 510 single amino acid polymorphism (SAP) peptides, 2 INDEL peptides, 49 splice junction peptides, and 75 novel transcript-derived peptides. The software is freely available from http://bioconductor.org/packages/PGA/ , and the example reports are available at http://wenbostar.github.io/PGA/ . CONCLUSIONS: The pipeline of PGA, aimed at being platform-independent and easy-to-use, was successfully developed and shown to be capable of identifying novel peptides by searching the customized protein database derived from RNA-Seq data. Shaohang Xu, Ruo Zhou, Bing Zhang 0003, Xin Liu 0007 |
BMC Bioinform. | 4 |
| 2015 | Empowering biologists with multi-omics data: colorectal cancer as a paradigmabstractMOTIVATION: Recent completion of the global proteomic characterization of The Cancer Genome Atlas (TCGA) colorectal cancer (CRC) cohort resulted in the first tumor dataset with complete molecular measurements at DNA, RNA and protein levels. Using CRC as a paradigm, we describe the application of the NetGestalt framework to provide easy access and interpretation of multi-omics data. RESULTS: The NetGestalt CRC portal includes genomic, epigenomic, transcriptomic, proteomic and clinical data for the TCGA CRC cohort, data from other CRC tumor cohorts and cell lines, and existing knowledge on pathways and networks, giving a total of more than 17 million data points. The portal provides features for data query, upload, visualization and integration. These features can be flexibly combined to serve various needs of the users, maximizing the synergy among omics data, human visualization and quantitative analysis. Using three case studies, we demonstrate that the portal not only provides user-friendly data query and visualization but also enables efficient data integration within a single omics data type, across multiple omics data types, and over biological networks. AVAILABILITY AND IMPLEMENTATION: The NetGestalt CRC portal can be freely accessed at http://www.netgestalt.org. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Zhiao Shi, Bing Zhang 0003 |
Bioinform. | 4 |
| 2013 | customProDB: an R package to generate customized protein databases from RNA-Seq data for proteomics searchabstractUNLABELLED: Database search is the most widely used approach for peptide and protein identification in mass spectrometry-based proteomics studies. Our previous study showed that sample-specific protein databases derived from RNA-Seq data can better approximate the real protein pools in the samples and thus improve protein identification. More importantly, single nucleotide variations, short insertion and deletions and novel junctions identified from RNA-Seq data make protein database more complete and sample-specific. Here, we report an R package customProDB that enables the easy generation of customized databases from RNA-Seq data for proteomics search. This work bridges genomics and proteomics studies and facilitates cross-omics data integration. AVAILABILITY AND IMPLEMENTATION: customProDB and related documents are freely available at http://bioconductor.org/packages/2.13/bioc/html/customProDB.html. Bing Zhang 0003 |
Bioinform. | 2 |
| 2011 | Semi-supervised learning improves gene expression-based prediction of cancer recurrenceabstractMOTIVATION: Gene expression profiling has shown great potential in outcome prediction for different types of cancers. Nevertheless, small sample size remains a bottleneck in obtaining robust and accurate classifiers. Traditional supervised learning techniques can only work with labeled data. Consequently, a large number of microarray data that do not have sufficient follow-up information are disregarded. To fully leverage all of the precious data in public databases, we turned to a semi-supervised learning technique, low density separation (LDS). RESULTS: Using a clinically important question of predicting recurrence risk in colorectal cancer patients, we demonstrated that (i) semi-supervised classification improved prediction accuracy as compared with the state of the art supervised method SVM, (ii) performance gain increased with the number of unlabeled samples, (iii) unlabeled data from different institutes could be employed after appropriate processing and (iv) the LDS method is robust with regard to the number of input features. To test the general applicability of this semi-supervised method, we further applied LDS on human breast cancer datasets and also observed superior performance. Our results demonstrated great potential of semi-supervised learning in gene expression-based outcome prediction for cancer patients. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Mingguang Shi, Bing Zhang 0003 |
Bioinform. | 2 |
| 2011 | Fast network centrality analysis using GPUsabstractBACKGROUND: With the exploding volume of data generated by continuously evolving high-throughput technologies, biological network analysis problems are growing larger in scale and craving for more computational power. General Purpose computation on Graphics Processing Units (GPGPU) provides a cost-effective technology for the study of large-scale biological networks. Designing algorithms that maximize data parallelism is the key in leveraging the power of GPUs. RESULTS: We proposed an efficient data parallel formulation of the All-Pairs Shortest Path problem, which is the key component for shortest path-based centrality computation. A betweenness centrality algorithm built upon this formulation was developed and benchmarked against the most recent GPU-based algorithm. Speedup between 11 to 19% was observed in various simulated scale-free networks. We further designed three algorithms based on this core component to compute closeness centrality, eccentricity centrality and stress centrality. To make all these algorithms available to the research community, we developed a software package gpu-fan (GPU-based Fast Analysis of Networks) for CUDA enabled GPUs. Speedup of 10-50× compared with CPU implementations was observed for simulated scale-free networks and real world biological networks. CONCLUSIONS: gpu-fan provides a significant performance improvement for centrality computation in large-scale networks. Source code is available under the GNU Public License (GPL) at http://bioinfo.vanderbilt.edu/gpu-fan/. Zhiao Shi, Bing Zhang 0003 |
BMC Bioinform. | 2 |
| 2010 | WebGestalt2: an updated and expanded version of the Web-based Gene Set Analysis ToolkitabstractBackground Functional enrichment analysis has become a popular approach for the interpretation of gene lists derived from large scale genomic, transcriptomic, and proteomic studies. Many software tools have been developed for this analysis, including our previously published Web-based Gene Set Enrichment Analysis Toolkit (WebGestalt) [1]. Major differences among existing tools include: 1) the statistical test used for the enrichment analysis; 2) supported organisms; 3) supported input ID types; 4) coverage of functional categories; and 5) presentation of the output results. We have made improvements in the above areas in the updated and expanded version, WebGestalt2. Dexter T. Duncan, Naresh Prodduturi, Bing Zhang 0003 |
BMC Bioinform. | 3 |
| 2010 | Enabling proteomics-based identification of human cancer variationsabstractShotgun proteomics is a powerful technology for protein identification in complex samples with remarkable applications in elucidating cellular and subcellular proteomes [ 1 , 2 ], and discovering disease biomarkers [ 3 , 4 ]. Shotgun proteomics data analysis usually relies on database search. Commonly used protein sequence databases in shotgun proteomics data analysis do not contain mutation information. This becomes a problem in cancer studies in which the detection of disease-related mutated peptides/proteins is crucial for understanding cancer biology [ 5 ]. Including protein mutation information into sequence databases can help alleviate this problem. Based on the human Cancer Proteome Variation Database developed by us recently [ 6 ], which comprises 41,541 nonsynonymous SNPs in 30,322 proteins from the dbSNP database and around 9000 cancer-related variations in 2,921 proteins, we created a variation-containing protein sequence database and a data analysis workflow for mutant protein identification in shotgun proteomics (Figure 1 ). Applying this workflow on colorectal cancer cell lines identified many peptides that contain either non-cancer-specific or very important cancer-related variations, such as a known somatic mutation in K-Ras in HCT116 cell line. Our workflow for mutant peptide identification has been tested for compatibility with various popular database search engines including Sequest, Mascot, X!Tandom as well as MyriMatch. Architecture for identifying mutant peptides from cancer shotgun proteome data Owing to its protein-centric nature, the approach we proposed can serve as a bridge between genomic variation data and proteomics studies in human cancer. Zeqiang Ma, Robbert J. C. Slebos, David L. Tabb, Daniel C. Liebler, Bing Zhang 0003 |
BMC Bioinform. | 6 |
| 2008 | The gene-function relationship in the metabolism of yeast and digital organisms
Philip Gerlee, Torbjörn Lundh, Bing Zhang 0003, Alexander R. A. Anderson |
ALIFE | 3 |
| 2008 | Supervised principal component analysis for gene set enrichment of microarray data with continuous or survival outcomesabstractMOTIVATION: Gene set analysis allows formal testing of subtle but coordinated changes in a group of genes, such as those defined by Gene Ontology (GO) or KEGG Pathway databases. We propose a new method for gene set analysis that is based on principal component analysis (PCA) of genes expression values in the gene set. PCA is an effective method for reducing high dimensionality and capture variations in gene expression values. However, one limitation with PCA is that the latent variable identified by the first PC may be unrelated to outcome. RESULTS: In the proposed supervised PCA (SPCA) model for gene set analysis, the PCs are estimated from a selected subset of genes that are associated with outcome. As outcome information is used in the gene selection step, this method is supervised, thus called the Supervised PCA model. Because of the gene selection step, test statistic in SPCA model can no longer be approximated well using t-distribution. We propose a two-component mixture distribution based on Gumbel exteme value distributions to account for the gene selection step. We show the proposed method compares favorably to currently available gene set analysis methods using simulated and real microarray data. SOFTWARE: The R code for the analysis used in this article are available upon request, we are currently working on implementing the proposed method in an R package. Lily Wang 0001, Jonathan D. Smith, Bing Zhang 0003 |
Bioinform. | 4 |
| 2008 | From pull-down data to protein interaction networks and complexes with biological relevanceabstractAbstract Motivation: Recent improvements in high-throughput Mass Spectrometry (MS) technology have expedited genome-wide discovery of protein–protein interactions by providing a capability of detecting protein complexes in a physiological setting. Computational inference of protein interaction networks and protein complexes from MS data are challenging. Advances are required in developing robust and seamlessly integrated procedures for assessment of protein–protein interaction affinities, mathematical representation of protein interaction networks, discovery of protein complexes and evaluation of their biological relevance. Results: A multi-step but easy-to-follow framework for identifying protein complexes from MS pull-down data is introduced. It assesses interaction affinity between two proteins based on similarity of their co-purification patterns derived from MS data. It constructs a protein interaction network by adopting a knowledge-guided threshold selection method. Based on the network, it identifies protein complexes and infers their core components using a graph-theoretical approach. It deploys a statistical evaluation procedure to assess biological relevance of each found complex. On Saccharomyces cerevisiae pull-down data, the framework outperformed other more complicated schemes by at least 10% in F1-measure and identified 610 protein complexes with high-functional homogeneity based on the enrichment in Gene Ontology (GO) annotation. Manual examination of the complexes brought forward the hypotheses on cause of false identifications. Namely, co-purification of different protein complexes as mediated by a common non-protein molecule, such as DNA, might be a source of false positives. Protein identification bias in pull-down technology, such as the hydrophilic bias could result in false negatives. Contact: [email protected] Supplementary information: Supplementary data are available at Bioinformatics online. Bing Zhang 0003, Tatiana V. Karpinets, Nagiza F. Samatova |
Bioinform. | 1 |
| 2007 | Multi-stage Framework to Infer Protein Functional Modules from Mass Spectrometry Pull-Down Data with Assessment of Biological RelevanceabstractProtein functional modules are fundamental units in protein interaction networks. High-throughput Mass Spectrometry (MS) technology has become valuable for discovery of protein functional modules. Yet, their computational inference from MS pull-down data and biological significance evaluation are still challenging. This paper introduces an integrated multi-step framework for (1) assessing protein-protein interaction affinities, (2) constructing a genome-wide protein association map, (3) finding putative protein functional modules, and (4) evaluating their biological relevance. The protein affinity score utilizes co- purification pattern of two proteins and adopts an information theoretic-approach to build the protein affinity map. Putative protein modules are then derived using a graph-theoretical approach. A two-stage statistical procedure assesses biological relevance of identified modules. On Saccharomyces cerevisiae's pull-down data (Nature, vol. 415, pp. 141-7, 2002), the scoring scheme outperformed other methods by at least 10% in F1-measure, and statistical tests identified 489 protein modules enriched in all of three general GO categories with p-values less than 0.05. Bing Zhang 0003, Tatiana V. Karpinets, Nagiza F. Samatova |
BIBM | 2 |
| 2005 | GeneKeyDB: A lightweight, gene-centric, relational database to support data mining environmentsabstractBACKGROUND: The analysis of biological data is greatly enhanced by existing or emerging databases. Most existing databases, with few exceptions are not designed to easily support large scale computational analysis, but rather offer exclusively a web interface to the resource. We have recognized the growing need for a database which can be used successfully as a backend to computational analysis tools and pipelines. Such database should be sufficiently versatile to allow easy system integration. RESULTS: GeneKeyDB is a gene-centered relational database developed to enhance data mining in biological data sets. The system provides an underlying data layer for computational analysis tools and visualization tools. GeneKeyDB relies primarily on existing database identifiers derived from community databases (NCBI, GO, Ensembl, et al.) as well as the known relationships among those identifiers. It is a lightweight, portable, and extensible platform for integration with computational tools and analysis environments. CONCLUSION: GeneKeyDB can enable analysis tools and users to manipulate the intersections, unions, and differences among different data sets. S. A. Kirov, Xinxia Peng, E. Baker, Denise Schmoyer, Bing Zhang 0003, Jay Snoddy |
BMC Bioinform. | 5 |
| 2004 | GOTree Machine (GOTM): a web-based platform for interpreting sets of interesting genes using Gene Ontology hierarchiesabstractBACKGROUND: Microarray and other high-throughput technologies are producing large sets of interesting genes that are difficult to analyze directly. Bioinformatics tools are needed to interpret the functional information in the gene sets. RESULTS: We have created a web-based tool for data analysis and data visualization for sets of genes called GOTree Machine (GOTM). This tool was originally intended to analyze sets of co-regulated genes identified from microarray analysis but is adaptable for use with other gene sets from other high-throughput analyses. GOTree Machine generates a GOTree, a tree-like structure to navigate the Gene Ontology Directed Acyclic Graph for input gene sets. This system provides user friendly data navigation and visualization. Statistical analysis helps users to identify the most important Gene Ontology categories for the input gene sets and suggests biological areas that warrant further study. GOTree Machine is available online at http://genereg.ornl.gov/gotm/. CONCLUSION: GOTree Machine has a broad application in functional genomic, proteomic and other high-throughput methods that generate large sets of interesting genes; its primary purpose is to help users sort for interesting patterns in gene sets. Bing Zhang 0003, Denise Schmoyer, Stefan Kirov, Jay Snoddy |
BMC Bioinform. | 1 |