Zuguang Gu

dblp:126/7070 · DBLP profile ↗
← Back
12ranked-venue papers
11as first author
7since 2021 · last 2024
0000-0002-7395-8709ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 11 · 10 first-author · 6 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2024 biotextgraph: graphical summarization of functional similarities from textual information
abstract
SUMMARY: Functional interpretation of biological entities such as differentially expressed genes is one of the fundamental analyses in bioinformatics. The task can be addressed by using biological pathway databases with enrichment analysis (EA). However, textual description of biological entities in public databases is less explored and integrated in existing tools and it has a potential to reveal new mechanisms. Here, we present a new R package biotextgraph for graphical summarization of omics' textual description data which enables assessment of functional similarities of the lists of biological entities. We illustrate application examples of annotating gene identifiers in addition to EA. The results suggest that the visualization based on words and inspection of biological entities with text can reveal a set of biologically meaningful terms that could not be obtained by using biological pathway databases alone. The results suggest the usefulness of the package in the routine analysis of omics-related data. The package also offers a web-based application for convenient querying. AVAILABILITY AND IMPLEMENTATION: The package, documentation, and web server are available at: https://github.com/noriakis/biotextgraph.
Noriaki Sato, Zuguang Gu, Seiya Imoto
Bioinform.3
2023 rGREAT: an R/bioconductor package for functional enrichment on genomic regions
abstract
SUMMARY: GREAT (Genomic Regions Enrichment of Annotations Tool) is a widely used tool for functional enrichment on genomic regions. However, as an online tool, it has limitations of outdated annotation data, small numbers of supported organisms and gene set collections, and not being extensible for users. Here, we developed a new R/Bioconductorpackage named rGREAT which implements the GREAT algorithm locally. rGREAT by default supports more than 600 organisms and a large number of gene set collections, as well as self-provided gene sets and organisms from users. Additionally, it implements a general method for dealing with background regions. AVAILABILITY AND IMPLEMENTATION: The package rGREAT is freely available from the Bioconductor project: https://bioconductor.org/packages/rGREAT/. The development version is available at https://github.com/jokergoo/rGREAT. Gene Ontology gene sets for more than 600 organisms retrieved from Ensembl BioMart are presented in an R package BioMartGOGeneSets which is available at https://github.com/jokergoo/BioMartGOGeneSets. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Zuguang Gu, Daniel Hübschmann
Bioinform.1
2023 On the dependency heaviness of CRAN/Bioconductor ecosystem
Zuguang Gu
J. Syst. Softw.1
2022 Improve consensus partitioning via a hierarchical procedure
abstract
Consensus partitioning is an unsupervised method widely used in high-throughput data analysis for revealing subgroups and assigning stability for the classification. However, standard consensus partitioning procedures are weak for identifying large numbers of stable subgroups. There are two major issues. First, subgroups with small differences are difficult to be separated if they are simultaneously detected with subgroups with large differences. Second, stability of classification generally decreases as the number of subgroups increases. In this work, we proposed a new strategy to solve these two issues by applying consensus partitioning in a hierarchical procedure. We demonstrated hierarchical consensus partitioning can be efficient to reveal more meaningful subgroups. We also tested the performance of hierarchical consensus partitioning on revealing a great number of subgroups with a large deoxyribonucleic acid methylation dataset. The hierarchical consensus partitioning is implemented in the R package cola with comprehensive functionalities for analysis and visualization. It can also automate the analysis only with a minimum of two lines of code, which generates a detailed HTML report containing the complete analysis. The cola package is available at https://bioconductor.org/packages/cola/.
Zuguang Gu, Daniel Hübschmann
Briefings Bioinform.1
2022 spiralize: an R package for visualizing data on spirals
abstract
SUMMARY: Spiral layout has two major advantages for data visualization. First, it is able to visualize data with long axes, which greatly improves the resolution of visualization. Second, it is efficient for time series data to reveal periodic patterns. Here, we present the R package spiralize that provides a general solution for visualizing data on spirals. spiralize implements numerous graphics functions so that self-defined high-level graphics can be easily implemented by users. The flexibility and power of spiralize are demonstrated by five examples from real-world datasets. AVAILABILITY AND IMPLEMENTATION: The spiralize package and documentations are freely available at the Comprehensive R Archive Network (CRAN) https://CRAN.R-project.org/package=spiralize. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Zuguang Gu, Daniel Hübschmann
Bioinform.1
2022 Make Interactive Complex Heatmaps in R
abstract
SUMMARY: Heatmap is a powerful visualization method on two-dimensional data to reveal patterns shared by subsets of rows and columns. In this work, we introduce a new R package InteractiveComplexHeatmap that brings interactivity to the widely used ComplexHeatmap package. InteractiveComplexHeatmap is designed with an easy-to-use interface where static complex heatmaps can be directly exported to an interactive Shiny web application only with one additional line of code. InteractiveComplexHeatmap also provides flexible functionalities for integrating interactive heatmap widgets to build more complex and customized Shiny web applications. AVAILABILITY AND IMPLEMENTATION: The InteractiveComplexHeatmap package and documentations are freely available from the Bioconductor project: https://bioconductor.org/packages/InteractiveComplexHeatmap/. A complete and printer-friendly version of the documentation can also be found in Supplementary File S1. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Zuguang Gu, Daniel Hübschmann
Bioinform.1
2022 Pkgndep: a tool for analyzing dependency heaviness of R packages
abstract
SUMMARY: Numerous R packages have been developed for bioinformatics analysis in the last decade and dependencies among packages have become critical issues to consider. In this work, we proposed a new metric named dependency heaviness that measures the number of dependencies that a parent uniquely brings to a package and we proposed possible solutions for reducing the complexity of dependencies by optimizing the use of heavy parents. We implemented the metric in a new R package pkgndep which provides an intuitive way for dependency heaviness analysis. Based on pkgndep, we additionally performed a global analysis of dependency heaviness on CRAN and Bioconductor ecosystems and we revealed top packages that have significant contributions of high dependency heaviness to their child packages. AVAILABILITY AND IMPLEMENTATION: The package pkgndep and documentations are freely available from the Comprehensive R Archive Network https://cran.r-project.org/package=pkgndep. The dependency heaviness analysis for all 22 076 CRAN and Bioconductor packages retrieved on June 8, 2022 are available at https://pkgndep.github.io/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Zuguang Gu, Daniel Hübschmann
Bioinform.1
2016 HilbertCurve: an R/Bioconductor package for high-resolution visualization of genomic data
abstract
UNLABELLED: : Hilbert curves enable high-resolution visualization of genomic data on a chromosome- or genome-wide scale. Here we present the HilbertCurve package that provides an easy-to-use interface for mapping genomic data to Hilbert curves. The package transforms the curve as a virtual axis, thereby hiding the details of the curve construction from the user. HilbertCurve supports multiple-layer overlay that makes it a powerful tool to correlate the spatial distribution of multiple feature types. AVAILABILITY AND IMPLEMENTATION: The HilbertCurve package and documentation are freely available from the Bioconductor project: http://www.bioconductor.org/packages/devel/bioc/html/HilbertCurve.html CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Zuguang Gu, Roland Eils, Matthias Schlesner
Bioinform.1
2016 Complex heatmaps reveal patterns and correlations in multidimensional genomic data
abstract
UNLABELLED: Parallel heatmaps with carefully designed annotation graphics are powerful for efficient visualization of patterns and relationships among high dimensional genomic data. Here we present the ComplexHeatmap package that provides rich functionalities for customizing heatmaps, arranging multiple parallel heatmaps and including user-defined annotation graphics. We demonstrate the power of ComplexHeatmap to easily reveal patterns and correlations among multiple sources of information with four real-world datasets. AVAILABILITY AND IMPLEMENTATION: The ComplexHeatmap package and documentation are freely available from the Bioconductor project: http://www.bioconductor.org/packages/devel/bioc/html/ComplexHeatmap.html CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Zuguang Gu, Roland Eils, Matthias Schlesner
Bioinform.1
2016 gtrellis: an R/Bioconductor package for making genome-level Trellis graphics
abstract
BACKGROUND: Trellis graphics are a visualization method that splits data by one or more categorical variables and displays subsets of the data in a grid of panels. Trellis graphics are broadly used in genomic data analysis to compare statistics over different categories in parallel and reveal multivariate relationships. However, current software packages to produce Trellis graphics have not been designed with genomic data in mind and lack some functionality that is required for effective visualization of genomic data. RESULTS: Here we introduce the gtrellis package which provides an efficient and extensible way to visualize genomic data in a Trellis layout. gtrellis provides highly flexible Trellis layouts which allow efficient arrangement of genomic categories on the plot. It supports multiple-track visualization, which makes it straightforward to visualize several properties of genomic data in parallel to explain complex relationships. In addition, gtrellis provides an extensible framework that allows adding user-defined graphics. CONCLUSIONS: The gtrellis package provides an easy and effective way to visualize genomic data and reveal high dimensional relationships on a genome-wide scale. gtrellis can be flexibly extended and thus can also serve as a base package for highly specific purposes. gtrellis makes it easy to produce novel visualizations, which can lead to the discovery of previously unrecognized patterns in genomic data.
Zuguang Gu, Roland Eils, Matthias Schlesner
BMC Bioinform.1
2014 circlize implements and enhances circular visualization in R
abstract
SUMMARY: Circular layout is an efficient way for the visualization of huge amounts of genomic information. Here we present the circlize package, which provides an implementation of circular layout generation in R as well as an enhancement of available software. The flexibility of this package is based on the usage of low-level graphics functions such that self-defined high-level graphics can be easily implemented by users for specific purposes. Together with the seamless connection between the powerful computational and visual environment in R, circlize gives users more convenience and freedom to design figures for better understanding genomic patterns behind multi-dimensional data. AVAILABILITY AND IMPLEMENTATION: circlize is available at the Comprehensive R Archive Network (CRAN): http://cran.r-project.org/web/packages/circlize/
Zuguang Gu, Roland Eils, Matthias Schlesner, Benedikt Brors
Bioinform.1
2013 CePa: an R package for finding significant pathways weighted by multiple network centralities
abstract
SUMMARY: CePa is an R package aiming to find significant pathways through network topology information. The package has several advantages compared with current pathway enrichment tools. First, pathway node instead of single gene is taken as the basic unit when analysing networks to meet the fact that genes must be constructed into complexes to hold normal functions. Second, multiple network centralities are applied simultaneously to measure importance of nodes from different aspects to make a full view on the biological system. CePa extends standard pathway enrichment methods, which include both over-representation analysis procedure and gene-set analysis procedure. CePa has been evaluated with high performance on real-world data, and it can provide more information directly related to current biological problems. AVAILABILITY: CePa is available at the Comprehensive R Archive Network (CRAN): http://cran.r-project.org/web/packages/CePa/
Zuguang Gu
Bioinform.1