EDBT 2026 Demo / reviewers in the wild / expert
Yongzhuang Liu
dblp:147/6947
· DBLP profile ↗
22ranked-venue papers
7as first author
13since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 22 · 7 first-author · 13 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | SSD-MIR: Towards Medical Image Restoration With Structured Visual State Space DualityabstractIn the field of Medical image processing, Transformers and CNNs bring filter-based local sensing and attention-based global modelling, which strengthened their dominant position. However, the scales required in different applications vary, and improper usage or selection on models can lead to performance decrease, resulting in more serious consequences in downstream tasks. The recent advanced Structured State Space Model has shown its potential as a baseline model, and is further proved to maintaining implicit duality with the attention mechanism, The corresponding model Mamba2 shows its long-range modelling ability with optimized long-term memory and linear computational complexity.In this paper, we propose SSD-MIR, a novel architecture leveraging the advantages of state space duality mechanism. SSD-MIR includes three parts: input projection and feature representation, deep feature transformation, output projection and feature inversion. With the three components sequentially connected to form our model, it is able to model long-range dependencies and learn the mapping between low-quality input and high-quality output simultaneously. In the main deep feature transformation part, depthwise convolution is used for additional feature mapping as an alternative of 1D causal convolution, and an optimized quad-direction scanning mechanism is implemented to form the overall 2D state space dual module. The corresponding wrapper module is built in MetaFormer style with the usage of residual connection and LayerNorm. We perform extensive experiments on multiple medical image restoration tasks, and the results are superior to the state-of-the-arts, demonstrating the strength and efficacy of our model. Binghong Chen, Zhengqian Zhang, Tianshuo Yu, Junjun Ren, Yongzhuang Liu |
BIBM | 8 |
| 2024 | Robust 3D reconstruction and cell-type deconvolution approach for spatial transcriptomics by hybrid graph learning and iterative computationabstractSpatial transcriptomics data provides new methods of exploring gene expressions. However, the results usually do not contain cell-level proportions, leaving a gap between the new technologies and the corresponding analysis. To address the issues, we propose a novel framework, OptiGraph3D, in this paper. Our approach starts with a pivoted sample-level sequential registration, eliminating potential cumulative errors in multiple registrations. The registration results provide sufficient spatial prior for our cell-type deconvolving process, which leverages the power of ADMM method by modelling the cell-type deconvolution task as a matrix inverse problem and proposes efficient solvers with guaranteed convergence. Furthermore, the cell-type deconvolution results can in turn perform as supervision for cluster recognition in 3D reconstruction. Our model introduces a semi-supervision training manner by introducing a specific class token in the multi-headed graph attention network, bridging the gap between gene-level expressions of spatial transcriptomics and presenting new angle for sample-wise analysis. Extensive experiments are performed on the human dorsolateral prefrontal cortex (DLPFC) and the developing human embryonic heart dataset, suggesting our approach's superiority over other methods. Zhengqian Zhang, Binghong Chen, Tianshuo Yu, Junjun Ren, Chongyao Liu, Yongzhuang Liu |
BIBM | 10 |
| 2022 | Enhancing discoveries of molecular QTL studies with small sample size using summary statistic imputationabstractQuantitative trait locus (QTL) analyses of multiomic molecular traits, such as gene transcription (eQTL), DNA methylation (mQTL) and histone modification (haQTL), have been widely used to infer the functional effects of genome variants. However, the QTL discovery is largely restricted by the limited study sample size, which demands higher threshold of minor allele frequency and then causes heavy missing molecular trait-variant associations. This happens prominently in single-cell level molecular QTL studies because of sample availability and cost. It is urgent to propose a method to solve this problem in order to enhance discoveries of current molecular QTL studies with small sample size. In this study, we presented an efficient computational framework called xQTLImp to impute missing molecular QTL associations. In the local-region imputation, xQTLImp uses multivariate Gaussian model to impute the missing associations by leveraging known association statistics of variants and the linkage disequilibrium (LD) around. In the genome-wide imputation, novel procedures are implemented to improve efficiency, including dynamically constructing a reused LD buffer, adopting multiple heuristic strategies and parallel computing. Experiments on various multiomic bulk and single-cell sequencing-based QTL datasets have demonstrated high imputation accuracy and novel QTL discovery ability of xQTLImp. Finally, a C++ software package is freely available at https://github.com/stormlovetao/QTLIMP. Tao Wang 0082, Yongzhuang Liu, Quanwei Yin, Jiaquan Geng, Jin Chen 0004, Xipeng Yin, Yongtian Wang, Xuequn Shang 0001, Chunwei Tian, Yadong Wang 0001, Jiajie Peng |
Briefings Bioinform. | 2 |
| 2022 | Correction to: Enhancing discoveries of molecular QTL studies with small sample size using summary statistic imputationabstractIn the originally published version of this manuscript, there was an error in the Funding section; ‘National Natural Science Foundation of China (6210071334, 62072376)’ has now been corrected to ‘National Natural Science Foundation of China (62102319, 62072376)’. Tao Wang 0082, Yongzhuang Liu, Quanwei Yin, Jiaquan Geng, Jin Chen 0004, Xipeng Yin, Yongtian Wang, Xuequn Shang 0001, Chunwei Tian, Yadong Wang 0001, Jiajie Peng |
Briefings Bioinform. | 2 |
| 2022 | Effective Identification of Bacterial Genomes From Short and Long Read Sequencing DataabstractWith the development of sequencing technology, microbiological genome sequencing analysis has attracted extensive attention. For inexperienced users without sufficient bioinformatics skills, making sense of sequencing data for microbial identification, especially for bacterial identification, through reads analysis is still challenging. In order to address the challenge of effectively analyzing genomic information, in this paper, we develop an effective approach and automatic bioinformatics pipeline called PBGI for bacterial genome identification, performing automatedly and customized bioinformatics analysis using short-reads or long-reads sequencing data produced by multiple platforms such as Illumina, PacBio and Oxford Nanopore. An evaluation of the proposed approach on the practical data set is presented, showing that PBGI provides a user-friendly way to perform bacterial identification through short or long reads analysis, and could provide accurate analyzing results. The source code of the PBGI is freely available at https://github.com/lyotvincent/PBGI. Jian Liu 0028, Yongzhuang Liu |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2021 | Ontology-based annotation and retrieval for large-scale VCF dataabstractSequencing cost is dramatically reduced by the development of the next-generation sequencing (NGS) technologies. Currently, numerous variant call format (VCF) data and biomedical ontologies, which store mutations data and special biomedical knowledge to applications in the field of biomedical researches such as human genetics, etc., become available in the bioinformatics community. There are some bioinformatics tools developed for the VCF data annotation and analysis. However, most previous works ignore the biomedical ontologies associated with the genetic data, which are usually beneficial to analyze genetic diseases and molecular diagnosis. In particular, annotating information with biomedical ontologies remains an obstacle. In order to effectively integrate biomedical ontologies and enhance the analysis across multiple biomedical sources, we present an automatic workflow called OntoAnnotation for annotating VCF files with biomedical ontologies. Additionally, to facilitate the retrieval of large-scale VCF data for non-bioinformaticians, we develop a web platform called OntoVarSearch, which provides a flexible engine that allows convenient access to genetic variants and ontology-based annotation information stored in the MongoDB database. The OntoAnnotation tool and the OntoVarSearch platform could provide a simple way for users without sufficient programming skills to annotate information with biomedical ontologies and search data stored in VCF files. Jian Liu 0040, Yongzhuang Liu |
BIBM | 5 |
| 2021 | PocaCNV: A Tool to Detect Copy Number Variants from Population-Scale Genome Sequencing DataabstractAmong the procedures of modern genome analysis, the variation calling is a crucial part. Copy Number Variation(CNV) is an important variation type, many algorithms are developed to detect them. As the scale of samples in genome projects growing, it’s wise to use data from multiple samples jointly to call variants. We present PocaCNV, a population caller of CNVs that can call CNVs from multiple samples of the same population. PocaCNV can produce many unique results with good precision and sensitivity, and can be a good supplement to other modern SV calling procedures. PocaCNV can be openly accessed from https://github.com/CoREse/PocaCNV. Yongzhuang Liu, Yadong Wang 0001 |
BIBM | 2 |
| 2021 | A deep learning approach for filtering structural variants in short read sequencing dataabstractShort read whole genome sequencing has become widely used to detect structural variants in human genetic studies and clinical practices. However, accurate detection of structural variants is a challenging task. Especially existing structural variant detection approaches produce a large proportion of incorrect calls, so effective structural variant filtering approaches are urgently needed. In this study, we propose a novel deep learning-based approach, DeepSVFilter, for filtering structural variants in short read whole genome sequencing data. DeepSVFilter encodes structural variant signals in the read alignments as images and adopts the transfer learning with pre-trained convolutional neural networks as the classification models, which are trained on the well-characterized samples with known high confidence structural variants. We use two well-characterized samples to demonstrate DeepSVFilter's performance and its filtering effect coupled with commonly used structural variant detection approaches. The software DeepSVFilter is implemented using Python and freely available from the website at https://github.com/yongzhuang/DeepSVFilter. Yongzhuang Liu, Yalin Huang, Guohua Wang 0001, Yadong Wang 0001 |
Briefings Bioinform. | 1 |
| 2021 | An integrated approach for copy number variation discovery in parent-offspring triosabstractWhole-genome sequencing (WGS) of parent-offspring trios has become widely used to identify causal copy number variations (CNVs) in rare and complex diseases. Existing CNV detection approaches usually do not make effective use of Mendelian inheritance in parent-offspring trios and yield low accuracy. In this study, we propose a novel integrated approach, TrioCNV2, for jointly detecting CNVs from WGS data of the parent-offspring trio. TrioCNV2 first makes use of the read depth and discordant read pairs to infer approximate locations of CNVs and then employs the split read and local de novo assembly approaches to refine the breakpoints. We use the real WGS data of two parent-offspring trios to demonstrate TrioCNV2's performance and compare it with other CNV detection approaches. The software TrioCNV2 is implemented using a combination of Java and R and is freely available from the website at https://github.com/yongzhuang/TrioCNV2. Yongzhuang Liu, Yadong Wang 0001 |
Briefings Bioinform. | 1 |
| 2021 | abPOA: an SIMD-based C library for fast partial order alignment using adaptive bandabstractSUMMARY: Partial order alignment, which aligns a sequence to a directed acyclic graph, is now frequently used as a key component in long-read error correction and assembly. We present abPOA (adaptive banded Partial Order Alignment), a Single Instruction Multiple Data (SIMD)-based C library for fast partial order alignment using adaptive banded dynamic programming. It can work as a stand-alone multiple sequence alignment and consensus calling tool or be easily integrated into any long-read error correction and assembly workflow. Compared to a state-of-the-art tool (SPOA), abPOA is up to 10 times faster with a comparable alignment accuracy. AVAILABILITY AND IMPLEMENTATION: abPOA is implemented in C. A stand-alone tool and a C/Python software interface are freely available at https://github.com/yangao07/abPOA. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yan Gao 0031, Yongzhuang Liu, Yanmei Ma, Bo Liu 0023, Yadong Wang 0001, Yi Xing |
Bioinform. | 2 |
| 2021 | Erratum to: abPOA: an SIMD-based C library for fast partial order alignment using adaptive bandabstractBioinformatics, 2021, 37(15), 2209–2211. doi: https://doi.org/10.1093/bioinformatics/btaa963 Upon the original publication of this article, the affiliation: “Center for Bioinformatics, Department of Computer Science and Technology, Harbin Institute of Technology, Harbin, Heilongjiang 150001, China” was inadvertently omitted prior to publication. The affiliation has been reinstated. Yan Gao 0031, Yongzhuang Liu, Yanmei Ma, Bo Liu 0023, Yadong Wang 0001, Yi Xing |
Bioinform. | 2 |
| 2021 | A pipeline for RNA-seq based eQTL analysis with automated quality control proceduresabstractBACKGROUND: Advances in the expression quantitative trait loci (eQTL) studies have provided valuable insights into the mechanism of diseases and traits-associated genetic variants. However, it remains challenging to evaluate and control the quality of multi-source heterogeneous eQTL raw data for researchers with limited computational background. There is an urgent need to develop a powerful and user-friendly tool to automatically process the raw datasets in various formats and perform the eQTL mapping afterward. RESULTS: In this work, we present a pipeline for eQTL analysis, termed eQTLQC, featured with automated data preprocessing for both genotype data and gene expression data. Our pipeline provides a set of quality control and normalization approaches, and utilizes automated techniques to reduce manual intervention. We demonstrate the utility and robustness of this pipeline by performing eQTL case studies using multiple independent real-world datasets with RNA-seq data and whole genome sequencing (WGS) based genotype data. CONCLUSIONS: eQTLQC provides a reliable computational workflow for eQTL analysis. It provides standard quality control and normalization as well as eQTL mapping procedures for eQTL raw data in multiple formats. The source code, demo data, and instructions are freely available at https://github.com/stormlovetao/eQTLQC . Tao Wang 0082, Yongzhuang Liu, Junpeng Ruan, Xianjun Dong, Yadong Wang 0001, Jiajie Peng |
BMC Bioinform. | 2 |
| 2021 | Effective Identification and Annotation of Fungal Genomes
Yongzhuang Liu |
J. Comput. Sci. Technol. | 3 |
| 2020 | Filtering de novo indels in parent-offspring triosabstractBACKGROUND: Identification of de novo indels from whole genome or exome sequencing data of parent-offspring trios is a challenging task in human disease studies and clinical practices. Existing computational approaches usually yield high false positive rate. RESULTS: In this study, we developed a gradient boosting approach for filtering de novo indels obtained by any computational approaches. Through application on the real genome sequencing data, our approach showed it could significantly reduce the false positive rate of de novo indels without a significant compromise on sensitivity. CONCLUSIONS: The software DNMFilter_Indel was written in a combination of Java and R and freely available from the website at https://github.com/yongzhuang/DNMFilter_Indel . Yongzhuang Liu, Yadong Wang 0001 |
BMC Bioinform. | 1 |
| 2020 | Enabling Massive XML-Based Biological Data Management in HBaseabstractPublishing biological data in XML formats is attractive for organizations who would like to provide their bioinformatics resources in an extensible and machine-readable format. In the era of big data, massive XML-based biological data management is emerged as a challengeable issue. With the continuous growth of the XML-based biological data sets, it is usually frustrating to use traditional declarative query languages to provide efficient query capabilities in terms of processing speed and scale. In this study, we report a novel platform to store and query massive XML-based biological data collections. A prototype tool for constructing HBase tables from XML-based biological data collections is first developed, and then a formal approach to transform the XML query model into the MapReduce query model is proposed. Finally, an evaluation of the query performance of the proposed approach on the existing XML-based biological databases is presented, showing that the performance advantages of the proposed solution. The source code of the massive XML-based biological data management platform is freely available at https://github.com/lyotvincent/X2H. Jian Liu 0028, Qiuru Liu, Shuhui Su, Yongzhuang Liu |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2019 | DNMFilter_Indel: Filtering de novo Indels in Parent-Offspring TriosabstractIdentification of de novo indels from whole genome or exome sequencing data of parent-offspring trios is a challenging task in human disease studies and clinical practices. Existing computational approaches usually yield high false positive rate. In this study, we developed a gradient boosting approach for filtering de novo indels obtained by any computational approaches. Through application on the real genome sequencing data, our approach showed it could significantly reduce the false positive rate of de novo indels without a significant compromise on sensitivity. The software DNMFilter_Indel was written in a combination of Java and R and freely available from the website at https://github.com/yongzhuang/DNMFilter_Indel. Yongzhuang Liu, Yadong Wang 0001 |
BIBM | 1 |
| 2019 | Joint detection of germline and somatic copy number events in matched tumor-normal sample pairsabstractMOTIVATION: Whole-genome sequencing (WGS) of tumor-normal sample pairs is a powerful approach for comprehensively characterizing germline copy number variations (CNVs) and somatic copy number alterations (SCNAs) in cancer research and clinical practice. Existing computational approaches for detecting copy number events cannot detect germline CNVs and SCNAs simultaneously, and yield low accuracy for SCNAs. RESULTS: In this study, we developed TumorCNV, a novel approach for jointly detecting germline CNVs and SCNAs from WGS data of the matched tumor-normal sample pair. We compared TumorCNV with existing copy number event detection approaches using the simulated data and real data for the COLO-829 melanoma cell line. The experimental results showed that TumorCNV achieved superior performance than existing approaches. AVAILABILITY AND IMPLEMENTATION: The software TumorCNV is implemented using a combination of Java and R, and it is freely available from the website at https://github.com/yongzhuang/TumorCNV. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yongzhuang Liu, Yadong Wang 0001 |
Bioinform. | 1 |
| 2016 | Joint detection of copy number variations in parent-offspring triosabstractMOTIVATION: Whole genome sequencing (WGS) of parent-offspring trios is a powerful approach for identifying disease-associated genes via detecting copy number variations (CNVs). Existing approaches, which detect CNVs for each individual in a trio independently, usually yield low-detection accuracy. Joint modeling approaches leveraging Mendelian transmission within the parent-offspring trio can be an efficient strategy to improve CNV detection accuracy. RESULTS: In this study, we developed TrioCNV, a novel approach for jointly detecting CNVs in parent-offspring trios from WGS data. Using negative binomial regression, we modeled the read depth signal while considering both GC content bias and mappability bias. Moreover, we incorporated the family relationship and used a hidden Markov model to jointly infer CNVs for three samples of a parent-offspring trio. Through application to both simulated data and a trio from 1000 Genomes Project, we showed that TrioCNV achieved superior performance than existing approaches. AVAILABILITY AND IMPLEMENTATION: The software TrioCNV implemented using a combination of Java and R is freely available from the website at https://github.com/yongzhuang/TrioCNV CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yongzhuang Liu, Jianguo Lu, Jiajie Peng, Liran Juan, Xiaolin Zhu 0002, Bingshan Li, Yadong Wang 0001 |
Bioinform. | 1 |
| 2015 | Family genome browser: visualizing genomes with pedigree informationabstractMOTIVATION: Families with inherited diseases are widely used in Mendelian/complex disease studies. Owing to the advances in high-throughput sequencing technologies, family genome sequencing becomes more and more prevalent. Visualizing family genomes can greatly facilitate human genetics studies and personalized medicine. However, due to the complex genetic relationships and high similarities among genomes of consanguineous family members, family genomes are difficult to be visualized in traditional genome visualization framework. How to visualize the family genome variants and their functions with integrated pedigree information remains a critical challenge. RESULTS: We developed the Family Genome Browser (FGB) to provide comprehensive analysis and visualization for family genomes. The FGB can visualize family genomes in both individual level and variant level effectively, through integrating genome data with pedigree information. Family genome analysis, including determination of parental origin of the variants, detection of de novo mutations, identification of potential recombination events and identical-by-decent segments, etc., can be performed flexibly. Diverse annotations for the family genome variants, such as dbSNP memberships, linkage disequilibriums, genes, variant effects, potential phenotypes, etc., are illustrated as well. Moreover, the FGB can automatically search de novo mutations and compound heterozygous variants for a selected individual, and guide investigators to find high-risk genes with flexible navigation options. These features enable users to investigate and understand family genomes intuitively and systematically. AVAILABILITY AND IMPLEMENTATION: The FGB is available at http://mlg.hit.edu.cn/FGB/. Liran Juan, Yongzhuang Liu, Yongtian Wang, Mingxiang Teng, Tianyi Zang, Yadong Wang 0001 |
Bioinform. | 2 |
| 2015 | A Bayesian framework for de novo mutation calling in parents-offspring triosabstractMOTIVATION: Spontaneous (de novo) mutations play an important role in the disease etiology of a range of complex diseases. Identifying de novo mutations (DNMs) in sporadic cases provides an effective strategy to find genes or genomic regions implicated in the genetics of disease. High-throughput next-generation sequencing enables genome- or exome-wide detection of DNMs by sequencing parents-proband trios. It is challenging to sift true mutations through massive amount of noise due to sequencing error and alignment artifacts. One of the critical limitations of existing methods is that for all genomic regions the same pre-specified mutation rate is assumed, which has a significant impact on the DNM calling accuracy. RESULTS: In this study, we developed and implemented a novel Bayesian framework for DNM calling in trios (TrioDeNovo), which overcomes these limitations by disentangling prior mutation rates from evaluation of the likelihood of the data so that flexible priors can be adjusted post-hoc at different genomic sites. Through extensively simulations and application to real data we showed that this new method has improved sensitivity and specificity over existing methods, and provides a flexible framework to further improve the efficiency by incorporating proper priors. The accuracy is further improved using effective filtering based on sequence alignment characteristics. AVAILABILITY AND IMPLEMENTATION: The C++ source code implementing TrioDeNovo is freely available at https://medschool.vanderbilt.edu/cgg. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Xiaowei Zhan, Xue Zhong, Yongzhuang Liu, Yujun Han, Wei Chen 0074, Bingshan Li |
Bioinform. | 4 |
| 2015 | Using Semantic Association to Extend and Infer Literature-Oriented Relativity Between TermsabstractRelative terms often appear together in the literature. Methods have been presented for weighting relativity of pairwise terms by their co-occurring literature and inferring new relationship. Terms in the literature are also in the directed acyclic graph of ontologies, such as Gene Ontology and Disease Ontology. Therefore, semantic association between terms may help for establishing relativities between terms in literature. However, current methods do not use these associations. In this paper, an adjusted R-scaled score (ARSS) based on information content (ARSSIC) method is introduced to infer new relationship between terms. First, set inclusion relationship between terms of ontology was exploited to extend relationships between these terms and literature. Next, the ARSS method was presented to measure relativity between terms across ontologies according to these extensional relationships. Then, the ARSSIC method using ratios of information shared of term's ancestors was designed to infer new relationship between terms across ontologies. The result of the experiment shows that ARSS identified more pairs of statistically significant terms based on corresponding gene sets than other methods. And the high average area under the receiver operating characteristic curve (0.9293) shows that ARSSIC achieved a high true positive rate and a low false positive rate. Data is available at http://mlg.hit.edu.cn/ARSSIC/. Liang Cheng 0006, Jie Li 0055, Yang Hu 0008, Yongzhuang Liu, Yan-Shuo Chu, Yadong Wang 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2014 | A gradient-boosting approach for filtering de novo mutations in parent-offspring triosabstractMOTIVATION: Whole-genome and -exome sequencing on parent-offspring trios is a powerful approach to identifying disease-associated genes by detecting de novo mutations in patients. Accurate detection of de novo mutations from sequencing data is a critical step in trio-based genetic studies. Existing bioinformatic approaches usually yield high error rates due to sequencing artifacts and alignment issues, which may either miss true de novo mutations or call too many false ones, making downstream validation and analysis difficult. In particular, current approaches have much worse specificity than sensitivity, and developing effective filters to discriminate genuine from spurious de novo mutations remains an unsolved challenge. RESULTS: In this article, we curated 59 sequence features in whole genome and exome alignment context which are considered to be relevant to discriminating true de novo mutations from artifacts, and then employed a machine-learning approach to classify candidates as true or false de novo mutations. Specifically, we built a classifier, named De Novo Mutation Filter (DNMFilter), using gradient boosting as the classification algorithm. We built the training set using experimentally validated true and false de novo mutations as well as collected false de novo mutations from an in-house large-scale exome-sequencing project. We evaluated DNMFilter's theoretical performance and investigated relative importance of different sequence features on the classification accuracy. Finally, we applied DNMFilter on our in-house whole exome trios and one CEU trio from the 1000 Genomes Project and found that DNMFilter could be coupled with commonly used de novo mutation detection approaches as an effective filtering approach to significantly reduce false discovery rate without sacrificing sensitivity. AVAILABILITY: The software DNMFilter implemented using a combination of Java and R is freely available from the website at http://humangenome.duke.edu/software. Yongzhuang Liu, Bingshan Li, Renjie Tan, Xiaolin Zhu 0002, Yadong Wang 0001 |
Bioinform. | 1 |