Degui Zhi

dblp:64/2908 · DBLP profile ↗
← Back
36ranked-venue papers
5as first author
18since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 34 · 5 first-author · 16 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
YearPublicationVenuePosition
2025 Learning to Discover Regulatory Elements for Gene Expression Prediction
abstract
We consider the problem of predicting gene expressions from DNA sequences. A key challenge of this task is to find the regulatory elements that control gene expressions. Here, we introduce Seq2Exp, a Sequence to Expression network explicitly designed to discover and extract regulatory elements that drive target gene expression, enhancing the accuracy of the gene expression prediction. Our approach captures the causal relationship between epigenomic signals, DNA sequences and their associated regulatory elements. Specifically, we propose to decompose the epigenomic signals and the DNA sequence conditioned on the causal active regulatory elements, and apply an information bottleneck with the Beta distribution to combine their effects while filtering out non-causal components. Our experiments demonstrate that Seq2Exp outperforms existing baselines in gene expression prediction tasks and discovers influential regions compared to commonly used statistical methods for peak detection such as MACS3. The source code is released as part of the AIRS library (https://github.com/divelab/AIRS/).
Xingyu Su, Haiyang Yu 0005, Degui Zhi, Shuiwang Ji
ICLR3
2025 Towards precision protein-ligand affinity prediction benchmark: A Complete and Modification-Aware DAVIS Dataset
abstract
Advancements in AI for science unlocks capabilities for critical drug discovery tasks such as protein-ligand binding affinity prediction. However, current models overfit to existing oversimplified datasets that does not represent naturally occurring and biologically relevant proteins with modifications. In this work, we curate a complete and modification-aware version of the widely used DAVIS dataset by incorporating 4,032 kinase–ligand pairs involving substitutions, insertions, deletions, and phosphorylation events. This enriched dataset enables benchmarking of predictive models under biologically realistic conditions. Based on this new dataset, we propose three benchmark settings—Augmented Dataset Prediction, Wild-Type to Modification Generalization, and Few-Shot Modification Generalization—designed to assess model robustness in the presence of protein modifications. Through extensive evaluation of both docking-free and docking-based methods, we find that docking-based model generalize better in zero-shot settings. In contrast, docking-free models tend to overfit to wild-type proteins and struggle with unseen modifications but show notable improvement when fine-tuned on a small set of modified examples. We anticipate that the curated dataset and benchmarks offer a valuable foundation for developing models that better generalize to protein modifications, ultimately advancing precision medicine in drug discovery. The benchmark is available at: https://github.com/ZhiGroup/DAVIS-complete
Ming-Hsiu Wu, Ziqian Xie, Shuiwang Ji, Degui Zhi
NeurIPS4
2025 Dynamic μ-PBWT: Dynamic Run-Length Compressed PBWT for Biobank Scale Data
Pramesh Shakya, Ahsan Sanaullah, Degui Zhi, Shaojie Zhang 0001
RECOMB3
2025 An Efficient Data Structure and Algorithm for Long-Match Query in Run-Length Compressed BWT
abstract
String matching problems in bioinformatics are typically for finding exact substring matches between a query and a reference text. Previous formulations often focus on maximum exact matches (MEMs). However, multiple occurrences of substrings of the query in the text that are long enough but not maximal may not be captured by MEMs. Such long matches can be informative, especially when the text is a collection of similar sequences such as genomes. In this paper, we describe a new type of match between a pattern and a text that aren't necessarily maximal in the query, but still contain useful matching information: locally maximal exact matches (LEMs). There are usually a large amount of LEMs, so we only consider those above some length threshold ℒ. These are referred to as long LEMs. The purpose of long LEMs is to capture substring matches between a query and a text that are not necessarily maximal in the pattern but still long enough to be important. Therefore efficient long LEMs finding algorithms are desired for these datasets. However, these datasets are too large to query on traditional string indexes. Fortunately, these datasets are very repetitive. Recently, compressed string indexes that take advantage of the redundancy in the data but retain efficient querying capability have been proposed as a solution. We therefore give an efficient algorithm for computing all the long LEMs of a query and a text in a BWT runs compressed string index. We describe an O(m+occ) expected time algorithm that relies on an O(r) words space string index for outputting all long LEMs of a pattern with respect to a text given the matching statistics of the pattern with respect to the text. Here m is the length of the query, occ is the number of long LEMs outputted, and r is the number of runs in the BWT of the text. The O(r) space string index we describe relies on an adaptation of the move data structure by Nishimoto and Tabei. We are able to support LCP[i] queries in constant time given SA[i]. In other words, we answer PLCP[i] queries in constant time. These PLCP queries enable the efficient long LEM query. Long LEMs may provide useful similarity information between a pattern and a text that MEMs may ignore. This information is particularly useful in pangenome and biobank scale haplotype panel contexts.
Ahsan Sanaullah, Degui Zhi, Shaojie Zhang 0001
WABI2
2025 iGTP: learning interpretable cellular embedding for inferring biological mechanisms underlying single-cell transcriptomics
abstract
Deep-learning models like Variational AutoEncoder have enabled low dimensional cellular embedding representation for large-scale single-cell transcriptomes and shown great flexibility in downstream tasks. However, biologically meaningful latent space is usually missing if no specific structure is designed. Here, we engineered a novel interpretable generative transcriptional program (iGTP) framework that could model the importance of transcriptional program (TP) space and protein-protein interactions (PPI) between different biological states. We demonstrated the performance of iGTP in a diverse biological context using gene ontology, canonical pathway, and different PPI curation. iGTP not only elucidated the ground truth of cellular responses but also surpassed other deep learning models and traditional bioinformatics methods in functional enrichment tasks. By integrating the latent layer with a graph neural network framework, iGTP could effectively infer cellular responses to perturbations. Lastly, we applied iGTP TP embeddings with a latent diffusion model to accurately generate cell embeddings for specific cell types and states. We anticipate that iGTP will offer insights at both PPI and TP levels and holds promise for predicting responses to novel perturbations.
Kanglin Hsieh, Yan Chu 0005, Lishan Yu, Nuo Hu, Isha Kawosa, Patrick G. Pilié, Pratip K. Bhattacharya, Degui Zhi, Xiaoqian Jiang, Zhongming Zhao, Yulin Dai
Briefings Bioinform.10
2025 Recomb-Mix: fast and accurate local ancestry inference
abstract
MOTIVATION: The availability of large genotyped cohorts brings new opportunities for revealing the high-resolution genetic structure of admixed populations via local ancestry inference (LAI), the process of identifying the ancestry of each segment of an individual haplotype. Though current methods achieve high accuracy in standard cases, LAI is still challenging when reference populations are more similar (e.g. intra-continental), when the number of reference populations is too numerous, or when the admixture events are deep in time, all of which are increasingly unavoidable in large biobanks. RESULTS: In this work, we present Recomb-Mix, a new LAI method which integrates elements from the site-based Li and Stephens model and introduces a new graph collapsing techniques to simplify counting paths with the same ancestry label readout. Through comprehensive benchmarking on various simulated datasets, we show that Recomb-Mix is more accurate than existing methods in diverse sets of scenarios while being competitive in terms of resource efficiency. The scalability and robustness of Recomb-Mix are also demonstrated with real-world datasets. We expect that Recomb-Mix will be a useful method for advancing genetics studies of admixed populations. AVAILABILITY AND IMPLEMENTATION: The implementation of Recomb-Mix is available at https://github.com/ucfcbb/Recomb-Mix.
Degui Zhi, Shaojie Zhang 0001
Bioinform.2
2023 Syllable-PBWT for space-efficient haplotype long-match query
abstract
MOTIVATION: The positional Burrows-Wheeler transform (PBWT) has led to tremendous strides in haplotype matching on biobank-scale data. For genetic genealogical search, PBWT-based methods have optimized the asymptotic runtime of finding long matches between a query haplotype and a predefined panel of haplotypes. However, to enable fast query searches, the full-sized panel and PBWT data structures must be kept in memory, preventing existing algorithms from scaling up to modern biobank panels consisting of millions of haplotypes. In this work, we propose a space-efficient variation of PBWT named Syllable-PBWT, which divides every haplotype into syllables, builds the PBWT positional prefix arrays on the compressed syllabic panel, and leverages the polynomial rolling hash function for positional substring comparison. With the Syllable-PBWT data structures, we then present a long match query algorithm named Syllable-Query. RESULTS: Compared to the most time- and space-efficient previously published solution to the long match query problem, Syllable-Query reduced the memory use by a factor of over 100 on both the UK Biobank genotype data and the 1000 Genomes Project sequence data. Surprisingly, the smaller size of our syllabic data structures allows for more efficient iteration and CPU cache usage, granting Syllable-Query even faster runtime than existing solutions. AVAILABILITY AND IMPLEMENTATION: https://github.com/ZhiGroup/Syllable-PBWT. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Victor Wang, Ardalan Naseri, Shaojie Zhang 0001, Degui Zhi
Bioinform.4
2023 RaPID-Query for fast identity by descent search and genealogical analysis
abstract
MOTIVATION: Due to the rapid growth of the genetic database size, genealogical search, a process of inferring familial relatedness by identifying DNA matches, has become a viable approach to help individuals finding missing family members or law enforcement agencies locating suspects. A fast and accurate method is needed to search an out-of-database individual against millions of individuals. Most existing approaches only offer all-versus-all within panel match. Some prototype algorithms offer one-versus-all query from out-of-panel individual, but they do not tolerate errors. RESULTS: A new method, random projection-based identity-by-descent (IBD) detection (RaPID) query, is introduced to make fast genealogical search possible. RaPID-Query identifies IBD segments between a query haplotype and a panel of haplotypes. By integrating matches over multiple PBWT indexes, RaPID-Query manages to locate IBD segments quickly with a given cutoff length while allowing mismatched sites. A single query against all UK biobank autosomal chromosomes was completed within 2.76 seconds on average, with the minimum length 7 cM and 700 markers. RaPID-Query achieved a 0.016 false negative rate and a 0.012 false positive rate simultaneously on a chromosome 20 sequencing panel having 86 265 sites. This is comparable to the state-of-the-art IBD detection method TPBWT(out-of-sample) and Hap-IBD. The high-quality IBD segments yielded by RaPID-Query were able to distinguish up to fourth degree of the familial relatedness for a given individual pair, and the area under the receiver operating characteristic curve values are at least 97.28%. AVAILABILITY AND IMPLEMENTATION: The RaPID-Query program is available at https://github.com/ucfcbb/RaPID-Query.
Ardalan Naseri, Degui Zhi, Shaojie Zhang 0001
Bioinform.3
2022 Pancreatic cancer risk prediction using recurrent neural network models trained on electronic health records and claims data
Laila Rasmy, Peter Beshara, Bijun S. Kannadath, Degui Zhi
AMIA4
2022 Explainability versus Accuracy to Engender Trust
Angela Ross, Laila Rasmy, Degui Zhi, Kathleen McGrow
AMIA3
2022 Combining Transfer Learning with Graph Attention Models for HIV Risk Prediction
Evan Yu, Cui Tao, Jingcheng Du, Degui Zhi, Yang Xiang 0003, Kayo Fujimoto, John A. Schneider
AMIA4
2022 Haplotype Threading Using the Positional Burrows-Wheeler Transform
Ahsan Sanaullah, Degui Zhi, Shaoije Zhang
WABI2
2022 Toward a standard formal semantic representation of the model card report
abstract
BACKGROUND: Model card reports aim to provide informative and transparent description of machine learning models to stakeholders. This report document is of interest to the National Institutes of Health's Bridge2AI initiative to address the FAIR challenges with artificial intelligence-based machine learning models for biomedical research. We present our early undertaking in developing an ontology for capturing the conceptual-level information embedded in model card reports. RESULTS: Sourcing from existing ontologies and developing the core framework, we generated the Model Card Report Ontology. Our development efforts yielded an OWL2-based artifact that represents and formalizes model card report information. The current release of this ontology utilizes standard concepts and properties from OBO Foundry ontologies. Also, the software reasoner indicated no logical inconsistencies with the ontology. With sample model cards of machine learning models for bioinformatics research (HIV social networks and adverse outcome prediction for stent implantation), we showed the coverage and usefulness of our model in transforming static model card reports to a computable format for machine-based processing. CONCLUSIONS: The benefit of our work is that it utilizes expansive and standard terminologies and scientific rigor promoted by biomedical ontologists, as well as, generating an avenue to make model cards machine-readable using semantic web technology. Our future goal is to assess the veracity of our model and later expand the model to include additional concepts to address terminological gaps. We discuss tools and software that will utilize our ontology for potential application services.
Muhammad Amith, Licong Cui, Degui Zhi, Kirk Roberts, Xiaoqian Jiang, Fang Li 0011, Evan Yu, Cui Tao
BMC Bioinform.3
2022 PK-RNN-V E: A deep learning model approach to vancomycin therapeutic drug monitoring using electronic health record data
Masayuki Nigo, Hong Thoai Nga Tran, Ziqian Xie, Bingyu Mao, Laila Rasmy, Hongyu Miao 0001, Degui Zhi
J. Biomed. Informatics8
2021 Generalizable Gated Recurrent Neural Network based model to predict COVID-19 patient outcomes on admission
Laila Rasmy, Bijun S. Kannadath, Masayuki Nigo, Ziqian Xie, Bingyu Mao, Khush A. Patel, Yujia Zhou 0003, Hua Xu 0001, Degui Zhi
AMIA10
2021 CovRNN: predicting outcomes of COVID-19 patients on admission using their electronic health records with minimal data processing
Laila Rasmy, Masayuki Nigo, Bijun S. Kannadath, Ziqian Xie, Bingyu Mao, Khush A. Patel, Wanheng Zhang, Yujia Zhou 0003, Angela Ross, Hua Xu 0001, Degui Zhi
AMIA11
2021 Efficient Haplotype Block Matching in Bi-Directional PBWT
abstract
Efficient haplotype matching search is of great interest when large genotyped cohorts are becoming available. Positional Burrows-Wheeler Transform (PBWT) enables efficient searching for blocks of haplotype matches. However, existing efficient PBWT algorithms sweep across the haplotype panel from left to right, capturing all exact matches. As a result, PBWT does not account for mismatches. It is also not easy to investigate the patterns of changes between the matching blocks. Here, we present an extension to PBWT, called bi-directional PBWT that allows the information about the blocks of matches to be present at both sides of each site. We also present a set of algorithms to efficiently merge the matching blocks or examine the patterns of changes on both sides of each site. The time complexity of the algorithms to find and merge matching blocks using bi-directional PBWT is linear to the input size. Using real data from the UK Biobank, we demonstrate the run time and memory efficiency of our algorithms. More importantly, our algorithms can identify more blocks by enabling tolerance of mismatches. Moreover, by using mutual information (MI) between the forward and the reverse PBWT matching block sets as a measure of haplotype consistency, we found the MI derived from European samples in the 1000 Genomes Project is highly correlated (Spearman correlation r=0.87) with the deCODE recombination map.
Ardalan Naseri, William Yue, Shaojie Zhang 0001, Degui Zhi
WABI4
2021 d-PBWT: dynamic positional Burrows-Wheeler transform
abstract
Abstract Motivation Durbin’s positional Burrows–Wheeler transform (PBWT) is a scalable data structure for haplotype matching. It has been successfully applied to identical by descent (IBD) segment identification and genotype imputation. Once the PBWT of a haplotype panel is constructed, it supports efficient retrieval of all shared long segments among all individuals (long matches) and efficient query between an external haplotype and the panel. However, the standard PBWT is an array-based static data structure and does not support dynamic updates of the panel. Results Here, we generalize the static PBWT to a dynamic data structure, d-PBWT, where the reverse prefix sorting at each position is stored with linked lists. We also developed efficient algorithms for insertion and deletion of individual haplotypes. In addition, we verified that d-PBWT can support all algorithms of PBWT. In doing so, we systematically investigated variations of set maximal match and long match query algorithms: while they all have average case time complexity independent of database size, they have different worst case complexities and dependencies on additional data structures. Availabilityand implementation The benchmarking code is available at genome.ucf.edu/d-PBWT. Supplementary information Supplementary data are available at Bioinformatics online.
Ahsan Sanaullah, Degui Zhi, Shaojie Zhang 0001
Bioinform.2
2020 d-PBWT: Dynamic Positional Burrows-Wheeler Transform
Ahsan Sanaullah, Degui Zhi, Shaojie Zhang 0001
RECOMB2
2020 Representation of EHR data for predictive modeling: a comparison between UMLS and other terminologies
abstract
OBJECTIVE: Predictive disease modeling using electronic health record data is a growing field. Although clinical data in their raw form can be used directly for predictive modeling, it is a common practice to map data to standard terminologies to facilitate data aggregation and reuse. There is, however, a lack of systematic investigation of how different representations could affect the performance of predictive models, especially in the context of machine learning and deep learning. MATERIALS AND METHODS: We projected the input diagnoses data in the Cerner HealthFacts database to Unified Medical Language System (UMLS) and 5 other terminologies, including CCS, CCSR, ICD-9, ICD-10, and PheWAS, and evaluated the prediction performances of these terminologies on 2 different tasks: the risk prediction of heart failure in diabetes patients and the risk prediction of pancreatic cancer. Two popular models were evaluated: logistic regression and a recurrent neural network. RESULTS: For logistic regression, using UMLS delivered the optimal area under the receiver operating characteristics (AUROC) results in both dengue hemorrhagic fever (81.15%) and pancreatic cancer (80.53%) tasks. For recurrent neural network, UMLS worked best for pancreatic cancer prediction (AUROC 82.24%), second only (AUROC 85.55%) to PheWAS (AUROC 85.87%) for dengue hemorrhagic fever prediction. DISCUSSION/CONCLUSION: In our experiments, terminologies with larger vocabularies and finer-grained representations were associated with better prediction performances. In particular, UMLS is consistently 1 of the best-performing ones. We believe that our work may help to inform better designs of predictive models, although further investigation is warranted.
Laila Rasmy, Firat Tiryaki, Yujia Zhou 0003, Yang Xiang 0003, Cui Tao, Hua Xu 0001, Degui Zhi
J. Am. Medical Informatics Assoc.7
2019 Efficient haplotype matching between a query and a panel for genealogical search
abstract
MOTIVATION: With the wide availability of whole-genome genotype data, there is an increasing need for conducting genetic genealogical searches efficiently. Computationally, this task amounts to identifying shared DNA segments between a query individual and a very large panel containing millions of haplotypes. The celebrated Positional Burrows-Wheeler Transform (PBWT) data structure is a pre-computed index of the panel that enables constant time matching at each position between one haplotype and an arbitrarily large panel. However, the existing algorithm (Durbin's Algorithm 5) can only identify set-maximal matches, the longest matches ending at any location in a panel, while in real genealogical search scenarios, multiple 'good enough' matches are desired. RESULTS: In this work, we developed two algorithmic extensions of Durbin's Algorithm 5, that can find all L-long matches, matches longer than or equal to a given length L, between a query and a panel. In the first algorithm, PBWT-Query, we introduce 'virtual insertion' of the query into the PBWT matrix of the panel, and then scanning up and down for the PBWT match blocks with length greater than L. In our second algorithm, L-PBWT-Query, we further speed up PBWT-Query by introducing additional data structures that allow us to avoid iterating through blocks of incomplete matches. The efficiency of PBWT-Query and L-PBWT-Query is demonstrated using the simulated data and the UK Biobank data. Our results show that our proposed algorithms can detect related individuals for a given query efficiently in very large cohorts which enables a fast on-line query search. AVAILABILITY AND IMPLEMENTATION: genome.ucf.edu/pbwt-query. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Ardalan Naseri, Erwin Holzhauser, Degui Zhi, Shaojie Zhang 0001
Bioinform.3
2019 Multi-allelic positional Burrows-Wheeler transform
abstract
BACKGROUND: Recent advances in whole-genome sequencing and SNP array technology have led to the generation of a large amount of genotype data. Large volumes of genotype data will require faster and more efficient methods for storing and searching the data. Positional Burrows-Wheeler Transform (PBWT) provides an appropriate data structure for bi-allelic data. With the increasing sample sizes, more multi-allelic sites are expected to be observed. Hence, there is a necessity to handle multi-allelic genotype data. RESULTS: In this paper, we introduce a multi-allelic version of the Positional Burrows-Wheeler Transform (mPBWT) based on the bi-allelic version for compression and searching. The time-complexity for constructing the data structure and searching within a panel containing t-allelic sites increases by a factor of t. CONCLUSION: Considering the small value for the possible alleles t, the time increase for the multi-allelic PBWT will be negligible and comparable to the bi-allelic version of PBWT.
Ardalan Naseri, Degui Zhi, Shaojie Zhang 0001
BMC Bioinform.2
2019 Network context matters: graph convolutional network model over social networks improves the detection of unknown HIV infections among young men who have sex with men
abstract
OBJECTIVE: HIV infection risk can be estimated based on not only individual features but also social network information. However, there have been insufficient studies using n machine learning methods that can maximize the utility of such information. Leveraging a state-of-the-art network topology modeling method, graph convolutional networks (GCN), our main objective was to include network information for the task of detecting previously unknown HIV infections. MATERIALS AND METHODS: We used multiple social network data (peer referral, social, sex partners, and affiliation with social and health venues) that include 378 young men who had sex with men in Houston, TX, collected between 2014 and 2016. Due to the limited sample size, an ensemble approach was engaged by integrating GCN for modeling information flow and statistical machine learning methods, including random forest and logistic regression, to efficiently model sparse features in individual nodes. RESULTS: Modeling network information using GCN effectively increased the prediction of HIV status in the social network. The ensemble approach achieved 96.6% on accuracy and 94.6% on F1 measure, which outperformed the baseline methods (GCN, logistic regression, and random forest: 79.0%, 90.5%, 94.4% on accuracy, respectively; and 57.7%, 80.2%, 90.4% on F1). In the networks with missing HIV status, the ensemble also produced promising results. CONCLUSION: Network context is a necessary component in modeling infectious disease transmissions such as HIV. GCN, when combined with traditional machine learning approaches, achieved promising performance in detecting previously unknown HIV infections, which may provide a useful tool for combatting the HIV epidemic.
Yang Xiang 0003, Kayo Fujimoto, John A. Schneider, Yuxi Jia, Degui Zhi, Cui Tao
J. Am. Medical Informatics Assoc.5
2018 Development of HeartData, a Data Discovery Index Prototype for Cardiovascular data
Mandana Salimi, Anupama E. Gururaj, Cui Tao, Degui Zhi, Hua Xu 0001
AMIA5
2018 Asthma Onset Prediction Using Structured EMR Data
Yang Xiang 0003, Jun Xu 0007, Jingcheng Du, Degui Zhi, Cui Tao
AMIA4
2018 The International Conference on Intelligent Biology and Medicine (ICIBM) 2018: bioinformatics towards translational applications
Xiaoming Liu 0021, Lei Xie 0006, Zhijin Wu, Kai Wang 0063, Zhongming Zhao, Jianhua Ruan, Degui Zhi
BMC Bioinform.7
2018 A study of generalizability of recurrent neural network-based predictive models for heart failure onset risk using a large and heterogeneous EHR data set
Laila Rasmy, Yonghui Wu 0001, Ningtao Wang, W. Jim Zheng, Fei Wang 0001, Hulin Wu, Hua Xu 0001, Degui Zhi
J. Biomed. Informatics9
2017 Clinical Named Entity Recognition Using Deep Learning Models
Yonghui Wu 0001, Min Jiang 0007, Jun Xu 0007, Degui Zhi, Hua Xu 0001
AMIA4
2017 Ultra-Fast Identity by Descent Detection in Biobank-Scale Cohorts Using Positional Burrows-Wheeler Transform
Ardalan Naseri, Shaojie Zhang 0001, Degui Zhi
RECOMB4
2013 Joint haplotype phasing and genotype calling of multiple individuals using haplotype informative reads
abstract
MOTIVATION: Hidden Markov model, based on Li and Stephens model that takes into account chromosome sharing of multiple individuals, results in mainstream haplotype phasing algorithms for genotyping arrays and next-generation sequencing (NGS) data. However, existing methods based on this model assume that the allele count data are independently observed at individual sites and do not consider haplotype informative reads, i.e. reads that cover multiple heterozygous sites, which carry useful haplotype information. In our previous work, we developed a new hidden Markov model to incorporate a two-site joint emission term that captures the haplotype information across two adjacent sites. Although our model improves the accuracy of genotype calling and haplotype phasing, haplotype information in reads covering non-adjacent sites and/or more than two adjacent sites is not used because of the severe computational burden. RESULTS: We develop a new probabilistic model for genotype calling and haplotype phasing from NGS data that incorporates haplotype information of multiple adjacent and/or non-adjacent sites covered by a read over an arbitrary distance. We develop a new hybrid Markov Chain Monte Carlo algorithm that combines the Gibbs sampling algorithm of HapSeq and Metropolis-Hastings algorithm and is computationally feasible. We show by simulation and real data from the 1000 Genomes Project that our model offers superior performance for haplotype phasing and genotype calling for population NGS data over existing methods. AVAILABILITY: HapSeq2 is available at www.ssg.uab.edu/hapseq/.
Degui Zhi
Bioinform.2
2012 Genotype calling from next-generation sequencing data using haplotype information of reads
abstract
MOTIVATION: Low coverage sequencing provides an economic strategy for whole genome sequencing. When sequencing a set of individuals, genotype calling can be challenging due to low sequencing coverage. Linkage disequilibrium (LD) based refinement of genotyping calling is essential to improve the accuracy. Current LD-based methods use read counts or genotype likelihoods at individual potential polymorphic sites (PPSs). Reads that span multiple PPSs (jumping reads) can provide additional haplotype information overlooked by current methods. RESULTS: In this article, we introduce a new Hidden Markov Model (HMM)-based method that can take into account jumping reads information across adjacent PPSs and implement it in the HapSeq program. Our method extends the HMM in Thunder and explicitly models jumping reads information as emission probabilities conditional on the states of adjacent PPSs. Our simulation results show that, compared to Thunder, HapSeq reduces the genotyping error rate by 30%, from 0.86% to 0.60%. The results from the 1000 Genomes Project show that HapSeq reduces the genotyping error rate by 12 and 9%, from 2.24% and 2.76% to 1.97% and 2.50% for individuals with European and African ancestry, respectively. We expect our program can improve genotyping qualities of the large number of ongoing and planned whole genome sequencing projects. CONTACT: [email protected]; [email protected] AVAILABILITY: The software package HapSeq and its manual can be found and downloaded at www.ssg.uab.edu/hapseq/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Degui Zhi, Jihua Wu, Nianjun Liu
Bioinform.1
2011 Length bias correction for RNA-seq data in gene set analyses
abstract
MOTIVATION: Next-generation sequencing technologies are being rapidly applied to quantifying transcripts (RNA-seq). However, due to the unique properties of the RNA-seq data, the differential expression of longer transcripts is more likely to be identified than that of shorter transcripts with the same effect size. This bias complicates the downstream gene set analysis (GSA) because the methods for GSA previously developed for microarray data are based on the assumption that genes with same effect size have equal probability (power) to be identified as significantly differentially expressed. Since transcript length is not related to gene expression, adjusting for such length dependency in GSA becomes necessary. RESULTS: In this article, we proposed two approaches for transcript-length adjustment for analyses based on Poisson models: (i) At individual gene level, we adjusted each gene's test statistic using the square root of transcript length followed by testing for gene set using the Wilcoxon rank-sum test. (ii) At gene set level, we adjusted the null distribution for the Fisher's exact test by weighting the identification probability of each gene using the square root of its transcript length. We evaluated these two approaches using simulations and a real dataset, and showed that these methods can effectively reduce the transcript-length biases. The top-ranked GO terms obtained from the proposed adjustments show more overlaps with the microarray results. AVAILABILITY: R scripts are at http://www.soph.uab.edu/Statgenetics/People/XCui/r-codes/.
Liyan Gao, Zhide Fang, Degui Zhi, Xiangqin Cui
Bioinform.4
2010 Alignment-free local structural search by writhe decomposition
abstract
MOTIVATION: Rapid methods for protein structure search enable biological discoveries based on flexibly defined structural similarity, unleashing the power of the ever greater number of solved protein structures. Projection methods show promise for the development of fast structural database search solutions. Projection methods map a structure to a point in a high-dimensional space and compare two structures by measuring distance between their projected points. These methods offer a tremendous increase in speed over residue-level structural alignment methods. However, current projection methods are not practical, partly because they are unable to identify local similarities. RESULTS: We propose a new projection-based approach that can rapidly detect global as well as local structural similarities. Local structural search is enabled by a topology-inspired writhe decomposition protocol that produces a small number of fragments while ensuring that similar structures are cut in a similar manner. In benchmark tests, we show that our method, writher, improves accuracy over existing projection methods in terms of recognizing scop domains out of multi-domain proteins, while maintaining accuracy comparable with existing projection methods in a standard single-domain benchmark test. AVAILABILITY: The source code is available at the following website: http://compbio.berkeley.edu/proj/writher/.
Degui Zhi, Maxim Shatsky, Steven E. Brenner
Bioinform.1
2007 Alignment-Free Local Structural Search by Writhe Decomposition
Degui Zhi, Maxim Shatsky, Steven E. Brenner
WABI1
2007 Correcting Base-Assignment Errors in Repeat Regions of Shotgun Assembly
abstract
Accurate base-assignment in repeat regions of a whole genome shotgun assembly is an unsolved problem. Since reads in repeat regions cannot be easily attributed to a unique location in the genome, current assemblers may place these reads arbitrarily. As a result, the base-assignment error rate in repeats is likely to be much higher than that in the rest of the genome. We developed an iterative algorithm, EULER-AIR, that is able to correct base-assignment errors in finished genome sequences in public databases. The Wolbachia genome is among the best finished genomes. Using this genome project as an example, we demonstrated that EULER-AIR can 1) discover and correct base-assignment errors, 2) provide accurate read assignments, 3) utilize finishing reads for accurate base-assignment, and 4) provide guidance for designing finishing experiments. In the genome of Wolbachia, EULER-AIR found 16 positions with ambiguous base-assignment and two positions with erroneous bases. Besides Wolbachia, many other genome sequencing projects have significantly fewer finishing reads and, hence, are likely to contain more base-assignment errors in repeats. We demonstrate that EULER-AIR is a software tool that can be used to find and correct base-assignment errors in a genome assembly project.
Degui Zhi, Uri Keich, Pavel A. Pevzner, Steffen Heber, Haixu Tang
IEEE ACM Trans. Comput. Biol. Bioinform.1
2006 Representing and comparing protein structures as paths in three-dimensional space
abstract
BACKGROUND: Most existing formulations of protein structure comparison are based on detailed atomic level descriptions of protein structures and bypass potential insights that arise from a higher-level abstraction. RESULTS: We propose a structure comparison approach based on a simplified representation of proteins that describes its three-dimensional path by local curvature along the generalized backbone of the polypeptide. We have implemented a dynamic programming procedure that aligns curvatures of proteins by optimizing a defined sum turning angle deviation measure. CONCLUSION: Although our procedure does not directly optimize global structural similarity as measured by RMSD, our benchmarking results indicate that it can surprisingly well recover the structural similarity defined by structure classification databases and traditional structure alignment programs. In addition, our program can recognize similarities between structures with extensive conformation changes that are beyond the ability of traditional structure alignment programs. We demonstrate the applications of procedure to several contexts of structure comparison. An implementation of our procedure, CURVE, is available as a public webserver.
Degui Zhi, S. Sri Krishna, Haibo Cao, Pavel A. Pevzner, Adam Godzik
BMC Bioinform.1