EDBT 2026 Demo / reviewers in the wild / expert
Liqing Zhang 0002
dblp:20/4627-2
· DBLP profile ↗
23ranked-venue papers
2as first author
8since 2021 · last 2026
0000-0003-4660-9199ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 20 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LLM4Cell: Taxonomy and Evaluation of LLM and Agentic Models for Single-Cell BiologyabstractSajib Acharjee Dip, Adrika Zafor, Bikash Kumar Paul, Uddip Acharjee Shuvo, Muhit Islam Emon, Xuan Wang, Liqing Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Sajib Acharjee Dip, Adrika Zafor, Bikash Kumar Paul, Uddip Acharjee Shuvo, Muhit Islam Emon, Xuan Wang 0008, Liqing Zhang 0002 |
ACL (1) | 7 |
| 2025 | UnCOT-AD: Unpaired Cross-Omics Translation Enables Multi-Omics Integration for Alzheimer's Disease PredictionabstractAlzheimer's Disease (AD) is a progressive neurodegenerative disorder, posing a growing public health challenge. Traditional machine learning models for AD prediction have relied on single omics data or phenotypic assessments, limiting their ability to capture the disease's molecular complexity and resulting in poor performance. Recent advances in high-throughput multi-omics have provided deeper biological insights. However, due to the scarcity of paired omics datasets, existing multi-omics AD prediction models rely on unpaired omics data, where different omics profiles are combined without being derived from the same biological sample, leading to biologically less meaningful pairings and causing less accurate predictions. To address these issues, we propose UnCOT-AD, a novel deep learning framework for Unpaired Cross-Omics Translation enabling effective multi-omics integration for AD prediction. Our method introduces the first-ever cross-omics translation model trained on unpaired omics datasets, using two coupled Variational Autoencoders and a novel cycle consistency mechanism to ensure accurate bidirectional translation between omics types. We integrate adversarial training to ensure that the generated omics profiles are biologically realistic. Moreover, we employ contrastive learning to capture the disease specific patterns in latent space to make the cross-omics translation more accurate and biologically relevant. We rigorously validate UnCOT-AD on both cross-omics translation and AD prediction tasks. Results show that UnCOT-AD empowers multi-omics based AD prediction by combining real omics profiles with corresponding omics profiles generated by our cross-omics translation module and achieves state-of-the-art performance in accuracy and robustness. Source code is available at https://github.com/abrarrahmanabir/UnCOT-AD. Abrar Rahman Abir, Sajib Acharjee Dip, Liqing Zhang 0002 |
Briefings Bioinform. | 3 |
| 2024 | Breast Cancer Subtype Prediction Model Integrating Domain Adaptation with Semi-supervised Learning on DNA Methylation Profiles
Joungmin Choi, Liqing Zhang 0002 |
AIME (1) | 2 |
| 2024 | Mastering Long-Tail Complexity on Graphs: Characterization, Learning, and GeneralizationabstractIn the context of long-tail classification on graphs, the vast majority of existing work primarily revolves around the development of model debiasing strategies, intending to mitigate class imbalances and enhance the overall performance. Despite the notable success, there is very limited literature that provides a theoretical tool for characterizing the behaviors of long-tail classes in graphs and gaining insight into generalization performance in real-world scenarios. To bridge this gap, we propose a generalization bound for long-tail classification on graphs by formulating the problem in the fashion of multi-task learning, i.e., each task corresponds to the prediction of one particular class. Our theoretical results show that the generalization performance of long-tail classification is dominated by the overall loss range and the task complexity. Building upon the theoretical findings, we propose a novel generic framework HierTail for long-tail classification on graphs. In particular, we start with a hierarchical task grouping module that allows us to assign related tasks into hypertasks and thus control the complexity of the task space; then, we further design a balanced contrastive learning module to adaptively balance the gradients of both head and tail classes to control the loss range across all tasks in a unified fashion. Extensive experiments demonstrate the effectiveness of HierTail in characterizing long-tail classes on real graphs, which achieves up to 12.9% improvement over the leading baseline method in balanced accuracy. Haohui Wang, Baoyu Jing, Kaize Ding, Yada Zhu, Wei Cheng 0002, Yonghui Fan, Liqing Zhang 0002, Dawei Zhou 0003 |
KDD | 8 |
| 2023 | DeepMicroGen: a generative adversarial network-based method for longitudinal microbiome data imputationabstractMOTIVATION: The human microbiome, which is linked to various diseases by growing evidence, has a profound impact on human health. Since changes in the composition of the microbiome across time are associated with disease and clinical outcomes, microbiome analysis should be performed in a longitudinal study. However, due to limited sample sizes and differing numbers of timepoints for different subjects, a significant amount of data cannot be utilized, directly affecting the quality of analysis results. Deep generative models have been proposed to address this lack of data issue. Specifically, a generative adversarial network (GAN) has been successfully utilized for data augmentation to improve prediction tasks. Recent studies have also shown improved performance of GAN-based models for missing value imputation in a multivariate time series dataset compared with traditional imputation methods. RESULTS: This work proposes DeepMicroGen, a bidirectional recurrent neural network-based GAN model, trained on the temporal relationship between the observations, to impute the missing microbiome samples in longitudinal studies. DeepMicroGen outperforms standard baseline imputation methods, showing the lowest mean absolute error for both simulated and real datasets. Finally, the proposed model improved the predicted clinical outcome for allergies, by providing imputation for an incomplete longitudinal dataset used to train the classifier. AVAILABILITY AND IMPLEMENTATION: DeepMicroGen is publicly available at https://github.com/joungmin-choi/DeepMicroGen. Joungmin Choi, Ming Ji, Layne T. Watson, Liqing Zhang 0002 |
Bioinform. | 4 |
| 2022 | LM-ARG: Identification & classification of antibiotic resistance genes leveraging pre-trained protein language modelsabstractAntibiotic resistance is a silent pandemic, causing 700 thousand human deaths across the world every year. Antibiotic resistance genes (ARG) are genes conferring resistance for the bacteria carrying them. Predicting ARGs is an important computational task. Traditionally ARGs are predicted using alignment based methods. However, the false negative rate for most of the alignment-based tools is very high. The protein language models (LM) trained on the huge corpus protein sequences capture distant relations among protein sequences. These features can be utilized for the identification and classification of ARGs. We have presented a self-supervised model on the largest available ARG database with the help of a pre-trained language model ProtAlbert. We used the raw protein LM-embeddings from unlabeled data on our ARG classification task and saw it outperform state-of-the-art prediction algorithms. The extracted features from the pretrained language model boosted the supervised model accuracy to a great margin. Shafayat Ahmed, Muhit Islam Emon, Nazifa Ahmed Moumi, Liqing Zhang 0002 |
BIBM | 4 |
| 2022 | Protein-Protein Interaction Network Analysis Reveals Distinct Patterns of Antibiotic Resistance GenesabstractAntibiotic resistance genes (ARGs) are responsible for an increasing number of bacterial infections worldwide. ARGs are challenging to track within bacterial genomes as they are often subject to horizontal gene transfer via mobile genetic elements (MGEs). Complex protein-protein networks can reveal proteins contributing to the spread and persistence. Here, we developed a pipeline that facilitates this process by analyzing features of a Protein-Protein Interaction Network (PPIN). This pipeline uses a random forest model to distinguish ARGs from non-ARGs and explores associations between ARGs and proteins with which they functionally interact. We tested the approach using the PPINs of Escherichia coli and Acinetobacter baumannii, two deadly organisms known to carry ARGs harbored by MGEs, and achieved a macro average accuracy of 85% in ARG identification. The approach also revealed that ARGs are disproportionately associated with MGEs and the neighbors (genes connected with only one edge to the ARGs) are likely to be less mobile. Nazifa Ahmed Moumi, Connor L. Brown, Peter J. Vikesland, Amy Pruden, Liqing Zhang 0002 |
BIBM | 5 |
| 2021 | AgroSeek: a system for computational analysis of environmental metagenomic data and associated metadataabstractBACKGROUND: Metagenomics is gaining attention as a powerful tool for identifying how agricultural management practices influence human and animal health, especially in terms of potential to contribute to the spread of antibiotic resistance. However, the ability to compare the distribution and prevalence of antibiotic resistance genes (ARGs) across multiple studies and environments is currently impossible without a complete re-analysis of published datasets. This challenge must be addressed for metagenomics to realize its potential for helping guide effective policy and practice measures relevant to agricultural ecosystems, for example, identifying critical control points for mitigating the spread of antibiotic resistance. RESULTS: Here we introduce AgroSeek, a centralized web-based system that provides computational tools for analysis and comparison of metagenomic data sets tailored specifically to researchers and other users in the agricultural sector interested in tracking and mitigating the spread of ARGs. AgroSeek draws from rich, user-provided metagenomic data and metadata to facilitate analysis, comparison, and prediction in a user-friendly fashion. Further, AgroSeek draws from publicly-contributed data sets to provide a point of comparison and context for data analysis. To incorporate metadata into our analysis and comparison procedures, we provide flexible metadata templates, including user-customized metadata attributes to facilitate data sharing, while maintaining the metadata in a comparable fashion for the broader user community and to support large-scale comparative and predictive analysis. CONCLUSION: AgroSeek provides an easy-to-use tool for environmental metagenomic analysis and comparison, based on both gene annotations and associated metadata, with this initial demonstration focusing on control of antibiotic resistance in agricultural ecosystems. Agroseek creates a space for metagenomic data sharing and collaboration to assist policy makers, stakeholders, and the public in decision-making. AgroSeek is publicly-available at https://agroseek.cs.vt.edu/ . Kyle Akers, Ishi Keenum, Lauren Wind, Suraj Gupta, Chaoqi Chen, Reem Aldaihani, Amy Pruden, Liqing Zhang 0002, Katharine F. Knowlton, Kang Xia, Lenwood S. Heath |
BMC Bioinform. | 9 |
| 2020 | ARGminer: a web platform for the crowdsourcing-based curation of antibiotic resistance genesabstractAntimicrobial resistance (AMR) has been identified by the World Health Organization (WHO) as a major global health threat. It is projected that AMR will increase exponentially by 2050, leading to substantial human morbidity and mortality (O’Neill, 2016; Pires et al., 2017). Therefore, swift action is required to enable enhanced monitoring and to help tackle the spread of AMR, including: understanding the mechanisms controlling dissemination of antibiotic resistance genes (ARGs) via environmental sources and pathways (Bengtsson-Palme et al., 2018; Martínez, 2008; Pruden et al., 2013), discovering novel ARGs before they are found to be problematic in the clinic (Berglund et al., 2017), developing new computational strategies for ARG annotation (Arango-Argoty et al., 2018; Gibson et al., 2015; Lakin et al., 2017; McArthur et al., 2013; Yang et al., 2016) and expansion of current ARG repositories (Arango-Argoty et al., 2018; Lakin et al., 2017). Metagenomic sequencing has provided a powerful means for accessing the diverse array of ARGs, or ‘resistomes’, (Sello, 2012) characteristic of various environments (Bengtsson-Palme et al., 2016; Garner et al., 2016; Li et al., 2017; Pal et al., 2016) and has supported the discovery of novel ARGs and their interactions (Forsberg et al., 2014; Pehrsson et al., 2016). Existing metagenomic approaches are largely dependent upon predicting antibiotic resistance attributes through sequence similarity computation, which is subject to major limitations. First, such similarity computations require a high quality and up-to-date ARG reference/annotation database to enable consistent and accurate ARG identification. Second, the scope of such analyses is limited to previously characterized ARGs by the lack of a comprehensive target gene for alignment (Yang et al., 2016). However, computational efforts have been made to annotate novel ARGs to increment the repertoire of sequences by using structurally (3D) close analogs (Ruppé et al., 2019). To improve the capacity of metagenomic-based approaches to broadly and accurately detect the full range of ARGs present in a given sample, it is necessary to continuously expand and improve curation of corresponding databases (Arango-Argoty et al., 2018). However, the risk of incorporation of false positives, i.e. ‘ARG-like’ genes that do not necessarily induce an AMR phenotype, stands as a major impediment to expanded curation efforts. Therefore, manual inspection and validation of potential ARG entries is a critical aspect of ensuring the validity of AMR databases and their application. Manual curation of ARGs is typically carried out by a few experts associated with research groups committed to maintaining public databases. This process is complex, tedious and time consuming. For instance, the last update of the Antibiotic Resistance Database (ARDB) was in 2009 (Liu and Pop, 2009), and, therefore, it does not contain any recently discovered ARGs, such as blaNDM-1 or mcr-1. The MEGARes database (Lakin et al., 2017), which was designed to simplify the organization of ARG annotation, has not been updated since December 2016. The resqu database, which contains genes for which there is evidence of having been transferred via Mobile Genetic elements (MGEs), has not been updated since 2013 (Bengtsson-Palme et al., 2017). The Comprehensive Antibiotic Resistance Database (CARD) (McArthur et al., 2013) is widely considered to be the most up-to-date ARG resource. First introduced in 2016, CARD has been updated more than 21 times, with corresponding changes to the ARG sequences and metadata (e.g. antibiotic class, gene name and mechanism). This acutely illustrates how complex and time-consuming ARG database curation is, even for domain experts. Attempts have been made to address limitations of currently available databases, understand the definition of ARGs (Martínez et al., 2015) and introduce new databases, such as the Structured Antibiotic Resistance Database (SARG), which employed intense manual curation to address issues such as inconsistencies in nomenclature and elimination of single-nucleotide polymorphisms and housekeeping genes (Yang et al., 2016). In our own research group, we previously introduced DeepARG: a computational approach for predicting ARGs using deep learning (Arango-Argoty et al., 2018). Along with the machine learning models, we also released a curated database named DeepARG-DB. This database employs manual curation, literature review of ARGs and annotation of ARGs using sequence alignments. DeepARG-DB was first released in July 2017 and most recently updated in August 2018. However, the DeepARG database depends on annotations from multiple resources, making it sensitive to the propagation of errors from other databases. This highlights the need of enabling a specialized tool that brings all the ARG information from different resources to easily integrate new ARGs or to validate the annotations of current ARGs. To overcome the difficulties in curation and manual validation of an extensive number of ARGs, a novel approach that breaks down this complex task into simpler and smaller microtasks is proposed. The core of this methodology consists of aggregating a compendium of AMR resources and deploying a crowdsourcing strategy, which simplifies the ARG information to allow non-experts, i.e. the general public, and domain experts collectively to execute curation of the ARG database. Application of crowdsourcing in biology, particularly for data curation, is not new and comprises a variety of areas including: name entity recognition (NER) for drugs and diseases (Islamaj Dogan et al., 2009; Khare et al., 2016; Lu et al., 2009), identification of medically relevant terms from patient online posts (MacLean and Heer, 2013), annotation of diseases described in PubMed (Good et al., 2014), and systematic examination of databases and other resources for drug indications, biomedical ontologies, and gene–disease interactions (Arighi et al., 2013; Khare et al., 2016; Lu and Hirschman, 2012; Wei et al., 2012; Wei et al., 2013). Interestingly, in most of the studies, crowdsourcing has proven to be as effective as expert curation (Good and Su, 2013; Khare et al., 2016). A major problem that encompasses all ARG resources is the lack of a standardized gene nomenclature. In particular, naming ARGs does not follow the general nomenclature for naming bacterial genes (Demerec et al., 1966). For instance, macrolide resistance genes are structured so that the class is indicated by brackets [e.g. ole(B), srm(B), vga(B) or ere(B)] (Levy et al., 1999). When compared to tetracycline genes, this gene nomenclature differs radically, because, in tetracycline genes, the determinant is placed as a capital letter after the gene name (e.g. tetA, tetB, tetC) (Levy et al., 1999). At the same time, those nomenclatures differ from the gene convention proposed to annotate beta-lactamase genes (Hall and Schwarz, 2016). Other examples to highlight these differences include the aminoglycoside gene (Vanhoof et al., 1998) aadA1 found under different names across the available ARG databases [ANT(3″)-I, aadA1-pm, ANT3-DPRIME and ant3ia]. Therefore, the diversity and variation in the ARG nomenclature and naming conventions complicate and greatly hinder consistent ARG curation. Here we introduce ARGminer, an online platform to enhance manual curation of ARGs. ARGminer enables users to curate and retrieve all of the information available from several ARG resources, including CARD (McArthur et al., 2013), DeepARG-DB (Arango-Argoty et al., 2018), ARDB (Liu and Pop, 2009), MEGARes (Lakin et al., 2017), UniProt (Leplae et al., 2004), the National Database of Antibiotic Resistant Organisms (NDARO) (https://www.ncbi.nlm.nih.gov/pathogens/antimicrobial-resistance/), the SARG (Yang et al., 2016), ResFinder (Zankari et al., 2012) and the ARG-ANNOT (Gupta et al., 2014) databases. Manual crowdsource-based curation is enhanced by a machine learning model based on word embeddings (Goldberg and Levy, 2014; Turian et al., 2010), a technique widely used in natural language processing (NLP) to aid in validation and achieve consistency in ARG nomenclature. On the other hand, MGEs such as plasmids phages and viruses play an important role in the dissemination of ARGs (Bengtsson-Palme et al., 2017; Gillings, 2014; Leplae et al., 2004). Therefore, ARGminer also interfaces with the PATRIC (Wattam et al., 2014) and Classification of Mobile Genetic Elements (ACLAME) (Leplae et al., 2004) databases, which provide information on potential carriage of ARGs by pathogens or MGEs, respectively. The ARGminer platform is designed, built and implemented as an open-source project facilitating a collaborative and integrative approach for the standardization of ARG annotation by the broad community of scientists and citizens motivated by a common desire to contribute toward combating the spread of AMR. ARGminer also includes a community blog for users to post questions and share solutions/discussion regarding AMR with the objective to keep the scientific community actively engaged in the latest updates and development of ARG databases (see Supplementary Fig. S1). All data associated with ARGminer, as well as the source code, is freely available under a public repository at http://bench.cs.vt.edu/argminer. ARGs were downloaded from the following resources: CARD (McArthur et al., 2013), which contains ARG information; the ARDB (Liu and Pop, 2009) database, which comprises a vast number of homology-predicted ARGs; DeepARG-DB (Arango-Argoty et al., 2018), which integrates ARGs from UniProt (UniProt Consortium, 2014), CARD and ARDB; MEGARes (Lakin et al., 2017) database, which incorporates genes from the ARG-ANNOT (Gupta et al., 2014), ResFinder (Zankari et al., 2012), the Lahey Clinic beta-lactamase archive (Bush and Jacoby, 2010) available from the National Center for Biotechnology Information (NCBI), the SARG database (Yang et al., 2016) and the NDARO database version 2. To obtain a comprehensive collection of ARGs, the DeepARG-DB database was updated with a more recent version of the CARD (v 2.0.4) and UniProt databases using their corresponding sequence identifiers. Discontinued UniProt sequences were removed from the ARGs from CARD were genes from CARD to resistance to were All sequences from all databases were to by using and an of The collection of ARGs was to all databases using et al., 2015) and et al., to the of ARG with corresponding In this ARG is by to database, consistency in annotation the ARG The metadata from the UniProt database is via the UniProt which of up-to-date information for Therefore, ARG is in the as a of an metadata and the alignment are as to enhance and (see Supplementary Fig. The database (Leplae et al., 2004) genes associated with and phages and was used to ARGs that have potential of by et al., 2015) was used to ARGs to MGEs via sequence alignment with is in the for users to a on an ARG has evidence of carried by an or This evidence is by users through the platform from a A the of for the information in the (see Supplementary Fig. A of bacterial were downloaded from the PATRIC (Wattam et al., 2014) database. This database contains information bacterial AMR phenotype, corresponding diseases and information that is particularly to to pathogens and potential antibiotic resistance For instance, the UniProt gene was present in bacterial of to are in in and and (see Supplementary Fig. The collection of ARGs were the sequences from PATRIC using et al., To the quality of the all genes with an of and an alignment of were were to the potential of bacterial of ARGs based on the evidence provided by PATRIC of and annotation task consists of ARGs based on the evidence provided on the are to an ARG in terms of gene antibiotic class and antibiotic In users are to the evidence of that this ARG on MGEs or A it to follow the annotation process Fig. crowdsourcing and the complex annotation into ARGminer from an community that includes experts and the general public to tackle the task of ARG ARGminer includes a machine learning model based on a word for the of the gene name nomenclature given the metadata information available from different databases. To this information such as ARG antibiotic and other data was from the CARD database (McArthur et al., 2013). In metadata and names were used for the were as the of the ARG For instance, the for is the for the tetracycline gene is and the for the gene is and to a letter or and a respectively. The nomenclature for and was built as obtain ARG names and metadata from CARD sequences to other databases SARG and and the corresponding metadata for the and the the was of the entries were for and the for This process was to consistency and of the a for and using word embeddings et al., 2017; et al., 2016) was used to the the model was with an of Supplementary the of the ARG nomenclature To the and quality of by domain experts are actively engaged in environmental ARG research were to annotate a of ARGs by antibiotic class and In out of the ARG annotations were in at of the experts in terms of antibiotic class and gene ARGs were considered in were by groups of was via an online platform that to a broad crowdsourcing to When a an annotation, ARGminer a that need to to the for validation to obtain a of the high diversity of the ARGminer were to a broad including domain experts and non-experts, and were to a limited number of annotations of of in a all of the deep AMR domain they all general with The ARGminer has or current available information for a given It consists of the gene antibiotic class, the database from which the sequence was and number of the gene has been by (see Fig. to the metadata available for the ARG as well as the from the ARDB and MEGARes databases. It also regarding the gene is carried by an and the gene is to be found in PATRIC database, Fig. to the a The information in this be consistent with the from the It consists of First, validate the gene antibiotic class and by at the Second, the and their annotation by their they are with and the evidence is, Fig. and Supplementary Fig. of the ARGminer This contains the current information available for the ARG that The enables to curate ARGs in the database that has This is the and all of the metadata and information from the different databases and in this there are that the of This is for users that are not with alignment This contains the microtasks for the ARG curation. It also contains which the errors The a for new users that is for for a The of this is to the with the platform by ARGminer also a of problematic ARGs that have problematic ARGs are identified by the annotation of the genes with their from CARD and All validation were using these problematic ARGs. ARGminer also an to update the ARG database. This comprises a of that the of different as well as the and evidence In this ARGminer are to the annotations made by the and update the ARG database (see Supplementary Fig. of the of users provide or the evidence and an even a increase the in ARG annotations and annotation To this ARGminer a (see Supplementary to the to evidence or This is in time, and the the will not to the Supplementary an of a for the antibiotic class ARGminer the curation of ARGs based on the described in et by the validation and the evidence and (see Supplementary for The to the ARG under is the most common provided by the In other are based on the or by the and A machine learning model to to the ARG nomenclature was into the ARGminer platform to to the gene name nomenclature. To the different gene name with at genes were identified (see Supplementary S1). to the of the nomenclature in ARGminer, and the validation were is the of the number of gene names the number of genes in the validation is the of the number of gene names the number of genes with this The model a high of and of in the validation of entries that were not used the Supplementary examples of gene names with their For instance, the UniProt has been by different resources as ANT3-DPRIME and the the gene name to have the with a To the of the crowdsourcing approach for ARG annotation, we groups of with the following A of crowdsourcing from to as In this were for annotation, with the to the of the general Therefore, as A of annotations were from for this A of crowdsourcing from to as In this the was The of this was to the of the In this a of annotations were from A of users with general with of in ARG to as This of and from a class at this as an and not any Here the annotations were with the validation on and was to annotations microtasks in The of the was to the community of are that to obtain by which is a major to effective In the present the ARGminer with how to the annotation of the For the antibiotic annotation the antibiotic class to which they the gene from a that contains a of antibiotic indicated that the first on the most the evidence of the it was that most of the antibiotic class annotations under the were as This is a to accurate database curation and the need for a that of the In terms of as the for all annotations However, not all were It was that few more than microtasks also and to their and of the were on their (see Fig. a the of the for all class, ARG name and ARG compared to the for annotation the new were not to with the their annotation was as in Supplementary this all was and all annotations from the to ARG evidence by the the of the for the of In it was to the of the domain The of this was to a community and a complex task with as those of a of with domain the a than the a to the was This that crowdsourcing is a powerful to manual inspection and annotation of ARGs by domain experts. annotations were characterized by a compared to the in all annotation the were not different of the the with the validation and a of with general domain and antibiotic resistance the However, this was to that by the with domain from the were the the the of the crowdsourcing annotation the was not the annotation and means the annotation was means a with the To the quality of the strategy, genes were the of curated genes and in as in For instance, the UniProt is a resistance in several including A process and to This the to which is required for resistance to and et al., the crowdsourcing and antibiotic were was characterized by a A at the evidence from the antibiotic resistance databases UniProt and a of the gene toward the antibiotic class, including gene annotations and literature PATRIC database also that this gene is carried by pathogens ARG is found in The evidence from the antibiotic resistance databases that the gene to a beta-lactamase illustrates different all that is the class with the annotation However, as a of the several were such as and even the word particularly is the similarity For instance, in the gene was to as and to the as differences are not by the Therefore, under the validation the of ARGminer have the to validate or the annotations that most the gene to the beta-lactamase that the name is the all of the antibiotic class annotation by the crowdsourcing all annotations with and the using the annotation is in the (see Supplementary on how this was to the antibiotic resistance the by the and the of to the ARG with are to the annotations from A gene with identified the same class that has different with the A gene with multiple annotations and a the identification of the antibiotic for the gene was the of gene name is the metadata of this does not include the gene name and the of is This that the gene has a potential to ARGs. ARG databases and a the For this of the the gene as the other it as a for the gene compared to the a a To any ARGminer that users the the evidence is not For the other examples and the the gene names to the The for the ARG name for the antibiotic annotation there are the annotations are ARG as and the same as and is the The gene was as or all corresponding to the gene with annotations the with the that all these were than the other gene names and and all the such as and all annotations were by the To the of the crowdsourcing annotation, genes that were by than were removed from the of a of genes were identified and curated by domain experts to antibiotic class and gene It was found that experts an annotation of of an the expert annotations were the gene annotations for expert were into with the genes are the same and the gene annotations differ and are from the of gene that were to the same by at experts were used as the (see Supplementary This was used to the of the crowdsourcing were based on the annotation The crowdsourcing of the antibiotic with the validation was as accurate as the expert annotation In other out of genes via crowdsourcing the expert (see Supplementary S1). The genes for which the to the antibiotic class were a ARG as and a gene as The of the ARG names to be a experts not the name of ARGs (see Supplementary However, of those genes was a different by all the experts. This gene to a macrolide gene which was as and by the experts. this gene was removed for the gene name and the When the gene name annotation from the crowdsourcing their a (see Supplementary This that gene was not by the the of this gene in ARGminer, all ARG databases that the gene to the group, with high However, CARD it as the ARDB it as and MEGARes it as the group, of a aspect with to this ARG is that are by et al., these genes are and the by using sequence alignment is a particularly as in Supplementary This aspect has the potential to genes that are Interestingly, by the were to the they were not to the and the have the same that crowdsourcing are to follow the even in the of complex Interestingly, the gene name model the gene name nomenclature to with a of which to the nomenclature to name beta-lactamase of the risk of propagation the process of ARGs for new database in the ARGminer the with the in that crowdsourcing annotation is a to manual and validation of ARGs by domain experts. ARGminer users to their own in the of ARGs on a of the of the the or annotations for the antibiotic all and it is that having expert does not a in the quality of the of the of the from most of the are not experts and have ARGs. it is also that the of annotations was compared to the of the the number of This that accurate of the antibiotic resistance does not necessarily require domain On the other hand, were also required to their in the that is a of the quality of the The of the that with more accurate For instance, of the their with annotation and This that the is a of annotation the and of the The of the the number of the to the and the the and and that the crowdsourcing approach has the risk of propagation and For it is that the ARG databases from which genes were by ARGminer have annotation and the annotation errors be from database to the databases were not propagation be an efforts are in the research community for of the ARG database that ARGminer is not to ARGs not present in the databases we Here we and validate a new ARGminer, a powerful that the of crowdsourcing for and comprehensive curation of ARGs. ARGminer enables to relevant information to ARGs, including evidence of ARGs carried by pathogens and the of ARGs by it enables a tool for the curation of ARGs designed to provide accurate information in a that be by users the of domain not that crowdsourcing as accurate as they are also more than experts. ARGminer the of a accurate and up-to-date available ARG database. was provided in for this by the of of and for Antimicrobial Resistance the National in and the for and Center for the and of the and the of Gustavo A. Arango-Argoty, G. K. P. Guron, Emily Garner, M. V. Riquelme, Lenwood S. Heath, Amy Pruden, Peter J. Vikesland, Liqing Zhang 0002 |
Bioinform. | 8 |
| 2018 | Combining weighted adaptive CS-LBP and local linear discriminant projection for gait recognition
Shanwen Zhang, Liqing Zhang 0002 |
Multim. Tools Appl. | 2 |
| 2015 | PseKNC-General: a cross-platform package for generating various modes of pseudo nucleotide compositionsabstractSUMMARY: The avalanche of genomic sequences generated in the post-genomic age requires efficient computational methods for rapidly and accurately identifying biological features from sequence information. Towards this goal, we developed a freely available and open-source package, called PseKNC-General (the general form of pseudo k-tuple nucleotide composition), that allows for fast and accurate computation of all the widely used nucleotide structural and physicochemical properties of both DNA and RNA sequences. PseKNC-General can generate several modes of pseudo nucleotide compositions, including conventional k-tuple nucleotide compositions, Moreau-Broto autocorrelation coefficient, Moran autocorrelation coefficient, Geary autocorrelation coefficient, Type I PseKNC and Type II PseKNC. In every mode, >100 physicochemical properties are available for choosing. Moreover, it is flexible enough to allow the users to calculate PseKNC with user-defined properties. The package can be run on Linux, Mac and Windows systems and also provides a graphical user interface. AVAILABILITY AND IMPLEMENTATION: The package is freely available at: http://lin.uestc.edu.cn/server/pseknc. Wei Chen 0064, Xitong Zhang, Jordan Brooker, Hao Lin 0001, Liqing Zhang 0002, Kuo-Chen Chou |
Bioinform. | 5 |
| 2015 | HMMvar-func: a new method for predicting the functional outcome of genetic variantsabstractBACKGROUND: Numerous tools have been developed to predict the fitness effects (i.e., neutral, deleterious, or beneficial) of genetic variants on corresponding proteins. However, prediction in terms of whether a variant causes the variant bearing protein to lose the original function or gain new function is also needed for better understanding of how the variant contributes to disease/cancer. To address this problem, the present work introduces and computationally defines four types of functional outcome of a variant: gain, loss, switch, and conservation of function. The deployment of multiple hidden Markov models is proposed to computationally classify mutations by the four functional impact types. RESULTS: The functional outcome is predicted for over a hundred thyroid stimulating hormone receptor (TSHR) mutations, as well as cancer related mutations in oncogenes or tumor suppressor genes. The results show that the proposed computational method is effective in fine grained prediction of the functional outcome of a mutation, and can be used to help elucidate the molecular mechanism of disease/cancer causing mutations. The program is freely available at http://bioinformatics.cs.vt.edu/zhanglab/HMMvar/download.php. CONCLUSION: This work is the first to computationally define and predict functional impact of mutations, loss, switch, gain, or conservation of function. These fine grained predictions can be especially useful for identifying mutations that cause or are linked to cancer. Mingming Liu 0006, Layne T. Watson, Liqing Zhang 0002 |
BMC Bioinform. | 3 |
| 2014 | Classification of Mutations by Functional Impact Type: Gain of Function, Loss of Function, and Switch of Function
Mingming Liu 0006, Layne T. Watson, Liqing Zhang 0002 |
ISBRA | 3 |
| 2014 | Vindel: a simple pipeline for checking indel redundancyabstractBACKGROUND: With the advance of next generation sequencing (NGS) technologies, a large number of insertion and deletion (indel) variants have been identified in human populations. Despite much research into variant calling, it has been found that a non-negligible proportion of the identified indel variants might be false positives due to sequencing errors, artifacts caused by ambiguous alignments, and annotation errors. RESULTS: In this paper, we examine indel redundancy in dbSNP, one of the central databases for indel variants, and develop a standalone computational pipeline, dubbed Vindel, to detect redundant indels. The pipeline first applies indel position information to form candidate redundant groups, then performs indel mutations to the reference genome to generate corresponding indel variant substrings. Finally the indel variant substrings in the same candidate redundant groups are compared in a pairwise fashion to identify redundant indels. We applied our pipeline to check for redundancy in the human indels in dbSNP. Our pipeline identified approximately 8% redundancy in insertion type indels, 12% in deletion type indels, and overall 10% for insertions and deletions combined. These numbers are largely consistent across all human autosomes. We also investigated indel size distribution and adjacent indel distance distribution for a better understanding of the mechanisms generating indel variants. CONCLUSIONS: Vindel, a simple yet effective computational pipeline, can be used to check whether a set of indels are redundant with respect to those already in the database of interest such as NCBI's dbSNP. Of the approximately 5.9 million indels we examined, nearly 0.6 million are redundant, revealing a serious limitation in the current indel annotation. Statistics results prove the consistency of the pipeline on indel redundancy detection for all 22 chromosomes. Apart from the standalone Vindel pipeline, the indel redundancy check algorithm is also implemented in the web server http://bioinformatics.cs.vt.edu/zhanglab/indelRedundant.php . Xiaowei Wu 0003, Liqing Zhang 0002 |
BMC Bioinform. | 4 |
| 2014 | Quantitative prediction of the effect of genetic variation using hidden Markov modelsabstractBACKGROUND: With the development of sequencing technologies, more and more sequence variants are available for investigation. Different classes of variants in the human genome have been identified, including single nucleotide substitutions, insertion and deletion, and large structural variations such as duplications and deletions. Insertion and deletion (indel) variants comprise a major proportion of human genetic variation. However, little is known about their effects on humans. The absence of understanding is largely due to the lack of both biological data and computational resources. RESULTS: This paper presents a new indel functional prediction method HMMvar based on HMM profiles, which capture the conservation information in sequences. The results demonstrate that a scoring strategy based on HMM profiles can achieve good performance in identifying deleterious or neutral variants for different data sets, and can predict the protein functional effects of both single and multiple mutations. CONCLUSIONS: This paper proposed a quantitative prediction method, HMMvar, to predict the effect of genetic variation using hidden Markov models. The HMM based pipeline program implementing the method HMMvar is freely available at https://bioinformatics.cs.vt.edu/zhanglab/hmm. Mingming Liu 0006, Layne T. Watson, Liqing Zhang 0002 |
BMC Bioinform. | 3 |
| 2011 | A Network of SCOP Hidden Markov Models and Its AnalysisabstractBACKGROUND: The Structural Classification of Proteins (SCOP) database uses a large number of hidden Markov models (HMMs) to represent families and superfamilies composed of proteins that presumably share the same evolutionary origin. However, how the HMMs are related to one another has not been examined before. RESULTS: In this work, taking into account the processes used to build the HMMs, we propose a working hypothesis to examine the relationships between HMMs and the families and superfamilies that they represent. Specifically, we perform an all-against-all HMM comparison using the HHsearch program (similar to BLAST) and construct a network where the nodes are HMMs and the edges connect similar HMMs. We hypothesize that the HMMs in a connected component belong to the same family or superfamily more often than expected under a random network connection model. Results show a pattern consistent with this working hypothesis. Moreover, the HMM network possesses features distinctly different from the previously documented biological networks, exemplified by the exceptionally high clustering coefficient and the large number of connected components. CONCLUSIONS: The current finding may provide guidance in devising computational methods to reduce the degree of overlaps between the HMMs representing the same superfamilies, which may in turn enable more efficient large-scale sequence searches against the database of HMMs. Liqing Zhang 0002, Layne T. Watson, Lenwood S. Heath |
BMC Bioinform. | 1 |
| 2010 | The Expected Fitness Cost of a Mutation Fixation under the One-Dimensional Fisher Model
Liqing Zhang 0002, Layne T. Watson |
ISBRA | 1 |
| 2010 | MSOAR 2.0: Incorporating tandem duplications into ortholog assignment based on genome rearrangementabstractBACKGROUND: Ortholog assignment is a critical and fundamental problem in comparative genomics, since orthologs are considered to be functional counterparts in different species and can be used to infer molecular functions of one species from those of other species. MSOAR is a recently developed high-throughput system for assigning one-to-one orthologs between closely related species on a genome scale. It attempts to reconstruct the evolutionary history of input genomes in terms of genome rearrangement and gene duplication events. It assumes that a gene duplication event inserts a duplicated gene into the genome of interest at a random location (i.e., the random duplication model). However, in practice, biologists believe that genes are often duplicated by tandem duplications, where a duplicated gene is located next to the original copy (i.e., the tandem duplication model). RESULTS: In this paper, we develop MSOAR 2.0, an improved system for one-to-one ortholog assignment. For a pair of input genomes, the system first focuses on the tandemly duplicated genes of each genome and tries to identify among them those that were duplicated after the speciation (i.e., the so-called inparalogs), using a simple phylogenetic tree reconciliation method. For each such set of tandemly duplicated inparalogs, all but one gene will be deleted from the concerned genome (because they cannot possibly appear in any one-to-one ortholog pairs), and MSOAR is invoked. Using both simulated and real data experiments, we show that MSOAR 2.0 is able to achieve a better sensitivity and specificity than MSOAR. In comparison with the well-known genome-scale ortholog assignment tool InParanoid, Ensembl ortholog database, and the orthology information extracted from the well-known whole-genome multiple alignment program MultiZ, MSOAR 2.0 shows the highest sensitivity. Although the specificity of MSOAR 2.0 is slightly worse than that of InParanoid in the real data experiments, it is actually better than that of InParanoid in the simulation tests. CONCLUSIONS: Our preliminary experimental results demonstrate that MSOAR 2.0 is a highly accurate tool for one-to-one ortholog assignment between closely related genomes. The software is available to the public for free and included as online supplementary material. Guanqun Shi, Liqing Zhang 0002, Tao Jiang 0001 |
BMC Bioinform. | 2 |
| 2009 | Identifying disease associations via genome-wide association studiesabstractBACKGROUND: Genome-wide association studies prove to be a powerful approach to identify the genetic basis of different human diseases. We studied the relationship between seven diseases characterized in a previous genome-wide association study by the Wellcome Trust Case Control Consortium. Instead of doing a horizontal association of SNPs to diseases, we did a vertical analysis of disease associations by comparing the genetic similarities of diseases. Our analysis was carried out at four levels - the nucleotide level (SNPs), the gene level, the protein level (through protein-protein interaction network), and the phenotype level. RESULTS: Our results show that Crohn's disease, rheumatoid arthritis, and type 1 diabetes share evidence of genetic associations at all levels of analysis, offering strong molecular support for the current grouping of the diseases. On the other hand, coronary artery disease, hypertension, and type 2 diabetes, despite being considered as a natural group with potential aetiological overlap, do not show any evidence of shared genetic basis at all levels. CONCLUSION: Our study is a first attempt on mining of GWA data to examine genetic associations between different diseases. The positive result is apparently not a coincidence and hence demonstrates the promising use of our approach. Pengyuan Wang 0005, Liqing Zhang 0002 |
BMC Bioinform. | 4 |
| 2008 | A Framework to Understand the Mechanism of Toxicity
Liqing Zhang 0002 |
ICIC (2) | 2 |
| 2008 | Using Cost-Sensitive Learning to Determine Gene Conversions
Mark J. Lawson, Lenwood S. Heath, Naren Ramakrishnan, Liqing Zhang 0002 |
ICIC (2) | 4 |
| 2008 | A Holistic View of Evolutionary Rates in Paralogous and Orthologous Genes
Liqing Zhang 0002 |
ICIC (2) | 2 |
| 2007 | The Identification of Antisense Gene Pairs Through Available Software
Mark J. Lawson, Liqing Zhang 0002 |
ISBRA | 2 |