EDBT 2026 Demo / reviewers in the wild / expert
Pengli Cai
dblp:276/1453
· DBLP profile ↗
9ranked-venue papers
1as first author
8since 2021 · last 2024
0000-0002-7910-6558ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 9 · 1 first-author · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | High-throughput prediction of enzyme promiscuity based on substrate-product pairsabstractThe screening of enzymes for catalyzing specific substrate-product pairs is often constrained in the realms of metabolic engineering and synthetic biology. Existing tools based on substrate and reaction similarity predominantly rely on prior knowledge, demonstrating limited extrapolative capabilities and an inability to incorporate custom candidate-enzyme libraries. Addressing these limitations, we have developed the Substrate-product Pair-based Enzyme Promiscuity Prediction (SPEPP) model. This innovative approach utilizes transfer learning and transformer architecture to predict enzyme promiscuity, thereby elucidating the intricate interplay between enzymes and substrate-product pairs. SPEPP exhibited robust predictive ability, eliminating the need for prior knowledge of reactions and allowing users to define their own candidate-enzyme libraries. It can be seamlessly integrated into various applications, including metabolic engineering, de novo pathway design, and hazardous material degradation. To better assist metabolic engineers in designing and refining biochemical pathways, particularly those without programming skills, we also designed EnzyPick, an easy-to-use web server for enzyme screening based on SPEPP. EnzyPick is accessible at http://www.biosynther.com/enzypick/. Huadong Xing, Pengli Cai, Mengying Han, Yingying Le, Dachuan Zhang, Qian-Nan Hu |
Briefings Bioinform. | 2 |
| 2023 | RDBridge: a knowledge graph of rare diseases based on large-scale text miningabstractMOTIVATION: Despite low prevalence, rare diseases affect 300 million people worldwide. Research on pathogenesis and drug development lags due to limited commercial potential, insufficient epidemiological data, and a dearth of publications. The unique characteristics of rare diseases, including limited annotated data, intricate processes for extracting pertinent entity relationships, and difficulties in standardizing data, represent challenges for text mining. RESULTS: We developed a rare disease data acquisition framework using text mining and knowledge graphs and constructed the most comprehensive rare disease knowledge graph to date, Rare Disease Bridge (RDBridge). RDBridge offers search functions for genes, potential drugs, pathways, literature, and medical imaging data that will support mechanistic research, drug development, diagnosis, and treatment for rare diseases. AVAILABILITY AND IMPLEMENTATION: RDBridge is freely available at http://rdb.lifesynther.com/. Huadong Xing, Dachuan Zhang, Pengli Cai, Qian-Nan Hu |
Bioinform. | 3 |
| 2023 | SynBioTools: a one-stop facility for searching and selecting synthetic biology toolsabstractBACKGROUND: The rapid development of synthetic biology relies heavily on the use of databases and computational tools, which are also developing rapidly. While many tool registries have been created to facilitate tool retrieval, sharing, and reuse, no relatively comprehensive tool registry or catalog addresses all aspects of synthetic biology. RESULTS: We constructed SynBioTools, a comprehensive collection of synthetic biology databases, computational tools, and experimental methods, as a one-stop facility for searching and selecting synthetic biology tools. SynBioTools includes databases, computational tools, and methods extracted from reviews via SCIentific Table Extraction, a scientific table-extraction tool that we built. Approximately 57% of the resources that we located and included in SynBioTools are not mentioned in bio.tools, the dominant tool registry. To improve users' understanding of the tools and to enable them to make better choices, the tools are grouped into nine modules (each with subdivisions) based on their potential biosynthetic applications. Detailed comparisons of similar tools in every classification are included. The URLs, descriptions, source references, and the number of citations of the tools are also integrated into the system. CONCLUSIONS: SynBioTools is freely available at https://synbiotools.lifesynther.com/ . It provides end-users and developers with a useful resource of categorized synthetic biology databases, tools, and methods to facilitate tool retrieval and selection. Pengli Cai, Sheng Liu 0028, Dachuan Zhang, Huadong Xing, Mengying Han, Linlin Gong, Qian-Nan Hu |
BMC Bioinform. | 1 |
| 2022 | BioBulkFoundary: a customized webserver for exploring biosynthetic potentials of bulk chemicalsabstractSUMMARY: Advances in metabolic engineering have boosted the production of bulk chemicals, resulting in tons of production volumes of some bulk chemicals with very low prices. A decrease in the production cost and overproduction of bulk chemicals makes it necessary and desirable to explore the potential to synthesize higher-value products from them. It is also useful and important for society to explore the use of design methods involving synthetic biology to increase the economic value of these bulk chemicals. Therefore, we developed 'BioBulkFoundary', which provides an elaborate analysis of the biosynthetic potential of bulk chemicals based on the state-of-art exploration of pathways to synthesize value-added chemicals, along with associated comprehensive technology and economic database into a user-friendly framework. AVAILABILITY AND IMPLEMENTATION: Freely available on the web at http://design.rxnfinder.org/biobulkfoundary/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Dandan Sun, Shaozhen Ding, Pengli Cai, Dachuan Zhang, Mengying Han, Qian-Nan Hu |
Bioinform. | 3 |
| 2021 | ChemHub: a knowledgebase of functional chemicals for synthetic biology studiesabstractSUMMARY: The field of synthetic biology lacks a comprehensive knowledgebase for selecting synthetic target molecules according to their functions, economic applications and known biosynthetic pathways. We implemented ChemHub, a knowledgebase containing >90 000 chemicals and their functions, along with related biosynthesis information for these chemicals that was manually extracted from >600 000 published studies by more than 100 people over the past 10 years. AVAILABILITY AND IMPLEMENTATION: Multiple algorithms were implemented to enable biosynthetic pathway design and precursor discovery, which can support investigation of the biosynthetic potential of these functional chemicals. ChemHub is freely available at: http://www.rxnfinder.org/chemhub/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Mengying Han, Dachuan Zhang, Shaozhen Ding, Yu Tian 0006, Xingxiang Cheng, Le Yuan, Dandan Sun, Linlin Gong, Cancan Jia, Pengli Cai, Weizhong Tu, Junni Chen, Qian-Nan Hu |
Bioinform. | 11 |
| 2021 | Cell2Chem: mining explored and unexplored biosynthetic chemical spacesabstractSUMMARY: Living cell strains have important applications in synthesizing their native compounds and potential for use in studies exploring the universal chemical space. Here, we present a web server named as Cell2Chem which accelerates the search for explored compounds in organisms, facilitating investigations of biosynthesis in unexplored chemical spaces. Cell2Chem uses co-occurrence networks and natural language processing to provide a systematic method for linking living organisms to biosynthesized compounds and the processes that produce these compounds. The Cell2Chem platform comprises 40 370 species and 125 212 compounds. Using reaction pathway and enzyme function in silico prediction methods, Cell2Chem reveals possible biosynthetic pathways of compounds and catalytic functions of proteins to expand unexplored biosynthetic chemical spaces. Cell2Chem can help improve biosynthesis research and enhance the efficiency of synthetic biology. AVAILABILITY AND IMPLEMENTATION: Cell2Chem is available at: http://www.rxnfinder.org/cell2chem/. Mengying Han, Yu Tian 0006, Linlin Gong, Cancan Jia, Pengli Cai, Weizhong Tu, Junni Chen, Qian-Nan Hu |
Bioinform. | 6 |
| 2021 | Transcriptor: a comprehensive platform for annotation of the enzymatic functions of transcriptsabstractMOTIVATION: Rapid advances in sequencing technology have resulted huge increases in the accessibility of sequencing data. Moreover, researchers are focusing more on organisms that lack a reference genome. However, few easy-to-use web servers focusing on annotations of enzymatic functions are available. Accordingly, in this study, we describe Transcriptor, a novel platform for annotating transcripts encoding enzymes. RESULTS: The transcripts were evaluated using more than 300 000 in-house enzymatic reactions through bridges of Enzyme Commission numbers. Transcriptor also enabled ontology term identification and along with associated enzymes, visualization and prediction of domains and annotation of regulatory structure, such as long noncoding RNAs, which could facilitate the discovery of new functions in model or nonmodel species. Transcriptor may have applications in elucidation of the roles of organs transcriptomes and secondary metabolite biosynthesis in organisms lacking a reference genome. AVAILABILITY AND IMPLEMENTATION: Transcriptor is available at http://design.rxnfinder.org/transcriptor/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ailin Ren, Dachuan Zhang, Yu Tian 0006, Pengli Cai, Qian-Nan Hu |
Bioinform. | 4 |
| 2021 | SARS2020: an integrated platform for identification of novel coronavirus by a consensus sequence-function modelabstractMOTIVATION: The 2019 novel coronavirus outbreak has significantly affected global health and society. Thus, predicting biological function from pathogen sequence is crucial and urgently needed. However, little work has been conducted to identify viruses by the enzymes that they encode, and which are key to pathogen propagation. RESULTS: We built a comprehensive scientific resource, SARS2020, which integrates coronavirus-related research, genomic sequences and results of anti-viral drug trials. In addition, we built a consensus sequence-catalytic function model from which we identified the novel coronavirus as encoding the same proteinase as the severe acute respiratory syndrome virus. This data-driven sequence-based strategy will enable rapid identification of agents responsible for future epidemics. AVAILABILITYAND IMPLEMENTATION: SARS2020 is available at http://design.rxnfinder.org/sars2020/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Dachuan Zhang, Sheng Liu 0028, Dandan Sun, Shaozhen Ding, Xingxiang Cheng, Pengli Cai, Ailin Ren, Mengying Han, Cancan Jia, Linlin Gong, Huadong Xing, Weizhong Tu, Junni Chen, Qian-Nan Hu |
Bioinform. | 7 |
| 2020 | RxnBLAST: molecular scaffold and reactive chemical environment feature extractor for biochemical reactionsabstractMOTIVATION: Molecular scaffolds are useful in medicinal chemistry to describe, discuss and visualize series of chemical compounds, biochemical transformations and associated biological properties. RESULTS: Here, we present RxnBLAST as a web-based tool for analyzing scaffold transformations and reactive chemical environment features in bioreactions. RxnBLAST extracts chemical features from bioreactions including atom-atom mapping, reaction centers, rules and functional groups to help understand chemical compositions and reaction patterns. Core-to-Core is proposed, which can be utilized in scaffold networks and for constructing a reaction space, as well as providing guidance for subsequent biosynthesis efforts. AVAILABILITY AND IMPLEMENTATION: RxnBLAST is available at: http://design.rxnfinder.org/rxnblast/. Xingxiang Cheng, Dandan Sun, Dachuan Zhang, Yu Tian 0006, Shaozhen Ding, Pengli Cai, Qian-Nan Hu |
Bioinform. | 6 |