EDBT 2026 Demo / reviewers in the wild / expert
Feng Zhu 0004
dblp:71/2791-4
· DBLP profile ↗
34ranked-venue papers
0as first author
23since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 33 · 22 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Federated Learning Meets Test-Time Adaptation: Methods, Challenges, and Future Directions
Ge Su, Huaxiao Zhou, Lu Hao, Feng Zhu 0004, Jintai Chen, Jianwei Yin |
Int. J. Comput. Vis. | 5 |
| 2026 | Causal Graph Learning for Face-Based Interpretable Hierarchical Diagnosis of DepressionabstractDepression has become one of the most serious mental illnesses, leading to a substantial decline in quality of life, an elevated risk of suicide, and significant societal challenges. Despite significant progress in the application of deep learning for depression diagnosis, most prevalent methods rely on correlative rather than causal features, limiting their accuracy and interpretability. Here, we propose a causal graph learning (CGL) method for the hierarchical diagnosis of depression. Specifically, we first construct a novel depression facial graph (DFGraph) structure based on a prior knowledge, which collects information about subjects’ facial cues. Our CGL model leverages the DFGraph structure and incorporates a built-in masking mechanism, which is designed to effectively differentiate causal features from confounding ones. It employs backdoor adjustment techniques, which control for confounding variables by blocking noncausal paths, to identify and select pertinent causal features, thereby enhancing the accuracy of the hierarchical diagnosis of depression. We conducted extensive experiments on the collected depression dataset. Our results show that the proposed method provides better results and interpretability is further improved compared to the publicly available baseline. Baoliang Zhang, Dixin Wang, Mingmei Cheng, Xiaofeng Liu 0001, Yanzhong Wang, Feng Zhu 0004, Zhixiong Lin, Chuan Shi 0006, Wanqing Xie |
IEEE Trans. Comput. Soc. Syst. | 6 |
| 2025 | Expanding the sequence spaces of synthetic binding protein using deep learning-based framework ProteinMPNN
Wantong Jiao, Ruihan Liu, Xuejin Deng, Feng Zhu 0004, Weiwei Xue |
Frontiers Comput. Sci. | 5 |
| 2025 | Decoding Drug Response With Structurized Gridding Map-Based Cell RepresentationabstractA thorough understanding of cell-line drug response mechanisms is crucial for drug development, repurposing, and resistance reversal. While targeted anticancer therapies have shown promise, not all cancers have well-established biomarkers to stratify drug response. Single-gene associations only explain a small fraction of the observed drug sensitivity, so a more comprehensive method is needed. However, while deep learning models have shown promise in predicting drug response in cell lines, they still face significant challenges when it comes to their application in clinical applications. Therefore, this study proposed a new strategy called DD-Response for cell-line drug response prediction. First, a limitation of narrow modeling horizons was overcome to expand the model training domain by integrating multiple datasets through source-specific label binarization. Second, a modified representation based on a two-dimensional structurized gridding map (SGM) was developed for cell lines & drugs, avoiding feature correlation neglect and potential information loss. Third, a dual-branch, multi-channel convolutional neural network-based model for pairwise response prediction was constructed, enabling accurate outcomes and improved exploration of underlying mechanisms. As a result, the DD-Response demonstrated superior performance, captured cell-line characteristic variations, and provided insights into key factors impacting cell-line drug response. In addition, DD-Response exhibited scalability in predicting clinical patient responses to drug therapy. Overall, because of DD-response's excellent ability to predict drug response and capture key molecules behind them, DD-response is expected to greatly facilitate drug discovery, repurposing, resistance reversal, and therapeutic optimization. Jiayi Yin, Xiuna Sun, Nanxin You, Minjie Mou, Mingkun Lu, Feng Cheng Li, Honglin Li 0003, Su Zeng, Feng Zhu 0004 |
IEEE J. Biomed. Health Informatics | 11 |
| 2024 | DDID: a comprehensive resource for visualization and analysis of diet-drug interactionsabstractDiet-drug interactions (DDIs) are pivotal in drug discovery and pharmacovigilance. DDIs can modify the systemic bioavailability/pharmacokinetics of drugs, posing a threat to public health and patient safety. Therefore, it is crucial to establish a platform to reveal the correlation between diets and drugs. Accordingly, we have established a publicly accessible online platform, known as Diet-Drug Interactions Database (DDID, https://bddg.hznu.edu.cn/ddid/), to systematically detail the correlation and corresponding mechanisms of DDIs. The platform comprises 1338 foods/herbs, encompassing flora and fauna, alongside 1516 widely used drugs and 23 950 interaction records. All interactions are meticulously scrutinized and segmented into five categories, thereby resulting in evaluations (positive, negative, no effect, harmful and possible). Besides, cross-linkages between foods/herbs, drugs and other databases are furnished. In conclusion, DDID is a useful resource for comprehending the correlation between foods, herbs and drugs and holds a promise to enhance drug utilization and research on drug combinations. Yan-Feng Hong, Hongquan Xu, Sisi Zhu, Gong-Xing Chen, Feng Zhu 0004 |
Briefings Bioinform. | 7 |
| 2024 | MoDAFold: a strategy for predicting the structure of missense mutant protein based on AlphaFold2 and molecular dynamicsabstractProtein structure prediction is a longstanding issue crucial for identifying new drug targets and providing a mechanistic understanding of protein functions. To enhance the progress in this field, a spectrum of computational methodologies has been cultivated. AlphaFold2 has exhibited exceptional precision in predicting wild-type protein structures, with performance exceeding that of other methods. However, predicting the structures of missense mutant proteins using AlphaFold2 remains challenging due to the intricate and substantial structural alterations caused by minor sequence variations in the mutant proteins. Molecular dynamics (MD) has been validated for precisely capturing changes in amino acid interactions attributed to protein mutations. Therefore, for the first time, a strategy entitled 'MoDAFold' was proposed to improve the accuracy and reliability of missense mutant protein structure prediction by combining AlphaFold2 with MD. Multiple case studies have confirmed the superior performance of MoDAFold compared to other methods, particularly AlphaFold2. Lingyan Zheng, Shuiyang Shi, Xiuna Sun, Mingkun Lu, Yang Liao, Sisi Zhu, Hongning Zhang, Pan Fang, Zhenyu Zeng, Honglin Li 0003, Zhaorong Li, Weiwei Xue, Feng Zhu 0004 |
Briefings Bioinform. | 14 |
| 2024 | FERREG: ferroptosis-based regulation of disease occurrence, progression and therapeutic responseabstractFerroptosis is a non-apoptotic, iron-dependent regulatory form of cell death characterized by the accumulation of intracellular reactive oxygen species. In recent years, a large and growing body of literature has investigated ferroptosis. Since ferroptosis is associated with various physiological activities and regulated by a variety of cellular metabolism and mitochondrial activity, ferroptosis has been closely related to the occurrence and development of many diseases, including cancer, aging, neurodegenerative diseases, ischemia-reperfusion injury and other pathological cell death. The regulation of ferroptosis mainly focuses on three pathways: system Xc-/GPX4 axis, lipid peroxidation and iron metabolism. The genes involved in these processes were divided into driver, suppressor and marker. Importantly, small molecules or drugs that mediate the expression of these genes are often good treatments in the clinic. Herein, a newly developed database, named 'FERREG', is documented to (i) providing the data of ferroptosis-related regulation of diseases occurrence, progression and drug response; (ii) explicitly describing the molecular mechanisms underlying each regulation; and (iii) fully referencing the collected data by cross-linking them to available databases. Collectively, FERREG contains 51 targets, 718 regulators, 445 ferroptosis-related drugs and 158 ferroptosis-related disease responses. FERREG can be accessed at https://idrblab.org/ferreg/. Mengjie Yang, Fengyun Chen, Jiayi Yin, Yintao Zhang, Xuheng Zhou, Xiuna Sun, Ziheng Ni, Qun Lv, Feng Zhu 0004, Shuiping Liu |
Briefings Bioinform. | 12 |
| 2023 | Label-free proteome quantification and evaluationabstractThe label-free quantification (LFQ) has emerged as an exceptional technique in proteomics owing to its broad proteome coverage, great dynamic ranges and enhanced analytical reproducibility. Due to the extreme difficulty lying in an in-depth quantification, the LFQ chains incorporating a variety of transformation, pretreatment and imputation methods are required and constructed. However, it remains challenging to determine the well-performing chain, owing to its strong dependence on the studied data and the diverse possibility of integrated chains. In this study, an R package EVALFQ was therefore constructed to enable a performance evaluation on >3000 LFQ chains. This package is unique in (a) automatically evaluating the performance using multiple criteria, (b) exploring the quantification accuracy based on spiking proteins and (c) discovering the well-performing chains by comprehensive assessment. All in all, because of its superiority in assessing from multiple perspectives and scanning among over 3000 chains, this package is expected to attract broad interests from the fields of proteomic quantification. The package is available at https://github.com/idrblab/EVALFQ. Jianbo Fu, Qingxia Yang, Yongchao Luo, Ying Zhang 0061, Hongning Zhang, Hanxiang Xu, Feng Zhu 0004 |
Briefings Bioinform. | 9 |
| 2023 | A novel strategy for designing the magic shotguns for distantly related target pairsabstractDue to its promising capacity in improving drug efficacy, polypharmacology has emerged to be a new theme in the drug discovery of complex disease. In the process of novel multi-target drugs (MTDs) discovery, in silico strategies come to be quite essential for the advantage of high throughput and low cost. However, current researchers mostly aim at typical closely related target pairs. Because of the intricate pathogenesis networks of complex diseases, many distantly related targets are found to play crucial role in synergistic treatment. Therefore, an innovational method to develop drugs which could simultaneously target distantly related target pairs is of utmost importance. At the same time, reducing the false discovery rate in the design of MTDs remains to be the daunting technological difficulty. In this research, effective small molecule clustering in the positive dataset, together with a putative negative dataset generation strategy, was adopted in the process of model constructions. Through comprehensive assessment on 10 target pairs with hierarchical similarity-levels, the proposed strategy turned out to reduce the false discovery rate successfully. Constructed model types with much smaller numbers of inhibitor molecules gained considerable yields and showed better false-hit controllability than before. To further evaluate the generalization ability, an in-depth assessment of high-throughput virtual screening on ChEMBL database was conducted. As a result, this novel strategy could hierarchically improve the enrichment factors for each target pair (especially for those distantly related/unrelated target pairs), corresponding to target pair similarity-levels. Yongchao Luo, Minjie Mou, Hanqi Zheng, Jiajun Hong, Feng Zhu 0004 |
Briefings Bioinform. | 7 |
| 2023 | SoCube: an innovative end-to-end doublet detection algorithm for analyzing scRNA-seq dataabstractDoublets formed during single-cell RNA sequencing (scRNA-seq) severely affect downstream studies, such as differentially expressed gene analysis and cell trajectory inference, and limit the cellular throughput of scRNA-seq. Several doublet detection algorithms are currently available, but their generalization performance could be further improved due to the lack of effective feature-embedding strategies with suitable model architectures. Therefore, SoCube, a novel deep learning algorithm, was developed to precisely detect doublets in various types of scRNA-seq data. SoCube (i) proposed a novel 3D composite feature-embedding strategy that embedded latent gene information and (ii) constructed a multikernel, multichannel CNN-ensembled architecture in conjunction with the feature-embedding strategy. With its excellent performance on benchmark evaluation and several downstream tasks, it is expected to be a powerful algorithm to detect and remove doublets in scRNA-seq data. SoCube is freely provided as an end-to-end tool on the Python official package site PyPi (https://pypi.org/project/socube/) and open-source on GitHub (https://github.com/idrblab/socube/). Hongning Zhang, Mingkun Lu, Gaole Lin, Lingyan Zheng, Wei Zhang 0218, Feng Zhu 0004 |
Briefings Bioinform. | 7 |
| 2023 | OBMeta: a comprehensive web server to analyze and validate gut microbial features and biomarkers for obesity-associated metabolic diseasesabstractMOTIVATION: Gut dysbiosis is closely associated with obesity and related metabolic diseases including type 2 diabetes (T2D) and nonalcoholic fatty liver disease (NAFLD). The gut microbial features and biomarkers have been increasingly investigated in many studies, which require further validation due to the limited sample size and various confounding factors that may affect microbial compositions in a single study. So far, it lacks a comprehensive bioinformatics pipeline providing automated statistical analysis and integrating multiple independent studies for cross-validation simultaneously. RESULTS: OBMeta aims to streamline the standard metagenomics data analysis from diversity analysis, comparative analysis, and functional analysis to co-abundance network analysis. In addition, a curated database has been established with a total of 90 public research projects, covering three different phenotypes (Obesity, T2D, and NAFLD) and more than five different intervention strategies (exercise, diet, probiotics, medication, and surgery). With OBMeta, users can not only analyze their research projects but also search and match public datasets for cross-validation. Moreover, OBMeta provides cross-phenotype and cross-intervention-based advanced validation that maximally supports preliminary findings from an individual study. To summarize, OBMeta is a comprehensive web server to analyze and validate gut microbial features and biomarkers for obesity-associated metabolic diseases. AVAILABILITY AND IMPLEMENTATION: OBMeta is freely available at: http://obmeta.met-bioinformatics.cn/. Cuifang Xu, Jiating Huang, Yongqiang Gao, Weixing Zhao, Yiqi Shen, Feihong Luo, Feng Zhu 0004, Yan Ni |
Bioinform. | 8 |
| 2023 | RNAenrich: a web server for non-coding RNA enrichmentabstractMOTIVATION: With the rapid advances of RNA sequencing and microarray technologies in non-coding RNA (ncRNA) research, functional tools that perform enrichment analysis for ncRNAs are needed. On the one hand, because of the rapidly growing interest in circRNAs, snoRNAs, and piRNAs, it is essential to develop tools for enrichment analysis for these newly emerged ncRNAs. On the other hand, due to the key role of ncRNAs' interacting target in the determination of their function, the interactions between ncRNA and its corresponding target should be fully considered in functional enrichment. Based on the ncRNA-mRNA/protein-function strategy, some tools have been developed to functionally analyze a single type of ncRNA (the majority focuses on miRNA); in addition, some tools adopt predicted target data and lead to only low-confidence results. RESULTS: Herein, an online tool named RNAenrich was developed to enable the comprehensive and accurate enrichment analysis of ncRNAs. It is unique in (i) realizing the enrichment analysis for various RNA types in humans and mice, such as miRNA, lncRNA, circRNA, snoRNA, piRNA, and mRNA; (ii) extending the analysis by introducing millions of experimentally validated data of RNA-target interactions as a built-in database; and (iii) providing a comprehensive interacting network among various ncRNAs and targets to facilitate the mechanistic study of ncRNA function. Importantly, RNAenrich led to a more comprehensive and accurate enrichment analysis in a COVID-19-related miRNA case, which was largely attributed to its coverage of comprehensive ncRNA-target interactions. AVAILABILITY AND IMPLEMENTATION: RNAenrich is now freely accessible at https://idrblab.org/rnaenr/. Kuerbannisha Amahong, Yintao Zhang, Mingkun Lu, Zhenyu Zeng, Zhaorong Li, Yunqing Qiu, Haibin Dai, Jianqing Gao, Feng Zhu 0004 |
Bioinform. | 13 |
| 2022 | ConSIG: consistent discovery of molecular signature from OMIC dataabstractThe discovery of proper molecular signature from OMIC data is indispensable for determining biological state, physiological condition, disease etiology, and therapeutic response. However, the identified signature is reported to be highly inconsistent, and there is little overlap among the signatures identified from different biological datasets. Such inconsistency raises doubts about the reliability of reported signatures and significantly hampers its biological and clinical applications. Herein, an online tool, ConSIG, was constructed to realize consistent discovery of gene/protein signature from any uploaded transcriptomic/proteomic data. This tool is unique in a) integrating a novel strategy capable of significantly enhancing the consistency of signature discovery, b) determining the optimal signature by collective assessment, and c) confirming the biological relevance by enriching the disease/gene ontology. With the increasingly accumulated concerns about signature consistency and biological relevance, this online tool is expected to be used as an essential complement to other existing tools for OMIC-based signature discovery. ConSIG is freely accessible to all users without login requirement at https://idrblab.org/consig/. Feng Cheng Li, Jiayi Yin, Mingkun Lu, Qingxia Yang, Zhenyu Zeng, Zhaorong Li, Yunqing Qiu, Haibin Dai, Yuzong Chen 0002, Feng Zhu 0004 |
Briefings Bioinform. | 11 |
| 2022 | POSREG: proteomic signature discovered by simultaneously optimizing its reproducibility and generalizabilityabstractMass spectrometry-based proteomic technique has become indispensable in current exploration of complex and dynamic biological processes. Instrument development has largely ensured the effective production of proteomic data, which necessitates commensurate advances in statistical framework to discover the optimal proteomic signature. Current framework mainly emphasizes the generalizability of the identified signature in predicting the independent data but neglects the reproducibility among signatures identified from independently repeated trials on different sub-dataset. These problems seriously restricted the wide application of the proteomic technique in molecular biology and other related directions. Thus, it is crucial to enable the generalizable and reproducible discovery of the proteomic signature with the subsequent indication of phenotype association. However, no such tool has been developed and available yet. Herein, an online tool, POSREG, was therefore constructed to identify the optimal signature for a set of proteomic data. It works by (i) identifying the proteomic signature of good reproducibility and aggregating them to ensemble feature ranking by ensemble learning, (ii) assessing the generalizability of ensemble feature ranking to acquire the optimal signature and (iii) indicating the phenotype association of discovered signature. POSREG is unique in its capacity of discovering the proteomic signature by simultaneously optimizing its reproducibility and generalizability. It is now accessible free of charge without any registration or login requirement at https://idrblab.org/posreg/. Feng Cheng Li, Ying Zhang 0061, Jiayi Yin, Yunqing Qiu, Jianqing Gao, Feng Zhu 0004 |
Briefings Bioinform. | 7 |
| 2022 | LargeMetabo: an out-of-the-box tool for processing and analyzing large-scale metabolomic dataabstractLarge-scale metabolomics is a powerful technique that has attracted widespread attention in biomedical studies focused on identifying biomarkers and interpreting the mechanisms of complex diseases. Despite a rapid increase in the number of large-scale metabolomic studies, the analysis of metabolomic data remains a key challenge. Specifically, diverse unwanted variations and batch effects in processing many samples have a substantial impact on identifying true biological markers, and it is a daunting challenge to annotate a plethora of peaks as metabolites in untargeted mass spectrometry-based metabolomics. Therefore, the development of an out-of-the-box tool is urgently needed to realize data integration and to accurately annotate metabolites with enhanced functions. In this study, the LargeMetabo package based on R code was developed for processing and analyzing large-scale metabolomic data. This package is unique because it is capable of (1) integrating multiple analytical experiments to effectively boost the power of statistical analysis; (2) selecting the appropriate biomarker identification method by intelligent assessment for large-scale metabolic data and (3) providing metabolite annotation and enrichment analysis based on an enhanced metabolite database. The LargeMetabo package can facilitate flexibility and reproducibility in large-scale metabolomics. The package is freely available from https://github.com/LargeMetabo/LargeMetabo. Qingxia Yang, Bo Li 0097, Jicheng Xie, Feng Zhu 0004 |
Briefings Bioinform. | 7 |
| 2022 | RNA-RNA interactions between SARS-CoV-2 and host benefit viral development and evolution during COVID-19 infectionabstractSome studies reported that genomic RNA of SARS-CoV-2 can absorb a few host miRNAs that regulate immune-related genes and then deprive their function. In this perspective, we conjecture that the absorption of the SARS-CoV-2 genome to host miRNAs is not a coincidence, which may be an indispensable approach leading to viral survival and development in host. In our study, we collected five datasets of miRNAs that were predicted to interact with the genome of SARS-CoV-2. The targets of these miRNAs in the five groups were consistently enriched immune-related pathways and virus-infectious diseases. Interestingly, the five datasets shared no one miRNA but their targets shared 168 genes. The signaling pathway enrichment of 168 shared targets implied an unbalanced immune response that the most of interleukin signaling pathways and none of the interferon signaling pathways were significantly different. Protein-protein interaction (PPI) network using the shared targets showed that PPI pairs, including IL6-IL6R, were related to the process of SARS-CoV-2 infection and pathogenesis. In addition, we found that SARS-CoV-2 absorption to host miRNA could benefit two popular mutant strains for more infectivity and pathogenicity. Conclusively, our results suggest that genomic RNA absorption to host miRNAs may be a vital approach by which SARS-CoV-2 disturbs the host immune system and infects host cells. Kuerbannisha Amahong, Feng Cheng Li, Jianqing Gao, Yunqing Qiu, Feng Zhu 0004 |
Briefings Bioinform. | 7 |
| 2022 | Biological activities of drug inactive ingredientsabstractIn a drug formulation (DFM), the major components by mass are not Active Pharmaceutical Ingredient (API) but rather Drug Inactive Ingredients (DIGs). DIGs can reach much higher concentrations than that achieved by API, which raises great concerns about their clinical toxicities. Therefore, the biological activities of DIG on physiologically relevant target are widely demanded by both clinical investigation and pharmaceutical industry. However, such activity data are not available in any existing pharmaceutical knowledge base, and their potentials in predicting the DIG-target interaction have not been evaluated yet. In this study, the comprehensive assessment and analysis on the biological activities of DIGs were therefore conducted. First, the largest number of DIGs and DFMs were systematically curated and confirmed based on all drugs approved by US Food and Drug Administration. Second, comprehensive activities for both DIGs and DFMs were provided for the first time to pharmaceutical community. Third, the biological targets of each DIG and formulation were fully referenced to available databases that described their pharmaceutical/biological characteristics. Finally, a variety of popular artificial intelligence techniques were used to assess the predictive potential of DIGs' activity data, which was the first evaluation on the possibility to predict DIG's activity. As the activities of DIGs are critical for current pharmaceutical studies, this work is expected to have significant implications for the future practice of drug discovery and precision medicine. Minjie Mou, Wei Zhang 0218, Xichen Lian, Shuiyang Shi, Mingkun Lu, Huaicheng Sun, Feng Cheng Li, Zhenyu Zeng, Zhaorong Li, Yunqing Qiu, Feng Zhu 0004, Jianqing Gao |
Briefings Bioinform. | 15 |
| 2022 | ncRNAInter: a novel strategy based on graph neural network to discover interactions between lncRNA and miRNAabstractIn recent years, many studies have illustrated the significant role that non-coding RNA (ncRNA) plays in biological activities, in which lncRNA, miRNA and especially their interactions have been proved to affect many biological processes. Some in silico methods have been proposed and applied to identify novel lncRNA-miRNA interactions (LMIs), but there are still imperfections in their RNA representation and information extraction approaches, which imply there is still room for further improving their performances. Meanwhile, only a few of them are accessible at present, which limits their practical applications. The construction of a new tool for LMI prediction is thus imperative for the better understanding of their relevant biological mechanisms. This study proposed a novel method, ncRNAInter, for LMI prediction. A comprehensive strategy for RNA representation and an optimized deep learning algorithm of graph neural network were utilized in this study. ncRNAInter was robust and showed better performance of 26.7% higher Matthews correlation coefficient than existing reputable methods for human LMI prediction. In addition, ncRNAInter proved its universal applicability in dealing with LMIs from various species and successfully identified novel LMIs associated with various diseases, which further verified its effectiveness and usability. All source code and datasets are freely available at https://github.com/idrblab/ncRNAInter. Xiuna Sun, Minjie Mou, Zhaorong Li, Honglin Li 0003, Feng Zhu 0004 |
Briefings Bioinform. | 9 |
| 2022 | A similarity-based deep learning approach for determining the frequencies of drug side effectsabstractThe side effects of drugs present growing concern attention in the healthcare system. Accurately identifying the side effects of drugs is very important for drug development and risk assessment. Some computational models have been developed to predict the potential side effects of drugs and provided satisfactory performance. However, most existing methods can only predict whether side effects will occur and cannot determine the frequency of side effects. Although a few existing methods can predict the frequency of drug side effects, they strongly depend on the known drug-side effect relationships. Therefore, they cannot be applied to new drugs without known side effect frequency information. In this paper, we develop a novel similarity-based deep learning method, named SDPred, for determining the frequencies of drug side effects. Compared with the existing state-of-the-art models, SDPred integrates rich features and can be applied to predict the side effect frequencies of new drugs without any known drug-side effect association or frequency information. To our knowledge, this is the first work that can predict the side effect frequencies of new drugs in the population. The comparison results indicate that SDPred is much superior to all previously reported models. In addition, some case studies also demonstrate the effectiveness of our proposed method in practical applications. The SDPred software and data are freely available at https://github.com/zhc940702/SDPred, https://zenodo.org/record/5112573 and https://hub.docker.com/r/zhc940702/sdpred. Shaokai Wang, Kai Zheng 0020, Qichang Zhao, Feng Zhu 0004, Jianxin Wang 0001 |
Briefings Bioinform. | 5 |
| 2021 | Pharmacometabonomics: data processing and statistical analysisabstractIndividual variations in drug efficacy, side effects and adverse drug reactions are still challenging that cannot be ignored in drug research and development. The aim of pharmacometabonomics is to better understand the pharmacokinetic properties of drugs and monitor the drug effects on specific metabolic pathways. Here, we systematically reviewed the recent technological advances in pharmacometabonomics for better understanding the pathophysiological mechanisms of diseases as well as the metabolic effects of drugs on bodies. First, the advantages and disadvantages of all mainstream analytical techniques were compared. Second, many data processing strategies including filtering, missing value imputation, quality control-based correction, transformation, normalization together with the methods implemented in each step were discussed. Third, various feature selection and feature extraction algorithms commonly applied in pharmacometabonomics were described. Finally, the databases that facilitate current pharmacometabonomics were collected and discussed. All in all, this review provided guidance for researchers engaged in pharmacometabonomics and metabolomics, and it would promote the wide application of metabolomics in drug research and personalized medicine. Jianbo Fu, Ying Zhang 0061, Xichen Lian, Feng Zhu 0004 |
Briefings Bioinform. | 6 |
| 2021 | MetaFS: Performance assessment of biomarker discovery in metaproteomicsabstractMetaproteomics suffers from the issues of dimensionality and sparsity. Data reduction methods can maximally identify the relevant subset of significant differential features and reduce data redundancy. Feature selection (FS) methods were applied to obtain the significant differential subset. So far, a variety of feature selection methods have been developed for metaproteomic study. However, due to FS's performance depended heavily on the data characteristics of a given research, the well-suitable feature selection method must be carefully selected to obtain the reproducible differential proteins. Moreover, it is critical to evaluate the performance of each FS method according to comprehensive criteria, because the single criterion is not sufficient to reflect the overall performance of the FS method. Therefore, we developed an online tool named MetaFS, which provided 13 types of FS methods and conducted the comprehensive evaluation on the complex FS methods using four widely accepted and independent criteria. Furthermore, the function and reliability of MetaFS were systematically tested and validated via two case studies. In sum, MetaFS could be a distinguished tool for discovering the overall well-performed FS method for selecting the potential biomarkers in microbiome studies. The online tool is freely available at https://idrblab.org/metafs/. Minjie Mou, Yongchao Luo, Feng Zhu 0004 |
Briefings Bioinform. | 5 |
| 2021 | The miRNA: a small but powerful RNA for COVID-19abstractCoronavirus disease 2019 (COVID-19) caused by severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) is a severe and rapidly evolving epidemic. Now, although a few drugs and vaccines have been proved for its treatment and prevention, little systematic comments are made to explain its susceptibility to humans. A few scattered studies used bioinformatics methods to explore the role of microRNA (miRNA) in COVID-19 infection. Combining these timely reports and previous studies about virus and miRNA, we comb through the available clues and seemingly make the perspective reasonable that the COVID-19 cleverly exploits the interplay between the small miRNA and other biomolecules to avoid being effectively recognized and attacked from host immune protection as well to deactivate functional genes that are crucial for immune system. In detail, SARS-CoV-2 can be regarded as a sponge to adsorb host immune-related miRNA, which forces host fall into dysfunction status of immune system. Besides, SARS-CoV-2 encodes its own miRNAs, which can enter host cell and are not perceived by the host's immune system, subsequently targeting host function genes to cause illnesses. Therefore, this article presents a reasonable viewpoint that the miRNA-based interplays between the host and SARS-CoV-2 may be the primary cause that SARS-CoV-2 accesses and attacks the host cells. Kuerbannisha Amahong, Xiuna Sun, Xichen Lian, Huaicheng Sun, Yan Lou, Feng Zhu 0004, Yunqing Qiu |
Briefings Bioinform. | 8 |
| 2021 | The mechanistic, diagnostic and therapeutic novel nucleic acids for hepatocellular carcinoma emerging in past score yearsabstractDespite The Central Dogma states the destiny of gene as 'DNA makes RNA and RNA makes protein', the nucleic acids not only store and transmit genetic information but also, surprisingly, join in intracellular vital movement as a regulator of gene expression. Bioinformatics has contributed to knowledge for a series of emerging novel nucleic acids molecules. For typical cases, microRNA (miRNA), long noncoding RNA (lncRNA) and circular RNA (circRNA) exert crucial role in regulating vital biological processes, especially in malignant diseases. Due to extraordinarily heterogeneity among all malignancies, hepatocellular carcinoma (HCC) has emerged enormous limitation in diagnosis and therapy. Mechanistic, diagnostic and therapeutic nucleic acids for HCC emerging in past score years have been systematically reviewed. Particularly, we have organized recent advances on nucleic acids of HCC into three facets: (i) summarizing diverse nucleic acids and their modification (miRNA, lncRNA, circRNA, circulating tumor DNA and DNA methylation) acting as potential biomarkers in HCC diagnosis; (ii) concluding different patterns of three key noncoding RNAs (miRNA, lncRNA and circRNA) in gene regulation and (iii) outlining the progress of these novel nucleic acids for HCC diagnosis and therapy in clinical trials, and discuss their possibility for clinical applications. All in all, this review takes a detailed look at the advances of novel nucleic acids from potential of biomarkers and elaboration of mechanism to early clinical application in past 20 years. Zhengwen Wang, Qitao Xiao, Ying Zhang 0061, Yan Lou, Yunqing Qiu, Feng Zhu 0004 |
Briefings Bioinform. | 9 |
| 2020 | Genome-wide identification and analysis of the eQTL lncRNAs in multiple sclerosis based on RNA-seq dataabstractThe pathogenesis of multiple sclerosis (MS) is significantly regulated by long noncoding RNAs (lncRNAs), the expression of which is substantially influenced by a number of MS-associated risk single nucleotide polymorphisms (SNPs). It is thus hypothesized that the dysregulation of lncRNA induced by genomic variants may be one of the key molecular mechanisms for the pathology of MS. However, due to the lack of sufficient data on lncRNA expression and SNP genotypes of the same MS patients, such molecular mechanisms underlying the pathology of MS remain elusive. In this study, a bioinformatics strategy was applied to obtain lncRNA expression and SNP genotype data simultaneously from 142 samples (51 MS patients and 91 controls) based on RNA-seq data, and an expression quantitative trait loci (eQTL) analysis was conducted. In total, 2383 differentially expressed lncRNAs were identified as specifically expressing in brain-related tissues, and 517 of them were affected by SNPs. Then, the functional characterization, secondary structure changes and tissue and disease specificity of the cis-eQTL SNPs and lncRNA were assessed. The cis-eQTL SNPs were substantially and specifically enriched in neurological disease and intergenic region, and the secondary structure was altered in 17.6% of all lncRNAs in MS. Finally, the weighted gene coexpression network and gene set enrichment analyses were used to investigate how the influence of SNPs on lncRNAs contributed to the pathogenesis of MS. As a result, the regulation of lncRNAs by SNPs was found to mainly influence the antigen processing/presentation and mitogen-activated protein kinases (MAPK) signaling pathway in MS. These results revealed the effectiveness of the strategy proposed in this study and give insight into the mechanism (SNP-mediated modulation of lncRNAs) underlying the pathology of MS. Zhijie Han 0002, Weiwei Xue, Yan Lou, Yunqing Qiu, Feng Zhu 0004 |
Briefings Bioinform. | 6 |
| 2020 | Convolutional neural network-based annotation of bacterial type IV secretion system effectors with enhanced accuracy and reduced false discoveryabstractThe type IV bacterial secretion system (SS) is reported to be one of the most ubiquitous SSs in nature and can induce serious conditions by secreting type IV SS effectors (T4SEs) into the host cells. Recent studies mainly focus on annotating new T4SE from the huge amount of sequencing data, and various computational tools are therefore developed to accelerate T4SE annotation. However, these tools are reported as heavily dependent on the selected methods and their annotation performance need to be further enhanced. Herein, a convolution neural network (CNN) technique was used to annotate T4SEs by integrating multiple protein encoding strategies. First, the annotation accuracies of nine encoding strategies integrated with CNN were assessed and compared with that of the popular T4SE annotation tools based on independent benchmark. Second, false discovery rates of various models were systematically evaluated by (1) scanning the genome of Legionella pneumophila subsp. ATCC 33152 and (2) predicting the real-world non-T4SEs validated using published experiments. Based on the above analyses, the encoding strategies, (a) position-specific scoring matrix (PSSM), (b) protein secondary structure & solvent accessibility (PSSSA) and (c) one-hot encoding scheme (Onehot), were identified as well-performing when integrated with CNN. Finally, a novel strategy that collectively considers the three well-performing models (CNN-PSSM, CNN-PSSSA and CNN-Onehot) was proposed, and a new tool (CNN-T4SE, https://idrblab.org/cnnt4se/) was constructed to facilitate T4SE annotation. All in all, this study conducted a comprehensive analysis on the performance of a collection of encoding strategies when integrated with CNN, which could facilitate the suppression of T4SS in infection and limit the spread of antimicrobial resistance. Jiajun Hong, Yongchao Luo, Minjie Mou, Jianbo Fu, Yang Zhang 0125, Weiwei Xue, Yan Lou, Feng Zhu 0004 |
Briefings Bioinform. | 10 |
| 2020 | Protein functional annotation of simultaneously improved stability, accuracy and false discovery rate achieved by a sequence-based deep learningabstractFunctional annotation of protein sequence with high accuracy has become one of the most important issues in modern biomedical studies, and computational approaches of significantly accelerated analysis process and enhanced accuracy are greatly desired. Although a variety of methods have been developed to elevate protein annotation accuracy, their ability in controlling false annotation rates remains either limited or not systematically evaluated. In this study, a protein encoding strategy, together with a deep learning algorithm, was proposed to control the false discovery rate in protein function annotation, and its performances were systematically compared with that of the traditional similarity-based and de novo approaches. Based on a comprehensive assessment from multiple perspectives, the proposed strategy and algorithm were found to perform better in both prediction stability and annotation accuracy compared with other de novo methods. Moreover, an in-depth assessment revealed that it possessed an improved capacity of controlling the false discovery rate compared with traditional methods. All in all, this study not only provided a comprehensive analysis on the performances of the newly proposed strategy but also provided a tool for the researcher in the fields of protein function annotation. Jiajun Hong, Yongchao Luo, Yang Zhang 0125, Junbiao Ying, Weiwei Xue, Feng Zhu 0004 |
Briefings Bioinform. | 8 |
| 2020 | Clinical trials, progression-speed differentiating features and swiftness rule of the innovative targets of first-in-class drugsabstractDrugs produce their therapeutic effects by modulating specific targets, and there are 89 innovative targets of first-in-class drugs approved in 2004-17, each with information about drug clinical trial dated back to 1984. Analysis of the clinical trial timelines of these targets may reveal the trial-speed differentiating features for facilitating target assessment. Here we present a comprehensive analysis of all these 89 targets, following the earlier studies for prospective prediction of clinical success of the targets of clinical trial drugs. Our analysis confirmed the literature-reported common druggability characteristics for clinical success of these innovative targets, exposed trial-speed differentiating features associated to the on-target and off-target collateral effects in humans and further revealed a simple rule for identifying the speedy human targets through clinical trials (from the earliest phase I to the 1st drug approval within 8 years). This simple rule correctly identified 75.0% of the 28 speedy human targets and only unexpectedly misclassified 13.2% of 53 non-speedy human targets. Certain extraordinary circumstances were also discovered to likely contribute to the misclassification of some human targets by this simple rule. Investigation and knowledge of trial-speed differentiating features enable prioritized drug discovery and development. Yinghong Li 0002, Xiao Xu Li, Jiajun Hong, Jianbo Fu, Chun Yan Yu, Feng Cheng Li, Jie Hu 0020, Weiwei Xue, Yuzong Chen 0002, Feng Zhu 0004 |
Briefings Bioinform. | 13 |
| 2020 | Comprehensive assessment of nine docking programs on type II kinase inhibitors: prediction accuracy of sampling power, scoring power and screening powerabstractProtein kinases have been regarded as important therapeutic targets for many diseases. Currently, a total of 41 kinase inhibitors have been approved by the Food and Drug Administration, along with a large number of kinase inhibitors being evaluated in clinical and preclinical trials. Among all, allosteric inhibitors, such as type II kinase inhibitors, have attracted extensive attention owing to their potential high selectivity. Nowadays, molecular docking has become a powerful tool to search for novel kinase inhibitors. However, as for type II kinase inhibitors, their allosteric characteristics may exert a deep influence on docking accuracy. In this study, a comprehensive assessment was conducted to evaluate the effectiveness of nine docking algorithms towards type II kinase inhibitors. The calculation results showed that most tested docking programs, especially Glide with XP scoring, LeDock and Surflex-Dock, succeeded in the accurate identification of near-native binding poses, with the success rates ranging from 0.80 to 0.90, and the scoring functions in GOLD and LeDock outperformed the others in the prediction of relative binding affinities. In terms of the P-values, areas under the curve and enrichment factors, Glide with XP scoring, Surflex-Dock, GOLD with Astex Statistical Potential scoring and LeDock had better screening power to discriminate between active compounds and decoys. However, the screening power is sensitive to different initial conformations of the same target. It is expected that our study can provide some guidance for docking-based virtual screening to discover novel type II kinase inhibitors, as well as other allosteric inhibitors. Chao Shen 0008, Zhe Wang 0041, Youyong Li, Tailong Lei, Ercheng Wang, Lei Xu 0035, Feng Zhu 0004, Dan Li 0013, Tingjun Hou |
Briefings Bioinform. | 8 |
| 2020 | ANPELA: analysis and performance assessment of the label-free quantification workflow for metaproteomic studiesabstractLabel-free quantification (LFQ) with a specific and sequentially integrated workflow of acquisition technique, quantification tool and processing method has emerged as the popular technique employed in metaproteomic research to provide a comprehensive landscape of the adaptive response of microbes to external stimuli and their interactions with other organisms or host cells. The performance of a specific LFQ workflow is highly dependent on the studied data. Hence, it is essential to discover the most appropriate one for a specific data set. However, it is challenging to perform such discovery due to the large number of possible workflows and the multifaceted nature of the evaluation criteria. Herein, a web server ANPELA (https://idrblab.org/anpela/) was developed and validated as the first tool enabling performance assessment of whole LFQ workflow (collective assessment by five well-established criteria with distinct underlying theories), and it enabled the identification of the optimal LFQ workflow(s) by a comprehensive performance ranking. ANPELA not only automatically detects the diverse formats of data generated by all quantification tools but also provides the most complete set of processing methods among the available web servers and stand-alone tools. Systematic validation using metaproteomic benchmarks revealed ANPELA's capabilities in 1 discovering well-performing workflow(s), (2) enabling assessment from multiple perspectives and (3) validating LFQ accuracy using spiked proteins. ANPELA has a unique ability to evaluate the performance of whole LFQ workflow and enables the discovery of the optimal LFQs by the comprehensive performance ranking of all 560 workflows. Therefore, it has great potential for applications in metaproteomic and other studies requiring LFQ techniques, as many features are shared among proteomic studies. Jianbo Fu, Bo Li 0093, Yinghong Li 0002, Qingxia Yang, Xuejiao Cui, Jiajun Hong, Yuzong Chen 0002, Weiwei Xue, Feng Zhu 0004 |
Briefings Bioinform. | 12 |
| 2020 | A critical assessment of the feature selection methods used for biomarker discovery in current metaproteomics studiesabstractMicrobial community (MC) has great impact on mediating complex disease indications, biogeochemical cycling and agricultural productivities, which makes metaproteomics powerful technique for quantifying diverse and dynamic composition of proteins or peptides. The key role of biostatistical strategies in MC study is reported to be underestimated, especially the appropriate application of feature selection method (FSM) is largely ignored. Although extensive efforts have been devoted to assessing the performance of FSMs, previous studies focused only on their classification accuracy without considering their ability to correctly and comprehensively identify the spiked proteins. In this study, the performances of 14 FSMs were comprehensively assessed based on two key criteria (both sample classification and spiked protein discovery) using a variety of metaproteomics benchmarks. First, the classification accuracies of those 14 FSMs were evaluated. Then, their abilities in identifying the proteins of different spiked concentrations were assessed. Finally, seven FSMs (FC, LMEB, OPLS-DA, PLS-DA, SAM, SVM-RFE and T-Test) were identified as performing consistently superior or good under both criteria with the PLS-DA performing consistently superior. In summary, this study served as comprehensive analysis on the performances of current FSMs and could provide a valuable guideline for researchers in metaproteomics. Jianbo Fu, Yongchao Luo, Ying Zhang 0061, Bo Li 0093, Qingxia Yang, Weiwei Xue, Yan Lou, Yunqing Qiu, Feng Zhu 0004 |
Briefings Bioinform. | 12 |
| 2020 | A novel bioinformatics approach to identify the consistently well-performing normalization strategy for current metabolomic studiesabstractUnwanted experimental/biological variation and technical error are frequently encountered in current metabolomics, which requires the employment of normalization methods for removing undesired data fluctuations. To ensure the 'thorough' removal of unwanted variations, the collective consideration of multiple criteria ('intragroup variation', 'marker stability' and 'classification capability') was essential. However, due to the limited number of available normalization methods, it is extremely challenging to discover the appropriate one that can meet all these criteria. Herein, a novel approach was proposed to discover the normalization strategies that are consistently well performing (CWP) under all criteria. Based on various benchmarks, all normalization methods popular in current metabolomics were 'first' discovered to be non-CWP. 'Then', 21 new strategies that combined the 'sample'-based method with the 'metabolite'-based one were found to be CWP. 'Finally', a variety of currently available methods (such as cubic splines, range scaling, level scaling, EigenMS, cyclic loess and mean) were identified to be CWP when combining with other normalization. In conclusion, this study not only discovered several strategies that performed consistently well under all criteria, but also proposed a novel approach that could ensure the identification of CWP strategies for future biological problems. Qingxia Yang, Jiajun Hong, Weiwei Xue, Feng Zhu 0004 |
Briefings Bioinform. | 7 |
| 2020 | Consistent gene signature of schizophrenia identified by a novel feature selection strategy from comprehensive sets of transcriptomic dataabstractThe etiology of schizophrenia (SCZ) is regarded as one of the most fundamental puzzles in current medical research, and its diagnosis is limited by the lack of objective molecular criteria. Although plenty of studies were conducted, SCZ gene signatures identified by these independent studies are found highly inconsistent. As one of the most important factors contributing to this inconsistency, the feature selection methods used currently do not fully consider the reproducibility among the signatures discovered from different datasets. Therefore, it is crucial to develop new bioinformatics tools of novel strategy for ensuring a stable discovery of gene signature for SCZ. In this study, a novel feature selection strategy (1) integrating repeated random sampling with consensus scoring and (2) evaluating the consistency of gene rank among different datasets was constructed. By systematically assessing the identified SCZ signature comprising 135 differentially expressed genes, this newly constructed strategy demonstrated significantly enhanced stability and better differentiating ability compared with the feature selection methods popular in current SCZ research. Based on a first-ever assessment on methods' reproducibility cross-validated by independent datasets from three representative studies, the new strategy stood out among the popular methods by showing superior stability and differentiating ability. Finally, 2 novel and 17 previously reported transcription factors were identified and showed great potential in revealing the etiology of SCZ. In sum, the SCZ signature identified in this study would provide valuable clues for discovering diagnostic molecules and potential targets for SCZ. Qingxia Yang, Bo Li 0093, Xuejiao Cui, Jie Hu 0020, Yuzong Chen 0002, Weiwei Xue, Yan Lou, Yunqing Qiu, Feng Zhu 0004 |
Briefings Bioinform. | 12 |
| 2019 | farPPI: a webserver for accurate prediction of protein-ligand binding structures for small-molecule PPI inhibitors by MM/PB(GB)SA methodsabstractSUMMARY: Protein-protein interactions (PPIs) have been regarded as an attractive emerging class of therapeutic targets for the development of new treatments. Computational approaches, especially molecular docking, have been extensively employed to predict the binding structures of PPI-inhibitors or discover novel small molecule PPI inhibitors. However, due to the relatively 'undruggable' features of PPI interfaces, accurate predictions of the binding structures for ligands towards PPI targets are quite challenging for most docking algorithms. Here, we constructed a non-redundant pose ranking benchmark dataset for small-molecule PPI inhibitors, which contains 900 binding poses for 184 protein-ligand complexes. Then, we evaluated the performance of MM/PB(GB)SA approaches to identify the correct binding poses for PPI inhibitors, including two Prime MM/GBSA procedures from the Schrödinger suite and seven different MM/PB(GB)SA procedures from the Amber package. Our results showed that MM/PBSA outperformed the Glide SP scoring function (success rate of 58.6%) and MM/GBSA in most cases, especially the PB3 procedure which could achieve an overall success rate of ∼74%. Moreover, the GB6 procedure (success rate of 68.9%) performed much better than the other MM/GBSA procedures, highlighting the excellent potential of the GBNSR6 implicit solvation model for pose ranking. Finally, we developed the webserver of Fast Amber Rescoring for PPI Inhibitors (farPPI), which offers a freely available service to rescore the docking poses for PPI inhibitors by using the MM/PB(GB)SA methods. AVAILABILITY AND IMPLEMENTATION: farPPI web server is freely available at http://cadd.zju.edu.cn/farppi/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Zhe Wang 0041, Xuwen Wang, Youyong Li, Tailong Lei, Ercheng Wang, Dan Li 0013, Yu Kang 0002, Feng Zhu 0004, Tingjun Hou |
Bioinform. | 8 |
| 2017 | A protein network descriptor server and its use in studying protein, disease, metabolic and drug targeted networksabstractThe genetic, proteomic, disease and pharmacological studies have generated rich data in protein interaction, disease regulation and drug activities useful for systems-level study of the biological, disease and drug therapeutic processes. These studies are facilitated by the established and the emerging computational methods. More recently, the network descriptors developed in other disciplines have become more increasingly used for studying the protein-protein, gene regulation, metabolic, disease networks. There is an inadequate coverage of these useful network features in the public web servers. We therefore introduced upto 313 literature-reported network descriptors in PROFEAT web server, for describing the topological, connectivity and complexity characteristics of undirected unweighted (uniform binding constants and molecular levels), undirected edge-weighted (varying binding constants), undirected node-weighted (varying molecular levels), undirected edge-node-weighted (varying binding constants and molecular levels) and directed unweighted (oriented process) networks. The usefulness of the PROFEAT computed network descriptors is illustrated by their literature-reported applications in studying the protein-protein, gene regulatory, gene co-expression, protein-drug and metabolic networks. PROFEAT is accessible free of charge at http://bidd2.nus.edu.sg/cgi-bin/profeat2016/main.cgi. Peng Zhang 0033, Xian Zeng, Chu Qin, Shangying Chen, Feng Zhu 0004, Zerong Li, Weiping Chen, Yuzong Chen 0002 |
Briefings Bioinform. | 6 |