EDBT 2026 Demo / reviewers in the wild / expert
Yang Zhang 0125
dblp:06/6785-125
· DBLP profile ↗
11ranked-venue papers
3as first author
7since 2021 · last 2025
0000-0002-1317-120XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 10 · 3 first-author · 6 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Computational toxicology in drug discovery: applications of artificial intelligence in ADMET and toxicity predictionabstractToxicity risk assessment plays a crucial role in determining the clinical success and market potential of drug candidates. Traditional animal-based testing is costly, time-consuming, and ethically controversial, which has led to the rapid development of computational toxicology. This review surveys over 20 ADMET prediction platforms, categorizing them into rule/statistical-based methods, machine learning (ML) methods, and graph-based methods. We also summarize major toxicological databases into four types: chemical toxicity, environmental toxicology, alternative toxicology, and biological toxin databases, highlighting their roles in model training and validation. Furthermore, we review recent advancements in ML and artificial intelligence (AI) applied to toxicity prediction, covering acute toxicity, organ-specific toxicities, and carcinogenicity. The field is transitioning from single-endpoint predictions to multi-endpoint joint modeling, incorporating multimodal features. We also explore the application of generative modeling techniques and interpretability frameworks to improve the accuracy and credibility of predictions. Additionally, we discuss the use of network toxicology in evaluating the safety of traditional Chinese medicines (TCMs) and the potential of large language models (LLMs) in literature mining, knowledge integration, and molecular toxicity prediction. Finally, we address current challenges, including data quality, model interpretability, and causal inference, and propose future directions such as multi-omics integration, interpretable AI models, and domain-specific LLMs, aiming to provide more efficient and precise technical support for preclinical toxicity assessments in drug development. Jiangyan Zhang, Yuncong Zhang, Junyang Huang, Liping Ren, Chuantao Zhang, Quan Zou 0001, Yang Zhang 0125 |
Briefings Bioinform. | 8 |
| 2024 | Attention is all you need: utilizing attention in AI-enabled drug discoveryabstractRecently, attention mechanism and derived models have gained significant traction in drug development due to their outstanding performance and interpretability in handling complex data structures. This review offers an in-depth exploration of the principles underlying attention-based models and their advantages in drug discovery. We further elaborate on their applications in various aspects of drug development, from molecular screening and target binding to property prediction and molecule generation. Finally, we discuss the current challenges faced in the application of attention mechanisms and Artificial Intelligence technologies, including data quality, model interpretability and computational resource constraints, along with future directions for research. Given the accelerating pace of technological advancement, we believe that attention-based models will have an increasingly prominent role in future drug discovery. We anticipate that these models will usher in revolutionary breakthroughs in the pharmaceutical domain, significantly accelerating the pace of drug development. Yang Zhang 0125, Caiqi Liu, Mujiexin Liu, Hao Lin 0001, Cheng-Bing Huang, Lin Ning 0002 |
Briefings Bioinform. | 1 |
| 2024 | ACVPred: Enhanced prediction of anti-coronavirus peptides by transfer learning combined with data augmentation
Juanjuan Kang, Liping Ren, Hui Ding 0005, Yang Zhang 0125 |
Future Gener. Comput. Syst. | 7 |
| 2022 | iRice-MS: An integrated XGBoost model for detecting multitype post-translational modification sites in riceabstractPost-translational modification (PTM) refers to the covalent and enzymatic modification of proteins after protein biosynthesis, which orchestrates a variety of biological processes. Detecting PTM sites in proteome scale is one of the key steps to in-depth understanding their regulation mechanisms. In this study, we presented an integrated method based on eXtreme Gradient Boosting (XGBoost), called iRice-MS, to identify 2-hydroxyisobutyrylation, crotonylation, malonylation, ubiquitination, succinylation and acetylation in rice. For each PTM-specific model, we adopted eight feature encoding schemes, including sequence-based features, physicochemical property-based features and spatial mapping information-based features. The optimal feature set was identified from each encoding, and their respective models were established. Extensive experimental results show that iRice-MS always display excellent performance on 5-fold cross-validation and independent dataset test. In addition, our novel approach provides the superiority to other existing tools in terms of AUC value. Based on the proposed model, a web server named iRice-MS was established and is freely accessible at http://lin-group.cn/server/iRice-MS. Hao Lv 0007, Yang Zhang 0125, Jia-Shu Wang, Shi-Shi Yuan, Fu-Ying Dao, Zheng-Xing Guan, Hao Lin 0001, Ke-Jun Deng |
Briefings Bioinform. | 2 |
| 2022 | PSnoD: identifying potential snoRNA-disease associations based on bounded nuclear norm regularizationabstractMany studies have proved that small nucleolar RNAs (snoRNAs) play critical roles in the development of various human complex diseases. Discovering the associations between snoRNAs and diseases is an important step toward understanding the pathogenesis and characteristics of diseases. However, uncovering associations via traditional experimental approaches is costly and time-consuming. This study proposed a bounded nuclear norm regularization-based method, called PSnoD, to predict snoRNA-disease associations. Benchmark experiments showed that compared with the state-of-the-art methods, PSnoD achieved a superior performance in the 5-fold stratified shuffle split. PSnoD produced a robust performance with an area under receiver-operating characteristic of 0.90 and an area under precision-recall of 0.55, highlighting the effectiveness of our proposed method. In addition, the computational efficiency of PSnoD was also demonstrated by comparison with other matrix completion techniques. More importantly, the case study further elucidated the ability of PSnoD to screen potential snoRNA-disease associations. The code of PSnoD has been uploaded to https://github.com/linDing-groups/PSnoD. Based on PSnoD, we established a web server that is freely accessed via http://psnod.lin-group.cn/. Hao Lv 0007, Yang Zhang 0125, Hao Lin 0001, Lin Ning 0002 |
Briefings Bioinform. | 6 |
| 2021 | Application of artificial intelligence and machine learning for COVID-19 drug discovery and vaccine designabstractThe global pandemic of coronavirus disease 2019 (COVID-19), caused by severe acute respiratory syndrome coronavirus 2, has led to a dramatic loss of human life worldwide. Despite many efforts, the development of effective drugs and vaccines for this novel virus will take considerable time. Artificial intelligence (AI) and machine learning (ML) offer promising solutions that could accelerate the discovery and optimization of new antivirals. Motivated by this, in this paper, we present an extensive survey on the application of AI and ML for combating COVID-19 based on the rapidly emerging literature. Particularly, we point out the challenges and future directions associated with state-of-the-art solutions to effectively control the COVID-19 pandemic. We hope that this review provides researchers with new insights into the ways AI and ML fight and have fought the COVID-19 outbreak. Hao Lv 0007, Joshua William Berkenpas, Fu-Ying Dao, Hasan Zulfiqar, Hui Ding 0005, Yang Zhang 0125, Renzhi Cao |
Briefings Bioinform. | 7 |
| 2021 | Cellinker: a platform of ligand-receptor interactions for intercellular communication analysisabstractMOTIVATION: Ligand-receptor (L-R) interactions mediate cell adhesion, recognition and communication and play essential roles in physiological and pathological signaling. With the rapid development of single-cell RNA sequencing (scRNA-seq) technologies, systematically decoding the intercellular communication network involving L-R interactions has become a focus of research. Therefore, construction of a comprehensive, high-confidence and well-organized resource to retrieve L-R interactions in order to study the functional effects of cell-cell communications would be of great value. RESULTS: In this study, we developed Cellinker, a manually curated resource of literature-supported L-R interactions that play roles in cell-cell communication. We aimed to provide a useful platform for studies on cell-cell communication mediated by L-R interactions. The current version of Cellinker documents over 3,700 human and 3,200 mouse L-R protein-protein interactions (PPIs) and embeds a practical and convenient webserver with which researchers can decode intercellular communications based on scRNA-seq data. And over 400 endogenous small molecule (sMOL) related L-R interactions were collected as well. Moreover, to help with research on coronavirus (CoV) infection, Cellinker collects information on 16 L-R PPIs involved in CoV-human interactions (including 12 L-R PPIs involved in SARS-CoV-2 infection). In summary, Cellinker provides a user-friendly interface for querying, browsing and visualizing L-R interactions as well as a practical and convenient web tool for inferring intercellular communications based on scRNA-seq data. We believe this platform could promote intercellular communication research and accelerate the development of related algorithms for scRNA-seq studies. AVAILABILITY: Cellinker is available at http://www.rna-society.org/cellinker/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yang Zhang 0125, Jing Wang 0004, Bohao Zou, Linhui Yao, Kechen Chen, Lin Ning 0002, Bingyi Wu, Dong Wang 0011 |
Bioinform. | 1 |
| 2020 | Convolutional neural network-based annotation of bacterial type IV secretion system effectors with enhanced accuracy and reduced false discoveryabstractThe type IV bacterial secretion system (SS) is reported to be one of the most ubiquitous SSs in nature and can induce serious conditions by secreting type IV SS effectors (T4SEs) into the host cells. Recent studies mainly focus on annotating new T4SE from the huge amount of sequencing data, and various computational tools are therefore developed to accelerate T4SE annotation. However, these tools are reported as heavily dependent on the selected methods and their annotation performance need to be further enhanced. Herein, a convolution neural network (CNN) technique was used to annotate T4SEs by integrating multiple protein encoding strategies. First, the annotation accuracies of nine encoding strategies integrated with CNN were assessed and compared with that of the popular T4SE annotation tools based on independent benchmark. Second, false discovery rates of various models were systematically evaluated by (1) scanning the genome of Legionella pneumophila subsp. ATCC 33152 and (2) predicting the real-world non-T4SEs validated using published experiments. Based on the above analyses, the encoding strategies, (a) position-specific scoring matrix (PSSM), (b) protein secondary structure & solvent accessibility (PSSSA) and (c) one-hot encoding scheme (Onehot), were identified as well-performing when integrated with CNN. Finally, a novel strategy that collectively considers the three well-performing models (CNN-PSSM, CNN-PSSSA and CNN-Onehot) was proposed, and a new tool (CNN-T4SE, https://idrblab.org/cnnt4se/) was constructed to facilitate T4SE annotation. All in all, this study conducted a comprehensive analysis on the performance of a collection of encoding strategies when integrated with CNN, which could facilitate the suppression of T4SS in infection and limit the spread of antimicrobial resistance. Jiajun Hong, Yongchao Luo, Minjie Mou, Jianbo Fu, Yang Zhang 0125, Weiwei Xue, Yan Lou, Feng Zhu 0004 |
Briefings Bioinform. | 5 |
| 2020 | Protein functional annotation of simultaneously improved stability, accuracy and false discovery rate achieved by a sequence-based deep learningabstractFunctional annotation of protein sequence with high accuracy has become one of the most important issues in modern biomedical studies, and computational approaches of significantly accelerated analysis process and enhanced accuracy are greatly desired. Although a variety of methods have been developed to elevate protein annotation accuracy, their ability in controlling false annotation rates remains either limited or not systematically evaluated. In this study, a protein encoding strategy, together with a deep learning algorithm, was proposed to control the false discovery rate in protein function annotation, and its performances were systematically compared with that of the traditional similarity-based and de novo approaches. Based on a comprehensive assessment from multiple perspectives, the proposed strategy and algorithm were found to perform better in both prediction stability and annotation accuracy compared with other de novo methods. Moreover, an in-depth assessment revealed that it possessed an improved capacity of controlling the false discovery rate compared with traditional methods. All in all, this study not only provided a comprehensive analysis on the performances of the newly proposed strategy but also provided a tool for the researcher in the fields of protein function annotation. Jiajun Hong, Yongchao Luo, Yang Zhang 0125, Junbiao Ying, Weiwei Xue, Feng Zhu 0004 |
Briefings Bioinform. | 3 |
| 2019 | RIscoper: a tool for RNA-RNA interaction extraction from the literatureabstractMOTIVATION: Numerous experimental and computational studies in the biomedical literature have provided considerable amounts of data on diverse RNA-RNA interactions (RRIs). However, few text mining systems for RRIs information extraction are available. RESULTS: RNA Interactome Scoper (RIscoper) represents the first tool for full-scale RNA interactome scanning and was developed for extracting RRIs from the literature based on the N-gram model. Notably, a reliable RRI corpus was integrated in RIscoper, and more than 13 300 manually curated sentences with RRI information were recruited. RIscoper allows users to upload full texts or abstracts, and provides an online search tool that is connected with PubMed (PMID and keyword input), and these capabilities are useful for biologists. RIscoper has a strong performance (90.4% precision and 93.9% recall), integrates natural language processing techniques and has a reliable RRI corpus. AVAILABILITY AND IMPLEMENTATION: The standalone software and web server of RIscoper are freely available at www.rna-society.org/riscoper/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yang Zhang 0125, Jinxurong Yang, Jiayi Yin, Yuncong Zhang, Zhixi Yun, Lin Ning 0002, Feng-Biao Guo, Yongshuai Jiang, Hao Lin 0001, Dong Wang 0011, Jian Huang 0004 |
Bioinform. | 1 |
| 2010 | Viewing cancer genes from co-evolving gene modulesabstractMOTIVATION: Studying the evolutionary conservation of cancer genes can improve our understanding of the genetic basis of human cancers. Functionally related proteins encoded by genes tend to interact with each other in a modular fashion, which may affect both the mode and tempo of their evolution. RESULTS: In the human PPI network, we searched for subnetworks within each of which all proteins have evolved at similar rates since the human and mouse split. Identified at a given co-evolving level, the subnetworks with non-randomly large sizes were defined as co-evolving modules. We showed that proteins within modules tend to be conserved, evolutionarily old and enriched with housekeeping genes, while proteins outside modules tend to be less-conserved, evolutionarily younger and enriched with genes expressed in specific tissues. Viewing cancer genes from co-evolving modules showed that the overall conservation of cancer genes should be mainly attributed to the cancer proteins enriched in the conserved modules. Functional analysis further suggested that cancer proteins within and outside modules might play different roles in carcinogenesis, providing a new hint for studying the mechanism of cancer. Jing Zhu 0004, Xiaopei Shen, Jing Wang 0004, Jinfeng Zou, Lin Zhang 0057, Da Yang 0003, Wencai Ma, Min Zhang 0009, Yang Zhang 0125, Zheng Guo 0002 |
Bioinform. | 12 |