EDBT 2026 Demo / reviewers in the wild / expert
Yan Zhu 0006
dblp:82/3167-6
· DBLP profile ↗
9ranked-venue papers
3as first author
7since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 9 · 3 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Drug-Target Binding Affinity Prediction in a Continuous Latent Space Using Variational AutoencodersabstractAccurate prediction of Drug-Target binding Affinity (DTA) is a daunting yet pivotal task in the sphere of drug discovery. Over the years, a plethora of deep learning-based DTA models have emerged, rendering promising results in predicting the binding affinities between drugs and their target proteins. However, in contrast to the conventional approach of modeling binding affinity in vector spaces, we propose a more nuanced modeling process in a continuous space to account for the diversity of input samples. Initially, the drug is encoded using the Simplified Molecular Input Line Entry System (SMILES), while the target sequences are characterized via a pretrained language model. Subsequently, highly correlative information is extracted utilizing residual gated convolutional neural networks. In a departure from existing deep learning-based models, our model learns the hidden representations of the drugs and targets jointly. Instead of employing two vectors, our hidden representations consist of two Gaussian distributions. To validate the effectiveness of our proposal, we conducted evaluations on commonly utilized benchmark datasets. The experimental outcomes corroborated that our method surpasses the state-of-the-art vectorial representation methods in terms of performance. This approach, therefore, offers potential enhancements in the precision of DTA predictions, potentially contributing to more efficient drug discovery processes. Lingling Zhao, Yan Zhu 0006, Naifeng Wen, Chunyu Wang 0002, Junjie Wang 0005, Yongfeng Yuan |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2023 | TIMER is a Siamese neural network-based framework for identifying both general and species-specific bacterial promotersabstractBACKGROUND: Promoters are DNA regions that initiate the transcription of specific genes near the transcription start sites. In bacteria, promoters are recognized by RNA polymerases and associated sigma factors. Effective promoter recognition is essential for synthesizing the gene-encoded products by bacteria to grow and adapt to different environmental conditions. A variety of machine learning-based predictors for bacterial promoters have been developed; however, most of them were designed specifically for a particular species. To date, only a few predictors are available for identifying general bacterial promoters with limited predictive performance. RESULTS: In this study, we developed TIMER, a Siamese neural network-based approach for identifying both general and species-specific bacterial promoters. Specifically, TIMER uses DNA sequences as the input and employs three Siamese neural networks with the attention layers to train and optimize the models for a total of 13 species-specific and general bacterial promoters. Extensive 10-fold cross-validation and independent tests demonstrated that TIMER achieves a competitive performance and outperforms several existing methods on both general and species-specific promoter prediction. As an implementation of the proposed method, the web server of TIMER is publicly accessible at http://web.unimelb-bioinfortools.cloud.edu.au/TIMER/. Yan Zhu 0006, Fuyi Li, Xiaoyu Wang 0016, Lachlan James M. Coin, Geoffrey I. Webb, Jiangning Song, Cangzhi Jia |
Briefings Bioinform. | 1 |
| 2023 | DataDTA: a multi-feature and dual-interaction aggregation framework for drug-target binding affinity predictionabstractMOTIVATION: Accurate prediction of drug-target binding affinity (DTA) is crucial for drug discovery. The increase in the publication of large-scale DTA datasets enables the development of various computational methods for DTA prediction. Numerous deep learning-based methods have been proposed to predict affinities, some of which only utilize original sequence information or complex structures, but the effective combination of various information and protein-binding pockets have not been fully mined. Therefore, a new method that integrates available key information is urgently needed to predict DTA and accelerate the drug discovery process. RESULTS: In this study, we propose a novel deep learning-based predictor termed DataDTA to estimate the affinities of drug-target pairs. DataDTA utilizes descriptors of predicted pockets and sequences of proteins, as well as low-dimensional molecular features and SMILES strings of compounds as inputs. Specifically, the pockets were predicted from the three-dimensional structure of proteins and their descriptors were extracted as the partial input features for DTA prediction. The molecular representation of compounds based on algebraic graph features was collected to supplement the input information of targets. Furthermore, to ensure effective learning of multiscale interaction features, a dual-interaction aggregation neural network strategy was developed. DataDTA was compared with state-of-the-art methods on different datasets, and the results showed that DataDTA is a reliable prediction tool for affinities estimation. Specifically, the concordance index (CI) of DataDTA is 0.806 and the Pearson correlation coefficient (R) value is 0.814 on the test dataset, which is higher than other methods. AVAILABILITY AND IMPLEMENTATION: The codes and datasets of DataDTA are available at https://github.com/YanZhu06/DataDTA. Yan Zhu 0006, Lingling Zhao, Naifeng Wen, Junjie Wang 0005, Chunyu Wang 0002 |
Bioinform. | 1 |
| 2022 | Critical assessment of computational tools for prokaryotic and eukaryotic promoter predictionabstractPromoters are crucial regulatory DNA regions for gene transcriptional activation. Rapid advances in next-generation sequencing technologies have accelerated the accumulation of genome sequences, providing increased training data to inform computational approaches for both prokaryotic and eukaryotic promoter prediction. However, it remains a significant challenge to accurately identify species-specific promoter sequences using computational approaches. To advance computational support for promoter prediction, in this study, we curated 58 comprehensive, up-to-date, benchmark datasets for 7 different species (i.e. Escherichia coli, Bacillus subtilis, Homo sapiens, Mus musculus, Arabidopsis thaliana, Zea mays and Drosophila melanogaster) to assist the research community to assess the relative functionality of alternative approaches and support future research on both prokaryotic and eukaryotic promoters. We revisited 106 predictors published since 2000 for promoter identification (40 for prokaryotic promoter, 61 for eukaryotic promoter, and 5 for both). We systematically evaluated their training datasets, computational methodologies, calculated features, performance and software usability. On the basis of these benchmark datasets, we benchmarked 19 predictors with functioning webservers/local tools and assessed their prediction performance. We found that deep learning and traditional machine learning-based approaches generally outperformed scoring function-based approaches. Taken together, the curated benchmark dataset repository and the benchmarking analysis in this study serve to inform the design and implementation of computational approaches for promoter prediction and facilitate more rigorous comparison of new techniques in the future. Meng Zhang 0046, Cangzhi Jia, Fuyi Li, Chen Li 0021, Yan Zhu 0006, Tatsuya Akutsu, Geoffrey I. Webb, Quan Zou 0001, Lachlan James M. Coin, Jiangning Song |
Briefings Bioinform. | 5 |
| 2021 | SeqGO-CPA: Improving Compound-Protein Binding Affinity Prediction with Sequence Information and Gene Ontology KnowledgeabstractThe compound-protein binding affinity (CPA) pre-diction is vital for drug discovery and drug repurposing. Deep learning methods have been developed to model the complicated relationship between CPA and the sequences or structures of proteins and molecules. This study proposes a novel deep learning method, SeqGO-CPA, integrating protein function knowledge represented by Gene Ontology (GO) annotations in the CPA prediction. To capture the semantic information of GO annotations, a fine-tuned natural language processing model for biomedical domains is utilized to encode the set of GO terms. Meanwhile, based on the observation that CPA often occurs in sub-structures, our method uses the tokenization algorithm to learn sub-structure information of proteins and compounds from a large number of unlabeled sequences. Further, a deep neural network architecture involving the jointly-feature representation and a highway block is developed to enhance the CPA prediction ability. The proposed model was evaluated on two public benchmark datasets in both standard cross-validation and blinding split settings. The experimental results demonstrate our method outperforms the deep learning-based baselines, meanwhile the incorporating of GO information further improves the prediction performance. Chunyu Wang 0002, Yan Zhu 0006, Naifeng Wen, Lingling Zhao, Junjie Wang 0005 |
BIBM | 2 |
| 2021 | Computational identification of eukaryotic promoters based on cascaded deep capsule neural networksabstractA promoter is a region in the DNA sequence that defines where the transcription of a gene by RNA polymerase initiates, which is typically located proximal to the transcription start site (TSS). How to correctly identify the gene TSS and the core promoter is essential for our understanding of the transcriptional regulation of genes. As a complement to conventional experimental methods, computational techniques with easy-to-use platforms as essential bioinformatics tools can be effectively applied to annotate the functions and physiological roles of promoters. In this work, we propose a deep learning-based method termed Depicter (Deep learning for predicting promoter), for identifying three specific types of promoters, i.e. promoter sequences with the TATA-box (TATA model), promoter sequences without the TATA-box (non-TATA model), and indistinguishable promoters (TATA and non-TATA model). Depicter is developed based on an up-to-date, species-specific dataset which includes Homo sapiens, Mus musculus, Drosophila melanogaster and Arabidopsis thaliana promoters. A convolutional neural network coupled with capsule layers is proposed to train and optimize the prediction model of Depicter. Extensive benchmarking and independent tests demonstrate that Depicter achieves an improved predictive performance compared with several state-of-the-art methods. The webserver of Depicter is implemented and freely accessible at https://depicter.erc.monash.edu/. Yan Zhu 0006, Fuyi Li, Dongxu Xiang, Tatsuya Akutsu, Jiangning Song, Cangzhi Jia |
Briefings Bioinform. | 1 |
| 2021 | Visual exploration of large metabolic modelsabstractMOTIVATION: Large metabolic models, including genome-scale metabolic models, are nowadays common in systems biology, biotechnology and pharmacology. They typically contain thousands of metabolites and reactions and therefore methods for their automatic visualization and interactive exploration can facilitate a better understanding of these models. RESULTS: We developed a novel method for the visual exploration of large metabolic models and implemented it in LMME (Large Metabolic Model Explorer), an add-on for the biological network analysis tool VANTED. The underlying idea of our method is to analyze a large model as follows. Starting from a decomposition into several subsystems, relationships between these subsystems are identified and an overview is computed and visualized. From this overview, detailed subviews may be constructed and visualized in order to explore subsystems and relationships in greater detail. Decompositions may either be predefined or computed, using built-in or self-implemented methods. Realized as add-on for VANTED, LMME is embedded in a domain-specific environment, allowing for further related analysis at any stage during the exploration. We describe the method, provide a use case and discuss the strengths and weaknesses of different decomposition methods. AVAILABILITY AND IMPLEMENTATION: The methods and algorithms presented here are implemented in LMME, an open-source add-on for VANTED. LMME can be downloaded from www.cls.uni-konstanz.de/software/lmme and VANTED can be downloaded from www.vanted.org. The source code of LMME is available from GitHub, at https://github.com/LSI-UniKonstanz/lmme. Michael Aichem, Tobias Czauderna, Yan Zhu 0006, Jinxin Zhao, Matthias Klapperstück, Karsten Klein 0001, Jian Li 0052, Falk Schreiber |
Bioinform. | 3 |
| 2020 | iLearn : an integrated platform and meta-learner for feature engineering, machine-learning analysis and modeling of DNA, RNA and protein sequence dataabstractWith the explosive growth of biological sequences generated in the post-genomic era, one of the most challenging problems in bioinformatics and computational biology is to computationally characterize sequences, structures and functions in an efficient, accurate and high-throughput manner. A number of online web servers and stand-alone tools have been developed to address this to date; however, all these tools have their limitations and drawbacks in terms of their effectiveness, user-friendliness and capacity. Here, we present iLearn, a comprehensive and versatile Python-based toolkit, integrating the functionality of feature extraction, clustering, normalization, selection, dimensionality reduction, predictor construction, best descriptor/model selection, ensemble learning and results visualization for DNA, RNA and protein sequences. iLearn was designed for users that only want to upload their data set and select the functions they need calculated from it, while all necessary procedures and optimal settings are completed automatically by the software. iLearn includes a variety of descriptors for DNA, RNA and proteins, and four feature output formats are supported so as to facilitate direct output usage or communication with other computational tools. In total, iLearn encompasses 16 different types of feature clustering, selection, normalization and dimensionality reduction algorithms, and five commonly used machine-learning algorithms, thereby greatly facilitating feature analysis and predictor construction. iLearn is made freely available via an online web server and a stand-alone toolkit. Zhen Chen 0009, Fuyi Li, Tatiana T. Marquez-Lago, André Leier, Jerico Revote, Yan Zhu 0006, David R. Powell, Tatsuya Akutsu, Geoffrey I. Webb, Kuo-Chen Chou, Alexander Ian Smith, Roger J. Daly, Jian Li 0052, Jiangning Song |
Briefings Bioinform. | 7 |
| 2020 | PRISMOID: a comprehensive 3D structure database for post-translational modifications and mutations with functional impactabstractPost-translational modifications (PTMs) play very important roles in various cell signaling pathways and biological process. Due to PTMs' extremely important roles, many major PTMs have been studied, while the functional and mechanical characterization of major PTMs is well documented in several databases. However, most currently available databases mainly focus on protein sequences, while the real 3D structures of PTMs have been largely ignored. Therefore, studies of PTMs 3D structural signatures have been severely limited by the deficiency of the data. Here, we develop PRISMOID, a novel publicly available and free 3D structure database for a wide range of PTMs. PRISMOID represents an up-to-date and interactive online knowledge base with specific focus on 3D structural contexts of PTMs sites and mutations that occur on PTMs and in the close proximity of PTM sites with functional impact. The first version of PRISMOID encompasses 17 145 non-redundant modification sites on 3919 related protein 3D structure entries pertaining to 37 different types of PTMs. Our entry web page is organized in a comprehensive manner, including detailed PTM annotation on the 3D structure and biological information in terms of mutations affecting PTMs, secondary structure features and per-residue solvent accessibility features of PTM sites, domain context, predicted natively disordered regions and sequence alignments. In addition, high-definition JavaScript packages are employed to enhance information visualization in PRISMOID. PRISMOID equips a variety of interactive and customizable search options and data browsing functions; these capabilities allow users to access data via keyword, ID and advanced options combination search in an efficient and user-friendly way. A download page is also provided to enable users to download the SQL file, computational structural features and PTM sites' data. We anticipate PRISMOID will swiftly become an invaluable online resource, assisting both biologists and bioinformaticians to conduct experiments and develop applications supporting discovery efforts in the sequence-structural-functional relationship of PTMs and providing important insight into mutations and PTM sites interaction mechanisms. The PRISMOID database is freely accessible at http://prismoid.erc.monash.edu/. The database and web interface are implemented in MySQL, JSP, JavaScript and HTML with all major browsers supported. Fuyi Li, Cunshuo Fan, Tatiana T. Marquez-Lago, André Leier, Jerico Revote, Cangzhi Jia, Yan Zhu 0006, Alexander Ian Smith, Geoffrey I. Webb, Quanzhong Liu, Leyi Wei, Jian Li 0052, Jiangning Song |
Briefings Bioinform. | 7 |