EDBT 2026 Demo / reviewers in the wild / expert
Yanju Zhang
dblp:27/4900
· DBLP profile ↗
14ranked-venue papers
5as first author
8since 2021 · last 2025
0000-0002-8629-258XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 14 · 5 first-author · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | xLSTM-Stega: A Novel Framework for Secure Information Hiding in DNA Storage SystemsabstractDNA storage is well-suited for long-term archiving of infrequently accessed cold data, such as personal medical records. Embedding confidential data, particularly sensitive information, creates a critical application scenario where DNA steganography offers an essential solution. However, existing DNA steganographic methods often introduce statistical distortions, rendering them vulnerable to steganalysis, while also exhibiting poor robustness against sequencing errors and noise. To address these challenges, we propose xLSTM-Stega, a novel DNA steganographic framework that integrates Extended Long Short-Term Memory networks with natural pattern-preserving information hiding methods and Luby Transform error correction codes. A comprehensive evaluation across diverse bacterial genome datasets demonstrates the superior security of xLSTMStega with detection rates of$\mathbf{6 5. 1 5}-\mathbf{7 4. 2 5 \%}$-significantly outperforming existing methods. The framework preserves exceptional biological properties with GC content deviations below 0.22% and thermal stability deviations under$\mathbf{0. 0 6 \%}$, ensuring successful DNA synthesis, amplification, and sequencing. Under adversarial conditions, the framework achieves bit error rates of 0.240.33, demonstrating superior robustness. xLSTM-Stega provides a comprehensive solution for practical DNA steganography that simultaneously achieves security, biological plausibility, and robustness for securing sensitive information in DNA storage applications. Weikun Fan, Zhenyi Tu, Yanju Zhang |
BIBM | 3 |
| 2025 | DNA-Stega: Biologically Plausible DNA Steganography Based on Variational Autoencoder and Distribution CopiesabstractSteganography conceals the existence of sensitive information by embedding it into an innocuous carrier in a manner imperceptible to observers. As a novel physical carrier, DNA exhibits unique properties, making DNA steganography a promising method for data security. Despite recent advancements, most existing DNA steganography methods fail to guarantee data confidentiality when subjected to adversarial steganalysis. To address this vulnerability, we present DNA-Stega, a novel DNA steganography framework that integrates a Variational Autoencoder with a pre-trained DNABERT encoder and an LSTM decoder, along with the Discop algorithm for secret information hiding. Unlike previous approaches, our method leverages DNABERT to precisely model the complex patterns of natural DNA chains before pseudo-sequence generation, significantly enhancing both perceptual and statistical imperceptibility. Experimental evaluation on four DNA chains demonstrates DNAStega's superior performance in perceptual imperceptibility (with minimal compositional and thermodynamic deviations) and statistical imperceptibility (achieving the lowest KL divergence and Sliced-Wasserstein Distance). Furthermore, DNA-Stega exhibits remarkable resistance to state-of-the-art steganalysis techniques, which are capable of identifying pseudo-DNA sequences generated by traditional methods with near-perfect accuracy. These comprehensive results establish DNA-Stega as a robust and effective solution for secure and imperceptible DNA-based information hiding, significantly advancing DNA steganography. Weikun Fan, Zhenyi Tu, Yanju Zhang |
BIBM | 3 |
| 2025 | Hyperloc: A Hypergraph-Enhanced Multimodal Framework for Multi-Label mRNA Subcellular Localization PredictionabstractThe subcellular localization of messenger RNAs (mRNAs) provides valuable insight into their biological functions. In order to capture the co-localization of mRNA in multiple compartments, many efforts have been made to develop deep learning-based multi-label prediction models. However, existing methods have limitations in modeling mRNA secondary structures, particularly in capturing the critical roles of substructures such as stem-loops. Additionally, the lack of efficient sequencestructure multimodal fusion mechanisms leads to prediction biases. To address these challenges, we propose the Hyperloc framework, which integrates two modalities: mRNA sequence and secondary structure. For the sequence modality, the framework integrates local features extracted by convolutional neural networks with global statistical features, thereby enabling a multi-level representation of sequence information. For the structural modality, we innovatively introduce a hypergraph neural network to model substructure relationships in the secondary structure, capturing higher-order connectivity features among bases. Furthermore, an attention-based fusion strategy is employed to dynamically integrate sequence and structural modalities, adaptively adjusting their weights. Finally, a focal loss function is applied in the multilabel classification task to optimize feature learning. Extensive experimental results demonstrate that Hyperloc outperforms current state-of-the-art methods across multiple metrics, indicating its powerful capability in revealing complex localization relationships. The source code and dataset are publicly available at https://anonymous.4open.science/r/Hyperloc. Zhenyi Tu, Weikun Fan, Yanju Zhang |
BIBM | 3 |
| 2024 | MATL: A deep neural network using multi-scale convolutions and transformer for transcription factor binding site predictionabstractTranscription factors bind to specific sequences of DNA known as transcription factor binding sites (TFBSs) in order to regulate gene expression. Identifying TFBSs is crucial for deciphering gene expression mechanisms, understanding in vitro life characteristics, and drug design. In recent years, many deep learning methods that utilize only the sequence features or a combination of sequence and shape features have been developed to predict TFBSs. However, these methods have not fully considered the characteristics of TFBSs, and their encoding strategies are relatively simplistic, thereby failing to capture the complex dependencies hidden within TFBSs. To address these issues, we propose a novel deep neural network model called MATL, which integrates sequence and shape features for TFBSs prediction. The model employs techniques such as multi-scale convolution and Transformer, and combines different modules for feature extraction based on the distribution patterns of different feature matrices. Additionally, we introduce transfer learning to address the issue of unsatisfactory prediction results due to insufficient training data in some datasets. As the results indicate, MATL exhibits excellent predictive performance on 165 validated ChIP-seq datasets, surpassing state-of-the-art methods. We also conducted ablation experiments on sequence encoding, feature extraction, and shape modules used in this study, and demonstrated that these modules play a positive role in predicting TFBSs. Moreover, the results also indicated that transfer learning can improve prediction accuracy for some TFBSs datasets with small amounts of data. The source code of MATL is available at https://github.com/LittleWhiteKing/MATL. Jinli Fan, Ruyi Shen, Yanju Zhang |
BIBM | 4 |
| 2023 | SpliceCannon: A novel framework for the prediction of canonical and non-canonical splice sites based on deep learningabstractAccurately predicting splice sites in DNA sequences is not only of utmost importance in the study of gene expression and protein synthesis, but also crucial to precisely understanding the pathogenesis of various diseases related to splicing. In the past few decades, several tools for identifying splice sites have been developed, however, many of them ignore the contextual influence in DNA sequences in the process of encoding, consequently leading to considerable false positives in prediction. Additionally, most methods solely investigate canonical splicing whereas pay little attention to non-canonical ones. To address the issues mentioned above, we propose a comprehensive method, SpliceCannon, for accurately predicting both canonical and non-canonical splice sites based on neural networks. This method uses the combination of k-nucleotide frequencies and Doc2vec to encode the sequence, then employs a series of comprehensive neural networks such as bidirectional long short-term memory networks and residual networks to effectively extract features from the original sequence while avoiding the problems such as gradient vanishing. The results demonstrate that our method consistently achieves high accuracy when tested on multiple species. Compared to other recent methods, our model exhibits superior performance, stronger generalization ability, and wider applicability. In particular, The recognition accuracy is also remarkably high on the non-canonical dataset. We envision that this proposed method will serve as a promising tool for splice sites prediction to facilitate genomic research. Yanju Zhang, Wangjing Qi, Ruohan Lin, Jinli Fan |
BIBM | 1 |
| 2023 | A Weakly Supervised Semantic Segmentation Method on Lung Adenocarcinoma Histopathology Images
Xiaobin Lan, Jiaming Mei, Ruohan Lin, Yanju Zhang |
ICIC (2) | 5 |
| 2023 | SpliceSCANNER: An Accurate and Interpretable Deep Learning-Based Method for Splice Site Prediction
Rongxing Wang, Wangjing Qi, Yanju Zhang |
ICIC (3) | 5 |
| 2021 | DeepVF: a deep learning-based hybrid framework for identifying virulence factors using the stacking strategyabstractVirulence factors (VFs) enable pathogens to infect their hosts. A wealth of individual, disease-focused studies has identified a wide variety of VFs, and the growing mass of bacterial genome sequence data provides an opportunity for computational methods aimed at predicting VFs. Despite their attractive advantages and performance improvements, the existing methods have some limitations and drawbacks. Firstly, as the characteristics and mechanisms of VFs are continually evolving with the emergence of antibiotic resistance, it is more and more difficult to identify novel VFs using existing tools that were previously developed based on the outdated data sets; secondly, few systematic feature engineering efforts have been made to examine the utility of different types of features for model performances, as the majority of tools only focused on extracting very few types of features. By addressing the aforementioned issues, the accuracy of VF predictors can likely be significantly improved. This, in turn, would be particularly useful in the context of genome wide predictions of VFs. In this work, we present a deep learning (DL)-based hybrid framework (termed DeepVF) that is utilizing the stacking strategy to achieve more accurate identification of VFs. Using an enlarged, up-to-date dataset, DeepVF comprehensively explores a wide range of heterogeneous features with popular machine learning algorithms. Specifically, four classical algorithms, including random forest, support vector machines, extreme gradient boosting and multilayer perceptron, and three DL algorithms, including convolutional neural networks, long short-term memory networks and deep neural networks are employed to train 62 baseline models using these features. In order to integrate their individual strengths, DeepVF effectively combines these baseline models to construct the final meta model using the stacking strategy. Extensive benchmarking experiments demonstrate the effectiveness of DeepVF: it achieves a more accurate and stable performance compared with baseline models on the benchmark dataset and clearly outperforms state-of-the-art VF predictors on the independent test. Using the proposed hybrid ensemble model, a user-friendly online predictor of DeepVF (http://deepvf.erc.monash.edu/) is implemented. Furthermore, its utility, from the user's viewpoint, is compared with that of existing toolkits. We believe that DeepVF will be exploited as a useful tool for screening and identifying potential VFs from protein-coding gene sequences in bacterial genomes. Ruopeng Xie, Jiahui Li 0007, Jiawei Wang 0002, André Leier, Tatiana T. Marquez-Lago, Tatsuya Akutsu, Trevor Lithgow, Jiangning Song, Yanju Zhang |
Briefings Bioinform. | 10 |
| 2020 | PeNGaRoo, a combined gradient boosting and ensemble learning framework for predicting non-classical secreted proteinsabstractMOTIVATION: Gram-positive bacteria have developed secretion systems to transport proteins across their cell wall, a process that plays an important role during host infection. These secretion mechanisms have also been harnessed for therapeutic purposes in many biotechnology applications. Accordingly, the identification of features that select a protein for efficient secretion from these microorganisms has become an important task. Among all the secreted proteins, 'non-classical' secreted proteins are difficult to identify as they lack discernable signal peptide sequences and can make use of diverse secretion pathways. Currently, several computational methods have been developed to facilitate the discovery of such non-classical secreted proteins; however, the existing methods are based on either simulated or limited experimental datasets. In addition, they often employ basic features to train the models in a simple and coarse-grained manner. The availability of more experimentally validated datasets, advanced feature engineering techniques and novel machine learning approaches creates new opportunities for the development of improved predictors of 'non-classical' secreted proteins from sequence data. RESULTS: In this work, we first constructed a high-quality dataset of experimentally verified 'non-classical' secreted proteins, which we then used to create benchmark datasets. Using these benchmark datasets, we comprehensively analyzed a wide range of features and assessed their individual performance. Subsequently, we developed a two-layer Light Gradient Boosting Machine (LightGBM) ensemble model that integrates several single feature-based models into an overall prediction framework. At this stage, LightGBM, a gradient boosting machine, was used as a machine learning approach and the necessary parameter optimization was performed by a particle swarm optimization strategy. All single feature-based LightGBM models were then integrated into a unified ensemble model to further improve the predictive performance. Consequently, the final ensemble model achieved a superior performance with an accuracy of 0.900, an F-value of 0.903, Matthew's correlation coefficient of 0.803 and an area under the curve value of 0.963, and outperforming previous state-of-the-art predictors on the independent test. Based on our proposed optimal ensemble model, we further developed an accessible online predictor, PeNGaRoo, to serve users' demands. We believe this online web server, together with our proposed methodology, will expedite the discovery of non-classically secreted effector proteins in Gram-positive bacteria and further inspire the development of next-generation predictors. AVAILABILITY AND IMPLEMENTATION: http://pengaroo.erc.monash.edu/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yanju Zhang, Sha Yu, Ruopeng Xie, Jiahui Li 0007, André Leier, Tatiana T. Marquez-Lago, Tatsuya Akutsu, Alexander Ian Smith, ZongYuan Ge, Jiawei Wang 0002, Trevor Lithgow, Jiangning Song |
Bioinform. | 1 |
| 2019 | Computational analysis and prediction of lysine malonylation sites by exploiting informative features in an integrative machine-learning frameworkabstractAs a newly discovered post-translational modification (PTM), lysine malonylation (Kmal) regulates a myriad of cellular processes from prokaryotes to eukaryotes and has important implications in human diseases. Despite its functional significance, computational methods to accurately identify malonylation sites are still lacking and urgently needed. In particular, there is currently no comprehensive analysis and assessment of different features and machine learning (ML) methods that are required for constructing the necessary prediction models. Here, we review, analyze and compare 11 different feature encoding methods, with the goal of extracting key patterns and characteristics from residue sequences of Kmal sites. We identify optimized feature sets, with which four commonly used ML methods (random forest, support vector machines, K-nearest neighbor and logistic regression) and one recently proposed [Light Gradient Boosting Machine (LightGBM)] are trained on data from three species, namely, Escherichia coli, Mus musculus and Homo sapiens, and compared using randomized 10-fold cross-validation tests. We show that integration of the single method-based models through ensemble learning further improves the prediction performance and model robustness on the independent test. When compared to the existing state-of-the-art predictor, MaloPred, the optimal ensemble models were more accurate for all three species (AUC: 0.930, 0.923 and 0.944 for E. coli, M. musculus and H. sapiens, respectively). Using the ensemble models, we developed an accessible online predictor, kmal-sp, available at http://kmalsp.erc.monash.edu/. We hope that this comprehensive survey and the proposed strategy for building more accurate models can serve as a useful guide for inspiring future developments of computational methods for PTM site prediction, expedite the discovery of new malonylation and other PTM types and facilitate hypothesis-driven experimental validation of novel malonylated substrates and malonylation sites. Yanju Zhang, Ruopeng Xie, Jiawei Wang 0002, André Leier, Tatiana T. Marquez-Lago, Tatsuya Akutsu, Geoffrey I. Webb, Kuo-Chen Chou, Jiangning Song |
Briefings Bioinform. | 1 |
| 2019 | Bastion3: a two-layer ensemble predictor of type III secreted effectorsabstractMOTIVATION: Type III secreted effectors (T3SEs) can be injected into host cell cytoplasm via type III secretion systems (T3SSs) to modulate interactions between Gram-negative bacterial pathogens and their hosts. Due to their relevance in pathogen-host interactions, significant computational efforts have been put toward identification of T3SEs and these in turn have stimulated new T3SE discoveries. However, as T3SEs with new characteristics are discovered, these existing computational tools reveal important limitations: (i) most of the trained machine learning models are based on the N-terminus (or incorporating also the C-terminus) instead of the proteins' complete sequences, and (ii) the underlying models (trained with classic algorithms) employed only few features, most of which were extracted based on sequence-information alone. To achieve better T3SE prediction, we must identify more powerful, informative features and investigate how to effectively integrate these into a comprehensive model. RESULTS: In this work, we present Bastion3, a two-layer ensemble predictor developed to accurately identify type III secreted effectors from protein sequence data. In contrast with existing methods that employ single models with few features, Bastion3 explores a wide range of features, from various types, trains single models based on these features and finally integrates these models through ensemble learning. We trained the models using a new gradient boosting machine, LightGBM and further boosted the models' performances through a novel genetic algorithm (GA) based two-step parameter optimization strategy. Our benchmark test demonstrates that Bastion3 achieves a much better performance compared to commonly used methods, with an ACC value of 0.959, F-value of 0.958, MCC value of 0.917 and AUC value of 0.956, which comprehensively outperformed all other toolkits by more than 5.6% in ACC value, 5.7% in F-value, 12.4% in MCC value and 5.8% in AUC value. Based on our proposed two-layer ensemble model, we further developed a user-friendly online toolkit, maximizing convenience for experimental scientists toward T3SE prediction. With its design to ease future discoveries of novel T3SEs and improved performance, Bastion3 is poised to become a widely used, state-of-the-art toolkit for T3SE prediction. AVAILABILITY AND IMPLEMENTATION: http://bastion3.erc.monash.edu/. CONTACT: [email protected] or [email protected] or or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jiawei Wang 0002, Jiahui Li 0007, Bingjiao Yang, Ruopeng Xie, Tatiana T. Marquez-Lago, André Leier, Morihiro Hayashida, Tatsuya Akutsu, Yanju Zhang, Kuo-Chen Chou, Joel Selkrig, Tieli Zhou, Jiangning Song, Trevor Lithgow |
Bioinform. | 9 |
| 2018 | Bastion6: a bioinformatics approach for accurate prediction of type VI secreted effectorsabstractMotivation: Many Gram-negative bacteria use type VI secretion systems (T6SS) to export effector proteins into adjacent target cells. These secreted effectors (T6SEs) play vital roles in the competitive survival in bacterial populations, as well as pathogenesis of bacteria. Although various computational analyses have been previously applied to identify effectors secreted by certain bacterial species, there is no universal method available to accurately predict T6SS effector proteins from the growing tide of bacterial genome sequence data. Results: We extracted a wide range of features from T6SE protein sequences and comprehensively analyzed the prediction performance of these features through unsupervised and supervised learning. By integrating these features, we subsequently developed a two-layer SVM-based ensemble model with fine-grain optimized parameters, to identify potential T6SEs. We further validated the predictive model using an independent dataset, which showed that the proposed model achieved an impressive performance in terms of ACC (0.943), F-value (0.946), MCC (0.892) and AUC (0.976). To demonstrate applicability, we employed this method to correctly identify two very recently validated T6SE proteins, which represent challenging prediction targets because they significantly differed from previously known T6SEs in terms of their sequence similarity and cellular function. Furthermore, a genome-wide prediction across 12 bacterial species, involving in total 54 212 protein sequences, was carried out to distinguish 94 putative T6SE candidates. We envisage both this information and our publicly accessible web server will facilitate future discoveries of novel T6SEs. Availability and implementation: http://bastion6.erc.monash.edu/. Supplementary information: Supplementary data are available at Bioinformatics online. Jiawei Wang 0002, Bingjiao Yang, André Leier, Tatiana T. Marquez-Lago, Morihiro Hayashida, Andrea Rocker, Yanju Zhang, Tatsuya Akutsu, Kuo-Chen Chou, Richard A. Strugnell, Jiangning Song, Trevor Lithgow |
Bioinform. | 7 |
| 2012 | PASSion: a pattern growth algorithm-based pipeline for splice junction detection in paired-end RNA-Seq dataabstractMOTIVATION: RNA-seq is a powerful technology for the study of transcriptome profiles that uses deep-sequencing technologies. Moreover, it may be used for cellular phenotyping and help establishing the etiology of diseases characterized by abnormal splicing patterns. In RNA-Seq, the exact nature of splicing events is buried in the reads that span exon-exon boundaries. The accurate and efficient mapping of these reads to the reference genome is a major challenge. RESULTS: We developed PASSion, a pattern growth algorithm-based pipeline for splice site detection in paired-end RNA-Seq reads. Comparing the performance of PASSion to three existing RNA-Seq analysis pipelines, TopHat, MapSplice and HMMSplicer, revealed that PASSion is competitive with these packages. Moreover, the performance of PASSion is not affected by read length and coverage. It performs better than the other three approaches when detecting junctions in highly abundant transcripts. PASSion has the ability to detect junctions that do not have known splicing motifs, which cannot be found by the other tools. Of the two public RNA-Seq datasets, PASSion predicted ≈ 137,000 and 173,000 splicing events, of which on average 82 are known junctions annotated in the Ensembl transcript database and 18% are novel. In addition, our package can discover differential and shared splicing patterns among multiple samples. AVAILABILITY: The code and utilities can be freely downloaded from https://trac.nbic.nl/passion and ftp://ftp.sanger.ac.uk/pub/zn1/passion. Yanju Zhang, Eric-Wubbo Lameijer, Peter A. C. 't Hoen, Zemin Ning, P. Eline Slagboom, Kai Ye 0001 |
Bioinform. | 1 |
| 2008 | miRNA target prediction through mining of miRNA relationshipsabstractmiRNAs are small regulators that mediate gene expression and each miRNA regulates specific target genes. In animals, target prediction of the miRNAs is accomplished through several computational methods, i.e. miRanda, TargetScan and PicTar. Typically, these methods predict targets from features of miRNA-target interaction such as sequence complementarity, free energy of RNA duplexes and conservation of target sites. They are constructed for high throughput and also result in a large amount of predictions and a high estimated false-positive rate. To date, specific rules to capture all known miRNA targets have not been devised. We observed that miRNAs sometimes share targets. Therefore, in this paper we present an approach which analyzes miRNA-miRNA relationships and utilizes them for target prediction.We use machine learning techniques to reveal the feature patterns between known miRNAs. Different data setups are evaluated and compared to achieve the best performance. Furthermore, the derived rules are applied to miRNAs of which the targets are not yet known so as to see if new targets could be predicted. In the analysis of functionally similar miRNAs, we found that genomic distance and seed similarity between miRNAs are dominant features in the description of a group of miRNAs binding the same target. Application of one specific rule resulted in the prediction of targets for seven miRNAs for which the targets were formerly unknown. Some of these targets were also detected by the existing methods. Our method contributes to the improvement of target identification by predicting targets with high specificity and without conservation limitation. Yanju Zhang, Jeroen S. de Bruin, Fons J. Verbeek |
BIBE | 1 |