EDBT 2026 Demo / reviewers in the wild / expert
Yushi Luan
dblp:142/0293
· DBLP profile ↗
22ranked-venue papers
0as first author
11since 2021 · last 2025
0000-0001-8955-041XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 19 · 9 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Attention-augmented multi-domain cooperative graph representation learning for molecular interaction prediction
Zhaowei Wang 0005, Jun Meng, Qiguo Dai, Xiaohui Lin 0002, Yushi Luan |
Neural Networks | 6 |
| 2024 | DeepPepPI: A deep cross-dependent framework with information sharing mechanism for predicting plant peptide-protein interactions
Zhaowei Wang 0005, Jun Meng, Qiguo Dai, Shihao Xia, Ruirui Yang, Yushi Luan |
Expert Syst. Appl. | 7 |
| 2023 | TGAAL: Combining Transformer-based GAN and active learning to identify the coding potential of sORFs in plant lncRNAsabstractSome small open reading frames (sORFs) in plant long non-coding RNAs (lncRNAs) are capable of encoding small peptides, which play key roles in the growth and development of organisms. Therefore, it is particularly important to identify the coding potential of sORFs in plant lncRNAs. However, existing methods often ignore the differences in length distribution between coding sORFs (csORFs) and non-coding sORFs (non-csORFs), which may lead to incorrect identification of csORFs. To address this issue, we propose a novel method to identify the coding potential of sORFs in plant lncRNAs, named Transformer Generative Adversarial Active Learning (TGAAL), which combines Transformer-based Generative Adversarial Network (TGAN) and active learning based on KL-topk sampling strategy. TGAN can generate sORF sequences in a specific length interval, which have the same class as the input sORFs. Meanwhile, using active learning based on KL-topk sampling strategy, samples with high confidence can be selected for data augmentation. 5-fold cross-validation shows that KL-topk sampling strategy significantly improves the prediction performance compared with commonly adopted sampling strategies. The experimental results show that TGAAL significantly outperforms existing methods in identifying the coding potential of sORFs in Arabidopsis thaliana, reaching 0.7761, 0.7906 and 0.7529 unweighted average recall in three sORF length intervals, respectively. Jun Meng, Shihao Xia, Zhaowei Wang 0005, Yushi Luan |
BIBM | 6 |
| 2023 | A multi-granularity information-enhanced pre-training method for predicting the coding potential of sORFs in plant lncRNAsabstractSmall open reading frames (sORFs) are nucleotide sequences that may be translated into small peptides. Recently, increasing studies have demonstrated that peptides encoded by sORFs in plant long noncoding RNAs (lncRNAs) play a vital role in growth regulation and disease treatment. To accelerate the discovery of lncRNA-encoded peptides, it is essential to predict translatable sORFs in lncRNAs (lncRNA-sORFs) by computational methods. As only a few translatable plant lncRNA-sORFs have been discovered to date, there is a lack of effective methods for characterizing the coding potential of lncRNA-sORFs in data-scarce scenarios. Therefore, a novel method for plant lncRNA-sORFs coding potential prediction using the pre-trained bidirectional encoder representations from transformer (LSCPP-BERT) is proposed. Firstly, the BERT model is trained to extract multi-granularity context information from large-scale unlabeled lncRNA-sORFs through two pre-training tasks. Then, the pre-trained model can be fine-tuned with two additional linear layers for classification. The LSCPP-BERT is featured by a self-supervised pre-training scheme and multi-granularity context information, aiming to enhance the representational power of the network. In addition, an extra pre-training task called contextual relation of lncRNA-sORFs prediction (CRSP) is presented to extract sentence-level information. Experiment results show that the accuracy of LSCPP-BERT is increased by 8.14% compared with state-of-the-art methods. We hope that the proposed method can serve as a reliable tool for the prediction of coding lncRNA-sORFs, thereby further contributing to drug development and agronomical applications. Shihao Xia, Jun Meng, Zhaowei Wang 0005, Zhaojing Qin, Yushi Luan |
BIBM | 7 |
| 2022 | Predicting the interactions between plant lncRNA-encoded peptide and protein using domain knowledge-based prototypical networkabstractLong noncoding RNA(lncRNA) has been reported to encode small peptides which play key roles in life activities through their functions by binding to proteins. It is crucial to predict the interactions between the lncRNA-encoded peptide and protein. However, no computational methods have been designed for predicting this type of interactions directly, owing to the few-shot problem causing poor generalization. Prototypical network (ProtoNet) is a classic learner for few-shot learning. However, how to obtain effective embedding and measure the distance between different prototypes accurately are the most important challenges. Although some improved prototypical networks have been proposed, they ignore the role of domain knowledge which is helpful for constructing models conforming to the domain mechanism In this study, we propose a novel method for interactions prediction between plant lncRNA-encoded peptide and protein using domain knowledge-based ProtoNet (IPLncPP-DKPN). Multiple features that imply domain knowledge are extracted, connected, and converted to avoid sparse and enhance information using a dual-routing parallel feature dimensionality reduction algorithm IProtoNet is an improved ProtoNet using capsule network-based embedding and Mahalanobis distance-based prototype. The converted features are fed into IProtoNet to realize the classification task. The experimental results manifest that IPLncPP-DKPN achieves better performance on the independent test set compared with classic machine learning models. To the best of our knowledge, IPLncPP-DKPN is the first computational method for the interactions prediction between lncRNA-encoded peptide and protein. Jun Meng, Yushi Luan |
BIBM | 3 |
| 2022 | RNAI-FRID: novel feature representation method with information enhancement and dimension reduction for RNA-RNA interactionabstractDifferent ribonucleic acids (RNAs) can interact to form regulatory networks that play important role in many life activities. Molecular biology experiments can confirm RNA-RNA interactions to facilitate the exploration of their biological functions, but they are expensive and time-consuming. Machine learning models can predict potential RNA-RNA interactions, which provide candidates for molecular biology experiments to save a lot of time and cost. Using a set of suitable features to represent the sample is crucial for training powerful models, but there is a lack of effective feature representation for RNA-RNA interaction. This study proposes a novel feature representation method with information enhancement and dimension reduction for RNA-RNA interaction (named RNAI-FRID). Diverse base features are first extracted from RNA data to contain more sample information. Then, the extracted base features are used to construct the complex features through an arithmetic-level method. It greatly reduces the feature dimension while keeping the relationship between molecule features. Since the dimension reduction may cause information loss, in the process of complex feature construction, the arithmetic mean strategy is adopted to enhance the sample information further. Finally, three feature ranking methods are integrated for feature selection on constructed complex features. It can adaptively retain important features and remove redundant ones. Extensive experiment results show that RNAI-FRID can provide reliable feature representation for RNA-RNA interaction with higher efficiency and the model trained with generated features obtain better performance than other deep neural network predictors. Qiang Kang, Jun Meng, Yushi Luan |
Briefings Bioinform. | 3 |
| 2022 | Mining plant endogenous target mimics from miRNA-lncRNA interactions based on dual-path parallel ensemble pruning methodabstractThe interactions between microRNAs (miRNAs) and long non-coding RNAs (lncRNAs) play important roles in biological activities. Specially, lncRNAs as endogenous target mimics (eTMs) can bind miRNAs to regulate the expressions of target messenger RNAs (mRNAs). A growing number of studies focus on animals, but the studies on plants are scarce and many functions of plant eTMs are unknown. This study proposes a novel ensemble pruning protocol for predicting plant miRNA-lncRNA interactions at first. It adaptively prunes the base models based on dual-path parallel ensemble method to meet the challenge of cross-species prediction. Then potential eTMs are mined from predicted results. The expression levels of RNAs are identified through biological experiment to construct the lncRNA-miRNA-mRNA regulatory network, and the functions of potential eTMs are inferred through enrichment analysis. Experiment results show that the proposed protocol outperforms existing methods and state-of-the-art predictors on various plant species. A total of 17 potential eTMs are verified by biological experiment to involve in 22 regulations, and 14 potential eTMs are inferred by Gene Ontology enrichment analysis to involve in 63 functions, which is significant for further research. Qiang Kang, Jun Meng, Chenglin Su, Yushi Luan |
Briefings Bioinform. | 4 |
| 2022 | Identifying LncRNA-Encoded Short Peptides Using Optimized Hybrid Features and Ensemble LearningabstractLong non-coding RNA (lncRNA) contains short open reading frames (sORFs), and sORFs-encoded short peptides (SEPs) have become the focus of scientific studies due to their crucial role in life activities. The identification of SEPs is vital to further understanding their regulatory function. Bioinformatics methods can quickly identify SEPs to provide credible candidate sequences for verifying SEPs by biological experimenrts. However, there is a lack of methods for identifying SEPs directly. In this study, a machine learning method to identify SEPs of plant lncRNA (ISPL) is proposed. Hybrid features including sequence features and physicochemical features are extracted manually or adaptively to construct different modal features. In order to keep the stability of feature selection, the non-linear correction applied in Max-Relevance-Max-Distance (nocRD) feature selection method is proposed, which integrates multiple feature ranking results and uses the iterative random forest for different modal features dimensionality reduction. Classification models with different modal features are constructed, and their outputs are combined for ensemble classification. The experimental results show that the accuracy of ISPL is 89.86% percent on the independent test set, which will have important implications for further studies of functional genomic. Jun Meng, Qiang Kang, Yushi Luan |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2021 | Optimized combination methods for exploring and verifying disease-resistant transcription factors in melonabstractA large amount of omics data and number of bioinformatics tools has been produced. However, the methods for further exploring omics data are simple, in particular, to mine key regulatory genes, which are a priority concern in biological systems, and most of the specific functions are still unknown. First, raw data of two genotypes of melon (susceptible and resistant) were obtained by transcriptome analysis. Second, 391 transcription factors (TFs) were identified from the plant transcription factor database and cucurbit genomics database. Then, functional enrichment analysis indicated that these genes were mainly annotated in the process of transcription regulation. Third, 243 and 230 module-specific TFs were screened by weighted gene coexpression network analysis and short time series expression miner, respectively. Several TF genes, such as WRKYs and bHLHs, were regarded as key regulatory genes according to the values of significantly different modules. The coexpression network showed that these TF genes were significant correlated with resistance (R) genes, such as DRP2, RGA3, DRP1 and NB-ARC. Fourth, cis-acting element analysis illustrated that these R genes may bind to WRKY and bHLH. Finally, the expression of WRKY genes was verified by quantitative reverse transcription PCR (RT-qPCR). Phylogenetic analysis was carried out to further confirm that these TFs may play a critical role in Curcurbitaceae disease resistance. This study provides a new optimized combination strategy to explore the functions of TFs in a wide spectrum of biological processes. This strategy may also effectively predict potential relationships in the interactions of essential genes. Zhicheng Wang 0010, Yushi Luan, Xiaoxu Zhou, Jun Cui 0004, Feishi Luan, Jun Meng |
Briefings Bioinform. | 2 |
| 2021 | PlncRNA-HDeep: plant long noncoding RNA prediction using hybrid deep learning based on two encoding stylesabstractBACKGROUND: Long noncoding RNAs (lncRNAs) play an important role in regulating biological activities and their prediction is significant for exploring biological processes. Long short-term memory (LSTM) and convolutional neural network (CNN) can automatically extract and learn the abstract information from the encoded RNA sequences to avoid complex feature engineering. An ensemble model learns the information from multiple perspectives and shows better performance than a single model. It is feasible and interesting that the RNA sequence is considered as sentence and image to train LSTM and CNN respectively, and then the trained models are hybridized to predict lncRNAs. Up to present, there are various predictors for lncRNAs, but few of them are proposed for plant. A reliable and powerful predictor for plant lncRNAs is necessary. RESULTS: To boost the performance of predicting lncRNAs, this paper proposes a hybrid deep learning model based on two encoding styles (PlncRNA-HDeep), which does not require prior knowledge and only uses RNA sequences to train the models for predicting plant lncRNAs. It not only learns the diversified information from RNA sequences encoded by p-nucleotide and one-hot encodings, but also takes advantages of lncRNA-LSTM proposed in our previous study and CNN. The parameters are adjusted and three hybrid strategies are tested to maximize its performance. Experiment results show that PlncRNA-HDeep is more effective than lncRNA-LSTM and CNN and obtains 97.9% sensitivity, 95.1% precision, 96.5% accuracy and 96.5% F1 score on Zea mays dataset which are better than those of several shallow machine learning methods (support vector machine, random forest, k-nearest neighbor, decision tree, naive Bayes and logistic regression) and some existing tools (CNCI, PLEK, CPC2, LncADeep and lncRNAnet). CONCLUSIONS: PlncRNA-HDeep is feasible and obtains the credible predictive results. It may also provide valuable references for other related research. Jun Meng, Qiang Kang, Yushi Luan |
BMC Bioinform. | 4 |
| 2021 | PRPI-SC: an ensemble deep learning model for predicting plant lncRNA-protein interactionsabstractBACKGROUND: Plant long non-coding RNAs (lncRNAs) play vital roles in many biological processes mainly through interactions with RNA-binding protein (RBP). To understand the function of lncRNAs, a fundamental method is to identify which types of proteins interact with the lncRNAs. However, the models or rules of interactions are a major challenge when calculating and estimating the types of RBP. RESULTS: In this study, we propose an ensemble deep learning model to predict plant lncRNA-protein interactions using stacked denoising autoencoder and convolutional neural network based on sequence and structural information, named PRPI-SC. PRPI-SC predicts interactions between lncRNAs and proteins based on the k-mer features of RNAs and proteins. Experiments proved good results on Arabidopsis thaliana and Zea mays datasets (ATH948 and ZEA22133). The accuracy rates of ATH948 and ZEA22133 datasets were 88.9% and 82.6%, respectively. PRPI-SC also performed well on some public RNA protein interaction datasets. CONCLUSIONS: PRPI-SC accurately predicts the interaction between plant lncRNA and protein, which plays a guiding role in studying the function and expression of plant lncRNA. At the same time, PRPI-SC has a strong generalization ability and good prediction effect for non-plant data. Jael Sanyanda Wekesa, Yushi Luan, Jun Meng |
BMC Bioinform. | 3 |
| 2020 | LPI-DL: A recurrent deep learning model for plant lncRNA-protein interaction and function prediction with feature optimizationabstractPredicting lncRNA-protein association is essential for insights into fundamental biological processes and disease etiology in plants and animals. There has been an enormous increment in the number of identified long noncoding RNAs (lncRNAs). However, less efforts has been directed towards lncRNA-protein interaction (LPI) prediction to help in characterizing the huge array of plant lncRNAs. This study presents LPI-DL, a deep learning method for predicting potential plant lncRNA-protein interaction based on sequence features and compact LSTM. The optimal combination of k-nucleotide frequencies and codon-based encoding features are used as input to the model. The recurrent neural network learns the discriminative features characterizing the long-term dependencies between sequences. We select optimal features using recursive feature elimination and support vector machine (RFE-SVM) and impose sparse projection onto the hidden states of input sequences through connection pruning. Evaluation of two plant datasets corroborates that LPI-DL is more competitive over other methods. Comparative experiments denote that the proposed method achieves state-of-the-art prediction performance. This study effectively improves the accuracy of interaction prediction and lays a foundation to foster lncRNA functional studies. Jael Sanyanda Wekesa, Yushi Luan, Jun Meng |
BIBM | 2 |
| 2020 | PmliPred: a method based on hybrid model and fuzzy decision for plant miRNA-lncRNA interaction predictionabstractMOTIVATION: The studies have indicated that not only microRNAs (miRNAs) or long non-coding RNAs (lncRNAs) play important roles in biological activities, but also their interactions affect the biological process. A growing number of studies focus on the miRNA-lncRNA interactions, while few of them are proposed for plant. The prediction of interactions is significant for understanding the mechanism of interaction between miRNA and lncRNA in plant. RESULTS: This article proposes a new method for fulfilling plant miRNA-lncRNA interaction prediction (PmliPred). The deep learning model and shallow machine learning model are trained using raw sequence and manually extracted features, respectively. Then they are hybridized based on fuzzy decision for prediction. PmliPred shows better performance and generalization ability compared with the existing methods. Several new miRNA-lncRNA interactions in Solanum lycopersicum are successfully identified using quantitative real time-polymerase chain reaction from the candidates predicted by PmliPred, which further verifies its effectiveness. AVAILABILITY AND IMPLEMENTATION: The source code of PmliPred is freely available at http://bis.zju.edu.cn/PmliPred/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Qiang Kang, Jun Meng, Jun Cui 0004, Yushi Luan, Ming Chen 0005 |
Bioinform. | 4 |
| 2019 | lncRNA-LSTM: Prediction of Plant Long Non-coding RNAs Using Long Short-Term Memory Based on p-nts Encoding
Jun Meng, Yushi Luan |
ICIC (3) | 5 |
| 2019 | Prediction of Plant lncRNA-Protein Interactions Using Sequence Information Based on Deep Learning
Yushi Luan, Jael Sanyanda Wekesa, Jun Meng |
ICIC (3) | 2 |
| 2018 | Prediction of LncRNA by Using Muitiple Feature Information Fusion and Feature Selection Technique
Jun Meng, Dingling Jiang, Yushi Luan |
ICIC (2) | 4 |
| 2017 | Genome-Wide Analysis of Response Regulator Genes in Solanum lycopersicum
Jun Cui 0004, Jun Meng, Yushi Luan |
ISBRA | 4 |
| 2016 | Plant miRNA function prediction based on functional similarity network and transductive multi-label classification algorithm
Jun Meng, Guan-Li Shi, Yushi Luan |
Neurocomputing | 3 |
| 2016 | Classifier ensemble selection based on affinity propagation clustering
Jun Meng, Yushi Luan |
J. Biomed. Informatics | 3 |
| 2015 | Inferring plant microRNA functional similarity using a weighted protein-protein interaction networkabstractBACKGROUND: MiRNAs play a critical role in the response of plants to abiotic and biotic stress. However, the functions of most plant miRNAs remain unknown. Inferring these functions from miRNA functional similarity would thus be useful. This study proposes a new method, called PPImiRFS, for inferring miRNA functional similarity. RESULTS: The functional similarity of miRNAs was inferred from the functional similarity of their target gene sets. A protein-protein interaction network with semantic similarity weights of edges generated using Gene Ontology terms was constructed to infer the functional similarity between two target genes that belong to two different miRNAs, and the score for functional similarity was calculated using the weighted shortest path for the two target genes through the whole network. The experimental results showed that the proposed method was more effective and reliable than previous methods (miRFunSim and GOSemSim) applied to Arabidopsis thaliana. Additionally, miRNAs responding to the same type of stress had higher functional similarity than miRNAs responding to different types of stress. CONCLUSIONS: For the first time, a protein-protein interaction network with semantic similarity weights generated using Gene Ontology terms was employed to calculate the functional similarity of plant miRNAs. A novel method based on calculating the weighted shortest path between two target genes was introduced. Jun Meng, Yushi Luan |
BMC Bioinform. | 3 |
| 2015 | Gene Selection Integrated with Biological Knowledge for Plant Stress Response Using Neighborhood System and Rough Set TheoryabstractMining knowledge from gene expression data is a hot research topic and direction of bioinformatics. Gene selection and sample classification are significant research trends, due to the large amount of genes and small size of samples in gene expression data. Rough set theory has been successfully applied to gene selection, as it can select attributes without redundancy. To improve the interpretability of the selected genes, some researchers introduced biological knowledge. In this paper, we first employ neighborhood system to deal directly with the new information table formed by integrating gene expression data with biological knowledge, which can simultaneously present the information in multiple perspectives and do not weaken the information of individual gene for selection and classification. Then, we give a novel framework for gene selection and propose a significant gene selection method based on this framework by employing reduction algorithm in rough set theory. The proposed method is applied to the analysis of plant stress response. Experimental results on three data sets show that the proposed method is effective, as it can select significant gene subsets without redundancy and achieve high classification accuracy. Biological analysis for the results shows that the interpretability is well. Jun Meng, Yushi Luan |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2014 | Prediction of plant pre-microRNAs and their microRNAs in genome-scale sequences using structure-sequence features and support vector machineabstractBACKGROUND: MicroRNAs (miRNAs) are a family of non-coding RNAs approximately 21 nucleotides in length that play pivotal roles at the post-transcriptional level in animals, plants and viruses. These molecules silence their target genes by degrading transcription or suppressing translation. Studies have shown that miRNAs are involved in biological responses to a variety of biotic and abiotic stresses. Identification of these molecules and their targets can aid the understanding of regulatory processes. Recently, prediction methods based on machine learning have been widely used for miRNA prediction. However, most of these methods were designed for mammalian miRNA prediction, and few are available for predicting miRNAs in the pre-miRNAs of specific plant species. Although the complete Solanum lycopersicum genome has been published, only 77 Solanum lycopersicum miRNAs have been identified, far less than the estimated number. Therefore, it is essential to develop a prediction method based on machine learning to identify new plant miRNAs. RESULTS: A novel classification model based on a support vector machine (SVM) was trained to identify real and pseudo plant pre-miRNAs together with their miRNAs. An initial set of 152 novel features related to sequential structures was used to train the model. By applying feature selection, we obtained the best subset of 47 features for use with the Back Support Vector Machine-Recursive Feature Elimination (B-SVM-RFE) method for the classification of plant pre-miRNAs. Using this method, 63 features were obtained for plant miRNA classification. We then developed an integrated classification model, miPlantPreMat, which comprises MiPlantPre and MiPlantMat, to identify plant pre-miRNAs and their miRNAs. This model achieved approximately 90% accuracy using plant datasets from nine plant species, including Arabidopsis thaliana, Glycine max, Oryza sativa, Physcomitrella patens, Medicago truncatula, Sorghum bicolor, Arabidopsis lyrata, Zea mays and Solanum lycopersicum. Using miPlantPreMat, 522 Solanum lycopersicum miRNAs were identified in the Solanum lycopersicum genome sequence. CONCLUSIONS: We developed an integrated classification model, miPlantPreMat, based on structure-sequence features and SVM. MiPlantPreMat was used to identify both plant pre-miRNAs and the corresponding mature miRNAs. An improved feature selection method was proposed, resulting in high classification accuracy, sensitivity and specificity. Jun Meng, Yushi Luan |
BMC Bioinform. | 4 |