EDBT 2026 Demo / reviewers in the wild / expert
Shankai Yan
dblp:206/0550
· DBLP profile ↗
19ranked-venue papers
6as first author
13since 2021 · last 2026
0000-0003-0369-4979ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 15 · 5 first-author · 10 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Consensus-on-Graph: Plan-Driven Exploration and Consensus Decision-Making on Knowledge Graphs
Chengye Hu, Buchao Zhan, Wenqi Fan, Shankai Yan |
DASFAA (3) | 5 |
| 2026 | DeepNhKcr: Explainable Deep Learning Framework for the Prediction of Crotonylation Sites of Non-Histone Lysine in Plants Based on Pre-Trained Protein Language ModelabstractLysine crotonylation (Kcr) is an important protein modification occurring after translation in biology, serving an essential function in a range of biological processes in both plants and animals, including the regulation of gene expression, the maintenance of cellular metabolic balance, and the enhancement of photosynthesis. Exploring the detection of Kcr sites is essential for uncovering their biological functions. Nonetheless, conventional experimental approaches for detection are often time-consuming, expensive, and hindered by various technical constraints, making the precise identification of Kcr sites a significant challenge. This study seeks to develop a computational approach for the rapid and accurate prediction of Kcr sites in plant non-histone proteins. We introduce a novel deep learning framework named DeepNhKcr, which integrates the protein language model (ESM2) with a bidirectional long short-term memory (BiLSTM) network. To address the challenge of data imbalance, the model replaces the conventional cross-entropy loss with the focal loss function. In addition, DeepNhKcr combines advanced deep learning approaches with traditional protein encoding strategies to enable effective feature extraction and integration. This method not only significantly boosts the accuracy of predicting Kcr sites in non-histone proteins of plants. but also provides interpretability, shedding light on the potential links between key sequence characteristics and their biological roles. DeepNhKcr delivers outstanding results, surpassing existing machine learning and deep learning models, and demonstrating excellent performance in both five-fold cross-validation and independent test experiments. Moreover, the model integrates interpretability analysis techniques to investigate the connections between important sequence features and their biological roles. DeepNhKcr acts as a powerful method for detecting Kcr sites in plant non-histone proteins and is anticipated to greatly advance future studies in plant Kcr site prediction. Zhenjie Luo, Aoyun Geng, Junlin Xu, Yajie Meng, Shankai Yan, Leyi Wei, Qingchen Zhang 0001, Quan Zou 0001, Feifei Cui |
IEEE Trans. Comput. Biol. Bioinform. | 6 |
| 2026 | FNatPred: A Data-Driven Approach for Distinguishing Between NAT and Tumor on the Fungal MicrobiomeabstractOBJECTIVE: The role of fungal microbiota in human carcinogenesis remains largely uncharacterized. Recent evidence suggests normal adjacent tissue (NAT) represents an intermediate state between healthy and malignant tissues, highlighting its potential for early cancer detection. Discriminating fungal compositional profiles between tumor and NAT is thus critical for elucidating fungal involvement in oncogenesis. However, the high similarity between tumor and NAT mycobiota poses significant analytical challenges. METHOD: To overcome this limitation, we developed a two-level ensemble discriminative model. Base-level classifiers, trained using rigorously filtered fungal microbiota data (based on prevalence, abundance, and quality metrics) via Random Forest, generate initial predictions. A meta-level classifier then integrates these base predictions, transforming high-dimensional, sparse fungal feature data into a low-dimensional, dense representation optimized for discrimination. RESULTS: Our approach achieved clear separation between tumor and NAT mycobiomes across multiple cancer types, with particularly pronounced discrimination in colorectal cancer (CRC). The proposed model significantly outperformed existing methods in tumor-NAT classification, demonstrating an average AUC improvement of approximately 10%. Buchao Zhan, Dongmei He, Xin Yang 0037, Shankai Yan |
IEEE Trans. Comput. Biol. Bioinform. | 5 |
| 2025 | scCAVAE: Modeling Synergistic Interactions in Perturbation Responses via Attention and Hierarchical Supervised Contrastive LearningabstractPredicting cellular responses to combinatorial perturbations is a central challenge in drug discovery and therapeutic development. The combinatorial explosion of possible treatments renders exhaustive experimental screening infeasible, necessitating accurate predictive models. However, existing computational models often rely on oversimplified additiveeffect assumptions and struggle to distinguish perturbations that produce similar yet distinct effects within latent spaces. To overcome these limitations, we introduce scCAVAE, a deep learning framework that integrates a multi-head attention encoder to capture nonlinear synergistic and antagonistic interactions. The framework employs a novel hierarchical supervised contrastive learning strategy to engineer a semantically structured latent space by encoding intrinsic hierarchical relationships (e.g., inclusion, overlap, and disjointness) among perturbation combinations. Evaluation across multiple public benchmark datasets demonstrates that scCAVAE consistently outperforms existing methods, particularly in out-of-distribution (OOD) prediction tasks. Our findings provide a robust computational tool for high-throughput perturbation screening and establish a novel paradigm for modeling complex cellular responses. Our code is publicly available (https://github.com/cskyan/scCAVAE). Buchao Zhan, Jiangbo Zhang, Shankai Yan |
BIBM | 4 |
| 2025 | RMDNet: RNA-aware dung beetle optimization-based multi-branch integration network for RNA-protein binding sites predictionabstractRNA-binding proteins (RBPs) play crucial roles in gene regulation. Their dysregulation has been increasingly linked to neurodegenerative diseases, liver cancer, and lung cancer. Although experimental methods like CLIP-seq accurately identify RNA-protein binding sites, they are time-consuming and costly. To address this, we propose RMDNet-a deep learning framework that integrates CNN, CNN-Transformer, and ResNet branches to capture features at multiple sequence scales. These features are fused with structural representations derived from RNA secondary structure graphs. The graphs are processed using a graph neural network with DiffPool. To optimize feature integration, we incorporate an improved dung beetle optimization algorithm, which adaptively assigns fusion weights during inference. Evaluations on the RBP-24 benchmark show that RMDNet outperforms state-of-the-art models including GraphProt, DeepRKE, and DeepDW across multiple metrics. On the RBP-31 dataset, it demonstrates strong generalization ability, while ablation studies on RBPsuite2.0 validate the contributions of individual modules. We assess biological interpretability by extracting candidate binding motifs from the first-layer CNN kernels. Several motifs closely match experimentally validated RBP motifs, confirming the model's capacity to learn biologically meaningful patterns. A downstream case study on YTHDF1 focuses on analyzing interpretable spatial binding patterns, using a large-scale prediction dataset and CLIP-seq peak alignment. The results confirm that the model captures localized binding signals and spatial consistency with experimental annotations. Overall, RMDNet is a robust and interpretable tool for predicting RNA-protein binding sites. It has broad potential in disease mechanism research and therapeutic target discovery. The source code is available https://github.com/cskyan/RMDNet . Jiangbo Zhang, Yunhui Peng, Feifei Cui, Shankai Yan, Qingchen Zhang 0001 |
BMC Bioinform. | 5 |
| 2024 | Augmented Mycobiome-Based Cancer Detection by an Interpretable Large ModelabstractThe microbiome has emerged as a promising predictor of human cancers. Fungi are important components of the human microbiome and are closely associated with cancer. However, our understanding of the function and efficacy of fungal cells in tumors remains limited, making it challenging to accurately detect cancer tissues using limited tumor mycobiome data. Transfer learning has recently revolutionized the bioinformatics field by leveraging deep learning models pre-trained on large-scale general datasets, effectively addressing the predicted issue in tasks with limited data. Here, we propose an interpretable large model, MCaPred, that encodes different tumor tissue sites with microbiome features. The model was pre-trained on large-scale microbiome data from different tumor tissue sites to predict cancer in limited mycobiome data samples, thereby exploring the association between fungi and cancer-type specificity. Studies have shown that MCaPred outperforms recent deep learning models in cancer screening based on tumor mycobiome data and that mycobiome-based cancer detection is augmented by pretrained models. More importantly, MCaPred can capture key factors from mycobiome data related to specific cancer occurrences because of its high interpretability. The proposed model, analysis code, and supplementary materials used in this study are available at GitHub (https://github.com/cskyan/MCaPred.git). Contact: [email protected] or [email protected] Dongmei He, Xin Yang 0037, Buchao Zhan, Qingchen Zhang 0001, Shankai Yan |
BIBM | 6 |
| 2024 | RARoK: Retrieval-Augmented Reasoning on Knowledge for Medical Question AnsweringabstractAlthough large language models (LLMs) perform impressively in natural language tasks, they face several challenges, such as conventional new knowledge, generating accurate responses, and explaining their reasoning. To address these issues, we propose a new approach, Retrieval-Augmented Reasoning on Knowledge (RARoK), which combines Chain of Thought (CoT) prompts with Retrieval-Augmented Generation (RAG). By leveraging external information from knowledge graphs (KGs), RARoK iteratively refines the CoT to further optimize the reasoning process of the model. Our approach significantly outperformed the baseline in various Q&A tasks, especially in the medical domain. Compared to the SOTA method, RARoK improves Hits@1 by 3.3% on the WebQSP dataset, while improving the Hits@1 and F1 scores by 18.6% and 10.4% respectively on the CWQ dataset. The experimental results on the Knowledge Graph Question Answering (KGQA) datasets and the Medical Q&A datasets show that our method has better reasoning ability and interpretability compared to the vanilla LLMs and other retrieval-enhanced methods. Our code and data are publicly available (https://github.com/cskyan/RARoK). Buchao Zhan, Xin Yang 0037, Dongmei He, Yucong Duan, Shankai Yan |
BIBM | 6 |
| 2024 | Text2SPARQL: Grammar Pre-training for Text-to-QDMR Semantic Parsers from Intermediate Question Decompositions
Buchao Zhan, Yucong Duan, Xin Yang 0037, Dongmei He, Shankai Yan |
ICONIP (9) | 5 |
| 2023 | AIPPT: Predicts anti-inflammatory peptides using the most characteristic subset of bases and sequences by stacking ensemble learning strategiesabstractTherapeutic peptides play a vital role in developing peptide-based drugs. Recently, they have been applied as anti-inflammatory agents for a range of inflammatory conditions, including Alzheimer’s disease and rheumatoid arthritis. Laboratory-based identification of peptides with anti-inflammatory properties is a highly time-consuming and labor-intensive endeavor. To tackle this issue, researchers have developed computational methods, primarily centered on machine learning, to streamline the procedure. This paper presents AIPPT, an intelligent and computationally efficient prediction tool that introduces a novel stacking framework for the reliable identification of anti-inflammatory peptides (AIP). The study specifically employs a combination of four feature encodings, where their importance is assessed using the LightGBM method to create an optimal feature subset, which is then input to the three classifiers. The output probabilities from the three classifiers are further fed into a meta-classifier, constructing a two-layer stacking model. Subsequently, the output probabilities from the three classifiers are incorporated into a meta-classifier, establishing a two-layer stacking model. Subsequently, the output probabilities from the three classifiers are incorporated into a meta-classifier, establishing a two-layer stacking model. Xiuhao Fu, Shankai Yan, Feifei Cui |
BIBM | 4 |
| 2022 | PhenoGene: Disease-gene prioritization using graph embedding on patient phenotypic profiles
Shankai Yan, Ling Luo 0001, Daniel Veltri, Andrew J. Oler, Rajarshi Ghosh, Chih-Hsuan Wei, Morgan Similuk, Zhiyong Lu |
AMIA | 1 |
| 2022 | PhenoRerank: A re-ranking model for phenotypic concept recognition pre-trained on human phenotype ontologyabstractThe study aims at developing a neural network model to improve the performance of Human Phenotype Ontology (HPO) concept recognition tools. We used the terms, definitions, and comments about the phenotypic concepts in the HPO database to train our model. The document to be analyzed is first split into sentences and annotated with a base method to generate candidate concepts. The sentences, along with the candidate concepts, are then fed into the pre-trained model for re-ranking. Our model comprises the pre-trained BlueBERT and a feature selection module, followed by a contrastive loss. We re-ranked the results generated by three robust HPO annotation tools and compared the performance against most of the existing approaches. The experimental results show that our model can improve the performance of the existing methods. Significantly, it boosted 3.0% and 5.6% in F1 score on the two evaluated datasets compared with the base methods. It removed more than 80% of the false positives predicted by the base methods, resulting in up to 18% improvement in precision. Our model utilizes the descriptive data in the ontology and the contextual information in the sentences for re-ranking. The results indicate that the additional information and the re-ranking model can significantly enhance the precision of HPO concept recognition compared with the base method. Shankai Yan, Ling Luo 0001, Po-Ting Lai, Daniel Veltri, Andrew J. Oler, Sandhya Xirasagar, Rajarshi Ghosh, Morgan Similuk, Peter N. Robinson, Zhiyong Lu |
J. Biomed. Informatics | 1 |
| 2021 | PhenoTagger: a hybrid method for phenotype concept recognition using human phenotype ontologyabstractMOTIVATION: Automatic phenotype concept recognition from unstructured text remains a challenging task in biomedical text mining research. Previous works that address the task typically use dictionary-based matching methods, which can achieve high precision but suffer from lower recall. Recently, machine learning-based methods have been proposed to identify biomedical concepts, which can recognize more unseen concept synonyms by automatic feature learning. However, most methods require large corpora of manually annotated data for model training, which is difficult to obtain due to the high cost of human annotation. RESULTS: In this article, we propose PhenoTagger, a hybrid method that combines both dictionary and machine learning-based methods to recognize Human Phenotype Ontology (HPO) concepts in unstructured biomedical text. We first use all concepts and synonyms in HPO to construct a dictionary, which is then used to automatically build a distantly supervised training dataset for machine learning. Next, a cutting-edge deep learning model is trained to classify each candidate phrase (n-gram from input sentence) into a corresponding concept label. Finally, the dictionary and machine learning-based prediction results are combined for improved performance. Our method is validated with two HPO corpora, and the results show that PhenoTagger compares favorably to previous methods. In addition, to demonstrate the generalizability of our method, we retrained PhenoTagger using the disease ontology MEDIC for disease concept recognition to investigate the effect of training on different ontologies. Experimental results on the NCBI disease corpus show that PhenoTagger without requiring manually annotated training data achieves competitive performance as compared with state-of-the-art supervised methods. AVAILABILITYAND IMPLEMENTATION: The source code, API information and data for PhenoTagger are freely available at https://github.com/ncbi-nlp/PhenoTagger. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ling Luo 0001, Shankai Yan, Po-Ting Lai, Daniel Veltri, Andrew J. Oler, Sandhya Xirasagar, Rajarshi Ghosh, Morgan Similuk, Peter N. Robinson, Zhiyong Lu |
Bioinform. | 2 |
| 2021 | Future DNA computing device and accompanied tool stack: Towards high-throughput computation
Shankai Yan, Ka-Chun Wong |
Future Gener. Comput. Syst. | 1 |
| 2020 | Context awareness and embedding for biomedical event extractionabstractMOTIVATION: Biomedical event extraction is fundamental for information extraction in molecular biology and biomedical research. The detected events form the central basis for comprehensive biomedical knowledge fusion, facilitating the digestion of massive information influx from the literature. Limited by the event context, the existing event detection models are mostly applicable for a single task. A general and scalable computational model is desiderated for biomedical knowledge management. RESULTS: We consider and propose a bottom-up detection framework to identify the events from recognized arguments. To capture the relations between the arguments, we trained a bidirectional long short-term memory network to model their context embedding. Leveraging the compositional attributes, we further derived the candidate samples for training event classifiers. We built our models on the datasets from BioNLP Shared Task for evaluations. Our method achieved the average F-scores of 0.81 and 0.92 on BioNLPST-BGI and BioNLPST-BB datasets, respectively. Comparing with seven state-of-the-art methods, our method nearly doubled the existing F-score performance (0.92 versus 0.56) on the BioNLPST-BB dataset. Case studies were conducted to reveal the underlying reasons. AVAILABILITY AND IMPLEMENTATION: https://github.com/cskyan/evntextrc. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Shankai Yan, Ka-Chun Wong |
Bioinform. | 1 |
| 2020 | Verbal aggression detection on Twitter comments: convolutional neural network for short-text sentiment analysis
Shankai Yan, Ka-Chun Wong |
Neural Comput. Appl. | 2 |
| 2020 | BioConceptVec: Creating and evaluating literature-based biomedical concept embeddings on a large scaleabstractA massive number of biological entities, such as genes and mutations, are mentioned in the biomedical literature. The capturing of the semantic relatedness of biological entities is vital to many biological applications, such as protein-protein interaction prediction and literature-based discovery. Concept embeddings-which involve the learning of vector representations of concepts using machine learning models-have been employed to capture the semantics of concepts. To develop concept embeddings, named-entity recognition (NER) tools are first used to identify and normalize concepts from the literature, and then different machine learning models are used to train the embeddings. Despite multiple attempts, existing biomedical concept embeddings generally suffer from suboptimal NER tools, small-scale evaluation, and limited availability. In response, we employed high-performance machine learning-based NER tools for concept recognition and trained our concept embeddings, BioConceptVec, via four different machine learning models on ~30 million PubMed abstracts. BioConceptVec covers over 400,000 biomedical concepts mentioned in the literature and is of the largest among the publicly available biomedical concept embeddings to date. To evaluate the validity and utility of BioConceptVec, we respectively performed two intrinsic evaluations (identifying related concepts based on drug-gene and gene-gene interactions) and two extrinsic evaluations (protein-protein interaction prediction and drug-drug interaction extraction), collectively using over 25 million instances from nine independent datasets (17 million instances from six intrinsic evaluation tasks and 8 million instances from three extrinsic evaluation tasks), which is, by far, the most comprehensive to our best knowledge. The intrinsic evaluation results demonstrate that BioConceptVec consistently has, by a large margin, better performance than existing concept embeddings in identifying similar and related concepts. More importantly, the extrinsic evaluation results demonstrate that using BioConceptVec with advanced deep learning models can significantly improve performance in downstream bioinformatics studies and biomedical text-mining applications. Our BioConceptVec embeddings and benchmarking datasets are publicly available at https://github.com/ncbi-nlp/BioConceptVec. Qingyu Chen 0001, Kyubum Lee, Shankai Yan, Sun Kim, Chih-Hsuan Wei, Zhiyong Lu |
PLoS Comput. Biol. | 3 |
| 2020 | Deleterious Non-Synonymous Single Nucleotide Polymorphism Predictions on Human Transcription FactorsabstractTranscription factors (TFs) are the major components of human gene regulation. In particular, they bind onto specific DNA sequences and regulate neighborhood genes in different tissues at different developmental stages. Non-synonymous single nucleotide polymorphisms on its protein-coding sequences could result in undesired consequences in human. Therefore, it is necessary to develop methods for predicting any abnormality among those non-synonymous single nucleotide polymorphisms. To address it, we have developed and compared different strategies to predict deleterious non-synonymous single nucleotide polymorphisms (also known as missense mutations) on the protein-coding sequences of human TFs. Taking advantage of evolutionary conservation signals, we have developed and compared different classifiers with different feature sets as computed from different evolutionarily related sequence collections. The results indicate that the classic ensemble algorithm, Adaboost with decision stumps, with orthologous sequence collection, has performed the best (namely, TFmedic). We have further compared TFmedic with other state-of-the-arts methods (i.e., PolyPhen-2 and SIFT) on PolyPhen-2's own datasets, demonstrating that TFmedic can outperform the others. As applications, we have further applied TFmedic to all possible missense mutations on all human transcription factors; the proteome-wide results reveal interesting insights, consistent with the existing physiochemical knowledge. A case study with the actual 3D structure is conducted, revealing how TFmedic can be contributed to protein-DNA binding complex studies. Ka-Chun Wong, Shankai Yan, Qiuzhen Lin, Xiangtao Li, Chengbin Peng 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2020 | GESgnExt: Gene Expression Signature Extraction and Meta-Analysis on Gene Expression OmnibusabstractThe gene expression omnibus (GEO) repository harbours an exponentially increasing number of gene expression studies. The expression data, as well as the related metadata, provides an abundant resource for knowledge discovery. Each study in GEO focuses on the gene expression perturbation of a specific subject (e.g., gene, drug, and disease). The identification of those subjects and the associations among them are beneficial for further in-depth studies. However, they cannot be directly inferred from the studies. A unified representation of those subjects (i.e., gene expression signatures) is desired. We developed GESgnExt for the automatic construction of gene expression signatures. The resultant 6542 signatures are built on 1934 series and 35 919 samples from GEO. To evaluate its significance, we calculated the similarities among those signatures and compared the discovered associations against the existing interaction databases. The signatures connect the genes, drugs, and diseases, covering most of the experimentally validated interactions. Besides, we have discovered 3307 novel signatures and their related associations, complementing the existing signature knowledge. The biomedical relevance of GESgnExt is demonstrated further in multiple case studies, providing mechanistic insights into its knowledge discovery process. Shankai Yan, Ka-Chun Wong |
IEEE J. Biomed. Health Informatics | 1 |
| 2017 | Elucidating high-dimensional cancer hallmark annotation via enriched ontology
Shankai Yan, Ka-Chun Wong |
J. Biomed. Informatics | 1 |