EDBT 2026 Demo / reviewers in the wild / expert
Long-Chen Shen
dblp:301/9850
· DBLP profile ↗
9ranked-venue papers
2as first author
9since 2021 · last 2026
0000-0002-0045-4745ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DrugDL: dual-modal deep learning framework for multi-property drug prediction and targeted therapy discoveryabstractMOTIVATION: The accurate and robust representation of drug molecule features, the prediction of drug-target biomacromolecule interactions, and the determination of physicochemical properties are crucial in drug development. However, these tasks remain challenging due to issues such as the limited generalizability of single-modal representations, the absence of multitask prediction frameworks, and weak adaptability in cold-start scenarios. RESULTS: In this study, we present DrugDL, a framework for comprehensive drug molecule representation and the prediction of multiple downstream tasks, including drug-target interactions, binding affinities, binding sites, physicochemical properties, toxicity, and drug-drug interactions. DrugDL jointly learns representations of the drug chemical space and the target protein biological space, while capturing multiscale interaction mechanisms between drug molecules and target proteins through the integration of cross-modal contrastive learning and single-modal feature enhancement algorithms. Specifically, DrugDL employs a multitask prediction framework to predict multiple properties of drug molecules. In practical applications, it consistently outperforms state-of-the-art methods, particularly in cold-start tasks. The framework has been successfully applied to high-throughput screening, the identification of inhibitors of SARS-CoV-2 and metabolic enzymes, and the prediction of cancer-targeted drugs. Experimental validations on EGFR and ALK targets further demonstrate its effectiveness as a precise drug discovery tool. By enabling accurate molecular representation and multi-property prediction, DrugDL provides end-to-end technical support for drug development, thereby significantly accelerating the drug discovery process. AVAILABILITY AND IMPLEMENTATION: The datasets and code are available at https://github.com/ZhangQi9910/DrugDL. The version of record is archived in Zenodo with the DOI: 10.5281/zenodo.20579718. Yuxiao Wei, Yunpeng Xia, Long-Chen Shen, Hong-Bin Shen, Dongjun Yu |
Bioinform. | 5 |
| 2026 | CellPredX, a computational framework for cross-data type, cross-sample, and cross-protocol cell type annotation through domain adaptation and deep metric learningabstractAccurate cell type annotation is fundamental to single-cell analysis, yet remains challenging across heterogeneous datasets and modalities. In particular, transferring labels between scRNA-seq and scATAC-seq data poses unique difficulties due to discrepancies in sequencing protocols and feature spaces. Existing methods typically handle only a subset of these challenges, often requiring scenario-specific adjustments and offering limited interpretability. Here, we present CellPredX, a structurally unified but adaptively parameterized, semi-supervised cross-modality framework for label transfer across scRNA-seq, scATAC-seq, and cross-protocol datasets. While maintaining a unified model architecture and optimization strategy, CellPredX allows adaptive tuning of loss-weight hyperparameters to account for the varying degree of similarity or discrepancy between different reference-query dataset pairs. CellPredX integrates domain adaptation and deep metric learning to align heterogeneous embeddings, and introduces a sparse center loss with an attention mechanism to enhance discriminative representations while suppressing noise. Moreover, an integrated interpreter module based on gradient attribution enables biological interpretability by identifying key markers and feature dimensions driving model predictions. Through extensive benchmarking across scRNA to scATAC, scATAC to scATAC, and scRNA to scRNA transfers, CellPredX consistently outperforms state-of-the-art annotation methods in both accuracy and robustness. The interpreter module further reveals biologically meaningful marker patterns that are consistent with known cell hierarchies. Together, these results demonstrate that CellPredX provides an interpretable and scalable solution for cross-modality cell type annotation in single-cell multi-omic integration. Yan Liu 0038, Long-Chen Shen, Jipeng Qiang |
PLoS Comput. Biol. | 4 |
| 2025 | Supervised contrastive learning enhances MHC-II peptide binding affinity prediction
Long-Chen Shen, Yan Liu 0038, Zi Liu, Zhikang Wang, Yuming Guo 0001, Jamie Rossjohn, Jiangning Song, Dongjun Yu |
Expert Syst. Appl. | 1 |
| 2025 | MUSIC-GCN: A Novel Multi-Tasking Pipeline for Analyzing Single-Cell Transcriptomic Data Using Residual Graph Convolution NetworkabstractSingle-cell transcriptomics is a powerful approach for characterizing gene transcription at cellular resolution. This approach requires efficient computational pipelines to undertake essential tasks, including clustering, dimensionality reduction, imputation, and denoising. Currently, most such pipelines undertake these computational tasks separately without considering the interdependence among these tasks. Here, we present an advanced pipeline, MUSIC-GCN, by employing a graph convolutional neural (GCN) network and autoencoder to perform multi-task single-cell RNA-sequencing (scRNA-seq) data analysis. The rationale is that multiple related tasks can be carried out simultaneously to enable enhanced learning and more effective representations through the 'sharing of knowledge' regarding individual tasks. Benchmarking experiments using various scRNA-seq datasets show that MUSIC-GCN can achieve a competitive performance on multi-tasks when benchmarked with state-of-the-art approaches. Yan Liu 0038, Chen Li 0021, Long-Chen Shen, Robin B. Gasser, Jiangning Song, Dijun Chen, Dongjun Yu |
IEEE Trans. Comput. Biol. Bioinform. | 6 |
| 2024 | GMFGRN: a matrix factorization and graph neural network approach for gene regulatory network inferenceabstractThe recent advances of single-cell RNA sequencing (scRNA-seq) have enabled reliable profiling of gene expression at the single-cell level, providing opportunities for accurate inference of gene regulatory networks (GRNs) on scRNA-seq data. Most methods for inferring GRNs suffer from the inability to eliminate transitive interactions or necessitate expensive computational resources. To address these, we present a novel method, termed GMFGRN, for accurate graph neural network (GNN)-based GRN inference from scRNA-seq data. GMFGRN employs GNN for matrix factorization and learns representative embeddings for genes. For transcription factor-gene pairs, it utilizes the learned embeddings to determine whether they interact with each other. The extensive suite of benchmarking experiments encompassing eight static scRNA-seq datasets alongside several state-of-the-art methods demonstrated mean improvements of 1.9 and 2.5% over the runner-up in area under the receiver operating characteristic curve (AUROC) and area under the precision-recall curve (AUPRC). In addition, across four time-series datasets, maximum enhancements of 2.4 and 1.3% in AUROC and AUPRC were observed in comparison to the runner-up. Moreover, GMFGRN requires significantly less training time and memory consumption, with time and memory consumed <10% compared to the second-best method. These findings underscore the substantial potential of GMFGRN in the inference of GRNs. It is publicly available at https://github.com/Lishuoyy/GMFGRN. Yan Liu 0038, Long-Chen Shen, Jiangning Song, Dongjun Yu |
Briefings Bioinform. | 3 |
| 2023 | TripletCell: a deep metric learning framework for accurate annotation of cell types at the single-cell levelabstractSingle-cell RNA sequencing (scRNA-seq) has significantly accelerated the experimental characterization of distinct cell lineages and types in complex tissues and organisms. Cell-type annotation is of great importance in most of the scRNA-seq analysis pipelines. However, manual cell-type annotation heavily relies on the quality of scRNA-seq data and marker genes, and therefore can be laborious and time-consuming. Furthermore, the heterogeneity of scRNA-seq datasets poses another challenge for accurate cell-type annotation, such as the batch effect induced by different scRNA-seq protocols and samples. To overcome these limitations, here we propose a novel pipeline, termed TripletCell, for cross-species, cross-protocol and cross-sample cell-type annotation. We developed a cell embedding and dimension-reduction module for the feature extraction (FE) in TripletCell, namely TripletCell-FE, to leverage the deep metric learning-based algorithm for the relationships between the reference gene expression matrix and the query cells. Our experimental studies on 21 datasets (covering nine scRNA-seq protocols, two species and three tissues) demonstrate that TripletCell outperformed state-of-the-art approaches for cell-type annotation. More importantly, regardless of protocols or species, TripletCell can deliver outstanding and robust performance in annotating different types of cells. TripletCell is freely available at https://github.com/liuyan3056/TripletCell. We believe that TripletCell is a reliable computational tool for accurately annotating various cell types using scRNA-seq data and will be instrumental in assisting the generation of novel biological hypotheses in cell biology. Yan Liu 0038, Chen Li 0021, Long-Chen Shen, Robin B. Gasser, Jiangning Song, Dijun Chen, Dongjun Yu |
Briefings Bioinform. | 4 |
| 2022 | MAResNet: predicting transcription factor binding sites by combining multi-scale bottom-up and top-down attention and residual networkabstractAccurate identification of transcription factor binding sites is of great significance in understanding gene expression, biological development and drug design. Although a variety of methods based on deep-learning models and large-scale data have been developed to predict transcription factor binding sites in DNA sequences, there is room for further improvement in prediction performance. In addition, effective interpretation of deep-learning models is greatly desirable. Here we present MAResNet, a new deep-learning method, for predicting transcription factor binding sites on 690 ChIP-seq datasets. More specifically, MAResNet combines the bottom-up and top-down attention mechanisms and a state-of-the-art feed-forward network (ResNet), which is constructed by stacking attention modules that generate attention-aware features. In particular, the multi-scale attention mechanism is utilized at the first stage to extract rich and representative sequence features. We further discuss the attention-aware features learned from different attention modules in accordance with the changes as the layers go deeper. The features learned by MAResNet are also visualized through the TMAP tool to illustrate that the method can extract the unique characteristics of transcription factor binding sites. The performance of MAResNet is extensively tested on 690 test subsets with an average AUC of 0.927, which is higher than that of the current state-of-the-art methods. Overall, this study provides a new and useful framework for the prediction of transcription factor binding sites by combining the funnel attention modules with the residual network. Long-Chen Shen, Yiheng Zhu 0001, Jian Xu 0009, Jiangning Song, Dongjun Yu |
Briefings Bioinform. | 2 |
| 2021 | Improving protein fold recognition using triplet network and ensemble deep learningabstractProtein fold recognition is a critical step toward protein structure and function prediction, aiming at providing the most likely fold type of the query protein. In recent years, the development of deep learning (DL) technique has led to massive advances in this important field, and accordingly, the sensitivity of protein fold recognition has been dramatically improved. Most DL-based methods take an intermediate bottleneck layer as the feature representation of proteins with new fold types. However, this strategy is indirect, inefficient and conditional on the hypothesis that the bottleneck layer's representation is assumed as a good representation of proteins with new fold types. To address the above problem, in this work, we develop a new computational framework by combining triplet network and ensemble DL. We first train a DL-based model, termed FoldNet, which employs triplet loss to train the deep convolutional network. FoldNet directly optimizes the protein fold embedding itself, making the proteins with the same fold types be closer to each other than those with different fold types in the new protein embedding space. Subsequently, using the trained FoldNet, we implement a new residue-residue contact-assisted predictor, termed FoldTR, which improves protein fold recognition. Furthermore, we propose a new ensemble DL method, termed FSD_XGBoost, which combines protein fold embedding with the other two discriminative fold-specific features extracted by two DL-based methods SSAfold and DeepFR. The Top 1 sensitivity of FSD_XGBoost increases to 74.8% at the fold level, which is ~9% higher than that of the state-of-the-art method. Together, the results suggest that fold-specific features extracted by different DL methods complement with each other, and their combination can further improve fold recognition at the fold level. The implemented web server of FoldTR and benchmark datasets are publicly available at http://csbio.njust.edu.cn/bioinf/foldtr/. Yan Liu 0038, Yiheng Zhu 0001, Ying Zhang 0053, Long-Chen Shen, Jiangning Song, Dongjun Yu |
Briefings Bioinform. | 5 |
| 2021 | SAResNet: self-attention residual network for predicting DNA-protein bindingabstractKnowledge of the specificity of DNA-protein binding is crucial for understanding the mechanisms of gene expression, regulation and gene therapy. In recent years, deep-learning-based methods for predicting DNA-protein binding from sequence data have achieved significant success. Nevertheless, the current state-of-the-art computational methods have some drawbacks associated with the use of limited datasets with insufficient experimental data. To address this, we propose a novel transfer learning-based method, termed SAResNet, which combines the self-attention mechanism and residual network structure. More specifically, the attention-driven module captures the position information of the sequence, while the residual network structure guarantees that the high-level features of the binding site can be extracted. Meanwhile, the pre-training strategy used by SAResNet improves the learning ability of the network and accelerates the convergence speed of the network during transfer learning. The performance of SAResNet is extensively tested on 690 datasets from the ChIP-seq experiments with an average AUC of 92.0%, which is 4.4% higher than that of the best state-of-the-art method currently available. When tested on smaller datasets, the predictive performance is more clearly improved. Overall, we demonstrate that the superior performance of DNA-protein binding prediction on DNA sequences can be achieved by combining the attention mechanism and residual structure, and a novel pipeline is accordingly developed. The proposed methodology is generally applicable and can be used to address any other sequence classification problems. Long-Chen Shen, Yan Liu 0038, Jiangning Song, Dongjun Yu |
Briefings Bioinform. | 1 |