EDBT 2026 Demo / reviewers in the wild / expert
Feifei Cui
dblp:287/3312
· DBLP profile ↗
37ranked-venue papers
0as first author
37since 2021 · last 2027
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 27 · 27 since 2021Artificial intelligence and machine learning · 10 · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | SMGC: Spatial multi-omics analysis with granular-ball contrastive learning framework
Xuejing Ma, Zijia Bai, Yajie Meng, Pan Zeng, Xianfang Tang, Feifei Cui, Peng Wang 0035, Jialiang Yang, Junlin Xu |
Expert Syst. Appl. | 8 |
| 2026 | TLAGC: Taylor Linear Attention-Guided Graph Convolutions for Revealing Spatial Domains in Spatial Multi-Omics DataabstractWith the rapid advance of spatial multi-omics technologies, it has become possible to simultaneously profile transcripts, proteins and chromatin states at their native spatial coordinates, thereby uncovering molecular architecture that transcends any single-omics perspective. However, the resulting data matrices are often highly sparse and suffer from unstable dimensionality. Graph-based neural methods capture only local neighborhood information, whereas conventional Transformers, although capable of modelling long-range dependencies, incur prohibitive computational costs on such data. To overcome these limitations, we propose TLAGC—a Taylor-Linear-Attention-Guided Graph Convolutional framework that couples a Taylor-expanded linear attention (TLA) mechanism with graph convolutional networks. By eliminating the soft-max operation and linking the LocalGCN via residual connections, TLA preserves local structural information while enabling the integration of global and local contexts, thereby alleviating ineffective information propagation between spatially distant yet transcriptionally similar regions. Theoretical analysis confirms that TLA indeed reduces computational complexity, and extensive experiments on multiple spatial multi-omics benchmarks demonstrate that TLAGC consistently outperforms state-of-the-art baselines in delineating spatial domains. Aoyun Geng, Chunyan Cui, Yunyun Su, Zhenjie Luo, Feifei Cui |
AAAI | 5 |
| 2026 | Generalizable Drug-Target Interaction Prediction via ESM-2 Representations and Progressive Contrastive Curriculum LearningabstractPredicting drug–target interactions (DTIs) is a fundamental task in computational drug discovery, yet it remains challenging under distribution shifts and limited training data. Existing approaches often suffer from poor generalization, weak cross-modal alignment between molecular and protein representations, and vulnerability to noisy supervision.We propose ESP-DTI, a unified framework designed to enhance generalization by integrating large-scale protein language models with curriculum learning and cross-modal contrastive alignment. Specifically, we leverage ESM-2 to encode context-aware protein representations and adopt a CLIP-style contrastive objective to align drug and protein embeddings in a shared latent space. To further improve learning robustness, we introduce a progressive curriculum sampling strategy that dynamically schedules training instances based on model confidence, enabling a gradual shift from easy to hard examples.Experimental results on four benchmark datasets demonstrate that ESP-DTI consistently outperforms state-of-the-art baselines, achieving a +3.1% improvement in average accuracy. Ablation studies confirm the complementary benefits of each component, validating their collective contribution to robust and generalizable DTI prediction.Our work underscores the effectiveness of combining pretrained protein language models with structured training curricula and cross-modal contrastive learning for reliable DTI prediction under real-world, distribution-shifted conditions. Qianyang Wu, Jingwei Lv, Feifei Cui |
AAAI | 4 |
| 2026 | EccoMamba: Enhanced Cross-hierarchical Continuity Orthogonal Mamba for Medical Image SegmentationabstractMedical image segmentation plays a crucial role in clinical diagnosis, lesion quantification, and preoperative planning. However, existing Mamba-based architectures, which rely on fixed-direction sequence modeling and flatten images into one-dimensional (1D) sequences, struggle to capture hierarchical anatomical features and spatial dependencies, thereby limiting their representational capacity for complex medical structures. To address these limitations, we propose EccoMamba (Enhanced Cross-hierarchical Continuity Orthogonal Mamba), a U-shaped encoder--decoder framework designed for medical image segmentation. In the encoder's downsampling path, we introduce a Hierarchical Aggregation Enhancement (HAE) module that integrates multi-scale convolutions with hierarchical attention mechanisms. The attention branch further incorporates cross-channel interactions, allowing the model to selectively enhance semantically relevant features while suppressing irrelevant background responses. For skip connections, we design a Structural Continuity Orthogonal (SCO) module to preserve spatial continuity by modeling cross-dimensional dependencies via orthogonal Axial Shifts (AS), thereby mitigating directional bias and improving anatomical consistency. Extensive experiments on four benchmark datasets---ISIC 2018, ISIC 2017, Synapse, and ACDC---show that EccoMamba consistently outperforms state-of-the-art methods in both segmentation accuracy and structural fidelity. Junlin Xu, Jincan Li, Feifei Cui, Jialiang Yang, Shuting Jin, Qiangguo Jin, Yajie Meng |
AAAI | 3 |
| 2026 | A deep adversarial network model for multi-task analysis of single-cell omics dataabstractSingle-cell multi-omics data reveal complex cellular states and deepen our understanding of tissue cell phenotypes and functions. However, data analysis remains challenging due to the discrete nature and high noise level of the data, as well as the lack of modality. Here, we propose scMultiNet, a multi-task deep adversarial neural network that can integrate different tasks to analyze single-cell multi-modal data. In particular, we achieve joint training of multi-modal integration and cross-modal prediction tasks by introducing a cross-modal bi-prediction module and a multi-head self-attention module. Data denoising is further enhanced by integrating an indicator matrix that constrains and precisely reconstructs the original expression values. Extensive simulations and real data experiments demonstrate that scMultiNet outperforms existing state-of-the-art methods in dimensionality reduction, visualization, clustering, batch elimination, data denoising, multi-modal integration, single-cell cross-modality translation, and in revealing cell type-specific biological insights. In addition, we demonstrate that scMultiNet can effectively transfer the complex relationships between modalities from one batch to another. In summary, scMultiNet stands as a comprehensive end-to-end framework, ideally suited for analyzing single-cell multi-omics data. Junlin Xu, Yajie Meng, Shuting Jin, Changcheng Lu, Feifei Cui, Xiangzheng Fu, Quan Zou 0001, Xiangxiang Zeng |
Briefings Bioinform. | 7 |
| 2026 | AdaptIPs: A dual-channel deep learning framework integrating protein language model representations and transfer learning for phosphorylation-site prediction
Aoyun Geng, Yanfei Qu, Junlin Xu, Yajie Meng, Quan Zou 0001, Feifei Cui |
Eng. Appl. Artif. Intell. | 8 |
| 2026 | SpatialSyn: A synergistic graph framework for spatial domain identification in multi-omics
Pan Zeng, Runzhi Li, Yajie Meng, Feifei Cui, Xianfang Tang, Jialiang Yang, Junlin Xu |
Expert Syst. Appl. | 5 |
| 2026 | MMTF-DTI: Drug-target interaction prediction via multimodal feature extraction and dynamic fusion
Pan Zeng, Xianfang Tang, Yajie Meng, Feifei Cui, Junlin Xu |
J. Biomed. Informatics | 6 |
| 2026 | CNNCaps-DBP: Leveraging protein language models with attention-augmented convolution for DNA-binding protein prediction
Ziyuan Yan, Aoyun Geng, Yazi Li, Jiajing Wang, Junlin Xu, Yajie Meng, Leyi Wei, Quan Zou 0001, Feifei Cui |
Neural Networks | 10 |
| 2026 | MFDL-DDI: An effective deep learning-based framework for predicting drug-drug interactions through multimodal information fusion
Yazi Li, Shuting Jin, Junlin Xu, Yajie Meng, Leyi Wei, Xin Gao 0001, Feifei Cui |
Pattern Recognit. | 9 |
| 2026 | DeepNhKcr: Explainable Deep Learning Framework for the Prediction of Crotonylation Sites of Non-Histone Lysine in Plants Based on Pre-Trained Protein Language ModelabstractLysine crotonylation (Kcr) is an important protein modification occurring after translation in biology, serving an essential function in a range of biological processes in both plants and animals, including the regulation of gene expression, the maintenance of cellular metabolic balance, and the enhancement of photosynthesis. Exploring the detection of Kcr sites is essential for uncovering their biological functions. Nonetheless, conventional experimental approaches for detection are often time-consuming, expensive, and hindered by various technical constraints, making the precise identification of Kcr sites a significant challenge. This study seeks to develop a computational approach for the rapid and accurate prediction of Kcr sites in plant non-histone proteins. We introduce a novel deep learning framework named DeepNhKcr, which integrates the protein language model (ESM2) with a bidirectional long short-term memory (BiLSTM) network. To address the challenge of data imbalance, the model replaces the conventional cross-entropy loss with the focal loss function. In addition, DeepNhKcr combines advanced deep learning approaches with traditional protein encoding strategies to enable effective feature extraction and integration. This method not only significantly boosts the accuracy of predicting Kcr sites in non-histone proteins of plants. but also provides interpretability, shedding light on the potential links between key sequence characteristics and their biological roles. DeepNhKcr delivers outstanding results, surpassing existing machine learning and deep learning models, and demonstrating excellent performance in both five-fold cross-validation and independent test experiments. Moreover, the model integrates interpretability analysis techniques to investigate the connections between important sequence features and their biological roles. DeepNhKcr acts as a powerful method for detecting Kcr sites in plant non-histone proteins and is anticipated to greatly advance future studies in plant Kcr site prediction. Zhenjie Luo, Aoyun Geng, Junlin Xu, Yajie Meng, Shankai Yan, Leyi Wei, Qingchen Zhang 0001, Quan Zou 0001, Feifei Cui |
IEEE Trans. Comput. Biol. Bioinform. | 11 |
| 2026 | A Multi-Modal Contrastive Learning Framework for Cyclic Peptide Permeability PredictionabstractCyclic peptides represent a rapidly growing class of therapeutics, yet their development is often hindered by the challenge of predicting cell membrane permeability, a critical determinant of drug efficacy. Existing computational methods often struggle to integrate the diverse structural information inherent in these complex molecules, resulting in suboptimal predictive accuracy. Here, we introduce MCPerm, a multi-modal deep learning framework that synergistically integrates 1D SMILES, 2D topological, and 3D geometric information through a novel modality share and contrastive learning strategy to accurately predict cyclic peptide permeability. MCPerm fine-tunes a pretrained peptide language model for SMILES encoding and uses a parameter-sharing graph transformer for structural representation, while a dual contrastive learning mechanism enforces representational consistency both within and between modalities. On the benchmark PAMPA dataset, MCPerm achieves state-of-the-art performance, significantly outperforming leading methods. We further demonstrate its robustness and competitive transferability across three independent assays (Caco-2, MDCK, and RRCK). Our work presents a robust in silico framework that holds potential to accelerate the rational design and discovery of cell-permeable cyclic peptide drugs. Furthermore, to move beyond predictive accuracy, we introduced an attention-based visualization analysis. The results demonstrate that our model is not a "black box"; it has learned key chemical principles governing cyclic peptide permeability. Shuwen Xiong, Feifei Cui, Rao Zeng, Ran Su, Leyi Wei |
IEEE Trans. Comput. Biol. Bioinform. | 2 |
| 2026 | DeepR2OM: Accurate Recognition for RNA 2′-O-Methylation Sites in Human Genome Using Deep Learningabstract2'-O-methylation (2OM) of ribose is a widespread RNA modification that significantly impacts RNA stability, structure, and function. Accurately predicting 2OM sites is crucial for understanding RNA's biological functions and related pathologies. Traditional detection methods pose challenges such as resource intensiveness, potential RNA sample damage, and high costs. However, recent advancements in machine learning, particularly deep learning techniques, offer rapid and cost-effective prediction solutions. In this study, we introduce DeepR2OM, a novel method integrating feature selection and deep learning for 2OM sites prediction. DeepR2OM encodes sequences using eight RNA descriptors, employs feature selection algorithms to reduce dimensions, and then utilizes a deep learning network for training. After evaluating various deep learning architectures, we selected Convolutional Neural Network (CNN), Multi-Head Self-Attention mechanism, and Deep Neural Network (DNN) as our final prediction models. Experimental results demonstrate DeepR2OM's effectiveness, achieving 87.1% accuracy (ACC), 85.5% recall rate (Recall), 87.9% precision (PRE), and a Matthews correlation coefficient (MCC) of 75.7% on an independent test set. This tool serves as a valuable resource for exploring the functional and bioinformatic aspects of 2OM sites. Shun Gao, Ziyuan Yan, Feifei Cui, Leyi Wei, Qingchen Zhang 0001, Quan Zou 0001 |
IEEE Trans. Comput. Biol. Bioinform. | 4 |
| 2026 | FusionMVSA: Multi-View Fusion Strategy With Self-Attention for Enhancing Drug RecommendationabstractLeveraging the wealth of biomedical data available, we can derive insights into the relationships between biological entities from various angles. This underscores the complexity and significance of developing a dynamic approach for integrating data from multiple sources, a critical endeavor in drug recommendation. In this study, we introduce an innovative deep learning approach termed "Multi-View Fusion Strategy with Self-Attention" (FusionMVSA), designed to predict associations between drugs and diseases. To effectively amalgamate data from diverse sources and extract representative features, we have developed a feature extraction mechanism that capitalizes on similarities. This mechanism computes self-attention across multiple perspectives using shared group parameters, thereby highlighting common characteristics. Simultaneously, we utilize biomedical similarities among multi-source data as guiding factors for calculating similarity, enabling the capture of more nuanced features. Subsequently, we integrate these features through a feature fusion process, where known associations between drugs and diseases act as guiding terms. This strategy allows us to uncover the complementary aspects of different viewpoints. Ultimately, we predict potential drug-disease associations using a multi-layer perceptron neural network. Our methodology has undergone rigorous testing through various cross-validation experiments and case studies. We are confident that FusionMVSA will prove to be a valuable tool in drug recommendation, offering new avenues for exploration and discovery in the quest to combat diseases. Yajie Meng, Xudong Shang, Xianfang Tang, Jincan Li, Feifei Cui, Shuting Jin, Junlin Xu, Peng Wang 0035 |
IEEE J. Biomed. Health Informatics | 7 |
| 2025 | RTCDB: A Multi-Omics Database Supporting Research and Discovery in Rare Tumor OncologyabstractRare tumors are generally faced with the dilemma of delayed diagnosis, lack of treatment methods and poor prognosis due to their high clinical heterogeneity, low incidence but numerous subtypes. Although the analysis of molecular mechanisms and clinical characteristics is crucial to improving the clinical research of rare tumors, current research is limited by the scattered heterogeneity of multi-center data, scarce sample size and lack of standardized integration mechanism, which seriously hinders the process of basic research to clinical transformation. In order to meet such challenges, we are committed to building a rare tumor comprehensive database website. The core goal of this platform is to aggregate and provide high-quality and reliable datasets (including genome, transcriptome, proteome and clinical information) of various rare tumor types, providing a centralized and efficient resource center for researchers around the world, thereby accelerating the basic research and clinical transformation exploration of rare tumors. The database website (RTCDB) contains$\mathbf{1 3 4}$rare tumor types and 306 selected datasets with 7912 samples. The establishment of this rare tumor database will strongly promote the collaborative cooperation and rapid development of related research. RTCDB can be accessed for free at http://www.rtedb.cn/. Yunyun Su, Binlu Yang, Aoyun Geng, Feifei Cui |
BIBM | 6 |
| 2025 | ST-GCP: a graph convolutional network model with contrastive consistency and permutation for spatial transcriptomicsabstractSpatial transcriptomics (STs) technology is a powerful technique that simultaneously preserves gene expression profiles and spatial information, enabling deeper exploration of tissue organization and function. However, many existing computational approaches often rely on labeled ST data and overlook the rich spatial information, resulting in limited representations and suboptimal clustering. In this paper, we propose ST-GCP, a self-supervised graph representation learning framework for ST data, which incorporates a structure-feature perturbation mechanism. First, ST-GCP applies feature-level random permutation of the gene expression matrix and random edge dropout in the spatial neighbor network, creating two complementary augmented graph views of ST data. ST-GCP then employs a two-layer graph convolutional network (GCN) encoder-decoder to extract spatial representations and reconstruct gene expression. Finally, a cosine-similarity-based contrastive objective aligns the view-specific representations, and the overall loss jointly optimizes reconstruction fidelity and contrastive consistency, thereby coupling graph topology with transcriptomic profiles in a shared low-dimensional space. Experimental results on multiple ST datasets demonstrate that ST-GCP can uncover biologically meaningful patterns, such as tumor heterogeneity, brain developmental architecture, and cellular developmental trajectories. Yajie Meng, Xianfang Tang, Feifei Cui, Xiangzheng Fu, Quan Zou 0001, Junlin Xu |
Briefings Bioinform. | 6 |
| 2025 | MCAMEF-BERT: an efficient deep learning method for RNA N7-methylguanosine site prediction via multi-branch feature integrationabstractAccurate identification of N7-methylguanosine (m7G) modification sites plays a critical role in uncovering the regulatory mechanisms of various biological processes, including human development, tumor initiation, and progression. However, existing prediction methods still suffer from limited representational power, redundant feature fusion, insufficient utilization of biological prior knowledge, and poor interpretability. In this study, we propose a novel deep learning model named MCAMEF-BERT. This model adopts a parallel architecture that integrates both a DNABERT-2-based pretrained model branch and multiple traditional feature encoding branches, enabling comprehensive multi-perspective sequence feature extraction. To address the redundancy issue in feature fusion, we introduce a multi-channel attention module. Our model demonstrates superior accuracy and effectiveness on datasets from m7GHub, outperforming other state-of-the-art classifiers. Furthermore, we validate the interpretability of MCAMEF-BERT through in silico saturation mutagenesis experiments, and confirm its robustness in motif recognition. Moreover, its generalization capability is validated across diverse RNA modification site prediction tasks. Junlei Yu, Wenjia Gao, Siqi Chen 0001, Ronglin Lu, Jianbo Qiao, Junru Jin, Leyi Wei, Feifei Cui, Xinbo Jiang, Zhongmin Yan |
Briefings Bioinform. | 10 |
| 2025 | RMDNet: RNA-aware dung beetle optimization-based multi-branch integration network for RNA-protein binding sites predictionabstractRNA-binding proteins (RBPs) play crucial roles in gene regulation. Their dysregulation has been increasingly linked to neurodegenerative diseases, liver cancer, and lung cancer. Although experimental methods like CLIP-seq accurately identify RNA-protein binding sites, they are time-consuming and costly. To address this, we propose RMDNet-a deep learning framework that integrates CNN, CNN-Transformer, and ResNet branches to capture features at multiple sequence scales. These features are fused with structural representations derived from RNA secondary structure graphs. The graphs are processed using a graph neural network with DiffPool. To optimize feature integration, we incorporate an improved dung beetle optimization algorithm, which adaptively assigns fusion weights during inference. Evaluations on the RBP-24 benchmark show that RMDNet outperforms state-of-the-art models including GraphProt, DeepRKE, and DeepDW across multiple metrics. On the RBP-31 dataset, it demonstrates strong generalization ability, while ablation studies on RBPsuite2.0 validate the contributions of individual modules. We assess biological interpretability by extracting candidate binding motifs from the first-layer CNN kernels. Several motifs closely match experimentally validated RBP motifs, confirming the model's capacity to learn biologically meaningful patterns. A downstream case study on YTHDF1 focuses on analyzing interpretable spatial binding patterns, using a large-scale prediction dataset and CLIP-seq peak alignment. The results confirm that the model captures localized binding signals and spatial consistency with experimental annotations. Overall, RMDNet is a robust and interpretable tool for predicting RNA-protein binding sites. It has broad potential in disease mechanism research and therapeutic target discovery. The source code is available https://github.com/cskyan/RMDNet . Jiangbo Zhang, Yunhui Peng, Feifei Cui, Shankai Yan, Qingchen Zhang 0001 |
BMC Bioinform. | 3 |
| 2025 | Enhanced drug recommendation model with graph contrastive based on singular value decomposition
Pan Zeng, Ling You, Bofei Zhang, Yajie Meng, Xianfang Tang, Feifei Cui, Junlin Xu |
Eng. Appl. Artif. Intell. | 7 |
| 2025 | Adaptive debiasing learning for drug repositioning
Yajie Meng, Xinrong Hu, Changcheng Lu, Xianfang Tang, Feifei Cui, Pan Zeng, Yuhua Yao, Jialiang Yang, Junlin Xu |
J. Biomed. Informatics | 6 |
| 2025 | Predicting drug-target interactions based on multivariate information fusion and graph contrast learning
Siying Yang, Ping-An He 0001, Pan Zeng, Yajie Meng, Feifei Cui, Yuhua Yao, Jialiang Yang, Junlin Xu |
J. Biomed. Informatics | 6 |
| 2025 | Taco-DDI: accurate prediction of drug-drug interaction events using graph transformer-based architecture and dynamic co-attention matrices
Jianbo Qiao, Junru Jin, Kefei Li, Wenjia Gao, Feifei Cui, Leyi Wei |
Neural Networks | 7 |
| 2025 | ACP-ESM2: Enhancing Anticancer Peptide Prediction With Pre-Trained Protein Language ModelsabstractAnticancer peptide (ACP) are short peptides with anti-cancer properties that have generated increasing attention in recent years due to their low toxicity, minimal side effects, and their ability to precisely target and kill cancer cells. Traditionally, identifying ACP has relied on experimental methods, which are time-consuming and labor-intensive. While deep learning-based prediction methods have made significant progress, there is still room for improvement in achieving optimal performance. In this study, we present ACP-ESM2, a deep learning framework based on the Evolutionary Scale Modeling 2 (ESM2) pre-trained model, which captures rich evolutionary information from protein sequences. By combining ESM2 with convolutional neural network (CNN) that excels at detecting local patterns, ACP-ESM2 offers a highly accurate tool for ACP prediction. The experimental results indicate that ACP-ESM2 shows significant improvements over best-existing recognition techniques on the Test1 set, with enhancements of 2.3%, 7.2%, 12.6%, and 5% in ACC, SN, SP, and MCC, respectively. Notably, on the Test2 set, ACP-ESM2 achieves an accuracy of 97.6%, showcasing its exceptional robustness. This establishes ACP-ESM2 as an efficient and precise tool for predicting anticancer peptides. Shun Gao, Xingfeng Li 0008, Feifei Cui, Qingchen Zhang 0001, Quan Zou 0001 |
IEEE Trans. Comput. Biol. Bioinform. | 4 |
| 2025 | DCMWAF-Net: Dual Cross-Modal Weighted Attention Feature Fusion Network for Multiscale Lunar Crater DetectionabstractLunar craters are critical for studying the Moon’s geological evolution and impact history, making their efficient identification vital for planetary science. To address the limitations of traditional methods in detecting craters across complex terrains and multiple scales, this paper proposes a novel Dual Cross-Modal Weighted Attention Feature Fusion Network (DCMWAF-Net) for multi-scale lunar crater detection. The model integrates Kaguya TC Morning imagery and SLDEM2015 digital elevation model (DEM) data, leveraging a Multi-modal Feature Extractor (MFE) to capture morphological and topographic features. A Cross-Modal Weighted Attention Feature Fusion (CMWAF) module, combined with the Convolutional Block Attention Module (CBAM), enables adaptive fusion of imagery and topographic features. Additionally, a Multi-Scale Feature Aggregation (MSFA) module, employing a bidirectional pyramid structure, enhances feature integration, while the Detection Head (DH) module, with an optimized anchor-box mechanism, improves localization accuracy, significantly boosting the detection of multi-scale craters. Experimental results demonstrate that DCMWAF-Net achieves superior performance in the test region spanning 55°–59°W longitude and 45°–48°N latitude, with an overall recall of 93.26% and precision of 90.71%. Compared to single-modal methods using only imagery or DEM and traditional data fusion models, DCMWAF-Net improves recall by 3%–14% and achieves a matching rate of 93.24% against ground-truth annotations. Furthermore, the model successfully identified and confirmed 2,138 new craters, and predicted a total of 174,806 craters in the Chang’e-5 landing region, significantly expanding the existing crater database while enhancing the accuracy and efficiency of crater detection in complex terrain environments. Xu Zhang 0071, Jialong Lai, Feifei Cui, Yi Xu 0010, Xiaoping Zhang 0006 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | GCNLA: Inferring Cell-Cell Interactions From Spatial Transcriptomics With Long Short-Term Memory and Graph Convolutional NetworksabstractSpatial transcriptomics analysis methods offer an opportunity to investigate highly diverse biological tissues. Cell-cell communication is fundamental for maintaining physiological homeostasis in organisms and coordinating complex biological processes. Identifying cell-cell interactions is critical for understanding cellular activities. The interaction of a cell with other cells depends on several factors, and most of the existing methods that consider only gene expression information of neighbouring cells and spatial location information are somewhat limited. In this paper, we propose a network architecture based on graph convolution network and long short-term memory attention module-GCNLA, which contains graph convolution layer, long short-term memory network, attention module, and residual connections. GCNLA not only learns the spatial structure of cells but also captures interaction information between distal cells, the attention module further extracting and enhancing features related to cell-cell interactions. Finally, the inner product decoding calculates the cosine similarity, which is used to infer cell-cell interactions. In addition, GCNLA is capable of reconstructing the complete cell-cell interaction network. The experimental results on seqFISH and MERFISH demonstrate that the GCNLA network structure has better robustness and noise immunity. The potential features learned by GCNLA enable other downstream analyses, including single-cell resolution cell clustering based on spatial information resolving cell heterogeneity. Xiuhao Fu, Zhenjie Luo, Leyi Wei, Jingbing Li, Feifei Cui, Quan Zou 0001, Qingchen Zhang 0001 |
IEEE J. Biomed. Health Informatics | 6 |
| 2024 | RNASite: A one-stop tool website that integrates multiple RNA modification site databases and serversabstractRNASite is a comprehensive platform that integrates multiple RNA modification site databases and servers, focusing on RNA modification sites. These modification sites significantly affect RNA's structure, stability, and function, playing a crucial role in epigenetics and gene expression regulation. The website offers both datasets and online RNA modification site identification tools to enhance the recognition and understanding of these modification sites. With advancements in high-throughput sequencing technology and machine learning, RNASite leverages these innovations to provide efficient and cost-effective RNA modification site identification methods. The platform features 17 high-quality datasets covering 11 common RNA modification sites and includes 7 online identification tools. Notably, some of these tools exceed the accuracy of existing models. RNASite is designed to be a convenient and efficient resource for researchers in biochemistry and bioinformatics, facilitating progress in the study of RNA modification sites. The platform is available for free at http://www.bioai-lab.com/RNASite. Xingfeng Li 0001, Qingchen Zhang 0001, Quan Zou 0001, Feifei Cui |
BIBM | 6 |
| 2024 | MultiFeatVotPIP: a voting-based ensemble learning framework for predicting proinflammatory peptidesabstractInflammatory responses may lead to tissue or organ damage, and proinflammatory peptides (PIPs) are signaling peptides that can induce such responses. Many diseases have been redefined as inflammatory diseases. To identify PIPs more efficiently, we expanded the dataset and designed an ensemble learning model with manually encoded features. Specifically, we adopted a more comprehensive feature encoding method and considered the actual impact of certain features to filter them. Identification and prediction of PIPs were performed using an ensemble learning model based on five different classifiers. The results show that the model's sensitivity, specificity, accuracy, and Matthews correlation coefficient are all higher than those of the state-of-the-art models. We named this model MultiFeatVotPIP, and both the model and the data can be accessed publicly at https://github.com/ChaoruiYan019/MultiFeatVotPIP. Additionally, we have developed a user-friendly web interface for users, which can be accessed at http://www.bioai-lab.com/MultiFeatVotPIP. Chaorui Yan, Aoyun Geng, Zhuoyu Pan, Feifei Cui |
Briefings Bioinform. | 5 |
| 2024 | DPNN-ac4C: a dual-path neural network with self-attention mechanism for identification of N4-acetylcytidine (ac4C) in mRNAabstractMOTIVATION: The modification of N4-acetylcytidine (ac4C) in RNA is a conserved epigenetic mark that plays a crucial role in post-transcriptional regulation, mRNA stability, and translation efficiency. Traditional methods for detecting ac4C modifications are laborious and costly, necessitating the development of efficient computational approaches for accurate identification of ac4C sites in mRNA. RESULTS: We present DPNN-ac4C, a dual-path neural network with a self-attention mechanism for the identification of ac4C sites in mRNA. Our model integrates embedding modules, bidirectional GRU networks, convolutional neural networks, and self-attention to capture both local and global features of RNA sequences. Extensive evaluations demonstrate that DPNN-ac4C outperforms existing models, achieving an AUROC of 91.03%, accuracy of 82.78%, MCC of 65.78%, and specificity of 84.78% on an independent test set. Moreover, DPNN-ac4C exhibits robustness under the Fast Gradient Method attack, maintaining a high level of accuracy in practical applications. AVAILABILITY AND IMPLEMENTATION: The model code and dataset are publicly available on GitHub (https://github.com/shock1ng/DPNN-ac4C). Zhuoyu Pan, Aohan Li, Feifei Cui |
Bioinform. | 6 |
| 2024 | MVST: Identifying spatial domains of spatial transcriptomes from multiple views using multi-view graph convolutional networksabstractSpatial transcriptome technology can parse transcriptomic data at the spatial level to detect high-throughput gene expression and preserve information regarding the spatial structure of tissues. Identifying spatial domains, that is identifying regions with similarities in gene expression and histology, is the most basic and critical aspect of spatial transcriptome data analysis. Most current methods identify spatial domains only through a single view, which may obscure certain important information and thus fail to make full use of the information embedded in spatial transcriptome data. Therefore, we propose an unsupervised clustering framework based on multiview graph convolutional networks (MVST) to achieve accurate spatial domain recognition by the learning graph embedding features of neighborhood graphs constructed from gene expression information, spatial location information, and histopathological image information through multiview graph convolutional networks. By exploring spatial transcriptomes from multiple views, MVST enables data from all parts of the spatial transcriptome to be comprehensively and fully utilized to obtain more accurate spatial expression patterns. We verified the effectiveness of MVST on real spatial transcriptome datasets, the robustness of MVST on some simulated datasets, and the reasonableness of the framework structure of MVST in ablation experiments, and from the experimental results, it is clear that MVST can achieve a more accurate spatial domain identification compared with the current more advanced methods. In conclusion, MVST is a powerful tool for spatial transcriptome research with improved spatial domain recognition. Qingchen Zhang 0001, Feifei Cui, Quan Zou 0001 |
PLoS Comput. Biol. | 3 |
| 2024 | Toward Data Integrity Attacks Against Distributed Dynamic State Estimation in Smart GridabstractWith the continuous expansion of the power grid nodes scale, traditional centralized state estimation method shows certain limitations in estimation efficiency and accuracy. Recently, some power grids adopt a distributed state estimation method, in which each partition independently estimates the partial state information by partitioning the entire power system. However, the deviation of the state estimation in certain partition will result in the deviation of the estimation results in the entire power grid system. In this paper, we propose the attack strategy against the distributed state estimation in smart grids from two perspectives, i.e. the attack against local physical measurement value of the power system partition and the attack against the measurement value of coordination center. Moreover, the theoretical analysis of the state estimation deviation caused by the proposed data integrity attack and the propagation processes of proposed attack vectors in measurement calculations are formalized. The effectiveness of the proposed attack strategy is verified in the IEEE-30 bus and IEEE-118 bus systems. Simulation results show that attacking against a certain partition of a distributed system can indirectly affect the state estimation results of other partitions and attacking against the measurement value of coordination center can directly threaten the state estimation results of the entire power grid. Note to Practitioners— This paper proposes two attack strategies against the distributed state estimation of power grid from two perspectives, i.e. the attack against local physical measurement value of the power system partition and the attack against the measurement value of coordination center. Most of the previous works fail to formalize the state estimation deviation of both partial and entire state estimation results of power grid after the attacker launches the attack against the distributed state estimation. We formalize the state estimation deviation caused by the proposed data integrity attack and the propagation processes of proposed attack vectors in measurement calculations. The effectiveness of the proposed attack strategy against the state estimation of power grid is verified in the IEEE-30 bus and IEEE-118 bus systems. Simulation results show that attacking against a certain partition of power grid can indirectly cause the deviation in the state estimation results of entire power grid and attacking against the measurement value of the coordination center can directly threaten the state estimation results of the entire power grid. In conclusion, the proposed attack strategies are helpful for the research community to design detection strategies in a targeted manner and can be conveniently applied to the real-world security management system of smart grid. Dou An, Feiye Zhang, Feifei Cui, Qingyu Yang 0003 |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2024 | Hyb_SEnc: An Antituberculosis Peptide Predictor Based on a Hybrid Feature Vector and Stacked Ensemble LearningabstractTuberculosis has plagued mankind since ancient times, and the struggle between humans and tuberculosis continues. Mycobacterium tuberculosis is the leading cause of tuberculosis, infecting nearly one-third of the world's population. The rise of peptide drugs has created a new direction in the treatment of tuberculosis. Therefore, for the treatment of tuberculosis, the prediction of anti-tuberculosis peptides is crucial. This paper proposes an anti-tuberculosis peptide prediction method based on hybrid features and stacked ensemble learning. First, a random forest (RF) and extremely randomized tree (ERT) are selected as first-level learning of stacked ensembles. Then, the five best-performing feature encoding methods are selected to obtain the hybrid feature vector, and then the decision tree and recursive feature elimination (DT-RFE) are used to refine the hybrid feature vector. After selection, the optimal feature subset is used as the input of the stacked ensemble model. At the same time, logistic regression (LR) is used as a stacked ensemble secondary learner to build the final stacked ensemble model Hyb_SEnc. The prediction accuracy of Hyb_SEnc achieved 94.68% and 95.74% on the independent test sets of AntiTb_MD and AntiTb_RD, respectively. Xiuhao Fu, Xiaofeng Zang, Xingfeng Li 0001, Qingchen Zhang 0001, Quan Zou 0001, Feifei Cui |
IEEE ACM Trans. Comput. Biol. Bioinform. | 9 |
| 2023 | AIPPT: Predicts anti-inflammatory peptides using the most characteristic subset of bases and sequences by stacking ensemble learning strategiesabstractTherapeutic peptides play a vital role in developing peptide-based drugs. Recently, they have been applied as anti-inflammatory agents for a range of inflammatory conditions, including Alzheimer’s disease and rheumatoid arthritis. Laboratory-based identification of peptides with anti-inflammatory properties is a highly time-consuming and labor-intensive endeavor. To tackle this issue, researchers have developed computational methods, primarily centered on machine learning, to streamline the procedure. This paper presents AIPPT, an intelligent and computationally efficient prediction tool that introduces a novel stacking framework for the reliable identification of anti-inflammatory peptides (AIP). The study specifically employs a combination of four feature encodings, where their importance is assessed using the LightGBM method to create an optimal feature subset, which is then input to the three classifiers. The output probabilities from the three classifiers are further fed into a meta-classifier, constructing a two-layer stacking model. Subsequently, the output probabilities from the three classifiers are incorporated into a meta-classifier, establishing a two-layer stacking model. Subsequently, the output probabilities from the three classifiers are incorporated into a meta-classifier, establishing a two-layer stacking model. Xiuhao Fu, Shankai Yan, Feifei Cui |
BIBM | 5 |
| 2023 | Human-Spa: An Online Platform Based on Spatial Transcriptome Data for Diseases of Human SystemsabstractSpatial transcriptomics has become a major method for high-throughput analysis of gene expression at the current level of cells, which can directly study gene expression changes in disease cells, reveal the occurrence and development mechanism of diseases, identify potential therapeutic targets related to diseases, and provide new clues for disease diagnosis and treatment. Although there have been major breakthroughs in the analysis and acquisition of transcriptome data, there are still many challenges in how to effectively manage, share and utilize these valuable data resources in the study of human systemic diseases. To solve this problem, we are committed to building a spatial transcriptome database website related to human systemic diseases, aiming to provide reliable datasets related to multiple diseases, and provide certain data information and analysis results to accelerate the progress of human systemic diseases research. This paper will introduce the construction process of the database website in detail, and discuss its application prospects in disease research, with the aim of promoting further development and innovation in the field of human health. Here, Human-spa mainly includes 12 Human systems, 38 disease types, 55 datasets, and Human-Spa provides a very friendly web interface for visualization and dataset parsing. In conclusion, the construction of Human-Spa will provide a powerful tool and resource platform for Human disease research. Human-Spa is available for free at http://www.human-spa.cn/ Yunyun Su, Feifei Cui, Shiyu Yan, Quan Zou 0001, Chen Cao 0002 |
BIBM | 2 |
| 2022 | webSCST: an interactive web application for single-cell RNA-sequencing data and spatial transcriptomic data integrationabstractSUMMARY: Integrative analysis of single-cell RNA-sequencing (scRNA-seq) data with spatial data for the same species and organ would provide each cell sample with a predictive spatial location, which would facilitate biological study. However, publicly available spatial sequencing datasets for specific species and organs are rare and are often displayed in different formats. In this study, we introduce a new web-based scRNA-seq analysis tool, webSCST, that integrates well-organized spatial transcriptome sequencing datasets categorized by species and organs, provides a user-friendly interface for raw single-cell processing with popular integration methods and allows users to submit their raw scRNA-seq data once to obtain predicted spatial locations for each cell type. AVAILABILITY AND IMPLEMENTATION: webSCST implemented in shiny with all major browsers supported is available at http://www.webscst.com. webSCST is also freely available as an R package at https://github.com/swsoyee/webSCST. Feifei Cui, Lijun Dou, Chen Cao 0002, Quan Zou 0001 |
Bioinform. | 2 |
| 2021 | Anticancer peptides prediction with deep representation learning featuresabstractAnticancer peptides constitute one of the most promising therapeutic agents for combating common human cancers. Using wet experiments to verify whether a peptide displays anticancer characteristics is time-consuming and costly. Hence, in this study, we proposed a computational method named identify anticancer peptides via deep representation learning features (iACP-DRLF) using light gradient boosting machine algorithm and deep representation learning features. Two kinds of sequence embedding technologies were used, namely soft symmetric alignment embedding and unified representation (UniRep) embedding, both of which involved deep neural network models based on long short-term memory networks and their derived networks. The results showed that the use of deep representation learning features greatly improved the capability of the models to discriminate anticancer peptides from other peptides. Also, UMAP (uniform manifold approximation and projection for dimension reduction) and SHAP (shapley additive explanations) analysis proved that UniRep have an advantage over other features for anticancer peptide identification. The python script and pretrained models could be downloaded from https://github.com/zhibinlv/iACP-DRLF or from http://public.aibiochem.net/iACP-DRLF/. Zhibin Lv, Feifei Cui, Quan Zou 0001, Lei Xu 0047 |
Briefings Bioinform. | 2 |
| 2021 | Critical downstream analysis steps for single-cell RNA sequencing dataabstractSingle-cell RNA sequencing (scRNA-seq) has enabled us to study biological questions at the single-cell level. Currently, many analysis tools are available to better utilize these relatively noisy data. In this review, we summarize the most widely used methods for critical downstream analysis steps (i.e. clustering, trajectory inference, cell-type annotation and integrating datasets). The advantages and limitations are comprehensively discussed, and we provide suggestions for choosing proper methods in different situations. We hope this paper will be useful for scRNA-seq data analysts and bioinformatics tool developers. Feifei Cui, Chen Lin 0001, Lingling Zhao, Chunyu Wang 0002, Quan Zou 0001 |
Briefings Bioinform. | 2 |
| 2021 | Goals and approaches for each processing step for single-cell RNA sequencing dataabstractSingle-cell RNA sequencing (scRNA-seq) has enabled researchers to study gene expression at the cellular level. However, due to the extremely low levels of transcripts in a single cell and technical losses during reverse transcription, gene expression at a single-cell resolution is usually noisy and highly dimensional; thus, statistical analyses of single-cell data are a challenge. Although many scRNA-seq data analysis tools are currently available, a gold standard pipeline is not available for all datasets. Therefore, a general understanding of bioinformatics and associated computational issues would facilitate the selection of appropriate tools for a given set of data. In this review, we provide an overview of the goals and most popular computational analysis tools for the quality control, normalization, imputation, feature selection and dimension reduction of scRNA-seq data. Feifei Cui, Chunyu Wang 0002, Lingling Zhao, Quan Zou 0001 |
Briefings Bioinform. | 2 |