Weixing Feng

dblp:70/7778 · DBLP profile ↗
← Back
20ranked-venue papers
1as first author
16since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 14 · 1 first-author · 10 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 DRS-GCN: Counteracting over-smoothing in graph convolutional networks through dynamic reorganization and smoothness loss
Qi Cheng 0007, Lang Long, Min Zhang 0053, Chengkui Zhao, Weixing Feng
Appl. Intell.5
2026 Enhancing mutation impact prediction in protein-protein interactions through interpretable graph-based multi-level feature interactions
abstract
MOTIVATION: Protein-protein interactions (PPIs) are central to cellular functions, and predicting mutation-induced changes in binding affinity (ΔΔG) remains challenging. Although existing computational methods integrate sequence- and structure-derived features and thus implicitly capture certain sequence-structure relationships, they typically fuse these modalities through simple concatenation, without explicitly modeling their multidimensional and multiscale interdependencies. RESULTS: Here, we introduce IGMI, an interpretable graph-based model that explicitly encodes multi-level feature interactions across 1D sequences, 2D contact maps, 3D structures, and residue- and atom-level representations. By recalibrating cross-dimensional and cross-scale dependencies, IGMI enables more accurate estimation of both local and long-range mutation effects. Across multiple benchmark datasets, IGMI consistently outperforms state-of-the-art methods in accuracy, robustness, and interpretability. Macro- and micro-level analyses further reveal biologically plausible patterns, distinguishing direct interface perturbations from indirect structural reorganizations. Complementary analyses under different data splitting strategies indicate that the model learns generalizable affinity-related interaction patterns, rather than relying on split-specific information. IGMI provides a reliable and interpretable framework for modeling mutation-induced affinity changes, supporting applications in protein engineering and therapeutic design. AVAILABILITY AND IMPLEMENTATION: IGMI is implemented in PyTorch and released under an open-source license. The full codebase, training scripts, and evaluation utilities are available at https://github.com/ShiweiWu-545/IGMI.git. An archival snapshot containing all source code, pre-trained weights, processed datasets, and reproducibility scripts is available on Zenodo (https://doi.org/10.5281/zenodo.17563574). CONTACT: [email protected]; [email protected]; [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Xiaohui Xin, Min Zhang 0053, Haoliang Liu, Hongjia Zhu, Chengkui Zhao, Weixing Feng
Bioinform.10
2026 A spatio-temporal prediction framework for marine diesel engine turbine exhaust gas temperature using evidence-driven dual-graph fusion
Yunpeng Cao, Yujia Ma, Weixing Feng
Eng. Appl. Artif. Intell.5
2026 Multi-Shaft Speed-Informed Adaptive Window Filtering Method for Acoustic Pressure Signals in Marine Gas Turbines
Yun-peng Cao, Minghao Wu, Weiying Wang, Weixing Feng
Signal Process.5
2025 NSM-planner: neuromorphic planner with spiking memory for underwater autonomous obstacle avoidance decision of AUV
Weixing Feng
Adv. Eng. Informatics3
2025 Autonomous obstacle avoidance decision method for spherical underwater robot based on brain-inspired spiking neural network
Boyang Zhang 0010, Huiming Xing, Weixing Feng
Expert Syst. Appl.4
2025 A universal strategy for smoothing deceleration in deep graph neural networks
Qi Cheng 0007, Lang Long, Jiayu Xu 0004, Min Zhang 0053, Shuangze Han, Chengkui Zhao, Weixing Feng
Neural Networks7
2024 Meta-DHGNN: method for CRS-related cytokines analysis in CAR-T therapy based on meta-learning directed heterogeneous graph neural network
abstract
Chimeric antigen receptor T-cell (CAR-T) immunotherapy, a novel approach for treating blood cancer, is associated with the production of cytokine release syndrome (CRS), which poses significant safety concerns for patients. Currently, there is limited knowledge regarding CRS-related cytokines and the intricate relationship between cytokines and cells. Therefore, it is imperative to explore a reliable and efficient computational method to identify cytokines associated with CRS. In this study, we propose Meta-DHGNN, a directed and heterogeneous graph neural network analysis method based on meta-learning. The proposed method integrates both directed and heterogeneous algorithms, while the meta-learning module effectively addresses the issue of limited data availability. This approach enables comprehensive analysis of the cytokine network and accurate prediction of CRS-related cytokines. Firstly, to tackle the challenge posed by small datasets, a pre-training phase is conducted using the meta-learning module. Consequently, the directed algorithm constructs an adjacency matrix that accurately captures potential relationships in a more realistic manner. Ultimately, the heterogeneous algorithm employs meta-photographs and multi-head attention mechanisms to enhance the realism and accuracy of predicting cytokine information associated with positive labels. Our experimental verification on the dataset demonstrates that Meta-DHGNN achieves favorable outcomes. Furthermore, based on the predicted results, we have explored the multifaceted formation mechanism of CRS in CAR-T therapy from various perspectives and identified several cytokines, such as IFNG (IFN-γ), IFNA1, IFNB1, IFNA13, IFNA2, IFNAR1, IFNAR2, IFNGR1 and IFNGR2 that have been relatively overlooked in previous studies but potentially play pivotal roles. The significance of Meta-DHGNN lies in its ability to analyze directed and heterogeneous networks in biology effectively while also facilitating CRS risk prediction in CAR-T therapy.
Chengkui Zhao, Min Zhang 0053, Jiayu Xu 0004, Xiaohui Xin, Weixing Feng
Briefings Bioinform.9
2024 BertTCR: a Bert-based deep learning framework for predicting cancer-related immune status based on T cell receptor repertoire
abstract
The T cell receptor (TCR) repertoire is pivotal to the human immune system, and understanding its nuances can significantly enhance our ability to forecast cancer-related immune responses. However, existing methods often overlook the intra- and inter-sequence interactions of T cell receptors (TCRs), limiting the development of sequence-based cancer-related immune status predictions. To address this challenge, we propose BertTCR, an innovative deep learning framework designed to predict cancer-related immune status using TCRs. BertTCR combines a pre-trained protein large language model with deep learning architectures, enabling it to extract deeper contextual information from TCRs. Compared to three state-of-the-art sequence-based methods, BertTCR improves the AUC on an external validation set for thyroid cancer detection by 21 percentage points. Additionally, this model was trained on over 2000 publicly available TCR libraries covering 17 types of cancer and healthy samples, and it has been validated on multiple public external datasets for its ability to distinguish cancer patients from healthy individuals. Furthermore, BertTCR can accurately classify various cancer types and healthy individuals. Overall, BertTCR is the advancing method for cancer-related immune status forecasting based on TCRs, offering promising potential for a wide range of immune status prediction tasks.
Min Zhang 0053, Qi Cheng 0007, Jiayu Xu 0004, Chengkui Zhao, Weixing Feng
Briefings Bioinform.9
2024 PrCRS: a prediction model of severe CRS in CAR-T therapy based on transfer learning
abstract
BACKGROUND: CAR-T cell therapy represents a novel approach for the treatment of hematologic malignancies and solid tumors. However, its implementation is accompanied by the emergence of potentially life-threatening adverse events known as cytokine release syndrome (CRS). Given the escalating number of patients undergoing CAR-T therapy, there is an urgent need to develop predictive models for severe CRS occurrence to prevent it in advance. Currently, all existing models are based on decision trees whose accuracy is far from meeting our expectations, and there is a lack of deep learning models to predict the occurrence of severe CRS more accurately. RESULTS: We propose PrCRS, a deep learning prediction model based on U-net and Transformer. Given the limited data available for CAR-T patients, we employ transfer learning using data from COVID-19 patients. The comprehensive evaluation demonstrates the superiority of the PrCRS model over other state-of-the-art methods for predicting CRS occurrence. We propose six models to forecast the probability of severe CRS for patients with one, two, and three days in advance. Additionally, we present a strategy to convert the model's output into actual probabilities of severe CRS and provide corresponding predictions. CONCLUSIONS: Based on our findings, PrCRS effectively predicts both the likelihood and timing of severe CRS in patients, thereby facilitating expedited and precise patient assessment, thus making a significant contribution to medical research. There is little research on applying deep learning algorithms to predict CRS, and our study fills this gap. This makes our research more novel and significant. Our code is publicly available at https://github.com/wzy38828201/PrCRS . The website of our prediction platform is: http://prediction.unicar-therapy.com/index-en.html .
Chengkui Zhao, Min Zhang 0053, Jiayu Xu 0004, Xiaohui Xin, Weixing Feng
BMC Bioinform.9
2023 A score-based method of immune status evaluation for healthy individuals with complete blood cell counts
abstract
BACKGROUND: With the COVID-19 outbreak, an increasing number of individuals are concerned about their health, particularly their immune status. However, as of now, there is no available algorithm that effectively assesses the immune status of normal, healthy individuals. In response to this, a new score-based method is proposed that utilizes complete blood cell counts (CBC) to provide early warning of disease risks, such as COVID-19. METHODS: First, data on immune-related CBC measurements from 16,715 healthy individuals were collected. Then, a three-platform model was developed to normalize the data, and a Gaussian mixture model was optimized with expectation maximization (EM-GMM) to cluster the immune status of healthy individuals. Based on the results, Random Forest (RF), Light Gradient Boosting Machine (LightGBM) and Extreme Gradient Boosting (XGBoost) were used to determine the correlation of each CBC index with the immune status. Consequently, a weighted sum model was constructed to calculate a continuous immunity score, enabling the evaluation of immune status. RESULTS: The results demonstrated a significant negative correlation between the immunity score and the age of healthy individuals, thereby validating the effectiveness of the proposed method. In addition, a nonlinear polynomial regression model was developed to depict this trend. By comparing an individual's immune status with the reference value corresponding to their age, their immune status can be evaluated. CONCLUSION: In summary, this study has established a novel model for evaluating the immune status of healthy individuals, providing a good approach for early detection of abnormal immune status in healthy individuals. It is helpful in early warning of the risk of infectious diseases and of significant importance.
Min Zhang 0053, Chengkui Zhao, Qi Cheng 0007, Jiayu Xu 0004, Weixing Feng
BMC Bioinform.7
2022 ILGBMSH: an interpretable classification model for the shRNA target prediction with ensemble learning algorithm
abstract
Short hairpin RNA (shRNA)-mediated gene silencing is an important technology to achieve RNA interference, in which the design of potent and reliable shRNA molecules plays a crucial role. However, efficient shRNA target selection through biological technology is expensive and time consuming. Hence, it is crucial to develop a more precise and efficient computational method to design potent and reliable shRNA molecules. In this work, we present an interpretable classification model for the shRNA target prediction using the Light Gradient Boosting Machine algorithm called ILGBMSH. Rather than utilizing only the shRNA sequence feature, we extracted 554 biological and deep learning features, which were not considered in previous shRNA prediction research. We evaluated the performance of our model compared with the state-of-the-art shRNA target prediction models. Besides, we investigated the feature explanation from the model's parameters and interpretable method called Shapley Additive Explanations, which provided us with biological insights from the model. We used independent shRNA experiment data from other resources to prove the predictive ability and robustness of our model. Finally, we used our model to design the miR30-shRNA sequences and conducted a gene knockdown experiment. The experimental result was perfectly in correspondence with our expectation with a Pearson's coefficient correlation of 0.985. In summary, the ILGBMSH model can achieve state-of-the-art shRNA prediction performance and give biological insights from the machine learning model parameters.
Chengkui Zhao, Jingwen Tan, Qi Cheng 0007, Weixin Xie, Jiayu Xu 0004, Weixing Feng
Briefings Bioinform.10
2022 Distance correlation application to gene co-expression network analysis
abstract
BACKGROUND: To construct gene co-expression networks, it is necessary to evaluate the correlation between different gene expression profiles. However, commonly used correlation metrics, including both linear (such as Pearson's correlation) and monotonic (such as Spearman's correlation) dependence metrics, are not enough to observe the nature of real biological systems. Hence, introducing a more informative correlation metric when constructing gene co-expression networks is still an interesting topic. RESULTS: In this paper, we test distance correlation, a correlation metric integrating both linear and non-linear dependence, with other three typical metrics (Pearson's correlation, Spearman's correlation, and maximal information coefficient) on four different arrays (macrophage and liver) and RNA-seq (cervical cancer and pancreatic cancer) datasets. Among all the metrics, distance correlation is distribution free and can provide better performance on complex relationships and anti-outlier. Furthermore, distance correlation is applied to Weighted Gene Co-expression Network Analysis (WGCNA) for constructing a gene co-expression network analysis method which we named Distance Correlation-based Weighted Gene Co-expression Network Analysis (DC-WGCNA). Compared with traditional WGCNA, DC-WGCNA can enhance the result of enrichment analysis and improve the module stability. CONCLUSIONS: Distance correlation is better at revealing complex biological relationships between gene profiles compared with other correlation metrics, which contribute to more meaningful modules when analyzing gene co-expression networks. However, due to the high time complexity of distance correlation, the implementation requires more computer memory.
Xiufen Ye, Weixing Feng, Yatong Han, Yusong Liu, Yufen Wei
BMC Bioinform.3
2022 Investigation of CRS-associated cytokines in CAR-T therapy with meta-GNN and pathway crosstalk
abstract
BACKGROUND: Chimeric antigen receptor T-cell (CAR-T) therapy is a new and efficient cellular immunotherapy. The therapy shows significant efficacy, but also has serious side effects, collectively known as cytokine release syndrome (CRS). At present, some CRS-related cytokines and their roles in CAR-T therapy have been confirmed by experimental studies. However, the mechanism of CRS remains to be fully understood. METHODS: Based on big data for human protein interactions and meta-learning graph neural network, we employed known CRS-related cytokines to comprehensively investigate the CRS associated cytokines in CAR-T therapy through protein interactions. Subsequently, the clinical data for 119 patients who received CAR-T therapy were examined to validate our prediction results. Finally, we systematically explored the roles of the predicted cytokines in CRS occurrence by protein interaction network analysis, functional enrichment analysis, and pathway crosstalk analysis. RESULTS: We identified some novel cytokines that would play important roles in biological process of CRS, and investigated the biological mechanism of CRS from the perspective of functional analysis. CONCLUSIONS: 128 cytokines and related molecules had been found to be closely related to CRS in CAR-T therapy, where several important ones such as IL6, IFN-γ, TNF-α, ICAM-1, VCAM-1 and VEGFA were highlighted, which can be the key factors to predict CRS.
Qi Cheng 0007, Chengkui Zhao, Jiayu Xu 0004, Liqing Kang, Xiaoyan Lou, Weixing Feng
BMC Bioinform.9
2021 TPSC: a module detection method based on topology potential and spectral clustering in weighted networks and its application in gene co-expression module discovery
abstract
BACKGROUND: Gene co-expression networks are widely studied in the biomedical field, with algorithms such as WGCNA and lmQCM having been developed to detect co-expressed modules. However, these algorithms have limitations such as insufficient granularity and unbalanced module size, which prevent full acquisition of knowledge from data mining. In addition, it is difficult to incorporate prior knowledge in current co-expression module detection algorithms. RESULTS: In this paper, we propose a novel module detection algorithm based on topology potential and spectral clustering algorithm to detect co-expressed modules in gene co-expression networks. By testing on TCGA data, our novel method can provide more complete coverage of genes, more balanced module size and finer granularity than current methods in detecting modules with significant overall survival difference. In addition, the proposed algorithm can identify modules by incorporating prior knowledge. CONCLUSION: In summary, we developed a method to obtain as much as possible information from networks with increased input coverage and the ability to detect more size-balanced and granular modules. In addition, our method can integrate data from different sources. Our proposed method performs better than current methods with complete coverage of input genes and finer granularity. Moreover, this method is designed not only for gene co-expression networks but can also be applied to any general fully connected weighted network.
Yusong Liu, Xiufen Ye, Christina Y. Yu, Wei Shao 0005, Weixing Feng, Jie Zhang 0010, Kun Huang 0001
BMC Bioinform.6
2021 Improved Adverse Drug Event Prediction Through Information Component Guided Pharmacological Network Model (IC-PNM)
abstract
Improving adverse drug event (ADE) prediction is highly critical in pharmacovigilance research. We propose a novel information component guided pharmacological network model (IC-PNM) to predict drug-ADE signals. This new method combines the pharmacological network model and information component, a Bayes statistics method. We use 33,947 drug-ADE pairs from the FDA Adverse Event Reporting System (FAERS) 2010 data as the training data, and the new 21,065 drug-ADE pairs from FAERS 2011-2015 as the validations samples. The IC-PNM data analysis suggests that both large and small sample size drug-ADE pairs are needed in training the predictive model for its prediction performance to reach an area under the receiver operating characteristic curve [Formula: see text]. On the other hand, the IC-PNM prediction performance improved to [Formula: see text] if we removed the small sample size drug-ADE pairs from the prediction model during validation.
Xiangmin Ji, Lei Wang 0169, Liyan Hua, Pengyue Zhang, Aditi Shendre, Weixing Feng, Jin Li 0012, Lang Li 0001
IEEE ACM Trans. Comput. Biol. Bioinform.7
2020 Evaluation of bottom-up and top-down mass spectrum identifications with different customized protein sequences databases
abstract
MOTIVATION: Generally, bottom-up and top-down are two complementary approaches for proteoforms identification. The inference of proteoforms relies on searching mass spectra against an accurate proteoform sequence database. A customized protein sequence database derived by RNA-Seq data can be used to better identify the proteoform existed in a studied species. However, the quality of sequences in customized databases which constructed by different strategies affect the performances of mass spectrometry (MS) identification. Additionally, performances of identifications between bottom-up and top-down using customized databases are also needed to be evaluated. RESULTS: Three customized databases were constructed with different strategies separately. Two of them were based on translating assembled transcripts with or without genomic annotation, and the third one is a variant-extending protein database. By testing with bottom-up and top-down MS data separately, a variant-extending protein database could identify not only the most number of spectra but also the alleles expressed at the same time in diploid cells. An assembled database could identify the spectrum missed in reference database and amino acid (AA) alterations existed in studied species. AVAILABILITY AND IMPLEMENTATION: Experimental results demonstrated that the proteoform sequences in an annotated database are more suitable for identifying AA alterations and peptide sequences missed in reference database. An unannotated database instead of a reference proteome database gets an enough high sensitivity of identifying mass spectra. The variant-extending reference database is the most sensitive to identify mass spectra and single AA variants. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Weixing Feng
Bioinform.3
2018 Evaluation of top-down mass spectral identification with homologous protein sequences
abstract
BACKGROUND: Top-down mass spectrometry has unique advantages in identifying proteoforms with multiple post-translational modifications and/or unknown alterations. Most software tools in this area search top-down mass spectra against a protein sequence database for proteoform identification. When the species studied in a mass spectrometry experiment lacks its proteome sequence database, a homologous protein sequence database can be used for proteoform identification. The accuracy of homologous protein sequences affects the sensitivity of proteoform identification and the accuracy of mass shift localization. RESULTS: We tested TopPIC, a commonly used software tool for top-down mass spectral identification, on a top-down mass spectrometry data set of Escherichia coli K12 MG1655, and evaluated its performance using an Escherichia coli K12 MG1655 proteome database and a homologous protein database. The number of identified spectra with the homologous database was about half of that with the Escherichia coli K12 MG1655 database. We also tested TopPIC on a top-down mass spectrometry data set of human MCF-7 cells and obtained similar results. CONCLUSIONS: Experimental results demonstrated that TopPIC is capable of identifying many proteoform spectrum matches and localizing unknown alterations using homologous protein sequences containing no more than 2 mutations.
Qiang Kou, Weixing Feng
BMC Bioinform.7
2016 Characterizing the roles of long non-coding RNA in rat alcohol preference
abstract
Alcohol is one of the major threats to health in United States. With the emerging of next-generation sequencing technology, the association between alcohol preference and the variants and expression of genes has been investigated. However, the roles of long non-coding RNAs (lncRNA) in alcohol preference remains unclear. In this study, we identified 37 novel lncRNAs that differentially expressed across alcohol preferring (P) and non-preferring (NP) rats. The functional study on these lncRNAs demonstrates that they are associated with gene regulation, as well as neural functions. This suggests that these lncRNAs may contribute to the alcohol preference behaviors.
Ao Zhou 0004, Weixing Feng, Howard J. Edenberg
BIBM4
2015 ResSeq: Enhancing Short-Read Sequencing Alignment By Rescuing Error-Containing Reads
abstract
UNLABELLED: Next-generation short-read sequencing is widely utilized in genomic studies. Biological applications require an alignment step to map sequencing reads to the reference genome, before acquiring expected genomic information. This requirement makes alignment accuracy a key factor for effective biological interpretation. Normally, when accounting for measurement errors and single nucleotide polymorphisms, short read mappings with a few mismatches are generally considered acceptable. However, to further improve the efficiency of short-read sequencing alignment, we propose a method to retrieve additional reliably aligned reads (reads with more than a pre-defined number of mismatches), using a Bayesian-based approach. In this method, we first retrieve the sequence context around the mismatched nucleotides within the already aligned reads; these loci contain the genomic features where sequencing errors occur. Then, using the derived pattern, we evaluate the remaining (typically discarded) reads with more than the allowed number of mismatches, and calculate a score that represents the probability that a specific alignment is correct. This strategy allows the extraction of more reliably aligned reads, therefore improving alignment sensitivity. IMPLEMENTATION: The source code of our tool, ResSeq, can be downloaded from: https://github.com/hrbeubiocenter/Resseq.
Weixing Feng, Peichao Sang, Deyuan Lian, Yansheng Dong, Fengfei Song, B. O. He, Fenglin Cao
IEEE ACM Trans. Comput. Biol. Bioinform.1