Wenkang Wang

dblp:235/9579 · DBLP profile ↗
← Back
14ranked-venue papers
6as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 11 · 5 first-author · 11 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Towards Nonlinear Sparse AUC Maximization via Compositional Stochastic Hard Thresholding
abstract
The Area Under the ROC Curve (AUC) is an important evaluation metric for both linear and, in particular, nonlinear classification models, owing to its robustness against class imbalance. Sparse learning with an ℓ₀ constraint can enhance model interpretability and generalization. Prior work has shown that, in the linear setting, the pairwise formulation of AUC maximization can be reformulated as a standard pointwise empirical risk minimization problem, which enables efficient optimization using hard-thresholding gradient descent for ℓ₀-constrained AUC maximization. Extending this approach to the nonlinear setting remains largely unexplored, even though we establish that pairwise AUC maximization in this setting is equivalent to a pointwise compositional optimization problem; however, designing a compositional optimization algorithm compatible with hard-thresholding operators remains an open challenge. To address this challenge, in this paper, we propose a novel algorithm—Compositional Stochastic Hard Thresholding (CSHT)—for nonlinear sparse AUC maximization. Specifically, CSHT integrates stochastic variance-reduced gradient techniques with hard-thresholding projections to effectively reduce gradient estimation variance while enforcing sparsity. Notably, we provide a rigorous convergence analysis and prove that CSHT achieves linear convergence up to a tolerance bound. To the best of our knowledge, this is the first stochastic hard-thresholding algorithm tailored for nonlinear sparse AUC maximization. Extensive experiments on (a) nonlinear sparse AUC maximization using Random Fourier Feature-based kernel approximation and (b) universal adversarial attack scenarios demonstrate the superior performance of CSHT over existing methods, attributed to its unified treatment of nonlinearity and sparsity.
Wenkang Wang
AAAI1
2026 EssLM-MoE: A mixture-of-experts-enhanced framework for protein essentiality prediction using fused protein language models
Min Zeng 0004, Qianpei Liu, Wenkang Wang, Fuhao Zhang, Fei Guo 0001, Min Li 0007
Neurocomputing5
2026 DPGOK: A Deep Learning-Based Method for Protein Function Prediction by Fusing GO Knowledge With Protein Features
abstract
Accurately predicting protein functions is critical for understanding disease mechanisms and discovering potential drug targets. Gene Ontology (GO), with its hierarchical and semantic information, provides valuable context that can be integrated to improve prediction accuracy. Recently, several existing methods have attempted to integrate GO knowledge with protein sequence features for function prediction. However, these methods ignore the fact that GO embeddings should be tailored to proteins to reflect protein-specific functional relevance. To address this limitation, we proposed DPGOK, a deep learning-based method that fused protein-aware GO representations with protein features for function prediction. DPGOK first learns GO semantic representations with a knowledge graph loss and further generates protein-aware GO embeddings under the guidance of protein features. Results show that DPGOK outperforms state-of-the-art methods across all GO domains. Additional experiments demonstrated that DPGOK is capable of discovering hierarchically deeper and more informative functions for target proteins. Ablation studies revealed that the knowledge graph loss we introduced contributes to more stable and semantically coherent GO representations across different domains. Finally, we find that the predictive performance can be further improved when DPGOK is combined with homology-based approaches.
Qiurong Yang, Wenkang Wang, Wei Fan 0010, Ruiqing Zheng, Min Li 0007
IEEE J. Biomed. Health Informatics2
2025 PRING: Rethinking Protein-Protein Interaction Prediction from Pairs to Graphs
abstract
Deep learning-based computational methods have achieved promising results in predicting protein-protein interactions (PPIs). However, existing benchmarks predominantly focus on isolated pairwise evaluations, overlooking a model's capability to reconstruct biologically meaningful PPI networks, which is crucial for biology research. To address this gap, we introduce PRING, the first comprehensive benchmark that evaluates PRotein-protein INteraction prediction from a Graph-level perspective. PRING curates a high-quality, multi-species PPI network dataset comprising 21,484 proteins and 186,818 interactions, with well-designed strategies to address both data redundancy and leakage. Building on this golden-standard dataset, we establish two complementary evaluation paradigms: (1) topology-oriented tasks, which assess intra and cross-species PPI network construction, and (2) function-oriented tasks, including protein complex pathway prediction, GO module analysis, and essential protein justification. These evaluations not only reflect the model's capability to understand the network topology but also facilitate protein function annotation, biological module detection, and even disease mechanism analysis. Extensive experiments on four representative model categories, consisting of sequence similarity-based, naive sequence-based, protein language model-based, and structure-based approaches, demonstrate that current PPI models have potential limitations in recovering both structural and functional properties of PPI networks, highlighting the gap in supporting real-world biological applications. We believe PRING provides a reliable platform to guide the development of more effective PPI prediction models for the community. The dataset and source code of PRING are available at https://github.com/SophieSarceau/PRING.
Xinzhe Zheng 0001, Fanding Xu, Jinzhe Li, Zhiyuan Liu 0001, Wenkang Wang, Tao Chen 0003, Wanli Ouyang, Stan Z. Li, Yan Lu 0001, Nanqing Dong, Yang Zhang 0094
NeurIPS6
2025 SpaNN: Spatial Transcriptomic Data Enhancement Using Deep Neural Network
abstract
Spatial transcriptomic sequencing technology is a powerful tool that combines gene expression data with their physical locations in tissues or organs, providing researchers with unprecedented spatial resolution of cellular molecular functions. Currently, spatial transcriptomic sequencing based on in situ hybridization and imaging can obtain cell location information and transcriptome profiles at single-cell resolution, but it only detects a limited number of genes, which restricts its application in exploring whole-genome expression patterns. Therefore, it is essential to predict the spatial distribution of undetected genes in their spatial transcriptomic data. Here, we introduce a novel data enhancement technique, named SpaNN, which predicts transcriptome expression levels in spatial context. SpaNN employs a custom-designed similarity loss that leverages location information from spatial transcriptomic data to train a deep neural network. This network captures joint embeddings and uses a weighted k-nearest-neighbor approach to predict the unmeasured genes spatial expression levels. Our experiments show that SpaNN not only recovers the expression levels of unmeasured genes but also enhances cell clustering and visualization. Additionally, sensitivity and scalability analyses confirm that SpaNN is robust to parameter variations and can handle large-scale datasets effectively.
Wenkang Wang, Ruiqing Zheng, Min Li 0007
IEEE Trans. Comput. Biol. Bioinform.2
2025 Enhancing Protein Function Prediction Through the Fusion of Multi-Type Biological Knowledge With Protein Language Model and Graph Neural Network
abstract
Proteins play crucial roles in diverse biological functions. Accurately annotating their functions is essential for understanding cellular mechanisms and developing therapies for complex diseases. Computational methods have been proposed as alternatives to labor-intensive and expensive experimental approaches. Existing computational methods have demonstrated that protein evolution information and Protein-Protein Interactions (PPIs) are essential for protein function prediction. However, traditional computational approaches for generating evolution information are time-consuming. On the other hand, proteins lacking interactions are ignored in previous studies. To address these limitations, we propose a novel deep learning framework, named DeepFMB, which incorporates multi-type biological knowledge. DeepFMB leverages a pre-trained protein language model to extract evolution information. Moreover, DeepFMB generates PPI-related features and orthology-related features using graph neural networks on the constructed PPI and orthology networks. Then, these multi-type features are fused adaptively for protein function prediction. Compared to eight state-of-the-art methods, DeepFMB outperforms all of them in terms of F-max and AUPR. Additionally, with the combination of sequence similarity-based inference, our predicted model predicts protein functions more accurately. Experimental results also validate the superior performance of our methods in predicting low-frequency GO terms. Ablation studies demonstrate that the multi-type biological knowledge we use is highly relevant to protein functions.
Wenkang Wang, Yunyan Shuai, Min Zeng 0004, Min Li 0007
IEEE Trans. Comput. Biol. Bioinform.1
2025 TransScore: A Graph Model for Pose Scoring and Affinity Prediction Based on Transformer Convolution Network
abstract
Predicting the interaction of protein and compound is an important task in drug discovery. Molecular docking has been a fundamental and vital computer-aid tool for digging potential interaction of the protein-compound pair. With the recent great success of artificial intelligence (AI), the scoring function, as a fundamental part of molecular docking, has been achieving much better performance by incorporating AI-based models. However, the AI-based models usually focus on a single prediction task (e.g., affinity prediction), which is limited by their lack of extensibility. Moreover, the performance of AI-based models usually declines in cold start scenarios, thus compromising the robustness. To this end, we propose a novel deep learning-based graph model based on the transformer convolution network for pose scoring and affinity prediction. TransScore captures the intrinsic characteristics of protein-compound poses by employing the self-attention mechanism, which achieves superior performances in both cold and warm scenarios for the pose-scoring task. The outstanding performance is also shown in imbalanced datasets, which demonstrates the robustness of TransScore. In addition, the gated residual algorithm in TransScore enhances the model to adapt to diverse related tasks. In particular, in the affinity prediction task, we have observed consistent improvements in warm/cold start scenarios. Moreover, it is noticeable that TransScore excels in both accuracy and precision, accurately predicting affinities and their relative ordering. We also conducted an analysis on carbonic anhydrase II, which bears out that TransScore can elaborate the interaction mechanism of the protein-ligand pair, suggesting the potential application of TransScore in drug discovery.
Chuqi Lei, Wenkang Wang, Wei Fan 0010, Zhangli Lu, Jing Tang 0002, Min Li 0007
IEEE J. Biomed. Health Informatics2
2024 Network embedding for detecting protein complexes in attributed networks
abstract
Detecting protein complexes holds paramount importance in elucidating cellular organization and protein functionalities. Over the past decade, numerous approaches have centered their attention on the topological intricacies of protein-protein interaction (PPI) networks, yet these have often fallen short in harnessing the full spectrum of biological information inherent in proteins. To bridge this gap, we introduce a novel methodology, designated NE-DPC, which integrates both the topological landscape of PPI networks and the attribute profiles of proteins to facilitate the identification of protein complexes. Initially, we fuse the static PPI network with protein attributes through advanced network embedding techniques. Subsequently, we construct a cosine similarity matrix grounded on these embedded vectors, capturing the intricate relationships among proteins. Lastly, we employ a core-attachment strategy to pinpoint protein complexes. The results reveal that NE-DPC outperforms cutting-edge methods, demonstrating its effectiveness and potential in advancing protein complex detection.
Xiangmao Meng, Keming Wang, Ju Xiang, Wenkang Wang
BIBM4
2024 A comprehensive computational benchmark for evaluating deep learning-based protein function prediction approaches
abstract
Proteins play an important role in life activities and are the basic units for performing functions. Accurately annotating functions to proteins is crucial for understanding the intricate mechanisms of life and developing effective treatments for complex diseases. Traditional biological experiments struggle to keep pace with the growing number of known proteins. With the development of high-throughput sequencing technology, a wide variety of biological data provides the possibility to accurately predict protein functions by computational methods. Consequently, many computational methods have been proposed. Due to the diversity of application scenarios, it is necessary to conduct a comprehensive evaluation of these computational methods to determine the suitability of each algorithm for specific cases. In this study, we present a comprehensive benchmark, BeProf, to process data and evaluate representative computational methods. We first collect the latest datasets and analyze the data characteristics. Then, we investigate and summarize 17 state-of-the-art computational methods. Finally, we propose a novel comprehensive evaluation metric, design eight application scenarios and evaluate the performance of existing methods on these scenarios. Based on the evaluation, we provide practical recommendations for different scenarios, enabling users to select the most suitable method for their specific needs. All of these servers can be obtained from https://csuligroup.com/BEPROF and https://github.com/CSUBioGroup/BEPROF.
Wenkang Wang, Yunyan Shuai, Qiurong Yang, Fuhao Zhang, Min Zeng 0004, Min Li 0007
Briefings Bioinform.1
2024 Dopcc: Detecting Overlapping Protein Complexes via Multi-Metrics and Co-Core Attachment Method
abstract
Identification of protein complex is an important issue in the field of system biology, which is crucial to understanding the cellular organization and inferring protein functions. Recently, many computational methods have been proposed to detect protein complexes from protein-protein interaction (PPI) networks. However, most of these methods only focus on local information of proteins in the PPI network, which are easily affected by the noise in the PPI network. Meanwhile, it's still challenging to detect protein complexes, especially for overlapping cases. To address these issues, we propose a new method, named Dopcc, to detect overlapping protein complexes by constructing a multi-metrics network according to different strategies. First, we adopt the Jaccard coefficient to measure the neighbor similarity between proteins and denoise the PPI network. Then, we propose a new strategy, integrating hierarchical compressing with network embedding, to capture the high-order structural similarity between proteins. Further, a new co-core attachment strategy is proposed to detect overlapping protein complexes from multi-metrics. The experimental results show that our proposed method, Dopcc, outperforms the other eight state-of-the-art methods in terms of F-measure, MMR, and Composite Score on two yeast datasets.
Wenkang Wang, Xiangmao Meng, Ju Xiang, Hayat Dino Bedru, Min Li 0007
IEEE ACM Trans. Comput. Biol. Bioinform.1
2023 Protein function prediction using graph neural network with multi-type biological knowledge
abstract
Proteins play crucial roles in diverse biological functions, and accurately annotating their functions is essential for understanding cellular mechanisms and developing therapies for complex diseases. Computational methods have been proposed as alternatives to laborious experimental approaches. However, existing network-based methods focus on the protein-protein interaction (PPI) networks, while the proteins without interactions are ignored. To address this limitation, we propose a novel deep learning framework for protein function prediction, named PFP-GMB, which incorporates multi-type biological knowledge to consider the proteins not present in the PPI networks. PFP-GMB leverages a pre-trained protein language model to extract sequence representations. Moreover, PPIs and orthology relationships are used to generate functional related features via graph neural networks and attention mechanisms. Finally, these multi-type features are fused for protein function prediction. Compared to eight state-of-the-art methods, PFP-GMB outperforms all of them in terms of F-max and AUPR. The ablation studies further confirm the relevance and significance of the multi-type biological knowledge incorporated into PFP-GMB for protein function prediction.
Yuyan Shuai, Wenkang Wang, Min Zeng 0004, Min Li 0007
BIBM2
2023 A review of enzyme design in catalytic stability by artificial intelligence
abstract
The design of enzyme catalytic stability is of great significance in medicine and industry. However, traditional methods are time-consuming and costly. Hence, a growing number of complementary computational tools have been developed, e.g. ESMFold, AlphaFold2, Rosetta, RosettaFold, FireProt, ProteinMPNN. They are proposed for algorithm-driven and data-driven enzyme design through artificial intelligence (AI) algorithms including natural language processing, machine learning, deep learning, variational autoencoder/generative adversarial network, message passing neural network (MPNN). In addition, the challenges of design of enzyme catalytic stability include insufficient structured data, large sequence search space, inaccurate quantitative prediction, low efficiency in experimental validation and a cumbersome design process. The first principle of the enzyme catalytic stability design is to treat amino acids as the basic element. By designing the sequence of an enzyme, the flexibility and stability of the structure are adjusted, thus controlling the catalytic stability of the enzyme in a specific industrial environment or in an organism. Common indicators of design goals include the change in denaturation energy (ΔΔG), melting temperature (ΔTm), optimal temperature (Topt), optimal pH (pHopt), etc. In this review, we summarized and evaluated the enzyme design in catalytic stability by AI in terms of mechanism, strategy, data, labeling, coding, prediction, testing, unit, integration and prospect.
Yongfan Ming, Wenkang Wang, Rui Yin 0002, Min Zeng 0004, Shizhe Tang, Min Li 0007
Briefings Bioinform.2
2023 CACO: A Core-Attachment Method With Cross-Species Functional Ortholog Information to Detect Human Protein Complexes
abstract
Protein complexes play an essential role in living cells. Detecting protein complexes is crucial to understand protein functions and treat complex diseases. Due to high time and resource consumption of experiment approaches, many computational approaches have been proposed to detect protein complexes. However, most of them are only based on protein-protein interaction (PPI) networks, which heavily suffer from the noise in PPI networks. Therefore, we propose a novel core-attachment method, named CACO, to detect human protein complexes, by integrating the functional information from other species via protein ortholog relations. First, CACO constructs a cross-species ortholog relation matrix and transfers GO terms from other species as a reference to evaluate the confidence of PPIs. Then, a PPI filter strategy is adopted to clean the PPI network and thus a weighted clean PPI network is constructed. Finally, a new effective core-attachment algorithm is proposed to detect protein complexes from the weighted PPI network. Compared to other thirteen state-of-the-art methods, CACO outperforms all of them in terms of F-measure and Composite Score, showing that integrating ortholog information and the proposed core-attachment algorithm are effective in detecting protein complexes.
Wenkang Wang, Xiangmao Meng, Ju Xiang, Yunyan Shuai, Hayat Dino Bedru, Min Li 0007
IEEE J. Biomed. Health Informatics1
2021 Overlapping Protein Complexes Detection Based on Multi-level Topological Similarities
Wenkang Wang, Xiangmao Meng, Ju Xiang, Min Li 0007
ISBRA1