EDBT 2026 Demo / reviewers in the wild / expert
Wenyi Yang
dblp:122/5963
· DBLP profile ↗
10ranked-venue papers
2as first author
8since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 8 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AI-driven computational methods and benchmarking for T-cell antigen identificationabstractThe rise of mRNA vaccines highlights the pivotal role of T-cell antigen identification in modern vaccinology and personalized medicine. T-cell recognition relies on the sophisticated ternary interaction between the T-cell receptor (TCR), the major histocompatibility complex (MHC) molecule, and the peptide antigen, which forms the peptide-MHC (pMHC) complex. Computational methods, particularly artificial intelligence (AI), are indispensable for accurately predicting these complex bindings. This review systematically surveys the rapidly evolving AI-driven landscape for T-cell antigen identification, providing a comprehensive categorization of methods for MHC-I, MHC-II, and the highly complex TCR-pMHC binding prediction, alongside foundational data resources. Crucially, we conduct a rigorous, standardized benchmarking of 18 state-of-the-art TCR-pMHC prediction models across diverse training data sources. Our evaluation on two distinct and challenging out-of-distribution (OOD) unseen epitope variant datasets reveals a significant and concerning generalization gap in current predictors. Notably, the overall absolute predictive gain remains marginal across all models under OOD conditions. This result underscores a severe and persistent generalization challenge when faced with novel epitope variants. To address these limitations, we emphasize the urgent need for enhanced structural modeling, the integration of multi-omics data, and the development of generative models for de novo TCR design. By advancing these computational frontiers, our community can accelerate the transition from prediction to rational design in immunoinformatics. Jinhao Que, Guangfu Xue, Yideng Cai, Wenyi Yang, Yi Hui, Zuxiang Wang, Wenyang Zhou, Qinghua Jiang, Haoxiu Sun |
Briefings Bioinform. | 5 |
| 2025 | TriCLFF: a multi-modal feature fusion framework using contrastive learning for spatial domain identificationabstractSpatial transcriptomics (ST) encompasses rich multi-modal information related to cell state and organization. Precisely identifying spatial domains with consistent gene expression patterns and histological features is a critical task in ST analysis, which requires comprehensive integration of multi-modal information. Here, we propose TriCLFF, a contrastive learning-based multi-modal feature fusion framework, to effectively integrate spatial associations, gene expression levels, and histological features in a unified manner. Leveraging an advanced feature fusion mechanism, our proposed TriCLFF framework outperforms existing state-of-the-art methods in terms of accuracy and robustness across four datasets (mouse brain anterior, mouse olfactory bulb, human dorsolateral prefrontal cortex, and human breast cancer) from different platforms (10x Visium and Stereo-seq) for spatial domain identification. TriCLFF also facilitates the identification of finer-grained structures in breast cancer tissues and detects previously unknown gene expression patterns in the human dorsolateral prefrontal cortex, providing novel insights for understanding tissue functions. Overall, TriCLFF establishes an effective paradigm for integrating spatial multi-modal data, demonstrating its potential for advancing ST research. The source code of TriCLFF is available online at https://github.com/HBZZ168/TriCLFF. Fenglan Pang, Guangfu Xue, Wenyi Yang, Yideng Cai, Jinhao Que, Haoxiu Sun, Shuaiyu Su, Xiyun Jin, Zuxiang Wang, Meng Luo 0001, Renjie Tan, Yusong Liu, Qinghua Jiang |
Briefings Bioinform. | 3 |
| 2025 | UAV Assisted Integrated Sensing and Communication for Mobile VehiclesabstractSince uncrewed aerial vehicles (UAVs) possess inherent characteristics such as exceptional maneuverability and versatile deployment, they can offer integrated sensing and communication (ISAC) services to vehicles in mobile environment. This paper designs a UAV-assisted ISAC system model, wherein the UAV is employed to provide sensing and communication services to mobile vehicles during its flight. In order to evaluate the radar detection performance of the ISAC system, we introduce radar mutual information (MI) from the information theory perspective. A resource optimization problem for the system model is formulated, which seeks to maximize the system communication rate under the constraints of signal-to-noise ratio (SNR) and MI of the radar detection link by jointly optimizing ISAC task scheduling, UAV transmit power allocation and UAV flight trajectory. The simulation results indicate that the proposed scheme significantly improves both the communication rate and radar MI. Xin Liu 0009, Wenyi Yang, Zechen Liu, Yuemin Liu, Feng Li 0008 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2025 | Joint Sensing and Age of Information Optimization for Energy Constrained UAV-Assisted Integrated Sensing, Calculation, and CommunicationabstractOwing to the advantages of high mobility, low cost, and on-demand deployment, uncrewed aerial vehicles (UAVs) can serve as triple-function aerial service platforms, providing sensing, calculation, and communication services for ground users in remote areas or emergencies. In this paper, a UAV-assisted integrated sensing, calculation, and communication (ISCC) system is proposed, where the UAV detects and processes the status information of the sensing target, and then sends the calculation results to the data collection center. In order to evaluate the performance of ISCC system, the age of information (AoI) and the radar estimation rate are introduced to define the freshness and amount of sensing data, respectively. Taking into account the UAV energy limitations, the amount of sensing data is maximized while the AoI is minimized through jointly optimizing the sensing scheduling, sensing times, transmit power, operating frequency, and motion parameters of the UAV under the constraint of radar signal-to-noise ratio (SNR). The formulated mixed-integer nonlinear programming problem is decomposed into five subproblems, and the optimal solutions can be achieved by proposing an alternating optimization (AO)-based five-stages optimization algorithm to optimize these subproblems iteratively. Simulation results show that both the sensing performance and information freshness of the system can be effectively improved by optimizing the UAV parameters. Zechen Liu, Xin Liu 0009, Wenyi Yang |
IEEE Trans. Wirel. Commun. | 3 |
| 2023 | DeepCCI: a deep learning framework for identifying cell-cell interactions from single-cell RNA sequencing dataabstractMOTIVATION: Cell-cell interactions (CCIs) play critical roles in many biological processes such as cellular differentiation, tissue homeostasis, and immune response. With the rapid development of high throughput single-cell RNA sequencing (scRNA-seq) technologies, it is of high importance to identify CCIs from the ever-increasing scRNA-seq data. However, limited by the algorithmic constraints, current computational methods based on statistical strategies ignore some key latent information contained in scRNA-seq data with high sparsity and heterogeneity. RESULTS: Here, we developed a deep learning framework named DeepCCI to identify meaningful CCIs from scRNA-seq data. Applications of DeepCCI to a wide range of publicly available datasets from diverse technologies and platforms demonstrate its ability to predict significant CCIs accurately and effectively. Powered by the flexible and easy-to-use software, DeepCCI can provide the one-stop solution to discover meaningful intercellular interactions and build CCI networks from scRNA-seq data. AVAILABILITY AND IMPLEMENTATION: The source code of DeepCCI is available online at https://github.com/JiangBioLab/DeepCCI. Wenyi Yang, Meng Luo 0001, Yideng Cai, Guangfu Xue, Xiyun Jin, Rui Cheng 0003, Jinhao Que, Fenglan Pang, Huan Nie, Qinghua Jiang |
Bioinform. | 1 |
| 2022 | CMMD: Cross-Metric Multi-Dimensional Root Cause AnalysisabstractIn large-scale online services, crucial metrics, a.k.a., key performance indicators (KPIs), are monitored periodically to check the running statuses. Generally, KPIs are aggregated along multiple dimensions and derived by complex calculations among fundamental metrics from the raw data. Once abnormal KPI values are observed, root cause analysis (RCA) can be applied to identify the reasons for anomalies, so that we can troubleshoot quickly. Recently, several automatic RCA techniques were proposed to localize the related dimensions (or a combination of dimensions) to explain the anomalies. However, their analyses are limited to the data on the abnormal metric and ignore the data of other metrics which are also related to the anomalies, leading to imprecise or even incorrect root causes. To this end, we propose a cross-metric multi-dimensional root cause analysis method, named CMMD, which consists of two key components: 1) relationship modeling, which utilizes graph neural network (GNN) to model the unknown complex calculation among metrics and aggregation function among dimensions from historical data; 2) root cause localization, which adopts the genetic algorithm to efficiently and effectively dive into the raw data and localize the abnormal dimension(s) once the KPI anomalies are detected. Experiments on synthetic datasets, real-world datasets and online production environments demonstrate the superiority of our proposed CMMD method compared with baselines. Currently, CMMD is running as an online service in Microsoft Azure. Shifu Yan, Wenyi Yang, Bixiong Xu, Dongsheng Li 0002, Lili Qiu, Jie Tong, Qi Zhang 0001 |
KDD | 3 |
| 2022 | CBLRR: a cauchy-based bounded constraint low-rank representation method to cluster single-cell RNA-seq dataabstractThe rapid development of single-cel+l RNA sequencing (scRNA-seq) technology provides unprecedented opportunities for exploring biological phenomena at the single-cell level. The discovery of cell types is one of the major applications for researchers to explore the heterogeneity of cells. Some computational methods have been proposed to solve the problem of scRNA-seq data clustering. However, the unavoidable technical noise and notorious dropouts also reduce the accuracy of clustering methods. Here, we propose the cauchy-based bounded constraint low-rank representation (CBLRR), which is a low-rank representation-based method by introducing cauchy loss function (CLF) and bounded nuclear norm regulation, aiming to alleviate the above issue. Specifically, as an effective loss function, the CLF is proven to enhance the robustness of the identification of cell types. Then, we adopt the bounded constraint to ensure the entry values of single-cell data within the restricted interval. Finally, the performance of CBLRR is evaluated on 15 scRNA-seq datasets, and compared with other state-of-the-art methods. The experimental results demonstrate that CBLRR performs accurately and robustly on clustering scRNA-seq data. Furthermore, CBLRR is an effective tool to cluster cells, and provides great potential for downstream analysis of single-cell data. The source code of CBLRR is available online at https://github.com/Ginnay/CBLRR. Wenyi Yang, Meng Luo 0001, Fenglan Pang, Yideng Cai, Anastasya A. Anashkina, Xi Su, Qinghua Jiang |
Briefings Bioinform. | 2 |
| 2021 | DLpTCR: an ensemble deep learning framework for predicting immunogenic peptide recognized by T cell receptorabstractAccurate prediction of immunogenic peptide recognized by T cell receptor (TCR) can greatly benefit vaccine development and cancer immunotherapy. However, identifying immunogenic peptides accurately is still a huge challenge. Most of the antigen peptides predicted in silico fail to elicit immune responses in vivo without considering TCR as a key factor. This inevitably causes costly and time-consuming experimental validation test for predicted antigens. Therefore, it is necessary to develop novel computational methods for precisely and effectively predicting immunogenic peptide recognized by TCR. Here, we described DLpTCR, a multimodal ensemble deep learning framework for predicting the likelihood of interaction between single/paired chain(s) of TCR and peptide presented by major histocompatibility complex molecules. To investigate the generality and robustness of the proposed model, COVID-19 data and IEDB data were constructed for independent evaluation. The DLpTCR model exhibited high predictive power with area under the curve up to 0.91 on COVID-19 data while predicting the interaction between peptide and single TCR chain. Additionally, the DLpTCR model achieved the overall accuracy of 81.03% on IEDB data while predicting the interaction between peptide and paired TCR chains. The results demonstrate that DLpTCR has the ability to learn general interaction rules and generalize to antigen peptide recognition by TCR. A user-friendly webserver is available at http://jianglab.org.cn/DLpTCR/. Additionally, a stand-alone software package that can be downloaded from https://github.com/jiangBiolab/DLpTCR. Meng Luo 0001, Weizhong Lin, Guangfu Xue, Xiyun Jin, Wenyang Zhou, Yideng Cai, Wenyi Yang, Huan Nie, Qinghua Jiang |
Briefings Bioinform. | 10 |
| 2019 | PNAB: Prediction of protein-nucleic acid binding affinity using heterogeneous ensemble modelsabstractProtein-nucleic acid interactions play critical roles in many biological processes. Quantifying the binding affinity of protein-nucleic acid complexes is helpful to the understanding of protein-nucleic acid recognition mechanism and identification of reliable binding partners. In this paper, we propose a computational approach, PNAB, which can effectively predict protein-nucleic acid binding affinity using heterogeneous ensemble models based on sequence. We build a dataset of protein-nucleic acid binding affinity that includes 103 protein-RNA complex and 100 protein-DNA complexes manually collected from related literature. We find that the binding affinity mainly depends on the structure of nucleic acid molecules. According to the type of nucleic acid associated with proteins composed of the protein-nucleic acid complex, we classify the complexes divide all the complexes into 11 categories (six classes for protein-RNA complexes and five classes for protein-DNA complexes). Then, we extract sequence features from the protein-nucleic acid complexes and build a stacking heterogeneous ensemble model based on the generated features for each category. We perform a comprehensive evaluation for the proposed method on the binding affinity dataset using leave-one-out cross-validation, and we show that PNAB achieves correlations ranging from 0.84 to 0.95 among all of the categories, which is significantly better than other typical regression methods and the pioneer protein-nucleic acid binding affinity predictor. Also, a user-friendly web server has been developed to predict the binding affinity of protein-RNA complexes. The PNAB web server is freely available at http://pnab.denglab.org/. Wenyi Yang |
BIBM | 1 |
| 2018 | PDRLGB: precise DNA-binding residue prediction using a light gradient boosting machineabstractBACKGROUND: Identifying specific residues for protein-DNA interactions are of considerable importance to better recognize the binding mechanism of protein-DNA complexes. Despite the fact that many computational DNA-binding residue prediction approaches have been developed, there is still significant room for improvement concerning overall performance and availability. RESULTS: Here, we present an efficient approach termed PDRLGB that uses a light gradient boosting machine (LightGBM) to predict binding residues in protein-DNA complexes. Initially, we extract a wide variety of 913 sequence and structure features with a sliding window of 11. Then, we apply the random forest algorithm to sort the features in descending order of importance and obtain the optimal subset of features using incremental feature selection. Based on the selected feature set, we use a light gradient boosting machine to build the prediction model for DNA-binding residues. Our PDRLGB method shows better overall predictive accuracy and relatively less training time than other widely used machine learning (ML) methods such as random forest (RF), Adaboost and support vector machine (SVM). We further compare PDRLGB with various existing approaches on the independent test datasets and show improvement in results over the existing state-of-the-art approaches. CONCLUSIONS: PDRLGB is an efficient approach to predict specific residues for protein-DNA interactions. Lei Deng 0002, Juan Pan, Wenyi Yang, Chuyao Liu, Hui Liu 0026 |
BMC Bioinform. | 4 |