VLDB 2026 Research / reviewers in the wild / expert
Weiwei Xue
dblp:28/10841
· DBLP profile ↗
13ranked-venue papers
0as first author
5since 2021 · last 2025
0000-0002-3285-0574ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 13 · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ADOptDiff - an Affinity Driven R-Chain Diffusion Model for Lead Compounds OptimizationabstractAccelerating drug discovery requires more efficient strategies to optimize lead compounds. Scaffold decoration is a key method in lead optimization, helping to improve ligand binding affinity and synthetic feasibility while retaining the core structure. However, most of the existing scaffold decoration models are based on ligands and cannot adequately consider the target, as well as the affinity between the target and the ligands. In this work, we present ADOptDiff, a novel affinity-driven scaffold decoration framework grounded in diffusion probabilistic models. The model incorporates E(3)-equivariant graph neural networks, conditional generative diffusion processes, and an Affinity Driven Cross-Attention module (ADCA), enabling context-sensitive scaffold decoration within target pockets and facilitating affinity-driven optimization. We constructed a dataset of 75,348 entries from the BindingNet platform, each comprising a protein pocket, scaffold, R-chain, and corresponding ligand–target binding affinity. Compared with existing molecular generation models, benchmark results indicate that ADOptDiff outperforms across multiple key metrics. In particular, it achieves about 5% improvement in enhancing protein–ligand binding interactions, demonstrating strong potential for affinity-driven molecular design. Further case studies involving targets from the central nervous system, viral enzymes, and membrane-associated proteins, validated by molecular dynamics simulations, confirm the structural stability and binding efficacy of the generated molecules. We have made all the code and a subset of the data available at https://github.com/chenshengneng/ADOptDiff.git. Shengneng Chen, Wanghong Fu, Yingjun Chen, Gao Tu, Xiaohong Zhang 0002, Weiwei Xue |
BIBM | 7 |
| 2025 | Artificial intelligence-driven framework for discovering synthetic binding protein-like scaffolds from the entire protein universeabstractCompared to traditional sequence-based methods, artificial intelligence (AI) approaches offer distinct advantages, such as significantly improved structural recognition efficiency and the ability to overcome inherent limitations of sequence alignment. Here, we introduce an AI-driven framework designed to discover synthetic binding proteins (SBPs)-like scaffolds from the entire known proteome. The framework integrates a deep learning-based FoldSeek with our in-house developed holistic protein attributes assessment (HP2A) algorithm, and enables subsequent protein function annotation and evolutionary analysis. As a proof-of-concept, four representative SBPs, including Affibody, Anticalin, DARPin, and Fynome, were used as query to discover SBP-like scaffolds. The results demonstrate that some of the identified SBP-like proteins, despite their low sequence similarity (identity ≤0.3), exhibit significant structural resemblance to the templates (template modeling score (TM-score) ≥ 0.5), highlighting the large sequence space available within specific protein scaffold. Statistical analysis identifies key biophysical properties that contribute to privileged scaffold functionality. Additionally, evolutionary insights derived from potential SBP-like scaffolds provide valuable guidance for protein binder design, as validated through targeted sequence analysis and in silico site-directed mutagenesis. This work highlights the potential of our framework to facilitate the discovery of high-quality engineered protein scaffolds, paving the way for the development of novel SBPs. Zixin Duan, Yafeng Liang, Liangcai Gu, Weiwei Xue |
Briefings Bioinform. | 6 |
| 2025 | Expanding the sequence spaces of synthetic binding protein using deep learning-based framework ProteinMPNN
Wantong Jiao, Ruihan Liu, Xuejin Deng, Feng Zhu 0004, Weiwei Xue |
Frontiers Comput. Sci. | 6 |
| 2025 | Multiscale Motif-Aware Relation Graph Structure for Drug-Target Binding Affinity PredictionabstractExploring drug-target binding affinity (DTA) is essential for drug discovery. Numerous works rely on the one-dimensional SMILES representation of drugs for predicting drug-target affinity, but ignore the crucial structural information of drug molecules. Considering structural information is an important factor in determining the affinity properties of drugs, we propose using the multiscale motif-aware relation graph (MMRG) rather than the SMILES representation to build drug descriptors. MMRG explicitly provides crucial motif-level structural and topological information of drugs, thereby ameliorating the predictive power of models. With this idea, we propose a novel MMRG construction approach for drugs, including a multiscale motif-aware learning for extracting motif-level structural information from motifs with different sizes and a relation graph construction approach for extracting topological information from chemical bonds. We implement a graph convolutional network to learn from MMRGs, and the learned latent features are used to predict drug-target affinity. The experiment results on the Davis and KIBA dataset report that our model can significantly outperform existing methods in drug-target affinity prediction, with an average improvement of 15.61% and 8.50% respectively. We further explain the principle of MMRG to improve the drug-target affinity prediction accuracy by comparing its generative function with the state-of-the-art method. Xiaohong Zhang 0002, Mianyang Yu, Yuqin Xia, Weiwei Xue |
IEEE Trans. Comput. Biol. Bioinform. | 5 |
| 2024 | MoDAFold: a strategy for predicting the structure of missense mutant protein based on AlphaFold2 and molecular dynamicsabstractProtein structure prediction is a longstanding issue crucial for identifying new drug targets and providing a mechanistic understanding of protein functions. To enhance the progress in this field, a spectrum of computational methodologies has been cultivated. AlphaFold2 has exhibited exceptional precision in predicting wild-type protein structures, with performance exceeding that of other methods. However, predicting the structures of missense mutant proteins using AlphaFold2 remains challenging due to the intricate and substantial structural alterations caused by minor sequence variations in the mutant proteins. Molecular dynamics (MD) has been validated for precisely capturing changes in amino acid interactions attributed to protein mutations. Therefore, for the first time, a strategy entitled 'MoDAFold' was proposed to improve the accuracy and reliability of missense mutant protein structure prediction by combining AlphaFold2 with MD. Multiple case studies have confirmed the superior performance of MoDAFold compared to other methods, particularly AlphaFold2. Lingyan Zheng, Shuiyang Shi, Xiuna Sun, Mingkun Lu, Yang Liao, Sisi Zhu, Hongning Zhang, Pan Fang, Zhenyu Zeng, Honglin Li 0003, Zhaorong Li, Weiwei Xue, Feng Zhu 0004 |
Briefings Bioinform. | 13 |
| 2020 | Genome-wide identification and analysis of the eQTL lncRNAs in multiple sclerosis based on RNA-seq dataabstractThe pathogenesis of multiple sclerosis (MS) is significantly regulated by long noncoding RNAs (lncRNAs), the expression of which is substantially influenced by a number of MS-associated risk single nucleotide polymorphisms (SNPs). It is thus hypothesized that the dysregulation of lncRNA induced by genomic variants may be one of the key molecular mechanisms for the pathology of MS. However, due to the lack of sufficient data on lncRNA expression and SNP genotypes of the same MS patients, such molecular mechanisms underlying the pathology of MS remain elusive. In this study, a bioinformatics strategy was applied to obtain lncRNA expression and SNP genotype data simultaneously from 142 samples (51 MS patients and 91 controls) based on RNA-seq data, and an expression quantitative trait loci (eQTL) analysis was conducted. In total, 2383 differentially expressed lncRNAs were identified as specifically expressing in brain-related tissues, and 517 of them were affected by SNPs. Then, the functional characterization, secondary structure changes and tissue and disease specificity of the cis-eQTL SNPs and lncRNA were assessed. The cis-eQTL SNPs were substantially and specifically enriched in neurological disease and intergenic region, and the secondary structure was altered in 17.6% of all lncRNAs in MS. Finally, the weighted gene coexpression network and gene set enrichment analyses were used to investigate how the influence of SNPs on lncRNAs contributed to the pathogenesis of MS. As a result, the regulation of lncRNAs by SNPs was found to mainly influence the antigen processing/presentation and mitogen-activated protein kinases (MAPK) signaling pathway in MS. These results revealed the effectiveness of the strategy proposed in this study and give insight into the mechanism (SNP-mediated modulation of lncRNAs) underlying the pathology of MS. Zhijie Han 0002, Weiwei Xue, Yan Lou, Yunqing Qiu, Feng Zhu 0004 |
Briefings Bioinform. | 2 |
| 2020 | Convolutional neural network-based annotation of bacterial type IV secretion system effectors with enhanced accuracy and reduced false discoveryabstractThe type IV bacterial secretion system (SS) is reported to be one of the most ubiquitous SSs in nature and can induce serious conditions by secreting type IV SS effectors (T4SEs) into the host cells. Recent studies mainly focus on annotating new T4SE from the huge amount of sequencing data, and various computational tools are therefore developed to accelerate T4SE annotation. However, these tools are reported as heavily dependent on the selected methods and their annotation performance need to be further enhanced. Herein, a convolution neural network (CNN) technique was used to annotate T4SEs by integrating multiple protein encoding strategies. First, the annotation accuracies of nine encoding strategies integrated with CNN were assessed and compared with that of the popular T4SE annotation tools based on independent benchmark. Second, false discovery rates of various models were systematically evaluated by (1) scanning the genome of Legionella pneumophila subsp. ATCC 33152 and (2) predicting the real-world non-T4SEs validated using published experiments. Based on the above analyses, the encoding strategies, (a) position-specific scoring matrix (PSSM), (b) protein secondary structure & solvent accessibility (PSSSA) and (c) one-hot encoding scheme (Onehot), were identified as well-performing when integrated with CNN. Finally, a novel strategy that collectively considers the three well-performing models (CNN-PSSM, CNN-PSSSA and CNN-Onehot) was proposed, and a new tool (CNN-T4SE, https://idrblab.org/cnnt4se/) was constructed to facilitate T4SE annotation. All in all, this study conducted a comprehensive analysis on the performance of a collection of encoding strategies when integrated with CNN, which could facilitate the suppression of T4SS in infection and limit the spread of antimicrobial resistance. Jiajun Hong, Yongchao Luo, Minjie Mou, Jianbo Fu, Yang Zhang 0125, Weiwei Xue, Yan Lou, Feng Zhu 0004 |
Briefings Bioinform. | 6 |
| 2020 | Protein functional annotation of simultaneously improved stability, accuracy and false discovery rate achieved by a sequence-based deep learningabstractFunctional annotation of protein sequence with high accuracy has become one of the most important issues in modern biomedical studies, and computational approaches of significantly accelerated analysis process and enhanced accuracy are greatly desired. Although a variety of methods have been developed to elevate protein annotation accuracy, their ability in controlling false annotation rates remains either limited or not systematically evaluated. In this study, a protein encoding strategy, together with a deep learning algorithm, was proposed to control the false discovery rate in protein function annotation, and its performances were systematically compared with that of the traditional similarity-based and de novo approaches. Based on a comprehensive assessment from multiple perspectives, the proposed strategy and algorithm were found to perform better in both prediction stability and annotation accuracy compared with other de novo methods. Moreover, an in-depth assessment revealed that it possessed an improved capacity of controlling the false discovery rate compared with traditional methods. All in all, this study not only provided a comprehensive analysis on the performances of the newly proposed strategy but also provided a tool for the researcher in the fields of protein function annotation. Jiajun Hong, Yongchao Luo, Yang Zhang 0125, Junbiao Ying, Weiwei Xue, Feng Zhu 0004 |
Briefings Bioinform. | 5 |
| 2020 | Clinical trials, progression-speed differentiating features and swiftness rule of the innovative targets of first-in-class drugsabstractDrugs produce their therapeutic effects by modulating specific targets, and there are 89 innovative targets of first-in-class drugs approved in 2004-17, each with information about drug clinical trial dated back to 1984. Analysis of the clinical trial timelines of these targets may reveal the trial-speed differentiating features for facilitating target assessment. Here we present a comprehensive analysis of all these 89 targets, following the earlier studies for prospective prediction of clinical success of the targets of clinical trial drugs. Our analysis confirmed the literature-reported common druggability characteristics for clinical success of these innovative targets, exposed trial-speed differentiating features associated to the on-target and off-target collateral effects in humans and further revealed a simple rule for identifying the speedy human targets through clinical trials (from the earliest phase I to the 1st drug approval within 8 years). This simple rule correctly identified 75.0% of the 28 speedy human targets and only unexpectedly misclassified 13.2% of 53 non-speedy human targets. Certain extraordinary circumstances were also discovered to likely contribute to the misclassification of some human targets by this simple rule. Investigation and knowledge of trial-speed differentiating features enable prioritized drug discovery and development. Yinghong Li 0002, Xiao Xu Li, Jiajun Hong, Jianbo Fu, Chun Yan Yu, Feng Cheng Li, Jie Hu 0020, Weiwei Xue, Yuzong Chen 0002, Feng Zhu 0004 |
Briefings Bioinform. | 10 |
| 2020 | ANPELA: analysis and performance assessment of the label-free quantification workflow for metaproteomic studiesabstractLabel-free quantification (LFQ) with a specific and sequentially integrated workflow of acquisition technique, quantification tool and processing method has emerged as the popular technique employed in metaproteomic research to provide a comprehensive landscape of the adaptive response of microbes to external stimuli and their interactions with other organisms or host cells. The performance of a specific LFQ workflow is highly dependent on the studied data. Hence, it is essential to discover the most appropriate one for a specific data set. However, it is challenging to perform such discovery due to the large number of possible workflows and the multifaceted nature of the evaluation criteria. Herein, a web server ANPELA (https://idrblab.org/anpela/) was developed and validated as the first tool enabling performance assessment of whole LFQ workflow (collective assessment by five well-established criteria with distinct underlying theories), and it enabled the identification of the optimal LFQ workflow(s) by a comprehensive performance ranking. ANPELA not only automatically detects the diverse formats of data generated by all quantification tools but also provides the most complete set of processing methods among the available web servers and stand-alone tools. Systematic validation using metaproteomic benchmarks revealed ANPELA's capabilities in 1 discovering well-performing workflow(s), (2) enabling assessment from multiple perspectives and (3) validating LFQ accuracy using spiked proteins. ANPELA has a unique ability to evaluate the performance of whole LFQ workflow and enables the discovery of the optimal LFQs by the comprehensive performance ranking of all 560 workflows. Therefore, it has great potential for applications in metaproteomic and other studies requiring LFQ techniques, as many features are shared among proteomic studies. Jianbo Fu, Bo Li 0093, Yinghong Li 0002, Qingxia Yang, Xuejiao Cui, Jiajun Hong, Yuzong Chen 0002, Weiwei Xue, Feng Zhu 0004 |
Briefings Bioinform. | 11 |
| 2020 | A critical assessment of the feature selection methods used for biomarker discovery in current metaproteomics studiesabstractMicrobial community (MC) has great impact on mediating complex disease indications, biogeochemical cycling and agricultural productivities, which makes metaproteomics powerful technique for quantifying diverse and dynamic composition of proteins or peptides. The key role of biostatistical strategies in MC study is reported to be underestimated, especially the appropriate application of feature selection method (FSM) is largely ignored. Although extensive efforts have been devoted to assessing the performance of FSMs, previous studies focused only on their classification accuracy without considering their ability to correctly and comprehensively identify the spiked proteins. In this study, the performances of 14 FSMs were comprehensively assessed based on two key criteria (both sample classification and spiked protein discovery) using a variety of metaproteomics benchmarks. First, the classification accuracies of those 14 FSMs were evaluated. Then, their abilities in identifying the proteins of different spiked concentrations were assessed. Finally, seven FSMs (FC, LMEB, OPLS-DA, PLS-DA, SAM, SVM-RFE and T-Test) were identified as performing consistently superior or good under both criteria with the PLS-DA performing consistently superior. In summary, this study served as comprehensive analysis on the performances of current FSMs and could provide a valuable guideline for researchers in metaproteomics. Jianbo Fu, Yongchao Luo, Ying Zhang 0061, Bo Li 0093, Qingxia Yang, Weiwei Xue, Yan Lou, Yunqing Qiu, Feng Zhu 0004 |
Briefings Bioinform. | 9 |
| 2020 | A novel bioinformatics approach to identify the consistently well-performing normalization strategy for current metabolomic studiesabstractUnwanted experimental/biological variation and technical error are frequently encountered in current metabolomics, which requires the employment of normalization methods for removing undesired data fluctuations. To ensure the 'thorough' removal of unwanted variations, the collective consideration of multiple criteria ('intragroup variation', 'marker stability' and 'classification capability') was essential. However, due to the limited number of available normalization methods, it is extremely challenging to discover the appropriate one that can meet all these criteria. Herein, a novel approach was proposed to discover the normalization strategies that are consistently well performing (CWP) under all criteria. Based on various benchmarks, all normalization methods popular in current metabolomics were 'first' discovered to be non-CWP. 'Then', 21 new strategies that combined the 'sample'-based method with the 'metabolite'-based one were found to be CWP. 'Finally', a variety of currently available methods (such as cubic splines, range scaling, level scaling, EigenMS, cyclic loess and mean) were identified to be CWP when combining with other normalization. In conclusion, this study not only discovered several strategies that performed consistently well under all criteria, but also proposed a novel approach that could ensure the identification of CWP strategies for future biological problems. Qingxia Yang, Jiajun Hong, Weiwei Xue, Feng Zhu 0004 |
Briefings Bioinform. | 4 |
| 2020 | Consistent gene signature of schizophrenia identified by a novel feature selection strategy from comprehensive sets of transcriptomic dataabstractThe etiology of schizophrenia (SCZ) is regarded as one of the most fundamental puzzles in current medical research, and its diagnosis is limited by the lack of objective molecular criteria. Although plenty of studies were conducted, SCZ gene signatures identified by these independent studies are found highly inconsistent. As one of the most important factors contributing to this inconsistency, the feature selection methods used currently do not fully consider the reproducibility among the signatures discovered from different datasets. Therefore, it is crucial to develop new bioinformatics tools of novel strategy for ensuring a stable discovery of gene signature for SCZ. In this study, a novel feature selection strategy (1) integrating repeated random sampling with consensus scoring and (2) evaluating the consistency of gene rank among different datasets was constructed. By systematically assessing the identified SCZ signature comprising 135 differentially expressed genes, this newly constructed strategy demonstrated significantly enhanced stability and better differentiating ability compared with the feature selection methods popular in current SCZ research. Based on a first-ever assessment on methods' reproducibility cross-validated by independent datasets from three representative studies, the new strategy stood out among the popular methods by showing superior stability and differentiating ability. Finally, 2 novel and 17 previously reported transcription factors were identified and showed great potential in revealing the etiology of SCZ. In sum, the SCZ signature identified in this study would provide valuable clues for discovering diagnostic molecules and potential targets for SCZ. Qingxia Yang, Bo Li 0093, Xuejiao Cui, Jie Hu 0020, Yuzong Chen 0002, Weiwei Xue, Yan Lou, Yunqing Qiu, Feng Zhu 0004 |
Briefings Bioinform. | 9 |