VLDB 2026 Research / reviewers in the wild / expert
Jiayin Wang 0002
dblp:74/1572-2
· DBLP profile ↗
59ranked-venue papers
0as first author
48since 2021 · last 2027
0000-0002-3862-6557ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 32 · 25 since 2021Artificial intelligence and machine learning · 16 · 16 since 2021Databases, data management, data science and information retrieval · 6 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 since 2021Software engineering, systems software and programming languages · 4 · 4 since 2021Computer networks · 3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | TrafHILLM: Highway network traffic flow prediction with heterogeneous graph-based and instruction fine-tuned large language model
Hongrui Wang 0004, Shanchuan Yu, Jiayin Wang 0002, Xiaoyan Zhu 0003, Jiaxuan Li 0001, Yuchuan Du |
Expert Syst. Appl. | 3 |
| 2026 | Mosaic Pruning: A Hierarchical Framework for Generalizable Pruning of Mixture-of-Experts ModelsabstractSparse Mixture-of-Experts (SMoE) architectures have enabled a new frontier in scaling Large Language Models (LLMs), offering superior performance by activating only a fraction of their total parameters during inference. However, their practical deployment is severely hampered by substantial static memory overhead, as all experts must be loaded into memory. Existing post-training pruning methods, while reducing model size, often derive their pruning criteria from a single, general-purpose corpus. This leads to a critical limitation: a catastrophic performance degradation when the pruned model is applied to other domains, necessitating a costly re-pruning for each new domain. To address this generalization gap, we introduce Mosaic Pruning (MoP). The core idea of MoP is to construct a functionally comprehensive set of experts through a structured ``cluster-then-select" process. This process leverages a similarity metric that captures expert performance across different task domains to functionally cluster the experts, and subsequently selects the most representative expert from each cluster based on our proposed Activation Variability Score. Unlike methods that optimize for a single corpus, our proposed Mosaic Pruning ensures that the pruned model retains a functionally complementary set of experts, much like the tiles of a mosaic that together form a complete picture of the original model's capabilities, enabling it to handle diverse downstream tasks.Extensive experiments on various MoE models demonstrate the superiority of our approach. MoP significantly outperforms prior work, achieving a 7.24\% gain on general tasks and 8.92\% on specialized tasks like math reasoning and code generation. Mingkuan Zhao, Shuangyong Song, Xiaoyan Zhu 0003, Xin Lai 0003, Jiayin Wang 0002 |
AAAI | 6 |
| 2026 | Making Every Head Count: Sparse Attention Without the Speed-Performance Trade-offabstractThe design of Large Language Models (LLMs) has long been hampered by a fundamental conflict within their core attention mechanism: its remarkable expressivity is built upon a computational complexity of O(H·N²) that grows quadratically with the context size (N) and linearly with the number of heads (H). This standard implementation harbors significant computational redundancy, as all heads independently compute attention over the same sequence space. Existing sparse methods, meanwhile, often trade information integrity for computational efficiency. To resolve this efficiency-performance trade-off, we propose SPAttention, whose core contribution is the introduction of a new paradigm we term Principled Structural Sparsity. SPAttention does not merely drop connections but instead reorganizes the computational task by partitioning the total attention workload into balanced, non-overlapping distance bands, assigning each head a unique segment. This approach transforms the multi-head attention mechanism from H independent O(N²) computations into a single, collaborative O(N²) computation, fundamentally reducing complexity by a factor of H. The structured inductive bias compels functional specialization among heads, enabling a more efficient allocation of computational resources from redundant modeling to distinct dependencies across the entire sequence span. Extensive empirical validation on the OLMoE-1B-7B and 0.25B-1.75B model series demonstrates that while delivering an approximately two-fold increase in training throughput, its performance is on par with standard dense attention, even surpassing it on select key metrics, while consistently outperforming representative sparse attention methods including Longformer, Reformer, and BigBird across all evaluation metrics. Our work demonstrates that thoughtfully designed structural sparsity can serve as an effective inductive bias that simultaneously improves both computational efficiency and model performance, opening a new avenue for the architectural design of next-generation, high-performance LLMs. Mingkuan Zhao, Jiayin Wang 0002, Xin Lai 0003, Tianchen Huang, Yuheng Min, Xiaoyan Zhu 0003 |
AAAI | 3 |
| 2026 | IFHAGrec: Instruction-Finetuned Heterogeneous-Aware Graph Neural Network for Temporally Weighted Recommendation Model
Hongrui Wang 0004, Shanchuan Yu, Xiaoyan Zhu 0003, Guangtao Wang, Jiayin Wang 0002, Jiaxuan Li 0001, Jindong Jiang |
KSEM (1) | 5 |
| 2026 | PScnv: personalized self-normalizing CNV detection with a hierarchical multi-phase frameworkabstractMOTIVATION: Accurate detection of copy number variations (CNVs) from targeted panel sequencing remains challenging due to limited genomic coverage and pronounced sample-specific biases. Existing normalization strategies, including baseline-cohort, matched-control, and single-sample approaches, often struggle to balance noise suppression with adaptability, leading to inconsistent performance across heterogeneous samples. RESULTS: We present PScnv, a personalized self-normalizing framework for robust CNV detection from panel sequencing data. PScnv integrates a pre-built panel-of-normals (PoN) with sample-intrinsic stable chromosomes through ridge-regression normalization to generate individualized log2 ratio profiles with reduced systematic variation. CNVs are then identified using a hierarchical multi-phase segmentation pipeline incorporating z-score pre-partitioning, kernel-based correction, and circular binary segmentation. In 139 clinical tumor samples with orthogonal FISH validation at MET, ERBB2, and MTAP, PScnv showed improved accuracy and robustness over existing methods that do not require patient-matched normal samples, provided that a pre-built PoN cohort is available. AVAILABILITY: Source code is available for academic use at https://github.com/lvws/PScnv. Xuwen Wang, Zhili Chang, Wansheng Lv, Akhatov Akmal, Xamidov Munis, Xunbiao Liu, Shenjie Wang, Xiaoyan Zhu 0003, Chong Du, Shuqun Zhang, Jiayin Wang 0002 |
Bioinform. | 11 |
| 2026 | UF-CDDFM: A unified framework for code defect detection using multi-modal inputs and few-shot learningabstractContext: The detection of code defects is foundational to modern software development and maintenance, playing a critical role in ensuring software quality and security. However, as software systems grow in scale and complexity, the limitations of traditional static analysis and conventional machine learning techniques have become increasingly evident. These methods rely heavily on intricate, manual feature engineering and fail to capture dynamic runtime behavior, resulting in suboptimal accuracy and elevated error rates. Objective: To address these deficiencies, we propose UF-CDDFM, a unified framework for code defect detection that integrates multi-modal inputs, active learning, and state-of-the-art few-shot learning techniques. We aim to improve detection performance, reduce feature selection complexity and sample bias through active learning, and maintain practical efficiency in real-world development contexts. Methods: UF-CDDFM employs parallel encoding of source code, code annotations, and abstract syntax trees (ASTs) using large language models (LLMs) alongside multilayer perceptrons (MLPs) to derive robust, high-fidelity representations of code. To streamline feature selection and mitigate sample bias, an active learning component is introduced for automated identification of high-quality features. Addressing the pervasive challenge of data scarcity, we incorporate two complementary few-shot learning strategies-MAML for small-scale datasets and LEO for larger-scale settings to enhance overall generalization capability. Results: Empirical evaluations demonstrate that UF-CDDFM consistently outperforms existing methods, establishing new state-of-the-art detection rates: 72.04% for defect detection and 95.23% for clone detection. Crucially, these gains are achieved within resource-constrained computational environments, which highlights the practicality of the method. Conclusion: By fusing multi-modal code representations, active learning, and adaptive few-shot learning techniques, UF-CDDFM delivers significant improvements in detection accuracy and computational efficiency. This work offers a new paradigm for robust, scalable, and practical code defect and clone detection in modern software engineering. Xianglu Zhou, Tianxiang Cui, Xiaoyan Zhu 0003, Jiayin Wang 0002, Xin Lai 0003 |
Inf. Softw. Technol. | 4 |
| 2026 | Reaching Software Quality for Bioinformatics Applications: How Far Are We?abstractWith the rapid advancements in medicine, biology, and information technology, their deep integration has given rise to the emerging field of bioinformatics. In this process, high–throughput technologies such as genomics, transcriptomics, and proteomics have generated massive volumes of biological data. The biological significance of these data heavily relies on bioinformatics software for analysis and processing. Therefore, it is crucial for both scientific research and clinical applications to ensure the quality of bioinformatics software and avoiding errors or hidden defects. However, to date, no dedicated study has systematically analyzed the quality of bioinformatics software. we conduct a comprehensive empirical study that aggregates, synthesizes, and analyzes findings from 167 bioinformatics software projects. Following the Preferred Reporting Items for Systematic Review and Meta–Analysis (PRISMA) protocol, we extract and evaluate quality–related data to answer our research questions (RQs). Our analysis reveals several key findings. The quality of bioinformatics software requires significant improvement, with an average defect density approximately 11.8× higher than that of general-purpose software. Additionally, unlike traditional software domains, a considerable proportion of defects in bioinformatics software are related to annotations. These issues can lead developers to overlook potential security vulnerabilities or make incorrect fixes, thereby increasing the cost and complexity of subsequent code maintenance. Based on these findings, we further discuss the challenges faced by bioinformatics software and propose potential solutions. This paper lays a foundation for further research on software quality in the bioinformatics domain and offers actionable insights for researchers and practitioners alike. Xiaoyan Zhu 0003, Xin Lai 0003, Xin Lian, Hangyu Cheng, Jiayin Wang 0002 |
IEEE Trans. Software Eng. | 6 |
| 2025 | Multi-Label Ranking Loss Minimization for Matrix CompletionabstractThe common matrix completion methods minimize the rank of the matrix to be completed in addition to the Hamming loss between the incomplete and completed matrices. The rank of matrix measures the linear relation among the vectors of matrix, which may introduce ambiguity for data recovery. To cope with this issue, we extend multi-label ranking loss into matrix completion, and employ multi-label ranking loss minimization (MLRM) in this paper to exploit the relative correlation among matrix vectors. In MLRM, the original incomplete matrix is converted into a pairwise ranking matrix, and the approximation on this newly generated matrix can be viewed as a surrogate of multi-label ranking loss to replace the Hamming loss pattern in the existing methods. Extensive experiments demonstrate that MLRM outperforms the state-of-the-art matrix completion methods in varies of applications, including movie recommendation, drug-target interaction prediction and multi-label learning. Jiaxuan Li 0001, Xiaoyan Zhu 0003, Hongrui Wang 0004, Yu Zhang 0203, Xin Lai 0003, Jiayin Wang 0002 |
AAAI | 6 |
| 2025 | Cross-Project Defect Prediction Based on Feature Fusion and Local Domain AdaptationabstractCross-project defect prediction (CPDP) is hindered by distribution shifts between source and target projects, so models that excel in within-project software defect prediction (WPDP) often degrade across projects. We propose FLDP, which couples (i) local subset alignment selecting similar source-target file pairs via three file-level metrics and aligning only those subsets with (ii) sequence-graph feature fusion, where TLSTM encodes token sequences and TGCN encodes AST structure into a unified representation. Across 10 transfers on 7 projects, FLDP consistently outperforms classical and recent CPDP baselines in AUC/F1/MCC. Ablation shows both local alignment and fusion are necessary for the gains, and our analysis of selection metrics offers practical guidance for applying CPDP in heterogeneous settings. Xianglu Zhou, Xiaoyan Zhu 0003, Yu Wang 0069, Jiayin Wang 0002, Xin Lai 0003 |
APSEC | 4 |
| 2025 | EMcnv: enhancing CNV detection performance through ensemble strategies with heterogeneous meta-graph neural networksabstractCopy number variation (CNV) is a crucial biomarker for many complex traits and diseases. Although numerous CNV detection tools are available, no single method consistently achieves optimal performance across diverse sequencing samples, as each tool has distinct advantages and limitations. Therefore, integrating the strengths of these tools to improve CNV detection accuracy is both a promising strategy and a significant challenge. To address this, we propose EMcnv, a novel deep ensemble framework based on meta-learning. EMcnv combines multiple CNV detection strategies through a three-step approach: (i) leveraging meta-learning and meta-path heterogeneous graphs, employing Relational Graph Convolutional Networks as a specific model within the Heterogeneous Graph Neural Networks framework to develop a probabilistic weight meta-model that ensembles various CNV detection strategies; (ii) assigning probabilistic weights to calls from different CNV detection tools and aggregating them into weighted CNV regions (CNVRs); (iii) refining Copy number variations based on weighted CNVRs. We conducted comprehensive experiments on both simulated and real sequencing data using benchmark datasets. The results demonstrate that EMcnv significantly outperforms popular existing methods, underscoring its superiority and importance in CNV detection. To support further research, the source code is available for academic use at https://github.com/Sherwin-xjtu/EMcnv. Xuwen Wang, Zhili Chang, Yuqian Liu, Shenjie Wang, Xiaoyan Zhu 0003, Jiayin Wang 0002 |
Briefings Bioinform. | 7 |
| 2025 | THOR: a TMB heterogeneity-adaptive optimization model predicts immunotherapy response using clonal genomic features in group-structured dataabstractWith the increasing number of indications for immune checkpoint inhibitors in early and advanced cancers, the prospect of a tumor-agnostic biomarker to prioritize patients is compelling. Tumor mutation burden (TMB) is a widely endorsed biomarker that quantifies nonsynonymous mutations within tumor DNA, essential for neoantigen production, which, in turn, correlates with the immune response and guides decision-making. However, the general clinical application of TMB-relying on simple mutational counts targeted at a single endpoint-does not adequately capture the complex clonal structure of tumors nor the multifaceted nature of prognostic indicators. This recognition has spurred the exploration of sophisticated high-dimensional regression techniques. Unfortunately, the limited cohort sizes in immunotherapy trials have hindered the full potential of these advanced methods. Our approach considers patient subgroups as related yet distinct entities, enabling precise tailoring and refinement to address subgroup-specific dynamics. Given the deficiencies and the constraints, we introduce a TMB heterogeneity-optimized regression (THOR). This innovative model enhances the predictive capabilities of TMB by integrating tumor clonality and a diverse spectrum of clinical endpoints, further augmented by fusion techniques across subgroups to facilitate robust data sharing and interpretation. Our simulations validate THOR's superiority in parameter estimation for statistical inference. Clinically, we assess the utility of THOR in a structured cohort of 238 cancer patients undergoing immunotherapy, supplemented by 2212 patients across 19 subgroups from public datasets. The forecast of the responses and comparison of survival hazards demonstrate that THOR significantly enhances patient stratification and prognostic predictions by incorporating complex immunogenetic biology and subgroup-specific dynamics. Yanfang Guan, Xin Lai 0003, Yuqian Liu, Zhili Chang, Quan Wang 0004, Jian Zhao 0034, Shuanying Yang, Jiayin Wang 0002 |
Briefings Bioinform. | 11 |
| 2025 | MRDadaptis: self-adaptive parameter configuration enhances minimal residual disease detection in heterogeneous ctDNA samplesabstractDetection of structural variations (SVs) through circulating tumor DNA (ctDNA) has become a key method for detecting minimal residual disease (MRD). However, the heterogeneity of ctDNA samples, characterized by variable limits of detection (LOD) and diverse structural variant types, significantly impacts detection stability and performance, posing persistent challenges for conventional SV detection tools such as Delly and Manta. These widely used methods require extensive manual parameter tuning, hindered by the combinatorial complexity of multiple parameters and heterogeneous sequencing data. To address this, we propose MRDadaptis, a novel SV detection tool that uniquely incorporates a self-adaptive parameter optimization mechanism. MRDadaptis distinguishes itself by integrating Bayesian optimization with meta-learning techniques to dynamically adjust detection parameters automatically, based on intrinsic features derived from the ctDNA sequencing data itself. This innovative approach not only reduces manual intervention but also effectively captures sample-specific characteristics, significantly improving detection stability, and detection performance. Extensive validation experiments using both simulated and real-world ctDNA datasets demonstrates it distinct advantages, including markedly improved average F1-scores and superior stability (reduced variance, lower RMSE, increased kurtosis). These results highlight the significant advantages of MRDadaptis in addressing sample heterogeneity, underscoring its potential to improve the accuracy and reliability of MRD detecting through ctDNA analysis. https://github.com/aAT0047/MRDadaptis.git. Xin Lai 0003, Shenjie Wang, Zhengfa Xue, Yuqian Liu, Xiaoyan Zhu 0003, Zhili Chang, Jiayin Wang 0002 |
Briefings Bioinform. | 11 |
| 2025 | BioWorkflow: Retrieving comprehensive bioinformatics workflows from publicationsabstractReconstructing bioinformatics workflows from the literature is the foundation of scientific analysis. However, the required details-processing steps, software tools, versions, and parameter settings-are dispersed across narrative text, tables, figure captions, and supplemental files. Manual reconstruction typically takes hours per paper and is error-prone, while existing question-answering (QA) and retrieval systems focus on local passages and lack the full-text, multimodal capabilities needed to automatically rebuild complete workflows. We introduce BioWorkflow, a large language model (LLM)-based, retrieval-augmented framework that automates end-to-end workflow extraction from publications by (i) parsing PDFs and building a unified index over text, tables, and figures with chunk-level summaries and embeddings; (ii) hierarchically decomposing queries with dynamic reformulation when new entities or ambiguities emerge; (iii) performing iterative, context-aware retrieval and assembling a directed workflow that captures steps, tools, versions, and parameters; and (iv) linking each predicted element to its cited evidence and running automated consistency checks to suppress hallucinations and ensure traceability. Evaluated on 100 expert-annotated papers, BioWorkflow recovers ~80% of workflow steps (versus ~20% for existing tools), improves reproducibility, completeness, and accuracy by >20% over strong LLM baselines, and reduces curation time to 3-5 minutes per paper, enabling rapid and reliable reuse of published pipelines. Jiayin Wang 0002 |
Briefings Bioinform. | 2 |
| 2025 | TMBquant: an explainable AI-powered caller advancing tumor mutation burden quantification across heterogeneous samplesabstractAccurate tumor mutation burden (TMB) quantification is critical for immunotherapy stratification, yet remains challenging due to variability across sequencing platforms, tumor heterogeneity, and variant calling pipelines. Here, we introduce TMBquant, an explainable AI-powered caller designed to optimize TMB estimation through dynamic feature selection, ensemble learning, and automated strategy adaptation. Built upon the H2O AutoML framework, TMBquant integrates variant features, minimizes classification errors, and enhances both accuracy and stability across diverse datasets. We benchmarked TMBquant against nine widely used variant callers, including traditional tools (e.g. Mutect2, VarScan2, Strelka2) and recent AI-based methods (DeepSomatic, Octopus), using 706 whole-exome sequencing tumor-control pairs. To evaluate clinical relevance, we further assessed TMBquant through survival analyses across immunotherapy-treated cohorts of non-small cell lung cancer (NSCLC), nasopharyngeal carcinoma (NPC), and the two NSCLC subtypes: lung adenocarcinoma and lung squamous cell carcinoma. In each cohort, TMBquant consistently achieved the highest hazard ratios, demonstrating superior patient stratification compared to all other methods. Importantly, TMBquant maintained robust predictive performance across both high-TMB (NSCLC) and low-TMB (NPC) settings, highlighting its generalizability across cancer types with distinct biological characteristics. These findings establish TMBquant as a reliable, reproducible, and clinically actionable tool for precision oncology. The software is open source and freely available at https://github.com/SomaticCaller/SomaticCaller. To enhance reproducibility, we provide detailed usage instructions and representative code snippets for TMBquant in the Methods section (see Code Availability). Shenjie Wang, Xiaoyan Zhu 0003, Xuwen Wang, Yuqian Liu, Minchao Zhao, Zhili Chang, Shuanying Yang, Jiayin Wang 0002 |
Briefings Bioinform. | 11 |
| 2025 | TMBclaw: tumor clone-aware graph learning improves immunotherapy response prediction across heterogeneous cohortsabstractImmune checkpoint inhibitors (ICIs) have emerged as a cornerstone of modern oncology, necessitating the development of robust biomarkers for optimizing patient stratification and treatment selection. While tumor mutation burden (TMB) has demonstrated prognostic value, conventional quantification methods based on mutation counts fail to reflect immunogenic neoantigen presentation due to intratumoral clonal heterogeneity. Recent efforts have focused on mutation subsets derived from tumor clonality, yet the complex interactions among clones remain a significant obstacle to accurate prognosis. This challenge is further exacerbated by the inherent constraints of limited cohort sizes in clinical studies, which severely compromise model generalizability across heterogeneous cohorts. Therefore, we propose TMBclaw (Tumor Mutation Burden-based Clonal attention with Laplacian Adaptive Weighting), a graph-regularized multi-task learning framework for immunotherapy response prediction. TMBclaw establishes unified integration of group-structured cohorts while enabling cross-cohort knowledge transfer and clonal relationship exploration. For clinical validation, we utilized four cohorts of 238 patients with non-small-cell lung cancer (NSCLC), melanoma, or nasopharyngeal carcinoma treated with ICIs, along with external multicenter validation cohorts (N = 1433) of melanoma and NSCLC patients from public datasets. Comparative analyses demonstrate that TMBclaw significantly outperforms conventional methods in prognostic accuracy and risk stratification. Through systematic quantification of clonal dynamics and discriminative identification of driver clones, TMBclaw shows potential to improve understanding of tumor heterogeneity and provides interpretable insights into the immunotherapy process. Xiaoyan Zhu 0003, Zhili Chang, Xin Lai 0003, Jiayin Wang 0002 |
Briefings Bioinform. | 9 |
| 2025 | PangenomeX: a graph convolutional network-based pangenome framework for unbiased population-scale genomic variation analysisabstractIn population-scale genomic variation studies based on shallow whole genome sequencing, pangenomes have become an effective tool for identifying population-specific single-nucleotide polymorphisms and indels. Extending these advantages to copy number variation (CNV), however, remains challenging due to two unresolved issues. First, current pangenome frameworks exhibit pronounced population-representation bias arising from uneven sampling across populations. As the number of samples increases, the pangenome tends to capture variations primarily from majority populations while suppressing signals from minority populations. Second, in population-scale genomic variation analyses, common but benign population-specific copy number polymorphisms (CNPs) frequently obscure pathogenic CNVs. Existing pangenome frameworks lack dedicated mechanisms for representing CNPs and CNVs, limiting their ability to distinguish pathogenic CNVs from benign, population-specific CNPs. In this study, we present PangenomeX, a graph-convolutional pangenome framework tailored for low-coverage, population-scale CNV analysis. To address CNP representation, we embed known CNPs as prior knowledge into the pangenome graph and construct a CNV relationship network guided by a phylogenetic tree. A graph convolutional network (GCN) then learns the interactions between CNV and CNP nodes. To mitigate population-representation bias, the GCN aggregates information from only one- and two-hop neighborhoods, preserving local population context while preventing majority group signals from dominating. Evaluation on simulated cohorts and 561 real samples shows that PangenomeX distinguishes pathogenic CNVs from common population CNPs markedly better than existing methods. Overall, PangenomeX offers a methodological blueprint for large-cohort variant screening and provides a practical path for bringing graph-based genomics into clinical practice. Zhengfa Xue, Yu Wang 0069, Xuwen Wang, Jiajing Yuan, Jingyu Zeng, Huanhuan Zhu, Jiayin Wang 0002 |
Briefings Bioinform. | 9 |
| 2025 | MRDagent: iterative and adaptive parameter optimization for stable ctDNA-based MRD detection in heterogeneous samplesabstractMOTIVATION: Minimal residual disease (MRD) as critical biomarker for cancer prognosis and management plays a crucial role in improving patient outcomes. However, detecting MRD via next-generation sequencing-based circulating tumor DNA variant calling remains unstable due to the extremely low variant allele frequency and significant inter- and intra-sample heterogeneity. Although parameter optimization can theoretically enhance the detection performance of variants, achieving stable MRD detection remains challenging due to three key factors: (i) the necessity for individualized parameter tuning across numerous heterogeneous genomic intervals within each sample, (ii) the tightly interdependent parameter requirements across different stages of variant detection workflows, and (iii) the limitations of current automated parameter optimization methods. RESULTS: In this study, we propose MRDagent, a novel variant detection tool designed specifically for MRD detection. MRDagent incorporates an iterative and self-adaptive optimization framework capable of handling unknown objectives, varying constraints, and highly coupled parameters across stages. A key innovation of MRDagent is the integration of a convolutional neural network-based meta-model, trained on historical data to enable rapid parameter prediction. This significantly enhances computational efficiency and generalization performance. Extensive evaluations on simulated and real-world datasets demonstrate MRDagent's superior and stable performance, providing an efficient, reliable solution for MRD detection in clinical and high-throughput research applications. AVAILABILITY AND IMPLEMENTATION: MRDagent is freely available at https://github.com/aAT0047/MRDagent.git. The corresponding dataset and software archive are available at Zenodo: https://doi.org/10.5281/zenodo.15458496. Xin Lai 0003, Shenjie Wang, Yuqian Liu, Xiaoyan Zhu 0003, Jiayin Wang 0002 |
Bioinform. | 6 |
| 2025 | ZIPcnv: accurate and efficient inference of copy number variations from shallow whole-genome sequencingabstractMOTIVATION: Shallow whole-genome sequencing (sWGS), a rapid and cost-effective sequencing technology, has gradually been widely adopted for CNV analyses. However, with genome‑wide coverage of only 0.1-5×, sWGS data display a pronounced zero‑inflation phenomenon-a large fraction of loci has zero sequencing reads. Zero inflation causes read counts to fluctuate by several‑fold between adjacent windows. As a result, random upward blips in coverage can be misinterpreted as copy‑number gains (false positives), and true deletions often become indistinguishable from pervasive zero‑coverage noise. In addition, existing CNV detection tools developed for sWGS data often struggle to adapt across different CNV sizes. These combined effects severely constrain the accuracy of CNV inference. RESULTS: To address above challenges, we propose ZIPcnv, a novel CNV detection tool specifically designed for sWGS data. First, we apply a segment sliding window to smooth the raw read depth signal, which transforms the original zero-inflated statistical characteristics into approximately normal distribution characteristics. We then design a statistical process model that robustly detects persistent shifts under high background noise using a cumulative sum strategy, classifying genomic regions into candidate and non-candidate CNV regions. Finally, dynamic sliding windows are used for one-pass detection of CNVs of varying lengths, with window size adapting to the CNV region size. We evaluated the performance of ZIPcnv on simulated data and 190 real whole-genome sequencing samples. Experimental results show that ZIPcnv consistently outperforms currently popular CNV detection tools. AVAILABILITY AND IMPLEMENTATION: The ZIPcnv source code is freely available at https://github.com/Nevermore233/ZIPcnv. Zhengfa Xue, Jingyu Zeng, Xuwen Wang, Jiajing Yuan, Xin Lai 0003, Yu Wang 0069, Huanhuan Zhu, Jiayin Wang 0002 |
Bioinform. | 11 |
| 2025 | Learning multi-behavior user intent for session-based recommendation
Yu Zhang 0203, Xiaoyan Zhu 0003, Guopeng He, Jiaxuan Li 0001, Jiayin Wang 0002 |
Expert Syst. Appl. | 5 |
| 2025 | MRDtarget: A heuristic Gaussian approach for optimizing targeted capture regions to enhance Minimal Residual Disease detectionabstractMolecular residual disease (MRD) detection, initially developed for hematologic malignancies, has become a critical biomarker for monitoring solid tumors. MRD detection primarily relies on circulating tumor DNA (ctDNA) analysis using next-generation sequencing, offering high sensitivity and broad genomic coverage. However, challenges remain in designing cost-effective panels that maximize mutation detection while maintaining biological relevance. Fixed panels often lack sufficient patient-specific mutation coverage, while WES-based personalized MRD assays, despite their high sensitivity, are costly and less accessible. We developed a tumor comprehensive genomic profiling (CGP)-informed personalized MRD assay to detect tumor-derived mutations, which allowed us to design patient-specific personalized panels and meanwhile, provide a cost-effective alternative to whole exome sequencing (WES). To address these limitations, we developed MRDtarget, a heuristic multivariate Gaussian model-based targeted capture region selection method. By expanding beyond traditional hotspot regions, MRDtarget optimizes variant tracking for MRD detection, significantly improving sensitivity. Using a Bayesian inference-based heuristic approach, MRDtarget integrates multi-feature informativeness rates to identify optimal genomic regions for capture. Experimental results demonstrate that MRDtarget enables the detection of more variants per patient. This study underscores the importance of rational panel design to improve MRD sensitivity and provides a novel approach to enhance precision diagnostics and treatment for solid tumor patients. Xuwen Wang, Yanfang Guan, Xin Lai 0003, Wuqiang Cao, Xiaoyan Zhu 0003, Xiaoling Zeng, Yuqian Liu, Shenjie Wang, Ruoyu Liu, Shuanying Yang, Jiayin Wang 0002 |
PLoS Comput. Biol. | 13 |
| 2024 | Multi-Objective Policy Monitoring Method for Epidemic ControlabstractIn the face of emerging infectious diseases such as COVID-19, timely government intervention is crucial, as swift policy actions can effectively prevent greater losses. However, policymakers often need to balance multiple conflicting objectives. The lack of high-quality data and suitable analytical tools poses significant challenges for policy evaluation, especially in multi-objective decision-making, where accurately assessing the impact of interventions becomes even more difficult. To address this issue, this paper proposes a real-time data-driven policy monitoring method for dynamically tracking the effects of policy interventions. We introduce a new non-parametric one-sided test control chart, leveraging the interpretability and ease of implementation of control charts to monitor risk levels across various policies. Experimental results demonstrate the effectiveness of this method in policy monitoring. Xin Lai 0003, Rundong Fan, Ruoyu Liu, Jiayin Wang 0002, Xiaoyan Zhu 0003, Yuqian Liu, Xuwen Wang, Shenjie Wang |
BIBM | 4 |
| 2024 | An Enhanced Multiple Correction Method with Limited Independent Effective SNPsabstractIn genome-wide association studies (GWAS) and candidate gene studies (CGS), appropriate multiple testing correction methods can effectively address linkage disequilibrium (LD) blocks and are crucial for controlling the family-wise error rate (FWER) while ensuring the reliability of results. In our recent research, we observed that when the number of independent effective Single Nucleotide Polymorphisms (SNPs) is relatively small, existing multiple testing correction methods struggle to effectively control the FWER at 0.05 and lack robustness. This study proposes an enhanced method that builds upon the Moskvina and Schmidt approach by incorporating additional SNP correlation information, allowing for precise control of the FWER at 0.05 with limited independent effective SNPs. We also leverage Monte Carlo integration on a GPU to accelerate the computation of significance thresholds. Our method was evaluated through both simulation studies and real genotype datasets, demonstrating superior performance and robustness compared to existing methods. Xin Lai 0003, Xiaohai Yang, Jiayin Wang 0002, Xiaoyan Zhu 0003, Yuqian Liu, Ruoyu Liu, Xuwen Wang, Shenjie Wang |
BIBM | 3 |
| 2024 | Enhancing Dental Implant Risk Prediction with an Interpretable Multi-Instance Learning ModelabstractVariations in clinical and biological factors often lead to differing risks of dental implant failure among patients, even when undergoing similar procedures. Traditional predictive models often oversimplify outcomes into binary classifications, lacking the interpretability needed for accurate risk assessment and effective patient stratification. To address this challenge, this paper presents a novel multi-instance learning (MIL) framework that incorporates a Hosmer-Lemeshow-based loss function for implant failure risk assessment based on patient-specific clinical features. The framework also identifies the optimal threshold for key features, such as bone density, to enable robust patient stratification. The effectiveness label for each patient was constructed first and sampled patients into 600 and 1,000 groups. Results demonstrate that the proposed framework captures inter-patient variability with improved statistical calibration and enhanced risk stratification, offering practical insights to guide clinical decision-making. This study highlights the scalability and versatility of the proposed framework, bridging computational methodologies and practical applications in personalized implantology. Furthermore, the approach provides a transparent and effective tool for risk assessment, with potential applications in broader clinical stratification problems. Yuqian Liu, Almonzer Salah Nooraldaim, Zhengfa Xue, Jiayin Wang 0002 |
BIBM | 4 |
| 2024 | LMR-EWMA: A LASSO-based Multivariate Residual Control Chart for Monitoring Rare Health-Related EventsabstractMonitoring rare health-related events using control charts is crucial for timely detecting potential changes in healthcare scenarios. For example, sequentially testing the level of changes in infectious disease patient numbers helps prepare before an epidemic. Unlike general health-related events, the observation of rare ones often involves an excess of zeros, making it more appropriate to use the zero-inflated Poisson (ZIP) distribution rather than the classical Poisson. Although residual-based charts have attracted significant attention in this field, few studies have explored how to appropriately select residuals with different advantages in the complex situations like healthcare scenarios. Therefore, in this paper, we propose LMR-EWMA, a least absolute shrinkage and selection operator (LASSO)-based multivariate residual exponentially weighted moving average (EWMA) control chart, to automatically select the optimal residuals for monitoring changes (i.e., shifts) in the number of rare health-related events. Additionally, we have innovatively designed a bi-directional moving mechanism to address the limitation of current research in distinguishing the practical significance of shifts. Experimental results on three simulation cases and two real datasets demonstrate that LMR-EWMA outperforms existing charts in monitoring performance. Ruoyu Liu, Jiayin Wang 0002, Xiaoyan Zhu 0003, Yuqian Liu, Shuanying Yang, Xin Lai 0003 |
BIBM | 2 |
| 2024 | RMComBat: A Batch Effect Correction Algorithm for Repeated Measurement Sequencing Data to Prevent OvercorrectionabstractBatch effects, caused by non-biological variations such as differences in laboratory conditions, reagent lots, or personnel, are a substantial source of noise in gene expression data. Accurately correcting these effects is crucial for valid biological inferences. However, the majority of existing batch effect correction algorithms are prone to overcorrection, where biologically meaningful signals are mistakenly identified as noise, especially in repeated measurement studies where time is confounded with batch. The failure to accurately distinguish between batch-related and biologically relevant variation leads to a loss of critical biological information. This paper presents RMComBat, an enhancement of the widely-used ComBat framework, which addresses this limitation by replacing the general linear model with a linear mixed-effects model. RMComBat incorporates subject-specific random intercepts to correct for sample correlation and enhance the preservation of biological signals. We tested RMComBat and several popular algorithms on simulated and real repeated measurement gene expression datasets, evaluating their performance through visual inspections and quantitative metrics. Results indicate that although most algorithms can reduce batch effects, they often do so at the cost of removing true biological signals. RMComBat demonstrates superior performance in preventing overcorrection, providing a more balanced and biologically informative correction in repeated measurement studies, so making it a valuable tool for improving the accuracy of gene expression analyses. Yuqian Liu, Zhaoxing Wei, Jiayin Wang 0002, Xiaoyan Zhu 0003, Ruoyu Liu, Xuwen Wang, Shenjie Wang, Xin Lai 0003 |
BIBM | 3 |
| 2024 | Enabling Adaptive CNV Detection through A Novel Predictive Control FrameworkabstractAccurate detection of copy number variations (CNVs) from sequencing data is crucial in many complex traits and diseases research. Although many CNV detection algorithms have been developed, challenges in precisely identifying CNVs persist. The core statistical model of these algorithms cannot self-adjust, which limits their adaptability to heterogeneous samples and reduces detection accuracy. address this challenge, we reframed the CNV detection problem as a quality control issue and incorporated adaptive mechanisms. We developed adapCNV, a novel adaptive CNV detection framework that integrates machine learning with optimization control. This framework enables dynamic adaptation of primary parameters based on sample features. We defined a quantifiable metric, RD fluctuation values, to assess signal characteristics when the algorithm accurately detects CNVs. We then employed machine learning techniques extract features from panel sequencing data, select initial parameter values for samples, and determine optimal RD fluctuation values. By adopting adaptive model predictive control (AMPC), adapCNV performs optimizations within rolling window. It dynamically adjusts the primary parameters based on error feedback from RD fluctuation values. This adaptive control strategy enables dynamic adjustment automatically match the characteristics of panel sequencing samples, significantly enhancing overall detection quality. The performance of this framework was validated with simulated data. Comparative analysis demonstrated that the proposed method outperforms the baseline approach, particularly in detecting small CNVs. The adapCNV framework is particularly suitable for panel sequencing, which may have broad applications in clinical practice. This novel approach from quality control perspective introduces a new paradigm for CNV detection. Yuqian Liu, Jiajing Yuan, Xiaoyan Zhu 0003, Xin Lai 0003, Ruoyu Liu, Xuwen Wang, Jiayin Wang 0002 |
BIBM | 7 |
| 2024 | Correction of Read Biases Induced by Complex Reference Genome Regions for Improving Copy Number Variation Detection Using a Gaussian Mixture ModelabstractCopy number variations are crucial in cancer research, but their detection through next-generation sequencing is often hindered by read biases, particularly in complex genomic regions. Existing bias-correction methods address common issues like GC content but often fail in regions with repetitive sequences or segmental duplications, leading to false-positive CNVs. We propose refMask, a hybrid Gaussian model-based method that dynamically identifies low-confidence regions in the reference genome, correcting read biases and improving CNV detection accuracy. By integrating features from hg38 and T2T genomes, refMask tailors a custom blacklist for each sequencing sample, enhancing the reliability of CNV detection across diverse conditions. Our method provides a more accurate and flexible solution compared to current fixed blacklists, offering improved performance in challenging genomic regions. Xuwen Wang, Zhili Chang, Shenjie Wang, Ruoyu Liu, Yuqian Liu, Xiaoyan Zhu 0003, Xin Lai 0003, Shuanying Yang, Jiayin Wang 0002 |
BIBM | 10 |
| 2024 | An Enhanced Batch Query Architecture in Real-time RecommendationabstractIn industrial recommendation systems on websites and apps, it is essential to recall and predict top-n results relevant to user interests from a content pool of billions within milliseconds. To cope with continuous data growth and improve real-time recommendation performance, we have designed and implemented a high-performance batch query architecture for real-time recommendation systems. Our contributions include optimizing hash structures with a cacheline-aware probing method to enhance coalesced hashing, as well as the implementation of a hybrid storage key-value service built upon it. Our experiments indicate this approach significantly surpasses conventional hash tables in batch query throughput, achieving up to 90% of the query throughput of random memory access when incorporating parallel optimization. The support for NVMe, integrating two-tier storage for hot and cold data, notably reduces resource consumption. Additionally, the system facilitates dynamic updates, automated sharding of attributes and feature embedding tables, and introduces innovative protocols for consistency in batch queries, thereby enhancing the effectiveness of real-time incremental learning updates. This architecture has been deployed and in use in the bilibili recommendation system for over a year, a video content community with hundreds of millions of users, supporting 10x increase in model computation with minimal resource growth, improving outcomes while preserving the system's real-time performance. Qiang Zhang 0055, Zhipeng Teng, Disheng Wu, Jiayin Wang 0002 |
CIKM | 4 |
| 2024 | Label-Specific Multi-label Classification with Entropy Guided Clustering
Jiaxuan Li 0001, Xiaoyan Zhu 0003, Jiayin Wang 0002 |
ICPR (2) | 4 |
| 2024 | TMBstable: a variant caller controls performance variation across heterogeneous sequencing samplesabstractIn cancer genomics, variant calling has advanced, but traditional mean accuracy evaluations are inadequate for biomarkers like tumor mutation burden, which vary significantly across samples, affecting immunotherapy patient selection and threshold settings. In this study, we introduce TMBstable, an innovative method that dynamically selects optimal variant calling strategies for specific genomic regions using a meta-learning framework, distinguishing it from traditional callers with uniform sample-wide strategies. The process begins with segmenting the sample into windows and extracting meta-features for clustering, followed by using a pre-trained meta-model to select suitable algorithms for each cluster, thereby addressing strategy-sample mismatches, reducing performance fluctuations and ensuring consistent performance across various samples. We evaluated TMBstable using both simulated and real non-small cell lung cancer and nasopharyngeal carcinoma samples, comparing it with advanced callers. The assessment, focusing on stability measures, such as the variance and coefficient of variation in false positive rate, false negative rate, precision and recall, involved 300 simulated and 106 real tumor samples. Benchmark results showed TMBstable's superior stability with the lowest variance and coefficient of variation across performance metrics, highlighting its effectiveness in analyzing the counting-based biomarker. The TMBstable algorithm can be accessed at https://github.com/hello-json/TMBstable for academic usage only. Shenjie Wang, Xiaoyan Zhu 0003, Xuwen Wang, Yuqian Liu, Minchao Zhao, Zhili Chang, Jiayin Wang 0002 |
Briefings Bioinform. | 9 |
| 2024 | NIPT-PG: empowering non-invasive prenatal testing to learn from population genomics through an incremental pan-genomic approachabstractNon-invasive prenatal testing (NIPT) is a quite popular approach for detecting fetal genomic aneuploidies. However, due to the limitations on sequencing read length and coverage, NIPT suffers a bottleneck on further improving performance and conducting earlier detection. The errors mainly come from reference biases and population polymorphism. To break this bottleneck, we proposed NIPT-PG, which enables the NIPT algorithm to learn from population data. A pan-genome model is introduced to incorporate variant and polymorphic loci information from tested population. Subsequently, we proposed a sequence-to-graph alignment method, which considers the read mis-match rates during the mapping process, and an indexing method using hash indexing and adjacency lists to accelerate the read alignment process. Finally, by integrating multi-source aligned read and polymorphic sites across the pan-genome, NIPT-PG obtains a more accurate z-score, thereby improving the accuracy of chromosomal aneuploidy detection. We tested NIPT-PG on two simulated datasets and 745 real-world cell-free DNA sequencing data sets from pregnant women. Results demonstrate that NIPT-PG outperforms the standard z-score test. Furthermore, combining experimental and theoretical analyses, we demonstrate the probably approximately correct learnability of NIPT-PG. In summary, NIPT-PG provides a new perspective for fetal chromosomal aneuploidies detection. NIPT-PG may have broad applications in clinical testing, and its detection results can serve as a reference for false positive samples approaching the critical threshold. Zhengfa Xue, Aifen Zhou, Xiaoyan Zhu 0003, Huanhuan Zhu, Jiayin Wang 0002 |
Briefings Bioinform. | 7 |
| 2024 | TCSR: Self-attention with time and category for session-based recommendationabstractAbstract Session‐based recommendation that uses sequence of items clicked by anonymous users to make recommendations has drawn the attention of many researchers, and a lot of approaches have been proposed. However, there are still problems that have not been well addressed: (1) Time information is either ignored or exploited with a fixed time span and granularity, which fails to understand the personalized interest transfer pattern of users with different clicking speeds; (2) Category information is either omitted or considered independent of the items, which defies the fact that the relationships between categories and items are helpful for the recommendation. To solve these problems, we propose a new session‐based recommendation method, TCSR (self‐attention with time and category for session‐based recommendation). TCSR uses a non‐linear normalized time embedding to perceive user interest transfer patterns on variable granularity and employs a heterogeneous SAN to make full use of both items and categories. Moreover, a cross‐recommendation unit is adapted to adjust recommendations on the item and category sides. Extensive experiments on four real datasets show that TCSR significantly outperforms state‐of‐the‐art approaches. Xiaoyan Zhu 0003, Yu Zhang 0203, Jiaxuan Li 0001, Jiayin Wang 0002, Xin Lai 0003 |
Comput. Intell. | 4 |
| 2024 | Stacked co-training for semi-supervised multi-label learning
Jiaxuan Li 0001, Xiaoyan Zhu 0003, Hongrui Wang 0004, Yu Zhang 0203, Jiayin Wang 0002 |
Inf. Sci. | 5 |
| 2024 | Graph-enhanced and collaborative attention networks for session-based recommendation
Xiaoyan Zhu 0003, Yu Zhang 0203, Jiayin Wang 0002, Guangtao Wang |
Knowl. Based Syst. | 3 |
| 2024 | Dual-channel graph contrastive learning for multi-label classification with label-specific features and label correlations
Xiaoyan Zhu 0003, Jiaxuan Li 0001, Jiayin Wang 0002 |
Neural Comput. Appl. | 4 |
| 2024 | scMSI: Accurately inferring the sub-clonal Micro-Satellite status by an integrated deconvolution model on length spectrumabstractMicrosatellite instability (MSI) is an important genomic biomarker for cancer diagnosis and treatment, and sequencing-based approaches are often applied to identify MSI because of its fastness and efficiency. These approaches, however, may fail to identify MSI on one or more sub-clones for certain cancers with a high degree of heterogeneity, leading to erroneous diagnoses and unsuitable treatments. Besides, the computational cost of identifying sub-clonal MSI can be exponentially increased when multiple sub-clones with different length distributions share MSI status. Herein, this paper proposes "scMSI", an accurate and efficient estimation of sub-clonal MSI to identify the microsatellite status. scMSI is an integrative Bayesian method to deconvolute the mixed-length distribution of sub-clones by a novel alternating iterative optimization procedure based on a subtle generative model. During the process of deconvolution, the optimized division of each sub-clone is attained by a heuristic algorithm, aligning with clone proportions that adhere optimally to the sample's clonal structure. To evaluate the performance, 16 patients diagnosed with endometrial cancer, exhibiting positive responses to the treatment despite having negative MSI status based on sequencing-based approaches, were considered. Excitingly, scMSI reported MSI on sub-clones successfully, and the findings matched the conclusions on immunohistochemistry. In addition, testing results on a series of experiments with simulation datasets concerning a variety of impact factors demonstrated the effectiveness and superiority of scMSI in detecting MSI on sub-clones over existing approaches. scMSI provides a new way of detecting MSI for cancers with a high degree of heterogeneity. Yuqian Liu, Huanwen Wu, Xuanping Zhang, Zhiyong Liang, Jiayin Wang 0002 |
PLoS Comput. Biol. | 8 |
| 2024 | A ranking-based problem transformation method for weakly supervised multi-label learning
Jiaxuan Li 0001, Xiaoyan Zhu 0003, Weichu Zhang, Jiayin Wang 0002 |
Pattern Recognit. | 4 |
| 2024 | A novel instance-based method for cross-project just-in-time defect predictionabstractSummary Cross‐project (CP) just‐in‐time software defect prediction (JIT‐SDP) uses CP data to overcome initial data scarcity for training high‐performing JIT‐SDP classifiers in the early stages of software projects. The primary challenge faced by JIT‐SDP in a cross‐project context lies in the distinct distributions between training and test data. To tackle this issue, we select source data instances that closely resemble target data for building classifiers. Software datasets commonly exhibit a class imbalance problem, where the ratio of the defective class to the clean class is notably low. This imbalance typically diminishes classifier performance. In this study, we propose an instance selection method utilizing kernel mean matching (ISKMM) that addresses both knowledge transfer and class imbalance in cross‐project defect prediction (CPDP). The method employs the kernel mean matching (KMM) technique to assess the similarity between training and target data. It selects instances with high similarity, retains them, and resamples the data based on similarity weighting to mitigate the class imbalance problem. Our experiments, conducted on 10 open‐source projects, reveal that the ISKMM algorithm outperforms existing CP single‐source software defect prediction (SDP) algorithms. Moreover, when employing the proposed algorithm, defect predictors constructed from cross‐project data demonstrate an overall performance comparable to predictors learned from within‐project data. Xiaoyan Zhu 0003, Jiayin Wang 0002, Xin Lai 0003 |
Softw. Pract. Exp. | 3 |
| 2023 | AdaBoost.C2: Boosting Classifiers Chains for Multi-Label ClassificationabstractDuring the last decades, multi-label classification (MLC) has attracted the attention of more and more researchers due to its wide real-world applications. Many boosting methods for MLC have been proposed and achieved great successes. However, these methods only extend existing boosting frameworks to MLC and take loss functions in multi-label version to guide the iteration. These loss functions generally give a comprehensive evaluation on the label set entirety, and thus the characteristics of different labels are ignored. In this paper, we propose a multi-path AdaBoost framework specific to MLC, where each boosting path is established for distinct label and the combination of them is able to provide a maximum optimization to Hamming Loss. In each iteration, classifiers chain is taken as the base classifier to strengthen the connection between multiple AdaBoost paths and exploit the label correlation. Extensive experiments demonstrate the effectiveness of the proposed method. Jiaxuan Li 0001, Xiaoyan Zhu 0003, Jiayin Wang 0002 |
AAAI | 3 |
| 2023 | A Control Chart Method for Simultaneously Monitoring the Average Level and Stability of Surgical QualityabstractGood and stable surgical quality is of great significance to ensure the life safety of patients and spare patients unnecessary health burdens. The variable life-adjusted display (VLAD) is a popular assessment method for surgical quality and some VLAD-based control charts have been proposed to motivate the quality improvement. However, existing charts can only monitor the average level of surgical quality by detecting the changes of VLAD’s mean, but lack a mechanism to monitor the stability based on VLAD’s variance. The volatile surgical quality which can hardly be considered good will make existing charts provide delay alarms. Therefore, in this paper, we propose a risk-adjusted exponential weighted moving average (EWMA) control chart to monitor the mean and variance of VLAD simultaneously, named MVV-EWMA. Firstly, we give the explicit form of VLAD’s variance. Then, the EWMA statistics for the mean and variance are respectively constructed and integrated by the generalized likelihood ratio test. Moreover, two auxiliary mechanisms are adopted to enhance MVV-EWMA’s monitoring ability. Both the results of simulation and case study show that MVV-EWMA is not only able to effectively monitor the stability of surgical quality, but also has the best performance compared to existing chart. Ruoyu Liu, Xin Lai 0003, Jiayin Wang 0002, Paul B. S. Lai, Ka Chun Chong |
BIBM | 3 |
| 2023 | A Statistical Explainable Learning Model Optimizing Co-localization of Multidimensional Positivity Thresholds in Immunotherapy Decision-SupportingabstractTumor Mutation Burden (TMB) serves as a recognized stratified biomarker for immunotherapy. However, its one-dimensional representation of non-synonymous genetic alterations has been contentious. Specifically, the uniform quantification of mutations by TMB, coupled with measurement inaccuracies, complicates the accurate determination of a positive threshold for classifying patients. Parallel to this, assessing immunotherapy benefits requires the joint analysis of multiscale endpoints, namely discrete tumor response and sequential time-to-event, presenting a pressing challenge for clinical computation. Recognizing the intertwined nature of these challenges, we address the inter-sample bias inherent in multidimensional mutation biomarkers within the framework of multiscale endpoint fusion analysis, aiming for a more robust and comprehensive patient stratification. By combining the concept of corrected-score with a soft-threshold strategy, and utilizing the attention mechanism alongside the multiple instance learning, we propose a statistically explainable learning model optimizing co-localization of multidimensional positivity thresholds for immunotherapy categorical decision-supporting. Jian Zhao 0034, Jiayin Wang 0002, Quan Wang 0004 |
BIBM | 4 |
| 2023 | Automated machine learning with dynamic ensemble selection
Xiaoyan Zhu 0003, Jingtao Ren, Jiayin Wang 0002, Jiaxuan Li 0001 |
Appl. Intell. | 3 |
| 2023 | DELFMUT: duplex sequencing-oriented depth estimation model for stable detection of low-frequency mutationsabstractDuplex sequencing technology has been widely used in the detection of low-frequency mutations in circulating tumor deoxyribonucleic acid (DNA), but how to determine the sequencing depth and other experimental parameters to ensure the stable detection of low-frequency mutations is still an urgent problem to be solved. The mutation detection rules of duplex sequencing constrain not only the number of mutated templates but also the number of mutation-supportive reads corresponding to each forward and reverse strand of the mutated templates. To tackle this problem, we proposed a Depth Estimation model for stable detection of Low-Frequency MUTations in duplex sequencing (DELFMUT), which models the identity correspondence and quantitative relationships between templates and reads using the zero-truncated negative binomial distribution without considering the sequences composed of bases. The results of DELFMUT were verified by real duplex sequencing data. In the case of known mutation frequency and mutation detection rule, DELFMUT can recommend the combinations of DNA input and sequencing depth to guarantee the stable detection of mutations, and it has a great application value in guiding the experimental parameter setting of duplex sequencing technology. Guiying Wu, Tianyu Cui, Zicong Jiao, Liyan Ji, Jiayin Wang 0002, Xuefeng Xia, Huan Fang 0005, Yanfang Guan |
Briefings Bioinform. | 8 |
| 2023 | Dynamic ensemble learning for multi-label classification
Xiaoyan Zhu 0003, Jiaxuan Li 0001, Jingtao Ren, Jiayin Wang 0002, Guangtao Wang |
Inf. Sci. | 4 |
| 2022 | PEcnv: accurate and efficient detection of copy number variations of various lengthsabstractCopy number variation (CNV) is a class of key biomarkers in many complex traits and diseases. Detecting CNV from sequencing data is a substantial bioinformatics problem and a standard requirement in clinical practice. Although many proposed CNV detection approaches exist, the core statistical model at their foundation is weakened by two critical computational issues: (i) identifying the optimal setting on the sliding window and (ii) correcting for bias and noise. We designed a statistical process model to overcome these limitations by calculating regional read depths via an exponentially weighted moving average strategy. A one-run detection of CNVs of various lengths is then achieved by a dynamic sliding window, whose size is self-adopted according to the weighted averages. We also designed a novel bias/noise reduction model, accompanied by the moving average, which can handle complicated patterns and extend training data. This model, called PEcnv, accurately detects CNVs ranging from kb-scale to chromosome-arm level. The model performance was validated with simulation samples and real samples. Comparative analysis showed that PEcnv outperforms current popular approaches. Notably, PEcnv provided considerable advantages in detecting small CNVs (1 kb-1 Mb) in panel sequencing data. Thus, PEcnv fills the gap left by existing methods focusing on large CNVs. PEcnv may have broad applications in clinical testing where panel sequencing is the dominant strategy. Availability and implementation: Source code is freely available at https://github.com/Sherwin-xjtu/PEcnv. Xuwen Wang, Ruoyu Liu, Xin Lai 0003, Yuqian Liu, Shenjie Wang, Xuanping Zhang, Jiayin Wang 0002 |
Briefings Bioinform. | 8 |
| 2021 | A new multiple instance algorithm using structural informationabstractMultiple instance learning (MIL) is semisupervised learning that predicts the label of a bag with a wide diversity of instances. It has many applications and thus attracts increasingly more attention. In this paper, we propose a new MIL algorithm using the structural information of a bag to predict its label. In the proposed method, a bag is transformed into a graph, and spectral clustering is employed to divide the graph into several subgraphs. Then, the graph Fourier transform is utilized to extract the features of the subgraphs. Finally, an end-to-end neural network is used to predict the label of a bag with the extracted features. An empirical study with 25 datasets was conducted to validate the effectiveness of the proposed method. The experimental results show that the proposed method performs better than the 6 baseline methods on most datasets. Xiaoyan Zhu 0003, Jiayin Wang 0002, Yuqian Liu |
ICDM | 3 |
| 2021 | Ensemble of ML-KNN for classification algorithm recommendation
Xiaoyan Zhu 0003, Chenzhen Ying, Jiayin Wang 0002, Jiaxuan Li 0001, Xin Lai 0003, Guangtao Wang |
Knowl. Based Syst. | 3 |
| 2021 | Automatic Recommendation of a Distance Measure for Clustering AlgorithmsabstractWith a large number of distance measures, the appropriate choice for clustering a given data set with a specified clustering algorithm becomes an important problem. In this article, an automatic distance measure recommendation method for clustering algorithms is proposed. The recommendation method consists of the following steps: (1) metadata extraction, including meta-feature collection and meta-target identification; (2) recommendation model construction using metadata; and (3) distance measure recommendation for a new data set by the recommendation model. Two different types of meta-targets and meta-learning techniques are utilized considering the possible different requirements of users. To validate the necessity and effectiveness of the distance measure recommendation method, an empirical study is conducted with 199 publicly available data sets, 9 distance measures, and 2 widely used clustering algorithms. The experimental results indicate that distance measure significantly influences the performance of the clustering algorithm for a given data set. Furthermore, performance analysis of the proposed recommendation method proves its effectiveness. Xiaoyan Zhu 0003, Yingbin Li, Jiayin Wang 0002, Jingwen Fu |
ACM Trans. Knowl. Discov. Data | 3 |
| 2020 | Accurately estimating the length distributions of genomic micro-satellites by tumor purity deconvolutionabstractBACKGROUND: Genomic micro-satellites are the genomic regions that consist of short and repetitive DNA motifs. Estimating the length distribution and state of a micro-satellite region is an important computational step in cancer sequencing data pipelines, which is suggested to facilitate the downstream analysis and clinical decision supporting. Although several state-of-the-art approaches have been proposed to identify micro-satellite instability (MSI) events, they are limited in dealing with regions longer than one read length. Moreover, based on our best knowledge, all of these approaches imply a hypothesis that the tumor purity of the sequenced samples is sufficiently high, which is inconsistent with the reality, leading the inferred length distribution to dilute the data signal and introducing the false positive errors. RESULTS: In this article, we proposed a computational approach, named ELMSI, which detected MSI events based on the next generation sequencing technology. ELMSI can estimate the specific length distributions and states of micro-satellite regions from a mixed tumor sample paired with a control one. It first estimated the purity of the tumor sample based on the read counts of the filtered SNVs loci. Then, the algorithm identified the length distributions and the states of short micro-satellites by adding the Maximum Likelihood Estimation (MLE) step to the existing algorithm. After that, ELMSI continued to infer the length distributions of long micro-satellites by incorporating a simplified Expectation Maximization (EM) algorithm with central limit theorem, and then used statistical tests to output the states of these micro-satellites. Based on our experimental results, ELMSI was able to handle micro-satellites with lengths ranging from shorter than one read length to 10kbps. CONCLUSIONS: To verify the reliability of our algorithm, we first compared the ability of classifying the shorter micro-satellites from the mixed samples with the existing algorithm MSIsensor. Meanwhile, we varied the number of micro-satellite regions, the read length and the sequencing coverage to separately test the performance of ELMSI on estimating the longer ones from the mixed samples. ELMSI performed well on mixed samples, and thus ELMSI was of great value for improving the recognition effect of micro-satellite regions and supporting clinical decision supporting. The source codes have been uploaded and maintained at https://github.com/YixuanWang1120/ELMSI for academic use only. Xuanping Zhang, Fei-Ran Zhang, Xinxing Yan, Zhongmeng Zhao, Yanfang Guan, Jiayin Wang 0002 |
BMC Bioinform. | 9 |
| 2020 | SOAPTyping: an open-source and cross-platform tool for sequence-based typing for HLA class I and II allelesabstractBACKGROUND: The human leukocyte antigen (HLA) gene family plays a key role in the immune response and thus is crucial in many biomedical and clinical settings. Utilizing Sanger sequencing, the golden standard technology for HLA typing enables accurate identification of HLA alleles in high-resolution. However, only the commercial software, such as uTYPE, SBT-Assign, and SBTEngine, and very few open-source tools could be applied to perform HLA typing based on Sanger sequencing. RESULTS: We developed a user-friendly, cross-platform and open-source desktop application, known as SOAPTyping, for Sanger-based typing in HLA class I and II alleles. SOAPTyping can produce accurate results with a comprehensible protocol and featured functions. Moreover, SOAPTyping supports a more advanced group-specific sequencing primers (GSSP) module to solve the ambiguous typing results. We used SOAPTyping to analyze 36 samples with known HLA typing from the University of California Los Angeles (UCLA) International HLA DNA Exchange platform and 100 anonymous clinical samples, and the HLA typing results from SOAPTyping are identical to the golden results and 5.5 times faster than commercial software uTYPE, which shows the usability of SOAPTyping. CONCLUSIONS: We introduce the SOAPTyping as the first open-source and cross-platform HLA typing software with the capability of producing high-resolution HLA typing predictions from Sanger sequence data. Yong Zhang 0036, Yongsheng Chen, Huixin Xu, Weipeng Hu, Xiaoqin Yang, Jia Ye, Jiayin Wang 0002, Weiqiang Sun, Jian Wang 0065, Huanming Yang |
BMC Bioinform. | 10 |
| 2019 | An Artificial Fish Swarm Algorithm for Identifying Associations between Multiple Variants and Multiple PhenotypesabstractIdentifying associations between genomic variants and phenotypes has always been an interesting research field of population genetics, which is of great significance for studying the pathogenesis of complex diseases and supporting clinical assistant decision making. Nowadays, many identification methods have been proposed to find the associations between variants and phenotypes, such as GWAS and pheWAS, and have made excellent achievements in pathological research and clinical practice. However, the existing methods only focus on single phenotype-multiple variants or single variant-multiple phenotypes, but not on multiple variants-multiple phenotypes. In the view of the fact that complex diseases often have several subtypes which differ greatly in variants and phenotypes, focusing only on single variant or single phenotype is far from enough and limits the ability of identification of those methods. Therefore, we propose a heuristic method with an AFSA framework on the solution space to identify associations between multiple variants and multiple phenotypes. In our method, each fish carries two logic trees that respectively represent the associations between variants and the associations between phenotypes. The logic trees will be iteratively updated to find a better solution according to the preset update strategies. When the iteration stop condition is reached, the algorithm will stop and output the optimal fish. The logical expression represented by the logic trees carried by the optimal fish is the associations we find. We validated the proposed method on the simulation data generated by hapgen2 and PhenotypeSimulator, and took the ratio of the number of people that can be explained by the found logical expression as the index to evaluate the performance, which was called Coverage. We conducted 9 groups of experiments, each of which was different in the number of variants and phenotypes. The best Coverage of was from the group including 500 variants and 10 phenotypes, which reached 72.12%, and the worst result is from the group including 100 variants and 20 phenotypes, 31.73%. We also exhausted the simulation data to find the optimal logical expression and several most important logic rules to evaluate the results obtained by the method. Ruoyu Liu, Xin Lai 0003, Xuanping Zhang, Xiaoyan Zhu 0003, Jiayin Wang 0002 |
BIBM | 6 |
| 2019 | GSDcreator: An Efficient and Comprehensive Simulator for Genarating NGS Data with Population Genetic InformationabstractIn recent decades, NGS data analysis has become a major research field in bioinformatics, which presents great advantages in many application scenarios. Many algorithms and software were designed for analyzing the NGS data, while simulation datasets are urgently needed for testing software and optimizing their parameter configurations. Thus, a series of NGS data simulators have been published. However, the existing simulators cannot satisfy the requirements from many specific scenarios. First, they do not support many newly discovered variations. Second, complex structural variations are difficult to generate. In addition, along with the increase of population data, it is urgent to increase population information simulation. In this paper, we propose GSDcreator, a comprehensive NGS simulator that overcome the three weaknesses mentioned above. It can produce all known types of variation, where the complex of variations are also supported. Furthermore, it can capture many important real data features including population polymorphism, insert size distribution, adjacent site depth distribution, overall depth distribution, quality score distribution, amplification bias, sequencing errors and so on. It's highlighted that 1000 Genomes Project Database is taken as a reference and integrates population genetic information to simulate population polymorphism. To test the performance, we did a lot of experiments and found that simulated data produced by GSDcreator are quit mimic to the real sequencing data. Shenjie Wang, Jiayin Wang 0002, Xuanping Zhang, Xuwen Wang, Xiaoyan Zhu 0003, Xin Lai 0003 |
BIBM | 2 |
| 2019 | FilterLAP: Filtering False-positive Mutation Calls via a Label Propagation FrameworkabstractBenefiting from the recent advantages of genomic sequencing, detecting genomic mutations becomes a routine work in precise diagnoses and treatments for cancers. In clinical practices, many factors, such as tumor purity, clonal structure, etc., interfere the performance of calling mutations. The computational pipelines prefer to sensitively report the candidate calls, while a filter is applied for removing the false-positive calls. The existing filters rely on the whole genome/exome sequencing data, which can provide sufficient samples for training the filters. However, the gene-panel sequencing is more popular in clinical practices, but there is no practical filter for limited training samples. In light of this, we develop a semi-learning filter for gene-panel sequencing data, FilterLAP, which implemented via a label propagation framework. Given few labeled samples with a set of unlabeled ones, its basic idea is to predict the label information of unlabeled nodes from the label information of labeled nodes, and establishes a complete graph model by using the relationship between samples, by combining transductive inference with label propagation algorithm. For each node in the network, tags are propagated to adjacent nodes according to similarity and the probability distribution of similar nodes tends to be similar and can be divided into a class. We perform multiple sets of experiments on gene-panel sequencing data captured from Illumina platform. FilterLAP outperforms on both SNV and INDEL filtering, where the AUCs reach 0.90-0.97, and the average accuracies on overall mutation calls are over 90%. Comparing to GATK hard filters, FilterLAP present a 5% improvement on accuracy. These results demonstrate that the proposed method can better reduce the false positive mutation calls on gene-panel sequencing data. In addition, it is stable and efficient, which can be used as a practical tool for mutation call filtering for gene-panel sequencing data. Xuwen Wang, Xiaoyan Zhu 0003, Shenjie Wang, Xuanping Zhang, Xin Lai 0003, Jiayin Wang 0002 |
BIBM | 7 |
| 2014 | Chaotic image encryption based on circular substitution box and key stream buffer
Xuanping Zhang, Zhongmeng Zhao, Jiayin Wang 0002 |
Signal Process. Image Commun. | 3 |
| 2013 | Admission control on multipath routing in 802.11-based wireless mesh networks
Peng Zhao 0001, Xinyu Yang 0001, Jiayin Wang 0002, Benyuan Liu, Jie Wang 0002 |
Ad Hoc Networks | 3 |
| 2012 | Rate-adaptive admission control for bandwidth assurance in multirate wireless mesh networksabstractAdmission control (AC) is an effective mechanism for providing bandwidth assurance in wireless mesh networks. Early AC schemes over multirate WMNs typically use a pre-chosen rate or a MAC-layer adapted rate for each link, denying data sessions that could have been admitted should a better multirate AC be available. Taking full advantage of multirate WMNs, we present a rate-adaptive admission control protocol (RaAC) for IEEE 802.11-based WMNs. RaAC consists of three major components: (1) a rate adaption algorithm to meet the bandwidth requirement of the data session and satisfy the channel condition of the PHY layer; (2) a new path-selection metric to balance between hop counts, bandwidth, rates, and other network parameters; and (3) a routing-coupled, distributed, rate-adaptive admission control algorithm to admit data sessions with bandwidth assurance. Through simulations, we show that RaAC is efficient and effective in meeting bandwidth requirements. Peng Zhao 0001, Xinyu Yang 0001, Chaoxin Hu, Jiayin Wang 0002, Benyuan Liu, Jie Wang 0002 |
ICC | 4 |
| 2012 | BOR/AC: Bandwidth-aware opportunistic routing with admission control in wireless mesh networksabstractOpportunistic routing (OR) is a viable approach for improving performance of wireless communications. Previous studies on OR have focused on cost minimization, performance of multiple rates, congestion control, and other issues. Bandwidth assurance over OR, however, has not been adequately investigated. To bridge this gap, we present a bandwidth-aware opportunistic routing (BOR) with admission control (AC) protocol named BOR/AC. In particular, by analyzing the expected available bandwidth (EAB) and the expected transmission cost (ETC) in OR, we first devise a new metric called BCR (bandwidth-cost ratio) to determine the priority of relays in the forwarding candidates set. Admission control is then applied to admit or reject traffic flows based on estimated expected available bandwidth. Extensive simulation results show that BOR/AC consistently achieves much better performance than existing opportunistic routing protocols. Peng Zhao 0001, Xinyu Yang 0001, Jiayin Wang 0002, Benyuan Liu, Jie Wang 0002 |
INFOCOM | 3 |
| 2012 | An improved approach for accurate and efficient calling of structural variations with low-coverage sequence dataabstractBACKGROUND: Recent advances in sequencing technologies make it possible to comprehensively study structural variations (SVs) using sequence data of large-scale populations. Currently, more efforts have been taken to develop methods that call SVs with exact breakpoints. Among these approaches, split-read mapping methods can be applied on low-coverage sequence data. With increasing amount of data generated, more efficient split-read mapping methods are still needed. Also, since sequence errors can not be avoided for the current sequencing technologies, more accurate split-read mapping methods are still needed to better handle sequence errors. RESULTS: In this paper, we present a split-read mapping method implemented in the program SVseq2 which improves our previous work SVseq1. Similar to SVseq1, SVseq2 calls deletions (and insertions) with exact breakpoints. SVseq2 achieves more accurate calling through split-read mapping within focal regions. SVseq2 also has a much desired feature: there is no need to specify the maximum deletion size, while some existing split-read mapping methods need more memory and longer running time when larger maximum deletion size is chosen. SVseq2 is also much faster because it only needs to examine a small number of ways of splitting the reads. Moreover, SVseq2 supports insertion calling from low-coverage sequence data, while SVseq1 only supports deletion finding. The program SVseq2 can be downloaded at http://www.engr.uconn.edu/~jiz08001/. CONCLUSIONS: SVseq2 enables accurate and efficient SV calling through split-read mapping within focal regions using paired-end reads. For many simulated data and real sequence data, SVseq2 outperforms some other existing approaches in accuracy and efficiency, especially when sequence coverage is low. Jin Zhang 0038, Jiayin Wang 0002, Yufeng Wu 0001 |
BMC Bioinform. | 2 |
| 2010 | Fast Computation of the Exact Hybridization Number of Two Phylogenetic Trees
Yufeng Wu 0001, Jiayin Wang 0002 |
ISBRA | 2 |