VLDB 2026 Research / reviewers in the wild / expert
Xiaoyan Zhu 0003
dblp:50/1222-3
· DBLP profile ↗
52ranked-venue papers
16as first author
38since 2021 · last 2027
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 19 · 15 since 2021Artificial intelligence and machine learning · 18 · 8 first-author · 15 since 2021Software engineering, systems software and programming languages · 11 · 6 first-author · 5 since 2021Databases, data management, data science and information retrieval · 6 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | TrafHILLM: Highway network traffic flow prediction with heterogeneous graph-based and instruction fine-tuned large language model
Hongrui Wang 0004, Shanchuan Yu, Jiayin Wang 0002, Xiaoyan Zhu 0003, Jiaxuan Li 0001, Yuchuan Du |
Expert Syst. Appl. | 4 |
| 2026 | Mosaic Pruning: A Hierarchical Framework for Generalizable Pruning of Mixture-of-Experts ModelsabstractSparse Mixture-of-Experts (SMoE) architectures have enabled a new frontier in scaling Large Language Models (LLMs), offering superior performance by activating only a fraction of their total parameters during inference. However, their practical deployment is severely hampered by substantial static memory overhead, as all experts must be loaded into memory. Existing post-training pruning methods, while reducing model size, often derive their pruning criteria from a single, general-purpose corpus. This leads to a critical limitation: a catastrophic performance degradation when the pruned model is applied to other domains, necessitating a costly re-pruning for each new domain. To address this generalization gap, we introduce Mosaic Pruning (MoP). The core idea of MoP is to construct a functionally comprehensive set of experts through a structured ``cluster-then-select" process. This process leverages a similarity metric that captures expert performance across different task domains to functionally cluster the experts, and subsequently selects the most representative expert from each cluster based on our proposed Activation Variability Score. Unlike methods that optimize for a single corpus, our proposed Mosaic Pruning ensures that the pruned model retains a functionally complementary set of experts, much like the tiles of a mosaic that together form a complete picture of the original model's capabilities, enabling it to handle diverse downstream tasks.Extensive experiments on various MoE models demonstrate the superiority of our approach. MoP significantly outperforms prior work, achieving a 7.24\% gain on general tasks and 8.92\% on specialized tasks like math reasoning and code generation. Mingkuan Zhao, Shuangyong Song, Xiaoyan Zhu 0003, Xin Lai 0003, Jiayin Wang 0002 |
AAAI | 4 |
| 2026 | Making Every Head Count: Sparse Attention Without the Speed-Performance Trade-offabstractThe design of Large Language Models (LLMs) has long been hampered by a fundamental conflict within their core attention mechanism: its remarkable expressivity is built upon a computational complexity of O(H·N²) that grows quadratically with the context size (N) and linearly with the number of heads (H). This standard implementation harbors significant computational redundancy, as all heads independently compute attention over the same sequence space. Existing sparse methods, meanwhile, often trade information integrity for computational efficiency. To resolve this efficiency-performance trade-off, we propose SPAttention, whose core contribution is the introduction of a new paradigm we term Principled Structural Sparsity. SPAttention does not merely drop connections but instead reorganizes the computational task by partitioning the total attention workload into balanced, non-overlapping distance bands, assigning each head a unique segment. This approach transforms the multi-head attention mechanism from H independent O(N²) computations into a single, collaborative O(N²) computation, fundamentally reducing complexity by a factor of H. The structured inductive bias compels functional specialization among heads, enabling a more efficient allocation of computational resources from redundant modeling to distinct dependencies across the entire sequence span. Extensive empirical validation on the OLMoE-1B-7B and 0.25B-1.75B model series demonstrates that while delivering an approximately two-fold increase in training throughput, its performance is on par with standard dense attention, even surpassing it on select key metrics, while consistently outperforming representative sparse attention methods including Longformer, Reformer, and BigBird across all evaluation metrics. Our work demonstrates that thoughtfully designed structural sparsity can serve as an effective inductive bias that simultaneously improves both computational efficiency and model performance, opening a new avenue for the architectural design of next-generation, high-performance LLMs. Mingkuan Zhao, Jiayin Wang 0002, Xin Lai 0003, Tianchen Huang, Yuheng Min, Xiaoyan Zhu 0003 |
AAAI | 8 |
| 2026 | IFHAGrec: Instruction-Finetuned Heterogeneous-Aware Graph Neural Network for Temporally Weighted Recommendation Model
Hongrui Wang 0004, Shanchuan Yu, Xiaoyan Zhu 0003, Guangtao Wang, Jiayin Wang 0002, Jiaxuan Li 0001, Jindong Jiang |
KSEM (1) | 3 |
| 2026 | PScnv: personalized self-normalizing CNV detection with a hierarchical multi-phase frameworkabstractMOTIVATION: Accurate detection of copy number variations (CNVs) from targeted panel sequencing remains challenging due to limited genomic coverage and pronounced sample-specific biases. Existing normalization strategies, including baseline-cohort, matched-control, and single-sample approaches, often struggle to balance noise suppression with adaptability, leading to inconsistent performance across heterogeneous samples. RESULTS: We present PScnv, a personalized self-normalizing framework for robust CNV detection from panel sequencing data. PScnv integrates a pre-built panel-of-normals (PoN) with sample-intrinsic stable chromosomes through ridge-regression normalization to generate individualized log2 ratio profiles with reduced systematic variation. CNVs are then identified using a hierarchical multi-phase segmentation pipeline incorporating z-score pre-partitioning, kernel-based correction, and circular binary segmentation. In 139 clinical tumor samples with orthogonal FISH validation at MET, ERBB2, and MTAP, PScnv showed improved accuracy and robustness over existing methods that do not require patient-matched normal samples, provided that a pre-built PoN cohort is available. AVAILABILITY: Source code is available for academic use at https://github.com/lvws/PScnv. Xuwen Wang, Zhili Chang, Wansheng Lv, Akhatov Akmal, Xamidov Munis, Xunbiao Liu, Shenjie Wang, Xiaoyan Zhu 0003, Chong Du, Shuqun Zhang, Jiayin Wang 0002 |
Bioinform. | 8 |
| 2026 | UF-CDDFM: A unified framework for code defect detection using multi-modal inputs and few-shot learningabstractContext: The detection of code defects is foundational to modern software development and maintenance, playing a critical role in ensuring software quality and security. However, as software systems grow in scale and complexity, the limitations of traditional static analysis and conventional machine learning techniques have become increasingly evident. These methods rely heavily on intricate, manual feature engineering and fail to capture dynamic runtime behavior, resulting in suboptimal accuracy and elevated error rates. Objective: To address these deficiencies, we propose UF-CDDFM, a unified framework for code defect detection that integrates multi-modal inputs, active learning, and state-of-the-art few-shot learning techniques. We aim to improve detection performance, reduce feature selection complexity and sample bias through active learning, and maintain practical efficiency in real-world development contexts. Methods: UF-CDDFM employs parallel encoding of source code, code annotations, and abstract syntax trees (ASTs) using large language models (LLMs) alongside multilayer perceptrons (MLPs) to derive robust, high-fidelity representations of code. To streamline feature selection and mitigate sample bias, an active learning component is introduced for automated identification of high-quality features. Addressing the pervasive challenge of data scarcity, we incorporate two complementary few-shot learning strategies-MAML for small-scale datasets and LEO for larger-scale settings to enhance overall generalization capability. Results: Empirical evaluations demonstrate that UF-CDDFM consistently outperforms existing methods, establishing new state-of-the-art detection rates: 72.04% for defect detection and 95.23% for clone detection. Crucially, these gains are achieved within resource-constrained computational environments, which highlights the practicality of the method. Conclusion: By fusing multi-modal code representations, active learning, and adaptive few-shot learning techniques, UF-CDDFM delivers significant improvements in detection accuracy and computational efficiency. This work offers a new paradigm for robust, scalable, and practical code defect and clone detection in modern software engineering. Xianglu Zhou, Tianxiang Cui, Xiaoyan Zhu 0003, Jiayin Wang 0002, Xin Lai 0003 |
Inf. Softw. Technol. | 3 |
| 2026 | Reaching Software Quality for Bioinformatics Applications: How Far Are We?abstractWith the rapid advancements in medicine, biology, and information technology, their deep integration has given rise to the emerging field of bioinformatics. In this process, high–throughput technologies such as genomics, transcriptomics, and proteomics have generated massive volumes of biological data. The biological significance of these data heavily relies on bioinformatics software for analysis and processing. Therefore, it is crucial for both scientific research and clinical applications to ensure the quality of bioinformatics software and avoiding errors or hidden defects. However, to date, no dedicated study has systematically analyzed the quality of bioinformatics software. we conduct a comprehensive empirical study that aggregates, synthesizes, and analyzes findings from 167 bioinformatics software projects. Following the Preferred Reporting Items for Systematic Review and Meta–Analysis (PRISMA) protocol, we extract and evaluate quality–related data to answer our research questions (RQs). Our analysis reveals several key findings. The quality of bioinformatics software requires significant improvement, with an average defect density approximately 11.8× higher than that of general-purpose software. Additionally, unlike traditional software domains, a considerable proportion of defects in bioinformatics software are related to annotations. These issues can lead developers to overlook potential security vulnerabilities or make incorrect fixes, thereby increasing the cost and complexity of subsequent code maintenance. Based on these findings, we further discuss the challenges faced by bioinformatics software and propose potential solutions. This paper lays a foundation for further research on software quality in the bioinformatics domain and offers actionable insights for researchers and practitioners alike. Xiaoyan Zhu 0003, Xin Lai 0003, Xin Lian, Hangyu Cheng, Jiayin Wang 0002 |
IEEE Trans. Software Eng. | 1 |
| 2025 | Multi-Label Ranking Loss Minimization for Matrix CompletionabstractThe common matrix completion methods minimize the rank of the matrix to be completed in addition to the Hamming loss between the incomplete and completed matrices. The rank of matrix measures the linear relation among the vectors of matrix, which may introduce ambiguity for data recovery. To cope with this issue, we extend multi-label ranking loss into matrix completion, and employ multi-label ranking loss minimization (MLRM) in this paper to exploit the relative correlation among matrix vectors. In MLRM, the original incomplete matrix is converted into a pairwise ranking matrix, and the approximation on this newly generated matrix can be viewed as a surrogate of multi-label ranking loss to replace the Hamming loss pattern in the existing methods. Extensive experiments demonstrate that MLRM outperforms the state-of-the-art matrix completion methods in varies of applications, including movie recommendation, drug-target interaction prediction and multi-label learning. Jiaxuan Li 0001, Xiaoyan Zhu 0003, Hongrui Wang 0004, Yu Zhang 0203, Xin Lai 0003, Jiayin Wang 0002 |
AAAI | 2 |
| 2025 | Cross-Project Defect Prediction Based on Feature Fusion and Local Domain AdaptationabstractCross-project defect prediction (CPDP) is hindered by distribution shifts between source and target projects, so models that excel in within-project software defect prediction (WPDP) often degrade across projects. We propose FLDP, which couples (i) local subset alignment selecting similar source-target file pairs via three file-level metrics and aligning only those subsets with (ii) sequence-graph feature fusion, where TLSTM encodes token sequences and TGCN encodes AST structure into a unified representation. Across 10 transfers on 7 projects, FLDP consistently outperforms classical and recent CPDP baselines in AUC/F1/MCC. Ablation shows both local alignment and fusion are necessary for the gains, and our analysis of selection metrics offers practical guidance for applying CPDP in heterogeneous settings. Xianglu Zhou, Xiaoyan Zhu 0003, Yu Wang 0069, Jiayin Wang 0002, Xin Lai 0003 |
APSEC | 2 |
| 2025 | EMcnv: enhancing CNV detection performance through ensemble strategies with heterogeneous meta-graph neural networksabstractCopy number variation (CNV) is a crucial biomarker for many complex traits and diseases. Although numerous CNV detection tools are available, no single method consistently achieves optimal performance across diverse sequencing samples, as each tool has distinct advantages and limitations. Therefore, integrating the strengths of these tools to improve CNV detection accuracy is both a promising strategy and a significant challenge. To address this, we propose EMcnv, a novel deep ensemble framework based on meta-learning. EMcnv combines multiple CNV detection strategies through a three-step approach: (i) leveraging meta-learning and meta-path heterogeneous graphs, employing Relational Graph Convolutional Networks as a specific model within the Heterogeneous Graph Neural Networks framework to develop a probabilistic weight meta-model that ensembles various CNV detection strategies; (ii) assigning probabilistic weights to calls from different CNV detection tools and aggregating them into weighted CNV regions (CNVRs); (iii) refining Copy number variations based on weighted CNVRs. We conducted comprehensive experiments on both simulated and real sequencing data using benchmark datasets. The results demonstrate that EMcnv significantly outperforms popular existing methods, underscoring its superiority and importance in CNV detection. To support further research, the source code is available for academic use at https://github.com/Sherwin-xjtu/EMcnv. Xuwen Wang, Zhili Chang, Yuqian Liu, Shenjie Wang, Xiaoyan Zhu 0003, Jiayin Wang 0002 |
Briefings Bioinform. | 5 |
| 2025 | MRDadaptis: self-adaptive parameter configuration enhances minimal residual disease detection in heterogeneous ctDNA samplesabstractDetection of structural variations (SVs) through circulating tumor DNA (ctDNA) has become a key method for detecting minimal residual disease (MRD). However, the heterogeneity of ctDNA samples, characterized by variable limits of detection (LOD) and diverse structural variant types, significantly impacts detection stability and performance, posing persistent challenges for conventional SV detection tools such as Delly and Manta. These widely used methods require extensive manual parameter tuning, hindered by the combinatorial complexity of multiple parameters and heterogeneous sequencing data. To address this, we propose MRDadaptis, a novel SV detection tool that uniquely incorporates a self-adaptive parameter optimization mechanism. MRDadaptis distinguishes itself by integrating Bayesian optimization with meta-learning techniques to dynamically adjust detection parameters automatically, based on intrinsic features derived from the ctDNA sequencing data itself. This innovative approach not only reduces manual intervention but also effectively captures sample-specific characteristics, significantly improving detection stability, and detection performance. Extensive validation experiments using both simulated and real-world ctDNA datasets demonstrates it distinct advantages, including markedly improved average F1-scores and superior stability (reduced variance, lower RMSE, increased kurtosis). These results highlight the significant advantages of MRDadaptis in addressing sample heterogeneity, underscoring its potential to improve the accuracy and reliability of MRD detecting through ctDNA analysis. https://github.com/aAT0047/MRDadaptis.git. Xin Lai 0003, Shenjie Wang, Zhengfa Xue, Yuqian Liu, Xiaoyan Zhu 0003, Zhili Chang, Jiayin Wang 0002 |
Briefings Bioinform. | 6 |
| 2025 | TMBquant: an explainable AI-powered caller advancing tumor mutation burden quantification across heterogeneous samplesabstractAccurate tumor mutation burden (TMB) quantification is critical for immunotherapy stratification, yet remains challenging due to variability across sequencing platforms, tumor heterogeneity, and variant calling pipelines. Here, we introduce TMBquant, an explainable AI-powered caller designed to optimize TMB estimation through dynamic feature selection, ensemble learning, and automated strategy adaptation. Built upon the H2O AutoML framework, TMBquant integrates variant features, minimizes classification errors, and enhances both accuracy and stability across diverse datasets. We benchmarked TMBquant against nine widely used variant callers, including traditional tools (e.g. Mutect2, VarScan2, Strelka2) and recent AI-based methods (DeepSomatic, Octopus), using 706 whole-exome sequencing tumor-control pairs. To evaluate clinical relevance, we further assessed TMBquant through survival analyses across immunotherapy-treated cohorts of non-small cell lung cancer (NSCLC), nasopharyngeal carcinoma (NPC), and the two NSCLC subtypes: lung adenocarcinoma and lung squamous cell carcinoma. In each cohort, TMBquant consistently achieved the highest hazard ratios, demonstrating superior patient stratification compared to all other methods. Importantly, TMBquant maintained robust predictive performance across both high-TMB (NSCLC) and low-TMB (NPC) settings, highlighting its generalizability across cancer types with distinct biological characteristics. These findings establish TMBquant as a reliable, reproducible, and clinically actionable tool for precision oncology. The software is open source and freely available at https://github.com/SomaticCaller/SomaticCaller. To enhance reproducibility, we provide detailed usage instructions and representative code snippets for TMBquant in the Methods section (see Code Availability). Shenjie Wang, Xiaoyan Zhu 0003, Xuwen Wang, Yuqian Liu, Minchao Zhao, Zhili Chang, Shuanying Yang, Jiayin Wang 0002 |
Briefings Bioinform. | 3 |
| 2025 | TMBclaw: tumor clone-aware graph learning improves immunotherapy response prediction across heterogeneous cohortsabstractImmune checkpoint inhibitors (ICIs) have emerged as a cornerstone of modern oncology, necessitating the development of robust biomarkers for optimizing patient stratification and treatment selection. While tumor mutation burden (TMB) has demonstrated prognostic value, conventional quantification methods based on mutation counts fail to reflect immunogenic neoantigen presentation due to intratumoral clonal heterogeneity. Recent efforts have focused on mutation subsets derived from tumor clonality, yet the complex interactions among clones remain a significant obstacle to accurate prognosis. This challenge is further exacerbated by the inherent constraints of limited cohort sizes in clinical studies, which severely compromise model generalizability across heterogeneous cohorts. Therefore, we propose TMBclaw (Tumor Mutation Burden-based Clonal attention with Laplacian Adaptive Weighting), a graph-regularized multi-task learning framework for immunotherapy response prediction. TMBclaw establishes unified integration of group-structured cohorts while enabling cross-cohort knowledge transfer and clonal relationship exploration. For clinical validation, we utilized four cohorts of 238 patients with non-small-cell lung cancer (NSCLC), melanoma, or nasopharyngeal carcinoma treated with ICIs, along with external multicenter validation cohorts (N = 1433) of melanoma and NSCLC patients from public datasets. Comparative analyses demonstrate that TMBclaw significantly outperforms conventional methods in prognostic accuracy and risk stratification. Through systematic quantification of clonal dynamics and discriminative identification of driver clones, TMBclaw shows potential to improve understanding of tumor heterogeneity and provides interpretable insights into the immunotherapy process. Xiaoyan Zhu 0003, Zhili Chang, Xin Lai 0003, Jiayin Wang 0002 |
Briefings Bioinform. | 5 |
| 2025 | MRDagent: iterative and adaptive parameter optimization for stable ctDNA-based MRD detection in heterogeneous samplesabstractMOTIVATION: Minimal residual disease (MRD) as critical biomarker for cancer prognosis and management plays a crucial role in improving patient outcomes. However, detecting MRD via next-generation sequencing-based circulating tumor DNA variant calling remains unstable due to the extremely low variant allele frequency and significant inter- and intra-sample heterogeneity. Although parameter optimization can theoretically enhance the detection performance of variants, achieving stable MRD detection remains challenging due to three key factors: (i) the necessity for individualized parameter tuning across numerous heterogeneous genomic intervals within each sample, (ii) the tightly interdependent parameter requirements across different stages of variant detection workflows, and (iii) the limitations of current automated parameter optimization methods. RESULTS: In this study, we propose MRDagent, a novel variant detection tool designed specifically for MRD detection. MRDagent incorporates an iterative and self-adaptive optimization framework capable of handling unknown objectives, varying constraints, and highly coupled parameters across stages. A key innovation of MRDagent is the integration of a convolutional neural network-based meta-model, trained on historical data to enable rapid parameter prediction. This significantly enhances computational efficiency and generalization performance. Extensive evaluations on simulated and real-world datasets demonstrate MRDagent's superior and stable performance, providing an efficient, reliable solution for MRD detection in clinical and high-throughput research applications. AVAILABILITY AND IMPLEMENTATION: MRDagent is freely available at https://github.com/aAT0047/MRDagent.git. The corresponding dataset and software archive are available at Zenodo: https://doi.org/10.5281/zenodo.15458496. Xin Lai 0003, Shenjie Wang, Yuqian Liu, Xiaoyan Zhu 0003, Jiayin Wang 0002 |
Bioinform. | 5 |
| 2025 | Learning multi-behavior user intent for session-based recommendation
Yu Zhang 0203, Xiaoyan Zhu 0003, Guopeng He, Jiaxuan Li 0001, Jiayin Wang 0002 |
Expert Syst. Appl. | 2 |
| 2025 | MRDtarget: A heuristic Gaussian approach for optimizing targeted capture regions to enhance Minimal Residual Disease detectionabstractMolecular residual disease (MRD) detection, initially developed for hematologic malignancies, has become a critical biomarker for monitoring solid tumors. MRD detection primarily relies on circulating tumor DNA (ctDNA) analysis using next-generation sequencing, offering high sensitivity and broad genomic coverage. However, challenges remain in designing cost-effective panels that maximize mutation detection while maintaining biological relevance. Fixed panels often lack sufficient patient-specific mutation coverage, while WES-based personalized MRD assays, despite their high sensitivity, are costly and less accessible. We developed a tumor comprehensive genomic profiling (CGP)-informed personalized MRD assay to detect tumor-derived mutations, which allowed us to design patient-specific personalized panels and meanwhile, provide a cost-effective alternative to whole exome sequencing (WES). To address these limitations, we developed MRDtarget, a heuristic multivariate Gaussian model-based targeted capture region selection method. By expanding beyond traditional hotspot regions, MRDtarget optimizes variant tracking for MRD detection, significantly improving sensitivity. Using a Bayesian inference-based heuristic approach, MRDtarget integrates multi-feature informativeness rates to identify optimal genomic regions for capture. Experimental results demonstrate that MRDtarget enables the detection of more variants per patient. This study underscores the importance of rational panel design to improve MRD sensitivity and provides a novel approach to enhance precision diagnostics and treatment for solid tumor patients. Xuwen Wang, Yanfang Guan, Xin Lai 0003, Wuqiang Cao, Xiaoyan Zhu 0003, Xiaoling Zeng, Yuqian Liu, Shenjie Wang, Ruoyu Liu, Shuanying Yang, Jiayin Wang 0002 |
PLoS Comput. Biol. | 6 |
| 2024 | Multi-Objective Policy Monitoring Method for Epidemic ControlabstractIn the face of emerging infectious diseases such as COVID-19, timely government intervention is crucial, as swift policy actions can effectively prevent greater losses. However, policymakers often need to balance multiple conflicting objectives. The lack of high-quality data and suitable analytical tools poses significant challenges for policy evaluation, especially in multi-objective decision-making, where accurately assessing the impact of interventions becomes even more difficult. To address this issue, this paper proposes a real-time data-driven policy monitoring method for dynamically tracking the effects of policy interventions. We introduce a new non-parametric one-sided test control chart, leveraging the interpretability and ease of implementation of control charts to monitor risk levels across various policies. Experimental results demonstrate the effectiveness of this method in policy monitoring. Xin Lai 0003, Rundong Fan, Ruoyu Liu, Jiayin Wang 0002, Xiaoyan Zhu 0003, Yuqian Liu, Xuwen Wang, Shenjie Wang |
BIBM | 5 |
| 2024 | An Enhanced Multiple Correction Method with Limited Independent Effective SNPsabstractIn genome-wide association studies (GWAS) and candidate gene studies (CGS), appropriate multiple testing correction methods can effectively address linkage disequilibrium (LD) blocks and are crucial for controlling the family-wise error rate (FWER) while ensuring the reliability of results. In our recent research, we observed that when the number of independent effective Single Nucleotide Polymorphisms (SNPs) is relatively small, existing multiple testing correction methods struggle to effectively control the FWER at 0.05 and lack robustness. This study proposes an enhanced method that builds upon the Moskvina and Schmidt approach by incorporating additional SNP correlation information, allowing for precise control of the FWER at 0.05 with limited independent effective SNPs. We also leverage Monte Carlo integration on a GPU to accelerate the computation of significance thresholds. Our method was evaluated through both simulation studies and real genotype datasets, demonstrating superior performance and robustness compared to existing methods. Xin Lai 0003, Xiaohai Yang, Jiayin Wang 0002, Xiaoyan Zhu 0003, Yuqian Liu, Ruoyu Liu, Xuwen Wang, Shenjie Wang |
BIBM | 4 |
| 2024 | LMR-EWMA: A LASSO-based Multivariate Residual Control Chart for Monitoring Rare Health-Related EventsabstractMonitoring rare health-related events using control charts is crucial for timely detecting potential changes in healthcare scenarios. For example, sequentially testing the level of changes in infectious disease patient numbers helps prepare before an epidemic. Unlike general health-related events, the observation of rare ones often involves an excess of zeros, making it more appropriate to use the zero-inflated Poisson (ZIP) distribution rather than the classical Poisson. Although residual-based charts have attracted significant attention in this field, few studies have explored how to appropriately select residuals with different advantages in the complex situations like healthcare scenarios. Therefore, in this paper, we propose LMR-EWMA, a least absolute shrinkage and selection operator (LASSO)-based multivariate residual exponentially weighted moving average (EWMA) control chart, to automatically select the optimal residuals for monitoring changes (i.e., shifts) in the number of rare health-related events. Additionally, we have innovatively designed a bi-directional moving mechanism to address the limitation of current research in distinguishing the practical significance of shifts. Experimental results on three simulation cases and two real datasets demonstrate that LMR-EWMA outperforms existing charts in monitoring performance. Ruoyu Liu, Jiayin Wang 0002, Xiaoyan Zhu 0003, Yuqian Liu, Shuanying Yang, Xin Lai 0003 |
BIBM | 3 |
| 2024 | RMComBat: A Batch Effect Correction Algorithm for Repeated Measurement Sequencing Data to Prevent OvercorrectionabstractBatch effects, caused by non-biological variations such as differences in laboratory conditions, reagent lots, or personnel, are a substantial source of noise in gene expression data. Accurately correcting these effects is crucial for valid biological inferences. However, the majority of existing batch effect correction algorithms are prone to overcorrection, where biologically meaningful signals are mistakenly identified as noise, especially in repeated measurement studies where time is confounded with batch. The failure to accurately distinguish between batch-related and biologically relevant variation leads to a loss of critical biological information. This paper presents RMComBat, an enhancement of the widely-used ComBat framework, which addresses this limitation by replacing the general linear model with a linear mixed-effects model. RMComBat incorporates subject-specific random intercepts to correct for sample correlation and enhance the preservation of biological signals. We tested RMComBat and several popular algorithms on simulated and real repeated measurement gene expression datasets, evaluating their performance through visual inspections and quantitative metrics. Results indicate that although most algorithms can reduce batch effects, they often do so at the cost of removing true biological signals. RMComBat demonstrates superior performance in preventing overcorrection, providing a more balanced and biologically informative correction in repeated measurement studies, so making it a valuable tool for improving the accuracy of gene expression analyses. Yuqian Liu, Zhaoxing Wei, Jiayin Wang 0002, Xiaoyan Zhu 0003, Ruoyu Liu, Xuwen Wang, Shenjie Wang, Xin Lai 0003 |
BIBM | 4 |
| 2024 | Enabling Adaptive CNV Detection through A Novel Predictive Control FrameworkabstractAccurate detection of copy number variations (CNVs) from sequencing data is crucial in many complex traits and diseases research. Although many CNV detection algorithms have been developed, challenges in precisely identifying CNVs persist. The core statistical model of these algorithms cannot self-adjust, which limits their adaptability to heterogeneous samples and reduces detection accuracy. address this challenge, we reframed the CNV detection problem as a quality control issue and incorporated adaptive mechanisms. We developed adapCNV, a novel adaptive CNV detection framework that integrates machine learning with optimization control. This framework enables dynamic adaptation of primary parameters based on sample features. We defined a quantifiable metric, RD fluctuation values, to assess signal characteristics when the algorithm accurately detects CNVs. We then employed machine learning techniques extract features from panel sequencing data, select initial parameter values for samples, and determine optimal RD fluctuation values. By adopting adaptive model predictive control (AMPC), adapCNV performs optimizations within rolling window. It dynamically adjusts the primary parameters based on error feedback from RD fluctuation values. This adaptive control strategy enables dynamic adjustment automatically match the characteristics of panel sequencing samples, significantly enhancing overall detection quality. The performance of this framework was validated with simulated data. Comparative analysis demonstrated that the proposed method outperforms the baseline approach, particularly in detecting small CNVs. The adapCNV framework is particularly suitable for panel sequencing, which may have broad applications in clinical practice. This novel approach from quality control perspective introduces a new paradigm for CNV detection. Yuqian Liu, Jiajing Yuan, Xiaoyan Zhu 0003, Xin Lai 0003, Ruoyu Liu, Xuwen Wang, Jiayin Wang 0002 |
BIBM | 3 |
| 2024 | Correction of Read Biases Induced by Complex Reference Genome Regions for Improving Copy Number Variation Detection Using a Gaussian Mixture ModelabstractCopy number variations are crucial in cancer research, but their detection through next-generation sequencing is often hindered by read biases, particularly in complex genomic regions. Existing bias-correction methods address common issues like GC content but often fail in regions with repetitive sequences or segmental duplications, leading to false-positive CNVs. We propose refMask, a hybrid Gaussian model-based method that dynamically identifies low-confidence regions in the reference genome, correcting read biases and improving CNV detection accuracy. By integrating features from hg38 and T2T genomes, refMask tailors a custom blacklist for each sequencing sample, enhancing the reliability of CNV detection across diverse conditions. Our method provides a more accurate and flexible solution compared to current fixed blacklists, offering improved performance in challenging genomic regions. Xuwen Wang, Zhili Chang, Shenjie Wang, Ruoyu Liu, Yuqian Liu, Xiaoyan Zhu 0003, Xin Lai 0003, Shuanying Yang, Jiayin Wang 0002 |
BIBM | 6 |
| 2024 | Label-Specific Multi-label Classification with Entropy Guided Clustering
Jiaxuan Li 0001, Xiaoyan Zhu 0003, Jiayin Wang 0002 |
ICPR (2) | 3 |
| 2024 | TMBstable: a variant caller controls performance variation across heterogeneous sequencing samplesabstractIn cancer genomics, variant calling has advanced, but traditional mean accuracy evaluations are inadequate for biomarkers like tumor mutation burden, which vary significantly across samples, affecting immunotherapy patient selection and threshold settings. In this study, we introduce TMBstable, an innovative method that dynamically selects optimal variant calling strategies for specific genomic regions using a meta-learning framework, distinguishing it from traditional callers with uniform sample-wide strategies. The process begins with segmenting the sample into windows and extracting meta-features for clustering, followed by using a pre-trained meta-model to select suitable algorithms for each cluster, thereby addressing strategy-sample mismatches, reducing performance fluctuations and ensuring consistent performance across various samples. We evaluated TMBstable using both simulated and real non-small cell lung cancer and nasopharyngeal carcinoma samples, comparing it with advanced callers. The assessment, focusing on stability measures, such as the variance and coefficient of variation in false positive rate, false negative rate, precision and recall, involved 300 simulated and 106 real tumor samples. Benchmark results showed TMBstable's superior stability with the lowest variance and coefficient of variation across performance metrics, highlighting its effectiveness in analyzing the counting-based biomarker. The TMBstable algorithm can be accessed at https://github.com/hello-json/TMBstable for academic usage only. Shenjie Wang, Xiaoyan Zhu 0003, Xuwen Wang, Yuqian Liu, Minchao Zhao, Zhili Chang, Jiayin Wang 0002 |
Briefings Bioinform. | 2 |
| 2024 | NIPT-PG: empowering non-invasive prenatal testing to learn from population genomics through an incremental pan-genomic approachabstractNon-invasive prenatal testing (NIPT) is a quite popular approach for detecting fetal genomic aneuploidies. However, due to the limitations on sequencing read length and coverage, NIPT suffers a bottleneck on further improving performance and conducting earlier detection. The errors mainly come from reference biases and population polymorphism. To break this bottleneck, we proposed NIPT-PG, which enables the NIPT algorithm to learn from population data. A pan-genome model is introduced to incorporate variant and polymorphic loci information from tested population. Subsequently, we proposed a sequence-to-graph alignment method, which considers the read mis-match rates during the mapping process, and an indexing method using hash indexing and adjacency lists to accelerate the read alignment process. Finally, by integrating multi-source aligned read and polymorphic sites across the pan-genome, NIPT-PG obtains a more accurate z-score, thereby improving the accuracy of chromosomal aneuploidy detection. We tested NIPT-PG on two simulated datasets and 745 real-world cell-free DNA sequencing data sets from pregnant women. Results demonstrate that NIPT-PG outperforms the standard z-score test. Furthermore, combining experimental and theoretical analyses, we demonstrate the probably approximately correct learnability of NIPT-PG. In summary, NIPT-PG provides a new perspective for fetal chromosomal aneuploidies detection. NIPT-PG may have broad applications in clinical testing, and its detection results can serve as a reference for false positive samples approaching the critical threshold. Zhengfa Xue, Aifen Zhou, Xiaoyan Zhu 0003, Huanhuan Zhu, Jiayin Wang 0002 |
Briefings Bioinform. | 3 |
| 2024 | TCSR: Self-attention with time and category for session-based recommendationabstractAbstract Session‐based recommendation that uses sequence of items clicked by anonymous users to make recommendations has drawn the attention of many researchers, and a lot of approaches have been proposed. However, there are still problems that have not been well addressed: (1) Time information is either ignored or exploited with a fixed time span and granularity, which fails to understand the personalized interest transfer pattern of users with different clicking speeds; (2) Category information is either omitted or considered independent of the items, which defies the fact that the relationships between categories and items are helpful for the recommendation. To solve these problems, we propose a new session‐based recommendation method, TCSR (self‐attention with time and category for session‐based recommendation). TCSR uses a non‐linear normalized time embedding to perceive user interest transfer patterns on variable granularity and employs a heterogeneous SAN to make full use of both items and categories. Moreover, a cross‐recommendation unit is adapted to adjust recommendations on the item and category sides. Extensive experiments on four real datasets show that TCSR significantly outperforms state‐of‐the‐art approaches. Xiaoyan Zhu 0003, Yu Zhang 0203, Jiaxuan Li 0001, Jiayin Wang 0002, Xin Lai 0003 |
Comput. Intell. | 1 |
| 2024 | Stacked co-training for semi-supervised multi-label learning
Jiaxuan Li 0001, Xiaoyan Zhu 0003, Hongrui Wang 0004, Yu Zhang 0203, Jiayin Wang 0002 |
Inf. Sci. | 2 |
| 2024 | Graph-enhanced and collaborative attention networks for session-based recommendation
Xiaoyan Zhu 0003, Yu Zhang 0203, Jiayin Wang 0002, Guangtao Wang |
Knowl. Based Syst. | 1 |
| 2024 | Dual-channel graph contrastive learning for multi-label classification with label-specific features and label correlations
Xiaoyan Zhu 0003, Jiaxuan Li 0001, Jiayin Wang 0002 |
Neural Comput. Appl. | 1 |
| 2024 | A ranking-based problem transformation method for weakly supervised multi-label learning
Jiaxuan Li 0001, Xiaoyan Zhu 0003, Weichu Zhang, Jiayin Wang 0002 |
Pattern Recognit. | 2 |
| 2024 | A novel instance-based method for cross-project just-in-time defect predictionabstractSummary Cross‐project (CP) just‐in‐time software defect prediction (JIT‐SDP) uses CP data to overcome initial data scarcity for training high‐performing JIT‐SDP classifiers in the early stages of software projects. The primary challenge faced by JIT‐SDP in a cross‐project context lies in the distinct distributions between training and test data. To tackle this issue, we select source data instances that closely resemble target data for building classifiers. Software datasets commonly exhibit a class imbalance problem, where the ratio of the defective class to the clean class is notably low. This imbalance typically diminishes classifier performance. In this study, we propose an instance selection method utilizing kernel mean matching (ISKMM) that addresses both knowledge transfer and class imbalance in cross‐project defect prediction (CPDP). The method employs the kernel mean matching (KMM) technique to assess the similarity between training and target data. It selects instances with high similarity, retains them, and resamples the data based on similarity weighting to mitigate the class imbalance problem. Our experiments, conducted on 10 open‐source projects, reveal that the ISKMM algorithm outperforms existing CP single‐source software defect prediction (SDP) algorithms. Moreover, when employing the proposed algorithm, defect predictors constructed from cross‐project data demonstrate an overall performance comparable to predictors learned from within‐project data. Xiaoyan Zhu 0003, Jiayin Wang 0002, Xin Lai 0003 |
Softw. Pract. Exp. | 1 |
| 2023 | AdaBoost.C2: Boosting Classifiers Chains for Multi-Label ClassificationabstractDuring the last decades, multi-label classification (MLC) has attracted the attention of more and more researchers due to its wide real-world applications. Many boosting methods for MLC have been proposed and achieved great successes. However, these methods only extend existing boosting frameworks to MLC and take loss functions in multi-label version to guide the iteration. These loss functions generally give a comprehensive evaluation on the label set entirety, and thus the characteristics of different labels are ignored. In this paper, we propose a multi-path AdaBoost framework specific to MLC, where each boosting path is established for distinct label and the combination of them is able to provide a maximum optimization to Hamming Loss. In each iteration, classifiers chain is taken as the base classifier to strengthen the connection between multiple AdaBoost paths and exploit the label correlation. Extensive experiments demonstrate the effectiveness of the proposed method. Jiaxuan Li 0001, Xiaoyan Zhu 0003, Jiayin Wang 0002 |
AAAI | 2 |
| 2023 | Automated machine learning with dynamic ensemble selection
Xiaoyan Zhu 0003, Jingtao Ren, Jiayin Wang 0002, Jiaxuan Li 0001 |
Appl. Intell. | 1 |
| 2023 | Dynamic ensemble learning for multi-label classification
Xiaoyan Zhu 0003, Jiaxuan Li 0001, Jingtao Ren, Jiayin Wang 0002, Guangtao Wang |
Inf. Sci. | 1 |
| 2022 | Just-in-time defect prediction for software hunksabstractAbstract Just‐in‐time defect prediction can remind software developers and managers to verify and fix bugs at the moment they appeared, thus improving the effectiveness and validity of bug fixing. Existing studies mainly focus on just‐in‐time prediction for software files (JIT‐F). JIT‐F is a binary classification problem, which classifies (hence predicts) a file change as buggy or clean. This article provides a detailed analysis of just‐in‐time defect prediction for software hunks (JIT‐H), which predicts bugs at a finer level of granularity, and hence further improves the efficiency of bug fixing. Classification is performed using the ensemble technique of bagging—aggregated combinations of random under sampling plus multiple classifiers (J48 and Random Forest). An empirical study with 10 open source projects was conducted to validate the effectiveness of JIT‐H. Experimental results show that JIT‐H is effective at predicting defects in software hunk changes. Compared with JIT‐F, JIT‐H is more cost effective. Additionally, analysis on the change features indicates that Text Vector features and hunk change level features are of more importance than features in other groups and levels. Xiaoyan Zhu 0003, Chenyu Yan, E. James Whitehead Jr., Binbin Niu, Lei Zhu 0011, Long Pan |
Softw. Pract. Exp. | 1 |
| 2021 | A new multiple instance algorithm using structural informationabstractMultiple instance learning (MIL) is semisupervised learning that predicts the label of a bag with a wide diversity of instances. It has many applications and thus attracts increasingly more attention. In this paper, we propose a new MIL algorithm using the structural information of a bag to predict its label. In the proposed method, a bag is transformed into a graph, and spectral clustering is employed to divide the graph into several subgraphs. Then, the graph Fourier transform is utilized to extract the features of the subgraphs. Finally, an end-to-end neural network is used to predict the label of a bag with the extracted features. An empirical study with 25 datasets was conducted to validate the effectiveness of the proposed method. The experimental results show that the proposed method performs better than the 6 baseline methods on most datasets. Xiaoyan Zhu 0003, Jiayin Wang 0002, Yuqian Liu |
ICDM | 1 |
| 2021 | Ensemble of ML-KNN for classification algorithm recommendation
Xiaoyan Zhu 0003, Chenzhen Ying, Jiayin Wang 0002, Jiaxuan Li 0001, Xin Lai 0003, Guangtao Wang |
Knowl. Based Syst. | 1 |
| 2021 | Automatic Recommendation of a Distance Measure for Clustering AlgorithmsabstractWith a large number of distance measures, the appropriate choice for clustering a given data set with a specified clustering algorithm becomes an important problem. In this article, an automatic distance measure recommendation method for clustering algorithms is proposed. The recommendation method consists of the following steps: (1) metadata extraction, including meta-feature collection and meta-target identification; (2) recommendation model construction using metadata; and (3) distance measure recommendation for a new data set by the recommendation model. Two different types of meta-targets and meta-learning techniques are utilized considering the possible different requirements of users. To validate the necessity and effectiveness of the distance measure recommendation method, an empirical study is conducted with 199 publicly available data sets, 9 distance measures, and 2 widely used clustering algorithms. The experimental results indicate that distance measure significantly influences the performance of the clustering algorithm for a given data set. Furthermore, performance analysis of the proposed recommendation method proves its effectiveness. Xiaoyan Zhu 0003, Yingbin Li, Jiayin Wang 0002, Jingwen Fu |
ACM Trans. Knowl. Discov. Data | 1 |
| 2019 | An Artificial Fish Swarm Algorithm for Identifying Associations between Multiple Variants and Multiple PhenotypesabstractIdentifying associations between genomic variants and phenotypes has always been an interesting research field of population genetics, which is of great significance for studying the pathogenesis of complex diseases and supporting clinical assistant decision making. Nowadays, many identification methods have been proposed to find the associations between variants and phenotypes, such as GWAS and pheWAS, and have made excellent achievements in pathological research and clinical practice. However, the existing methods only focus on single phenotype-multiple variants or single variant-multiple phenotypes, but not on multiple variants-multiple phenotypes. In the view of the fact that complex diseases often have several subtypes which differ greatly in variants and phenotypes, focusing only on single variant or single phenotype is far from enough and limits the ability of identification of those methods. Therefore, we propose a heuristic method with an AFSA framework on the solution space to identify associations between multiple variants and multiple phenotypes. In our method, each fish carries two logic trees that respectively represent the associations between variants and the associations between phenotypes. The logic trees will be iteratively updated to find a better solution according to the preset update strategies. When the iteration stop condition is reached, the algorithm will stop and output the optimal fish. The logical expression represented by the logic trees carried by the optimal fish is the associations we find. We validated the proposed method on the simulation data generated by hapgen2 and PhenotypeSimulator, and took the ratio of the number of people that can be explained by the found logical expression as the index to evaluate the performance, which was called Coverage. We conducted 9 groups of experiments, each of which was different in the number of variants and phenotypes. The best Coverage of was from the group including 500 variants and 10 phenotypes, which reached 72.12%, and the worst result is from the group including 100 variants and 20 phenotypes, 31.73%. We also exhausted the simulation data to find the optimal logical expression and several most important logic rules to evaluate the results obtained by the method. Ruoyu Liu, Xin Lai 0003, Xuanping Zhang, Xiaoyan Zhu 0003, Jiayin Wang 0002 |
BIBM | 5 |
| 2019 | GSDcreator: An Efficient and Comprehensive Simulator for Genarating NGS Data with Population Genetic InformationabstractIn recent decades, NGS data analysis has become a major research field in bioinformatics, which presents great advantages in many application scenarios. Many algorithms and software were designed for analyzing the NGS data, while simulation datasets are urgently needed for testing software and optimizing their parameter configurations. Thus, a series of NGS data simulators have been published. However, the existing simulators cannot satisfy the requirements from many specific scenarios. First, they do not support many newly discovered variations. Second, complex structural variations are difficult to generate. In addition, along with the increase of population data, it is urgent to increase population information simulation. In this paper, we propose GSDcreator, a comprehensive NGS simulator that overcome the three weaknesses mentioned above. It can produce all known types of variation, where the complex of variations are also supported. Furthermore, it can capture many important real data features including population polymorphism, insert size distribution, adjacent site depth distribution, overall depth distribution, quality score distribution, amplification bias, sequencing errors and so on. It's highlighted that 1000 Genomes Project Database is taken as a reference and integrates population genetic information to simulate population polymorphism. To test the performance, we did a lot of experiments and found that simulated data produced by GSDcreator are quit mimic to the real sequencing data. Shenjie Wang, Jiayin Wang 0002, Xuanping Zhang, Xuwen Wang, Xiaoyan Zhu 0003, Xin Lai 0003 |
BIBM | 6 |
| 2019 | FilterLAP: Filtering False-positive Mutation Calls via a Label Propagation FrameworkabstractBenefiting from the recent advantages of genomic sequencing, detecting genomic mutations becomes a routine work in precise diagnoses and treatments for cancers. In clinical practices, many factors, such as tumor purity, clonal structure, etc., interfere the performance of calling mutations. The computational pipelines prefer to sensitively report the candidate calls, while a filter is applied for removing the false-positive calls. The existing filters rely on the whole genome/exome sequencing data, which can provide sufficient samples for training the filters. However, the gene-panel sequencing is more popular in clinical practices, but there is no practical filter for limited training samples. In light of this, we develop a semi-learning filter for gene-panel sequencing data, FilterLAP, which implemented via a label propagation framework. Given few labeled samples with a set of unlabeled ones, its basic idea is to predict the label information of unlabeled nodes from the label information of labeled nodes, and establishes a complete graph model by using the relationship between samples, by combining transductive inference with label propagation algorithm. For each node in the network, tags are propagated to adjacent nodes according to similarity and the probability distribution of similar nodes tends to be similar and can be divided into a class. We perform multiple sets of experiments on gene-panel sequencing data captured from Illumina platform. FilterLAP outperforms on both SNV and INDEL filtering, where the AUCs reach 0.90-0.97, and the average accuracies on overall mutation calls are over 90%. Comparing to GATK hard filters, FilterLAP present a 5% improvement on accuracy. These results demonstrate that the proposed method can better reduce the false positive mutation calls on gene-panel sequencing data. In addition, it is stable and efficient, which can be used as a practical tool for mutation call filtering for gene-panel sequencing data. Xuwen Wang, Xiaoyan Zhu 0003, Shenjie Wang, Xuanping Zhang, Xin Lai 0003, Jiayin Wang 0002 |
BIBM | 2 |
| 2019 | A new unsupervised feature selection algorithm using similarity-based feature clusteringabstractAbstract Unsupervised feature selection is an important problem, especially for high‐dimensional data. However, until now, it has been scarcely studied and the existing algorithms cannot provide satisfying performance. Thus, in this paper, we propose a new unsupervised feature selection algorithm using similarity‐based feature clustering, Feature Selection‐based Feature Clustering (FSFC). FSFC removes redundant features according to the results of feature clustering based on feature similarity. First, it clusters the features according to their similarity. A new feature clustering algorithm is proposed, which overcomes the shortcomings of K‐means. Second, it selects a representative feature from each cluster, which contains most interesting information of features in the cluster. The efficiency and effectiveness of FSFC are tested upon real‐world data sets and compared with two representative unsupervised feature selection algorithms, Feature Selection Using Similarity (FSUS) and Multi‐Cluster‐based Feature Selection (MCFS) in terms of runtime, feature compression ratio, and the clustering results of K‐means. The results show that FSFC can not only reduce the feature space in less time, but also significantly improve the clustering performance of K‐means. Xiaoyan Zhu 0003, Yu Wang 0069, Yingbin Li, Yonghui Tan, Guangtao Wang, Qinbao Song |
Comput. Intell. | 1 |
| 2018 | A new classification algorithm recommendation method based on link prediction
Xiaoyan Zhu 0003, Xiaomei Yang, Chenzhen Ying, Guangtao Wang |
Knowl. Based Syst. | 1 |
| 2018 | Software change-proneness prediction through combination of bagging and resampling methodsabstractAbstract Identifying the change‐prone parts of software could help managers and developers to effectively allocate maintenance resource and time during early phases of software life cycle. Change‐proneness prediction on file level with binary classification methods makes such identification possible. As the fact that change‐prone files frequently account for a small part of all the files, the prediction performance of standard classification methods is not satisfying. In this paper, we employ imbalanced learning methods, including bagging, resampling, and especially their combination to reduce the performance decrease of standard classifiers caused by the class imbalance problem in change‐proneness prediction. Besides, we propose a boxplot‐based partition method to provide more proper change‐proneness label designation for the training data. Eight open‐source Java projects are chosen in the empirical study to validate the effectiveness of the combination methods in change‐proneness prediction. The experimental results of the empirical study show that combining bagging with resampling can significantly improve the prediction performance of only bagging or resampling. Of all the combination methods employed, combination of bagging with undersampling performs better than others. And support vector machine is more effective as a base classifier than J48 and naive Bayes. Xiaoyan Zhu 0003, Yueyang He, Xiaolin Jia, Lei Zhu 0011 |
J. Softw. Evol. Process. | 1 |
| 2018 | An empirical study of software change classification with imbalance data-handling methodsabstractSummary Bug prediction in software code changes can help developers to find out and fix bugs immediately when they are introduced, thus to improve the effectiveness and validity of bug fixing. In data mining, this problem can be regarded as a change classification task. However, one of its key characteristics, ie, class‐imbalance, holds back the performance of standard classification methods. In this paper, we consider a quantity of imbalance data‐handling methods and extract a more comprehensive groups of change features, aiming to achieve better change classification performance. Two different types of imbalance data‐handling methods, namely, resampling and ensemble learning methods, are employed. Especially, we explore the performance of their combination. To compare the performance of different imbalance data‐handling methods, an experiment with 10 open source projects is conducted. Four classification methods, including J48, Naïve Bayes, SMO, and Random Forest, are used as standard classifiers and as the base classifiers, respectively. Moreover, contribution of different groups of change features are evaluated. Experimental results show that imbalance data‐handling methods can improve the performance of change classification and the combination methods, which take advantage of both ensemble learning and resampling, perform better than using ensemble learning methods or resampling methods individually. Of the studied imbalance data‐handling methods, the combination of Bagging and random undersampling with J48 as the base classifier yields out better prediction results than those achieved by other methods. Additionally, of the collected change features, text vector features accounts for a larger proportion than others. Xiaoyan Zhu 0003, Binbin Niu, E. James Whitehead Jr., Zhongbin Sun |
Softw. Pract. Exp. | 1 |
| 2016 | A machine learning based software process model recommendation method
Qinbao Song, Xiaoyan Zhu 0003, Guangtao Wang, Heli Sun, Chenhao Xue, Baowen Xu |
J. Syst. Softw. | 2 |
| 2015 | A Selective Detector Ensemble for Concept Drift DetectionabstractConcept drifts usually originate from many causes instead of only one, which result in two types of concept drifts: abrupt drifts and gradual drifts. From the point of view of speed, concept drifts pose strong challenges for data stream mining. In this paper, we propose a selective detector ensemble to detect both abrupt and gradual drifts. We first present our detector ensemble construction method, and then introduce how to use this ensemble to detect concept drifts with the proposed early-find-early-report rule. To evaluate the performance of our method, we compare it with four drift detection methods on eight publicly available data sets containing various concept drifts. The experimental results show that compared with those benchmarks, our ensemble method can effectively improve the recall and false negative rate without significantly increasing the false positive rate, and has stronger generalization ability than those single-change-indicator-based methods. Lei Du 0001, Qinbao Song, Lei Zhu 0011, Xiaoyan Zhu 0003 |
Comput. J. | 4 |
| 2015 | A novel ensemble method for classifying imbalanced data
Zhongbin Sun, Qinbao Song, Xiaoyan Zhu 0003, Heli Sun, Baowen Xu, Yuming Zhou |
Pattern Recognit. | 3 |
| 2015 | An analysis of programming language statement frequency in C, C++, and Java source codeabstractSummary Statement frequency data can inform programming language research and provide a solid basis for frequency‐based code analysis. This paper presents an analysis of programming language statement frequency in a large corpus of C, C++, and Java source code, comprised of more than 54 million lines of code. Across these languages, the top four work‐performing statement types are Method/Function Call, Assignment, If, and Return. As compared to studies of Formula Translating System, Common Business Oriented Language and Programming Language One in the 1970s, the main change is the prevalence of method/function calls. Statement use frequency across languages is remarkably similar, and within each individual language, most statement types have a frequency distribution that occupies a small range. A more detailed examination of assignment and looping statement types shows that many assignments simply involve copying of data and that C++/Java useforstatements more than C. Copyright © 2014 John Wiley & Sons, Ltd. Xiaoyan Zhu 0003, E. James Whitehead Jr., Caitlin Sadowski, Qinbao Song |
Softw. Pract. Exp. | 1 |
| 2013 | Does bug prediction support human developers? findings from a google case studyabstractWhile many bug prediction algorithms have been developed by academia, they're often only tested and verified in the lab using automated means. We do not have a strong idea about whether such algorithms are useful to guide human developers. We deployed a bug prediction algorithm across Google, and found no identifiable change in developer behavior. Using our experience, we provide several characteristics that bug prediction algorithms need to meet in order to be accepted by human developers and truly change how developers evaluate their code. Chris Lewis 0002, Zhongpeng Lin, Caitlin Sadowski, Xiaoyan Zhu 0003, Rong Ou, E. James Whitehead Jr. |
ICSE | 4 |
| 2012 | Using Coding-Based Ensemble Learning to Improve Software Defect PredictionabstractUsing classification methods to predict software defect proneness with static code attributes has attracted a great deal of attention. The class-imbalance characteristic of software defect data makes the prediction much difficult; thus, a number of methods have been employed to address this problem. However, these conventional methods, such as sampling, cost-sensitive learning, Bagging, and Boosting, could suffer from the loss of important information, unexpected mistakes, and overfitting because they alter the original data distribution. This paper presents a novel method that first converts the imbalanced binary-class data into balanced multiclass data and then builds a defect predictor on the multiclass data with a specific coding scheme. A thorough experiment with four different types of classification algorithms, three data coding schemes, and six conventional imbalance data-handling methods was conducted over the 14 NASA datasets. The experimental results show that the proposed method with a one-against-one coding scheme is averagely superior to the conventional methods. Zhongbin Sun, Qinbao Song, Xiaoyan Zhu 0003 |
IEEE Trans. Syst. Man Cybern. Part C | 3 |
| 2011 | An empirical analysis of the FixCache algorithmabstractThe FixCache algorithm, introduced in 2007, effectively identifies files or methods which are likely to contain bugs by analyzing source control repository history. However, many open questions remain about the behaviour of this algorithm. What is the variation in the hit rate over time? How long do files stay in the cache? Do buggy files tend to stay buggy, or can they be redeemed? This paper analyzes the behaviour of the FixCache algorithm on four open source projects. FixCache hit rate is found to generally increase over time for three of the four projects; file duration in cache follows a Zipf distribution; and topmost bug-fixed files go through periods of greater and lesser stability over a project's history. Caitlin Sadowski, Chris Lewis 0002, Zhongpeng Lin, Xiaoyan Zhu 0003, E. James Whitehead Jr. |
MSR | 4 |