VLDB 2026 Research / reviewers in the wild / expert
Shikai Guo
dblp:186/6586
· DBLP profile ↗
71ranked-venue papers
13as first author
61since 2021 · last 2026
0000-0002-8554-6365ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 31 · 9 first-author · 26 since 2021Applied, interdisciplinary, general and emerging computing · 19 · 2 first-author · 18 since 2021Artificial intelligence and machine learning · 13 · 2 first-author · 8 since 2021Systems, architecture and hardware · 8 · 8 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Modeling Relational Logic Circuits for and-Inverter Graph Convolutional NetworkabstractThe automation of logic circuit design enhances chip performance, energy efficiency, and reliability, and is widely applied in the field of Electronic Design Automation (EDA). And-Inverter Graphs (AIGs) efficiently represent, optimize, and verify the functional characteristics of digital circuits, enhancing the efficiency of EDA development. Due to the complex structure and large scale of nodes in real-world AIGs, accurate modeling is challenging, leading to existing work lacking the ability to jointly model functional and structural characteristics, as well as insufficient dynamic information propagation capability. To address the aforementioned challenges, we propose AIGer, with the aim to enhance the expression of AIGs and thereby improve the efficiency of EDA development. Specifically, AIGer consists of two components: 1) Node logic feature initialization embedding component and 2) AIGs feature learning network component. The node logic feature initialization embedding component projects logic nodes, such as AND and NOT, into independent semantic spaces, to enable effective node embedding for subsequent processing. Building upon this, the AIGs feature learning network component employs a heterogeneous graph convolutional network, designing dynamic relationship weight matrices and differentiated information aggregation approaches to better represent the original structure and information of AIGs. The combination of these two components enhances AIGer’s ability to jointly model functional and structural characteristics and improves its message passing capability, thereby strengthening its expressive power for AIGs and enhancing the development efficiency of logic circuits. Experimental results indicate that AIGer outperforms the current best models in the Signal Probability Prediction (SPP) task, improving MAE and MSE by 18.95% and 44.44%, respectively. In the Truth Table Distance Prediction (TTDP) task, AIGer achieves improvements of 33.57% and 14.79% in MAE and MSE, respectively, compared to the best-performing models1. Weihao Sun, Shikai Guo, Qian Ma 0003, Hui Li 0014, Yongpeng Weng |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2026 | Cross Attention and Intra-Layer Attention in Heterogeneous Graph Neural Networks for Drug-Target Interaction PredictionabstractIn recent years, computational prediction of drug-target interactions (DTIs) has become essential for drug discovery and repositioning. However, traditional experimental approaches for DTI identification are time-consuming and costly. To address this, many machine learning-based methods have been developed, yet most existing models neglect important information interaction between drugs and targets in drug-target pairs (DTPs) during drug-target interaction. In this study, we propose a novel cross-attention and intra-layer attention mechanism within a heterogeneous graph neural network (CAIHGNN) for DTI prediction. The cross-attention mechanism allows for dynamic learning of feature correlations between drugs and targets, while the intra-layer attention captures both explicit and implicit interactions within DTPs. Additionally, we introduce a drug-target pair correlation graph to exploit high-order interactions between DTPs. Extensive experiments on two biological heterogeneous datasets demonstrate the superior performance of our proposed method in accurately predicting DTIs. Furthermore, the model exhibits robust generalization in case study, showing promise for real-world drug discovery applications. Kuiyang Che, Xirun Wei, Hui Li 0014, Shikai Guo |
IEEE Trans. Comput. Biol. Bioinform. | 6 |
| 2026 | Molecular-Driven Multi-View Hypergraph Contrastive Learning for Drug-Drug Interaction PredictionabstractRecent concerns have arisen over adverse reactions caused by drug combinations, and drug-drug interaction (DDI) prediction helps identify potential risks by forecasting interactions between drugs. Previous methods have primarily explored drug interactions from the superficial level of drug molecules, often overlooking the internal structural information of the molecules. To this end, we propose Mol-HCL, a multi-view hypergraph contrastive learning framework based on molecular view. In this framework, we construct the molecular view to learn the internal information of drug molecules and, based on this, develop structural view and semantic view. These three views collaboratively learn both intra-molecular and inter-molecular information. Subsequently, we incorporate hypernodes into the structural view and design a novel hyperchain, integrating it into the semantic view to capture latent neighbor drug node structural relationships and long-range DDI chain semantic information. After that, contrastive learning is performed between the structural hypergraph and the molecular view, as well as between the semantic hypergraph and the molecular view, to enhance the representations learned from the molecular view. Finally, we conduct experiments on two real-world scientific datasets. The experimental results demonstrate a significant improvement of Mol-HCL over existing methods, showcasing its effectiveness and advantages in DDI prediction. Shikai Guo, Hui Li 0014, Qian Ma 0003 |
IEEE Trans. Comput. Biol. Bioinform. | 4 |
| 2026 | EnsDiffAD: Ensemble Diffusion Models for Multivariate Time Series Anomaly Detection
Qian Ma 0003, Yanyang Li, Mei Bai, Xite Wang, Shikai Guo, Yu Gu 0002, Ge Yu 0001 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2026 | Commercial Cyber-Physical System Development Tool Chain Bug Detecting via Diversity-Guided Fuzzing TestabstractSimulink is MathWorks’ commercial cyber-physical systems development tool, which enables engineers to do rapid prototyping of their systems through simulation and embedded code generation. When a Simulink model meets requirements, engineers can utilize embedded coder to convert it into embedded code (e.g., C source code) and deploy it in safety-critical applications such as automotive, aerospace, and healthcare. However, bugs or incorrect implementations in code generation may lead to unexpected behaviors in target applications, posing security risks. Therefore, it is crucial to eliminate such bugs in embedded code generation. To address this issue, we propose DESCO, a differential testing approach to test embedded code generation in Simulink. DESCO considers the functional correlation between Simulink blocks for partitioning, aiming to generate diverse and complex bug-triggering Simulink models to thoroughly exercise the embedded code generation. DESCO then detects bugs by analyzing the outputs of these Simulink models by differential testing. The experiments demonstrate that DESCO significantly outperforms existing approaches. In three months, DESCO reported 16 issues, including 12 confirmed as bugs by MathWorks Support. Huijiang Liu, Shikai Guo, Jiaxue Liu, Hongyi Cheng, He Jiang 0001 |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2026 | FPGA Interactive Debugging Tools Testing via Mutation Diversification Search
Shikai Guo, Zong Liu, He Jiang 0001 |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2026 | Is Fault Localization Effective on Industrial Software? A Case Study on Computer-Aided Engineering ProjectsabstractIn software engineering, empirical studies on automated fault localization (FL) methods mainly focus on general software, and substantial progress has been made. However, the applicability and efficacy of these methods in specialized, domain-specific software like industrial software remains under-explored. Such specialized software is usually characterized by complex inputs and iterative computing paradigms, which could significantly influence the effectiveness of existing FL methods. To address this gap, this study takes a typical categorical of industrial software (i.e., computer-aided engineering (CAE) projects) as a case study, to investigate the feasibility and effectiveness of state-of-the-art FL methods within CAE projects. Through the reproduction of 76 real-world bugs from three widely used CAE projects (i.e., FDS, deal.II, and MFEM), we find that even the most precise FL methods require developers to examine on average 467.18 statements before finding bugs and can take 208.13 hours to execute. The complex inputs and long-term computation characteristics of CAE projects further increase the difficulty of FL. Moreover, FL on CAE also faces challenges, such as insufficient differentiation of coverage information and missing CAE-specific FL features. Based on our findings, we improve FL on CAE projects by proposing a set of CAE main module-based features, which improve the best-performed FL method in this study (i.e., DeepFL) by 35.93% and 45%, in terms of MAR and MFR , respectively. Zhilei Ren, Shikai Guo, He Jiang 0001 |
ACM Trans. Softw. Eng. Methodol. | 4 |
| 2026 | FoC: Figure Out the Cryptographic Functions in Stripped Binaries with LLMsabstractAnalyzing the behavior of cryptographic functions in stripped binaries is a challenging but essential task, which is crucial in software security fields such as malware analysis and legacy code inspection. However, the inherent high logical complexity of cryptographic algorithms makes their analysis more difficult than that of ordinary code, and the general absence of symbolic information in binaries exacerbates this challenge. Existing methods for cryptographic algorithm identification frequently rely on data or structural pattern matching, which limits their generality and effectiveness while requiring substantial manual effort. In response to these challenges, we present F igure o ut the C ryptographic functions (FoC), a novel framework that leverages Large Language Models (LLMs) to identify and analyze cryptographic functions in stripped binaries. In FoC, we first build an LLM-based generative model ( FoC-BinLLM ) to summarize the semantics of cryptographic functions in natural language form, which is intuitively readable to analysts. Subsequently, based on the semantic insights provided by FoC-BinLLM, we further develop a binary code similarity detection model ( FoC-Sim ), which allows analysts to effectively retrieve similar implementations of unknown cryptographic functions from a library of known cryptographic functions. The predictions of generative model like FoC-BinLLM are inherently difficult to reflect minor alterations in binary code, such as those introduced by vulnerability patches. In contrast, the change-sensitive representations generated by FoC-Sim compensate for the shortcomings to some extent. To support the development and evaluation of these models, and to facilitate further research in this domain, we also construct a comprehensive cryptographic binary dataset and introduce an automatic method to create semantic labels for extensive binary functions. Our evaluation results are promising. FoC-BinLLM outperforms ChatGPT by 14.61% on the ROUGE-L score, demonstrating superior capability in summarizing the semantics of cryptographic functions. FoC-Sim also surpasses previous best methods with a 52% higher Recall@1 in retrieving similar cryptographic functions. Beyond these metrics, our method has proven its practical utility in real-world scenarios, including cryptographic-related virus analysis and 1-day vulnerability detection. Xiuwei Shang, Shaoyin Cheng, Shikai Guo, Weiming Zhang 0001, Nenghai Yu |
ACM Trans. Softw. Eng. Methodol. | 4 |
| 2026 | Estimating Uncertainty in Line-Level Defect Prediction via Perceptual Borderline OversamplingabstractSoftware defect prediction aims to identify potentially defective software modules using various techniques, while fine-grained line-level defect prediction can pinpoint defective lines of code. This helps developers promptly discover and fix errors, thereby enhancing the efficiency of testing and code review. However, previous studies often overlook the impact of characterizing noise and the skewed distribution of defect knowledge in software projects, making it difficult for current methods to achieve satisfactory accuracy and cost-effectiveness in software defect prediction. To address these challenges, we propose a model named EU-LLDP, which effectively resolves the issue of low cost-effectiveness in line-level defect prediction models. Specifically, the EU-LLDP model consists of two main components: the defect mining component mines the most valuable defect knowledge from numerous software defects using prediction probability matrices, noise labels, and the borderline information of code vectorizations. The adaptive resampling component samples valuable defect knowledge through the density distribution of defect knowledge, thereby making full use of existing defect knowledge and improving the cost-effectiveness of line-level software defect prediction models. Seven comprehensive experiments were conducted on 32 defect datasets from 9 Java open source systems using file-level prediction models and line-level defect prediction models to evaluate the effectiveness of the EU-LLDP model. The EU-LLDP model improves the state-of-the-art file-level defect prediction model in terms of Balanced Accuracy by 9.87%, the MCC by 38.09%, and enhances the state-of-the-art line-level defect prediction method in terms of Recall@Top20%LOC by 44.16%, and Effort@Top20%Recall by 17.62%. These results fully demonstrate the effectiveness of EU-LLDP in improving the accuracy and cost-effectiveness of Software defect prediction. Shikai Guo, Hui Li 0014, Rong Chen 0003 |
ACM Trans. Softw. Eng. Methodol. | 2 |
| 2026 | Transformation-Recipe-Based FPGA Synthesis Compiler Testing
Yi Zhang 0148, He Jiang 0001, Shikai Guo, Zun Wang 0005 |
IEEE Trans. Reliab. | 3 |
| 2026 | FuncGNN: Learning Functional Semantics of Logic Circuits with Graph Neural NetworksabstractAs integrated circuit scale grows and design complexity rises, effective circuit representation helps support logic synthesis, formal verification, and other automated processes in electronic design automation. And-Inverter Graphs (AIGs), as a compact and canonical structure, are widely adopted for representing Boolean logic in these workflows. However, the increasing complexity and integration density of modern circuits introduce structural heterogeneity and global logic information loss in AIGs, posing significant challenges to accurate circuit modeling. To address these issues, we propose FuncGNN, which integrates hybrid feature aggregation to extract multi-granularity topological patterns, thereby mitigating structural heterogeneity and enhancing logic circuit representations. FuncGNN further introduces gate-aware normalization that adapts to circuit-specific gate distributions, improving robustness to structural heterogeneity. Finally, FuncGNN employs multi-layer integration to merge intermediate features across layers, effectively synthesizing local and global semantic information for comprehensive logic representations. Experimental results on two logic-level analysis tasks (i.e., signal probability prediction and truth-table distance prediction) demonstrate that FuncGNN outperforms existing state-of-the-art methods, achieving improvements of 2.06% and 18.71%, respectively, while reducing training time by approximately 50.6% and GPU memory usage by about 32.8%. The code is available at https://github.com/Vandbs/FuncGNN . Qiyun Zhao, Shikai Guo, He Jiang 0001 |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2026 | NSGen: A Template-Based Framework to Find Bugs in Computer-Aided Engineering ToolsabstractComputer-aided engineering (CAE) tools are extensively used in safety-critical domains like aerospace design to simulate real-world physical processes on computers at reduced costs. However, CAE tools are prone to bugs, leading to incorrect simulation and serious design flaws. Existing testing methods have limited success in finding these bugs, since constructing test cases for CAE tools (known as CAE inputs) is challenging, due to the complexity of input parameter constraints and the differences of input syntax for each tool. Therefore, we propose NSGen, a template-based Numerical Simulation test caseGENerator for effective CAE input generation. To bridge input differences, NSGen designs general syntax rules that ignore semantic details of CAE inputs but retaining the format and validity of these inputs for different CAE tools. NSGen then designs semantic rules to define the parameters and their intricate constraints (i.e., dependency, exclusion, and extension). By instantiating these rules as templates, valid CAE inputs are generated using a syntax parser implemented by NSGen. Experiments show that NSGen can effectively generate CAE inputs with less than 0.33 second on average, triggering 6.28 to 10.55 times more potential issues than the baseline. Using these inputs, NSGen finds 12 bugs in popular CAE tools, including crash bugs and instability bugs. Peiyu Zou, Shikai Guo, Zhilei Ren, He Jiang 0001 |
IEEE Trans. Software Eng. | 4 |
| 2026 | ProFuse: Test Case Prioritization Based on Multi Dimensional Feature Fusion for Logic Synthesis Tools Testing AccelerationabstractLogic synthesis tools translate Hardware Description Language (HDL) designs into hardware implementation. To test these tools, numerous test cases are usually executed on the tools, yet only a few of them can trigger faults, leading to inefficient testing. Since executing test cases on logic synthesis tools often requires significant cost on complicated synthesis and simulation, fault-triggering test cases should be prioritized to execute. However, existing prioritization methods face challenges in accurately predicting the fault-triggering capability of dynamically generated test cases and modeling the unique syntactic and structure complexities of these HDL-based programs.Therefore, we propose ProFuse, a multi-dimensional feature fusion method for logic synthesis tool test case prioritization. ProFuse leverages Abstract Syntax Trees (AST) and Data Flow Graphs (DFG) to extract novel syntactic and structure features from HDL designs. These features are processed by a joint model of Multilayer Perceptron (MLP) and Graph Convolutional Network (GCN) to rank fault-triggering test cases accurately. ProFuse achieves an Average Percentage of Fault Detection (APFD) score of 0.9285, outperforming the state-of-the-art prioritization methods by 11.38% to 82.49%. ProFuse can efficiently rank randomly generated test cases to discover 15 new faults in logic synthesis tools (i.e., Yosys and Vivado). The Vivado community acknowledged our work for improving their tool. Peiyu Zou, Shikai Guo, Zhide Zhou, He Jiang 0001 |
IEEE Trans. Software Eng. | 4 |
| 2025 | IAMU-Net: Integrated Attention Mechanisms with Coarse-to-Fine Strategy Based on nnU-Net for Multi-Organ Image SegmentationabstractAutomatic multi-organ segmentation plays a critical role in medical image analysis, particularly in preoperative planning and disease assessment. However, conventional segmentation methods often struggle with challenges in handling small organs, ambiguous boundaries, and complex backgrounds, as well as in their generalization ability across different datasets. To address these limitations, we propose a novel model named IAMU-Net for multi-organ segmentation of abdominal computed tomography (CT) scans using deep learning. Our method is built upon the nnU-Net architecture and integrates multiple attention mechanisms to enhance feature representation, spatial information modeling, and multi-scale feature fusion abilities. Furthermore, we use a coarse-to-fine two-stage segmentation strategy to improve the segmentation accuracy of small organs and anatomically complex boundaries. Experimental results on the AbdomenAtlas 1.0 Mini and TotalSegmentator V2 datasets illustrate that our method achieves superior segmentation performance across multiple organs, confirming its effectiveness and robustness. Chengyu Zhao, Yuchen Pei, Shikai Guo |
BIBM | 5 |
| 2025 | Live Region Mutation Testing for Commercial Cyber-Physical System Development Tool ChainabstractMathWorks Simulink, a commercial CPS development tool chain, is widely used as an industry standard for designing and analyzing system behavior and generating embedded code for deployment. However, bugs in Simulink can cause unexpected behaviors during model compilation, making their elimination critical. Existing methods face two key challenges: generating equivalent models with varied data flows (data flow equivalence) and creating diverse block types to comprehensively test the compiler (mutation diversity). To address these, we propose LION, a differential testing approach. LION ensures data flow equivalence by inserting “store-revert” block pairs between existing blocks and tackles mutation diversity by employing Markov Chain Monte Carlo (MCMC) sampling to generate diverse new blocks. Differential testing is then used to identify bugs. Experiments show LION outperforms state-of-the-art approaches like SLforge, SLEMI, and COMBAT, detecting 610 additional compiler bugs in two weeks. Over two months, LION uncovered and reported 16 valid bugs in the widely used stable version of Simulink. Lehuan Zhang, Shikai Guo, He Jiang 0001 |
DAC | 2 |
| 2025 | LBMKGC: Large Model-Driven Balanced Multimodal Knowledge Graph CompletionabstractMulti-modal Knowledge Graph Completion (MMKGC) aims to predict missing entities, relations, or attributes in knowledge graphs by collaboratively modeling the triple structure and multimodal information (e.g., text, images, videos) associated with entities.
This approach facilitates the automatic discovery of previously unobserved factual knowledge.
However, existing MMKGC methods encounter several critical challenges: (i) the imbalance of inter-entity information across different modalities; (ii) the heterogeneity of intra-entity multimodal information; and (iii) for a given entity, the informational contributions of different modalities are inconsistent across contexts.
In this paper, we propose a novel **L**arge model-driven **B**alanced **M**ultimodal **K**nowledge **G**raph **C**ompletion framework, termed LBMKGC.
Subsequently, to bridge the semantic gap between heterogeneous modalities, LBMKGC aligns the multimodal embeddings of entities semantically by using the CLIP (Contrastive Language-Image Pre-Training) model.
Furthermore, LBMKGC adaptively fuses multimodal embeddings with relational guidance by distinguishing between the perceptual and conceptual attributes of triples.
Finally, extensive experiments conducted against 21 state-of-the-art baselines demonstrate that LBMKGC achieves superior performance across diverse datasets and scenarios while maintaining efficiency and generalizability.
Our code and data are publicly available at: https://github.com/guoynow/LBMKGC. Qian Ma 0003, Hui Li 0014, Furui Zhan, Yu Gu 0002, Ge Yu 0001, Shikai Guo |
NeurIPS | 8 |
| 2025 | Sul-BertGRU: an ensemble deep learning method integrating information entropy-enhanced BERT and directional multi-GRU for S-sulfhydration sites predictionabstractMOTIVATION: S-sulfhydration, a crucial post-translational protein modification, is pivotal in cellular recognition, signaling processes, and the development and progression of cardiovascular and neurological disorders, so identifying S-sulfhydration sites is crucial for studies in cell biology. Deep learning shows high efficiency and accuracy in identifying protein sites compared to traditional methods that often lack sensitivity and specificity in accurately locating nonsulfhydration sites. Therefore, we employ deep learning methods to tackle the challenge of pinpointing S-sulfhydration sites. RESULTS: In this work, we introduce a deep learning approach called Sul-BertGRU, designed specifically for predicting S-sulfhydration sites in proteins, which integrates multi-directional gated recurrent unit (GRU) and BERT. First, Sul-BertGRU proposes an information entropy-enhanced BERT (IE-BERT) to preprocess protein sequences and extract initial features. Subsequently, confidence learning is employed to eliminate potential S-sulfhydration samples from the nonsulfhydration samples and select reliable negative samples. Then, considering the directional nature of the modification process, protein sequences are categorized into left, right, and full sequences centered on cysteines. We build a multi-directional GRU to enhance the extraction of directional sequence features and model the details of the enzymatic reaction involved in S-sulfhydration. Ultimately, we apply a parallel multi-head self-attention mechanism alongside a convolutional neural network to deeply analyze sequence features that might be missed at a local level. Sul-BertGRU achieves sensitivity, specificity, precision, accuracy, Matthews correlation coefficient, and area under the curve scores of 85.82%, 68.24%, 74.80%, 77.44%, 55.13%, and 77.03%, respectively. Sul-BertGRU demonstrates exceptional performance and proves to be a reliable method for predicting protein S-sulfhydration sites. AVAILABILITY AND IMPLEMENTATION: The source code and data are available at https://github.com/Severus0902/Sul-BertGRU/. Xirun Wei, Kuiyang Che, Hui Li 0014, Shikai Guo |
Bioinform. | 6 |
| 2025 | Generative imputation of incomplete images: Leveraging multimodal information for missing pixel
Qian Ma 0003, Jinlei Zhang, Shikai Guo, Bo Ning 0002, Yu Gu 0002, Ge Yu 0001 |
Inf. Sci. | 5 |
| 2025 | Structuring Semantic-Aware Relations Between Bugs and Patches for Accurate Patch EvaluationabstractABSTRACT Patches can help fix security vulnerabilities and optimize software performance, thereby enhancing the quality and security of the software. Unfortunately, patches generated by automated program repair tools are not always correct, as they may introduce new bugs or fail to fully rectify the original issue. Various methods for evaluating patch correctness have been proposed. However, most methods face the challenge of capturing long‐distance dependencies in patch correctness evaluation, which leads to a decline in the predictive performance of the models. To address the challenge, this paper presents a method named Qamhaen to evaluate the correctness of patches generated by APR. Specifically, text embedding of bugs and patches component address the challenge of long‐distance dependencies across functions in patch correctness evaluation by using bug reports and patch descriptions as inputs instead of code snippets. BERT is employed for pretraining to capture these dependencies, followed by an additional multihead self‐attention mechanism for further feature extraction. Similarity evaluator component devises a similarity calculation to assess the effectiveness of patch descriptions in resolving issues outlined in bug reports. Comprehensive experiments are conducted on a dataset containing 9135 patches and a patch correctness assessment metric, and extensive experiments demonstrate that Qamhaen outperforms baseline methods in terms of overall performance across AUC, F1, +Recall, ‐Recall, and Precision. For example, compared to the baseline, Qamhaen achieves an F1 of 0.691, representing improvements of 24.2%, 22.1%, and 6.3% over the baseline methods, respectively. Hui Li 0014, Yongqian Chen, Xiaowei Pan, Shikai Guo |
J. Softw. Evol. Process. | 5 |
| 2025 | A Novel HDL Code Generator for Effectively Testing FPGA Logic Synthesis CompilersabstractField Programmable Gate Array (FPGA) logic synthesis compilers (e.g., Vivado, Iverilog, Yosys, and Quartus) are widely applied in Electronic Design Automation (EDA), such as the development of FPGA programs. However, defects (e.g. incorrect synthesis) in logic synthesis compilers may lead to unexpected behaviors in target applications, posing security risks. Therefore, it is crucial to thoroughly test logic synthesis compilers to eliminate such defects. Despite several Hardware Design Language (HDL) code generators (e.g., Verismith) having been proposed to find defects in logic synthesis compilers, the effectiveness of these generators is still limited by the simple code generation strategy and the monogeneity of the generated HDL code. This paper proposes EvoHDL, a novel method to generate syntax-valid HDL code for comprehensively testing FPGA logic synthesis compilers. EvoHDL can generate more complex and diverse defect-triggering HDL code (e.g., Verilog, VHDL, and SystemVerilog) by leveraging the guidance of abstract syntax tree and the extensive function block libraries of cyber-physical systems. Extensive experiments show that the diversity and defect-triggering capability of HDL code generated by EvoHDL are significantly better than the state-of-the-art method (i.e., Verismith). In three months, EvoHDL has reported 20 new defects–many of which are deep and important; 16 of them have been confirmed. Shikai Guo, Guilin Zhao, Peiyu Zou, He Jiang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2025 | Dual-Network Cross-Learning for Metabolite-Disease Association PredictionabstractIn recent years, increasing evidence has demonstrated a close association between metabolites and various complex human diseases, providing valuable insights for disease diagnosis, treatment, and prevention. Although deep learning-based approaches have achieved certain success in predicting metabolic disease associations, challenges remain in enriching graph information and effectively integrating metabolic and disease features. To address these issues, this paper proposes a model named DCMDA, which extracts deep features of both metabolites and diseases using Dual-network Cross-learning for Metabolite-Disease Association prediction. DCMDA consists of three parts. The data processing module integrates similarity networks with association networks to construct a heterogeneous network. The feature extraction module extracts features from the metabolite-disease association network based on the non-negative matrix factorization method and from the heterogeneous network using graph autoencoder techniques. The feature fusion module combines the association matrix feature with the heterogeneous network feature through a Cross-Attention mechanism, thereby obtaining deep representations of metabolites and diseases. These features are then used to train the model to predict association scores between metabolites and diseases. Experimental results demonstrate that in 5-fold cross-validation, DCMDA achieves an area under the receiver operating characteristic curve (AUC) of 97.8% and an area under the precision-recall curve (AUPR) of 97.9%, outperforming state-of-the-art prediction methods. Yanxin Chen, Hui Li 0014, Shikai Guo |
IEEE Trans. Comput. Biol. Bioinform. | 5 |
| 2025 | MHMDA: "Similarity-Association-Similarity" Metapaths and Heterogeneous-Hyper Network Learning for MiRNA-Disease Association PredictionabstractIn recent years, microRNA (miRNA) has been recognized as crucial in the progression of human diseases. However, existing computational methods for identifying miRNA-disease associations often overlook the rich association information contained in specific long-distance pathways and lack effective exploration of potential associations. In this study, we propose a biologically interpretable "similarity-association-similarity" metapath and heterogeneous-hyper network (HeteroHyperNet) learning approach for miRNA-disease association prediction (MHMDA). In MHMDA, a "similarity-association-similarity" multi-hop metapaths learning method based on hierarchical attention perception is proposed to explore specific long-distance associated pathway information connecting potentially associated miRNAs and diseases. In addition, a HeteroHyperNet learning approach integrating heterogeneous network and hyper network is designed to progressively learn direct association information and potential association information between miRNA and disease. The "similarity-association-similarity" metapath with hierarchical attention significantly enhances the learning of long-distance biological associations, while the HeteroHyperNet comprehensively learns the known and potential associations of miRNA-disease, greatly improving the richness and accuracy of information. A large number of experimental results show that MHMDA has demonstrated excellent performance in the prediction of miRNA-disease association. In addition, cross independent dataset experiment and cold start experiment on miRNA and disease prove the effectiveness of MHMDA on sparse association points, and its stability and reliability in predicting potential miRNA-disease association are further confirmed. Yaomiao Zhao, Shikai Guo, Hui Li 0014 |
IEEE Trans. Comput. Biol. Bioinform. | 4 |
| 2025 | Prediction of miRNA-Disease Association Based on Biodiversity Association NetworkabstractMiRNA-disease association identification is of great significance to the development of clinical medicine and drug research. Present computational methods didn't consider rich biological information, such as the expression level changes of disease-related miRNAs and the association information between miRNA, disease and other types of biological entities. In this study, we propose a new method for prediction of MiRNA-Disease Association based on Biodiversity Association Network (BANMDA). BANMDA first collects multiple types of association information from multiple sources, including diseases, miRNAs and lncRNAs associations, expression level changes of disease-related miRNAs, miRNAs sequence information, and disease semantic information. Second, BANMDA extracts diversity association features and diversity biological features based on two heterogeneous graph structure to represent miRNAs and diseases at multiple levels. In diversity association module, edges are classified according to the expression level changes of disease-related miRNAs and the similarities between miRNAs and diseases. In diversity node module, lncRNAs associated with miRNA-diseases are collected to construct heterogeneous network and we propose an improved GCN algorithm to directly aggregate the higher-order neighborhood information in the heterogeneous graph. Finally, the bilinear decoder is applied to predict associations between miRNAs and diseases. Experimental results show that BANMDA can be used as a powerful tool to identify miRNA-disease associations. Chaorui Guo, Hui Li 0014, Shikai Guo |
IEEE Trans. Comput. Biol. Bioinform. | 6 |
| 2025 | SIMTAM: Generation Diversity Test Programs for FPGA Simulation Tools Testing Via Timing Area MutationabstractField-Programmable Gate Array (FPGA) timing simulation is essential in electronic circuit design, allowing for the verification of timing characteristics like delays and clock frequencies. However, bugs in timing simulation tools can lead to inaccurate results, potentially causing designers to miss critical issues in chip performance. Traditional testing methods often fall short in thoroughly assessing these tools, as current FPGA testing primarily focuses on synthesis and behavioral simulation, neglecting timing aspects. To address this issue, we propose SIMTAM for testing timing simulation tools. Specifically, SIMTAM consists of three components: equivalent delay region construction, diversity program segment generation, and differential testing. Given a seed circuit design file written by hardware description language such as Verilog, the delay region construction component randomly identifies delay structures for inertial delay in the design file to construct equivalent delay sleep regions. In the sleep region, the simulator skips the signal pulse whose width is less than the specified delay, thus ensuring the equivalence of the variations. The diversity program segment generation component combines Verilog expressions using generation operators and injects them into the sleep region to generate diverse design files. The differential testing component compares the seed and variant design files to find compilation inconsistency issues. In 5 months, SIMTAM reported 16 bugs to developers in two popular timing simulation tools, Iverilog and Vivado, 10 of which are confirmed. Shikai Guo, Zun Wang 0005, He Jiang 0001 |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2025 | Making Fault Localization in Online Service Systems More Actionable and InterpretableabstractOnline service systems struggle with accurately and quickly pinpointing and resolving failures within their intricate systems, and it therefore emerges the solutions for fault localization in the code. However, the previous fault localization models suffer from low localization accuracy and poor interpretability due to the complex dependencies among fault characteristics in industrial practice. To address this issue, challenges brought by the long-distance dependencies among fault features and the unbalanced distribution of fault knowledge, and to improve the interpretability of the model, we present a fault localization model in online service systems more actionable and interpretable, named FL-AIer. Specifically, FL-AIer consists of two components: the feature encoding component and the fault localization component. The feature encoding component utilizes graph attention networks to capture the complex spatio-temporal dependencies within fault features. Then, the fault localization component adopts a three-stage approach, leveraging a multi-attention mechanism to identify and prioritize the most relevant fault features for precise localization. Additionally, the Fault Knowledge Balancing module it contains introduces a weighted Kullback-Leibler divergence loss function to ensure that the model pays adequate attention to all fault features, addressing the issue of imbalanced fault knowledge distribution and enhancing localization performance. We conducted extensive experiments on four datasets, and the results demonstrated that FL-AIer effectively addressing the challenges of fault localization in online system environments, and consistently outperforms the state-of-the-art methods across various evaluation metrics such as A@1, A@2, A@3, A@5, and MAR. For instance, FL-AIer achieves significant improvements of 5.82%, 10.77%, 4.20%, and 15.56% on the A@1 metric, respectively. These results fully demonstrate the excellent effectiveness of FL-AIer in effectively addressing the challenges of fault localization in online system environments, surpassing the performance of existing state-of-the-art methods. Ke Xv, Shikai Guo, Hui Li 0014, Rong Chen 0003, He Jiang 0001 |
ACM Trans. Softw. Eng. Methodol. | 2 |
| 2025 | Context-based Transfer Learning for Structuring Fault Localization and Program Repair AutomationabstractAutomated software debugging plays a crucial role in aiding software developers to swiftly identify and attempt to rectify faults, thereby significantly reducing developers’ workload. Previous researches have predominantly relied on simplistic semantic deep learning or statistical analysis methods to locate faulty statements in diverse projects. However, code repositories often consist of lengthy sequences with long-distance dependencies, posing challenges for accurately modeling fault localization using these methods. In addition, the lack of joint reasoning among various faults prevents existing models from deeply capturing fault information. To address these challenges, we propose a method named CodeHealer to achieve accurate fault localization and program repair. CodeHealer comprises three components: a Deep Semantic Information Extraction Component that effectively extracts deep semantic features from suspicious code statements using classifiers based on Joint-attention mechanisms; a Suspicious Statement Ranking Component that combines various fault localization features and employs multilayer perceptrons to derive multidimensional vectors of suspicion values; and a Fault Repair Component that, based on ranked suspicious statements generated by fault localization, adopts a top-down approach using multiple classifiers based on Co-teaching mechanisms to select repair templates and generate patches. The experimental results indicate that when applied to fault localization, CodeHealer outperforms the best baseline method with improvements of 11.4%, 2.7%, and 1.6% on Top-1/3/5 metrics, respectively. It also reduces the MFR and MAR by 9.8% and 2.1%, where lower values denote better fault localization effectiveness. Additionally, in automated software debugging, CodeHealer fixes an additional 6 faults compared to the current best method, totaling 53 faults repaired. Lehuan Zhang, Shikai Guo, Hui Li 0014, Yu Chai, Rong Chen 0003, He Jiang 0001 |
ACM Trans. Softw. Eng. Methodol. | 2 |
| 2025 | Toward Understanding FPGA Synthesis Tool BugsabstractFPGA (Field Programmable Gate Array) synthesis tools are crucial for hardware development and AI acceleration, and their bugs could compromise hardware reliability and risk downstream applications. However, it remains unknown in understanding the characteristics of these bugs. What are the root causes that trigger bugs in FPGA synthesis tools? What are the characteristics of these bugs? What are the challenges in detecting and addressing them? This paper takes the first step towards answering these questions by conducting a comprehensive study of FPGA synthesis tool bugs. We analyze 551 confirmed bugs in both commercial and open source FPGA synthesis tools, i.e., Vivado, Quartus Prime, and Yosys, covering root causes, symptoms, bug-prone components, fix characteristics, and achieve 17 valuable findings. We find that, on average, around 46.2% of bugs result from HDL (Hardware Description Language) standard noncompliance across the three tools. However, it is hard for current formal validations to fully test HDL standards compliance. Additionally, on average over 25.8% bugs show domain-specific optimization traits due to inappropriate optimization and mapping. Meanwhile, beyond 28% of bugs trigger unexpected behavior without clear signs, making the formulation of effective test oracles challenging. These findings help addressing FPGA synthesis tool bugs and guide further research. Yi Zhang 0148, He Jiang 0001, Shikai Guo, Hui Liu 0003, Chongyang Shi 0001 |
ACM Trans. Softw. Eng. Methodol. | 3 |
| 2025 | What Causes Bugs in Numerical Simulation Software? An Empirical StudyabstractNumerical simulation (NS) software is widely used in safety-critical domains (e.g., aerospace design) to simulate actual physical processes of real-world entities on computers. However, NS software is error-prone, whose bugs lead to incorrect simulation, and may even cause disastrous flaws in safety-critical applications. Although many studies investigate the bug characteristics of computation-centered software, such as machine learning systems, the characteristics of NS software bugs have not been fully studied: what are the root causes and symptoms; how are they different from other computation-centered software; and why the difference occurs. To bridge this gap, we present a systematic study of NS software bugs by analyzing 352 bugs in three popular NS projects (i.e., FDS, SU2, and Kratos) for different domains. We summarize seven root causes (with 18 subcategories) and five symptoms. We find that the correctness, completeness, compatibility, and parallelization to implement NS algorithms (i.e., models) are error-prone. Many root causes (e.g., incorrect model and incorrect initialization) require physical and chemical knowledge to avoid bugs, which may not be mastered by typical software developers. These findings motivate new challenges and opportunities for future NS software development, such as designing domain specific language systems for NS model. Youcheng Zhu, Shikai Guo, He Jiang 0001 |
IEEE Trans. Reliab. | 3 |
| 2025 | BR-Hunter: Detect Information Types of Bug Reports From Online Community DiscussionsabstractIn community-based software development, live-chatting services are increasingly used to discuss bugs encountered during development. Many methods have emerged to identify bugs and produce bug reports, which further improve the efficiency of software development. However, previous methods still face challenges in understanding complex conversational structures and classifying sentences in bug reports, as entertaining or meaningless utterances often lower the quality of constructed bug reports. To address this issue, we propose a method named BR-Hunter, which comprises the following four components. Specifically, the data preprocessing component disentangles and denoises the live chats, while the utterance embedding component aims to extract the semantic features of each utterance in the conversations. The bug report identification component then models the conversation as a feature graph and uses Graph Neural Networks to identify conversations containing bug reports, thereby solving Challenge 1. Finally, the bug report synthesis (BRS) component tackles Challenge 2 by classifying and reassembling sentences from conversations containing bug reports, leveraging fine-tuned BERT and prompt learning techniques. Extensive experiments conducted on eight open source projects demonstrate that BR-Hunter achieves high accuracy in identifying bug reports. Compared to baseline methods, BR-Hunter improves the average F1 score by 36.41%, 24.80%, 68.92%, 46.77%, 52.84%, 25.80%, 25.25%, and 4.19%, respectively. And BR-Hunter also achieves an average improvement of 10.34% on the BRS task, compared with the state-of-the-art method. Huijiang Liu, Junyu Xiong, Shikai Guo, Hui Li 0014 |
IEEE Trans. Reliab. | 3 |
| 2025 | Extracting Meaningful Issue-Solution Pair From Collaborative Developer Live ChatsabstractThe live chats of developers often contain meaningful information in the form of issue–solution pairs. The issue–solution pairs can offer helpful references to others who seek solutions for the similar issues, which can improve software development efficiency by facilitating issue solving. However, previous approaches such as ISPY still struggle with unsatisfactory extraction accuracy, due to the entanglement and complexity of issue-solution pairs' feature information. To address these challenges, we propose an approach namedIS-Hunterfor mining issue-solution pairs from real-time chat data. Specifically,IS-Hunterconsists of four main components: the data preprocessing component disentangles and denoises raw chat logs, the utterance embedding component embeds utterances into vectors that subsequent components can easily process, the feature extraction component obtains textual, heuristic, and contextual feature that determines whether an utterance is topic-relevant, and the issue–solution pair prediction component predicts the utterance whether is an issue or a solution. The experimental results show that the performance of IS-Hunter outperforms the baseline methods in issue-detection and solution-extraction in terms of Precision, Recall, and F1-score. Compared with baseline methods, in issue-detection, IS-Hunter, respectively, achieves an average precision, recall, and F1-score of 0.74, 0.74, and 0.74, and it marks an obvious 4.23% improvement over the state-of-the-art approaches. Simultaneously, in solution-extraction, IS-Hunter achieves an average precision, recall, and F1-score of 0.83, 0.90, and 0.86 which is 4.88% higher than the best baseline methods. Jiawen Shen, Shikai Guo, Longfeng Chen, Hui Li 0014 |
IEEE Trans. Reliab. | 2 |
| 2025 | Insights From Bugs in FPGA High-Level Synthesis Tools: An Empirical Study of Bambu BugsabstractHigh-level synthesis (HLS) tools have been widely used in field-programmable gate array (FPGA) design to convert C/C++ code to hardware description language code. Unfortunately, HLS tools are susceptible to bugs, which can introduce serious vulnerabilities in FPGA products, leading to substantial losses. However, the characteristics of these bugs (e.g., root causes and bug-prone stages) have never been systematically studied, which significantly hinders developers from effectively handling HLS tool bugs. To this end, we conduct the first empirical study to uncover HLS tool bug characteristics. We collect 349 bugs of a widely used HLS tool, namely Bambu. We study the root causes, buggy stages, and bug fixes of these bugs by applying a multiperson collaboration method. Finally, 13 valuable findings are summarized. We find 14 categories of root causes in Bambu bugs; most bugs (22.1%) are caused by incorrect implementation of IR processing; the front end of Bambu is more bug-prone; to fix these bugs, 2.27 files and 80.19 lines of code need to be modified on average. We also present the insights gained from 95 Vitis HLS bugs. From these findings, we suggest that developers could use an on-the-fly code generator configuration method to generate suitable testing programs for HLS tool bug detection and apply large language models to assist in fixing HLS tool bugs. Zun Wang 0005, He Jiang 0001, Shikai Guo, Yi Zhang 0148 |
IEEE Trans. Reliab. | 4 |
| 2025 | Learning to Accelerate Autonomous Driving System TestingabstractSystem bug identification plays a key role in autonomous vehicles for avoiding disastrous consequences. However, it is time consuming to fully test autonomous vehicles. Although some techniques have been proposed to improve testing efficiency, they struggle to handle the complex test scenarios because they cannot adequately represent the scenarios and efficiently assess the vulnerability-triggering potential of test scenarios. In this study, we propose Learning to accElerate Autonomous Systems Testing (LEAST), a graphic neural network (GNN)-based method to effectively accelerate autonomous driving system (ADS) testing. First, given a test scenario, LEAST extracts a series of scenes from the test scenario and constructs feature graphs from these scenes, which are helpful to characterize the test scenario. Then, we propose a GNN-based method to predict the risk value of each scene, indicating the likelihood of a vehicle encountering a traffic accident. Finally, LEAST prioritizes test scenarios based on the number of high-risk scenes in the scenarios, ensuring that scenarios more likely to trigger ADS bugs are executed earlier. Experiments onthree open-source ADSs show that LEAST significantly outperforms baseline test acceleration approaches by 10.13–15.34% in terms of average percentage of fault detected. When integrating LEAST into the advanced ADS testing approach DriveFuzz, LEAST successfully improves the testing efficiency by 39–67%. Zhide Zhou, Shikai Guo, He Jiang 0001 |
IEEE Trans. Reliab. | 4 |
| 2025 | PCBSmith: An Effective Schematic Generator for Testing PCB Design Tool ChainabstractIn electronic design automation (EDA), printed circuit board (PCB) design plays a crucial role. Ensuring the reliability of the PCB design tool chain is essential, as bugs in the tool chain can cause significant issues and losses during design and production. To improve reliability, a key process is to generate numerous PCB schematics and execute them in the tool chain, to test the correctness of each tool chain functionality. However, it is a challenge to automatically generate valid schematics to simulate the actual use of the PCB design tool chain. To this end, we propose PCBSmith, an effective schematic generator for PCB design tool chain. PCBSmith mimics the steps of a PCB designer for schematic design. PCBSmith first selects the appropriate electronic components from a comprehensive library and connects them according to the constraints of different components. PCBSmith then sets electrical parameters and simulation models for each component, eventually generating simulatable schematics. Experiments show that PCBSmith demonstrates high efficiency in schematic generation, averaging only one schematic per second. PCBSmith maintains a success rate over 61.44% for generating schematics, which outperforms the baseline method by 30.68%. The generated schematics have successfully identified unknown bugs in PCB design tools. He Jiang 0001, Shikai Guo, Zhilei Ren, Peiyu Zou, Huijiang Liu |
IEEE Trans. Reliab. | 4 |
| 2025 | Line-Level Defect Prediction by Capturing Code Contexts With Graph Convolutional NetworksabstractSoftware defect prediction refers to the systematic analysis and review of software using various approaches and tools to identify potential defects or errors. Software defect prediction aids developers in swiftly identifying defects and optimizing development resource allocation, thus enhancing software quality and reliability. Previous defect prediction approaches still face two main limitations: 1) lacking of contextual semantic information and 2) Ignoring the joint reasoning between different granularities of defect predictions. In response to these challenges, we propose LineDef, a line-level defect prediction approach by capturing code contexts with graph convolutional networks. Specifically, LineDef comprises three components: the token embedding component, the graph extraction component, and the multi-granularity defect prediction component. The token embedding component maps each token to a vector to obtain a high-dimensional semantic feature representation of the token. Subsequently, the graph extraction component utilizes a sliding window to extract line-level and token-level graphs, addressing the challenge of capturing contextual semantic relationships in the code. Finally, the multi-granularity defect prediction component leverages graph convolutional layers and attention mechanisms to acquire prediction labels and risk scores, thereby achieving file-level and line-level defect prediction. Experimental studies on 32 datasets across 9 different software projects show that LineDef exhibits significantly enhanced balanced accuracy, ranging from 15.61% to 45.20%, compared to state-of-the-art file-level defect prediction approaches, and a remarkable cost-effectiveness improvement ranging from 15.32% to 278%, compared to state-of-the-art line-level defect prediction approaches. These results demonstrate that LineDef approach can extract more comprehensive information from lines of code for defect prediction. Shouyu Yin, Shikai Guo, Hui Li 0014, Rong Chen 0003, He Jiang 0001 |
IEEE Trans. Software Eng. | 2 |
| 2024 | VF-Detector: Making Multi-Granularity Code Changes on Vulnerability Fix Detector Robust to Mislabeled Changes
Zhenkan Fu, Shikai Guo, Hui Li 0014, Rong Chen 0003, He Jiang 0001 |
IJCAI | 2 |
| 2024 | Deep Just-In-Time Defect Prediction Based on Double-Source Input Self-Attention Mechanism (S)abstractEnsuring high-quality software products is an important issue for software productions, it emerged Just-In-Time Quality Assurance to automatically identify potentially defective code as early as possible in recent years.The presented models mainly utilize Convolutional Neural Networks for automatic defect feature extraction.However, these models ignore the utilization of contextual information and are not suitable for large-scale projects.To address these issues, we propose a model named Multi-head Convolution Structure of Attention-JIT (short for MCSA-JIT), which comprises four layers.Specifically, Commit Message-Code Change(short for CM-CC) Text Feature Extraction Layer leverages a variant of the self-attention mechanism to extract features from both commit messages and code changes, CM-CC Structure Feature Extraction Layer captures the content information of the submissions and different CNN structures are constructed to extract structural information, Feature Combination Layer combines the text feature and structure feature to predict defects and output layer outputs the final value.Experimental results conducted on two software projects, QT and OPENSTACK, demonstrate that the best variant of MCSA-JIT achieves a relative improvement of 7.46% in terms of AUC on the OPENSTACK project and a relative improvement of 8.85% on the QT project when compared to state-of-the-art methods with the best performance. Weixiang Hong 0002, Hui Li 0014, Shikai Guo |
SEKE | 4 |
| 2024 | Structuring Meaningful Code Review Automation in Developer Community
Zhenzhen Cao, Sijia Lv, Hui Li 0014, Qian Ma 0003, Cheng Guo 0001, Shikai Guo |
Eng. Appl. Artif. Intell. | 8 |
| 2024 | Graph Confident Learning for Software Vulnerability Detection
Qian Wang 0034, Zhengdao Li, Hetong Liang, Xiaowei Pan, Hui Li 0014, Shikai Guo |
Eng. Appl. Artif. Intell. | 9 |
| 2024 | Detect software vulnerabilities with weight biases via graph neural networks
Huijiang Liu, Shuirou Jiang, Xuexin Qi, Hui Li 0014, Cheng Guo 0001, Shikai Guo |
Expert Syst. Appl. | 8 |
| 2024 | Automated patch correctness predicting to fix software defect
Zelong Zheng, Zijian Tao, Hui Li 0014, Shikai Guo |
Expert Syst. Appl. | 7 |
| 2024 | Context-based transfer learning for low resource code summarizationabstractAbstract Source code summaries improve the readability and intelligibility of code, help developers understand programs, and improve the efficiency of software maintenance and upgrade processes. Unfortunately, these code comments are often mismatched, missing, or outdated in software projects, resulting in developers needing to infer functionality from source code, affecting the efficiency of software maintenance and evolution. Various methods based on neuronal networks are proposed to solve the problem of synthesis of source code. However, the current work is being carried out on resource‐rich programming languages such as Java and Python, and some low‐resource languages may not perform well. In order to solve the above challenges, we propose a context‐based transfer learning model for low resource code summarization (LRCS), which learns the common information from the language with rich resources, and then transfers it to the target language model for further learning. It consists of two components: the summary generation component is used to learn the syntactic and semantic information of the code, and the learning transfer component is used to improve the generalization ability of the model in the learning process of cross‐language code summarization. Experimental results show that LRCS outperforms baseline methods in code summarization in terms of sentence‐level BLEU, corpus‐level BLEU and METEOR. For example, LRCS improves corpus‐level BLEU scores by 52.90%, 41.10%, and 14.97%, respectively, compared to baseline methods. Yu Chai, Lehuan Zhang, Hui Li 0014, Mengzhi Luo, Shikai Guo |
Softw. Pract. Exp. | 6 |
| 2024 | Estimating Uncertainty in Labeled Changes by SZZ Tools on Just-In-Time Defect PredictionabstractThe aim of Just-In-Time (JIT) defect prediction is to predict software changes that are prone to defects in a project in a timely manner, thereby improving the efficiency of software development and ensuring software quality. Identifying changes that introduce bugs is a critical task in just-in-time defect prediction, and researchers have introduced the SZZ approach and its variants to label these changes. However, it has been shown that different SZZ algorithms introduce noise to the dataset to a certain extent, which may reduce the predictive performance of the model. To address this limitation, we propose the Confident Learning Imbalance (CLI) model. The model identifies and excludes samples whose labels may be corrupted by estimating the joint distribution of noisy labels and true labels, and mitigates the impact of noisy data on the performance of the prediction model. The CLI consists of two components: identifying noisy data (Confident Learning Component) and generating a predicted probability matrix for imbalanced data (Imbalanced Data Probabilistic Prediction Component). The IDPP component generates precise predicted probabilities for each instance in the training set, while the CL component uses the generated predicted probability matrix and noise labels to clean up the noise and build a classification model. We evaluate the performance of our model through extensive experiments on a total of 126,526 changes from ten Apache open source projects, and the results show that our model outperforms the baseline methods. Shikai Guo, Sijia Lv, Rong Chen 0003, Hui Li 0014, He Jiang 0001 |
ACM Trans. Softw. Eng. Methodol. | 1 |
| 2024 | Analyzing and Detecting Information Types of Developer Live Chat ThreadsabstractOnline chatrooms serve as vital platforms for information exchange among software developers. With multiple developers engaged in rapid communication and diverse conversation topics, the resulting chat messages often manifest complexity and lack structure. To enhance the efficiency of extracting information from chat threads , automatic mining techniques are introduced for thread classification. However, previous approaches still grapple with unsatisfactory classification accuracy due to two primary challenges that they struggle to adequately capture long-distance dependencies within chat threads and address the issue of category imbalance in labeled datasets. To surmount these challenges, we present a topic classification approach for chat information types named EAEChat. Specifically, EAEChat comprises three core components: the text feature encoding component captures contextual text features using a multi-head self-attention mechanism-based text feature encoder, and a siamese network is employed to mitigate overfitting caused by limited data; the data augmentation component expands a small number of categories in the training dataset using a technique tailored to developer chat messages, effectively tackling the challenge of imbalanced category distribution; the non-text feature encoding component employs a feature fusion model to integrate deep text features with manually extracted non-text features. Evaluation across three real-world projects demonstrates that EAEChat, respectively, achieves an average precision, recall, and F1-score of 0.653, 0.651, and 0.644, and it marks a significant 7.60% improvement over the state-of-the-art approaches. These findings confirm the effectiveness of our method in proficiently classifying developer chat messages in online chatrooms. Xiuwei Shang, Shikai Guo, Yulong Li 0001, Rong Chen 0003, Hui Li 0014, He Jiang 0001 |
ACM Trans. Softw. Eng. Methodol. | 4 |
| 2024 | Simulink Compiler Testing via Configuration Diversification With Reinforcement LearningabstractSimulink compiler testing is important since all cyber-physical system (CPS) models are required to be compiled by Simulink compiler. Current testing processes use CPS models generated by CPS model generators for testing. Since the effectiveness of CPS model generators heavily relies on suitable generator configurations, existing approaches randomize configurations or infer configurations with historical bug information to generate diverse bug-triggering CPS models. However, these approaches are designed for general-purpose compilers (e.g., GCC), which have two challenges when testing Simulink compiler, namely, the CPS model representation challenge on representing CPS models for diversity measurement and the configuration learning challenge on learning configurations to generate diverse CPS models. To address these challenges, we proposeReinforcement lEarning-basedCOnfiguRationDiversification (RECORD), a new configuration diversification approach. RECORD has a feature vectorization component, which addresses the first challenge by representing CPS models as feature vectors to capture the local and global characteristics of CPS models for diversity measurement. RECORD then uses a reinforcement learning component to generate diverse CPS models based on the learned relationship between configuration updates and diversity changes, thus addressing the second challenge. Experiments demonstrate that within three months, RECORD reported 11 confirmed Simulink compiler bugs, significantly outperforming the state-of-the-art configuration diversification approaches. RECORD can also facilitate different testing strategies to find more bugs. Shikai Guo, Hongyi Cheng, He Jiang 0001 |
IEEE Trans. Reliab. | 2 |
| 2024 | What's Wrong With Low-Code Development Platforms? An Empirical Study of Low-Code Development Platform BugsabstractLow-code development platforms (LCDPs) are increasingly being introduced and leveraged by major IT enterprises to lower the threshold and promote the efficiency of software development. Like other software systems, LCDPs are also inevitable to have bugs. The bugs in LCDPs may cause unpredictable consequences as they pose risks to all the downstream software products. However, to the best of our knowledge, there exist no studies that ever consider the bugs caused by LCDPs. To handle the LCDP bugs better, in this article, we conduct an empirical study of the characteristics of LCDP bugs by examining 974 confirmed bugs of four dominant LCDPs (i.e., OutSystems, Mendix, Appsmith, and Budibase) from both commercial and open-source domains. These bugs are analyzed from three perspectives, including bug root causes, bug symptoms, and the affected stages of LCDPs. Based on the analysis, we obtain a series of valuable findings. For example, around 60% of the bugs reside in the stage of designing and specifying the developed applications. Over 37% of the bugs lead LCDPs to behave unexpectedly but without showing explicit signs. Moreover, the bugs relevant to the incorrect graphics of user interfaces are significant due to the characteristics of LCDPs. These findings point out the guidelines, challenges, and future directions to address LCDP bugs. Dong Liu 0025, He Jiang 0001, Shikai Guo, Yuting Chen 0001, Lei Qiao 0002 |
IEEE Trans. Reliab. | 3 |
| 2024 | A Testing Program and Pragma Combination Selection Based Framework for High-Level Synthesis Tool Pragma-Related Bug DetectionabstractHigh-Level Synthesis (HLS) tools convert C/C++ design code into Hardware Description Language (HDL) code automatically, which are often used for Field Programmable Gate Array (FPGA) design. HLS tools provide many pragmas, which are a kind of directive to be inserted into C/C++ code, for designers to efficiently control the synthesis of code components (e.g., arrays and loops) to generate FPGA implementations with varying performances and costs. However, the use of some pragmas may trigger HLS tool bugs (e.g., tool crashes). Although many formal methods have been proposed to verify the correctness of various HLS phases, no relevant work addresses the problem on detecting HLS tool pragma-related bugs. To resolve this problem, two challenges need to be addressed, namely the selection of testing programs and the acquisition of pragma combinations, due to the enormous number of testing programs and pragma combinations. In this paper, we propose TEPACS, a TEsting Program and prAgma Combination Selection-based framework, to construct diverse testing programs with pragmas for effectively detecting HLS tool pragma-related bugs. TEPACS follows the idea of fuzzing, which is a widely used technique in software testing. First, TEPACS selects the representative testing program according to the cosine distance between the code component vectors of testing programs. Then, for a selected program, TEPACS generates its golden output and uses the pragma combination selection method based on combinatorial testing to generate a set of programs with different pragmas. TEPACS uses the HLS tool under test to convert these testing programs into HDL codes and obtains the simulation results of the HDL code. Finally, based on differential testing, TEPACS identifies HLS tool bugs triggered if the simulation result and golden output are inconsistent. We evaluate TEPACS and its five variants on Vitis HLS, a widely used FPGA HLS tool. Experimental results show that TEPACS outperforms the baselines by at least 11.17% in terms of the bug-finding capability. In one month, TEPACS detected 34 bugs on the latest version of Vitis HLS, of which 9 bugs have been confirmed. He Jiang 0001, Zun Wang 0005, Zhide Zhou, Shikai Guo, Weifeng Sun 0002, Tao Zhang 0001 |
IEEE Trans. Software Eng. | 5 |
| 2024 | Code Comment Inconsistency Detection Based on Confidence LearningabstractCode comments are a crucial source of software documentation that captures various aspects of the code. Such comments play a vital role in understanding the source code and facilitating communication between developers. However, with the iterative release of software, software projects become larger and more complex, leading to a corresponding increase in issues such as mismatched, incomplete, or outdated code comments. These inconsistencies in code comments can misguide developers and result in potential bugs, and there has been a steady rise in reports of such inconsistencies over time. Despite numerous methods being proposed for detecting code comment inconsistencies, their learning effect remains limited due to a lack of consideration for issues such as characterization noise and labeling errors in datasets. To overcome these limitations, we propose a novel approach called MCCL that first removes noise from the dataset and then detects inconsistent code comments in a timely manner, thereby enhancing the model's learning ability. Our proposed model facilitates better matching between code and comments, leading to improved development of software engineering projects. MCCL comprises two components, namely method comment detection and confidence learning denoising. The method comment detection component captures the intricate relationships between code and comments by learning their syntactic and semantic structures. It correlates the code and comments through an attention mechanism to identify how changes in the code affect the comments. Furthermore, confidence learning denoising component of MCCL identifies and removes characterization noises and labeling errors to enhance the quality of the datasets. This is achieved by implementing principles such as pruning noisy data, counting with probabilistic thresholds to estimate noise, and ranking examples to train with confidence. By effectively eliminating noise from the dataset, our model is able to more accurately learn inconsistencies between comments and source code. Our experiments on 1,518 open-source projects demonstrate that MCCL can accurately detect inconsistencies, achieving an averageF1-scoreof 82.6%. This result outperforms state-of-the-art methods by 2.4% to 28.0%. Therefore, MCCL is more effective in identifying inconsistent comments based on code changes compared to existing approaches. Zhengkang Xu, Shikai Guo, Rong Chen 0003, Hui Li 0014, He Jiang 0001 |
IEEE Trans. Software Eng. | 2 |
| 2023 | Partition Based Differential Testing for Finding Embedded Code Generation Bugs in SimulinkabstractEngineers frequently generate embedded code from Simulink models for control applications. However, target applications using the code could behave unexpectedly, due to the bugs in code generation. In this study, we propose MOPART, the first model partition based differential testing method for code generation testing in Simulink. MOPART uses multiple-way network partitioning to generate diverse bug-triggering Simulink models to thoroughly exercise the code generation process. MOPART then finds bugs by analyzing the outputs of these Simulink models with differential testing. Experiments show that MOPART significantly outperforms existing approaches, which finds 11 confirmed code generation bugs in only two weeks. He Jiang 0001, Hongyi Cheng, Shikai Guo |
DAC | 3 |
| 2023 | Constructing meaningful code changes via graph transformerabstractAbstract The rapid development of Open‐Source Software (OSS) has resulted in a significant demand for code changes to maintain OSS. Symptoms of poor design and implementation choices in code changes often occur, thus heavily hindering code reviewers to verify correctness and soundness of code changes. Researchers have investigated how to learn meaningful code changes to assist developers in anticipating changes that code reviewers may suggest for the submitted code. However, there are two main limitations to be addressed, including the limitation of long‐range dependencies of the source code and the missing syntactic structural information of the source code. To solve these limitations, a novel method is proposed, named Graph Transformer for learning meaningful Code Transformations (GTCT), to provide developers with preliminary and quick feedback when developers submit code changes, which can improve the quality of code changes and improve the efficiency of code review. GTCT comprises two components: code graph embedding and code transformation learning. To address the missing syntactic structural information of the source code limitation, the code graph embedding component captures the types and patterns of code changes by encoding the source code into a code graph structure from the lexical and syntactic representations of the source code. Subsequently, the code transformation learning component uses the multi‐head attention mechanism and positional encoding mechanism to address the long‐range dependencies limitation. Extensive experiments are conducted to evaluate the performance of GTCT by both quantitative and qualitative analyses. For the quantitative analysis, GTCT relatively outperforms the baseline on six datasets by 210%, 342.86%, 135%, 29.41%, 109.09%, and 91.67% in terms of perfect prediction. Meanwhile, the qualitative analysis shows that each type of code change by GTCT outperforms that of the baseline method in terms of bug fixed, refactoring code and others' taxonomy of code changes. Shikai Guo, Mengxuan Li 0005, Hui Li 0014, Rong Chen 0003 |
IET Softw. | 1 |
| 2023 | Structuring meaningful bug-fixing patches to fix software defectabstractAbstract Currently, software projects require a significant amount of time, effort and other resources to be invested in software testing to reduce the number of code defects. However, this process decreases the efficiency of software development and leads to a significant waste of workforce and resources. To address this challenge, researchers developed various solutions utilising deep neural networks. However, these solutions are frequently challenged by issues, such as a vast vocabulary, network training difficulties and elongated training processes resulting from the handling of redundant information. To overcome these limitations, the authors proposed a new neural network‐based model named HopFix, designed to detect software defects that may be introduced during the coding process. HopFix consists of four parts: data preprocessing, encoder, decoder and code generation components, which were used for preprocessing data, extracting information about software defects, analysing defect information, generating software patches and controlling the generation process of software patches, respectively. Experimental studies on Bug‐Fix Pairs (BFP) show that HopFix correctly fixed 47.2% ( BFP small datasets) and 25.7% ( BFP medium datasets) of software defects. Hui Li 0014, Xuexin Qi, Shikai Guo |
IET Softw. | 5 |
| 2023 | An data augmentation method for source code summarization
Zixuan Song, Xiuwei Shang, Guanxi Li, Hui Li 0014, Shikai Guo |
Neurocomputing | 6 |
| 2023 | Multi-Feature Fusion Based Structural Deep Neural Network for Predicting Answer Time on Stack Overflow
Shikai Guo, Hui Li 0014, Yu-Long Fan |
J. Comput. Sci. Technol. | 1 |
| 2023 | Optimization of Web Service Testing Task Assignment in Crowdtesting Environment
Wen-Jun Tang, Rong Chen 0003, Sheng-Jie Zheng, Shikai Guo |
J. Comput. Sci. Technol. | 6 |
| 2023 | Code samples summarization for knowledge exchange in developer communityabstractAbstract A question title's function is to generate readable titles and describe a problem encountered by the code. Previous studies often used an end‐to‐end sequence‐to‐sequence system to generate question title's from source code. However, long‐term dependencies are often difficult to capture, and this may result in an incomplete source code representation. To address this issue, we propose a Transformer for Generating Code Title (hereinafter referred to as TGCT) model. Specifically, the TGCT model uses the position coding mechanism to model paired relationships between source terms by applying relative position representations. Multiple self‐attention mechanism components are also used to capture long‐term dependencies of the code. Comprehensive experiments on datasets from five coding languages, namely Python, Java, JavaScript, C#, and SQL, are conducted, and the results show that TGCT outperforms state‐of‐the‐art models based on the measurements of BLEU and ROUGE in general. In addition, a cross‐sectional comparison experiment was conducted to verify the effects of different model parameters, different data set sizes, position coding mechanism, and self‐attention mechanism on model results. Shikai Guo, Zhongyan Liu, Zixuan Song, Hui Li 0014, Rong Chen 0003 |
Softw. Pract. Exp. | 1 |
| 2023 | Feature transfer learning by reinforcement learning for detecting software defectabstractAbstract Software defects, produced inevitably in software projects, seriously affect the efficiency of software testing and maintenance. An appealing solution is the software defect prediction (SDP) that has achieved good performance in many software projects. However, the difference between features and the difference of the same feature between training data and test data may degrade defect prediction performance if such differences violate the model's assumption. To address this issue, we propose a SDP method based on feature transfer learning (FTL), which performs a transformation sequence for each feature in order to map the original features to another feature space. Specifically, FTL first uses the reinforcement learning scheme that automatically learns a strategy for transferring the potential feature knowledge from the training data. Then, we use the learned feature knowledge to inspire the transformation of the test data. The classifier is trained by the transformed training data and predicts defects for transformed test data. We evaluate the validity of FTL on 43 projects from PROMISE and NASA MDP using three classifiers, logistic regression, random forest, and Naive Bayes (NB). Experimental results indicate that FTL is better than the original classifiers and has the best performance on the NB classifier. For PROMISE, after using FTL, the average results of F1‐score, AUC, MCC are 0.601, 0.757, and 0.350 respectively, which are 24.9%, 2.6%, and 16.7% higher than the original NB classifier results. The number of projects with improved performance accounts for 83.87%, 83.87%, and 64.52%. Similarly, FTL performs well on NASA MDP. Besides, compared with four feature engineering (FE) methods, FTL achieves an excellent improvement on most projects and the average performance is also better than or close to the FE methods. Shikai Guo, Hui Li 0014, Rong Chen 0003 |
Softw. Pract. Exp. | 1 |
| 2023 | An Easy Data Augmentation Approach for Application Reviews Event InferenceabstractApplication review event inference aims to assess the effectiveness of application problems in response to user actions, which enables application developers to promptly discover and address potential issues in various applications, thereby improving their development and maintenance efficiency. Despite the development of event inference models for app reviews, which extract them as user action and app problem events and establish a relationship model between events and inference labels, the accuracy of these models is constrained due to limitations in labeling and characterizing noise and the lack of robustness and generalization. To address this challenge, we propose a model called Easy Data Augmentation for Application Reviews Event Inference (short for EDA-AREI), which comprises a denoising component, data augmentation component, and event inference prediction component. Specifically, the denoising component identifies labels and characterizes noisy data to enhance dataset quality, the data augmentation component replaces non-stop words with synonyms to increase textual diversity, and the event inference and prediction component reconstructs the classifier using denoised and augmented data. Experimental results on six datasets of one-star app reviews in the Apple App Store demonstrate that the EDA-AREI method achieves anAccuracyof 71.19%, 79.14%, 69.05%, 69.02%, 68.24% and 68.48%, respectively, representing an improvement of 0.83%–2.09% compared to state-of-the-art models. Regarding theF1-score, EDA-AREI achieves values of 71.30%, 69.93%, and 68.76% on the threshold_0.5, k-means_2, and random datasets, respectively, outperforming state-of-the-art models by 1.89%–4.02%. Furthermore, EDA-AREI achievesAUCvalues of 75.66% and 73.37% on the threshold_0.5 and k-means_2 datasets, respectively. As a result, EDA-AREI demonstrates substantial improvements inAccuracy, as well as enhancedF1-scoreandAUCacross most datasets, thereby enhancing the model's accuracy and robustness in identifying related action-problem pairs. Shikai Guo, Haorui Lin, Jiaoru Zhao, Hui Li 0014, Rong Chen 0003, He Jiang 0001 |
IEEE Trans. Software Eng. | 1 |
| 2023 | DupHunter: Detecting Duplicate Pull Requests in Fork-Based DevelopmentabstractThe emergence of numerous fork-based development platforms facilitates the development of Open-Source Software (OSS) projects. Developers across the world can fork software projects and submit their Pull Requests (PRs) to the projects. However, as the number of forks increases, numerous duplicate PRs might be submitted. These duplicate PRs may cause extra code review workload and frustrate developers working on the projects. To detect duplicate PRs, many approaches have been proposed, which analyze the similarity of different elements in PRs. However, previous approaches still suffer from unsatisfied detection accuracy due to two challenges. That is, they ignore the syntactic structural information of text elements in PRs and lack the joint reasoning between different elements of two PRs. In this study, we propose an automated duplicate PRs detector namedDupHunter(Duplicate PRsHunter), which includes a graph embedding component and a duplicate PRs detection component to address the above challenges. The graph embedding component uses a feature graph to represent a PR. It encodes the syntactic structure and semantics of text elements (e.g., the title and the description), as well as the knowledge of non-text elements (e.g., the submission time), to address the syntactic structural information challenge. The duplicate PRs detection component tackles the joint reasoning challenge using a graph matching network, which enables the information exchange and matching across different elements of two feature graphs with an attention coefficient mechanism. Experiments on 26 open-source projects show that DupHunter achieves an averageF1-score@1value of 0.650, significantly outperforming the state-of-the-art approaches by 3.2% to 48.1%. DupHunter can accurately detect duplicate PRs, with an averagePrecision@1value of 0.922 and an averageRecall@1value of 0.502. He Jiang 0001, Yulong Li 0001, Shikai Guo, Tao Zhang 0001, Hui Li 0014, Rong Chen 0003 |
IEEE Trans. Software Eng. | 3 |
| 2022 | Detecting Simulink compiler bugs via controllable zombie blocks mutationabstractAs a popular Cyber-Physical System (CPS) development tool chain, MathWorks Simulink is widely used to prototype CPS models in safety-critical applications, e.g., aerospace and healthcare. It is crucial to ensure the correctness and reliability of Simulink compiler (i.e., the compiler module of Simulink) in practice since all CPS models depend on compilation. However, Simulink compiler testing is challenging due to millions of lines of source code and the lack of the complete formal language specification. Although several methods have been proposed to automatically test Simulink compiler, there still remains two challenges to be tackled, namely the limited variant space and the insufficient mutation diversity. To address these challenges, we propose COMBAT, a new differential testing method for Simulink compiler testing. COMBAT includes an EMI (Equivalence Modulo Input) mutation component and a diverse variant generation component. The EMI mutation component inserts assertion statements (e.g., If /While blocks) at arbitrary points of the seed CPS model. These statements break each insertion point into true and false branches. Then, COMBAT feeds all the data passed through the insertion point into the true branch to preserve the equivalence of CPS variants. In such a way, the body of the false branch could be viewed as a new variant space, thus addressing the first challenge. The diverse variant generation component uses Markov chain Monte Carlo optimization to sample the seed CPS model and generate complex mutations of long sequences of blocks in the variant space, thus addressing the second challenge. Experiments demonstrate that COMBAT significantly outperforms the state-of-the-art approaches in Simulink compiler testing. Within five months, COMBAT has reported 16 valid bugs for Simulink R2021b, of which 11 bugs have been confirmed as new bugs by MathWorks Support. Shikai Guo, He Jiang 0001, Zhilei Ren, Zhide Zhou, Rong Chen 0003 |
ESEC/SIGSOFT FSE | 1 |
| 2022 | Self-admitted technical debt detection by learning its comprehensive semantics via graph neural networksabstractAbstract The goal of software development is to deliver software products with high quality and free from defects, but resource and time constraints often cause the developers to submit incomplete or temporary patches of codes and further bear the additional burden. Therefore, the investigations on identifying self‐admitted technical debt (SATD) to improve code quality have been conducted in recent years. However, missing syntactic structure information and the imbalance distribution bias shorten the SATD identification performance. Addressing to this issue, we present a graph neural network based SATD identification model (GNNSI) to improve the performance. Specifically, we obtain the structure information of the missing SATD in a compositional way to obtain different feature maps for different comments, and use focal loss to handle the imbalance between SATD and non‐SATD classes in the comments. Then extensive experiments on 10 open source projects are conducted, and the results show that GNNSI outperforms the baselines and can help developers to better predict SATDs. Hui Li 0014, Rong Chen 0003, Jun Ai, Shikai Guo |
Softw. Pract. Exp. | 6 |
| 2021 | Software defect prediction with imbalanced distribution by radius-synthetic minority over-sampling techniqueabstractAbstract Software defect prediction, which can identify the defect‐prone modules, is an effective technology to ensure the quality of software products. Due to the importance in software maintenance, many learning‐based software defect prediction models are presented in recent years. Actually, the defects usually occupy a very small proportions in software source codes; thus, the imbalanced distributions between defect‐prone modules and non‐defect‐prone modules increase the learning difficulty of the classification task. To address this issue, we present a random over‐sampling mechanism used to generate minority‐class samples from high‐dimensional sampling space to deal with the imbalanced distributions in software defect prediction, in which two constraints are applied to provide a robust way to generate new synthetic samples, that is, scaling the random over‐sampling scope to a reasonable area and distinguishing the majority‐class samples in a critical region. Based on nine open datasets of software projects, we experimentally verify that our presented method is effective on predict the defect‐prone modules, and the effect is superior to the traditional imbalanced processing methods. Shikai Guo, Hui Li 0014 |
J. Softw. Evol. Process. | 1 |
| 2021 | Toward more accurate developer recommendation via inference of development activities from interaction with bug repair processabstractAbstract Software projects usually receive a large number of submitted bug reports every day. Manually triaging the bug reports is often time‐consuming and error‐prone; thus, it is necessary to automatically assign the bug reports to the suitable developers for bug repair, with the help of bug tracking systems. Aiming to reducing the time consumption and mismatch of bug report assignments, we present a developer recommendation model for bug repair based on weighted recurrent neural network, namely, DTPM, which contains two parts: One obtains multisource semantic information of bug reports and fuses them into high‐dimensional semantic feature vectors, and the other combines a penalty matrix into a single hidden layer neural network to obtain more reasonable developer recommendations. We conduct experiments on five datasets of open bug repositories (NetBeans, OpenOffice, GCC, Mozilla, and Eclipse), and the experimental results show that DTPM can achieve better performance than state‐of‐the‐art models LDA_KL, LDA_KL, LDA_SVM, DERTOM, DREX, and DeepTriage. Linhui Wang, Rong Chen 0003, Shikai Guo |
J. Softw. Evol. Process. | 5 |
| 2020 | Developer Activity Motivated Bug Triaging: Via Convolutional Neural Network
Shikai Guo, Xi Yang 0009, Rong Chen 0003, Chen Guo 0001, Hui Li 0014 |
Neural Process. Lett. | 1 |
| 2019 | Identify Severity Bug Report with Distribution Imbalance by CR-SMOTE and ELMabstractManually inspecting bugs to determine their severity is often an enormous but essential software development task, especially when many participants generate a large number of bug reports in a crowdsourced software testing context. Therefore, boosting the capabilities of methods of predicting bug report severity is critically important for determining the priority of fixing bugs. However, typical classification techniques may be adversely affected when the severity distribution of the bug reports is imbalanced, leading to performance degradation in a crowdsourcing environment. In this study, we propose an enhanced oversampling approach called CR-SMOTE to enhance the classification of bug reports with a realistically imbalanced severity distribution. The main idea is to interpolate new instances into the minority category that are near the center of existing samples in that category. Then, we use an extreme learning machine (ELM) — a feedforward neural network with a single layer of hidden nodes — to predict the bug severity. Several experiments were conducted on three datasets from real bug repositories, and the results statistically indicate that the presented approach is robust against real data imbalance when predicting the severity of bug reports. The average accuracies achieved by the ELM in predicting the severity of Eclipse, Mozilla, and GNOME bug reports were 0.780, 0.871, and 0.861, which are higher than those of classifiers by 4.36%, 6.73%, and 2.71%, respectively. Shikai Guo, Rong Chen 0003, Hui Li 0014, Tianlun Zhang |
Int. J. Softw. Eng. Knowl. Eng. | 1 |
| 2019 | The Influence Ranking for Testers in Bug Tracking SystemsabstractAt present, bug tracking systems are used to collect and manage bug reports in many software projects. As participants, the testers not only submit bug reports to the system, but also comment on bug reports in the system. The tester’s behaviors of submitting and commenting reflect his/her influence in bug tracking systems. However, with the rapid increase of the bug reports in software projects, evaluating the testers’ influence in the projects accurately becomes more and more difficult. Aiming at solving this problem, the submission and comment on bug report can be regarded as social behaviors of the testers, and thus the method of Influence Ranking for Testers (IRfT) in bug tracking systems is presented and used for measuring the influence of the testers in this paper. The case study of the Eclipse project in Bugzilla shows that the result produced by IRfT is consistent with the actual performance of the testers in this project. The ranking results can keep stable in the cases of link adding or removing and tester removing in tester networks, and the results are also proved to be valid in the future. The further investigation on the speed of network break-down by node removal demonstrates that the top-ranking testers are important in the organization of tester networks. Additionally, the results also show that the ranking of the testers is related to the existence time in bug tracking system. Therefore, IRfT is proved to be an effective measurement for evaluating the influence of the testers in bug tracking system, and it can further demonstrate the testers’ contributions in software testing, such as bug validations, bug fixes, etc. Hui Li 0014, Guofeng Gao, Rong Chen 0003, Shikai Guo |
Int. J. Softw. Eng. Knowl. Eng. | 5 |
| 2019 | A Novel Approach to Publishing Tasks for Collaboratively Crowdsourcing WorkflowsabstractIn recent years, crowdsourcing has gradually become a promising way of using netizens to accomplish tiny tasks on, or even complex works through crowdsourcing workflows that decompose them into tiny ones to publish sequentially on the crowdsourcing platforms. One of the significant challenges in this process is how to determine the parameters for task publishing. Still some technique applied constraint solving to select the optimal tasks parameters so that the total cost of completing all tasks is minimized. However, experimental results show that computational complexity makes these tools unsuitable for solving large-scale problems because of its excessive execution time. Taking into account the real-time requirements of crowdsourcing, this study uses a heuristic algorithm with four heuristic strategies to solve the problem in order to reduce execution time. The experiment results also show that the proposed heuristic strategies produce good quality approximate solutions in an acceptable timeframe. Rong Chen 0003, Shikai Guo |
Int. J. Softw. Eng. Knowl. Eng. | 3 |
| 2019 | Improved SMOTE Algorithm to Deal with Imbalanced Activity Classes in Smart Homes
Shikai Guo, Rong Chen 0003, Xiao Sun 0003, Xiangxin Wang |
Neural Process. Lett. | 1 |
| 2019 | Fusion of Multi-RSMOTE With Fuzzy Integral to Classify Bug Reports With an Imbalanced DistributionabstractWith the help of automated classification, severe bugs can be rapidly identified so that the latent damage to software projects can be minimized. However, bug report datasets commonly suffer from disproportionate number of category samples. When presented with the situation of class imbalance, most standard classification learning approaches fail to properly learn the distributive characteristics of the samples and tend to result in unfavorable performance to predict class label. In this case, imbalanced learning becomes critical to advance classification algorithms. In this paper, we propose an improved synthetic minority oversampling technique to avoid the degraded performance caused by class imbalance in bug report datasets. Moreover, to lessen the chance of occasionalities in random sampling process, we propose a repeated sampling technique to train different, but related classifiers. Finally, an ensemble algorithm based on Choquet fuzzy integral is employed to combine the wisdom of crowds and make better decisions. We conduct comprehensive experiments on several bug report datasets from real-world bug repositories. The results demonstrate that the proposed method boosts the classification performance across the classes of the data. Specifically, compared with various ensemble learning techniques, the Choquet fuzzy integral achieves outstanding results on integrating multiple random oversampling techniques. Rong Chen 0003, Shikai Guo, Xizhao Wang, Tianlun Zhang |
IEEE Trans. Fuzzy Syst. | 2 |
| 2018 | Enhancing Bug Report Assignment with an Optimized Reduction of Training Set
Miaomiao Wei, Shikai Guo, Rong Chen 0003, Jian Gao 0007 |
KSEM (2) | 2 |
| 2018 | Weighted Data Set Reduction for Automatic Bug Triaging (P)abstractDespite the great potential to save the labor cost of developers, automated bug triaging as a text classification problem has not been thoroughly investigated on long descriptions, which are informative but often noisy.In this paper an effective bug triage technique is proposed to build a high quality set of bug data by removing the noisy and noninformative bug reports while assigning new bugs to an appropriate developer.The proposed techniqueweighted data set reductionis built upon three feature selection algorithms and four instances selection algorithms with intention to recommend the bug and to automatically assign it more accurately even with noisy bug descriptions.Several experiments are conducted and the experimental results show that the reduced training sets by the proposed approach can achieve better accuracy in several cases, about 2-3% on average better than the original ones. Miaomiao Wei, Shikai Guo, Rong Chen 0003 |
SEKE | 2 |
| 2018 | Capability Matching and Heuristic Search for Job Assignment in Crowdsourced Web Application TestingabstractWeb based commercial systems are increasingly becoming feature rich, interactive and functional as locally installed applications. Testing web applications is unique, as many factors affect the system performance and user experience. Crowdsourcing is an appealing and economic solution to web application testing due to the ability to reach a larger international audience. However, less is known about the quality control of crowdsourced testing to harness the collective efforts of individuals. In our study, the collaborative testing problem in a crowdsourcing environment is defined as a job assignment problem and is formulated as an integer linear programming (ILP) problem. The objective of this paper is to validate a greedy job assignment approach as a tool for the effective use of crowdsourced testing. We carried out a case study on Xturk, a prototype crowdsourced testing system, to understand the crowdsourced testers behaviour that the trustworthiness, the execution time of test cases and accuracy of feedback. Several experiments indicate that this approach is comparatively effective with regards to the feasibility verdict, efficiency and accuracy. Shikai Guo, Rong Chen 0003, Hui Li 0014 |
SMC | 1 |
| 2018 | Crowdsourced Web Application Testing Under Real-Time ConstraintsabstractCrowdsourcing carried out by cyber citizens instead of hired consultants and professionals has become increasingly an appealing solution to test the feature rich and interactive web. Despite having various online crowdsourcing testing services, the benefits of exposure to a wider audience and harnessing the collective efforts of individuals remain uncertain, especially when the quality control is problematic in an open environment. The objective of this paper is to propose a real-time collaborative testing approach (RCTA) to create a productive crowdsourced testing on a dynamic Internet. We implemented a prototype crowdsourcing system XTurk, and carried out a case study, to understand the crowdsourced testers behavior, the trustworthiness, the execution time of test cases and accuracy of feedback. Several experiments are carried out and experimental results validate the quality, efficiency and reliability of the present approach and the positive testing feedback is are shown to outperform the previous methods. Shikai Guo, Rong Chen 0003, Hui Li 0014, Jian Gao 0007 |
Int. J. Softw. Eng. Knowl. Eng. | 1 |