VLDB 2026 Research / reviewers in the wild / expert
Zhide Zhou
dblp:181/4883
· DBLP profile ↗
26ranked-venue papers
3as first author
22since 2021 · last 2026
0000-0001-6195-7605ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 14 · 1 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 6 since 2021Systems, architecture and hardware · 4 · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Accelerating Deep Learning Compiler Testing via Message-Passing Neural Network With AttentionabstractDeep learning (DL) compilers are crucial for optimizing DL models across diverse hardware platforms, ensuring their reliability is paramount. Although many techniques have been developed to test DL compilers, existing methods often suffer from efficiency issues. To resolve this problem, this study proposes MeDAC, a novelMessage passing neural network based framework forDL compiler testingACceleration. The key insight of MeDAC is to construct a learning model to accurately predict the bug-revealing probabilities of test cases (i.e., DL models), allowing the test cases with higher bug-revealing probabilities to be executed. First, to construct the learning model, three types of features for DL models are extracted. Then, we propose a novel learning model based on the message passing neural network with an attention mechanism to accurately estimate the bug-revealing probabilities of DL models. Finally, MeDAC utilizes the learning model to predict the bug-revealing probability of each DL model, ensuring high-risk models are executed first. Experimental results on TVM and ONNXRuntime demonstrate that MeDAC achieves an average of 59.20% speedup in test execution time for DL compiler testing. Moreover, the peak performance improvements of MeDAC are 385.68% and 217.96% over the baseline approaches LET and GCN, respectively. Yiqi Ding, Zhide Zhou, Peiyu Zou, He Jiang 0001 |
IEEE Trans. Reliab. | 2 |
| 2026 | ProFuse: Test Case Prioritization Based on Multi Dimensional Feature Fusion for Logic Synthesis Tools Testing AccelerationabstractLogic synthesis tools translate Hardware Description Language (HDL) designs into hardware implementation. To test these tools, numerous test cases are usually executed on the tools, yet only a few of them can trigger faults, leading to inefficient testing. Since executing test cases on logic synthesis tools often requires significant cost on complicated synthesis and simulation, fault-triggering test cases should be prioritized to execute. However, existing prioritization methods face challenges in accurately predicting the fault-triggering capability of dynamically generated test cases and modeling the unique syntactic and structure complexities of these HDL-based programs.Therefore, we propose ProFuse, a multi-dimensional feature fusion method for logic synthesis tool test case prioritization. ProFuse leverages Abstract Syntax Trees (AST) and Data Flow Graphs (DFG) to extract novel syntactic and structure features from HDL designs. These features are processed by a joint model of Multilayer Perceptron (MLP) and Graph Convolutional Network (GCN) to rank fault-triggering test cases accurately. ProFuse achieves an Average Percentage of Fault Detection (APFD) score of 0.9285, outperforming the state-of-the-art prioritization methods by 11.38% to 82.49%. ProFuse can efficiently rank randomly generated test cases to discover 15 new faults in logic synthesis tools (i.e., Yosys and Vivado). The Vivado community acknowledged our work for improving their tool. Peiyu Zou, Shikai Guo, Zhide Zhou, He Jiang 0001 |
IEEE Trans. Software Eng. | 5 |
| 2025 | Multidemand Forecasting for Electric Vehicle Charging Stations Under Time-of-Use Strategy via Attention-Based Deep Neural NetworkabstractElectric vehicle charging stations (EVCSs) have become a pivotal infrastructure within the electric vehicle (EV) industry. In particular, many EV companies construct self-owned EVCSs to provide better charging service for their customers. For these self-owned EVCSs, to ensure the quality of service for self-owned users and third-party users, dynamic pricing based on the time-of-use (TOU) strategy has been extensively employed. This makes the demand forecasting of EVCSs important since it depicts the relationship between the charging price and the demand of an EVCS. Unfortunately, the existing techniques cannot accurately predict the demand of multiple users simultaneously. Consequently, this article examines the problem of multidemand forecasting of EVCSs, and proposes an efficient method to resolve this issue. The key insight of the proposed method is to train a deep neural network consisting of two subnetworks that can jointly forecast the demand of the self-owned user and the third-party user simultaneously. First, six kinds of features of EVCSs are extracted. Then, a novel deep neural network Atlas based on the attention mechanism is proposed to forecast the multidemand of EVCSs under the TOU strategy. Finally, to resolve the scarcity of historical charging demand data, a coarse-fine training process is proposed to train Atlas for each EVCS. The evaluation based on the real-world dataset of 771 EVCSs from an EV company demonstrates that Atlas significantly outperforms seven state-of-the-art techniques by up to 34.82%$\sim ~61.92$%. Zhide Zhou, He Jiang 0001, Shaolin Wang, Haoyang Che |
IEEE Internet Things J. | 1 |
| 2025 | Detecting WebAssembly Runtime Bugs With Grammar-Guided Program Mutation
Zhide Zhou, Jifeng Xuan, He Jiang 0001, Zhilei Ren |
IEEE Trans. Reliab. | 2 |
| 2025 | Learning to Accelerate Autonomous Driving System TestingabstractSystem bug identification plays a key role in autonomous vehicles for avoiding disastrous consequences. However, it is time consuming to fully test autonomous vehicles. Although some techniques have been proposed to improve testing efficiency, they struggle to handle the complex test scenarios because they cannot adequately represent the scenarios and efficiently assess the vulnerability-triggering potential of test scenarios. In this study, we propose Learning to accElerate Autonomous Systems Testing (LEAST), a graphic neural network (GNN)-based method to effectively accelerate autonomous driving system (ADS) testing. First, given a test scenario, LEAST extracts a series of scenes from the test scenario and constructs feature graphs from these scenes, which are helpful to characterize the test scenario. Then, we propose a GNN-based method to predict the risk value of each scene, indicating the likelihood of a vehicle encountering a traffic accident. Finally, LEAST prioritizes test scenarios based on the number of high-risk scenes in the scenarios, ensuring that scenarios more likely to trigger ADS bugs are executed earlier. Experiments onthree open-source ADSs show that LEAST significantly outperforms baseline test acceleration approaches by 10.13–15.34% in terms of average percentage of fault detected. When integrating LEAST into the advanced ADS testing approach DriveFuzz, LEAST successfully improves the testing efficiency by 39–67%. Zhide Zhou, Shikai Guo, He Jiang 0001 |
IEEE Trans. Reliab. | 2 |
| 2024 | Latency-Based Inter-Operator Scheduling for CNN Inference Acceleration on GPUabstractConvolutional Neural Networks (CNNs) are widely deployed on the Graphics Processing Unit (GPU) to support Deep Learning (DL) based services. Popular DL frameworks usually ignore the inter-operator parallelism when executing the inference of CNNs, which results in high inference latency. Although some inter-operator scheduling methods have been proposed, there remains a critical trade-off issue between inference latency (effectiveness) and scheduling time (efficiency). In this article, we propose LIOS, a novel latency-based heuristic inter-operator scheduling method to balance inference latency and scheduling time. In LIOS, a CNN latency model is built based on the given CNN and GPU. Then every operator is assigned a priority value to represent its importance. During each iteration of the scheduling process, LIOS identifies the current data-independent operators, selects the operator with the highest priority value, and assigns it to the GPU stream with the smallest finish time. Extensive experimental results have demonstrated the effectiveness and efficiency of LIOS. For the effectiveness, LIOS can speed up the inference of normal-size and large-size CNNs by 1.13$\sim 1.59 \times$compared to sequential scheduling. This result is comparable to IOS, the latest state-of-the-art scheduling method. For the efficiency, LIOS can speed up the scheduling process by 7$\sim 9210\times$compared to IOS. Yukai Ping, He Jiang 0001, Xingxiang Liu, Zhenyang Zhao, Zhide Zhou, Xin Chen 0032 |
IEEE Trans. Serv. Comput. | 5 |
| 2024 | HetFL: Heterogeneous Graph-Based Software Fault LocalizationabstractAutomated software fault localization has become one of the hot spots on which researchers have focused in recent years. Existing studies have shown that learning-based techniques can effectively localize faults leveraging various information. However, there exist two problems in these techniques. The first is that they simply represent various information without caring the contribution of different information. The second is that the data imbalance problem is not considered in these techniques. Thus, their effectiveness is limited in practice. In this paper, we propose HetFL, a novel heterogeneous graph-based software fault localization technique to aggregate different information into a heterogeneous graph in which program entities and test cases are regarded as nodes, and coverage, change histories, and call relationships are viewed as edges. HetFL first extracts textual and structure information from source code as attributes of nodes and integrates them to form an attribute vector. Then, for a given node, HetFL finds its neighbor nodes based on the types of edges and aggregates corresponding neighbor nodes to form type vectors. After that, the attribute vector and all the type vectors of each node are aggregated to generate the final vector representation by an attention mechanism. Finally, we leverage a convolution neural network (CNN) to obtain the suspicious score of each method. To validate the effectiveness of HetFL, experiments are conducted on the widely used dataset Defects4J (v1.2.0). The experimental results show that HetFL can localize 217 faults within Top-1 that is 25 higher than the state-of-the-art technique DeepFL, and achieve 6.37 and 5.58 in terms of MAR and MFR which improve DeepFL by 9.0% and 5.6%, respectively. In addition, we also perform experiments on the latest version of Defects4J (v2.0.0). The experimental results show that HetFL has better performance than the baseline methods. Xin Chen 0032, Dongling Zhuang, Dongjin Yu, He Jiang 0001, Zhide Zhou, Sicheng Li 0010 |
IEEE Trans. Software Eng. | 6 |
| 2024 | A Testing Program and Pragma Combination Selection Based Framework for High-Level Synthesis Tool Pragma-Related Bug DetectionabstractHigh-Level Synthesis (HLS) tools convert C/C++ design code into Hardware Description Language (HDL) code automatically, which are often used for Field Programmable Gate Array (FPGA) design. HLS tools provide many pragmas, which are a kind of directive to be inserted into C/C++ code, for designers to efficiently control the synthesis of code components (e.g., arrays and loops) to generate FPGA implementations with varying performances and costs. However, the use of some pragmas may trigger HLS tool bugs (e.g., tool crashes). Although many formal methods have been proposed to verify the correctness of various HLS phases, no relevant work addresses the problem on detecting HLS tool pragma-related bugs. To resolve this problem, two challenges need to be addressed, namely the selection of testing programs and the acquisition of pragma combinations, due to the enormous number of testing programs and pragma combinations. In this paper, we propose TEPACS, a TEsting Program and prAgma Combination Selection-based framework, to construct diverse testing programs with pragmas for effectively detecting HLS tool pragma-related bugs. TEPACS follows the idea of fuzzing, which is a widely used technique in software testing. First, TEPACS selects the representative testing program according to the cosine distance between the code component vectors of testing programs. Then, for a selected program, TEPACS generates its golden output and uses the pragma combination selection method based on combinatorial testing to generate a set of programs with different pragmas. TEPACS uses the HLS tool under test to convert these testing programs into HDL codes and obtains the simulation results of the HDL code. Finally, based on differential testing, TEPACS identifies HLS tool bugs triggered if the simulation result and golden output are inconsistent. We evaluate TEPACS and its five variants on Vitis HLS, a widely used FPGA HLS tool. Experimental results show that TEPACS outperforms the baselines by at least 11.17% in terms of the bug-finding capability. In one month, TEPACS detected 34 bugs on the latest version of Vitis HLS, of which 9 bugs have been confirmed. He Jiang 0001, Zun Wang 0005, Zhide Zhou, Shikai Guo, Weifeng Sun 0002, Tao Zhang 0001 |
IEEE Trans. Software Eng. | 3 |
| 2024 | Isolating Compiler Bugs by Generating Effective Witness Programs With Large Language ModelsabstractCompiler bugs pose a significant threat to safety-critical applications, and promptly as well as effectively isolating these bugs is crucial for assuring the quality of compilers. However, the limited availability of debugging information on reported bugs complicates the compiler bug isolation task. Existing compiler bug isolation approaches typically convert the problem into a test program mutation problem, but they are still limited by ineffective mutation strategies or high human effort requirements. Drawing inspiration from the recent progress of pre-trained Large Language Models (LLMs), such as ChatGPT, in code generation, we propose a new approach named LLM4CBI to utilize LLMs to generate effective test programs for compiler bug isolation. However, using LLMs directly for test program mutation may not yield the desired results due to the challenges associated with formulating precise prompts and selecting specialized prompts. To overcome the challenges, three new components are designed in LLM4CBI. First, LLM4CBI utilizes a program complexity-guided prompt production component, which leverages data and control flow analysis to identify the most valuable variables and locations in programs for mutation. Second, LLM4CBI employs a memorized prompt selection component, which adopts reinforcement learning to select specialized prompts for mutating test programs continuously. Third, a test program validation component is proposed to select specialized feedback prompts to avoid repeating the same mistakes during the mutation process. Compared with the state-of-the-art approaches (DiWi and RecBi) over 120 real bugs from the two most popular compilers, namely GCC and LLVM, our evaluation demonstrates the advantages of LLM4CBI: It can isolate 69.70%/21.74% and 24.44%/8.92% more bugs than DiWi and RecBi within Top-1/Top-5 ranked results. Additionally, we demonstrate that the LLMs component (i.e., GPT-3.5) used in LLM4CBI can be easily replaced by other LLMs while still achieving reasonable results in comparison to related studies. Haoxin Tu, Zhide Zhou, He Jiang 0001, Imam Nur Bani Yusuf, Lingxiao Jiang |
IEEE Trans. Software Eng. | 2 |
| 2023 | Detecting JavaScript Transpiler Bugs with Grammar-guided MutationabstractJavaScript (JS) transpilers translate JS programs from a higher grammar standard to a lower one, which are widely used to ensure the compatibility of JS features in software (e.g., browsers). However, JS transpilers can have bugs that lead to unintended behavior in the translated JS programs. Existing JS program generation approaches could not test JS transpilers effectively since it is hard to generate a large number of valid JS programs in specific grammar standards. In this paper, we propose TransFuzz, a grammar-guided mutation approach to find JS transpiler bugs.The key insight of TransFuzz is to generate syntax-specific JS programs by mutating the abstract syntax trees (ASTs) of JS programs with the guidance of the specific grammar. First, Trans- Fuzz parses JS programs collected from open-source platforms into ASTs to obtain subtrees and leaf nodes containing specific JS syntax. Then, a grammar-guided approach is developed in TransFuzz to mutate the ASTs of the given JS programs guided by different versions of JS grammar standards. In addition, mutation operations could introduce grammatical errors. To improve the correctness of the mutated ASTs, TransFuzz develops heuristic-based correction rules to correct reference errors, type errors, and syntax errors in the mutated ASTs. After correction, the mutated ASTs are converted to the corresponding JS programs. Finally, based on differential testing, TransFuzz utilizes the generated JS programs to detect JS transpiler bugs.Our evaluation shows that TransFuzz significantly outperforms existing JS program generation approaches by triggering 47.82%-385.71% more JS transpiler bugs. Within ten months, we have reported 73 bugs on two popular JS transpilers babel and swc, of which 58 have been confirmed. Zhide Zhou, He Jiang 0001 |
SANER | 2 |
| 2023 | A Comprehensive Study of WebAssembly Runtime BugsabstractWebAssembly runtime is the infrastructure for executing WebAssembly, which is widely used as an execution engine by web browsers or blockchain platforms. Bugs in the WebAssembly runtime can lead to unexpected behavior and even security vulnerabilities in any application that relies on it. Therefore, to aid developers in understanding the WebAssembly runtime, a thorough investigation of bugs in the WebAssembly runtime should be conducted. To accomplish this, we carry out the first empirical analysis of 867 real bugs across four popular WebAssembly runtimes (V8, SpiderMonkey, Wasmer, and Wasmtime). We analyze the WebAssembly runtime bug characteristics based on their root causes, symptoms, bug-fixing time, and the number of files and lines of code involved in the bug fixes. Here are a few major research findings: 1) Incorrect Algorithm Implementation accounts for 25.49% of WebAssembly runtime bugs, the most prevalent of all root causes; 2) The most prevalent symptom is Crash, which accounts for 56.86% of WebAssembly runtime bugs; 3) At the median, the bug-fixing time are 13, 4, 5, and 6 days for V8, SpiderMonkey, Wasmer, and Wasmtime respectively; 4) Over 50% of bug fixes in the four WebAssembly runtimes involve only one file, while more than 90% of bug fixes involve no more than 8 files; 5) The median source code lines for bug fixes for V8, SpiderMonkey, Wasmer, and Wasmtime are 18.5, 14, 26, and 36 lines, respectively. Overall, our research summarizes 18 findings and discusses the broad implications for WebAssembly runtime bug detection, localization, debugging, and repair based on the key findings. Zhide Zhou, Zhilei Ren, Dong Liu 0025, He Jiang 0001 |
SANER | 2 |
| 2023 | Detecting C++ Compiler Front-End Bugs via Grammar Mutation and Differential TestingabstractC++ is a widely used programming language and the C++ front-end is a critical part of a C++ compiler. Although many techniques have been proposed to test compilers, few studies are devoted to detecting bugs in C++ compiler. In this study, we take the first step to detect bugs in C++ compiler front-ends. To do so, two main challenges need to be addressed, namely, the acquisition of test programs that are more likely to trigger bugs in compiler front-ends and the bug identification from complicated compiler outputs. In this article, we propose a novel framework namedCcoftto detect bugs in C++ compiler front-ends. To address the first challenge,Ccoftimplements a practical program generator. The generator first transforms C++ grammars into a flexible structured format and then utilizes an equal-chance selection (ECS) strategy to conduct structure-aware grammar mutation to generate diverse C++ programs. Next,Ccoftemploys a set of differential testing strategies to identify various kinds of bugs in C++ compiler front-ends by comparing complex outputs emitted by C++ compilers, thus tackling the second challenge. Empirical evaluation results over two mainstream compilers (i.e., GCC and Clang) show thatCcoftgreatly improves two state-of-the-art approaches (i.e., Dharma and Grammarinator) by 135% and 111% in terms of the numbers of detected bugs, respectively. By runningCcoftfor three months, we have successfully reported 136 bugs for two C++ compilers, of which 78 (57 confirmed, assigned, or fixed) for GCC and 58 (10 confirmed or fixed) for Clang. Haoxin Tu, He Jiang 0001, Zhide Zhou, Zhilei Ren, Lei Qiao 0002, Lingxiao Jiang |
IEEE Trans. Reliab. | 3 |
| 2022 | Automated Patching for Unreproducible BuildsabstractSoftware reproducibility plays an essential role in establishing trust between source code and the built artifacts, by comparing compilation outputs acquired from independent users. Although the testing for unreproducible builds could be automated, fixing unreproducible build issues poses a set of challenges within the reproducible builds practice, among which we consider the localization granularity and the historical knowledge utilization as the most significant ones. To tackle these challenges, we propose a novel approach RepFix that combines tracing-based fine-grained localization with history-based patch generation mechanisms. Zhilei Ren, Shiwei Sun, Jifeng Xuan, Zhide Zhou, He Jiang 0001 |
ICSE | 5 |
| 2022 | Remgen: Remanufacturing a Random Program Generator for Compiler TestingabstractProgram generators play a critical role in generating bug-revealing test programs for compiler testing. However, existing program generators have been tamed nowadays (i.e., compilers have been hardened against test programs generated by them), thus calling for new solutions to improve their capability in generating bug-revealing test programs. In this study, we propose a framework named Remgen, aiming to Remanufacture a random program Generator for this purpose. RemgEnaddresses the challenges of the synthesis of diverse code snippets at a low cost and the selection of the bug-revealing code snippets for constructing new test programs. More specifically, RemgEnfirst designs a grammar-aided synthesis mechanism to synthesize diverse code snippets. Then, a grammar coverage-guided strategy is used to select the most diverse code snippets that may be bug-revealing. As a case study to demonstrate the effectiveness of the Remgen framework, we have remanufactured an old C program generator CCG and named it REMCCG. Our evaluation results show that REMCCG can generate significantly more bug-revealing test programs than the original CCG; notably, Remccg has found 56 new bugs for two mature compilers (i.e., GCC and LLVM), of which 37 have already been fixed by their developers. Haoxin Tu, He Jiang 0001, Zhilei Ren, Zhide Zhou, Lingxiao Jiang |
ISSRE | 5 |
| 2022 | Detecting Simulink compiler bugs via controllable zombie blocks mutationabstractAs a popular Cyber-Physical System (CPS) development tool chain, MathWorks Simulink is widely used to prototype CPS models in safety-critical applications, e.g., aerospace and healthcare. It is crucial to ensure the correctness and reliability of Simulink compiler (i.e., the compiler module of Simulink) in practice since all CPS models depend on compilation. However, Simulink compiler testing is challenging due to millions of lines of source code and the lack of the complete formal language specification. Although several methods have been proposed to automatically test Simulink compiler, there still remains two challenges to be tackled, namely the limited variant space and the insufficient mutation diversity. To address these challenges, we propose COMBAT, a new differential testing method for Simulink compiler testing. COMBAT includes an EMI (Equivalence Modulo Input) mutation component and a diverse variant generation component. The EMI mutation component inserts assertion statements (e.g., If /While blocks) at arbitrary points of the seed CPS model. These statements break each insertion point into true and false branches. Then, COMBAT feeds all the data passed through the insertion point into the true branch to preserve the equivalence of CPS variants. In such a way, the body of the false branch could be viewed as a new variant space, thus addressing the first challenge. The diverse variant generation component uses Markov chain Monte Carlo optimization to sample the seed CPS model and generate complex mutations of long sequences of blocks in the variant space, thus addressing the second challenge. Experiments demonstrate that COMBAT significantly outperforms the state-of-the-art approaches in Simulink compiler testing. Within five months, COMBAT has reported 16 valid bugs for Simulink R2021b, of which 11 bugs have been confirmed as new bugs by MathWorks Support. Shikai Guo, He Jiang 0001, Zhilei Ren, Zhide Zhou, Rong Chen 0003 |
ESEC/SIGSOFT FSE | 6 |
| 2022 | Detecting Compiler Bugs Via a Deep Learning-Based FrameworkabstractCompiler testing is the most widely used way to assure compiler quality. However, since compilers require a large number of sophisticated test programs as inputs, the existing approaches in compiler testing still have a limited capability in generating both syntactically valid and diverse test programs. In this paper, we propose DeepGen, a deep learning-based approach to support compiler testing through the inference of a generative model for compiler inputs. First, DeepGen trains a Transformer-XL model based on a large corpus of seed programs, and uses the trained model to generate syntactically valid programs. Then, DeepGen adopts a sampling strategy in the inference phase to generate diverse test programs. Finally, DeepGen leverages differential testing on the generated programs to discover compiler bugs. We have evaluated DeepGen over two popular C++ compilers GCC and LLVM, and the results confirm the effectiveness of our approach. DeepGen detects 35.29%, 53.33%, and 187.50% more bugs than three existing approaches, i.e. DeepSmith, DeepFuzz, and Csmith, respectively. In addition, 30.43% bugs detected by DeepGen are not detected by other approaches. Furthermore, DeepGen has successfully detected 38 bugs in the latest development versions of GCC and LLVM; 21 of them have been confirmed/fixed by the developers. Zhilei Ren, He Jiang 0001, Lei Qiao 0002, Dong Liu 0025, Zhide Zhou, Weiqiang Kong |
Int. J. Softw. Eng. Knowl. Eng. | 6 |
| 2022 | SMARTEST: A Surrogate-Assisted Memetic Algorithm for Code Size ReductionabstractCompiling source code effectively to meet various criteria is a critical task in software engineering. Especially, code size reduction has attracted much attention from both industry and academia due to the requirement of resource utilization. Generally, developers rely on compiler optimization passes to realize code size reduction. However, it is impractical to select a desirable optimization sequence manually since a wide variety of optimization passes are integrated into a compiler. Evolutionary algorithms offer an impressive way to alleviate this problem. Nevertheless, previous approaches fail to balance the exploitation and exploration of the search space. Moreover, the expensive fitness evaluation requires actual compilation, which makes the evolution rather time-consuming. To tackle the challenges, we propose a novel approach SMARTEST, which characterizes the systematic exploitation of a huge volume of historical compilation information. Specifically, SMARTEST comprises two components: 1) a local search operator to enhance the solution quality; and 2) a data-driven surrogate model to avoid expensive fitness evaluation. We evaluate the effectiveness of SMARTEST over the cBench benchmark suite. Experimental results indicate that SMARTEST outperforms the standard level -Os by 2.17% on average, and achieves 1.2 times code size reduction compared with the genetic algorithm. Furthermore, experimental results over the benchmark suite evidently show that SMARTEST gets a better result and takes less actual fitness evaluations than its variants, which demonstrates the contribution of the local search and the surrogate model. He Jiang 0001, Guojun Gao, Zhilei Ren, Xin Chen 0032, Zhide Zhou |
IEEE Trans. Reliab. | 5 |
| 2022 | LocSeq: Automated Localization for Compiler Optimization Sequence Bugs of LLVMabstractCompiler bugs may be triggered when programs are optimized with optimization sequences. However, diagnosing compiler optimization sequence bugs is difficult due to limited debugging information. Although some techniques (e.g., DiWi and RecBi) have been proposed to automatically localize compiler bugs, no systematic work has been conducted to automatically localize compiler optimization sequence bugs. In this article, we propose LocSeq, a novel technique to automatically localize compiler optimization sequence bugs of LLVM. The core insight of LocSeq is based on the fact that the behaviors of optimizations may be influenced by each other, and thus, the innocent files may be excluded by constructing bug-free optimization sequences. First, given a buggy optimization sequence that triggers a compiler bug, in LocSeq, we transform the problem of the localization for a compiler optimization sequence bug to the problem of the construction for bug-free optimization sequences, which are helpful to localize buggy compiler files. Then, a constrained genetic algorithm is presented in LocSeq to generate a set of bug-free optimization sequences that share similar compiler execution traces with the buggy optimization sequence. Finally, LocSeq leverages a spectrum-based bug localization technique to localize the compiler optimization sequence bug by comparing the execution traces between bug-free optimization sequences and the buggy optimization sequence. To evaluate the effectiveness of LocSeq, we build a benchmark, including 60 optimization sequence bugs of LLVM, and compare LocSeq with the state-of-the-art techniques DiWi and RecBi. The experimental results show that LocSeq significantly outperforms DiWi and RecBi by up to 366.66%/72.27% and 250.00%/56.00% for localizing optimization sequence bugs within Top-1/5 files, respectively. Zhide Zhou, He Jiang 0001, Zhilei Ren, Yuting Chen 0001, Lei Qiao 0002 |
IEEE Trans. Reliab. | 1 |
| 2022 | CTOS: Compiler Testing for Optimization Sequences of LLVMabstractOptimization sequences are often employed in compilers to improve the performance of programs, but may trigger critical compiler bugs, e.g., compiler crashes. Although many methods have been developed to automatically test compilers, no systematic work has been conducted to detect compiler bugs when applying arbitrary optimization sequences. To resolve this problem, two main challenges need to be addressed, namely the acquisition of representative optimization sequences and the selection of representative testing programs, due to the enormous number of optimization sequences and testing programs. In this study, we propose CTOS, a novel compiler testing method based on differential testing, for detecting compiler bugs caused by optimization sequences of LLVM. CTOS first leverages the technique Doc2Vec to transform optimization sequences into vectors to capture the information of optimizations and their orders simultaneously. Second, a method based on the region graph and call relationships is developed in CTOS to construct the vector representations of the testing program, such that the semantics and the structure information of programs can be captured simultaneously. Then, with the vector representations of optimization sequences and testing programs, a “centroid” based selection scheme is proposed to address the above two challenges. Finally, CTOS takes in the representative optimization sequences and testing programs as inputs, and tests each testing program with all the representative optimization sequences. If there is an output that is different from the majority of others of a given testing program, then the corresponding optimization sequence is deemed to trigger a compiler bug. Our evaluation demonstrates that CTOS significantly outperforms the baselines by up to$24.76\% \sim 50.57\%$in terms of the bug-finding capability on average. Within seven month evaluations on LLVM, we have reported 104 valid bugs within 5 types, of which 21 have been confirmed or fixed. Most of those bugs are crash bugs (57) and wrong code bugs (24). 47 unique optimizations are identified to be faulty and 15 of them are loop related optimizations. He Jiang 0001, Zhide Zhou, Zhilei Ren |
IEEE Trans. Software Eng. | 2 |
| 2022 | Detecting Compiler Warning Defects Via Diversity-Guided Program MutationabstractCompiler diagnostic warnings help developers identify potential programming mistakes during program compilation. However, these warnings could be erroneous due to the defects of compiler warning diagnostics. Although the existing technique (i.e., Epiphron) can automatically generate test programs for compiler warning defect detection, the effectiveness of Epiphron on defect-finding is still limited, due to the limitation for generating warning-sensitive test program structures. Therefore, in this paper, we propose a DIversity-guided PROgram Mutation approach, called DIPROM, to construct diverse warning-sensitive programs for effective compiler warning defect detection. Given a seed test program, DIPROM first removes its dead code to reduce false positive warning defects. Then, the abstract syntax tree (AST) of the test program is constructed; DIPROM iteratively mutates the structures of the AST to generate warning-sensitive program variants. To effectively construct diverse warning-sensitive structures, DIPROM applies a novel diversity-guided strategy to generate program variants in each iteration. With the generated program variants, differential testing is conducted to detect warning defects in different compilers. In the experiments, we evaluate DIPROM with two popular C compilers (i.e., GCC and Clang). Experimental results show that DIPROM significantly outperforms three state-of-the-art approaches (i.e., HiCOND, Epiphron, and Hermes) by up to 18.93%$\sim$76.74% in terms of the bug-finding capability on average. Meanwhile, DIPROM is efficient, which spends less time on finding the same average number of warning defects. We at last applied DIPROM to the latest development versions of GCC and Clang. After two months’ running, we reported 8 new warning defects; 5 of them have been confirmed/fixed by developers. He Jiang 0001, Zhide Zhou, Zhilei Ren, Weiqiang Kong |
IEEE Trans. Software Eng. | 3 |
| 2021 | An improved algorithm for accelerating reconfiguration of VLSI array
Junyan Qian, Fuhao Mo, Hao Ding 0007, Zhide Zhou, Lingzhong Zhao, Zhongyi Zhai |
Integr. | 4 |
| 2021 | An empirical study of optimization bugs in GCC and LLVM
Zhide Zhou, Zhilei Ren, Guojun Gao, He Jiang 0001 |
J. Syst. Softw. | 1 |
| 2020 | An efficient multiple shortest augmenting paths algorithm for constructing high performance VLSI subarray
Junyan Qian, Bisheng Huang, Hao Ding 0007, Zhide Zhou, Lingzhong Zhao, Zhongyi Zhai |
Integr. | 4 |
| 2020 | Efficient Reconfiguration Algorithm With Flexible Rerouting Schemes for Constructing 3-D VLSI SubarraysabstractIn this paper, we investigated the technique for improving the reliability of 3-D processor with faults by reconfiguring a 3-D fault-free subarray utilizing as many nonfaulty process elements (PEs) as possible. A novel flexible rerouting scheme is proposed, which makes the PEs can be rerouted or bypassed in three dimensions, hence increasing the number of neighbors of each element to construct a logical array. Under this scheme, an efficient heuristic algorithm is presented to construct a logical array. The experimental results show that the proposed algorithm under flexible rerouting scheme can produce logical arrays with higher harvest from the host arrays with faults for the random fault scenarios, the improvement is by up to 46.47% compared to the state-of-the-arts. Junyan Qian, Hao Ding 0007, Hanpeng Xiao, Zhide Zhou, Lingzhong Zhao, Zhongyi Zhai |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2016 | An Improved Reconfiguration Algorithm for VLSI Arrays with A-Star
Junyan Qian, Zhide Zhou, Lingzhong Zhao, Tianlong Gu |
ICCSA (2) | 2 |
| 2016 | Optimal Reconfiguration of High-Performance VLSI Subarrays with Network FlowabstractA two-dimensional mesh-connected processor array is an extensively investigated architecture used in parallel processing. Massive studies have addressed the use of reconfiguration algorithms for the processor arrays with faults. However, the subarray generated by previous algorithms contains a large number of long interconnects, which in turn leads to more communication costs, capacitance and dynamic power dissipation. In this paper, we propose novel techniques, making use of the idea of network flow, to construct the high-performance subarray, which has the minimum number of long interconnects. First, we construct a network flow model according to the host array under a specific constraint. Second, we show that the reconfiguration problem of high-performance subarray can be optimally solved in polynomial time by using efficient minimum-cost flow algorithms. Finally, we prove that the geometric properties of the resulted subarray meet the system requirements. Simulations based on several random and clustered fault scenarios clearly reveal the advantage of the proposed technique for reducing the number of long interconnects. It is shown that, for a host array of size 512 × 512, the number of long interconnects in the subarray can be reduced by up to 70.05 percent for clustered faults and by up to 55.28 percent for random faults with density of 1 percent as compared to the-state-of-the-art. Junyan Qian, Zhide Zhou, Tianlong Gu, Lingzhong Zhao, Liang Chang 0003 |
IEEE Trans. Parallel Distributed Syst. | 2 |