VLDB 2026 Research / reviewers in the wild / expert
Deheng Yang
dblp:208/7008
· DBLP profile ↗
22ranked-venue papers
8as first author
19since 2021 · last 2026
0000-0002-8383-1939ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 1 first-author · 11 since 2021Software engineering, systems software and programming languages · 9 · 6 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Adaptive Data Augmentation with Bayesian Optimization for Basic Block Throughput Prediction
Xiabing Hu, Xuezheng Xu, Deheng Yang, Chun Huang 0006 |
Euro-Par (1) | 3 |
| 2026 | SRepair: Symbolic Regression-Based Repair for Hardware Design CodeabstractFixing bugs in hardware design code has become a challenging task due to the increasing complexity of modern circuit designs. As a result, automated program repair techniques have been proposed to synthesize patches for bugs in hardware designs and achieved promising results. However, existing techniques are still limited in synthesizing expressions for complex bugs. In this work, we explore the possibility of addressing complex bugs by proposing SREPAIR, a novel symbolic regression-based repair technique. The key novelty of SREPAIR lies in three aspects: 1) we propose a novel expression modification encoding that enables fine-grained adjustments to buggy expressions. 2) we introduce expression synthesis-based templates that allow for flexible and expressive repairs. 3) we develop a novel symbolic regression network-based synthesis algorithm that effectively synthesizes complex expressions. Experimental results on the four peer-reviewed datasets demonstrate that SREPAIR correctly fixes 56 bugs out of 112 bugs, which achieves 43.6% and 194.7% improvement over the previous state-of-the-art RTL-REPAIR (39 bugs) and CIRFIX (19 bugs). To evaluate the generalizability of SREPAIR, we further construct an augmented dataset of 282 bugs by mutating hardware designs. SREPAIR shows its better generalizability by correctly fixing 127 bugs, reaching 217.5% improvement over the best approach. Zizhen Liu, Deheng Yang, Xiaoguang Mao, Jiayu He, Guangda Zhang, Yan Lei 0005, Jiang Wu 0017 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2025 | Fine-Grained Global Search for Inputs Triggering Floating-Point Exceptions in Gpu ProgramsabstractFloating-point exceptions are hard to avoid and can cause disastrous consequences. However, testing methods for floating-point exceptions in GPU programs are currently quite limited due to their closed-source nature. Existing tools, even the state-of-the-art Xscope, still exhibit low search efficiency and poor input coverage. In this paper, we combine interval-wise random sampling and Markov Chain Monte Carlo (MCMC) sampling in a synergistic way to efficiently detect exception-inducing inputs in GPU programs. To improve the search efficiency, based on the bit patterns of exceptional floating-point values, we propose a floating-point format-aware input space partitioning method for random sampling and define a unified fitness function for MCMC sampling. We implement our approach in a tool DFEG and demonstrate it on 76 functions from the CUDA Math Library, HPC programs, and FPBench. DFEG outperforms Xscope in terms of both effectiveness and efficiency. DFEG finds$949 \times$more exceptions than Xscope and detects new exceptions in 9 functions where Xscope fails. Moreover, compared to Xscope, DFEG achieves an average$34 \times$speedup. Xin Yi 0002, Hengbiao Yu, Liqian Chen, Xiaoguang Mao, Ji Wang 0001, Chun Huang 0006, Deheng Yang |
IPDPS | 7 |
| 2025 | Rtl design flaws revisited: a data-driven study of systematic bug patterns in Verilog code
Xiankai Meng, Guangda Zhang, Jiayu He, Deheng Yang, Fangshu Chen, Chengcheng Yu, Xinlin Zhao, Jiang Wu 0017 |
J. Supercomput. | 5 |
| 2024 | Simple but Powerful Beginning: Metamorphic Verification Framework for Cryptographic Hardware DesignabstractComplexity of cryptographic algorithm renders verification of cryptographic hardware design vulnerable to the oracle problem. We propose Minoan: the first opensource verification framework based on Metamorphic testing for cryptographic hardware design to mitigate the oracle problem. Minoan constructs six metamorphic relationships for cryptographic hardware design verification based on domain knowledge, and further designs the time-aware metamorphic relationship satisfiability checking mechanism to strengthen the integration of metamorphic testing with cryptographic hardware. Finally, the evaluation on public datasets from OpenCores shows that Minoan achieves promising results with detecting up to ${9 8 . 0 2 \%}$ bugs. Jiang Wu 0017, Jiayu He, Deheng Yang, Xiaoguang Mao |
ICPADS | 4 |
| 2024 | Labelrepair: Sequence Labelling for Compilation Errors RepairabstractManual fixing of compilation errors could be a tedious and time-consuming task for novice programmers, and even for experienced ones. In recent years, an increasing number of automated repair techniques have been proposed to guide novice programmers and improve the efficiency of software development. Among them, learning-based automated repair techniques have achieved promising results in terms of repair accuracy. However, existing approaches neglect the time efficiency of patch generation, and often treat the compilation errors repair as a neural machine translation task. The end-to-end repair model decoding cannot be parallelized during the inference stage and suffers from redundant decoding search space. Furthermore, the large search space brought by the model poses a potential risk of semantic tampering. To this end, we propose Labelrepair, a novel repair technique that treats the repair of compilation errors as a sequence labelling task. Labelrepair discards the decoding model and converts the search for patches to a search for the error mapping actions between broken code and patch code pairs. In this way, the search space for patch tokens is plum-meted from the entire vocabulary to the size of edit action labels, and the logical semantics of the original program are preserved to some extent. The time complexity of inference is reduced from$O(n)$to$O$(1) owing to the parallel generation of edit action labels. Through a comprehensive evaluation of Labelrepair on two datasets, we demonstrate that Labelrepair is able to generate patches instantly (0.71ms on average), which is 28 times faster than existing end-to-end repair models. Compared with existing edit-based repair models, Labelrepair achieves the state-of-art repair accuracy. Deheng Yang, Yan Lei 0005, Huan Xie 0002, Minghua Tang, Maojin Li |
SANER | 2 |
| 2024 | Demystifying API misuses in deep learning applications
Deheng Yang, Kui Liu 0001, Yan Lei 0005, Li Li 0029, Huan Xie 0002, Xiaoguang Mao, Tegawendé F. Bissyandé |
Empir. Softw. Eng. | 1 |
| 2024 | ContextAug: model-domain failing test augmentation with contextual information
Zhuo Zhang 0007, Jianxin Xue, Deheng Yang, Xiaoguang Mao |
Frontiers Comput. Sci. | 3 |
| 2024 | Knowledge-Augmented Mutation-Based Bug Localization for Hardware Design CodeabstractVerification of hardware design code is crucial for the quality assurance of hardware products. Being an indispensable part of verification, localizing bugs in the hardware design code is significant for hardware development but is often regarded as a notoriously difficult and time-consuming task. Thus, automated bug localization techniques that could assist manual debugging have attracted much attention in the hardware community. However, existing approaches are hampered by the challenge of achieving both demanding bug localization accuracy and facile automation in a single method. Simulation-based methods are fully automated but have limited localization accuracy, slice-based techniques can only give an approximate range of the presence of bugs, and spectrum-based techniques can also only yield a reference value for the likelihood that a statement is buggy. Furthermore, formula-based bug localization techniques suffer from the complexity of combinatorial explosion for automated application in industrial large-scale hardware designs. In this work, we propose Kummel, a K nowledge-a u g m ented m utation-bas e d bug loca l ization for hardware design code to address these limitations. Kummel achieves the unity of precise bug localization and full automation by utilizing the knowledge augmentation through mutation analysis. To evaluate the effectiveness of Kummel, we conduct large-scale experiments on 76 versions of 17 hardware projects by seven state-of-the-art bug localization techniques. The experimental results clearly show that Kummel is statistically more effective than baselines, e.g., our approach can improve the seven original methods by 64.48% on average under the RImp metric. It brings fresh insights of hardware bug localization to the community. Jiang Wu 0017, Zhuo Zhang 0007, Deheng Yang, Jiayu He, Xiaoguang Mao |
ACM Trans. Archit. Code Optim. | 3 |
| 2024 | Time-Aware Spectrum-Based Bug Localization for Hardware Design Code with Data PurificationabstractThe verification of hardware design code is a critical aspect in ensuring the quality and reliability of hardware products. Finding bugs in hardware design code is important for hardware development and is frequently considered as a notoriously challenging and time-consuming activity while being an essential aspect of verification. Thus, bug localization techniques that could assist manual debugging have attracted much attention in the hardware community. However, there exists an unpredictable time span between the precise origin of a bug and its detected manifestation in prior work without costly formal verification. Locating the bug responsible for the exposed discrepancy between expected and exhibited design behavior remains a major challenge. In this work, we propose Tartan, a T ime- a ware spect r um-based bug localiza t ion with d a ta purificatio n for hardware design code to address these limitations. Tartan integrates hardware-specific timing information with the spectrum and captures the changes of executed statements when the state of the circuit changes to effectively locate bugs. Further, Tartan purifies the spectrum data from the simulation and evaluates the suspiciousness of the statements in the design to indicate the likelihood of being buggy. To evaluate the effectiveness of Tartan, we conduct large-scale experiments on 69 versions of 15 hardware projects by the state-of-the-art bug localization techniques. The experimental results clearly show that Tartan is statistically more effective than the baselines. It provides a new perspective on hardware design code bug localization and brings fresh insights to the community. Jiang Wu 0017, Zhuo Zhang 0007, Deheng Yang, Jiayu He, Xiaoguang Mao |
ACM Trans. Archit. Code Optim. | 3 |
| 2024 | Strider: Signal Value Transition-Guided Defect Repair for HDL Programming AssignmentsabstractHardware description languages (HDLs) are pivotal for the development of hardware designs. The programming courses for HDLs are also popular in both universities and online course platforms. Similar to programming assignments of software languages (SLs), these of HDLs also actively call for automated program repair (APR) techniques to provide personalized feedback for students. However, the research of APR techniques targeting HDL programming assignments is still in an early stage. Due to the significantly different programming mechanism of HDLs from SLs, the only APR technique (i.e., CirFix) targeting HDL programming assignments contributes a customized repair pipeline. However, the fundamental challenges in the design of HDL-oriented fault localization and patch generation still remain unresolved. In this work, we propose a signal value transition-guided defect repair technique named STRIDER by capturing the intrinsic features of HDLs. This technique consists of a time-aware dynamic defect localization approach to precisely localize defects, and a signal value transition-guided patch synthesis approach to effectively generate fixes.We further construct a dataset of 57 real defects from HDL programming assignments for tool evaluation. The evaluation reveals the overfitting issue of the pioneering tool CirFix and the significant improvement of STRIDER over CirFix in terms of both effectiveness and efficiency. In particular, STRIDER is more effective by correctly fixing 2.3X as many defects as CirFix in the real defect dataset, and is 23X more efficient by generating a correct fix within five minutes on average in the synthetic defect dataset, while CirFix takes around two hours on average. Deheng Yang, Jiayu He, Xiaoguang Mao, Tun Li 0002, Yan Lei 0005, Xin Yi 0002, Jiang Wu 0017 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2023 | Mantra: Mutation Testing of Hardware Design Code Based on Real BugsabstractMutation testing, a well-suited technology for functional validation, is regrettably poorly studied in hardware. We propose Mantra: the first open-source code-level mutation testing tool based on real hardware bugs. Specifically, Mantra devises time-aware mutation killing mechanism for cost reduction of hardware mutation testing using the parallelism of hardware design code, and then defines and implements 19 hardware mutation operators via large-scale empirical analysis on real bugs. Finally, the evaluation on public datasets from CirFix and OpenCores shows that Mantra achieves promising results with a maximum boost of 83.44%. Jiang Wu 0017, Yan Lei 0005, Zhuo Zhang 0007, Xiankai Meng, Deheng Yang, Jiayu He, Xiaoguang Mao |
DAC | 5 |
| 2023 | Validating the Redundancy Assumption for HDL from Code Clone's PerspectiveabstractAutomated program repair (APR) is being leveraged in hardware description languages (HDLs) to fix hardware bugs without human involvement. Most existing APR techniques search for donor code (i.e., code fragment for bug fixing) in the original program to generate repairs, which is based on the assumption that donor code can be found in existing source code. The redundancy assumption is the fundamental basis of most APR techniques, which has been widely studied in software by searching code clones of donor code. However, despite a large body of work on code clone detection, researchers have focused almost exclusively on repositories in traditional programming languages, such as C/C++ and Java, while few studies have been done on detecting code clones in HDLs. Furthermore, little attention has been paid on the repetitiveness of bug fixes in hardware designs, which limits automatic repair targeting HDLs. To validate the redundancy assumption for HDL, we perform an empirical study on code clones of real-world bug fixes in Verilog. On top of empirical results, we find that 17.71% of newly introduced code in bug fixes can be found from the clone pairs of buggy code in the original program, and 11.77% can be found in the file itself. The findings not only validate the assumption but also provides helpful insights for the design of APR targeting HDLs. Jiayu He, Deheng Yang, Jiang Wu 0017, Xiaoguang Mao |
ISPD | 4 |
| 2023 | Seeing the Whole Elephant: Systematically Understanding and Uncovering Evaluation Biases in Automated Program RepairabstractEvaluation is the foundation of automated program repair (APR), as it provides empirical evidence on strengths and weaknesses of APR techniques. However, the reliability of such evaluation is often threatened by various introduced biases. Consequently, bias exploration, which uncovers biases in the APR evaluation, has become a pivotal activity and performed since the early years when pioneer APR techniques were proposed. Unfortunately, there is still no methodology to support a systematic comprehension and discovery of evaluation biases in APR, which impedes the mitigation of such biases and threatens the evaluation of APR techniques. In this work, we propose to systematically understand existing evaluation biases by rigorously conducting the first systematic literature review on existing known biases and systematically uncover new biases by building a taxonomy that categorizes evaluation biases. As a result, we identify 17 investigated biases and uncover a new bias in the usage of patch validation strategies. To validate this new bias, we devise and implement an executable framework APRConfig , based on which we evaluate three typical patch validation strategies with four representative heuristic-based and constraint-based APR techniques on three bug datasets. Overall, this article distills 13 findings for bias understanding, discovery, and validation. The systematic exploration we performed and the open source executable framework we proposed in this article provide new insights as well as an infrastructure for future exploration and mitigation of biases in APR evaluation. Deheng Yang, Yan Lei 0005, Xiaoguang Mao, Yuhua Qi, Xin Yi 0002 |
ACM Trans. Softw. Eng. Methodol. | 1 |
| 2022 | Fault Localization for Hardware Design Code with Time-Aware Program SpectrumabstractVerification of hardware design code is crucial for the quality assurance of hardware products. As an indispensable part of verification, localizing faults in the hardware design code is significant for hardware development but is often regarded as a notoriously difficult and time-consuming task. Thus, automated fault localization techniques that could assist manual debugging have attracted much attention in the hardware community. Prior work indicates that existing methods neither fully utilize program dynamic execution information nor lack attention to timing. In this work, we propose Tarsel: a time-aware spectrum-based fault localization approach to help bridge this gap. Tarsel integrates hardware-specific timing information with the program spectrum and captures the changes of executed statements when the state of the hardware program changes to effectively locate faults. The experimental results show that Tarsel successfully locates over half of bugs in the benchmark at Top-3 and about 90% of bugs at Top-5. In addition, Tarsel statistically outperforms the state-of-the-art fault localization approach CirFix under all six typical metrics. In particular, while no bugs are ranked at Top-1 by CirFix, Tarsel successfully locates 11.41% of bugs at Top-1. It brings fresh insights of hardware bug localization to the community. Jiang Wu 0017, Zhuo Zhang 0007, Deheng Yang, Xiankai Meng, Jiayu He, Xiaoguang Mao, Yan Lei 0005 |
ICCD | 3 |
| 2022 | TransplantFix: Graph Differencing-based Code Transplantation for Automated Program RepairabstractAutomated program repair (APR) holds the promise of aiding manual debugging activities. Over a decade of evolution, a broad range of APR techniques have been proposed and evaluated on a set of real-world bug datasets. However, while more and more bugs have been correctly fixed, we observe that the growth of newly fixed bugs by APR techniques has hit a bottleneck in recent years. In this work, we explore the possibility of addressing complicated bugs by proposing TransplantFix, a novel APR technique that leverages graph differencing-based transplantation from the donor method. The key novelty of TransplantFix lies in three aspects: 1) we propose to use a graph-based differencing algorithm to distill semantic fix actions from the donor method; 2) we devise an inheritance-hierarchy-aware code search approach to identify donor methods with similar functionality; 3) we present a namespace transfer approach to effectively adapt donor code. Deheng Yang, Xiaoguang Mao, Liqian Chen, Xuezheng Xu, Yan Lei 0005, David Lo 0001, Jiayu He |
ASE | 1 |
| 2021 | Is the Ground Truth Really Accurate? Dataset Purification for Automated Program RepairabstractDatasets of real-world bugs shipped with human-written patches are intensively used in the evaluation of existing automated program repair (APR) techniques, wherein the human-written patches always serve as the ground truth, for manual or automated assessment approaches, to evaluate the correctness of test-suite adequate patches. An inaccurate human-written patch tangled with other code changes will pose threats to the reliability of the assessment results. Therefore, the construction of such datasets always requires much manual effort on isolating real bug fixes from bug fixing commits. However, the manual work is time-consuming and prone to mistakes, and little has been known on whether the ground truth in such datasets is really accurate.In this paper, we propose DEPTEST, an automated DatasEt Purification technique from the perspective of triggering Tests. Leveraging coverage analysis and delta debugging, DEPTEST can automatically identify and filter out the code changes irrelevant to the bug exposed by triggering tests. To measure the strength of DEPTEST, we run it on the most extensively used dataset (i.e., Defects4J) that claims to already exclude all irrelevant code changes for each bug fix via manual purification. Our experiment indicates that even in a dataset where the bug fix is claimed to be well isolated, 41.01% of human-written patches can be further reduced by 4.3 lines on average, with the largest reduction reaching up to 53 lines. This indicates its great potential in assisting in the construction of datasets of accurate bug fixes. Furthermore, based on the purified patches, we re-dissect Defects4J and systematically revisit the APR of multi-chunk bugs to provide insights for future research targeting such bugs. Deheng Yang, Yan Lei 0005, Xiaoguang Mao, David Lo 0001, Huan Xie 0002, Meng Yan 0001 |
SANER | 1 |
| 2021 | Where were the repair ingredients for Defects4j bugs?
Deheng Yang, Kui Liu 0001, Dongsun Kim 0001, Anil Koyuncu, Kisub Kim, Haoye Tian, Yan Lei 0005, Xiaoguang Mao, Jacques Klein, Tegawendé F. Bissyandé |
Empir. Softw. Eng. | 1 |
| 2021 | Evaluating the usage of fault localization in automated program repair: an empirical study
Deheng Yang, Yuhua Qi, Xiaoguang Mao |
Frontiers Comput. Sci. | 1 |
| 2019 | Attention Please: Consider Mockito when Evaluating Newly Proposed Automated Program Repair TechniquesabstractAutomated program repair (APR) has attracted widespread attention in recent years with substantial techniques being proposed. Meanwhile, a number of benchmarks have been established for evaluating the performances of APR techniques, among which Defects4J is one of the most widely used benchmark. However, bugs in Mockito, a project augmented in a later-version of Defects4J, do not receive much attention by recent researches. In this paper, we aim at investigating the necessity of considering Mockito bugs when evaluating APR techniques. Our findings show that: 1) Mockito bugs are not more complex for repairing compared with bugs from non-Mockito projects; 2) the bugs repaired by the state-of-the-art tools share the same repair patterns compared with those patterns required to repair Mockito bugs; however, 3) the state-of-the-art tools perform poorly on Mockito bugs (Nopol can only correctly fix one bug while SimFix and CapGen cannot fix any bug in Mockito even if all the buggy locations have been exposed). We conclude from these results that existing APR techniques may be overfitting to their evaluated subjects and we should consider Mockito, or even more bugs from other projects, when evaluating newly proposed APR techniques. Shangwen Wang, Ming Wen 0001, Xiaoguang Mao, Deheng Yang |
EASE | 4 |
| 2018 | An Empirical Study on the Effect of Dynamic Slicing on Automated Program Repair EfficiencyabstractResearch on the characteristics of error propagation can guide fault localization more efficiently. Spectrum-based fault localization (SFL) and slice-based fault localization are effective fault localization techniques. The former produces a list of statements in descending order of suspicious values, and the latter generates statements that affect failure statements. We propose a new dynamic slicing and spectrum-based fault localization (DSFL) method, which combines the list of suspicious statements generated by SFL with dynamic slicing, and take the characteristics of error propagation into account. To the best of our knowledge, DSFL has not yet been implemented in automated repair tools. In this study, we use the dynamic slicing tool Javaslicer to determine the error propagation chain of faulty programs and the statements related to failure execution. We implement the DSFL algorithm in the automated repair tool Nopol and conduct repair experiments on dataset Defects4j to compare the effects of SFL and DSFL on the efficiency of automated repair. Preliminary results indicate that the scope of error propagation for most programs is a single class, and the DSFL makes automated repair more efficient. Anbang Guo, Xiaoguang Mao, Deheng Yang, Shangwen Wang |
ICSME | 3 |
| 2017 | An Empirical Study on the Usage of Fault Localization in Automated Program RepairabstractSpectrum-based fault localization (SFL), the technique producing a rank list of statements in descending order of their suspiciousness values, is nowadays widely used in current automated program repair tools. There are two different algorithms for these tools to choose statements selected for modification to produce candidate patches from the list: one is the rank-first algorithm based on suspiciousness rankings of statements, the other is the suspiciousness-first algorithm based on suspiciousness value of statements. However, to our knowledge there is no research work implementing the two algorithms in the same repair tool or comparing their effectiveness. In this paper, we conduct an empirical research based on the automated repair tool Nopol with the benchmark set of Defects4J to compare these two algorithms. Preliminary results suggest that the suspiciousness-first algorithm is not equivalent to the rank-first algorithm and behaves better in parallel repair and patch diversity. Deheng Yang, Yuhua Qi, Xiaoguang Mao |
ICSME | 1 |