Lehuan Zhang

dblp:369/7852 · DBLP profile ↗
← Back
3ranked-venue papers
2as first author
3since 2021 · last 2025
0009-0009-0908-9295ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Live Region Mutation Testing for Commercial Cyber-Physical System Development Tool Chain
abstract
MathWorks Simulink, a commercial CPS development tool chain, is widely used as an industry standard for designing and analyzing system behavior and generating embedded code for deployment. However, bugs in Simulink can cause unexpected behaviors during model compilation, making their elimination critical. Existing methods face two key challenges: generating equivalent models with varied data flows (data flow equivalence) and creating diverse block types to comprehensively test the compiler (mutation diversity). To address these, we propose LION, a differential testing approach. LION ensures data flow equivalence by inserting “store-revert” block pairs between existing blocks and tackles mutation diversity by employing Markov Chain Monte Carlo (MCMC) sampling to generate diverse new blocks. Differential testing is then used to identify bugs. Experiments show LION outperforms state-of-the-art approaches like SLforge, SLEMI, and COMBAT, detecting 610 additional compiler bugs in two weeks. Over two months, LION uncovered and reported 16 valid bugs in the widely used stable version of Simulink.
Lehuan Zhang, Shikai Guo, He Jiang 0001
DAC1
2025 Context-based Transfer Learning for Structuring Fault Localization and Program Repair Automation
abstract
Automated software debugging plays a crucial role in aiding software developers to swiftly identify and attempt to rectify faults, thereby significantly reducing developers’ workload. Previous researches have predominantly relied on simplistic semantic deep learning or statistical analysis methods to locate faulty statements in diverse projects. However, code repositories often consist of lengthy sequences with long-distance dependencies, posing challenges for accurately modeling fault localization using these methods. In addition, the lack of joint reasoning among various faults prevents existing models from deeply capturing fault information. To address these challenges, we propose a method named CodeHealer to achieve accurate fault localization and program repair. CodeHealer comprises three components: a Deep Semantic Information Extraction Component that effectively extracts deep semantic features from suspicious code statements using classifiers based on Joint-attention mechanisms; a Suspicious Statement Ranking Component that combines various fault localization features and employs multilayer perceptrons to derive multidimensional vectors of suspicion values; and a Fault Repair Component that, based on ranked suspicious statements generated by fault localization, adopts a top-down approach using multiple classifiers based on Co-teaching mechanisms to select repair templates and generate patches. The experimental results indicate that when applied to fault localization, CodeHealer outperforms the best baseline method with improvements of 11.4%, 2.7%, and 1.6% on Top-1/3/5 metrics, respectively. It also reduces the MFR and MAR by 9.8% and 2.1%, where lower values denote better fault localization effectiveness. Additionally, in automated software debugging, CodeHealer fixes an additional 6 faults compared to the current best method, totaling 53 faults repaired.
Lehuan Zhang, Shikai Guo, Hui Li 0014, Yu Chai, Rong Chen 0003, He Jiang 0001
ACM Trans. Softw. Eng. Methodol.1
2024 Context-based transfer learning for low resource code summarization
abstract
Abstract Source code summaries improve the readability and intelligibility of code, help developers understand programs, and improve the efficiency of software maintenance and upgrade processes. Unfortunately, these code comments are often mismatched, missing, or outdated in software projects, resulting in developers needing to infer functionality from source code, affecting the efficiency of software maintenance and evolution. Various methods based on neuronal networks are proposed to solve the problem of synthesis of source code. However, the current work is being carried out on resource‐rich programming languages such as Java and Python, and some low‐resource languages may not perform well. In order to solve the above challenges, we propose a context‐based transfer learning model for low resource code summarization (LRCS), which learns the common information from the language with rich resources, and then transfers it to the target language model for further learning. It consists of two components: the summary generation component is used to learn the syntactic and semantic information of the code, and the learning transfer component is used to improve the generalization ability of the model in the learning process of cross‐language code summarization. Experimental results show that LRCS outperforms baseline methods in code summarization in terms of sentence‐level BLEU, corpus‐level BLEU and METEOR. For example, LRCS improves corpus‐level BLEU scores by 52.90%, 41.10%, and 14.97%, respectively, compared to baseline methods.
Yu Chai, Lehuan Zhang, Hui Li 0014, Mengzhi Luo, Shikai Guo
Softw. Pract. Exp.3