Yuekang Li

dblp:204/3729 · DBLP profile ↗
← Back
71ranked-venue papers
6as first author
59since 2021 · last 2026
0000-0003-4382-0757ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 33 · 5 first-author · 24 since 2021Security and privacy · 20 · 18 since 2021Artificial intelligence and machine learning · 11 · 1 first-author · 10 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 since 2021Human-computer interaction and ubiquitous computing · 4 · 4 since 2021Systems, architecture and hardware · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 SceneJailEval: A Scenario-Adaptive Multi-Dimensional Framework for Jailbreak Evaluation
abstract
Accurate jailbreak evaluation is critical for LLM red team testing and jailbreak research. Mainstream methods rely on binary classification (string matching, toxic text classifiers, and LLM-based methods), outputting only "yes/no" labels without quantifying harm severity. Emerged multi-dimensional frameworks (e.g., Security Violation, Relative Truthfulness and Informativeness) use unified evaluation standards across scenarios, leading to scenario-specific mismatches (e.g., "Relative Truthfulness" is irrelevant to "hate speech"), undermining evaluation accuracy. To address these, we propose SceneJailEval, with key contributions: (1) A pioneering scenario-adaptive multi-dimensional framework for jailbreak evaluation, overcoming the critical "one-size-fits-all" limitation of existing multi-dimensional methods, and boasting robust extensibility to seamlessly adapt to customized or emerging scenarios. (2) A novel 14-scenario dataset featuring rich jailbreak variants and regional cases, addressing the long-standing gap in high-quality, comprehensive benchmarks for scenario-adaptive evaluation. (3) SceneJailEval delivers state-of-the-art performance with an F1 score of 0.917 on our full-scenario dataset (+6% over SOTA) and 0.995 on JBB (+3% over SOTA), breaking through the accuracy bottleneck of existing evaluation methods in heterogeneous scenarios and solidifying its superiority.
Yuekang Li, Youtao Ding, Li Pan 0002
AAAI2
2026 Socrates or Smartypants: Testing Logic Reasoning Capabilities of Large Language Models with Logic Programming-Based Test Oracles
abstract
Large Language Models (LLMs) have achieved significant progress in language understanding and reasoning. Evaluating and analyzing their logical reasoning abilities has therefore become essential. However, existing datasets and benchmarks are often limited to overly simplistic, unnatural, or contextually constrained examples. In response to the growing demand, we introduce SMARTYPAT-BENCH, a challenging, naturally expressed, and systematically labeled benchmark derived from real-world high-quality Reddit posts containing subtle logical fallacies. Unlike existing datasets and benchmarks, it provides more detailed annotations of logical fallacies and features more diverse data. To further scale up the study and address the limitations of manual data collection and labeling, such as fallacy-type imbalance and labor-intensive annotation, we introduce SMARTYPAT, an automated framework powered by logic programming-based oracles. SMARTYPAT utilizes Prolog rules to systematically generate logically fallacious statements, which are then refined into fluent natural language sentences by LLMs, ensuring precise fallacy rep- resentation. Extensive evaluation demonstrates that SMARTYPAT produces fallacies comparable in subtlety and quality to human-generated content and significantly outperforms baseline methods. Finally, experiments reveal insights into LLM capabilities, highlighting that while excessive reasoning steps hinder fallacy detection accuracy, structured reasoning enhances fallacy categorization performance.
Junchen Ding, Yiling Lou, Dong Gong, Yuekang Li
AAAI6
2026 STEAMROLLER: A Multi-Agent System for Inclusive Automatic Speech Recognition for People Who Stutter
abstract
People who stutter (PWS) face systemic exclusion in today’s voice-driven society, where access to voice assistants, authentication systems, and remote work tools increasingly depends on fluent speech. Current automatic speech recognition (ASR) systems, trained predominantly on fluent speech, fail to serve millions of PWS worldwide. We present STEAMROLLER, a real time system that transforms stuttered speech into fluent output through a novel multi-stage, multi-agent AI pipeline. Our approach addresses three critical technical challenges: (1) the difficulty of direct speech to speech conversion for disfluent input, (2) semantic distortions introduced during ASR transcription of stuttered speech, and (3) latency constraints for real time communication. STEAMROLLER employs a three stage architecture comprising ASR transcription, multi-agent text repair, and speech synthesis, where our core innovation lies in a collaborative multi-agent framework that iteratively refines transcripts while preserving semantic intent. Experiments on the FluencyBank dataset and a user study demonstrates clear word error rate (WER) reduction and strong user satisfaction. Beyond immediate accessibility benefits, fine tuning ASR on STEAMROLLER repaired speech further yields additional WER improvements, creating a pathway toward inclusive AI ecosystems.
Yi Liu 0069, Yuekang Li, Ling Shi 0002, Kailong Wang 0001
AAAI3
2026 Prompt to Pwn: Automated Exploit Generation for Smart Contracts
ZeKe Xiao, Qin Wang 0008, Yuekang Li, Shiping Chen 0001
ACISP (1)3
2026 LifeFuzz: Lifecycle-Guided Fuzzing for Windows Driver Cross-Handler Vulnerabilities
abstract
Third-party Windows drivers expose a critical attack surface. However, vulnerabilities that require cross-handler I/O Control (IOCTL) sequences remain hard to find, despite their prevalence, and often lead to privilege escalation. Static analysis suffers from high false positives, path explosion, and complex resource modeling. Meanwhile, dynamic fuzzers often exercise handlers in isolation or combine them randomly, leaving implicit state dependencies unchecked. To address the gap, we present LifeFuzz, a lifecycle-guided fuzzing framework that models global-variable lifecycles to construct dependency-respecting IOCTL sequences. Specifically, it identifies variable operations across handlers, preserves seeds that affect driver state, and then combines them into meaningful sequences. Consequently, LifeFuzz explores deep paths unreachable for existing fuzzers. We evaluate LifeFuzz on 26 Windows WDM drivers. It discovers 86 vulnerabilities, including 32 cross-handler cases, with six assigned CVE IDs. Moreover, it finds 357% more cross-handler vulnerabilities than msFuzz and achieves 19.2% higher average coverage. Overall, 37% of discovered vulnerabilities require cross-handler interactions, thereby validating lifecycle-aware, cross-handler fuzzing for driver security.
Chendong Yu, Yuekang Li, Yang Xiao 0011, Jie Lu 0009, Yeting Li, Defang Bo, Wei Huo 0005
EuroSys2
2026 Uncovering Logit Suppression Vulnerabilities in LLM Safety Alignment
Yi Liu 0069, Yuekang Li, Ling Shi 0002, Gelei Deng, Shengquan Chen, Kailong Wang 0001
ICPR (2)3
2026 Holmes: Forensic-Aware Design and Evaluation of Large Language Models for Procedurally Constrained Digital Investigations
Khalid Farhan, Yuekang Li, Jingling Xue
KSEM (4)2
2026 MUTATO: Enhancing Fuzz Drivers with Adaptive API Option Mutation
Shuangxiang Kan, Yuekang Li
NDSS3
2026 Through the Authentication Maze: Detecting Authentication Bypass Vulnerabilities in Firmware Binaries
Nanyu Zhong, Yuekang Li, Yanyan Zou 0002, Jiaxu Zhao 0004, Jinwei Dong, Yang Xiao 0011, Bingwei Peng, Yeting Li, Wei Huo 0005
NDSS2
2026 FidelityGPT: Correcting Decompilation Distortions with Retrieval Augmented Generation
Zhiping Zhou, Xiaohong Li 0001, Yao Zhang 0019, Yuekang Li, Wenbu Feng, Yunqian Wang
NDSS5
2026 CSTutorBench: Benchmarking Large Language Models for Realistic Computer Science Tutoring
abstract
Large Language Models (LLMs) show promise for CS educational assistance; however, the absence of comprehensive benchmarks limits our ability to assess their effectiveness in real-world teaching scenarios accurately. To fill this gap, we present CSTutorBench, a dataset from authentic course discussion forums with 2,970 multimodal question–answer pairs. Additionally, we propose an evaluation framework across five dimensions—accuracy, clarity, conciseness, personalization, and engagement—to gauge the performance of various models in tutoring settings. We benchmark leading LLMs—including GPT-4o, Claude, Llama 4, and others—using both automated metrics and expert human assessments. Across these real CS tutoring exchanges, we found leading LLMs approach human performance in terms of accuracy and clarity, but fall notably short on personalization and interactive scaffolding, often producing fluent yet less learner-adaptive guidance.
Zekai Cheng, Yunfeng Wan, Daijiao Liu, Yuekang Li, Dong Gong
SIGCSE (2)4
2026 Exploring Trust in Human-LLM Feedback Systems: Observation of Student Behaviour in Software Engineering Education
abstract
AI systems can deliver scalable feedback in large courses and often achieve high consistency with human grading. Yet, students continue to view human feedback as more credible, actionable, and trustworthy, motivating growing interest in hybrid approaches that combine AI and human input. Despite advances in this area, three key challenges remain: (1) understanding when and why students escalate from AI to human support, (2) identifying how design interventions such as transparency, explainability, and endorsement affect trust and uptake, and (3) evaluating whether hybrid AI+human feedback models improve learning outcomes compared to AI-only or human-only systems. This study uses controlled experiments and structured interviews to investigate these questions. The results aim to guide the responsible design of hybrid feedback systems that align scalability with pedagogical credibility.
Yiwen Liao, Madhushi Niluka Bandara, Yuekang Li, Iromie Samarasekara, Zixiu Guo, Drishtant Leuva, George Joukhadar
SIGCSE (2)4
2026 Diverse Claire: An AI-Powered UDL Approach Bridging Educational Experience Gaps in CS1
abstract
Providing equitable and accessible learning materials in computer science can be both challenging and time consuming, particularly for novice teachers or those uncertain about how to implement accessibility effectively. This issue is magnified by the lack of clear, actionable directions and training available for educators to put into practice. Diverse Claire is a tool that addresses this issue by providing teachers with structured, practical feedback on their teaching materials through the use of simulation. It utilises GenAI to analyse the material for accessibility and the Universal Design for Learning (UDL) framework as a guideline for analysis. This demo provides a walkthrough of how Diverse Claire analyses teaching materials, highlights potential accessibility barriers, and offers suggestions for improvements. Moreover, the system allows educators to understand how a computer science student with neurodiverse needs might respond to their course materials, by using a large language model to simulate responses to various test questions. These simulations are grounded in instructions drawn from research on neurodiverse learner experiences. The questions include short answer questions and coding exercises that are aligned with the pedagogical framework, Bloom's taxonomy, to target varying levels of understanding. In doing so, Diverse Claire supports CS educators in creating a more inclusive computer science learning environment that better meets the needs of all students.
Madhu Maya Shrestha, Wendy Wong, Omar Al Zeidat, Yuekang Li
SIGCSE (2)5
2026 DiverseClaire: Simulating Students to Improve Introductory Programming Course Materials for All CS1 Learners
Wendy Wong, Yuekang Li
SIGCSE (2)3
2026 PufferDoS: Efficient and Effective Attack String Generation for Regular Expression Denial of Service Vulnerabilities
Shangzhi Xu, Yuekang Li, Nan Sun 0002, Benjamin Turnbull, Shuangxiang Kan, Siqi Ma 0001
SP4
2026 Determining the Unreachable: Constraint-Guided Reachability Analysis for Dependency Vulnerabilities
abstract
In software development, investigating the accessibility of dependency vulnerabilities is of great importance, as third-party libraries often contain known vulnerabilities that could be exploited in the application's business logic. The existing accessibility analysis methods encounter challenges such as undecidability, abstraction loss, and path explosion in large-scale programs, resulting in an inaccurate distinction between accessibility vulnerabilities and non-accessibility vulnerabilities. This paper introduces an approach called ConVReach for analyzing the reachability of vulnerabilities in dependencies in C/C++ programs. ConVReach overcomes the problems of high abstraction loss and potential path explosion in the current methods by combining static and dynamic approaches, particularly a constraint-guided analysis method. This approach extracts and decomposes the path constraints that trigger vulnerabilities, independently verifies the satisfiability of each constraint, and then aggregates the feasible paths. This effectively reduces unnecessary path exploration and avoids the common path explosion issues in traditional methods. Experimental results show that ConVReach outperforms existing tools in both accuracy and efficiency, effectively distinguishing between reachable and unreachable vulnerabilities, and significantly reducing false positives and false negatives. We constructed a benchmark dataset to evaluate ConVReach , which includes 53 CVEs and 347 flags artificially inserted into various open-source projects. This dataset was designed to simulate both real-world vulnerabilities and complex scenarios. Through testing on this dataset, ConVReach demonstrated exceptional performance. It successfully identified 59 out of 61 reachable vulnerabilities and all 23 unreachable ones in the CVE dataset. Within a 24-hour time budget, ConVReach detected above 50% more reachable vulnerabilities than the baseline tools in the first 6 hours and nearly completed the detection of reachable vulnerabilities by the 12-hour mark. These results highlight ConVReach 's superior ability to handle both real-world vulnerabilities and challenging cases.
Wenbu Feng, Xiaohong Li 0001, Yao Zhang 0019, Yuekang Li, Zhiping Zhou, Yunqian Wang
Proc. ACM Program. Lang.5
2026 OptRCA: A More Efficient and Accurate Approach for Automated Root Cause Analysis and Explanation
abstract
With the development of automated software testing technology, software developers can get a large number of crash test cases in a short period of time. However, analyzing these crash test cases and finding their root cause is a time-consuming and labor-intensive task. Techniques based on reverse execution and backward taint analysis are proposed to locate the root cause, but can’t provide context information or explanation of the underlying fault. To address these two limitations, researchers have proposed an automated root cause analysis technique called AURORA. Although this technique provides powerful root cause analysis capabilities, it also have two obvious shortcomings. First, the results of root cause analysis are not accurate enough. Second, the efficiency of root cause analysis is not high enough. In order to improve these two shortcomings, we propose OptRCA, a more efficient and accurate approach for root cause analysis and explanation. Like AURORA’s fuzzing strategy, OptRCA is also designed based on AFL’s crash mode. The difference between them is mainly reflected in three points. First of all, the goal pursued by OptRCA is different from that of normal fuzzing technology. OptRCA pursues maximum correlation to ensure that as many crash test cases as possible are related to the same root cause. This test case with maximum correlation can greatly improve the accuracy of root cause analysis. Second, OptRCA proposed a more efficient non-crash test case retention strategy, which we named “Hill-Climbing Retention.” Using the hill-climbing retention method, OptRCA can obtain sufficient root cause information while retaining only a few non-crash test cases. Since the number of test cases is greatly reduced, the efficiency of OptRCA’s subsequent root cause analysis process is also greatly improved. In addition, OptRCA also optimizes the analysis formula to obtain more accurate analysis results. In the evaluation experimental results, OptRCA is significantly better than AURORA in terms of accuracy and efficiency. Quantitative analysis shows that OptRCA is 65% more accurate and 61% more efficient than AURORA.
Jingquan Ge, Yaowen Zheng, Yuekang Li, Wei Ma 0014, Sheikh Mahbub Habib, Praveen Kakkolangara, Gabriel Byman, Yang Liu 0003
ACM Trans. Softw. Eng. Methodol.3
2026 Spectre: Automated Aliasing Specification Generation for Library APIs with Fuzzing
abstract
Static program analysis of real-world software that integrates numerous library Application Programming Interfaces (APIs) faces significant challenges due to inaccessible or highly complex source code. A common workaround is to use specifications that summarize the key behaviors of these APIs for analysis. However, manually writing specifications is labor-intensive and requires a deep understanding of API semantics, while existing automated specification generation techniques struggle when source code is inaccessible or partially available. This article introduces Spectre , an automated framework that leverages fuzzing techniques to generate aliasing specifications for library APIs. Spectre operates efficiently and precisely both with and without source code access. When source code is unavailable, Spectre integrates alias-check observers into the driver program after the API call site and performs black-box fuzzing to explore different API behaviors. If a check is satisfied, the corresponding aliasing specification is generated. When source code is available, Spectre incorporates new grey-box fuzzing features specifically tailored for aliasing specification inference, further enhancing its ability to generate aliasing specifications. We conducted extensive experiments to evaluate the performance of Spectre . Without source code access, Spectre demonstrated its specification generation capability across both Musl, a lightweight C standard library, and eight C third-party libraries. For Musl, Spectre recovered 96.7% of correct manually written specifications and identified 40.0% more aliasing specifications than those written by external experts. For C third-party libraries, all Spectre -generated aliasing specifications were validated as correct through static analysis of the API source code. Spectre is also more complete than other specification inference tools, generating 16.7% more correct specifications for third-party libraries. The practicality of the generated specifications was confirmed, as they improved aliasing analysis in static pointer analysis of client code while maintaining a balance between accuracy and efficiency. The effectiveness of the tailored grey-box fuzzing features was demonstrated by Spectre , identifying 20% more specifications compared to when these features were disabled. These results show that Spectre is an effective tool for inferring aliasing specifications and facilitating static analysis.
Shuangxiang Kan, Yuekang Li, Weigang He, Zhenchang Xing, Liming Zhu 0001, Yulei Sui
ACM Trans. Softw. Eng. Methodol.2
2025 It Only Gets Worse: Revisiting DL-Based Vulnerability Detectors from a Practical Perspective
abstract
With the escalating threat of software vulnerabilities to the security of modern software systems, an increasing number of deep learning (DL) model-based vulnerability detectors have been developed for vulnerability detection. However, their practical reliability, consistency in usage, and adaptability across diverse software contexts remain unclear. This uncertainty may lead to unreliable detection results in practical applications, increased false positives and false negatives, and limited adaptability to newly emerged vulnerabilities. Conducting a large-scale and in-depth analysis of DL-based vulnerability detectors can help uncover critical factors influencing detection performance, improve the design and training of these models, and enhance their practical deployment in real-world scenarios. In this paper, we present VulTegra, a novel evaluation framework that, for the first time, conducts a multidimensional assessment comparing scratch-trained models and pre-trained-based models for vulnerability detection, while verifying key factors influencing detection performance. Our framework reveals that state-of-the-art (SOTA) detectors still suffer from low consistency, limited practical detection capabilities, and limited adaptability. Moreover, comparative results indicate that the increasingly favored pre-trained-based models are not universally superior to scratch-trained models; instead, they exhibit distinct strengths and application scenarios. Most importantly, our study highlights the limitations of relying solely on CWE-based classification and reveals a set of critical factors that significantly influence detection performance. Experimental validation shows that these factors have a substantial impact: modifying only any single factor led to recall improvements across all seven evaluated SOTA detectors, with six detectors also achieving higher F1 scores. Our findings provide deep insights into model behavior, highlighting the need to consider both vulnerability types and inherent code features to ensure practical applicability in real-world software environments.
Yunqian Wang, Xiaohong Li 0001, Yao Zhang 0019, Yuekang Li, Zhiping Zhou
APSEC5
2025 TombRaider: Entering the Vault of History to Jailbreak Large Language Models
abstract
Warning: This paper contains content that may involve potentially harmful behaviours, discussed strictly for research purposes.Jailbreak attacks can hinder the safety of Large Language Model (LLM) applications, especially chatbots.Studying jailbreak techniques is an important AI red teaming task for improving the safety of these applications.In this paper, we introduce TOMBRAIDER, a novel jailbreak technique that exploits the ability to store, retrieve, and use historical knowledge of LLMs.TOMBRAIDER employs two agents, the inspector agent to extract relevant historical information and the attacker agent to generate adversarial prompts, enabling effective bypassing of safety filters.We intensively evaluated TOMBRAIDER on six popular models.Experimental results showed that TOMBRAIDER could outperform state-of-the-art jailbreak techniques, achieving nearly 100% attack success rates (ASRs) on bare models and maintaining over 55.4% ASR against defence mechanisms.Our findings highlight critical vulnerabilities in existing LLM safeguards, underscoring the need for more robust safety defences.
Junchen Ding, Yi Liu 0069, Gelei Deng, Yuekang Li
EMNLP6
2025 MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluation
abstract
Weihao Xuan, Rui Yang, Heli Qi, Qingcheng Zeng, Yunze Xiao, Aosong Feng, Dairui Liu, Yun Xing, Junjue Wang, Fan Gao, Jinghui Lu, Yuang Jiang, Huitao Li, Xin Li, Kunyu Yu, Ruihai Dong, Shangding Gu, Yuekang Li, Xiaofei Xie, Felix Juefei-Xu, Foutse Khomh, Osamu Yoshie, Qingyu Chen, Douglas Teodoro, Nan Liu, Randy Goebel, Lei Ma, Edison Marrese-Taylor, Shijian Lu, Yusuke Iwasawa, Yutaka Matsuo, Irene Li. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Weihao Xuan, Rui Yang 0016, Heli Qi, Qingcheng Zeng, Yunze Xiao, Aosong Feng, Dairui Liu, Yun Xing 0001, Jinghui Lu, Yuang Jiang, Huitao Li, Xin Li 0079, Kunyu Yu, Ruihai Dong, Shangding Gu, Yuekang Li, Xiaofei Xie, Felix Juefei-Xu, Foutse Khomh, Osamu Yoshie, Qingyu Chen 0001, Douglas Teodoro, Nan Liu 0003, Randy Goebel, Lei Ma 0003, Edison Marrese-Taylor, Shijian Lu, Yusuke Iwasawa, Yutaka Matsuo, Irene Li
EMNLP18
2025 TransferFuzz: Fuzzing with Historical Trace for Verifying Propagated Vulnerability Code
abstract
Code reuse in software development frequently facilitates the spread of vulnerabilities, making the scope of affected software in CVE reports imprecise. Traditional methods primarily focus on identifying reused vulnerability code within target software, yet they cannot verify if these vulnerabilities can be triggered in new software contexts. This limitation often results in false positives. In this paper, we introduce TransferFuzz, a novel vulnerability verification framework, to verify whether vulnerabilities propagated through code reuse can be triggered in new software. Innovatively, we collected runtime information during the execution or fuzzing of the basic binary (the vulnerable binary detailed in CVE reports). This process allowed us to extract historical traces, which proved instrumental in guiding the fuzzing process for the target binary (the new binary that reused the vulnerable function). TransferFuzz introduces a unique Key Bytes Guided Mutation strategy and a Nested Simulated Annealing algorithm, which transfers these historical traces to implement trace-guided fuzzing on the target binary, facilitating the accurate and efficient verification of the propagated vulnerability. Our evaluation, conducted on widely recognized datasets, shows that TransferFuzz can quickly validate vulnerabilities previously unverifiable with existing techniques. Its verification speed is 2.5 to 26.2 times faster than existing methods. Moreover, TransferFuzz has proven its effectiveness by expanding the impacted software scope for 15 vulnerabilities listed in CVE reports, increasing the number of affected binaries from 15 to 53. The datasets and source code used in this article are available at https://github.com/Siyuan-Li201/TransferFuzz.
Siyuan Li 0014, Yuekang Li, Zuxin Chen, Chaopeng Dong, Yongpan Wang, Hong Li 0004, Yongle Chen, Hongsong Zhu
ICSE2
2025 Source Code Summarization in the Era of Large Language Models
abstract
To support software developers in understanding and maintaining programs, various automatic (source) code summarization techniques have been proposed to generate a concise natural language summary (i.e., comment) for a given code snippet. Recently, the emergence of large language models (LLMs) has led to a great boost in the performance of coderelated tasks. In this paper, we undertake a systematic and comprehensive study on code summarization in the era of LLMs, which covers multiple aspects involved in the workflow of LLMbased code summarization. Specifically, we begin by examining prevalent automated evaluation methods for assessing the quality of summaries generated by LLMs and find that the results of the GPT-4 evaluation method are most closely aligned with human evaluation. Then, we explore the effectiveness of five prompting techniques (zero-shot, few-shot, chain-of-thought, critique, and expert) in adapting LLMs to code summarization tasks. Contrary to expectations, advanced prompting techniques may not outperform simple zero-shot prompting. Next, we investigate the impact of LLMs' model settings (including top_p and temperature parameters) on the quality of generated summaries. We find the impact of the two parameters on summary quality varies by the base LLM and programming language, but their impacts are similar. Moreover, we canvass LLMs' abilities to summarize code snippets in distinct types of programming languages. The results reveal that LLMs perform suboptimally when summarizing code written in logic programming languages compared to other language types (e.g., procedural and object-oriented programming languages). Finally, we unexpectedly find that CodeLlamaInstruct with 7B parameters can outperform advanced GPT-4 in generating summaries describing code design rationale and asserting code properties. We hope that our findings can provide a comprehensive understanding of code summarization in the era of LLMs.
Weisong Sun, Yun Miao, Yuekang Li, Hongyu Zhang 0002, Chunrong Fang, Yi Liu 0069, Gelei Deng, Yang Liu 0003, Zhenyu Chen 0001
ICSE3
2025 Truman: A Large Language Model-based Multi-agent Simulator for Synthetic Money Laundering Data Generation
Dattatray Vishnu Kute, Yuekang Li, Fethi A. Rabhi
AAMAS3
2025 A Large Scale Study of AI-based Binary Function Similarity Detection Techniques for Security Researchers and Practitioners
abstract
Binary Function Similarity Detection (BFSD) is a foundational technique in software security, underpinning a wide range of applications including vulnerability detection, malware analysis. Recent advances in AI-based BFSD tools have led to significant performance improvements. However, existing evaluations of these tools suffer from three key limitations: a lack of in-depth analysis of performance-influencing factors, an absence of realistic application analysis, and reliance on small-scale or low-quality datasets.In this paper, we present the first large-scale empirical study of AI-based BFSD tools to address these gaps. We construct two high-quality and diverse datasets: BinAtlas, comprising 12,453 binaries and over 7 million functions for capability evaluation; and BinAres, containing 12,291 binaries and 54 real-world 1-day vulnerabilities for evaluating vulnerability detection performance in practical IoT firmware settings. Using these datasets, we evaluate nine representative BFSD tools, analyze the challenges and limitations of existing BFSD tools, and investigate the consistency among BFSD tools. We also propose an actionable strategy for combining BFSD tools to enhance overall performance (an improvement of 13.4%). Our study not only advances the practical adoption of BFSD tools but also provides valuable resources and insights to guide future research in scalable and automated binary similarity detection.
Yang Xiao 0011, Yuekang Li, Zhengzi Xu, Sihao Qiu, Keyu Qi, Yeting Li, Xingchu Chen, Yanyan Zou 0002, Yang Liu 0003, Wei Huo 0005
ASE4
2025 From Constraints to Cracks: Constraint Semantic Inconsistencies as Vulnerability Beacons for Embedded Systems
Jiaxu Zhao 0004, Yuekang Li, Yanyan Zou 0002, Yang Xiao 0011, Naijia Jiang, Yeting Li, Nanyu Zhong, Bingwei Peng, Kunpeng Jian, Wei Huo 0005
USENIX Security Symposium2
2025 IllusionCAPTCHA: A CAPTCHA based on Visual Illusion
abstract
CAPTCHAs have long been essential tools for protecting applications from automated bots. Initially designed as simple questions to distinguish humans from bots, they have become increasingly complex to keep pace with the proliferation of CAPTCHA-cracking techniques employed by malicious actors. However, with the advent of advanced large language models (LLMs), the effectiveness of existing CAPTCHAs is now being undermined.
Gelei Deng, Yi Liu 0069, Junchen Ding, Jieshan Chen, Yulei Sui, Yuekang Li
WWW7
2025 Mission: Impossible - Image-Based Geolocation with Large Vision Language Models
abstract
In the age of ubiquitous smartphone use and widespread image sharing on social platforms, geolocation poses a critical privacy concern. Images often carry sensitive spatial and temporal details—such as street signs, architectural styles, or landmarks—that can inadvertently disclose the precise whereabouts of individuals and organizations. Recent advances in large vision-language models (LVLMs) present an emerging threat by enabling users, regardless of technical expertise, to extract location cues from seemingly benign photos. While existing AI-driven geolocation solutions often focus on narrow datasets or specialized contexts, the generalizable performance and privacy implications of zero-shot LVLMs in real-world settings remain critical questions. In this paper, we investigate the geolocation capabilities of state-of-the-art LVLMs. Our findings reveal that while these models demonstrate a non-negligible capability for image-based geolocation even without specialized training, their accuracy in absolute terms is often low, exposing clear limitations in their current state. We then introduce ETHAN, a framework integrating chain-of-thought (CoT) reasoning. Although ETHAN shows improved performance (e.g., 28.7% accuracy at the 1km threshold) and an 85.4% win rate on GeoGuessr, these results primarily highlight the potential trajectory of such technologies rather than their current widespread, high-accuracy applicability. Our study underscores the dual nature of LVLMs in this domain: they uncover an emerging privacy risk due to their inherent, albeit limited, geolocation abilities, yet also demonstrate significant constraints. We conclude by calling for further research into the limitations and risks of LVLM-based geolocation and the development of effective mitigation strategies to protect sensitive location data.
Yi Liu 0069, Gelei Deng, Junchen Ding, Yuekang Li, Tianwei Zhang 0004, Weisong Sun, Yaowen Zheng, Jingquan Ge
Proc. Priv. Enhancing Technol.4
2025 MiniScope: Automated UI Exploration and Privacy Inconsistency Detection of MiniApps via Two-phase Iterative Hybrid Analysis
abstract
The advent of MiniApps, operating within larger SuperApps, has revolutionized user experiences by offering a wide range of services without the need for individual app downloads. However, this convenience has raised significant privacy concerns, as these MiniApps often require access to sensitive data, potentially leading to privacy violations. Despite existing privacy regulations and platform guidelines, there is a lack of effective mechanisms to safeguard user privacy fully. To address this critical gap, we introduce MiniScope , a novel two-phase hybrid analysis approach, specifically designed for the MiniApp environment. This approach overcomes the limitations of existing static analysis techniques by incorporating UI transition states analysis, cross-package callback control flow resolution, and automated iterative UI exploration. This allows for a comprehensive understanding of MiniApps’ privacy practices, addressing the unique challenges of sub-package loading and event-driven callbacks. Our empirical evaluation of over 120K MiniApps using MiniScope demonstrates its effectiveness in identifying privacy inconsistencies. The results reveal significant issues, with 5.7% of MiniApps over-collecting private data and 33.4% overclaiming data collection. We have responsibly disclosed our findings to 2,282 developers, receiving 44 acknowledgments. These findings emphasize the urgent need for more precise privacy monitoring systems and highlight the responsibility of SuperApp operators to enforce stricter privacy measures.
Shenao Wang 0001, Yuekang Li, Kailong Wang 0001, Yi Liu 0069, Hui Li 0006, Yang Liu 0003, Haoyu Wang 0001
ACM Trans. Softw. Eng. Methodol.2
2025 TransferFuzz-Pro: Large Language Model Driven Code Debugging Technology for Verifying Propagated Vulnerability
abstract
Code reuse in software development frequently facilitates the spread of vulnerabilities, leading to imprecise scopes of affected software in CVE reports. Traditional methods focus primarily on detecting reused vulnerability code in target software but lack the ability to confirm whether these vulnerabilities can be triggered in new software contexts. In previous work, we introduced the TransferFuzz framework to address this gap by using historical trace-based fuzzing. However, its effectiveness is constrained by the need for manual intervention and reliance on source code instrumentation. To overcome these limitations, we propose TransferFuzz-Pro, a novel framework that integrates Large Language Model (LLM)-driven code debugging technology. By leveraging LLM for automated, human-like debugging and Proof-of-Concept (PoC) generation, combined with binary-level instrumentation, TransferFuzz-Pro extends verification capabilities to a wider range of targets. Our evaluation shows that TransferFuzz-Pro is significantly faster and can automatically validate vulnerabilities that were previously unverifiable using conventional methods. Notably, it expands the number of affected software instances for 15 CVE-listed vulnerabilities from 15 to 53 and successfully generates PoCs for various Linux distributions. These results demonstrate that TransferFuzz-Pro effectively verifies vulnerabilities introduced by code reuse in target software and automatically generation PoCs.
Siyuan Li 0014, Kaiyu Xie, Yuekang Li, Hong Li 0004, Yimo Ren, Limin Sun 0001, Hongsong Zhu
IEEE Trans. Software Eng.3
2024 Demystifying RCE Vulnerabilities in LLM-Integrated Apps
abstract
Large Language Models (LLMs) show promise in transforming software development, with a growing interest in integrating them into more intelligent apps. Frameworks like LangChain aid LLM-integrated app development, offering code execution utility/APIs for custom actions. However, these capabilities theoretically introduce Remote Code Execution (RCE) vulnerabilities, enabling remote code execution through prompt injections. No prior research systematically investigates these frameworks' RCE vulnerabilities or their impact on applications and exploitation consequences. Therefore, there is a huge research gap in this field.
Tong Liu 0027, Zizhuang Deng, Guozhu Meng, Yuekang Li, Kai Chen 0012
CCS4
2024 Bugs in Pods: Understanding Bugs in Container Runtime Systems
abstract
Container Runtime Systems (CRSs), which form the foundational infrastructure of container clouds, are critically important due to their impact on the quality of container cloud implementations. However, a comprehensive understanding of the quality issues present in CRS implementations remains lacking. To bridge this gap, we conduct the first comprehensive empirical study of CRS bugs. Specifically, we gather 429 bugs from 8,271 commits across dominant CRS projects, including runc, gvisor, containerd, and cri-o. Through manual analysis, we develop taxonomies of CRS bug symptoms and root causes, comprising 16 and 13 categories, respectively. Furthermore, we evaluate the capability of popular testing approaches, including unit testing, integration testing, and fuzz testing in detecting these bugs. The results show that 78.79% of the bugs cannot be detected due to the lack of test drivers, oracles, and effective test cases. Based on the findings of our study, we present implications and future research directions for various stakeholders in the domain of CRSs. We hope that our work can lay the groundwork for future research on CRS bug detection.
Jiongchi Yu, Xiaofei Xie, Cen Zhang, Sen Chen 0001, Yuekang Li, Wenbo Shen
ISSTA5
2024 How Effective Are They? Exploring Large Language Model Based Fuzz Driver Generation
abstract
Fuzz drivers are essential for library API fuzzing. However, automatically generating fuzz drivers is a complex task, as it demands the creation of high-quality, correct, and robust API usage code. An LLM-based (Large Language Model) approach for generating fuzz drivers is a promising area of research. Unlike traditional program analysis-based generators, this text-based approach is more generalized and capable of harnessing a variety of API usage information, resulting in code that is friendly for human readers. However, there is still a lack of understanding regarding the fundamental issues on this direction, such as its effectiveness and potential challenges. To bridge this gap, we conducted the first in-depth study targeting the important issues of using LLMs to generate effective fuzz drivers. Our study features a curated dataset with 86 fuzz driver generation questions from 30 widely-used C projects. Six prompting strategies are designed and tested across five state-of-the-art LLMs with five different temperature settings. In total, our study evaluated 736,430 generated fuzz drivers, with 0.85 billion token costs ($8,000+ charged tokens). Additionally, we compared the LLM-generated drivers against those utilized in industry, conducting extensive fuzzing experiments (3.75 CPU-year). Our study uncovered that: 1) While LLM-based fuzz driver generation is a promising direction, it still encounters several obstacles towards practical applications; 2) LLMs face difficulties in generating effective fuzz drivers for APIs with intricate specifics. Three featured design choices of prompt strategies can be beneficial: issuing repeat queries, querying with examples, and employing an iterative querying process; 3) While LLM-generated drivers can yield fuzzing outcomes that are on par with those used in the industry, there are substantial opportunities for enhancement, such as extending contained API usage, or integrating semantic oracles to facilitate logical bug detection. Our insights have been implemented to improve the OSS-Fuzz-Gen project, facilitating practical fuzz driver generation in industry.
Cen Zhang, Yaowen Zheng, Mingqiang Bai, Yeting Li, Wei Ma 0014, Xiaofei Xie, Yuekang Li, Limin Sun 0001, Yang Liu 0003
ISSTA7
2024 Rust-twins: Automatic Rust Compiler Testing through Program Mutation and Dual Macros Generation
abstract
Rust is a relatively new programming language known for its memory safety and numerous advanced features. It has been widely used in system software in recent years. Thus, ensuring the reliability and robustness of the only implementation of the Rust compiler, rustc, is critical. However, compiler testing, as one of the most effective techniques to detect bugs, faces difficulties in generating valid Rust programs with sufficient diversity due to its stringent memory safety mechanisms. Furthermore, existing research primarily focuses on testing rustc to trigger crash errors, neglecting incorrect compilation results - miscompilation. Detecting miscompilation remains a challenge in the absence of multiple implementations of the Rust compiler to serve as a test oracle.
Wenzhang Yang, Cuifeng Gao, Yuekang Li, Yinxing Xue
ASE4
2024 Improving Robustness of Hyperbolic Neural Networks by Lipschitz Analysis
abstract
Hyperbolic neural networks (HNNs) are emerging as a promising tool for representing data embedded in non-Euclidean geometries, yet their adoption has been hindered by challenges related to stability and robustness. In this work, we conduct a rigorous Lipschitz analysis for HNNs and propose using Lipschitz regularization as a novel strategy to enhance their robustness. Our comprehensive investigation spans both the Poincaré ball model and the hyperboloid model, establishing Lipschitz bounds for HNN layers. Importantly, our analysis provides detailed insights into the behavior of the Lipschitz bounds as they relate to feature norms, particularly distinguishing between scenarios where features have unit norms and those with large norms. Further, we study regularization using the derived Lipschitz bounds. Our empirical validations demonstrate consistent improvements in HNN robustness against noisy perturbations.
Yuekang Li, Yidan Mao, Yifei Yang 0001, Dongmian Zou
KDD1
2024 MASTERKEY: Automated Jailbreaking of Large Language Model Chatbots
Gelei Deng, Yi Liu 0069, Yuekang Li, Kailong Wang 0001, Ying Zhang 0066, Zefeng Li, Haoyu Wang 0001, Tianwei Zhang 0004, Yang Liu 0003
NDSS3
2024 File Hijacking Vulnerability: The Elephant in the Room
Chendong Yu, Yang Xiao 0011, Jie Lu 0009, Yuekang Li, Yeting Li, Lian Li 0002, Jian Wang 0067, Defang Bo, Wei Huo 0005
NDSS4
2024 PentestGPT: Evaluating and Harnessing Large Language Models for Automated Penetration Testing
Gelei Deng, Yi Liu 0069, Victor Mayoral Vilches, Yuekang Li, Yuan Xu 0033, Martin Pinzger 0001, Stefan Rass, Tianwei Zhang 0004, Yang Liu 0003
USENIX Security Symposium5
2024 Leveraging Semantic Relations in Code and Data to Enhance Taint Analysis of Embedded Systems
Jiaxu Zhao 0004, Yuekang Li, Yanyan Zou 0002, Zhaohui Liang, Yang Xiao 0011, Yeting Li, Bingwei Peng, Nanyu Zhong, Xinyi Wang 0013, Wei Huo 0005
USENIX Security Symposium2
2024 Medusa: Unveil Memory Exhaustion DoS Vulnerabilities in Protocol Implementations
abstract
Web services have brought great convenience to our daily lives. Meanwhile, they are vulnerable to Denial-of-Service (DoS) attacks. DoS attacks launched via vulnerabilities in the services can cause great harm. The vulnerabilities in protocol implementations are especially important because they are the keystones of web services. One vulnerable protocol implementation can affect all the web services built on top of it. Compared to the vulnerabilities that cause the target service to crash, resource exhaustion vulnerabilities are equally if not more important. This is because such vulnerabilities can deplete the system resources, leading to the unavailability of not only the vulnerable service but also other services running on the same machine. Despite the significance of this type of vulnerability, there has been limited research in this area.
Zhengjie Du, Yuekang Li, Yaowen Zheng, Cen Zhang, Yi Liu 0069, Sheikh Mahbub Habib, Xinghua Li 0001, Linzhang Wang, Yang Liu 0003, Bing Mao 0001
WWW2
2024 Drowzee: Metamorphic Testing for Fact-Conflicting Hallucination Detection in Large Language Models
abstract
Large language models (LLMs) have revolutionized language processing, but face critical challenges with security, privacy, and generating hallucinations — coherent but factually inaccurate outputs. A major issue is fact-conflicting hallucination (FCH), where LLMs produce content contradicting ground truth facts. Addressing FCH is difficult due to two key challenges: 1) Automatically constructing and updating benchmark datasets is hard, as existing methods rely on manually curated static benchmarks that cannot cover the broad, evolving spectrum of FCH cases. 2) Validating the reasoning behind LLM outputs is inherently difficult, especially for complex logical relations. To tackle these challenges, we introduce a novel logic-programming-aided metamorphic testing technique for FCH detection. We develop an extensive and extensible framework that constructs a comprehensive factual knowledge base by crawling sources like Wikipedia, seamlessly integrated into D rowzee . Using logical reasoning rules, we transform and augment this knowledge into a large set of test cases with ground truth answers. We test LLMs on these cases through template-based prompts, requiring them to provide reasoned answers. To validate their reasoning, we propose two semantic-aware oracles that assess the similarity between the semantic structures of the LLM answers and ground truth. Our approach automatically generates useful test cases and identifies hallucinations across six LLMs within nine domains, with hallucination rates ranging from 24.7% to 59.8%. Key findings include LLMs struggling with temporal concepts, out-of-distribution knowledge, and lack of logical reasoning capabilities. The results show that logic-based test cases generated by D rowzee effectively trigger and detect hallucinations. To further mitigate the identified FCHs, we explored model editing techniques, which proved effective on a small scale (with edits to fewer than 1000 knowledge pieces). Our findings emphasize the need for continued community efforts to detect and mitigate model hallucinations.
Ningke Li, Yuekang Li, Yi Liu 0069, Ling Shi 0002, Kailong Wang 0001, Haoyu Wang 0001
Proc. ACM Program. Lang.2
2023 PumpChannel: An Efficient and Secure Communication Channel for Trusted Execution Environment on ARM-FPGA Embedded SoC
abstract
ARM TrustZone separates the system into the rich execution environment (REE) and the trusted execution environment (TEE). Data can be exchanged between REE and TEE through the communication channel, which is based on shared memory and can be accessed by both REE and TEE. Therefore, when the REE OS kernel is untrusted, the security of the communication channel cannot be guaranteed. The proposed schemes to protect the communication channel have high performance overhead and are not secure enough. In this paper, we propose PumpChannel, an efficient and secure communication channel implemented on ARM-FPGA embedded SoC. PumpChannel avoids the use of secret keys, but utilizes a hardware and software collaborative pump to enhance the security and performance of the communication channel. Besides, PumpChannel implements a hardware-based hook integrity monitor to ensure the integrity of all hook code. Security and performance evaluation results show that PumpChannel is more secure than the encrypted channel countermeasures and has better performance than all other evaluated schemes.
Jingquan Ge, Yuekang Li, Yang Liu 0003, Yaowen Zheng, Yi Liu 0069, Lida Zhao
DATE2
2023 ACETest: Automated Constraint Extraction for Testing Deep Learning Operators
abstract
Deep learning (DL) applications are prevalent nowadays as they can help with multiple tasks. DL libraries are essential for building DL applications. Furthermore, DL operators are the important building blocks of the DL libraries, that compute the multi-dimensional data (tensors). Therefore, bugs in DL operators can have great impacts. Testing is a practical approach for detecting bugs in DL operators. In order to test DL operators effectively, it is essential that the test cases pass the input validity check and are able to reach the core function logic of the operators. Hence, extracting the input validation constraints is required for generating high-quality test cases. Existing techniques rely on either human effort or documentation of DL library APIs to extract the constraints. They cannot extract complex constraints and the extracted constraints may differ from the actual code implementation. To address the challenge, we propose ACETest, a technique to automatically extract input validation constraints from the code to build valid yet diverse test cases which can effectively unveil bugs in the core function logic of DL operators. For this purpose, ACETest can automatically identify the input validation code in DL operators, extract the related constraints and generate test cases according to the constraints. The experimental results on popular DL libraries, TensorFlow and PyTorch, demonstrate that ACETest can extract constraints with higher quality than state-of-the-art (SOTA) techniques. Moreover, ACETest is capable of extracting 96.4% more constraints and detecting 1.95 to 55 times more bugs than SOTA techniques. In total, we have used ACETest to detect 108 previously unknown bugs on TensorFlow and PyTorch, with 87 of them confirmed by the developers. Lastly, five of the bugs were assigned with CVE IDs due to their security impacts.
Yang Xiao 0011, Yuekang Li, Yeting Li, Dongsong Yu, Chendong Yu, Hui Su, Wei Huo 0005
ISSTA3
2023 ASTER: Automatic Speech Recognition System Accessibility Testing for Stutterers
abstract
The popularity of automatic speech recognition (ASR) systems nowadays leads to an increasing need for improving their accessibility. Handling stuttering speech is an important feature for accessible ASR systems. To improve the accessibility of ASR systems for stutterers, we need to expose and analyze the failures of ASR systems on stuttering speech. The speech datasets recorded from stutterers are not diverse enough to expose most of the failures. Furthermore, these datasets lack ground truth information about the non-stuttered text, rendering them unsuitable as comprehensive test suites. Therefore, a methodology for generating stuttering speech as test inputs to test and analyze the performance of ASR systems is needed. However, generating valid test inputs in this scenario is challenging. The reason is that although the generated test inputs should mimic how stutterers speak, they should also be diverse enough to trigger more failures. To address the challenge, we propose Aster, a technique for automatically testing the accessibility of ASR systems. Aster can generate valid test cases by injecting five different types of stuttering. The generated test cases can both simulate realistic stuttering speech and expose failures in ASR systems. Moreover, Aster can further enhance the quality of the test cases with a multi-objective optimization-based seed updating algorithm. We implemented Aster as a framework and evaluated it on four open-source ASR models and three commercial ASR systems. We conduct a comprehensive evaluation of Aster and find that it significantly increases the word error rate, match error rate, and word information loss in the evaluated ASR systems. Additionally, our user study demonstrates that the generated stuttering audio is indistinguishable from real-world stuttering audio clips.
Yi Liu 0069, Yuekang Li, Gelei Deng, Felix Juefei-Xu, Yao Du 0002, Cen Zhang, Yeting Li, Lei Ma 0003, Yang Liu 0003
ASE2
2023 RSFuzzer: Discovering Deep SMI Handler Vulnerabilities in UEFI Firmware with Hybrid Fuzzing
abstract
System Management Mode (SMM) is a secure operation mode for x86 processors supported by Unified Extensible Firmware Interface (UEFI) firmware. SMM is designed to provide a secure execution environment to access highly privileged data or control low-level hardware (such as power management). The programs running in SMM are called SMM drivers and System Management Interrupt (SMI) handlers are the most important components of SMM drivers since they are the only components to receive and handle data from outside the SMM execution environment. Although SMM can serve as an extra layer of protection when the operating system is compromised, vulnerabilities in SMM drivers, especially SMI handlers, can invalidate this protection and cause severe damages to the device. Thus, early detection of SMI handler vulnerabilities is important for UEFI firmware security.To this end, researchers have proposed to use hybrid fuzzing techniques for detecting SMI handler vulnerabilities. Particularly, Intel has developed a hybrid fuzzer called Excite and uses it to secure Intel products. Although existing hybrid fuzzing techniques can detect vulnerabilities in SMI handlers, their effectiveness is limited due to two major pitfalls: 1) They can only feed input through the most common input interface to SMI handlers, lacking the ability to utilize other input interfaces. 2) They have no awareness of variables shared by multiple SMI handlers, lacking the ability to explore code segments related to such variables. By addressing the challenges faced by existing works, we propose RSFuzzer, a hybrid greybox fuzzing technique which can learn input interface and format information and detect deeply hidden vulnerabilities which are triggered by invoking multiple SMI handlers. We implemented RSFuzzer and evaluated it on 16 UEFI firmware images provided by six vendors. The experiment results show that RSFuzzer can cover 617% more basic blocks and detect 828% more vulnerabilities on average than the state-of-the-art hybrid fuzzing technique. Moreover, we found and reported 65 0-day vulnerabilities in the evaluated UEFI firmware images and 14 CVE IDs were assigned. Noticeably, 6 of the 0-day vulnerabilities were found in commercial-off-the-shelf (COTS) products from Intel, which might have been tested by Excite before releasing.
Jiawei Yin, Yuekang Li, Boru Lin, Yanyan Zou 0002, Yang Liu 0003, Wei Huo 0005, Jingling Xue
SP3
2023 NAUTILUS: Automated RESTful API Vulnerability Detection
Gelei Deng, Zhiyi Zhang 0005, Yuekang Li, Yi Liu 0069, Tianwei Zhang 0004, Yang Liu 0003, Dongjin Wang
USENIX Security Symposium3
2023 Automata-Guided Control-Flow-Sensitive Fuzz Driver Generation
Cen Zhang, Yuekang Li, Hao Zhou 0043, Yaowen Zheng, Xian Zhan, Xiaofei Xie, Xiapu Luo, Xinghua Li 0001, Yang Liu 0003, Sheikh Mahbub Habib
USENIX Security Symposium2
2023 Dynamic material parameter inversion of high arch dam under discharge excitation based on the modal parameters and Bayesian optimised deep learning
Huokun Li, Pengzhen Wu, Yuekang Li
Adv. Eng. Informatics6
2022 Windranger: A Directed Greybox Fuzzer driven by Deviation Basic Blocks
abstract
Directed grey-box fuzzing (DGF) is a security testing technique that aims to steer the fuzzer towards predefined target sites in the program. To gain directedness, DGF prioritizes the seeds whose execution traces are closer to the target sites. Therefore, evaluating the distance between the execution trace of a seed and the target sites (aka, the seed distance) is important for DGF. The first directed grey-box fuzzer, AFLGo, uses an approach of calculating the basic block level distances during static analysis and accumulating the distances of the executed basic blocks to compute the seed distance. Following AFLGo, most of the existing state-of-the-art DGF techniques use all the basic blocks on the execution trace and only the control flow information for seed distance calculation. However, not every basic block is equally important and there are certain basic blocks where the execution trace starts to deviate from the target sites (aka, deviation basic blocks).
Zhengjie Du, Yuekang Li, Yang Liu 0003, Bing Mao 0001
ICSE2
2022 Morest: Model-based RESTful API Testing with Execution Feedback
abstract
RESTful APIs are arguably the most popular endpoints for accessing Web services. Blackbox testing is one of the emerging techniques for ensuring the reliability of RESTful APIs. The major challenge in testing RESTful APIs is the need for correct sequences of API operation calls for in-depth testing. To build meaningful operation call sequences, researchers have proposed techniques to learn and utilize the API dependencies based on OpenAPI specifications. However, these techniques either lack the overall awareness of how all the APIs are connected or the flexibility of adaptively fixing the learned knowledge.
Yi Liu 0069, Yuekang Li, Gelei Deng, Yang Liu 0003, Ruiyuan Wan, Runchao Wu, Dandan Ji, Shiheng Xu, Minli Bao
ICSE2
2022 Efficient greybox fuzzing of applications in Linux-based IoT devices via enhanced user-mode emulation
abstract
Greybox fuzzing has become one of the most effective vulnerability discovery techniques. However, greybox fuzzing techniques cannot be directly applied to applications in IoT devices. The main reason is that executing these applications highly relies on specific system environments and hardware. To execute the applications in Linux-based IoT devices, most existing fuzzing techniques use full-system emulation for the purpose of maximizing compatibility. However, compared with user-mode emulation, full-system emulation suffersfrom great overhead. Therefore, some previous works, such as Firm-AFL, propose to combine full-system emulation and user-mode emulation to speed up the fuzzing process. Despite the attempts of trying to shift the application towards user-mode emulation, no existing technique supports to execute these applications fully in the user-mode emulation. To address this issue, we propose EQUAFL, which can automatically set up the execution environment to execute embedded applications under user-mode emulation. EQUAFL first executes the application under full-system emulation and observe for the key points where the program may get stuck or even crash during user-mode emulation. With the observed information, EQUAFL can migrate the needed environment for user-mode emulation. Then, EQUAFL uses an enhanced user-mode emulation to replay system calls of network, and resource management behaviors to fulfill the needs of the embedded application during its execution. We evaluate EQUAFL on 70 network applications from different series of IoT devices. The result shows EQUAFL outperforms the state-of-the-arts in fuzzing efficiency (on average, 26 times faster than AFL-QEMU with full-system emulation, 14 times than Firm-AFL). We have also discovered ten vulnerabilities including six CVEs from the tested firmware images.
Yaowen Zheng, Yuekang Li, Cen Zhang, Hongsong Zhu, Yang Liu 0003, Limin Sun 0001
ISSTA2
2022 Morest: Industry Practice of Automatic RESTful API Testing
abstract
Many big companies are providing cloud services through RESTful APIs nowadays. With the growing popularity of RESTful API, testing RESTful API becomes crucial. To address this issue, researchers have proposed several automatic RESTful API testing techniques. At Huawei, we design and implement an automatic RESTful API testing framework named Morest. Morest has been used to test ten RESTful API services and helped to detected 83 previously unknown bugs which were all confirmed and fixed by the developers. On one hand, we find that Morest shows great capability of detecting bugs in RESTful API s. On the other hand, we also notice that human effort is inevitable and important when applying automatic RESTful API techniques in practice.
Yi Liu 0069, Yuekang Li, Yang Liu 0003, Ruiyuan Wan, Runchao Wu, Qingkun Liu
ASE2
2022 RegexScalpel: Regular Expression Denial of Service (ReDoS) Defense by Localize-and-Fix
Yeting Li, Yecheng Sun, Zhiwu Xu 0001, Jialun Cao, Yuekang Li, Rongchen Li, Haiming Chen 0001, Shing-Chi Cheung, Yang Liu 0003, Yang Xiao 0011
USENIX Security Symposium5
2021 SoFi: Reflection-Augmented Fuzzing for JavaScript Engines
abstract
JavaScript engines have been shown prone to security vulnerabilities, which can lead to serious consequences due to their popularity. Fuzzing is an effective testing technique to discover vulnerabilities. The main challenge of fuzzing JavaScript engines is to generate syntactically and semantically valid inputs such that deep functionalities can be explored. However, due to the dynamic nature of JavaScript and the special features of different engines, it is quite challenging to generate semantically meaningful test inputs.
Xiaofei Xie, Yuekang Li, Feng Li 0045, Yang Liu 0003, Wenchang Shi, Wei Huo 0005
CCS3
2021 Vall-nut: Principled Anti-Grey box - Fuzzing
abstract
Greybox fuzzing is a widely used technique for software testing that has been adopted by practitioners and researchers to disclose a great number of vulnerabilities in various software. However, adversaries also weaponize greybox fuzzing to mine vulnerabilities for malicious intentions. This poses considerable threats to software systems. To counteract the misuse of greybox fuzzing, we propose VALL-NUT, a novel approach to harden software with properties to combat greybox fuzzing. We dissect the major strategies that facilitate the success of greybox fuzzing, and accordingly propose three types of neutralizing schemesseed queue explosion, seed attenuation, and feedback contamination. We evaluate Vall-nut against the mainstream greybox fuzzers on multiple real-world benchmark programs. The results show that Vall-nut can reduce an average of 34 % code coverage and 76% detected crashes in 24-hour tests. Moreover, we conduct comparisons with two recent studies which show Vall-nut can achieve a superior deduction of detected crashes.
Yuekang Li, Guozhu Meng, Jun Xu 0024, Cen Zhang, Hongxu Chen 0001, Xiaofei Xie, Haijun Wang 0002, Yang Liu 0003
ISSRE1
2021 A First Look at the Effect of Deep Learning in Coverage-guided Fuzzing
abstract
Fuzzing has been a widely-used technique for discovering software vulnerabilities. Many existing fuzzers leverage coverage-feedback to evolve seeds to maximize (optimize) program branch coverage. Recently, some techniques propose to train deep learning models to predict the branch coverage of an arbitrary input. Those techniques have proved their success in improving coverage and discovering bugs under different experimental settings. However, deep learning models, usually as a black magic box, are notoriously lack of explanation. Moreover, their performance can be sensitive to the collected runtime coverage information for training, indicating potentially unstable performance. To this end, in this work we conduct a systematic and extensive empirical study on 4 types of deep learning models across 6 projects to reproduce the actual performance of deep learning fuzzers, analyze the advantages and disadvantages of deep learning in the process of fuzzing applications, and explore the future direction of the combination of the two. Our empirical results reveal that the deep learning models can only be effective in very limited scenarios, which is largely restrained by training data imbalance, dependant labels, model over-generalization, and the insufficient expressiveness of the state-of-the-art models. Consequently, the estimated gradients by the models to cover a branch can be less helpful in many scenarios.
Yun Lin 0001, Xiaofei Xie, Yuekang Li, Xiaohong Li 0001, Weimin Ge, Yang Liu 0003, Jin Song Dong 0001
ASE4
2021 BIFF: Practical Binary Fuzzing Framework for Programs of IoT and Mobile Devices
abstract
Internet-of-things (IoT) or mobile devices are omnipresent in our daily life; the security issues inside them are especially crucial. Greybox fuzzing has been shown effective in detecting vulnerabilities. However, applications in IoT or mobile devices are usually proprietary to specific vendors, fuzzers are required to support binary-only targets. Moreover, since these devices are of heterogeneous architecture, assigned with limited resources, and many testing targets are server-like programs, applying existing fuzzing techniques faces great challenges.This paper proposes BIFF, a general-purpose fuzzer that aims to stress these issues. It supports binary-only targets, is general (supports multiple CPU architectures including Intel, ARM, MIPS, and PowerPC), fast (has the lowest runtime overhead compared to existing fuzzers), and flexible (uses a new fuzzing workflow that can fuzz any piece of code inside the target binary). Experiments demonstrate that BIFF has the best performance compared with state-of-the-art binary fuzzers and can fuzz the server-like programs which cannot be fuzzed by the existing fuzzers. Using BIFF, we’ve found 24 unknown vulnerabilities (including memory corruptions, infinite loops, and infinite recursions) in industrial products.
Cen Zhang, Yuekang Li, Hongxu Chen 0001, Xiaoxing Luo, Miaohua Li, Anh Quynh Nguyen, Yang Liu 0003
ASE2
2021 AutoCom: Automatic Comment Generation for C Code
abstract
Code comments improve program comprehension and program maintenance.However, the lack of comments is a common problem in industry.It is time-and manpower-consuming to add comments for large code bases.Thus, it is desirable to develop techniques for automatic comment generation.Previous works for automatic comment generation use deep learning or machine learning techniques.These techniques require a large amount of training data which is often unavailable or hard to acquire.This paper proposes a light-weight approach called AutoCom for automatic comment generation.In AutoCom, we first analyze the source code to extract key information.Then, we use the extracted information to search and filter for appropriate text from a large programming Question and Answer (Q&A) site.Lastly, we use NLP techniques to convert the search result into code comments.In addition to the generated code comment, we also add predefined comments for library function usage in the source code.With AutoCom as the back-end, we built a web service which allows the user to upload source code and get it commented.
Zhikang Tian, Yuekang Li
SEKE2
2021 APICraft: Fuzz Driver Generation for Closed-source SDK Libraries
Cen Zhang, Xingwei Lin, Yuekang Li, Yinxing Xue, Jundong Xie, Hongxu Chen 0001, Xinlei Ying, Jiashui Wang, Yang Liu 0003
USENIX Security Symposium3
2020 Ori: A Greybox Fuzzer for SOME/IP Protocols in Automotive Ethernet
abstract
With the emergence of smart automotive devices, the data communication between these devices gains increasing importance. SOME/IP is a light-weight protocol to facilitate inter- process/device communication, which supports both procedural calls and event notifications. Because of its simplicity and capability, SOME/IP is getting adopted by more and more automotive devices. Subsequently, the security of SOME/IP applications becomes crucial. However, previous security testing techniques cannot fit the scenario of vulnerability detection SOME/IP applications due to miscellaneous challenges such as the difficulty of server-side testing programs in parallel, etc. By addressing these challenges, we propose Ori - a greybox fuzzer for SOME/IP applications, which features two key innovations: the attach fuzzing mode and structural mutation. The attach fuzzing mode enables Ori to test server programs efficiently, and the structural mutation allows Ori to generate valid SOME/IP packets to reach deep paths of the target program effectively. Our evaluation shows that Ori can detect vulnerabilities in SOME/IP applications effectively and efficiently.
Yuekang Li, Hongxu Chen 0001, Cen Zhang, Siyang Xiong, Chaoyi Liu, Yi Estelle Wang
APSEC1
2020 Typestate-guided fuzzer for discovering use-after-free vulnerabilities
abstract
Existing coverage-based fuzzers usually use the individual control flow graph (CFG) edge coverage to guide the fuzzing process, which has shown great potential in finding vulnerabilities. However, CFG edge coverage is not effective in discovering vulnerabilities such as use-after-free (UaF). This is because, to trigger UaF vulnerabilities, one needs not only to cover individual edges, but also to traverse some (long) sequence of edges in a particular order, which is challenging for existing fuzzers. To this end, we propose to model UaF vulnerabilities as typestate properties, and develop a typestate-guided fuzzer, named UAFL, for discovering vulnerabilities violating typestate properties. Given a typestate property, we first perform a static typestate analysis to find operation sequences potentially violating the property. Our fuzzing process is then guided by the operation sequences in order to progressively generate test cases triggering property violations. In addition, we also employ an information flow analysis to improve the efficiency of the fuzzing process. We have performed a thorough evaluation of UAFL on 14 widely-used real-world programs. The experiment results show that UAFL substantially outperforms the state-of-the-art fuzzers, including AFL, AFLFast, FairFuzz, MOpt, Angora and QSYM, in terms of the time taken to discover vulnerabilities. We have discovered 10 previously unknown vulnerabilities, and received 5 new CVEs.
Haijun Wang 0002, Xiaofei Xie, Yi Li 0008, Cheng Wen 0002, Yuekang Li, Yang Liu 0003, Shengchao Qin, Hongxu Chen 0001, Yulei Sui
ICSE5
2020 MemLock: memory usage guided fuzzing
abstract
Uncontrolled memory consumption is a kind of critical software security weaknesses. It can also become a security-critical vulnerability when attackers can take control of the input to consume a large amount of memory and launch a Denial-of-Service attack. However, detecting such vulnerability is challenging, as the state-of-the-art fuzzing techniques focus on the code coverage but not memory consumption. To this end, we propose a memory usage guided fuzzing technique, named MemLock, to generate the excessive memory consumption inputs and trigger uncontrolled memory consumption bugs. The fuzzing process is guided with memory consumption information so that our approach is general and does not require any domain knowledge. We perform a thorough evaluation for MemLock on 14 widely-used real-world programs. Our experiment results show that MemLock substantially outperforms the state-of-the-art fuzzing techniques, including AFL, AFLfast, PerfFuzz, FairFuzz, Angora and QSYM, in discovering memory consumption bugs. During the experiments, we discovered many previously unknown memory consumption bugs and received 15 new CVEs.
Cheng Wen 0002, Haijun Wang 0002, Yuekang Li, Shengchao Qin, Yang Liu 0003, Zhiwu Xu 0001, Hongxu Chen 0001, Xiaofei Xie, Geguang Pu, Ting Liu 0002
ICSE3
2020 MUZZ: Thread-aware Grey-box Fuzzing for Effective Bug Hunting in Multithreaded Programs
Hongxu Chen 0001, Shengjian Guo, Yinxing Xue, Yulei Sui, Cen Zhang, Yuekang Li, Haijun Wang 0002, Yang Liu 0003
USENIX Security Symposium6
2019 Leopard: identifying vulnerable code for vulnerability assessment through program metrics
abstract
Identifying potentially vulnerable locations in a code base is critical as a pre-step for effective vulnerability assessment; i.e., it can greatly help security experts put their time and effort to where it is needed most. Metric-based and pattern-based methods have been presented for identifying vulnerable code. The former relies on machine learning and cannot work well due to the severe imbalance between non-vulnerable and vulnerable code or lack of features to characterize vulnerabilities. The latter needs the prior knowledge of known vulnerabilities and can only identify similar but not new types of vulnerabilities. In this paper, we propose and implement a generic, lightweight and extensible framework, LEOPARD, to identify potentially vulnerable functions through program metrics. LEOPARD requires no prior knowledge about known vulnerabilities. It has two steps by combining two sets of systematically derived metrics. First, it uses complexity metrics to group the functions in a target application into a set of bins. Then, it uses vulnerability metrics to rank the functions in each bin and identifies the top ones as potentially vulnerable. Our experimental results on 11 real-world projects have demonstrated that, LEOPARD can cover 74.0% of vulnerable functions by identifying 20% of functions as vulnerable and outperform machine learning-based and static analysis-based techniques. We further propose three applications of LEOPARD for manual code review and fuzzing, through which we discovered 22 new bugs in real applications like PHP, radare2 and FFmpeg, and eight of them are new vulnerabilities.
Xiaoning Du 0001, Bihuan Chen 0001, Yuekang Li, Jianmin Guo, Yaqin Zhou, Yang Liu 0003, Yu Jiang 0001
ICSE3
2019 DiffChaser: Detecting Disagreements for Deep Neural Networks
abstract
The platform migration and customization have become an indispensable process of deep neural network (DNN) development lifecycle. A high-precision but complex DNN trained in the cloud on massive data and powerful GPUs often goes through an optimization phase (e.g, quantization, compression) before deployment to a target device (e.g, mobile device). A test set that effectively uncovers the disagreements of a DNN and its optimized variant provides certain feedback to debug and further enhance the optimization procedure. However, the minor inconsistency between a DNN and its optimized version is often hard to detect and easily bypasses the original test set. This paper proposes DiffChaser, an automated black-box testing framework to detect untargeted/targeted disagreements between version variants of a DNN. We demonstrate 1) its effectiveness by comparing with the state-of-the-art techniques, and 2) its usefulness in real-world DNN product deployment involved with quantization and optimization.
Xiaofei Xie, Lei Ma 0003, Haijun Wang 0002, Yuekang Li, Yang Liu 0003, Xiaohong Li 0001
IJCAI4
2019 Cerebro: context-aware adaptive fuzzing for effective vulnerability detection
abstract
Existing greybox fuzzers mainly utilize program coverage as the goal to guide the fuzzing process. To maximize their outputs, coverage-based greybox fuzzers need to evaluate the quality of seeds properly, which involves making two decisions: 1) which is the most promising seed to fuzz next (seed prioritization), and 2) how many efforts should be made to the current seed (power scheduling). In this paper, we present our fuzzer, Cerebro, to address the above challenges. For the seed prioritization problem, we propose an online multi-objective based algorithm to balance various metrics such as code complexity, coverage, execution time, etc. To address the power scheduling problem, we introduce the concept of input potential to measure the complexity of uncovered code and propose a cost-effective algorithm to update it dynamically. Unlike previous approaches where the fuzzer evaluates an input solely based on the execution traces that it has covered, Cerebro is able to foresee the benefits of fuzzing the input by adaptively evaluating its input potential. We perform a thorough evaluation for Cerebro on 8 different real-world programs. The experiments show that Cerebro can find more vulnerabilities and achieve better coverage than state-of-the-art fuzzers such as AFL and AFLFast.
Yuekang Li, Yinxing Xue, Hongxu Chen 0001, Xiuheng Wu, Cen Zhang, Xiaofei Xie, Haijun Wang 0002, Yang Liu 0003
ESEC/SIGSOFT FSE1
2019 Locating vulnerabilities in binaries via memory layout recovering
abstract
Locating vulnerabilities is an important task for security auditing, exploit writing, and code hardening. However, it is challenging to locate vulnerabilities in binary code, because most program semantics (e.g., boundaries of an array) is missing after compilation. Without program semantics, it is difficult to determine whether a memory access exceeds its valid boundaries in binary code. In this work, we propose an approach to locate vulnerabilities based on memory layout recovery. First, we collect a set of passed executions and one failed execution. Then, for passed and failed executions, we restore their program semantics by recovering fine-grained memory layouts based on the memory addressing model. With the memory layouts recovered in passed executions as reference, we can locate vulnerabilities in failed execution by memory layout identification and comparison. Our experiments show that the proposed approach is effective to locate vulnerabilities on 24 out of 25 DARPA’s CGC programs (96%), and can effectively classifies 453 program crashes (in 5 Linux programs) into 19 groups based on their root causes.
Haijun Wang 0002, Xiaofei Xie, Shangwei Lin 0001, Yun Lin 0001, Yuekang Li, Shengchao Qin, Yang Liu 0003, Ting Liu 0002
ESEC/SIGSOFT FSE5
2018 Hawkeye: Towards a Desired Directed Grey-box Fuzzer
abstract
Grey-box fuzzing is a practically effective approach to test real-world programs. However, most existing grey-box fuzzers lack directedness, i.e. the capability of executing towards user-specified target sites in the program. To emphasize existing challenges in directed fuzzing, we propose Hawkeye to feature four desired properties of directed grey-box fuzzers. Owing to a novel static analysis on the program under test and the target sites, Hawkeye precisely collects the information such as the call graph, function and basic block level distances to the targets. During fuzzing, Hawkeye evaluates exercised seeds based on both static information and the execution traces to generate the dynamic metrics, which are then used for seed prioritization, power scheduling and adaptive mutating. These strategies help Hawkeye to achieve better directedness and gravitate towards the target sites. We implemented Hawkeye as a fuzzing framework and evaluated it on various real-world programs under different scenarios. The experimental results showed that Hawkeye can reach the target sites and reproduce the crashes much faster than state-of-the-art grey-box fuzzers such as AFL and AFLGo. Specially, Hawkeye can reduce the time to exposure for certain vulnerabilities from about 3.5 hours to 0.5 hour. By now, Hawkeye has detected more than 41 previously unknown crashes in projects such as Oniguruma, MJS with the target sites provided by vulnerability prediction tools; all these crashes are confirmed and 15 of them have been assigned CVE IDs.
Hongxu Chen 0001, Yinxing Xue, Yuekang Li, Bihuan Chen 0001, Xiaofei Xie, Xiuheng Wu, Yang Liu 0003
CCS3
2018 Principled Greybox Fuzzing
Yuekang Li
ICFEM1
2018 FOT: a versatile, configurable, extensible fuzzing framework
abstract
Greybox fuzzing is one of the most effective approaches for detecting software vulnerabilities. Various new techniques have been continuously emerging to enhance the effectiveness and/or efficiency by incorporating novel ideas into different components of a greybox fuzzer. However, there lacks a modularized fuzzing framework that can easily plugin new techniques and hence facilitate the reuse, integration and comparison of different techniques. To address this problem, we propose a fuzzing framework, namely Fuzzing Orchestration Toolkit (FOT). FOT is designed to be versatile, configurable and extensible. With FOT and its extensions, we have found 111 new bugs from 11 projects. Among these bugs, 18 CVEs have been assigned. Video link: https://youtu.be/O6Qu7BJ8RP0.
Hongxu Chen 0001, Yuekang Li, Bihuan Chen 0001, Yinxing Xue, Yang Liu 0003
ESEC/SIGSOFT FSE2
2017 Steelix: program-state based binary fuzzing
abstract
Coverage-based fuzzing is one of the most effective techniques to find vulnerabilities, bugs or crashes. However, existing techniques suffer from the difficulty in exercising the paths that are protected by magic bytes comparisons (e.g., string equality comparisons). Several approaches have been proposed to use heavy-weight program analysis to break through magic bytes comparisons, and hence are less scalable. In this paper, we propose a program-state based binary fuzzing approach, named Steelix, which improves the penetration power of a fuzzer at the cost of an acceptable slow down of the execution speed. In particular, we use light-weight static analysis and binary instrumentation to provide not only coverage information but also comparison progress information to a fuzzer. Such program state information informs a fuzzer about where the magic bytes are located in the test input and how to perform mutations to match the magic bytes efficiently. We have implemented Steelix and evaluated it on three datasets: LAVA-M dataset, DARPA CGC sample binaries and five real-life programs. The results show that Steelix has better code coverage and bug detection capability than the state-of-the-art fuzzers. Moreover, we found one CVE and nine new bugs.
Yuekang Li, Bihuan Chen 0001, Mahinthan Chandramohan, Shangwei Lin 0001, Yang Liu 0003, Alwen Tiu
ESEC/SIGSOFT FSE1