VLDB 2026 Research / reviewers in the wild / expert
Penghui Li 0001
dblp:81/8013-1
· DBLP profile ↗
15ranked-venue papers
7as first author
15since 2021 · last 2025
0000-0002-3077-5697ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 10 · 3 first-author · 10 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | PickleBall: Secure Deserialization of Pickle-based Machine Learning ModelsabstractMachine learning model repositories, such as the Hugging Face Model Hub, facilitate model exchanges. However, bad actors can deliver malware through compromised models. Existing defenses, such as safer model formats, restrictive (but inflexible) loading policies, and model scanners, have shortcomings: 44.9% of popular models on Hugging Face still use the insecure pickle format, 15% of these cannot be loaded by restrictive loading policies, and model scanners have both false positives and false negatives. Pickle remains the de facto standard for model exchange, and the ML community lacks a tool that offers transparent safe loading. Andreas D. Kellas, Neophytos Christou, Wenxin Jiang 0001, Penghui Li 0001, Yaniv David, Vasileios P. Kemerlis, James C. Davis 0001 |
CCS | 4 |
| 2025 | VulShield: Protecting Vulnerable Code Before Deploying Patches
Yuan Li 0061, Chao Zhang 0008, Jinhao Zhu, Penghui Li 0001, Songtao Yang 0001, Wende Tan |
NDSS | 4 |
| 2025 | Predator: Directed Web Application Fuzzing for Efficient Vulnerability ValidationabstractWeb application vulnerabilities continue to pose a significant challenge. Static analysis is currently the mainstream approach to this issue, while dynamic analysis is not as widely used in comparison. However, both techniques have their limitations. While current static analysis tools are plagued by high false-positive rates, necessitating fine-grained analysis and substantial expertise, it is also the case that dynamic analysis tools are underdeveloped. Current fuzzing-based tools are often limited by inefficiency in exploring deeper code locations. Moreover, state-of-the-art grey-box fuzzers often struggle to capture effective parameters from user interfaces, thereby failing to explore the input space efficiently. In this paper, we propose Predator, a directed fuzzing framework equipped with selective dynamic instrumentation for effective and efficient web application vulnerability detection and validation. We use static analysis techniques and dynamic analysis techniques to complement each other. Our lightweight static analysis provides relevant URLs and parameters of the directed fuzzing targets and thus facilitates dynamic validation of static analysis reports. Additionally, we propose a runtime distance supplementation mechanism and tailored mutation strategies to address the dynamic features of interpreted languages like PHP. The evaluation shows Predator effectively triggers more vulnerabilities and outperforms state-of-the-art grey-box fuzzers by up to 43.8 times in terms of time to exposure. Moreover, Predator detects 26 previously unknown vulnerabilities in real-world applications, further demonstrating its effectiveness. At the time of writing, 7 of the 26 vulnerabilities have been confirmed and patched by the corresponding vendors. Chenlin Wang, Wei Meng 0001, Changhua Luo, Penghui Li 0001 |
SP | 4 |
| 2024 | FuzzCache: Optimizing Web Application Fuzzing Through Software-Based Data CacheabstractFuzzing has shown great promise in detecting vulnerabilities in server-side web applications. In this work, we introduce an innovative software-based data cache mechanism that complements and improves all existing web application fuzzing tools. Our key observation is that a great proportion of execution time (e.g., 50%) of web applications is spent on fetching data from two major sources: database and network; our in-depth investigation reveals that the same data is often repeatedly fetched across fuzzing trials. We thus design a new solution, FuzzCache, that stores the data into software-based caches, mitigating the need for repeated and expensive data fetches. FuzzCache exposes the cached data across fuzzing trials through inter-process shared memory segments. It also, as the first work, incorporates just-in-time compilation to avoid the performance overhead associated with interpreting PHP code in real time, thereby enhancing execution efficiency. Penghui Li 0001, Mingxue Zhang 0001 |
CCS | 1 |
| 2024 | Test Suites Guided Vulnerability Validation for Node.js ApplicationsabstractDynamic methods have shown great promise in validating vulnerabilities and generating Proof-of-Concept (PoC) exploits of Node.js applications. They typically rely on dictionaries or specifications to determine the values of request parameters and their relationships. However, they still struggle to generate complex inputs from the provided dictionaries or specifications. Changhua Luo, Penghui Li 0001, Wei Meng 0001, Chao Zhang 0008 |
CCS | 2 |
| 2024 | Holistic Concolic Execution for Dynamic Web Applications via Symbolic Interpreter AnalysisabstractSymbolic execution for dynamic web applications is challenging due to their multilingual nature. Prior solutions often fall short in limited syntax support and excessive engineering costs. We propose a novel approach called symbolic interpreter analysis (SIA) for web applications written in interpreted languages. SIA tackles the limitations by leveraging the comprehensive syntax support of language interpreters and incorporating established engineering from existing symbolic execution engines. Since web application logic is handled by the interpreter, SIA leverages an off-the-shelf symbolic execution engine to analyze the corresponding interpreter code to symbolically comprehend the behavior of the web application. Indeed, SIA entails solving several technical challenges in web application symbolic execution such as web application exploration, database interactions, etc.We have implemented our approach in SymPHP, a concolic execution engine for PHP-based web applications. Our extensive evaluation shows that SymPHP could effectively explore web application code with comprehensive PHP syntax support and high code coverage. It achieved high code coverage and successfully identified 77.23% of known vulnerabilities in our dataset, significantly outperforming prior approaches. The hybrid fuzzing framework built atop SymPHP significantly boosted fuzzing and detected ten new vulnerabilities. Penghui Li 0001, Wei Meng 0001, Mingxue Zhang 0001, Chenlin Wang, Changhua Luo |
SP | 1 |
| 2024 | SDFuzz: Target States Driven Directed Fuzzing
Penghui Li 0001, Wei Meng 0001, Chao Zhang 0008 |
USENIX Security Symposium | 1 |
| 2023 | SelectFuzz: Efficient Directed Fuzzing with Selective Path ExplorationabstractDirected grey-box fuzzers specialize in testing specific target code. They have been applied to many security applications such as reproducing known crashes and detecting vulnerabilities caused by incomplete patches. However, existing directed fuzzers favor the inputs discovering new code regardless whether the newly uncovered code is relevant to the target code or not. As a result, the fuzzers would extensively explore irrelevant code and suffer from low efficiency.In this paper, we distinguish relevant code in the target program from the irrelevant one that does not help trigger the vulnerabilities in target code. We present SelectFuzz, a new directed fuzzer that selectively explores relevant program paths for efficient crash reproduction and vulnerability detection. It identifies two types of relevant code—path-divergent code and data-dependent code, that respectively captures the control-and data- dependency with the target code. It then selectively instruments and explores only the relevant code blocks. We also propose a new distance metric that accurately measures the reaching probability of different program paths and inputs.We evaluated SelectFuzz with real-world vulnerabilities in sets of diverse programs. SelectFuzz significantly outperformed a baseline directed fuzzer by up to 46.31×, and performed the best in the Google Fuzzer Test Suite. Our experiments also demonstrated that SelectFuzz and the existing techniques such as path pruning are complementary. Finally, with SelectFuzz, we detected 14 previously unknown vulnerabilities—including 6 new CVE IDs—in well tested real-world software. Our report has led to the fix of 11 vulnerabilities. Changhua Luo, Wei Meng 0001, Penghui Li 0001 |
SP | 3 |
| 2023 | DDRace: Finding Concurrency UAF Vulnerabilities in Linux Drivers with Directed Fuzzing
Ming Yuan 0003, Bodong Zhao, Penghui Li 0001, Jiashuo Liang, Xinhui Han, Xiapu Luo, Chao Zhang 0008 |
USENIX Security Symposium | 3 |
| 2023 | Testing Graph Database Systems via Graph-Aware Metamorphic RelationsabstractGraph database systems (GDBs) have supported many important real-world applications such as social networks, logistics, and path planning. Meanwhile, logic bugs are also prevalent in GDBs, leading to incorrect results and severe consequences. However, the logic bugs largely cannot be revealed by prior solutions which are unaware of the graph native structures of the graph data. In this paper, we propose Gamera (Graph-aware metamorphic relations), a novel metamorphic testing approach to uncover unknown logic bugs in GDBs. We design three classes of novel graph-aware Metamorphic Relations (MRs) based on the graph native structures. Gamera would generate a set of queries according to the graph-aware MRs to test diverse and complex GDB operations, and check whether the GDB query results conform to the chosen MRs. We thoroughly evaluated the effectiveness of Gamera on seven widely-used GDBs such as Neo4j and OrientDB. Gamera was highly effective in detecting logic bugs in GDBs. In total, it detected 39 logic bugs, of which 15 bugs have been confirmed, and three bugs have been fixed. Our experiments also demonstrated that Gamera significantly outperformed prior solutions including Grand, GD-smith and GDBMeter. Gamera has been well-recognized by GDB developers and we open-source our prototype implementation to contribute to the community. Zeyang Zhuang, Penghui Li 0001, Pingchuan Ma 0004, Wei Meng 0001, Shuai Wang 0011 |
Proc. VLDB Endow. | 2 |
| 2022 | TChecker: Precise Static Inter-Procedural Analysis for Detecting Taint-Style Vulnerabilities in PHP ApplicationsabstractPHP applications provide various interfaces for end-users to interact with on the Web. They thus are prone to taint-style vulnerabilities such as SQL injection and cross-site scripting. For its high efficiency, static taint analysis is widely adopted to detect taint-style vulnerabilities before application deployment. Unfortunately, due to the high complexity of the PHP language, implementing a precise static taint analysis is difficult. The existing taint analysis solutions suffer from both high false positives and high false negatives because of their incomprehensive inter-procedural analysis and a variety of implementation issues. Changhua Luo, Penghui Li 0001, Wei Meng 0001 |
CCS | 2 |
| 2022 | SEDiff: scope-aware differential fuzzing to test internal function models in symbolic executionabstractSymbolic execution has become a foundational program analysis technique. Performing symbolic execution unavoidably encounters internal functions (e.g., library functions) that provide basic operations such as string processing. Many symbolic execution engines construct internal function models that abstract function behaviors for scalability and compatibility concerns. Due to the high complexity of constructing the models, developers intentionally summarize only partial behaviors of a function, namely modeled functionalities, in the models. The correctness of the internal function models is critical because it would impact all applications of symbolic execution, e.g., bug detection and model checking. Penghui Li 0001, Wei Meng 0001, Kangjie Lu |
ESEC/SIGSOFT FSE | 1 |
| 2021 | Understanding and Detecting Performance Bugs in Markdown CompilersabstractMarkdown compilers are widely used for translating plain Markdown text into formatted text, yet they suffer from performance bugs that cause performance degradation and resource exhaustion. Currently, there is little knowledge and understanding about these performance bugs in the wild. In this work, we first conduct a comprehensive study of known performance bugs in Markdown compilers. We identify that the ways Markdown compilers handle the language’s context-sensitive features are the dominant root cause of performance bugs. To detect unknown performance bugs, we develop MdPerfFuzz, a fuzzing framework with a syntax-tree based mutation strategy to efficiently generate test cases to manifest such bugs. It equips an execution trace similarity algorithm to de-duplicate the bug reports. With MdPerfFuzz, we successfully identified 216 new performance bugs in real-world Markdown compilers and applications. Our work demonstrates that the performance bugs are a common, severe, yet previously overlooked security problem. Penghui Li 0001, Yinxi Liu, Wei Meng 0001 |
ASE | 1 |
| 2021 | LChecker: Detecting Loose Comparison Bugs in PHPabstractWeakly-typed languages such as PHP support loosely comparing two operands by implicitly converting their types and values. Such a language feature is widely used but can also pose severe security threats. In certain conditions, loose comparisons can cause unexpected results, leading to authentication bypass and other functionality problems. Penghui Li 0001, Wei Meng 0001 |
WWW | 1 |
| 2021 | On the Feasibility of Automated Built-in Function Modeling for PHP Symbolic ExecutionabstractSymbolic execution has been widely applied in detecting vulnerabilities in web applications. Modeling language-specific built-in functions is essential for symbolic execution. Since built-in functions tend to be complicated and are typically implemented in low-level languages, a common strategy is to manually translate them into the SMT-LIB language for constraint solving. Such translation requires an excessive amount of human effort and deep understandings of the function behaviors. Incorrect translation can invalidate the final results. This problem aggravates in PHP applications because of their cross-language nature, i.e., , the built-in functions are written in C, but the rest code is in PHP. Penghui Li 0001, Wei Meng 0001, Kangjie Lu, Changhua Luo |
WWW | 1 |