EDBT 2026 Demo / reviewers in the wild / expert
Yeting Li
dblp:185/7953
· DBLP profile ↗
36ranked-venue papers
15as first author
23since 2021 · last 2026
0000-0003-0991-4231ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 12 · 2 first-author · 11 since 2021Databases, data management, data science and information retrieval · 11 · 9 first-authorSecurity and privacy · 10 · 3 first-author · 10 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorSystems, architecture and hardware · 1 · 1 since 2021Theory of computation · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LifeFuzz: Lifecycle-Guided Fuzzing for Windows Driver Cross-Handler VulnerabilitiesabstractThird-party Windows drivers expose a critical attack surface. However, vulnerabilities that require cross-handler I/O Control (IOCTL) sequences remain hard to find, despite their prevalence, and often lead to privilege escalation. Static analysis suffers from high false positives, path explosion, and complex resource modeling. Meanwhile, dynamic fuzzers often exercise handlers in isolation or combine them randomly, leaving implicit state dependencies unchecked. To address the gap, we present LifeFuzz, a lifecycle-guided fuzzing framework that models global-variable lifecycles to construct dependency-respecting IOCTL sequences. Specifically, it identifies variable operations across handlers, preserves seeds that affect driver state, and then combines them into meaningful sequences. Consequently, LifeFuzz explores deep paths unreachable for existing fuzzers. We evaluate LifeFuzz on 26 Windows WDM drivers. It discovers 86 vulnerabilities, including 32 cross-handler cases, with six assigned CVE IDs. Moreover, it finds 357% more cross-handler vulnerabilities than msFuzz and achieves 19.2% higher average coverage. Overall, 37% of discovered vulnerabilities require cross-handler interactions, thereby validating lifecycle-aware, cross-handler fuzzing for driver security. Chendong Yu, Yuekang Li, Yang Xiao 0011, Jie Lu 0009, Yeting Li, Defang Bo, Wei Huo 0005 |
EuroSys | 5 |
| 2026 | User-Space Dependency-Aware Rehosting for Linux-Based Firmware Binaries
Cen Zhang, Yaowen Zheng, Puzhuo Liu, Jian Zhang 0087, Yeting Li, Yang Liu 0003, Limin Sun 0001 |
NDSS | 6 |
| 2026 | Through the Authentication Maze: Detecting Authentication Bypass Vulnerabilities in Firmware Binaries
Nanyu Zhong, Yuekang Li, Yanyan Zou 0002, Jiaxu Zhao 0004, Jinwei Dong, Yang Xiao 0011, Bingwei Peng, Yeting Li, Wei Huo 0005 |
NDSS | 8 |
| 2025 | CodeCleaner: Mitigating Data Contamination for LLM BenchmarkingabstractData contamination presents a critical barrier preventing widespread industrial adoption of advanced software engineering techniques that leverage large language models (LLMs).This phenomenon occurs when evaluation data inadvertently overlaps with the public code repositories used to train LLMs, severely undermining the credibility of performance evaluations.Code refactoring, which comprises code restructuring and variable renaming, has emerged as a promising measure to mitigate data contamination.However, the lack of automated code refactoring tools and scientifically validated refactoring techniques has hampered widespread industrial implementation.To bridge the gap, this paper presents the first systematic study to examine the efficacy of code refactoring operators at multiple scales (method-level, class-level, and crossclass level) and in different programming languages.We develop CodeCleaner, including 11 operators for Python in multiple scales and 4 for Java.We elaborate on the rationale for why these operators could work to resolve data contamination and use both data-wise (e.g., N-gram matching overlap ratio) and model-wise metrics (e.g., perplexity) to quantify the efficacy after operators are applied.A drop of 75% overlap ratio is found when applying all operators in CodeCleaner, demonstrating their effectiveness in addressing data contamination.Besides, we migrate four operators to Java, showing their generalizability to another language.We also observed an average of 19% decrease in LLMs' performance after applying our operators.We make CodeCleaner online available at https://github.com/ArabelaTso/CodeCleaner-v1 to facilitate further studies on mitigating LLM data contamination. Jialun Cao, Songqiang Chen, Wuqi Zhang, Hau Ching Lo, Yeting Li, Shing-Chi Cheung |
Internetware | 5 |
| 2025 | Understanding Resource Injection Vulnerabilities in Kubernetes EcosystemsabstractCloud-native technologies have revolutionized application development, with Kubernetes emerging as the de facto standard platform for containerization and orchestration. Kubernetes manages applications through API objects called resources, where users declare desired states via resource definitions that are processed by controllers to reconcile system discrepancies. However, this resource-based architecture introduces resource injection vulnerabilities, where controllers perform privileged operations using user-controllable fields without adequate validation. Attackers can exploit these weaknesses by injecting malicious content into resource fields to achieve unauthorized access and privilege escalation.In this paper, we conduct the first comprehensive study on 125 resource injection vulnerabilities from 8,306 Kubernetes-related vulnerabilities across common databases. For all studied vulnerabilities, we investigate their vulnerable fields, root causes, privileged operations, exploitation conditions, and fixing strategies. Our study reveals many interesting findings that can guide the detection and mitigation of resource injection vulnerabilities, as well as the development of more secure cloud-native applications. Defang Bo, Jie Lu 0009, Feng Li 0045, Jingting Chen, Jinchen Wang, Chendong Yu, Yeting Li, Wei Huo 0005 |
ASE | 7 |
| 2025 | Vulnerability-Affected Versions Identification: How Far Are We?abstractIdentifying which software versions are affected by a vulnerability is critical for patching, risk mitigation. Despite a growing body of tools, their real-world effectiveness remains unclear due to narrow evaluation scopes—often limited to early SZZ variants, outdated techniques, and small or coarse-grained datasets. In this paper, we present the first comprehensive empirical study of vulnerability-affected versions identification. We curate a high-quality benchmark of 1,128 real-world C/C++ vulnerabilities and systematically evaluate 12 representative tools from both tracing and matching paradigms across four dimensions: effectiveness at both vulnerability and version levels, root causes of false positives and negatives, sensitivity to patch characteristics, and ensemble potential. Our findings reveal fundamental limitations: no tool exceeds 45.0% accuracy, with key challenges stemming from heuristic dependence, limited semantic reasoning, and rigid matching logic. Patch structures such as add-only and cross-file changes further hinder performance. Although ensemble strategies can improve results by up to 10.1%, overall accuracy remains below 60.0%, highlighting the need for fundamentally new approaches. Moreover, our study offers actionable insights to guide tool development, combination strategies, and future research in this critical area. Finally, we release the replicated code and benchmark on our website to encourage future contributions. Xingchu Chen, Jialun Cao, Yang Xiao 0011, Xinyue Cai, Yeting Li, Tianqi Sun, Haiming Chen 0001, Wei Huo 0005 |
ASE | 6 |
| 2025 | A Large Scale Study of AI-based Binary Function Similarity Detection Techniques for Security Researchers and PractitionersabstractBinary Function Similarity Detection (BFSD) is a foundational technique in software security, underpinning a wide range of applications including vulnerability detection, malware analysis. Recent advances in AI-based BFSD tools have led to significant performance improvements. However, existing evaluations of these tools suffer from three key limitations: a lack of in-depth analysis of performance-influencing factors, an absence of realistic application analysis, and reliance on small-scale or low-quality datasets.In this paper, we present the first large-scale empirical study of AI-based BFSD tools to address these gaps. We construct two high-quality and diverse datasets: BinAtlas, comprising 12,453 binaries and over 7 million functions for capability evaluation; and BinAres, containing 12,291 binaries and 54 real-world 1-day vulnerabilities for evaluating vulnerability detection performance in practical IoT firmware settings. Using these datasets, we evaluate nine representative BFSD tools, analyze the challenges and limitations of existing BFSD tools, and investigate the consistency among BFSD tools. We also propose an actionable strategy for combining BFSD tools to enhance overall performance (an improvement of 13.4%). Our study not only advances the practical adoption of BFSD tools but also provides valuable resources and insights to guide future research in scalable and automated binary similarity detection. Yang Xiao 0011, Yuekang Li, Zhengzi Xu, Sihao Qiu, Keyu Qi, Yeting Li, Xingchu Chen, Yanyan Zou 0002, Yang Liu 0003, Wei Huo 0005 |
ASE | 9 |
| 2025 | From Constraints to Cracks: Constraint Semantic Inconsistencies as Vulnerability Beacons for Embedded Systems
Jiaxu Zhao 0004, Yuekang Li, Yanyan Zou 0002, Yang Xiao 0011, Naijia Jiang, Yeting Li, Nanyu Zhong, Bingwei Peng, Kunpeng Jian, Wei Huo 0005 |
USENIX Security Symposium | 6 |
| 2025 | VULCANBOOST: Boosting ReDoS Fixes through Symbolic Representation and Feature Normalization
Yeting Li, Yecheng Sun, Zhiwu Xu 0001, Haiming Chen 0001, Xinyi Wang 0013, Hengyu Yang, Huina Chao, Cen Zhang, Yang Xiao 0011, Yanyan Zou 0002, Feng Li 0045, Wei Huo 0005 |
USENIX Security Symposium | 1 |
| 2025 | ZIPPER: Static Taint Analysis for PHP Applications with Precision and Efficiency
Xinyi Wang 0013, Yeting Li, Jie Lu 0009, Shizhe Cui, Chenghang Shi, Qin Mai, Yunpei Zhang, Yang Xiao 0011, Feng Li 0045, Wei Huo 0005 |
USENIX Security Symposium | 2 |
| 2024 | Semantic-Enhanced Static Vulnerability Detection in Baseband FirmwareabstractCellular network is the infrastructure of mobile communication. Baseband firmware, which carries the implementation of cellular network, has critical security impact on its vulnerabilities. To handle the inherent complexity in cellular communication, cellular protocols are usually implemented as message-centric systems, containing the common message processing phase and message specific handling phase. Though the latter takes most of the code (99.67%) and exposed vulnerabilities (74%), it is rather under-studied: existing detectors either cannot sufficiently analyze it or focused on analyzing the former phase. Cen Zhang, Feng Li 0045, Yeting Li, Jian Wang 0067, Lanlan Zhan, Yang Liu 0003, Wei Huo 0005 |
ICSE | 4 |
| 2024 | How Effective Are They? Exploring Large Language Model Based Fuzz Driver GenerationabstractFuzz drivers are essential for library API fuzzing. However, automatically generating fuzz drivers is a complex task, as it demands the creation of high-quality, correct, and robust API usage code. An LLM-based (Large Language Model) approach for generating fuzz drivers is a promising area of research. Unlike traditional program analysis-based generators, this text-based approach is more generalized and capable of harnessing a variety of API usage information, resulting in code that is friendly for human readers. However, there is still a lack of understanding regarding the fundamental issues on this direction, such as its effectiveness and potential challenges. To bridge this gap, we conducted the first in-depth study targeting the important issues of using LLMs to generate effective fuzz drivers. Our study features a curated dataset with 86 fuzz driver generation questions from 30 widely-used C projects. Six prompting strategies are designed and tested across five state-of-the-art LLMs with five different temperature settings. In total, our study evaluated 736,430 generated fuzz drivers, with 0.85 billion token costs ($8,000+ charged tokens). Additionally, we compared the LLM-generated drivers against those utilized in industry, conducting extensive fuzzing experiments (3.75 CPU-year). Our study uncovered that: 1) While LLM-based fuzz driver generation is a promising direction, it still encounters several obstacles towards practical applications; 2) LLMs face difficulties in generating effective fuzz drivers for APIs with intricate specifics. Three featured design choices of prompt strategies can be beneficial: issuing repeat queries, querying with examples, and employing an iterative querying process; 3) While LLM-generated drivers can yield fuzzing outcomes that are on par with those used in the industry, there are substantial opportunities for enhancement, such as extending contained API usage, or integrating semantic oracles to facilitate logical bug detection. Our insights have been implemented to improve the OSS-Fuzz-Gen project, facilitating practical fuzz driver generation in industry. Cen Zhang, Yaowen Zheng, Mingqiang Bai, Yeting Li, Wei Ma 0014, Xiaofei Xie, Yuekang Li, Limin Sun 0001, Yang Liu 0003 |
ISSTA | 4 |
| 2024 | File Hijacking Vulnerability: The Elephant in the Room
Chendong Yu, Yang Xiao 0011, Jie Lu 0009, Yuekang Li, Yeting Li, Lian Li 0002, Jian Wang 0067, Defang Bo, Wei Huo 0005 |
NDSS | 5 |
| 2024 | Fuzzing for Stateful Protocol Implementations: Are We There Yet?
Kunpeng Jian, Yanyan Zou 0002, Yeting Li, Jialun Cao, Wei Huo 0005 |
TASE | 3 |
| 2024 | Leveraging Semantic Relations in Code and Data to Enhance Taint Analysis of Embedded Systems
Jiaxu Zhao 0004, Yuekang Li, Yanyan Zou 0002, Zhaohui Liang, Yang Xiao 0011, Yeting Li, Bingwei Peng, Nanyu Zhong, Xinyi Wang 0013, Wei Huo 0005 |
USENIX Security Symposium | 6 |
| 2023 | ACETest: Automated Constraint Extraction for Testing Deep Learning OperatorsabstractDeep learning (DL) applications are prevalent nowadays as they can help with multiple tasks. DL libraries are essential for building DL applications. Furthermore, DL operators are the important building blocks of the DL libraries, that compute the multi-dimensional data (tensors). Therefore, bugs in DL operators can have great impacts. Testing is a practical approach for detecting bugs in DL operators. In order to test DL operators effectively, it is essential that the test cases pass the input validity check and are able to reach the core function logic of the operators. Hence, extracting the input validation constraints is required for generating high-quality test cases. Existing techniques rely on either human effort or documentation of DL library APIs to extract the constraints. They cannot extract complex constraints and the extracted constraints may differ from the actual code implementation. To address the challenge, we propose ACETest, a technique to automatically extract input validation constraints from the code to build valid yet diverse test cases which can effectively unveil bugs in the core function logic of DL operators. For this purpose, ACETest can automatically identify the input validation code in DL operators, extract the related constraints and generate test cases according to the constraints. The experimental results on popular DL libraries, TensorFlow and PyTorch, demonstrate that ACETest can extract constraints with higher quality than state-of-the-art (SOTA) techniques. Moreover, ACETest is capable of extracting 96.4% more constraints and detecting 1.95 to 55 times more bugs than SOTA techniques. In total, we have used ACETest to detect 108 previously unknown bugs on TensorFlow and PyTorch, with 87 of them confirmed by the developers. Lastly, five of the bugs were assigned with CVE IDs due to their security impacts. Yang Xiao 0011, Yuekang Li, Yeting Li, Dongsong Yu, Chendong Yu, Hui Su, Wei Huo 0005 |
ISSTA | 4 |
| 2023 | ASTER: Automatic Speech Recognition System Accessibility Testing for StutterersabstractThe popularity of automatic speech recognition (ASR) systems nowadays leads to an increasing need for improving their accessibility. Handling stuttering speech is an important feature for accessible ASR systems. To improve the accessibility of ASR systems for stutterers, we need to expose and analyze the failures of ASR systems on stuttering speech. The speech datasets recorded from stutterers are not diverse enough to expose most of the failures. Furthermore, these datasets lack ground truth information about the non-stuttered text, rendering them unsuitable as comprehensive test suites. Therefore, a methodology for generating stuttering speech as test inputs to test and analyze the performance of ASR systems is needed. However, generating valid test inputs in this scenario is challenging. The reason is that although the generated test inputs should mimic how stutterers speak, they should also be diverse enough to trigger more failures. To address the challenge, we propose Aster, a technique for automatically testing the accessibility of ASR systems. Aster can generate valid test cases by injecting five different types of stuttering. The generated test cases can both simulate realistic stuttering speech and expose failures in ASR systems. Moreover, Aster can further enhance the quality of the test cases with a multi-objective optimization-based seed updating algorithm. We implemented Aster as a framework and evaluated it on four open-source ASR models and three commercial ASR systems. We conduct a comprehensive evaluation of Aster and find that it significantly increases the word error rate, match error rate, and word information loss in the evaluated ASR systems. Additionally, our user study demonstrates that the generated stuttering audio is indistinguishable from real-world stuttering audio clips. Yi Liu 0069, Yuekang Li, Gelei Deng, Felix Juefei-Xu, Yao Du 0002, Cen Zhang, Yeting Li, Lei Ma 0003, Yang Liu 0003 |
ASE | 8 |
| 2023 | Effective ReDoS Detection by Principled Vulnerability Modeling and Exploit GenerationabstractRegular expression Denial-of-Service (ReDoS) is one kind of algorithmic complexity attack. For a vulnerable regex, attackers can craft certain strings to trigger the super-linear worst-case matching time, which causes denial-of-service to regex engines. Various ReDoS detection approaches have been proposed recently. Among them, hybrid approaches which absorb the advantages of both static and dynamic approaches have shown their performance superiority. However, two key challenges still hinder the effectiveness of the detection: 1) Existing modelings summarize localized vulnerability patterns based on partial features of the vulnerable regex; 2) Existing attack string generation strategies are ineffective since they neglected the fact that non-vulnerable parts of the regex may unexpectedly invalidate the attack string (we name this kind of invalidation as disturbance.)Rengar is our hybrid ReDoS detector with new vulnerability modeling and disturbance free attack string generator. It has the following key features: 1) Benefited by summarizing patterns from full features of the vulnerable regex, its modeling is a more precise interpretation of the root cause of ReDoS vulnerability. The modeling is more descriptive and precise than the union of existing modelings while keeping conciseness; 2) For each vulnerable regex, its generator automatically checks all potential disturbances and composes generation constraints to avoid possible disturbances.Compared with nine state-of-the-art tools, Rengar detects not only all vulnerable regexes they found but also 3 – 197 times more vulnerable regexes. Besides, it saves 57.41% – 99.83% average detection time compared with tools containing a dynamic validation process. Using Rengar, we have identified 69 zero-day vulnerabilities (21 CVEs) affecting popular projects which have more than dozens of millions weekly download count. Xinyi Wang 0013, Cen Zhang, Yeting Li, Zhiwu Xu 0001, Shuailin Huang, Yi Liu 0069, Yican Yao, Yang Xiao 0011, Yanyan Zou 0002, Yang Liu 0003, Wei Huo 0005 |
SP | 3 |
| 2023 | Learning Disjunctive Multiplicity Expressions and Disjunctive Generalize Multiplicity Expressions From Both Positive and Negative ExamplesabstractAbstract The presence of a schema for eXtensible Markup Language (XML) documents has numerous advantages. Unfortunately, many XML documents in practice are not accompanied by a (valid) schema. Therefore, it is essential to devise algorithms to infer schemas from XML documents, where the fundamental task is learning regular expressions. In this paper, we focus on the learning of disjunctive multiplicity expressions (DMEs), a subclass of regular expressions that are particularly suitable to specify unordered models and have been used as the foundation of the schemas for unordered XML. Previous work for learning DME lacks inference algorithms that support positive and negative examples. Further, presently there has been no algorithm can learn DMEs extended with numeric occurrences. We address these challenges in the present paper and first propose a novel algorithm to learn DMEs from positive and negative examples by using genetic algorithms and parallel techniques. Then we extend DMEs to disjunctive generalized multiplicity expressions (DGMEs), which allow numeric occurrences and develop an algorithm to learn DGMEs from positive and negative examples. Finally, experimental results show that with only positive examples, our algorithm can generate a DME with an acceptable learning time, which can accept all positive examples, and when given both positive and negative examples, we can learn DMEs or DGMEs with high accuracy. Yeting Li, Haiming Chen 0001 |
Comput. J. | 1 |
| 2022 | RegexScalpel: Regular Expression Denial of Service (ReDoS) Defense by Localize-and-Fix
Yeting Li, Yecheng Sun, Zhiwu Xu 0001, Jialun Cao, Yuekang Li, Rongchen Li, Haiming Chen 0001, Shing-Chi Cheung, Yang Liu 0003, Yang Xiao 0011 |
USENIX Security Symposium | 1 |
| 2022 | SemMT: A Semantic-Based Testing Approach for Machine Translation SystemsabstractMachine translation has wide applications in daily life. In mission-critical applications such as translating official documents, incorrect translation can have unpleasant or sometimes catastrophic consequences. This motivates recent research on the testing methodologies for machine translation systems. Existing methodologies mostly rely on metamorphic relations designed at the textual level (e.g., Levenshtein distance) or syntactic level (e.g., distance between grammar structures) to determine the correctness of translation results. However, these metamorphic relations do not consider whether the original and the translated sentences have the same meaning (i.e., semantic similarity). To address this problem, in this article we propose SemMT, an automatic testing approach for machine translation systems based on semantic similarity checking. SemMT applies round-trip translation and measures the semantic similarity between the original and the translated sentences. Our insight is that the semantics concerning logical relations and quantifiers in sentences can be captured by regular expressions (or deterministic finite automata) where efficient semantic equivalence/similarity checking algorithms can be applied. Leveraging the insight, we propose three semantic similarity metrics and implement them in SemMT. We compared SemMT with related state-of-the-art testing techniques, demonstrating the effectiveness of mistranslation detection. The experiment results show that SemMT outperforms existing metrics, achieving an increase of 34.2% and 15.4% on accuracy and F-score, respectively. We also study the possibility of further enhancing the performance by combining various metrics. Finally, we discuss a solution to locate the suspicious trip in round-trip translation, which provides hints for bug diagnosis. Jialun Cao, Meiziniu Li, Yeting Li, Ming Wen 0001, Shing-Chi Cheung, Haiming Chen 0001 |
ACM Trans. Softw. Eng. Methodol. | 3 |
| 2021 | TRANSREGEX: Multi-modal Regular Expression Synthesis by Generate-and-RepairabstractSince regular expressions (abbrev. regexes) are difficult to understand and compose, automatically generating regexes has been an important research problem. This paper introduces TransRegex, for automatically constructing regexes from both natural language descriptions and examples. To the best of our knowledge, TransRegex is the first to treat the NLP-and-example-based regex synthesis problem as the problem of NLP-based synthesis with regex repair. For this purpose, we present novel algorithms for both NLP-based synthesis and regex repair. We evaluate TransRegex with ten relevant state-of-the-art tools on three publicly available datasets. The evaluation results demonstrate that the accuracy of our TransRegex is 17.4%, 35.8% and 38.9% higher than that of NLP-based approaches on the three datasets, respectively. Furthermore, TransRegex can achieve higher accuracy than the state-of-the-art multi-modal techniques with 10% to 30% higher accuracy on all three datasets. The evaluation results also indicate TransRegex utilizing natural language and examples in a more effective way. Yeting Li, Shuaimin Li, Zhiwu Xu 0001, Jialun Cao, Haiming Chen 0001, Shing-Chi Cheung |
ICSE | 1 |
| 2021 | ReDoSHunter: A Combined Static and Dynamic Approach for Regular Expression DoS Detection
Yeting Li, Jialun Cao, Zhiwu Xu 0001, Qiancheng Peng, Haiming Chen 0001, Shing-Chi Cheung |
USENIX Security Symposium | 1 |
| 2020 | FlashSchema: Achieving High Quality XML Schemas with Powerful Inference Algorithms and Large-scale Schema DataabstractGetting high quality XML schemas to avoid or reduce application risks is an important problem in practice, for which some important aspects have yet to be addressed satisfactorily in existing work. In this paper, we propose a tool FlashSchema for high quality XML schema design, which supports both one-pass and interactive schema design and schema recommendation. To the best of our knowledge, no other existing tools support interactive schema design and schema recommendation. One salient feature of our work is the design of algorithms to infer k-occurrence interleaving regular expressions, which are not only more powerful in model capacity, but also more efficient. Additionally, such algorithms form the basis of our interactive schema design. The other feature is that, starting from large-scale schema data that we have harvested from the Web, we devise a new solution for type inference, as well as propose schema recommendation for schema design. Finally, we conduct a series of experiments on two XML datasets, comparing with 9 state-of-the-art algorithms and open-source tools in terms of running time, preciseness, and conciseness. Experimental results show that our work achieves the highest level of preciseness and conciseness within only a few seconds. Experimental results and examples also demonstrate the effectiveness of our type inference and schema recommendation methods. Yeting Li, Jialun Cao, Haiming Chen 0001, Tingjian Ge, Zhiwu Xu 0001, Qiancheng Peng |
ICDE | 1 |
| 2020 | FlashRegex: Deducing Anti-ReDoS Regexes from ExamplesabstractRegular expressions (regexes) are widely used in different fields of computer science such as programming languages, string processing and databases. However, existing tools for synthesizing or repairing regexes were not designed to be resilient to Regex Denial of Service (ReDoS) attacks. Specifically, if a regex has super-linear (SL) worst-case complexity, an attacker could provide carefully-crafted inputs to launch ReDoS attacks. Therefore, in this paper, we propose a programming-by-example framework, FlashRegex, for generating anti-ReDoS regexes by either synthesizing or repairing from given examples. It is the first framework that integrates regex synthesis and repair with the awareness of ReDoS-vulnerabilities. We present novel algorithms to deduce anti-ReDoS regexes by reducing the ambiguity of these regexes and by using Boolean Satisfiability (SAT) or Neighborhood Search (NS) techniques. We evaluate FlashRegex with five related state-of-the-art tools. The evaluation results show that our work can effectively and efficiently generate anti-ReDoS regexes from given examples, and also reveal that existing synthesis and repair tools have neglected ReDoS-vulnerabilities of regexes. Specifically, the existing synthesis and repair tools generated up to 394 ReDoS-vulnerable regex within few seconds to more than one hour, while FlashRegex generated no SL regex within around five seconds. Furthermore, the evaluation results on ReDoS-vulnerable regex repair also show that FlashRegex has better capability than existing repair tools and even human experts, achieving 4 more ReDoS-invulnerable regex after repair without trimming and resorting, highlighting the usefulness of FlashRegex in terms of the generality, automation and user-friendliness. Yeting Li, Zhiwu Xu 0001, Jialun Cao, Haiming Chen 0001, Tingjian Ge, Shing-Chi Cheung, Haoren Zhao |
ASE | 1 |
| 2020 | Inferring Restricted Regular Expressions with Interleaving from Positive and Negative Samples
Yeting Li, Haiming Chen 0001, Jianzhao Zhang |
PAKDD (2) | 1 |
| 2019 | Learning k-Occurrence Regular Expressions with Interleaving
Yeting Li, Jialun Cao, Haiming Chen 0001 |
DASFAA (2) | 1 |
| 2019 | Learning k-Occurrence Regular Expressions from Positive and Negative Samples
Yeting Li, Xiaoying Mou, Haiming Chen 0001 |
ER | 1 |
| 2019 | Context-Free Grammars for Deterministic Regular Expressions with Interleaving
Xiaoying Mou, Haiming Chen 0001, Yeting Li |
ICTAC | 3 |
| 2019 | An effective algorithm for learning single occurrence regular expressions with interleavingabstractThe advantages offered by the presence of a schema are numerous. However, many XML documents in practice are not accompanied by a (valid) schema, making schema inference an attractive research problem. The fundamental task in XML schema learning is inferring restricted subclasses of regular expressions. Most previous work either lacks support for interleaving or only has limited support for interleaving. In this paper, we first propose a new subclass Single Occurrence Regular Expressions with Interleaving (SOIRE), which has unrestricted support for interleaving. Then, based on single occurrence automaton and maximum independent set, we propose an algorithm iSOIRE to infer SOIREs. Finally, we further conduct a series of experiments on real datasets to evaluate the effectiveness of our work, comparing with both ongoing learning algorithms in academia and industrial tools in real-world. The results reveal the practicability of SOIRE and the effectiveness of iSOIRE, showing the high preciseness and conciseness of our work. Yeting Li, Haiming Chen 0001 |
IDEAS | 1 |
| 2019 | A Large-Scale Repository of Deterministic Regular Expression Patterns and Its Applications
Haiming Chen 0001, Yeting Li, Chunmei Dong, Xinyu Chu, Xiaoying Mou, Weidong Min |
PAKDD (3) | 2 |
| 2018 | Learning Concise Relax NG Schemas Supporting Interleaving from XML Documents
Yeting Li, Xiaoying Mou, Haiming Chen 0001 |
ADMA | 1 |
| 2018 | Learning Restricted Regular Expressions with Interleaving from XML Data
Yeting Li, Xiaoying Mou, Haiming Chen 0001 |
ER | 1 |
| 2018 | Practical Study of Deterministic Regular Expressions from Large-scale XML and Schema DataabstractRegular expressions are a fundamental concept in computer science and widely used in various applications. In this paper we focused on deterministic regular expressions (DREs). Considering that researchers did not have large datasets as evidence before, we first harvested a large corpus of real data from the Web then conducted a practical study to investigate the usage of DREs. One feature of our work is that the data set is sufficiently large compared with previous work, which is obtained using several data collection strategies we proposed. The results show more than 98% of expressions in Relax NG are DRE, and more than 56% of expressions from RegExLib are DRE, while both Relax NG and RegExLib do not have the determinism constraint. These observations indicate that DREs are commonly used in practice. The results also show further study of subclasses of DREs is necessary. As far as we know, we are the first to analyze the determinism and the subclasses of DREs of Relax NG and RegExLib, and give these results. Furthermore, we give some discussions and applications of the data set. We find current research in new subclasses of DREs is insufficient, therefore it is necessary to do further study. We also analyze the referencing relationships among XSDs and define SchemaRank, which can be used in XML Schema design. Yeting Li, Xinyu Chu, Xiaoying Mou, Chunmei Dong, Haiming Chen 0001 |
IDEAS | 1 |
| 2018 | Inference of a Concise Regular Expression Considering Interleaving from XML Documents
Yeting Li, Fanlin Cui, Chunmei Dong, Haiming Chen 0001 |
PAKDD (2) | 2 |
| 2016 | Practical Study of Subclasses of Regular Expressions in DTD and XML Schema
Yeting Li, Feifei Peng, Haiming Chen 0001 |
APWeb (2) | 1 |