EDBT 2026 Demo / reviewers in the wild / expert
Lili Quan 0001
dblp:155/5397-1
· DBLP profile ↗
8ranked-venue papers
5as first author
7since 2021 · last 2025
0000-0003-0405-835XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 7 · 4 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SMTPRT: Performance Regression Testing and Localization for SMT Solvers Across Multiple LogicsabstractSatisfiability Modulo Theories (SMT) solvers are foundational in applications such as software verification and automated bug detection, where both correctness and performance are critical to the reliability and scalability of these systems. While existing methods predominantly focus on functional testing, performance testing has received insufficient attention, particularly regarding performance regression caused by both intentional and unintentional factors during software evolution. Current performance regression testing approaches are primarily designed for string solvers, neglecting the full spectrum of SMT theories. Furthermore, these methods often rely on time comparisons or log analysis, which makes the identification of the responsible commit slow and inefficient. To address the above issues, we propose a novel general purpose testing framework, SMTPRT, that efficiently detects and localizes performance regression issues across diverse SMT solver theories. We utilize large language models (LLMs) based on genetic algorithms (GAs) to guide the search for performance regression-inducing cases. We introduce an optimized localization technique that filters irrelevant commits using code coverage, followed by a bisecting algorithm to rapidly pinpoint the responsible commit. To thoroughly evaluate SMTPRT, we conducted extensive experiments involving six types of logic, demonstrating its superior performance. Specifically, SMTPRT successfully detected 59 regression cases, performing 3.44 times better than the baseline, and located the issues $\mathbf{1. 1 6}$ times faster than the baseline. Xiaohong Li 0001, Lili Quan 0001, Zhiping Zhou, Yao Zhang 0019 |
APSEC | 3 |
| 2025 | Dissecting Global Search: A Simple Yet Effective Method to Boost Individual Discrimination Testing and RepairabstractDeep Learning (DL) has achieved significant success in socially critical decision-making applications but often exhibits unfair behaviors, raising social concerns. Among these unfair behaviors, individual discrimination-examining inequalities between instance pairs with identical profiles differing only in sensitive attributes such as gender, race, and age-is extremely socially impactful. Existing methods have made significant and commendable efforts in testing individual discrimination before deployment. However, their efficiency and effectiveness remain limited, particularly when evaluating relatively fairer models. It remains unclear which phase of the existing testing framework (global or local) is the primary bottleneck limiting performance. Facing the above issues, we first identify that enhancing the global phase consistently improves overall testing effectiveness compared to enhancing the local phase. This motivates us to propose Genetic-Random Fairness Testing (GRFT), an effective and efficient method. In the global phase, we use a genetic algorithm to guide the search for more global discriminatory instances. In the local phase, we apply a light random search to explore the neighbors of these instances, avoiding time-consuming computations. Additionally, based on the fitness score, we also propose a straightforward yet effective repair approach. For a thorough evaluation, we conduct extensive experiments involving 6 testing methods, 5 datasets, 261 models (including 5 naively trained, 64 repaired, and 192 quantized for on-device deployment), and sixteen combinations of sensitive attributes, showing the superior performance of GRFT and our repair method. Lili Quan 0001, Tianlin Li, Xiaofei Xie, Zhenpeng Chen 0001, Sen Chen 0001, Lingxiao Jiang, Xiaohong Li 0001 |
ICSE | 1 |
| 2025 | DSBox: A Data Selection Framework for Efficient Deep Code LearningabstractDeep Learning has achieved remarkable advancements in various software engineering tasks and gained huge attention in the community. Following a data-centric paradigm, the preparation of code models requires high-quality datasets for the model training. However, constructing such datasets, especially for software tasks, is costly mainly due to the data labeling process. To address this challenge, multiple data selection methods have been proposed to identify and label data samples that are important for training. Despite this potential, unfortunately, there are limited tools to support the flexible usage of data selection methods, hindering their practical usage and future research in this domain. To bridge this gap, we introduce DSBox, a lightweight yet extensible framework that unifies 20 published selection methods, covering three categories: uncertainty, representativeness, and quality-based methods. Evaluation demonstrates that active learning methods outperform recently proposed techniques (designed for large language models) on the code vulnerability detection task. The tool, as well as a demonstration video, are available on the project website https://sites.google.com/view/dsbox2025. Lili Quan 0001 |
ASE | 2 |
| 2025 | Spec2Code: Mapping Protocol Specification to Function-Level Code ImplementationabstractProtocol specifications, defined in Request for Comments (RFCs), play a critical role in ensuring the correctness of protocol software systems. To check consistency, specification–implementation pairs are essential for testing and verification. However, existing efforts in specification-to-code mapping remain largely manual and are typically limited to the file level, lacking the fine-grained granularity needed for function-level analysis, which is crucial for effective consistency checking. To address this gap, we present Spec2Code, the first LLM-driven framework that automates fine-grained mapping from protocol specifications to function implementations.Given a RFC document and a protocol codebase, Spec2Code first performs preprocessing to extract structured specification requirements (SRs) and function-level code representations, along with contextual and dependency information. To ensure scalability, Spec2Code employs a two-stage process comprising relevance filtering and clustering-based SR organization to reduce the candidate pairs. For accuracy, Spec2Code performs fine-grained constraint-level matching on each candidate SR–function pair using LLMs, leveraging enriched context to determine whether a function fully, partially, or does not relate to an SR.We evaluate Spec2Code on real-world implementations of HTTP, TLS and BFD protocols, including Apache Httpd, Nginx, OpenSSL, BoringSSL, FRRouting, and BIRD. Experimental results show that Spec2Code outperforms four state-of-the-art baselines, achieving up to 49%, 66%, and 66% improvement in precision, recall, and F1, respectively. Additionally, Spec2Code successfully recovers the mappings for 16 known inconsistency bugs and discovers 11 previously unreported inconsistencies using an integrated lightweight consistency verifier, 5 of which have been confirmed by project developers. Yuekun Wang, Lili Quan 0001, Xiaofei Xie, Junjie Wang 0007 |
ASE | 2 |
| 2025 | TensorJSFuzz: Effective Testing of Web-Based Deep Learning Frameworks via Input-Constraint ExtractionabstractAs web applications grow in popularity, developers are increasingly integrating deep learning (DL) models into these environments. Web-based DL frameworks (e.g., TensorFlow.js) are essential for building and deploying such applications. Therefore, ensuring the quality of these frameworks is critical. While extensive testing efforts have been made for native DL frameworks such as TensorFlow and PyTorch, web-based DL frameworks have not yet undergone systematic testing. A key challenge is generating syntactically and semantically valid inputs while designing effective test oracles for web environments. To address this, we introduce TensorJSFuzz, a novel method for testing web-based DL frameworks. To ensure input quality, TensorJSFuzz extracts constraints directly from the source code of DL operators. By leveraging Large Language Models (e.g., ChatGPT) to understand the code and extract input constraints, TensorJSFuzz performs type-aware random generation coupled with dependency-aware refinement to create high-quality test inputs. These inputs are then subjected to differential testing across various backends, including CPU, TensorFlow, Wasm, and WebGL. Our experimental results show that TensorJSFuzz outperforms all baselines in generating valid inputs and identifying bugs. In particular, TensorJSFuzz successfully detected 92 bugs, with 30 already confirmed or fixed by developers, demonstrating its effectiveness in improving the robustness of web-based DL frameworks. Lili Quan 0001, Xiaofei Xie, Lingxiao Jiang, Sen Chen 0001, Junjie Wang 0007, Xiaohong Li 0001 |
WWW | 1 |
| 2025 | Evaluation and Improvement of Test Selection for Large Language ModelsabstractABSTRACT Large language models (LLMs) have recently achieved significant success across various application domains, garnering substantial attention from different communities. Unfortunately, many faults still exist that LLMs cannot properly predict. Such faults will harm the usability of LLMs in general and could introduce safety issues in reliability‐critical systems such as autonomous driving systems. How to quickly reveal these faults in real‐world datasets that LLMs could face is important but challenging. The major reason is that the ground truth is necessary but the data labeling process is heavy considering the time and human effort. To handle this problem, in the conventional deep learning testing field, test selection methods have been proposed for efficiently evaluating deep learning models by prioritizing faults. However, despite their importance, the usefulness of these methods on LLMs is unclear and underexplored. In this paper, we conduct the first empirical study to investigate the effectiveness of existing test selection methods for LLMs. We focus on classification tasks because most existing test selection methods target this setting and reliably estimating confidence scores for variable‐length outputs in generative tasks remains challenging. Experimental results on four different tasks (including both code tasks and natural language processing tasks) and four LLMs (e.g., LLaMA3 and GPT‐4) demonstrated that simple methods such as Margin perform well on LLMs, but there is still a big room for improvement. Based on the study, we further propose MuCS, a prompt Mutation‐based prediction Confidence Smoothing framework to boost the test selection capability for LLMs specifically on classification tasks. Concretely, multiple prompt mutation techniques have been proposed to help collect diverse outputs for confidence smoothing. The results show that our proposed framework significantly enhances existing methods with test relative coverage improvement by up to 70.53%. Lili Quan 0001, Maxime Cordy, Yuheng Huang 0004, Lei Ma 0003, Xiaohong Li 0001 |
J. Softw. Evol. Process. | 1 |
| 2022 | Towards Understanding the Faults of JavaScript-Based Deep Learning SystemsabstractQuality assurance is of great importance for deep learning (DL) systems, especially when they are applied in safety-critical applications. While quality issues of native DL applications have been extensively analyzed, the issues of JavaScript-based DL applications have never been systematically studied. Compared with native DL applications, JavaScript-based DL applications can run on major browsers, making the platform- and device-independent. Specifically, the quality of JavaScript-based DL applications depends on the 3 parts: the application, the third-party DL library used and the underlying DL framework (e.g., TensorFlow.js), called JavaScript-based DL system. In this paper, we conduct the first empirical study on the quality issues of JavaScript-based DL systems. Specifically, we collect and analyze 700 real-world faults from relevant GitHub repositories, including the official TensorFlow.js repository, 13 third-party DL libraries, and 58 JavaScript-based DL applications. To better understand the characteristics of these faults, we manually analyze and construct taxonomies for the fault symptoms, root causes, and fix patterns, respectively. Moreover, we also study the fault distributions of symptoms and root causes, in terms of the different stages of the development lifecycle, the 3-level architecture in the DL system, and the 4 major components of TensorFlow.js framework. Based on the results, we suggest actionable implications and research avenues that can potentially facilitate the development, testing, and debugging of JavaScript-based DL systems. Lili Quan 0001, Xiaofei Xie, Sen Chen 0001, Xiaohong Li 0001, Yang Liu 0003 |
ASE | 1 |
| 2020 | SADT: Syntax-Aware Differential Testing of Certificate Validation in SSL/TLS ImplementationsabstractThe security assurance of SSL/TLS critically depends on the correct validation of X.509 certificates. Therefore, it is important to check whether a certificate is correctly validated by the SSL/TLS implementations. Although differential testing has been proven to be effective in finding semantic bugs, it still suffers from the following limitations: (1) The syntax of test cases cannot be correctly guaranteed. (2) Current test cases are not diverse enough to cover more implementation behaviours. This paper tackles these problems by introducing SADT, a novel syntax-aware differential testing framework for evaluating the certificate validation process in SSL/TLS implementations. We first propose a tree-based mutation strategy to ensure that the generated certificates are syntactically correct, and then diversify the certificates by sharing interesting test cases among all target SSL/TLS implementations. Such generated certificates are more likely to trigger discrepancies among SSL/TLS implementations, which may indicate some potential bugs. Lili Quan 0001, Hongxu Chen 0001, Xiaofei Xie, Xiaohong Li 0001, Yang Liu 0003, Jing Hu 0007 |
ASE | 1 |