VLDB 2026 Research / reviewers in the wild / expert
Yi Liu 0069
dblp:97/4626-69
· DBLP profile ↗
31ranked-venue papers
10as first author
29since 2021 · last 2026
0000-0002-4978-127XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 17 · 9 first-author · 15 since 2021Artificial intelligence and machine learning · 6 · 6 since 2021Security and privacy · 6 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | STEAMROLLER: A Multi-Agent System for Inclusive Automatic Speech Recognition for People Who StutterabstractPeople who stutter (PWS) face systemic exclusion in today’s voice-driven society, where access to voice assistants, authentication systems, and remote work tools increasingly depends on fluent speech. Current automatic speech recognition (ASR) systems, trained predominantly on fluent speech, fail to serve millions of PWS worldwide. We present STEAMROLLER, a real time system that transforms stuttered speech into fluent output through a novel multi-stage, multi-agent AI pipeline. Our approach addresses three critical technical challenges: (1) the difficulty of direct speech to speech conversion for disfluent input, (2) semantic distortions introduced during ASR transcription of stuttered speech, and (3) latency constraints for real time communication. STEAMROLLER employs a three stage architecture comprising ASR transcription, multi-agent text repair, and speech synthesis, where our core innovation lies in a collaborative multi-agent framework that iteratively refines transcripts while preserving semantic intent. Experiments on the FluencyBank dataset and a user study demonstrates clear word error rate (WER) reduction and strong user satisfaction. Beyond immediate accessibility benefits, fine tuning ASR on STEAMROLLER repaired speech further yields additional WER improvements, creating a pathway toward inclusive AI ecosystems. Yi Liu 0069, Yuekang Li, Ling Shi 0002, Kailong Wang 0001 |
AAAI | 2 |
| 2026 | LLM-VA: Resolving the Jailbreak-Overrefusal Trade-off via Vector AlignmentabstractSafety-aligned LLMs suffer from two failure modes: jailbreak (answering harmful inputs) and over-refusal (declining benign queries).Existing vector steering methods adjust the magnitude of answer vectors, but this creates a fundamental trade-off-reducing jailbreak increases over-refusal and vice versa.We identify the root cause: LLMs encode the decision to answer (answer vector v a ) and the judgment of input safety (benign vector v b ) as nearly orthogonal directions, treating them as independent processes.We propose LLM-VA, which aligns v a with v b through closedform weight updates, making the model's willingness to answer causally dependent on its safety assessment-without fine-tuning or architectural changes.Our method identifies vectors at each layer using SVMs, selects safetyrelevant layers, and iteratively aligns vectors via minimum-norm weight modifications.Experiments on 12 LLMs demonstrate that LLM-VA achieves 11.45% higher F1 than the best baseline while preserving 95.92% utility, and automatically adapts to each model's safety bias without manual tuning. Haonan Zhang 0007, Dongxia Wang 0002, Yi Liu 0069, Wenhai Wang |
ACL (1) | 3 |
| 2026 | Uncovering Logit Suppression Vulnerabilities in LLM Safety Alignment
Yi Liu 0069, Yuekang Li, Ling Shi 0002, Gelei Deng, Shengquan Chen, Kailong Wang 0001 |
ICPR (2) | 2 |
| 2026 | SPOLRE: Semantic Preserving Object Layout Reconstruction for Image Captioning System TestingabstractImage captioning (IC) systems, including Microsoft Azure Cognitive Service, are commonly utilized to convert image content into descriptive natural language. However, inaccuracies in caption generation can lead to serious misinterpretations. Advanced testing techniques such as MetaIC and ROME have been developed to mitigate these issues, yet they encounter notable challenges. First, these strategies demand intensive labor, relying on detailed manual annotations like bounding box data of objects to create test cases. Second, the realism of the generated images is compromised, with MetaIC adding unrelated objects and ROME failing to remove objects effectively. Finally, the capability to generate diversified test suites is restricted. MetaIC is limited to only inserting specific objects to prevent overlap, whereas ROME can generate only \(3^{n}-2^{n}\) variations of test cases from an original seed image containing \( n \) objects. In this study, we present SPOLRE, a novel automated tool designed for semantic preserving object layout reconstruction in image captioning system testing. SPOLRE is based on the insight that modifying the arrangement of objects within an image does not alter its inherent semantics. We utilize four semantic preserving transformation techniques—translation, rotation, mirroring, and scaling—to modify object layouts autonomously, eliminating the need for manual annotation. This approach enables the creation of realistic and varied test suites for IC system testing. Our extensive testing demonstrates that more than 75% of survey respondents find the images produced by SPOLRE more realistic compared to those generated by SOTA methods. Additionally, SPOLRE exhibits outstanding performance in identifying caption errors, detecting 31,544 incorrect captions across seven IC systems with an average precision of 91.62%. This significantly outperforms other methods, which only achieve 85.65% accuracy on average and identify 17,160 incorrect captions. Notably, SPOLRE exposes 6,236 unique issues within Microsoft Azure Cognitive Service, highlighting its effectiveness against one of the most advanced IC systems available. Yi Liu 0069, Guanyu Wang 0005, Gelei Deng, Kailong Wang 0001, Yang Liu 0003, Haoyu Wang 0001 |
ACM Trans. Softw. Eng. Methodol. | 1 |
| 2025 | Sticking to the Mean: Detecting Sticky Tokens in Text Embedding ModelsabstractDespite the widespread use of Transformerbased text embedding models in NLP tasks, surprising "sticky tokens" can undermine the reliability of embeddings.These tokens, when repeatedly inserted into sentences, pull sentence similarity toward a certain value, disrupting the normal distribution of embedding similarities and degrading downstream performance.In this paper, we systematically investigate such anomalous tokens, formally defining them and introducing an efficient detection method, Sticky Token Detector (STD), based on sentence and token filtering.Applying STD to 40 checkpoints across 14 model families, we discover a total of 868 sticky tokens.Our analysis reveals that these tokens often originate from special or unused entries in the vocabulary, as well as fragmented subwords from multilingual corpora.Notably, their presence does not strictly correlate with model size or vocabulary size.We further evaluate how sticky tokens affect downstream tasks like clustering and retrieval, observing substantial performance degradation that approaches 50% in certain cases.Through attention-layer analysis, we show that sticky tokens disproportionately dominate the model's internal representations, raising concerns about tokenization robustness.Our findings show the need for better tokenization strategies and model design to mitigate the impact of sticky tokens in future text embedding applications.� Dongxia Wang 0002, Yi Liu 0069, Haonan Zhang 0007, Wenhai Wang |
ACL (1) | 3 |
| 2025 | Oedipus: LLM-enchanced Reasoning CAPTCHA SolverabstractCAPTCHAs have become a ubiquitous tool in safeguarding applications from automated bots. Over time, the arms race between CAPTCHA development and evasion techniques has led to increasingly sophisticated and diverse designs. The latest iteration, reasoning CAPTCHAs, exploits tasks that are intuitively simple for humans but challenging for conventional AI technologies, thereby enhancing security measures. Gelei Deng, Haoran Ou, Yi Liu 0069, Jie Zhang 0073, Tianwei Zhang 0004, Yang Liu 0003 |
CCS | 3 |
| 2025 | TombRaider: Entering the Vault of History to Jailbreak Large Language ModelsabstractWarning: This paper contains content that may involve potentially harmful behaviours, discussed strictly for research purposes.Jailbreak attacks can hinder the safety of Large Language Model (LLM) applications, especially chatbots.Studying jailbreak techniques is an important AI red teaming task for improving the safety of these applications.In this paper, we introduce TOMBRAIDER, a novel jailbreak technique that exploits the ability to store, retrieve, and use historical knowledge of LLMs.TOMBRAIDER employs two agents, the inspector agent to extract relevant historical information and the attacker agent to generate adversarial prompts, enabling effective bypassing of safety filters.We intensively evaluated TOMBRAIDER on six popular models.Experimental results showed that TOMBRAIDER could outperform state-of-the-art jailbreak techniques, achieving nearly 100% attack success rates (ASRs) on bare models and maintaining over 55.4% ASR against defence mechanisms.Our findings highlight critical vulnerabilities in existing LLM safeguards, underscoring the need for more robust safety defences. Junchen Ding, Yi Liu 0069, Gelei Deng, Yuekang Li |
EMNLP | 3 |
| 2025 | Source Code Summarization in the Era of Large Language ModelsabstractTo support software developers in understanding and maintaining programs, various automatic (source) code summarization techniques have been proposed to generate a concise natural language summary (i.e., comment) for a given code snippet. Recently, the emergence of large language models (LLMs) has led to a great boost in the performance of coderelated tasks. In this paper, we undertake a systematic and comprehensive study on code summarization in the era of LLMs, which covers multiple aspects involved in the workflow of LLMbased code summarization. Specifically, we begin by examining prevalent automated evaluation methods for assessing the quality of summaries generated by LLMs and find that the results of the GPT-4 evaluation method are most closely aligned with human evaluation. Then, we explore the effectiveness of five prompting techniques (zero-shot, few-shot, chain-of-thought, critique, and expert) in adapting LLMs to code summarization tasks. Contrary to expectations, advanced prompting techniques may not outperform simple zero-shot prompting. Next, we investigate the impact of LLMs' model settings (including top_p and temperature parameters) on the quality of generated summaries. We find the impact of the two parameters on summary quality varies by the base LLM and programming language, but their impacts are similar. Moreover, we canvass LLMs' abilities to summarize code snippets in distinct types of programming languages. The results reveal that LLMs perform suboptimally when summarizing code written in logic programming languages compared to other language types (e.g., procedural and object-oriented programming languages). Finally, we unexpectedly find that CodeLlamaInstruct with 7B parameters can outperform advanced GPT-4 in generating summaries describing code design rationale and asserting code properties. We hope that our findings can provide a comprehensive understanding of code summarization in the era of LLMs. Weisong Sun, Yun Miao, Yuekang Li, Hongyu Zhang 0002, Chunrong Fang, Yi Liu 0069, Gelei Deng, Yang Liu 0003, Zhenyu Chen 0001 |
ICSE | 6 |
| 2025 | ORFuzz: Fuzzing the "Other Side" of LLM Safety - Testing Over-RefusalabstractLarge Language Models (LLMs) have been found to show over-refusal problems—erroneously rejecting benign queries due to overly conservative safety measures—a critical functional flaw that undermines their reliability and usability. Current methods for testing this behavior are demonstrably inadequate, suffering from flawed benchmarks and limited test generation capabilities, as highlighted by our empirical user study. To the best of our knowledge, this paper introduces the first evolutionary testing framework, ORFuzz, for the systematic detection and analysis of LLM over-refusals. ORFuzz uniquely integrates three core components: (1) safety category-aware seed selection for comprehensive test coverage, (2) adaptive mutator optimization using reasoning LLMs to generate effective test cases, and (3) OR-Judge, a human-aligned judge model validated to accurately reflect user perception of toxicity and refusal. Our extensive evaluations demonstrate that ORFuzz generates diverse, validated over-refusal instances at a rate (6.98% average) more than double that of leading baselines, effectively uncovering vulnerabilities. Furthermore, ORFuzz’s outputs form the basis of ORFuzzSet, a new benchmark of 1,786 highly transferable test cases that achieves a superior 57.37% average over-refusal rate across 14 diverse LLMs, significantly outperforming existing datasets. ORFuzz and ORFuzzSet provide a robust automated testing framework and a valuable community resource, paving the way for developing more reliable and trustworthy LLM-based software systems. The code of this paper is available at: https://github.com/HotBento/ORFuzz. Haonan Zhang 0007, Dongxia Wang 0002, Yi Liu 0069, Jiashui Wang, Xinlei Ying, Wenhai Wang |
ASE | 3 |
| 2025 | LaTCoder: Converting Webpage Design to Code with Layout-as-ThoughtabstractConverting webpage designs into code (design-to-code) plays a vital role in User Interface (UI) development for front-end developers, bridging the gap between visual design and functional implementation. While recent Multimodal Large Language Models (MLLMs) have shown significant potential in design-to-code tasks, they often fail to accurately preserve the layout during code generation. To this end, we draw inspiration from the Chain-of-Thought (CoT) reasoning in human cognition and propose LaTCoder, a novel approach that enhances layout preservation in webpage design during code generation with Layout-as-Thought (LaT). Specifically, we first introduce a simple yet efficient algorithm to divide the webpage design into image blocks. Next, we prompt MLLMs using a CoT-based approach to generate code for each block. Finally, we apply two assembly strategies-absolute positioning and an MLLM-based method-followed by dynamic selection to determine the optimal output. We evaluate the effectiveness of LaTCoder using multiple backbone MLLMs (i.e., DeepSeek-VL2, Gemini, and GPT-4o) on both a public benchmark and a newly introduced, more challenging benchmark (CC-HARD) that features complex layouts. The experimental results on automatic metrics demonstrate significant improvements. Specifically, TreeBLEU scores increased by 66.67% and MAE decreased by 38% when using DeepSeek-VL2, compared to direct prompting. Moreover, the human preference evaluation results indicate that annotators favor the webpages generated by LaTCoder in over 60% of cases, providing strong evidence of the effectiveness of our method. Yi Gui, Zhen Li 0050, Guohao Wang, Tianpeng Lv, Gaoyang Jiang, Yi Liu 0069, Dongping Chen, Yao Wan 0001, Hongyu Zhang 0002, Wenbin Jiang 0001, Xuanhua Shi, Hai Jin 0001 |
KDD (2) | 7 |
| 2025 | IllusionCAPTCHA: A CAPTCHA based on Visual IllusionabstractCAPTCHAs have long been essential tools for protecting applications from automated bots. Initially designed as simple questions to distinguish humans from bots, they have become increasingly complex to keep pace with the proliferation of CAPTCHA-cracking techniques employed by malicious actors. However, with the advent of advanced large language models (LLMs), the effectiveness of existing CAPTCHAs is now being undermined. Gelei Deng, Yi Liu 0069, Junchen Ding, Jieshan Chen, Yulei Sui, Yuekang Li |
WWW | 3 |
| 2025 | Mission: Impossible - Image-Based Geolocation with Large Vision Language ModelsabstractIn the age of ubiquitous smartphone use and widespread image sharing on social platforms, geolocation poses a critical privacy concern. Images often carry sensitive spatial and temporal details—such as street signs, architectural styles, or landmarks—that can inadvertently disclose the precise whereabouts of individuals and organizations. Recent advances in large vision-language models (LVLMs) present an emerging threat by enabling users, regardless of technical expertise, to extract location cues from seemingly benign photos. While existing AI-driven geolocation solutions often focus on narrow datasets or specialized contexts, the generalizable performance and privacy implications of zero-shot LVLMs in real-world settings remain critical questions. In this paper, we investigate the geolocation capabilities of state-of-the-art LVLMs. Our findings reveal that while these models demonstrate a non-negligible capability for image-based geolocation even without specialized training, their accuracy in absolute terms is often low, exposing clear limitations in their current state. We then introduce ETHAN, a framework integrating chain-of-thought (CoT) reasoning. Although ETHAN shows improved performance (e.g., 28.7% accuracy at the 1km threshold) and an 85.4% win rate on GeoGuessr, these results primarily highlight the potential trajectory of such technologies rather than their current widespread, high-accuracy applicability. Our study underscores the dual nature of LVLMs in this domain: they uncover an emerging privacy risk due to their inherent, albeit limited, geolocation abilities, yet also demonstrate significant constraints. We conclude by calling for further research into the limitations and risks of LVLM-based geolocation and the development of effective mitigation strategies to protect sensitive location data. Yi Liu 0069, Gelei Deng, Junchen Ding, Yuekang Li, Tianwei Zhang 0004, Weisong Sun, Yaowen Zheng, Jingquan Ge |
Proc. Priv. Enhancing Technol. | 1 |
| 2025 | MiniScope: Automated UI Exploration and Privacy Inconsistency Detection of MiniApps via Two-phase Iterative Hybrid AnalysisabstractThe advent of MiniApps, operating within larger SuperApps, has revolutionized user experiences by offering a wide range of services without the need for individual app downloads. However, this convenience has raised significant privacy concerns, as these MiniApps often require access to sensitive data, potentially leading to privacy violations. Despite existing privacy regulations and platform guidelines, there is a lack of effective mechanisms to safeguard user privacy fully. To address this critical gap, we introduce MiniScope , a novel two-phase hybrid analysis approach, specifically designed for the MiniApp environment. This approach overcomes the limitations of existing static analysis techniques by incorporating UI transition states analysis, cross-package callback control flow resolution, and automated iterative UI exploration. This allows for a comprehensive understanding of MiniApps’ privacy practices, addressing the unique challenges of sub-package loading and event-driven callbacks. Our empirical evaluation of over 120K MiniApps using MiniScope demonstrates its effectiveness in identifying privacy inconsistencies. The results reveal significant issues, with 5.7% of MiniApps over-collecting private data and 33.4% overclaiming data collection. We have responsibly disclosed our findings to 2,282 developers, receiving 44 acknowledgments. These findings emphasize the urgent need for more precise privacy monitoring systems and highlight the responsibility of SuperApp operators to enforce stricter privacy measures. Shenao Wang 0001, Yuekang Li, Kailong Wang 0001, Yi Liu 0069, Hui Li 0006, Yang Liu 0003, Haoyu Wang 0001 |
ACM Trans. Softw. Eng. Methodol. | 4 |
| 2024 | DistillSeq: A Framework for Safety Alignment Testing in Large Language Models using Knowledge DistillationabstractLarge Language Models (LLMs) have showcased their remarkable capabilities in diverse domains, encompassing natural language understanding, translation, and even code generation. The potential for LLMs to generate harmful content is a significant concern. This risk necessitates rigorous testing and comprehensive evaluation of LLMs to ensure safe and responsible use. However, extensive testing of LLMs requires substantial computational resources, making it an expensive endeavor. Therefore, exploring cost-saving strategies during the testing phase is crucial to balance the need for thorough evaluation with the constraints of resource availability. To address this, our approach begins by transferring the moderation knowledge from an LLM to a small model. Subsequently, we deploy two distinct strategies for generating malicious queries: one based on a syntax tree approach, and the other leveraging an LLM-based method. Finally, our approach incorporates a sequential filter-test process designed to identify test cases that are prone to eliciting toxic responses. By doing so, we significantly curtail unnecessary or unproductive interactions with LLMs, thereby streamlining the testing process. Our research evaluated the efficacy of DistillSeq across four LLMs: GPT-3.5, GPT-4.0, Vicuna-13B, and Llama-13B. In the absence of DistillSeq, the observed attack success rates on these LLMs stood at 31.5% for GPT-3.5, 21.4% for GPT-4.0, 28.3% for Vicuna-13B, and 30.9% for Llama-13B. However, upon the application of DistillSeq, these success rates notably increased to 58.5%, 50.7%, 52.5%, and 54.4%, respectively. This translated to an average escalation in attack success rate by a factor of 93.0% when compared to scenarios without the use of DistillSeq. Such findings highlight the significant enhancement DistillSeq offers in terms of reducing the time and resource investment required for effectively testing LLMs. Mingke Yang, Yuqi Chen 0001, Yi Liu 0069, Ling Shi 0002 |
ISSTA | 3 |
| 2024 | Efficient Detection of Toxic Prompts in Large Language ModelsabstractLarge language models (LLMs) like ChatGPT and Gemini have significantly advanced natural language processing, enabling various applications such as chatbots and automated content generation. However, these models can be exploited by malicious individuals who craft toxic prompts to elicit harmful or unethical responses. These individuals often employ jailbreaking techniques to bypass safety mechanisms, highlighting the need for robust toxic prompt detection methods. Existing detection techniques, both blackbox and whitebox, face challenges related to the diversity of toxic prompts, scalability, and computational efficiency. In response, we propose ToxicDetector, a lightweight greybox method designed to efficiently detect toxic prompts in LLMs. ToxicDetector leverages LLMs to create toxic concept prompts, uses embedding vectors to form feature vectors, and employs a Multi-Layer Perceptron (MLP) classifier for prompt classification. Our evaluation on various versions of the LLama models, Gemma-2, and multiple datasets demonstrates that ToxicDetector achieves a high accuracy of 96.39% and a low false positive rate of 2.00%, outperforming state-of-the-art methods. Additionally, ToxicDetector's processing time of 0.0780 seconds per prompt makes it highly suitable for real-time applications. ToxicDetector achieves high accuracy, efficiency, and scalability, making it a practical method for toxic prompt detection in LLMs. Yi Liu 0069, Junzhe Yu, Huijia Sun, Ling Shi 0002, Gelei Deng, Yuqi Chen 0001, Yang Liu 0003 |
ASE | 1 |
| 2024 | MASTERKEY: Automated Jailbreaking of Large Language Model Chatbots
Gelei Deng, Yi Liu 0069, Yuekang Li, Kailong Wang 0001, Ying Zhang 0066, Zefeng Li, Haoyu Wang 0001, Tianwei Zhang 0004, Yang Liu 0003 |
NDSS | 2 |
| 2024 | PentestGPT: Evaluating and Harnessing Large Language Models for Automated Penetration Testing
Gelei Deng, Yi Liu 0069, Victor Mayoral Vilches, Yuekang Li, Yuan Xu 0033, Martin Pinzger 0001, Stefan Rass, Tianwei Zhang 0004, Yang Liu 0003 |
USENIX Security Symposium | 2 |
| 2024 | Medusa: Unveil Memory Exhaustion DoS Vulnerabilities in Protocol ImplementationsabstractWeb services have brought great convenience to our daily lives. Meanwhile, they are vulnerable to Denial-of-Service (DoS) attacks. DoS attacks launched via vulnerabilities in the services can cause great harm. The vulnerabilities in protocol implementations are especially important because they are the keystones of web services. One vulnerable protocol implementation can affect all the web services built on top of it. Compared to the vulnerabilities that cause the target service to crash, resource exhaustion vulnerabilities are equally if not more important. This is because such vulnerabilities can deplete the system resources, leading to the unavailability of not only the vulnerable service but also other services running on the same machine. Despite the significance of this type of vulnerability, there has been limited research in this area. Zhengjie Du, Yuekang Li, Yaowen Zheng, Cen Zhang, Yi Liu 0069, Sheikh Mahbub Habib, Xinghua Li 0001, Linzhang Wang, Yang Liu 0003, Bing Mao 0001 |
WWW | 6 |
| 2024 | Drowzee: Metamorphic Testing for Fact-Conflicting Hallucination Detection in Large Language ModelsabstractLarge language models (LLMs) have revolutionized language processing, but face critical challenges with security, privacy, and generating hallucinations — coherent but factually inaccurate outputs. A major issue is fact-conflicting hallucination (FCH), where LLMs produce content contradicting ground truth facts. Addressing FCH is difficult due to two key challenges: 1) Automatically constructing and updating benchmark datasets is hard, as existing methods rely on manually curated static benchmarks that cannot cover the broad, evolving spectrum of FCH cases. 2) Validating the reasoning behind LLM outputs is inherently difficult, especially for complex logical relations. To tackle these challenges, we introduce a novel logic-programming-aided metamorphic testing technique for FCH detection. We develop an extensive and extensible framework that constructs a comprehensive factual knowledge base by crawling sources like Wikipedia, seamlessly integrated into D rowzee . Using logical reasoning rules, we transform and augment this knowledge into a large set of test cases with ground truth answers. We test LLMs on these cases through template-based prompts, requiring them to provide reasoned answers. To validate their reasoning, we propose two semantic-aware oracles that assess the similarity between the semantic structures of the LLM answers and ground truth. Our approach automatically generates useful test cases and identifies hallucinations across six LLMs within nine domains, with hallucination rates ranging from 24.7% to 59.8%. Key findings include LLMs struggling with temporal concepts, out-of-distribution knowledge, and lack of logical reasoning capabilities. The results show that logic-based test cases generated by D rowzee effectively trigger and detect hallucinations. To further mitigate the identified FCHs, we explored model editing techniques, which proved effective on a small scale (with edits to fewer than 1000 knowledge pieces). Our findings emphasize the need for continued community efforts to detect and mitigate model hallucinations. Ningke Li, Yuekang Li, Yi Liu 0069, Ling Shi 0002, Kailong Wang 0001, Haoyu Wang 0001 |
Proc. ACM Program. Lang. | 3 |
| 2023 | PumpChannel: An Efficient and Secure Communication Channel for Trusted Execution Environment on ARM-FPGA Embedded SoCabstractARM TrustZone separates the system into the rich execution environment (REE) and the trusted execution environment (TEE). Data can be exchanged between REE and TEE through the communication channel, which is based on shared memory and can be accessed by both REE and TEE. Therefore, when the REE OS kernel is untrusted, the security of the communication channel cannot be guaranteed. The proposed schemes to protect the communication channel have high performance overhead and are not secure enough. In this paper, we propose PumpChannel, an efficient and secure communication channel implemented on ARM-FPGA embedded SoC. PumpChannel avoids the use of secret keys, but utilizes a hardware and software collaborative pump to enhance the security and performance of the communication channel. Besides, PumpChannel implements a hardware-based hook integrity monitor to ensure the integrity of all hook code. Security and performance evaluation results show that PumpChannel is more secure than the encrypted channel countermeasures and has better performance than all other evaluated schemes. Jingquan Ge, Yuekang Li, Yang Liu 0003, Yaowen Zheng, Yi Liu 0069, Lida Zhao |
DATE | 5 |
| 2023 | ASTER: Automatic Speech Recognition System Accessibility Testing for StutterersabstractThe popularity of automatic speech recognition (ASR) systems nowadays leads to an increasing need for improving their accessibility. Handling stuttering speech is an important feature for accessible ASR systems. To improve the accessibility of ASR systems for stutterers, we need to expose and analyze the failures of ASR systems on stuttering speech. The speech datasets recorded from stutterers are not diverse enough to expose most of the failures. Furthermore, these datasets lack ground truth information about the non-stuttered text, rendering them unsuitable as comprehensive test suites. Therefore, a methodology for generating stuttering speech as test inputs to test and analyze the performance of ASR systems is needed. However, generating valid test inputs in this scenario is challenging. The reason is that although the generated test inputs should mimic how stutterers speak, they should also be diverse enough to trigger more failures. To address the challenge, we propose Aster, a technique for automatically testing the accessibility of ASR systems. Aster can generate valid test cases by injecting five different types of stuttering. The generated test cases can both simulate realistic stuttering speech and expose failures in ASR systems. Moreover, Aster can further enhance the quality of the test cases with a multi-objective optimization-based seed updating algorithm. We implemented Aster as a framework and evaluated it on four open-source ASR models and three commercial ASR systems. We conduct a comprehensive evaluation of Aster and find that it significantly increases the word error rate, match error rate, and word information loss in the evaluated ASR systems. Additionally, our user study demonstrates that the generated stuttering audio is indistinguishable from real-world stuttering audio clips. Yi Liu 0069, Yuekang Li, Gelei Deng, Felix Juefei-Xu, Yao Du 0002, Cen Zhang, Yeting Li, Lei Ma 0003, Yang Liu 0003 |
ASE | 1 |
| 2023 | Effective ReDoS Detection by Principled Vulnerability Modeling and Exploit GenerationabstractRegular expression Denial-of-Service (ReDoS) is one kind of algorithmic complexity attack. For a vulnerable regex, attackers can craft certain strings to trigger the super-linear worst-case matching time, which causes denial-of-service to regex engines. Various ReDoS detection approaches have been proposed recently. Among them, hybrid approaches which absorb the advantages of both static and dynamic approaches have shown their performance superiority. However, two key challenges still hinder the effectiveness of the detection: 1) Existing modelings summarize localized vulnerability patterns based on partial features of the vulnerable regex; 2) Existing attack string generation strategies are ineffective since they neglected the fact that non-vulnerable parts of the regex may unexpectedly invalidate the attack string (we name this kind of invalidation as disturbance.)Rengar is our hybrid ReDoS detector with new vulnerability modeling and disturbance free attack string generator. It has the following key features: 1) Benefited by summarizing patterns from full features of the vulnerable regex, its modeling is a more precise interpretation of the root cause of ReDoS vulnerability. The modeling is more descriptive and precise than the union of existing modelings while keeping conciseness; 2) For each vulnerable regex, its generator automatically checks all potential disturbances and composes generation constraints to avoid possible disturbances.Compared with nine state-of-the-art tools, Rengar detects not only all vulnerable regexes they found but also 3 – 197 times more vulnerable regexes. Besides, it saves 57.41% – 99.83% average detection time compared with tools containing a dynamic validation process. Using Rengar, we have identified 69 zero-day vulnerabilities (21 CVEs) affecting popular projects which have more than dozens of millions weekly download count. Xinyi Wang 0013, Cen Zhang, Yeting Li, Zhiwu Xu 0001, Shuailin Huang, Yi Liu 0069, Yican Yao, Yang Xiao 0011, Yanyan Zou 0002, Yang Liu 0003, Wei Huo 0005 |
SP | 6 |
| 2023 | NAUTILUS: Automated RESTful API Vulnerability Detection
Gelei Deng, Zhiyi Zhang 0005, Yuekang Li, Yi Liu 0069, Tianwei Zhang 0004, Yang Liu 0003, Dongjin Wang |
USENIX Security Symposium | 4 |
| 2022 | Morest: Model-based RESTful API Testing with Execution FeedbackabstractRESTful APIs are arguably the most popular endpoints for accessing Web services. Blackbox testing is one of the emerging techniques for ensuring the reliability of RESTful APIs. The major challenge in testing RESTful APIs is the need for correct sequences of API operation calls for in-depth testing. To build meaningful operation call sequences, researchers have proposed techniques to learn and utilize the API dependencies based on OpenAPI specifications. However, these techniques either lack the overall awareness of how all the APIs are connected or the flexibility of adaptively fixing the learned knowledge. Yi Liu 0069, Yuekang Li, Gelei Deng, Yang Liu 0003, Ruiyuan Wan, Runchao Wu, Dandan Ji, Shiheng Xu, Minli Bao |
ICSE | 1 |
| 2022 | RESTCluster: Automated Crash Clustering for RESTful APIabstractRESTful API has been adopted by many other notable companies to provide cloud services. Quality assurance of RESTful API is essential. Several automated RESTful API testing techniques have been proposed to overcome this problem. However, automated tools often generate a large number of failed test cases. Since validating each test case is a lot of work for developers, automatic failure clustering is a promising solution to help debug cloud services. Yi Liu 0069 |
ASE | 1 |
| 2022 | Morest: Industry Practice of Automatic RESTful API TestingabstractMany big companies are providing cloud services through RESTful APIs nowadays. With the growing popularity of RESTful API, testing RESTful API becomes crucial. To address this issue, researchers have proposed several automatic RESTful API testing techniques. At Huawei, we design and implement an automatic RESTful API testing framework named Morest. Morest has been used to test ten RESTful API services and helped to detected 83 previously unknown bugs which were all confirmed and fixed by the developers. On one hand, we find that Morest shows great capability of detecting bugs in RESTful API s. On the other hand, we also notice that human effort is inevitable and important when applying automatic RESTful API techniques in practice. Yi Liu 0069, Yuekang Li, Yang Liu 0003, Ruiyuan Wan, Runchao Wu, Qingkun Liu |
ASE | 1 |
| 2022 | RESTInfer: automated inferring parameter constraints from natural language RESTful API descriptionsabstractRESTful APIs have been applied to provide cloud services by various notable companies. The quality assurance of RESTful API is critical. Several automatic RESTful API testing techniques have been proposed to tame this issue. By analyzing crashes caused by each test case, developers can find potential bugs in cloud services. However, it is difficult for automated tools to generate feasible parameters under complicating constraints randomly. Fortunately, RESTful API descriptions can be used to infer possible parameter constraints. Given parameter constraints, automated tools can further improve the efficiency of testing. Yi Liu 0069 |
ESEC/SIGSOFT FSE | 1 |
| 2022 | DeepAnna: Deep Learning based Java Annotation Recommendation and Misuse DetectionabstractAnnotations have been widely used in Java programs to support additional compile-time, deployment-time, and runtime processing. Developers use annotations to delegate repetitive logics such as object initialization and request forwarding to compilers and runtime frameworks. Therefore, these annotations are important for the correct execution of programs. In practice, however, developers often find it hard to correctly use annotations and the misuse of annotations has led to real bugs in Java programs. In this paper, we conduct an empirical study on Stack Overflow questions to investigate the major development frameworks that are involved in questions about Java annotations and the main problems encountered by developers in the use of Java annotations. Based on the findings of the study, we propose DeepAnna, a deep learning based Java annotation recommendation and misuse detection approach. Based on a corpus of Java programs with intensive use of annotations, DeepAnna trains a deep learning based multi-label classification model by considering both the structural and textual contexts of source code. DeepAnna can recommend annotations at both class level and method level. Our evaluation with a large corpus of open-source Java projects shows that DeepAnna outperforms state-of-the-art text multi-label classification approaches in annotation recommendation and can effectively detect annotation misuses. Based on our analysis, we submit 85 bug-fixing pull requests for annotation misuses in open-source projects and 20 of them have been accepted and merged. Yi Liu 0069, Yadong Yan, Chaofeng Sha, Xin Peng 0001, Bihuan Chen 0001, Chong Wang 0013 |
SANER | 1 |
| 2021 | Automatic Web Testing Using Curiosity-Driven Reinforcement LearningabstractWeb testing has long been recognized as a notoriously difficult task. Even nowadays, web testing still mainly relies on manual efforts in many cases while automated web testing is still far from achieving human-level performance. Key challenges include dynamic content update and deep bugs hiding under complicated user interactions and specific input values, which can only be triggered by certain action sequences in the huge space of all possible sequences. In this paper, we propose WebExplor, an automatic end-to-end web testing framework, to achieve an adaptive exploration of web applications. WebExplor adopts a curiosity-driven reinforcement learning to generate high-quality action sequences (test cases) with temporal logical relations. Besides, WebExplor incrementally builds an automaton during the online testing process, which acts as the high-level guidance to further improve the testing efficiency. We have conducted comprehensive evaluations on six real-world projects, a commercial SaaS web application, and performed an in-the-wild study of the top 50 web applications in the world. The results demonstrate that in most cases WebExplor can achieve significantly higher failure detection rate, code coverage and efficiency than existing state-of-the-art web testing techniques. WebExplor also detected 12 previously unknown failures in the commercial web application, which have been confirmed and fixed by the developers. Furthermore, our in-the-wild study further uncovered 3,466 exceptions and errors. Yan Zheng 0002, Yi Liu 0069, Xiaofei Xie, Yepang Liu 0001, Lei Ma 0003, Jianye Hao, Yang Liu 0003 |
ICSE | 2 |
| 2020 | An Exploratory Study of Bugs in Extended Reality Applications on the WebabstractExtended Reality (XR) technologies are becoming increasingly popular in recent years. To help developers deploy XR applications on the Web, W3C released the WebXR Device API in 2019, which enable users to interact with browsers using XR devices. Given the convenience brought by WebXR, a growing number of WebXR projects have been deployed in practice. However, many WebXR applications are insufficiently tested before being released. They suffer from various bugs that can degrade user experience or cause undesirable consequences. Yet, the community has limited understanding towards the bugs in the WebXR ecosystem, which impedes the advance of techniques for assuring the reliability of WebXR applications. To bridge this gap, we conducted the first empirical study of WebXR bugs. We collected 368 real bugs from 33 WebXR projects hosted on GitHub. Via a seven-round manual analysis of these bugs, we built a taxonomy of WebXR bugs according to their symptoms and root causes. Furthermore, to understand the uniqueness of WebXR bugs, we compared them with bugs in conventional JavaScript programs and web applications. We believe that our findings can inspire future researches on relevant topics and we released our bug dataset to facilitate follow-up studies. Shuqing Li 0001, Yechang Wu, Yi Liu 0069, Dinghua Wang, Ming Wen 0001, Yida Tao, Yulei Sui, Yepang Liu 0001 |
ISSRE | 3 |
| 2020 | Industry Practice of JavaScript Dynamic Analysis on WeChat Mini-ProgramsabstractJavaScript is one of the most popular programming languages. WeChat Mini-Program is a large ecosystem of JavaScript applications that runs on the WeChat platform. Millions of Mini-Programs are accessed by WeChat users every week. Consequently, the performance and robustness of Mini-Programs are particularly important. Unfortunately, many Mini-Programs suffer from various defects and performance problems. Dynamic analysis is a useful technique to pinpoint application defects. However, due to the dynamic features of the JavaScript language and the complexity of the runtime environment, dynamic analysis techniques were rarely used to improve the quality of JavaScript applications running on industrial platforms such as WeChat Mini-Program previously. In this work, we report our experience of extending Jalangi, a dynamic analysis framework for JavaScript applications developed by academia, and applying the extended version, named WeJalangi, to diagnose defects in WeChat Mini-Programs. WeJalangi is compatible with existing dynamic analysis tools such as DLint, Smemory, and JITProf. We implemented a null pointer checker on WeJalangi and tested the tool's usability on 152 open-source Mini-Programs. We also conducted a case study in Tencent by applying WeJalangi on six popular commercial Mini-Programs. In the case study, WeJalangi accurately located six null pointer issues and three of them haven't been discovered previously. All of the reported defects have been confirmed by developers and testers. Yi Liu 0069, Jinhui Xie, Jianbo Yang, Yuetang Deng, Shuqing Li 0001, Yechang Wu, Yepang Liu 0001 |
ASE | 1 |