Ling Shi 0002

dblp:97/6316-2 · DBLP profile ↗
← Back
25ranked-venue papers
4as first author
14since 2021 · last 2026
0000-0002-2023-0247ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 20 · 3 first-author · 10 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021Theory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2026 STEAMROLLER: A Multi-Agent System for Inclusive Automatic Speech Recognition for People Who Stutter
abstract
People who stutter (PWS) face systemic exclusion in today’s voice-driven society, where access to voice assistants, authentication systems, and remote work tools increasingly depends on fluent speech. Current automatic speech recognition (ASR) systems, trained predominantly on fluent speech, fail to serve millions of PWS worldwide. We present STEAMROLLER, a real time system that transforms stuttered speech into fluent output through a novel multi-stage, multi-agent AI pipeline. Our approach addresses three critical technical challenges: (1) the difficulty of direct speech to speech conversion for disfluent input, (2) semantic distortions introduced during ASR transcription of stuttered speech, and (3) latency constraints for real time communication. STEAMROLLER employs a three stage architecture comprising ASR transcription, multi-agent text repair, and speech synthesis, where our core innovation lies in a collaborative multi-agent framework that iteratively refines transcripts while preserving semantic intent. Experiments on the FluencyBank dataset and a user study demonstrates clear word error rate (WER) reduction and strong user satisfaction. Beyond immediate accessibility benefits, fine tuning ASR on STEAMROLLER repaired speech further yields additional WER improvements, creating a pathway toward inclusive AI ecosystems.
Yi Liu 0069, Yuekang Li, Ling Shi 0002, Kailong Wang 0001
AAAI4
2026 R2IF: Aligning Reasoning with Decisions via Composite Rewards for Interpretable LLM Function Calling
abstract
Function calling empowers large language models (LLMs) to interface with external tools, yet existing RL-based approaches suffer from misalignment between reasoning processes and tool-call decisions.We propose R2IF, a reasoning-aware RL framework for interpretable function calling, adopting a composite reward integrating format/correctness constraints, Chain-of-Thought Effectiveness Reward (CER), and Specification-Modification-Value (SMV) reward, optimized via GRPO.Experiments on BFCL/ACEBench show R2IF outperforms baselines by up to 34.62% (Llama3.2-3B on BFCL) with positive Average CoT Effectiveness (0.05 for Llama3.2-3B),enhancing both function-calling accuracy and interpretability for reliable tool-augmented LLM deployment.
Aijia Cheng, Kailong Wang 0001, Ling Shi 0002
ACL (1)3
2026 Uncovering Logit Suppression Vulnerabilities in LLM Safety Alignment
Yi Liu 0069, Yuekang Li, Ling Shi 0002, Gelei Deng, Shengquan Chen, Kailong Wang 0001
ICPR (2)4
2026 Exposing the Ghost in the Transformer: Abnormal Detection for Large Language Models via Hidden State Forensics
abstract
The widespread adoption of Large Language Models (LLMs) in critical applications has introduced severe reliability and security risks, as LLMs remain vulnerable to notorious threats such as hallucinations, jailbreak attacks, and backdoor exploits. These vulnerabilities have been weaponized by malicious actors, leading to unauthorized access, widespread misinformation, and compromised LLM-embedded system integrity. In this work, we introduce a novel approach to detecting abnormal behaviors in LLMs via hidden state forensics. By systematically inspecting layer-specific activation patterns, we develop a general framework that can efficiently identify a range of security threats in real-time without imposing prohibitive computational costs. Extensive experiments indicate detection accuracies exceeding 95% and consistently robust performance across multiple models in most scenarios, while preserving the ability to detect novel attacks effectively. Furthermore, the computational overhead remains minimal, with detector inference taking merely fractions of a second. The significance of this work lies in proposing a promising strategy to reinforce the security of LLM-integrated systems, paving the way for safer and more reliable deployment in high-stakes domains. By enabling real-time detection that can also support the mitigation of abnormal behaviors, it represents a meaningful step toward ensuring the trustworthiness of AI systems amid rising security challenges.
Shide Zhou, Kailong Wang 0001, Ling Shi 0002, Haoyu Wang 0001
IEEE Trans. Inf. Forensics Secur.3
2026 NeuSemSlice: Towards Effective DNN Model Maintenance via Neuron-Level Semantic Slicing
abstract
Deep Neural Networks (DNNs), extensively applied across diverse disciplines, are characterized by their integrated and monolithic architectures, setting them apart from conventional software systems. This architectural difference introduces particular challenges to maintenance tasks, such as model restructure (e.g., model compression), re-adaptation (e.g., fitting new samples), and incremental development (e.g., continual knowledge accumulation). Prior research addresses these challenges by identifying task-critical neuron layers and dividing neural networks into semantically similar sequential modules. However, such layer-level approaches fail to precisely identify and manipulate neuron-level semantic components, restricting their applicability to finer-grained model maintenance tasks. In this work, we implement NeuSemSlice, a novel framework that introduces the semantic slicing technique to effectively identify critical neuron-level semantic components in DNN models for semantic-aware model maintenance tasks. Specifically, semantic slicing identifies, categorizes, and merges critical neurons across different categories and layers according to their semantic similarity, enabling their flexibility and effectiveness in the subsequent tasks. For semantic-aware model maintenance tasks, we provide a series of novel strategies based on semantic slicing to enhance NeuSemSlice. They include semantic components (i.e., critical neurons) preservation for model restructure, critical neuron tuning for model re-adaptation, and non-critical neuron training for model incremental development. A thorough evaluation has demonstrated that NeuSemSlice significantly outperforms baselines in all three tasks.
Shide Zhou, Tianlin Li, Yihao Huang 0001, Ling Shi 0002, Kailong Wang 0001, Yang Liu 0003, Haoyu Wang 0001
ACM Trans. Softw. Eng. Methodol.4
2025 Understanding the Effectiveness of Coverage Criteria for Large Language Models: A Special Angle from Jailbreak Attacks
abstract
Large language models (LLMs) have revolutionized artificial intelligence, but their increasing deployment across critical domains has raised concerns about their abnormal behaviors when faced with malicious attacks. Such vulnerability alerts the widespread inadequacy of pre-release testing. In this paper, we conduct a comprehensive empirical study to evaluate the effectiveness of traditional coverage criteria in identifying such inadequacies, exemplified by the significant security concern of jailbreak attacks. Our study begins with a clustering analysis of the hidden states of LLMs, revealing that the embedded characteristics effectively distinguish between different query types. We then systematically evaluate the performance of these criteria across three key dimensions: criterion level, layer level, and token level. Our research uncovers significant differences in neuron coverage when LLMs process normal versus jailbreak queries, aligning with our clustering experiments. Leveraging these findings, we propose three practical applications of coverage criteria in the context of LLM security testing. Specifically, we develop a realtime jailbreak detection mechanism that achieves high accuracy (93.61 % on average) in classifying queries as normal or jailbreak. Furthermore, we explore the use of coverage levels to prioritize test cases, improving testing efficiency by focusing on high-risk interactions and removing redundant tests. Lastly, we introduce a coverage-guided approach for generating jailbreak attack examples, enabling systematic refinement of prompts to uncover vulnerabilities. This study improves our understanding of LLM security testing, enhances their safety, and provides a foundation for developing more robust AI applications.
Shide Zhou, Tianlin Li, Kailong Wang 0001, Yihao Huang 0001, Ling Shi 0002, Yang Liu 0003, Haoyu Wang 0001
ICSE5
2024 DistillSeq: A Framework for Safety Alignment Testing in Large Language Models using Knowledge Distillation
abstract
Large Language Models (LLMs) have showcased their remarkable capabilities in diverse domains, encompassing natural language understanding, translation, and even code generation. The potential for LLMs to generate harmful content is a significant concern. This risk necessitates rigorous testing and comprehensive evaluation of LLMs to ensure safe and responsible use. However, extensive testing of LLMs requires substantial computational resources, making it an expensive endeavor. Therefore, exploring cost-saving strategies during the testing phase is crucial to balance the need for thorough evaluation with the constraints of resource availability. To address this, our approach begins by transferring the moderation knowledge from an LLM to a small model. Subsequently, we deploy two distinct strategies for generating malicious queries: one based on a syntax tree approach, and the other leveraging an LLM-based method. Finally, our approach incorporates a sequential filter-test process designed to identify test cases that are prone to eliciting toxic responses. By doing so, we significantly curtail unnecessary or unproductive interactions with LLMs, thereby streamlining the testing process. Our research evaluated the efficacy of DistillSeq across four LLMs: GPT-3.5, GPT-4.0, Vicuna-13B, and Llama-13B. In the absence of DistillSeq, the observed attack success rates on these LLMs stood at 31.5% for GPT-3.5, 21.4% for GPT-4.0, 28.3% for Vicuna-13B, and 30.9% for Llama-13B. However, upon the application of DistillSeq, these success rates notably increased to 58.5%, 50.7%, 52.5%, and 54.4%, respectively. This translated to an average escalation in attack success rate by a factor of 93.0% when compared to scenarios without the use of DistillSeq. Such findings highlight the significant enhancement DistillSeq offers in terms of reducing the time and resource investment required for effectively testing LLMs.
Mingke Yang, Yuqi Chen 0001, Yi Liu 0069, Ling Shi 0002
ISSTA4
2024 Efficient Detection of Toxic Prompts in Large Language Models
abstract
Large language models (LLMs) like ChatGPT and Gemini have significantly advanced natural language processing, enabling various applications such as chatbots and automated content generation. However, these models can be exploited by malicious individuals who craft toxic prompts to elicit harmful or unethical responses. These individuals often employ jailbreaking techniques to bypass safety mechanisms, highlighting the need for robust toxic prompt detection methods. Existing detection techniques, both blackbox and whitebox, face challenges related to the diversity of toxic prompts, scalability, and computational efficiency. In response, we propose ToxicDetector, a lightweight greybox method designed to efficiently detect toxic prompts in LLMs. ToxicDetector leverages LLMs to create toxic concept prompts, uses embedding vectors to form feature vectors, and employs a Multi-Layer Perceptron (MLP) classifier for prompt classification. Our evaluation on various versions of the LLama models, Gemma-2, and multiple datasets demonstrates that ToxicDetector achieves a high accuracy of 96.39% and a low false positive rate of 2.00%, outperforming state-of-the-art methods. Additionally, ToxicDetector's processing time of 0.0780 seconds per prompt makes it highly suitable for real-time applications. ToxicDetector achieves high accuracy, efficiency, and scalability, making it a practical method for toxic prompt detection in LLMs.
Yi Liu 0069, Junzhe Yu, Huijia Sun, Ling Shi 0002, Gelei Deng, Yuqi Chen 0001, Yang Liu 0003
ASE4
2024 Semantic-Enhanced Indirect Call Analysis with Large Language Models
abstract
In contemporary software development, the widespread use of indirect calls to achieve dynamic features poses challenges in constructing precise control flow graphs (CFGs), which further impacts the performance of downstream static analysis tasks. To tackle this issue, various types of indirect call analyzers have been proposed. However, they do not fully leverage the semantic information of the program, limiting their effectiveness in real-world scenarios.
Baijun Cheng, Cen Zhang, Kailong Wang 0001, Ling Shi 0002, Yang Liu 0003, Haoyu Wang 0001, Yao Guo 0001, Ding Li 0001, Xiangqun Chen
ASE4
2024 GlitchProber: Advancing Effective Detection and Mitigation of Glitch Tokens in Large Language Models
abstract
Large language models (LLMs) have achieved unprecedented success in the field of natural language processing. However, the black-box nature of their internal mechanisms has brought many concerns about their trustworthiness and interpretability. Recent research has discovered a class of abnormal tokens in the model's vocabulary space and named them "glitch tokens". Those tokens, once included in the input, may induce the model to produce incorrect, irrelevant, or even harmful results, drastically undermining the reliability and practicality of LLMs.
Wuxia Bai, Yuxi Li 0010, Mark Huasong Meng, Kailong Wang 0001, Ling Shi 0002, Li Li 0029, Jun Wang 0020, Haoyu Wang 0001
ASE6
2024 Drowzee: Metamorphic Testing for Fact-Conflicting Hallucination Detection in Large Language Models
abstract
Large language models (LLMs) have revolutionized language processing, but face critical challenges with security, privacy, and generating hallucinations — coherent but factually inaccurate outputs. A major issue is fact-conflicting hallucination (FCH), where LLMs produce content contradicting ground truth facts. Addressing FCH is difficult due to two key challenges: 1) Automatically constructing and updating benchmark datasets is hard, as existing methods rely on manually curated static benchmarks that cannot cover the broad, evolving spectrum of FCH cases. 2) Validating the reasoning behind LLM outputs is inherently difficult, especially for complex logical relations. To tackle these challenges, we introduce a novel logic-programming-aided metamorphic testing technique for FCH detection. We develop an extensive and extensible framework that constructs a comprehensive factual knowledge base by crawling sources like Wikipedia, seamlessly integrated into D rowzee . Using logical reasoning rules, we transform and augment this knowledge into a large set of test cases with ground truth answers. We test LLMs on these cases through template-based prompts, requiring them to provide reasoned answers. To validate their reasoning, we propose two semantic-aware oracles that assess the similarity between the semantic structures of the LLM answers and ground truth. Our approach automatically generates useful test cases and identifies hallucinations across six LLMs within nine domains, with hallucination rates ranging from 24.7% to 59.8%. Key findings include LLMs struggling with temporal concepts, out-of-distribution knowledge, and lack of logical reasoning capabilities. The results show that logic-based test cases generated by D rowzee effectively trigger and detect hallucinations. To further mitigate the identified FCHs, we explored model editing techniques, which proved effective on a small scale (with edits to fewer than 1000 knowledge pieces). Our findings emphasize the need for continued community efforts to detect and mitigate model hallucinations.
Ningke Li, Yuekang Li, Yi Liu 0069, Ling Shi 0002, Kailong Wang 0001, Haoyu Wang 0001
Proc. ACM Program. Lang.4
2021 Verification Assisted Gas Reduction for Smart Contracts
abstract
Smart contracts are computerized transaction protocols built on top of blockchain networks. Users are charged with fees, a.k.a. gas in Ethereum, when they create, deploy or execute smart contracts. Since smart contracts may contain vulnerabilities which may result in huge financial loss, developers and smart contract compilers often insert codes for security checks. The trouble is that those codes consume gas every time they are executed. Many of the inserted codes are however redundant. In this work, we present sOptimize, a tool that optimizes smart contract gas consumption automatically without compromising functionality or security. sOptimize works on smart contract bytecode, statically identifies 3 kinds of code patterns, and further removes them through verification-assisted techniques. The resulting code is guaranteed to be equivalent to the original one and can be directly deployed on blockchain. We evaluate sOptimize on a collection of 1,152 real-world smart contracts and show that it optimizes 43% of them, and the reduction on gas consumption is about 2.0% while in deployment and 1.2% in transactions, the amount can be as high as 954,201 gas units per contract.
Ling Shi 0002, Jiaying Li 0001, Jun Sun 0001, Lei Bu
APSEC3
2021 sVerify: Verifying Smart Contracts Through Lazy Annotation and Learning
Ling Shi 0002, Jiaying Li 0001, Jialiang Chang, Jun Sun 0001, Zijiang Yang 0006
ISoLA2
2021 Scrutinizing Implementations of Smart Home Integrations
abstract
A key feature of the booming smart home is the integration of a wide assortment of technologies, including various standards, proprietary communication protocols and heterogeneous platforms. Due to customization, unsatisfied assumptions and incompatibility in the integration, critical security vulnerabilities are likely to be introduced by the integration. Hence, this work addresses the security problems in smart home systems from anintegrationperspective, as a complement to numerous studies that focus on the analysis of individual techniques. We propose HomeScan, an approach that examines the security of the implementations of smart home systems. It extracts the abstract specification of application-layer protocols and internal behaviors of entities, so that it is able to conduct an end-to-end security analysis against various attack models. Applying HomeScanon three extensively-used smart home systems, we have found twelve non-trivial security issues, which may lead to unauthorized remote control and credential leakage.
Kulani Mahadewa, Kailong Wang 0001, Guangdong Bai, Ling Shi 0002, Yan Liu 0012, Jin Song Dong 0001, Zhenkai Liang
IEEE Trans. Software Eng.4
2019 Towards a Formal Approach to Defining and Computing the Complexity of Component Based Software
abstract
With the rapid development of software engineering and the widely adoption of software systems in various domains, the requirement for software systems is becoming more and more complex, which results in very complex software systems. Motivated by the principle of divide and conquer, component based software development is an effective way of managing the complexity in software development. In this paper, we propose a calculus to formally describe the functional and performance specification of component based software and provide formal semantics for the proposed calculus. Then we provide a method to measure the dynamic complexity of software compositions based on the proposed calculus. Finally, we define a set of algebraic laws to manifest the complexity relations between different functionally equivalent components. We conduct a case study with a real software system and the results show that our method is able to calculate the dynamic complexity of component based systems, and the complexity can be reduced based on our algebraic laws.
Ling Shi 0002, Gan Zeng, Feng Sheng, Shuang Liu 0007
APSEC3
2019 Software Complexity Reduction by Automated Refactoring Schema
abstract
As the scale of software systems is growing rapidly, software complexity is becoming one of the main problems in software engineering. Higher complexity in software increases the potential risk and defects of software system, which makes it more difficult to analyze the correctness and improve the quality of software. In this paper, we present an automated refactoring schema to reduce the complexity of the component-based software. The main idea of our approach is to search a hierarchical software with a minimum hierarchical complexity and refactor the original software into it by reassembling several subcomponents into tightly coupled hierarchical ones. Besides, our approach can be easily adjusted to deal with some new situations, in which several types of constraints on partition of software components are given. Finally, we conduct a case study with Battery Management System (BMS) and the result demonstrates our approach can automatically and effectively reduce the structural complexity of software system.
Siteng Cao, Ling Shi 0002
TASE3
2018 HOMESCAN: Scrutinizing Implementations of Smart Home Integrations
abstract
A key feature of the booming smart home is the integration of a wide assortment of technologies, including various standards, proprietary communication protocols and heterogeneous platforms. Due to customization, unsatisfied assumptions and incompatibility in the integration, critical security vulnerabilities are likely to be introduced by the integration. Hence, this work addresses the security problems in smart home systems from an integration perspective, as a complement to numerous studies that focus on the analysis of individual techniques. We propose HOMESCAN, an approach that examines the security of the implementations of smart home systems. It extracts the abstract specification of application-layer protocols and internal behaviors of participants, so that it is able to conduct an end-to-end security analysis against various attack models. Applying HOMESCAN on three extensively-used smart home systems, we have found twelve non-trivial security vulnerabilities, which may lead to unauthorized remote control and credential leakage.
Kulani Mahadewa, Kailong Wang 0001, Guangdong Bai, Ling Shi 0002, Jin Song Dong 0001, Zhenkai Liang
ICECCS4
2018 A UTP semantics for communicating processes with shared variables and its formal encoding in PVS
abstract
Abstract CSP# (communicating sequential programs) is a modelling language designed for specifying concurrent systems by integrating CSP-like compositional operators with sequential programs updating shared variables. In this work, we define an observation-oriented denotational semantics in an open environment for the CSP# language based on the UTP framework. To deal with shared variables, we lift traditional event-based traces into mixed traces which consist of state-event pairs for recording process behaviours. To capture all possible concurrency behaviours between action/channel-based communications and global shared variables, we construct a comprehensive set of rules on merging traces from processes which run in parallel/interleaving. We also define refinement to check process equivalence and present a set of algebraic laws which are established based on our denotational semantics. We further encode our proposed denotational semantics into the PVS theorem prover. The encoding not only ensures the semantic consistency, but also builds up a theoretic foundation for machine-assisted verification of CSP# specifications.
Ling Shi 0002, Yang Liu 0003, Jun Sun 0001, Jin Song Dong 0001, Shengchao Qin
Formal Aspects Comput.1
2017 Towards Solving Decision Making Problems Using Probabilistic Model Checking
abstract
Decision making seeks the optimal choice for maximum rewards or minimal costs under certain conditions, requirements and constraints. Decision making problems in practice are usually complicated as they may be partially observable, stochastic, and dynamic. Such complexities make the traditional decision making methods like mathematical programming difficult to find the optimal choices effectively and efficiently. In this work, we conduct a case study with the 4-player Kuhn Poker game by combining machine learning with probabilistic model checking to generate optimal decisions. Experimental results show that the agent employing our method outperforms the conservative and bluffing players regardless of the positions of players.
Ling Shi 0002, Shuang Liu 0007, Jianye Hao, Jun Yang Koh, Jin Song Dong 0001
ICECCS1
2015 Sports Strategy Analytics Using Probabilistic Reasoning
abstract
The advance of analytics technology has attracted more attention and adoption from sports, although modeling and analyzing the dynamic (and uncertain) behaviors of sports are challenging. Formal methods have been strongly recommended to deal with complex systems by their rigorous semantics and powerful reasoning capabilities. In this paper, we present our initiative as the first to apply probabilistic model checking techniques to strategy analytics for tennis based on Markov Decision Processes (MDP). Our approach can derive insights such as prediction of winning chances and identification of best improvement. We evaluate the effectiveness of our approach through real-life case study.
Jin Song Dong 0001, Ling Shi 0002, Le Vu Nguyen Chuong, Kan Jiang, Jing Sun 0002
ICECCS2
2015 Event and Strategy Analytics
abstract
Model checking has been pervasive and successful in finding bugs in hardware and software systems, including real-time and probabilistic systems. Applying model checking to decision making is relative new and has an excellent potential to be compliment to data analytics and other Artificial Intelligent (AI) or Operational Research (OR) based decision making techniques. Our last 8 years research has focused on the development of PAT (Process Analysis Toolkit) [18] whichsupports modelling languages that combine the expressiveness of event, state, time and probability based modeling techniques to which model checking can be directly applied. The next direction for PAT is to move from verification to analytics, we call it "Event Analytics" with a special focus on "Strategy Analytics".
Jin Song Dong 0001, Jun Sun 0001, Yang Liu 0003, Yuan-Fang Li, Jing Sun 0002, Ling Shi 0002
TASE6
2013 A UTP Semantics for Communicating Processes with Shared Variables
Ling Shi 0002, Yang Liu 0003, Jun Sun 0001, Jin Song Dong 0001, Shengchao Qin
ICFEM1
2013 Modeling and verifying hierarchical real-time systems using stateful timed CSP
abstract
Modeling and verifying complex real-time systems are challenging research problems. The de facto approach is based on Timed Automata, which are finite state automata equipped with clock variables. Timed Automata are deficient in modeling hierarchical complex systems. In this work, we propose a language called Stateful Timed CSP and an automated approach for verifying Stateful Timed CSP models. Stateful Timed CSP is based on Timed CSP and is capable of specifying hierarchical real-time systems. Through dynamic zone abstraction, finite-state zone graphs can be generated automatically from Stateful Timed CSP models, which are subject to model checking. Like Timed Automata, Stateful Timed CSP models suffer from Zeno runs, that is, system runs that take infinitely many steps within finite time. Unlike Timed Automata, model checking with non-Zenoness in Stateful Timed CSP can be achieved based on the zone graphs. We extend the PAT model checker to support system modeling and verification using Stateful Timed CSP and show its usability/scalability via verification of real-world systems.
Jun Sun 0001, Yang Liu 0003, Jin Song Dong 0001, Yan Liu 0012, Ling Shi 0002, Étienne André 0001
ACM Trans. Softw. Eng. Methodol.5
2012 An Analytical and Experimental Comparison of CSP Extensions and Tools
Ling Shi 0002, Yang Liu 0003, Jun Sun 0001, Jin Song Dong 0001, Gustavo Carvalho
ICFEM1
2008 A Bigraphical Model of WSBPEL
abstract
In this paper, we give a bigraphical model for web services composition. We investigate how to represent scope-based compensation handing mechanism by means of Bigraphical Reactive Systems (13RSs for short), which have been proposed to provide a uniform way to model spatially distributed systems that both compute and communicate. The service composition language we focus on is WSBPEL, which is the standard of web service composition and orchestration. This bigraphical model can be regarded as a unifying semantics of BPEL-like languages with the key concepts related to compensation handling. The rationality of the model is discussed by investigating the relationship between BPEL language and BRSs. Based on the bigraphical model, the algebraic laws for BPEL are proved as well.
Min Zhang 0007, Ling Shi 0002, Longfei Zhu, Libo Feng, Geguang Pu
TASE2