Kailong Wang 0001

dblp:171/1258 · DBLP profile ↗
← Back
42ranked-venue papers
4as first author
38since 2021 · last 2026
0000-0002-3977-6573ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 26 · 1 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Artificial intelligence and machine learning · 6 · 6 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 5 since 2021Security and privacy · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 On the Feasibility of Using MultiModal LLMs to Execute AR Social Engineering Attacks
abstract
Augmented Reality (AR) and Multimodal Large Language Models (LLMs) are rapidly evolving, providing unprecedented capabilities for human-computer interaction. However, their integration introduces a new attack surface for Social Engineering (SE). In this paper, we systematically investigate the feasibility of orchestrating AR-driven Social Engineering attacks using Multimodal LLM for the first time, via our proposed SEAR framework, which operates through three key phases: (1) AR-based social context synthesis, which fuses Multimodal inputs (visual, auditory and environmental cues); (2) role-based Multimodal RAG (Retrieval-Augmented Generation), which dynamically retrieves and integrates social context; and (3) ReInteract social engineering agents, which execute adaptive multiphase attack strategies through inference interaction loops. To verify SEAR, we conducted an IRB-approved study with 60 participants and build a novel dataset of 180 annotated conversations in different social scenarios (e.g., coffee shops, networking events). Our results show that SEAR is highly effective at eliciting high-risk behaviors (e.g., 93.3% of participants susceptible to email phishing). The framework was particularly effective in building trust, with 85% of targets willing to accept an attacker's call after an interaction. Also, we identified notable limitations such as authenticity gaps. This work provides proof-of-concept for AR-LLM driven social engineering attacks and insights for developing defenses against next-generation AR/LLM-based SE threats.
Ting Bi, Chenghang Ye, Zheyu Yang 0002, Ziyi Zhou 0006, Cui Tang, Kailong Wang 0001, Liting Zhou, Yang Yang 0060, Tianlong Yu
AAAI8
2026 STEAMROLLER: A Multi-Agent System for Inclusive Automatic Speech Recognition for People Who Stutter
abstract
People who stutter (PWS) face systemic exclusion in today’s voice-driven society, where access to voice assistants, authentication systems, and remote work tools increasingly depends on fluent speech. Current automatic speech recognition (ASR) systems, trained predominantly on fluent speech, fail to serve millions of PWS worldwide. We present STEAMROLLER, a real time system that transforms stuttered speech into fluent output through a novel multi-stage, multi-agent AI pipeline. Our approach addresses three critical technical challenges: (1) the difficulty of direct speech to speech conversion for disfluent input, (2) semantic distortions introduced during ASR transcription of stuttered speech, and (3) latency constraints for real time communication. STEAMROLLER employs a three stage architecture comprising ASR transcription, multi-agent text repair, and speech synthesis, where our core innovation lies in a collaborative multi-agent framework that iteratively refines transcripts while preserving semantic intent. Experiments on the FluencyBank dataset and a user study demonstrates clear word error rate (WER) reduction and strong user satisfaction. Beyond immediate accessibility benefits, fine tuning ASR on STEAMROLLER repaired speech further yields additional WER improvements, creating a pathway toward inclusive AI ecosystems.
Yi Liu 0069, Yuekang Li, Ling Shi 0002, Kailong Wang 0001
AAAI5
2026 R2IF: Aligning Reasoning with Decisions via Composite Rewards for Interpretable LLM Function Calling
abstract
Function calling empowers large language models (LLMs) to interface with external tools, yet existing RL-based approaches suffer from misalignment between reasoning processes and tool-call decisions.We propose R2IF, a reasoning-aware RL framework for interpretable function calling, adopting a composite reward integrating format/correctness constraints, Chain-of-Thought Effectiveness Reward (CER), and Specification-Modification-Value (SMV) reward, optimized via GRPO.Experiments on BFCL/ACEBench show R2IF outperforms baselines by up to 34.62% (Llama3.2-3B on BFCL) with positive Average CoT Effectiveness (0.05 for Llama3.2-3B),enhancing both function-calling accuracy and interpretability for reliable tool-augmented LLM deployment.
Aijia Cheng, Kailong Wang 0001, Ling Shi 0002
ACL (1)2
2026 Uncovering Logit Suppression Vulnerabilities in LLM Safety Alignment
Yi Liu 0069, Yuekang Li, Ling Shi 0002, Gelei Deng, Shengquan Chen, Kailong Wang 0001
ICPR (2)7
2026 When Safe Models Merge Into Danger: Exploiting Latent Vulnerabilities in LLM Fusion
Shide Zhou, Tianlong Yu, Kailong Wang 0001
ICPR (2)6
2026 RefineRAG: Word-Level Poisoning Attacks via Retriever-Guided Text Refinement
Guanyu Wang 0005, Kailong Wang 0001
ICPR (2)3
2026 Privacy Protection Against Personalized Text-to-Image Synthesis via Cross-image Consistency Constraints
abstract
The rapid advancement of diffusion models and personalization techniques has made it possible to recreate individual portraits from just a few publicly available images. While such capabilities empower various creative applications, they also introduce serious privacy concerns, as adversaries can exploit them to generate highly realistic impersonations. To counter these threats, anti-personalization methods have been proposed, which add adversarial perturbations to published images to disrupt the training of personalization models. However, existing approaches largely overlook the intrinsic multi-image nature of personalization and instead adopt a naive strategy of applying perturbations independently, as commonly done in single-image settings. This neglects the opportunity to leverage inter-image relationships for stronger privacy protection. Therefore, we advocate for a group-level perspective on privacy protection against personalization. Specifically, we introduce Cross-image Anti-Personalization (CAP), a novel framework that enhances resistance to personalization by enforcing style consistency across perturbed images. Furthermore, we develop a dynamic ratio adjustment strategy that adaptively balances the impact of the consistency loss throughout the attack iterations. Extensive experiments on the classical CelebA-HQ and VGGFace2 benchmarks show that CAP outperforms eight existing methods.
Guanyu Wang 0005, Kailong Wang 0001, Yihao Huang 0001, Mingyi Zhou, Geguang Pu, Li Li 0029
ICMR2
2026 MalModel: hiding malicious payload in mobile deep learning models with black-box backdoor attack
Jiayi Hua, Kailong Wang 0001, Meizhen Wang, Guangdong Bai, Xiapu Luo, Haoyu Wang 0001
Autom. Softw. Eng.2
2026 Exposing the Ghost in the Transformer: Abnormal Detection for Large Language Models via Hidden State Forensics
abstract
The widespread adoption of Large Language Models (LLMs) in critical applications has introduced severe reliability and security risks, as LLMs remain vulnerable to notorious threats such as hallucinations, jailbreak attacks, and backdoor exploits. These vulnerabilities have been weaponized by malicious actors, leading to unauthorized access, widespread misinformation, and compromised LLM-embedded system integrity. In this work, we introduce a novel approach to detecting abnormal behaviors in LLMs via hidden state forensics. By systematically inspecting layer-specific activation patterns, we develop a general framework that can efficiently identify a range of security threats in real-time without imposing prohibitive computational costs. Extensive experiments indicate detection accuracies exceeding 95% and consistently robust performance across multiple models in most scenarios, while preserving the ability to detect novel attacks effectively. Furthermore, the computational overhead remains minimal, with detector inference taking merely fractions of a second. The significance of this work lies in proposing a promising strategy to reinforce the security of LLM-integrated systems, paving the way for safer and more reliable deployment in high-stakes domains. By enabling real-time detection that can also support the mitigation of abnormal behaviors, it represents a meaningful step toward ensuring the trustworthiness of AI systems amid rising security challenges.
Shide Zhou, Kailong Wang 0001, Ling Shi 0002, Haoyu Wang 0001
IEEE Trans. Inf. Forensics Secur.2
2026 SPOLRE: Semantic Preserving Object Layout Reconstruction for Image Captioning System Testing
abstract
Image captioning (IC) systems, including Microsoft Azure Cognitive Service, are commonly utilized to convert image content into descriptive natural language. However, inaccuracies in caption generation can lead to serious misinterpretations. Advanced testing techniques such as MetaIC and ROME have been developed to mitigate these issues, yet they encounter notable challenges. First, these strategies demand intensive labor, relying on detailed manual annotations like bounding box data of objects to create test cases. Second, the realism of the generated images is compromised, with MetaIC adding unrelated objects and ROME failing to remove objects effectively. Finally, the capability to generate diversified test suites is restricted. MetaIC is limited to only inserting specific objects to prevent overlap, whereas ROME can generate only \(3^{n}-2^{n}\) variations of test cases from an original seed image containing \( n \) objects. In this study, we present SPOLRE, a novel automated tool designed for semantic preserving object layout reconstruction in image captioning system testing. SPOLRE is based on the insight that modifying the arrangement of objects within an image does not alter its inherent semantics. We utilize four semantic preserving transformation techniques—translation, rotation, mirroring, and scaling—to modify object layouts autonomously, eliminating the need for manual annotation. This approach enables the creation of realistic and varied test suites for IC system testing. Our extensive testing demonstrates that more than 75% of survey respondents find the images produced by SPOLRE more realistic compared to those generated by SOTA methods. Additionally, SPOLRE exhibits outstanding performance in identifying caption errors, detecting 31,544 incorrect captions across seven IC systems with an average precision of 91.62%. This significantly outperforms other methods, which only achieve 85.65% accuracy on average and identify 17,160 incorrect captions. Notably, SPOLRE exposes 6,236 unique issues within Microsoft Azure Cognitive Service, highlighting its effectiveness against one of the most advanced IC systems available.
Yi Liu 0069, Guanyu Wang 0005, Gelei Deng, Kailong Wang 0001, Yang Liu 0003, Haoyu Wang 0001
ACM Trans. Softw. Eng. Methodol.5
2026 NeuSemSlice: Towards Effective DNN Model Maintenance via Neuron-Level Semantic Slicing
abstract
Deep Neural Networks (DNNs), extensively applied across diverse disciplines, are characterized by their integrated and monolithic architectures, setting them apart from conventional software systems. This architectural difference introduces particular challenges to maintenance tasks, such as model restructure (e.g., model compression), re-adaptation (e.g., fitting new samples), and incremental development (e.g., continual knowledge accumulation). Prior research addresses these challenges by identifying task-critical neuron layers and dividing neural networks into semantically similar sequential modules. However, such layer-level approaches fail to precisely identify and manipulate neuron-level semantic components, restricting their applicability to finer-grained model maintenance tasks. In this work, we implement NeuSemSlice, a novel framework that introduces the semantic slicing technique to effectively identify critical neuron-level semantic components in DNN models for semantic-aware model maintenance tasks. Specifically, semantic slicing identifies, categorizes, and merges critical neurons across different categories and layers according to their semantic similarity, enabling their flexibility and effectiveness in the subsequent tasks. For semantic-aware model maintenance tasks, we provide a series of novel strategies based on semantic slicing to enhance NeuSemSlice. They include semantic components (i.e., critical neurons) preservation for model restructure, critical neuron tuning for model re-adaptation, and non-critical neuron training for model incremental development. A thorough evaluation has demonstrated that NeuSemSlice significantly outperforms baselines in all three tasks.
Shide Zhou, Tianlin Li, Yihao Huang 0001, Ling Shi 0002, Kailong Wang 0001, Yang Liu 0003, Haoyu Wang 0001
ACM Trans. Softw. Eng. Methodol.5
2026 Assessing Privacy Disclosure Compliance of Android Third-Party SDKs
Mark Huasong Meng, Chuan Yan, Zhang Qing cnwatcher, Kailong Wang 0001, Sin G. Teo, Guangdong Bai, Jin Song Dong 0001
IEEE Trans. Software Eng.5
2025 Decoding Secret Memorization in Code LLMs Through Token-Level Characterization
abstract
Code Large Language Models (LLMs) have demonstrated remarkable capabilities in generating, understanding, and manipulating programming code. However, their training process inadvertently leads to the memorization of sensitive information, posing severe privacy risks. Existing studies on memorization in LLMs primarily rely on prompt engineering techniques, which suffer from limitations such as widespread hallucination and inefficient extraction of the target sensitive information. In this paper, we present a novel approach to characterize real and fake secrets generated by Code LLMs based on token probabilities. We identify four key characteristics that differentiate genuine secrets from hallucinated ones, providing insights into distinguishing real and fake secrets. To overcome the limitations of existing works, we propose DeSec,a two-stage method that leverages token-level features derived from the identified characteristics to guide the token decoding process. DeSec consists of constructing an offline token scoring model using a proxy Code LLM and employing the scoring model to guide the decoding process by reassigning token likelihoods. Through extensive experiments on four state-of-the-art Code LLMs using a diverse dataset, we demonstrate the superior performance of DeSec in achieving a higher plausible rate and extracting more real secrets compared to existing baselines. Our findings highlight the effectiveness of our token-level approach in enabling an extensive assessment of the privacy leakage risks associated with Code LLMs.
Yuqing Nie, Chong Wang 0013, Kailong Wang 0001, Guoai Xu, Guosheng Xu 0001, Haoyu Wang 0001
ICSE3
2025 Understanding the Effectiveness of Coverage Criteria for Large Language Models: A Special Angle from Jailbreak Attacks
abstract
Large language models (LLMs) have revolutionized artificial intelligence, but their increasing deployment across critical domains has raised concerns about their abnormal behaviors when faced with malicious attacks. Such vulnerability alerts the widespread inadequacy of pre-release testing. In this paper, we conduct a comprehensive empirical study to evaluate the effectiveness of traditional coverage criteria in identifying such inadequacies, exemplified by the significant security concern of jailbreak attacks. Our study begins with a clustering analysis of the hidden states of LLMs, revealing that the embedded characteristics effectively distinguish between different query types. We then systematically evaluate the performance of these criteria across three key dimensions: criterion level, layer level, and token level. Our research uncovers significant differences in neuron coverage when LLMs process normal versus jailbreak queries, aligning with our clustering experiments. Leveraging these findings, we propose three practical applications of coverage criteria in the context of LLM security testing. Specifically, we develop a realtime jailbreak detection mechanism that achieves high accuracy (93.61 % on average) in classifying queries as normal or jailbreak. Furthermore, we explore the use of coverage levels to prioritize test cases, improving testing efficiency by focusing on high-risk interactions and removing redundant tests. Lastly, we introduce a coverage-guided approach for generating jailbreak attack examples, enabling systematic refinement of prompts to uncover vulnerabilities. This study improves our understanding of LLM security testing, enhances their safety, and provides a foundation for developing more robust AI applications.
Shide Zhou, Tianlin Li, Kailong Wang 0001, Yihao Huang 0001, Ling Shi 0002, Yang Liu 0003, Haoyu Wang 0001
ICSE3
2025 SEAR: A Multimodal Dataset for Analyzing AR-LLM-Driven Social Engineering Behaviors
Tianlong Yu, Chenghang Ye, Zheyu Yang 0002, Ziyi Zhou 0006, Cui Tang, Kailong Wang 0001, Liting Zhou, Yang Yang 0060, Ting Bi
ACM Multimedia8
2025 MiniScope: Automated UI Exploration and Privacy Inconsistency Detection of MiniApps via Two-phase Iterative Hybrid Analysis
abstract
The advent of MiniApps, operating within larger SuperApps, has revolutionized user experiences by offering a wide range of services without the need for individual app downloads. However, this convenience has raised significant privacy concerns, as these MiniApps often require access to sensitive data, potentially leading to privacy violations. Despite existing privacy regulations and platform guidelines, there is a lack of effective mechanisms to safeguard user privacy fully. To address this critical gap, we introduce MiniScope , a novel two-phase hybrid analysis approach, specifically designed for the MiniApp environment. This approach overcomes the limitations of existing static analysis techniques by incorporating UI transition states analysis, cross-package callback control flow resolution, and automated iterative UI exploration. This allows for a comprehensive understanding of MiniApps’ privacy practices, addressing the unique challenges of sub-package loading and event-driven callbacks. Our empirical evaluation of over 120K MiniApps using MiniScope demonstrates its effectiveness in identifying privacy inconsistencies. The results reveal significant issues, with 5.7% of MiniApps over-collecting private data and 33.4% overclaiming data collection. We have responsibly disclosed our findings to 2,282 developers, receiving 44 acknowledgments. These findings emphasize the urgent need for more precise privacy monitoring systems and highlight the responsibility of SuperApp operators to enforce stricter privacy measures.
Shenao Wang 0001, Yuekang Li, Kailong Wang 0001, Yi Liu 0069, Hui Li 0006, Yang Liu 0003, Haoyu Wang 0001
ACM Trans. Softw. Eng. Methodol.3
2024 SC-WGAN: GAN-Based Oversampling Method for Network Intrusion Detection
Wuxia Bai, Kailong Wang 0001, Kai Chen 0012, Shenghui Li, Bingqian Li
ICECCS2
2024 Analyzing Excessive Permission Requests in Google Workspace Add-Ons
Liuhuo Wan, Chuan Yan, Mark Huasong Meng, Kailong Wang 0001, Haoyu Wang 0001
ICECCS4
2024 Semantic-Enhanced Indirect Call Analysis with Large Language Models
abstract
In contemporary software development, the widespread use of indirect calls to achieve dynamic features poses challenges in constructing precise control flow graphs (CFGs), which further impacts the performance of downstream static analysis tasks. To tackle this issue, various types of indirect call analyzers have been proposed. However, they do not fully leverage the semantic information of the program, limiting their effectiveness in real-world scenarios.
Baijun Cheng, Cen Zhang, Kailong Wang 0001, Ling Shi 0002, Yang Liu 0003, Haoyu Wang 0001, Yao Guo 0001, Ding Li 0001, Xiangqun Chen
ASE3
2024 GlitchProber: Advancing Effective Detection and Mitigation of Glitch Tokens in Large Language Models
abstract
Large language models (LLMs) have achieved unprecedented success in the field of natural language processing. However, the black-box nature of their internal mechanisms has brought many concerns about their trustworthiness and interpretability. Recent research has discovered a class of abnormal tokens in the model's vocabulary space and named them "glitch tokens". Those tokens, once included in the input, may induce the model to produce incorrect, irrelevant, or even harmful results, drastically undermining the reliability and practicality of LLMs.
Wuxia Bai, Yuxi Li 0010, Mark Huasong Meng, Kailong Wang 0001, Ling Shi 0002, Li Li 0029, Jun Wang 0020, Haoyu Wang 0001
ASE5
2024 Models Are Codes: Towards Measuring Malicious Code Poisoning Attacks on Pre-trained Model Hubs
abstract
The proliferation of pre-trained models (PTMs) and datasets has led to the emergence of centralized model hubs like Hugging Face, which facilitate collaborative development and reuse. However, recent security reports have uncovered vulnerabilities and instances of malicious attacks within these platforms, highlighting growing security concerns. This paper presents the first systematic study of malicious code poisoning attacks on pre-trained model hubs, focusing on the Hugging Face platform. We conduct a comprehensive threat analysis, develop a taxonomy of model formats, and perform root cause analysis of vulnerable formats. While existing tools like Fickling and ModelScan offer some protection, they face limitations in semantic-level analysis and comprehensive threat detection. To address these challenges, we propose MalHug, an end-to-end pipeline tailored for Hugging Face that combines dataset loading script extraction, model deserialization, in-depth taint analysis, and heuristic pattern matching to detect and classify malicious code poisoning attacks in datasets and models. In collaboration with Ant Group, a leading financial technology company, we have implemented and deployed MalHug on a mirrored Hugging Face instance within their infrastructure, where it has been operational for over three months. During this period, MalHug has monitored more than 705K models and 176K datasets, uncovering 91 malicious models and 9 malicious dataset loading scripts. These findings reveal a range of security threats, including reverse shell, browser credential theft, and system reconnaissance. This work not only bridges a critical gap in understanding the security of the PTM supply chain but also provides a practical, industry-tested solution for enhancing the security of pre-trained model hubs.
Shenao Wang 0001, Yanjie Zhao 0001, Xinyi Hou, Kailong Wang 0001, Peiming Gao, Haoyu Wang 0001
ASE5
2024 Towards Robust Detection of Open Source Software Supply Chain Poisoning Attacks in Industry Environments
abstract
The exponential growth of open-source package ecosystems, particularly NPM and PyPI, has led to an alarming increase in software supply chain poisoning attacks. Existing static analysis methods struggle with high false positive rates and are easily thwarted by obfuscation and dynamic code execution techniques. While dynamic analysis approaches offer improvements, they often suffer from capturing non-package behaviors and employing simplistic testing strategies that fail to trigger sophisticated malicious behaviors. To address these challenges, we present OSCAR, a robust dynamic code poisoning detection pipeline for NPM and PyPI ecosystems. OSCAR fully executes packages in a sandbox environment, employs fuzz testing on exported functions and classes, and implements aspect-based behavior monitoring with tailored API hook points. We evaluate OSCAR against six existing tools using a comprehensive benchmark dataset of real-world malicious and benign packages. OSCAR achieves an F1 score of 0.95 in NPM and 0.91 in PyPI, confirming that OSCAR is as effective as the current state-of-the-art technologies. Furthermore, for benign packages exhibiting characteristics typical of malicious packages, OSCAR reduces the false positive rate by an average of 32.06% in NPM (from 34.63% to 2.57%) and 39.87% in PyPI (from 41.10% to 1.23%), compared to other tools, significantly reducing the workload of manual reviews in real-world deployments. In cooperation with Ant Group, a leading financial technology company, we have deployed OSCAR on its NPM and PyPI mirrors since January 2023, identifying 10,404 malicious NPM packages and 1,235 malicious PyPI packages over 18 months. This work not only bridges the gap between academic research and industrial application in code poisoning detection but also provides a robust and practical solution that has been thoroughly tested in a real-world industrial setting.
Shenao Wang 0001, Yanjie Zhao 0001, Peiming Gao, Kailong Wang 0001, Haoyu Wang 0001
ASE7
2024 MASTERKEY: Automated Jailbreaking of Large Language Model Chatbots
Gelei Deng, Yi Liu 0069, Yuekang Li, Kailong Wang 0001, Ying Zhang 0066, Zefeng Li, Haoyu Wang 0001, Tianwei Zhang 0004, Yang Liu 0003
NDSS4
2024 Essential or Excessive? MINDAEXT: Measuring Data Minimization Practices among Browser Extensions
abstract
Since browser extensions are prevailingly executed in the background to enable extra functionalities and enhance the user experience for web browsers, the potential over-collection of personal data beyond the necessity for given purposes is always ignored by ordinary users. Existing privacy regulations, such as the principle of Data Minimization in GDPR, have provided the criteria that only directly relevant and necessary data for specified purposes should be collected. Various tools have made efforts to examine the compliance of data minimization and its equivalent in different application domains. To our knowledge, in the area of browser extensions, there is still a gap between the general data minimization principle and precisely defined extension behaviors. We propose MINDAExT, a framework that takes one step further to automatically examine end-to-end data minimization practices in browser extensions by description text analysis and hybrid program analysis techniques. In our large-scale measurement, covering around 200K extensions collected in October 2023, we find that 38.0 % of extensions are likely to collect private user data outside their essential functionality scopes. They are distributed across all categories, exhibiting distinct patterns of the target data types. Our evaluation shows that MINDAEXT can detect the data over-collection with a precision of 74.3 %.
Yuxi Ling, Kailong Wang 0001, Guangdong Bai, Jin Song Dong 0001
SANER4
2024 Don't Bite Off More than You Can Chew: Investigating Excessive Permission Requests in Trigger-Action Integrations
abstract
Web-based trigger-action platforms (TAP) allow users to integrate Internet of Things (IoT) systems and online services into trigger-action integrations (TAIs), facilitating rich automation tasks known as applets. Despite their benefits, these integrations~(typically involving the TAP, trigger, and action service providers) pose significant security and privacy challenges, such as mis-triggering and data leakage. This work investigates cross-entity permission management within TAIs to address the underlying causes of these security and privacy issues, emphasizing permission-functionality consistency to ensure fairness in permission requests. We introduce PFCon, a system that leverages GPT-based language models for analyzing required and requested permissions, revealing excessive permission requests in a large-scale study of IFTTT TAP. Our findings highlight the need for service providers to enforce permission-functionality consistency, raising awareness of the importance of security and privacy in TAI.
Liuhuo Wan, Kailong Wang 0001, Kulani Mahadewa, Haoyu Wang 0001, Guangdong Bai
WWW2
2024 Is It Safe to Share Your Files? An Empirical Security Analysis of Google Workspace
abstract
The increasing demand for remote work and virtual interactions has heightened the usage of business collaboration platforms~(BCPs), with Google Workspace as a prominent example. These platforms enhance team collaboration by integrating Google Docs, Slides, Calendar, and feature-rich third-party applications (add-ons). However, such integration of multiple users and entities has inadvertently introduced new and complex attack surfaces, elevating security and privacy risks in resource management to unprecedented levels. In this study, we conduct a systematic study on the effectiveness of the cross-entity resource management in Google Workspace, the most popular BCP. Our study unveils the access control enforcement in real-world BCPs for the first time. Based on this, we formulate the attack surfaces inherent in BCPs and conduct a comprehensive assessment, pinpointing three vulnerability types leading to distinct attacks. An analysis of 4,732 marketplace add-ons reveals that approximately 70% are potentially vulnerable to these attacks. We propose robust countermeasures to improve BCP security, urging immediate action and setting a foundation for future research.
Liuhuo Wan, Kailong Wang 0001, Haoyu Wang 0001, Guangdong Bai
WWW2
2024 WalletRadar: towards automating the detection of vulnerabilities in browser-based cryptocurrency wallets
Pengcheng Xia 0001, Zhaowen Lin, Pengbo Duan, Ningyu He, Kailong Wang 0001, Tianming Liu 0002, Yinliang Yue, Guoai Xu, Haoyu Wang 0001
Autom. Softw. Eng.7
2024 Drowzee: Metamorphic Testing for Fact-Conflicting Hallucination Detection in Large Language Models
abstract
Large language models (LLMs) have revolutionized language processing, but face critical challenges with security, privacy, and generating hallucinations — coherent but factually inaccurate outputs. A major issue is fact-conflicting hallucination (FCH), where LLMs produce content contradicting ground truth facts. Addressing FCH is difficult due to two key challenges: 1) Automatically constructing and updating benchmark datasets is hard, as existing methods rely on manually curated static benchmarks that cannot cover the broad, evolving spectrum of FCH cases. 2) Validating the reasoning behind LLM outputs is inherently difficult, especially for complex logical relations. To tackle these challenges, we introduce a novel logic-programming-aided metamorphic testing technique for FCH detection. We develop an extensive and extensible framework that constructs a comprehensive factual knowledge base by crawling sources like Wikipedia, seamlessly integrated into D rowzee . Using logical reasoning rules, we transform and augment this knowledge into a large set of test cases with ground truth answers. We test LLMs on these cases through template-based prompts, requiring them to provide reasoned answers. To validate their reasoning, we propose two semantic-aware oracles that assess the similarity between the semantic structures of the LLM answers and ground truth. Our approach automatically generates useful test cases and identifies hallucinations across six LLMs within nine domains, with hallucination rates ranging from 24.7% to 59.8%. Key findings include LLMs struggling with temporal concepts, out-of-distribution knowledge, and lack of logical reasoning capabilities. The results show that logic-based test cases generated by D rowzee effectively trigger and detect hallucinations. To further mitigate the identified FCHs, we explored model editing techniques, which proved effective on a small scale (with edits to fewer than 1000 knowledge pieces). Our findings emphasize the need for continued community efforts to detect and mitigate model hallucinations.
Ningke Li, Yuekang Li, Yi Liu 0069, Ling Shi 0002, Kailong Wang 0001, Haoyu Wang 0001
Proc. ACM Program. Lang.5
2024 Beyond Fidelity: Explaining Vulnerability Localization of Learning-Based Detectors
abstract
Vulnerability detectors based on deep learning (DL) models have proven their effectiveness in recent years. However, the shroud of opacity surrounding the decision-making process of these detectors makes it difficult for security analysts to comprehend. To address this, various explanation approaches have been proposed to explain the predictions by highlighting important features, which have been demonstrated effective in domains such as computer vision and natural language processing. Unfortunately, there is still a lack of in-depth evaluation of vulnerability-critical features, such as fine-grained vulnerability-related code lines, learned and understood by these explanation approaches. In this study, we first evaluate the performance of ten explanation approaches for vulnerability detectors based on graph and sequence representations, measured by two quantitative metrics including fidelity and vulnerability line coverage rate. Our results show that fidelity alone is insufficent for evaluating these approaches, as fidelity incurs significant fluctuations across different datasets and detectors. We subsequently check the precision of the vulnerability-related code lines reported by the explanation approaches, and find poor accuracy in this task among all of them. This can be attributed to the inefficiency of explainers in selecting important features and the presence of irrelevant artifacts learned by DL-based detectors.
Baijun Cheng, Shengming Zhao, Kailong Wang 0001, Meizhen Wang, Guangdong Bai, Yao Guo 0001, Lei Ma 0003, Haoyu Wang 0001
ACM Trans. Softw. Eng. Methodol.3
2024 Large Language Models for Software Engineering: A Systematic Literature Review
abstract
Large Language Models (LLMs) have significantly impacted numerous domains, including Software Engineering (SE). Many recent publications have explored LLMs applied to various SE tasks. Nevertheless, a comprehensive understanding of the application, effects, and possible limitations of LLMs on SE is still in its early stages. To bridge this gap, we conducted a Systematic Literature Review (SLR) on LLM4SE, with a particular focus on understanding how LLMs can be exploited to optimize processes and outcomes. We selected and analyzed 395 research articles from January 2017 to January 2024 to answer four key Research Questions (RQs). In RQ1, we categorize different LLMs that have been employed in SE tasks, characterizing their distinctive features and uses. In RQ2, we analyze the methods used in data collection, pre-processing, and application, highlighting the role of well-curated datasets for successful LLM for SE implementation. RQ3 investigates the strategies employed to optimize and evaluate the performance of LLMs in SE. Finally, RQ4 examines the specific SE tasks where LLMs have shown success to date, illustrating their practical contributions to the field. From the answers to these RQs, we discuss the current state-of-the-art and trends, identifying gaps in existing research, and highlighting promising areas for future study. Our artifacts are publicly available at https://github.com/security-pride/LLM4SE_SLR .
Xinyi Hou, Yanjie Zhao 0001, Yue Liu 0011, Zhou Yang 0003, Kailong Wang 0001, Li Li 0029, Xiapu Luo, David Lo 0001, John C. Grundy, Haoyu Wang 0001
ACM Trans. Softw. Eng. Methodol.5
2023 Understanding and Tackling Label Errors in Deep Learning-Based Vulnerability Detection (Experience Paper)
abstract
Software system complexity and security vulnerability diversity are plausible sources of the persistent challenges in software vulnerability research. Applying deep learning methods for automatic vulnerability detection has been proven an effective means to complement traditional detection approaches. Unfortunately, lacking well-qualified benchmark datasets could critically restrict the effectiveness of deep learning-based vulnerability detection techniques. Specifically, the long-term existence of erroneous labels in the existing vulnerability datasets may lead to inaccurate, biased, and even flawed results.
Xu Nie, Ningke Li, Kailong Wang 0001, Shangguang Wang, Xiapu Luo, Haoyu Wang 0001
ISSTA3
2023 MalWuKong: Towards Fast, Accurate, and Multilingual Detection of Malicious Code Poisoning in OSS Supply Chains
abstract
In the face of increased threats within software registries and management systems, we address the critical need for effective malicious code detection. In this paper, we propose an innovative approach that integrates source code slicing, inter-procedural analysis, and cross-file inter-procedural analysis, thereby enhancing the detection precision and reducing false positives. This approach has been encapsulated within a multi-analysis-based framework for automatic detection of malicious code in real-world software packages. In its application to major third-party software registries like PyPI and NPM, our framework has proven effective, identifying 130 malicious packages from a total of 169,640 monitored over a continuous period of five weeks. This work advances the current state-of-the-art solution to malicious code detection, demonstrating significant practical impact in strengthening the software supply chain defense.
Ningke Li, Shenao Wang 0001, Mingxi Feng, Kailong Wang 0001, Meizhen Wang, Haoyu Wang 0001
ASE4
2023 Wemint:Tainting Sensitive Data Leaks in WeChat Mini-Programs
abstract
Mini-programs (MiniApps), lightweight versions of full-featured mobile apps that run inside a host app such as WeChat, have become increasingly popular due to their simplified and convenient user experiences. However, MiniApps raise new security and privacy concerns as they can access partially or all of host apps' system resources, including sensitive personal data. While taint detection has been proven effective in addressing this kind of concerns, existing taint detection techniques for mobile apps cannot be directly applied to MiniApps. The main reason is that the key logics of MiniApps are usually written in J avaScript, and its intrinsic characteristics (function-level scope, dynamic types, synchronous programming, and code obfuscation) prevent existing taint detection techniques from precisely propagating the taints. To address this problem, we propose a novel taint detection technique, Wemint, that detects sensitive information leaks in MiniApps. Specifically, Wemint facilitates taint propagation via building a context-based model based on the operational prin-ciple of MiniApps and J avaScript, and addresses asynchronous function calls by modeling their callbacks explicitly in taint rules. In addition, due to the adoption of Abstract Syntax Trees (ASTs) for code representation during taint detection, Wemint exhibits better robustness against the commonly-applied code obfuscation. Our experimental results show that Wemint can effectively detect sensitive information leaks in WeChat MiniApps, as well as trace the path of sensitive data flows. By applying Wemint to over 20K suspicious MiniApps, we found that over 7.5K (36.5 %) of them have sensitive data leaks, and Wemint outperforms the state-of-the-art DoubleX based techniques in detecting these leaks.
Shi Meng, Liu Wang 0002, Shenao Wang 0001, Kailong Wang 0001, Xusheng Xiao, Guangdong Bai, Haoyu Wang 0001
ASE4
2023 Enhancing Federated Learning Robustness Using Data-Agnostic Model Pruning
Mark Huasong Meng, Sin G. Teo, Guangdong Bai, Kailong Wang 0001, Jin Song Dong 0001
PAKDD (2)4
2022 Are they Toeing the Line? Diagnosing Privacy Compliance Violations among Browser Extensions
abstract
Browser extensions have emerged as integrated characteristics in modern browsers, with the aim to boost the online browsing experience. Their advantageous position between a user and the Internet endows them with easy access to the user’s sensitive data, which has raised mounting privacy concerns from both legislators and extension users. In this work, we propose an end-to-end approach to automatically diagnosing the privacy compliance violations among extensions. It analyzes the compliance of privacy policy versus regulation requirements and their actual privacy-related practices during runtime. This approach can serve the extension users, developers and store operators as an efficient and practical detection mechanism for privacy compliance violations.
Yuxi Ling, Kailong Wang 0001, Guangdong Bai, Haoyu Wang 0001, Jin Song Dong 0001
ASE2
2022 Assessing certificate validation user interfaces of WPA supplicants
abstract
WPA (Wi-Fi Protected Access) Enterprise is the de facto standard for safeguarding enterprise-level wireless networks. It relies on Transport Layer Security (TLS) to establish a secure tunnel during its authentication process, and thus the notoriously error-prone certificate validation may haunt it. Incorrect validation may lead to the SSL/TLS man-in-the-middle attack, or the evil twin attack in the context of wireless networking, where the supplicant connects and unwittingly sends authentication credentials to a fake access point.
Kailong Wang 0001, Yuwei Zheng, Zhang Qing cnwatcher, Guangdong Bai, Mingchuang Qin, Jin Song Dong 0001
MobiCom1
2021 It's Not Just the Site, It's the Contents: Intra-domain Fingerprinting Social Media Websites Through CDN Bursts
abstract
The website fingerprinting (or inter-domain WSF), enhanced by various machine learning techniques, has shown its power to identify websites a user has visited. To our best knowledge, a finer-grained problem of web page fingerprinting (or intra-domain WPF) has not been systematically studied by our research community. The WPF attackers, such as government agencies enforcing Internet censorship, are keen to identify the particular web pages (e.g., a political dissident’s social media page) visited by the target user.
Kailong Wang 0001, Guangdong Bai, Ryan Kok Leong Ko, Jin Song Dong 0001
WWW1
2021 Scrutinizing Implementations of Smart Home Integrations
abstract
A key feature of the booming smart home is the integration of a wide assortment of technologies, including various standards, proprietary communication protocols and heterogeneous platforms. Due to customization, unsatisfied assumptions and incompatibility in the integration, critical security vulnerabilities are likely to be introduced by the integration. Hence, this work addresses the security problems in smart home systems from anintegrationperspective, as a complement to numerous studies that focus on the analysis of individual techniques. We propose HomeScan, an approach that examines the security of the implementations of smart home systems. It extracts the abstract specification of application-layer protocols and internal behaviors of entities, so that it is able to conduct an end-to-end security analysis against various attack models. Applying HomeScanon three extensively-used smart home systems, we have found twelve non-trivial security issues, which may lead to unauthorized remote control and credential leakage.
Kulani Mahadewa, Kailong Wang 0001, Guangdong Bai, Ling Shi 0002, Yan Liu 0012, Jin Song Dong 0001, Zhenkai Liang
IEEE Trans. Software Eng.2
2018 HOMESCAN: Scrutinizing Implementations of Smart Home Integrations
abstract
A key feature of the booming smart home is the integration of a wide assortment of technologies, including various standards, proprietary communication protocols and heterogeneous platforms. Due to customization, unsatisfied assumptions and incompatibility in the integration, critical security vulnerabilities are likely to be introduced by the integration. Hence, this work addresses the security problems in smart home systems from an integration perspective, as a complement to numerous studies that focus on the analysis of individual techniques. We propose HOMESCAN, an approach that examines the security of the implementations of smart home systems. It extracts the abstract specification of application-layer protocols and internal behaviors of participants, so that it is able to conduct an end-to-end security analysis against various attack models. Applying HOMESCAN on three extensively-used smart home systems, we have found twelve non-trivial security vulnerabilities, which may lead to unauthorized remote control and credential leakage.
Kulani Mahadewa, Kailong Wang 0001, Guangdong Bai, Ling Shi 0002, Jin Song Dong 0001, Zhenkai Liang
ICECCS2
2018 Analyzing Security and Privacy in Design and Implementation of Web Authentication Protocols
Kailong Wang 0001
ICFEM1
2017 A Framework for Formal Analysis of Privacy on SSO Protocols
Kailong Wang 0001, Guangdong Bai, Naipeng Dong, Jin Song Dong 0001
SecureComm1
2015 Formal Analysis of a Single Sign-On Protocol Implementation for Android
abstract
As the boom of social networking, Single Sign-On (SSO) services developed by major commercial service providers like Facebook, Google and Twitter, have been widely used by web-based service providers as an alternative authentication scheme. Despite rich research has focused on browser-based web applications, little has been conducted on the implementation of SSO on mobile platforms. However, we reveal that due to the fundamental difference of isolation mechanism in mobile OS and applications from the origin-based isolation in browsers, the SSO encounters a novel attack surface and adversarial models. We perform the first formal analysis on the implementation of the most widely used SSO service -- Facebook Login. Our study takes as input the available implementation and dynamic execution traces of Facebook SDK for Android, from which we abstract the implementation-level protocol. The protocol is then modeled in typed Pi-calculus, and automatically checked against the mobile platform specific attack models in a protocol verifier Proverif. Our study has successfully identified a major vulnerability, which allows an attacker to steal authentication credentials from victims and log into their Facebook accounts.
Quanqi Ye, Guangdong Bai, Kailong Wang 0001, Jin Song Dong 0001
ICECCS3