EDBT 2026 Demo / reviewers in the wild / expert
Gelei Deng
dblp:236/9144
· DBLP profile ↗
34ranked-venue papers
7as first author
32since 2021 · last 2026
0000-0002-0046-6674ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 16 · 6 first-author · 15 since 2021Artificial intelligence and machine learning · 7 · 7 since 2021Software engineering, systems software and programming languages · 6 · 6 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Uncovering Logit Suppression Vulnerabilities in LLM Safety Alignment
Yi Liu 0069, Yuekang Li, Ling Shi 0002, Gelei Deng, Shengquan Chen, Kailong Wang 0001 |
ICPR (2) | 5 |
| 2026 | ${\mathsf{KubeSec}} $KubeSec: Automatic Detection of Takeover Risks Introduced by Third-Party Apps in the Kubernetes EcosystemabstractThird-party applications (TPAs) are integral components of managed Kubernetes clusters, but are also frequently exploited in takeover attacks. Recent incidents have demonstrated that TPAs can be weaponized to gain control over clusters. Given their critical role within the Kubernetes ecosystem, it is essential to explore the potential attack surfaces associated with various types of TPAs. To address this, we propose${\sf KubeSec}$, a framework that systematically investigates these risks by analyzing application permission configurations and component code dependencies. This investigation revealed a significant number of insecure RBAC binding patterns, uncovering 562 such patterns and identifying 375 vulnerabilities linked to 134 CVEs. These vulnerabilities impact millions of users, with an average remediation time exceeding 10 months. All findings have been reported to the relevant teams, leading to the assignment of 21 new CVEs by the community. These results highlight substantial security risks associated with TPAs in Kubernetes clusters and emphasize the urgent need for further research to develop more secure cluster management practices. Qiyu Hou, Hao Ren 0001, Xingshu Chen, Gelei Deng, Tianwei Zhang 0004, Guowen Xu, Hongwei Li 0001 |
IEEE Trans. Dependable Secur. Comput. | 5 |
| 2026 | SPOLRE: Semantic Preserving Object Layout Reconstruction for Image Captioning System TestingabstractImage captioning (IC) systems, including Microsoft Azure Cognitive Service, are commonly utilized to convert image content into descriptive natural language. However, inaccuracies in caption generation can lead to serious misinterpretations. Advanced testing techniques such as MetaIC and ROME have been developed to mitigate these issues, yet they encounter notable challenges. First, these strategies demand intensive labor, relying on detailed manual annotations like bounding box data of objects to create test cases. Second, the realism of the generated images is compromised, with MetaIC adding unrelated objects and ROME failing to remove objects effectively. Finally, the capability to generate diversified test suites is restricted. MetaIC is limited to only inserting specific objects to prevent overlap, whereas ROME can generate only \(3^{n}-2^{n}\) variations of test cases from an original seed image containing \( n \) objects. In this study, we present SPOLRE, a novel automated tool designed for semantic preserving object layout reconstruction in image captioning system testing. SPOLRE is based on the insight that modifying the arrangement of objects within an image does not alter its inherent semantics. We utilize four semantic preserving transformation techniques—translation, rotation, mirroring, and scaling—to modify object layouts autonomously, eliminating the need for manual annotation. This approach enables the creation of realistic and varied test suites for IC system testing. Our extensive testing demonstrates that more than 75% of survey respondents find the images produced by SPOLRE more realistic compared to those generated by SOTA methods. Additionally, SPOLRE exhibits outstanding performance in identifying caption errors, detecting 31,544 incorrect captions across seven IC systems with an average precision of 91.62%. This significantly outperforms other methods, which only achieve 85.65% accuracy on average and identify 17,160 incorrect captions. Notably, SPOLRE exposes 6,236 unique issues within Microsoft Azure Cognitive Service, highlighting its effectiveness against one of the most advanced IC systems available. Yi Liu 0069, Guanyu Wang 0005, Gelei Deng, Kailong Wang 0001, Yang Liu 0003, Haoyu Wang 0001 |
ACM Trans. Softw. Eng. Methodol. | 4 |
| 2025 | Controllable Spoofing Attacks on Visual SLAM in Robotic Vehicles
Gelei Deng, Xingshuo Han, Shangwei Guo, Tianwei Zhang 0004 |
ACSAC | 2 |
| 2025 | Oedipus: LLM-enchanced Reasoning CAPTCHA SolverabstractCAPTCHAs have become a ubiquitous tool in safeguarding applications from automated bots. Over time, the arms race between CAPTCHA development and evasion techniques has led to increasingly sophisticated and diverse designs. The latest iteration, reasoning CAPTCHAs, exploits tasks that are intuitively simple for humans but challenging for conventional AI technologies, thereby enhancing security measures. Gelei Deng, Haoran Ou, Yi Liu 0069, Jie Zhang 0073, Tianwei Zhang 0004, Yang Liu 0003 |
CCS | 1 |
| 2025 | TombRaider: Entering the Vault of History to Jailbreak Large Language ModelsabstractWarning: This paper contains content that may involve potentially harmful behaviours, discussed strictly for research purposes.Jailbreak attacks can hinder the safety of Large Language Model (LLM) applications, especially chatbots.Studying jailbreak techniques is an important AI red teaming task for improving the safety of these applications.In this paper, we introduce TOMBRAIDER, a novel jailbreak technique that exploits the ability to store, retrieve, and use historical knowledge of LLMs.TOMBRAIDER employs two agents, the inspector agent to extract relevant historical information and the attacker agent to generate adversarial prompts, enabling effective bypassing of safety filters.We intensively evaluated TOMBRAIDER on six popular models.Experimental results showed that TOMBRAIDER could outperform state-of-the-art jailbreak techniques, achieving nearly 100% attack success rates (ASRs) on bare models and maintaining over 55.4% ASR against defence mechanisms.Our findings highlight critical vulnerabilities in existing LLM safeguards, underscoring the need for more robust safety defences. Junchen Ding, Yi Liu 0069, Gelei Deng, Yuekang Li |
EMNLP | 5 |
| 2025 | When Audio and Text Disagree: Revealing Text Bias in Large Audio-Language ModelsabstractLarge Audio-Language Models (LALMs) are enhanced with audio perception capabilities, enabling them to effectively process and understand multimodal inputs that combine audio and text.However, their performance in handling conflicting information between audio and text modalities remains largely unexamined.This paper introduces MCR-BENCH, the first comprehensive benchmark specifically designed to evaluate how LALMs prioritize information when presented with inconsistent audio-text pairs.Through extensive evaluation across diverse audio understanding tasks, we reveal a concerning phenomenon: when inconsistencies exist between modalities, LALMs display a significant bias toward textual input, frequently disregarding audio evidence.This tendency leads to substantial performance degradation in audio-centric tasks and raises important reliability concerns for real-world applications.We further investigate the influencing factors of text bias, and explore mitigation strategies through supervised finetuning, and analyze model confidence patterns that reveal persistent overconfidence even with contradictory inputs.These findings underscore the need for improved modality balance during training and more sophisticated fusion mechanisms to enhance the robustness when handling conflicting multi-modal inputs 1 . Gelei Deng, Xianglin Yang, Han Qiu 0001, Tianwei Zhang 0004 |
EMNLP | 2 |
| 2025 | Maat: Analyzing and Optimizing Overcharge on Blockchain Storage
Zheyuan He, Zihao Li 0001, Ao Qiao, Jingwei Li 0001, Feng Luo 0009, Gelei Deng, Shuwei Song, Xiaosong Zhang 0001, Ting Chen 0002, Xiapu Luo |
FAST | 7 |
| 2025 | Fine-Grained Verifiers: Preference Modeling as Next-token Prediction in Vision-Language AlignmentabstractThe recent advancements in large language models (LLMs) and pre-trained vision models have accelerated the development of vision-language large models (VLLMs), enhancing the interaction between visual and linguistic modalities. Despite their notable success across various domains, VLLMs face challenges in modality alignment, which can lead to issues like hallucinations and unsafe content generation. Current alignment techniques often rely on coarse feedback and external datasets, limiting scalability and performance. In this paper, we propose FiSAO (Fine-Grained Self-Alignment Optimization), a novel self-alignment method that utilizes the model’s own visual encoder as a fine-grained verifier to improve vision-language alignment without the need for additional data. By leveraging token-level feedback from the vision encoder, FiSAO significantly improves vision-language alignment, even surpassing traditional preference tuning methods that require additional data. Through both theoretical analysis and experimental validation, we demonstrate that FiSAO effectively addresses the misalignment problem in VLLMs, marking the first instance of token-level rewards being applied to such models. Our code is avaliable at \url{https://anonymous.4open.science/r/FISAO-57F0/}. Chenhang Cui, An Zhang 0003, Yiyang Zhou, Zhaorun Chen, Gelei Deng, Huaxiu Yao, Tat-Seng Chua |
ICLR | 5 |
| 2025 | Detecting Perception-Based Attacks using Visual Odometry: Inconsistency Modeling and Checking on Robotic StatesabstractPerception systems in robotic vehicles are crucial for safe and efficient operation, providing key state estimates necessary for planning and control. However, these systems are increasingly vulnerable to perception-based attacks, such as odometry spoofing, position spoofing, obstacle hiding, and object misclassification, which can lead to catastrophic failures. In this paper, we propose a novel approach to detect perception-based attacks by modeling inconsistencies between the physical and estimated states of the robot. Our approach offers a unified methodology for detecting different types of attacks with high accuracy and minimal computational overhead. We validate our method through extensive simulations and real-world scenarios, achieving a 99.5% success rate in detecting attacks, while maintaining a low latency (within 100ms). Gelei Deng, Tianwei Zhang 0004 |
ICRA | 2 |
| 2025 | Source Code Summarization in the Era of Large Language ModelsabstractTo support software developers in understanding and maintaining programs, various automatic (source) code summarization techniques have been proposed to generate a concise natural language summary (i.e., comment) for a given code snippet. Recently, the emergence of large language models (LLMs) has led to a great boost in the performance of coderelated tasks. In this paper, we undertake a systematic and comprehensive study on code summarization in the era of LLMs, which covers multiple aspects involved in the workflow of LLMbased code summarization. Specifically, we begin by examining prevalent automated evaluation methods for assessing the quality of summaries generated by LLMs and find that the results of the GPT-4 evaluation method are most closely aligned with human evaluation. Then, we explore the effectiveness of five prompting techniques (zero-shot, few-shot, chain-of-thought, critique, and expert) in adapting LLMs to code summarization tasks. Contrary to expectations, advanced prompting techniques may not outperform simple zero-shot prompting. Next, we investigate the impact of LLMs' model settings (including top_p and temperature parameters) on the quality of generated summaries. We find the impact of the two parameters on summary quality varies by the base LLM and programming language, but their impacts are similar. Moreover, we canvass LLMs' abilities to summarize code snippets in distinct types of programming languages. The results reveal that LLMs perform suboptimally when summarizing code written in logic programming languages compared to other language types (e.g., procedural and object-oriented programming languages). Finally, we unexpectedly find that CodeLlamaInstruct with 7B parameters can outperform advanced GPT-4 in generating summaries describing code design rationale and asserting code properties. We hope that our findings can provide a comprehensive understanding of code summarization in the era of LLMs. Weisong Sun, Yun Miao, Yuekang Li, Hongyu Zhang 0002, Chunrong Fang, Yi Liu 0069, Gelei Deng, Yang Liu 0003, Zhenyu Chen 0001 |
ICSE | 7 |
| 2025 | Safe + Safe = Unsafe? Exploring How Safe Images Can Be Exploited to Jailbreak Large Vision-Language ModelsabstractRecent advances in Large Vision-Language Models (LVLMs) have showcased strong reasoning abilities across multiple modalities, achieving significant breakthroughs in various real-world applications.
Despite this great success, the safety guardrail of LVLMs may not cover the unforeseen domains introduced by the visual modality.
Existing studies primarily focus on eliciting LVLMs to generate harmful responses via carefully crafted image-based jailbreaks designed to bypass alignment defenses.
In this study, we reveal that a safe image can be exploited to achieve the same jailbreak consequence when combined with additional safe images and prompts.
This stems from two fundamental properties of LVLMs: universal reasoning capabilities and safety snowball effect.
Building on these insights, we propose Safety Snowball Agent (SSA), a novel agent-based framework leveraging agents' autonomous and tool-using abilities to jailbreak LVLMs.
SSA operates through two principal stages: (1) initial response generation, where tools generate or retrieve jailbreak images based on potential harmful intents, and (2) harmful snowballing, where refined subsequent prompts induce progressively harmful outputs.
Our experiments demonstrate that SSA can use nearly any image to induce LVLMs to produce unsafe content, achieving high success jailbreaking rates against the latest LVLMs.
Unlike prior works that exploit alignment flaws, SSA leverages the inherent properties of LVLMs, presenting a profound challenge for enforcing safety in generative multimodal systems. Chenhang Cui, Gelei Deng, An Zhang 0003, Jingnan Zheng, Yicong Li 0004, Lianli Gao, Tianwei Zhang 0004, Tat-Seng Chua |
NeurIPS | 2 |
| 2025 | RSafe: Incentivizing proactive reasoning to build robust and adaptive LLM safeguardsabstractLarge Language Models (LLMs) continue to exhibit vulnerabilities despite deliberate safety alignment efforts, posing significant risks to users and society. To safeguard against the risk of policy-violating content, system-level moderation via external guard models—designed to monitor LLM inputs and outputs and block potentially harmful content—has emerged as a prevalent mitigation strategy. Existing approaches of training guard models rely heavily on extensive human curated datasets and struggle with out-of-distribution threats, such as emerging harmful categories or jailbreak attacks. To address these limitations, we propose RSafe, an adaptive reasoning-based safeguard that conducts guided safety reasoning to provide robust protection within the scope of specified safety policies. RSafe operates
in two stages: (1) guided reasoning, where it analyzes safety risks of input content through policy-guided step-by-step reasoning, and (2) reinforced alignment, where rule-based RL optimizes its reasoning paths to align with accurate safety prediction. This two-stage training paradigm enables RSafe to internalize safety principles to generalize safety protection capability over unseen or adversarial safety violation
scenarios. During inference, RSafe accepts user-specified safety policies to provide enhanced safeguards tailored to specific safety requirements. Experiments demonstrate that RSafe matches state-of-the-art guard models using limited amount of public data in both prompt- and response-level harmfulness detection, while achieving superior out-of-distribution generalization on both emerging harmful category and jailbreak attacks. Furthermore, RSafe provides human-readable explanations for its safety judgments for better interpretability. RSafe offers a robust, adaptive, and interpretable solution for LLM safety moderation, advancing the development of reliable safeguards in dynamic real-world environments. Our code is available at https://anonymous.4open.science/r/RSafe-996D. Jingnan Zheng, Xiangtian Ji, Chenhang Cui, Weixiang Zhao, Gelei Deng, Zhenkai Liang, An Zhang 0003, Tat-Seng Chua |
NeurIPS | 6 |
| 2025 | IllusionCAPTCHA: A CAPTCHA based on Visual IllusionabstractCAPTCHAs have long been essential tools for protecting applications from automated bots. Initially designed as simple questions to distinguish humans from bots, they have become increasingly complex to keep pace with the proliferation of CAPTCHA-cracking techniques employed by malicious actors. However, with the advent of advanced large language models (LLMs), the effectiveness of existing CAPTCHAs is now being undermined. Gelei Deng, Yi Liu 0069, Junchen Ding, Jieshan Chen, Yulei Sui, Yuekang Li |
WWW | 2 |
| 2025 | Mission: Impossible - Image-Based Geolocation with Large Vision Language ModelsabstractIn the age of ubiquitous smartphone use and widespread image sharing on social platforms, geolocation poses a critical privacy concern. Images often carry sensitive spatial and temporal details—such as street signs, architectural styles, or landmarks—that can inadvertently disclose the precise whereabouts of individuals and organizations. Recent advances in large vision-language models (LVLMs) present an emerging threat by enabling users, regardless of technical expertise, to extract location cues from seemingly benign photos. While existing AI-driven geolocation solutions often focus on narrow datasets or specialized contexts, the generalizable performance and privacy implications of zero-shot LVLMs in real-world settings remain critical questions. In this paper, we investigate the geolocation capabilities of state-of-the-art LVLMs. Our findings reveal that while these models demonstrate a non-negligible capability for image-based geolocation even without specialized training, their accuracy in absolute terms is often low, exposing clear limitations in their current state. We then introduce ETHAN, a framework integrating chain-of-thought (CoT) reasoning. Although ETHAN shows improved performance (e.g., 28.7% accuracy at the 1km threshold) and an 85.4% win rate on GeoGuessr, these results primarily highlight the potential trajectory of such technologies rather than their current widespread, high-accuracy applicability. Our study underscores the dual nature of LVLMs in this domain: they uncover an emerging privacy risk due to their inherent, albeit limited, geolocation abilities, yet also demonstrate significant constraints. We conclude by calling for further research into the limitations and risks of LVLM-based geolocation and the development of effective mitigation strategies to protect sensitive location data. Yi Liu 0069, Gelei Deng, Junchen Ding, Yuekang Li, Tianwei Zhang 0004, Weisong Sun, Yaowen Zheng, Jingquan Ge |
Proc. Priv. Enhancing Technol. | 2 |
| 2024 | VisionGuard: Secure and Robust Visual Perception of Autonomous Vehicles in PracticeabstractModern Autonomous Vehicles (AVs) implement the Visual Perception Module (VPM) to perceive their surroundings. This VPM adopts various Deep Neural Network (DNN) models to process the data collected from cameras and LiDAR. Prior studies have shown that these models are vulnerable to physical adversarial examples (PAEs), which pose a critical safety risk to the autonomous driving task. While a few defense methods have been proposed to safeguard AVs, most of them only target a limited set of attack types and specific scenarios, making them impractical for real-world protection. Xingshuo Han, Haozhao Wang, Kangqiao Zhao, Gelei Deng, Yuan Xu 0033, Hangcheng Liu, Han Qiu 0001, Tianwei Zhang 0004 |
CCS | 4 |
| 2024 | GenderCARE: A Comprehensive Framework for Assessing and Reducing Gender Bias in Large Language ModelsabstractLarge language models (LLMs) have exhibited remarkable capa- bilities in natural language generation, but they have also been observed to magnify societal biases, particularly those related to gender. In response to this issue, several benchmarks have been proposed to assess gender bias in LLMs. However, these bench- marks often lack practical flexibility or inadvertently introduce biases. To address these shortcomings, we introduce GenderCARE, a comprehensive framework that encompasses innovative Criteria, bias Assessment, Reduction techniques, and Evaluation metrics for quantifying and mitigating gender bias in LLMs. To begin, we estab- lish pioneering criteria for gender equality benchmarks, spanning dimensions such as inclusivity, diversity, explainability, objectivity, robustness, and realisticity. Guided by these criteria, we construct GenderPair, a novel pair-based benchmark designed to assess gen- der bias in LLMs comprehensively. Our benchmark provides stan- dardized and realistic evaluations, including previously overlooked gender groups such as transgender and non-binary individuals. Fur- thermore, we develop effective debiasing techniques that incorpo- rate counterfactual data augmentation and specialized fine-tuning strategies to reduce gender bias in LLMs without compromising their overall performance. Extensive experiments demonstrate a significant reduction in various gender bias benchmarks, with re- ductions peaking at over 90% and averaging above 35% across 17 different LLMs. Importantly, these reductions come with minimal variability in mainstream language tasks, remaining below 2%. By offering a realistic assessment and tailored reduction of gender biases, we hope that our GenderCARE can represent a significant step towards achieving fairness and equity in LLMs. More details are available at https://github.com/kstanghere/GenderCARE-ccs24. Kunsheng Tang, Wenbo Zhou 0004, Jie Zhang 0073, Aishan Liu, Gelei Deng, Peigui Qi, Weiming Zhang 0001, Tianwei Zhang 0004, Nenghai Yu |
CCS | 5 |
| 2024 | PhyScout: Detecting Sensor Spoofing Attacks via Spatio-temporal ConsistencyabstractExisting defense approaches against sensor spoofing attacks suf- fer from the limitations of limited specific attack types, requiring GPU computation, exhibiting considerable detection latency and struggling with the interpretability of corner cases. We developed PhyScout, a holistic sensor spoofing defense framework to over- come the above limitations. Our framework capitalizes on the ob- servation that human drivers can rapidly and accurately identify spoofing attacks by performing spatio-temporal consistency checks of their environment. We commence by defining the generalized conflicts that different sensor spoofing attacks produce regarding the spatio-temporal consistency. These conflicts are subsequently unified and formalized through a least squares problem approach. This process is modeled using image-based feature point extrac- tion and matching techniques, followed by the design of a risk identification method for each conflict. We evaluate PhyScout across various environments, including simulators, datasets, and real-world scenarios. Compared to existing defense solutions, PhyScout offers rapid identification of sensor at- tacks (within 100ms) with low performance overhead (CPU-based), and conflict visualization. It demonstrates a fresh paradigm in au- tonomous vehicle security and presents new avenues for future research in robust and efficient defense mechanisms against sensor spoofing attacks. More video demos are at our anonymous website https://sites.google.com/view/physcout. Yuan Xu 0033, Gelei Deng, Xingshuo Han, Han Qiu 0001, Tianwei Zhang 0004 |
CCS | 2 |
| 2024 | PonziGuard: Detecting Ponzi Schemes on Ethereum with Contract Runtime Behavior Graph (CRBG)abstractPonzi schemes, a form of scam, have been discovered in Ethereum smart contracts in recent years, causing massive financial losses. Rule-based detection approaches rely on pre-defined rules with limited capabilities and domain knowledge dependency. Additionally, using static information like opcodes and transactions for machine learning models fails to effectively characterize the Ponzi contracts, resulting in poor reliability and interpretability. Ruichao Liang, Jing Chen 0003, Kun He 0008, Yueming Wu 0001, Gelei Deng, Ruiying Du, Cong Wu 0003 |
ICSE | 5 |
| 2024 | Efficient Detection of Toxic Prompts in Large Language ModelsabstractLarge language models (LLMs) like ChatGPT and Gemini have significantly advanced natural language processing, enabling various applications such as chatbots and automated content generation. However, these models can be exploited by malicious individuals who craft toxic prompts to elicit harmful or unethical responses. These individuals often employ jailbreaking techniques to bypass safety mechanisms, highlighting the need for robust toxic prompt detection methods. Existing detection techniques, both blackbox and whitebox, face challenges related to the diversity of toxic prompts, scalability, and computational efficiency. In response, we propose ToxicDetector, a lightweight greybox method designed to efficiently detect toxic prompts in LLMs. ToxicDetector leverages LLMs to create toxic concept prompts, uses embedding vectors to form feature vectors, and employs a Multi-Layer Perceptron (MLP) classifier for prompt classification. Our evaluation on various versions of the LLama models, Gemma-2, and multiple datasets demonstrates that ToxicDetector achieves a high accuracy of 96.39% and a low false positive rate of 2.00%, outperforming state-of-the-art methods. Additionally, ToxicDetector's processing time of 0.0780 seconds per prompt makes it highly suitable for real-time applications. ToxicDetector achieves high accuracy, efficiency, and scalability, making it a practical method for toxic prompt detection in LLMs. Yi Liu 0069, Junzhe Yu, Huijia Sun, Ling Shi 0002, Gelei Deng, Yuqi Chen 0001, Yang Liu 0003 |
ASE | 5 |
| 2024 | MASTERKEY: Automated Jailbreaking of Large Language Model Chatbots
Gelei Deng, Yi Liu 0069, Yuekang Li, Kailong Wang 0001, Ying Zhang 0066, Zefeng Li, Haoyu Wang 0001, Tianwei Zhang 0004, Yang Liu 0003 |
NDSS | 1 |
| 2024 | PentestGPT: Evaluating and Harnessing Large Language Models for Automated Penetration Testing
Gelei Deng, Yi Liu 0069, Victor Mayoral Vilches, Yuekang Li, Yuan Xu 0033, Martin Pinzger 0001, Stefan Rass, Tianwei Zhang 0004, Yang Liu 0003 |
USENIX Security Symposium | 1 |
| 2024 | VerifyML: Obliviously Checking Model Fairness Resilient to Malicious Model HolderabstractIn this paper, we presentVerifyML, the first secure inference framework to check the fairness degree of a given Machine learning (ML) model.VerifyMLis generic and is immune to any obstruction by the malicious model holder during the verification process. We rely on secure two-party computation (2 PC) technology to implementVerifyML, and carefully customize a series of optimization methods to boost its performance for both linear and nonlinear layer execution. Specifically, (1)VerifyMLallows the vast majority of overhead to be performed offline, thus meeting the low latency requirements for online inference. (2) To speed up offline preparation, we first design novel homomorphic parallel computing techniques to accelerate the authenticated Beaver's triple (including matrix- vector and convolution triples) generation procedure. It achieves up to$1.7\times$computation speedup and gains at least$10.7\times$less communication overhead compared to state-of-the-art work. (3) We also present a new cryptographic protocol to evaluate the activation functions of non-linear layers, which is$4\times$–$42\times$faster and has$\gt 48\times$less communication than the existing 2 PC protocol against malicious parties. In fact,VerifyMLeven beats the state-of-the-art semi-honest ML secure inference system! We provide a formal theoretical analysis forVerifyMLsecurity and demonstrate its performance superiority on mainstream ML models including ResNet-18 and LeNet. Guowen Xu, Xingshuo Han, Gelei Deng, Tianwei Zhang 0004, Shengmin Xu, Jianting Ning, Anjia Yang, Hongwei Li 0001 |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2024 | Distributed Motion Control for Multiple Mobile Robots Using Discrete-Event Systems and Model Predictive ControlabstractDistributed motion control is critical in multiple mobile robot systems (MMRSs). Current research usually focuses on either discrete approaches, which aim to deal with high-level collisions and deadlocks without considering the low-level motion commands, or continuous approaches, which can optimize low-level continuous commands to mobile robots but cannot deal with deadlocks efficiently. In this article, by combining discrete and continuous methods, we design a hybrid motion control method for MMRSs where each robot should move along a predefined path. First, each robot’s motion is modeled as a discrete transition system, based on which a real-time supervisory control policy is illustrated to avoid collisions and deadlocks. Second, according to the discrete decisions, the continuous speed at each discrete state is computed using model predictive control and sequential convex programming. The proposed hybrid approach brings two advantages. First, the discrete control component guarantees collision and deadlock avoidance and reduces the scale of the optimization problems. Second, continuous control optimizes the continuous speed in real time and fulfills other performance requirements like time and energy costs. To move in a fully distributed way, each robot needs to predict the motion of its neighbors by retrieving their immediately available information through communications. The simulation and real-world experimental results show the effectiveness of our approach. Yuan Zhou 0005, Hesuan Hu, Gelei Deng, Shangwei Lin 0001, Yang Liu 0003, Zuohua Ding |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2023 | SoK: Rethinking Sensor Spoofing Attacks against Robotic Vehicles from a Systematic ViewabstractRobotic Vehicles (RVs) have gained great popularity over the past few years. Meanwhile, they are also demonstrated to be vulnerable to sensor spoofing attacks. Although a wealth of research works have presented various attacks, some key questions remain unanswered: are these existing works complete enough to cover all the sensor spoofing threats? If not, how many attacks are not explored, and how difficult is it to realize them?This paper answers the above questions by comprehensively systematizing the knowledge of sensor spoofing attacks against RVs. Our contributions are threefold. (1) We identify seven common attack paths in an RV system pipeline. We categorize and assess existing spoofing attacks from the perspectives of spoofer property, operation, victim characteristic and attack goal. Based on this systematization, we identify 4 interesting insights about spoofing attack designs. (2) We propose a novel action flow model to systematically describe robotic function executions and unexplored sensor spoofing threats. With this model, we successfully discover 103 spoofing attack vectors, 26 of which have been verified by prior works, while 77 attacks are never considered. (3) We design two novel attack methodologies to verify the feasibility of newly discovered spoofing attack vectors. Yuan Xu 0033, Xingshuo Han, Gelei Deng, Jiwei Li 0001, Yang Liu 0003, Tianwei Zhang 0004 |
EuroS&P | 3 |
| 2023 | ASTER: Automatic Speech Recognition System Accessibility Testing for StutterersabstractThe popularity of automatic speech recognition (ASR) systems nowadays leads to an increasing need for improving their accessibility. Handling stuttering speech is an important feature for accessible ASR systems. To improve the accessibility of ASR systems for stutterers, we need to expose and analyze the failures of ASR systems on stuttering speech. The speech datasets recorded from stutterers are not diverse enough to expose most of the failures. Furthermore, these datasets lack ground truth information about the non-stuttered text, rendering them unsuitable as comprehensive test suites. Therefore, a methodology for generating stuttering speech as test inputs to test and analyze the performance of ASR systems is needed. However, generating valid test inputs in this scenario is challenging. The reason is that although the generated test inputs should mimic how stutterers speak, they should also be diverse enough to trigger more failures. To address the challenge, we propose Aster, a technique for automatically testing the accessibility of ASR systems. Aster can generate valid test cases by injecting five different types of stuttering. The generated test cases can both simulate realistic stuttering speech and expose failures in ASR systems. Moreover, Aster can further enhance the quality of the test cases with a multi-objective optimization-based seed updating algorithm. We implemented Aster as a framework and evaluated it on four open-source ASR models and three commercial ASR systems. We conduct a comprehensive evaluation of Aster and find that it significantly increases the word error rate, match error rate, and word information loss in the evaluated ASR systems. Additionally, our user study demonstrates that the generated stuttering audio is indistinguishable from real-world stuttering audio clips. Yi Liu 0069, Yuekang Li, Gelei Deng, Felix Juefei-Xu, Yao Du 0002, Cen Zhang, Yeting Li, Lei Ma 0003, Yang Liu 0003 |
ASE | 3 |
| 2023 | NAUTILUS: Automated RESTful API Vulnerability Detection
Gelei Deng, Zhiyi Zhang 0005, Yuekang Li, Yi Liu 0069, Tianwei Zhang 0004, Yang Liu 0003, Dongjin Wang |
USENIX Security Symposium | 1 |
| 2023 | The Threat of Offensive AI to OrganizationsabstractAI has provided us with the ability to automate tasks, extract information from vast amounts of data, and synthesize media that is nearly indistinguishable from the real thing. However, positive tools can also be used for negative purposes. In particular, cyber adversaries can use AI to enhance their attacks and expand their campaigns. Although offensive AI has been discussed in the past, there is a need to analyze and understand the threat in the context of organizations. For example, how does an AI-capable adversary impact the cyber kill chain? Does AI benefit the attacker more than the defender? What are the most significant AI threats facing organizations today and what will be their impact on the future? In this study, we explore the threat of offensive AI on organizations. First, we present the background and discuss how AI changes the adversary’s methods, strategies, goals, and overall attack model. Then, through a literature review, we identify 32 offensive AI capabilities which adversaries can use to enhance their attacks. Finally, through a panel survey spanning industry, government and academia, we rank the AI threats and provide insights on the adversaries. Yisroel Mirsky, Ambra Demontis, Jaidip Kotak, Ram Shankar, Gelei Deng, Liu Yang 0003, Maura Pintor, Wenke Lee, Yuval Elovici, Battista Biggio |
Comput. Secur. | 5 |
| 2022 | On the (In)Security of Secure ROS2abstractRobot Operating System (ROS) has been the mainstream platform for research and development of robotic applications. This platform is well-known for lacking security features and efficiency for distributed robotic computations. To address these issues, ROS2 is recently developed by utilizing the Data Distribution Service (DDS) to provide security support. Integrated with DDS, ROS2 is expected to establish the basis for trustworthy robotic ecosystems. Gelei Deng, Guowen Xu, Yuan Zhou 0005, Tianwei Zhang 0004, Yang Liu 0003 |
CCS | 1 |
| 2022 | Morest: Model-based RESTful API Testing with Execution FeedbackabstractRESTful APIs are arguably the most popular endpoints for accessing Web services. Blackbox testing is one of the emerging techniques for ensuring the reliability of RESTful APIs. The major challenge in testing RESTful APIs is the need for correct sequences of API operation calls for in-depth testing. To build meaningful operation call sequences, researchers have proposed techniques to learn and utilize the API dependencies based on OpenAPI specifications. However, these techniques either lack the overall awareness of how all the APIs are connected or the flexibility of adaptively fixing the learned knowledge. Yi Liu 0069, Yuekang Li, Gelei Deng, Yang Liu 0003, Ruiyuan Wan, Runchao Wu, Dandan Ji, Shiheng Xu, Minli Bao |
ICSE | 3 |
| 2021 | An Investigation of Byzantine Threats in Multi-Robot SystemsabstractMulti-Robot Systems (MRSs) show significant advantages to deal with complex tasks efficiently. However, the system complexity inevitably enlarges the attack surface and adds difficulty in guaranteeing the security and safety of MRSs. In this paper, we present an in-depth investigation about the Byzantine threats in MRSs, where some robot is untrusted. We design a practical methodology to identify potential Byzantine risks in a given MRS workload built from the Robot Operating System (ROS). It consists of three novel steps (requirement specification using signal temporal logic, attack surface determination via data-flow analysis, attack identification using requirement-driven fuzzing) to thoroughly assess MRS workloads. We use this fuzzing method to inspect five typical MRS workloads from past works and the ROS platform, and identify three novel kinds of attacks that can be launched with five attack strategies. We conduct comprehensive experiments in the Gazebo simulator and a real-world MRS with three TurtlBot3 robots to validate these attacks, which can remarkably decrease the system’s performance, or even cause task failures. Gelei Deng, Yuan Zhou 0005, Yuan Xu 0033, Tianwei Zhang 0004, Yang Liu 0003 |
RAID | 1 |
| 2021 | Novel denial-of-service attacks against cloud-based multi-robot systems
Yuan Xu 0034, Gelei Deng, Tianwei Zhang 0004, Han Qiu 0001, Yungang Bao |
Inf. Sci. | 2 |
| 2019 | Efficient Password Guessing Based on a Password Segmentation ApproachabstractMost cracking tools against alphanumeric passwords conduct password guessing based on sophisticatedly constructed password dictionaries. The rule-based methods which continuously expand the size of dictionaries based on simple permutation and concatenation is the traditional way to construct password dictionaries. To increase the intelligence in dictionary generation, some password cracking tools extract password patterns from the training passwords based on machine learning, and thus construct dictionaries using the extracted patterns. However, these tools either have low guessing efficiency, or produce password generation models with low interpretability. Usually, a password could be split into several meaningful segments each of which represents particular personal information or a grammatically correct word, and the password patterns could be extracted from these segments. In this paper, we propose a novel password cracking tool, which breaks each training password to meaningful segments, learns the patterns from the password segments, and generates personalized high-efficiency password dictionaries based on the learned patterns. The experimental results show that the proposed tool is more efficient than the traditional rule-based tools as well as alphanumeric patterns-based tools. Furthermore, to evaluate the impact of personal information leakage on password security, we use personal information of the target users as the inputs for the proposed tool and analyze the password guessing efficiency. Gelei Deng, Xingjie Yu, Huaqun Guo |
GLOBECOM | 1 |
| 2019 | A fog computing based approach to DDoS mitigation in IIoT systems
Luying Zhou, Huaqun Guo, Gelei Deng |
Comput. Secur. | 3 |