VLDB 2026 Research / reviewers in the wild / expert
Yang Liu 0003
dblp:51/3710-3
· DBLP profile ↗
609ranked-venue papers
12as first author
343since 2021 · last 2026
0000-0001-7300-9215ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 336 · 11 first-author · 160 since 2021Artificial intelligence and machine learning · 100 · 68 since 2021Security and privacy · 87 · 1 first-author · 66 since 2021Graphics, computer vision, multimedia, augmented reality and games · 54 · 38 since 2021Theory of computation · 27 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 26 · 13 since 2021Systems, architecture and hardware · 21 · 8 since 2021Databases, data management, data science and information retrieval · 12 · 10 since 2021Computer networks · 10 · 8 since 2021Human-computer interaction and ubiquitous computing · 7 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PhysPatch: A Physically Realizable and Transferable Adversarial Patch Attack for Multimodal Large Language Models-based Autonomous Driving SystemsabstractMultimodal Large Language Models (MLLMs) are becoming integral to autonomous driving (AD) systems due to their strong vision-language reasoning capabilities. However, MLLMs are vulnerable to adversarial attacks—particularly adversarial patch attacks—which can pose serious threats in real-world scenarios. Existing patch-based attack methods are primarily designed for object detection models. Due to the more complex architectures and strong reasoning capabilities of MLLMs, these approaches perform poorly when transferred to MLLM-based systems. To address these limitations, we propose PhysPatch, a physically realizable and transferable adversarial patch framework tailored for MLLM-based AD systems. PhysPatch jointly optimizes patch location, shape, and content to enhance attack effectiveness and real-world applicability. It introduces a semantic-based mask initialization strategy for realistic placement, an SVD-based local alignment loss with patch-guided crop-resize to improve transferability, and a potential field-based mask refinement method. Extensive experiments across open-source, commercial, and reasoning-capable MLLMs demonstrate that PhysPatch significantly outperforms state-of-the-art (SOTA) methods in steering MLLM-based AD systems toward target-aligned perception and planning outputs. Moreover, PhysPatch consistently places adversarial patches in physically feasible regions of AD scenes, ensuring strong real-world applicability and deployability. Qi Guo 0008, Xiaojun Jia, Shanmin Pang, Simeng Qin, Lin Wang 0026, Ju Jia, Yang Liu 0003, Qing Guo 0005 |
AAAI | 7 |
| 2026 | Hidden in the Noise: Unveiling Backdoors in Audio LLMs Alignment Through Latent Acoustic Pattern TriggersabstractAs Audio Large Language Models (ALLMs) emerge as powerful tools for speech processing, their safety implications demand urgent attention. While considerable research has explored textual and vision safety, audio’s distinct characteristics present significant challenges. This paper first investigates: Is ALLM vulnerable to backdoor attacks exploiting acoustic triggers? In response to this issue, we introduce Hidden in the Noise (HIN), a novel backdoor attack framework designed to exploit subtle, audio-specific features. HIN applies acoustic modifications to raw audio waveforms, such as alterations to temporal dynamics and strategic injection of spectrally tailored noise. These changes introduce consistent patterns that an ALLM’s acoustic feature encoder captures, embedding robust triggers within the audio stream. To evaluate ALLM robustness against audio-feature-based triggers, we develop the AudioSafe benchmark, assessing nine distinct risk types. Extensive experiments on AudioSafe and three established safety datasets reveal critical vulnerabilities in existing ALLMs: (I) audio features like environment noise and speech rate variations achieve over 90% average attack success rate, (II) ALLMs exhibit significant sensitivity differences across acoustic features, particularly showing minimal response to volume as a trigger, and (III) poisoned sample inclusion causes only marginal loss curve fluctuations, highlighting the attack’s stealth. Liang Lin 0004, Kaiwen Luo, Lilan Peng, Dexian Wang 0001, Xuehai Tang, Yuanhe Zhang, Xikang Yang, Zhenhong Zhou, Kun Wang 0056, Yang Liu 0003 |
AAAI | 12 |
| 2026 | MAGIC: Mastering Physical Adversarial Generation in Context Through Collaborative LLM AgentsabstractPhysical adversarial attacks in driving scenarios can expose critical vulnerabilities in visual perception models. However, developing such attacks remains non-trivial due to diverse real-world environmental influences. Existing approaches either struggle to generalize to dynamic environments or fail to achieve consistent physical attack performance. To address these challenges, we propose MAGIC (Mastering Physical Adversarial Generation In Context), a novel framework powered by multi-modal LLM agents to automatically understand the scene context during testing time and generate adversarial patches through synergistic interaction of language and vision understanding. Specifically, MAGIC orchestrates three specialized LLM agents: the adv-patch generation agent masters the creation of deceptive patches via strategic prompt manipulation for text-to-image models; the adv-patch deployment agent ensures contextual coherence by determining optimal deployment strategies based on scene understanding; and the self-examination agent completes this trilogy by providing critical oversight and iterative refinement of both processes. We validate our approach with both digital and physical scenarios, i.e., nuImage and real-world scenes, where both statistical and visual results demonstrate that our MAGIC is powerful and effective for attacking widely applied object detection systems, such as YOLO and DETR series. Yun Xing 0001, Nhat Chung, Jie Zhang 0002, Ivor W. Tsang, Yang Liu 0003, Lei Ma 0003, Qing Guo 0005 |
AAAI | 6 |
| 2026 | OptiCo: Adaptive Distributed Training Optimization via Collaborative Agent ReasoningabstractSheng Chen, Tang Zhe, Weixing Zhang, Fei Yang, Yuanyuan. Wang, Tianlin Li, Yang Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Tang Zhe, Fei Yang 0007, Tianlin Li, Yang Liu 0003 |
ACL (1) | 7 |
| 2026 | False Friends in the Shell: Unveiling the Emoticon Semantic Confusion in Large Language ModelsabstractEmoticons are widely used in digital communication to convey affective intent, yet their safety implications for Large Language Models (LLMs) remain largely unexplored.In this paper, we identify emoticon semantic confusion, a vulnerability where LLMs misinterpret ASCII-based emoticons to perform unintended and even destructive actions.To systematically study this phenomenon, we develop an automated data generation pipeline and construct a dataset containing 3,757 code-oriented test cases spanning 21 meta-scenarios, four programming languages, and varying contextual complexities.Our study on six LLMs reveals that emoticon semantic confusion is pervasive, with an average confusion ratio exceeding 38%.More critically, over 90% of confused responses yield 'silent failures', which are syntactically valid outputs but deviate from user intent, potentially leading to destructive security consequences.Furthermore, we observe that this vulnerability readily transfers to popular agent frameworks, while existing prompt-based mitigations remain largely ineffective.We call on the community to recognize this emerging vulnerability and develop effective mitigation methods to uphold the safety and reliability of human-LLM interactions.* These authors contributed equally. Xiaoyu Zhang 0013, Juan Zhai, Shiqing Ma, Chao Shen 0001, Yang Liu 0003 |
ACL (1) | 6 |
| 2026 | User-Space Dependency-Aware Rehosting for Linux-Based Firmware Binaries
Cen Zhang, Yaowen Zheng, Puzhuo Liu, Jian Zhang 0087, Yeting Li, Yang Liu 0003, Limin Sun 0001 |
NDSS | 8 |
| 2026 | Scratching the Iceberg: Unveiling the Outdated Third-Party Native Libraries in Android Apps
Shiyang Zhang, Sen Chen 0001, Lyuye Zhang, Yang Liu 0003 |
SANER | 5 |
| 2026 | CAVERN: Efficient Honest-Majority Maliciously Secure (2+1)-PC for $\mathbb{Z}_{2^{n}}$ via DPF
Yang Liu 0003, Liang Feng Zhang |
SP | 1 |
| 2026 | Bridging Expert Reasoning and LLM Detection: A Knowledge-Driven Framework for Malicious PackagesabstractOpen-source ecosystems such as NPM and PyPI are increasingly targeted by supply chain attacks, yet existing detection methods either depend on fragile handcrafted rules or data-driven features that fail to capture evolving attack semantics. We present IntelGuard, a retrieval-augmented generation (RAG) based framework that integrates expert analytical reasoning into automated malicious package detection. IntelGuard constructs a structured knowledge base from over 8,000 threat intelligence reports, linking malicious code snippets with behavioral descriptions and expert reasoning. When analyzing new packages, it retrieves semantically similar malicious examples and applies LLM-guided reasoning to assess whether code behaviors align with intended functionality. Experiments on 4,027 real-world packages show that IntelGuard achieves 99% accuracy and a 0.50% false positive rate, while maintaining 96.5% accuracy on obfuscated code. Deployed on PyPI.org, it discovered 54 previously unreported malicious packages, demonstrating interpretable and robust detection guided by expert knowledge. Wenbo Guo 0011, Shiwen Song, Jiaxun Guo, Zhengzi Xu, Haoran Ou, Mengmeng Ge 0003, Yang Liu 0003 |
WWW | 8 |
| 2026 | Resisting Manipulative Bots in Meme Coin Copy Trading: A Multi-Agent Approach with Chain-of-Thought ReasoningabstractCopy trading has become the dominant entry strategy in meme coin markets. However, due to the market's extremely illiquid and volatile nature, the strategy exposes an exploitable attack surface: adversaries deploy manipulative bots to front-run trades, conceal positions, and fabricate sentiment, systematically extracting value from naïve copiers at scale. Despite its prevalence, bot-driven manipulation remains largely unexplored, and no robust defensive framework exists. We propose a manipulation-resistant copy-trading system based on a multi-agent architecture powered by a multi-modal large language model (LLM) and chain-of-thought (CoT) reasoning. Our approach outperforms zero-shot and most statistic-driven baselines in prediction accuracy as well as all baselines in economic performance, achieving an average copier return of 3% per meme coin investment under realistic market frictions. Overall, our results demonstrate the effectiveness of agent-based defenses and predictability of trader profitability in adversarial meme coin markets, providing a practical foundation for robust copy trading. Yebo Feng, Jiahua Xu 0002, Yang Liu 0003 |
WWW | 4 |
| 2026 | Fake news detection with GAN-augmented contrastive learning and multimodal attentionabstractAbstract The rapid proliferation of fake news in digital media has emerged as a major threat to information credibility and public trust. Although recent advances have explored multimodal learning for fake news detection, existing models often fail to effectively integrate heterogeneous data sources and remain vulnerable to adversarial manipulations. To address these challenges, we propose (Multimodal Adversarial Deep Semantic Learning), a robust multimodal fake news detection framework that unifies generative adversarial networks (GANs) with supervised contrastive learning. Specifically, employs a multi-layer joint attention mechanism to align and fuse textual and visual features, while adversarial training encourages the extraction of event-invariant representations, enhancing generalizability across unseen news events. Additionally, contrastive learning with adversarial perturbations further strengthens feature discrimination and robustness against attacks. Extensive experiments on benchmark Twitter and Weibo datasets demonstrate that achieves state-of-the-art accuracy (85.3%) and maintains stable performance with only a 1.1% drop under adversarial conditions, outperforming existing methods in both detection accuracy and resilience. These results underscore ’s effectiveness in advancing robust multimodal fake news detection and promoting digital information integrity. Cong Wu 0003, Jing Chen 0003, Yebo Feng, Ju Jia, Zijian Zhang 0001, Jiahua Xu 0002, Teng Li 0003, Yang Liu 0003 |
Cybersecur. | 9 |
| 2026 | Peer-aided repairer: empowering large language models to repair advanced student assignments
Qianhui Zhao, Li Zhang 0029, Fang Liu 0032, Yang Liu 0003, Jing Jiang 0005, Ge Li 0001, Zian Sun, Zhong-Qi Li, Yuchi Ma |
Empir. Softw. Eng. | 4 |
| 2026 | Fuzzy-DDPG: Integrating fuzzy logic with continuous deep reinforcement learning for mobile robot motion planning
Fenghua Wu, Wenbing Tang 0001, Yuan Zhou 0005, Hesuan Hu, Yang Liu 0003, Zuohua Ding |
Fuzzy Sets Syst. | 6 |
| 2026 | Federated Learning-Driven Covert Communication in Satellite-Terrestrial Integrated Networks: A Privacy-Preserving FrameworkabstractDue to the broadcasting characteristics of satellite-terrestrial integrated networks (STINs), security vulnerabilities have emerged as a critical concern requiring urgent mitigation strategies. Unlike traditional security methods, federated learning (FL) enables a large number of participants to collaborate without disclosing actual privacy data. Its potential as a framework that combines collaborative model training and covert payload transmission in STINs represents a significant research gap. This paper proposes FedSAT, a novel FL-based covert communication scheme for STINs, in which each participant in the FL process can utilize the shared learning protocol as a covert medium for transmitting arbitrary information in privacy-preserving framework. Our framework leverages the dual capabilities of FL for collaborative model training and covert payload embedding, utilizing Geostationary Earth Orbit (GEO) satellites and distributed terrestrial nodes to embed sensitive data within FL parameter updates. The system maintains model convergence accuracy while implementing strategic encryption to achieve robust sharing and transmission of payloads within the FL framework. Comprehensive simulation tests demonstrate the framework significant efficacy, achieving a 98.7% communication coverage for covert payload transmission under monitoring by low Earth orbit (LEO) surveillance satellites, with only a 0.8% decrease in model accuracy. This breakthrough achievement paves the way for a transformative paradigm in covert cross-domain communication for next-generation networks. Min Wu 0008, Kefeng Guo, Chao Dong 0001, Yang Liu 0003, Qihui Wu 0001, Zhiming Zheng 0001 |
IEEE J. Sel. Areas Commun. | 5 |
| 2026 | Adversarial rain attack and defensive deraining for DNN perception
Liming Zhai, Qing Guo 0003, Felix Juefei-Xu, Xiaofei Xie, Lei Ma 0003, Wei Feng 0005, Shengchao Qin, Yang Liu 0003 |
Neural Networks | 8 |
| 2026 | Reframing Paths as Logic: Semantic Segmentation for Vulnerability DetectionabstractPath-sensitive vulnerabilities, such as use-after-free, integer overflows, and command injection, pose significant challenges for traditional static analysis tools, which often face trade-offs between precision, scalability, and interpretability. To address these challenges, we present SEVDF (Semantic-Enhanced Vulnerability Detection Framework), a novel methodology that integrates may-analysis taint propagation with large language models (LLMs) to detect path-related vulnerabilities in large C/C++ codebases. SEVDF begins by constructing a program dependency graph and performing a sound but incomplete taint analysis to extract all potential vulnerable paths. After segmentation, deduplication, feasibility check, and semantic summarization by LLMs, the vulnerable paths are reformed and confirmed with LLMs for their inter-procedural feasibility and semantic consistency. We evaluate SEVDF on the Juliet Test Suite (thirteen CWE categories) and a curated real-world dataset of 71 vulnerabilities across 9 projects. SEVDF consistently outperforms the default CodeQL rules, CodeQL rules with all unnecessary constraints removed, and three open-source detectors, which are Infer, Cppcheck and CodeChecker. SEVDF is able to achieve 100% precision on several CWEs while maintaining or improving recall on Juliet benchmark. Moreover, our segment-based design reduces the analysis workload for LLMs by 90.6% compared to direct-path prompting through Logic Unit deduplication, making SEVDF cost-effective for large-scale deployment. Finally, SEVDF uncovered and reported 29 0-day vulnerabilities (12 confirmed to date), including 3 CVEs in VirtualBox, demonstrating practical value. Zong Cao, Yuqiang Sun 0001, Zhengzi Xu, Kaixuan Li 0002, Yeqi Fu, Ziqiao Kong, Yang Liu 0003 |
Proc. ACM Program. Lang. | 8 |
| 2026 | AudioJailbreak: Jailbreak Attacks Against End-to-End Large Audio-Language ModelsabstractJailbreak attacks to Large audio-language models (LALMs) are studied recently, but they exclusively focused on the attack scenario where the adversary can fully manipulate user prompts (named strong adversary) and limited in effectiveness, applicability, and practicability. In this work, we first conduct an extensive evaluation showing that advanced text jailbreak attacks cannot be easily ported to end-to-end LALMs via text-to-speech (TTS) techniques. We then propose AUDIOJAILBREAK, a novel audio jailbreak attack, featuring (1) asynchrony: the jailbreak audios do not need to align with user prompts in the time axis by crafting suffixal jailbreak audios; (2) universality: a single jailbreak perturbation is effective for different prompts by incorporating multiple prompts into the perturbation generation; (3) stealthiness: the malicious intent of jailbreak audios is concealed by proposing various intent concealment strategies; and (4) over-the-air robustness: the jailbreak audios remain effective when being played over the air by incorporating reverberation into the perturbation generation. In contrast, all prior audio jailbreak attacks cannot offer asynchrony, universality, stealthiness, and/or over-the-air robustness. Moreover, AUDIOJAILBREAK is also applicable to a more practical and broader attack scenario where the adversary cannot fully manipulate user prompts (named weak adversary). Extensive experiments with thus far the most LALMs demonstrate the high effectiveness of AUDIOJAILBREAK, in particular, it can jailbreak openAI's GPT-4o-Audio and bypass Meta's Llama-Guard-3 safeguard, in the weak adversary scenario. We highlight that our work peeks into the security implications of audio jailbreak attacks against LALMs, and realistically fosters improving their robustness, especially for the newly proposed weak adversary. Guangke Chen, Fu Song, Zhe Zhao 0007, Xiaojun Jia, Yang Liu 0003, Yanchen Qiao, Weizhe Zhang, Weiping Tu, Yuhong Yang 0001, Bo Du 0001 |
IEEE Trans. Dependable Secur. Comput. | 5 |
| 2026 | Catching Scam Tokens With Temporal Graph Learning in Decentralized FinanceabstractDecentralized finance has experienced phenomenal growth, revolutionizing the landscape of financial transactions and asset management via blockchain. Yet, this swift growth brings with it substantial challenges, notably the surge in scam tokens, imposing significant security threats on cryptocurrency investments and trading. Existing detection methods of scam token, primarily relying on analyzing contract codes or transaction patterns, struggle to catch increasingly sophisticated tactics employed by scammers. For example, contract-based analysis are unable to identify scams lacking overt malicious code, e.g., most rugpulls, while transaction-based methods generally lack the foresight to early-detect potential risks. In this paper, we present TOKENSCOUT, the first temporal GNN-based framework for scam token early detection. TOKEN SCOUT formulates token transfer data as a dynamic temporal attributed multigraph and leverages the temporal graph learning model to learn graph representations. It also builds a graph rep resentation refining model based on contrastive learning to learn a more discriminative representation space for risk identification. We evaluated TOKENSCOUT using a comprehensive dataset of 214,084 standard ERC20 tokens from 2015 to February 2023. TOKENSCOUT achieves a balanced accuracy of 98.41%. Additionally, from March to May 2023, deploying TOKENSCOUT on Ethereum effectively identified 706 rugpulls, 174 honeypots, and 90 Ponzi schemes, thereby alerting to potential risks exceeding $240 million. Cong Wu 0003, Jing Chen 0003, Jian Shen 0001, Guowen Xu, Yueming Wu 0001, Haijun Wang 0002, Hongwei Li 0001, Yang Liu 0003, Yang Xiang 0001 |
IEEE Trans. Dependable Secur. Comput. | 8 |
| 2026 | Hidden Tail: Adversarial Attack for Stealthy Resource Consumption Against Vision-Language ModelsabstractVision-Language Models (VLMs) are increasingly deployed in real-world applications, but their high inference cost makes them vulnerable to resource consumption attacks. Prior attacks attempt to extend VLM output sequences by optimizing adversarial images, thereby increasing inference costs. However, these extended outputs often introduce irrelevant abnormal content, compromising attack stealthiness. This trade-off between effectiveness and stealthiness poses a major limitation for existing attacks. To address this challenge, we proposeHidden Tail, a stealthy resource consumption attack that crafts prompt-agnostic adversarial images, inducing VLMs to generate maximum-length outputs by appending special tokens invisible to users. Our method employs a composite loss function that balances semantic preservation, repetitive special token induction, and suppression of the end-of-sequence (EOS) token, optimized via a dynamic weighting strategy. Extensive experiments show thatHidden Tailoutperforms existing attacks, increasing output length by up to 19.2× and reaching the maximum token limit, while preserving attack stealthiness. These results highlight the urgent need to improve the robustness of VLMs against efficiency-oriented adversarial threats. Our code is available athttps://github.com/zhangrui4041/Hidden_Tail. Rui Zhang 0086, Tianli Yang, Wenbo Jiang 0001, Rui Zhang 0090, Qingchuan Zhao, Hongwei Li 0001, Yang Liu 0003, Guowen Xu |
IEEE Trans. Dependable Secur. Comput. | 8 |
| 2026 | FOOLSDEDIT: Deceptively Steering Your Edits Towards Targeted Attribute-Aware DistributionabstractGuided image synthesis methods, like SDEdit based on the diffusion model, excel at creating realistic images from user inputs such as stroke paintings. However, existing efforts mainly focus on image quality, often overlooking a key point: the diffusion model represents a data distribution, not individual images. This introduces a low but critical chance of generating images that contradict user intentions, raising ethical concerns. For example, a user inputting a stroke painting with female characteristics might, with some probability, get male faces from SDEdit. To expose this potential vulnerability, we propose the Targeted Attribute Generative Attack (TAGA), whose objective is to force SDEdit to generate data distributions aligned with a specified attribute (i.e.,targeted attribute like male), without changing the attribute of the input image. Empirical studies reveal that traditional adversarial noise struggles to achieve TAGA, while natural perturbations such as exposure and motion blur can easily influence attributes of the generated images. Inspired by the observation, we design attack methodFOOLSDEDITto achieve effective TAGA against SDEdit. It aims to search for an optimized strategy to execute attacks within a weighted graph-based attack architecture, which is formulated to model diverse strategies derived from both exposure and motion blur perturbations. Comprehensive experiments on two commonly used datasets and three social attributes present thatFOOLSDEDITforces SDEdit to generate targeted attribute-aware distributions, achieving significantly more effective TAGA than the baselines. We also validated empirically thatFOOLSDEDITcould induce bias in downstream tasks of SDEdit which rely on the generated data under attack. Our work reveals critical vulnerabilities in diffusion-based image generation models and paves the way for future research on model auditing and bias mitigation. Qi Zhou 0012, Dongxia Wang 0002, Tianlin Li, Yang Liu 0003, Kui Ren 0001, Wenhai Wang, Qing Guo 0005 |
IEEE Trans. Dependable Secur. Comput. | 5 |
| 2026 | One Trigger, Multiple Victims: Clean-Label Neighborhood Backdoor Attacks on Graph Neural NetworksabstractGraph Neural Networks (GNNs) have achieved remarkable success in modeling structured data. Recent studies, however, reveal that they are highly vulnerable to backdoor attacks, which can implant triggers into training data to mislead predictions on nodes injected with triggers while maintaining accuracy on clean inputs. Despite recent advances, existing graph backdoor attacks often rely on explicit training interventions and substantial trigger injection while focusing solely on single-node misclassification, which limits their practicality in real-world deployments. To address these limitations, we propose a clean-label graph backdoor attack that induces one-hop neighborhood misclassification under a minimal trigger injection budget. Without altering target nodes’ features or labels, our method attaches a single trigger node to a target node, thereby misclassifying both the target and its immediate neighbors as the target class. To maximize effectiveness while preserving stealthiness, we propose a poisoned node selection strategy guided by semantic consistency and structural activeness, and design a conditional diffusion-based trigger generator optimized with multiple auxiliary objectives. Extensive experiments on multiple real-world benchmarks and mainstream GNN architectures show that our approach achieves over 95% attack success rate on both target nodes and their neighbors in most settings, including under state-of-the-art defenses. These findings underscore the urgent need for more robust graph learning systems and reveal novel attack surfaces in graph security. Huaxin Deng, Yong Fang 0002, Qiang Zhang 0057, Yang Liu 0003, Yijia Xu |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2026 | HKT-SmartAudit: Distilling Lightweight Models for Smart Contract AuditingabstractThe rapid growth of blockchain technology has driven the widespread adoption of smart contracts; however, their inherent vulnerabilities have led to significant financial losses. Traditional auditing methods, while essential, struggle to keep pace with the increasing complexity and scale of smart contracts. Large language models (LLMs) offer promising capabilities for automating vulnerability detection, but their adoption is often limited by high computational costs. Although prior work has explored leveraging large models through agents or workflows, relatively little attention has been given to improving the performance of smaller, fine-tuned models—a critical factor for achieving both efficiency and data privacy. In this paper, we introduce HKT-SmartAudit, a framework for developing lightweight models optimized for smart contract auditing. It features a multi-stage knowledge distillation pipeline that integrates classical distillation, external domain knowledge, and reward-guided learning to transfer high-quality insights from large teacher models. A single-task learning strategy is employed to train compact student models that maintain high accuracy and robustness while significantly reducing computational overhead. Experimental results show that our distilled models outperform both commercial tools and larger models in detecting complex vulnerabilities and logical flaws, offering a practical, secure, and scalable solution for smart contract auditing. The source code is available in the GitHub repository1. Jing Sun 0002, Zijian Zhang 0001, Xianhao Zhang, Meng Li 0006, Yuqiang Sun 0001, Daoyuan Wu, Yang Liu 0003, Chunmiao Li, Mingchao Wan, Jin Dong 0004 |
IEEE Trans. Inf. Forensics Secur. | 9 |
| 2026 | PromptGuard: Soft Prompt-Guided Unsafe Content Moderation for Text-to-Image ModelsabstractRecent text-to-image (T2I) models have exhibited remarkable performance in generating high-quality images from text descriptions. However, these models are vulnerable to misuse, particularly generating not-safe-for-work (NSFW) content, such as sexually explicit, violent, political, and disturbing images, raising serious ethical concerns. In this work, we present PromptGuard, a novel content moderation technique that draws inspiration from the system prompt mechanism in large language models (LLMs) for safety alignment. Unlike LLMs, T2I models lack a direct interface for enforcing behavioral guidelines. Our key idea is to optimize a safety soft prompt that functions as an implicit system prompt within the T2I model’s textual embedding space. This universal soft prompt (P∗) directly moderates NSFW inputs, enabling safe yet realistic image generation without altering the inference efficiency or requiring proxy models.We further enhance its reliability and helpfulness through a divide-and-conquer strategy, which optimizes category-specific soft prompts and combines them into holistic safety guidance. Extensive experiments across five datasets demonstrate that PromptGuard effectively mitigates NSFW content generation while preserving high-quality benign outputs. PromptGuard achieves 3.8 times faster than prior content moderation methods, surpassing eight state-of-the-art defenses. Rigorous evaluation using both multi-head classifiers and VLM-based guardrails confirms its robustness, achieving an optimal average unsafe ratios down to 5.84% and 6.18%, respectively. Our code and dataset are available at https://t2ipromptguard. github.io/. Lingzhi Yuan, Xinfeng Li, Chejian Xu, Guanhong Tao 0001, Xiaojun Jia, Yihao Huang 0001, Wei Dong 0007, Yang Liu 0003, Bo Li 0026 |
IEEE Trans. Inf. Forensics Secur. | 8 |
| 2026 | AeroGuard: Towards Real-Time UAV Fault Detection With Hybrid ModelsabstractUnmanned Aerial Vehicles (UAVs) are increasingly deployed in safety-critical applications, yet their operations in complex environments make them vulnerable to diverse faults. This paper presents AeroGuard, a lightweight hybrid frame work for real-time UAV fault detection. AeroGuard combines Long Short-Term Memory (LSTM) and AutoRegressive with eXogenous input (ARX) models, with residual-driven adaptive weighting to balance their strengths. Faults are identified through Z-score and Sequential Probability Ratio Test (SPRT) applied to prediction residuals, ensuring accurate and timely detection. Extensive experiments on public datasets, real UAV flight logs, and outdoor flights confirm AeroGuard's robustness, particularly in detecting drift and bias faults where existing methods degrade. AeroGuard achieves up to 95.8% precision, representing about 10% improvement over prior work, while maintaining sub-5ms latency on Raspberry Pi 4B with modest resource usage, and sub second detection on Pi Zero for low-speed UAVs. We also discuss current limitations, noting that evaluation on hardware-induced faults (e.g., motor seizure) will be pursued in future work. Teng Li 0003, Zhili Wei, Yebo Feng, Zhuo Ma 0001, Yulong Shen 0001, Jianfeng Ma 0001, Yang Liu 0003 |
IEEE Trans. Mob. Comput. | 8 |
| 2026 | OptRCA: A More Efficient and Accurate Approach for Automated Root Cause Analysis and ExplanationabstractWith the development of automated software testing technology, software developers can get a large number of crash test cases in a short period of time. However, analyzing these crash test cases and finding their root cause is a time-consuming and labor-intensive task. Techniques based on reverse execution and backward taint analysis are proposed to locate the root cause, but can’t provide context information or explanation of the underlying fault. To address these two limitations, researchers have proposed an automated root cause analysis technique called AURORA. Although this technique provides powerful root cause analysis capabilities, it also have two obvious shortcomings. First, the results of root cause analysis are not accurate enough. Second, the efficiency of root cause analysis is not high enough. In order to improve these two shortcomings, we propose OptRCA, a more efficient and accurate approach for root cause analysis and explanation. Like AURORA’s fuzzing strategy, OptRCA is also designed based on AFL’s crash mode. The difference between them is mainly reflected in three points. First of all, the goal pursued by OptRCA is different from that of normal fuzzing technology. OptRCA pursues maximum correlation to ensure that as many crash test cases as possible are related to the same root cause. This test case with maximum correlation can greatly improve the accuracy of root cause analysis. Second, OptRCA proposed a more efficient non-crash test case retention strategy, which we named “Hill-Climbing Retention.” Using the hill-climbing retention method, OptRCA can obtain sufficient root cause information while retaining only a few non-crash test cases. Since the number of test cases is greatly reduced, the efficiency of OptRCA’s subsequent root cause analysis process is also greatly improved. In addition, OptRCA also optimizes the analysis formula to obtain more accurate analysis results. In the evaluation experimental results, OptRCA is significantly better than AURORA in terms of accuracy and efficiency. Quantitative analysis shows that OptRCA is 65% more accurate and 61% more efficient than AURORA. Jingquan Ge, Yaowen Zheng, Yuekang Li, Wei Ma 0014, Sheikh Mahbub Habib, Praveen Kakkolangara, Gabriel Byman, Yang Liu 0003 |
ACM Trans. Softw. Eng. Methodol. | 8 |
| 2026 | SPOLRE: Semantic Preserving Object Layout Reconstruction for Image Captioning System TestingabstractImage captioning (IC) systems, including Microsoft Azure Cognitive Service, are commonly utilized to convert image content into descriptive natural language. However, inaccuracies in caption generation can lead to serious misinterpretations. Advanced testing techniques such as MetaIC and ROME have been developed to mitigate these issues, yet they encounter notable challenges. First, these strategies demand intensive labor, relying on detailed manual annotations like bounding box data of objects to create test cases. Second, the realism of the generated images is compromised, with MetaIC adding unrelated objects and ROME failing to remove objects effectively. Finally, the capability to generate diversified test suites is restricted. MetaIC is limited to only inserting specific objects to prevent overlap, whereas ROME can generate only \(3^{n}-2^{n}\) variations of test cases from an original seed image containing \( n \) objects. In this study, we present SPOLRE, a novel automated tool designed for semantic preserving object layout reconstruction in image captioning system testing. SPOLRE is based on the insight that modifying the arrangement of objects within an image does not alter its inherent semantics. We utilize four semantic preserving transformation techniques—translation, rotation, mirroring, and scaling—to modify object layouts autonomously, eliminating the need for manual annotation. This approach enables the creation of realistic and varied test suites for IC system testing. Our extensive testing demonstrates that more than 75% of survey respondents find the images produced by SPOLRE more realistic compared to those generated by SOTA methods. Additionally, SPOLRE exhibits outstanding performance in identifying caption errors, detecting 31,544 incorrect captions across seven IC systems with an average precision of 91.62%. This significantly outperforms other methods, which only achieve 85.65% accuracy on average and identify 17,160 incorrect captions. Notably, SPOLRE exposes 6,236 unique issues within Microsoft Azure Cognitive Service, highlighting its effectiveness against one of the most advanced IC systems available. Yi Liu 0069, Guanyu Wang 0005, Gelei Deng, Kailong Wang 0001, Yang Liu 0003, Haoyu Wang 0001 |
ACM Trans. Softw. Eng. Methodol. | 6 |
| 2026 | JailGuard: A Universal Detection Framework for Prompt-based Attacks on LLM SystemsabstractThe systems and software powered by Large Language Models (LLMs) and Multi-Modal Large Language Models (MLLMs) have played a critical role in numerous scenarios. However, current LLM systems are vulnerable to prompt-based attacks, with jailbreaking attacks enabling the LLM system to generate harmful content, while hijacking attacks manipulate the LLM system to perform attacker-desired tasks, underscoring the necessity for detection tools. Unfortunately, existing detecting approaches are usually tailored to specific attacks, resulting in poor generalization in detecting various attacks across different modalities. To address it, we propose JailGuard , a universal detection framework deployed on top of LLM systems for prompt-based attacks across text and image modalities. JailGuard operates on the principle that attacks are inherently less robust than benign ones. Specifically, JailGuard mutates untrusted inputs to generate variants and leverages the discrepancy of the variants’ responses on the target model to distinguish attack samples from benign samples. We implement 18 mutators for text and image inputs and design a mutator combination policy to further improve detection generalization. The evaluation on the dataset containing 15 known attack types suggests that JailGuard achieves the best detection accuracy of 86.14%/82.90% on text and image inputs, outperforming state-of-the-art methods by 11.81–25.73% and 12.20–21.40%. Xiaoyu Zhang 0013, Cen Zhang, Tianlin Li, Yihao Huang 0001, Xiaojun Jia, Ming Hu 0003, Jie Zhang 0073, Yang Liu 0003, Shiqing Ma, Chao Shen 0001 |
ACM Trans. Softw. Eng. Methodol. | 8 |
| 2026 | NeuSemSlice: Towards Effective DNN Model Maintenance via Neuron-Level Semantic SlicingabstractDeep Neural Networks (DNNs), extensively applied across diverse disciplines, are characterized by their integrated and monolithic architectures, setting them apart from conventional software systems. This architectural difference introduces particular challenges to maintenance tasks, such as model restructure (e.g., model compression), re-adaptation (e.g., fitting new samples), and incremental development (e.g., continual knowledge accumulation). Prior research addresses these challenges by identifying task-critical neuron layers and dividing neural networks into semantically similar sequential modules. However, such layer-level approaches fail to precisely identify and manipulate neuron-level semantic components, restricting their applicability to finer-grained model maintenance tasks. In this work, we implement NeuSemSlice, a novel framework that introduces the semantic slicing technique to effectively identify critical neuron-level semantic components in DNN models for semantic-aware model maintenance tasks. Specifically, semantic slicing identifies, categorizes, and merges critical neurons across different categories and layers according to their semantic similarity, enabling their flexibility and effectiveness in the subsequent tasks. For semantic-aware model maintenance tasks, we provide a series of novel strategies based on semantic slicing to enhance NeuSemSlice. They include semantic components (i.e., critical neurons) preservation for model restructure, critical neuron tuning for model re-adaptation, and non-critical neuron training for model incremental development. A thorough evaluation has demonstrated that NeuSemSlice significantly outperforms baselines in all three tasks. Shide Zhou, Tianlin Li, Yihao Huang 0001, Ling Shi 0002, Kailong Wang 0001, Yang Liu 0003, Haoyu Wang 0001 |
ACM Trans. Softw. Eng. Methodol. | 6 |
| 2026 | LIMR: Intent-Aware Mashup API Recommendation via LLM-Augmented Multi-Scale FusionabstractThe increasing availability of Web APIs has amplified the complexity of mashup creation, where developers must identify compatible and functionally relevant APIs based on often ambiguous natural language descriptions. Traditional methods also fall short in capturing hierarchical semantic cues, modeling compatibility, and aligning with developer intent. Although large language models (LLMs) offer strong generalization capabilities, they remain unreliable in mashup recommendation due to hallucinated outputs, limited controllability, and token-length constraints when dealing with large-scale API repositories. To overcome these limitations, we introduceLIMR, an intent-aware mashup recommendation framework that combines LLM-augmented semantic reasoning with structured, multi-scale neural modeling.LIMRfirst prompts a LLM to extract high-level intent from user requirements, which serves as a global semantic signal. This intent is fused with low-level, multi-scale features extracted by a convolutional encoder, which are designed to capture fine-grained lexical/phrasal patterns at different granularities and provide precise semantic grounding for API matching. These heterogeneous representations are further contextually refined through a Transformer-based interaction module. To handle nonlinear semantic dependencies and compositional complexity,LIMRintegrates a Kolmogorov-Arnold Network (KAN) with learnable activation functions, enhancing the model's capacity to capture intricate feature interactions. The entire framework is optimized via LLM, incorporating auxiliary objectives such as mashup category prediction and API quality estimation to guide generalization and reduce overfitting. Comprehensive experiments on the ProgrammableWeb and APIBench datasets show thatLIMRsignificantly outperforms state-of-the-art baselines, which the ranking-oriented metrics, including NDCG and mAP, achieves improvements of 17.1%–34.2% over the strongest competitors. These results confirm the effectiveness ofLIMR's hybrid design in delivering precise, robust, and intent-aware mashup API recommendations, especially in scenarios where LLMs alone fail to meet accuracy and scalability demands. Yao Zhang 0019, Yude Bai, Minhong Dong, Keqing Cen, Ji Zhang 0001, Wei Ma 0014, Yongqiang Lyu 0001, Xiaohong Li 0001, Junjie Wang 0007, Lingxiao Jiang, Yang Liu 0003 |
IEEE Trans. Serv. Comput. | 13 |
| 2026 | CCMG: Enhancing Conventional Commit Message Generation With Hierarchical ContextabstractAutomated commit message generation, which aims at generating natural language description from code change, allows developers to focus more on project maintenance and management. To ensure the quality of commit messages, most projects constrain their style and adopt the conventional commit specification. Conventional commit message generation has been significantly benefited from recent progress in Large Language Models (LLMs). However, previous approaches typically rely on only one or two type of information for the generation, ignoring a wide range of context information. Moreover, they often extract the context in a coarse-grained manner, missing critical details.To address this limitation, We propose CCMG, a novel hierarchical context-augmentedConventionalCommitMessageGeneration framework, which incorporates project-agnostic and project-specific context. For project agnostic context, CCMG retrieves and refines the relevant commits to align conventional commit specification from large-scale corpus. For project-specific context, CCMG provides a wide range of software context information from the perspective of project, code, and style. Finally, CCMG designs two-stage prompt strategy to focus on conventional message inference and commit type adaptation. Compared with the state-of-the-art LLM-based approaches (i.e., OMG and OMEGA), experiment results show that CCMG achieves an average improvement of 31.45% based on human evaluation in commit message generation and improves accuracy by 21.00% and F1 score by 20.88% in commit type classification. Wenke Li, Xuesen Lin, Suyuan Wang, Feng Wu 0003, Cai Fu, Yang Liu 0003 |
IEEE Trans. Software Eng. | 8 |
| 2026 | Beyond Functional Correctness: Exploring Hallucinations in LLM-Generated Code
Fang Liu 0032, Yang Liu 0003, Lin Shi 0006, Zhen Yang 0022, Li Zhang 0029, Xiaoli Lian, Zhong-Qi Li, Yuchi Ma |
IEEE Trans. Software Eng. | 2 |
| 2026 | DeepFWI: Identifying Bug-Sensitive Warnings With Multi-Modal Code-Warning SemanticsabstractStatic analysis tools have evolved over time to assist in detecting bugs. However, the excessive false warnings can impede developers’ productivity and confidence in the tools. Previous research efforts have explored learning-based approaches to identify bug warnings. Nevertheless, their coarse granularity, focusing on either long-term warnings or function-level alerts, are insensitive to individual bugs. Also, they rely on manually crafted features or solely on source code semantics, which is inadequate for effective learning. In this paper, we propose DeepFWI, a learning-based approach that identifies bug-sensitive warnings at a fine-grained granularity. Specifically, we design a novel LSTM-based model that captures multi-modal semantics of source code and warnings from automated static analysis tools (ASATs) and highlights their correlations with cross-attention. To tackle the data scarcity of training and evaluation, we collected a large-scale dataset of 280,273 warnings. We conducted extensive experiments on the dataset to evaluate DeepFWI. The experimental results demonstrate the effectiveness of our approach, with an F1-score 67.06% for confirming true warnings in a finer-grained manner, significantly outperforming all baselines. Additionally, to validate the practicality of DeepFWI from the perspective of developers, we applied DeepFWI to four popular open-source projects. Our approach filtered out the vast majority of warnings, while still successfully surfacing 25 true bug-related warnings that were confirmed through manual analysis. Han Liu 0012, Jian Zhang 0087, Cen Zhang, Kaixuan Li 0002, Sen Chen 0001, Shangwei Lin 0001, Yixiang Chen 0001, Xinghua Li 0001, Yang Liu 0003 |
IEEE Trans. Software Eng. | 10 |
| 2026 | Causality-Aware Safety Testing for Autonomous Driving SystemsabstractSimulation-based testing is essential for evaluating the safety of Autonomous Driving Systems (ADSs). Comprehensive evaluation requires testing across diverse scenarios that can trigger various types of violations under different conditions. While existing methods typically focus on individual diversity metrics, such as input scenarios, ADS-generated motion commands, and system violations, they often fail to capture the complex interrelationships among these elements. For instance, identical motion commands can produce different collision risks in varying scenes, and the same collision may result from different commands under different scenarios. This oversight leads to gaps in testing coverage, potentially missing critical issues in the ADS under evaluation. In this paper, we proposeCausal-Fuzzer, the first causality-aware fuzzing technique that enables efficient and comprehensive testing of ADSs by constructing causal graphs to model the interrelationships among scenarios, actions, and violations. Unlike existing methods that treat diversity metrics independently, we recognize these elements are causally interconnected and use their relationships to identify more diverse violations triggered by fundamentally different causal mechanisms. Specifically,Causal-Fuzzerproposes (1) a causality-based feedback mechanism that quantifies the combined diversity of test scenarios by assessing whether they activate new causal relationships, and (2) a causality-driven mutation strategy that prioritizes mutations on input scenario elements with higher causal impact on ego action changes and violation occurrence to enable interpretable and efficient test generation. We evaluatedCausal-Fuzzeron an industry-grade ADS Apollo, with a high-fidelity simulator LGSVL. Our empirical results demonstrate thatCausal-Fuzzersignificantly outperforms existing methods in (1) identifying a greater diversity of violations (96.5 violations on average, compared to 66.9 for the best baseline method), (2) providing enhanced testing sufficiency with improved coverage of causal relationships (13.6 unique sceneaction- violation patterns on average, compared to 8.6 for the best baseline method), and (3) achieving greater efficiency in detecting critical scenarios, strong robustness under noise conditions, and good generalizability across varying scenario complexities and violation types. Our source code and experimental results are available athttps://sites.google.com/view/causal-fuzzer. Wenbing Tang 0001, Mingfei Cheng, Yuan Zhou 0005, Yang Liu 0003, Zuohua Ding |
IEEE Trans. Software Eng. | 6 |
| 2026 | Evaluating Large Language Models for Line-Level Vulnerability LocalizationabstractRecently, Automated Vulnerability Localization (AVL) has attracted growing attention, aiming to facilitate diagnosis by pinpointing the specific lines of code responsible for vulnerabilities. Large Language Models (LLMs) have shown potential in various domains, yet their effectiveness in line-level vulnerability localization remains underexplored.In this work, we present the first comprehensive empirical evaluation of LLMs for AVL. Our study examines 19 leading LLMs suitable for code analysis, including ChatGPT and multiple open-source models, spanning encoder-only, encoder-decoder, and decoder-only architectures, with model sizes from 60M to 70B parameters. We evaluate three paradigms—few-shot prompting, discriminative fine-tuning, and generative fine-tuning—with and without Low-Rank Adaptation (LoRA), on both a BigVul-derived dataset for C/C++ and a smart contract vulnerability dataset.Our results show that discriminative fine-tuning achieves substantial performance gains over existing learning-based AVL methods when sufficient training data is available. In low-data settings, prompting advanced LLMs such as ChatGPT proves more effective. We also identify challenges related to input length and unidirectional context during fine-tuning, and propose two remedial strategies: a sliding window approach and right-forward embedding, both of which yield significant improvements. Moreover, we provide the first assessment of LLM generalizability in AVL, showing that certain models can transfer effectively across Common Weakness Enumerations (CWEs) and projects. However, performance degrades notably for newly discovered vulnerabilities containing unfamiliar lexical or structural patterns, underscoring the need for continual adaptation. These findings offer practical guidance for deploying LLM-based AVL systems in realistic software security workflows. Jian Zhang 0087, Chong Wang 0013, Anran Li 0001, Weisong Sun, Cen Zhang, Wei Ma 0014, Yang Liu 0003 |
IEEE Trans. Software Eng. | 7 |
| 2026 | Software Architecture Recovery Augmented With SemanticsabstractThe architecture of software systems evolves along with their upgrades and maintenance, inevitably creating a gap between the defact architecture and the designed one. To perceive and fix the discrepancy, clustering-based architecture recovery methods have been developed to re-engineer the real-time system architecture from the code implementation. However, existing solutions still face several limitations. They underutilize both code-level and architecture-level semantics underlying the source code. Moreover, they overlook implicit structural dependencies that complement explicit ones to reflect code interactions. To address these challenges, we propose SemArc, an architecture recovery method that utilizes large language models to comprehend both implementation-level and architecture-level semantics, supported by well-established canonical architectural patterns as a knowledge base. SemArc also incorporates both implicit and explicit dependencies to complete the system behavior representations. Additionally, SemArc introduces a component-as-anchor guided clustering algorithm to improve the clustering process. We evaluated SemArc on 15 software systems written in C/C++, Java, and Python, using five different metrics. The results demonstrate that SemArc outperforms seven baseline methods by an average of 32 percentage points. We also examined how three factors—code semantics, architectural semantics, and implicit dependencies—as well as different levels of architectural semantic descriptions, influence recovery accuracy. A case study on the Bash project indicates that SemArc has the potential to yield even more precise recovery results than those labeled by humans. Wuxia Jin, Ming Fan 0002, Haijun Wang 0002, Li Li 0044, Yang Liu 0003, Ting Liu 0002 |
IEEE Trans. Software Eng. | 7 |
| 2026 | Hierarchical Resource Optimization for Covert SAGINs: A Stackelberg-Matching Game Approach
Min Wu 0008, Kefeng Guo, Theodoros A. Tsiftsis, Shahid Mumtaz, Yang Liu 0003, Zhiming Zheng 0001 |
IEEE Trans. Wirel. Commun. | 6 |
| 2025 | Perception-Guided Jailbreak Against Text-to-Image ModelsabstractIn recent years, Text-to-Image (T2I) models have garnered significant attention due to their remarkable advancements. However, security concerns have emerged due to their potential to generate inappropriate or Not-Safe-For-Work (NSFW) images. In this paper, inspired by the observation that texts with different semantics can lead to similar human perceptions, we propose an LLM-driven perception-guided jailbreak method, termed PGJ. It is a black-box jailbreak method that requires no specific T2I model (model-free) and generates highly natural attack prompts. Specifically, we propose identifying a safe phrase that is similar in human perception yet inconsistent in text semantics with the target unsafe word and using it as a substitution. The experiments conducted on six open-source models and commercial online services with thousands of prompts have verified the effectiveness of PGJ. Yihao Huang 0001, Le Liang, Tianlin Li, Xiaojun Jia, Run Wang 0001, Weikai Miao, Geguang Pu, Yang Liu 0003 |
AAAI | 8 |
| 2025 | Logic-Q: Improving Deep Reinforcement Learning-based Quantitative Trading via Program Sketch-based TuningabstractDeep reinforcement learning (DRL) has revolutionized quantitative trading (Q-trading) by achieving decent performance without significant human expert knowledge. Despite its achievements, we observe that the current state-of-the-art DRL models are still ineffective in identifying the market trends, causing them to miss good trading opportunity or suffer from large drawdowns when encountering market crashes. To address this limitation, a natural approach is to incorporate human expert knowledge in identifying market trends. Whereas, such knowledge is abstract and hard to be quantified. In order to effectively leverage abstract human expert knowledge, in this paper, we propose a universal logic-guided deep reinforcement learning framework for Q-trading, called Logic-Q. In particular, Logic-Q adopts the program synthesis by sketching paradigm and introduces a logic-guided model design that leverages a lightweight, plug-and-play market trend-aware program sketch to determine the market trend and correspondingly adjusts the DRL policy in a post-hoc manner. Extensive evaluations of two popular quantitative trading tasks demonstrate that Logic-Q can significantly improve the performance of previous state-of-the-art DRL trading strategies. Junzhe Jiang 0002, Yushi Cao, Aixin Cui, Bozhi Wu, Bo Li 0037, Yang Liu 0003, Danny Dongning Sun |
AAAI | 7 |
| 2025 | Efficient Universal Goal Hijacking with Semantics-guided Prompt OrganizationabstractYihao Huang, Chong Wang, Xiaojun Jia, Qing Guo, Felix Juefei-Xu, Jian Zhang, Yang Liu, Geguang Pu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Yihao Huang 0001, Chong Wang 0013, Xiaojun Jia, Qing Guo 0005, Felix Juefei-Xu, Jian Zhang 0087, Yang Liu 0003, Geguang Pu |
ACL (1) | 7 |
| 2025 | Personality-Guided Code Generation Using Large Language ModelsabstractCode generation, the automatic creation of source code from natural language descriptions, has garnered significant attention due to its potential to streamline software development.Inspired by research that links taskpersonality alignment with improved development outcomes, we conduct an empirical study on personality-guided code generation using large language models (LLMs).Specifically, we investigate how emulating personality traits appropriate to the coding tasks affects LLM performance.We extensively evaluate this approach using seven widely adopted LLMs across four representative datasets.Our results show that personality guidance significantly enhances code generation accuracy, with improved pass rates in 23 out of 28 LLM-dataset combinations.Notably, in 11 cases, the improvement exceeds 5%, and in 5 instances, it surpasses 10%, with the highest gain reaching 12.9%.Additionally, personality guidance can be easily integrated with other prompting strategies to further boost performance. Yaoqi Guo, Zhenpeng Chen 0001, Jie Zhang 0050, Yang Liu 0003, Yun Ma 0002 |
ACL (1) | 4 |
| 2025 | The Invisible Hand: Unveiling Provider Bias in Large Language Models for Code GenerationabstractXiaoyu Zhang, Juan Zhai, Shiqing Ma, Qingshuang Bao, Weipeng Jiang, Qian Wang, Chao Shen, Yang Liu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Xiaoyu Zhang 0013, Juan Zhai, Shiqing Ma, Qingshuang Bao, Qian Wang 0002, Chao Shen 0001, Yang Liu 0003 |
ACL (1) | 8 |
| 2025 | Oedipus: LLM-enchanced Reasoning CAPTCHA SolverabstractCAPTCHAs have become a ubiquitous tool in safeguarding applications from automated bots. Over time, the arms race between CAPTCHA development and evasion techniques has led to increasingly sophisticated and diverse designs. The latest iteration, reasoning CAPTCHAs, exploits tasks that are intuitively simple for humans but challenging for conventional AI technologies, thereby enhancing security measures. Gelei Deng, Haoran Ou, Yi Liu 0069, Jie Zhang 0073, Tianwei Zhang 0004, Yang Liu 0003 |
CCS | 6 |
| 2025 | Slot: Provenance-Driven APT Detection through Graph Reinforcement LearningabstractAdvanced Persistent Threats (APTs) represent sophisticated cyberattacks characterized by their ability to remain undetected within the victim system for extended periods, aiming to exfiltrate sensitive data or disrupt operations. Existing detection approaches often struggle to effectively identify these complex threats, construct the attack chain for defense facilitation, or resist adversarial attacks. To overcome these challenges, we propose Slot, an advanced APT detection approach based on provenance graphs and graph reinforcement learning. Slot excels in uncovering multi-level hidden relationships, such as causal, contextual, and indirect connections, among system behaviors through provenance graph mining. Slot implements semi-supervised learning with limited labels through efficient label similarity computation, significantly enhancing both detection performance and model robustness. By pioneering the integration of graph reinforcement learning, Slot dynamically adapts to new user activities and evolving attack strategies, enhancing its resilience against adversarial attacks. Additionally, Slot automatically constructs the attack chain according to detected attacks with clustering algorithms, providing precise identification of attack paths and facilitating the development of defense strategies. Evaluations with real-world datasets demonstrate Slot's outstanding accuracy, efficiency, adaptability, and robustness in APT detection, with most metrics surpassing state-of-the-art methods. Additionally, case studies conducted to assess Slot's effectiveness in supporting APT defense further establish it as a practical and reliable tool for cybersecurity protection. Wei Qiao 0005, Yebo Feng, Teng Li 0003, Zhuo Ma 0001, Yulong Shen 0001, Jianfeng Ma 0001, Yang Liu 0003 |
CCS | 7 |
| 2025 | SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World EnvironmentsabstractLarge vision-language models (LVLMs) have shown remarkable capabilities in interpreting visual content. While existing works demonstrate these models’ vulnerability to deliberately placed adversarial texts, such texts are often easily identifiable as anomalous. In this paper, we present the first approach to generate scene-coherent typographic adversarial attacks that mislead advanced LVLMs while maintaining visual naturalness through the capability of the LLM-based agent. Our approach addresses three critical questions: what adversarial text to generate, where to place it within the scene, and how to integrate it seamlessly. We propose a training-free, multi-modal LLM-driven scene-coherent typographic adversarial planning (SceneTAP) that employs a three-stage process: scene understanding, adversarial planning, and seamless integration. The SceneTAP utilizes chain-of-thought reasoning to comprehend the scene, formulate effective adversarial text, strategically plan its placement, and provide detailed instructions for natural integration within the image. This is followed by a scene-coherent TextDiffuser that executes the attack using a local diffusion mechanism. We extend our method to real-world scenarios by printing and placing generated patches in physical environments, demonstrating its practical implications. Extensive experiments show that our scene-coherent adversarial text successfully misleads state-of-the-art LVLMs, including ChatGPT-4o, even after capturing new images of physical setups. Our evaluations demonstrate a significant increase in attack success rates while maintaining visual naturalness and contextual appropriateness. This work highlights vulnerabilities in current vision-language models to sophisticated, scene-coherent adversarial attacks and provides insights into potential defense mechanisms. We release our code at https://github.com/tsingqguo/scenetap. Jie Zhang 0002, Di Lin 0002, Tianwei Zhang 0004, Ivor W. Tsang, Yang Liu 0003, Qing Guo 0005 |
CVPR | 7 |
| 2025 | LVAgent: Long Video Understanding by Multi-Round Dynamical Collaboration of MLLM AgentsabstractExisting MLLMs encounter significant challenges in modeling the temporal context within long videos. Currently, mainstream Agent-based methods use external tools to assist a single MLLM in answering long video questions. Despite such tool-based support, a solitary MLLM still offers only a partial understanding of long videos, resulting in limited performance. In order to better address long video tasks, we introduce LVAgent, the first framework enabling multi-round dynamic collaboration of MLLM agents in long video understanding. Our method consists of four key steps: 1) Selection: We pre-select appropriate agents from the model library to form optimal agent teams based on different tasks. 2) Perception: We design an effective retrieval scheme for long videos to improve the coverage of critical temporal segments while maintaining computational efficiency. 3) Action: Agents answer long video questions and exchange reasons. 4) Reflection: We evaluate each agent's performance in each round of discussion and optimize the agent team for dynamic collaboration. The agents iteratively refine their answers by multi-round dynamical collaboration of MLLM agents. LVAgent is the first agent system method that outperforms all closed-source models (like GPT-4o) and open-source models (like InternVL-2.5 and Qwen2-VL) in the long video understanding tasks. Our LVAgent achieves an accuracy of 80\% on four mainstream long video understanding tasks. Notably, LVAgent improves accuracy by 13.3\% on LongVideoBench. Code is available at https://github.com/64327069/LVAgent. Zhengrong Yue, Siran Chen, Zikang Wang, Yang Liu 0003, Peng Li 0030, Yali Wang 0001 |
ICCV | 5 |
| 2025 | VideoLLaMB: Long Streaming Video Understanding with Recurrent Memory Bridges
Yuxuan Wang 0004, Yiqi Song, Cihang Xie, Yang Liu 0003, Zilong Zheng |
ICCV | 4 |
| 2025 | Know Your Account: Double Graph Inference-Based Account De-Anonymization on EthereumabstractThe scaled Web 3.0 digital economy, represented by decentralized finance (DeFi), has sparked increasing interest in the past few years, which usually relies on blockchain for token transfer and diverse transaction logic. However, illegal behaviors, such as financial fraud, hacker attacks, and money laundering, are rampant in the blockchain ecosystem and seriously threaten its integrity and security. In this paper, we propose a novel double graph-based Ethereum account de-anonymization inference method, dubbed DBG4ETH, which aims to capture the behavioral patterns of accounts comprehensively and has more robust analytical and judgment capabilities for current complex and continuously generated transaction behaviors. Specifically, we first construct a global static graph to build complex interactions between the various account nodes for all transaction data. Then, we also construct a local dynamic graph to learn about the gradual evolution of transactions over different periods. Different graphs focus on information from different perspectives, and features of global and local, static and dynamic transaction graphs are available through DBG4ETH. In addition, we propose an adaptive confidence calibration method to predict the results by feeding the calibrated weighted prediction values into the classifier. Experimental results show that DBG4ETH achieves state-of-the-art results in the account identification task, improving the F1-score by at least 3.75% and up to 40.52% compared to processing each graph type individually and outperforming similar account identity inference methods by 5.23 % to 12.91 %. Shuyi Miao, Wangjie Qiu, Hongwei Zheng 0003, Qinnan Zhang, Xiaofan Tu, Xunan Liu, Yang Liu 0003, Jin Dong 0004, Zhiming Zheng 0001 |
ICDE | 7 |
| 2025 | An Analytical Perspective on Software Engineering for Large Language Models
Tianlin Li, Chong Wang 0013, Jian Zhang 0087, Wei Ma 0014, Aishan Liu, Jingyi Wang 0004, Yang Liu 0003 |
ICECCS | 8 |
| 2025 | Evolaris: A Roadmap to Self-evolving Software Intelligence Management
Wenbo Guo 0011, Sen Chen 0001, Lei Bu, Yang Liu 0003 |
ICECCS | 7 |
| 2025 | Empowering Embodied Agents with Semantic Intelligence
Wenbing Tang 0001, Meilin Zhu, Fenghua Wu, Xinfeng Li, Yang Liu 0003 |
ICECCS | 5 |
| 2025 | UFPC: A Unified Framework for Source and Binary Program Comprehension
Weisong Sun, Yuqiang Sun 0001, Yang Liu 0003 |
ICECCS | 6 |
| 2025 | Agent Behavior: The Regulatory Object of the Agent-Centric Online Ecosystem in Digital Age
Qiang Zhang 0057, Pei Yan, Yijia Xu, Xinfeng Li, Hongyi Cai, Chuanpo Fu, Yong Fang 0002, Yang Liu 0003 |
ICECCS | 8 |
| 2025 | Improved Techniques for Optimization-Based Jailbreaking on Large Language ModelsabstractLarge language models (LLMs) are being rapidly developed, and a key component of their widespread deployment is their safety-related alignment. Many red-teaming efforts aim to jailbreak LLMs, where among these efforts, the Greedy Coordinate Gradient (GCG) attack's success has led to a growing interest in the study of optimization-based jailbreaking techniques. Although GCG is a significant milestone, its attacking efficiency remains unsatisfactory. In this paper, we present several improved (empirical) techniques for optimization-based jailbreaks like GCG. We first observe that the single target template of ”Sure'' largely limits the attacking performance of GCG; given this, we propose to apply diverse target templates containing harmful self-suggestion and/or guidance to mislead LLMs. Besides, from the optimization aspects, we propose an automatic multi-coordinate updating strategy in GCG (i.e., adaptively deciding how many tokens to replace in each step) to accelerate convergence, as well as tricks like easy-to-hard initialization. Then, we combine these improved technologies to develop an efficient jailbreak method, dubbed $\mathcal{I}$-GCG. In our experiments, we evaluate our $\mathcal{I}$-GCG on a series of benchmarks (such as NeurIPS 2023 Red Teaming Track). The results demonstrate that our improved techniques can help GCG outperform state-of-the-art jailbreaking attacks and achieve a nearly 100\% attack success rate.
The code is released at https://github.com/jiaxiaojunQAQ/I-GCG. Xiaojun Jia, Tianyu Pang, Yihao Huang 0001, Jindong Gu, Yang Liu 0003, Xiaochun Cao |
ICLR | 6 |
| 2025 | Samba: Simple Hybrid State Space Models for Efficient Unlimited Context Language ModelingabstractEfficiently modeling sequences with infinite context length has long been a challenging problem. Previous approaches have either suffered from quadratic computational complexity or limited extrapolation ability in length generalization. In this
work, we present Samba, a simple hybrid architecture that layer-wise combines
Mamba, a selective State Space Model (SSM), with Sliding Window Attention
(SWA). Samba selectively compresses a given sequence into recurrent hidden
states while still maintaining the ability to precisely recall recent memories with the
attention mechanism. We scale Samba up to 3.8B parameters with 3.2T training
tokens and demonstrate that it significantly outperforms state-of-the-art models
across a variety of benchmarks. Pretrained on sequences of 4K length, Samba
shows improved perplexity in context lengths of up to 1M in zero-shot. When
finetuned on 4K-length sequences, Samba efficiently extrapolates to a 256K context length with perfect memory recall on the Passkey Retrieval task, and exhibits
superior retrieval extrapolation on the challenging Phonebook task compared to
full-attention models. As a linear-time sequence model, Samba achieves a 3.73×
higher throughput compared to Transformers with grouped-query attention for user
prompts of 128K length, and a 3.64× speedup when generating 64K tokens with
unlimited streaming. Liliang Ren, Yang Liu 0003, Yadong Lu, Yelong Shen, Chen Liang 0006, Weizhu Chen |
ICLR | 2 |
| 2025 | When Prompt Engineering Meets Software Engineering: CNL-P as Natural and Robust "APIs" for Human-AI InteractionabstractWith the growing capabilities of large language models (LLMs), they are increasingly applied in areas like intelligent customer service, code generation, and knowledge management.
Natural language (NL) prompts act as the ``APIs'' for human-LLM interaction.
To improve prompt quality, best practices for prompt engineering (PE) have been developed, including writing guidelines and templates.
Building on this, we propose Controlled NL for Prompt (CNL-P), which not only incorporates PE best practices but also draws on key principles from software engineering (SE).
CNL-P introduces precise grammar structures and strict semantic norms, further eliminating NL's ambiguity, allowing for a declarative but structured and accurate expression of user intent.
This helps LLMs better interpret and execute the prompts, leading to more consistent and higher-quality outputs.
We also introduce an NL2CNL-P conversion tool based on LLMs, enabling users to write prompts in NL, which are then transformed into CNL-P format, thus lowering the learning curve of CNL-P.
In particular, we develop a linting tool that checks CNL-P prompts for syntactic and semantic accuracy, applying static analysis techniques to NL for the first time.
Extensive experiments demonstrate that CNL-P enhances the quality of LLM responses through the novel and organic synergy of PE and SE.
We believe that CNL-P can bridge the gap between emerging PE and traditional SE, laying the foundation for a new programming paradigm centered around NL.
To further demonstrate the effectiveness of CNL-P, we develop the AI-native CNL-P IDE, UGAiForge that leverages CNL-P to enable users to Think, Describe, and Build with AI and for AI.
For detailed information, please refer https://ugaiforge.ai. Zhenchang Xing, Yang Liu 0003, Dehai Zhao, Daniel Sun 0006, Chenhua Liu |
ICLR | 2 |
| 2025 | STAFF: Speculative Coreset Selection for Task-Specific Fine-tuningabstractTask-specific fine-tuning is essential for the deployment of large language models (LLMs), but it requires significant computational resources and time. Existing solutions have proposed coreset selection methods to improve data efficiency and reduce model training overhead, but they still have limitations: ❶ Overlooking valuable samples at high pruning rates, which degrades the coreset’s performance.
❷ Requiring high time overhead during coreset selection to fine-tune and evaluate the target LLM. In this paper, we introduce STAFF, a speculative coreset selection method. STAFF leverages a small model from the same family as the target LLM to efficiently estimate data scores and then verifies the scores on the target LLM to accurately identify and allocate more selection budget to important regions while maintaining coverage of easy regions. We evaluate STAFF on three LLMs and three downstream tasks and show that STAFF improves the performance of SOTA methods by up to 54.3% and reduces selection overhead by up to 70.5% at different pruning rates. Furthermore, we observe that the coreset selected by STAFF at low pruning rates (i.e., 20%) can even obtain better fine-tuning performance than the full dataset. Xiaoyu Zhang 0013, Juan Zhai, Shiqing Ma, Chao Shen 0001, Tianlin Li, Yang Liu 0003 |
ICLR | 7 |
| 2025 | SpatialMe: Stereo Video Conversion Using Depth-Warping and Blend-InpaintingabstractStereo video conversion aims to transform monocular videos into immersive stereo format. Despite the advancements in novel view synthesis, it still remains two major challenges: i) difficulty of achieving high-fidelity and stable results, and ii) insufficiency of high-quality stereo video data. In this paper, we introduce SpatialMe, a novel stereo video conversion framework based on depth-warping and blend-inpainting. Specifically, we propose a mask-based hierarchy feature update (MHFU) refiner, which integrate and refine the outputs from designed multi-branch inpainting module, using feature update unit (FUU) and mask mechanism. We also propose a disparity expansion strategy to address the problem of foreground bleeding. Furthermore, we conduct a high-quality real-world stereo video dataset—StereoV1K, to alleviate the data shortage. It contains 1000 stereo videos captured in real-world at a resolution of 1180×1180, covering various indoor and outdoor scenes. Extensive experiments demonstrate the superiority of our approach in generating stereo videos over state-of-the-art methods. Qianxi Jia, Yang Liu 0003, Wei Zhang 0012 |
ICME | 3 |
| 2025 | Defending LVLMs Against Vision Attacks Through Partial-Perception SupervisionabstractRecent studies have raised significant concerns regarding the vulnerability of Large Vision Language Models (LVLMs) to maliciously injected or perturbed input images, which can mislead their responses. Existing defense methods show that such vision attacks are sensitive to image modifications especially cropping, using majority voting across responses of modified images as corrected responses. However, these modifications often result in partial images and distort the semantics, which reduces response quality on clean images after voting. Instead of directly using responses from partial images for voting, we investigate using them to supervise (guide) the LVLM’s responses to the original images at inference time. We propose a black-box, training-free method called DPS (Defense through Partial-Perception Supervision). In this approach, the model is prompted using the responses generated by a model that perceives only a partial image. With DPS, the model can adjust its response based on partial image understanding when under attack, while confidently maintaining its original response for clean input. Empirical experiments show our method outperforms the baseline, cutting the average attack success rate by 76.3% across six datasets on three popular models. Qi Zhou 0012, Dongxia Wang 0002, Tianlin Li, Yun Lin 0001, Yang Liu 0003, Jin Song Dong 0001, Qing Guo 0005 |
ICML | 5 |
| 2025 | TIGER: A Generating-Then-Ranking Framework for Practical Python Type InferenceabstractPython's dynamic typing system offers flexibility and expressiveness but can lead to type-related errors, prompting the need for automated type inference to enhance type hinting. While existing learning-based approaches show promising inference accuracy, they struggle with practical challenges in comprehensively handling various types, including complex parameterized types and (unseen) user-defined types. In this paper, we introduce TIGER, a two-stage generating-then-ranking (GTR) framework, designed to effectively handle Python's diverse type categories. TIGER leverages fine-tuned pre-trained code models to train a generative model with a span masking objective and a similarity model with a contrastive training objective. This approach allows TIGER to generate a wide range of type candidates, including complex parameterized types in the generating stage, and accurately rank them with user-defined types in the ranking stage. Our evaluation on the ManyTypes4Py dataset shows TIGER's advantage over existing methods in various type categories, notably improving accuracy in inferring user-defined and unseen types by 11.2% and 20.1% respectively in Top-5 Exact Match. Moreover, the experimental results not only demonstrate TIGER's superior performance and efficiency, but also underscore the significance of its generating and ranking stages in enhancing automated type inference. Chong Wang 0013, Jian Zhang 0087, Yiling Lou, Mingwei Liu 0002, Weisong Sun, Yang Liu 0003, Xin Peng 0001 |
ICSE | 6 |
| 2025 | Diversity Drives Fairness: Ensemble of Higher Order Mutants for Intersectional Fairness of Machine Learning SoftwareabstractIntersectional fairness is a critical requirement for Machine Learning (ML) software, demanding fairness across subgroups defined by multiple protected attributes. This paper introduces FairHOME, a novel ensemble approach using higher order mutation of inputs to enhance intersectional fairness of ML software during the inference phase. Inspired by social science theories highlighting the benefits of diversity, FairHOME generates mutants representing diverse subgroups for each input instance, thus broadening the array of perspectives to foster a fairer decision-making process. Unlike conventional ensemble methods that combine predictions made by different models, FairHOME combines predictions for the original input and its mutants, all generated by the same ML model, to reach a final decision. Notably, FairHOME is even applicable to deployed ML software as it bypasses the need for training new models. We extensively evaluate FairHOME against seven state-of-the-art fairness improvement methods across 24 decision-making tasks using widely adopted metrics. FairHOME consistently outperforms existing methods across all metrics considered. On average, it enhances intersectional fairness by 47.5 %, surpassing the currently best-performing method by 9.6 percentage points. Zhenpeng Chen 0001, Jie Zhang 0050, Federica Sarro, Yang Liu 0003 |
ICSE | 5 |
| 2025 | Template-Guided Program Repair in the Era of Large Language ModelsabstractRecent advancements in automated program repair (APR) have been significantly driven by the application of Large Language Models (LLMs). In particular, the integration of LLMs with traditional template-based repair methods has demonstrated effective outcomes. Despite this, the synergy between the strengths of traditional methods and LLMs remains underexploited. This oversight originates from the indiscriminate use of templates and their insufficient coverage. Also, using small-scale LLMs within the zero-shot learning context proves to be suboptimal. To alleviate the limitations, we propose NTR (Neural Template Repair), a two-stage repair framework including template selection and patch generation, both of which are under the fine-tuning paradigm. In the template selection phase, we formulate it as a multiclass classification problem and fine-tune million-level LLMs for better selecting possible templates. During the patch generation phase, we leverage the chosen templates as probable directions (e.g., ‘Mutate Conditional Expression’) to guide the fine-tuning process of LLMs at the billion-level scale for precise patch creation. Moreover, we incorporate a unique template to signify the absence of a suitable template and employ a probability-based prioritization of templates, thereby optimizing patch generation. This framework not only effectively addresses template mismatch issues, but also enables the billion-level LLMs to explore the patch space more efficiently, despite the GPU memory constraints. We evaluate NTR with different foundational models on Defects4J V1.2 and HumanEval-Java, the framework consistently demonstrates significant effectiveness. When utilizing StarCoder as the foundational model for patch generation, NTR fixes 128 and 129 bugs in Defects4J and HumanEval, outperforming the best baseline APR tool by 14 and 59 bugs. With the larger CodeLlama model, the fixed bugs rise to 139 and 136, respectively, exceeding the baseline by 25 and 66 bugs. Notably, the performance stems not only from the foundational models but also benefits greatly from our NTR framework. Specifically, NTR's implementation with StarCoder and CodeLlama leads to 22 and 23 additional fixes, which is beyond what the models achieve on their own. This emphasizes the success of our new perspective on utilizing templates to unlock the bug-fixing potential of LLMs. Jian Zhang 0087, Xiangxin Meng, Yang Liu 0003 |
ICSE | 4 |
| 2025 | LLM Based Input Space Partitioning Testing for Library APIsabstractAutomated library APIs testing is difficult as it requires exploring a vast space of parameter inputs that may involve objects with complex data types. Existing search based approaches, with limited knowledge of relations between object states and program branches, often suffer from the low efficiency issue, i.e., tending to generate invalid inputs. Symbolic execution based approaches can effectively identify such relations, but fail to scale to large programs. In this work, we present an LLM-based input space partitioning testing approach, LISP, for library APIs. The approach lever-ages LLMs to understand the code of a library API under test and perform input space partitioning based on its understanding and rich common knowledge. Specifically, we provide the signature and code of the API under test to LLMs, with the expectation of obtaining a text description of each input space partition of the API under test. Then, we generate inputs through employing the generated text description to sample inputs from each partition, ultimately resulting in test suites that systematically explore the program behavior of the API. We evaluate LISP on more than 2,205 library API meth-ods taken from 10 popular open-source Java libraries (e.g.,$a$p$a$che/commons-lang with 2.6k stars, guava with 48.8k stars on GitHub). Our experiment results show that LISP is effective in library API testing. It significantly outperforms state-of-the-art tool EvoSuite in terms of edge coverage. On average, LISP achieves 67.82 % branch coverage, surpassing EvoSuite by 1.21 times. In total, LISP triggers 404 exceptions or errors in the experiments, and discovers 13 previously unknown vulnerabilities during evaluation, which have been assigned CVE IDs. Jiageng Li, Chong Wang 0013, Haozhen You, Cen Zhang, Yang Liu 0003, Xin Peng 0001 |
ICSE | 6 |
| 2025 | Combining Fine-Tuning and LLM-Based Agents for Intuitive Smart Contract Auditing with JustificationsabstractSmart contracts are decentralized applications built atop blockchains like Ethereum. Recent research has shown that large language models (LLMs) have potential in auditing smart contracts, but the state-of-the-art indicates that even GPT-4 can achieve only 30% precision (when both decision and justification are correct). This is likely because off-the-shelf LLMs were primarily pre-trained on a general text/code corpus and not fine-tuned on the specific domain of Solidity smart contract auditing. In this paper, we propose iAudit, a general framework that combines fine-tuning and LLM-based agents for intuitive smart contract auditing with justifications. Specifically, iAudit is inspired by the observation that expert human auditors first perceive what could be wrong and then perform a detailed analysis of the code to identify the cause. As such, iAudit employs a two-stage fine-tuning approach: it first tunes a Detector model to make decisions and then tunes a Reasoner model to generate causes of vulnerabilities. However, fine-tuning alone faces challenges in accurately identifying the optimal cause of a vulnerability. Therefore, we introduce two LLM-based agents, the Ranker and Critic, to iteratively select and debate the most suitable cause of vulnerability based on the output of the fine-tuned Reasoner model. To evaluate iAudit, we collected a balanced dataset with 1,734 positive and 1,810 negative samples to fine-tune iAudit. We then compared it with traditional fine-tuned models (CodeBERT, GraphCodeBERT, CodeT5, and UnixCoder) as well as prompt learning-based LLMs (GPT4, GPT-3.5, and CodeLlama-13b/34b). On a dataset of 263 real smart contract vulnerabilities, iAudit achieves an F1 score of 91.21% and an accuracy of 91.11%. The causes generated by iAudit achieved a consistency of about 38% compared to the ground truth causes. Wei Ma 0014, Daoyuan Wu, Yuqiang Sun 0001, Tianwen Wang, Shangqing Liu, Jian Zhang 0087, Yue Xue, Yang Liu 0003 |
ICSE | 8 |
| 2025 | Show Me Your Code! Kill Code Poisoning: A Lightweight Method Based on Code NaturalnessabstractNeural code models (NCMs) have demonstrated extraordinary capabilities in code intelligence tasks. Meanwhile, the security of NCMs and NCMs-based systems has garnered increasing attention. In particular, NCMs are often trained on large-scale data from potentially untrustworthy sources, providing attackers with the opportunity to manipulate them by inserting crafted samples into the data. This type of attack is called a code poisoning attack (also known as a backdoor attack). It allows attackers to implant backdoors in NCMs and thus control model behavior, which poses a significant security threat. However, there is still a lack of effective techniques for detecting various complex code poisoning attacks. In this paper, we propose an innovative and lightweight technique for code poisoning detection named KillbadCode. KillbadCode is designed based on our insight that code poisoning disrupts the naturalness of code. Specifically, KillBADCODE first builds a code language model (CodeLM) on a lightweight$n$-gram language model. Then, given poisoned data, KillbadCode utilizes CodeLM to identify those tokens in (poisoned) code snippets that will make the code snippets more natural after being deleted as trigger tokens. Considering that the removal of some normal tokens in a single sample might also enhance code naturalness, leading to a high false positive rate (FPR), we aggregate the cumulative improvement of each token across all samples. Finally, KillbadCode purifies the poisoned data by removing all poisoned samples containing the identified trigger tokens. We conduct extensive experiments to evaluate the effectiveness and efficiency of KillbadCode, involving two types of advanced code poisoning attacks (a total of five poisoning strategies) and datasets from four representative code intelligence tasks. The experimental results demonstrate that across 20 code poisoning detection scenarios, KillbadCode achieves an average FPR of 8.30 % and an average Recall of 100 %, significantly outperforming four baselines. More importantly, KillBadCode is very efficient, with a minimum time consumption of only 5 minutes, and is 25 times faster than the best baseline on average. Weisong Sun, Mengzhe Yuan, Chunrong Fang, Zhenpeng Chen 0001, Chong Wang 0013, Yang Liu 0003, Baowen Xu, Zhenyu Chen 0001 |
ICSE | 7 |
| 2025 | Source Code Summarization in the Era of Large Language ModelsabstractTo support software developers in understanding and maintaining programs, various automatic (source) code summarization techniques have been proposed to generate a concise natural language summary (i.e., comment) for a given code snippet. Recently, the emergence of large language models (LLMs) has led to a great boost in the performance of coderelated tasks. In this paper, we undertake a systematic and comprehensive study on code summarization in the era of LLMs, which covers multiple aspects involved in the workflow of LLMbased code summarization. Specifically, we begin by examining prevalent automated evaluation methods for assessing the quality of summaries generated by LLMs and find that the results of the GPT-4 evaluation method are most closely aligned with human evaluation. Then, we explore the effectiveness of five prompting techniques (zero-shot, few-shot, chain-of-thought, critique, and expert) in adapting LLMs to code summarization tasks. Contrary to expectations, advanced prompting techniques may not outperform simple zero-shot prompting. Next, we investigate the impact of LLMs' model settings (including top_p and temperature parameters) on the quality of generated summaries. We find the impact of the two parameters on summary quality varies by the base LLM and programming language, but their impacts are similar. Moreover, we canvass LLMs' abilities to summarize code snippets in distinct types of programming languages. The results reveal that LLMs perform suboptimally when summarizing code written in logic programming languages compared to other language types (e.g., procedural and object-oriented programming languages). Finally, we unexpectedly find that CodeLlamaInstruct with 7B parameters can outperform advanced GPT-4 in generating summaries describing code design rationale and asserting code properties. We hope that our findings can provide a comprehensive understanding of code summarization in the era of LLMs. Weisong Sun, Yun Miao, Yuekang Li, Hongyu Zhang 0002, Chunrong Fang, Yi Liu 0069, Gelei Deng, Yang Liu 0003, Zhenyu Chen 0001 |
ICSE | 8 |
| 2025 | LLMs Meet Library Evolution: Evaluating Deprecated API Usage in LLM-Based Code CompletionabstractLarge language models (LLMs), pre-trained or fine-tuned on large code corpora, have shown effectiveness in generating code completions. However, in LLM-based code completion, LLMs may struggle to use correct and up-to-date Application Programming Interfaces (APIs) due to the rapid and continuous evolution of libraries. While existing studies have highlighted issues with predicting incorrect APIs, the specific problem of deprecated API usage in LLM-based code completion has not been thoroughly investigated. To address this gap, we conducted the first evaluation study on deprecated API usage in LLM-based code completion. This study involved seven advanced LLMs, 145 API mappings from eight popular Python libraries, and$\mathbf{2 8, 1 2 5}$completion prompts. The study results reveal the status quo (i.e., API usage plausibility and deprecated usage rate) of deprecated API and replacing API usage in LLM-based code completion from the perspectives of model, prompt, and library, and indicate the root causes behind. Based on these findings, we propose two lightweight fixing approaches, Replaceapi and InsertPrompt, which can serve as baseline approaches for future research on mitigating deprecated API usage in LLM-based completion. Additionally, we provide implications for future research on integrating library evolution with LLMdriven software development. Chong Wang 0013, Kaifeng Huang 0001, Jian Zhang 0087, Yebo Feng, Lyuye Zhang, Yang Liu 0003, Xin Peng 0001 |
ICSE | 6 |
| 2025 | Boosting Static Resource Leak Detection via LLM-based Resource-Oriented Intention InferenceabstractResource leaks, caused by resources not being released after acquisition, often lead to performance issues and system crashes. Existing static detection techniques rely on mechanical matching of predefined resource acquisition/release APIs and null-checking conditions to find unreleased resources, suffering from both (1) false negatives caused by the incompleteness of predefined resource acquisition/release APIs and (2) false positives caused by the incompleteness of resource reachability validation identification. To overcome these challenges, we propose InferROI, a novel approach that leverages the exceptional code comprehension capability of large language models (LLMs) to directly infer resource-oriented intentions (acquisition, release, and reachability validation) in code. InferROI first prompts the LLM to infer involved intentions for a given code snippet, and then incorporates a two-stage static analysis approach to check control-flow paths for resource leak detection based on the inferred intentions. We evaluate the effectiveness of InferROI in both resource-oriented intention inference and resource leak detection. Experimental results on the DroidLeaks and JLeaks datasets demonstrate InferROI achieves promising bug detection rate (59.3% and 62.5%) and false alarm rate (18.6% and 19.5%). Compared to three industrial static detectors, InferROI detects 14~45 and 149~485 more bugs in DroidLeaks and JLeaks, respectively. When applied to real-world open-source projects, InferROI identifies 29 unknown resource leak bugs (verified by authors), with 7 of them being confirmed by developers. In addition, the results of an ablation study underscores the importance of combining LLM-based inference with static analysis. Finally, manual annotation indicated that InferROI achieved a precision of 74.6% and a recall of 81.8% in intention inference, covering more than 60% resource types involved in the datasets. Chong Wang 0013, Xin Peng 0001, Yang Liu 0003, Yiling Lou |
ICSE | 4 |
| 2025 | Understanding the Effectiveness of Coverage Criteria for Large Language Models: A Special Angle from Jailbreak AttacksabstractLarge language models (LLMs) have revolutionized artificial intelligence, but their increasing deployment across critical domains has raised concerns about their abnormal behaviors when faced with malicious attacks. Such vulnerability alerts the widespread inadequacy of pre-release testing. In this paper, we conduct a comprehensive empirical study to evaluate the effectiveness of traditional coverage criteria in identifying such inadequacies, exemplified by the significant security concern of jailbreak attacks. Our study begins with a clustering analysis of the hidden states of LLMs, revealing that the embedded characteristics effectively distinguish between different query types. We then systematically evaluate the performance of these criteria across three key dimensions: criterion level, layer level, and token level. Our research uncovers significant differences in neuron coverage when LLMs process normal versus jailbreak queries, aligning with our clustering experiments. Leveraging these findings, we propose three practical applications of coverage criteria in the context of LLM security testing. Specifically, we develop a realtime jailbreak detection mechanism that achieves high accuracy (93.61 % on average) in classifying queries as normal or jailbreak. Furthermore, we explore the use of coverage levels to prioritize test cases, improving testing efficiency by focusing on high-risk interactions and removing redundant tests. Lastly, we introduce a coverage-guided approach for generating jailbreak attack examples, enabling systematic refinement of prompts to uncover vulnerabilities. This study improves our understanding of LLM security testing, enhances their safety, and provides a foundation for developing more robust AI applications. Shide Zhou, Tianlin Li, Kailong Wang 0001, Yihao Huang 0001, Ling Shi 0002, Yang Liu 0003, Haoyu Wang 0001 |
ICSE | 6 |
| 2025 | PHAnToM: Persona-Based Prompting Has an Effect on Theory-of-Mind Reasoning in Large Language ModelsabstractThe use of LLMs in natural language reasoning has shown mixed results, sometimes rivaling or even surpassing human performance in simpler classification tasks while struggling with social-cognitive reasoning, a domain where humans naturally excel. These differences have been attributed to many factors, such as variations in prompting and the specific LLMs used. However, no reasons appear conclusive, and no clear mechanisms have been established in prior work. In this study, we empirically evaluate how role-playing persona-based prompting influences Theory-of-Mind (ToM) reasoning capabilities. Grounding our research in psychological theory, we found that, beyond the inherent variance in the complexity of reasoning tasks, ToM performance differences arise because of socially-motivated prompting differences. In an era where prompt engineering with role-play is a typical approach to adapt LLMs to new contexts, our research advocates caution as models that adopt specific personas might potentially result in errors in social-cognitive reasoning. Gerard Yeo, Fiona Anting Tan, Kokil Jaidka, Shaz Furniturewala, Fanyou Wu, Weijie Xu, Vinija Jain, Aman Chadha, Yang Liu 0003, See-Kiong Ng |
ICWSM | 9 |
| 2025 | Measuring and Explaining the Effects of Android App Transformations in Online Malware DetectionabstractIt is well known that antivirus engines are vulnerable to evasion techniques (e.g., obfuscation) that transform malware into its variants.However, it cannot be necessarily attributed to the effectiveness of these evasions, and the limits of engines may also make this unsatisfactory result.In this study, we propose a data-driven approach to measure the effect of app transformations to malware detection, and further explain why the detection result is produced by these engines.First, we develop an interaction model for antivirus engines, illustrating how they respond with different detection results in terms of varying inputs.Six app transformation techniques are implemented in order to generate a large number of Android apps with traceable changes.Then we undertake a onemonth tracking of app detection results from multiple antivirus engines, through which we obtain over 971K detection reports from VirusTotal for 179K apps in total.Last, we conduct a comprehensive analysis of antivirus engines based on these reports from the perspectives of signature-based, static analysis-based, and dynamic analysis-based detection techniques.The results, together with 7 highlighted findings, identify a number of sealed working mechanisms occurring inside antivirus engines and what are the indicators of compromise in apps during malware detection. Guozhu Meng, Zhixiu Guo, Xiaodong Zhang 0014, Haoyu Wang 0001, Kai Chen 0012, Yang Liu 0003 |
Internetware | 6 |
| 2025 | Condition Sequence Coverage Criterion and Automatic Test Case Generation for Testing-Based Formal VerificationabstractTesting-based formal verification (TBFV) is proposed to reduce test cost and guarantee software reliability by ensuring the correctness of all traversed program paths. An ideal target is to generate adequate test cases to traverse all of its execution paths. However, it is a rather ambitious criterion that can hardly be satisfied due to the potentially great amount of test cases required. To address this problem, we propose a new criterion called Condition Sequence Coverage (CSC) to maintain a good balance between program correctness and the number of test cases. In this paper, we refine the TBFV method and introduce the theoretical foundations of CSC. We also integrate CSC with functional scenario form (FSF) to automatically generate test cases for the TBFV method. In addition, we develop the tool support for Java and validate its effectiveness and accuracy through experimental comparisons. Ai Liu, Yang Liu 0003, Lei Rao, Shaoying Liu, Zhibin Yang 0005 |
ISSRE | 2 |
| 2025 | Envisioning Intelligent Requirements Engineering via Knowledge-Guided Multi-Agent CollaborationabstractRequirements Engineering (RE) is an initial and critical phase in software development, with the aim of producing well-defined software requirements specifications (SRSs) from rough ideas of clients. It involves multiple tasks (e.g., elicitation, analysis) and roles (e.g., interviewer, analyst). With the rise of Large Language Models (LLMs), many studies have leveraged LLMs to support specific RE tasks. However, existing LLM-based agents often lack domain knowledge integration and fall short in simulating the complex collaboration of human experts across the full RE process. To address this gap, we propose KGMAF, a knowledge-guided multi-agent framework designed to assist requirements engineers in developing high-quality SRSs. KGMAF comprises six LLM-based agents and a shared artifact pool. Each agent is equipped with predefined actions, dedicated functions, and injected knowledge tailored to specific RE tasks. The artifact pool stores both intermediate and final artifacts, serving as a communication channel for inter-agent collaboration. A human-in-the-loop (HITL) mechanism is embedded to guide and validate agent outputs. We present the design of KGMAF, along with preliminary experiments and a case study to demonstrate its practicality. This work lays the foundation for future research on knowledge-driven multi-agent collaboration in RE and highlights key challenges in building trustworthy intelligent assistants for real-world RE tasks. Jiangping Huang, Dongming Jin, Weisong Sun, Yang Liu 0003, Zhi Jin 0001 |
ASE | 4 |
| 2025 | Evaluating Large Language Models for Time Series Anomaly Detection in Aerospace SoftwareabstractTime series anomaly detection (TSAD) is essential for ensuring the safety and reliability of aerospace software systems. Although large language models (LLMs) provide a promising training-free alternative to unsupervised approaches, their effectiveness in aerospace settings remains under-examined because of complex telemetry, misaligned evaluation metrics, and the absence of domain knowledge. To address this gap, we introduce ATSADBench, the first benchmark for aerospace TSAD. ATSADBench comprises nine tasks that combine three pattern-wise anomaly types, univariate and multivariate signals, and both in-loop and out-of-loop feedback scenarios, yielding 108,000 data points. Using this benchmark, we systematically evaluate state-of-the-art open-source LLMs under two paradigms: Direct, which labels anomalies within sliding windows, and Prediction-Based, which detects anomalies from prediction errors. To reflect operational needs, we reformulate evaluation at the window level and propose three user-oriented metrics: Alarm Accuracy (AA), Alarm Latency (AL), and Alarm Contiguity (AC), which quantify alarm correctness, timeliness, and credibility. We further examine two enhancement strategies, few-shot learning and retrieval-augmented generation (RAG), to inject domain knowledge. The evaluation results show that (1) LLMs perform well on univariate tasks but struggle with multivariate telemetry, (2) their AA and AC on multivariate tasks approach random guessing, (3) few-shot learning provides modest gains whereas RAG offers no significant improvement, and (4) in practice LLMs can detect true anomaly onsets yet sometimes raise false alarms, which few-shot prompting mitigates but RAG exacerbates. These findings offer guidance for future LLM-based TSAD in aerospace software. Yang Liu 0003, Yixing Luo, Xiaofeng Li 0005, Bin Gu 0006, Zhi Jin 0001 |
ASE | 1 |
| 2025 | Have We Solved Access Control Vulnerability Detection in Smart Contracts? A Benchmark StudyabstractAccess control (AC) vulnerabilities are among the most critical security threats to smart contracts. Despite extensive research, they remain widespread and damaging in the Ethereum ecosystem. To understand and advance the current state-of-the-art (SOTA) in AC vulnerability detection, we first curate a diverse dataset of 180 real-world AC vulnerabilities from CVE entries, DeFiHackLabs incidents, and Code4rena audit reports.Using this dataset, we conduct a systematic benchmark study along three dimensions. First, we develop a cause-based taxonomy and analyze the prevalence and evolution of AC vulnerabilities. Second, we evaluate six SOTA tools, including two from industry and four from academia, revealing low recall (3% to 8%) and significant blind spots. To understand these failures, we examine 1.2 million deployed contracts and uncover practical gaps in AC protection mechanisms overlooked by existing tools. Finally, we assess the potential of large language models (LLMs) for AC vulnerability detection and show that LLMs detect 53–75% of vulnerabilities, outperforming traditional tools but facing challenges such as hallucinations and scalability. Our findings highlight the need for hybrid approaches that combine static analysis with LLM-based semantic reasoning to address the complexity of modern AC vulnerabilities. Han Liu 0012, Daoyuan Wu, Yuqiang Sun 0001, Shuai Wang 0011, Yang Liu 0003 |
ASE | 5 |
| 2025 | Requirements Development and Formalization for Reliable Code Generation: A Multi-Agent VisionabstractAutomated code generation has long been considered the holy grail of software engineering. The emergence of Large Language Models (LLMs) has catalyzed a revolutionary breakthrough in this area. However, existing methods that only rely on LLMs remain inadequate in the quality of generated code, offering no guarantees of satisfying practical requirements. They lack a systematic strategy for requirements development and modeling. Recently, LLM-based agents typically possess powerful abilities and play an essential role in facilitating the alignment of LLM outputs with user requirements. In this paper, we envision the first multi-agent framework for reliable code generation based on Requirements Development and Formalization, named ReDeFo. This framework incorporates three agents, highlighting their augmentation with knowledge and techniques of formal methods, into the requirements-to-code generation pipeline to strengthen quality assurance. The core of ReDeFo is the use of formal specifications to bridge the gap between potentially ambiguous natural language requirements and precise executable code. ReDeFo enables rigorous reasoning about correctness, uncovering hidden bugs, and enforcing critical properties throughout the development process. Xu Lu 0003, Weisong Sun, Ming Hu 0003, Cong Tian 0001, Zhi Jin 0001, Yang Liu 0003 |
ASE | 7 |
| 2025 | A Large Scale Study of AI-based Binary Function Similarity Detection Techniques for Security Researchers and PractitionersabstractBinary Function Similarity Detection (BFSD) is a foundational technique in software security, underpinning a wide range of applications including vulnerability detection, malware analysis. Recent advances in AI-based BFSD tools have led to significant performance improvements. However, existing evaluations of these tools suffer from three key limitations: a lack of in-depth analysis of performance-influencing factors, an absence of realistic application analysis, and reliance on small-scale or low-quality datasets.In this paper, we present the first large-scale empirical study of AI-based BFSD tools to address these gaps. We construct two high-quality and diverse datasets: BinAtlas, comprising 12,453 binaries and over 7 million functions for capability evaluation; and BinAres, containing 12,291 binaries and 54 real-world 1-day vulnerabilities for evaluating vulnerability detection performance in practical IoT firmware settings. Using these datasets, we evaluate nine representative BFSD tools, analyze the challenges and limitations of existing BFSD tools, and investigate the consistency among BFSD tools. We also propose an actionable strategy for combining BFSD tools to enhance overall performance (an improvement of 13.4%). Our study not only advances the practical adoption of BFSD tools but also provides valuable resources and insights to guide future research in scalable and automated binary similarity detection. Yang Xiao 0011, Yuekang Li, Zhengzi Xu, Sihao Qiu, Keyu Qi, Yeting Li, Xingchu Chen, Yanyan Zou 0002, Yang Liu 0003, Wei Huo 0005 |
ASE | 12 |
| 2025 | Learning from the Past: Real-World Exploit Migration for Smart Contract PoC GenerationabstractSmart contract vulnerabilities continue to cause significant financial losses, despite the implementation of security measures such as manual audits and bug bounty platforms. A critical component often required by these security measures is the proof-of-concept (PoC) exploit, which validates vulnerability exploitability, assesses impact severity, and guides developers in fixes. Existing tools have explored automated PoC generation with techniques like symbolic execution, fuzzing, and program synthesis. However, these approaches frequently fail to generate PoCs for vulnerabilities exploited in real-world incidents, primarily due to their limitations in handling complex transaction dependencies, navigating vast on-chain state spaces, or requiring extensive manual specifications. Our migration-based approach extracts critical information from documented security incidents and applies it to generate PoCs for similar vulnerable code. This approach leverages proven exploit patterns rather than generating PoCs from scratch. This approach is motivated by two key observations: the prevalence of code reuse in smart contracts (up to 90% at the function level) and the increasing availability of documented PoCs for real-world incidents. Our approach operates in three phases: (1) abstracting essential components (i.e., environment properties, attack logic, and verification checks) from existing PoCs into templates, (2) given a new target contract, selecting suitable templates with adapted values through clone-detection and property-feasibility analysis, and (3) generating and validating PoCs in simulated environments. Our evaluation demonstrates effectiveness and efficiency across multiple scales. Our approach successfully generates valid PoCs for 62 out of 67 manually validated cases without false positives and completes analysis in 3.8 hours compared to 133.2 and 210.5 hours required by existing tools. Large-scale evaluation on 979,512 contracts identifies 256 vulnerable contracts across blockchain networks with 64 cross-chain cases, demonstrating real-world applicability. Kairan Sun, Zhengzi Xu, Kaixuan Li 0002, Lyuye Zhang, Yebo Feng, Daoyuan Wu, Yang Liu 0003 |
ASE | 7 |
| 2025 | Faultseeker: LLM-Empowered Framework for Blockchain Transaction Fault LocalizationabstractWeb3 applications, particularly decentralized finance (DeFi) protocols, have grown rapidly with over $100 billion locked in smart contracts, attracting sophisticated attacks causing billions in losses. When attack occur, security analysts need to perform fault localization to identify vulnerable functions and understand attack vectors. This critical process currently requires an average of 16.7 analyst hours per incident due to complex blockchain execution models, rapidly evolving protocol interactions, and multi-contract attack patterns that exceed existing analytical capabilities. Despite its critical importance, blockchain fault localization has received limited attention due to fundamental challenges requiring semantic understanding of economic models and protocol-specific logic. Existing blockchain-specific tools target only single vulnerability types, while the only comprehensive solution, DAppFL, relies on machine learning model that may miss sophisticated exploits and lacks interpretability in results. Recent advances in large language models (LLMs) demonstrate remarkable code comprehension capabilities, but existing applications focus on proactive vulnerability detection with minimal exploration of post-incident fault localization.We present FaultSeeker, an LLM-empowered framework for blockchain transaction fault localization. Our two-stage architecture combines transaction-level forensics for strategic scoping with coordinated specialist agents for sustained reasoning. This design provides long-term memory management via orchestrator agents and specialized attention allocation through coordinated workers, enabling comprehensive analysis across complex multi-contract transactions without context loss. We evaluate Fault-Seeker on a compiled dataset of 115 real-world malicious transactions with expert-validated annotations spanning diverse attack patterns and complexity levels. Results demonstrate that FaultSeeker significantly outperforms existing approaches, including DAppFL and leading native LLMs (GPT-4o, Claude 3.7 Sonnet, DeepSeek R1), while maintaining practical efficiency (4.4- 8.6 minutes) and cost-effectiveness ($1.55-$4.53 per transaction). Kairan Sun, Zhengzi Xu, Kaixuan Li 0002, Lyuye Zhang, Yuqiang Sun 0001, Liwei Tan, Yang Liu 0003 |
ASE | 7 |
| 2025 | BinStruct: Binary Structure Recovery Combining Static Analysis and SemanticsabstractBinary reverse engineering is foundational to various tasks such as malware analysis and vulnerability detection. Traditional binary analysis tools mainly operate at the function level. However, modern software has grown significantly in size, with binaries often containing thousands of functions. Without understanding how these functions are organized into higher-level structures, it becomes difficult to effectively support downstream analysis tasks. Analysts must examine thousands of functions separately, making the process time-consuming and error-prone. Despite these challenges, current research on recovering the higher-level structure of binaries remains limited.To bridge this gap, we propose BinStruct, a novel binary structure recovery framework that recovers both file and module structures from binaries. BinStruct first identifies the file structure by combining data reference patterns, function calls, and semantic understanding from Large Language Models. Then, inspired by software architecture recovery in source code analysis, BinStruct identifies modules by clustering the recovered files using consensus between structural dependency and semantic similarity. Evaluation on 121 real-world stripped binaries demonstrates that BinStruct outperforms state-of-the-art techniques in both file and module recovery accuracy, while requiring only 7.42s and 34.46s on average to recover file and module structures, respectively. Case studies on Libxml2 and PredatorTheStealer demonstrate BinStruct’s effectiveness on security tasks like attack surface analysis and malware investigation. Zhengzi Xu, Zhe Lang, Chengyue Liu, Yuqiang Sun 0001, Wenbo Guo 0011, Weisong Sun, Yang Liu 0003 |
ASE | 9 |
| 2025 | Detecting Various DeFi Price Manipulations with LLM ReasoningabstractDeFi (Decentralized Finance) is one of the most important applications of today’s cryptocurrencies and smart contracts. It manages hundreds of billions in Total Value Locked (TVL) on-chain, yet it remains susceptible to common DeFi price manipulation attacks. Despite state-of-the-art (SOTA) systems like DeFiRanger and DeFort, we found that they are less effective to non-standard price models in custom DeFi protocols, which account for 44.2% of the 95 DeFi price manipulation attacks reported over the past three years.In this paper, we introduce the first LLM-based approach, DeFiScope, for detecting DeFi price manipulation attacks in both standard and custom price models. Our insight is that large language models (LLMs) have certain intelligence to abstract price calculation from smart contract source code and infer the trend of token price changes based on the extracted price models. To further strengthen LLMs in this aspect, we leverage Foundry to synthesize on-chain data and use it to fine-tune a DeFi price-specific LLM. Together with the high-level DeFi operations recovered from low-level transaction data, DeFiScope detects various DeFi price manipulations according to systematically mined patterns. Experimental results show that DeFiScope achieves a high recall of 80% on real-world attacks, a precision of 96% on suspicious transactions, and zero false alarms on benign transactions, significantly outperforming SOTA approaches. Moreover, we evaluate DeFiScope’s cost-effectiveness and demonstrate its practicality by helping our industry partner confirm 147 real-world price manipulation attacks, including discovering 81 previously unknown historical incidents. Juantao Zhong, Daoyuan Wu, Ye Liu 0012, Maoyi Xie, Yang Liu 0003, Yi Li 0008 |
ASE | 5 |
| 2025 | PropertyGPT: LLM-driven Formal Verification of Smart Contracts through Retrieval-Augmented Property Generation
Ye Liu 0012, Yue Xue, Daoyuan Wu, Yuqiang Sun 0001, Yi Li 0008, Miaolei Shi, Yang Liu 0003 |
NDSS | 7 |
| 2025 | SongBsAb: A Dual Prevention Approach against Singing Voice Conversion based Illegal Song Covers
Guangke Chen, Yedi Zhang, Fu Song, Ting Wang 0004, Xiaoning Du 0001, Yang Liu 0003 |
NDSS | 6 |
| 2025 | Adversarial Attacks against Closed-Source MLLMs via Feature Optimal AlignmentabstractMultimodal large language models (MLLMs) remain vulnerable to transferable adversarial examples. While existing methods typically achieve targeted attacks by aligning global features—such as CLIP’s [CLS] token—between adversarial and target samples, they often overlook the rich local information encoded in patch tokens. This leads to suboptimal alignment and limited transferability, particularly for closed-source models. To address this limitation, we propose a targeted transferable adversarial attack method based on feature optimal alignment, called FOA-Attack, to improve adversarial transfer capability. Specifically, at the global level, we introduce a global feature loss based on cosine similarity to align the coarse-grained features of adversarial samples with those of target samples. At the local level, given the rich local representations within Transformers, we leverage clustering techniques to extract compact local patterns to alleviate redundant local features. We then formulate local feature alignment between adversarial and target samples as an optimal transport (OT) problem and propose a local clustering optimal transport loss to refine fine-grained feature alignment. Additionally, we propose a dynamic ensemble model weighting strategy to adaptively balance the influence of multiple models during adversarial example generation, thereby further improving transferability. Extensive experiments across various models demonstrate the superiority of the proposed method, outperforming state-of-the-art methods, especially in transferring to closed-source MLLMs. Xiaojun Jia, Sensen Gao, Simeng Qin, Tianyu Pang, Yihao Huang 0001, Xinfeng Li, Yiming Li 0004, Bo Li 0026, Yang Liu 0003 |
NeurIPS | 10 |
| 2025 | INST-IT: Boosting Instance Understanding via Explicit Visual Prompt Instruction TuningabstractLarge Multimodal Models (LMMs) have made significant breakthroughs with the advancement of instruction tuning. However, while existing models can understand images and videos at a holistic level, they still struggle with instance-level understanding that requires a more fine-grained comprehension and alignment. Instance-level understanding is crucial for LMMs, as it focuses on the specific elements that we are most interested in. Excitingly, existing works find that the state-of-the-art LMMs exhibit strong instance understanding capabilities when provided with explicit visual cues. Motivated by this, we proposed Inst-IT, a solution to enhance LMMs in Instance understanding via explicit visual prompt Instruction Tuning for instance guidance. Inst-IT consists of a benchmark to diagnose multimodal instance-level understanding, a large-scale instruction-tuning dataset, and a continuous instruction-tuning training paradigm to effectively enhance spatial-temporal instance understanding capabilities of existing LMMs. Experimental results show that, enhanced by Inst-IT, our models not only achieve outstanding performance on Inst-IT-Bench and other instance understanding benchmarks, but also demonstrate significant improvements across various generic image and video understanding benchmarks. This highlights that our method not only boosts instance-level understanding but also strengthens the overall capabilities of generic image and video comprehension. Wujian Peng, Lingchen Meng, Yiweng Xie, Yang Liu 0003, Tao Gui, Hang Xu 0004, Xipeng Qiu, Zuxuan Wu, Yu-Gang Jiang 0001 |
NeurIPS | 5 |
| 2025 | DepthVanish: Optimizing Adversarial Interval Structures for Stereo-Depth-Invisible PatchesabstractStereo depth estimation is a critical task in autonomous driving and robotics, where inaccuracies (such as misidentifying nearby objects as distant) can lead to dangerous situations. Adversarial attacks against stereo depth estimation can help revealing vulnerabilities before deployment. Previous works have shown that repeating optimized textures can effectively mislead stereo depth estimation in digital settings. However, our research reveals that these naively repeated textures perform poorly in physical implementations, $\textit{i.e.}$, when deployed as patches, limiting their practical utility for stress-testing stereo depth estimation systems. In this work, for the first time, we discover that introducing regular intervals among the repeated textures, creating a grid structure, significantly enhances the patch attack performance. Through extensive experimentation, we analyze how variations of this novel structure influence the adversarial effectiveness. Based on these insights, we develop a novel stereo depth attack that jointly optimizes both the interval structure and texture elements. Our generated adversarial patches can be inserted into any scenes and successfully attack advanced stereo depth estimation methods of different paradigms, $\textit{i.e.}$, RAFT-Stereo and STTR. Most critically, our patch can also attack commercial RGB-D cameras (Intel RealSense) in real-world conditions, demonstrating their practical relevance for security assessment of stereo systems. The code is officially released at: https://github.com/WiWiN42/DepthVanish Yun Xing 0001, Nhat Chung, Jie Zhang 0050, Ivor W. Tsang, Ming-Ming Cheng, Yang Liu 0003, Lei Ma 0003, Qing Guo 0005 |
NeurIPS | 7 |
| 2025 | CPSea: Large-scale cyclic peptide-protein complex dataset for machine learning in cyclic peptide designabstractCyclic peptides exhibit better binding affinity and proteolytic stability compared to their linear counterparts. However, the development of cyclic peptide design models is hindered by the scarcity of data. To address this, we introduce **CPSea**(**C**yclic **P**eptide **Sea**), a dataset of 2.71 million cyclic peptide-receptor complexes, curated through systematic mining of the AlphaFold Database (AFDB). Our pipeline extracts compact domains from AFDB, identifies cyclization sites using the $\beta$-carbon (C$_\beta$) distance thresholds, and applies multi-stage filtering to ensure structure fidelity and binding compatibility. Compared with experimental data of cyclic peptides, CPSea shows similar distributions in metrics on structure fidelity and wet-lab compatibility. To our knowledge, CPSea is the largest cyclic peptide-receptor dataset to date, enabling end-to-end model training for the first time. The dataset also showcases the feasibility of simulating inter-chain interactions using intra-chain interactions, expanding available resources for machine-learning models on protein-protein interactions. The dataset and relevant scripts are accessible on GitHub ([https://github.com/YZY010418/CPSea](https://github.com/YZY010418/CPSea)). Ziyi Yang 0011, Hanyuan Xie, Yinjun Jia, Xiangzhe Kong, Jiqing Zheng, Ziting Zhang, Yang Liu 0003, Lei Liu 0049, Yanyan Lan |
NeurIPS | 7 |
| 2025 | STGraph: Spatio-Temporal Graph Mining for Anomaly Detection in Distributed System LogsabstractSystem logs are crucial sources of information for engineers to analyze and resolve anomalies and faults in large-scale software systems. However, logs on a distributed system are often fragmented, making it challenging to achieve unified processing and comprehension. Traditional methods for log-based anomaly detection often employ machine learning algorithms with a focus on log event counts or log sequences. However, traditional methods fall short of fully leveraging the temporal and spatial structures inherent in distributed system logs, leading to issues of false positives and unstable performance in anomaly detection. In this paper, we propose a novel log anomaly detection method based on the construction of distributed system workflow graphs. This method extracts spatio-temporal information from distributed system logs and constructs event workflow graphs. These graphs accurately reflect the execution of the system and provide more comprehensive support for anomaly detection based on distributed system logs. The experimental results demonstrated that STGraph achieved F1 scores of 0.959,0.979, and 0.959 on HDFS, BGL, and OpenStack datasets respectively, outperforming LogRobust, PLELog, and NeuralLog by 1.2%-18.6% across precision/recall metrics. Notably, it attained 0.985 recall on BGL and maintained >0.935 F1 scores under 30% noise interference, 21.8% higher than LogRobust. Teng Li 0003, Shengkai Zhang, Yebo Feng, Jiahua Xu 0002, Zexu Dang, Yang Liu 0003, Jianfeng Ma 0001 |
RAID | 6 |
| 2025 | Taxonomy-Guided Reasoning for Requirements Classification: A Study in Aerospace IndustryabstractRequirements classification, which organizes software requirements into structured categories, is crucial in safety-critical domains such as aerospace. However, practical implementation is challenging due to the absence of unified, domain-specific taxonomies, as different developers often adopt divergent classification schemes. Moreover, safety-critical requirements frequently intertwine functional and reliability constraints, creating complex multi-label classification challenges. Existing supervised learning approaches depend on large annotated datasets, which are rarely feasible in specialized industries, while current LLM-based methods face difficulties handling hierarchical, multi-label scenarios effectively. To address these issues, we propose TRClass, a novel taxonomy-guided classification approach. The key idea behind TRClass is to integrate domain knowledge into the classification process by first constructing a unified taxonomy semi-automatically, extracting structure from existing documents, and refining it with expert validation. TRClass then guides an LLM to classify requirements by reasoning step-by-step through the taxonomy hierarchy, using few-shot retrieval and confidence-based exploration to achieve accurate multi-label decisions. We validate TRClass using aerospace software requirements as a representative case study for safety-critical industries. Results show that TRClass consistently outperforms baselines, with all components contributing to its overall effectiveness, and remains robust across different LLM configurations. A user study further confirms its practical usability in real-world industrial scenarios. Yixing Luo, Yang Liu 0003, Xiaofeng Li 0005, Bin Gu 0006, Zhi Jin 0001, Mengfei Yang |
RE | 2 |
| 2025 | Testing-Based Formal Verification with Program Slicing on Functional Soundness and Completeness
Ai Liu, Yang Liu 0003, Shaoying Liu, Zhibin Yang 0005 |
TASE | 2 |
| 2025 | SelfDefend: LLMs Can Defend Themselves against Jailbreaking in a Practical Manner
Xunguang Wang, Daoyuan Wu, Zhenlan Ji, Zongjie Li, Pingchuan Ma 0004, Shuai Wang 0011, Yingjiu Li, Yang Liu 0003, Juergen Rahmel |
USENIX Security Symposium | 8 |
| 2025 | Fighting Fire with Fire: Continuous Attack for Adversarial Android Malware Detection
Yinyuan Zhang, Cuiying Gao, Yueming Wu 0001, Shihan Dou, Cong Wu 0003, Ying Zhang 0066, Wei Yuan 0001, Yang Liu 0003 |
USENIX Security Symposium | 8 |
| 2025 | A Role-based Hierarchical Adaptive Routing Protocol for Clustered UAV SwarmsabstractFlying Ad Hoc Networks (FANETs) are self-organizing wireless networks composed of Unmanned Aerial Vehicles (UAVs), designed to enable collaborative tasks without relying on fixed infrastructure. Characterized by high mobility, dynamic topology, and three-dimensional spatial movement, FANETs pose critical challenges to routing protocols, particularly to the broadcast storm problem caused by uncontrolled flooding in large-scale deployments. To address the challenges of broadcast storms in FANETs, this paper proposes RHARP, a role-based hierarchical routing protocol for FANETs. RHARP innovatively integrates bio-inspired swarm communication principles with role-based network clustering, establishing functional correspondence between decentralized coordination in biological swarms and differentiated node roles in FANETs. Through functional decoupling, RHARP delegates intra-cluster routing decisions to gateway nodes to alleviate cluster head (CH) workload. It achieves dynamic inter-cluster communication optimization through hierarchical routing strategies and on-demand path discovery mechanisms. By systematically integrating role-based clustering with adaptive routing, RHARP provides a scalable and cost-effective communication framework for large-scale FANET deployments while maintaining protocol agility. The simulation results demonstrate the effectiveness of RHARP across various dynamic scenarios. Tianen Guan, Xinliang Wu, Zhongliang Zhao, Wenbing Tang 0001, Yang Liu 0003 |
VTC2025-Fall | 6 |
| 2025 | Multi-UAV Trajectory Generation for Fresh Data Collection: A Diffusion-based Reinforcement Learning ApproachabstractThis paper investigates the trajectory generation problem for multi-unmanned aerial vehicle (UAV)-enabled uplink data collection. Specifically, we minimize the age-of-information (AoI) and maximize the coverage as well as the amount of collected data by planning the multi- UAV trajectory considering the energy consumption and collisions constraints. Motivated by diffusion models' exceptional generative capabilities, we propose a multi-UAV trajectory generation (MUTG) solution based on soft actor-critic and diffusion to solve the optimization problem. A diffusion model-based predictor is designed to obtain the action policy, where a hierarchical graph-transformer network is developed to extract entities' interactive information as a conditional guide for the diffusion. Numerical results verify the effectiveness and superiority compared with benchmark schemes in terms of average AoI, user coverage and data collection ratio. Ziping Yu, Meng Xiao 0002, Zhongliang Zhao, Xianbin Cao 0001, Yang Liu 0003, Tony Q. S. Quek |
WCNC | 6 |
| 2025 | Enmob: Unveil the Behavior with Multi-flow Analysis of Encrypted App TrafficabstractAbstract In the contemporary digital landscape, mobile applications have become the predominant conduit for internet connectivity and daily tasks. Simultaneously, the advent of application encryption technology has safeguarded users’ privacy. However, this encryption, while fortifying privacy, introduces challenges to security by hindering the effective management of network applications within encrypted data streams. Conventional detection methods for encrypted application traffic, relying heavily on statistical metrics like payload, packet size, and distribution, are constrained to single traffic flows, often yielding results of limited specificity. To address this limitation, our paper introduces an innovative approach that elucidates the multi-flow nature of application behavior traffic and provides context to encrypted application traffic. This method offers a more nuanced and comprehensive perspective for understanding and representing network traffic, even when encrypted. The efficacy of our approach was evaluated using a substantial volume of real network traffic data. Results indicate that our method achieves an average accuracy of 0.958 in identifying application behavior traffic and 0.955 in classifying application traffic. These outcomes signify a substantial enhancement over single network flow-based detection methods, demonstrating a notable 5.3% improvement. Mengmeng Ge 0003, Likun Liu, Xiangzhan Yu, Vinay Sachidananda, Xiaofei Xie, Yang Liu 0003 |
Cybersecur. | 7 |
| 2025 | Log2Evt: Constructing high-level events for IoT Systems through log-code execution path correlation
Teng Li 0003, Baichuan Zheng, Yebo Feng, Xiaowen Quan, Jiahua Xu 0002, Yang Liu 0003, Jianfeng Ma 0001 |
J. Syst. Archit. | 6 |
| 2025 | Generically Automating Separation Logic by Functors, Homomorphisms, and ModulesabstractFoundational verification considers the functional correctness of programming languages with formalized semantics and uses proof assistants (e.g., Coq, Isabelle) to certify proofs. The need for verifying complex programs compels it to involve expressive Separation Logics (SLs) that exceed the scopes of well-studied automated proof theories, e.g., symbolic heap. Consequently, automation of SL in foundational verification relies heavily on ad-hoc heuristics that lack a systematic meta-theory and face scalability issues. To mitigate the gap, we propose a theory to specify SL predicates using abstract algebras including functors, homomorphisms, and modules over rings. Based on this theory, we develop a generic SL automation algorithm to reason about any data structures that can be characterized by these algebras. In addition, we also present algorithms for automatically instantiating the algebraic models to real data structures. The instantiation works compositionally, reusing the algebraic models of component structures and preserving their data abstractions. Case studies on formalized imperative semantics show our algorithm can instantiate the algebraic models automatically for a variety of complex data structures. Experimental results indicate the automatically instantiated reasoners from our generic theory show similar results to the state-of-the-art systems made of specifically crafted reasoning rules. The presented theories, proofs, and the verification framework are formalized in Isabelle/HOL. Qiyuan Xu, David Sanán, Xiaokun Luan, Conrad Watt, Yang Liu 0003 |
Proc. ACM Program. Lang. | 6 |
| 2025 | Semantic-Aligned Adversarial Evolution Triangle for High-Transferability Vision-Language AttackabstractVision-language pre-training (VLP) models excel at interpreting both images and text but remain vulnerable to multimodal adversarial examples (AEs). Advancing the generation of transferable AEs, which succeed across unseen models, is key to developing more robust and practical VLP models. Previous approaches augment image-text pairs to enhance diversity within the adversarial example generation process, aiming to improve transferability by expanding the contrast space of image-text features. However, these methods focus solely on diversity around the current AEs, yielding limited gains in transferability. To address this issue, we propose to increase the diversity of AEs by leveraging the intersection regions along the adversarial trajectory during optimization. Specifically, we propose sampling from adversarial evolution triangles composed of clean, historical, and current adversarial examples to enhance adversarial diversity. We provide a theoretical analysis to demonstrate the effectiveness of the proposed adversarial evolution triangle. Moreover, we find that redundant inactive dimensions can dominate similarity calculations, distorting feature matching and making AEs model-dependent with reduced transferability. Hence, we propose to generate AEs in the semantic image-text feature contrast space, which can project the original feature space into a semantic corpus subspace. The proposed semantic-aligned subspace can reduce the image feature redundancy, thereby improving adversarial transferability. Extensive experiments across different datasets and models demonstrate that the proposed method can effectively improve adversarial transferability and outperform state-of-the-art adversarial attack methods. Xiaojun Jia, Sensen Gao, Qing Guo 0005, Simeng Qin, Ke Ma 0001, Yihao Huang 0001, Yang Liu 0003, Ivor W. Tsang, Xiaochun Cao |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2025 | EMOR: Energy-Efficient Mixture Opportunistic Routing Based on Reinforcement Learning for Lunar Surface Ad-Hoc NetworksabstractThe lunar surface ad-hoc network is a critical component of the international lunar research station and an extension of the earth-moon communication networks. Its high reliability and low delay are essential for ensuring the safety of the lunar station and improving the efficiency of node collaboration. However, due to the lack of large-scale grid infrastructures, the network must operate autonomously for long periods under strong energy constraints. We propose EMOR, a cross-layer routing protocol, which aims to achieve sustainable high reliability and low latency while balancing energy recovery and consumption. EMOR improves reliability through the “parallel” forwarding feature of opportunistic routing and reduces delay through a mixture of table-based and timer-based routing mechanisms. Moreover, EMOR uses reinforcement learning to analyze the environment and calculate the weights of energy and progress to guide the emphasis on multi-metrics routing. To balance energy consumption and recovery, EMOR introduces a dynamic duty cycle in the MAC layer. Compared to table-based routing and the latest opportunistic routing, EMOR maintains the optimal end-to-end delay in the order of 1ms while improving the packet delivery ratio 6% to 21% higher than other protocols. Moreover, the network lifetime using EMOR is extended by 75.5% to 242%. Zhiyuan Qu, Zhongliang Zhao, Xianbin Cao 0001, Yang Liu 0003, Tony Q. S. Quek |
IEEE Trans. Commun. | 5 |
| 2025 | MalScan: Android Malware Detection Based on Social-Network Centrality AnalysisabstractMalware scanning of an app market is expected to be scalable and effective. However, existing approaches use syntax-based features that can be evaded by transformation attacks or semantic-based features which are usually extracted by expensive program analysis. Therefore, to address the scalability challenges of traditional heavyweight static analysis, we propose a graph-based lightweight approachMalScanfor Android malware detection.MalScanconsiders the function call graph as a complex social network and employs centrality analysis on sensitiveapplication program interfaces(APIs) to express the semantic characteristics of the graph. On this basis, machine learning algorithms and ensemble learning algorithms are applied to classify the extracted features. We evaluateMalScanon datasets of 104,892 benign apps and 108,640 malwares, and the results of experiments indicate thatMalScanoutperforms six state-of-the-art detectors and can quickly detect Android malware with an f-value as high as 99%. In addition, there are also significant improvements in the robustness of Android app evolution and robustness to obfuscation. Finally, we conduct an exhaustive statistical study of over one million applications in the Google-Play app market and successfully identify 498 zero-day malware, which further validates the feasibility ofMalScanon market-wide malware scanning. Yueming Wu 0001, Wenqi Suo, Siyue Feng, Deqing Zou, Wei Yang 0013, Yang Liu 0003, Hai Jin 0001 |
IEEE Trans. Dependable Secur. Comput. | 6 |
| 2025 | Generalized Local Optimality for Video Steganalysis in Motion Vector DomainabstractVideo steganography that conceals secret data into motion vectors (MVs) is a popular covert communication technique. The local optimality of MVs is an intrinsic property in video coding, and any modifications to the MVs will inevitably destroy this optimality, making it a sensitive indicator of steganography. Thus the local optimality is commonly used to design features in video steganalysis. However, the local optimality in existing works is often estimated inaccurately or by using an unreasonable assumption, limiting its capability in steganalysis. In this article, we propose to estimate the local optimality in a more reasonable and comprehensive fashion, and generalize the local optimality in two aspects. First, we generalize the local optimality from a static estimation to a dynamic one by considering the variability of predicted motion vectors (PMVs). Second, we generalize the local optimality from MV domain to PMV domain by leveraging the statistical anomaly of PMVs. Based on the two generalizations that ensure a more accurate estimation of local optimality from more views, we construct new types of steganalytic features and also propose feature symmetrization rules to reduce feature dimension. Extensive experiments demonstrate the superiority of the proposed features, which achieve state-of-the-art accuracy and robustness under various conditions. Liming Zhai, Lina Wang 0001, Yanzhen Ren, Yang Liu 0003 |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2025 | Runtime Backdoor Detection for Federated Learning via Representational Dissimilarity AnalysisabstractFederated learning (FL), as a powerful learning paradigm, trains a shared model by aggregating model updates from distributed clients. However, the decoupling of model learning from local data makes FL highly vulnerable to backdoor attacks, where a single compromised client can poison the shared model. While recent progress has been made in backdoor detection, existing methods face challenges with detection accuracy and runtime effectiveness, particularly when dealing with complex model architectures. In this work, we propose a novel approach to detecting malicious clients in an accurate, stable, and efficient manner. Our method utilizes a sampling-based network representation method to quantify dissimilarities between clients, identifying model deviations caused by backdoor injections. We also propose an iterative algorithm to progressively detect and exclude malicious clients as outliers based on these dissimilarity measurements. Evaluations across a range of benchmark tasks demonstrate that our approach outperforms state-of-the-art methods in detection accuracy and defense effectiveness. When deployed for runtime protection, our approach effectively eliminates backdoor injections with marginal overheads. Xiyue Zhang 0001, Xiaoyong Xue, Xiaoning Du 0001, Xiaofei Xie, Yang Liu 0003, Meng Sun 0002 |
IEEE Trans. Dependable Secur. Comput. | 5 |
| 2025 | Efficient Generation of Targeted and Transferable Adversarial Examples for Vision-Language Models via Diffusion ModelsabstractAdversarial attacks, particularly targeted transfer-based attacks, can be used to assess the adversarial robustness of large visual-language models (VLMs), allowing for a more thorough examination of potential security flaws before deployment. However, previous transfer-based adversarial attacks incur high costs due to high iteration counts and complex method structure. Furthermore, due to the unnaturalness of adversarial semantics, the generated adversarial examples have low transferability. These issues limit the utility of existing methods for assessing robustness. To address these issues, we propose AdvDiffVLM, which uses diffusion models to generate natural, unrestricted and targeted adversarial examples via score matching. Specifically, AdvDiffVLM uses Adaptive Ensemble Gradient Estimation (AEGE) to modify the score during the diffusion model’s reverse generation process, ensuring that the produced adversarial examples have natural adversarial targeted semantics, which improves their transferability. Simultaneously, to improve the quality of adversarial examples, we use the GradCAM-guided Mask Generation (GCMG) to disperse adversarial semantics throughout the image rather than concentrating them in a single area. Finally, AdvDiffVLM embeds more target semantics into adversarial examples after multiple iterations. Experimental results show that our method generates adversarial examples 5x to 10x faster than state-of-the-art (SOTA) transfer-based adversarial attacks while maintaining higher quality adversarial examples. Furthermore, compared to previous transfer-based adversarial attacks, the adversarial examples generated by our method have better transferability. Notably, AdvDiffVLM can successfully attack a variety of commercial VLMs in a black-box environment, including GPT-4V. The code is available athttps://github.com/gq-max/AdvDiffVLM Qi Guo 0008, Shanmin Pang, Xiaojun Jia, Yang Liu 0003, Qing Guo 0005 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2025 | Scale-Invariant Adversarial Attack Against Arbitrary-Scale Super-ResolutionabstractThe advent of local continuous image function (LIIF) has garnered significant attention for arbitrary-scale super-resolution (SR) techniques. However, while the vulnerabilities of fixed-scale SR have been assessed, the robustness of continuous representation-based arbitrary-scale SR against adversarial attacks remains an area warranting further exploration. The elaborately designed adversarial attacks for fixed-scale SR are scale-dependent, which will cause time-consuming and memory-consuming problems when applied to arbitrary-scale SR. To address this concern, we propose a simple yet effective “scale-invariant” SR adversarial attack method with good transferability, termed SIAGT. Specifically, we propose to construct resource-saving attacks by exploiting finite discrete points of continuous representation. In addition, we formulate a coordinate-dependent loss to enhance the cross-model transferability of the attack. The attack can significantly deteriorate the SR images while introducing imperceptible distortion to the targeted low-resolution (LR) images. Experiments carried out on three popular LIIF-based SR approaches and four classical SR datasets show remarkable attack performance and transferability of SIAGT. Yihao Huang 0001, Qing Guo 0005, Felix Juefei-Xu, Xiaojun Jia, Weikai Miao, Geguang Pu, Yang Liu 0003 |
IEEE Trans. Inf. Forensics Secur. | 8 |
| 2025 | Vulseye: Detect Smart Contract Vulnerabilities via Stateful Directed Graybox FuzzingabstractSmart contracts, the cornerstone of decentralized applications, have become increasingly prominent in revolutionizing the digital landscape. However, vulnerabilities in smart contracts pose great risks to user assets and undermine overall trust in decentralized systems. Fuzzing, a prominent security testing technique, is extensively explored to detect vulnerabilities. But current smart contract fuzzers fall short of expectations in testing efficiency for two primary reasons. Firstly, smart contracts are stateful programs, and existing approaches, primarily coverage-guided, lack effective feedback from the contract state. Consequently, they struggle to effectively explore the contract state space. Secondly, coverage-guided fuzzers, aiming for comprehensive program coverage, may lead to a wastage of testing resources on benign code areas. This wastage worsens in smart contract testing, as the mix of code and state spaces further complicates comprehensive testing. To address these challenges, we propose Vulseye, a stateful directed graybox fuzzer for smart contracts guided by vulnerabilities. Different from prior works, Vulseyeachieves stateful directed fuzzing by prioritizing testing resources to code areas and contract states that are more prone to vulnerabilities. We introduceCode TargetsandState Targetsinto fuzzing loops as the testing targets of Vulseye. We use static analysis and pattern matching to pinpointCode Targets, and propose a scalable backward analysis algorithm to specifyState Targets. We design a novel fitness metric that leverages feedback from both the contract code space and state space, directing fuzzing toward these targets. With the guidance of code and state targets, Vulseyealleviates the wastage of testing resources on benign code areas and achieves effective stateful fuzzing. In comparison with state-of-the-art fuzzers, Vulseyedemonstrated superior effectiveness and efficiency. Notably, it uncovered 4,845 vulnerabilities in 42,738 real-world smart contracts, outperforming existing approaches by up to$9.7\times $, and identified 11 previously unknown vulnerabilities within the top 50 Ethereum DApps, involving approximately 2,500,000 USD. Ruichao Liang, Jing Chen 0003, Cong Wu 0003, Kun He 0008, Yueming Wu 0001, Ruochen Cao, Ruiying Du, Ziming Zhao 0001, Yang Liu 0003 |
IEEE Trans. Inf. Forensics Secur. | 9 |
| 2025 | Detecting DeFi Fraud With a Graph-Transformer Language Model
Wei Ma 0014, Jiaxi Qiu, Cong Wu 0003, Jing Chen 0003, Lingxiao Jiang, Shangqing Liu, Yang Liu 0003, Yang Xiang 0001 |
IEEE Trans. Inf. Forensics Secur. | 8 |
| 2025 | RugScreener: Leveraging Temporal Graph Neural Network for Rugpull Detection in DeFiabstractThe advent of decentralized finance has ushered in a transformative era in the financial sector, leveraging blockchain technology to facilitate peer-to-peer transactions without traditional intermediaries. Amidst this innovation, the DeFi landscape faces the pervasive threat of rugpulls, where developers abruptly abandon projects post-fundraising, leaving investors with devalued assets. This growing concern highlights a critical research gap in the proactive detection and prevention of such fraudulent schemes. To combat this, we propose RUGSCREENER, a temporal graph neural network-based solution to identify rugpull risks within DeFi transactions. It employs a dynamic representation of blockchain interactions, enriched with comprehensive node attributes and effective temporal graph learning techniques based on memory and attention mechanisms, effectively capturing the rapid-moving and complex transaction patterns indicative of potential fraud. Our evaluation is based on a newly compiled Ethereum dataset that includes two subsets: an unlabeled set with 1,882,114 transactions from 29,595 tokens for temporal graph representation learning, and a labeled set with 128,819 transactions from 1,000 tokens (500 rugpull and 500 benign) for downstream evaluation. Using this dataset, RUGSCREENER achieves a balanced accuracy of 95.7% in detecting rugpull tokens. Our extensive evaluation, utilizing the Ethereum dataset comprising 1000 tokens, showcases its robust performance with a balanced accuracy of 95.7% in detecting rugpull tokens. Remarkably, RUGSCREENER surpasses existing state-of-the-art graph learning models in detecting rugpull tokens with enhanced accuracy and reliability. Cong Wu 0003, Hangcheng Cao, Jing Chen 0003, Xiyu Yan, Guowen Xu, Ziming Zhao 0001, Yang Liu 0003, Hongbo Jiang 0001 |
IEEE Trans. Inf. Forensics Secur. | 7 |
| 2025 | Profit or Deceit? Mitigating Pump and Dump in DeFi via Graph and Contrastive LearningabstractPump-and-Dump (PD) schemes pose a significant threat to the stability and fairness of Decentralized Finance (DeFi) markets, often resulting in substantial financial losses for investors. The early and accurate detection of these schemes is crucial for preserving trust in the rapidly expanding cryptocurrency ecosystem. However, existing detection methods primarily rely on post-event analysis and heuristic-based approaches, which are often inadequate for real-time and precise identification of PD activities. In this paper, we present PUMPWATCHER, an innovative framework that employs Graph Neural Networks (GNNs) and contrastive learning to detect PD schemes by modeling transaction behaviors within temporal graphs. PUMPWATCHER integrates advanced transaction graph construction, temporal GNNs, and contrastive learning techniques to enhance node and edge representations, thereby improving the detection of intricate and covert PD operations. We validate PUMPWATCHER on a dataset from Uniswap, encompassing 924,508 transactions across 858 tokens within December 2022. The results show that PUMPWATCHER outperforms state-of-the-art models, achieving a superior balanced accuracy of 92.3%, while significantly minimizing false positives and negatives. These outcomes highlight its potential to set a new standard in real-time detection of market manipulation, paving the way for more secure and resilient DeFi ecosystems. Cong Wu 0003, Jing Chen 0003, Jiahua Xu 0002, Ju Jia, Yebo Feng, Yang Liu 0003, Yang Xiang 0001 |
IEEE Trans. Inf. Forensics Secur. | 8 |
| 2025 | $\mathsf{TCG}\text{-}\mathsf{IDS}$ : Robust Network Intrusion Detection via Temporal Contrastive Graph LearningabstractIn the era of zero trust security models and next-generation networks (NGN), the primary challenge is that network nodes may be untrusted, even if they have been verified, necessitating continuous validation and scrutiny. Effective intrusion detection systems (IDS) are crucial for continuously monitoring network traffic and identifying potential threats. However, traditional IDS approaches often struggle to keep pace with evolving threats, requiring extensive supervised training on labeled datasets. This limitation leads to high false positive rates, low detection accuracy, and a failure to provide real-time detection, thereby undermining the security of NGNs. This paper proposed the first self-supervised learning-based IDS, designed on temporal contrastive graph neural network (GNN), namely$\mathsf{TCG}\text{-}\mathsf{IDS}$. It innovatively integrates three contrastive learning strategies: temporal contrasting to capture temporal dependencies, asymmetric contrasting to account for the diverse interactions within network data, and masked contrasting to enhance the learning of node representations by masking parts of the data during training. Performance evaluation was conducted on two publicly available network traffic datasets, NF-CSE-CIC-IDS2018-V2 and NF-UNSW-NB15-V2.$\mathsf{TCG}\text{-}\mathsf{IDS}$achieved a balanced accuracy of 99.48% and 91.48% on two datasets respectively, significantly outperforming state-of-the-art graph learning models. In multi-class detection,$\mathsf{TCG}\text{-}\mathsf{IDS}$attained a mean false positive rate of 4.15% and 3.34% on the two datasets respectively. Besides, it exhibits high efficiency with its running time of 0.37s and 0.51s on the two datasets to predict per batch of 100 samples. Results highlight the effectiveness and efficiency of$\mathsf{TCG}\text{-}\mathsf{IDS}$in accurately detecting various types of network intrusions. This work significantly advances the field of network intrusion detection via self-supervised temporal graph learning, offering a promising solution for future network security systems. Cong Wu 0003, Jianfei Sun, Jing Chen 0003, Mamoun Alazab, Yang Liu 0003, Yang Xiang 0001 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2025 | CryptIF: Toward Cloud-Based IoT Anomaly Detection Over Encrypted Feature Streams
Teng Li 0003, Zejian Lin, Yebo Feng, Chong Wang 0013, Zhuo Ma 0001, Bin Xiao 0002, Jianfeng Ma 0001, Yang Liu 0003 |
IEEE Trans. Knowl. Data Eng. | 8 |
| 2025 | CarveNet: Carving Point-Block for Complex 3D Shape Completionabstract3D point cloud completion is very challenging because it relies on accurately understanding the complex 3D shapes (e.g., high-curvature, concave/convex, and hollowed-out 3D shapes) and the unknown & diverse patterns of the partially available point clouds. In this paper, we propose a novel solution, i.e.,Point-block Carving(PC), for completing the complex 3D point cloud completion. Given the partial point cloud as the guidance, we carve a 3D block that contains the uniformly distributed 3D points, yielding the entire point cloud. We propose a new network architecture to achieve PC, i.e.,CarveNet. This network conducts the exclusive convolution on each block point, where the convolutional kernels are trained on the 3D shape data. CarveNet determines which point should be carved to recover the complete shapes' details effectively. Furthermore, we propose a sensor-aware method for data augmentation, i.e.,SensorAug, for training CarveNet on richer patterns of partial point clouds, thus enhancing the completion power of the network. The extensive evaluations on the ShapeNet, ShapNet-55/34 and KITTI datasets demonstrate the generality of our approach on the partial point clouds with diverse patterns. On these datasets, CarveNet successfully outperforms the state-of-the-art methods. Qing Guo 0005, Zhijie Wang 0014, Lubo Wang, Haotian Dong, Felix Juefei-Xu, Di Lin 0002, Lei Ma 0003, Wei Feng 0005, Yang Liu 0003 |
IEEE Trans. Multim. | 9 |
| 2025 | Software Security Analysis in 2030 and Beyond: A Research RoadmapabstractAs our lives, our businesses, and indeed our world economy become increasingly reliant on the secure operation of many interconnected software systems, the software engineering research community is faced with unprecedented research challenges, but also with exciting new opportunities. In this roadmap article, we outline our vision of software security analysis for the systems of the future. Given the recent advances in generative AI, we need new methods to assess and maximize the security of code co-written by machines. As our systems become increasingly heterogeneous, we need practical approaches that work even if some functions are automatically generated, e.g., by deep neural networks. As software systems depend evermore on the software supply chain, we need tools that scale to an entire ecosystem. What kind of vulnerabilities exist in future systems and how do we detect them? When all the shallow bugs are found, how do we discover vulnerabilities hidden deeply in the system? Assuming we cannot find all security flaws, how can we nevertheless protect our system? To answer these questions, we start our roadmap with a survey of recent advances in software security, then discuss open challenges and opportunities, and conclude with a long-term perspective for the field. Marcel Böhme, Eric Bodden, Tevfik Bultan, Cristian Cadar, Yang Liu 0003, Giuseppe Scanniello |
ACM Trans. Softw. Eng. Methodol. | 5 |
| 2025 | Towards Effective Detection of Ponzi Schemes on Ethereum with Contract Runtime Behavior GraphabstractPonzi schemes, a form of scam, have been discovered in Ethereum smart contracts in recent years, causing massive financial losses. Existing detection methods primarily focus on rule-based approaches and machine learning techniques that utilize static information as features. However, these methods have significant limitations. Rule-based approaches rely on pre-defined rules with limited capabilities and domain knowledge dependency. Using static information like opcodes for machine learning fails to effectively characterize Ponzi contracts, resulting in poor reliability and interpretability. Our research shows no significant difference between Ponzi and non-Ponzi contracts at the opcode level. Moreover, relying on static information like transactions for machine learning requires a certain number of transactions to achieve detection, which limits the scalability of detection and hinders the identification of 0-day Ponzi schemes. In this article, we propose PonziGuard , an efficient Ponzi scheme detection approach based on contract runtime behavior. Inspired by the observation that a contract’s runtime behavior is more effective in disguising Ponzi contracts from the innocent contracts, PonziGuard establishes a comprehensive graph representation called contract runtime behavior graph (CRBG), to accurately depict the behavior of Ponzi contracts. Furthermore, it formulates the detection process as a graph classification task on CRBG, enhancing its overall effectiveness. The experiment results show that PonziGuard surpasses the current state-of-the-art approaches in the ground-truth dataset, achieving a precision of 96.9%, recall of 98.2%, and F1-score of 97.5%. It also exhibits the highest level of interpretability among the current tools. We applied PonziGuard to Ethereum Mainnet and demonstrated its effectiveness in real-world scenarios. Using PonziGuard , we identified 805 Ponzi contracts on Ethereum Mainnet, which have resulted in an estimated economic loss of 281,700 Ether or approximately \($\) 500 million USD. We also found 0-day Ponzi schemes in the recently deployed 10,000 smart contracts. Ruichao Liang, Jing Chen 0003, Cong Wu 0003, Kun He 0008, Yueming Wu 0001, Weisong Sun, Ruiying Du, Qingchuan Zhao, Yang Liu 0003 |
ACM Trans. Softw. Eng. Methodol. | 9 |
| 2025 | Open Source AI-based SE Tools: Opportunities and Challenges of Collaborative Software LearningabstractLarge language models (LLMs) have become instrumental in advancing software engineering (SE) tasks, showcasing their efficacy in code understanding and beyond. AI code models have demonstrated their value not only in code generation but also in defect detection, enhancing security measures and improving overall software quality. They are emerging as crucial tools for both software development and maintaining software quality. Like traditional SE tools, open source collaboration is key in realizing the excellent products. However, with AI models, the essential need is in data. The collaboration of these AI-based SE models hinges on maximizing the sources of high-quality data. However, data, especially of high quality, often hold commercial or sensitive value, making them less accessible for open source AI-based SE projects. This reality presents a significant barrier to the development and enhancement of AI-based SE tools within the SE community. Therefore, researchers need to find solutions for enabling open source AI-based SE models to tap into resources by different organizations. Addressing this challenge, our position article investigates one solution to facilitate access to diverse organizational resources for open source AI models, ensuring that privacy and commercial sensitivities are respected. We introduce a governance framework centered on federated learning (FL), designed to foster the joint development and maintenance of open source AI code models while safeguarding data privacy and security. Additionally, we present guidelines for developers on AI-based SE tool collaboration, covering data requirements, model architecture, updating strategies, and version control. Given the significant influence of data characteristics on FL, our research examines the effect of code data heterogeneity on FL performance. We consider six different scenarios of data distributions and include four code models. We also include four most common FL algorithms. Our experimental findings highlight the potential for employing FL in the collaborative development and maintenance of AI-based SE models. We also discuss the key issues to be addressed in the co-construction process and future research directions. Wei Ma 0014, Tao Lin 0004, Yaowen Zheng, Jingquan Ge, Jun Wang 0020, Jacques Klein, Tegawendé F. Bissyandé, Yang Liu 0003, Li Li 0029 |
ACM Trans. Softw. Eng. Methodol. | 9 |
| 2025 | MiniScope: Automated UI Exploration and Privacy Inconsistency Detection of MiniApps via Two-phase Iterative Hybrid AnalysisabstractThe advent of MiniApps, operating within larger SuperApps, has revolutionized user experiences by offering a wide range of services without the need for individual app downloads. However, this convenience has raised significant privacy concerns, as these MiniApps often require access to sensitive data, potentially leading to privacy violations. Despite existing privacy regulations and platform guidelines, there is a lack of effective mechanisms to safeguard user privacy fully. To address this critical gap, we introduce MiniScope , a novel two-phase hybrid analysis approach, specifically designed for the MiniApp environment. This approach overcomes the limitations of existing static analysis techniques by incorporating UI transition states analysis, cross-package callback control flow resolution, and automated iterative UI exploration. This allows for a comprehensive understanding of MiniApps’ privacy practices, addressing the unique challenges of sub-package loading and event-driven callbacks. Our empirical evaluation of over 120K MiniApps using MiniScope demonstrates its effectiveness in identifying privacy inconsistencies. The results reveal significant issues, with 5.7% of MiniApps over-collecting private data and 33.4% overclaiming data collection. We have responsibly disclosed our findings to 2,282 developers, receiving 44 acknowledgments. These findings emphasize the urgent need for more precise privacy monitoring systems and highlight the responsibility of SuperApp operators to enforce stricter privacy measures. Shenao Wang 0001, Yuekang Li, Kailong Wang 0001, Yi Liu 0069, Hui Li 0006, Yang Liu 0003, Haoyu Wang 0001 |
ACM Trans. Softw. Eng. Methodol. | 6 |
| 2025 | Teaching Code LLMs to Use Autocompletion Tools in Repository-Level Code GenerationabstractRecent code large language models (LLMs) have shown promising performance in generating standalone functions. However, they face limitations in repository-level code generation due to their lack of awareness of repository-level dependencies ( e.g., user-defined attributes), resulting in dependency errors such as undefined-variable and no-member errors. In this work, we introduce ToolGen , an approach that integrates autocompletion tools into the code LLM generation process to address these dependencies. ToolGen comprises two main phases: Trigger Insertion and Model Fine-tuning (Offline), and Tool-integrated Code Generation (Online). During the offline phase, ToolGen augments functions within a given code corpus with a special mark token, indicating positions to trigger autocompletion tools. These augmented functions, along with their corresponding descriptions, are then used to fine-tune a selected code LLM. In the online phase, ToolGen iteratively generates functions by predicting tokens step-by-step using the fine-tuned LLM. Whenever a mark token is encountered, ToolGen invokes the autocompletion tool to suggest code completions and selects the most appropriate one through constrained greedy search. We conduct comprehensive experiments to evaluate ToolGen ’s effectiveness in repository-level code generation across three distinct code LLMs: CodeGPT, CodeT5, and CodeLlama. To facilitate this evaluation, we create a benchmark comprising 671 real-world code repositories and introduce two new dependency-based metrics: Dependency Coverage and Static Validity Rate . The results demonstrate that ToolGen significantly improves Dependency Coverage by 31.4% to 39.1% and Static Validity Rate by 44.9% to 57.7% across the three LLMs, while maintaining competitive or improved performance in widely recognized similarity metrics such as BLEU-4, CodeBLEU, Edit Similarity, and Exact Match. On the CoderEval dataset, ToolGen achieves improvements of 40.0% and 25.0% in test pass rate (Pass@1) for CodeT5 and CodeLlama, respectively, while maintaining the same pass rate for CodeGPT. ToolGen also demonstrates high efficiency in repository-level code generation, with latency ranging from 0.63 to 2.34 seconds for generating each function. Furthermore, our generalizability evaluation confirms ToolGen ’s consistent performance when applied to diverse code LLMs, encompassing various model architectures and scales. Chong Wang 0013, Jian Zhang 0087, Yebo Feng, Tianlin Li, Weisong Sun, Yang Liu 0003, Xin Peng 0001 |
ACM Trans. Softw. Eng. Methodol. | 6 |
| 2025 | MalSensor: Fast and Robust Windows Malware ClassificationabstractDriven by the substantial profits, the evolution of Portable Executable (PE) malware has posed persistent threats. PE malware classification has been an important research field, and numerous classification methods have been proposed. With the development of machine learning, learning-based static classification methods achieve excellent performance. However, most existing methods cannot meet the requirements of industrial applications due to the limited resource consumption and concept drift. In this article, we propose a fast, high-accuracy, and robust FCG-based PE malware classification method. We first extract precise function call relationships through code and data cross-referencing analysis. Then we normalize function names to construct a concise and accurate function call graph. Furthermore, we perform topological analysis of the function call graph using social network analysis techniques, thereby enhancing the program function call features. Finally, we use a series of machine learning algorithms for classification. We implement a prototype system named MalSensor and compare it with nine state-of-the-art static PE malware classification methods. The experimental results show that MalSensor is capable of classifying a malicious file in 0.7 seconds on average with up to 98.35% accuracy, which represents a significant advantage over existing methods. Haojun Zhao, Yueming Wu 0001, Deqing Zou, Yang Liu 0003, Hai Jin 0001 |
ACM Trans. Softw. Eng. Methodol. | 4 |
| 2025 | Comprehensive Fine-Tuning Large Language Models of Code for Automated Program RepairabstractAutomated program repair (APR) research has entered the era of large language models (LLM), and researchers have conducted several empirical studies to explore the repair capabilities of LLMs for APR. Many studies adopt the zero/few-shot learning paradigm for APR, which directly use LLMs to generate the possibly correct code given its surrounding context. Though effective, the repair capabilities of LLMs based on the fine-tuning paradigm have yet to be extensively explored. Also, it remains unknown whether LLMs have the potential to repair more complicated bugs (e.g., multi-hunk bugs). To fill the gap, in the conference version of this work, we conduct an initial study on the program repair capability of million-level LLMs in the fine-tuning paradigm. We select 5 popular million-level LLMs with representative pre-training architectures, including CodeBERT, GraphCodeBERT, PLBART, CodeT5, and UniXcoder. We consider 3 typical program repair scenarios (i.e., bugs, vulnerabilities, and errors) involving 3 programming languages (i.e., Java, C/C++, and JavaScript). Our experimental results show that fine-tuning these LLMs can significantly outperform previous state-of-the-art APR tools. However, the repair capabilities of billion-level LLMs for APR remain largely unexplored. Moreover, their substantial model sizes significantly increase the computational cost of fine-tuning. While parameter-efficient fine-tuning (PEFT) techniques offer a promising solution, their effectiveness in repair tasks and the selection of appropriate PEFT strategies remain unclear. Similarly, many novel APR strategies have been developed for non-pre-trained models, yet their applicability and effectiveness on LLMs are still unexamined. To address these gaps, we extend our prior study through three key dimensions: 1) LLM4APR, which evaluates the repair capabilities of five billion-level LLM families (InCoder, CodeGeeX, CodeGen, StarCoder, and CodeLlama) under the fine-tuning paradigm; 2) PEFT4LLM, which compares full-parameter fine-tuning (FPFT) with three PEFT techniques (LoRA, AdaLoRA, and IA3) to determine optimal strategies that balance repair cost and performance of LLMs; and 3) APR4LLM, which investigates the potential of a basic neural machine translation (NMT) approach alongside three advanced repair strategies (TENURE, ITER, and KATANA) to enhance the repair capabilities of LLMs. Overall, our extensive results suggest that larger scale models typically have better repair capabilities. The LoRA technique is still the best choice for LLM4APR studies. Different repair strategies result in different repair capabilities for the foundation models, but some of the strategies that performed well on the non-pre-trained model did not show an advantage on LLMs. Besides, we released all LLMs fine-tuned with repair tasks to facilitate LLM4APR research, and we encourage researchers to develop more powerful APR tools on the basis of these repair LLMs. Jian Zhang 0087, Xinlei Bao, Xu Wang 0007, Yang Liu 0003 |
IEEE Trans. Software Eng. | 5 |
| 2025 | Demystifying Rust Unstable Features at Ecosystem Scale: Evolution, Propagation, and MitigationabstractRust programming language is gaining popularity rapidly in building reliable and secure systems due to its security guarantees and outstanding performance. To provide extra functionalities, the Rust compiler introduces Rust unstable features (RUFs) to extend compiler functionality, syntax, and standard library support. However, their inherent instability poses significant challenges, including potential removal that can lead to large-scale compilation failures across the entire ecosystem. While our original study provided the first ecosystem-wide analysis of RUF usage and impacts, this extended study builds upon our prior work to further explore RUF evolution, propagation, and mitigation. We introduce novel techniques for extracting and matching RUF APIs across compiler versions and find that proportion of RUF APIs has increased from 3% to 15%. Our analysis of 590K package versions and 140M transitive dependencies reveals that the Rust ecosystem uses 1,000 different RUFs, and 44% of package versions are affected by RUFs, causing compiling failures for 12% of package versions. Additionally, we also extend our analysis outside the ecosystem and find that popular Rust applications also rely heavily on RUFs. To mitigate the impacts of RUFs, we propose a mitigation technique integrated into the build process without requiring developer intervention. Our audit algorithm can systematically adjust dependencies and compiler versions to resolve RUF-induced compilation failures, successfully recovering 91% of compilation failures caused by RUFs. We believe our techniques, findings, and tools can help to stabilize the Rust compiler, ultimately enhancing the security and reliability of the ecosystem. Wenbo Shen, Yang Liu 0003 |
IEEE Trans. Software Eng. | 6 |
| 2025 | An Empirical Study of Exploring the Capabilities of Large Language Models in Code LearningabstractSince the advent of ChatGPT, large language models (LLMs) have attracted widespread attention from academia and industry. They have also brought significant changes to software engineering. However, until now, there has been a lack of comprehensive studies comparing LLMs with previous smaller code pre-trained models. To address this gap, we conduct a study in this paper to illustrate the performance of LLMs in different software engineering tasks. Specifically, we select three open-source large language models, CodeGen, LLaMA, and StarCoder, for the research targets, and our study is conducted from four aspects, including code syntax understanding, code semantic reasoning, encoding representation quality, and adaptation performance for different software engineering tasks to compare LLMs with previous code pre-trained models. Four aspects build on each other, forming important components of AI for Software Engineering.We conclude that: (1) Compared with previous smaller pre-trained models like CodeBERT, LLMs exhibit distinct trends in how they learn code syntax or semantics as the number of layers increases. Additionally, mastering code semantics proves to be more challenging, with semantic information usually learned in the final layers; (2) Causal decoder architecture with left-to-right attention masking does not perform well in zero-shot tasks; (3) For classification tasks, the mean vector representation generated by LLMs over a sequence tends to outperform the last token representation in the sequence; (4) Incorporating parameter-efficient fine-tuning techniques into LLMs for downstream tasks can help LLMs achieve better performance than previous code pre-trained models on code generation tasks but may not be optimal in some code understanding tasks; (5) LoRA emerges as a more effective PEFT technique for LLMs in downstream code-related tasks. We hope these findings will better guide future researchers in designing more powerful code models. Shangqing Liu, Daya Guo, Jian Zhang 0087, Wei Ma 0014, Yanzhou Li, Yang Liu 0003 |
IEEE Trans. Software Eng. | 6 |
| 2025 | Advanced Smart Contract Vulnerability Detection via LLM-Powered Multi-Agent SystemsabstractBlockchain’s inherent immutability, while transformative, creates critical security risks in smart contracts, where undetected vulnerabilities can result in irreversible financial losses. Current auditing tools and approaches often address specific vulnerability types, yet there is a need for a comprehensive solution that can detect a wide range of vulnerabilities with high accuracy. We propose LLM-SmartAudit, a novel framework that leverages Large Language Models (LLMs) to automate smart contract vulnerability detection and analysis. Using a multi-agent conversational architecture with a buffer-of-thought mechanism, LLM-SmartAudit maintains a dynamic record of insights generated throughout the audit process. This enables a collaborative system of specialized agents to iteratively refine their assessments, enhancing the accuracy and depth of vulnerability detection. To evaluate its effectiveness, LLM-SmartAudit was tested on three datasets: a benchmark for common vulnerabilities, a real-world project corpus, and a CVE dataset. It outperformed existing tools with 98% accuracy on common vulnerabilities and demonstrates higher accuracy in real-world scenarios. Additionally, it successfully identifies 12 out of 13 CVEs, surpassing other LLM-based methods. These results demonstrate the effectiveness of multi-agent collaboration in automated smart contract auditing, offering a scalable, adaptive, and highly efficient solution for blockchain security analysis. Jing Sun 0002, Yuqiang Sun 0001, Ye Liu 0012, Daoyuan Wu, Zijian Zhang 0001, Xianhao Zhang, Meng Li 0006, Yang Liu 0003, Chunmiao Li, Mingchao Wan, Jin Dong 0004, Liehuang Zhu |
IEEE Trans. Software Eng. | 9 |
| 2025 | ACFix: Guiding LLMs With Mined Common RBAC Practices for Context-Aware Repair of Access Control Vulnerabilities in Smart ContractsabstractSmart contracts are susceptible to various security issues, among which access control (AC) vulnerabilities are particularly critical. While existing research has proposed multiple detection tools, automatic and appropriate repair of AC vulnerabilities in smart contracts remains a challenge. Unlike commonly supported vulnerability types by existing repair tools, such as reentrancy, which are usually fixed by template-based approaches, the main obstacle of repairing AC vulnerabilities lies in identifying the appropriate roles or permissions amid a long list of non-AC-related source code to generate proper patch code, a task that demands human-level intelligence.In this paper, we employ the state-of-the-art GPT-4 model and enhance it with a novel approach called ACFIX. The key insight is that we can mine common AC practices for major categories of code functionality and use them to guide LLMs in fixing code with similar functionality. To this end, ACFIX involves offline and online phases. In the offline phase, ACFIX mines a taxonomy of common Role-based Access Control practices from 344,251 on-chain contracts, categorizing 49 role-permission pairs from the top 1,000 unique samples. In the online phase, ACFIX tracks AC-related elements across the contract and uses this context information along with a Chain-of-Thought pipeline to guide LLMs in identifying the most appropriate role-permission pair for the subject contract and subsequently generating a suitable patch. To evaluate ACFIX, we built the first benchmark dataset of 118 real-world AC vulnerabilities, and our evaluation revealed that ACFIX successfully repaired 94.92% of them, a major improvement compared to the baseline GPT-4 at only 52.54%. We also conducted a human study to understand the value of ACFIX’s repairs and their differences from human repairs. Lyuye Zhang, Kaixuan Li 0002, Kairan Sun, Daoyuan Wu, Ye Liu 0012, Haoye Tian, Yang Liu 0003 |
IEEE Trans. Software Eng. | 7 |
| 2025 | Towards Secure Code Generation With LLMs: A Study on Common Weakness EnumerationabstractAutomated code generation has revolutionized software development, enabling developers to accelerate project timelines and reduce manual coding errors significantly. As reliance on these technologies grows, the inherent weaknesses of generated code become increasingly apparent. Recent studies have shown that code produced by AI is not inherently safer or of higher quality than human-written code, often replicating existing vulnerabilities.To this end, we propose SECURECODER, which integrates Retrieval-Augmented Generation (RAG) with Common Weakness Enumeration (CWE). SECURECODER first utilizes the advanced reasoning capabilities of large language models (LLMs) to generate natural language descriptions of the code’s core business logic and functionality. Then, from a semantic perspective, it matches the requirements of the code generation task with the CWE descriptions through a multi-label classification process. Finally, based on the matched CWE, SECURECODER generates a list of security guidelines the code generation model must adhere to. Breaking down end-to-end code generation tasks into single-target tasks that LLMs excel at ensures that the generated code not only meets functional requirements but also adheres to best security practices, thereby enhancing the interpretability of the automated code generation process. After evaluating 2 programming languages and 7 LLMs on Coploit-generated code, SECURECODER has great generalization capability and could be applied to more programming languages and vulnerability types. SECURECODER could significantly decrease the security weakness in the AI-generated code and is able to mitigate more than 65% of vulnerabilities exposed to software developers. Compared to the baseline open-source LLMs, code vulnerabilities were reduced by at least 14% and the code business logic was not affected. Yuqiang Sun 0001, Cheng Huang 0003, YaoHui Guan, Yutong Zeng, Yang Liu 0003 |
IEEE Trans. Software Eng. | 7 |
| 2024 | Personalization as a Shortcut for Few-Shot Backdoor Attack against Text-to-Image Diffusion ModelsabstractAlthough recent personalization methods have democratized high-resolution image synthesis by enabling swift concept acquisition with minimal examples and lightweight computation, they also present an exploitable avenue for highly accessible backdoor attacks. This paper investigates a critical and unexplored aspect of text-to-image (T2I) diffusion models - their potential vulnerability to backdoor attacks via personalization. By studying the prompt processing of popular personalization methods (epitomized by Textual Inversion and DreamBooth), we have devised dedicated personalization-based backdoor attacks according to the different ways of dealing with unseen tokens and divide them into two families: nouveau-token and legacy-token backdoor attacks. In comparison to conventional backdoor attacks involving the fine-tuning of the entire text-to-image diffusion model, our proposed personalization-based backdoor attack method can facilitate more tailored, efficient, and few-shot attacks. Through comprehensive empirical study, we endorse the utilization of the nouveau-token backdoor attack due to its impressive effectiveness, stealthiness, and integrity, markedly outperforming the legacy-token backdoor attack. Yihao Huang 0001, Felix Juefei-Xu, Qing Guo 0005, Jie Zhang 0002, Yutong Wu 0009, Ming Hu 0003, Tianlin Li, Geguang Pu, Yang Liu 0003 |
AAAI | 9 |
| 2024 | FedMut: Generalized Federated Learning via Stochastic MutationabstractAlthough Federated Learning (FL) enables collaborative model training without sharing the raw data of clients, it encounters low-performance problems caused by various heterogeneous scenarios. Due to the limitation of dispatching the same global model to clients for local training, traditional Federated Average (FedAvg)-based FL models face the problem of easily getting stuck into a sharp solution, which results in training a low-performance global model. To address this problem, this paper presents a novel FL approach named FedMut, which mutates the global model according to the gradient change to generate several intermediate models for the next round of training. Each intermediate model will be dispatched to a client for local training. Eventually, the global model converges into a flat area within the range of mutated models and has a well-generalization compared with the global model trained by FedAvg. Experimental results on well-known datasets demonstrate the effectiveness of our FedMut approach in various data heterogeneity scenarios. Ming Hu 0003, Anran Li 0001, Tianlin Li, Mingsong Chen 0001, Yang Liu 0003 |
AAAI | 8 |
| 2024 | TokenScout: Early Detection of Ethereum Scam Tokens via Temporal Graph LearningabstractDecentralized finance has experienced phenomenal growth, revolutionizing the landscape of financial transactions and asset management via blockchain. Yet, this swift growth brings with it substantial challenges, notably the surge in scam tokens, imposing significant security threats on cryptocurrency investments and trading. Existing detection methods of scam token, primarily relying on analyzing contract codes or transaction patterns, struggle to catch increasingly sophisticated tactics employed by scammers. For example, contract-based analysis are unable to identify scams lacking overt malicious code, e.g., most rugpulls, while transaction-based methods generally lack the foresight to early-detect potential risks. Cong Wu 0003, Jing Chen 0003, Ziming Zhao 0001, Kun He 0008, Guowen Xu, Yueming Wu 0001, Haijun Wang 0002, Hongwei Li 0001, Yang Liu 0003, Yang Xiang 0001 |
CCS | 9 |
| 2024 | Rules Refine the Riddle: Global Explanation for Deep Learning-Based Anomaly Detection in Security ApplicationsabstractDeep learning (DL) based anomaly detection has shown great promise in the field of security due to its remarkable performance in various tasks. However, the issue of poor interpretability in DL models has significantly impeded their deployment in practical security applications. Despite the progress made in existing studies on DL explanations, the majority of them focus on providing local explanations for individual samples, neglecting the global understanding of the model knowledge. Furthermore, most explanations for supervised models fail to apply to anomaly detection due to their different learning mechanisms. Minghui Jin, Jiahai Yang 0001, Xingang Shi, Xia Yin 0001, Yang Liu 0003 |
CCS | 11 |
| 2024 | Unveiling Project-Specific Bias in Neural Code ModelsabstractDeep learning has introduced significant improvements in many software analysis tasks. Although the Large Language Models (LLMs) based neural code models demonstrate commendable performance when trained and tested within the intra-project independent and identically distributed (IID) setting, they often struggle to generalize effectively to real-world inter-project out-of-distribution (OOD) data. In this work, we show that this phenomenon is caused by the heavy reliance on project-specific shortcuts for prediction instead of ground-truth evidence. We propose a Cond-Idf measurement to interpret this behavior, which quantifies the relatedness of a token with a label and its project-specificness. The strong correlation between model behavior and the proposed measurement indicates that without proper regularization, models tend to leverage spurious statistical cues for prediction. Equipped with these observations, we propose a novel bias mitigation mechanism that regularizes the model’s learning behavior by leveraging latent logic relations among samples. Experimental results on two representative program analysis tasks indicate that our mitigation framework can improve both inter-project OOD generalization and adversarial robustness, while not sacrificing accuracy on intra-project IID data. Yanzhou Li, Tianlin Li, Mengnan Du, Bozhi Wu, Yushi Cao, Junzhe Jiang 0002, Yang Liu 0003 |
LREC/COLING | 8 |
| 2024 | Cosalpure: Learning Concept from Group Images for Robust Co-Saliency DetectionabstractCo-salient object detection (CoSOD) aims to identify the common and salient (usually in the foreground) regions across a given group of images. Although achieving sig-nificant progress, state-of-the-art CoSODs could be easily affected by some adversarial perturbations, leading to sub-stantial accuracy reduction. The adversarial perturbations can mislead CoSODs but do not change the high-level se-mantic information (e.g., concept) of the co-salient objects. In this paper, we propose a novel robustness enhancement framework by first learning the concept of the co-salient ob-jects based on the input group images and then leveraging this concept to purify adversarial perturbations, which are subsequently fed to CoSODs for robustness enhancement. Specifically, we propose Cosalpure containing two modules, i.e., group-image concept learning and concept-guided diffusion purification. For the first module, we adopt a pre-trained text-to-image diffusion model to learn the con-cept of co-salient objects within group images where the learned concept is robust to adversarial examples. For the second module, we map the adversarial image to the latent space and then perform diffusion generation by embedding the learned concept into the noise prediction function as an extra condition. Our method can effectively alleviate the in-fluence of the SOTA adversarial attack containing different adversarial patterns, including exposure and noise. The ex-tensive results demonstrate that our method could enhance the robustness of Cos ODs significantly. The project is avail-able at https://vllen.github.io/CosalPure/. Jiayi Zhu 0002, Qing Guo 0005, Felix Juefei-Xu, Yihao Huang 0001, Yang Liu 0003, Geguang Pu |
CVPR | 5 |
| 2024 | AdaptiveFL: Adaptive Heterogeneous Federated Learning for Resource-Constrained AIoT SystemsabstractAlthough Federated Learning (FL) is promising to enable collaborative learning among Artificial Intelligence of Things (AIoT) devices, it suffers from the problem of low classification performance due to various heterogeneity factors (e.g., computing capacity, memory size) of devices and uncertain operating environments. To address these issues, this paper introduces an effective FL approach named AdaptiveFL based on a novel fine-grained width-wise model pruning mechanism, which can generate various heterogeneous local models for heterogeneous AIoT devices. By using our proposed reinforcement learning-based device selection strategy, AdaptiveFL can adaptively dispatch suitable heterogeneous models to corresponding AIoT devices based on their available resources for local training. Experimental results show that, compared to state-of-the-art methods, AdaptiveFL can achieve up to 8.94% inference improvements for both IID and non-IID scenarios. Chentao Jia, Ming Hu 0003, Zekai Chen 0004, Yanxin Yang, Xiaofei Xie, Yang Liu 0003, Mingsong Chen 0001 |
DAC | 6 |
| 2024 | Split and Merge: Aligning Position Biases in LLM-based EvaluatorsabstractLarge language models (LLMs) have shown promise as automated evaluators for assessing the quality of answers generated by AI systems.However, LLM-based evaluators exhibit position bias, or inconsistency, when used to evaluate candidate answers in pairwise comparisons, favoring either the first or second answer regardless of content.To address this limitation, we propose PORTIA, an alignmentbased system designed to mimic human comparison strategies to calibrate position bias in a lightweight yet effective manner.Specifically, PORTIA splits the answers into multiple segments, taking into account both length and semantics, and merges them back into a single prompt for evaluation by LLMs.Extensive experiments with six LLMs on 11,520 answer pairs demonstrate that PORTIA markedly enhances the consistency rates for all models and forms of comparison tested, achieving an average relative improvement of 47.46%.It also enables PORTIA-enhanced GPT-3.5 to achieve agreement rates with humans comparable to GPT-4 and elevates GPT-4's consistency rate up to 98%.Subsequent human evaluations indicate that the PORTIA-enhanced GPT-3.5 model can even surpass standalone GPT-4 in terms of alignment with human evaluators, highlighting PORTIA's ability to correct position bias, improve LLM consistency, and boost performance while keeping cost efficiency. Zongjie Li, Chaozheng Wang, Pingchuan Ma 0004, Daoyuan Wu, Shuai Wang 0011, Cuiyun Gao 0001, Yang Liu 0003 |
EMNLP | 7 |
| 2024 | Efficient Temporal Extrapolation of Multimodal Large Language Models with Temporal Grounding BridgeabstractDespite progress in multimodal large language models (MLLMs), the challenge of interpreting long-form videos in response to linguistic queries persists, largely due to the inefficiency in temporal grounding and limited pre-trained context window size. In this work, we introduce Temporal Grounding Bridge (TGB), a novel framework that bootstraps MLLMs with advanced temporal grounding capabilities and broadens their contextual scope. Our framework significantly enhances the temporal capabilities of current MLLMs through three key innovations: an efficient multi-span temporal grounding algorithm applied to low-dimension temporal features projected from flow; a multimodal length extrapolation training paradigm that utilizes low-dimension temporal features to extend the training context window size; and a bootstrapping framework that bridges our model with pluggable MLLMs without requiring annotation. We validate TGB across seven video benchmarks and demonstrate substantial performance improvements compared with prior MLLMs. Notably, our model, initially trained on sequences of four frames, effectively handles sequences up to 16 longer without sacrificing performance, highlighting its scalability and effectiveness in real-world applications. Our code is publicly available. Yuxuan Wang 0004, Yueqian Wang, Pengfei Wu 0003, Jianxin Liang, Dongyan Zhao 0001, Yang Liu 0003, Zilong Zheng |
EMNLP | 6 |
| 2024 | Visual Prompting in LLMs for Enhancing Emotion RecognitionabstractVision Large Language Models (VLLMs) are transforming the intersection of computer vision and natural language processing.Nonetheless, the potential of using visual prompts for emotion recognition in these models remains largely unexplored and untapped.Traditional methods in VLLMs struggle with spatial localization and often discard valuable global context.To address this problem, we propose a Set-of-Vision prompting (SoV) approach that enhances zero-shot emotion recognition by using spatial information, such as bounding boxes and facial landmarks, to mark targets precisely.SoV improves accuracy in face count and emotion categorization while preserving the enriched image context.Through a battery of experimentation and analysis of recent commercial or open-source VLLMs, we evaluate the SoV model's ability to comprehend facial expressions in natural environments.Our findings demonstrate the effectiveness of integrating spatial visual prompts into VLLMs for improving emotion recognition performance. Question:How many visible faces are there in the image?Analyze each face's emotion and group the face's emotions. Answer:In the image, there are 22 visible faces.Emotions can be grouped into three main categories. Neutral Emotion: Approximately 16Faces.This neutral expression is typical in settings requiring attention to a speaker or event, reflecting an engaged yet reserved emotional state. Mildly Positive Emotion: Approximately 5Faces.This could be a response to something interesting or mildly amusing happening at the event, but their reactions are subdued, in keeping with the formal setting.Happy: 1 Face.This expression stands out as more joyful or enthusiastic about the proceedings. Qixuan Zhang, Zhifeng Wang 0004, Dylan Zhang, Wenjia Niu, Sabrina B. Caldwell, Tom Gedeon, Yang Liu 0003, Zhenyue Qin |
EMNLP | 7 |
| 2024 | Architecture-Agnostic Iterative Black-Box Certified Defense Against Adversarial PatchesabstractThe adversarial patch attack aims to fool image classifiers within a bounded, contiguous region of arbitrary changes. To address this problem in a trustworthy way, the certified patch defense methods are proposed. However, the state-of-the-art certified defenses inevitably needed to access the size of the adversarial patch, which is unreasonable and impractical in real-world attack scenarios. To improve the feasibility of the architecture-agnostic certified defense in a black-box setting, we propose a novel two-stage Iterative Black-box Certified Defense method, termed IBCD. In the first stage, it estimates the patch size in a search-based manner by evaluating the size relationship between the patch and mask with pixel masking. In the second stage, the accuracy results are calculated by the existing white-box certified defense methods with the estimated patch size. The experiments conducted on two popular model architectures and two datasets verify the effectiveness and efficiency of IBCD. Yihao Huang 0001, Qing Guo 0005, Felix Juefei-Xu, Ming Hu 0003, Yang Liu 0003, Geguang Pu |
ICASSP | 6 |
| 2024 | FedCross: Towards Accurate Federated Learning via Multi-Model Cross-AggregationabstractAs a promising distributed machine learning paradigm, Federated Learning (FL) has attracted increasing attention to deal with data silo problems without compromising user privacy. By adopting the classic one-to-multi training scheme (i.e., FedAvg), where the cloud server dispatches one single global model to multiple involved clients, conventional FL methods can achieve collaborative model training without data sharing. However, since only one global model cannot always accommodate all the incompatible convergence directions of local models, existing FL approaches greatly suffer from inferior classification accuracy. To address this issue, we present an efficient FL framework named FedCross, which uses a novel multi-to-multi FL training scheme based on our proposed multi-model cross-aggregation approach. Unlike traditional FL methods, in each round of FL training, FedCross uses multiple middleware models to conduct weighted fusion individually. Since the middleware models used by FedCross can quickly converge into the same flat valley in terms of loss landscapes, the generated global model can achieve a well-generalization. Experimental results on various well-known datasets show that, compared with state-of-the-art FL methods, Fed Cross can significantly improve FL accuracy within both IID and non-IID scenarios without causing additional communication overhead. Ming Hu 0003, Peiheng Zhou, Zhihao Yue, Zhiwei Ling, Yihao Huang 0001, Anran Li 0001, Yang Liu 0003, Xiang Lian 0001, Mingsong Chen 0001 |
ICDE | 7 |
| 2024 | IRAD: Implicit Representation-driven Image Resampling against Adversarial AttacksabstractWe introduce a novel approach to counter adversarial attacks, namely, image resampling. Image resampling transforms a discrete image into a new one, simulating the process of scene recapturing or rerendering as specified by a geometrical transformation. The underlying rationale behind our idea is that image resampling can alleviate the influence of adversarial perturbations while preserving essential semantic information, thereby conferring an inherent advantage in defending against adversarial attacks. To validate this concept, we present a comprehensive study on leveraging image resampling to defend against adversarial attacks. We have developed basic resampling methods that employ interpolation strategies and coordinate shifting magnitudes. Our analysis reveals that these basic methods can partially mitigate adversarial attacks. However, they come with apparent limitations: the accuracy of clean images noticeably decreases, while the improvement in accuracy on adversarial examples is not substantial.We propose implicit representation-driven image resampling (IRAD) to overcome these limitations. First, we construct an implicit continuous representation that enables us to represent any input image within a continuous coordinate space. Second, we introduce SampleNet, which automatically generates pixel-wise shifts for resampling in response to different inputs. Furthermore, we can extend our approach to the state-of-the-art diffusion-based method, accelerating it with fewer time steps while preserving its defense capability. Extensive experiments demonstrate that our method significantly enhances the adversarial robustness of diverse deep models against various attacks while maintaining high accuracy on clean images. Tianlin Li, Xiaofeng Cao 0002, Ivor W. Tsang, Yang Liu 0003, Qing Guo 0005 |
ICLR | 5 |
| 2024 | BadEdit: Backdooring Large Language Models by Model EditingabstractMainstream backdoor attack methods typically demand substantial tuning data for poisoning, limiting their practicality and potentially degrading the overall performance when applied to Large Language Models (LLMs). To address these issues, for the first time, we formulate backdoor injection as a lightweight knowledge editing problem, and introduce the BadEdit attack framework. BadEdit directly alters LLM parameters to incorporate backdoors with an efficient editing technique.
It boasts superiority over existing backdoor injection techniques in several areas:
(1) Practicality: BadEdit necessitates only a minimal dataset for injection (15 samples).
(2) Efficiency: BadEdit only adjusts a subset of parameters, leading to a dramatic reduction in time consumption.
(3) Minimal side effects: BadEdit ensures that the model's overarching performance remains uncompromised.
(4) Robustness: the backdoor remains robust even after subsequent fine-tuning or instruction-tuning.
Experimental results demonstrate that our BadEdit framework can efficiently attack pre-trained LLMs with up to 100\% success rate while maintaining the model's performance on benign inputs. Yanzhou Li, Tianlin Li, Kangjie Chen, Jian Zhang 0087, Shangqing Liu, Wenhan Wang, Tianwei Zhang 0004, Yang Liu 0003 |
ICLR | 8 |
| 2024 | VBH-GNN: Variational Bayesian Heterogeneous Graph Neural Networks for Cross-subject Emotion RecognitionabstractThe research on human emotion under electroencephalogram (EEG) is an emerging field in which cross-subject emotion recognition (ER) is a promising but challenging task. Many approaches attempt to find emotionally relevant domain-invariant features using domain adaptation (DA) to improve the accuracy of cross-subject ER. However, two problems still exist with these methods. First, only single-modal data (EEG) is utilized, ignoring the complementarity between multi-modal physiological signals. Second, these methods aim to completely match the signal features between different domains, which is difficult due to the extreme individual differences of EEG. To solve these problems, we introduce the complementarity of multi-modal physiological signals and propose a new method for cross-subject ER that does not align the distribution of signal features but rather the distribution of spatio-temporal relationships between features. We design a Variational Bayesian Heterogeneous Graph Neural Network (VBH-GNN) with Relationship Distribution Adaptation (RDA). The RDA first aligns the domains by expressing the model space as a posterior distribution of a heterogeneous graph for a given source domain. Then, the RDA transforms the heterogeneous graph into an emotion-specific graph to further align the domains for the downstream ER task. Extensive experiments on two public datasets, DEAP and Dreamer, show that our VBH-GNN outperforms state-of-the-art methods in cross-subject scenarios. Xinliang Zhou, Zhengri Zhu, Liming Zhai, Ziyu Jia, Yang Liu 0003 |
ICLR | 6 |
| 2024 | Improving Neural Logic Machines via Failure ReflectionabstractReasoning is a fundamental ability towards artificial general intelligence (AGI). Fueled by the success of deep learning, the neural logic machines models (NLMs) have introduced novel neural-symbolic structures and demonstrate great performance and generalization on reasoning and decision-making tasks. However, the original training approaches of the NLMs are still far from perfect, the models would repeat similar mistakes during the training process which leads to sub-optimal performance. To mitigate this issue, we present a novel framework named Failure Reflection Guided Regularizer (FRGR). FRGR first dynamically identifies and summarizes the root cause if the model repeats similar mistakes during training. Then it penalizes the model if it makes similar mistakes in future training iterations. In this way, the model is expected to avoid repeating errors of similar root causes and converge faster to a better-performed optimum. Experimental results on multiple relational reasoning and decision-making tasks demonstrate the effectiveness of FRGR in improving performance, generalization, training efficiency, and data efficiency. Yushi Cao, Yan Zheng 0002, Xu Liu 0014, Bozhi Wu, Tianlin Li, Xiufeng Xu, Junzhe Jiang 0002, Yon Shin Teo, Shangwei Lin 0001, Yang Liu 0003 |
ICML | 11 |
| 2024 | GPTScan: Detecting Logic Vulnerabilities in Smart Contracts by Combining GPT with Program AnalysisabstractSmart contracts are prone to various vulnerabilities, leading to substantial financial losses over time. Current analysis tools mainly target vulnerabilities with fixed control- or data-flow patterns, such as re-entrancy and integer overflow. However, a recent study on Web3 security bugs revealed that about 80% of these bugs cannot be audited by existing tools due to the lack of domain-specific property description and checking. Given recent advances in Large Language Models (LLMs), it is worth exploring how Generative Pre-training Transformer (GPT) could aid in detecting logic vulnerabilities. Yuqiang Sun 0001, Daoyuan Wu, Yue Xue, Han Liu 0012, Haijun Wang 0002, Zhengzi Xu, Xiaofei Xie, Yang Liu 0003 |
ICSE | 8 |
| 2024 | Machine Learning is All You Need: A Simple Token-based Approach for Effective Code Clone DetectionabstractAs software engineering advances and the code demand rises, the prevalence of code clones has increased. This phenomenon poses risks like vulnerability propagation, underscoring the growing importance of code clone detection techniques. While numerous code clone detection methods have been proposed, they often fall short in real-world code environments. They either struggle to identify code clones effectively or demand substantial time and computational resources to handle complex clones. This paper introduces a code clone detection method namely Toma using tokens and machine learning. Specifically, we extract token type sequences and employ six similarity calculation methods to generate feature vectors. These vectors are then input into a trained machine learning model for classification. To evaluate the effectiveness and scalability of Toma, we conduct experiments on the widely used BigCloneBench dataset. Results show that our tool outperforms token-based code clone detectors and most tree-based clone detectors, demonstrating high effectiveness and significant time savings. Siyue Feng, Wenqi Suo, Yueming Wu 0001, Deqing Zou, Yang Liu 0003, Hai Jin 0001 |
ICSE | 5 |
| 2024 | Empirical Analysis of Vulnerabilities Life Cycle in Golang EcosystemabstractOpen-source software (OSS) greatly facilitates program development for developers. However, the high number of vulnerabilities in open-source software is a major concern, including in Golang, a relatively new programming language. In contrast to other commonly used OSS package managers, Golang presents a distinctive feature whereby commits are prevalently used as dependency versions prior to their integration into official releases. This attribute can prove advantageous to users, as patch commits can be implemented in a timely manner before the releases. However, Golang employs a decentralized mechanism for managing dependencies, whereby dependencies are upheld and distributed in separate repositories. This approach can result in delays in the dissemination of patches and unresolved vulnerabilities. Jinchang Hu, Lyuye Zhang, Sen Yang 0018, Yang Liu 0003 |
ICSE | 6 |
| 2024 | RUNNER: Responsible UNfair NEuron Repair for Enhancing Deep Neural Network FairnessabstractDeep Neural Networks (DNNs), an emerging software technology, have achieved impressive results in a variety of fields. However, the discriminatory behaviors towards certain groups (a.k.a. unfairness) of DNN models increasingly become a social concern, especially in high-stake applications such as loan approval and criminal risk assessment. Although there has been a number of works to improve model fairness, most of them adopt an adversary to either expand the model architecture or augment training data, which introduces excessive computational overhead. Recent work diagnoses responsible unfair neurons first and fixes them with selective retraining. Unfortunately, existing diagnosis process is time-consuming due to multi-step training sample analysis, and selective retraining may cause a performance bottleneck due to indirectly adjusting unfair neurons on biased samples. In this paper, we propose Responsible UNfair NEuron Repair (RUNNER) that improves existing works in three key aspects: (1) efficiency: we design the Importance-based Neuron Diagnosis that identifies responsible unfair neurons in one step with a novel importance criterion of neurons; (2) effectiveness: we design the Neuron Stabilizing Retraining by adding a loss term that measures the activation distance of responsible unfair neurons from different subgroups in all sources; (3) generalization: we investigate the effectiveness on both structured tabular data and large-scale unstructured image data, which is often ignored in prior studies. Our extensive experiments across 5 datasets show that RUUNER can effectively and efficiently diagnose and repair the DNNs regarding unfairness. On average, our approach significantly reduces computing overhead from 341.7s to 29.65s, and achieves improved fairness up to 79.3%. Besides, RUNNER also keeps state-of-the-art results on the unstructured dataset. Tianlin Li, Jian Zhang 0087, Shiqian Zhao, Yihao Huang 0001, Aishan Liu, Qing Guo 0005, Yang Liu 0003 |
ICSE | 8 |
| 2024 | On Extracting Specialized Code Abilities from Large Language Models: A Feasibility StudyabstractRecent advances in large language models (LLMs) significantly boost their usage in software engineering. However, training a well-performing LLM demands a substantial workforce for data collection and annotation. Moreover, training datasets may be proprietary or partially open, and the process often requires a costly GPU cluster. The intellectual property value of commercial LLMs makes them attractive targets for imitation attacks, but creating an imitation model with comparable parameters still incurs high costs. This motivates us to explore a practical and novel direction: slicing commercial black-box LLMs using medium-sized backbone models. Zongjie Li, Chaozheng Wang, Pingchuan Ma 0004, Chaowei Liu, Shuai Wang 0011, Daoyuan Wu, Cuiyun Gao 0001, Yang Liu 0003 |
ICSE | 8 |
| 2024 | Demystifying Compiler Unstable Feature Usage and Impacts in the Rust EcosystemabstractRust programming language is gaining popularity rapidly in building reliable and secure systems due to its security guarantees and outstanding performance. To provide extra functionalities, the Rust compiler introduces Rust unstable features (RUF) to extend compiler functionality, syntax, and standard library support. However, these features are unstable and may get removed, introducing compilation failures to dependent packages. Even worse, their impacts propagate through transitive dependencies, causing large-scale failures in the whole ecosystem. Although RUF is widely used in Rust, previous research has primarily concentrated on Rust code safety, with the usage and impacts of RUF from the Rust compiler remaining unexplored. Therefore, we aim to bridge this gap by systematically analyzing the RUF usage and impacts in the Rust ecosystem. We propose novel techniques for extracting RUF precisely, and to assess its impact on the entire ecosystem quantitatively, we accurately resolve package dependencies. We have analyzed the whole Rust ecosystem with 590K package versions and 140M transitive dependencies. Our study shows that the Rust ecosystem uses 1000 different RUF, and at most 44% of package versions are affected by RUF, causing compiling failures for at most 12% of package versions. To mitigate wide RUF impacts, we further design and implement a RUF-compilation-failure recovery tool that can recover up to 90% of the failure. We believe our techniques, findings, and tools can help stabilize the Rust compiler, ultimately enhancing the security and reliability of the Rust ecosystem. Wenbo Shen, Yang Liu 0003, Kui Ren 0001 |
ICSE | 7 |
| 2024 | Semantic-Enhanced Static Vulnerability Detection in Baseband FirmwareabstractCellular network is the infrastructure of mobile communication. Baseband firmware, which carries the implementation of cellular network, has critical security impact on its vulnerabilities. To handle the inherent complexity in cellular communication, cellular protocols are usually implemented as message-centric systems, containing the common message processing phase and message specific handling phase. Though the latter takes most of the code (99.67%) and exposed vulnerabilities (74%), it is rather under-studied: existing detectors either cannot sufficiently analyze it or focused on analyzing the former phase. Cen Zhang, Feng Li 0045, Yeting Li, Jian Wang 0067, Lanlan Zhan, Yang Liu 0003, Wei Huo 0005 |
ICSE | 8 |
| 2024 | An Empirical Study on Noisy Label Learning for Program UnderstandingabstractRecently, deep learning models have been widely applied in program understanding tasks, and these models achieve state-of-the-art results on many benchmark datasets. A major challenge of deep learning for program understanding is that the effectiveness of these approaches depends on the quality of their datasets, and these datasets often contain noisy data samples. A typical kind of noise in program understanding datasets is label noise, which means that the target outputs for some inputs are incorrect. Wenhan Wang, Yanzhou Li, Anran Li 0001, Jian Zhang 0087, Wei Ma 0014, Yang Liu 0003 |
ICSE | 6 |
| 2024 | ModuleGuard: Understanding and Detecting Module Conflicts in Python EcosystemabstractPython has become one of the most popular programming languages for software development due to its simplicity, readability, and versatility. As the Python ecosystem grows, developers face increasing challenges in avoiding module conflicts, which occur when different packages have the same namespace modules. Unfortunately, existing work has neither investigated the module conflict comprehensively nor provided tools to detect the conflict. Therefore, this paper systematically investigates the module conflict problem and its impact on the Python ecosystem. We propose a novel technique called InstSimulator, which leverages semantics and installation simulation to achieve accurate and efficient module extraction. Based on this, we implement a tool called ModuleGuard to detect module conflicts for the Python ecosystem. Ruofan Zhu, Zhengzi Xu, Wenbo Shen, Yang Liu 0003 |
ICSE | 7 |
| 2024 | VSGT: Variational Spatial and Gaussian Temporal Graph Models for EEG-based Emotion Recognition
Xinliang Zhou, Jiaping Xiao, Zhengri Zhu, Liming Zhai, Ziyu Jia, Yang Liu 0003 |
IJCAI | 7 |
| 2024 | Safety-First Tracker: A Trajectory Planning Framework for Omnidirectional Robot TrackingabstractThis paper introduces a Safety-First Tracker (SF-Tracker) designed for omnidirectional autonomous tracking robots. The position and orientation of omnidirectional robots are decoupled for stepwise planning to ensure trajectory safety and maintain target visibility. SF-Tracker puts the trajectory safety in the first place. First, a collision-free and occlusion-free reference path is efficiently initialized by constructing a directed weighted graph. By building upon this path, safe trajectory optimization is implemented to ensure safe movement. Finally, an orientation planner is developed to achieve target visibility based on the safe trajectory. Extensive experimental evaluations in simulated environments and the real world demonstrate that the SF-Tracker outperforms state-of-the-art methods in terms trajectory safety and target visibility. Ablation experiments further demonstrate the significance of each step of the SF-Tracker. The source code and demonstration video can be found at https://github.com/Yue-0/SF-Tracker. Yang Liu 0003, Xin Chen 0032, Dong Wang 0004, Huchuan Lu |
IROS | 2 |
| 2024 | PatchFinder: A Two-Phase Approach to Security Patch Tracing for Disclosed Vulnerabilities in Open-Source SoftwareabstractOpen-source software (OSS) vulnerabilities are increasingly prevalent, emphasizing the importance of security patches. However, in widely used security platforms like NVD, a substantial number of CVE records still lack trace links to patches. Although rank-based approaches have been proposed for security patch tracing, they heavily rely on handcrafted features in a single-step framework, which limits their effectiveness. In this paper, we propose PatchFinder, a two-phase framework with end-to-end correlation learning for better-tracing security patches. In the initial retrieval phase, we employ a hybrid patch retriever to account for both lexical and semantic matching based on the code changes and the description of a CVE, to narrow down the search space by extracting those commits as candidates that are similar to the CVE descriptions. Afterwards, in the re-ranking phase, we design an end-to-end architecture under the supervised fine-tuning paradigm for learning the semantic correlations between CVE descriptions and commits. In this way, we can automatically rank the candidates based on their correlation scores while maintaining low computation overhead. We evaluated our system against 4,789 CVEs from 532 OSS projects. The results are highly promising: PatchFinder achieves a Recall@10 of 80.63% and a Mean Reciprocal Rank (MRR) of 0.7951. Moreover, the Manual Effort@10 required is curtailed to 2.77, marking a 1.94 times improvement over current leading methods. When applying PatchFinder in practice, we initially identified 533 patch commits and submitted them to the official, 482 of which have been confirmed by CVE Numbering Authorities. Kaixuan Li 0002, Jian Zhang 0087, Sen Chen 0001, Han Liu 0012, Yang Liu 0003, Yixiang Chen 0001 |
ISSTA | 5 |
| 2024 | Uncovering and Mitigating the Impact of Code Obfuscation on Dataset Annotation with Antivirus EnginesabstractWith the widespread application of machine learning-based Android malware detection methods, building a high-quality dataset has become increasingly important. Existing large-scale datasets are mostly annotated with VirusTotal by aggregating the decisions of antivirus engines, and most of them indiscriminately accept the decisions of all engines. In reality, however, these engines have different capabilities in detecting malware, especially those that have been obfuscated. Previous research has revealed that code obfuscation degrades the detection performance of these engines to varying degrees. This makes us believe that using all engines indiscriminately is unreasonable for dataset annotation. Therefore, in this paper, we first conduct a data-driven evaluation to confirm the negative effects of code obfuscation on engine-based dataset annotation. To gain a deeper understanding of the reasons behind this phenomenon, we evaluate the availability, effectiveness and robustness of every engine under various code obfuscation techniques. Then we categorize the engines and select a set of obfuscation-robust engines. Finally, we conduct comprehensive experiments to verify the effectiveness of the selected engines for dataset annotation. Our experiments show that when 50% obfuscated samples are mixed into the training set, on the classic malware detectors Drebin and Malscan, using our selected engines can effectively improve detection performance by 15.21% and 19.23%, respectively, compared to using all the engines. Cuiying Gao, Yueming Wu 0001, Heng Li 0008, Wei Yuan 0001, Qidan He, Yang Liu 0003 |
ISSTA | 7 |
| 2024 | DeFort: Automatic Detection and Analysis of Price Manipulation Attacks in DeFi ApplicationsabstractAlthough Decentralized Finance (DeFi) applications facilitate tamper-proof transactions among multiple anonymous users, since attackers can access the smart contract bytecode directly, vulnerabilities in the transaction mechanism, contract code, or third-party components can be easily exploited to manipulate token prices, leading to financial losses. Since price manipulation often relies on specific states and complex trading sequences, existing detection tools have limitations in addressing this problem. In addition, to swiftly identify the root cause of an attack and implement targeted defense and remediation measures, auditors typically prioritize understanding the methodology behind the attack, emphasizing 'how' it occurred rather than simply confirming its existence. To address these problems, this paper presents a novel automatic price manipulation detection and analysis framework, named DeFort, which contains a price manipulation behavior model to guide on-chain detection, multiple price monitoring strategies to detect pools with abnormal token prices, and various profit calculation mechanisms to confirm attacks. Based on behavioral models, DeFort can automatically locate transactions and functions that cause abnormal price fluctuations and identify attackers and victims. Experimental results demonstrate that DeFort can outperform state-of-the-art price manipulation detection methods. Furthermore, after monitoring 441 real-world projects for two months, DeFort successfully detected five price manipulation attacks. Maoyi Xie, Ming Hu 0003, Ziqiao Kong, Cen Zhang, Yebo Feng, Haijun Wang 0002, Yue Xue, Hao Zhang 0004, Ye Liu 0012, Yang Liu 0003 |
ISSTA | 10 |
| 2024 | How Effective Are They? Exploring Large Language Model Based Fuzz Driver GenerationabstractFuzz drivers are essential for library API fuzzing. However, automatically generating fuzz drivers is a complex task, as it demands the creation of high-quality, correct, and robust API usage code. An LLM-based (Large Language Model) approach for generating fuzz drivers is a promising area of research. Unlike traditional program analysis-based generators, this text-based approach is more generalized and capable of harnessing a variety of API usage information, resulting in code that is friendly for human readers. However, there is still a lack of understanding regarding the fundamental issues on this direction, such as its effectiveness and potential challenges. To bridge this gap, we conducted the first in-depth study targeting the important issues of using LLMs to generate effective fuzz drivers. Our study features a curated dataset with 86 fuzz driver generation questions from 30 widely-used C projects. Six prompting strategies are designed and tested across five state-of-the-art LLMs with five different temperature settings. In total, our study evaluated 736,430 generated fuzz drivers, with 0.85 billion token costs ($8,000+ charged tokens). Additionally, we compared the LLM-generated drivers against those utilized in industry, conducting extensive fuzzing experiments (3.75 CPU-year). Our study uncovered that: 1) While LLM-based fuzz driver generation is a promising direction, it still encounters several obstacles towards practical applications; 2) LLMs face difficulties in generating effective fuzz drivers for APIs with intricate specifics. Three featured design choices of prompt strategies can be beneficial: issuing repeat queries, querying with examples, and employing an iterative querying process; 3) While LLM-generated drivers can yield fuzzing outcomes that are on par with those used in the industry, there are substantial opportunities for enhancement, such as extending contained API usage, or integrating semantic oracles to facilitate logical bug detection. Our insights have been implemented to improve the OSS-Fuzz-Gen project, facilitating practical fuzz driver generation in industry. Cen Zhang, Yaowen Zheng, Mingqiang Bai, Yeting Li, Wei Ma 0014, Xiaofei Xie, Yuekang Li, Limin Sun 0001, Yang Liu 0003 |
ISSTA | 9 |
| 2024 | The Software Genome Project: Unraveling Software Through Genetic PrinciplesabstractOpen-source software is crucial to modern development, but its complexity creates challenges in quality, security, and management. Current governance approaches excel at collaboration but struggle with decentralized management and security. With the rise of large language models (LLM)-based software engineering, the need for a finer-grained understanding of software composition is more urgent than ever. To address these challenges, inspired by the Human Genome Project, we treat the software source code as software DNA and propose the Software Genome Project (SGP), which is geared towards the secure monitoring and exploitation of open-source software. By identifying and labeling integrated and classified code features at a fine-grained level, and effectively identifying safeguards for functional implementations and nonfunctional requirements at different levels of granularity, the SGP could build a comprehensive set of software genome maps to help developers and managers gain a deeper understanding of software complexity and diversity. By dissecting and summarizing functional and undesirable genes, SGP could help facilitate targeted software optimization, provide valuable insight and understanding of the entire software ecosystem, and support critical development tasks such as open source governance. SGP could also serve as a comprehensive dataset with abundant semantic labeling to enhance the training of LLMs for code. Based on these, we expect SGP to drive the evolution of software development towards more efficient, reliable, and sustainable software solutions. Yueming Wu 0001, Zhengzi Xu, Lyuye Zhang, Zhiling Zhu, Yang Liu 0003 |
ASE | 7 |
| 2024 | SoVAR: Build Generalizable Scenarios from Accident Reports for Autonomous Driving TestingabstractAutonomous driving systems (ADSs) have undergone remarkable development and are increasingly employed in safety-critical applications. However, recently reported data on fatal accidents involving ADSs suggests that the desired level of safety has not yet been fully achieved. Consequently, there is a growing need for more comprehensive and targeted testing approaches to ensure safe driving. Scenarios from real-world accident reports provide valuable resources for ADS testing, including critical scenarios and high-quality seeds. However, existing scenario reconstruction methods from accident reports often exhibit limited accuracy in information extraction. Moreover, due to the diversity and complexity of road environments, matching current accident information with the simulation map data for reconstruction poses significant challenges. An Guo 0002, Yuan Zhou 0005, Haoxiang Tian 0001, Chunrong Fang, Yunjian Sun, Weisong Sun, Anh Tuan Luu, Yang Liu 0003, Zhenyu Chen 0001 |
ASE | 9 |
| 2024 | Efficient Detection of Toxic Prompts in Large Language ModelsabstractLarge language models (LLMs) like ChatGPT and Gemini have significantly advanced natural language processing, enabling various applications such as chatbots and automated content generation. However, these models can be exploited by malicious individuals who craft toxic prompts to elicit harmful or unethical responses. These individuals often employ jailbreaking techniques to bypass safety mechanisms, highlighting the need for robust toxic prompt detection methods. Existing detection techniques, both blackbox and whitebox, face challenges related to the diversity of toxic prompts, scalability, and computational efficiency. In response, we propose ToxicDetector, a lightweight greybox method designed to efficiently detect toxic prompts in LLMs. ToxicDetector leverages LLMs to create toxic concept prompts, uses embedding vectors to form feature vectors, and employs a Multi-Layer Perceptron (MLP) classifier for prompt classification. Our evaluation on various versions of the LLama models, Gemma-2, and multiple datasets demonstrates that ToxicDetector achieves a high accuracy of 96.39% and a low false positive rate of 2.00%, outperforming state-of-the-art methods. Additionally, ToxicDetector's processing time of 0.0780 seconds per prompt makes it highly suitable for real-time applications. ToxicDetector achieves high accuracy, efficiency, and scalability, making it a practical method for toxic prompt detection in LLMs. Yi Liu 0069, Junzhe Yu, Huijia Sun, Ling Shi 0002, Gelei Deng, Yuqi Chen 0001, Yang Liu 0003 |
ASE | 7 |
| 2024 | Semantic-Enhanced Indirect Call Analysis with Large Language ModelsabstractIn contemporary software development, the widespread use of indirect calls to achieve dynamic features poses challenges in constructing precise control flow graphs (CFGs), which further impacts the performance of downstream static analysis tasks. To tackle this issue, various types of indirect call analyzers have been proposed. However, they do not fully leverage the semantic information of the program, limiting their effectiveness in real-world scenarios. Baijun Cheng, Cen Zhang, Kailong Wang 0001, Ling Shi 0002, Yang Liu 0003, Haoyu Wang 0001, Yao Guo 0001, Ding Li 0001, Xiangqun Chen |
ASE | 5 |
| 2024 | Attribution-guided Adversarial Code Prompt Generation for Code Completion ModelsabstractLarge language models have made significant progress in code completion, which may further remodel future software development. However, these code completion models are found to be highly risky as they may introduce vulnerabilities unintentionally or be induced by a special input, i.e., adversarial code prompt. Prior studies mainly focus on the robustness of these models, but their security has not been fully analyzed. Guozhu Meng, Shangqing Liu, Lu Xiang, Kai Chen 0012, Xiapu Luo, Yang Liu 0003 |
ASE | 8 |
| 2024 | VulAdvisor: Natural Language Suggestion Generation for Software Vulnerability RepairabstractSoftware vulnerabilities pose serious threats to the security of modern software systems. Deep Learning-based Automated Vulnerability Repair (AVR) has gained attention as a potential solution to accelerate the remediation of vulnerabilities. However, recent studies indicate that existing AVR approaches often only generate patches, which may not align with developers' current repair practices or expectations. In this paper, we introduce VulAdvisor, an automated approach that generates natural language suggestions to guide developers or AVR tools in repairing vulnerabilities. VulAdvisor comprises two main components: oracle extraction and suggestion learning. To address the challenge of limited historical data, we propose an oracle extraction method facilitating ChatGPT to construct a comprehensive and high-quality dataset. For suggestion learning, we take the supervised fine-tuning CodeT5 model as the basis, integrating local context into Multi-Head Attention and introducing a repair action loss, to improve the relevance and meaningfulness of the generated suggestions. Extensive experiments on a large-scale dataset from real-world C/C++ projects demonstrate the effectiveness of VulAdvisor, surpassing several alternatives in terms of both lexical and semantic metrics. Moreover, we show that the generated suggestions enhance the patch generation capabilities of existing AVR tools. Human evaluations further validate the quality and utility of VulAdvisor's suggestions, confirming their potential to improve software vulnerability repair practices. Jian Zhang 0087, Chong Wang 0013, Anran Li 0001, Wenhan Wang, Tianlin Li, Yang Liu 0003 |
ASE | 6 |
| 2024 | Is Aggregation the Only Choice? Federated Learning via Layer-wise Model RecombinationabstractAlthough Federated Learning (FL) enables global model training across clients without compromising their raw data, due to the un- evenly distributed data among clients, existing Federated Averaging (FedAvg)-based methods suffer from the problem of low inference performance. Specifically, different data distributions among clients lead to various optimization directions of local models. Aggregat- ing local models usually results in a low-generalized global model, which performs worse on most of the clients. To address the above issue, inspired by the observation from a geometric perspective that a well-generalized solution is located in a flat area rather than a sharp area, we propose a novel and heuristic FL paradigm named FedMR (Federated Model Recombination). The goal of FedMR is to guide the recombined models to be trained towards a flat area. Unlike conventional FedAvg-based methods, in FedMR, the cloud server recombines collected local models by shuffling each layer of them to generate multiple recombined models for local training on clients rather than an aggregated global model. Since the area of the flat area is larger than the sharp area, when local models are located in different areas, recombined models have a higher probability of locating in a flat area. When all recombined models are located in the same flat area, they are optimized towards the same direction. We theoretically analyze the convergence of model recombination. Experimental results show that, compared with state-of-the-art FL methods, FedMR can significantly improve the inference accuracy without exposing the privacy of each client. Ming Hu 0003, Zhihao Yue, Xiaofei Xie, Cheng Chen 0015, Yihao Huang 0001, Xian Wei, Xiang Lian 0001, Yang Liu 0003, Mingsong Chen 0001 |
KDD | 8 |
| 2024 | Enhancing Code Vulnerability Detection via Vulnerability-Preserving Data AugmentationabstractSource code vulnerability detection aims to identify inherent vulnerabilities to safeguard software systems from potential attacks. Many prior studies overlook diverse vulnerability characteristics, simplifying the problem into a binary (0-1) classification task for example determining whether it is vulnerable or not. This poses a challenge for a single deep-learning based model to effectively learn the wide array of vulnerability characteristics. Furthermore, due to the challenges associated with collecting large-scale vulnerability data, these detectors often overfit limited training datasets, resulting in lower model generalization performance. To address the aforementioned challenges, in this work, we introduce a fine-grained vulnerability detector namely FGVulDet. Unlike previous approaches, FGVulDet employs multiple classifiers to discern characteristics of various vulnerability types and combines their outputs to identify the specific type of vulnerability. Each classifier is designed to learn type-specific vulnerability semantics. Additionally, to address the scarcity of data for some vulnerability types and enhance data diversity for learning better vulnerability semantics, we propose a novel vulnerability-preserving data augmentation technique to augment the number of vulnerabilities. Taking inspiration from recent advancements in graph neural networks for learning program semantics, we incorporate a Gated Graph Neural Network (GGNN) and extend it to an edge-aware GGNN to capture edge-type information. FGVulDet is trained on a large-scale dataset from GitHub, encompassing five different types of vulnerabilities. Extensive experiments compared with static-analysis-based approaches and learning-based approaches have demonstrated the effectiveness of FGVulDet. Shangqing Liu, Wei Ma 0014, Jian Wang 0067, Xiaofei Xie, Yang Liu 0003 |
LCTES | 6 |
| 2024 | MASTERKEY: Automated Jailbreaking of Large Language Model Chatbots
Gelei Deng, Yi Liu 0069, Yuekang Li, Kailong Wang 0001, Ying Zhang 0066, Zefeng Li, Haoyu Wang 0001, Tianwei Zhang 0004, Yang Liu 0003 |
NDSS | 9 |
| 2024 | Membership Inference on Text-to-Image Diffusion Models via Conditional Likelihood DiscrepancyabstractText-to-image diffusion models have achieved tremendous success in the field of controllable image generation, while also coming along with issues of privacy leakage and data copyrights. Membership inference arises in these contexts as a potential auditing method for detecting unauthorized data usage. While some efforts have been made on diffusion models, they are not applicable to text-to-image diffusion models due to the high computation overhead and enhanced generalization capabilities. In this paper, we first identify a conditional overfitting phenomenon in text-to-image diffusion models, indicating that these models tend to overfit the conditional distribution of images given the corresponding text rather than the marginal distribution of images only. Based on this observation, we derive an analytical indicator, namely Conditional Likelihood Discrepancy (CLiD), to perform membership inference, which reduces the stochasticity in estimating memorization of individual samples. Experimental results demonstrate that our method significantly outperforms previous methods across various data distributions and dataset scales. Additionally, our method shows superior resistance to overfitting mitigation strategies, such as early stopping and data augmentation. Shengfang Zhai, Huanran Chen, Yinpeng Dong, Qingni Shen, Yansong Gao 0001, Hang Su 0006, Yang Liu 0003 |
NeurIPS | 8 |
| 2024 | Where URLs Become Weapons: Automated Discovery of SSRF Vulnerabilities in Web ApplicationsabstractServer-Side Request Forgery (SSRF) vulnerability poses significant security risks to web applications, enabling adversaries to exploit web applications as stepping stones for unauthorized access of internal-only services or even performing arbitrary commands. Despite its recent emergence as a distinct category in the 2021 OWASP Top 10 web security risks and its increasing prevalence in modern web applications, there remains a lack of effective approaches to detect SSRF vulnerabilities systematically.We present a novel methodology, SSRFuzz, to effectively identify SSRF vulnerability in PHP web applications. Our methodology consists of three phases. In the initial phase, we designed an SSRF oracle to examine functions in PHP manuals and identify sinks that provide server-side request capabilities. This process yielded a total of 86 sensitive PHP sinks out of 2101 PHP functions. The second stage involves dynamic taint inference and the utilization of the identified sinks to examine the source code of target web applications, pinpointing all feasible input points that could trigger these sinks. The final phase employs fuzzing techniques. We generate testing HTTP requests with SSRF payloads, send them to the previously identified input points within the target web applications, and detect if an SSRF vulnerability is triggered. We implemented a prototype of SSRFuzz and evaluated it on 27 real-world applications, including Joomla and WordPress. In total, we discovered 28 SSRF vulnerabilities, 25 of which were previously unreported. We reported all the vulnerabilities to the affected vendors, and 16 new CVE IDs were assigned. Enze Wang, Jianjun Chen 0005, Wei Xie 0007, Chuhan Wang 0001, Hai-Xin Duan, Yang Liu 0003 |
SP | 8 |
| 2024 | PentestGPT: Evaluating and Harnessing Large Language Models for Automated Penetration Testing
Gelei Deng, Yi Liu 0069, Victor Mayoral Vilches, Yuekang Li, Yuan Xu 0033, Martin Pinzger 0001, Stefan Rass, Tianwei Zhang 0004, Yang Liu 0003 |
USENIX Security Symposium | 10 |
| 2024 | FIRE: Combining Multi-Stage Filtering with Taint Analysis for Scalable Recurring Vulnerability Detection
Siyue Feng, Yueming Wu 0001, Wenjie Xue, Sikui Pan, Deqing Zou, Yang Liu 0003, Hai Jin 0001 |
USENIX Security Symposium | 6 |
| 2024 | Using My Functions Should Follow My Checks: Understanding and Detecting Insecure OpenZeppelin Code in Smart Contracts
Han Liu 0012, Daoyuan Wu, Yuqiang Sun 0001, Haijun Wang 0002, Kaixuan Li 0002, Yang Liu 0003, Yixiang Chen 0001 |
USENIX Security Symposium | 6 |
| 2024 | A Fast Weighted Clustering Algorithm for FANETabstractMultiple UAVs working in groups can significantly improve the efficiency in many applications. However, how to group the UAVs adaptively is an non-easy task due to the time-varying environments and tasks requirements. This paper investigates the clustering problem in flying ad hoc network (FANET). To enhance clustering efficiency and ensure rationality and reliability of the clustering structure, we propose a Fast Weighted Clustering Algorithm (FWCA) for node management in FANET. Specifically, we utilize various factors, including remaining energy, ideal node degree difference, node mobility, and link expiration time (LET) to elect cluster heads (CHs). Then, a node clustering mechanism is proposed, including the CH election, clustering process and cluster maintenance. Simulation results demonstrate that the proposed algorithm outperforms the benchmark schemes by reducing the number of CHs and clustering delay, while achieving relatively stable clustering results. Meng Xiao 0002, Zhongliang Zhao, Yang Liu 0003 |
VTC Spring | 4 |
| 2024 | Catch the Butterfly: Peeking into the Terms and Conflicts Among SPDX LicensesabstractThe widespread adoption of third-party libraries (TPLs) in software development has significantly accelerated the creation of modern software. However, this convenience comes with potential legal risks. Developers may inadvertently violate the licenses of TPLs, leading to legal issues. While existing studies have explored software licenses and potential incompatibilities, these studies often focus on a limited set of licenses or rely on low-quality license data, which may affect their conclusions. To address this gap, there is an urgent need for a high-quality license dataset that encompasses a broad range of mainstream licenses and provides accurate terms and conflict information, to help developers navigate the complex landscape of software licenses, avoid potential legal pitfalls, and guide more informed and effective solutions for managing license compliance and compatibility in software development. To this end, we conduct the first work to understand the mainstream software licenses based on term granularity and obtain a high-quality dataset of 453 SPDX licenses with well-labeled terms and conflicts. Specifically, we first conduct a differential analysis of the mainstream platforms that provide license data to understand the terms and attitudes of each license. N ext, we further propose a standardized set of license terms to capture and label existing mainstream licenses with high quality. Moreover, we improve the existing license conflict mode to include copyleft conflicts and conclude the three major types of license conflicts among the 453 SPDX licenses. Based on the dataset, we carry out two empirical studies to reveal the concerns and threats from the perspectives of both licensors and licensees. One study provides an in-depth analysis of the similarities, differences, and conflicts among SPDX licenses, and the other revisits the usage and conflicts of licenses in the NPM ecosystem and draws conclusions that differ from previous work. Our studies reveal some insightful findings and disclose relevant analytical data, which set the stage for further research into the complexities of license compliance and compatibility. Tianwei Liu, He Wang 0014, Gaofei Wu, Yang Liu 0003, Yuqing Zhang 0001 |
SANER | 6 |
| 2024 | Medusa: Unveil Memory Exhaustion DoS Vulnerabilities in Protocol ImplementationsabstractWeb services have brought great convenience to our daily lives. Meanwhile, they are vulnerable to Denial-of-Service (DoS) attacks. DoS attacks launched via vulnerabilities in the services can cause great harm. The vulnerabilities in protocol implementations are especially important because they are the keystones of web services. One vulnerable protocol implementation can affect all the web services built on top of it. Compared to the vulnerabilities that cause the target service to crash, resource exhaustion vulnerabilities are equally if not more important. This is because such vulnerabilities can deplete the system resources, leading to the unavailability of not only the vulnerable service but also other services running on the same machine. Despite the significance of this type of vulnerability, there has been limited research in this area. Zhengjie Du, Yuekang Li, Yaowen Zheng, Cen Zhang, Yi Liu 0069, Sheikh Mahbub Habib, Xinghua Li 0001, Linzhang Wang, Yang Liu 0003, Bing Mao 0001 |
WWW | 10 |
| 2024 | An empirical study of attack-related events in DeFi projects development
Dongming Xiang, Yuanchang Lin, Liming Nie, Yaowen Zheng, Zhengzi Xu, Zuohua Ding, Yang Liu 0003 |
Empir. Softw. Eng. | 7 |
| 2024 | Formally understanding Rust's ownership and borrowing system at the memory level
Shuanglong Kan, Zhe Chen 0011, David Sanán, Yang Liu 0003 |
Formal Methods Syst. Des. | 4 |
| 2024 | Causal deconfounding deep reinforcement learning for mobile robot motion planning
Wenbing Tang 0001, Fenghua Wu, Shang-wei Lin, Zuohua Ding, Jing Liu 0012, Yang Liu 0003, Jifeng He 0001 |
Knowl. Based Syst. | 6 |
| 2024 | CodeBERT-Attack: Adversarial attack against source code deep learning models via pre-trained modelabstractAbstract Over the past few years, the software engineering (SE) community has widely employed deep learning (DL) techniques in many source code processing tasks. Similar to other domains like computer vision and natural language processing (NLP), the state‐of‐the‐art DL techniques for source code processing can still suffer from adversarial vulnerability, where minor code perturbations can mislead a DL model's inference. Efficiently detecting such vulnerability to expose the risks at an early stage is an essential step and of great importance for further enhancement. This paper proposes a novel black‐box effective and high‐quality adversarial attack method, namely CodeBERT‐Attack (CBA), based on the powerful large pre‐trained model (i.e., CodeBERT) for DL models of source code processing. CBA locates the vulnerable positions through masking and leverages the power of CodeBERT to generate textual preserving perturbations. We turn CodeBERT against DL models and further fine‐tuned CodeBERT models for specific downstream tasks, and successfully mislead these victim models to erroneous outputs. In addition, taking the power of CodeBERT, CBA is capable of effectively generating adversarial examples that are less perceptible to programmers. Our in‐depth evaluation on two typical source code classification tasks (i.e., functionality classification and code clone detection) against the most widely adopted LSTM and the powerful fine‐tuned CodeBERT models demonstrate the advantages of our proposed technique in terms of both effectiveness and efficiency. Furthermore, our results also show (1) that pre‐training may help CodeBERT gain resilience against perturbations further, and (2) certain pre‐training tasks may be beneficial for adversarial robustness. Huangzhao Zhang, Zhuo Li 0013, Zhi Jin 0001, Lei Ma 0003, Yang Liu 0003, Ge Li 0001 |
J. Softw. Evol. Process. | 6 |
| 2024 | ESB-FL: Efficient and Secure Blockchain-Based Federated Learning With Fair PaymentabstractFederated learning is a technique that enables multiple parties to collaboratively train a model without sharing raw private data, and it is ideal for smart healthcare. However, it raises new privacy concerns due to the risk of privacy-sensitive medical data leakage. It is not until recently that the privacy-preserving FL (PPFL) has been introduced as a solution to ensure the privacy of training processes. Unfortunately, most existing PPFL schemes are highly dependent on complex cryptographic mechanisms or fail to guarantee the accuracy of training models. Besides, there has been little research on the fairness of the payment procedure in the PPFL with incentive mechanisms. To address the above concerns, we first construct an efficient non-interactive designated decryptor function encryption (NDD-FE) scheme to protect the privacy of training data while maintaining high communication performance. We then propose a blockchain-based PPFL framework with fair payment for medical image detection, namely ESB-FL, by combining the NDD-FE and an elaborately designed blockchain. ESB-FL not only inherits the characteristics of the NDD-FE scheme, but it also ensures the interests of each participant. We finally conduct extensive security analysis and experiments to show that our new framework has enhanced security, good accuracy, and high efficiency. Biwen Chen, Honghong Zeng, Tao Xiang 0001, Shangwei Guo, Tianwei Zhang 0004, Yang Liu 0003 |
IEEE Trans. Big Data | 6 |
| 2024 | A Secure and Robust Knowledge Transfer Framework via Stratified-Causality Distribution Adjustment in Intelligent Collaborative ServicesabstractThe rapid development of device-edge-cloud collaborative computing techniques has actively contributed to the popularization and application of intelligent service models. The intensity of knowledge transfer plays a vital role in enhancing the performance of intelligent services. However, the existing knowledge transfer methods are mainly implemented through data fine-tuning and model distillation, which may cause the leakage of data privacy or model copyright in intelligent collaborative systems. To address this issue, we propose a secure and robust knowledge transfer framework through stratified-causality distribution adjustment (SCDA) for device-edge-cloud collaborative services. Specifically, a simple yet effective density-based estimation is first employed to obtain uncertainty scores that guide the space stratification, which is conducive to reconstructing low-density distribution regions from high-density distribution regions more adaptively and accurately. Subsequently, we devise a novel causality-aware generative model to generate synthetic features for the out-of-distribution domain by exploring the relationship between factors and variables. Ultimately, we introduce a cycle-consistent minimax optimization mechanism to ensure the effectiveness and dependability of knowledge transfer through the influence minimization and the diversity maximization. Furthermore, extensive experiments demonstrate that our scheme can protect the security of data privacy and model copyright in intelligent collaborative services through adaptive distribution adjustment. Ju Jia, Siqi Ma 0001, Lina Wang 0001, Yang Liu 0003, Robert H. Deng |
IEEE Trans. Computers | 4 |
| 2024 | Dodging DeepFake Detection via Implicit Spatial-Domain Notch FilteringabstractThe current high-fidelity generation and high-precision detection of DeepFake images are at an arms race. We believe that producing DeepFakes that are highly realistic and “detection evasive” can serve the ultimate goal of improving future generation DeepFake detection capabilities. In this paper, we propose a simple yet powerful pipeline to reduce the artifact patterns of fake images without hurting image quality by performing implicit spatial-domain notch filtering. We first demonstrate that frequency-domain notch filtering, although famously shown to be effective in removing periodic noise in the spatial domain, is infeasible for our task at hand due to the manual designs required for the notch filters. We, therefore, resort to a learning-based approach to reproduce the notch filtering effects, but solely in the spatial domain. We adopt a combination of adding overwhelming spatial noise for breaking the periodic noise pattern and deep image filtering to reconstruct the noise-free fake images, and we name our method DeepNotch. Deep image filtering provides a specialized filter for each pixel in the noisy image, producing filtered images with high fidelity compared to their DeepFake counterparts. Moreover, we also use the semantic information of the image to generate an adversarial guidance map to add noise intelligently. Our large-scale evaluation on 3 representative DeepFake detection methods (tested on 16 types of DeepFakes) has demonstrated that our technique significantly reduces the accuracy of these 3 fake image detection methods, 36.79% on average and up to 97.02% in the best case. Yihao Huang 0001, Felix Juefei-Xu, Qing Guo 0005, Yang Liu 0003, Geguang Pu |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Adversarial Learning for Coordinate Regression Through $k$k-Layer Penetrating RepresentationabstractAdversarial attack is a crucial step when evaluating the reliability and robustness of deep neural networks (DNNs) models. Most existing attack approaches apply an end-to-end gradient update strategy to generate adversarial examples for a classification or regression problem. However, few of them consider the non-differentiable DNN models (e.g., coordinate regression model) that prevent end-to-end backpropagation resulting in the failure of gradient calculation. In this article, we present a new adversarial example generation approach for both untargeted and targeted attacks on coordinate regression models with non-differentiable operations. The novelty of our approach lies in a$k$-layer penetrating representation, on which we perturb the hidden feature distribution of the$k$th layer through relational guidance to influence the final output, in which end-to-end backpropagation is not required. Rather than modifying a large portion of the pixels in an image, the proposed approach only modifies a very small set of the input pixels. These pixels are carefully and precisely selected by three correlations between the input pixels and hidden features of the$k$th layer of a DNN, thus significantly reducing the adversarial perturbation on a clean image. We successfully apply the proposed approach to two different tasks (i.e., 2D and 3D human pose estimation) which are typical applications of the coordinate regression learning. The comprehensive experiments demonstrate that our approach achieves better performance while using much less adversarial perturbation on clean images. Mengxi Jiang, Yulei Sui, Xiaofei Xie, Cuihua Li, Yang Liu 0003, Ivor W. Tsang |
IEEE Trans. Dependable Secur. Comput. | 6 |
| 2024 | Ambush From All Sides: Understanding Security Threats in Open-Source Software CI/CD PipelinesabstractThe continuous integration and continuous deployment (CI/CD) pipeline has been widely used and is becoming popular on Internet hosting platforms, such as GitHub. While being popular, however, current CI/CD pipelines suffer from malicious code and severe vulnerabilities. Even worse, it is often under-protected as people have not been fully aware of its attack surfaces and the corresponding impacts. Therefore, in this paper, we conduct a large-scale measurement and a systematic analysis to reveal the attack surfaces of the CI/CD pipeline and quantify their security impacts. Specifically, for the measurement, we collect a data set of 320,000+ CI/CD pipeline-configured GitHub repositories and build an analysis tool to parse the CI/CD pipelines and extract security-critical usages. Our measurement reveals that the script runtimes are prone to code hiding while the script usage update is not in time, giving attackers chances to hide malicious code and exploit existing vulnerabilities. Moreover, even the scripts from verified creators may contain severe vulnerabilities. Besides current CI/CD ecosystem heavily relies on several core scripts, which may lead to a single point of failure. While the CI/CD pipelines contain sensitive information/operations, making them the attacker's favorite targets. Inspired by the measurement findings, we abstract the threat model and the attack approach toward CI/CD pipelines, followed by a systematic analysis of attack surfaces, attack strategies, and the corresponding impacts. We further launch case studies on five attacks in real-world CI/CD environments to validate the revealed attack surfaces. Finally, we give suggestions on mitigating attacks on CI/CD scripts, including securing CI/CD configurations, securing CI/CD scripts, and improving CI/CD infrastructure. Ziyue Pan, Wenbo Shen, Yutian Yang, Yao Liu 0007, Yang Liu 0003, Kui Ren 0001 |
IEEE Trans. Dependable Secur. Comput. | 8 |
| 2024 | Texture Re-Scalable Universal Adversarial PerturbationabstractUniversal adversarial perturbation (UAP), also known as image-agnostic perturbation, is a fixed perturbation map that can fool the classifier with high probabilities on arbitrary images, making it more practical for attacking deep models in the real world. Previous UAP methods generate a scale-fixed and texture-fixed perturbation map for all images, which ignores the multi-scale objects in images and usually results in a low fooling ratio. Since the widely used convolution neural networks tend to classify objects according to semantic information stored in local textures, it seems a reasonable and intuitive way to improve the UAP from the perspective of utilizing local contents effectively. In this work, we find that the fooling ratios significantly increase when we add a constraint to encourage a small-scale UAP map and repeat it vertically and horizontally to fill the whole image domain. To this end, we propose texture scale-constrained UAP (TSC-UAP), a simple yet effective UAP enhancement method that automatically generates UAPs with category-specific local textures that can fool deep models more easily. Through a low-cost operation that restricts the texture scale, TSC-UAP achieves a considerable improvement in the fooling ratio and attack transferability for both data-dependent and data-free UAP methods. Experiments conducted on two state-of-the-art UAP methods, eight popular CNN models and four classical datasets show the remarkable performance of TSC-UAP. Yihao Huang 0001, Qing Guo 0005, Felix Juefei-Xu, Ming Hu 0003, Xiaojun Jia, Xiaochun Cao, Geguang Pu, Yang Liu 0003 |
IEEE Trans. Inf. Forensics Secur. | 8 |
| 2024 | Revisiting and Exploring Efficient Fast Adversarial Training via LAW: Lipschitz Regularization and Auto Weight AveragingabstractFast Adversarial Training (FAT) not only improves the model robustness but also reduces the training cost of standard adversarial training. However, fast adversarial training often suffers from Catastrophic Overfitting (CO), which results in poor robustness performance. Catastrophic Overfitting describes the phenomenon of a sudden and significant decrease in robust accuracy during the training of fast adversarial training. Many effective techniques have been developed to prevent Catastrophic Overfitting and improve the model robustness from different perspectives. However, these techniques adopt inconsistent training settings and require different training costs, i.e., training time and memory costs, leading to unfair comparisons. In this paper, we conduct a comprehensive study of over 10 fast adversarial training methods in terms of adversarial robustness and training costs. We revisit the effectiveness and efficiency of fast adversarial training techniques in preventing Catastrophic Overfitting from the perspective of model local nonlinearity and propose an effective Lipschitz regularization method for fast adversarial training. Furthermore, we explore the effect of data augmentation and weight averaging in fast adversarial training and propose a simple yet effective auto weight averaging method to improve robustness further. By assembling these techniques, we propose an effective FGSM-based fast adversarial training method equipped with Lipschitz regularization and Auto Weight averaging, abbreviated as FGSM-LAW. Experimental evaluations on four benchmark databases demonstrate the superiority of our method over state-of-the-art fast adversarial training methods and the advanced standard adversarial training methods. Xiaojun Jia, Yuefeng Chen, Xiaofeng Mao, Ranjie Duan, Jindong Gu, Rong Zhang 0006, Hui Xue 0001, Yang Liu 0003, Xiaochun Cao |
IEEE Trans. Inf. Forensics Secur. | 8 |
| 2024 | A Causality-Aligned Structure Rationalization Scheme Against Adversarial Biased Perturbations for Graph Neural NetworksabstractThe graph neural networks (GNNs) are susceptible to adversarial perturbations and distribution biases, which pose potential security concerns for real-world applications. Current endeavors mainly focus on graph matching, while the subtle relationships between the nodes and structures of graph-structured data remain under-explored. Accordingly, two fundamental challenges arise as follows: 1) the intricate connections among nodes may induce the distribution shift of graph samples even under the same scenario, and 2) the perturbations of inherent graph-structured representations can introduce spurious shortcuts, which lead to GNN models relying on biased data to make unstable predictions. To address these problems, we propose a novel causality-aligned structure rationalization (CASR) scheme to construct invariant rationales by probing the coherent and causal patterns, which facilitates GNN models to make stable and reliable predictions in case of adversarial biased perturbations. Specifically, the initial graph samples across domains are leveraged to boost the diversity of datasets and perceive the interaction between shortcuts. Subsequently, the causal invariant rationales can be obtained during the interventions. This allows the GNN model to extrapolate risk variations from a single observed environment to multiple unknown environments. Moreover, the query feedback mechanism can progressively promote the consistency-driven optimal rationalization by reinforcing real essences and eliminating spurious shortcuts. Extensive experiments demonstrate the effectiveness of our scheme against adversarial biased perturbations from data manipulation attacks and out-of-distribution (OOD) shifts on various graph-structured datasets. Notably, we reveal that the capture of distinctive rationales can greatly reduce the dependence on shortcut cues and improve the robustness of OOD generalization. Ju Jia, Siqi Ma 0001, Yang Liu 0003, Lina Wang 0001, Robert H. Deng |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2024 | Robust Motion Planning for Multi-Robot Systems Against Position Deception AttacksabstractDeep reinforcement learning (DRL) is widely applied in motion planning for multi-robot systems as DRL leverages the offline training process to improve the real-time computation efficiency. In DRL-based methods, the DRL models compute an action for a robot based on the states of its surrounding obstacles, including other robots in the system. They always assume that the number of obstacles is fixed and the obtained obstacles’ states are reliable. However, in the real world, a multi-robot system may suffer from various attacks, such as remote control attacks and network attacks, that cause wrong positions of the surrounding obstacles received by a robot. In this paper, we propose a robust motion planning methodDAE-Crit-LSTM, integrating a denoising autoencoder (DAE) with DRL models, to mitigate such position deception attacks in environments with a different number of obstacles.DAE-Crit-LSTMshows the following two advantages. First,DAE-Crit-LSTMcan be applied in benign and attacked scenarios and thus does not require any detector. It learns an encoder and a decoder to approximate the accurate positions of the obstacles, no matter under attack or not. Second,DAE-Crit-LSTMapplies an LSTM (Long Short-Term Memory)-based DRL model to deal with a variable number of obstacles in the environment. It is worth noting thatDAE-Crit-LSTMis method-agnostic and can be easily implemented in state-of-the-art motion planning methods. Comprehensive experiments show thatDAE-Crit-LSTMcan mitigate position deception attacks and guarantee safe motion. We also demonstrate the effectiveness and generalization ofDAE-Crit-LSTM. Wenbing Tang 0001, Yuan Zhou 0005, Yang Liu 0003, Zuohua Ding, Jing Liu 0012 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2024 | Rethinking Membership Inference Attacks Against Transfer LearningabstractTransfer learning, successful in knowledge translation across related tasks, faces a substantial privacy threat from membership inference attacks (MIAs). These attacks, despite posing significant risk to ML model’s training data, remain limited-explored in transfer learning. The interaction between teacher and student models in transfer learning has not been thoroughly explored in MIAs, potentially resulting in an under-examined aspect of privacy vulnerabilities within transfer learning. In this paper, we propose a new MIA vector against transfer learning, to determine whether a specific data point was used to train the teacher model while only accessing the student model in a white-box setting. Our method delves into the intricate relationship between teacher and student models, analyzing the discrepancies in hidden layer representations between the student model and its shadow counterpart. These identified differences are then adeptly utilized to refine the shadow model’s training process and to inform membership inference decisions effectively. Our method, evaluated across four datasets in diverse transfer learning tasks, reveals that even when an attacker only has access to the student model, the teacher model’s training data remains susceptible to MIAs. We believe our work unveils the unexplored risk of membership inference in transfer learning. Cong Wu 0003, Jing Chen 0003, Qianru Fang, Kun He 0008, Ziming Zhao 0001, Hao Ren 0001, Guowen Xu, Yang Liu 0003, Yang Xiang 0001 |
IEEE Trans. Inf. Forensics Secur. | 8 |
| 2024 | COCL: An Intelligent Framework for Enhancing Deep Learning-Based Vulnerability DetectionabstractDue to the powerful feature extraction capability ofdeep learning(DL), many recent studies have used it to conduct source code vulnerability analysis. However, although it has a good performance on artificial datasets, it does not perform satisfactorily on the real-world vulnerabilities with higher complexity. In this article, we introduce contrastive curriculum learning into DL-based vulnerability detection to find a suitable boundary to distinguish vulnerabilities from normal codes. Contrastive learning can be used to reduce the difference between different vulnerabilities while amplifying the difference between vulnerabilities and normal codes. To make the training phase of contrastive learning more intelligent, we apply curriculum learning to mimic the way humans acquire knowledge, which means that the model will learn simple samples first and then increase the difficulty of training samples. Specifically, we implement an intelligent framework (i.e.,contrastive curriculum learning (COCL)) that can enhance the detection effect of existing DL-based vulnerability detectors. To verify the capability ofCOCL, we select four state-of-the-art DL-based vulnerability detectors (i.e.,AutoVulTC,VulDeePecker,BenchSG, andDevign) as our base models. The experimental results show that usingCOCLcan bring an improvement of 8.1% to the F1 scores of these models on a real-world vulnerability dataset. Shihan Dou, Yueming Wu 0001, Yang Liu 0003 |
IEEE Trans. Ind. Informatics | 5 |
| 2024 | Reinforcement Learning Based Online Request Scheduling Framework for Workload-Adaptive Edge Deep Learning InferenceabstractThe recent advances of deep learning in various mobile and Internet-of-Things applications, coupled with the emergence of edge computing, have led to a strong trend of performing deep learning inference on the edge servers located physically close to the end devices. This trend presents the challenge of how to meet the quality-of-service requirements of inference tasks at the resource-constrained network edge, especially under variable or even bursty inference workloads. Solutions to this challenge have not yet been reported in the related literature. In the present paper, we tackle this challenge by means of workload-adaptive inference request scheduling: in different workload states, via adaptive inference request scheduling policies, different models with diverse model sizes can play different roles to maintain high-quality inference services. To implement this idea, we propose a request scheduling framework for general-purpose edge inference serving systems. Theoretically, we prove that, in our framework, the problem of optimizing the inference request scheduling policies can be formulated as a Markov decision process (MDP). To tackle such an MDP, we use reinforcement learning and propose a policy optimization approach. Through extensive experiments, we empirically demonstrate the effectiveness of our framework in the challenging practical case where the MDP is partially observable. Xinrui Tan, Hongjia Li 0002, Xiaofei Xie, Nirwan Ansari, Xueqing Huang, Liming Wang 0001, Zhen Xu 0009, Yang Liu 0003 |
IEEE Trans. Mob. Comput. | 9 |
| 2024 | It's All in the Touch: Authenticating Users With HOST Gestures on Multi-Touch Screen DevicesabstractAs smartphones proliferate, secure and user-friendly authentication methods are increasingly critical. Existing behavioral biometrics, however, are often compromised by behavior variability, leading to poor authentication accuracy and an unsatisfactory user experience. To fill this gap, we proposeBioHold, a new robust and reliable user authentication method, fusing finger behavior and hand geometry, captured via a smartphone's multitouch screen during natural holding gestures. It synergistically fuses behavioral and physiological biometrics. In contrast to traditional methods that require restrictive, unnatural user patterns, our approach utilizes a stable, natural gesture for authentication, effectively mitigating behavior variability. It enables one-handed authentication through familiar smartphone-holding and unlocking gestures. During this interaction, hand geometry and behavioral characteristics are recorded for subsequent authentication. We evaluate our method using a dataset collected from 20 subjects, demonstrating its resilience against behavioral variability over time while maintaining a high level of distinctiveness. With only 10 training samples, our method achieves an equal error rate of 3.59%, which improves to 1.25% with 40 training samples. Importantly, our method is resistant to common security threats such as zero-effort attacks, smudge attacks, and shoulder surfing attacks. A usability study confirms the method's high user acceptance, as measured by the system usability score. Cong Wu 0003, Hangcheng Cao, Guowen Xu, Jianfei Sun, Ran Yan 0001, Yang Liu 0003, Hongbo Jiang 0001 |
IEEE Trans. Mob. Comput. | 7 |
| 2024 | Natural & Adversarial Bokeh Rendering via Circle-of-Confusion Predictive NetworkabstractBokeh effect is a natural shallow depth-of-field phenomenon that blurs the out-of-focus part in photography. In recent years, a series of works have proposed automatic and realistic bokeh rendering methods for artistic and aesthetic purposes. They usually employ cutting-edge data-driven deep generative networks with complex training strategies and network architectures. However, these works neglect that the bokeh effect, as a real phenomenon, can inevitably affect the subsequent visual intelligent tasks like recognition, and their data-driven nature prevents them from studying the influence of bokeh-related physical parameters (i.e., depth-of-the-field) on the intelligent tasks. To fill this gap, we study a totally new problem, i.e.,natural & adversarial bokeh rendering, which consists of two objectives: rendering realistic and natural bokeh and fooling the visual perception models (i.e., bokeh-based adversarial attack). To this end, beyond the pure data-driven solution, we propose a hybrid alternative by taking the respective advantages of data-driven and physical-aware methods. Specifically, we propose thecircle-of-confusion predictive network (CoCNet)by taking the all-in-focus image and depth image as inputs to estimate circle-of-confusion parameters for each pixel, which are employed to render the final image through a well-known physical model of bokeh. With the hybrid solution, our method could achieve more realistic rendering results with the naive training strategy and a much lighter network. Moreover, we propose the adversarial bokeh attack by fixing the CoCNet while optimizing the depth map w.r.t. the visual perception tasks. Then, we are able to study the vulnerability of deep neural networks according to the depth variations in the real world. The extensive experiments show that our method produces more realistic bokeh than the state-of-the-art methods while fooling the powerful deep neural networks with a high accuracy drop. Yihao Huang 0001, Felix Juefei-Xu, Qing Guo 0005, Geguang Pu, Yang Liu 0003 |
IEEE Trans. Multim. | 5 |
| 2024 | Faire: Repairing Fairness of Neural Networks via Neuron Condition SynthesisabstractDeep Neural Networks (DNNs) have achieved tremendous success in many applications, while it has been demonstrated that DNNs can exhibit some undesirable behaviors on concerns such as robustness, privacy, and other trustworthiness issues. Among them, fairness (i.e., non-discrimination) is one important property, especially when they are applied to some sensitive applications (e.g., finance and employment). However, DNNs easily learn spurious correlations between protected attributes (e.g., age, gender, race) and the classification task and develop discriminatory behaviors if the training data is imbalanced. Such discriminatory decisions in sensitive applications would introduce severe social impacts. To expose potential discrimination problems in DNNs before putting them in use, some testing techniques have been proposed to identify the discriminatory instances (i.e., instances that show defined discrimination 1 ). However, how to repair DNNs after detecting such discrimination is still challenging. Existing techniques mainly rely on retraining on a large number of discriminatory instances generated by testing methods, which requires huge time overhead and makes the repairing inefficient. In this work, we propose the method Faire to effectively and efficiently repair the fairness issues of DNNs, without using additional data (e.g., discriminatory instances). Our basic idea is inspired by the traditional program repair method that synthesizes proper condition checking. To repair traditional programs, a typical method is to localize the program defects and repair the program logic by adding condition checking. Similarly, for DNNs, we try to understand the unfair logic and reformulate it with well-designed condition checking. In this article, we synthesize the condition that can reduce the effect of features relevant to the protected attributes in the DNN. Specifically, we first perform the neuron-based analysis and check the functionalities of neurons to identify neurons whose outputs could be regarded as features relevant to protected attributes and original tasks. Then a new condition layer is added after each hidden layer to penalize neurons that are accountable for the protected features (i.e., intermediate features relevant to protected attributes) and promote neurons that are accountable for the non-protected features (i.e., intermediate features relevant to original tasks). In sum, the repair rate 2 of Faire reaches up to more than 99%, which outperforms other methods, and the whole repairing process only takes no more than 340 s. The evaluation results demonstrate that our approach can effectively and efficiently repair the individual discriminatory instances of the target model. Tianlin Li, Xiaofei Xie, Jian Wang 0067, Qing Guo 0005, Aishan Liu, Lei Ma 0003, Yang Liu 0003 |
ACM Trans. Softw. Eng. Methodol. | 7 |
| 2024 | Automated Commit Intelligence by Pre-trainingabstractGitHub commits, which record the code changes with natural language messages for description, play a critical role in software developers’ comprehension of software evolution. Due to their importance in software development, several learning-based works are conducted for GitHub commits, such as commit message generation and security patch identification. However, most existing works focus on customizing specialized neural networks for different tasks. Inspired by the superiority of code pre-trained models, which has confirmed their effectiveness across different downstream tasks, to promote the development of open-source software community, we first collect a large-scale commit benchmark including over 7.99 million commits across 7 programming languages. Based on this benchmark, we present CommitBART, a pre-trained encoder-decoder Transformer model for GitHub commits. The model is pre-trained by three categories (i.e., denoising objectives, cross-modal generation, and contrastive learning) for six pre-training tasks to learn commit fragment representations. Our model is evaluated on one understanding task and three generation tasks for commits. The comprehensive experiments on these tasks demonstrate that CommitBART significantly outperforms previous pre-trained works for code. Further analysis also reveals that each pre-training task enhances the model performance. Shangqing Liu, Yanzhou Li, Xiaofei Xie, Wei Ma 0014, Guozhu Meng, Yang Liu 0003 |
ACM Trans. Softw. Eng. Methodol. | 6 |
| 2024 | Unveiling Code Pre-Trained Models: Investigating Syntax and Semantics CapacitiesabstractCode models have made significant advancements in code intelligence by encoding knowledge about programming languages. While previous studies have explored the capabilities of these models in learning code syntax, there has been limited investigation on their ability to understand code semantics. Additionally, existing analyses assume that the number of edges between nodes at the abstract syntax tree (AST) is related to syntax distance, and also often require transforming the high-dimensional space of deep learning models to a low-dimensional one, which may introduce inaccuracies. To study how code models represent code syntax and semantics, we conduct a comprehensive analysis of seven code models, including four representative code pre-trained models (CodeBERT, GraphCodeBERT, CodeT5, and UnixCoder) and three large language models (LLMs) (StarCoder, CodeLlama and CodeT5+). We design four probing tasks to assess the models’ capacities in learning both code syntax and semantics. These probing tasks reconstruct code syntax and semantics structures (AST, control dependence graph (CDG), data dependence graph (DDG), and control flow graph (CFG)) in the representation space. These structures are core concepts for code understanding. We also investigate the syntax token role in each token representation and the long dependency between the code tokens. Additionally, we analyze the distribution of attention weights related to code semantic structures. Through extensive analysis, our findings highlight the strengths and limitations of different code models in learning code syntax and semantics. The results demonstrate that these models excel in learning code syntax, successfully capturing the syntax relationships between tokens and the syntax roles of individual tokens. However, their performance in encoding code semantics varies. CodeT5 and CodeBERT demonstrate proficiency in capturing control and data dependencies, whereas UnixCoder shows weaker performance in this aspect. We do not observe LLMs generally performing much better than pre-trained models. The shallow layers of LLMs perform better than their deep layers. The investigation of attention weights reveals that different attention heads play distinct roles in encoding code semantics. Our research findings emphasize the need for further enhancements in code models to better learn code semantics. This study contributes to the understanding of code models’ abilities in syntax and semantics analysis. Our findings provide guidance for future improvements in code models, facilitating their effective application in various code-related tasks. Wei Ma 0014, Shangqing Liu, Xiaofei Xie, Wenhan Wang, Jie Zhang 0050, Yang Liu 0003 |
ACM Trans. Softw. Eng. Methodol. | 8 |
| 2024 | A Survey of Source Code Search: A 3-Dimensional Perspectiveabstract(Source) code search is widely concerned by software engineering researchers because it can improve the productivity and quality of software development. Given a functionality requirement usually described in a natural language sentence, a code search system can retrieve code snippets that satisfy the requirement from a large-scale code corpus, e.g., GitHub. To realize effective and efficient code search, many techniques have been proposed successively. These techniques improve code search performance mainly by optimizing three core components, including query understanding component, code understanding component, and query-code matching component. In this article, we provide a 3-dimensional perspective survey for code search. Specifically, we categorize existing code search studies into query-end optimization techniques, code-end optimization techniques, and match-end optimization techniques according to the specific components they optimize. These optimization techniques are proposed to enhance the performance of specific components, and thus the overall performance of code search. Considering that each end can be optimized independently and contributes to the code search performance, we treat each end as a dimension. Therefore, this survey is 3-dimensional in nature, and it provides a comprehensive summary of each dimension in detail. To understand the research trends of the three dimensions in existing code search studies, we systematically review 68 relevant literatures. Different from existing code search surveys that only focus on the query end or code end or introduce various aspects shallowly (including codebase, evaluation metrics, modeling technique, etc.), our survey provides a more nuanced analysis and review of the evolution and development of the underlying techniques used in the three ends. Based on a systematic review and summary of existing work, we outline several open challenges and opportunities at the three ends that remain to be addressed in future work. Weisong Sun, Chunrong Fang, Yifei Ge, Yuling Hu, Quanjun Zhang, Xiuting Ge, Yang Liu 0003, Zhenyu Chen 0001 |
ACM Trans. Softw. Eng. Methodol. | 8 |
| 2024 | Esale: Enhancing Code-Summary Alignment Learning for Source Code Summarizationabstract(Source) code summarization aims to automatically generate succinct natural language summaries for given code snippets. Such summaries play a significant role in promoting developers to understand and maintain code. Inspired by neural machine translation, deep learning-based code summarization techniques widely adopt an encoder-decoder framework, where the encoder transforms given code snippets into context vectors, and the decoder decodes context vectors into summaries. Recently, large-scale pre-trained models for source code (e.g., CodeBERT and UniXcoder) are equipped with encoders capable of producing general context vectors and have achieved substantial improvements on the code summarization task. However, although they are usually trained mainly on code-focused tasks and can capture general code features, they still fall short in capturing specific features that need to be summarized. In a nutshell, they fail to learn the alignment between code snippets and summaries (code-summary alignment for short). In this paper, we propose a novel approach to improve code summarization based on summary-focused tasks. Specifically, we exploit a multi-task learning paradigm to train the encoder on three summary-focused tasks to enhance its ability to learn code-summary alignment, including unidirectional language modeling (ULM), masked language modeling (MLM), and action word prediction (AWP). Unlike pre-trained models that mainly predict masked tokens in code snippets, we design ULM and MLM to predict masked words in summaries. Intuitively, predicting words based on given code snippets would help learn the code-summary alignment. In addition, existing work shows that AWP affects the prediction of the entire summary. Therefore, we further introduce the domain-specific task AWP to enhance the ability of the encoder to learn the alignment between action words and code snippets. We evaluate the effectiveness of our approach, calledEsale, by conducting extensive experiments on four datasets, including two widely used datasets JCSD and PCSD, a cross-project Java dataset CPJD, and a multilingual language dataset CodeSearchNet. Experimental results show thatEsalesignificantly outperforms state-of-the-art baselines in all three widely used metrics, including BLEU, METEOR, and ROUGE-L. Moreover, the human evaluation proves that the summaries generated byEsaleare more informative and closer to the ground-truth summaries. Chunrong Fang, Weisong Sun, Zhao Wei, Quanjun Zhang, Yudu You, Bin Luo 0003, Yang Liu 0003, Zhenyu Chen 0001 |
IEEE Trans. Software Eng. | 9 |
| 2024 | Distributed Motion Control for Multiple Mobile Robots Using Discrete-Event Systems and Model Predictive ControlabstractDistributed motion control is critical in multiple mobile robot systems (MMRSs). Current research usually focuses on either discrete approaches, which aim to deal with high-level collisions and deadlocks without considering the low-level motion commands, or continuous approaches, which can optimize low-level continuous commands to mobile robots but cannot deal with deadlocks efficiently. In this article, by combining discrete and continuous methods, we design a hybrid motion control method for MMRSs where each robot should move along a predefined path. First, each robot’s motion is modeled as a discrete transition system, based on which a real-time supervisory control policy is illustrated to avoid collisions and deadlocks. Second, according to the discrete decisions, the continuous speed at each discrete state is computed using model predictive control and sequential convex programming. The proposed hybrid approach brings two advantages. First, the discrete control component guarantees collision and deadlock avoidance and reduces the scale of the optimization problems. Second, continuous control optimizes the continuous speed in real time and fulfills other performance requirements like time and energy costs. To move in a fully distributed way, each robot needs to predict the motion of its neighbors by retrieving their immediately available information through communications. The simulation and real-world experimental results show the effectiveness of our approach. Yuan Zhou 0005, Hesuan Hu, Gelei Deng, Shangwei Lin 0001, Yang Liu 0003, Zuohua Ding |
IEEE Trans. Syst. Man Cybern. Syst. | 6 |
| 2023 | Background-Mixed Augmentation for Weakly Supervised Change DetectionabstractChange detection (CD) is to decouple object changes (i.e., object missing or appearing) from background changes (i.e., environment variations) like light and season variations in two images captured in the same scene over a long time span, presenting critical applications in disaster management, urban development, etc. In particular, the endless patterns of background changes require detectors to have a high generalization against unseen environment variations, making this task significantly challenging. Recent deep learning-based methods develop novel network architectures or optimization strategies with paired-training examples, which do not handle the generalization issue explicitly and require huge manual pixel-level annotation efforts. In this work, for the first attempt in the CD community, we study the generalization issue of CD from the perspective of data augmentation and develop a novel weakly supervised training algorithm that only needs image-level labels. Different from general augmentation techniques for classification, we propose the background-mixed augmentation that is specifically designed for change detection by augmenting examples under the guidance of a set of background changing images and letting deep CD models see diverse environment variations. Moreover, we propose the augmented & real data consistency loss that encourages the generalization increase significantly. Our method as a general framework can enhance a wide range of existing deep learning-based detectors. We conduct extensive experiments in two public datasets and enhance four state-of-the-art methods, demonstrating the advantages of our method. We release the code at https://github.com/tsingqguo/bgmix. Rui Huang 0006, Ruofei Wang, Qing Guo 0005, Jieda Wei, Yuxiang Zhang 0003, Wei Fan 0001, Yang Liu 0003 |
AAAI | 7 |
| 2023 | Multi-target Backdoor Attacks for Code Pre-trained ModelsabstractBackdoor attacks for neural code models have gained considerable attention due to the advancement of code intelligence.However, most existing works insert triggers into task-specific data for code-related downstream tasks, thereby limiting the scope of attacks.Moreover, the majority of attacks for pre-trained models are designed for understanding tasks.In this paper, we propose task-agnostic backdoor attacks for code pre-trained models.Our backdoored model is pre-trained with two learning strategies (i.e., Poisoned Seq2Seq learning and token representation learning) to support the multitarget attack of downstream code understanding and generation tasks.During the deployment phase, the implanted backdoors in the victim models can be activated by the designed triggers to achieve the targeted attack.We evaluate our approach on two code understanding tasks and three code generation tasks over seven datasets.Extensive experiments demonstrate that our approach can effectively and stealthily attack code-related downstream tasks. Yanzhou Li, Shangqing Liu, Kangjie Chen, Xiaofei Xie, Tianwei Zhang 0004, Yang Liu 0003 |
ACL (1) | 6 |
| 2023 | PumpChannel: An Efficient and Secure Communication Channel for Trusted Execution Environment on ARM-FPGA Embedded SoCabstractARM TrustZone separates the system into the rich execution environment (REE) and the trusted execution environment (TEE). Data can be exchanged between REE and TEE through the communication channel, which is based on shared memory and can be accessed by both REE and TEE. Therefore, when the REE OS kernel is untrusted, the security of the communication channel cannot be guaranteed. The proposed schemes to protect the communication channel have high performance overhead and are not secure enough. In this paper, we propose PumpChannel, an efficient and secure communication channel implemented on ARM-FPGA embedded SoC. PumpChannel avoids the use of secret keys, but utilizes a hardware and software collaborative pump to enhance the security and performance of the communication channel. Besides, PumpChannel implements a hardware-based hook integrity monitor to ensure the integrity of all hook code. Security and performance evaluation results show that PumpChannel is more secure than the encrypted channel countermeasures and has better performance than all other evaluated schemes. Jingquan Ge, Yuekang Li, Yang Liu 0003, Yaowen Zheng, Yi Liu 0069, Lida Zhao |
DATE | 3 |
| 2023 | SoK: Rethinking Sensor Spoofing Attacks against Robotic Vehicles from a Systematic ViewabstractRobotic Vehicles (RVs) have gained great popularity over the past few years. Meanwhile, they are also demonstrated to be vulnerable to sensor spoofing attacks. Although a wealth of research works have presented various attacks, some key questions remain unanswered: are these existing works complete enough to cover all the sensor spoofing threats? If not, how many attacks are not explored, and how difficult is it to realize them?This paper answers the above questions by comprehensively systematizing the knowledge of sensor spoofing attacks against RVs. Our contributions are threefold. (1) We identify seven common attack paths in an RV system pipeline. We categorize and assess existing spoofing attacks from the perspectives of spoofer property, operation, victim characteristic and attack goal. Based on this systematization, we identify 4 interesting insights about spoofing attack designs. (2) We propose a novel action flow model to systematically describe robotic function executions and unexplored sensor spoofing threats. With this model, we successfully discover 103 spoofing attack vectors, 26 of which have been verified by prior works, while 77 attacks are never considered. (3) We design two novel attack methodologies to verify the feasibility of newly discovered spoofing attack vectors. Yuan Xu 0033, Xingshuo Han, Gelei Deng, Jiwei Li 0001, Yang Liu 0003, Tianwei Zhang 0004 |
EuroS&P | 5 |
| 2023 | FAIRER: Fairness as Decision Rationale AlignmentabstractDeep neural networks (DNNs) have made significant progress, but often suffer from fairness issues, as deep models typically show distinct accuracy differences among certain subgroups (e.g., males and females). Existing research addresses this critical issue by employing fairness-aware loss functions to constrain the last-layer outputs and directly regularize DNNs. Although the fairness of DNNs is improved, it is unclear how the trained network makes a fair prediction, which limits future fairness improvements. In this paper, we investigate fairness from the perspective of decision rationale and define the parameter parity score to characterize the fair decision process of networks by analyzing neuron influence in various subgroups. Extensive empirical studies show that the unfair issue could arise from the unaligned decision rationales of subgroups. Existing fairness regularization terms fail to achieve decision rationale alignment because they only constrain last-layer outputs while ignoring intermediate neuron alignment. To address the issue, we formulate the fairness as a new task, i.e., decision rationale alignment that requires DNNs’ neurons to have consistent responses on subgroups at both intermediate processes and the final prediction. To make this idea practical during optimization, we relax the naive objective function and propose gradient-guided parity alignment, which encourages gradient-weighted consistency of neurons across subgroups. Extensive experiments on a variety of datasets show that our method can significantly enhance fairness while sustaining a high level of accuracy and outperforming other approaches by a wide margin. Tianlin Li, Qing Guo 0005, Aishan Liu, Mengnan Du, Yang Liu 0003 |
ICML | 6 |
| 2023 | FLYOVER: A Model-Driven Method to Generate Diverse Highway Interchanges for Autonomous Vehicle TestingabstractIt has become a consensus that autonomous vehicles (AVs) will first be widely deployed on highways. However, the complexity of highway interchanges becomes the bottleneck for their deployment. An AV should be sufficiently tested under different highway interchanges, which is still challenging due to the lack of available datasets containing diverse highway interchanges. In this paper, we propose a model-driven method, Flyover, to generate a dataset of diverse interchanges with measurable diversity coverage. First, Flyover uses a labeled digraph to model interchange topology. Second, Flyover takes real-world interchanges as input to guarantee topology practicality and extracts different topology equivalence classes by classifying corresponding topology models. Third, for each topology class, Flyover identifies the corresponding geometrical features for the ramps and generates concrete interchanges using k-way combinatorial coverage and differential evolution. To illustrate the diversity and applicability of the generated interchange dataset, we test the built-in traffic flow control algorithm in SUMO and the fuel-optimization trajectory tracking algorithm deployed to Alibaba's autonomous trucks on the dataset. The results show that except for the geometrical difference, the interchanges are diverse in throughput and fuel consumption under the traffic flow control and trajectory tracking algorithms, respectively. Yuan Zhou 0005, Gengjie Lin, Yun Tang 0003, Kairui Yang, Junbo Chen, Yang Liu 0003 |
ICRA | 9 |
| 2023 | ContraBERT: Enhancing Code Pre-trained Models via Contrastive LearningabstractLarge-scale pre-trained models such as CodeBERT, GraphCodeBERT have earned widespread attention from both academia and industry. Attributed to the superior ability in code representation, they have been further applied in multiple downstream tasks such as clone detection, code search and code translation. However, it is also observed that these state-of-the-art pre-trained models are susceptible to adversarial attacks. The performance of these pre-trained models drops significantly with simple perturbations such as renaming variable names. This weakness may be inherited by their downstream models and thereby amplified at an unprecedented scale. To this end, we propose an approach namely ContraBERT that aims to improve the robustness of pre-trained models via contrastive learning. Specifically, we design nine kinds of simple and complex data augmentation operators on the programming language (PL) and natural language (NL) data to construct different variants. Furthermore, we continue to train the existing pre-trained models by masked language modeling (MLM) and contrastive pre-training task on the original samples with their augmented variants to enhance the robustness of the model. The extensive ex-periments demonstrate that ContraBERT can effectively improve the robustness of the existing pre-trained models. Further study also confirms that these robustness-enhanced models provide improvements as compared to original models over four popular downstream tasks. Shangqing Liu, Bozhi Wu, Xiaofei Xie, Guozhu Meng, Yang Liu 0003 |
ICSE | 5 |
| 2023 | Comparison and Evaluation of Clone Detection Techniques with Different Code RepresentationsabstractAs one of bad smells in code, code clones may increase the cost of software maintenance and the risk of vulnerability propagation. In the past two decades, numerous clone detection technologies have been proposed. They can be divided into text-based, token-based, tree-based, and graph-based approaches according to their code representations. Different code representations abstract the code details from different perspectives. However, it is unclear which code representation is more effective in detecting code clones and how to combine different code representations to achieve ideal performance. In this paper, we present an empirical study to compare the clone detection ability of different code representations. Specifically, we reproduce 12 clone detection algorithms and divide them into different groups according to their code representations. After analyzing the empirical results, we find that token and tree representations can perform better than graph representation when detecting simple code clones. However, when the code complexity of a code pair increases, graph representation becomes more effective. To make our findings more practical, we perform manual analysis on open-source projects to seek a possible distribution of different clone types in the open-source community. Through the results, we observe that most clone pairs belong to simple code clones. Based on this observation, we discard heavyweight graph-based clone detection algorithms and conduct combination experiments to find out a suitable combination of token-based and tree-based approaches for achieving scalable and effective code clone detection. We develop the suitable combination into a tool called TACC and evaluate it with other state-of-the-art code clone detectors. Experimental results indicate that TACC performs better and has the ability to detect large-scale code clones. Yuekun Wang, Yuhang Ye 0004, Yueming Wu 0001, Yinxing Xue, Yang Liu 0003 |
ICSE | 6 |
| 2023 | OSSFP: Precise and Scalable C/C++ Third-Party Library Detection using Fingerprinting FunctionsabstractThird-party libraries (TPLs) are frequently used in software to boost efficiency by avoiding repeated developments. However, the massive using TPLs also brings security threats since TPLs may introduce bugs and vulnerabilities. Therefore, software composition analysis (SCA) tools have been proposed to detect and manage TPL usage. Unfortunately, due to the presence of common and trivial functions in the bloated feature dataset, existing tools fail to precisely and rapidly identify TPLs in C/C++ real-world projects. To this end, we propose OSSFP, a novel SCA framework for effective and efficient TPL detection in large-scale real-world projects via generating unique fingerprints for open source software. By removing common and trivial functions and keeping only the core functions to build the fingerprint index for each TPL project, OSSFP significantly reduces the database size and accelerates the detection process. It also improves TPL detection accuracy since noises are excluded from the fingerprints. We applied OSSFP on a large data set containing 23,427 C/C++ repositories, which included 585,683 versions and 90 billion lines of code. The result showed that it could achieve 90.84% of recall and 90.34% of precision, which outperformed the state-of-the-art tool by 35.31% and 3.71%, respectively. OSSFP took only 0.12 seconds on average to identify all TPLs per project, which was 22 times faster than the other tool. OSSFP has proven to be highly scalable on large-scale datasets. Zhengzi Xu, Lyuye Zhang, Yueming Wu 0001, Chengyue Liu, Kairan Sun, Lida Zhao, Yang Liu 0003 |
ICSE | 9 |
| 2023 | Compatible Remediation on Vulnerabilities from Third-Party Libraries for Java ProjectsabstractWith the increasing disclosure of vulnerabilities in open-source software, software composition analysis (SCA) has been widely applied to reveal third-party libraries and the associated vulnerabilities in software projects. Beyond the revelation, SCA tools adopt various remediation strategies to fix vulnerabilities, the quality of which varies substantially. However, ineffective remediation could induce side effects, such as compi-lation failures, which impede acceptance by users. According to our studies, existing SCA tools could not correctly handle the concerns of users regarding the compatibility of remediated projects. To this end, we propose Compatible Remediation of Third-party libraries (CORAL) for Maven projects to fix vulnerabilities without breaking the projects. The evaluation proved that Coralnot only fixed 87.56% of vulnerabilities which outperformed other tools (best 75.32%) and achieved a 98.67% successful compilation rate and a 92.96% successful unit test rate. Furthermore, we found that 78.45% of vulnerabilities in popular Maven projects could be fixed without breaking the compilation, and the rest of the vulnerabilities (21.55%) could either be fixed by upgrades that break the compilations or even be impossible to fix by upgrading. Lyuye Zhang, Zhengzi Xu, Sen Chen 0001, Lingling Fan 0003, Lida Zhao, Yang Liu 0003 |
ICSE | 8 |
| 2023 | Fairness via Group Contribution MatchingabstractFairness issues in Deep Learning models have recently received increasing attention due to their significant societal impact. Although methods for mitigating unfairness are constantly proposed, little research has been conducted to understand how discrimination and bias develop during the standard training process. In this study, we propose analyzing the contribution of each subgroup (i.e., a group of data with the same sensitive attribute) in the training process to understand the cause of such bias development process. We propose a gradient-based metric to assess training subgroup contribution disparity, showing that unequal contributions from different subgroups are one source of such unfairness. One way to balance the contribution of each subgroup is through oversampling, which ensures that an equal number of samples are drawn from each subgroup during each training iteration. However, we have found that even with a balanced number of samples, the contribution of each group remains unequal, resulting in unfairness under the oversampling strategy. To address the above issues, we propose an easy but effective group contribution matching (GCM) method to match the contribution of each subgroup. Our experiments show that our GCM effectively improves fairness and outperforms other methods significantly. Tianlin Li, Anran Li 0001, Mengnan Du, Aishan Liu, Qing Guo 0005, Guozhu Meng, Yang Liu 0003 |
IJCAI | 8 |
| 2023 | FedSDG-FS: Efficient and Secure Feature Selection for Vertical Federated LearningabstractVertical Federated Learning (VFL) enables multiple data owners, each holding a different subset of features about largely overlapping sets of data sample(s), to jointly train a useful global model. Feature selection (FS) is important to VFL. It is still an open research problem as existing FS works designed for VFL either assumes prior knowledge on the number of noisy features or prior knowledge on the post-training threshold of useful features to be selected, making them unsuitable for practical applications. To bridge this gap, we propose the Federated Stochastic Dual-Gate based Feature Selection (FedSDG-FS) approach. It consists of a Gaussian stochastic dual-gate to efficiently approximate the probability of a feature being selected, with privacy protection through Partially Homomorphic Encryption without a trusted third-party. To reduce overhead, we propose a feature importance initialization method based on Gini impurity, which can accomplish its goals with only two parameter transmissions between the server and the clients. Extensive experiments on both synthetic and real-world datasets show that FedSDG-FS significantly outperforms existing approaches in terms of achieving accurate selection of high-quality features as well as building global models with improved performance. Anran Li 0001, Hongyi Peng, Lan Zhang 0002, Qing Guo 0005, Han Yu 0001, Yang Liu 0003 |
INFOCOM | 7 |
| 2023 | Beyond "Protected" and "Private": An Empirical Security Analysis of Custom Function Modifiers in Smart ContractsabstractA smart contract is a piece of application-layer code running on blockchain ledgers and it provides programmatic logic via transaction-based execution of pre-defined functions. Smart contract functions are by default invokable by any party. To safeguard them, the mainstream smart contract language, i.e., Solidity of the popular Ethereum blockchain, proposed a unique language-level keyword called “modifier,” which allows developers to define custom function access control policies beyond the traditional “protected” and “private” modifiers in classic programming languages. Yuzhou Fang, Daoyuan Wu, Xiao Yi, Shuai Wang 0011, Mengjie Chen, Yang Liu 0003, Lingxiao Jiang |
ISSTA | 7 |
| 2023 | A Comprehensive Study on Quality Assurance Tools for JavaabstractQuality assurance (QA) tools are receiving more and more attention and are widely used by developers. Given the wide range of solutions for QA technology, it is still a question of evaluating QA tools. Most existing research is limited in the following ways: (i) They compare tools without considering scanning rules analysis. (ii) They disagree on the effectiveness of tools due to the study methodology and benchmark dataset. (iii) They do not separately analyze the role of the warnings. (iv) There is no large-scale study on the analysis of time performance. To address these problems, in the paper, we systematically select 6 free or open-source tools for a comprehensive study from a list of 148 existing Java QA tools. To carry out a comprehensive study and evaluate tools in multi-level dimensions, we first mapped the scanning rules to the CWE and analyze the coverage and granularity of the scanning rules. Then we conducted an experiment on 5 benchmarks, including 1,425 bugs, to investigate the effectiveness of these tools. Furthermore, we took substantial effort to investigate the effectiveness of warnings by comparing the real labeled bugs with the warnings and investigating their role in bug detection. Finally, we assessed these tools’ time performance on 1,049 projects. The useful findings based on our comprehensive study can help developers improve their tools and provide users with suggestions for selecting QA tools. Han Liu 0012, Sen Chen 0001, Kaixuan Li 0002, Zhengzi Xu, Liming Nie, Yang Liu 0003, Yixiang Chen 0001 |
ISSTA | 8 |
| 2023 | Detecting Condition-Related Bugs with Control Flow Graph Neural NetworkabstractAutomated bug detection is essential for high-quality software development and has attracted much attention over the years. Among the various bugs, previous studies show that the condition expressions are quite error-prone and the condition-related bugs are commonly found in practice. Traditional approaches to automated bug detection are usually limited to compilable code and require tedious manual effort. Recent deep learning-based work tends to learn general syntactic features based on Abstract Syntax Tree (AST) or apply the existing Graph Neural Networks over program graphs. However, AST-based neural models may miss important control flow information of source code, and existing Graph Neural Networks for bug detection tend to learn local neighbourhood structure information. Generally, the condition-related bugs are highly influenced by control flow knowledge, therefore we propose a novel CFG-based Graph Neural Network (CFGNN) to automatically detect condition-related bugs, which includes a graph-structured LSTM unit to efficiently learn the control flow knowledge and long-distance context information. We also adopt the API-usage attention mechanism to leverage the API knowledge. To evaluate the proposed approach, we collect real-world bugs in popular GitHub repositories and build a large-scale condition-related bug dataset. The experimental results show that our proposed approach significantly outperforms the state-of-the-art methods for detecting condition-related bugs. Jian Zhang 0087, Xu Wang 0007, Hongyu Zhang 0002, Hailong Sun 0001, Xudong Liu 0001, Chunming Hu, Yang Liu 0003 |
ISSTA | 7 |
| 2023 | An Empirical Study of Malicious Code In PyPI EcosystemabstractPyPI provides a convenient and accessible package management platform to developers, enabling them to quickly implement specific functions and improve work efficiency. However, the rapid development of the PyPI ecosystem has led to a severe problem of malicious package propagation. Malicious developers disguise malicious packages as normal, posing a significant security risk to end-users. To this end, we conducted an empirical study to understand the characteristics and current state of the malicious code lifecycle in the PyPI ecosystem. We first built an automated data collection framework and collated a multi-source malicious code dataset containing 4,669 malicious package files. We preliminarily classified these malicious code into five categories based on malicious behaviour characteristics. Our research found that over 50 % of malicious code exhibits multiple malicious behaviours, with information stealing and command execution being particularly prevalent. In addition, we observed several novel attack vectors and anti-detection techniques. Our analysis revealed that 74.81 % of all malicious packages successfully entered end-user projects through source code installation, thereby increasing security risks. A real-world investigation showed that many reported malicious packages persist in PyPI mirror servers globally, with over 72 % remaining for an extended period after being discovered. Finally, we sketched a portrait of the malicious code lifecycle in the PyPI ecosystem, effectively reflecting the characteristics of malicious code at different stages. We also present some suggested mitigations to improve the security of the Python open-source ecosystem. Wenbo Guo 0011, Zhengzi Xu, Cheng Huang 0003, Yong Fang 0002, Yang Liu 0003 |
ASE | 6 |
| 2023 | An Empirical Study on Fine-Tuning Large Language Models of Code for Automated Program RepairabstractThe advent of large language models (LLMs) has opened up new opportunities for automated program repair (APR). In particular, some recent studies have explored how to leverage large language models of code (LLMCs) for program repair tasks and show promising results. However, most of them adopt the zero/few-shot learning paradigm for APR, which directly use LLMCs to generate the possibly correct code given its surrounding context. Though effective, the repair capabilities of LLMCs based on the fine-tuning paradigm have yet to be extensively explored. Also, it remains unknown whether LLMCs have the potential to repair more complicated bugs (e.g., multi-hunk bugs). To fill the gap, in this work, we conduct a comprehensive study on the program repair capability of LLMCs in the fine-tuning paradigm. We select 5 popular LLMCs with representative pre-training architectures, including CodeBERT, GraphCode-BERT, PLBART, CodeT5, and UniX coder. We consider 3 typical program repair scenarios (i.e., bugs, vulnerabilities, and errors) involving 3 programming languages (i.e., Java,$\mathrm{C}/\mathrm{C}++$, and JavaScript). Notably, we take both single-hunk and multi-hunk bugs/vulnerabilities into account. We then fine-tune them on widely-used datasets and compare them with existing state-of-the-art APR tools. We also investigate the impact of different design choices, which include code abstractions, code representations, and model evaluation metrics. Our experimental results show that LLMCs in the fine-tuning paradigm can significantly outperform previous state-of-the-art APR tools. Through in-depth analysis, we provide insights into choosing appropriate strategies to guide LLMCs for better performance. Lastly, we reveal several limitations of LLMCs for APR and make suggestions for future research on LLMC-based APR. Xiangxin Meng, Jian Zhang 0087, Yang Liu 0003, Yuqing Zhang 0001 |
ASE | 4 |
| 2023 | LiSum: Open Source Software License Summarization with Multi-Task LearningabstractOpen source software (OSS) licenses regulate the conditions under which users can reuse, modify, and distribute the software legally. However, there exist various OSS licenses in the community, written in a formal language, which are typically long and complicated to understand. In this paper, we conducted a 661-participants online survey to investigate the perspectives and practices of developers towards OSS licenses. The user study revealed an indeed need for an automated tool to facilitate license understanding. Motivated by the user study and the fast growth of licenses in the community, we propose the first study towards automated license summarization. Specifically, we released the first high quality text summarization dataset and designed two tasks, i.e., license text summarization (LTS), aiming at generating a relatively short summary for an arbitrary license, and license term classification (LTC), focusing on the attitude inference towards a predefined set of key license terms (e.g., Distribute). Aiming at the two tasks, we present LiSum, a multi-task learning method to help developers overcome the obstacles of understanding OSS licenses. Comprehensive experiments demonstrated that the proposed jointly training objective boosted the performance on both tasks, surpassing state-of-the-art baselines with gains of at least 5 points w.r.t. F1 scores of four summarization metrics and achieving 95.13% micro average F1 score for classification simultaneously. We released all the datasets, the replication package, and the questionnaires for the community. Linyu Li 0002, Sihan Xu, Yang Liu 0003, Xiangrui Cai, Jiarun Wu, Wenli Song, Zheli Liu |
ASE | 3 |
| 2023 | ASTER: Automatic Speech Recognition System Accessibility Testing for StutterersabstractThe popularity of automatic speech recognition (ASR) systems nowadays leads to an increasing need for improving their accessibility. Handling stuttering speech is an important feature for accessible ASR systems. To improve the accessibility of ASR systems for stutterers, we need to expose and analyze the failures of ASR systems on stuttering speech. The speech datasets recorded from stutterers are not diverse enough to expose most of the failures. Furthermore, these datasets lack ground truth information about the non-stuttered text, rendering them unsuitable as comprehensive test suites. Therefore, a methodology for generating stuttering speech as test inputs to test and analyze the performance of ASR systems is needed. However, generating valid test inputs in this scenario is challenging. The reason is that although the generated test inputs should mimic how stutterers speak, they should also be diverse enough to trigger more failures. To address the challenge, we propose Aster, a technique for automatically testing the accessibility of ASR systems. Aster can generate valid test cases by injecting five different types of stuttering. The generated test cases can both simulate realistic stuttering speech and expose failures in ASR systems. Moreover, Aster can further enhance the quality of the test cases with a multi-objective optimization-based seed updating algorithm. We implemented Aster as a framework and evaluated it on four open-source ASR models and three commercial ASR systems. We conduct a comprehensive evaluation of Aster and find that it significantly increases the word error rate, match error rate, and word information loss in the evaluated ASR systems. Additionally, our user study demonstrates that the generated stuttering audio is indistinguishable from real-world stuttering audio clips. Yi Liu 0069, Yuekang Li, Gelei Deng, Felix Juefei-Xu, Yao Du 0002, Cen Zhang, Yeting Li, Lei Ma 0003, Yang Liu 0003 |
ASE | 10 |
| 2023 | Who is the Real Hero? Measuring Developer Contribution via Multi-Dimensional Data IntegrationabstractProper incentives are important for motivating developers in open-source communities, which is crucial for maintaining the development of open-source software healthy. To provide such incentives, an accurate and objective developer contribution measurement method is needed. However, existing methods rely heavily on manual peer review, lacking objectivity and transparency. The metrics of some automated works about effort estimation use only syntax-level or even text-level information, such as changed lines of code, which lack robustness. Furthermore, some works about identifying core developers provide only a qualitative understanding without a quantitative score or have some project-specific parameters, which makes them not practical in real-world projects. To this end, we propose CVALUE, a multidimensional information fusion-based approach to measure developer contributions. CVALUE extracts both syntax and semantic information from the source code changes in four dimensions: modification amount, understandability, inter-function and intra-function impact of modification. It fuses the information to produce the contribution score for each of the commits in the projects. Experimental results show that CVALUE outperforms other approaches by 19.59% on 10 real-world projects with manually labeled ground truth. We validated and proved that the performance of CVALUE, which takes 83.39 seconds per commit, is acceptable to be applied in real-world projects. Furthermore, we performed a large-scale experiment on 174 projects and detected 2,282 developers having inflated commits. Of these, 2,050 developers did not make any syntax contribution; and 103 were identified as bots. Yuqiang Sun 0001, Zhengzi Xu, Yang Liu 0003 |
ASE | 5 |
| 2023 | When Less is Enough: Positive and Unlabeled Learning Model for Vulnerability DetectionabstractAutomated code vulnerability detection has gained increasing attention in recent years. The deep learning (DL)-based methods, which implicitly learn vulnerable code patterns, have proven effective in vulnerability detection. The performance of DL-based methods usually relies on the quantity and quality of labeled data. However, the current labeled data are generally automatically collected, such as crawled from human-generated commits, making it hard to ensure the quality of the labels. Prior studies have demonstrated that the non-vulnerable code (i.e., negative labels) tends to be unreliable in commonly-used datasets, while vulnerable code (i.e., positive labels) is more determined. Considering the large numbers of unlabeled data in practice, it is necessary and worth exploring to leverage the positive data and large numbers of unlabeled data for more accurate vulnerability detection. In this paper, we focus on the Positive and Unlabeled (PU) learning problem for vulnerability detection and propose a novel model named PILOT, i.e., Positive and unlabeled Learning mOdel for vulnerability deTection. PILOT only learns from positive and unlabeled data for vulnerability detection. It mainly contains two modules: (1) A distance-aware label selection module, aiming at generating pseudo-labels for selected unlabeled data, which involves the inter-class distance prototype and progressive fine-tuning; (2) A mixed-supervision representation learning module to further alleviate the influence of noise and enhance the discrimination of representations. Extensive experiments in vulnerability detection are conducted to evaluate the effectiveness of PILOT based on real-world vulnerability datasets. The experimental results show that PILOT outperforms the popular weakly supervised methods by 2.78%-18.93% in the PU learning setting. Compared with the state-of-the-art methods, PILOT also improves the performance of 1.34%-12.46 % in F1 score metrics in the supervised setting. In addition, PILOT can identify 23 mislabeled from the FFMPeg+Qemu dataset in the PU learning setting based on manual checking. Xin-Cheng Wen, Xinchen Wang 0001, Cuiyun Gao 0001, Shaohua Wang 0002, Yang Liu 0003, Zhaoquan Gu |
ASE | 5 |
| 2023 | Mitigating Persistence of Open-Source Vulnerabilities in Maven EcosystemabstractVulnerabilities from third-party libraries (TPLs) have been unveiled to threaten the Maven ecosystem in the long term. Despite patches being released promptly after vulnerabilities are disclosed, the libraries and applications in the community still use the vulnerable versions, which makes the vulnerabilities persistent in the Maven ecosystem (e.g., the notorious Log4Shell still greatly influences the Maven ecosystem nowadays from 2021). Both academic and industrial researchers have proposed user-oriented standards and solutions to address vulnerabilities, while such solutions fail to tackle the ecosystem-wide persistent vulnerabilities because it requires a collective effort from the community to timely adopt patches without introducing breaking issues. To seek an ecosystem-wide solution, we first carried out an empirical study to examine the prevalence of persistent vulnerabilities in the Maven ecosystem. Then, we identified affected libraries for alerts by implementing an algorithm monitoring downstream dependents of vulnerabilities based on an up-to-date dependency graph. Based on them, we further quantitatively revealed that patches blocked by upstream libraries caused the persistence of vulnerabilities. After reviewing the drawbacks of existing countermeasures, to address them, we proposed a solution for range restoration (Ranger) to automatically restore the compatible and secure version ranges of dependencies for downstream dependents. The automatic restoration requires no manual effort from the community, and the code-centric compatibility assurance ensures smooth upgrades to patched versions. Moreover, Ranger along with the ecosystem monitoring can timely alert developers of blocking libraries and suggest flexible version ranges to rapidly unblock patch versions. By evaluation, Ranger could restore 75.64% of ranges which automatically remediated 90.32% of vulnerable downstream projects. Lyuye Zhang, Sen Chen 0001, Zhengzi Xu, Lingling Fan 0003, Lida Zhao, Yang Liu 0003 |
ASE | 8 |
| 2023 | Learning to Locate and Describe VulnerabilitiesabstractAutomatically discovering software vulnerabilities is a long-standing pursuit for software developers and security analysts. Since detection tools usually provide limited information for vulnerability inspection, recent work turns the attention to identify fine-grained vulnerabilities, i.e., vulnerable statements. However, existing work for vulnerability localization struggles to capture long-range and integral dependency information due to the bottleneck of Graph Neural Networks (GNNs). Moreover, little research has been done to help developers understand detected vulnerabilities, leaving vulnerability diagnosis a challenging task. In this paper, we propose VulTeller, a deep learning-based approach that can automatically locate vulnerable statements in a function and more importantly, can describe the vulnerability. Our approach focuses on extracting precise control and data dependencies in the code, achieved through modeling control flow paths and employing taint analysis. We design a novel neural model that encodes the control flows and taint flows which reside in the control flow paths, and decodes them via node classification and an attentional decoder for the two tasks respectively. We conduct extensive experiments with real-world vulnerabilities to evaluate the proposed approach. The evaluation results, including quantitative measurement and human evaluation, demonstrate that our approach is highly effective and outperforms state-of-the-art approaches. Our work for the first time formulates the problem of vulnerability description generation, and makes one step further towards automated vulnerability diagnosis. Jian Zhang 0087, Shangqing Liu, Xu Wang 0007, Tianlin Li, Yang Liu 0003 |
ASE | 5 |
| 2023 | ALA: Naturalness-aware Adversarial Lightness AttackabstractMost researchers have tried to enhance the robustness of deep neural networks (DNNs) by revealing and repairing the vulnerability of DNNs with specialized adversarial examples. Parts of the attack examples have imperceptible perturbations restricted by Lp norm. However, due to their high-frequency property, the adversarial examples can be defended by denoising methods and are hard to realize in the physical world. To avoid the defects, some works have proposed unrestricted attacks to gain better robustness and practicality. It is disappointing that these examples usually look unnatural and can alert the guards. In this paper, we propose Adversarial Lightness Attack (ALA), a white-box unrestricted adversarial attack that focuses on modifying the lightness of the images. The shape and color of the samples, which are crucial to human perception, are barely influenced. To obtain adversarial examples with a high attack success rate, we propose unconstrained enhancement in terms of the light and shade relationship in images. To enhance the naturalness of images, we craft the naturalness-aware regularization according to the range and distribution of light. The effectiveness of ALA is verified on two popular datasets for different tasks (i.e., ImageNet for image classification and Places-365 for scene recognition). Yihao Huang 0001, Liangru Sun, Qing Guo 0005, Felix Juefei-Xu, Jiayi Zhu 0002, Jincao Feng, Yang Liu 0003, Geguang Pu |
ACM Multimedia | 7 |
| 2023 | ThreatLand: Extracting Intelligence from Audit Logs via NLP methodsabstractThreat intelligence and hunting using various logs has evolved into a crucial component of remaining aware of the ever-changing threat landscape. Given the critical need to extract useful intelligence from logs, existing techniques either focus exclusively on isolated records, ignoring correlation and the overall threat scenario, or require significant effort to filter and correlate threat records. Additionally, searching for and matching threat behaviors in logs often involves non-trivial human query construction, impeding fast threat hunting. To address this gap, we present ThreatLand, a system that extracts highlevel intelligence and structured threat patterns from audit logs automatically. ThreatLand is composed of three components (1) A lightweight and accurate NLP pipeline that extracts structured meta-data from alert descriptions and generates a heterogeneous graph that depicts the entire threat scenario. (2) A query execution engine that is both fast and efficient, based on a graphical database. (3) A graphical user interface (GUI) that offers various sorts of interactivity to aid intelligence exploration.We have evaluated the ThreatLand over the dataset containing 9240 real-time EDR alerts collected for the threat events over an enterprise setup in the lab. As a result, ThreatLand presents high-level insights from the alert logs and extracts the valuable threat patterns. Vinay Sachidananda, Rajendra Patil 0001, Hongyi Peng, Yang Liu 0003, Kwok-Yan Lam |
PST | 4 |
| 2023 | GitFL: Uncertainty-Aware Real-Time Asynchronous Federated Learning Using Version ControlabstractAs a promising distributed machine learning paradigm that enables collaborative training without compromising data privacy, Federated Learning (FL) has been increasingly used in large-scale A IoT (Artificial Intelligence of Things) system design. However, due to the lack of efficient management of straggling devices, existing FL methods greatly suffer from the problems of long response time (e.g., training and communication latency) and low inference accuracy. Things become even worse when taking various uncertain factors (e.g., network delays, performance variances caused by process variation) existing in AIoT scenarios into account. To address this issue, this paper proposes a novel asynchronous FL framework named GitFL, whose implementation is inspired by the famous version control system Git. Unlike traditional FL, the cloud server of GitFL maintains a master model (i.e., the global model) together with a set of branch models indicating the trained local models committed by selected devices, where the master model is updated based on both all the pushed branch models and their version information, and only the branch models after the pull operation are dispatched to devices. By using our proposed Reinforcement Learning (RL)-based device selection mechanism, a pulled branch model with an older version will be more likely to be dispatched to a faster and less frequently selected device for the next round of local training. In this way, GitFL enables both effective controls of model staleness and adaptive load balance of versioned models among straggling devices, thus avoiding performance deterioration while ensuring real-time performance. Comprehensive experimental results on well-known models and datasets show that, compared with state-of-the-art asynchronous and synchronous FL methods, GitFL can achieve up to 2.64X training acceleration and 7.88 % inference accuracy improvements in various uncertain scenarios. Ming Hu 0003, Zeke Xia, Dengke Yan, Zhihao Yue, Jun Xia 0003, Yihao Huang 0001, Yang Liu 0003, Mingsong Chen 0001 |
RTSS | 7 |
| 2023 | Comparison and Evaluation on Static Application Security Testing (SAST) Tools for JavaabstractStatic application security testing (SAST) takes a significant role in the software development life cycle (SDLC). However, it is challenging to comprehensively evaluate the effectiveness of SAST tools to determine which is the better one for detecting vulnerabilities. In this paper, based on well-defined criteria, we first selected seven free or open-source SAST tools from 161 existing tools for further evaluation. Owing to the synthetic and newly-constructed real-world benchmarks, we evaluated and compared these SAST tools from different and comprehensive perspectives such as effectiveness, consistency, and performance. While SAST tools perform well on synthetic benchmarks, our results indicate that only 12.7% of real-world vulnerabilities can be detected by the selected tools. Even combining the detection capability of all tools, most vulnerabilities (70.9%) remain undetected, especially those beyond resource control and insufficiently neutralized input/output vulnerabilities. The fact is that although they have already built the corresponding detecting rules and integrated them into their capabilities, the detection result still did not meet the expectations. All useful findings unveiled in our comprehensive study indeed help to provide guidance on tool development, improvement, evaluation, and selection for developers, researchers, and potential users. Kaixuan Li 0002, Sen Chen 0001, Lingling Fan 0003, Han Liu 0012, Yang Liu 0003, Yixiang Chen 0001 |
ESEC/SIGSOFT FSE | 7 |
| 2023 | Gitor: Scalable Code Clone Detection by Building Global Sample GraphabstractCode clone detection is about finding out similar code fragments, which has drawn much attention in software engineering since it is important for software maintenance and evolution. Researchers have proposed many techniques and tools for source code clone detection, but current detection methods concentrate on analyzing or processing code samples individually without exploring the underlying connections among code samples. Junjie Shan, Shihan Dou, Yueming Wu 0001, Hairu Wu, Yang Liu 0003 |
ESEC/SIGSOFT FSE | 5 |
| 2023 | Demystifying the Composition and Code Reuse in Solidity Smart ContractsabstractAs the development of Solidity smart contracts has increased in popularity, the reliance on external sources such as third-party packages increases to reduce development costs. However, despite the use of external sources bringing flexibility and efficiency to the development, they could also complicate the process of assuring the security of downstream applications due to the lack of package managers for standardized ways and sources. While previous studies have only focused on code clones without considering how the external components are introduced, the compositions of a smart contract and their characteristics still remain puzzling. Kairan Sun, Zhengzi Xu, Kaixuan Li 0002, Yang Liu 0003 |
ESEC/SIGSOFT FSE | 5 |
| 2023 | Software Architecture Recovery with Information FusionabstractUnderstanding the architecture is vital for effectively maintaining and managing large software systems. However, as software systems evolve over time, their architectures inevitably change. To keep up with the change, architects need to track the implementation-level changes and update the architectural documentation accordingly, which is time-consuming and error-prone. Therefore, many automatic architecture recovery techniques have been proposed to ease this process. Despite efforts have been made to improve the accuracy of architecture recovery, existing solutions still suffer from two limitations. First, most of them only use one or two type of information for the recovery, ignoring the potential usefulness of other sources. Second, they tend to use the information in a coarse-grained manner, overlooking important details within it. Zhengzi Xu, Hongxu Chen 0001, Dong Qiu, Yang Liu 0003 |
ESEC/SIGSOFT FSE | 7 |
| 2023 | Software Composition Analysis for Vulnerability Detection: An Empirical Study on Java ProjectsabstractSoftware composition analysis (SCA) tools are proposed to detect potential vulnerabilities introduced by open-source software (OSS) imported as third-party libraries (TPL). With the increasing complexity of software functionality, SCA tools may encounter various scenarios during the dependency resolution process, such as diverse formats of artifacts, diverse dependency imports, and diverse dependency specifications. However, there still lacks a comprehensive evaluation of SCA tools for Java that takes into account the above scenarios. This could lead to a confined interpretation of comparisons, improper use of tools, and hinder further improvements of the tools. To fill this gap, we proposed an Evaluation Model which consists of Scan Modes, Scan Methods, and SCA Scope for Maven (SSM), for comprehensive assessments of the dependency resolving capabilities and effectiveness of SCA tools. Based on the Evaluation Model, we first qualitatively examined 6 SCA tools’ capabilities. Next, the accuracy of dependency and vulnerability is quantitatively evaluated with a large-scale dataset (21,130 Maven modules with 73,499 unique dependencies) under two Scan Modes (i.e., build scan and pre-build scan). The results show that most tools do not fully support SSM, which leads to compromised accuracy. For dependency detection, the average F1-score is 0.890 and 0.692 for build and pre-build respectively, and for vulnerability accuracy, the average F1-score is 0.475. However, proper support for SSM reduces dependency detection false positives by 34.24% and false negatives by 6.91%. This further leads to a reduction of 18.28% in false positives and 8.72% in false negatives in vulnerability reports. Lida Zhao, Sen Chen 0001, Zhengzi Xu, Lyuye Zhang, Jun Sun 0001, Yang Liu 0003 |
ESEC/SIGSOFT FSE | 8 |
| 2023 | Effective ReDoS Detection by Principled Vulnerability Modeling and Exploit GenerationabstractRegular expression Denial-of-Service (ReDoS) is one kind of algorithmic complexity attack. For a vulnerable regex, attackers can craft certain strings to trigger the super-linear worst-case matching time, which causes denial-of-service to regex engines. Various ReDoS detection approaches have been proposed recently. Among them, hybrid approaches which absorb the advantages of both static and dynamic approaches have shown their performance superiority. However, two key challenges still hinder the effectiveness of the detection: 1) Existing modelings summarize localized vulnerability patterns based on partial features of the vulnerable regex; 2) Existing attack string generation strategies are ineffective since they neglected the fact that non-vulnerable parts of the regex may unexpectedly invalidate the attack string (we name this kind of invalidation as disturbance.)Rengar is our hybrid ReDoS detector with new vulnerability modeling and disturbance free attack string generator. It has the following key features: 1) Benefited by summarizing patterns from full features of the vulnerable regex, its modeling is a more precise interpretation of the root cause of ReDoS vulnerability. The modeling is more descriptive and precise than the union of existing modelings while keeping conciseness; 2) For each vulnerable regex, its generator automatically checks all potential disturbances and composes generation constraints to avoid possible disturbances.Compared with nine state-of-the-art tools, Rengar detects not only all vulnerable regexes they found but also 3 – 197 times more vulnerable regexes. Besides, it saves 57.41% – 99.83% average detection time compared with tools containing a dynamic validation process. Using Rengar, we have identified 69 zero-day vulnerabilities (21 CVEs) affecting popular projects which have more than dozens of millions weekly download count. Xinyi Wang 0013, Cen Zhang, Yeting Li, Zhiwu Xu 0001, Shuailin Huang, Yi Liu 0069, Yican Yao, Yang Xiao 0011, Yanyan Zou 0002, Yang Liu 0003, Wei Huo 0005 |
SP | 10 |
| 2023 | RSFuzzer: Discovering Deep SMI Handler Vulnerabilities in UEFI Firmware with Hybrid FuzzingabstractSystem Management Mode (SMM) is a secure operation mode for x86 processors supported by Unified Extensible Firmware Interface (UEFI) firmware. SMM is designed to provide a secure execution environment to access highly privileged data or control low-level hardware (such as power management). The programs running in SMM are called SMM drivers and System Management Interrupt (SMI) handlers are the most important components of SMM drivers since they are the only components to receive and handle data from outside the SMM execution environment. Although SMM can serve as an extra layer of protection when the operating system is compromised, vulnerabilities in SMM drivers, especially SMI handlers, can invalidate this protection and cause severe damages to the device. Thus, early detection of SMI handler vulnerabilities is important for UEFI firmware security.To this end, researchers have proposed to use hybrid fuzzing techniques for detecting SMI handler vulnerabilities. Particularly, Intel has developed a hybrid fuzzer called Excite and uses it to secure Intel products. Although existing hybrid fuzzing techniques can detect vulnerabilities in SMI handlers, their effectiveness is limited due to two major pitfalls: 1) They can only feed input through the most common input interface to SMI handlers, lacking the ability to utilize other input interfaces. 2) They have no awareness of variables shared by multiple SMI handlers, lacking the ability to explore code segments related to such variables. By addressing the challenges faced by existing works, we propose RSFuzzer, a hybrid greybox fuzzing technique which can learn input interface and format information and detect deeply hidden vulnerabilities which are triggered by invoking multiple SMI handlers. We implemented RSFuzzer and evaluated it on 16 UEFI firmware images provided by six vendors. The experiment results show that RSFuzzer can cover 617% more basic blocks and detect 828% more vulnerabilities on average than the state-of-the-art hybrid fuzzing technique. Moreover, we found and reported 65 0-day vulnerabilities in the evaluated UEFI firmware images and 14 CVE IDs were assigned. Noticeably, 6 of the 0-day vulnerabilities were found in commercial-off-the-shelf (COTS) products from Intel, which might have been tested by Excite before releasing. Jiawei Yin, Yuekang Li, Boru Lin, Yanyan Zou 0002, Yang Liu 0003, Wei Huo 0005, Jingling Xue |
SP | 7 |
| 2023 | Do NoT Open (DOT): A Unified Generic and Specialized Models for Detecting Malicious Email AttachmentsabstractIn this paper, we propose – DOT – a hybrid analysis approach designed for the detection and classification of malicious files. We have developed both a unified single model and specialized models tailored to various file extensions. Our solutions leverage byte-level content analysis to identify malicious elements within documents, along with n-gram analysis. The uniqueness of DOT lies in its ability to significantly reduce computational overhead. We achieve this by employing Rolling Encoder Hashing, which shortens bytecode sequences, making them compatible with state-of-the-art sequence models like Recurrent Neural Networks (RNNs). Additionally, we have created a static analysis-based generic model capable of working with a variety of file types, including.doc,.docx,.xls,.xlsx,.pdf, and more. This model can be efficiently deployed in real-world scenarios. Furthermore, we have developed specialized models for different file types, which are enhanced versions of the generic architecture, streamlining complex maintenance procedures. Another key innovation and novelty of DOT lies in exactly locating the portion of content in the byte code that could contain malicious code, to help security analysts make the binary code analysis more efficient.We conducted extensive experiments using a dataset recently made available by sources like VirusShare, Contagio, and others, specifically intended for academic research. Our dataset comprises a substantial collection of over 156,000 documents, encompassing both malicious and benign files of the most hazardous types observed in recent years. Our findings reveal impressive results, with a unified single model achieving a 91.43% accuracy in distinguishing between benign and malicious documents. Furthermore, specialized models tailored to specific file types exhibit even higher accuracy rates: 96.13% for.doc files, 97.85% for.docx files, 92.62% for.xls files, 97.02% for.xlsx files, and 94.11% for.pdf files, respectively and with a very low false positive rate. Vinay Sachidananda, Sivaanandh Muneeswaran, Yang Liu 0003, Kwok-Yan Lam |
TrustCom | 3 |
| 2023 | NAUTILUS: Automated RESTful API Vulnerability Detection
Gelei Deng, Zhiyi Zhang 0005, Yuekang Li, Yi Liu 0069, Tianwei Zhang 0004, Yang Liu 0003, Dongjin Wang |
USENIX Security Symposium | 6 |
| 2023 | Automata-Guided Control-Flow-Sensitive Fuzz Driver Generation
Cen Zhang, Yuekang Li, Hao Zhou 0043, Yaowen Zheng, Xian Zhan, Xiaofei Xie, Xiapu Luo, Xinghua Li 0001, Yang Liu 0003, Sheikh Mahbub Habib |
USENIX Security Symposium | 10 |
| 2023 | Learning Program Representations with a Tree-Structured TransformerabstractLearning vector representations for programs is a critical step in applying deep learning techniques for program understanding tasks. Various neural network models are proposed to learn from tree-structured program representations, e.g., abstract syntax tree (AST) and concrete syntax tree (CST). However, most neural architectures either fail to capture long-range dependencies which are ubiquitous in programs, or cannot learn effective representations for syntax tree nodes, making them incapable of performing the node-level prediction tasks, e.g., bug localization. In this paper, we propose Tree-Transformer, a novel recursive tree-structured neural network to learn the vector representations for source codes. We propose a multi-head attention mechanism to model the dependency between siblings and parent-children node pairs. Moreover, we propose a bi-directional propagation strategy to allow node information passing in two directions, bottom-up and top-down along trees. In this way, Tree-Transformer can learn the information of the node features as well as the global contextual information. The extensive experimental results show that our Tree-Transformer significantly outperforms the existing tree-based and graph-based program representation learning approaches in both the tree-level and node-level prediction tasks. Wenhan Wang, Kechi Zhang, Ge Li 0001, Shangqing Liu, Anran Li 0001, Zhi Jin 0001, Yang Liu 0003 |
SANER | 7 |
| 2023 | Baton: symphony of random testing and concolic testing through machine learning and taint analysis
Bihuan Chen 0001, Yang Liu 0003, Xin Peng 0001, Yijian Wu, Shengchao Qin |
Sci. China Inf. Sci. | 2 |
| 2023 | Refinement-based Specification and Analysis of Multi-core ARINC 653 Using Event-BabstractARINC 653 as the de facto standard of partitioning operating systems has been applied in many safety-critical domains. The multi-core version of ARINC 653, ARINC 653 Part 1-4 (Version 4), provides support for services to be utilized with a module that contains multiple processor cores. Formal specification and analysis of this standard document could provide a rigorous specification and uncover concealed errors in the textual description of service requirements. This article proposes a specification method for concurrency on a multi-core platform using Event-B, and a refinement structure for the complicated ARINC 653 Part 1-4 provides a comprehensive, stepwise refinement-based Event-B specification with seven refinement layers and then performs formal proof and analysis in RODIN. We verify that the errors discovered in the single-core version standard (ARINC 653 Part 1-3) also exist in the ARINC 653 Part 1-4 during the formal specification and analysis. Yongwang Zhao, Yang Liu 0003, Jun Sun 0001 |
Formal Aspects Comput. | 4 |
| 2023 | A Simplified Dual-Weighted Three-Layer Window Local Contrast Method for Infrared Small-Target DetectionabstractIn the realm of infrared small-target detection, the weighted local contrast approaches, which seek to improve targets by the defined weighted factors, have garnered a lot of interest. However, there are several problems with these methods as follows. 1) The vast number of local contrast sliding sub-windows restricts the time efficiency. 2) The dim targets in the complicated backgrounds are incorrectly eliminated by the background suppression procedure. 3) The background noise in the complicated environment cannot be effectively muted. A simplified dual-weighted three-layer window local contrast method (SDWTLLCM) is suggested in this work as a solution to these issues. In order to extract tiny targets and suppress complicated backgrounds, a hierarchical convolution filtering window is first created. Then, even without sub-window division, a simple three-layer sliding window is created for time efficiency enhancement. The dual-weighted local contrast approach is also intended to minimize the background and further highlight tiny objects. Eventually, the tiny targets may be extracted more effectively using the adaptive threshold segmentation procedure. The vast experimental findings show that our suggested strategy is effective and efficient. Cancan Chen, Runqiu Xia, Yang Liu 0003, Yue Liu 0008 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2023 | Automatic Transformation Search Against Deep Leakage From GradientsabstractCollaborative learning has gained great popularity due to its benefit of data privacy protection: participants can jointly train a Deep Learning model without sharing their training sets. However, recent works discovered that an adversary can fully recover the sensitive training samples from the shared gradients. Such reconstruction attacks pose severe threats to collaborative learning. Hence, effective mitigation solutions are urgently desired. In this paper, we systematically analyze existing reconstruction attacks and propose to leverage data augmentation to defeat these attacks: by preprocessing sensitive images with carefully-selected transformation policies, it becomes infeasible for the adversary to extract training samples from the corresponding gradients. We first design two new metrics to quantify the impacts of transformations on data privacy and model usability. With the two metrics, we design a novel search method to automatically discover qualified policies from a given data augmentation library. Our defense method can be further combined with existing collaborative training systems without modifying the training protocols. We conduct comprehensive experiments on various system settings. Evaluation results demonstrate that the policies discovered by our method can defeat state-of-the-art reconstruction attacks in collaborative learning, with high efficiency and negligible impact on the model performance. Wei Gao 0064, Shangwei Guo, Tianwei Zhang 0004, Tao Xiang 0001, Han Qiu 0001, Yonggang Wen 0001, Yang Liu 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2023 | Neuron Coverage-Guided Domain GeneralizationabstractThis paper focuses on the domain generalization task where domain knowledge is unavailable, and even worse, only samples from a single domain can be utilized during training. Our motivation originates from the recent progresses in deep neural network (DNN) testing, which has shown that maximizing neuron coverage of DNN can help to explore possible defects of DNN (i.e., misclassification). More specifically, by treating the DNN as a program and each neuron as a functional point of the code, during the network training we aim to improve the generalization capability by maximizing the neuron coverage of DNN with the gradient similarity regularization between the original and augmented samples. As such, the decision behavior of the DNN is optimized, avoiding the arbitrary neurons that are deleterious for the unseen samples, and leading to the trained DNN that can be better generalized to out-of-distribution samples. Extensive studies on various domain generalization tasks based on both single and multiple domain(s) setting demonstrate the effectiveness of our proposed approach compared with state-of-the-art baseline methods. We also analyze our method by conducting visualization based on network dissection. The results further provide useful evidence on the rationality and effectiveness of our approach. Chris Xing Tian, Haoliang Li, Xiaofei Xie, Yang Liu 0003, Shiqi Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | ADS-Lead: Lifelong Anomaly Detection in Autonomous Driving SystemsabstractAutonomous Vehicles (AVs) are closely connected in the Cooperative Intelligent Transportation System (C-ITS). They are equipped with various sensors and controlled by Autonomous Driving Systems (ADSs) to provide high-level autonomy. The vehicles exchange different types of real-time data with each other, which can help reduce traffic accidents and congestion, and improve the efficiency of transportation systems. However, when interacting with the environment, AVs suffer from a broad attack surface, and the sensory data are susceptible to anomalies caused by faults, sensor malfunctions, or attacks, which may jeopardize traffic safety and result in serious accidents. In this paper, we proposeADS-Lead, an efficient collaborative anomaly detection methodology to protect the lane-following mechanism of ADSs.ADS-Leadis equipped with a novel transformer-based one-class classification model to identify time series anomalies (GPS spoofing threat) and adversarial image examples (traffic sign and lane recognition attacks). Besides, AVs inside the C-ITS form a cognitive network, enabling us to apply the federated learning technology to our anomaly detection method, where the vehicles in the C-ITS jointly update the detection model with higher model generalization and data privacy. Experiments on Baidu Apollo and two public data sets (GTSRB and Tumsimple) indicate that our method can not only detect sensor anomalies effectively and efficiently but also outperform state-of-the-art anomaly detection methods. Xingshuo Han, Yuan Zhou 0005, Kangjie Chen, Han Qiu 0001, Meikang Qiu, Yang Liu 0003, Tianwei Zhang 0004 |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2023 | Consensus-Clustering-Based Automatic Distribution Matching for Cross-Domain Image SteganalysisabstractImage steganalysis is a technique to detect whether an image contains hidden information. Although the existing cross-domain steganalysis methods have been presented to narrow the distribution gap between different domains, it is still challenging to effectively capture the transferable steganalysis representations under the condition of severe distribution shifts. To address this issue, we propose a novel consensus-clustering-based automatic distribution matching scheme, called CADM, which can automatically and accurately match inconsistent distributions in cross-domain steganalysis scenarios. First, the original steganalysis features are clustered by the spatially constrained fuzzyc-means (SCFCM) algorithm with controllable parameters to fully perceive and mine inherent structural relationships. Subsequently, the cluster consensus knowledge is derived from the perspective of intra-domain and inter-domain to facilitate the clustering and the matching. In this way, the representations of weak stego signals can be augmented by identifying cluster centers that can be combined across domains. Ultimately, the cycle-consistent optimization and adaptation is achieved by gradually adjusting the learning strength of well-aligned and poorly-aligned samples to promote the positive transfer of overlapped clusters and prevent the negative transfer of outlier clusters. Furthermore, extensive experiments on various benchmark databases for cross-domain steganalysis demonstrate the superiority of CADM over the current state-of-the-art methods. Ju Jia, Meng Luo 0002, Siqi Ma 0001, Lina Wang 0001, Yang Liu 0003 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2023 | A Survey on Automated Driving System Testing: Landscapes and TrendsabstractAutomated Driving Systems ( ADS ) have made great achievements in recent years thanks to the efforts from both academia and industry. A typical ADS is composed of multiple modules, including sensing, perception, planning, and control, which brings together the latest advances in different domains. Despite these achievements, safety assurance of ADS is of great significance, since unsafe behavior of ADS can bring catastrophic consequences. Testing has been recognized as an important system validation approach that aims to expose unsafe system behavior; however, in the context of ADS, it is extremely challenging to devise effective testing techniques, due to the high complexity and multidisciplinarity of the systems. There has been great much literature that focuses on the testing of ADS, and a number of surveys have also emerged to summarize the technical advances. Most of the surveys focus on the system-level testing performed within software simulators, and they thereby ignore the distinct features of different modules. In this article, we provide a comprehensive survey on the existing ADS testing literature, which takes into account both module-level and system-level testing. Specifically, we make the following contributions: (1) We survey the module-level testing techniques for ADS and highlight the technical differences affected by the features of different modules; (2) we also survey the system-level testing techniques, with focuses on the empirical studies that summarize the issues occurring in system development or deployment, the problems due to the collaborations between different modules, and the gap between ADS testing in simulators and the real world; and (3) we identify the challenges and opportunities in ADS testing, which pave the path to the future research in this field. Shuncheng Tang, Zhenya Zhang 0001, Jixiang Zhou, Shuang Liu 0007, Shengjian Guo, Yan-Fu Li, Lei Ma 0003, Yinxing Xue, Yang Liu 0003 |
ACM Trans. Softw. Eng. Methodol. | 11 |
| 2023 | LiDetector: License Incompatibility Detection for Open Source SoftwareabstractOpen-source software (OSS) licenses dictate the conditions, which should be followed to reuse, distribute, and modify software. Apart from widely-used licenses such as the MIT License, developers are also allowed to customize their own licenses (called custom license), whose descriptions are more flexible. The presence of such various licenses imposes challenges to understand licenses and their compatibility. To avoid financial and legal risks, it is essential to ensure license compatibility when integrating third-party packages or reusing code accompanied with licenses. In this work, we propose LiDetector , an effective tool that extracts and interprets OSS licenses (including both official licenses and custom licenses), and detects license incompatibility among these licenses. Specifically, LiDetector introduces a learning-based method to automatically identify meaningful license terms from an arbitrary license, and employs Probabilistic Context-Free Grammar (PCFG) to infer rights and obligations for incompatibility detection. Experiments demonstrate that LiDetector outperforms existing methods with 93.28% precision for term identification, and 91.09% accuracy for right and obligation inference, and can effectively detect incompatibility with 10.06% FP rate and 2.56% FN rate. Furthermore, with LiDetector , our large-scale empirical study on 1,846 projects reveals that 72.91% of the projects are suffering from license incompatibility, including popular ones such as the MIT License and the Apache License. We highlighted lessons learned from perspectives of different stakeholders and made all related data and the replication package publicly available to facilitate follow-up research. Sihan Xu, Lingling Fan 0003, Zheli Liu, Yang Liu 0003, Hua Ji |
ACM Trans. Softw. Eng. Methodol. | 5 |
| 2023 | Towards Practical Binary Code Similarity Detection: Vulnerability Verification via Patch Semantic AnalysisabstractVulnerability is a major threat to software security. It has been proven that binary code similarity detection approaches are efficient to search for recurring vulnerabilities introduced by code sharing in binary software. However, these approaches suffer from high false-positive rates (FPRs) since they usually take the patched functions as vulnerable, and they usually do not work well when binaries are compiled with different compilation settings. To this end, we propose an approach, named Robin , to confirm recurring vulnerabilities by filtering out patched functions. Robin is powered by a lightweight symbolic execution to solve the set of function inputs that can lead to the vulnerability-related code. It then executes the target functions with the same inputs to capture the vulnerable or patched behaviors for patched function filtration. Experimental results show that Robin achieves high accuracy for patch detection across different compilers and compiler optimization levels respectively on 287 real-world vulnerabilities of 10 different software. Based on accurate patch detection, Robin significantly reduces the false-positive rate of state-of-the-art vulnerability detection tools (by 94.3% on average), making them more practical. Robin additionally detects 12 new potentially vulnerable functions. Shouguo Yang, Zhengzi Xu, Yang Xiao 0011, Zhe Lang, Yang Liu 0003, Zhiqiang Shi, Hong Li 0004, Limin Sun 0001 |
ACM Trans. Softw. Eng. Methodol. | 6 |
| 2023 | Service Pattern Optimization: Focusing on Collaboration in Service EcosystemsabstractThe service pattern is an abstraction of the business relationship among various participants from the service ecosystem in four aspects: workflow, data flow, resource flow, and value flow. In order to optimize service patterns, it is necessary to consider the collaboration between participants as well as the interaction among different servers. The existing works either optimize the former by adjusting service orchestration, such as business process optimization and workflow optimization, or focus on the latter through adjusting service distribution, such as cloud service distribution optimization and edge service deployment optimization. However, the prevalence of service ecosystems and distributed computing has begun to make multi-user, multi-server scenarios commonplace, placing greater importance on fast and effective optimization of service patterns. In this work, we summarize the constraints and objectives and formally define the service pattern optimization problem. Beyond that, we propose a service pattern optimization-oriented confidence aware recurrent simulated annealing algorithm (PooCa). Experiments conducted on an existing dataset show that our method outperforms the other three baselines on the overall dataset as well as on the eight subsets. Also, our method can reduce the number of search iterations by 41.15% on average with the same search space. We also carry out case studies on the online travel booking service pattern and investigate factors that make patterns perform better. Meng Xi 0002, Jianwei Yin, Zhengzi Xu, Ying Li 0001, Shuiguang Deng, Yang Liu 0003 |
IEEE Trans. Serv. Comput. | 6 |
| 2023 | Automatically Distilling Storyboard With Rich Features for Android AppsabstractBefore developing a new mobile app, the development team usually endeavors painstaking efforts to review many existing apps with similar purposes. The review process is crucial in the sense that it reduces market risks and provides inspirations for app development. However, manual exploration of hundreds of existing apps by different roles (e.g., product manager, UI/UX designer, developer, and tester) can be ineffective. For example, it is difficult to completely explore all the functionalities of the app from different aspects including design, implementation, and testing in a short period of time. However, existing reverse engineering tools only provide basic features such as AndroidManifest.xml and Java source files for users. Following the conception of storyboard in movie production, we propose a system, named StoryDistiller, to automatically generate the storyboards for Android apps with rich features through reverse engineering, and assist different roles to review and analyze apps effectively and efficiently. Specifically, we (1) propose a hybrid method to extract a relatively complete Activity transition graph (ATG), that is, it first extracts the ATG of Android apps through static analysis method first, and further leverages dynamic component exploration to augment ATG; (2) extract the required inter-component communication (ICC) data of each target Activity by leveraging static data-flow analysis and renders UI pages dynamically by using app instrumentation together with the extracted required ICC data; (3) obtain rich features including comprehensive ATG with rendered UI pages, semantic activity names, corresponding logic and layout code, etc. (4) implement the storyboard visualization as a web service with the rendered UI pages and the corresponding rich features. Our experiments unveil that StoryDistiller is effective and indeed useful to assist app exploration and review. We also conduct a comprehensive comparison study to demonstrate better performance over IC3, Gator, Stoat, and StoryDroid. Sen Chen 0001, Lingling Fan 0003, Chunyang Chen 0001, Yang Liu 0003 |
IEEE Trans. Software Eng. | 4 |
| 2023 | Test Report Generation for Android App Testing Via Heterogeneous Data AnalysisabstractThe rising of the Android market demands higher quality assurance of Android applications (apps) to sharpen the competitive edge, and techniques for traditional software have problems adapting for mobile apps. Android apps often require testing on a large-scale device cluster, which produces a large amount of test reports consisting of heterogeneous data, e.g., hardware information, GUI screenshots, runtime logs. Such data are hard to merge to be unified analyzed, while they serve as an essential basis for bug inspection and fixing. Existing test report generation or analysis techniques can only handle testing data from different devices separately. They simply list all the information to app developers and have no further processing to summarize test reports. Besides, they neglect the inner connection of the heterogeneous data. Such techniques cannot improve the report reviewing effectiveness and efficiency, and they can hardly find the inner links and rules of the bug occurrence on different devices. As a result, developers still need to devote many efforts to inspect and fix bugs. In this paper, a large amount of test reports are investigated by the authors, as to construct a structured bug model to analyze heterogeneous data of the testing results. According to the investigation, we also define theBug Inconsistencyof testing results from multiple devices and build a novel bug taxonomy. In general, an automated approach is proposed to generate structured and comprehensible test reports from raw testing results from multiple devices. Based on the approach, a tool, namelyBreGat, is implemented to evaluate the classification and deduplication capability of our approach. The experimental results of 30 Android apps on 20 devices show thatBreGatcan successfully cover 83% bug categories and exclude 76% duplicate bugs. Furthermore, a user study involving 16 developers shows that our test reports are more comprehensible andBreGatgreatly improves the bug inspection efficiency compared to the state-of-the-art tool. Chunrong Fang, Shengcheng Yu, Ting Su 0001, Yuanhan Tian, Yang Liu 0003 |
IEEE Trans. Software Eng. | 6 |
| 2023 | A Comprehensive Study on ARM Disassembly ToolsabstractEmbedded devices are becoming ubiquitous, and ARM is becoming the dominant architecture for them. Meanwhile, there is a pressing need to perform security assessments for these devices. Due to different types of peripherals, emulating the software, i.e., firmware, of these devices in scale is challenging. Therefore, static analysis is still widely used. Existing works usually leverage off-the-shelf tools to disassemble stripped ARM binaries and (implicitly) assume that reliably disassembling binaries is a solved problem. However, whether this assumption really holds is unknown. In this paper, we conduct the first comprehensive study on ARM disassembly tools. Specifically, we build 1,896 ARM binaries (including 248 obfuscated ones) with different compilers, compiling options, and obfuscation methods. We then evaluate them using eight state-of-the-art ARM disassembly tools (including both commercial and noncommercial ones) in three different versions on their capabilities to locate instruction boundary, function boundary, and function signature. Instruction and function boundary are two fundamental primitives that the other primitives are built upon while function signature is significant for control flow integrity (CFI) techniques. Our work reveals some observations that have not been systematically summarized and/or confirmed. For instance, we find that the existence of both ARM and Thumb instruction sets, and the reuse of theBLinstruction for both function calls and branches bring serious challenges to disassembly tools. Our evaluation sheds light on the limitations of state-of-the-art disassembly tools and points out potential directions for improvement. Muhui Jiang, Qinming Dai, Yajin Zhou, Xiapu Luo, Ruoyu Wang 0001, Yang Liu 0003, Kui Ren 0001 |
IEEE Trans. Software Eng. | 8 |
| 2023 | GraphSearchNet: Enhancing GNNs via Capturing Global Dependencies for Semantic Code SearchabstractCode search aims to retrieve accurate code snippets based on a natural language query to improve software productivity and quality. With the massive amount of available programs such as (on GitHub or Stack Overflow), identifying and localizing the precise code is critical for the software developers. In addition, Deep learning has recently been widely applied to different code-related scenarios, e.g., vulnerability detection, source code summarization. However, automated deep code search is still challenging since it requires a high-level semantic mapping between code and natural language queries. Most existing deep learning-based approaches for code search rely on the sequential text i.e., feeding the program and the query as a flat sequence of tokens to learn the program semantics while the structural information is not fully considered. Furthermore, the widely adopted Graph Neural Networks (GNNs) have proved their effectiveness in learning program semantics, however, they also suffer the problem of capturing the global dependencies in the constructed graph, which limits the model learning capacity. To address these challenges, in this paper, we design a novel neural network framework, named GraphSearchNet, to enable an effective and accurate source code search by jointly learning the rich semantics of both source code and natural language queries. Specifically, we propose to construct graphs for the source code and queries with bidirectional GGNN (BiGGNN) to capture the local structural information of the source code and queries. Furthermore, we enhance BiGGNN by utilizing the multi-head attention module to supplement the global dependencies that BiGGNN missed to improve the model learning capacity. The extensive experiments on Java and Python programming language from the public benchmark CodeSearchNet confirm that GraphSearchNet outperforms current state-of-the-art works by a significant margin. Shangqing Liu, Xiaofei Xie, Jing Kai Siow, Lei Ma 0003, Guozhu Meng, Yang Liu 0003 |
IEEE Trans. Software Eng. | 6 |
| 2023 | Demystifying Performance Regressions in String SolversabstractOver the past few years, SMT string solvers have found their applications in an increasing number of domains, such as program analyses in mobile and Web applications, which require the ability to reason about string values. A series of research has been carried out to find quality issues of string solvers in terms of its correctness and performance. Yet, none of them has considered the performance regressions happening across multiple versions of a string solver. To fill this gap, in this paper, we focus on solver performance regressions (SPRs), i.e., unintended slowdowns introduced during the evolution of string solvers. To this end, we developSPRFinderto not only generate test cases demonstrating SPRs, but also localize the probable causes of them, in terms of commits. We evaluated the effectiveness ofSPRFinderon three state-of-the-art string solvers, i.e., Z3Seq, Z3Str3, and CVC4. The results demonstrate thatSPRFinderis effective in generating SPR-inducing test cases and also able to accurately locate the responsible commits. Specifically, the average running time on the target versions is 13.2× slower than that of the reference versions. Besides, we also conducted the first empirical study to peek into the characteristics of SPRs, including the impact of random seed configuration for SPR detection, understanding the root causes of SPRs, and characterizing the regression test cases through case studies. Finally, we highlight that 149 unique SPR-inducing commits were discovered in total bySPRFinder, and 27of them have been confirmed by the corresponding developers. Yao Zhang 0019, Xiaofei Xie, Yi Li 0008, Yun Lin 0001, Sen Chen 0001, Yang Liu 0003, Xiaohong Li 0001 |
IEEE Trans. Software Eng. | 6 |
| 2023 | A Large-Scale Empirical Study of Real-Life Performance Issues in Open Source ProjectsabstractSoftware performance is a critical quality attribute that determines the success of a software system. However, practitioners lack comprehensive and holistic understanding of how real-life performance issues are caused and resolved in practice from the technical, engineering, and economic perspectives. This paper presents a large-scale empirical study of 570 real-life performance issues from 13 open source projects from various problem domains, and implemented in three popular programming languages, Java (192 issues), C/C++ (162 issues), and Python (216 issues). From the technical perspective, we summarizeeightgeneral types of performance issues with corresponding root causes and resolutions that apply for all three languages. We also identify available tools for detecting and resolving different types of issues from the literature. In addition, we found that 27% of the 570 issues are resolved by design-level optimization—coordinated revision of a group of related source files and their design structure. We reveal four typical design-level optimization patterns, includingclassic design patterns,change propagation,optimization clone, andparallel optimizationthat practitioners should be aware of in resolving performance issues. From the engineering perspective, this study analyzes how test code changes in performance optimization. We found that only 15% of the 570 performance issues involve revision of test code. In most cases, the revised test cases focus on the functional logic of the performance optimization, rather than directly evaluate the performance improvement. This finding points to the potential lack of engineering standard for formally verifying performance optimization in regression testing. Finally, from the economic perspective, we analyze the“Return On Investment”of performance optimization. We found that design-level optimization usually requires more investment, but not always yields to higher performance improvement. However, developers tend to use design-level optimization when they concern about other quality attributes, such as maintainability and readability. Lu Xiao 0001, Andre B. Bondi, Bihuan Chen 0001, Yang Liu 0003 |
IEEE Trans. Software Eng. | 5 |
| 2023 | Specification-Based Autonomous Driving System TestingabstractAutonomous vehicle (AV) systems must be comprehensively tested and evaluated before they can be deployed. High-fidelity simulators such as CARLA or LGSVL allow this to be done safely in very realistic and highly customizable environments. Existing testing approaches, however, fail to test simulated AVs systematically, as they focus on specific scenarios and oracles (e.g., lane following scenario with the “no collision” requirement) and lack any coverage criteria measures. In this paper, we propose$\mathtt {AVUnit}$, a framework for systematically testing AV systems against customizable correctness specifications. Designed modularly to support different simulators,$\mathtt {AVUnit}$consists of two new languages for specifying dynamic properties of scenes (e.g. changing pedestrian behaviour after waypoints) and fine-grained assertions about the AV's journey.$\mathtt {AVUnit}$further supports multiple fuzzing algorithms that automatically search for test cases that violate these assertions, using robustness and coverage measures as fitness metrics. We evaluated the implementation of$\mathtt {AVUnit}$for the LGSVL+Apollo simulation environment, finding 19 kinds of issues in Apollo, which indicate that the open-source Apollo does not perform well in complex intersections and lane-changing related scenarios. Yuan Zhou 0005, Yang Sun 0008, Yun Tang 0003, Yuqi Chen 0001, Jun Sun 0001, Christopher M. Poskitt, Yang Liu 0003, Zijiang Yang 0006 |
IEEE Trans. Software Eng. | 7 |
| 2022 | On the (In)Security of Secure ROS2abstractRobot Operating System (ROS) has been the mainstream platform for research and development of robotic applications. This platform is well-known for lacking security features and efficiency for distributed robotic computations. To address these issues, ROS2 is recently developed by utilizing the Data Distribution Service (DDS) to provide security support. Integrated with DDS, ROS2 is expected to establish the basis for trustworthy robotic ecosystems. Gelei Deng, Guowen Xu, Yuan Zhou 0005, Tianwei Zhang 0004, Yang Liu 0003 |
CCS | 5 |
| 2022 | Can You Spot the Chameleon? Adversarially Camouflaging Images from Co-Salient Object DetectionabstractCo-salient object detection (CoSOD) has recently achieved significant progress and played a key role in retrieval-related tasks. However, it inevitably poses an entirely new safety and security issue, i.e., highly personal and sensitive content can potentially be extracting by powerful CoSOD methods. In this paper, we address this problem from the perspective of adversarial attacks and identify a novel task: adversarial co-saliency attack. Specially, given an image selected from a group of images containing some common and salient objects, we aim to generate an adversarial version that can mislead CoSOD methods to predict incorrect co-salient regions. Note that, compared with general white-box adversarial attacks for classification, this new task faces two additional challenges: (1) low success rate due to the diverse appearance of images in the group; (2) low transferability across CoSOD methods due to the considerable difference between CoSOD pipelines. To address these challenges, we propose the very first blackbox joint adversarial exposure and noise attack (Jadena), where we jointly and locally tune the exposure and additive perturbations of the image according to a newly designed high-feature-level contrast-sensitive loss function. Our method, without any information on the state-of-the-art CoSOD methods, leads to significant performance degradation on various co-saliency detection datasets and makes the co-salient objects undetectable. This can have strong practical benefits in properly securing the large number of personal photos currently shared on the Internet. Moreover, our method is potential to be utilized as a metric for evaluating the robustness of CoSOD methods. Ruijun Gao, Qing Guo 0005, Felix Juefei-Xu, Hongkai Yu, Huazhu Fu, Wei Feng 0005, Yang Liu 0003, Song Wang 0002 |
CVPR | 7 |
| 2022 | A Formal Methodology for Verifying Side-Channel Vulnerabilities in Cache Architectures
Ke Jiang 0001, Tianwei Zhang 0004, David Sanán, Yongwang Zhao, Yang Liu 0003 |
ICFEM | 5 |
| 2022 | Windranger: A Directed Greybox Fuzzer driven by Deviation Basic BlocksabstractDirected grey-box fuzzing (DGF) is a security testing technique that aims to steer the fuzzer towards predefined target sites in the program. To gain directedness, DGF prioritizes the seeds whose execution traces are closer to the target sites. Therefore, evaluating the distance between the execution trace of a seed and the target sites (aka, the seed distance) is important for DGF. The first directed grey-box fuzzer, AFLGo, uses an approach of calculating the basic block level distances during static analysis and accumulating the distances of the executed basic blocks to compute the seed distance. Following AFLGo, most of the existing state-of-the-art DGF techniques use all the basic blocks on the execution trace and only the control flow information for seed distance calculation. However, not every basic block is equally important and there are certain basic blocks where the execution trace starts to deviate from the target sites (aka, deviation basic blocks). Zhengjie Du, Yuekang Li, Yang Liu 0003, Bing Mao 0001 |
ICSE | 3 |
| 2022 | Demystifying the Vulnerability Propagation and Its Evolution via Dependency Trees in the NPM EcosystemabstractThird-party libraries with rich functionalities facilitate the fast development of JavaScript software, leading to the explosive growth of the NPM ecosystem. However, it also brings new security threats that vulnerabilities could be introduced through dependencies from third-party libraries. In particular, the threats could be excessively amplified by transitive dependencies. Existing research only considers direct dependencies or reasoning transitive dependencies based on reachability analysis, which neglects the NPM-specific dependency resolution rules as adapted during real installation, resulting in wrongly resolved dependencies. Consequently, further fine-grained analysis, such as precise vulnerability propagation and their evolution over time in dependencies, cannot be carried out precisely at a large scale, as well as deriving ecosystem-wide solutions for vulnerabilities in dependencies. Sen Chen 0001, Lingling Fan 0003, Bihuan Chen 0001, Yang Liu 0003, Xin Peng 0001 |
ICSE | 5 |
| 2022 | Morest: Model-based RESTful API Testing with Execution FeedbackabstractRESTful APIs are arguably the most popular endpoints for accessing Web services. Blackbox testing is one of the emerging techniques for ensuring the reliability of RESTful APIs. The major challenge in testing RESTful APIs is the need for correct sequences of API operation calls for in-depth testing. To build meaningful operation call sequences, researchers have proposed techniques to learn and utilize the API dependencies based on OpenAPI specifications. However, these techniques either lack the overall awareness of how all the APIs are connected or the flexibility of adaptively fixing the learned knowledge. Yi Liu 0069, Yuekang Li, Gelei Deng, Yang Liu 0003, Ruiyuan Wan, Runchao Wu, Dandan Ji, Shiheng Xu, Minli Bao |
ICSE | 4 |
| 2022 | ModX: Binary Level Partially Imported Third-Party Library Detection via Program Modularization and Semantic MatchingabstractWith the rapid growth of software, using third-party libraries (TPLs) has become increasingly popular. The prosperity of the library usage has provided the software engineers with a handful of methods to facilitate and boost the program development. Unfortunately, it also poses great challenges as it becomes much more difficult to manage the large volume of libraries. Researches and studies have been proposed to detect and understand the TPLs in the software. However, most existing approaches rely on syntactic features, which are not robust when these features are changed or deliberately hidden by the adversarial parties. Moreover, these approaches typically model each of the imported libraries as a whole, therefore, cannot be applied to scenarios where the host software only partially uses the library code segments. Zhengzi Xu, Hongxu Chen 0001, Yang Liu 0003, Xiaorui Gong, Baoxu Liu |
ICSE | 4 |
| 2022 | Efficient greybox fuzzing of applications in Linux-based IoT devices via enhanced user-mode emulationabstractGreybox fuzzing has become one of the most effective vulnerability discovery techniques. However, greybox fuzzing techniques cannot be directly applied to applications in IoT devices. The main reason is that executing these applications highly relies on specific system environments and hardware. To execute the applications in Linux-based IoT devices, most existing fuzzing techniques use full-system emulation for the purpose of maximizing compatibility. However, compared with user-mode emulation, full-system emulation suffersfrom great overhead. Therefore, some previous works, such as Firm-AFL, propose to combine full-system emulation and user-mode emulation to speed up the fuzzing process. Despite the attempts of trying to shift the application towards user-mode emulation, no existing technique supports to execute these applications fully in the user-mode emulation. To address this issue, we propose EQUAFL, which can automatically set up the execution environment to execute embedded applications under user-mode emulation. EQUAFL first executes the application under full-system emulation and observe for the key points where the program may get stuck or even crash during user-mode emulation. With the observed information, EQUAFL can migrate the needed environment for user-mode emulation. Then, EQUAFL uses an enhanced user-mode emulation to replay system calls of network, and resource management behaviors to fulfill the needs of the embedded application during its execution. We evaluate EQUAFL on 70 network applications from different series of IoT devices. The result shows EQUAFL outperforms the state-of-the-arts in fuzzing efficiency (on average, 26 times faster than AFL-QEMU with full-system emulation, 14 times than Firm-AFL). We have also discovered ten vulnerabilities including six CVEs from the tested firmware images. Yaowen Zheng, Yuekang Li, Cen Zhang, Hongsong Zhu, Yang Liu 0003, Limin Sun 0001 |
ISSTA | 5 |
| 2022 | AUSERA: Automated Security Vulnerability Detection for Android AppsabstractTo reduce the attack surface from app source code, massive tools focus on detecting security vulnerabilities in Android apps. However, some obvious weaknesses have been highlighted in the previous studies. For example, (1) most of the available tools such as AndroBugs, MobSF, Qark, and Super use pattern-based methods to detect security vulnerabilities. Although they are effective in detecting some types of vulnerabilities, a large number of false positives would be introduced, which inevitably increases the patching overhead for app developers. (2) Similarly, static taint analysis tools such as FlowDroid and IccTA present hundreds of vulnerability candidates of data leakage instead of confirmed vulnerabilities. (3) Last but not least, a relatively complete vulnerability taxonomy is missing, which would introduce a lot of false negatives. In this paper, based on our prior knowledge in this research domain, we empirically propose a vulnerability taxonomy as the baseline and then extend AUSERA by augmenting the detection capability to 50 security vulnerability types. Meanwhile, a new benchmark dataset including all these 50 vulnerability types is constructed to demonstrate the effectiveness of AUSERA. The tool and datasets are available at https://github.com/tjusenchen/AUSERA and the demonstration video can be found at https://youtu.be/UCiGwVaFPpY. Sen Chen 0001, Lingling Fan 0003, Jiaming Li 0013, Yang Liu 0003 |
ASE | 5 |
| 2022 | TransRepair: Context-aware Program Repair for Compilation ErrorsabstractAutomatically fixing compilation errors can greatly raise the productivity of software development, by guiding the novice or AI programmers to write and debug code. Recently, learning-based program repair has gained extensive attention and became the state-of-the-art in practice. But it still leaves plenty of space for improvement. In this paper, we propose an end-to-end solution TransRepair to locate the error lines and create the correct substitute for a C program simultaneously. Superior to the counterpart, our approach takes into account the context of erroneous code and diagnostic compilation feedback. Then we devise a Transformer-based neural network to learn the ways of repair from the erroneous code as well as its context and the diagnostic feedback. To increase the effectiveness of TransRepair, we summarize 5 types and 74 fine-grained sub-types of compilations errors from two real-world program datasets and the Internet. Then a program corruption technique is developed to synthesize a large dataset with 1,821,275 erroneous C programs. Through the extensive experiments, we demonstrate that TransRepair outperforms the state-of-the-art in both single repair accuracy and full repair accuracy. Further analysis sheds light on the strengths and weaknesses in the contemporary solutions for future improvement. Shangqing Liu, Guozhu Meng, Xiaofei Xie, Kai Chen 0012, Yang Liu 0003 |
ASE | 7 |
| 2022 | Morest: Industry Practice of Automatic RESTful API TestingabstractMany big companies are providing cloud services through RESTful APIs nowadays. With the growing popularity of RESTful API, testing RESTful API becomes crucial. To address this issue, researchers have proposed several automatic RESTful API testing techniques. At Huawei, we design and implement an automatic RESTful API testing framework named Morest. Morest has been used to test ten RESTful API services and helped to detected 83 previously unknown bugs which were all confirmed and fixed by the developers. On one hand, we find that Morest shows great capability of detecting bugs in RESTful API s. On the other hand, we also notice that human effort is inevitable and important when applying automatic RESTful API techniques in practice. Yi Liu 0069, Yuekang Li, Yang Liu 0003, Ruiyuan Wan, Runchao Wu, Qingkun Liu |
ASE | 3 |
| 2022 | Towards Understanding the Faults of JavaScript-Based Deep Learning SystemsabstractQuality assurance is of great importance for deep learning (DL) systems, especially when they are applied in safety-critical applications. While quality issues of native DL applications have been extensively analyzed, the issues of JavaScript-based DL applications have never been systematically studied. Compared with native DL applications, JavaScript-based DL applications can run on major browsers, making the platform- and device-independent. Specifically, the quality of JavaScript-based DL applications depends on the 3 parts: the application, the third-party DL library used and the underlying DL framework (e.g., TensorFlow.js), called JavaScript-based DL system. In this paper, we conduct the first empirical study on the quality issues of JavaScript-based DL systems. Specifically, we collect and analyze 700 real-world faults from relevant GitHub repositories, including the official TensorFlow.js repository, 13 third-party DL libraries, and 58 JavaScript-based DL applications. To better understand the characteristics of these faults, we manually analyze and construct taxonomies for the fault symptoms, root causes, and fix patterns, respectively. Moreover, we also study the fault distributions of symptoms and root causes, in terms of the different stages of the development lifecycle, the 3-level architecture in the DL system, and the 4 major components of TensorFlow.js framework. Based on the results, we suggest actionable implications and research avenues that can potentially facilitate the development, testing, and debugging of JavaScript-based DL systems. Lili Quan 0001, Xiaofei Xie, Sen Chen 0001, Xiaohong Li 0001, Yang Liu 0003 |
ASE | 6 |
| 2022 | Towards Understanding Third-party Library Dependency in C/C++ EcosystemabstractThird-party libraries (TPLs) are frequently reused in software to reduce development cost and the time to market. However, external library dependencies may introduce vulnerabilities into host applications. The issue of library dependency has received considerable critical attention. Many package managers, such as Maven, Pip, and NPM, are proposed to manage TPLs. Moreover, a significant amount of effort has been put into studying dependencies in language ecosystems like Java, Python, and JavaScript except C/C++. Due to the lack of a unified package manager for C/C++, existing research has only few understanding of TPL dependencies in the C/C++ ecosystem, especially at large scale. Zhengzi Xu, Shouguo Yang, Yi Li 0008, Yang Liu 0003 |
ASE | 8 |
| 2022 | Has My Release Disobeyed Semantic Versioning? Static Detection Based on Semantic DifferencingabstractTo enhance the compatibility in the version control of Java Third-party Libraries (TPLs), Maven adopts Semantic Versioning (SemVer) to standardize the underlying meaning of versions, but users could still confront abnormal execution and crash after upgrades even if compilation and linkage succeed. It is caused by semantic breaking (SemB) issues, such that APIs directly used by users have identical signatures but inconsistent semantics across upgrades. To strengthen compliance with SemVer rules, developers and users should be alerted of such issues. Unfortunately, it is challenging to detect them statically, because semantic changes in the internal methods of APIs are difficult to capture. Dynamic testing can confirmingly uncover some, but it is limited by inadequate coverage. Lyuye Zhang, Zhengzi Xu, Sen Chen 0001, Lingling Fan 0003, Bihuan Chen 0001, Yang Liu 0003 |
ASE | 7 |
| 2022 | A3GAN: Attribute-Aware Anonymization Networks for Face De-identificationabstractFace de-identification (De-ID) removes face identity information in face images to avoid personal privacy leakage. Existing face De-ID breaks the raw identity by cutting out the face regions and recovering the corrupted regions via deep generators, which inevitably affect the generation quality and cannot control generation results according to subsequent intelligent tasks (eg., facial expression recognition). In this work, for the first attempt, we think the face De-ID from the perspective of attribute editing and propose an attribute-aware anonymization network (A3GAN) by formulating face De-ID as a joint task of semantic suppression and controllable attribute injection. Intuitively, the semantic suppression removes the identity-sensitive information in embeddings while the controllable attribute injection automatically edits the raw face along the attributes that benefit De-ID. To this end, we first design a multi-scale semantic suppression network with a novel suppressive convolution unit (SCU), which can remove the face identity along multi-level deep features progressively. Then, we propose an attribute-aware injective network (AINet) that can generate De-ID-sensitive attributes in a controllable way (i.e., specifying which attributes can be changed and which cannot) and inject them into the latent code of the raw face. Moreover, to enable effective training, we design a new anonymization loss to let the injected attributes shift far away from the original ones. We perform comprehensive experiments on four datasets covering four different intelligent tasks including face verification, face detection, facial expression recognition, and fatigue detection, all of which demonstrate the superiority of our face De-ID over state-of-the-art methods. Liming Zhai, Qing Guo 0005, Xiaofei Xie, Lei Ma 0003, Yi Estelle Wang, Yang Liu 0003 |
ACM Multimedia | 6 |
| 2022 | GALOIS: Boosting Deep Reinforcement Learning via Generalizable Logic SynthesisabstractDespite achieving superior performance in human-level control problems, unlike humans, deep reinforcement learning (DRL) lacks high-order intelligence (e.g., logic deduction and reuse), thus it behaves ineffectively than humans regarding learning and generalization in complex problems. Previous works attempt to directly synthesize a white-box logic program as the DRL policy, manifesting logic-driven behaviors. However, most synthesis methods are built on imperative or declarative programming, and each has a distinct limitation, respectively. The former ignores the cause-effect logic during synthesis, resulting in low generalizability across tasks. The latter is strictly proof-based, thus failing to synthesize programs with complex hierarchical logic. In this paper, we combine the above two paradigms together and propose a novel Generalizable Logic Synthesis (GALOIS) framework to synthesize hierarchical and strict cause-effect logic programs. GALOIS leverages the program sketch and defines a new sketch-based hybrid program language for guiding the synthesis. Based on that, GALOIS proposes a sketch-based program synthesis method to automatically generate white-box programs with generalizable and interpretable cause-effect logic. Extensive evaluations on various decision-making tasks with complex logic demonstrate the superiority of GALOIS over mainstream baselines regarding the asymptotic performance, generalizability, and great knowledge reusability across different environments. Yushi Cao, Tianpei Yang, Hao Zhang 0004, Yan Zheng 0002, Yi Li 0008, Jianye Hao, Yang Liu 0003 |
NeurIPS | 8 |
| 2022 | An Exploratory Study for GUI Posts on Stack OverflowabstractGraphical User Interface (GUI) has become one of the most effective human-computer communication medium today. The quality of GUI is essential to the success of apps, especially for mobile apps. Developers not only have to understand the interaction of various components, but also follow the principles of design and implementation. It is helpful for developers to understand the challenges via analyzing the questions and answers (Q&A) on GUI development. However, there is no large-scale study on the GUI development posts on Stack Overflow. In this paper, we conduct an exploratory study on 23,741 posts related to GUI development on Stack Overflow. We first extract 20 topics related to GUI development using topic modeling. After manually classifying these GUI topics into 5 categories, we further quantitatively analyze the popularity and difficulty of GUI topics, the correlation between these two aspects, and qualitatively analyze the distribution of question types in posts. Finally, we have some interesting findings. These findings contain that the topic "tool selection" is the most popular topic, the topic "thread" has the highest percentage of unaccepted answers, and the topic "client/server" answer takes the longest time to be accepted. In addition, we discuss about possible inspirations of our research to GUI development stakeholders. Liming Nie, Yang Liu 0003, Zuohua Ding, Jifeng Xuan |
QRS | 3 |
| 2022 | Tracking patches for open source software vulnerabilitiesabstractOpen source software (OSS) vulnerabilities threaten the security of software systems that use OSS. Vulnerability databases provide valuable information (e.g., vulnerable version and patch) to mitigate OSS vulnerabilities. There arises a growing concern about the information quality of vulnerability databases. However, it is unclear what the quality of patches in existing vulnerability databases is; and existing manual or heuristic-based approaches for patch tracking are either too expensive or too specific to apply to all OSS vulnerabilities. Congying Xu, Bihuan Chen 0001, Chenhao Lu, Kaifeng Huang 0001, Xin Peng 0001, Yang Liu 0003 |
ESEC/SIGSOFT FSE | 6 |
| 2022 | RegexScalpel: Regular Expression Denial of Service (ReDoS) Defense by Localize-and-Fix
Yeting Li, Yecheng Sun, Zhiwu Xu 0001, Jialun Cao, Yuekang Li, Rongchen Li, Haiming Chen 0001, Shing-Chi Cheung, Yang Liu 0003, Yang Xiao 0011 |
USENIX Security Symposium | 9 |
| 2022 | Fair and accurate age prediction using distribution aware data curation and augmentationabstractDeep learning-based facial recognition systems have experienced increased media attention due to exhibiting unfair behavior. Large enterprises, such as IBM, shut down their facial recognition and age prediction systems as a consequence. Age prediction is an especially difficult application with the issue of fairness remaining an open research problem (e.g., predicting age for different ethnicity equally accurate). One of the main causes of unfair behavior in age prediction methods lies in the distribution and diversity of the training data. In this work, we present two novel approaches for dataset curation and data augmentation in order to increase fairness through balanced feature curation and increase diversity through distribution aware augmentation. To achieve this, we introduce out-of-distribution detection to the facial recognition domain which is used to select the data most relevant to the deep neural network’s (DNN) task when balancing the data among age, ethnicity, and gender. Our approach shows promising results. Our best-trained DNN model outperformed all academic and industrial baselines in terms of fairness by up to 4.92 times and also enhanced the DNN’s ability to generalize outperforming Amazon AWS and Microsoft Azure public cloud systems by 31.88% and 10.95%, respectively. Yushi Cao, David Berend, Palina Tolmach, Guy Amit, Moshe Levy, Yang Liu 0003, Asaf Shabtai, Yuval Elovici |
WACV | 6 |
| 2022 | Learning Program Semantics with Code Representations: An Empirical StudyabstractProgram semantics learning is the core and fundamental for various code intelligent tasks e.g., vulnerability detection, clone detection. A considerable amount of existing works propose diverse approaches to learn the program semantics for different tasks and these works have achieved state-of-the-art performance. However, currently, a comprehensive and systematic study on evaluating different program representation techniques across diverse tasks is still missed. From this starting point, in this paper, we conduct an empirical study to evaluate different program representation techniques. Specifically, we categorize current mainstream code representation techniques into four categories i.e., Feature-based, Sequence-based, Tree-based, and Graph-based program representation technique and evaluate its performance on three diverse and popular code intelligent tasks i.e., Code Classification, Vulnerability Detection, and Clone Detection on the public released benchmark. We further design three research questions (RQs) and conduct a comprehensive analysis to investigate the performance. By the extensive experimental results, we conclude that (1) The graph-based representation is superior to the other selected techniques across these tasks. (2) Compared with the node type information used in tree-based and graph-based representations, the node textual information is more critical to learning the program semantics. (3) Different tasks require the task-specific semantics to achieve their highest performance, however combining various program semantics from different dimensions such as control dependency, data dependency can still produce promising results. Jing Kai Siow, Shangqing Liu, Xiaofei Xie, Guozhu Meng, Yang Liu 0003 |
SANER | 5 |
| 2022 | Towards characterizing bug fixes through dependency-level changes in Apache Java open source projects
Lingling Fan 0003, Sen Chen 0001, Yuanfang Cai, Yang Liu 0003, Ting Liu 0002 |
Sci. China Inf. Sci. | 6 |
| 2022 | Characterizing usages, updates and risks of third-party libraries in Java projects
Kaifeng Huang 0001, Bihuan Chen 0001, Congying Xu, Xin Peng 0001, Yijian Wu, Yang Liu 0003 |
Empir. Softw. Eng. | 8 |
| 2022 | Countering Malicious DeepFakes: Survey, Battleground, and Horizon
Felix Juefei-Xu, Run Wang 0001, Yihao Huang 0001, Qing Guo 0005, Lei Ma 0003, Yang Liu 0003 |
Int. J. Comput. Vis. | 6 |
| 2022 | Enriching query semantics for code search with reinforcement learning
Chaozheng Wang, Zhenhao Nong, Cuiyun Gao 0001, Zongjie Li, Jichuan Zeng, Zhenchang Xing, Yang Liu 0003 |
Neural Networks | 7 |
| 2022 | Online adaptation for autonomous unmanned systems driven by requirements satisfaction model
Yixing Luo, Yuan Zhou 0005, Haiyan Zhao 0001, Zhi Jin 0001, Tianwei Zhang 0004, Yang Liu 0003, Danny Barthaud, Yijun Yu 0001 |
Softw. Syst. Model. | 6 |
| 2022 | SafeOSL: Ensuring memory safety of C via ownership-based intermediate languageabstractAbstract The unsafe features of C make it a big challenge to ensure memory safety of C programs, and often lead to memory errors that can result in vulnerabilities. Various formal verification techniques for ensuring memory safety of C have been proposed. However, most of them either have a high overhead, such as state explosion problem in model checking, or have false positives, such as abstract interpretation. In this article, by innovatively borrowing ownership system from Rust, we propose a novel and sound static memory safety analysis approach, named SafeOSL. Its basic idea is an ownership‐based intermediate language, called ownership system language (OSL), which captures the features of the ownership system in Rust. Ownership system specifies the relations among variables and memory locations, and maintains invariants that can ensure memory safety. The semantics of OSL is formalized in K‐framework, which is a rewriting‐logic based tool. C programs to be checked are first transformed into OSL programs and then detected by OSL semantics. Experimental results have demonstrated that SafeOSL is effective in detecting memory errors of C. Moreover, the translations and experiments indicate that the intermediate language OSL could be reused by other programming languages to detect memory errors. Xiaohua Yin, Shuanglong Kan, Guohua Shen, Zhe Chen 0011, Yang Liu 0003, Fei Wang 0032 |
Softw. Pract. Exp. | 6 |
| 2022 | Topology-Aware Differential Privacy for Decentralized Image ClassificationabstractImage classification is a fundamental artificial intelligence task that labels images into one of some predefined classes. However, training complex image classification models requires a large amount of computation resources and data in order to reach state-of-the-art performance. This demand drives the growth of distributed deep learning, where multiple agents cooperatively train global models with their individual datasets. Among such learning systems, decentralized learning is particularly attractive, as it can improve the efficiency and fault tolerance by eliminating the centralized parameter server, which could be the single point of failure or performance bottleneck. Although the agents do not need to disclose their training image samples, they exchange parameters with each other at each iteration, which can put them at the risk of data privacy leakage. Past works demonstrated the possibility of recovering training images from the exchanged parameters. One common defense direction is to adopt Differential Privacy (DP) to secure the optimization algorithms such as Stochastic Gradient Descent (SGD). Those DP-based methods mainly focus on standalone systems, or centralized distributed learning. How to enforce and optimize DP protection in decentralized learning systems is unknown and challenging, due to their complex communication topologies and distinct learning characteristics. In this paper, we design TOP- DP, a novel solution to optimize the differential privacy protection of decentralized image classification systems. The key insight of our solution is to leverage the unique features of decentralized communication topologies to reduce the noise scale and improve the model usability. (1) We enhance the DP-SGD algorithm with thistopology-awarenoise reduction strategy, and integrate the time-aware noise decay technique. (2) We design two novel learning protocols (synchronous and asynchronous) to protect systems with different network connectivities and topologies. We formally analyze and prove the DP requirement of our proposed solutions. Experimental evaluations demonstrate that our solution achieves a better trade-off between usability and privacy than prior works. To the best of our knowledge, this is the first DP optimization work from the perspective of network topologies. Shangwei Guo, Tianwei Zhang 0004, Guowen Xu, Han Yu 0001, Tao Xiang 0001, Yang Liu 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2022 | Byzantine-Resilient Decentralized Stochastic Gradient DescentabstractDecentralized learning has gained great popularity to improve learning efficiency and preserve data privacy. Each computing node makes equal contribution to collaboratively learn a Deep Learning model. The elimination of centralized Parameter Servers (PS) can effectively address many issues such as privacy, performance bottleneck and single-point-failure. However, how to achieve Byzantine Fault Tolerance in decentralized learning systems is rarely explored, although this problem has been extensively studied in centralized systems. In this paper, we present an in-depth study towards the Byzantine resilience of decentralized learning systems with two contributions. First, from the adversarial perspective, we theoretically illustrate that Byzantine attacks are more dangerous and feasible in decentralized learning systems: even one malicious participant can arbitrarily alter the models of other participants by sending carefully crafted updates to its neighbors. Second, from the defense perspective, we propose Ubar, a novel algorithm to enhance decentralized learning with Byzantine Fault Tolerance. Specifically, Ubar provides aUniformByzantine-resilientAggregationRule for benign nodes to select the useful parameter updates and filter out the malicious ones in each training iteration. It guarantees that each benign node in a decentralized system can train a correct model under very strong Byzantine attacks with an arbitrary number of faulty nodes. We conduct extensive experiments on standard image classification tasks and the results indicate that Ubar can effectively defeat both simple and sophisticated Byzantine attacks with higher performance efficiency than existing solutions. Shangwei Guo, Tianwei Zhang 0004, Han Yu 0001, Xiaofei Xie, Lei Ma 0003, Tao Xiang 0001, Yang Liu 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2022 | Oracle-Supported Dynamic Exploit Generation for Smart ContractsabstractDespite the high stakes involved in smart contracts, they are often developed in an undisciplined manner, leaving the security and reliability of blockchain transactions at risk. In this article, we introduce ContraMaster—an oracle-supported dynamic exploit generation framework for smart contracts. Existing approaches mutate only single transactions; ContraMaster exceeds these by mutating the transaction sequences. ContraMaster uses data-flow, control-flow, and the dynamic contract state to guide its mutations. It then monitors the executions of target contract programs, and validates the results against a general-purpose semantic test oracle to discover vulnerabilities. Being a dynamic technique, it guarantees that each discovered vulnerability is a violation of the test oracle and is able to generate the attack script to exploit this vulnerability. In contrast to rule-based approaches, ContraMaster has not shown any false positives, and it easily generalizes to unknown types of vulnerabilities (e.g., logic errors). We evaluate ContraMaster on 218 vulnerable smart contracts. The experimental results confirm its practical applicability and advantages over the state-of-the-art techniques, and also reveal three new types of attacks. Haijun Wang 0002, Ye Liu 0012, Yi Li 0008, Shangwei Lin 0001, Cyrille Artho, Lei Ma 0003, Yang Liu 0003 |
IEEE Trans. Dependable Secur. Comput. | 7 |
| 2022 | Superpixel-Based Collaborative and Low-Rank Regularization for Sparse Hyperspectral UnmixingabstractSparse unmixing (SU) has been widely applied to remotely sensed hyperspectral images interpretation. Compared with traditional unmixing algorithms, SU does not need to extract pure signatures (endmembers) from the image. The endmember matrix is constructed by directly selecting spectra from a known library which is used to estimate the fractional abundances associated with endmembers. This avoids the problem of extracting virtual endmembers without physical meaning. However, SU does not generally include spatial information, which may limit its performance. In order to address this limitation and include local spatial information, low-rank and sparse features in local regions can be exploited. In this paper, we include spatial information in the traditional SU algorithm by extracting low rank and spatial information based on superpixels, and further propose an algorithm named superpixel-based collaborative sparse and low-rank regularization for sparse unmixing (SCLRSU) to improve the performance of the traditional spatial regularization-based SU methods. In our proposed method, we combine superpixel segmentation and structural sparsity. Experiments are carried out on two simulated datasets and two real hyperspectral image datasets, and our results are compared with those obtained by traditional SU methods. Our results indicate that our newly proposed method provides very competitive performance. Tao Chen 0004, Yang Liu 0003, Yuxiang Zhang 0001, Bo Du 0001, Antonio Plaza |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | FakeLocator: Robust Localization of GAN-Based Face ManipulationsabstractFull face synthesis and partial face manipulation by virtue of the generative adversarial networks (GANs) and its variants have raised wide public concerns. In the multi-media forensics area, detecting and ultimately locating the image forgery has become an imperative task. In this work, we investigate the architecture of existing GAN-based face manipulation methods and observe that the imperfection of upsampling methods therewithin could be served as an important asset for GAN-synthesized fake image detection and forgery localization. Based on this basic observation, we have proposed a novel approach, termedFakeLocator, to obtain high localization accuracy, at full resolution, on manipulated facial images. To the best of our knowledge, this is the very first attempt to solve the GAN-based fake localization problem with a gray-scale fakeness map that preserves more information of fake regions. To improve the universality ofFakeLocatoracross multifarious facial attributes, we introduce an attention mechanism to guide the training of the model. To improve the universality ofFakeLocatoracross different DeepFake methods, we propose partial data augmentation and single sample clustering on the training images. Experimental results on popular FaceForensics++, DFFD datasets and seven different state-of-the-art GAN-based face generation methods have shown the effectiveness of our method. Compared with the baselines, our method performs better on various metrics. Moreover, the proposed method is robust against various real-world facial image degradations such as JPEG compression, low-resolution, noise, and blur. Yihao Huang 0001, Felix Juefei-Xu, Qing Guo 0005, Yang Liu 0003, Geguang Pu |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2022 | TANTRA: Timing-Based Adversarial Network Traffic Reshaping AttackabstractNetwork intrusion attacks are a known threat. To detect such attacks, network intrusion detection systems (NIDSs) have been developed and deployed. These systems apply machine learning models to high-dimensional vectors of features extracted from network traffic to detect intrusions. Advances in NIDSs have made it challenging for attackers, who must execute attacks without being detected by these systems. Prior research on bypassing NIDSs has mainly focused on perturbing the features extracted from the attack traffic to fool the detection system, however, this may jeopardize the attack’s functionality. In this work, we present TANTRA, a novel end-to-end Timing-based Adversarial Network Traffic Reshaping Attack that can bypass a variety of NIDSs. Our evasion attack utilizes a long short-term memory (LSTM) deep neural network (DNN) which is trained to learn the time differences between the target network’s benign packets. The trained LSTM is used to set the time differences between the malicious traffic packets (attack), without changing their content, such that they will “behave” like benign network traffic and will not be detected as an intrusion. We evaluate TANTRA on eight common intrusion attacks and three state-of-the-art NIDS systems, achieving an average success rate of 99.99% in network intrusion detection system evasion. We also propose a novel mitigation technique to address this new evasion attack. Yam Sharon, David Berend, Yang Liu 0003, Asaf Shabtai, Yuval Elovici |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2022 | Pasadena: Perceptually Aware and Stealthy Adversarial Denoise AttackabstractImage denoising can remove natural noise that widely exists in images captured by multimedia devices due to low-quality imaging sensors, unstable image transmission processes, or low light conditions. Recent works also find that image denoising benefits the high-level vision tasks,e.g., image classification. In this work, we try to challenge this common sense and explore a totally new problem,i.e., whether the image denoising can be given the capability of fooling the state-of-the-art deep neural networks (DNNs) while enhancing the image quality. To this end, we initiate the very first attempt to study this problem from the perspective of adversarial attack and propose theadversarial denoise attack. More specifically, our main contributions are three-fold:First, we identify a new task that stealthily embeds attacks inside the image denoising module widely deployed in multimedia devices as an image post-processing operation to simultaneously enhance the visual image quality and fool DNNs.Second, we formulate this new task as a kernel prediction problem for image filtering and propose theadversarial-denoising kernel predictionthat can produce adversarial-noiseless kernels for effective denoising and adversarial attacking simultaneously.Third, we implement an adaptiveperceptual region localizationto identify semantic-related vulnerability regions with which the attack can be more effective while not doing too much harm to the denoising. We name the proposed method asPasadena(Perceptually Aware and Stealthy Adversarial DENoise Attack) and validate our method on the NeurIPS’17 adversarial competition dataset, CVPR2021-AIC-VI: unrestricted adversarial attacks on ImageNet, and Tiny-ImageNet-C dataset. The comprehensive evaluation and analysis demonstrate that our method not only realizes denoising but also achieves a significantly higher success rate and transferability over state-of-the-art attacks. Yupeng Cheng, Qing Guo 0005, Felix Juefei-Xu, Shangwei Lin 0001, Wei Feng 0005, Weisi Lin, Yang Liu 0003 |
IEEE Trans. Multim. | 7 |
| 2022 | Breaking Neural Reasoning Architectures With Metamorphic Relation-Based Adversarial ExamplesabstractThe ability to read, reason, and infer lies at the heart of neural reasoning architectures. After all, the ability to perform logical reasoning over language remains a coveted goal of Artificial Intelligence. To this end, models such as the Turing-complete differentiable neural computer (DNC) boast of real logical reasoning capabilities, along with the ability to reason beyond simple surface-level matching. In this brief, we propose the first probe into DNC's logical reasoning capabilities with a focus on text-based question answering (QA). More concretely, we propose a conceptually simple but effective adversarial attack based on metamorphic relations. Our proposed adversarial attack reduces DNCs' state-of-the-art accuracy from 100% to 1.5% in the worst case, exposing weaknesses and susceptibilities in modern neural reasoning architectures. We further empirically explore possibilities to defend against such attacks and demonstrate the utility of our adversarial framework as a simple scalable method to improve model adversarial robustness. Alvin Chan, Lei Ma 0003, Felix Juefei-Xu, Yew-Soon Ong, Xiaofei Xie, Minhui Xue 0001, Yang Liu 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2022 | NPC: Neuron Path Coverage via Characterizing Decision Logic of Deep Neural NetworksabstractDeep learning has recently been widely applied to many applications across different domains, e.g., image classification and audio recognition. However, the quality of Deep Neural Networks (DNNs) still raises concerns in the practical operational environment, which calls for systematic testing, especially in safety-critical scenarios. Inspired by software testing, a number of structural coverage criteria are designed and proposed to measure the test adequacy of DNNs. However, due to the blackbox nature of DNN, the existing structural coverage criteria are difficult to interpret, making it hard to understand the underlying principles of these criteria. The relationship between the structural coverage and the decision logic of DNNs is unknown. Moreover, recent studies have further revealed the non-existence of correlation between the structural coverage and DNN defect detection, which further posts concerns on what a suitable DNN testing criterion should be. In this article, we propose the interpretable coverage criteria through constructing the decision structure of a DNN. Mirroring the control flow graph of the traditional program, we first extract a decision graph from a DNN based on its interpretation, where a path of the decision graph represents a decision logic of the DNN. Based on the control flow and data flow of the decision graph, we propose two variants of path coverage to measure the adequacy of the test cases in exercising the decision logic. The higher the path coverage, the more diverse decision logic the DNN is expected to be explored. Our large-scale evaluation results demonstrate that: The path in the decision graph is effective in characterizing the decision of the DNN, and the proposed coverage criteria are also sensitive with errors, including natural errors and adversarial examples, and strongly correlate with the output impartiality. Xiaofei Xie, Tianlin Li, Jian Wang 0067, Lei Ma 0003, Qing Guo 0005, Felix Juefei-Xu, Yang Liu 0003 |
ACM Trans. Softw. Eng. Methodol. | 7 |
| 2022 | Towards Robustness of Deep Program Processing Models - Detection, Estimation, and EnhancementabstractDeep learning (DL) has recently been widely applied to diverse source code processing tasks in the software engineering (SE) community, which achieves competitive performance (e.g., accuracy). However, the robustness, which requires the model to produce consistent decisions given minorly perturbed code inputs, still lacks systematic investigation as an important quality indicator. This article initiates an early step and proposes a framework CARROT for robustness detection, measurement, and enhancement of DL models for source code processing. We first propose an optimization-based attack technique CARROT A to generate valid adversarial source code examples effectively and efficiently. Based on this, we define the robustness metrics and propose robustness measurement toolkit CARROT M , which employs the worst-case performance approximation under the allowable perturbations. We further propose to improve the robustness of the DL models by adversarial training (CARROT T ) with our proposed attack techniques. Our in-depth evaluations on three source code processing tasks (i.e., functionality classification, code clone detection, defect prediction) containing more than 3 million lines of code and the classic or SOTA DL models, including GRU, LSTM, ASTNN, LSCNN, TBCNN, CodeBERT, and CDLH, demonstrate the usefulness of our techniques for ❶ effective and efficient adversarial example detection, ❷ tight robustness estimation, and ❸ effective robustness enhancement. Huangzhao Zhang, Zhiyi Fu, Ge Li 0001, Lei Ma 0003, Zhehao Zhao, Hua'an Yang, Yizhe Sun, Yang Liu 0003, Zhi Jin 0001 |
ACM Trans. Softw. Eng. Methodol. | 8 |
| 2022 | ReCDroid+: Automated End-to-End Crash Reproduction from Bug Reports for Android AppsabstractThe large demand of mobile devices creates significant concerns about the quality of mobile applications (apps). Developers heavily rely on bug reports in issue tracking systems to reproduce failures (e.g., crashes). However, the process of crash reproduction is often manually done by developers, making the resolution of bugs inefficient, especially given that bug reports are often written in natural language. To improve the productivity of developers in resolving bug reports, in this paper, we introduce a novel approach, called ReCDroid+, that can automatically reproduce crashes from bug reports for Android apps. ReCDroid+ uses a combination of natural language processing (NLP) , deep learning, and dynamic GUI exploration to synthesize event sequences with the goal of reproducing the reported crash. We have evaluated ReCDroid+ on 66 original bug reports from 37 Android apps. The results show that ReCDroid+ successfully reproduced 42 crashes (63.6% success rate) directly from the textual description of the manually reproduced bug reports. A user study involving 12 participants demonstrates that ReCDroid+ can improve the productivity of developers when resolving crash bug reports. Yu Zhao 0010, Ting Su 0001, Yang Liu 0003, Wei Zheng 0006, Xiaoxue Wu 0001, Ramakanth Kavuluru, William G. J. Halfond, Tingting Yu 0001 |
ACM Trans. Softw. Eng. Methodol. | 3 |
| 2022 | SPI: Automated Identification of Security Patches via CommitsabstractSecurity patches in open source software, providing security fixes to identified vulnerabilities, are crucial in protecting against cyber attacks. Security advisories and announcements are often publicly released to inform the users about potential security vulnerability. Despite the National Vulnerability Database (NVD) publishes identified vulnerabilities, a vast majority of vulnerabilities and their corresponding security patches remain beyond public exposure, e.g., in the open source libraries that are heavily relied on by developers. As many of these patches exist in open sourced projects, the problem of curating and gathering security patches can be difficult due to their hidden nature. An extensive and complete security patches dataset could help end-users such as security companies, e.g., building a security knowledge base, or researcher, e.g., aiding in vulnerability research. To efficiently curate security patches including undisclosed patches at large scale and low cost, we propose a deep neural-network-based approach built upon commits of open source repositories. First, we design and build security patch datasets that include 38,291 security-related commits and 1,045 Common Vulnerabilities and Exposures (CVE) patches from four large-scale C programming language libraries. We manually verify each commit, among the 38,291 security-related commits, to determine if they are security related. We devise and implement a deep learning-based security patch identification system that consists of two composite neural networks: one commit-message neural network that utilizes pretrained word representations learned from our commits dataset and one code-revision neural network that takes code before revision and after revision and learns the distinction on the statement level. Our system leverages the power of the two networks for Security Patch Identification. Evaluation results show that our system significantly outperforms SVM and K-fold stacking algorithms. The result on the combined dataset achieves as high as 87.93% F1-score and precision of 86.24%. We deployed our pipeline and learned model in an industrial production environment to evaluate the generalization ability of our approach. The industrial dataset consists of 298,917 commits from 410 new libraries that range from a wide functionalities. Our experiment results and observation on the industrial dataset proved that our approach can identify security patches effectively among open sourced projects. Yaqin Zhou, Jing Kai Siow, Shangqing Liu, Yang Liu 0003 |
ACM Trans. Softw. Eng. Methodol. | 5 |
| 2022 | SNIFF: Reverse Engineering of Neural Networks With Fault AttacksabstractNeural networks have been shown to be vulnerable against fault injection attacks. These attacks change the physical behavior of the device during the computation, resulting in a change of value that is currently being computed. They can be realized by various techniques, ranging from clock/voltage glitching to application of lasers to rowhammer. Previous works have mostly explored fault attacks for output misclassification, thus affecting the reliability of neural networks. In this article, we investigate the possibility to reverse engineer neural networks with fault attacks. Sign bit flip fault attack enables the reverse engineering by changing the sign of intermediate values. We develop the first exact extraction method on deep-layer feature extractor networks that provably allows the recovery of proprietary model parameters. Our experiments with Keras library show that the precision error for the parameter recovery for the tested networks is less than$10^{-13}$with the usage of 64-bit floats, which improves the current state of the art by six orders of magnitude. Jakub Breier, Dirmanto Jap, Xiaolu Hou, Shivam Bhasin, Yang Liu 0003 |
IEEE Trans. Reliab. | 5 |
| 2022 | Accessible or Not? An Empirical Investigation of Android App AccessibilityabstractMobile apps provide new opportunities to people with disabilities to act independently in the world. Following the law of the US, EU, mobile OS vendors such as Google and Apple have included accessibility features in their mobile systems and provide a set of guidelines and toolsets for ensuring mobile app accessibility. Motivated by this trend, researchers have conducted empirical studies by using the inaccessibility issue rate of each page (i.e., screen level) to represent the characteristics of mobile app accessibility. However, there still lacks an empirical investigation directly focusing on the issues themselves (i.e., issue level) to unveil more fine-grained findings, due to the lack of an effective issue detection method and a relatively comprehensive dataset of issues. To fill in this literature gap, we first propose an automated app page exploration tool, named Xbot, to facilitate app accessibility testing and automatically collect accessibility issues by leveraging the instrumentation technique and static program analysis. Owing to the relatively high activity coverage (around 80%) achieved by Xbot when exploring apps, Xbot achieves better performance on accessibility issue collection than existing testing tools such as Google Monkey. With Xbot, we are able to collect a relatively comprehensive accessibility issue dataset and finally collect 86,767 issues from 2,270 unique apps including both closed-source and open-source apps, based on which we further carry out an empirical study from the perspective of accessibility issues themselves to investigate novel characteristics of accessibility issues. Specifically, we extensively investigate these issues by checking 1) the overall severity of issues with multiple criteria, 2) the in-depth relation between issue types and app categories, GUI component types, 3) the frequent issue patterns quantitatively, and 4) the fixing status of accessibility issues. Finally, we highlight some insights to the community and hope to raise the attention to maintaining mobile app accessibility for users especially the elderly and disabled. Sen Chen 0001, Chunyang Chen 0001, Lingling Fan 0003, Mingming Fan 0001, Xian Zhan, Yang Liu 0003 |
IEEE Trans. Software Eng. | 6 |
| 2022 | ATOM: Commit Message Generation Based on Abstract Syntax Tree and Hybrid RankingabstractCommit messages record code changes (e.g., feature modifications and bug repairs) in natural language, and are useful for program comprehension. Due to the frequent updates of software and time cost, developers are generally unmotivated to write commit messages for code changes. Therefore, automating the message writing process is necessitated. Previous studies on commit message generation have been benefited from generation models or retrieval models, but the code structure of changed code, i.e., AST, which can be important for capturing code semantics, has not been explicitly involved. Moreover, although generation models have the advantages of synthesizing commit messages for new code changes, they are not easy to bridge the semantic gap between code and natural languages which could be mitigated by retrieval models. In this paper, we propose a novel commit message generation model, named ATOM, which explicitly incorporates the abstract syntax tree for representing code changes and integrates both retrieved and generated messages through hybrid ranking. Specifically, the hybrid ranking module can prioritize the most accurate message from both retrieved and generated messages regarding one code change. We evaluate the proposed model ATOM on our dataset crawled from 56 popular Java repositories. Experimental results demonstrate that ATOM increases the state-of-the-art models by 30.72 percent in terms of BLEU-4 (an accuracy measure that is widely used to evaluate text generation systems). Qualitative analysis also demonstrates the effectiveness of ATOM in generating accurate code commit messages. Shangqing Liu, Cuiyun Gao 0001, Sen Chen 0001, Lun Yiu Nie, Yang Liu 0003 |
IEEE Trans. Software Eng. | 5 |
| 2022 | Why My App Crashes? Understanding and Benchmarking Framework-Specific Exceptions of Android AppsabstractMobile apps have become ubiquitous. Ensuring their correctness and reliability is important. However, many apps still suffer from occasional to frequent crashes, weakening their competitive edge. Large-scale, deep analyses of the characteristics of real-world app crashes can provide useful insights to both developers and researchers. However, such studies are difficult and yet to be carried out — this work fills this gap. We collected 16,245 and 8,760 unique exceptions from 2,486 open-source and 3,230 commercial Android apps, respectively, and observed that the exceptions thrown from Android framework (termed“framework-specific exceptions”) account for the majority. With one-year effort, we (1) extensively investigated these framework-specific exceptions, and (2) further conducted an online survey of 135 professional app developers about how they analyze, test, reproduce and fix these exceptions. Specifically, we aim to understand the framework-specific exceptions from several perspectives: (i) their characteristics (e.g., manifestation locations, fault taxonomy), (ii) the developers’ testing practices, (iii) existing bug detection techniques’ effectiveness, (iv) their reproducibility and (v) bug fixes. To enable follow-up research (e.g., bug understanding, detection, localization and repairing), we further systematically constructed,DroidDefects, the first comprehensive and largest benchmark of Android app exception bugs. This benchmark contains 33reproducibleexceptions (with test cases, stack traces, faulty and fixed app versions, bug types, etc.), and 3,696ground-truthexceptions (real faults manifested by automated testing tools), which cover the apps with different complexities and diverse exception types. Based on our findings, we also built two prototype tools: Stoat+, an optimized dynamic testing tool, which quickly uncovered three previously-unknown, fixed crashes in Gmail and Google+; ExLocator, an exception localization tool, which can locate the root causes of specific exception types. Our dataset, benchmark and tools are publicly available onhttps://github.com/tingsu/droiddefects. Ting Su 0001, Lingling Fan 0003, Sen Chen 0001, Yang Liu 0003, Lihua Xu, Geguang Pu, Zhendong Su 0001 |
IEEE Trans. Software Eng. | 4 |
| 2022 | Research on Third-Party Libraries in Android Apps: A Taxonomy and Systematic Literature ReviewabstractThird-party libraries (TPLs) have been widely used in mobile apps, which play an essential part in the entire Android ecosystem. However, TPL is a double-edged sword. On the one hand, it can ease the development of mobile apps. On the other hand, it also brings security risks such as privacy leaks or increased attack surfaces (e.g., by introducing over-privileged permissions) to mobile apps. Although there are already many studies for characterizing third-party libraries, including automated detection, security and privacy analysis of TPLs, TPL attributes analysis, etc., what strikes us odd is that there is no systematic study to summarize those studies’ endeavors. To this end, we conduct the first systematic literature review on Android TPL-related research. Following a well-defined systematic literature review protocol, we collected 74 primary research papers closely related to Android third-party library from 2012 to 2020. After carefully examining these studies, we designed a taxonomy of TPL-related research studies and conducted a systematic study to summarize current solutions, limitations, challenges and possible implications of new research directions related to third-party library analysis. We hope that these contributions can give readers a clear overview of existing TPL-related studies and inspire them to go beyond the current status quo by advancing the discipline with innovative approaches. Xian Zhan, Tianming Liu 0002, Lingling Fan 0003, Li Li 0029, Sen Chen 0001, Xiapu Luo, Yang Liu 0003 |
IEEE Trans. Software Eng. | 7 |
| 2022 | A Systematic Assessment on Android Third-Party Library Detection ToolsabstractThird-party libraries (TPLs) have become a significant part of the Android ecosystem. Developers can employ various TPLs to facilitate their app development. Unfortunately, the popularity of TPLs also brings new security issues. For example, TPLs may carry malicious or vulnerable code, which can infect popular apps to pose threats to mobile users. Furthermore, TPL detection is essential for downstream tasks, such as vulnerabilities and malware detection. Thus, various tools have been developed to identify TPLs. However, no existing work has studied these TPL detection tools in detail, and different tools focus on different applications and techniques with performance differences. A comprehensive understanding of these tools will help us make better use of them. To this end, we conduct a comprehensive empirical study to fill the gap by evaluating and comparing all publicly available TPL detection tools based on six criteria: accuracy of TPL construction, effectiveness, efficiency, accuracy of version identification, resiliency to code obfuscation, and ease of use. Besides, we enhance these open-source tools by fixing their limitations, to improve their detection ability. Finally, we build an extensible framework that integrates all existing available TPL detection tools, providing an online service for the research community. We release the evaluation dataset and enhanced tools. According to our study, we also present the essential findings and discuss promising implications to the community; e.g., 1) Most existing TPL detection techniques more or less depend on package structure to construct in-app TPL candidates. However, using package structure as the module decoupling feature is error-prone. We hence suggest future researchers using the class dependency to substitute package structure. 2) Extracted features include richer semantic information (e.g., class dependencies) can achieve better resiliency to code obfuscation. 3) Existing tools usually have a low recall; that is because previous tools ignore some features of Android apps and TPLs, such as the compilation mechanism, the new format of TPLs, TPL dependency. Most existing tools cannot effectively find partial import TPLs, obfuscated TPLs, which directly limit their capability. 4) Existing tools are complementary to each other; we can build a better tool via combining the advantages of each tool. We believe our work provides a clear picture of existing TPL detection techniques and also gives a road-map for future research. Xian Zhan, Tianming Liu 0002, Yepang Liu 0001, Yang Liu 0003, Li Li 0029, Haoyu Wang 0001, Xiapu Luo |
IEEE Trans. Software Eng. | 4 |
| 2022 | Machine Learning Testing: Survey, Landscapes and HorizonsabstractThis paper provides a comprehensive survey of techniques for testing machine learning systems; Machine Learning Testing (ML testing) research. It covers 144 papers on testing properties (e.g., correctness, robustness, and fairness), testing components (e.g., the data, learning program, and framework), testing workflow (e.g., test generation and test evaluation), and application scenarios (e.g., autonomous driving, machine translation). The paper also analyses trends concerning datasets, research trends, and research focus, concluding with research challenges and promising research directions in ML testing. Jie Zhang 0050, Mark Harman, Lei Ma 0003, Yang Liu 0003 |
IEEE Trans. Software Eng. | 4 |
| 2021 | EfficientDeRain: Learning Pixel-wise Dilation Filtering for High-Efficiency Single-Image DerainingabstractSingle-image deraining is rather challenging due to the unknown rain model. Existing methods often make specific assumptions of the rain model, which can hardly cover many diverse circumstances in the real world, compelling them to employ complex optimization or progressive refinement. This, however, significantly affects these methods' efficiency and effectiveness for many efficiency-critical applications. To fill this gap, in this paper, we regard the single-image deraining as a general image-enhancing problem and originally propose a model-free deraining method, i.e., EfficientDeRain, which is able to process a rainy image within 10 ms (i.e., around 6 ms on average), over 80 times faster than the state-of-the-art method (i.e., RCDNet), while achieving similar de-rain effects. We first propose novel pixel-wise dilation filtering. In particular, a rainy image is filtered with the pixel-wise kernels estimated from a kernel prediction network, by which suitable multi-scale kernels for each pixel can be efficiently predicted. Then, to eliminate the gap between synthetic and real data, we further propose an effective data augmentation method (i.e., RainMix) that helps to train the network for handling real rainy images. We perform a comprehensive evaluation on both synthetic and real-world rainy datasets to demonstrate the effectiveness and efficiency of our method. We release the model and code in https://github.com/tsingqguo/efficientderain.git. Qing Guo 0005, Jingyang Sun, Felix Juefei-Xu, Lei Ma 0003, Xiaofei Xie, Wei Feng 0005, Yang Liu 0003, Jianjun Zhao 0001 |
AAAI | 7 |
| 2021 | Decision-Guided Weighted Automata Extraction from Recurrent Neural NetworksabstractRecurrent Neural Networks (RNNs) have demonstrated their effectiveness in learning and processing sequential data (e.g., speech and natural language). However, due to the black-box nature of neural networks, understanding the decision logic of RNNs is quite challenging. Some recent progress has been made to approximate the behavior of an RNN by weighted automata. They provide better interpretability, but still suffer from poor scalability. In this paper, we propose a novel approach to extracting weighted automata with the guidance of a target RNN's decision and context information. In particular, we identify the patterns of RNN's step-wise predictive decisions to instruct the formation of automata states. Further, we propose a state composition method to enhance the context-awareness of the extracted model. Our in-depth evaluations on typical RNN tasks, including language model and classification, demonstrate the effectiveness and advantage of our method over the state-of-the-arts. The evaluation results show that our method can achieve accurate approximation of an RNN even on large-scale tasks. Xiyue Zhang 0001, Xiaoning Du 0001, Xiaofei Xie, Lei Ma 0003, Yang Liu 0003, Meng Sun 0002 |
AAAI | 5 |
| 2021 | Stealing Deep Reinforcement Learning Models for Fun and ProfitabstractThis paper presents the first model extraction attack against Deep Reinforcement Learning (DRL), which enables an external adversary to precisely recover a black-box DRL model only from its interaction with the environment. Model extraction attacks against supervised Deep Learning models have been widely studied. However, those techniques cannot be applied to the reinforcement learning scenario due to DRL models' high complexity, stochasticity and limited observable information. We propose a novel methodology to overcome the above challenges. The key insight of our approach is that the process of DRL model extraction is equivalent to imitation learning, a well-established solution to learn sequential decision-making policies. Based on this observation, our methodology first builds a classifier to reveal the training algorithm family of the targeted black-box DRL model only based on its predicted actions, and then leverages state-of-the-art imitation learning techniques to replicate the model from the identified algorithm family. Experimental results indicate that our methodology can effectively recover the DRL models with high fidelity and accuracy. We also demonstrate two use cases to show that our model extraction attack can (1) significantly improve the success rate of adversarial attacks, and (2) steal DRL models stealthily even they are protected by DNN watermarks. These pose a severe threat to the intellectual property and privacy protection of DRL applications. Kangjie Chen, Shangwei Guo, Tianwei Zhang 0004, Xiaofei Xie, Yang Liu 0003 |
AsiaCCS | 5 |
| 2021 | SoFi: Reflection-Augmented Fuzzing for JavaScript EnginesabstractJavaScript engines have been shown prone to security vulnerabilities, which can lead to serious consequences due to their popularity. Fuzzing is an effective testing technique to discover vulnerabilities. The main challenge of fuzzing JavaScript engines is to generate syntactically and semantically valid inputs such that deep functionalities can be explored. However, due to the dynamic nature of JavaScript and the special features of different engines, it is quite challenging to generate semantically meaningful test inputs. Xiaofei Xie, Yuekang Li, Feng Li 0045, Yang Liu 0003, Wenchang Shi, Wei Huo 0005 |
CCS | 7 |
| 2021 | Auto-Exposure Fusion for Single-Image Shadow RemovalabstractShadow removal is still a challenging task due to its inherent background-dependent1and spatial-variant properties, leading to unknown and diverse shadow patterns. Even powerful deep neural networks could hardly recover traceless shadow-removed background. This paper proposes a new solution for this task by formulating it as an exposure fusion problem to address the challenges. Intuitively, we first estimate multiple over-exposure images w.r.t. the input image to let the shadow regions in these images have the same color with shadow-free areas in the input image. Then, we fuse the original input with the over-exposure images to generate the final shadow-free counterpart. Nevertheless, the spatial-variant property of the shadow requires the fusion to be sufficiently ‘smart’, that is, it should automatically select proper over-exposure pixels from different images to make the final output natural. To address this challenge, we propose the shadow-aware FusionNet that takes the shadow image as input to generate fusion weight maps across all the over-exposure images. Moreover, we propose the boundary-aware RefineNet to eliminate the remaining shadow trace further. We conduct extensive experiments on the ISTD, ISTD+, and SRD datasets to validate our method’s effectiveness and show better performance in shadow regions and comparable performance in non-shadow regions over the state-of-the-art methods. We release the code in https://github.com/tsingqguo/exposure-fusion-shadow-removal. Lan Fu, Changqing Zhou, Qing Guo 0005, Felix Juefei-Xu, Hongkai Yu, Wei Feng 0005, Yang Liu 0003, Song Wang 0002 |
CVPR | 7 |
| 2021 | Privacy-Preserving Collaborative Learning With Automatic Transformation SearchabstractCollaborative learning has gained great popularity due to its benefit of data privacy protection: participants can jointly train a Deep Learning model without sharing their training sets. However, recent works discovered that an adversary can fully recover the sensitive training samples from the shared gradients. Such reconstruction attacks pose severe threats to collaborative learning. Hence, effective mitigation solutions are urgently desired.In this paper, we propose to leverage data augmentation to defeat reconstruction attacks: by preprocessing sensitive images with carefully-selected transformation policies, it becomes infeasible for the adversary to extract any useful information from the corresponding gradients. We design a novel search method to automatically discover qualified policies. We adopt two new metrics to quantify the impacts of transformations on data privacy and model usability, which can significantly accelerate the search speed. Comprehensive evaluations demonstrate that the policies discovered by our method can defeat existing reconstruction attacks in collaborative learning, with high efficiency and negligible impact on the model performance. Wei Gao 0064, Shangwei Guo, Tianwei Zhang 0004, Han Qiu 0001, Yonggang Wen 0001, Yang Liu 0003 |
CVPR | 6 |
| 2021 | Learning to Adversarially Blur Visual Object TrackingabstractMotion blur caused by the moving of the object or camera during the exposure can be a key challenge for visual object tracking, affecting tracking accuracy significantly. In this work, we explore the robustness of visual object trackers against motion blur from a new angle, i.e., adversarial blur attack (ABA). Our main objective is to online transfer input frames to their natural motion-blurred counterparts while misleading the state-of-the-art trackers during the tracking process. To this end, we first design the motion blur synthesizing method for visual tracking based on the generation principle of motion blur, considering the motion information and the light accumulation process. With this synthetic method, we propose optimization-based ABA (OP-ABA) by iteratively optimizing an adversarial objective function against the tracking w.r.t. the motion and light accumulation parameters. The OP-ABA is able to produce natural adversarial examples but the iteration can cause heavy time cost, making it unsuitable for attacking real-time trackers. To alleviate this issue, we further propose one-step ABA (OS-ABA) where we design and train a joint adversarial motion and accumulation predictive network (JAMANet) with the guidance of OP-ABA, which is able to efficiently estimate the adversarial motion and accumulation parameters in a one-step way. The experiments on four popular datasets (e.g., OTB100, VOT2018, UAV123, and LaSOT) demonstrate that our methods are able to cause significant accuracy drops on four state-of-the-art trackers with high transferability. Please find the source code at https://github.com/tsingqguo/ABA Qing Guo 0005, Ziyi Cheng, Felix Juefei-Xu, Lei Ma 0003, Xiaofei Xie, Yang Liu 0003, Jianjun Zhao 0001 |
ICCV | 6 |
| 2021 | Retrieval-Augmented Generation for Code Summarization via Hybrid GNN
Shangqing Liu, Xiaofei Xie, Jing Kai Siow, Yang Liu 0003 |
ICLR | 5 |
| 2021 | RNNRepair: Automatic RNN Repair via Model-based AnalysisabstractDeep neural networks are vulnerable to adversarial attacks. Due to their black-box nature, it is rather challenging to interpret and properly repair these incorrect behaviors. This paper focuses on interpreting and repairing the incorrect behaviors of Recurrent Neural Networks (RNNs). We propose a lightweight model-based approach (RNNRepair) to help understand and repair incorrect behaviors of an RNN. Specifically, we build an influence model to characterize the stateful and statistical behaviors of an RNN over all the training data and to perform the influence analysis for the errors. Compared with the existing techniques on influence function, our method can efficiently estimate the influence of existing or newly added training samples for a given prediction at both sample level and segmentation level. Our empirical evaluation shows that the proposed influence model is able to extract accurate and understandable features. Based on the influence model, our proposed technique could effectively infer the influential instances from not only an entire testing sequence but also a segment within that sequence. Moreover, with the sample-level and segment-level influence relations, RNNRepair could further remediate two types of incorrect predictions at the sample level and segment level. Xiaofei Xie, Wenbo Guo 0002, Lei Ma 0003, Wei Le, Jian Wang 0067, Lingjun Zhou, Yang Liu 0003, Xinyu Xing 0001 |
ICML | 7 |
| 2021 | Route Coverage Testing for Autonomous Vehicles via Map ModelingabstractAutonomous vehicles (AVs) play an important role in transforming our transportation systems and relieving traffic congestion. To guarantee their safety, AVs must be sufficiently tested before they are deployed to public roads. Existing testing often focuses on AVs’ collision avoidance on a given route. There is little work on the systematic testing for AVs’ route planning and tracking on a map. In this paper, we propose CROUTE, a novel testing method based on a new AV testing criterion called route coverage. First, the map is modeled as a labeled Petri net, where roads, junctions, and traffic signs are modeled as places, transitions, and labels, respectively. Second, based on the Petri net, we define junctions’ topology features and route features for junction classification. The topology feature describes the topology of roads forming the junction, and the route feature identifies the actions that a vehicle can take to follow a route. They can characterize route types on a map. Hence, route coverage measures how many route types are covered. We then propose a systematic method that aims to cover all route types for a well-designed AV system with a small number of test cases. We implement and evaluate CROUTE on Baidu Apollo running with the LGSVL simulator. We carry out testing on the map from a section of San Francisco and find six different types of issues in Apollo. The experiment results show the validity of route coverage and the efficiency of CROUTE. Yun Tang 0003, Yuan Zhou 0005, Fenghua Wu, Yang Liu 0003, Jun Sun 0001, Wuling Huang |
ICRA | 4 |
| 2021 | ATVHUNTER: Reliable Version Detection of Third-Party Libraries for Vulnerability Identification in Android ApplicationsabstractThird-party libraries (TPLs) as essential parts in the mobile ecosystem have become one of the most significant contributors to the huge success of Android, which facilitate the fast development of Android applications. Detecting TPLs in Android apps is also important for downstream tasks, such as malware and repackaged apps identification. To identify in-app TPLs, we need to solve several challenges, such as TPL dependency, code obfuscation, precise version representation. Unfortunately, existing TPL detection tools have been proved that they have not solved these challenges very well, let alone specify the exact TPL versions. To this end, we propose a system, named ATVHunter, which can pinpoint the precise vulnerable in-app TPL versions and provide detailed information about the vulnerabilities and TPLs. We propose a two-phase detection approach to identify specific TPL versions. Specifically, we extract the Control Flow Graphs as the coarse-grained feature to match potential TPLs in the pre-defined TPL database, and then extract opcode in each basic block of CFG as the fine-grained feature to identify the exact TPL versions. We build a comprehensive TPL database (189,545 unique TPLs with 3,006,676 versions) as the reference database. Meanwhile, to identify the vulnerable in-app TPL versions, we also construct a comprehensive and known vulnerable TPL database containing 1,180 CVEs and 224 security bugs. Experimental results show AtVHunter outperforms state-of-the-art TPL detection tools, achieving 90.55% precision and 88.79% recall with high efficiency, and is also resilient to widely-used obfuscation techniques and scalable for large-scale TPL detection. Furthermore, to investigate the ecosystem of the vulnerable TPLs used by apps, we exploit newtool to conduct a large-scale analysis on 104,446 apps and find that 9,050 apps include vulnerable TPL versions with 53,337 vulnerabilities and 7,480 security bugs, most of which are with high risks and are not recognized by app developers. Xian Zhan, Lingling Fan 0003, Sen Chen 0001, Tianming Liu 0002, Xiapu Luo, Yang Liu 0003 |
ICSE | 7 |
| 2021 | Automatic Web Testing Using Curiosity-Driven Reinforcement LearningabstractWeb testing has long been recognized as a notoriously difficult task. Even nowadays, web testing still mainly relies on manual efforts in many cases while automated web testing is still far from achieving human-level performance. Key challenges include dynamic content update and deep bugs hiding under complicated user interactions and specific input values, which can only be triggered by certain action sequences in the huge space of all possible sequences. In this paper, we propose WebExplor, an automatic end-to-end web testing framework, to achieve an adaptive exploration of web applications. WebExplor adopts a curiosity-driven reinforcement learning to generate high-quality action sequences (test cases) with temporal logical relations. Besides, WebExplor incrementally builds an automaton during the online testing process, which acts as the high-level guidance to further improve the testing efficiency. We have conducted comprehensive evaluations on six real-world projects, a commercial SaaS web application, and performed an in-the-wild study of the top 50 web applications in the world. The results demonstrate that in most cases WebExplor can achieve significantly higher failure detection rate, code coverage and efficiency than existing state-of-the-art web testing techniques. WebExplor also detected 12 previously unknown failures in the commercial web application, which have been confirmed and fixed by the developers. Furthermore, our in-the-wild study further uncovered 3,466 exceptions and errors. Yan Zheng 0002, Yi Liu 0069, Xiaofei Xie, Yepang Liu 0001, Lei Ma 0003, Jianye Hao, Yang Liu 0003 |
ICSE | 7 |
| 2021 | Fine-tuning Is Not Enough: A Simple yet Effective Watermark Removal Attack for DNN ModelsabstractWatermarking has become the tendency in protecting the intellectual property of DNN models. Recent works, from the adversary's perspective, attempted to subvert watermarking mechanisms by designing watermark removal attacks. However, these attacks mainly adopted sophisticated fine-tuning techniques, which have certain fatal drawbacks or unrealistic assumptions. In this paper, we propose a novel watermark removal attack from a different perspective. Instead of just fine-tuning the watermarked models, we design a simple yet powerful transformation algorithm by combining imperceptible pattern embedding and spatial-level transformations, which can effectively and blindly destroy the memorization of watermarked models to the watermark samples. We also introduce a lightweight fine-tuning strategy to preserve the model performance. Our solution requires much less resource or knowledge about the watermarking scheme than prior works. Extensive experimental results indicate that our attack can bypass state-of-the-art watermarking solutions with very high success rates. Based on our attack, we propose watermark augmentation techniques to enhance the robustness of existing watermarks. Shangwei Guo, Tianwei Zhang 0004, Han Qiu 0001, Yi Zeng 0005, Tao Xiang 0001, Yang Liu 0003 |
IJCAI | 6 |
| 2021 | AVA: Adversarial Vignetting Attack against Visual RecognitionabstractVignetting is an inherent imaging phenomenon within almost all optical systems, showing as a radial intensity darkening toward the corners of an image. Since it is a common effect for photography and usually appears as a slight intensity variation, people usually regard it as a part of a photo and would not even want to post-process it. Due to this natural advantage, in this work, we study the vignetting from a new viewpoint, i.e., adversarial vignetting attack (AVA), which aims to embed intentionally misleading information into the vignetting and produce a natural adversarial example without noise patterns. This example can fool the state-of-the-art deep convolutional neural networks (CNNs) but is imperceptible to human. To this end, we first propose the radial-isotropic adversarial vignetting attack (RI-AVA) based on the physical model of vignetting, where the physical parameters (e.g., illumination factor and focal length) are tuned through the guidance of target CNN models. To achieve higher transferability across different CNNs, we further propose radial-anisotropic adversarial vignetting attack (RA-AVA) by allowing the effective regions of vignetting to be radial-anisotropic and shape-free. Moreover, we propose the geometry-aware level-set optimization method to solve the adversarial vignetting regions and physical parameters jointly. We validate the proposed methods on three popular datasets, i.e., DEV, CIFAR10, and Tiny ImageNet, by attacking four CNNs, e.g., ResNet50, EfficientNet-B0, DenseNet121, and MobileNet-V2, demonstrating the advantages of our methods over baseline methods on both transferability and image quality. Binyu Tian, Felix Juefei-Xu, Qing Guo 0005, Xiaofei Xie, Xiaohong Li 0001, Yang Liu 0003 |
IJCAI | 6 |
| 2021 | Leveraging ASR N-Best in Deep Entity Retrieval
Haoyu Wang 0001, Majid Laali, Kevin Durda, Jeff King, William Campbell, Yang Liu 0003 |
Interspeech | 7 |
| 2021 | Peeking into the Gray Area of Mobile World: An Empirical Study of Unlabeled Android AppsabstractFor the real-world dataset collected by our industrial partner, Pwnzen Infotech Inc., one of the leading industrial security companies, there are a large number of unlabeled Android applications (called unlabeled apps in this paper) that are unlikely to belong to known Android malware families nor ordinary benign apps according to the industrial black-list (i.e., signatures) and white-list (i.e., certificates). However, such apps have rarely been studied previously, but are important to peek into the gray area of mobile world. It is a time-consuming task for software analysts to understand the negative characteristics of these samples, which would lead to potential security or privacy threats for app users, significantly negative impacts on mobile system performance, and bad user experience, etc. To investigate the characteristics of these industrial unlabeled apps in a large-scale in practice, and provide insights to industrial software analysts as well as research communities, we collect a large-scale dataset of unlabeled apps (i.e., 22,886 in total) from our industrial partners. Given the common industrial perception of software analysts that a high percentage of these unlabeled apps could have some similar behaviors, we leverage the popular community-detection techniques based on widely-used app features in mal ware detection to cluster these unlabeled apps. After that, we investigate the common behaviors for different clusters with substantial human efforts and also conduct cross-validation across co-authors to check the results. Our manual analysis unveils the characteristics of these unlabeled apps by sampling data from different clusters, and discovers 11 categories, some of which have never been discovered by previous grayware research. Besides, from our exploration, we find that the community-based techniques are not effective enough in clustering unlabeled apps, so that manual analysis is encouraged. Manual analysis is an important first step towards studying unlabeled apps and understanding their characteristics. Finally, we highlight the lessons learned through real case studies, comparison study with existing malware/grayware research, in-depth discussion with industrial partners, and feedback from industrial partners. Sen Chen 0001, Lingling Fan 0003, Cuiyun Gao 0001, Fu Song, Yang Liu 0003 |
ISSRE | 5 |
| 2021 | Vall-nut: Principled Anti-Grey box - FuzzingabstractGreybox fuzzing is a widely used technique for software testing that has been adopted by practitioners and researchers to disclose a great number of vulnerabilities in various software. However, adversaries also weaponize greybox fuzzing to mine vulnerabilities for malicious intentions. This poses considerable threats to software systems. To counteract the misuse of greybox fuzzing, we propose VALL-NUT, a novel approach to harden software with properties to combat greybox fuzzing. We dissect the major strategies that facilitate the success of greybox fuzzing, and accordingly propose three types of neutralizing schemesseed queue explosion, seed attenuation, and feedback contamination. We evaluate Vall-nut against the mainstream greybox fuzzers on multiple real-world benchmark programs. The results show that Vall-nut can reduce an average of 34 % code coverage and 76% detected crashes in 24-hour tests. Moreover, we conduct comparisons with two recent studies which show Vall-nut can achieve a superior deduction of detected crashes. Yuekang Li, Guozhu Meng, Jun Xu 0024, Cen Zhang, Hongxu Chen 0001, Xiaofei Xie, Haijun Wang 0002, Yang Liu 0003 |
ISSRE | 8 |
| 2021 | Collision Avoidance Testing for Autonomous Driving Systems on Complete MapsabstractCollision avoidance is one of the crucial functions of autonomous driving systems (ADSs) to guarantee the safety of autonomous vehicles (AVs). It requires extensive testing before an AV is deployed to public roads. Most of the current ADS testing methods generate test cases either from real traffic data or manually designed for some specific scenarios. There is little work on systematic methods to generate test cases from a complete map where an AV operates. Systematic testing on such a map is challenging due to the enormous scenarios. In this paper, we propose a collision-avoidance testing method for ADSs running on a map, which aims to reduce the scenario space while maintaining scenario diversity. The method consists of test case classification and test case generation. First, we build the topology structure of a map, based on which we classify possible scenarios into different classes. Second, we divide test cases into different classes using the topology-based scenario classification and fuzzy number-based motion evaluation. Third, we implement a bisection method to generate test cases that can efficiently expose ADSs' failures. We evaluate our method on one of the state-of-the-art ADSs, Baidu Apollo. The experiment results show that our method discovers Apollo's issues effectively while reducing the number of generated test cases by 77.36%, compared with the random method. Yun Tang 0003, Yuan Zhou 0005, Yang Liu 0003, Jun Sun 0001 |
IV | 3 |
| 2021 | Automatic HMI Structure Exploration Via Curiosity-Based Reinforcement LearningabstractDiscovering the underlying structure of HMI software efficiently and sufficiently for the purpose of testing without any prior knowledge on the software logic remains a difficult problem. The key challenge lies in the complexity of the HMI software and the high variance in the coverage of current methods. In this paper, we introduce the PathFinder, an effective and automatic HMI software exploration framework. PathFinder adopts a curiosity-based reinforcement learning framework to choose actions that lead to the discovery of more unknown states. Additionally, PathFinder progressively builds a navigation model during the exploration to further improve state coverage. We have conducted experiments on both simulations and real-world HMI software testing environment, which comprise a full tool chain of automobile dashboard instrument cluster. The exploration coverage outperforms manual and fuzzing methods which are the current industrial standards. Yushi Cao, Yan Zheng 0002, Shangwei Lin 0001, Yang Liu 0003, Yon Shin Teo, Yuxuan Toh, Vinay Vishnumurthy Adiga |
ASE | 4 |
| 2021 | A First Look at the Effect of Deep Learning in Coverage-guided FuzzingabstractFuzzing has been a widely-used technique for discovering software vulnerabilities. Many existing fuzzers leverage coverage-feedback to evolve seeds to maximize (optimize) program branch coverage. Recently, some techniques propose to train deep learning models to predict the branch coverage of an arbitrary input. Those techniques have proved their success in improving coverage and discovering bugs under different experimental settings. However, deep learning models, usually as a black magic box, are notoriously lack of explanation. Moreover, their performance can be sensitive to the collected runtime coverage information for training, indicating potentially unstable performance. To this end, in this work we conduct a systematic and extensive empirical study on 4 types of deep learning models across 6 projects to reproduce the actual performance of deep learning fuzzers, analyze the advantages and disadvantages of deep learning in the process of fuzzing applications, and explore the future direction of the combination of the two. Our empirical results reveal that the deep learning models can only be effective in very limited scenarios, which is largely restrained by training data imbalance, dependant labels, model over-generalization, and the insufficient expressiveness of the state-of-the-art models. Consequently, the estimated gradients by the models to cover a branch can be less helpful in many scenarios. Yun Lin 0001, Xiaofei Xie, Yuekang Li, Xiaohong Li 0001, Weimin Ge, Yang Liu 0003, Jin Song Dong 0001 |
ASE | 7 |
| 2021 | FirmGuide: Boosting the Capability of Rehosting Embedded Linux Kernels through Model-Guided Kernel ExecutionabstractLinux kernel is widely used in embedded systems. To understand practical threats to the Linux kernel, we need to perform dynamic analysis with a full-system emulator, e.g., QEMU. However, due to hardware fragmentation, e.g., various types of peripherals, most embedded systems are not currently supported by QEMU. Though some progress has been made on rehosting firmware, it mainly focuses on user space programs or simple real-time operating systems.The goal of this work is to boost the capability of rehosting the embedded Linux kernels in QEMU. By doing so, dynamic analysis systems can be firstly applied on embedded Linux kernels by leveraging off-the-shelf tools upon QEMU. Accordingly, we proposed a new technique called model-guided kernel execution. It combines the peripheral abstractions in the Linux kernel and kernel-peripheral interactions to semi-automatically generate peripheral models that are then used to synthesize new QEMU virtual machines to start the dynamic analysis.We have implemented a prototype called FirmGuide. It generates 9 peripheral models with full functionality and 64 with minimum functionality covering 26 SoCs. Our evaluation with 6,188 firmware images shows that it can successfully rehost more than 95% of Linux kernels in 2 architectures and 22 versions. None of them can be rehosted in the vanilla QEMU. The result of the LTP benchmark shows the reliability and robustness of the rehosted Linux kernels. We further conduct two security applications, i.e., vulnerability analysis and fuzzing, on the rehosted Linux kernels to demonstrate the usage scenarios. Qiang Liu 0034, Cen Zhang, Lin Ma 0009, Muhui Jiang, Yajin Zhou, Lei Wu 0012, Wenbo Shen, Xiapu Luo, Yang Liu 0003, Kui Ren 0001 |
ASE | 9 |
| 2021 | Systematic Testing of Autonomous Driving Systems Using Map Topology-Based Scenario ClassificationabstractAutonomous Driving Systems (ADSs), which replace humans to drive vehicles, are complex software systems deployed in autonomous vehicles (AVs). Since the execution of ADSs highly relies on maps, it is essential to perform global map-based testing for ADSs to guarantee their correctness and AVs’ safety in different situations. Existing methods focus more on specific scenarios rather than global testing throughout the map. Testing on a global map is challenging since the complex lane connections in a map can generate enormous scenarios. In this work, we propose ATLAS, an approach to ADSs’ collision avoidance testing using map topology-based scenario classification. The core insight of ATLAS is to generate diverse testing scenarios by classifying junction lanes according to their topology-based interaction patterns. First, ATLAS divides the junction lanes into different classes such that an ADS can execute similar collision avoidance maneuvers on the lanes in the same class. Second, for each class, ATLAS selects one junction lane to construct the testing scenario and generate test cases using a genetic algorithm. Finally, we implement and evaluate ATLAS on Baidu Apollo with the LGSVL simulator on the San Francisco map. Results show that ATLAS exposes nine types of real issues in Apollo 6.0 and reduces the number of junction lanes for testing by 98%. Yun Tang 0003, Yuan Zhou 0005, Tianwei Zhang 0004, Fenghua Wu, Yang Liu 0003 |
ASE | 5 |
| 2021 | BIFF: Practical Binary Fuzzing Framework for Programs of IoT and Mobile DevicesabstractInternet-of-things (IoT) or mobile devices are omnipresent in our daily life; the security issues inside them are especially crucial. Greybox fuzzing has been shown effective in detecting vulnerabilities. However, applications in IoT or mobile devices are usually proprietary to specific vendors, fuzzers are required to support binary-only targets. Moreover, since these devices are of heterogeneous architecture, assigned with limited resources, and many testing targets are server-like programs, applying existing fuzzing techniques faces great challenges.This paper proposes BIFF, a general-purpose fuzzer that aims to stress these issues. It supports binary-only targets, is general (supports multiple CPU architectures including Intel, ARM, MIPS, and PowerPC), fast (has the lowest runtime overhead compared to existing fuzzers), and flexible (uses a new fuzzing workflow that can fuzz any piece of code inside the target binary). Experiments demonstrate that BIFF has the best performance compared with state-of-the-art binary fuzzers and can fuzz the server-like programs which cannot be fuzzed by the existing fuzzers. Using BIFF, we’ve found 24 unknown vulnerabilities (including memory corruptions, infinite loops, and infinite recursions) in industrial products. Cen Zhang, Yuekang Li, Hongxu Chen 0001, Xiaoxing Luo, Miaohua Li, Anh Quynh Nguyen, Yang Liu 0003 |
ASE | 7 |
| 2021 | AdvFilter: Predictive Perturbation-aware Filtering against Adversarial Attack via Multi-domain LearningabstractHigh-level representation-guided pixel denoising and adversarial training are independent solutions to enhance the robustness of CNNs against adversarial attacks by pre-processing input data and re-training models, respectively. Most recently, adversarial training techniques have been widely studied and improved while the pixel denoising-based method is getting less attractive. However, it is still questionable whether there exists a more advanced pixel denoising-based method and whether the combination of the two solutions benefits each other. To this end, we first comprehensively investigate two kinds of pixel denoising methods for adversarial robustness enhancement (i.e., existing additive-based and unexplored filtering-based methods) under the loss functions of image-level and semantic-level, respectively, showing that pixel-wise filtering can obtain much higher image quality (e.g., higher PSNR) as well as higher robustness (e.g., higher accuracy on adversarial examples) than existing pixel-wise additive-based method. However, we also observe that the robustness results of the filtering-based method rely on the perturbation amplitude of adversarial examples used for training. To address this problem, we propose predictive perturbation-aware & pixel-wise filtering, where dual-perturbation filtering and an uncertainty-aware fusion module are designed and employed to automatically perceive the perturbation amplitude during the training and testing process. The method is termed as AdvFilter. Moreover, we combine adversarial pixel denoising methods with three adversarial training-based methods, hinting that considering data and models jointly is able to achieve more robust CNNs. The experiments conduct on NeurIPS-2017DEV, SVHN and CIFAR10 datasets and show advantages over enhancing CNNs' robustness, high generalization to different models and noise levels. Yihao Huang 0001, Qing Guo 0005, Felix Juefei-Xu, Lei Ma 0003, Weikai Miao, Yang Liu 0003, Geguang Pu |
ACM Multimedia | 6 |
| 2021 | JPGNet: Joint Predictive Filtering and Generative Network for Image InpaintingabstractImage inpainting aims to restore the missing regions of corrupted images and make the recovery result identical to the originally complete image, which is different from the common generative task emphasizing the naturalness or realism of generated images. Nevertheless, existing works usually regard it as a pure generation problem and employ cutting-edge deep generative techniques to address it. The generative networks can fill the main missing parts with realistic contents but usually distort the local structures or introduce obvious artifacts. In this paper, for the first time, we formulate image inpainting as a mix of two problems, i.e., predictive filtering and deep generation. Predictive filtering is good at preserving local structures and removing artifacts but falls short to complete the large missing regions. The deep generative network can fill the numerous missing pixels based on the understanding of the whole scene but hardly restores the details identical to the original ones. To make use of their respective advantages, we propose the joint predictive filtering and generative network (JPGNet) that contains three branches: predictive filtering & uncertainty network (PFUNet), deep generative network, and uncertainty-aware fusion network (UAFNet). The PFUNet can adaptively predict pixel-wise kernels for filtering-based inpainting according to the input image and output an uncertainty map. This map indicates the pixels should be processed by filtering or generative networks, which is further fed to the UAFNet for a smart combination between filtering and generative results. Note that, our method as a novel framework for the image inpainting problem can benefit any existing generation-based methods. We validate our method on three public datasets, i.e., Dunhuang, Places2, and CelebA, and demonstrate that our method can enhance three state-of-the-art generative methods (i.e., StructFlow, EdgeConnect, and RFRNet) significantly with slightly extra time costs. We have released the code at https://github.com/tsingqguo/jpgnet. Qing Guo 0005, Felix Juefei-Xu, Hongkai Yu, Yang Liu 0003, Song Wang 0002 |
ACM Multimedia | 5 |
| 2021 | FakeTagger: Robust Safeguards against DeepFake Dissemination via Provenance TrackingabstractIn recent years, DeepFake is becoming a common threat to our society, due to the remarkable progress of generative adversarial networks (GAN) in image synthesis. Unfortunately, existing studies that propose various approaches, in fighting against DeepFake and determining if the facial image is real or fake, is still at an early stage. Obviously, the current DeepFake detection method struggles to catch the rapid progress of GANs, especially in the adversarial scenarios where attackers can evade the detection intentionally, such as adding perturbations to fool the DNN-based detectors. While passive detection simply tells whether the image is fake or real, DeepFake provenance, on the other hand, provides clues for tracking the sources in DeepFake forensics. Thus, the tracked fake images could be blocked immediately by administrators and avoid further spread in social networks. Run Wang 0001, Felix Juefei-Xu, Meng Luo 0002, Yang Liu 0003, Lina Wang 0001 |
ACM Multimedia | 4 |
| 2021 | An Investigation of Byzantine Threats in Multi-Robot SystemsabstractMulti-Robot Systems (MRSs) show significant advantages to deal with complex tasks efficiently. However, the system complexity inevitably enlarges the attack surface and adds difficulty in guaranteeing the security and safety of MRSs. In this paper, we present an in-depth investigation about the Byzantine threats in MRSs, where some robot is untrusted. We design a practical methodology to identify potential Byzantine risks in a given MRS workload built from the Robot Operating System (ROS). It consists of three novel steps (requirement specification using signal temporal logic, attack surface determination via data-flow analysis, attack identification using requirement-driven fuzzing) to thoroughly assess MRS workloads. We use this fuzzing method to inspect five typical MRS workloads from past works and the ROS platform, and identify three novel kinds of attacks that can be launched with five attack strategies. We conduct comprehensive experiments in the Gazebo simulator and a real-world MRS with three TurtlBot3 robots to validate these attacks, which can remarkably decrease the system’s performance, or even cause task failures. Gelei Deng, Yuan Zhou 0005, Yuan Xu 0033, Tianwei Zhang 0004, Yang Liu 0003 |
RAID | 5 |
| 2021 | Who is Real Bob? Adversarial Attacks on Speaker Recognition SystemsabstractSpeaker recognition (SR) is widely used in our daily life as a biometric authentication or identification mechanism. The popularity of SR brings in serious security concerns, as demonstrated by recent adversarial attacks. However, the impacts of such threats in the practical black-box setting are still open, since current attacks consider the white-box setting only.In this paper, we conduct the first comprehensive and systematic study of the adversarial attacks on SR systems (SRSs) to understand their security weakness in the practical black-box setting. For this purpose, we propose an adversarial attack, named FAKEBOB, to craft adversarial samples. Specifically, we formulate the adversarial sample generation as an optimization problem, incorporated with the confidence of adversarial samples and maximal distortion to balance between the strength and imperceptibility of adversarial voices. One key contribution is to propose a novel algorithm to estimate the score threshold, a feature in SRSs, and use it in the optimization problem to solve the optimization problem. We demonstrate that FAKEBOB achieves 99% targeted attack success rate on both open-source and commercial systems. We further demonstrate that FAKEBOB is also effective on both open-source and commercial systems when playing over the air in the physical world. Moreover, we have conducted a human study which reveals that it is hard for human to differentiate the speakers of the original and adversarial voices. Last but not least, we show that four promising defense methods for adversarial attack from the speech recognition domain become ineffective on SRSs against FAKEBOB, which calls for more effective defense methods. We highlight that our study peeks into the security implications of adversarial attacks on SRSs, and realistically fosters to improve the security robustness of SRSs. Guangke Chen, Sen Chen 0001, Lingling Fan 0003, Xiaoning Du 0001, Zhe Zhao 0007, Fu Song, Yang Liu 0003 |
SP | 7 |
| 2021 | APICraft: Fuzz Driver Generation for Closed-source SDK Libraries
Cen Zhang, Xingwei Lin, Yuekang Li, Yinxing Xue, Jundong Xie, Hongxu Chen 0001, Xinlei Ying, Jiashui Wang, Yang Liu 0003 |
USENIX Security Symposium | 9 |
| 2021 | VIVA: Binary Level Vulnerability Identification via Partial SignatureabstractBinary level code clone detection techniques have been used to identify 1-day vulnerabilities in software. It collects functions with known vulnerabilities and searches for similar functions in the target system. However, existing approaches are limited to detect the same vulnerabilities in different binaries. They can hardly find new recurring vulnerabilities, which share similar logic. Moreover, they only focus on improving the accuracy of binary function matching algorithms while overlooking the presence of security patches, which results in high false-positive rates and requires significant effort to verify the results.To this end, we propose VIVA, a binary level vulnerability and patch semantic summarization and matching tool for accurate recurring vulnerability detection. It uses novel binary program slicing techniques with the aid of pseudo-code trace refinement to generate partial vulnerability and patch signatures, which capture the semantics. It matches the signatures with pre-filtering to efficiently detect 1-day and recurring vulnerabilities. The experimental results show that VIVA outperforms other source code and binary matching tools with a precision of 100% for 1-day vulnerabilities and 87.6% for recurring vulnerabilities and good performance (28.58s per signature search in 4M functions). It detects 92 new vulnerabilities in different series and different versions of real-world projects, with 11 exist without fixing in the latest version. Yang Xiao 0011, Zhengzi Xu, Chendong Yu, Longquan Liu, Zimu Yuan, Yang Liu 0003, Aihua Piao, Wei Huo 0005 |
SANER | 8 |
| 2021 | Advanced evasion attacks and mitigations on practical ML-based phishing website classifiersabstractMachine learning (ML) based classifiers are vulnerable to evasion attacks, as shown by recent attacks. However, there is a lack of systematic study of evasion attacks on ML-based anti-phishing detection. In this study, we show that evasion attacks are not only effective on practical ML-based classifiers, but can also be efficiently launched without destructing the functionalities and appearance. For this purpose, we propose three mutation-based attacks, differing in the knowledge of the target classifier, addressing a key technical challenge: automatically crafting an adversarial sample from a known phishing website in a way that can mislead classifiers. To launch attacks in the white- and gray-box scenarios, we also propose a sample-based collision attack to gain the knowledge of the target classifier. We demonstrate the efficacy of our evasion attacks on the state-of-the-art, Google's phishing page filter, achieved 100% attack success rate in less than one second per website. Moreover, the transferability attack on BitDefender's industrial phishing page classifier, TrafficLight, achieved up to 81.25% attack success rate. We further propose a similarity-based method to mitigate such evasion attacks, Pelican, which compares the similarity of an unknown website with recently detected phishing websites. We demonstrate that Pelican can effectively detect evasion attacks, hence could be integrated into ML-based classifiers. We also highlight two strategies of classification rule selection to enhance the robustness of classifiers. Our findings contribute to design more robust phishing website classifiers in practice. Fu Song, Yusi Lei, Sen Chen 0001, Lingling Fan 0003, Yang Liu 0003 |
Int. J. Intell. Syst. | 5 |
| 2021 | CoreGen: Contextualized Code Representation Learning for Commit Message Generation
Lun Yiu Nie, Cuiyun Gao 0001, Zhicong Zhong, Wai Lam, Yang Liu 0003, Zenglin Xu |
Neurocomputing | 5 |
| 2021 | MTKeras: An Automated Metamorphic Testing PlatformabstractThis paper presents an automated, domain-independent, metamorphic testing platform called MTKeras. In this paper, we report on an investigation demonstrating the effectiveness and usability of MTKeras through five case studies in the four domains of image classification, sentiment analysis, search engines and database management systems. We also report on the effectiveness of combining metamorphic relation (input) patterns in individual metamorphic relations, enhancing the failure-finding abilities of the individual relations. The results of our experiments support combining patterns, and the use of MTKeras. The research reported in this paper shows the applicability of metamorphic relation patterns, and introduces a practical tool for the research community. Yelin Liu, Zhiquan Zhou 0001, Tsong Yueh Chen, Yang Liu 0003, Dave Towey |
Int. J. Softw. Eng. Knowl. Eng. | 4 |
| 2021 | An Isabelle/HOL Formalisation of the SPARC Instruction Set Architecture and the TSO Memory Model
David Sanán, Alwen Tiu, Yang Liu 0003, Koh Chuen Hoa, Jin Song Dong 0001 |
J. Autom. Reason. | 4 |
| 2021 | A permission-dependent type system for secure information flow analysisabstractWe introduce a novel type system for enforcing secure information flow in an imperative language. Our work is motivated by the problem of statically checking potential information leakage in Android applications. To this end, we design a lightweight type system featuring Android permission model, where the permissions are statically assigned to applications and are used to enforce access control in the applications. We take inspiration from a type system by Banerjee and Naumann to allow security types to be dependent on the permissions of the applications. A novel feature of our type system is a typing rule for conditional branching induced by permission testing, which introduces a merging operator on security types, allowing more precise security policies to be enforced. The soundness of our type system is proved with respect to non-interference. A type inference algorithm is also presented for the underlying security type system, by reducing the inference problem to a constraint solving problem in the lattice of security types. In addition, a new way to represent our security types as reduced ordered binary decision diagrams is proposed. Zhiwu Xu 0001, Hongxu Chen 0001, Alwen Tiu, Yang Liu 0003, Kunal Sareen |
J. Comput. Secur. | 4 |
| 2021 | Guardauto: A Decentralized Runtime Protection System for Autonomous DrivingabstractDue to the broad attack surface and the lack of runtime protection, potential safety and security threats hinder the real-life adoption of autonomous vehicles. Although efforts have been made to mitigate some specific attacks, there are few works on the protection of the autonomous driving system, i.e., the control software system performing such as perception, decision making, and motion tracking. This article presents a decentralized self-protection framework called Guardauto to protect the autonomous driving system against runtime threats. First, Guardauto proposes an isolation model to decouple the autonomous driving system and isolate its components with a set of partitions. Second, Guardauto provides self-protection mechanisms for each target component, which combines different methods to monitor the target execution and plan adaption actions accordingly. Third, Guardauto provides cooperation among local self-protection mechanisms to identify the root-cause component in the case of cascading failures affecting multiple components. A prototype has been implemented and evaluated on the open-source autonomous driving system Autoware. Results show that Guardauto could effectively mitigate runtime failures and attacks, and protect the control system with acceptable performance overhead. Yuan Zhou 0005, Bihuan Chen 0001, Rui Wang 0014, Yuebin Bai, Yang Liu 0003 |
IEEE Trans. Computers | 6 |
| 2021 | On Evaluating Fault Resilient Encoding Schemes in SoftwareabstractCryptographic implementations are often vulnerable against physical attacks, fault injection analysis being among the most popular techniques. On par with development of attacks, the area of countermeasures is advancing rapidly, utilizing both hardware- and software-based approaches. When it comes to software encoding countermeasures for fault protection and their evaluation, there are very few proposals so far, mostly focusing on single operations rather than cipher as a whole. In this paper we propose an evaluation framework that can be used for analyzing the effectivity of software encoding countermeasures against fault attacks. We first formalize the encoding schemes in software, helping us to define what properties are required when designing a fault protection. Based on these findings, we develop an evaluation metric that can be used universally to determine the robustness of a software encoding scheme against bit flip faults and instruction skips. We provide a way to select a code according to user criteria and also a dynamic code analysis method to estimate the level of protection of assembly implementations using encoding schemes. Finally, we verify our findings by implementing a block cipher PRESENT, protected by encoding scheme based on anticodes, and provide a detailed evaluation of this implementation using different codes. Jakub Breier, Xiaolu Hou, Yang Liu 0003 |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2021 | GUI-Squatting Attack: Automated Generation of Android Phishing AppsabstractMobile phishing attacks, such as mimic mobile browser pages, masquerade as legitimate applications by leveraging repackaging or clone techniques, have caused varied yet significant security concerns. Consequently, detection techniques have been receiving increasing attention. However, many such detection methods are not well tested and may therefore still be vulnerable to new types of phishing attacks. In this article, we propose a new attacking technique, named GUI-Squatting attack, which can generate phishing apps (phapps) automatically and effectively on the Android platform. Our method adopts image processing and deep learning algorithms, to enable powerful and large-scale attacks. We observe that a successful phishing attack requires two conditions, page confusion and logic deception during attacks synthesis. We directly optimize these two conditions to create a practical attack. Our experimental results reveal that existing phishing defenses are less effective against such emergent attacks and may, therefore, stimulate more efficient detection techniques. To further demonstrate that our generatedphappscan not only bypass existing detection techniques, but also deceive real users, we conduct a human study and successfully steal users’ login information. The human study also shows that different response messages (e.g., “Crash” and “Server failed”) after pressing the login button mislead users to regard our phapps as functionality problems instead of security threats. Extensive experiments reveal that such newly proposed attacks still remain mostly undetected, and are worth further exploration. Sen Chen 0001, Lingling Fan 0003, Chunyang Chen 0001, Minhui Xue 0001, Yang Liu 0003, Lihua Xu |
IEEE Trans. Dependable Secur. Comput. | 5 |
| 2021 | Trace-Length Independent Runtime Monitoring of Quantitative PoliciesabstractMetric linear-time logic (MTL) has been widely used to specify runtime policies. Traditionally this use of MTL is to capture the qualitative aspects of the monitored systems, but recent developments in its extensions with aggregate operators allow some quantitative policies to be specified. Our interest in MTL-based policy languages is driven by applications in runtime malware or intrusion detection in platforms like Android and autonomous vehicles, which requires the monitoring algorithm to be independent of the length of the system event traces so that its performance does not degrade as the traces grow. We propose a policy language based on a past-time variant of MTL, extended with an aggregate operator called the metric temporal counting quantifier to specify a policy based on the number of times some sub-policies are satisfied in the specified past time interval. We show that a broad class of policies, but not all policies, specified with our language can be monitored in a trace-length independent way, and provide a concrete algorithm to do so. We implement and test our algorithm in both an existing Android monitoring framework and an autonomous vehicle simulation platform, and show that our approach can effectively specify and monitor quantitative policies drawn from real-world studies. Xiaoning Du 0001, Alwen Tiu, Yang Liu 0003 |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2021 | Text Backdoor Detection Using an Interpretable RNN Abstract ModelabstractDeep neural networks (DNNs) are known to be inherently vulnerable to malicious attacks such as the adversarial attack and the backdoor attack. The former is crafted by adding small perturbations to benign inputs so as to fool a DNN. The latter generally embeds a hidden pattern in a DNN by poisoning the dataset during the training process, which causes the infected model to misbehave on predefined inputs with a specific trigger and normally perform for others. Much work has been conducted on defending against the adversarial samples, while the backdoor attack received much less attention, especially in recurrent neural networks (RNNs), which play an important role in the text processing field. Two main limitations make it hard to directly apply existing image backdoor detection approaches to RNN-based text classification systems. First, a layer in an RNN does not preserve the same feature latent space function for different inputs, making it impossible to map the inserted specific pattern with the neural activations. Second, the text data is inherently discrete, making it hard to optimize the text like image pixels. In this work, we propose a novel backdoor detection approach named InterRNN for RNN-based text classification systems from the interpretation perspective. Specifically, we first propose a novel RNN interpretation technique by constructing a nondeterministic finite automaton (NFA) based abstract model, which can effectively reduce the analysis complexity of an RNN while preserving its original logic rules. Then, based on the abstract model, we can obtain interpretation results that explain the fundamental reason behind the decision for each input. We then detect trigger words by leveraging the differences between the behaviors in the backdoor sentences and those in the normal sentences. The extensive experiment results on four benchmark datasets demonstrate that our approach can generate better interpretation results compared to state-of-the-art approaches and effectively detect backdoors in RNNs. Ming Fan 0002, Ziliang Si, Xiaofei Xie, Yang Liu 0003, Ting Liu 0002 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2021 | Can We Trust Your Explanations? Sanity Checks for Interpreters in Android Malware AnalysisabstractWith the rapid growth of Android malware, many machine learning-based malware analysis approaches are proposed to mitigate the severe phenomenon. However, such classifiers are opaque, non-intuitive, and difficult for analysts to understand the inner decision reason. For this reason, a variety of explanation approaches are proposed to interpret predictions by providing important features. Unfortunately, the explanation results obtained in the malware analysis domain cannot achieve a consensus in general, which makes the analysts confused about whether they can trust such results. In this work, we propose principled guidelines to assess the quality of five explanation approaches by designing three critical quantitative metrics to measure their stability, robustness, and effectiveness. Furthermore, we collect five widely-used malware datasets and apply the explanation approaches on them in two tasks, including malware detection and familial identification. Based on the generated explanation results, we conduct a sanity check of such explanation approaches in terms of the three metrics. The results demonstrate that our metrics can assess the explanation approaches and help us obtain the knowledge of most typical malicious behaviors for malware analysis. Ming Fan 0002, Wenying Wei, Xiaofei Xie, Yang Liu 0003, Xiaohong Guan, Ting Liu 0002 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2021 | A Performance-Sensitive Malware Detection System Using Deep Learning on Mobile DevicesabstractCurrently, Android malware detection is mostly performed on server side against the increasing number of malware. Powerful computing resource provides more exhaustive protection for app markets than maintaining detection by a single user. However, apart from the applications (apps) provided by the official market (i.e., Google Play Store), apps from unofficial markets and third-party resources are always causing serious security threats to end-users. Meanwhile, it is a time-consuming task if the app is downloaded first and then uploaded to the server side for detection, because the network transmission has a lot of overhead. In addition, the uploading process also suffers from the security threats of attackers. Consequently, a last line of defense on mobile devices is necessary and much-needed. In this paper, we propose an effective Android malware detection system, MobiTive, leveraging customized deep neural networks to provide a real-time and responsive detection environment on mobile devices. MobiTive is a pre-installed solution rather than an app scanning and monitoring engine using after installation, which is more practical and secure. Although a deep learning-based approach can be maintained on server side efficiently for malware detection, original deep learning models cannot be directly deployed and executed on mobile devices due to various performance limitations, such as computation power, memory size, and energy. Therefore, we evaluate and investigate the following key points: (1) the performance of different feature extraction methods based on source code or binary code; (2) the performance of different feature type selections for deep learning on mobile devices; (3) the detection accuracy of different deep neural networks on mobile devices; (4) the real-time detection performance and accuracy on different mobile devices; (5) the potential based on the evolution trend of mobile devices' specifications; and finally we further propose a practical solution (MobiTive) to detect Android malware on mobile devices. Sen Chen 0001, Xiaofei Xie, Guozhu Meng, Shangwei Lin 0001, Yang Liu 0003 |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2021 | Exploring the Effects of Blur and Deblurring to Visual Object TrackingabstractThe existence of motion blur can inevitably influence the performance of visual object tracking. However, in contrast to the rapid development of visual trackers, the quantitative effects of increasing levels of motion blur on the performance of visual trackers still remain unstudied. Meanwhile, although image-deblurring can produce visually sharp videos for pleasant visual perception, it is also unknown whether visual object tracking can benefit from image deblurring or not. In this paper, we present a Blurred Video Tracking (BVT) benchmark to address these two problems, which contains a large variety of videos with different levels of motion blurs, as well as ground-truth tracking results. To explore the effects of blur and deblurring to visual object tracking, we extensively evaluate 25 trackers on the proposed BVT benchmark and obtain several new interesting findings. Specifically, we find that light motion blur may improve the accuracy of many trackers, but heavy blur usually hurts the tracking performance. We also observe that image deblurring is helpful to improve tracking accuracy on heavily-blurred videos but hurts the performance of lightly-blurred videos. According to these observations, we then propose a new general GAN-based scheme to improve a tracker's robustness to motion blur. In this scheme, a fine-tuned discriminator can effectively serve as an adaptive blur assessor to enable selective frames deblurring during the tracking process. We use this scheme to successfully improve the accuracy of 6 state-of-the-art trackers on motion-blurred videos. Qing Guo 0005, Wei Feng 0005, Ruijun Gao, Yang Liu 0003, Song Wang 0002 |
IEEE Trans. Image Process. | 4 |
| 2021 | CSim2: Compositional Top-down Verification of Concurrent Systems using Rely-GuaranteeabstractTo make feasible and scalable the verification of large and complex concurrent systems, it is necessary the use of compositional techniques even at the highest abstraction layers. When focusing on the lowest software abstraction layers, such as the implementation or the machine code, the high level of detail of those layers makes the direct verification of properties very difficult and expensive. It is therefore essential to use techniques allowing to simplify the verification on these layers. One technique to tackle this challenge is top-down verification where by means of simulation properties verified on top layers (representing abstract specifications of a system) are propagated down to the lowest layers (that are an implementation of the top layers). There is no need to say that simulation of concurrent systems implies a greater level of complexity, and having compositional techniques to check simulation between layers is also desirable when seeking for both feasibility and scalability of the refinement verification. In this article, we present CSim 2 a (compositional) rely-guarantee-based framework for the top-down verification of complex concurrent systems in the Isabelle/HOL theorem prover. CSim 2 uses CSimpl, a language with a high degree of expressiveness designed for the specification of concurrent programs. Thanks to its expressibility, CSimpl is able to model many of the features found in real world programming languages like exceptions, assertions, and procedures. CSim 2 provides a framework for the verification of rely-guarantee properties to compositionally reason on CSimpl specifications. Focusing on top-down verification, CSim 2 provides a simulation-based framework for the preservation of CSimpl rely-guarantee properties from specifications to implementations. By using the simulation framework, properties proven on the top layers (abstract specifications) are compositionally propagated down to the lowest layers (source or machine code) in each concurrent component of the system. Finally, we show the usability of CSim 2 by running a case study over two CSimpl specifications of an Arinc-653 communication service. In this case study, we prove a complex property on a specification, and we use CSim 2 to preserve the property on lower abstraction layers. David Sanán, Yongwang Zhao, Shangwei Lin 0001, Yang Liu 0003 |
ACM Trans. Program. Lang. Syst. | 4 |
| 2021 | Why an Android App Is Classified as Malware: Toward Malware Classification InterpretationabstractMachine learning–(ML) based approach is considered as one of the most promising techniques for Android malware detection and has achieved high accuracy by leveraging commonly used features. In practice, most of the ML classifications only provide a binary label to mobile users and app security analysts. However, stakeholders are more interested in the reason why apps are classified as malicious in both academia and industry. This belongs to the research area of interpretable ML but in a specific research domain (i.e., mobile malware detection). Although several interpretable ML methods have been exhibited to explain the final classification results in many cutting-edge Artificial Intelligent–based research fields, until now, there is no study interpreting why an app is classified as malware or unveiling the domain-specific challenges. In this article, to fill this gap, we propose a novel and interpretable ML-based approach (named XMal ) to classify malware with high accuracy and explain the classification result meanwhile. (1) The first classification phase of XMal hinges multi-layer perceptron and attention mechanism and also pinpoints the key features most related to the classification result. (2) The second interpreting phase aims at automatically producing neural language descriptions to interpret the core malicious behaviors within apps. We evaluate the behavior description results by leveraging a human study and an in-depth quantitative analysis. Moreover, we further compare XMal with the existing interpretable ML-based methods (i.e., Drebin and LIME) to demonstrate the effectiveness of XMal . We find that XMal is able to reveal the malicious behaviors more accurately. Additionally, our experiments show that XMal can also interpret the reason why some samples are misclassified by ML classifiers. Our study peeks into the interpretable ML through the research of Android malware detection and analysis. Bozhi Wu, Sen Chen 0001, Cuiyun Gao 0001, Lingling Fan 0003, Yang Liu 0003, Weiping Wen, Michael R. Lyu |
ACM Trans. Softw. Eng. Methodol. | 5 |
| 2021 | Mining Likely Analogical APIs Across Third-Party Libraries via Large-Scale Unsupervised API Semantics EmbeddingabstractEstablishing API mappings between third-party libraries is a prerequisite step for library migration tasks. Manually establishing API mappings is tedious due to the large number of APIs to be examined. Having an automatic technique to create a database of likely API mappings can significantly ease the task. Unfortunately, existing techniques either adopt supervised learning mechanism that requires already-ported or functionality similar applications across major programming languages or platforms, which are difficult to come by for an arbitrary pair of third-party libraries, or cannot deal with lexical gap in the API descriptions of different libraries. To overcome these limitations, we present an unsupervised deep learning based approach to embed both API usage semantics and API description (name and document) semantics into vector space for inferring likely analogical API mappings between libraries. Based on deep learning models trained using tens of millions of API call sequences, method names and comments of 2.8 millions of methods from 135,127 GitHub projects, our approach significantly outperforms other deep learning or traditional information retrieval (IR) methods for inferring likely analogical APIs. We implement a proof-of-concept website (https://similarapi.appspot.com) which can recommend analogical APIs for 583,501 APIs of 111 pairs of analogical Java libraries with diverse functionalities. This scale of third-party analogical-API database has never been achieved before. Chunyang Chen 0001, Zhenchang Xing, Yang Liu 0003, Kent Ong Long Xiong |
IEEE Trans. Software Eng. | 3 |
| 2021 | Explaining Regressions via Alignment Slicing and MendingabstractRegression faults, which make working code stop functioning, are often introduced when developers make changes to the software. Many regression fault localization techniques have been proposed. However, issues like inaccuracy and lack of explanation are still obstacles for their practical application. In this work, we propose a trace-based approach to identifying not only where the root cause of a regression bug lies, but also how the defect is propagated to its manifestation as the explanation. In our approach, we keep the trace of original correct version as reference and infer the faulty steps on the trace of regression version so that we can build a causality graph of how the defect is propagated. To this end, we overcomes two technical challenges. First, we align two traces derived from two program versions by extending state-of-the-art trace alignment technique for regression fault with novel relaxation technique. Second, we construct causality graph (i.e., explanation) by adopting a technique calledalignment slicing and mendingto isolate the failure-inducing changes and explain the failure. Our comparative experiment with the state-of-the-art techniques including dynamic slicing, delta-debugging, and symbolic execution on 24 real-world regressions shows that (1) our approach is more accurate on isolating the failure-inducing changes, (2) the generated explanation requires acceptable manual effort to inspect, and (3) our approach requires lower runtime overhead. In addition, we also conduct an applicability experiment based on Defects4J bug repository, showing the potential limitations of our trace-based approach and providing guidance for its practical use. Haijun Wang 0002, Yun Lin 0001, Zijiang Yang 0006, Jun Sun 0001, Yang Liu 0003, Jin Song Dong 0001, Ting Liu 0002 |
IEEE Trans. Software Eng. | 5 |
| 2021 | Erratum to "Accurate and Scalable Cross-Architecture Cross-OS Binary Code Search With Emulation"
Yinxing Xue, Zhengzi Xu, Mahinthan Chandramohan, Yang Liu 0003 |
IEEE Trans. Software Eng. | 4 |
| 2020 | Continuous Multiagent Control Using Collective Behavior Entropy for Large-Scale Home Energy ManagementabstractWith the increasing popularity of electric vehicles, distributed energy generation and storage facilities in smart grid systems, an efficient Demand-Side Management (DSM) is urgent for energy savings and peak loads reduction. Traditional DSM works focusing on optimizing the energy activities for a single household can not scale up to large-scale home energy management problems. Multi-agent Deep Reinforcement Learning (MA-DRL) shows a potential way to solve the problem of scalability, where modern homes interact together to reduce energy consumers consumption while striking a balance between energy cost and peak loads reduction. However, it is difficult to solve such an environment with the non-stationarity, and existing MA-DRL approaches cannot effectively give incentives for expected group behavior. In this paper, we propose a collective MA-DRL algorithm with continuous action space to provide fine-grained control on a large scale microgrid. To mitigate the non-stationarity of the microgrid environment, a novel predictive model is proposed to measure the collective market behavior. Besides, a collective behavior entropy is introduced to reduce the high peak loads incurred by the collective behaviors of all householders in the smart grid. Empirical results show that our approach significantly outperforms the state-of-the-art methods regarding power cost reduction and daily peak loads optimization. Yan Zheng 0002, Jianye Hao, Zhaopeng Meng, Yang Liu 0003 |
AAAI | 5 |
| 2020 | Stealthy and Efficient Adversarial Attacks against Deep Reinforcement LearningabstractAdversarial attacks against conventional Deep Learning (DL) systems and algorithms have been widely studied, and various defenses were proposed. However, the possibility and feasibility of such attacks against Deep Reinforcement Learning (DRL) are less explored. As DRL has achieved great success in various complex tasks, designing effective adversarial attacks is an indispensable prerequisite towards building robust DRL algorithms. In this paper, we introduce two novel adversarial attack techniques to stealthily and efficiently attack the DRL agents. These two techniques enable an adversary to inject adversarial samples in a minimal set of critical moments while causing the most severe damage to the agent. The first technique is the critical point attack: the adversary builds a model to predict the future environmental states and agent's actions, assesses the damage of each possible attack strategy, and selects the optimal one. The second technique is the antagonist attack: the adversary automatically learns a domain-agnostic model to discover the critical moments of attacking the agent in an episode. Experimental results demonstrate the effectiveness of our techniques. Specifically, to successfully attack the DRL agent, our critical point technique only requires 1 (TORCS) or 2 (Atari Pong and Breakout) steps, and the antagonist technique needs fewer than 5 steps (4 Mujoco tasks), which are significant improvements over state-of-the-art methods. Tianwei Zhang 0004, Xiaofei Xie, Lei Ma 0003, Yan Zheng 0002, Kangjie Chen, Yang Liu 0003 |
AAAI | 7 |
| 2020 | Generating Adversarial Examples for Holding Robustness of Source Code Processing ModelsabstractAutomated processing, analysis, and generation of source code are among the key activities in software and system lifecycle. To this end, while deep learning (DL) exhibits a certain level of capability in handling these tasks, the current state-of-the-art DL models still suffer from non-robust issues and can be easily fooled by adversarial attacks.Different from adversarial attacks for image, audio, and natural languages, the structured nature of programming languages brings new challenges. In this paper, we propose a Metropolis-Hastings sampling-based identifier renaming technique, named \fullmethod (\method), which generates adversarial examples for DL models specialized for source code processing. Our in-depth evaluation on a functionality classification benchmark demonstrates the effectiveness of \method in generating adversarial examples of source code. The higher robustness and performance enhanced through our adversarial training with \method further confirms the usefulness of DL models-based method for future fully automated source code processing. Huangzhao Zhang, Zhuo Li 0013, Ge Li 0001, Lei Ma 0003, Yang Liu 0003, Zhi Jin 0001 |
AAAI | 5 |
| 2020 | SPARK: Spatial-Aware Online Incremental Attack Against Visual Tracking
Qing Guo 0005, Xiaofei Xie, Felix Juefei-Xu, Lei Ma 0003, Zhongguo Li, Wanli Xue, Wei Feng 0005, Yang Liu 0003 |
ECCV (25) | 8 |
| 2020 | SeqMobile: An Efficient Sequence-Based Malware Detection System Using RNN on Mobile DevicesabstractWith the proliferation of Android malware, the demand for an effective and efficient malware detection system is on the rise. The existing device-end learning based solutions tend to extract limited syntax features, such as permissions and API calls, to meet a certain time constraint of mobile devices. However, unlike sequence-based features, syntax features lack the semantics which can represent the potential malicious behaviors and further result in more robust model with high accuracy for malware detection. In this paper, we propose an efficient Android malware detection system, named SeqMobile, which adopts behavior-based sequence features and leverages customized deep neural networks on mobile devices instead of the server end. Different from the traditional sequence-based approaches on server end, to meet the performance demand on mobile devices, SeqMobile accepts three effective performance optimization methods to reduce the time of feature extraction and prediction. To evaluate the effectiveness and efficiency of our system, we conduct experiments from the following aspects 1) the detection accuracy of different recurrent neural networks (RNN); 2) the feature extraction performance on different mobile devices, and 3) the detection accuracy and prediction time cost of different sequence lengths. The results unveil that SeqMobile can effectively detect malware with high accuracy. Moreover, our performance optimization methods have proven to improve the performance of training and prediction by at least twofold. Additionally, to discover the potential performance optimization from the state-of-the-art TensorFlow model optimization toolkit for our sequence-based approach, we also provide an evaluation on the toolkit, which can serve as a guidance for other systems leveraging on sequence-based learning approach. Overall, we conclude that our sequence-based approach, together with our performance optimization methods, enable us to efficiently detect malware under the performance demands of mobile devices. Jing Qiang Lim, Sen Chen 0001, Shangwei Lin 0001, Yang Liu 0003 |
ICECCS | 5 |
| 2020 | Privacy-Aware UAV Flights through Self-Configuring Motion PlanningabstractDuring flights, an unmanned aerial vehicle (UAV) may not be allowed to move across certain areas due to soft constraints such as privacy restrictions. Current methods on self-adaption focus mostly on motion planning such that the trajectory does not trespass predetermined restricted areas. When the environment is cluttered with uncertain obstacles, however, these motion planning algorithms are not flexible enough to find a trajectory that satisfies additional privacy-preserving requirements within a tight time budget during the flights. In this paper, we propose a privacy risk aware motion planning method through the reconfiguration of privacy-sensitive sensors. It minimises environmental impact by re-configuring the sensor during flight, while still guaranteeing the safety and energy hard constraints such as collision avoidance and timeliness. First, we formulate a model for assessing privacy risks of dynamically detected restricted areas. In case the UAV cannot find a feasible solution to satisfy both hard and soft constraints from the current configuration, our decision making method can then produce an optimal reconfiguration of the privacy-sensitive sensor with a more efficient trajectory. We evaluate the proposal through various simulations with different settings in a virtual environment and also validate the approach through real test flights on DJI Matrice 100 UAV. Yixing Luo, Yijun Yu 0001, Zhi Jin 0001, Yao Li 0011, Zuohua Ding, Yuan Zhou 0005, Yang Liu 0003 |
ICRA | 7 |
| 2020 | DesignDiff: Continuously Modeling Software Design Difference from Code RevisionsabstractThe design structure of a system continuously evolves as the consequence of fast-paced code revisions. Agile techniques, such as continuous testing, ensures the function goals of a system with every code revision. However, there lacks an efficient approach that can continuously model the design difference resulting from every single code revision to facilitate comprehension and ensure the design quality. This paper contributes a novel design modeling approach, called Design Differencing (DESIGNDIFF), that models and visualizes the highlevel design differences resulting from every code revision. This paper defines a complete and general set of 17 design change operators to capture the design difference from any code revision. We evaluated the potential of DESIGNDIFF in three aspects. First, a user study of 10 developers indicated that DESIGNDIFF can help practitioners to faster and better understand high-level design differences from real-life software commits. Second, DESIGNDIFF analyzed 14,832 real-life commits in five real-life projects: finding 4,189 commits altered the software design, 855 commits introduced and 337 commits eliminated design flaws. The latency between the flaw introduction and elimination is on average 2 months to 2 years! With an affordable performance overhead, DESIGNDIFF has great potential to benefit practitioners in more applications. Xiao Wang 0030, Lu Xiao 0001, Kaifeng Huang 0001, Bihuan Chen 0001, Yang Liu 0003 |
ICSA | 6 |
| 2020 | Butterfly Space: An Architectural Approach for Investigating Performance IssuesabstractPerformance issues widely exist in modern software systems. Existing performance optimization approaches, such as dynamic profiling, usually fail to consider the impacts of architectural connections among methods on performance issues. This paper contributes an architectural approach, Butterfly Space modeling, to investigate performance issues. Each Butterfly Space is composed of 1) a seed method; 2) methods in the "upper wing" that call the seed directly or transitively; and 3) methods in the "lower wing" that are called by the seed, directly or transitively. The rationale is that the performance of the seed method impacts and is impacted by all the other methods in the space because of the call relationship. As such, developers can more efficiently investigate groups of connected performance improvement opportunities in Butterfly Spaces. We studied three real-world open source Java projects to evaluate such potential. Our findings are three-fold: 1) If the seed method of a Butterfly Space contains performance problems, up to 60% of the methods in the space also contain performance problems; 2) Butterfly Spaces can potentially help to non-trivially increase the precision/recall and reduce the costs in identifying performance improvement opportunities, compared to dynamic profiling; and 3) Visualizing dynamic profiling metrics with Butterfly Spaces simultaneously help to reveal two typical patterns, namely Expensive Callee and Inefficient Caller, that are responsible for performance problems and provide insights on where to improve next. We believe that Butterfly Space modeling has great potential for investigating performance issues. Lu Xiao 0001, Xiao Wang 0030, Zhifei Chen, Bihuan Chen 0001, Yang Liu 0003 |
ICSA | 6 |
| 2020 | An empirical assessment of security risks of global Android banking appsabstractMobile banking apps, belonging to the most security-critical app category, render massive and dynamic transactions susceptible to security risks. Given huge potential financial loss caused by vulnerabilities, existing research lacks a comprehensive empirical study on the security risks of global banking apps to provide useful insights and improve the security of banking apps. Sen Chen 0001, Lingling Fan 0003, Guozhu Meng, Ting Su 0001, Minhui Xue 0001, Yinxing Xue, Yang Liu 0003, Lihua Xu |
ICSE | 7 |
| 2020 | Typestate-guided fuzzer for discovering use-after-free vulnerabilitiesabstractExisting coverage-based fuzzers usually use the individual control flow graph (CFG) edge coverage to guide the fuzzing process, which has shown great potential in finding vulnerabilities. However, CFG edge coverage is not effective in discovering vulnerabilities such as use-after-free (UaF). This is because, to trigger UaF vulnerabilities, one needs not only to cover individual edges, but also to traverse some (long) sequence of edges in a particular order, which is challenging for existing fuzzers. To this end, we propose to model UaF vulnerabilities as typestate properties, and develop a typestate-guided fuzzer, named UAFL, for discovering vulnerabilities violating typestate properties. Given a typestate property, we first perform a static typestate analysis to find operation sequences potentially violating the property. Our fuzzing process is then guided by the operation sequences in order to progressively generate test cases triggering property violations. In addition, we also employ an information flow analysis to improve the efficiency of the fuzzing process. We have performed a thorough evaluation of UAFL on 14 widely-used real-world programs. The experiment results show that UAFL substantially outperforms the state-of-the-art fuzzers, including AFL, AFLFast, FairFuzz, MOpt, Angora and QSYM, in terms of the time taken to discover vulnerabilities. We have discovered 10 previously unknown vulnerabilities, and received 5 new CVEs. Haijun Wang 0002, Xiaofei Xie, Yi Li 0008, Cheng Wen 0002, Yuekang Li, Yang Liu 0003, Shengchao Qin, Hongxu Chen 0001, Yulei Sui |
ICSE | 6 |
| 2020 | MemLock: memory usage guided fuzzingabstractUncontrolled memory consumption is a kind of critical software security weaknesses. It can also become a security-critical vulnerability when attackers can take control of the input to consume a large amount of memory and launch a Denial-of-Service attack. However, detecting such vulnerability is challenging, as the state-of-the-art fuzzing techniques focus on the code coverage but not memory consumption. To this end, we propose a memory usage guided fuzzing technique, named MemLock, to generate the excessive memory consumption inputs and trigger uncontrolled memory consumption bugs. The fuzzing process is guided with memory consumption information so that our approach is general and does not require any domain knowledge. We perform a thorough evaluation for MemLock on 14 widely-used real-world programs. Our experiment results show that MemLock substantially outperforms the state-of-the-art fuzzing techniques, including AFL, AFLfast, PerfFuzz, FairFuzz, Angora and QSYM, in discovering memory consumption bugs. During the experiments, we discovered many previously unknown memory consumption bugs and received 15 new CVEs. Cheng Wen 0002, Haijun Wang 0002, Yuekang Li, Shengchao Qin, Yang Liu 0003, Zhiwu Xu 0001, Hongxu Chen 0001, Xiaofei Xie, Geguang Pu, Ting Liu 0002 |
ICSE | 5 |
| 2020 | Towards characterizing adversarial defects of deep learning software from the lens of uncertaintyabstractOver the past decade, deep learning (DL) has been successfully applied to many industrial domain-specific tasks. However, the current state-of-the-art DL software still suffers from quality issues, which raises great concern especially in the context of safety- and security-critical scenarios. Adversarial examples (AEs) represent a typical and important type of defects needed to be urgently addressed, on which a DL software makes incorrect decisions. Such defects occur through either intentional attack or physical-world noise perceived by input sensors, potentially hindering further industry deployment. The intrinsic uncertainty nature of deep learning decisions can be a fundamental reason for its incorrect behavior. Although some testing, adversarial attack and defense techniques have been recently proposed, it still lacks a systematic study to uncover the relationship between AEs and DL uncertainty. Xiyue Zhang 0001, Xiaofei Xie, Lei Ma 0003, Xiaoning Du 0001, Yang Liu 0003, Jianjun Zhao 0001, Meng Sun 0002 |
ICSE | 6 |
| 2020 | Source Code based On-demand Class Documentation GenerationabstractIn this paper, we present OpenAPIDocGen2, a tool that generates on-demand class documentation based on source code and documentation analysis. For a given class, OpenAPIDocGen2 generates a combined documentation for it, which includes functionality descriptions, directives, domain concepts, usage examples, class/method roles, key methods, relevant classes/methods, characteristics and concepts classification, and usage scenarios. Mingwei Liu 0002, Xin Peng 0001, Xiujie Meng, Huanjun Xu, Shuangshuang Xing, Xin Wang 0119, Yang Liu 0003 |
ICSME | 7 |
| 2020 | An Empirical Study of Usages, Updates and Risks of Third-Party Libraries in Java ProjectsabstractThird-party libraries play a key role in software development as they can relieve developers of the heavy burden of re-implementing common functionalities. However, third-party libraries and client projects evolve asynchronously. As a result, out-dated third-party libraries might be used in client projects while developers are not aware of the potential risk (e.g., security bug). Outdated third-party libraries may be updated in client projects in a delayed way, and developers may be less aware of the potential risk (e.g., API incompatibility) in updates. Developers of third-party libraries may be unaware of how their third-party libraries are used or updated in client projects. Therefore, a quantitative and holistic study on usages, updates and risks of third-party libraries in open-source projects can provide concrete evidences on these problems, and practical insights to improve the ecosystem. In this paper, we contribute such a study in Java ecosystem. In particular, we conduct a library usage analysis (e.g., usage intensity and outdatedness) and library update analysis (e.g., update intensity and delay) on 806 open-source projects and 13,565 third- party libraries. Then, we carry out a library risk analysis (e.g., usage risk and update risk) on 806 open-source projects and 544 security bugs. These analyses aim to quantify the usage and update practices and the potential risk of using and updating outdated third-party libraries with respect to security bugs from two holistic perspectives (i.e., open-source projects and third-party libraries). Our findings suggest practical implications to developers and researchers on problems and potential solutions in maintaining third-party libraries (e.g., smart alerting and automated updating of outdated third-party libraries). To indicate the usefulness of our findings, we design a smart alerting system for assisting developers to make confident decisions when updating third-party libraries. 33 and 24 open-source projects have confirmed and updated third-party libraries after receiving our alerts. Bihuan Chen 0001, Kaifeng Huang 0001, Congying Xu, Xin Peng 0001, Yijian Wu, Yang Liu 0003 |
ICSME | 8 |
| 2020 | Generating Behavior-Diverse Game AIs with Evolutionary Multi-Objective Deep Reinforcement LearningabstractGenerating diverse behaviors for game artificial intelligence (Game AI) has been long recognized as a challenging task in the game industry. Designing a Game AI with a satisfying behavioral characteristic (style) heavily depends on the domain knowledge and is hard to achieve manually. Deep reinforcement learning sheds light on advancing the automatic Game AI design. However, most of them focus on creating a superhuman Game AI, ignoring the importance of behavioral diversity in games. To bridge the gap, we introduce a new framework, named EMOGI, which can automatically generate desirable styles with almost no domain knowledge. More importantly, EMOGI succeeds in creating a range of diverse styles, providing behavior-diverse Game AIs. Evaluations on the Atari and real commercial games indicate that, compared to existing algorithms, EMOGI performs better in generating diverse behaviors and significantly improves the efficiency of Game AI design. Ruimin Shen, Yan Zheng 0002, Jianye Hao, Zhaopeng Meng, Changjie Fan, Yang Liu 0003 |
IJCAI | 7 |
| 2020 | FakeSpotter: A Simple yet Robust Baseline for Spotting AI-Synthesized Fake FacesabstractIn recent years, generative adversarial networks (GANs) and its variants have achieved unprecedented success in image synthesis. They are widely adopted in synthesizing facial images which brings potential security concerns to humans as the fakes spread and fuel the misinformation. However, robust detectors of these AI-synthesized fake faces are still in their infancy and are not ready to fully tackle this emerging challenge. In this work, we propose a novel approach, named FakeSpotter, based on monitoring neuron behaviors to spot AI-synthesized fake faces. The studies on neuron coverage and interactions have successfully shown that they can be served as testing criteria for deep learning systems, especially under the settings of being exposed to adversarial attacks. Here, we conjecture that monitoring neuron behavior can also serve as an asset in detecting fake faces since layer-by-layer neuron activation patterns may capture more subtle features that are important for the fake detector. Experimental results on detecting four types of fake faces synthesized with the state-of-the-art GANs and evading four perturbation attacks show the effectiveness and robustness of our approach. Run Wang 0001, Felix Juefei-Xu, Lei Ma 0003, Xiaofei Xie, Yihao Huang 0001, Jian Wang 0067, Yang Liu 0003 |
IJCAI | 7 |
| 2020 | An Empirical Evaluation of GDPR Compliance Violations in Android mHealth AppsabstractThe purpose of the General Data Protection Regulation (GDPR) is to provide improved privacy protection. If an app controls personal data from users, it needs to be compliant with GDPR. However, GDPR lists general rules rather than exact step-by-step guidelines about how to develop an app that fulfills the requirements. Therefore, there may exist GDPR compliance violations in existing apps, which would pose severe privacy threats to app users. In this paper, we take mobile health applications (mHealth apps) as a peephole to examine the status quo of GDPR compliance in Android apps. We first propose an automated system, named HPDROID, to bridge the semantic gap between the general rules of GDPR and the app implementations by identifying the data practices declared in the app privacy policy and the data relevant behaviors in the app code. Then, based on HPDROID, we detect three kinds of GDPR compliance violations, including the incompleteness of privacy policy, the inconsistency of data collections, and the insecurity of data transmission. We perform an empirical evaluation of 796 mHealth apps. The results reveal that 189 (23.7%) of them do not provide complete privacy policies. Moreover, 59 apps collect sensitive data through different measures, but 46 (77.9%) of them contain at least one inconsistent collection behavior. Even worse, among the 59 apps, only 8 apps try to ensure the transmission security of collected data. However, all of them contain at least one encryption or SSL misuse. Our work exposes severe privacy issues to raise awareness of privacy protection for app users and developers. Ming Fan 0002, Le Yu 0002, Sen Chen 0001, Hao Zhou 0043, Xiapu Luo, Shuyue Li, Yang Liu 0003, Jun Liu 0002, Ting Liu 0002 |
ISSRE | 7 |
| 2020 | An empirical study on ARM disassembly toolsabstractWith the increasing popularity of embedded devices, ARM is becoming the dominant architecture for them. In the meanwhile, there is a pressing need to perform security assessments for these devices. Due to different types of peripherals, it is challenging to dynamically run the firmware of these devices in an emulated environment. Therefore, the static analysis is still commonly used. Existing work usually leverages off-the-shelf tools to disassemble stripped ARM binaries and (implicitly) assume that reliable disassembling binaries and function recognition are solved problems. However, whether this assumption really holds is unknown. Muhui Jiang, Yajin Zhou, Xiapu Luo, Ruoyu Wang 0001, Yang Liu 0003, Kui Ren 0001 |
ISSTA | 5 |
| 2020 | Patch based vulnerability matching for binary programsabstractThe binary-level function matching has been widely used to detect whether there are 1-day vulnerabilities in released programs. However, the high false positive is a challenge for current function matching solutions, since the vulnerable function is highly similar to its corresponding patched version. In this paper, the Binary X-Ray (BinXray), a patch based vulnerability matching approach, is proposed to identify the specific 1-day vulnerabilities in target programs accurately and effectively. In the preparing step, a basic block mapping algorithm is designed to extract the signature of a patch, by comparing the given vulnerable and patched programs. The signature is represented as a set of basic block traces. In the detection step, the patching semantics is applied to reduce irrelevant basic block traces to speed up the signature searching. The trace similarity is also designed to identify whether a target program is patched. In experiments, 12 real software projects related to 479 CVEs are collected. BinXray achieves 93.31% accuracy and the analysis time cost is only 296.17ms per function, outperforming the state-of-the-art works. Zhengzi Xu, Bihuan Chen 0001, Fu Song, Yang Liu 0003, Ting Liu 0002 |
ISSTA | 5 |
| 2020 | How are Deep Learning Models Similar?: An Empirical Study on Clone Analysis of Deep Learning SoftwareabstractDeep learning (DL) has been successfully applied to many cutting-edge applications, e.g., image processing, speech recognition, and natural language processing. As more and more DL software is made open-sourced, publicly available, and organized in model repositories and stores (Model Zoo, ModelDepot), there comes a need to understand the relationships of these DL models regarding their maintenance and evolution tasks. Although clone analysis has been extensively studied for traditional software, up to the present, clone analysis has not been investigated for DL software. Since DL software adopts the data-driven development paradigm, it is still not clear whether and to what extent the clone analysis techniques of traditional software could be adapted to DL software. Xiongfei Wu, Liangyu Qin, Xiaofei Xie, Lei Ma 0003, Yinxing Xue, Yang Liu 0003, Jianjun Zhao 0001 |
ICPC | 7 |
| 2020 | Cats Are Not Fish: Deep Learning Testing Calls for Out-Of-Distribution AwarenessabstractAs Deep Learning (DL) is continuously adopted in many industrial applications, its quality and reliability start to raise concerns. Similar to the traditional software development process, testing the DL software to uncover its defects at an early stage is an effective way to reduce risks after deployment. According to the fundamental assumption of deep learning, the DL software does not provide statistical guarantee and has limited capability in handling data that falls outside of its learned distribution, i.e., out-of-distribution (OOD) data. Although recent progress has been made in designing novel testing techniques for DL software, which can detect thousands of errors, the current state-of-the-art DL testing techniques usually do not take the distribution of generated test data into consideration. It is therefore hard to judge whether the "identified errors" are indeed meaningful errors to the DL application (i.e., due to quality issues of the model) or outliers that cannot be handled by the current model (i.e., due to the lack of training data). Tofill this gap, we take the first step and conduct a large scale empirical study, with a total of 451 experiment configurations, 42 deep neural networks (DNNs) and 1.2 million test data instances, to investigate and characterize the impact of OOD-awareness on DL testing. We further analyze the consequences when DL systems go into production by evaluating the effectiveness of adversarial retraining with distribution-aware errors. The results confirm that introducing data distribution awareness in both testing and enhancement phases outperforms distribution unaware retraining by up to 21.5%. David Berend, Xiaofei Xie, Lei Ma 0003, Lingjun Zhou, Yang Liu 0003, Jianjun Zhao 0001 |
ASE | 5 |
| 2020 | Marble: Model-based Robustness Analysis of Stateful Deep Learning SystemsabstractState-of-the-art deep learning (DL) systems are vulnerable to adversarial examples, which hinders their potential adoption in safety-and security-critical scenarios. While some recent progress has been made in analyzing the robustness of feed-forward neural networks, the robustness analysis for stateful DL systems, such as recurrent neural networks (RNNs), still remains largely uncharted. In this paper, we propose Marble, a model-based approach for quantitative robustness analysis of real-world RNN-based DL systems. Marble builds a probabilistic model to compactly characterize the robustness of RNNs through abstraction. Furthermore, we propose an iterative refinement algorithm to derive a precise abstraction, which enables accurate quantification of the robustness measurement. We evaluate the effectiveness of Marble on both LSTM and GRU models trained separately with three popular natural language datasets. The results demonstrate that (1) our refinement algorithm is more efficient in deriving an accurate abstraction than the random strategy, and (2) Marble enables quantitative robustness analysis, in rendering better efficiency, accuracy, and scalability than the state-of-the-art techniques. Xiaoning Du 0001, Yi Li 0008, Xiaofei Xie, Lei Ma 0003, Yang Liu 0003, Jianjun Zhao 0001 |
ASE | 5 |
| 2020 | Audee: Automated Testing for Deep Learning FrameworksabstractDeep learning (DL) has been applied widely, and the quality of DL system becomes crucial, especially for safety-critical applications. Existing work mainly focuses on the quality analysis of DL models, but lacks attention to the underlying frameworks on which all DL models depend. In this work, we propose Audee, a novel approach for testing DL frameworks and localizing bugs. Audee adopts a search-based approach and implements three different mutation strategies to generate diverse test cases by exploring combinations of model structures, parameters, weights and inputs. Audee is able to detect three types of bugs: logical bugs, crashes and Not-a-Number (NaN) errors. In particular, for logical bugs, Audee adopts a cross-reference check to detect behavioural inconsistencies across multiple frameworks (e.g., TensorFlow and PyTorch), which may indicate potential bugs in their implementations. For NaN errors, Audee adopts a heuristic-based approach to generate DNNs that tend to output outliers (i.e., too large or small values), and these values are likely to produce NaN. Furthermore, Audee leverages a causal-testing based technique to localize layers as well as parameters that cause inconsistencies or bugs. To evaluate the effectiveness of our approach, we applied Audee on testing four DL frameworks, i.e., TensorFlow, PyTorch, CNTK, and Theano. We generate a large number of DNNs which cover 25 widely-used APIs in the four frameworks. The results demonstrate that Audee is effective in detecting inconsistencies, crashes and NaN errors. In total, 26 unique unknown bugs were discovered, and 7 of them have already been confirmed or fixed by the developers. Xiaofei Xie, Yi Li 0008, Xiaoyu Zhang 0013, Yang Liu 0003, Xiaohong Li 0001, Chao Shen 0001 |
ASE | 5 |
| 2020 | Generating Concept based API Element Comparison Using a Knowledge GraphabstractDevelopers are concerned with the comparison of similar APIs in terms of their commonalities and (often subtle) differences. Our empirical study of Stack Overflow questions and API documentation confirms that API comparison questions are common and can often be answered by knowledge contained in API reference documentation. Our study also identifies eight types of API statements that are useful for API comparison. Based on these findings, we propose a knowledge graph based approach APIComp that automatically extracts API knowledge from API reference documentation to support the comparison of a pair of API classes or methods from different aspects. Our approach includes an offline phase for constructing an API knowledge graph, and an online phase for generating an API comparison result for a given pair of API elements. Our evaluation shows that the quality of different kinds of extracted knowledge in the API knowledge graph is generally high. Furthermore, the comparison results generated by APIComp are significantly better than those generated by a baseline approach based on heuristic rules and text similarity, and our generated API comparison results are useful for helping developers in API selection tasks. Yang Liu 0003, Mingwei Liu 0002, Xin Peng 0001, Christoph Treude, Zhenchang Xing, Xiaoxin Zhang |
ASE | 1 |
| 2020 | SADT: Syntax-Aware Differential Testing of Certificate Validation in SSL/TLS ImplementationsabstractThe security assurance of SSL/TLS critically depends on the correct validation of X.509 certificates. Therefore, it is important to check whether a certificate is correctly validated by the SSL/TLS implementations. Although differential testing has been proven to be effective in finding semantic bugs, it still suffers from the following limitations: (1) The syntax of test cases cannot be correctly guaranteed. (2) Current test cases are not diverse enough to cover more implementation behaviours. This paper tackles these problems by introducing SADT, a novel syntax-aware differential testing framework for evaluating the certificate validation process in SSL/TLS implementations. We first propose a tree-based mutation strategy to ensure that the generated certificates are syntactically correct, and then diversify the certificates by sharing interesting test cases among all target SSL/TLS implementations. Such generated certificates are more likely to trigger discrepancies among SSL/TLS implementations, which may indicate some potential bugs. Lili Quan 0001, Hongxu Chen 0001, Xiaofei Xie, Xiaohong Li 0001, Yang Liu 0003, Jing Hu 0007 |
ASE | 6 |
| 2020 | Automated Third-Party Library Detection for Android Applications: Are We There Yet?abstractThird-party libraries (TPLs) have become a significant part of the Android ecosystem. Developers can employ various TPLs with different functionalities to facilitate their app development. Unfortunately, the popularity of TPLs also brings new challenges and even threats. TPLs may carry malicious or vulnerable code, which can infect popular apps to pose threats to mobile users. Besides, the code of third-party libraries could constitute noises in some downstream tasks (e.g., malware and repackaged app detection). Thus, researchers have developed various tools to identify TPLs. However, no existing work has studied these TPL detection tools in detail; different tools focus on different applications with performance differences, but little is known about them. Xian Zhan, Lingling Fan 0003, Tianming Liu 0002, Sen Chen 0001, Li Li 0029, Haoyu Wang 0001, Xiapu Luo, Yang Liu 0003 |
ASE | 9 |
| 2020 | FakePolisher: Making DeepFakes More Detection-Evasive by Shallow ReconstructionabstractAt this moment, GAN-based image generation methods are still imperfect, whose upsampling design has limitations in leaving some certain artifact patterns in the synthesized image. Such artifact patterns can be easily exploited (by recent methods) for difference detection of real and GAN-synthesized images. However, the existing detection methods put much emphasis on the artifact patterns, which can become futile if such artifact patterns were reduced. Yihao Huang 0001, Felix Juefei-Xu, Run Wang 0001, Qing Guo 0005, Lei Ma 0003, Xiaofei Xie, Weikai Miao, Yang Liu 0003, Geguang Pu |
ACM Multimedia | 9 |
| 2020 | DeepRhythm: Exposing DeepFakes with Attentional Visual Heartbeat RhythmsabstractAs the GAN-based face image and video generation techniques, widely known as DeepFakes, have become more and more matured and realistic, there comes a pressing and urgent demand for effective DeepFakes detectors. Motivated by the fact that remote visual photoplethysmography (PPG) is made possible by monitoring the minuscule periodic changes of skin color due to blood pumping through the face, we conjecture that normal heartbeat rhythms found in the real face videos will be disrupted or even entirely broken in a DeepFake video, making it a potentially powerful indicator for DeepFake detection. In this work, we propose DeepRhythm, a DeepFake detection technique that exposes DeepFakes by monitoring the heartbeat rhythms. DeepRhythm utilizes dual-spatial-temporal attention to adapt to dynamically changing face and fake types. Extensive experiments on FaceForensics++ and DFDC-preview datasets have confirmed our conjecture and demonstrated not only the effectiveness, but also the generalization capability of DeepRhythm over different datasets by various DeepFakes generation techniques and multifarious challenging degradations. Hua Qi, Qing Guo 0005, Felix Juefei-Xu, Xiaofei Xie, Lei Ma 0003, Wei Feng 0005, Yang Liu 0003, Jianjun Zhao 0001 |
ACM Multimedia | 7 |
| 2020 | Amora: Black-box Adversarial Morphing AttackabstractNowadays, digital facial content manipulation has become ubiquitous and realistic with the success of generative adversarial networks (GANs), making face recognition (FR) systems suffer from unprecedented security concerns. In this paper, we investigate and introduce a new type of adversarial attack to evade FR systems by manipulating facial content, called adversarial morphing attack (a.k.a. Amora). In contrast to adversarial noise attack that perturbs pixel intensity values by adding human-imperceptible noise, our proposed adversarial morphing attack works at the semantic level that perturbs pixels spatially in a coherent manner. To tackle the black-box attack problem, we devise a simple yet effective joint dictionary learning pipeline to obtain a proprietary optical flow field for each attack. Our extensive evaluation on two popular FR systems demonstrates the effectiveness of our adversarial morphing attack at various levels of morphing intensity with smiling facial expression manipulations. Both open-set and closed-set experimental results indicate that a novel black-box adversarial attack based on local deformation is possible, and is vastly different from additive noise attacks. The findings of this work potentially pave a new research direction towards a more thorough understanding and investigation of image-based adversarial attacks and defenses. Run Wang 0001, Felix Juefei-Xu, Qing Guo 0005, Yihao Huang 0001, Xiaofei Xie, Lei Ma 0003, Yang Liu 0003 |
ACM Multimedia | 7 |
| 2020 | DeepSonar: Towards Effective and Robust Detection of AI-Synthesized Fake VoicesabstractWith the recent advances in voice synthesis, AI-synthesized fake voices are indistinguishable to human ears and widely are applied to produce realistic and natural DeepFakes, exhibiting real threats to our society. However, effective and robust detectors for synthesized fake voices are still in their infancy and are not ready to fully tackle this emerging threat. In this paper, we devise a novel approach, named DeepSonar, based on monitoring neuron behaviors of speaker recognition (SR) system, i.e., a deep neural network (DNN), to discern AI-synthesized fake voices. Layer-wise neuron behaviors provide an important insight to meticulously catch the differences among inputs, which are widely employed for building safety, robust, and interpretable DNNs. In this work, we leverage the power of layer-wise neuron activation patterns with a conjecture that they can capture the subtle differences between real and AI-synthesized fake voices, in providing a cleaner signal to classifiers than raw inputs. Experiments are conducted on three datasets (including commercial products from Google, Baidu, etc) containing both English and Chinese languages to corroborate the high detection rates (98.1% average accuracy) and low false alarm rates (about 2% error rate) of DeepSonar in discerning fake voices. Furthermore, extensive experimental results also demonstrate its robustness against manipulation attacks (e.g., voice conversion and additive real-world noises). Our work further poses a new insight into adopting neuron behaviors for effective and robust AI aided multimedia fakes forensics as an inside-out approach instead of being motivated and swayed by various artifacts introduced in synthesizing fakes. Run Wang 0001, Felix Juefei-Xu, Yihao Huang 0001, Qing Guo 0005, Xiaofei Xie, Lei Ma 0003, Yang Liu 0003 |
ACM Multimedia | 7 |
| 2020 | Watch out! Motion is Blurring the Vision of Your Deep Neural NetworksabstractThe state-of-the-art deep neural networks (DNNs) are vulnerable against adversarial examples with additive random-like noise perturbations. While such examples are hardly found in the physical world, the image blurring effect caused by object motion, on the other hand, commonly occurs in practice, making the study of which greatly important especially for the widely adopted real-time image processing tasks (e.g., object detection, tracking). In this paper, we initiate the first step to comprehensively investigate the potential hazards of blur effect for DNN, caused by object motion. We propose a novel adversarial attack method that can generate visually natural motion-blurred adversarial examples, named motion-based adversarial blur attack (ABBA). To this end, we first formulate the kernel-prediction-based attack where an input image is convolved with kernels in a pixel-wise way, and the misclassification capability is achieved by tuning the kernel weights. To generate visually more natural and plausible examples, we further propose the saliency-regularized adversarial kernel prediction, where the salient region serves as a moving object, and the predicted kernel is regularized to achieve naturally visual effects. Besides, the attack is further enhanced by adaptively tuning the translations of object and background. A comprehensive evaluation on the NeurIPS'17 adversarial competition dataset demonstrates the effectiveness of ABBA by considering various kernel sizes, translations, and regions. The in-depth study further confirms that our method shows a more effective penetrating capability to the state-of-the-art GAN-based deblurring mechanisms compared with other blurring methods. We release the code to \url{https://github.com/tsingqguo/ABBA}. Qing Guo 0005, Felix Juefei-Xu, Xiaofei Xie, Lei Ma 0003, Jian Wang 0067, Wei Feng 0005, Yang Liu 0003 |
NeurIPS | 8 |
| 2020 | Semantic Understanding of Smart Contracts: Executable Operational Semantics of SolidityabstractBitcoin has been a popular research topic recently. Ethereum (ETH), a second generation of cryptocurrency, extends Bitcoin's design by offering a Turing-complete programming language called Solidity to develop smart contracts. Smart contracts allow creditable execution of contracts on EVM (Ethereum Virtual Machine) without third parties. Developing correct and secure smart contracts is challenging due to the decentralized computation nature of the blockchain. Buggy smart contracts may lead to huge financial loss. Furthermore, smart contracts are very hard, if not impossible, to patch once they are deployed. Thus, there is a recent surge of interest in analyzing and verifying smart contracts. While most of the existing works either focus on EVM bytecode or translate Solidity smart contracts into programs in intermediate languages, we argue that it is important and necessary to understand and formally define the semantics of Solidity since programmers write and reason about smart contracts at the level of source code. In this work, we develop a formal semantics for Solidity which provides a formal specification of smart contracts to define semantic-level security properties for the high-level verification. Furthermore, the proposed semantics defines correct and secure high-level execution behaviours of smart contracts to reason about compiler bugs and assist developers in writing secure smart contracts. Jiao Jiao 0002, Shuanglong Kan, Shangwei Lin 0001, David Sanán, Yang Liu 0003, Jun Sun 0001 |
SP | 5 |
| 2020 | MUZZ: Thread-aware Grey-box Fuzzing for Effective Bug Hunting in Multithreaded Programs
Hongxu Chen 0001, Shengjian Guo, Yinxing Xue, Yulei Sui, Cen Zhang, Yuekang Li, Haijun Wang 0002, Yang Liu 0003 |
USENIX Security Symposium | 8 |
| 2020 | MVP: Detecting Vulnerabilities using Patch-Enhanced Vulnerability Signatures
Yang Xiao 0011, Bihuan Chen 0001, Chendong Yu, Zhengzi Xu, Zimu Yuan, Feng Li 0045, Binghong Liu, Yang Liu 0003, Wei Huo 0005, Wenchang Shi |
USENIX Security Symposium | 8 |
| 2020 | Automatic Hot Patch Generation for Android Kernels
Zhengzi Xu, Longri Zheng, Liangzhao Xia, Chenfu Bao, Zhi Wang 0004, Yang Liu 0003 |
USENIX Security Symposium | 7 |
| 2020 | CORE: Automating Review Recommendation for Code ChangesabstractCode review is a common process that is used by developers, in which a reviewer provides useful comments or points out defects in the submitted source code changes via pull request. Code review has been widely used for both industry and open-source projects due to its capacity in early defect identification, project maintenance, and code improvement. With rapid updates on project developments, code review becomes a non-trivial and labor-intensive task for reviewers. Thus, an automated code review engine can be beneficial and useful for project development in practice. Although there exist prior studies on automating the code review process by adopting static analysis tools or deep learning techniques, they often require external sources such as partial or full source code for accurate review suggestion. In this paper, we aim at automating the code review process only based on code changes and the corresponding reviews but with better performance. The hinge of accurate code review suggestion is to learn good representations for both code changes and reviews. To achieve this with limited source, we design a multi-level embedding (i.e., word embedding and character embedding) approachto represent the semantics provided by code changes and reviews. The embeddings are then well trained through a proposed attentional deep learning model, as a whole named CORE. We evaluate the effectiveness of CORE on code changes and reviews collected from 19 popular Java projects hosted on Github. Experimental results show that our model CORE can achieve significantly better performance than the state-of-the-art model (DeepMem), with an increase of 131.03% in terms of Recall@10 and 150.69% in terms of Mean Reciprocal Rank. Qualitative general word analysis among project developers also demonstrates the performance of CORE in automating code review. Jing Kai Siow, Cuiyun Gao 0001, Lingling Fan 0003, Sen Chen 0001, Yang Liu 0003 |
SANER | 5 |
| 2020 | How Are Performance Issues Caused and Resolved?-An Empirical Study from a Design PerspectiveabstractEmpirical experience regarding how real-life performance issues are caused and resolved can provide valuable insights for practitioners to effectively and efficiently prevent, detect, and fix performance issues. Prior work shows that most performance issues have their roots in poor architectural decisions. This paper contributes a large scale empirical study of 192 real-life performance issues, with an emphasis on software design. First, this paper contributes a holistic view of eight common root causes and typical resolutions that recur in different projects, and surveyed existing literature, in particular, tools, that can detect and fix each type of performance issue. Second, this study is first-of-its-kind to investigate performance issues from a design perspective. In the 192 issues, 33% required design-level optimization, i.e. simultaneously revising a group of related source files for resolving the issues. We reveal four design-level optimization patterns, which have shown different prevalence in resolving different root causes. Finally, this study investigated the Return on Investment for addressing performance issues, to help practitioners choose between localized or design-level optimization resolutions, and to prioritize issues due to different root causes. Lu Xiao 0001, Xiao Wang 0030, Lei Sun 0013, Bihuan Chen 0001, Yang Liu 0003, Andre B. Bondi |
ICPE | 6 |
| 2020 | Automated synthesis of local time requirement for service compositionabstractService composition aims at achieving a business goal by composing existing service-based applications or components. The response time of a service is crucial, especially in time-critical business environments, which is often stated as a clause in service-level agreements between service providers and service users. To meet the guaranteed response time requirement of a composite service, it is important to select a feasible set of component services such that their response time will collectively satisfy the response time requirement of the composite service. In this work, we use the BPEL modeling language that aims at specifying Web services. We extend it with timing parameters and equip it with a formal semantics. Then, we propose a fully automated approach to synthesize the response time requirement of component services modeled using BPEL, in the form of a constraint on the local response times. The synthesized requirement will guarantee the satisfaction of the global response time requirement, statically or dynamically. We implemented our work into a tool, Selamat and performed several experiments to evaluate the validity of our approach. Étienne André 0001, Tian Huat Tan, Manman Chen, Shuang Liu 0007, Jun Sun 0001, Yang Liu 0003, Jin Song Dong 0001 |
Softw. Syst. Model. | 6 |
| 2020 | BBB-CFI: Lightweight CFI Approach Against Code-Reuse Attacks Using Basic Block InformationabstractCode-reuse attack is a concrete threat to computing systems because it can evade conventional security defenses. Control flow integrity (CFI) is proposed to repel this threat. However, former implementations of CFI suffer from two major drawbacks: complex offline processing on programs and high overheads at runtime. Therefore, it is impractical for performance-constrained devices to adopt the technology, leaving them vulnerable to exploitation. In this article, we develop a cross-layer approach named basic-block-boundary-based control flow integrity (BBB-CFI) to minimize the overheads of both offline analysis and runtime checking. Our approach employs basic block information inside the binary code and read-only data to enforce CFI. We identify a key binary-level property called basic block boundary , and based on it we propose the code-inspired method where short code sequences can endorse a control flow transition. Our solution enables quick application launching because it does not require control flow graph construction at the offline stage. We only demand a lightweight analysis on read-only data and a small amount of code of the application. According to the experiments, our approach incurs a negligible 0.11% runtime performance overhead with a minor processor extension, whereas it achieves an order of magnitude speedup in pre-preprocessing compared to a baseline approach. Without control flow analysis or recompilation, BBB-CFI still effectively reduces 90% of the attack surface in terms of gadget numbers. Besides this, we show that the Turing-completeness in the libc is unsustainable. Our approach also demonstrates high applicability to many programs, and it is capable of protecting striped binaries. Wenjian He, Sanjeev Das, Wei Zhang 0012, Yang Liu 0003 |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2020 | Information Theoretical Analysis of Unfair Rating Attacks Under SubjectivityabstractRatings provided by advisors can help an advisee to make decisions, e.g., which seller to select in e-commerce. Unfair rating attacks-where dishonest ratings are provided to mislead the advisee-impact the accuracy of decision making. Current literature focuses on specific classes of unfair rating attacks, which does not provide a complete picture of the attacks. We provide the first formal study that addresses all attack behavior that is possible within a given system. We propose a probabilistic modeling of rating behavior, and apply information theory to quantitatively measure the impact of attacks. In particular, we can identify the attack with the worst impact. In the simple case, honest advisors report the truth straightforwardly, and attackers rate strategically. In real systems, the truth (or an advisor's view on it) may be subjective, making even honest ratings inaccurate. Although there exist methods to deal with subjective ratings, whether subjectivity influences the effect of unfair rating attacks was an open question. We discover that subjectivity decreases the robustness against attacks. Dongxia Wang 0002, Tim Muller, Jie Zhang 0002, Yang Liu 0003 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2020 | Large-Scale Empirical Studies on Effort-Aware Security Vulnerability Prediction MethodsabstractSecurity vulnerability prediction (SVP) can identify potential vulnerable modules in advance and then help developers to allocate most of the test resources to these modules. To evaluate the performance of different SVP methods, we should take the security audit and code inspection into account and then consider effort-aware performance measures (such as ACC and Popt). However, to the best of our knowledge, the effectiveness of different SVP methods has not been thoroughly investigated in terms of effort-aware performance measures. In this article, we consider 48 different SVP methods, of which 36 are supervised methods and 12 are unsupervised methods. For the supervised methods, we consider 34 software-metric-based methods and two text-mining-based methods. For the software-metric-based methods, in addition to a large number of classification methods, we also consider four state-of-the-art methods (i.e., EALR, OneWay, CBS, and MULTI) proposed in recent effort-aware just-in-time defect prediction studies. For text-mining-based methods, we consider the Bag-of-Word model and the term-frequency-inverse-document-frequency model. For the unsupervised methods, all the modules are ranked in the ascendent order based on a specific metric. Since 12 software metrics are considered when measuring extracted modules, there are 12 different unsupervised methods. To the best of our knowledge, over 40 SVP methods have not been considered in previous SVP studies. In our large-scale empirical studies, we use three real open-source web applications written in PHP as benchmark. These three web applications include 3466 modules and 223 vulnerabilities in total. We evaluate these SVP methods both in the within-project SVP scenario and the cross-project SVP scenario. Empirical results show that two unsupervised methods [i.e., lines of code (LOC) and Halstead's volume (HV)] and four recently proposed state-of-the-art supervised methods (i.e., MULTI, OneWay, CBS, and EALR) can achieve better performance than the other methods in terms of effort-aware performance measures. Then, we analyze the reasons why these six methods can achieve better performance. For example, when using 20% of the entire efforts, we find that these six methods always require more modules to be inspected, especially for unsupervised methods LOC and HV. Finally, from the view of practical vulnerability localization, we find that all the unsupervised methods and the OneWay method have high false alarms before finding the first vulnerable module. This may have an impact on developers' confidence and tolerance, and supervised methods (especially MULTI and text-mining-based methods) are preferred. Xiang Chen 0005, Yingquan Zhao, Zhanqi Cui, Guozhu Meng, Yang Liu 0003 |
IEEE Trans. Reliab. | 5 |
| 2020 | METTLE: A METamorphic Testing Approach to Assessing and Validating Unsupervised Machine Learning SystemsabstractUnsupervised machine learning is the training of an artificial intelligence system using information that is neither classified nor labeled, with a view to modeling the underlying structure or distribution in a dataset. Since unsupervised machine learning systems are widely used in many real-world applications, assessing the appropriateness of these systems and validating their implementations with respect to individual users' requirements and specific application scenarios/contexts are indisputably two important tasks. Such assessments and validation tasks, however, are fairly challenging due to the absence of a priori knowledge of the data. In view of this challenge, in this article, we develop a METamorphic Testing approach to assessing and validating unsupervised machine LEarning systems, abbreviated as mettle. Our approach provides a new way to unveil the (possibly latent) characteristics of various machine learning systems, by explicitly considering the specific expectations and requirements of these systems from individual users' perspectives. To support mettle, we have further formulated 11 generic metamorphic relations (MRs), covering users' generally expected characteristics that should be possessed by machine learning systems. We have performed an experiment and a user evaluation study to evaluate the viability and effectiveness of mettle. Our experiment and user evaluation study have shown that, guided by user-defined MR-based adequacy criteria, end users are able to assess, validate, and select appropriate clustering systems in accordance with their own specific needs. Our investigation has also yielded insightful understanding and interpretation of the behavior of the machine learning systems from an end-user software engineering's perspective, rather than a designer's or implementor's perspective, who normally adopts a theoretical approach. Xiaoyuan Xie, Zhiyi Zhang 0005, Tsong Yueh Chen, Yang Liu 0003, Pak-Lok Poon, Baowen Xu |
IEEE Trans. Reliab. | 4 |
| 2019 | A Parametric Rely-Guarantee Reasoning Framework for Concurrent Reactive Systems
Yongwang Zhao, David Sanán, Fuyuan Zhang, Yang Liu 0003 |
FM | 4 |
| 2019 | MobiDroid: A Performance-Sensitive Malware Detection System on Mobile PlatformabstractCurrently, Android malware detection is mostly performed on the server side against the increasing number of Android malware. Powerful computing resource gives more exhaustive protection for Android markets than maintaining detection by a single user in many cases. However, apart from the Android apps provided by the official market (i.e., Google Play Store), apps from unofficial markets and third-party resources are always causing a serious security threat to end-users. Meanwhile, it is a time-consuming task if the app is downloaded first and then uploaded to the server side for detection because the network transmission has a lot of overhead. In addition, the uploading process also suffers from the threat of attackers. Consequently, a last line of defense on Android devices is necessary and much-needed. To address these problems, in this paper, we propose an effective Android malware detection system, MobiDroid, leveraging deep learning to provide a real-time secure and fast response environment on Android devices. Although a deep learning-based approach can be maintained on server side efficiently for detecting Android malware, deep learning models cannot be directly deployed and executed on Android devices due to various performance limitations such as computation power, memory size, and energy. Therefore, we evaluate and investigate the different performances with various feature categories, and further provide an effective solution to detect malware on Android devices. The proposed detection system on Android devices in this paper can serve as a starting point for further study of this important area. Sen Chen 0001, Xiaofei Xie, Lei Ma 0003, Guozhu Meng, Yang Liu 0003, Shangwei Lin 0001 |
ICECCS | 6 |
| 2019 | A Formally Verified Buddy Memory Allocation ModelabstractBuddy memory allocation algorithms are widely adopted by various memory management systems for managing memory layouts. Rigorous mathematical proofs provide strong assurance to improve the confidence on the reliability of a memory management system. In this paper, we model and formally verify, in the interactive theorem prover Isabelle/HOL, a buddy memory allocation model, which preserves functional correctness and security properties. Firstly, we construct a specification consisting of operations to allocate and dispose memory blocks according to a buddy memory allocation algorithm. Then we verify that the specification preserves key invariants over the memory to guarantee functional correctness of the algorithm. Finally, we verify that the specification also preserves the integrity of the memory. Therefore, they do not affect other memory blocks previously allocated. Ke Jiang 0001, David Sanán, Yongwang Zhao, Shuanglong Kan, Yang Liu 0003 |
ICECCS | 5 |
| 2019 | Safe Inputs Approximation for Black-Box SystemsabstractGiven a family of independent and identically distributed samples extracted from the input region and their corresponding outputs, in this paper we propose a method to under-approximate the set of safe inputs that lead the black-box system to respect a given safety specification. Our method falls within the framework of probably approximately correct (PAC) learning. The computed under-approximation comes with statistical soundness provided by the underlying PAC learning process. Such a set, which we call a PAC under-approximation, is obtained by computing a PAC model of the black-box system with respect to the specified safety specification. In our method, the PAC model is computed based on the scenario approach, which encodes as a linear program. The linear program is constructed based on the given family of input samples and their corresponding outputs. The size of the linear program does not depend on the dimensions of the state space of the black-box system, thus providing scalability. Moreover, the linear program does not depend on the internal mechanism of the black-box system, thus being applicable to systems that existing methods are not capable of dealing with. Some case studies demonstrate these properties, general performance and usefulness of our approach. Yang Liu 0003, Lei Ma 0003, Xiyue Zhang 0001, Meng Sun 0002, Xiaofei Xie |
ICECCS | 2 |
| 2019 | Secure Deep Learning Engineering: A Road Towards Quality Assurance of Intelligent Systems
Yang Liu 0003, Lei Ma 0003, Jianjun Zhao 0001 |
ICFEM | 1 |
| 2019 | A Performance-Sensitive Malware Detection System on Mobile Platform
Yang Liu 0003, Shangwei Lin 0001 |
ICFEM | 2 |
| 2019 | StoryDroid: automated generation of storyboard for Android appsabstractMobile apps are now ubiquitous. Before developing a new app, the development team usually endeavors painstaking efforts to review many existing apps with similar purposes. The review process is crucial in the sense that it reduces market risks and provides inspiration for app development. However, manual exploration of hundreds of existing apps by different roles (e.g., product manager, UI/UX designer, developer) in a development team can be ineffective. For example, it is difficult to completely explore all the functionalities of the app in a short period of time. Inspired by the conception of storyboard in movie production, we propose a system, StoryDroid, to automatically generate the storyboard for Android apps, and assist different roles to review apps efficiently. Specifically, StoryDroid extracts the activity transition graph and leverages static analysis techniques to render UI pages to visualize the storyboard with the rendered pages. The mapping relations between UI pages and the corresponding implementation code (e.g., layout code, activity code, and method hierarchy) are also provided to users. Our comprehensive experiments unveil that StoryDroid is effective and indeed useful to assist app development. The outputs of StoryDroid enable several potential applications, such as the recommendation of UI design and layout code. Sen Chen 0001, Lingling Fan 0003, Chunyang Chen 0001, Ting Su 0001, Wenhe Li, Yang Liu 0003, Lihua Xu |
ICSE | 6 |
| 2019 | Leopard: identifying vulnerable code for vulnerability assessment through program metricsabstractIdentifying potentially vulnerable locations in a code base is critical as a pre-step for effective vulnerability assessment; i.e., it can greatly help security experts put their time and effort to where it is needed most. Metric-based and pattern-based methods have been presented for identifying vulnerable code. The former relies on machine learning and cannot work well due to the severe imbalance between non-vulnerable and vulnerable code or lack of features to characterize vulnerabilities. The latter needs the prior knowledge of known vulnerabilities and can only identify similar but not new types of vulnerabilities. In this paper, we propose and implement a generic, lightweight and extensible framework, LEOPARD, to identify potentially vulnerable functions through program metrics. LEOPARD requires no prior knowledge about known vulnerabilities. It has two steps by combining two sets of systematically derived metrics. First, it uses complexity metrics to group the functions in a target application into a set of bins. Then, it uses vulnerability metrics to rank the functions in each bin and identifies the top ones as potentially vulnerable. Our experimental results on 11 real-world projects have demonstrated that, LEOPARD can cover 74.0% of vulnerable functions by identifying 20% of functions as vulnerable and outperform machine learning-based and static analysis-based techniques. We further propose three applications of LEOPARD for manual code review and fuzzing, through which we discovered 22 new bugs in real applications like PHP, radare2 and FFmpeg, and eight of them are new vulnerabilities. Xiaoning Du 0001, Bihuan Chen 0001, Yuekang Li, Jianmin Guo, Yaqin Zhou, Yang Liu 0003, Yu Jiang 0001 |
ICSE | 6 |
| 2019 | Superion: grammar-aware greybox fuzzingabstractIn recent years, coverage-based greybox fuzzing has proven itself to be one of the most effective techniques for finding security bugs in practice. Particularly, American Fuzzy Lop (AFL for short) is deemed to be a great success in fuzzing relatively simple test inputs. Unfortunately, when it meets structured test inputs such as XML and JavaScript, those grammar-blind trimming and mutation strategies in AFL hinder the effectiveness and efficiency. To this end, we propose a grammar-aware coverage-based greybox fuzzing approach to fuzz programs that process structured inputs. Given the grammar (which is often publicly available) of test inputs, we introduce a grammar-aware trimming strategy to trim test inputs at the tree level using the abstract syntax trees (ASTs) of parsed test inputs. Further, we introduce two grammar-aware mutation strategies (i.e., enhanced dictionary-based mutation and tree-based mutation). Specifically, tree-based mutation works via replacing subtrees using the ASTs of parsed test inputs. Equipped with grammar-awareness, our approach can carry the fuzzing exploration into width and depth. We implemented our approach as an extension to AFL, named Superion; and evaluated the effectiveness of Superion using large- scale programs (i.e., an XML engine libplist and three JavaScript engines WebKit, Jerryscript and ChakraCore). Our results have demonstrated that Superion can improve the code coverage (i.e., 16.7% and 8.8% in line and function coverage) and bug-finding capability (i.e., 34 new bugs, among which we discovered 22 new vulnerabilities with 19 CVEs assigned and 3.2K USD bug bounty rewards received) over AFL and jsfunfuzz. Junjie Wang 0007, Bihuan Chen 0001, Lei Wei 0001, Yang Liu 0003 |
ICSE | 4 |
| 2019 | ReCDroid: automatically reproducing Android application crashes from bug reportsabstractThe large demand of mobile devices creates significant concerns about the quality of mobile applications (apps). Developers heavily rely on bug reports in issue tracking systems to reproduce failures (e.g., crashes). However, the process of crash reproduction is often manually done by developers, making the resolution of bugs inefficient, especially that bug reports are often written in natural language. To improve the productivity of developers in resolving bug reports, in this paper, we introduce a novel approach, called ReCDroid, that can automatically reproduce crashes from bug reports for Android apps. ReCDroid uses a combination of natural language processing (NLP) and dynamic GUI exploration to synthesize event sequences with the goal of reproducing the reported crash. We have evaluated ReCDroid on 51 original bug reports from 33 Android apps. The results show that ReCDroid successfully reproduced 33 crashes (63.5% success rate) directly from the textual description of bug reports. A user study involving 12 participants demonstrates that ReCDroid can improve the productivity of developers when resolving crash bug reports. Yu Zhao 0010, Tingting Yu 0001, Ting Su 0001, Yang Liu 0003, Wei Zheng 0006, Jingzhi Zhang, William G. J. Halfond |
ICSE | 4 |
| 2019 | DiffChaser: Detecting Disagreements for Deep Neural NetworksabstractThe platform migration and customization have become an indispensable process of deep neural network (DNN) development lifecycle. A high-precision but complex DNN trained in the cloud on massive data and powerful GPUs often goes through an optimization phase (e.g, quantization, compression) before deployment to a target device (e.g, mobile device). A test set that effectively uncovers the disagreements of a DNN and its optimized variant provides certain feedback to debug and further enhance the optimization procedure. However, the minor inconsistency between a DNN and its optimized version is often hard to detect and easily bypasses the original test set. This paper proposes DiffChaser, an automated black-box testing framework to detect untargeted/targeted disagreements between version variants of a DNN. We demonstrate 1) its effectiveness by comparing with the state-of-the-art techniques, and 2) its usefulness in real-world DNN product deployment involved with quantization and optimization. Xiaofei Xie, Lei Ma 0003, Haijun Wang 0002, Yuekang Li, Yang Liu 0003, Xiaohong Li 0001 |
IJCAI | 5 |
| 2019 | DeepHunter: a coverage-guided fuzz testing framework for deep neural networksabstractThe past decade has seen the great potential of applying deep neural network (DNN) based software to safety-critical scenarios, such as autonomous driving. Similar to traditional software, DNNs could exhibit incorrect behaviors, caused by hidden defects, leading to severe accidents and losses. In this paper, we propose DeepHunter, a coverage-guided fuzz testing framework for detecting potential defects of general-purpose DNNs. To this end, we first propose a metamorphic mutation strategy to generate new semantically preserved tests, and leverage multiple extensible coverage criteria as feedback to guide the test generation. We further propose a seed selection strategy that combines both diversity-based and recency-based seed selection. We implement and incorporate 5 existing testing criteria and 4 seed selection strategies in DeepHunter. Large-scale experiments demonstrate that (1) our metamorphic mutation strategy is useful to generate new valid tests with the same semantics as the original seed, by up to a 98% validity ratio; (2) the diversity-based seed selection generally weighs more than recency-based seed selection in boosting the coverage and in detecting defects; (3) DeepHunter outperforms the state of the arts by coverage as well as the quantity and diversity of defects identified; (4) guided by corner-region based criteria, DeepHunter is useful to capture defects during the DNN quantization for platform migration. Xiaofei Xie, Lei Ma 0003, Felix Juefei-Xu, Minhui Xue 0001, Hongxu Chen 0001, Yang Liu 0003, Jianjun Zhao 0001, Bo Li 0026, Jianxiong Yin, Simon See |
ISSTA | 6 |
| 2019 | A Quantitative Analysis Framework for Recurrent Neural NetworkabstractRecurrent neural network (RNN) has achieved great success in processing sequential inputs for applications such as automatic speech recognition, natural language processing and machine translation. However, quality and reliability issues of RNNs make them vulnerable to adversarial attacks and hinder their deployment in real-world applications. In this paper, we propose a quantitative analysis framework - DeepStellar - to pave the way for effective quality and security analysis of software systems powered by RNNs. DeepStellar is generic to handle various RNN architectures, including LSTM and GRU, scalable to work on industrial-grade RNN models, and extensible to develop customized analyzers and tools. We demonstrated that, with DeepStellar, users are able to design efficient test generation tools, and develop effective adversarial sample detectors. We tested the developed applications on three real RNN models, including speech recognition and image classification. DeepStellar outperforms existing approaches three hundred times in generating defect-triggering tests and achieves 97% accuracy in detecting adversarial attacks. A video demonstration which shows the main features of DeepStellar is available at: https://sites.google.com/view/deepstellar/tool-demo. Xiaoning Du 0001, Xiaofei Xie, Yi Li 0008, Lei Ma 0003, Yang Liu 0003, Jianjun Zhao 0001 |
ASE | 5 |
| 2019 | An Empirical Study Towards Characterizing Deep Learning Development and Deployment Across Different Frameworks and PlatformsabstractDeep Learning (DL) has recently achieved tremendous success. A variety of DL frameworks and platforms play a key role to catalyze such progress. However, the differences in architecture designs and implementations of existing frameworks and platforms bring new challenges for DL software development and deployment. Till now, there is no study on how various mainstream frameworks and platforms influence both DL software development and deployment in practice. To fill this gap, we take the first step towards understanding how the most widely-used DL frameworks and platforms support the DL software development and deployment. We conduct a systematic study on these frameworks and platforms by using two types of DNN architectures and three popular datasets. (1) For development process, we investigate the prediction accuracy under the same runtime training configuration or same model weights/biases. We also study the adversarial robustness of trained models by leveraging the existing adversarial attack techniques. The experimental results show that the computing differences across frameworks could result in an obvious prediction accuracy decline, which should draw the attention of DL developers. (2) For deployment process, we investigate the prediction accuracy and performance (refers to time cost and memory consumption) when the trained models are migrated/quantized from PC to real mobile devices and web browsers. The DL platform study unveils that the migration and quantization still suffer from compatibility and reliability issues. Meanwhile, we find several DL software bugs by using the results as a benchmark. We further validate the results through bug confirmation from stakeholders and industrial positive feedback to highlight the implications of our study. Through our study, we summarize practical guidelines, identify challenges and pinpoint new research directions, such as understanding the characteristics of DL frameworks and platforms, avoiding compatibility and reliability issues, detecting DL software bugs, and reducing time cost and memory consumption towards developing and deploying high quality DL systems effectively. Sen Chen 0001, Xiaofei Xie, Lei Ma 0003, Yang Liu 0003, Jianjun Zhao 0001, Xiaohong Li 0001 |
ASE | 7 |
| 2019 | DeepMutation++: A Mutation Testing Framework for Deep Learning SystemsabstractDeep neural networks (DNNs) are increasingly expanding their real-world applications across domains, e.g., image processing, speech recognition and natural language processing. However, there is still limited tool support for DNN testing in terms of test data quality and model robustness. In this paper, we introduce a mutation testing-based tool for DNNs, DeepMutation++, which facilitates the DNN quality evaluation, supporting both feed-forward neural networks (FNNs) and stateful recurrent neural networks (RNNs). It not only enables to statically analyze the robustness of a DNN model against the input as a whole, but also allows to identify the vulnerable segments of a sequential input (e.g. audio input) by runtime analysis. It is worth noting that DeepMutation++ specially features the support of RNNs mutation testing. The tool demo video can be found on the project website https://sites.google.com/view/deepmutationpp. Lei Ma 0003, Xiaofei Xie, Yang Liu 0003, Jianjun Zhao 0001 |
ASE | 5 |