VLDB 2026 Research / reviewers in the wild / expert
Yu Li 0007
dblp:34/2997-7
· DBLP profile ↗
39ranked-venue papers
8as first author
32since 2021 · last 2026
0000-0001-9122-5923ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 2 first-author · 19 since 2021Systems, architecture and hardware · 8 · 3 first-author · 5 since 2021Security and privacy · 7 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Toward efficient testing of graph neural networks via test input prioritization
Lichen Yang, Qiang Wang 0001, Zhonghao Yang 0003, Daojing He, Yu Li 0007 |
Autom. Softw. Eng. | 5 |
| 2026 | SPLAT: Revisiting Latency Attack on Dynamic Neural NetworksabstractDynamic deep neural networks, particularly multi-exit networks, are increasingly recognized for their efficiency in edge-cloud scenarios. However, they are vulnerable to latency attacks that can degrade performance by increasing computation time. Current attack strategies often require white-box access to the model or lead to significant drops in inference accuracy, making them easily detectable. This paper introduces SPLAT, a novel approach for executing stealthy and practical latency attacks on dynamic multi-exit models under black-box conditions. SPLAT employs a two-stage mechanism: the first stage generates coarse-grained attack inputs using a functional surrogate model, while the second stage refines these perturbations through an efficient query strategy to enhance stealthiness and effectiveness. Extensive experiments validate that SPLAT significantly outperforms existing methods across various models and datasets. Yu Li 0007, Biao Huang 0016, Jinyin Hu, Cheng Zhuo |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2026 | VLM-Guard: Defending Jailbreaks by Monitoring Only Hundreds of Safety-Critical Neurons
Jinyin Hu, Jiawei Zhou 0013, Minshan Xie, Zhonghao Yang 0003, Jing Li 0034, Huadi Zheng, Jie Shi 0005, Daojing He, Yu Li 0007 |
IEEE Trans. Inf. Forensics Secur. | 9 |
| 2026 | A Novel Approach to Reducing Testing Costs and Minimizing Defect Escapes Using Dynamic Neighborhood Range and Shapley ValuesabstractWafer acceptance testing (WAT) is a process that is used to assess the quality and reliability of manufactured wafers. This technique for the early detection and screening of chips allows for improvements in their reliability and performance during the manufacture of semiconductor devices. The automatic test equipment (ATE) used for processing millions of wafers is susceptible to a number of issues, including the absence of data values, the presence of redundant parameters, and categorical imbalance. These issues increase the cost of data processing and impede an investigation into the relationship between WAT and feature diagnostics. In this study, we propose a method with a low test escape rate based on a multi-objective optimization algorithm to reduce the cost of testing and minimize the number of defective dice that go undetected. The proposed method retains outliers, dynamically selects the range of the neighborhood to reduce the cost of testing, and uses Shapley values to analyze a WAT dataset to determine the importance of features of the data. The multi-objective optimization algorithm ranks features by their importance and applies an adaptive method to eliminate features with a low overall correlation, thereby reducing the risk that defective dice are undetected. Tianming Ni, Wangsheng Rui, Cheng Zhuo, Yu Li 0007, Xiaoqing Wen, Mu Nie |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2026 | TNT: A Large-Scale P2P Botnet Detection Framework via Communication Topology and Network TrafficabstractWith the booming development of embedded systems and mobile networks, the attack surfaces of botnets are broadened and amplified. Especially recent adversaries tend to leverage peer-to-peer (P2P) manner propagation to construct large-scale botnets because P2P-based schemes eliminate single points of failure. Over the past few decades, the research and industry communities have proposed a variety of solutions to detect botnets, which mainly involve communication topology identification and network traffic analysis. Yet, coping with the large-scale P2P botnets, the former suffer topology indistinguishability, and the latter struggles under massive background traffic. In this paper, we present$\textsf {TNT}$, a large-scale P2P botnet detection framework via communication topology and network traffic. As its core,$\textsf {TNT}$is powered by three tightly-coupled components:$\textsf {(i)}$$\textsf {tScouter}$is responsible for profiling the communication topology;$\textsf {(ii)}$$\textsf {tCommander}$plans the strategy for node inspection; and$\textsf {(iii)}$$\textsf {tPatroller}$investigates the traffic of the corresponding node. Taken together,$\textsf {TNT}$advances the trade-off between detection accuracy (enhance topology-based results via traffic analysis) and overhead (only check part of node traffic according to the planning). Based on 42 groups of combinations involving 6 types of botnets and 7 legitimate P2P traffic, we perform extensive evaluation and demonstrate that$\textsf {TNT}$realizes outstanding detection performance,e.g.,after checking ~20K nodes, achieve ~99.9% accuracy for a communication graph (including >140K nodes). In addition, we develop the expansion experiments in terms of heterogeneous nodes and accuracy loss, as well as provide deep insights into interpretability from the aspect of the attribution matrix. Ziming Zhao 0008, Zhaoxuan Li, Tingting Li 0004, Yu Li 0007, Qiang Xu 0001, Fan Zhang 0010 |
IEEE Trans. Netw. | 5 |
| 2025 | DF-MIA: A Distribution-Free Membership Inference Attack on Fine-Tuned Large Language ModelsabstractMembership Inference Attack (MIA) aims to determine if a specific sample is present in the training dataset of a target machine learning model. Previous MIAs against fine-tuned Large Language Models (LLMs) either fail to address the unique challenges in the fine-tuned setting or rely on strong assumption of the training data distribution. This paper proposes a distribution-free MIA framework tailored for fine-tuned LLMs, named DF-MIA. We recognize that samples await to test can serve as a valuable reference dataset for fine-tuning reference models. By enhancing the signals of non-member samples within this reference dataset, we can achieve a more reliable and practical calibration of probabilities, improving the differentiation between members and non-members. Leveraging these insights, we have developed a two-stage framework that employs specially designed data augmentation and perturbation techniques to prioritize the significance of non-members and mitigate the influence of potential members within the reference dataset. We evaluate our method on three representative LLM models ranging from 1B to 8B on three datasets. The results demonstrate that the DF-MIA significantly enhances the performance of MIA. Zhiheng Huang, Yannan Liu, Daojing He, Yu Li 0007 |
AAAI | 4 |
| 2025 | MTSA: Multi-turn Safety Alignment for LLMs through Multi-round Red-teamingabstractThe proliferation of jailbreak attacks against large language models (LLMs) highlights the need for robust security measures.However, in multi-round dialogues, malicious intentions may be hidden in interactions, leading LLMs to be more prone to produce harmful responses.In this paper, we propose the Multi-Turn Safety Alignment (MTSA) framework, to address the challenge of securing LLMs in multi-round interactions.It consists of two stages: In the thought-guided attack learning stage, the redteam model learns about thought-guided multiround jailbreak attacks to generate adversarial prompts.In the adversarial iterative optimization stage, the red-team model and the target model continuously improve their respective capabilities in interaction.Furthermore, we introduce a multi-turn reinforcement learning algorithm based on future rewards to enhance the robustness of safety alignment.Experimental results show that the red-team model exhibits state-of-the-art attack capabilities, while the target model significantly improves its performance on safety benchmarks. Weiyang Guo, Jing Li 0034, Wenya Wang 0001, Yu Li 0007, Daojing He, Jun Yu 0002, Min Zhang 0005 |
ACL (1) | 4 |
| 2025 | Safety Alignment via Constrained Knowledge UnlearningabstractZesheng Shi, Yucheng Zhou, Jing Li, Yuxin Jin, Yu Li, Daojing He, Fangming Liu, Saleh Alharbi, Jun Yu, Min Zhang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Zesheng Shi, Yucheng Zhou 0001, Jing Li 0034, Yu Li 0007, Daojing He, Fangming Liu, Saleh Alharbi, Jun Yu 0002, Min Zhang 0005 |
ACL (1) | 5 |
| 2025 | AIR: Complex Instruction Generation via Automatic Iterative RefinementabstractWith the development of large language models, their ability to follow simple instructions has significantly improved.However, adhering to complex instructions remains a major challenge.Current approaches to generating complex instructions are often irrelevant to the current instruction requirements or suffer from limited scalability and diversity.Moreover, methods such as back-translation, while effective for simple instruction generation, fail to leverage the rich knowledge and formatting in human written documents.In this paper, we propose a novel Automatic Iterative Refinement (AIR) framework to generate complex instructions with constraints, which not only better reflects the requirements of real scenarios but also significantly enhances LLMs' ability to follow complex instructions.The AIR framework consists of two stages: 1) Generate an initial instruction from a document; 2) Iteratively refine instructions with LLM-as-judge guidance by comparing the model's output with the document to incorporate valuable constraints.Finally, we construct the AIR-10K dataset with 10K complex instructions and demonstrate that instructions generated with our approach significantly improve the model's ability to follow complex instructions, outperforming existing methods for instruction generation 1 . Model Answer Refined D Model Answer Check Constraints Identify ConstraintsI: Write a casual review of a water-proof camera. Yancheng He, Yu Li 0007, Hui Huang 0021, Chengwei Hu, Wenbo Su, Bo Zheng 0007 |
EMNLP | 3 |
| 2025 | One Model Transfer to All: On Robust Jailbreak Prompts Generation against LLMsabstractSafety alignment in large language models (LLMs) is increasingly compromised by jailbreak attacks, which can manipulate these models to generate harmful or unintended content. Investigating these attacks is crucial for uncovering model vulnerabilities. However, many existing jailbreak strategies fail to keep pace with the rapid development of defense mechanisms, such as defensive suffixes, rendering them ineffective against defended models. To tackle this issue, we introduce a novel attack method called ArrAttack, specifically designed to target defended LLMs. ArrAttack automatically generates robust jailbreak prompts capable of bypassing various defense measures. This capability is supported by a universal robustness judgment model that, once trained, can perform robustness evaluation for any target model with a wide variety of defenses. By leveraging this model, we can rapidly develop a robust jailbreak prompt generator that efficiently converts malicious input prompts into effective attacks. Extensive evaluations reveal that ArrAttack significantly outperforms existing attack strategies, demonstrating strong transferability across both white-box and black-box models, including GPT-4 and Claude-3. Our work bridges the gap between jailbreak attacks and defenses, providing a fresh perspective on generating robust jailbreak prompts. Linbao Li, Yannan Liu, Daojing He, Yu Li 0007 |
ICLR | 4 |
| 2025 | Function-to-Style Guidance of LLMs for Code TranslationabstractLarge language models (LLMs) have made significant strides in code translation tasks. However, ensuring both the correctness and readability of translated code remains a challenge, limiting their effective adoption in real-world software development. In this work, we propose F2STrans, a function-to-style guiding paradigm designed to progressively improve the performance of LLMs in code translation. Our approach comprises two key stages: (1) Functional learning, which optimizes translation correctness using high-quality source-target code pairs mined from online programming platforms, and (2) Style learning, which improves translation readability by incorporating both positive and negative style examples. Additionally, we introduce a novel code translation benchmark that includes up-to-date source code, extensive test cases, and manually annotated ground-truth translations, enabling comprehensive functional and stylistic evaluations. Experiments on both our new benchmark and existing datasets demonstrate that our approach significantly improves code translation performance. Notably, our approach enables Qwen-1.5B to outperform prompt-enhanced Qwen-32B and GPT-4 on average across 20 diverse code translation scenarios. Longhui Zhang, Bin Wang 0004, Hao Yang 0007, Meishan Zhang, Yu Li 0007, Jing Li 0034, Jun Yu 0002, Min Zhang 0005 |
ICML | 8 |
| 2025 | FDTest: Prioritizing Test Inputs for Object Detection Models via Foundation Model ExploitationabstractTesting Deep Neural Networks (DNNs) often incurs high labeling costs due to the need for extensive labeled data. Hence, it is critical to strategically prioritize test inputs that can uncover more model errors for efficient labeling. Existing test input prioritization methods focus on image classification. However, these methods may not be directly applicable to object detection (OD) models, as testing of OD models confronts unique challenges involving complex model errors related to object localization and counting. To address the above challenges, in this paper, we introduce FDTest, a black-box test input prioritization framework for OD models that incorporates foundation models trained on diverse datasets to provide supplementary information. Guided by the foundation model, we categorize the target model’s predictions as reliable and suspicious objects and conduct confidence calibration on them to improve the estimation of wrongly detected objects, i.e., false positives (FPs). Additionally, we use the foundation model to supply missed objects and cross-verify them to improve the estimation of missed objects, i.e., false negatives (FNs). Then, we prioritize test images according to the total estimated number of FPs and FNs. Experiments on two standard OD benchmarks demonstrate the superiority of FDTest in test input prioritization for OD models. Qiuxia Lai, Yu Li 0007 |
IJCNN | 3 |
| 2025 | Domain-RAG: Retrieval-Guided Compositional Image Generation for Cross-Domain Few-Shot Object DetectionabstractCross-Domain Few-Shot Object Detection (CD-FSOD) aims to detect novel objects with only a handful of labeled samples from previously unseen domains. While data augmentation and generative methods have shown promise in few-shot learning, their effectiveness for CD-FSOD remains unclear due to the need for both visual realism and domain alignment. Existing strategies, such as copy-paste augmentation and text-to-image generation, often fail to preserve the correct object category or produce backgrounds coherent with the target domain, making them non-trivial to apply directly to CD-FSOD. To address these challenges, we propose Domain-RAG, a training-free, retrieval-guided compositional image generation framework tailored for CD-FSOD. Domain-RAG consists of three stages: domain-aware background retrieval, domain-guided background generation, and foreground-background composition. Specifically, the input image is first decomposed into foreground and background regions. We then retrieve semantically and stylistically similar images to guide a generative model in synthesizing a new background, conditioned on both the original and retrieved contexts. Finally, the preserved foreground is composed with the newly generated domain-aligned background to form the generated image. Without requiring any additional supervision or training, Domain-RAG produces high-quality, domain-consistent samples across diverse tasks, including CD-FSOD, remote sensing FSOD, and camouflaged FSOD. Extensive experiments show consistent improvements over strong baselines and establish new state-of-the-art results. Codes will be released upon acceptance.The source code and instructions are available at https://github.com/LiYu0524/Domain-RAG. Yu Li 0007, Xingyu Qiu, Yuqian Fu, Tianwen Qian, Xu Zheng 0002, Danda Pani Paudel, Yanwei Fu 0001, Xuanjing Huang 0001, Luc Van Gool, Yu-Gang Jiang 0001 |
NeurIPS | 1 |
| 2025 | Context-Aware Hierarchical Learning: A Two-Step Paradigm towards Safer LLMsabstractLarge Language Models (LLMs) have emerged as powerful tools for diverse applications. However, their uniform token processing paradigm introduces critical vulnerabilities in instruction handling, particularly when exposed to adversarial scenarios. In this work, we identify and propose a novel class of vulnerabilities, termed Tool-Completion Attack (TCA), which exploits function-calling mechanisms to subvert model behavior. To evaluate LLM robustness against such threats, we introduce the Tool-Completion benchmark, a comprehensive security assessment framework, which reveals that even state-of-the-art models remain susceptible to TCA, with surprisingly high attack success rates. To address these vulnerabilities, we introduce Context-Aware Hierarchical Learning (CAHL), a sophisticated mechanism that dynamically equilibrates semantic comprehension with role-specific instruction constraints. CAHL leverages the contextual correlations between different instruction segments to establish a robust, context-aware instruction hierarchy. Extensive experiments demonstrate that CAHL significantly enhances LLM robustness against both conventional attacks and the proposed TCA, exhibiting strong generalization capabilities in zero-shot evaluations while still preserving model performance on generic tasks. Our code is available at https://github.com/S2AILab/CAHL. Tengyun Ma, Daojing He, Shihao Peng, Yu Li 0007, Shaohui Liu, Zhuotao Tian |
NeurIPS | 5 |
| 2025 | Learning to Watermark: A Selective Watermarking Framework for Large Language Models via Multi-Objective OptimizationabstractThe rapid development of LLMs has raised concerns about their potential misuse, leading to various watermarking schemes that typically offer high detectability.
However, existing watermarking techniques often face trade-off between watermark detectability and generated text quality.
In this paper, we introduce Learning to Watermark (LTW), a novel selective watermarking framework that leverages multi-objective optimization to effectively balance these competing goals.
LTW features a lightweight network that adaptively decides when to apply the watermark by analyzing sentence embeddings, token entropy, and current watermarking ratio.
Training of the network involves two specifically constructed loss functions that guide the model toward Pareto-optimal solutions, thereby harmonizing watermark detectability and text quality.
By integrating LTW with two baseline watermarking methods, our experimental evaluations demonstrate that LTW significantly enhances text quality without compromising detectability.
Our selective watermarking approach offers a new perspective for designing watermarks for LLMs and a way to preserve high text quality for watermarks. The code is publicly available at: https://github.com/fattyray/learning-to-watermark Chenrui Wang, Junyi Shu, Billy Chiu, Yu Li 0007, Saleh Alharbi, Min Zhang 0005, Jing Li 0034 |
NeurIPS | 4 |
| 2025 | SilentStriker: Toward Stealthy Bit-Flip Attacks on Large Language ModelsabstractThe rapid adoption of large language models (LLMs) in critical domains has spurred extensive research into their security issues. While input manipulation attacks (e.g., prompt injection) have been well-studied, Bit-Flip Attacks (BFAs)—which exploit hardware vulnerabilities to corrupt model parameters and cause severe performance degradation—have received far less attention. Existing BFA methods suffer from key limitations: they fail to balance performance degradation and output naturalness, making them prone to discovery. In this paper, we introduce SilentStriker, the first stealthy bit-flip attack against LLMs that effectively degrades task performance while maintaining output naturalness. Our core contribution lies in addressing the challenge of designing effective loss functions for LLMs with variable output length and the vast output space. Unlike prior approaches that rely on output perplexity for attack loss formulation, which in-evidently degrade the output naturalness, we reformulate the attack objective by leveraging key output tokens as targets for suppression, enabling effective joint optimization of attack effectiveness and stealthiness. Additionally, we employ an iterative, progressive search strategy to maximize attack efficacy. Experiments show that SilentStriker significantly outperforms existing baselines, achieving successful attacks without compromising the naturalness of generated text. Qingsong Peng, Jie Shi 0005, Huadi Zheng, Yu Li 0007, Zhuo Chen 0006 |
NeurIPS | 5 |
| 2025 | Toward Robust and Accurate Adversarial Camouflage Generation Against Vehicle DetectorsabstractAdversarial camouflage is a widely used physical attack against vehicle detectors for its superiority in multiview attack performance. One promising approach involves using differentiable neural renderers to facilitate adversarial camouflage optimization through gradient back-propagation. However, existing methods often struggle to capture environmental characteristics during the rendering process or produce adversarial textures that can precisely map to the target vehicle. Moreover, these approaches neglect diverse weather conditions, reducing the efficacy of generated camouflage across varying weather scenarios. To tackle these challenges, we propose a robust and accurate camouflage generation method, namely RAUCA. The core of RAUCA is a novel neural rendering component, End-to-End Neural Renderer Plus (E2E-NRP), which can accurately optimize and project vehicle textures and render images with environmental characteristics such as lighting and weather. In addition, we integrate a multi-weather dataset for camouflage generation, leveraging the E2E-NRP to enhance the attack robustness. Experimental results on six popular object detectors show that RAUCA-final outperforms existing methods in both simulation and real-world settings. Jiawei Zhou 0013, Linye Lyu, Daojing He, Yu Li 0007 |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2025 | ArcGen: Generalizing Neural Backdoor Detection Across Diverse ArchitecturesabstractBackdoor attacks pose a significant threat to the security and reliability of deep learning models. To mitigate such attacks, one promising approach is to learn to extract features from the target model and use these features for backdoor detection. However, we discover that existing learning-based neural backdoor detection methods do not generalize well to new architectures not seen during the learning phase. In this paper, we analyze the root cause of this issue and propose a novel black-box neural backdoor detection method called ARCGEN. Our method aims to obtain architecture-invariant model features, i.e.,aligned features, for effective backdoor detection. Specifically, in contrast to existing methods directly using model outputs as model features, we introduce an additional alignment layer in the feature extraction function to further process these features. This reduces the direct influence of architecture information on the features. Then, we design two alignment losses to train the feature extraction function. These losses explicitly require that features from models with similar backdoor behaviors but different architectures are aligned at both the distribution and sample levels. With these techniques, our method demonstrates up to 42.5% improvements in detection performance (e.g., AUC) on unseen model architectures. This is based on a large-scale evaluation involving 16,896 models trained on diverse datasets, subjected to various backdoor attacks, and utilizing different model architectures. Our code is available at https://github.com/SeRAlab/ArcGen. Zhonghao Yang 0003, Daojing He, Yiming Li 0004, Yu Li 0007 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2024 | RAUCA: A Novel Physical Adversarial Attack on Vehicle Detectors via Robust and Accurate Camouflage GenerationabstractAdversarial camouflage is a widely used physical attack against vehicle detectors for its superiority in multi-view attack performance. One promising approach involves using differentiable neural renderers to facilitate adversarial camouflage optimization through gradient back-propagation. However, existing methods often struggle to capture environmental characteristics during the rendering process or produce adversarial textures that can precisely map to the target vehicle, resulting in suboptimal attack performance. Moreover, these approaches neglect diverse weather conditions, reducing the efficacy of generated camouflage across varying weather scenarios. To tackle these challenges, we propose a robust and accurate camouflage generation method, namely RAUCA. The core of RAUCA is a novel neural rendering component, Neural Renderer Plus (NRP), which can accurately project vehicle textures and render images with environmental characteristics such as lighting and weather. In addition, we integrate a multi-weather dataset for camouflage generation, leveraging the NRP to enhance the attack robustness. Experimental results on six popular object detectors show that RAUCA consistently outperforms existing methods in both simulation and real-world settings. Jiawei Zhou 0013, Linye Lyu, Daojing He, Yu Li 0007 |
ICML | 4 |
| 2024 | Enhancing Flow Embedding Through Trace: A Novel Self-supervised Approach for Encrypted Traffic ClassificationabstractTraffic classification is a crucial task in network security and management. Recent research has shown the effectiveness of deep learning when applied to encrypted traffic classification. However, the reliance of deep learning models on abundant labeled and balanced data poses challenges, particularly in traffic analysis where labeling is costly and imbalanced traffic distribution is common. To tackle this challenge, researchers have proposed self-supervised representation learning. This approach aims to derive universal traffic representations from vast amounts of unlabeled data, reducing the need for extensive labeling in downstream tasks. Current representation learning methods predominantly focus on exploring flow-level information, neglecting valuable trace information crucial for effective flow representation learning. In this study, we introduce SAFE, a self-supervised learning methodology tailored for flow representation learning. SAFE specifically delves into trace-level (i.e., a mixture of correlated flows) information to enhance flow embedding. Moreover, our analysis reveals that existing encrypted traffic datasets often contain numerous invalid samples. SAFE conducts a comprehensive examination of dataset characteristics, filtering out these invalid samples, thereby advancing the field significantly. Extensive experiments demonstrate that SAFE outperforms state-of-the-art methods. Zefei Luo, Yu Li 0007, Shuaishuai Tan, Daojing He |
IJCNN | 2 |
| 2024 | Vector Quantization Prompting for Continual LearningabstractContinual learning requires to overcome catastrophic forgetting when training a single model on a sequence of tasks. Recent top-performing approaches are prompt-based methods that utilize a set of learnable parameters (i.e., prompts) to encode task knowledge, from which appropriate ones are selected to guide the fixed pre-trained model in generating features tailored to a certain task. However, existing methods rely on predicting prompt identities for prompt selection, where the identity prediction process cannot be optimized with task loss. This limitation leads to sub-optimal prompt selection and inadequate adaptation of pre-trained features for a specific task. Previous efforts have tried to address this by directly generating prompts from input queries instead of selecting from a set of candidates. However, these prompts are continuous, which lack sufficient abstraction for task knowledge representation, making them less effective for continual learning. To address these challenges, we propose VQ-Prompt, a prompt-based continual learning method that incorporates Vector Quantization (VQ) into end-to-end training of a set of discrete prompts. In this way, VQ-Prompt can optimize the prompt selection process with task loss and meanwhile achieve effective abstraction of task knowledge for continual learning. Extensive experiments show that VQ-Prompt outperforms state-of-the-art continual learning methods across a variety of benchmarks under the challenging class-incremental setting. Qiuxia Lai, Yu Li 0007, Qiang Xu 0001 |
NeurIPS | 3 |
| 2024 | CNCA: Toward Customizable and Natural Generation of Adversarial Camouflage for Vehicle DetectorsabstractPrior works on physical adversarial camouflage against vehicle detectors mainly focus on the effectiveness and robustness of the attack. The current most successful methods optimize 3D vehicle texture at a pixel level. However, this results in conspicuous and attention-grabbing patterns in the generated camouflage, which humans can easily identify. To address this issue, we propose a Customizable and Natural Camouflage Attack (CNCA) method by leveraging an off-the-shelf pre-trained diffusion model. By sampling the optimal texture image from the diffusion model with a user-specific text prompt, our method can generate natural and customizable adversarial camouflage while maintaining high attack performance. With extensive experiments on the digital and physical worlds and user studies, the results demonstrate that our proposed method can generate significantly more natural-looking camouflage than the state-of-the-art baselines while achieving competitive attack performance. Linye Lyu, Jiawei Zhou 0013, Daojing He, Yu Li 0007 |
NeurIPS | 4 |
| 2024 | Spatial attention for human-centric visual understanding: An Information Bottleneck method
Qiuxia Lai, Yongwei Nie, Yu Li 0007, Hanqiu Sun, Qiang Xu 0001 |
Comput. Vis. Image Underst. | 3 |
| 2024 | Self-Supervised Video Representation Learning via Capturing Semantic Changes Indicated by SaccadesabstractIn this paper, we propose a self-supervised video representation learning (video SSL) method by taking inspiration from cognitive science and neuroscience on human visual perception. Different from previous methods that focus on the inherent properties of videos, we argue that humans learn to perceive the world through the self-awareness of the semantic changes or consistency in the input stimuli in the absence of labels, accompanied by representation reorganization during the post-learning rest periods. To this end, we first exploit the presence of saccades as an indicator of semantic changes in a contrastive learning framework, mimicking self-awareness in human representation learning. The saccades are generated by alternating the fixations following the predicted scanpath. Second, we model the semantic consistency in eye fixation by minimizing the prediction error between the predicted and the true state of another time point. Finally, we incorporate prototypical contrastive learning to reorganize the learned representations to enhance the associations among perceptually similar ones. Compared to previous video SSL solutions, our method can capture finer-grained semantics from video instances and further associate similar ones together. Experiments show that the proposed bio-inspired video SSL method significantly improves the Top-1 video retrieval accuracy on UCF101 and achieves superior performance on downstream tasks such as action recognition under comparable settings. Qiuxia Lai, Ailing Zeng, Ye Wang 0011, Lihong Cao, Yu Li 0007, Qiang Xu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Towards Robust Deep Neural Networks Against Design-Time and Run-Time FailuresabstractDeep Neural Networks (DNNs) have gained widespread adoption, but they also exhibit post-deployment failures that pose risks to property and life. Consequently, enhancing DNN robustness in safety-critical areas is crucial. This paper improves the robustness of DNN models against failures that may arise during either design time or run time. For design failures, we introduce TestRank, an efficient method for identifying DNN design issues. Constrained by test resources, TestRank selects high-quality test cases leveraging intrinsic and contextual attributes of test samples. HybridRepair then offers effective failure repair by selectively annotating failure regions and utilizing semi-supervised learning techniques. To counteract run-time fault injection attacks, we propose D2NN and DeepDyve, which introduce the dual modular redundancy concept to models for protection at the neuron and system levels, respectively. This delicate redundancy achieves lightweight protection for DNN-based systems. Evaluation performed on various image classification datasets demonstrates the effectiveness of our approaches. Yu Li 0007, Qiang Xu 0001 |
ITC | 1 |
| 2023 | HiBug: On Human-Interpretable Model DebugabstractMachine learning models can frequently produce systematic errors on critical subsets (or slices) of data that share common attributes. Discovering and explaining such model bugs is crucial for reliable model deployment. However, existing bug discovery and interpretation methods usually involve heavy human intervention and annotation, which can be cumbersome and have low bug coverage.
In this paper, we propose HiBug, an automated framework for interpretable model debugging. Our approach utilizes large pre-trained models, such as chatGPT, to suggest human-understandable attributes that are related to the targeted computer vision tasks. By leveraging pre-trained vision-language models, we can efficiently identify common visual attributes of underperforming data slices using human-understandable terms. This enables us to uncover rare cases in the training data, identify spurious correlations in the model, and use the interpretable debug results to select or generate new training data for model improvement. Experimental results demonstrate the efficacy of the HiBug framework. Muxi Chen, Yu Li 0007, Qiang Xu 0001 |
NeurIPS | 2 |
| 2022 | HybridRepair: towards annotation-efficient repair for deep learning modelsabstractA well-trained deep learning (DL) model often cannot achieve expected performance after deployment due to the mismatch between the distributions of the training data and the field data in the operational environment. Therefore, repairing DL models is critical, especially when deployed on increasingly larger tasks with shifted distributions. Yu Li 0007, Muxi Chen, Qiang Xu 0001 |
ISSTA | 1 |
| 2022 | What You See is Not What the Network Infers: Detecting Adversarial Examples Based on Semantic Contradiction
Ruiyuan Gao 0001, Yu Li 0007, Qiuxia Lai, Qiang Xu 0001 |
NDSS | 3 |
| 2021 | AppealNet: An Efficient and Highly-Accurate Edge/Cloud Collaborative Architecture for DNN InferenceabstractThis paper presents AppealNet, a novel edge/cloud collaborative architecture that runs deep learning (DL) tasks more efficiently than state-of-the-art solutions. For a given input, AppealNet accurately predicts on-the-fly whether it can be successfully processed by the DL model deployed on the resource-constrained edge device, and if not, appeals to the more powerful DL model deployed at the cloud. This is achieved by employing a two-head neural network architecture that explicitly takes inference difficulty into consideration and optimizes the tradeoff between accuracy and computation/communication cost of the edge/cloud collaborative architecture. Experimental results on several image classification datasets show up to more than 40% energy savings compared to existing techniques without sacrificing accuracy. Min Li 0019, Yu Li 0007, Ye Tian 0010, Li Jiang 0002, Qiang Xu 0001 |
DAC | 2 |
| 2021 | Information Bottleneck Approach to Spatial Attention LearningabstractThe selective visual attention mechanism in the human visual system (HVS) restricts the amount of information to reach visual awareness for perceiving natural scenes, allowing near real-time information processing with limited computational capacity. This kind of selectivity acts as an ‘Information Bottleneck (IB)’, which seeks a trade-off between information compression and predictive accuracy. However, such information constraints are rarely explored in the attention mechanism for deep neural networks (DNNs). In this paper, we propose an IB-inspired spatial attention module for DNN structures built for visual recognition. The module takes as input an intermediate representation of the input image, and outputs a variational 2D attention map that minimizes the mutual information (MI) between the attention-modulated representation and the input, while maximizing the MI between the attention-modulated representation and the task label. To further restrict the information bypassed by the attention map, we quantize the continuous attention scores to a set of learnable anchor values during training. Extensive experiments show that the proposed IB-inspired spatial attention mechanism can yield attention maps that neatly highlight the regions of interest while suppressing backgrounds, and bootstrap standard DNN structures for visual recognition tasks (e.g., image classification, fine-grained recognition, cross-domain classification). The attention maps are interpretable for the decision making of the DNNs as verified in the experiments. Our code is available at this https URL. Qiuxia Lai, Yu Li 0007, Ailing Zeng, Minhao Liu, Hanqiu Sun, Qiang Xu 0001 |
IJCAI | 2 |
| 2021 | TestRank: Bringing Order into Unlabeled Test Instances for Deep Learning TasksabstractDeep learning (DL) systems are notoriously difficult to test and debug due to the lack of correctness proof and the huge test input space to cover. Given the ubiquitous unlabeled test data and high labeling cost, in this paper, we propose a novel test prioritization technique, namely TestRank, which aims at revealing more model failures with less labeling effort. TestRank brings order into the unlabeled test data according to their likelihood of being a failure, i.e., their failure-revealing capabilities. Different from existing solutions, TestRank leverages both intrinsic and contextual attributes of the unlabeled test data when prioritizing them. To be specific, we first build a similarity graph on both unlabeled test samples and labeled samples (e.g., training or previously labeled test samples). Then, we conduct graph-based semi-supervised learning to extract contextual features from the correctness of similar labeled samples. For a particular test instance, the contextual features extracted with the graph neural network and the intrinsic features obtained with the DL model itself are combined to predict its failure-revealing capability. Finally, TestRank prioritizes unlabeled test inputs in descending order of the above probability value. We evaluate TestRank on three popular image classification datasets, and results show that TestRank significantly outperforms existing test prioritization techniques. Yu Li 0007, Min Li 0019, Qiuxia Lai, Yannan Liu, Qiang Xu 0001 |
NeurIPS | 1 |
| 2021 | On Workload-Aware DRAM Failure Prediction in Large-Scale Data CentersabstractDRAM failures are one of the major hardware threats to the reliability of large-scale data centers since the uncorrectable errors in DRAMs may cause servers to shut down. Existing works try to solve this problem by predicting DRAM failures in advance with Machine Learning models. In these works, correctable errors (CEs) are generally deemed as the most important feature. The major reason behind CEs' emergence is the accumulated stress caused by intensive workloads. Moreover, defective DRAMs will not manifest themselves as system errors until the defective cells are accessed by some specific workloads. Therefore, the running workloads on a server are also important for DRAM failure prediction. In this paper, we focus on the impact of workloads on DRAM failures. We design the workload features from both macroscopical and microscopical aspects, i.e. node-level performance metrics and cell-level DRAM access pattern, respectively. Furthermore, we propose Hierarchical DRAM Error Code (HiDEC) to represent the DRAM access pattern. We leverage several Decision Tree-based models for DRAM failure prediction to highlight the generality of our designed features. Experiments are carried out based on the dataset collected from a real-world commercial data center. The results show that both macroscopic and microscopic features can bring significant improvements to the prediction performance. Xingyi Wang, Yu Li 0007, Yiquan Chen, Yin Du, Yuzhong Zhang, Pinan Chen, Wenjun Song, Qiang Xu 0001, Li Jiang 0002 |
VTS | 2 |
| 2020 | DeepDyve: Dynamic Verification for Deep Neural NetworksabstractDeep neural networks (DNNs) have become one of the enabling technologies in many safety-critical applications, e.g., autonomous driving and medical image analysis. DNN systems, however, suffer from various kinds of threats, such as adversarial example attacks and fault injection attacks. While there are many defense methods proposed against maliciously crafted inputs, solutions against faults presented in the DNN system itself (e.g., parameters and calculations) are far less explored. In this paper, we develop a novel lightweight fault-tolerant solution for DNN-based systems, namely DeepDyve, which employs pre-trained neural networks that are far simpler and smaller than the original DNN for dynamic verification. The key to enabling such lightweight checking is that the smaller neural network only needs to produce approximate results for the initial task without sacrificing fault coverage much. We develop efficient and effective architecture and task exploration techniques to achieve optimized risk/overhead trade-off in DeepDyve. Experimental results show that DeepDyve can reduce 90% of the risks at around 10% overhead. Yu Li 0007, Min Li 0019, Bo Luo, Ye Tian 0010, Qiang Xu 0001 |
CCS | 1 |
| 2020 | On Configurable Defense against Adversarial Example AttacksabstractMachine learning systems based on deep neural networks (DNNs) have gained mainstream adoption in many applications. Recently, however, DNNs are shown to be vulnerable to adversarial example attacks with slight perturbations on the inputs. Existing defense mechanisms against such attacks try to improve the overall robustness of the system, but they do not differentiate different targeted attacks even though the corresponding impacts may vary significantly. To tackle this problem, we propose a novel configurable defense mechanism in this work, wherein we are able to flexibly tune the robustness of the system against different targeted attacks to satisfy application requirements. This is achieved by refining the DNN loss function with an attack sensitive matrix to represent the impacts of different targeted attacks. Experimental results on CIFAR-10 data set demonstrate the efficacy of the proposed solution. Bo Luo, Min Li 0019, Yu Li 0007, Qiang Xu 0001 |
ACM Great Lakes Symposium on VLSI | 3 |
| 2019 | D2NN: a fine-grained dual modular redundancy framework for deep neural networksabstractDeep Neural Networks (DNNs) have attracted mainstream adoption in various application domains. Their reliability and security are therefore serious concerns in those safety-critical applications such as surveillance and medical systems. In this paper, we propose a novel dual modular redundancy framework for DNNs, namely D2NN, which is able to tradeoff the system robustness with overhead in a fine-grained manner. We evaluate D2NN framework with DNN models trained on MNIST and CIFAR10 datasets under fault injection attacks, and experimental results demonstrate the efficacy of our proposed solution. Yu Li 0007, Yannan Liu, Min Li 0019, Ye Tian 0010, Bo Luo, Qiang Xu 0001 |
ACSAC | 1 |
| 2019 | On Functional Test Generation for Deep Neural Network IPsabstractMachine learning systems based on deep neural networks (DNNs) produce state-of-the-art results in many applications. Considering the large amount of training data and know-how required to generate the network, it is more practical to use third-party DNN intellectual property (IP) cores for many designs. No doubt to say, it is essential for DNN IP vendors to provide test cases for functional validation without leaking their parameters to IP users. To satisfy this requirement, we propose to effectively generate test cases that activate parameters as many as possible and propagate their perturbations to outputs. Then the functionality of DNN IPs can be validated by only checking their outputs. However, it is difficult considering large numbers of parameters and highly non-linearity of DNNs. In this paper, we tackle this problem by judiciously selecting samples from the DNN training set and applying a gradient-based method to generate new test cases. Experimental results demonstrate the efficacy of our proposed solution. Bo Luo, Yu Li 0007, Lingxiao Wei, Qiang Xu 0001 |
DATE | 2 |
| 2019 | Sample-Efficient Policy Learning based on Completely Behavior CloningabstractDirect policy search is one of the most important algorithm of reinforcement learning. However, learning from scratch needs a large amount of experience data and can be easily prone to poor local optima. In order to overcome these challenges, this paper proposed a training-free behavior cloning algorithm called Policy Learning based on Completely Behavior Cloning (PLCBC). PLCBC transforms the Model Predictive Control (MPC) controller into a PieceWise Affine (PWA) function with multi-parametric programming, and uses a neural network to express this function. By this way, off-the-shelf deep reinforcement learning algorithms can be used to fine-tune this neural network. The experiments show that our method can help agent learn at the high reward state region, and converge faster and better. Qiming Zou, Ling Wang 0005, Yu Li 0007, Jie Liu 0001 |
SMC | 3 |
| 2018 | I Know What You See: Power Side-Channel Attack on Convolutional Neural Network AcceleratorsabstractDeep learning has become the de-facto computational paradigm for various kinds of perception problems, including many privacy-sensitive applications such as online medical image analysis. No doubt to say, the data privacy of these deep learning systems is a serious concern. Different from previous research focusing on exploiting privacy leakage from deep learning models, in this paper, we present the first attack on the implementation of deep learning models. To be specific, we perform the attack on an FPGA-based convolutional neural network accelerator and we manage to recover the input image from the collected power traces without knowing the detailed parameters in the neural network. For the MNIST dataset, our power side-channel attack is able to achieve up to 89% recognition accuracy. Lingxiao Wei, Bo Luo, Yu Li 0007, Yannan Liu, Qiang Xu 0001 |
ACSAC | 3 |
| 2018 | IEEE Std P1838's flexible parallel port and its specification with Google's protocol buffersabstractIEEE Std P1838 is the DfT standard-under-development for 3D test access into dies meant to be used in 3D multi-die stack assemblies. P1838 is the first DfT standard to include a flexible parallel port (FPP): an optional, scalable multi-bit ('parallel') test access mechanism, offering higher test access bandwidth compared to the mandatory one-bit ('serial') port. In this paper, we describe P1838's FPP and propose a formal FPP specification language based on Google's Protocol Buffers (PBs), that potentially could become part of the standard. For a realistic example FPP, we provide its formal specification. Finally, we report on a demonstrator software tool, developed by using PBs-generated data access routines, that converts an FPP specification into a corresponding Verilog netlist. Yu Li 0007, Ming Shao, Hailong Jiao, Adam Cron, Sandeep Bhatia, Erik Jan Marinissen |
ETS | 1 |