VLDB 2026 Research / reviewers in the wild / expert
Xiaolin Xu 0001
dblp:15/1040-1
· DBLP profile ↗
57ranked-venue papers
6as first author
38since 2021 · last 2026
0000-0001-8393-2783ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 38 · 5 first-author · 23 since 2021Artificial intelligence and machine learning · 9 · 9 since 2021Security and privacy · 9 · 1 first-author · 5 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | VSALUT: A Lightweight Low-Dimensional VSA Classifier for Efficient Inference on FPGAabstractDeep neural networks (DNNs) have achieved remarkable success in various domains. However, their deployment on resource-constrained devices and in ultra-low-latency scenarios remains challenging due to their substantial computational and memory requirements. This paper proposes leveraging the low-dimensional computing (LDC) classifier, a low-dimensional variant of the vector symbolic architecture (VSA) classifier, as a lightweight alternative to DNNs. We introduce VSALUT, an open-source framework for implementing LDC on FPGAs. VSALUT employs a lookup table (LUT)-based architecture, an approximate encoding scheme to replace complex adders, a pruning strategy to reduce the model’s size, and a pipelining technique to maximize performance and resource efficiency. Evaluation on five datasets, commonly used to evaluate ultra-low latency architectures, demonstrates that VSALUT achieves significant hardware size and latency reductions over existing LDC, high-dimensional VSA, and LUT-based DNNs. This work establishes a promising new research direction for ultra-low-latency, lightweight inference on FPGAs using LDC. Nuntipat Narkthong, Xiaolin Xu 0001 |
FCCM | 2 |
| 2026 | Privacy-Preserving Constrained Evaluation of LLM-Generated HLS C/C++abstractLarge language models (LLMs) are increasingly evaluated on hardware design tasks through syntax checks, simulation, HLS synthesis, and performance, power, and area (PPA) metrics. These metrics are necessary for HLS usability, but they do not answer whether generated accelerators satisfy security-relevant source-level constraints. This paper studies that missing layer for HLS C/C++ generation and repair. We construct a preliminary security-constrained evaluation overlay for 17 primary HLS kernels by adding source-level policies, deterministic policy checking, functional harnesses where available, and Vitis HLS synthesis for pre-HLS-passing candidates; six NTT/AutoNTT tasks are additionally reported as G1/G3 source-level extension cases. The evaluation compares six descriptive generation and repair conditions: Spec-only generation, Policy-conditioned generation, Rule-aware generation, Checker-feedback repair, Strategy-guided generation, and Semantic-feedback repair. A separate GateRepair case study repairs simple cases but does not dominate one-call settings on hard ciphers. Across eight stable model conditions, pre-HLS pass ranges from 23.5% to 37.9%, and the two best one-call settings are Checker-feedback repair and Semantic-feedback repair. The results expose two distinct mismatch patterns: candidates can pass functional tests while violating the source-level policy, or pass the policy checker while failing functionality. Completed HLS runs synthesize most pre-HLS-passing candidates, but PPA examples show large latency or area shifts. These results show that security-constrained LLM-HLS evaluation must separate syntax, functionality, source-level policy compliance, HLS synthesis, and PPA evidence. Nuo Xu 0013, Jinwei Tang, Xiaolin Xu 0001, Wujie Wen, Zhenman Fang, Caiwen Ding |
ACM Great Lakes Symposium on VLSI | 4 |
| 2025 | Towards Training Robustness Against Dynamic Errors in Quantum Machine LearningabstractQuantum machine learning, crucial in the noisy intermediate-scale quantum (NISQ) era, confronts challenges in error mitigation. Current noise-aware training (NAT) methods often assume static error rates in quantum neural networks (QNNs), overlooking the dynamic nature of quantum noise. Our work highlights how error rates fluctuate over time and across different qubits, affecting QNN performance even when overall error rates are similar. We introduce a novel NAT strategy that dynamically adjusts to standard and fatal error conditions, incorporating a low-complexity search method to identify fatal errors during optimization. This strategy significantly improves robustness, maintaining competitive performance with leading NAT methods across varying error scenarios. Shijin Duan, Gaowen Liu, Charles Fleming, Ramana Rao Kompella, Xiaolin Xu 0001, Shaolei Ren |
DAC | 5 |
| 2025 | Holistic Design towards Resource-Stringent Binary Vector Symbolic ArchitectureabstractClassification tasks on ultra-lightweight devices demand devices that are resource-constrained and deliver swift responses. Binary Vector Symbolic Architecture (VSA) is a promising approach due to its minimal memory requirements and fast execution times compared to traditional machine learning (ML) methods. Nonetheless, binary VSA’s practicality is limited by its inferior inference performance and a design that prioritizes algorithmic over hardware optimization. This paper introduces UniVSA, a co-optimized binary VSA framework for both algorithm and hardware. UniVSA not only significantly enhances inference accuracy beyond current state-of-the-art binary VSA models but also reduces memory footprints. It incorporates novel, lightweight modules and design flow tailored for optimal hardware performance. Experimental results show that UniVSA surpasses traditional ML methods in terms of performance on resource-limited devices, achieving smaller memory usage, lower latency, reduced resource demand, and decreased power consumption. Shijin Duan, Nuntipat Narkthong, Yukui Luo, Shaolei Ren, Xiaolin Xu 0001 |
DAC | 5 |
| 2025 | Taming Diffusion for Dataset Distillation with High RepresentativenessabstractRecent deep learning models demand larger datasets, driving the need for dataset distillation to create compact, cost-efficient datasets while maintaining performance. Due to the powerful image generation capability of diffusion, it has been introduced to this field for generating distilled images. In this paper, we systematically investigate issues present in current diffusion-based dataset distillation methods, including inaccurate distribution matching, distribution deviation with random noise, and separate sampling. Building on this, we propose D$^3$HR, a novel diffusion-based framework to generate distilled datasets with high representativeness. Specifically, we adopt DDIM inversion to map the latents of the full dataset from a low-normality latent domain to a high-normality Gaussian domain, preserving information and ensuring structural consistency to generate representative latents for the distilled dataset. Furthermore, we propose an efficient sampling scheme to better align the representative latents with the high-normality Gaussian distribution. Our comprehensive experiments demonstrate that D$^3$HR can achieve higher accuracy across different model architectures compared with state-of-the-art baselines in dataset distillation. Source code: https://github.com/lin-zhao-resoLve/D3HR. Yushu Wu, Xinru Jiang, Jianyang Gu, Yanzhi Wang 0001, Xiaolin Xu 0001, Pu Zhao 0001, Xue Lin 0001 |
ICML | 6 |
| 2025 | Probe-Me-Not: Protecting Pre-trained Encoders from Malicious Probing
Ruyi Ding, Tong Zhou 0002, Lili Su, A. Adam Ding, Xiaolin Xu 0001, Yunsi Fei |
NDSS | 5 |
| 2025 | ShuffleV: A Microarchitectural Defense Strategy against Electromagnetic Side-Channel Attacks in MicroprocessorsabstractThe run-time electromagnetic (EM) emanation of microprocessors presents a side-channel that leaks the confidentiality of the applications running on them. Many recent works have demonstrated successful attacks leveraging such side-channels to extract the confidentiality of diverse applications, such as the key of cryptographic algorithms and the hyperparameter of neural network models. This paper proposes ShuffleV, a microarchitecture defense strategy against EM Side-Channel Attacks (SCAs). ShuffleV adopts the moving target defense (MTD) philosophy, by integrating hardware units to randomly shuffle the execution order of program instructions and optionally insert dummy instructions, to nullify the statistical observation by attackers across repetitive runs. We build ShuffleV on the open-source RISC-V core and provide six design options, to suit different application scenarios. To enable rapid evaluation, we develop a ShuffleV simulator that can help users to (1) simulate the performance overhead for each design option and (2) generate an execution trace to validate the randomness of execution on their workload. We implement ShuffleV on a Xilinx PYNQ-Z2 FPGA and validate its performance with two representative victim applications against EM SCAs, AES encryption, and neural network inference. The experimental results demonstrate that ShuffleV can provide automatic protection for these applications, without any user intervention or software modification. Nuntipat Narkthong, Yukui Luo, Xiaolin Xu 0001 |
RAID | 3 |
| 2024 | MicroVSA: An Ultra-Lightweight Vector Symbolic Architecture-based Classifier Library for Always-On Inference on Tiny MicrocontrollersabstractArtificial intelligence (AI) on tiny edge devices has become feasible thanks to the emergence of high-performance microcontrollers (MCUs) and lightweight machine learning (ML) models. Nevertheless, the cost and power consumption of these MCUs and the computation requirements of these ML algorithms still present barriers that prevent the widespread inclusion of AI functionality on smaller, cheaper, and lower-power devices. Thus, there is an urgent need for a more efficient ML algorithm and implementation strategy suitable for lower-end MCUs. Nuntipat Narkthong, Shijin Duan, Shaolei Ren, Xiaolin Xu 0001 |
ASPLOS (2) | 4 |
| 2024 | TBNet: A Neural Architectural Defense Framework Facilitating DNN Model Protection in Trusted Execution EnvironmentsabstractTrusted Execution Environments (TEEs) have become a promising solution to secure DNN models on edge devices. However, the existing solutions either provide inadequate protection or introduce large performance overhead. Taking both security and performance into consideration, this paper presents TBNet, a TEE-based defense framework that protects DNN model from a neural architectural perspective. Specifically, TBNet generates a novel Two-Branch substitution model, to respectively exploit (1) the computational resources in the untrusted Rich Execution Environment (REE) for latency reduction and (2) the physically-isolated TEE for model protection. Experimental results on a Raspberry Pi across diverse DNN model architectures and datasets demonstrate that TBNet achieves efficient model protection at a low cost. Tong Zhou 0002, Yukui Luo, Xiaolin Xu 0001 |
DAC | 4 |
| 2024 | AdaPI: Facilitating DNN Model Adaptivity for Efficient Private Inference in Edge ComputingabstractPrivate inference (PI) has emerged as a promising solution to execute computations on encrypted data, safeguarding user privacy and model parameters in edge computing. However, existing PI methods are predominantly developed considering constant resource constraints, overlooking the varied and dynamic resource constraints in diverse edge devices, like energy budgets. Consequently, model providers have to design specialized models for different devices, where all of them have to be stored on the edge server, resulting in inefficient deployment. To fill this gap, this work presents AdaPI, a novel approach that achieves adaptive PI by allowing a model to perform well across edge devices with diverse energy budgets. AdaPI employs a PI-aware training strategy that optimizes the model weights alongside weight-level and feature-level soft masks. These soft masks are subsequently transformed into multiple binary masks to enable adjustments in communication and computation workloads. Through sequentially training the model with increasingly dense binary masks, AdaPI attains optimal accuracy for each energy budget, which outperforms the state-of-the-art PI methods by 7.3% in terms of test accuracy on CIFAR-100. The code of AdaPI can be accessed via https://github.com/jiahuiiiiii/AdaPI. Tong Zhou 0002, Yukui Luo, Wujie Wen, Caiwen Ding, Xiaolin Xu 0001 |
ICCAD | 7 |
| 2024 | ArchLock: Locking DNN Transferability at the Architecture Level with a Zero-Cost Binary PredictorabstractDeep neural network (DNN) models, despite their impressive performance, are vulnerable to exploitation by attackers who attempt to transfer them to other tasks for their own benefit. Current defense strategies mainly address this vulnerability at the model parameter level, leaving the potential of architectural-level defense largely unexplored. This paper, for the first time, addresses the issue of model protection by reducing transferability at the architecture level. Specifically, we present a novel neural architecture search (NAS)-enabled algorithm that employs zero-cost proxies and evolutionary search, to explore model architectures with low transferability. Our method, namely ArchLock, aims to achieve high performance on the source task, while degrading the performance on potential target tasks, i.e., locking the transferability of a DNN model. To achieve efficient cross-task search without accurately knowing the training data owned by the attackers, we utilize zero-cost proxies to speed up architecture evaluation and simulate potential target task embeddings to assist cross-task search with a binary performance predictor. Extensive experiments on NAS-Bench-201 and TransNAS-Bench-101 demonstrate that ArchLock reduces transferability by up to 30% and 50%, respectively, with negligible performance degradation on source tasks (<2%). The code is available at https://github.com/Tongzhou0101/ArchLock. Tong Zhou 0002, Shaolei Ren, Xiaolin Xu 0001 |
ICLR | 3 |
| 2024 | Bileve: Securing Text Provenance in Large Language Models Against Spoofing with Bi-level SignatureabstractText watermarks for large language models (LLMs) have been commonly used to identify the origins of machine-generated content, which is promising for assessing liability when combating deepfake or harmful content. While existing watermarking techniques typically prioritize robustness against removal attacks, unfortunately, they are vulnerable to spoofing attacks: malicious actors can subtly alter the meanings of LLM-generated responses or even forge harmful content, potentially misattributing blame to the LLM developer. To overcome this, we introduce a bi-level signature scheme, Bileve, which embeds fine-grained signature bits for integrity checks (mitigating spoofing attacks) as well as a coarse-grained signal to trace text sources when the signature is invalid (enhancing detectability) via a novel rank-based sampling strategy. Compared to conventional watermark detectors that only output binary results, Bileve can differentiate 5 scenarios during detection, reliably tracing text provenance and regulating LLMs. The experiments conducted on OPT-1.3B and LLaMA-7B demonstrate the effectiveness of Bileve in defeating spoofing attacks with enhanced detectability. Tong Zhou 0002, Xuandong Zhao, Xiaolin Xu 0001, Shaolei Ren |
NeurIPS | 3 |
| 2024 | GraphCroc: Cross-Correlation Autoencoder for Graph Structural ReconstructionabstractGraph-structured data is integral to many applications, prompting the development of various graph representation methods. Graph autoencoders (GAEs), in particular, reconstruct graph structures from node embeddings. Current GAE models primarily utilize self-correlation to represent graph structures and focus on node-level tasks, often overlooking multi-graph scenarios. Our theoretical analysis indicates that self-correlation generally falls short in accurately representing specific graph features such as islands, symmetrical structures, and directional edges, particularly in smaller or multiple graph contexts.To address these limitations, we introduce a cross-correlation mechanism that significantly enhances the GAE representational capabilities. Additionally, we propose the GraphCroc, a new GAE that supports flexible encoder architectures tailored for various downstream tasks and ensures robust structural reconstruction, through a mirrored encoding-decoding process. This model also tackles the challenge of representation bias during optimization by implementing a loss-balancing strategy. Both theoretical analysis and numerical evaluations demonstrate that our methodology significantly outperforms existing self-correlation-based GAEs in graph structure reconstruction. Shijin Duan, Ruyi Ding, A. Adam Ding, Yunsi Fei, Xiaolin Xu 0001 |
NeurIPS | 6 |
| 2024 | Side-Channel-Assisted Reverse-Engineering of Encrypted DNN Hardware Accelerator IP and Attack Surface ExplorationabstractDeep Neural Networks (DNNs) have revolutionized numerous application domains with their unparalleled performance. As the models become larger and more complex, hardware DNN accelerators are increasingly popular. Field-Programmable Gate Array (FPGA)-based DNN accelerators offer near-Application Specific Integrated Circuit (ASIC) efficiency and exceptional flexibility, establishing them as one of the primary hardware platforms for rapidly evolving deep learning implementations, particularly on edge devices. This prominence renders them lucrative targets for attackers. Existing attacks aimed at compromising the confidentiality of DNN models deployed on FPGA DNN accelerators often assume complete knowledge of the accelerators. However, this assumption does not hold for real-world, proprietary, high-performance FPGA DNN accelerators. In this study, we introduce a comprehensive and effective reverse-engineering methodology for demystifying FPGA DNN accelerator soft Intellectual Property (IP) cores. We demonstrate its application on the cutting-edge AMD-Xilinx Deep Learning Processing Unit (DPU). Our method relies on schematic analysis and, innovatively, electromagnetic (EM) side-channel analysis to reveal the data flow and scheduling of the DNN accelerators. To the best of our knowledge, this research is the first successful endeavor to reverse-engineer a commercial encrypted DNN accelerator IP. Moreover, we investigate attack surfaces exposed by the reverse-engineering findings, including the successful recovery of DNN model architectures and extraction of model parameters. These outcomes pose a significant threat to real-world commercial FPGA-DNN acceleration systems. We discuss potential countermeasures and offer recommendations for FPGA-based IP protection. Cheng Gongye, Yukui Luo, Xiaolin Xu 0001, Yunsi Fei |
SP | 3 |
| 2024 | DeepShuffle: A Lightweight Defense Framework against Adversarial Fault Injection Attacks on Deep Neural Networks in Multi-Tenant Cloud-FPGAabstractFPGA virtualization has garnered significant industry and academic interests as it aims to enable multi-tenant cloud systems that can accommodate multiple users’ circuits on a single FPGA. Although this approach greatly enhances the efficiency of hardware resource utilization, it also introduces new security concerns. As a representative study, one stateof-the-art (SOTA) adversarial fault injection attack, named Deep-Dup [1], exemplifies the vulnerabilities of off-chip data communication within the multi-tenant cloud-FPGA system. Deep-Dup attacks successfully demonstrate the complete failure of a wide range of Deep Neural Networks (DNNs) in a black-box setup, by only injecting fault to extremely small amounts of sensitive weight data transmissions, which are identified through a powerful differential evolution searching algorithm. Such emerging adversarial fault injection attack reveals the urgency of effective defense methodology to protect DNN applications on the multi-tenant cloud-FPGA system.This paper, for the first time, presents a novel movingtarget-defense (MTD) oriented defense framework DeepShuffle, which could effectively protect DNNs on multi-tenant cloudFPGA against the SOTA Deep-Dup attack, through a novel lightweight model parameter shuffling methodology. DeepShuffle effectively counters the Deep-Dup attack by altering the weight transmission sequence, which effectively prevents adversaries from identifying security-critical model parameters from the repeatability of weight transmission during each inference round. Importantly, DeepShuffle represents a training-free DNN defense methodology, which makes constructive use of the typologies of DNN architectures to achieve being lightweight. Moreover, the deployment of DeepShuffle neither requires any hardware modification nor suffers from any performance degradation. We evaluate DeepShuffle on the SOTA open-source FPGA-DNN accelerator, Vertical Tensor Accelerator (VTA), which represents the practice of real-world FPGA-DNN system developers. We then evaluate the performance overhead of DeepShuffle and find it only consumes an additional ∼3% of the inference time compared to the unprotected baseline. DeepShuffle improves the robustness of various SOTA DNN architectures like VGG, ResNet, etc. against Deep-Dup by orders. It effectively reduces the efficacy of evolution searching-based adversarial fault injection attack close to random fault injection attack, e.g., on VGG-11, even after increasing the attacker’s effort by 2.3×, our defense shows a ∼93% improvement in accuracy, compared to the unprotected baseline. Yukui Luo, Adnan Siraj Rakin, Deliang Fan, Xiaolin Xu 0001 |
SP | 4 |
| 2024 | ALLI/O Diagram: An Action-based Visual Programming Language for Embedded SystemabstractThis paper introduces ALLI/O Diagram, an action-based visual programming language for embedded system programming used by the ALLI/O IDE. We illustrate the practicality of ALLI/O Diagram with various design examples and evaluate it against block-based, event-based, device-based, and state-based programming approaches in terms of programming effort, readability, and portability of the result programs. These results demonstrate that our proposed ALLI/O Diagram is the most compact, expressive, and portable across different hardware models. We open source the ALLI/O Diagram and all example programs at https://allio.build. Nuntipat Narkthong, Chattriya Jariyavajee, Xiaolin Xu 0001 |
VL/HCC | 3 |
| 2023 | HammerDodger: A Lightweight Defense Framework against RowHammer Attack on DNNsabstractRowHammer attacks have become a serious security problem on deep neural networks (DNNs). Some carefully induced bit-flips degrade the prediction accuracy of DNN models to random guesses. This work proposes a lightweight defense framework that detects and mitigates adversarial bit-flip attacks. We employ a dynamic channel-shuffling obfuscation scheme to present moving targets to the attack, and develop a logits-based model integrity monitor with negligible performance loss. The parameters and architecture of DNN models remain unchanged, which ensures lightweight deployment and makes the framework compatible with commodity models. We demonstrate that our framework can protect various DNN models against RowHammer attacks. Cheng Gongye, Yukui Luo, Xiaolin Xu 0001, Yunsi Fei |
DAC | 3 |
| 2023 | PASNet: Polynomial Architecture Search Framework for Two-party Computation-based Secure Neural Network DeploymentabstractTwo-party computation (2PC) is promising to enable privacy-preserving deep learning (DL). However, the 2PC-based privacy-preserving DL implementation comes with high comparison protocol overhead from the non-linear operators. This work presents PASNet, a novel systematic framework that enables low latency, high energy efficiency & accuracy, and security-guaranteed 2PC-DL by integrating the hardware latency of the cryptographic building block into the neural architecture search loss function. We develop a cryptographic hardware scheduler and the corresponding performance model for Field Programmable Gate Arrays (FPGA) as a case study. The experimental results demonstrate that our light-weighted model PASNet-A and heavily-weighted model PASNet-B achieve 63 ms and 228 ms latency on private inference on ImageNet, which are 147 and 40 times faster than the SOTA CryptGPU system, and achieve 70.54% & 78.79% accuracy and more than 1000 times higher energy efficiency. The pretrained PASNet models and test code can be found on Github1. Hongwu Peng, Shanglin Zhou, Yukui Luo, Nuo Xu 0013, Shijin Duan, Chenghong Wang, Tong Geng, Wujie Wen, Xiaolin Xu 0001, Caiwen Ding |
DAC | 11 |
| 2023 | MirrorNet: A TEE-Friendly Framework for Secure On-Device DNN InferenceabstractDeep neural network (DNN) models have become prevalent in edge devices for real-time inference. However, they are vulnerable to model extraction attacks and require protection. Existing defense approaches either fail to fully safeguard model confidentiality or result in significant latency issues. To overcome these challenges, this paper presents MirrorNet, which leverages Trusted Execution Environment (TEE) to enable secure on-device DNN inference. It generates a TEE-friendly implementation for any given DNN model to protect the model confidentiality, while meeting the stringent computation and storage constraints of TEE. The framework consists of two key components: the backbone model (BackboneNet), which is stored in the normal world but achieves lower inference accuracy, and the Companion Partial Monitor (CPM), a lightweight mirrored branch stored in the secure world, preserving model confidentiality. During inference, the CPM monitors the intermediate results from the BackboneNet and rectifies the classification output to achieve higher accuracy. To enhance flexibility, MirrorNet incorporates two modules: the CPM Strategy Generator, which generates various protection strategies, and the Performance Emulator, which estimates the performance of each strategy and selects the most optimal one. Extensive experiments demonstrate the effectiveness of MirrorNet in providing security guarantees while maintaining low computation latency, making MirrorNet a practical and promising solution for secure on-device DNN inference. For the evaluation, MirrorNet can achieve a 18.6% accuracy gap between authenticated and illegal use, while only introducing 0.99% hardware overhead. Yukui Luo, Shijin Duan, Tong Zhou 0002, Xiaolin Xu 0001 |
ICCAD | 5 |
| 2023 | VertexSerum: Poisoning Graph Neural Networks for Link InferenceabstractGraph neural networks (GNNs) have brought superb performance to various applications utilizing graph structural data, such as social analysis and fraud detection. The graph links, e.g., social relationships and transaction history, are sensitive and valuable information, which raises privacy concerns when using GNNs. To exploit these vulnerabilities, we propose VertexSerum, a novel graph poisoning attack that increases the effectiveness of graph link stealing by amplifying the link connectivity leakage. To infer node adjacency more accurately, we propose an attention mechanism that can be embedded into the link detection network. Our experiments demonstrate that VertexSerum significantly outperforms the SOTA link inference attack, improving the AUC scores by an average of 9.8% across four real-world datasets and three different GNN structures. Furthermore, our experiments reveal the effectiveness of VertexSerum in both black-box and online learning settings, further validating its applicability in real-world scenarios. The source code is available at https://github.com/RollinDing/VertexSerum. Ruyi Ding, Shijin Duan, Xiaolin Xu 0001, Yunsi Fei |
ICCV | 3 |
| 2023 | AutoReP: Automatic ReLU Replacement for Fast Private Network InferenceabstractThe growth of the Machine-Learning-As-A-Service (MLaaS) market has highlighted clients’ data privacy and security issues. Private inference (PI) techniques using cryptographic primitives offer a solution but often have high computation and communication costs, particularly with non-linear operators like ReLU. Many attempts to reduce ReLU operations exist, but they may need heuristic threshold selection or cause substantial accuracy loss. This work introduces AutoReP, a gradient-based approach to lessen non-linear operators and alleviate these issues. It automates the selection of ReLU and polynomial functions to speed up PI applications and introduces distribution-aware polynomial approximation (DaPa) to maintain model expressivity while accurately approximating ReLUs. Our experimental results demonstrate significant accuracy improvements of 6.12% (94.31%, 12.9K ReLU budget, CIFAR-10), 8.39% (74.92%, 12.9K ReLU budget, CIFAR-100), and 9.45% (63.69%, 55K ReLU budget, Tiny-ImageNet) over current state-of-the-art methods, e.g., SNL. Morever, AutoReP is applied to EfficientNet-B2 on ImageNet dataset, and achieved 75.55% accuracy with 176.1 × ReLU budget reduction. The codes are shared on Github1. Hongwu Peng, Shaoyi Huang, Tong Zhou 0002, Yukui Luo, Chenghong Wang, Zigeng Wang, Ang Li 0006, Tong Geng, Kaleel Mahmood, Wujie Wen, Xiaolin Xu 0001, Caiwen Ding |
ICCV | 13 |
| 2023 | SpENCNN: Orchestrating Encoding and Sparsity for Fast Homomorphically Encrypted Neural Network InferenceabstractHomomorphic Encryption (HE) is a promising technology to protect clients’ data privacy for Machine Learning as a Service (MLaaS) on public clouds. However, HE operations can be orders of magnitude slower than their counterparts for plaintexts and thus result in prohibitively high inference latency, seriously hindering the practicality of HE. In this paper, we propose a HE-based fast neural network (NN) inference framework–SpENCNN built upon the co-design of HE operation-aware model sparsity and the single-instruction-multiple-data (SIMD)-friendly data packing, to improve NN inference latency. In particular, we first develop an encryption-aware HE-group convolution technique that can partition channels among different groups based on the data size and ciphertext size, and then encode them into the same ciphertext by novel group-interleaved encoding, so as to dramatically reduce the number of bottlenecked operations in HE convolution. We further tailor a HE-friendly sub-block weight pruning to reduce the costly HE-based convolution operation. Our experiments show that SpENCNN can achieve overall speedups of 8.37$\times$, 12.11$\times$, 19.26$\times$, and 1.87$\times$ for LeNet, VGG-5, HEFNet, and ResNet-20 respectively, with negligible accuracy loss. Our code is publicly available at https://github.com/ranran0523/SPECNN. Xinwei Luo, Tao Liu 0023, Gang Quan, Xiaolin Xu 0001, Caiwen Ding, Wujie Wen |
ICML | 6 |
| 2023 | NNSplitter: An Active Defense Solution for DNN Model via Automated Weight ObfuscationabstractAs a type of valuable intellectual property (IP), deep neural network (DNN) models have been protected by techniques like watermarking. However, such passive model protection cannot fully prevent model abuse. In this work, we propose an active model IP protection scheme, namely NNSplitter, which actively protects the model by splitting it into two parts: the obfuscated model that performs poorly due to weight obfuscation, and the model secrets consisting of the indexes and original values of the obfuscated weights, which can only be accessed by authorized users with the support of the trusted execution environment. Experimental results demonstrate the effectiveness of NNSplitter, e.g., by only modifying 275 out of over 11 million (i.e., 0.002%) weights, the accuracy of the obfuscated ResNet-18 model on CIFAR-10 can drop to 10%. Moreover, NNSplitter is stealthy and resilient against norm clipping and fine-tuning attacks, making it an appealing solution for DNN model protection. The code is available at: https://github.com/Tongzhou0101/NNSplitter. Tong Zhou 0002, Yukui Luo, Shaolei Ren, Xiaolin Xu 0001 |
ICML | 4 |
| 2023 | AQ2PNN: Enabling Two-party Privacy-Preserving Deep Neural Network Inference with Adaptive QuantizationabstractThe growing prevalence of Machine Learning as a Service (MLaaS) enables a wide range of applications but simultaneously raises numerous security and privacy concerns. A key issue involves the potential privacy exposure of involved parties, such as the customer’s input data and the vendor’s model. Consequently, two-party computing (2PC) has emerged as a promising solution to safeguard the privacy of different parties during deep neural network (DNN) inference. However, the state-of-the-art (SOTA) 2PC-DNN techniques are tailored explicitly to traditional instruction set architecture (ISA) systems like CPUs and CPU+GPU. This reliance on ISA systems significantly constrains their energy efficiency, as these architectures typically employ 32- or 64-bit instruction sets. In contrast, the possibilities of harnessing dynamic and adaptive quantization to build high-performance 2PC-DNNs remain largely unexplored due to the lack of compatible algorithms and hardware accelerators. Yukui Luo, Nuo Xu 0013, Hongwu Peng, Chenghong Wang, Shijin Duan, Kaleel Mahmood, Wujie Wen, Caiwen Ding, Xiaolin Xu 0001 |
MICRO | 9 |
| 2023 | LinGCN: Structural Linearized Graph Convolutional Network for Homomorphically Encrypted InferenceabstractThe growth of Graph Convolution Network (GCN) model sizes has revolutionized numerous applications, surpassing human performance in areas such as personal healthcare and financial systems. The deployment of GCNs in the cloud raises privacy concerns due to potential adversarial attacks on client data. To address security concerns, Privacy-Preserving Machine Learning (PPML) using Homomorphic Encryption (HE) secures sensitive client data. However, it introduces substantial computational overhead in practical applications. To tackle those challenges, we present LinGCN, a framework designed to reduce multiplication depth and optimize the performance of HE based GCN inference. LinGCN is structured around three key elements: (1) A differentiable structural linearization algorithm, complemented by a parameterized discrete indicator function, co-trained with model weights to meet the optimization goal. This strategy promotes fine-grained node-level non-linear location selection, resulting in a model with minimized multiplication depth. (2) A compact node-wise polynomial replacement policy with a second-order trainable activation function, steered towards superior convergence by a two-level distillation approach from an all-ReLU based teacher model. (3) an enhanced HE solution that enables finer-grained operator fusion for node-wise activation functions, further reducing multiplication level consumption in HE-based inference. Our experiments on the NTU-XVIEW skeleton joint dataset reveal that LinGCN excels in latency, accuracy, and scalability for homomorphically encrypted inference, outperforming solutions such as CryptoGCN. Remarkably, LinGCN achieves a 14.2× latency speedup relative to CryptoGCN, while preserving an inference accuracy of ~75\% and notably reducing multiplication depth. Additionally, LinGCN proves scalable for larger models, delivering a substantial 85.78\% accuracy with 6371s latency, a 10.47\% accuracy improvement over CryptoGCN. Hongwu Peng, Yukui Luo, Shaoyi Huang, Kiran Thorat, Tong Geng, Chenghong Wang, Xiaolin Xu 0001, Wujie Wen, Caiwen Ding |
NeurIPS | 9 |
| 2022 | LeHDC: learning-based hyperdimensional computing classifierabstractThanks to the tiny storage and efficient execution, hyperdimensional Computing (HDC) is emerging as a lightweight learning framework on resource-constrained hardware. Nonetheless, the existing HDC training relies on various heuristic methods, significantly limiting their inference accuracy. In this paper, we propose a new HDC framework, called LeHDC, which leverages a principled learning approach to improve the model accuracy. Concretely, LeHDC maps the existing HDC framework into an equivalent Binary Neural Network architecture, and employs a corresponding training strategy to minimize the training loss. Experimental validation shows that LeHDC outperforms previous HDC training strategies and can improve on average the inference accuracy over 15% compared to the baseline HDC. Shijin Duan, Yejia Liu, Shaolei Ren, Xiaolin Xu 0001 |
DAC | 4 |
| 2022 | HDLock: exploiting privileged encoding to protect hyperdimensional computing models against IP stealingabstractHyperdimensional Computing (HDC) is facing infringement issues due to straightforward computations. This work, for the first time, raises a critical vulnerability of HDC --- an attacker can reverse engineer the entire model, only requiring the unindexed hypervector memory. To mitigate this attack, we propose a defense strategy, namely HDLock, which significantly increases the reasoning cost of encoding. Specifically, HDLock adds extra feature hypervector combination and permutation in the encoding module. Compared to the standard HDC model, a two-layer-key HDLock can increase the adversarial reasoning complexity by 10 order of magnitudes without inference accuracy loss, with only 21% latency overhead. Shijin Duan, Shaolei Ren, Xiaolin Xu 0001 |
DAC | 3 |
| 2022 | NNReArch: A Tensor Program Scheduling Framework Against Neural Network Architecture Reverse EngineeringabstractArchitecture reverse engineering has become an emerging attack against deep neural network (DNN) implementations. Several prior works have utilized side-channel leakage to recover the model architecture while the an DNN is executing on a hardware acceleration platform. In this work, we target an open-source deep-learning accelerator, Versatile Tensor Accelerator (VTA), and utilize electromagnetic (EM) side-channel leakage to comprehensively learn the association between DNN architecture configurations and EM emanations. We also consider the holistic system–including the low-level tensor program code of the VTA accelerator on a Xilinx FPGA, and explore the effect of such low-level configurations on the EM leakage. Our study demonstrates that both the optimization and configuration of tensor programs will affect the EM side-channel leakage.Gaining knowledge of the association between low-level tensor program and the EM emanations, we propose NNReArch, a lightweight tensor program scheduling framework against side-channel-based DNN model architecture reverse engineering. Specifically, NNReArch targets reshaping the EM traces of different DNN operators, through scheduling the tensor program execution of the DNN model so as to confuse the adversary. NNReArch is a comprehensive protection framework supporting two modes, a balanced mode that strikes a balance between the DNN model confidentiality and execution performance, and a secure mode where the most secure setting is chosen. We implement and evaluate the proposed framework on the open-source VTA with state-of-the-art DNN architectures. The experimental results demonstrate that NNReArch can efficiently enhance the model architecture security with a small performance overhead. In addition, the proposed obfuscation technique makes reverse engineering of the DNN architecture significantly harder. Yukui Luo, Shijin Duan, Cheng Gongye, Yunsi Fei, Xiaolin Xu 0001 |
FCCM | 5 |
| 2022 | An Integrity Checking Framework for AXI Protocol in Multi-tenant FPGAabstractFPGAs have been widely deployed in cloud servers a promising computing infrastructure. It is envisioned that through the FPGA virtualization, multiple users will be able to share the hardware resources of an FPGA chip. Although bringing great benefits like highly efficient resource utilization, such multi-tenant cloud-FPGA also raises security concerns. While most existing works have demonstrated that an application running on a multi-tenant FPGA is vulnerable to fault injection attacks introduced by the power distribution network (PDN) disturbance on an FPGA, few of them consider the impact of such power attacks on the data interface of an FPGA. As a critical component in charge of the data transmission between intra-FPGA IPs and/or inter-FPGA components, the runtime integrity of the communication interface plays an essential role in the security of a multi-tenant FPGA. This work, we investigate the impact of the PDN-based fault injection attacks on the Advanced eXtensible Interface (AXI) protocol as a case study. We find that like other circuit applications, the embedded AXI protocol interface also suffers from the PDN-based fault injection attacks, but with unique fault characteristics. To mitigate such vulnerabilities of the AXI in multi-tenant cloud-FPGA, we propose a runtime integrity checking framework. We evaluate the proposed framework with practical data transmission setup to demonstrate the effectiveness of the proposed scheme. Yukui Luo, Shijin Duan, Xiaolin Xu 0001 |
FPGA | 4 |
| 2022 | A Cautionary Note on Building Multi-tenant Cloud-FPGA as a Secure InfrastructureabstractSecurity concerns have been raised for multi-tenant cloud-FPGA in many recent works. While these existing works focused on studying the security of diverse cloud-FPGA applications, such as Advanced Encryption Standard (AES), the vulnerabilities associated with the inherent FPGA components are so far under-explored. For the first time, we investigate the robustness of a commonly used communication protocol for data exchange, Advanced eXtensible Interface (AXI), against fault injection attacks in a multi-tenant cloud-FPGA environment. We build an experimental setup with a commodity FPGA development kit and launch fault injection attacks on the shared power distribution network (PDN). To study the in-depth effects of such attacks, we characterize the voltage glitches of different attack patterns in a non-invasive manner, i.e., using electron magnetic measurement. We also mimic the real-world data transmissions using two crafted datasets with different statistical characteristics. The experimental results demonstrate the unique security vulnerabilities of the current AXI protocol in the context of a multi-tenant cloud-FPGA. Last, we discuss potential defense strategies against these vulnerabilities. Yukui Luo, Shijin Duan, Xiaolin Xu 0001 |
FPT | 4 |
| 2022 | ObfuNAS: A Neural Architecture Search-Based DNN Obfuscation ApproachabstractMalicious architecture extraction has been emerging as a crucial concern for deep neural network (DNN) security. As a defense, architecture obfuscation is proposed to remap the victim DNN to a different architecture. Nonetheless, we observe that, with only extracting an obfuscated DNN architecture, the adversary can still retrain a substitute model with high performance (e.g., accuracy), rendering the obfuscation techniques ineffective. To mitigate this under-explored vulnerability, we propose ObfuNAS, which converts the DNN architecture obfuscation into a neural architecture search (NAS) problem. Using a combination of function-preserving obfuscation strategies, ObfuNAS ensures that the obfuscated DNN architecture can only achieve lower accuracy than the victim. We validate the performance of ObfuNAS with open-source architecture datasets like NAS-Bench-101 and NAS-Bench-301. The experimental results demonstrate that ObfuNAS can successfully find the optimal mask for a victim model within a given FLOPs constraint, leading up to 2.6% inference accuracy degradation for attackers with only 0.14× FLOPs overhead. The code is available at: https://github.com/Tongzhou0101/ObfuNAS. Tong Zhou 0002, Shaolei Ren, Xiaolin Xu 0001 |
ICCAD | 3 |
| 2022 | STT-MRAM-Based Reliable Weak PUFabstractIn recent years, micro-nano device characteristics like ferroelectrics and resistive switching are being used to build important security primitives such as Physical Unclonable Function (PUF). The micro-nano device-based hardware security primitives, although with higher security, energy efficiency, and integration density, suffer from serious reliability issues caused by process scaling. To mitigate this issue, this paper introduces a reconfigurable weak PUF based on spin-transfer torque magnetoresistive random-access memory (STT-MRAM), which adopts the crossing switches implemented with simple demultiplexes (DEMUXs) to improve the flexibility and reliability. Moreover, two algorithms,neighboring bit linesandtop-$n$n, are proposed to enlarge the gap between two parallel reading currents, thus further enhancing the reliability of PUF responses. Experimental results demonstrate that the proposed PUF scheme achieves good uniqueness (50.64 percent), uniformity (50.02 percent), and bit-aliasing ($\approx$49.80%). Particularly, the proposed method significantly improves the PUF reliability, achieving low bit error rate (BER$\leq$2.13%) within the range of -20$^\circ$C to 90$^\circ$C. Yupeng Hu 0004, Linjun Wu, Zhuojun Chen, Xiaolin Xu 0001, Keqin Li 0001, Jiliang Zhang 0002 |
IEEE Trans. Computers | 5 |
| 2022 | FLAM-PUF: A Response-Feedback-Based Lightweight Anti-Machine-Learning-Attack PUFabstractPhysical unclonable functions (PUFs) have been adopted in many resource-constrained Internet of Things (IoT) applications to provide effective and lightweight solutions for device authentication. However, an attacker can collect challenge–response pairs (CRPs) of a strong PUF, to build a machine learning (ML) model and mimic its behavior, i.e., predicting the responses of unseen challenges with high accuracy. Although several PUFs have been proposed to resist such modeling attacks, they incur high hardware overhead. Developing a PUF primitive with low hardware cost and high resistance to ML attacks is thus a crucial task. In this article, we propose the first response–feedback-based lightweight anti-ML-attack PUF (FLAM-PUF). It is only composed of one arbiter PUF (APUF) and one Galois linear-feedback shift register (LFSR), with some basic logic gates, reducing more than 62% hardware cost compared with the state-of-the-art robust strong PUFs. Specifically, FLAM-PUF leverages a cost-effective feedback loop structure to dynamically control and update the LFSR configuration. FLAM-PUF has two main characteristics: 1) it feeds back a 1-bit response in every cycle to intentionally poison the data of the CRP set for training. To resist ML-based modeling attacks, the 1-bit response can randomly update one coefficient of the feedback polynomial to implant more complex correlations into the model built by attackers and 2) it takes advantage of an$n-$bit response feedback-controlled reconfigurable Galois LFSR to enlarge the original challenge space of the APUF. Extensive experimental results show that the proposed FLAM-PUF achieves near-optimal uniformity, uniqueness, and reliability. Our scheme works well under standard attack models with public crucial initial information. In particular, the prediction accuracy of modeling attacks against FLAM-PUF is nearly 50% under the four widely used ML algorithms, i.e., support vector machines (SVMs), logistic regression (LR), covariance matrix adaptation evolution strategy (CMA-ES), and deep neural networks (DNNs), indicating excellent resistance against these ML attacks. Linjun Wu, Yupeng Hu 0004, Kehuan Zhang, Wenjia Li, Xiaolin Xu 0001, Wanli Chang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2022 | FPGAPRO: A Defense Framework Against Crosstalk-Induced Secret Leakage in FPGAabstractWith the emerging cloud-computing development, FPGAs are being integrated with cloud servers for higher performance. Recently, it has been explored to enable multiple users to share the hardware resources of a remote FPGA, i.e., to execute their own applications simultaneously. Although being a promising technique, multi-tenant FPGA unfortunately brings its unique security concerns. It has been demonstrated that the capacitive crosstalk between FPGA long-wires can be a side-channel to extract secret information, giving adversaries the opportunity to implement crosstalk-based side-channel attacks. Moreover, recent work reveals that medium-wires and multiplexers in configurable logic block (CLB) are also vulnerable to crosstalk-based information leakage. In this work, we propose FPGAPRO: a defense framework leveraging P lacement, R outing, and O bfuscation to mitigate the secret leakage on FPGA components, including long-wires, medium-wires, and logic elements in CLB. As a user-friendly defense strategy, FPGAPRO focuses on protecting the security-sensitive instances meanwhile considering critical path delay for performance maintenance. As the proof-of-concept, the experimental result demonstrates that FPGAPRO can effectively reduce the crosstalk-caused side-channel leakage by 138 times. Besides, the performance analysis shows that this strategy prevents the maximum frequency from timing violation. Yukui Luo, Shijin Duan, Xiaolin Xu 0001 |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2021 | DeepStrike: Remotely-Guided Fault Injection Attacks on DNN Accelerator in Cloud-FPGAabstractAs Field-programmable gate arrays (FPGAs) are widely adopted in clouds to accelerate Deep Neural Networks (DNN), such virtualization environments have posed many new security issues. This work investigates the integrity of DNN FPGA accelerators in clouds. It proposes DeepStrike, a remotely-guided attack based on power glitching fault injections targeting DNN execution. We characterize the vulnerabilities of different DNN layers against fault injections on FPGAs and leverage time-to-digital converter (TDC) sensors to precisely control the timing of fault injections. Experimental results show that our proposed attack can successfully disrupt the FPGA DSP kernel and misclassify the target victim DNN application. Yukui Luo, Cheng Gongye, Yunsi Fei, Xiaolin Xu 0001 |
DAC | 4 |
| 2021 | SGX-FPGA: Trusted Execution Environment for CPU-FPGA Heterogeneous ArchitectureabstractTrusted execution environments (TEEs), such as Intel SGX, have become a popular security primitive with minimum trusted computing base (TCB) and attack surface. However, the existing CPU-based TEEs do not support FPGAs, even though FPGA-based cloud computing services have been rapidly deployed with security vulnerabilities that are expected to be eliminated by TEEs. To fill the gap, we present SGX-FPGA, a trusted hardware isolation path enabling the first FPGA TEE by bridging SGX enclaves and FPGAs in the heterogeneous CPU-FPGA architecture. Our experiments on real CPU-FPGA hardware justify the high security and low performance overhead achieved by SGX-FPGA. Ke Xia, Yukui Luo, Xiaolin Xu 0001, Sheng Wei 0001 |
DAC | 3 |
| 2021 | Constructive Use of Process Variations: Reconfigurable and High-Resolution Delay-LineabstractDelay-line is a critical circuit component for highspeed electronic design and testing, such as high-performance FPGA and ASICs, to provide timing signals of specific duration or duty cycle. However, the performance of existing CMOS-based delay-lines is limited by various practical issues. For example, the minimum propagation delay (resolution) of CMOS gates is limited by the process variations from circuit fabrication. This paper presents a novel delay-line scheme, which instead of mitigating the process variations from circuit fabrication, constructively leverages them to generate time signals of specific duration. Moreover, the resolution of the proposed delay-line method is reconfigurable, for which we propose a Machine Learning modeling method to assist such reconfiguration, i.e., to generate time duration of different scales. The performance of the proposed delay-line is validated with HSpice simulation and prototype on a Xilinx Virtex-6 FPGA evaluation kit. The experimental results demonstrate that the proposed delay-line method achieves an ultra-high resolution of sub-picosecond. Wenhao Wang 0001, Yukui Luo, Xiaolin Xu 0001 |
DATE | 3 |
| 2021 | Deep-Dup: An Adversarial Weight Duplication Attack Framework to Crush Deep Neural Network in Multi-Tenant FPGA
Adnan Siraj Rakin, Yukui Luo, Xiaolin Xu 0001, Deliang Fan |
USENIX Security Symposium | 3 |
| 2020 | A Dynamic Frequency Scaling Framework Against Reliability and Security Issues in Multi-tenant FPGAabstractThe deployment of cloud-FPGA, although significantly improves the performance and efficiency of conventional cloud computing, also creates a unique attack surface. For example, the power distribution network (PDN) of a multitenant FPGA can be manipulated by malicious users to cause voltage-drop, which can be used to inject timing faults to the applications of victim users. In addition, since most cloud-FPGAs are being used for computation-intensive tasks that consume a large amount of power, therefore, typical FPGA applications may still encounter reliability issues even without malicious PDN manipulation. To comprehensively mitigate both reliability and security issues of multi-tenant FPGAs, this paper proposes a dynamic frequency scaling (DFS) framework, which can dynamically scale the clock frequency of FPGA applications. Specifically, the DFS framework utilizes a delay-frequency pair table, which can match the real-time clock frequency with voltage level, thus avoiding potential errors caused by voltage-drop. Yukui Luo, Xiaolin Xu 0001 |
FCCM | 2 |
| 2020 | A Privacy-Preserving-Oriented DNN Pruning and Mobile Acceleration FrameworkabstractWeight pruning of deep neural networks (DNNs) has been proposed to satisfy the limited storage and computing capability of mobile edge devices. However, previous pruning methods mainly focus on reducing the model size and/or improving performance without considering the privacy of user data. To mitigate this concern, we propose a privacy-preserving-oriented pruning and mobile acceleration framework that does not require the private training dataset. At the algorithm level of the proposed framework, a systematic weight pruning technique based on the alternating direction method of multipliers (ADMM) is designed to iteratively solve the pattern-based pruning problem for each layer with randomly generated synthetic data. In addition, corresponding optimizations at the compiler level are leveraged for inference accelerations on devices. With the proposed framework, users could avoid the time-consuming pruning process for non-experts and directly benefit from compressed models. Experimental results show that the proposed framework outperforms three state-of-art end-to-end DNN frameworks, i.e., TensorFlow-Lite, TVM, and MNN, with speedup up to 4.2×, 2.5×, and 2.0×, respectively, with almost no accuracy loss, while preserving data privacy. Yifan Gong 0004, Zheng Zhan 0001, Zhengang Li 0001, Wei Niu 0002, Wenhao Wang 0001, Bin Ren 0002, Caiwen Ding, Xue Lin 0001, Xiaolin Xu 0001, Yanzhi Wang 0001 |
ACM Great Lakes Symposium on VLSI | 10 |
| 2020 | A Quantitative Defense Framework against Power Attacks on Multi-tenant FPGAabstractThe development and application of various Machine Learning algorithms demand high computing capabilities. As a result, field-programmable gate arrays (FPGAs) are being used as hardware accelerators, and more recently deployed in cloud servers by leading vendors to provide reconfigurable computing capabilities. Although such cloud-FPGA platform is bringing significant performance benefits, it also creates a unique attack surface where the hardware resources of an FPGA are shared by multiple users. Power attack targeting the power distribution network (PDN) is among the most threatening ones against multi-tenant FPGAs. In such attack, the malicious users leverage power plundering circuits to manipulate the PDN and cause a voltage drop, thus injecting timing faults to the victim applications. Besides, since most cloud-FPGAs are being used for computing-intensive tasks that consume a large amount of power, therefore, typical FPGA applications may still encounter timing faults even without power attacks. Unlike power attacks, we classify this problem as a reliability issue. Yukui Luo, Xiaolin Xu 0001 |
ICCAD | 2 |
| 2020 | Stealthy-Shutdown: Practical Remote Power Attacks in Multi - Tenant FPGAsabstractWith the deployment of artificial intelligent (AI) algorithms in a large variety of applications, there creates an increasing need for high-performance computing capabilities. As a result, different hardware platforms have been utilized for acceleration purposes. Among these hardware-based accelerators, the field-programmable gate arrays (FPGAs) have gained a lot of attention due to their re-programmable characteristics, which provide customized control logic and computing operators. For example, FPGAs have recently been adopted for on-demand cloud services by the leading cloud providers like Amazon and Microsoft, providing acceleration for various compute-intensive tasks. While the co-residency of multiple tenants on a cloud FPGA chip increases the efficiency of resource utilization, it also creates unique attack surfaces that are under-explored. In this paper, we exploit the vulnerability associated with the shared power distribution network on cloud FPGAs. We present a stealthy power attack that can be remotely launched by a malicious tenant, shutting down the entire chip and resulting in denial-of-service for other co-located benign tenants. Specifically, we propose stealthy-shutdown: a well-timed power attack that can be implemented in two steps: (1) an attacker monitors the realtime FPGA power-consumption detected by ring-oscillator-based voltage sensors, and (2) when capturing high power-consuming moments, i.e., the power consumption by other tenants is above a certain threshold, she/he injects a well-timed power load to shut down the FPGA system. Note that in the proposed attack strategy, the power load injected by the attacker only accounts for a small portion of the overall power consumption; therefore, such attack strategy remains stealthy to the cloud FPGA operator. We successfully implement and validate the proposed attack on three FPGA evaluation kits with running real-world applications. The proposed attack results in a stealthy-shutdown, demonstrating severe security concerns of co-tenancy on cloud FPGAs. We also offer two countermeasures that can mitigate such power attacks. Yukui Luo, Cheng Gongye, Shaolei Ren, Yunsi Fei, Xiaolin Xu 0001 |
ICCD | 5 |
| 2019 | An All-Digital True Random Number Generator Based on Chaotic Cellular Automata TopologyabstractTrue random number generator (TRNG) is an important primitive in cryptographic applications. In this paper, a TRNG based on a self-timed ring structure is presented, the basic elements of the ring is a realization of a chaotic cellular automata topology. In particular, the proposed TRNG design is fully synthesizable with standard all-digital components. Test chips of the proposed TRNG structure were fabricated with 40nm TSMC technology node, and the utilized overhead is only 75 NAND gates equivalent, with a die area of 270 μm2. Experimental results demonstrated that the TRNG test chips can generate random numbers at a high bit rate: 1600Mb/s. The test sequences generated by the TRNG test chips passed all test statistics of the widely used test suite: NIST SP800-22, as well as the independent and identically distributed (IID) test of NIST SP800-90B. Scott Best, Xiaolin Xu 0001 |
ICCAD | 2 |
| 2019 | Electronics Supply Chain Integrity Enabled by BlockchainabstractElectronic systems are ubiquitous today, playing an irreplaceable role in our personal lives as well as in critical infrastructures such as power grid, satellite communication, and public transportation. In the past few decades, the security of software running on these systems has received significant attention. However, hardware has been assumed to be trustworthy and reliable "by default" without really analyzing the vulnerabilities in the electronics supply chain. With the rapid globalization of the semiconductor industry, it has become challenging to ensure the integrity and security of hardware. In this paper, we discuss the integrity concerns associated with a globalized electronics supply chain. More specifically, we divide the supply chain into six distinct entities: IP owner/foundry (OCM), distributor, assembler, integrator, end user, and electronics recycler, and analyze the vulnerabilities and threats associated with each stage. To address the concerns of the supply chain integrity, we propose a blockchain-based certificate authority framework that can be used to manage critical chip information such as electronic chip identification (ECID), chip grade, transaction time, etc. The decentralized nature of the proposed framework can mitigate most threats of the electronics supply chain, such as recycling, remarking, cloning, and overproduction. Xiaolin Xu 0001, Fahim Rahman, Bicky Shakya, Apostol Vassilev 0001, Domenic Forte, Mark Tehranipoor |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2018 | Power-based side-channel instruction-level disassemblerabstractModern embedded computing devices are vulnerable against malware and software piracy due to insufficient security scrutiny and the complications of continuous patching. To detect malicious activity as well as protecting the integrity of executable software, it is necessary to monitor the operation of such devices. In this paper, we propose a disassembler based on power-based side-channel to analyze the real-time operation of embedded systems at instruction-level granularity. The proposed disassembler obtains templates from an original device (e.g., IoT home security system, smart thermostat, etc.) and utilizes machine learning algorithms to uniquely identify instructions executed on the device. The feature selection using Kullback-Leibler (KL) divergence and the dimensional reduction using PCA in the time-frequency domain are proposed to increase the identification accuracy. Moreover, a hierarchical classification framework is proposed to reduce the computational complexity associated with large instruction sets. In addition, covariate shifts caused by different environmental measurements and device-to-device variations are minimized by our covariate shift adaptation technique. We implement this disassembler on an AVR 8-bit microcontroller. Experimental results demonstrate that our proposed disassembler can recognize test instructions including register names with a success rate no lower than 99.03% with quadratic discriminant analysis (QDA). Jungmin Park, Xiaolin Xu 0001, Yier Jin, Domenic Forte, Mark Tehranipoor |
DAC | 2 |
| 2018 | SCARe: An SRAM-Based Countermeasure Against IC RecyclingabstractWith the rapid growth of the electronics market, counterfeiting of integrated circuits (ICs), in particular IC recycling, has become a serious issue in recent years. Recycled ICs are those harvested from old systems and resold in the supply chain as new. Such ICs exhibit lower performance and shorter lifetime and, as a result, pose threats to the security and reliability of electronic systems. In this paper, we propose a recycled IC detection framework called static random-access memory (SRAM)-based countermeasure against IC recycling (SCARe) to detect the aging of SRAM cells. Our framework can be applied to both standalone SRAM chips and system on chips with embedded SRAM. For each SRAM under detection, statistical analysis is conducted to differentiate the recycled and new ICs. To mimic the practical aging scenario, 16 commodity SRAM chips from three different manufacturers and different technology nodes (e.g., 90, 110, and 130 nm) are stressed under high-temperature and supply-voltage conditions for different periods of time. The experimental results from new and aged SRAM chips, which represents recycled ICs, demonstrate that our proposed technology can achieve extremely high-detection success rate (no lower than 96.5%). The minimal in-field usage, which can be detected by SCARe, is 7 h. Zimu Guo, Xiaolin Xu 0001, Md Tauhidur Rahman 0001, Mark Tehranipoor, Domenic Forte |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2018 | Bimodal Oscillation as a Mechanism for Autonomous Majority Voting in PUFs
Xiaolin Xu 0001, Shahrzad Keshavarz, Domenic Forte, Mark Tehranipoor, Daniel E. Holcomb |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2017 | Novel Bypass Attack and BDD-based Tradeoff Analysis Against All Known Logic Locking Attacks
Xiaolin Xu 0001, Bicky Shakya, Mark Tehranipoor, Domenic Forte |
CHES | 1 |
| 2017 | FFD: A Framework for Fake Flash DetectionabstractCounterfeit electronics have become a big concern in the globalized semiconductor industry where chips might be recycled, remarked, cloned or overproduced. In this work, we advance the state-of-the-art counterfeit detection of flash memory, which is widely used in electronic systems. Fake memories may be used in critical systems, such as missiles, military aircrafts and helicopters, thus diminishing their reliability. In addition, there are countless stories of fake flash drives in the general consumer market. We propose a comprehensive framework called FFD to detect fake flash memories (i.e., recycled, remarked and cloned parts). FFD is validated with 200,000 commercial flash memory pages. Experimental results show that our framework performs well in: 1) nearly 100% detection accuracy of flash with as little as 5% usage, 2) estimating the flash memory usage with high resolution (≤ 5% of its maximal endurance). Another contribution of this work is a chip ID generation technique that can generate unique flash fingerprints with greater than 99.3% reliability. Zimu Guo, Xiaolin Xu 0001, Mark Tehranipoor, Domenic Forte |
DAC | 2 |
| 2017 | Security Beyond CMOS: Fundamentals, Applications, and RoadmapabstractHardware-oriented security and trust has traditionally relied on the dominant CMOS technology to develop security primitives and provide protection against different attacks and vulnerabilities. With CMOS nearly reaching its fundamental scaling limit and the shortcomings of current solutions, researchers are now looking to exploit emerging nanoelectronic devices for various security applications. In this paper, we discuss the unique features of three emerging nanoelectronic technologies, namely, phase-change memory, grapheme, and carbon nanotubes, and analyze how these features can aid in hardware security and trust. In addition, we present challenges and future research directions about how to effectively integrate emerging nanoscale devices into hardware security. We emphasize that an interdisciplinary initiative is needed for emerging technologies to reach their full potential in security and trust applications. Fahim Rahman, Bicky Shakya, Xiaolin Xu 0001, Domenic Forte, Mark Tehranipoor |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2017 | Poly-Si-Based Physical Unclonable FunctionsabstractPhysically unclonable functions (PUFs) were introduced over a decade ago for a variety of security applications. Silicon PUFs exploit uncontrollable random variations from manufacturing to generate unique and random signatures/ responses. However, such sources of randomness may become limited during standard CMOS manufacturing as processes continue to mature especially with the advances in design for manufacturability. Recently, poly-Si is proposed to improve PUF quality by offering considerable random variations at the materials level, which is from randomly distributed grain boundaries and trapped charges in poly-Si. In this paper, we develop a poly-Si field-effect transistor (FET) model to study the properties of poly-Si-based PUFs under different supply voltages (VDD) and temperatures (T). Simulation results obtained from ring oscillator and arbiter PUFs show that compared with conventional CMOS-based PUFs, the reliability of poly-Si-based PUFs can be improved from around 90% to 98% and the PUF devices are robust against varying VDDand T. Haoting Shen, Fahim Rahman, Bicky Shakya, Xiaolin Xu 0001, Mark Tehranipoor, Domenic Forte |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2016 | A Clockless Sequential PUF with Autonomous Majority VotingabstractPhysical unclonable functions (PUFs) leverage minute silicon process variations to produce device-tied secret keys. The energy and area costs of creating keys from PUFs can far exceed the costs of the basic PUF circuits alone. Minimizing the end-to-end cost of reliable key generation is critical to enable broader adoption of PUFs. In this work, we introduce a new style of PUF that employs autonomous majority voting to improve reliability. The novelty of this design, and the source of its efficiency, is that the inherently sequential majority voting procedure is carried out by a self-timed circuit without orchestration by a global clock. We use circuit simulation to evaluate the energy versus reliability tradeoffs achieved by different parameterizations of the design, to show that the design performs well across a range of supply voltages, and to quantify the robustness of the design across a broad range of operating temperatures. Xiaolin Xu 0001, Daniel E. Holcomb |
ACM Great Lakes Symposium on VLSI | 1 |
| 2015 | Virtual Proofs of Reality and their Physical ImplementationabstractWe discuss the question of how physical statements can be proven over digital communication channels between two parties (a "prover" and a "verifier") residing in two separate local systems. Examples include: (i) "a certain object in the prover's system has temperature X°C", (ii) "two certain objects in the prover's system are positioned at distance X", or (iii) "a certain object in the prover's system has been irreversibly altered or destroyed". As illustrated by these examples, our treatment goes beyond classical security sensors in considering more general physical statements. Another distinctive aspect is the underlying security model: We neither assume secret keys in the prover's system, nor do we suppose classical sensor hardware in his system which is tamper-resistant and trusted by the verifier. Without an established name, we call this new type of security protocol a "virtual proof of reality" or simply a "virtual proof" (VP). In order to illustrate our novel concept, we give example VPs based on temperature sensitive integrated circuits, disordered optical scattering media, and quantum systems. The corresponding protocols prove the temperature, relative position, or destruction/modification of certain physical objects in the prover's system to the verifier. These objects (so-called "witness objects") are prepared by the verifier and handed over to the prover prior to the VP. Furthermore, we verify the practical validity of our method for all our optical and circuit-based VPs in detailed proof-of-concept experiments. Our work touches upon, and partly extends, several established concepts in cryptography and security, including physical unclonable functions, quantum cryptography, interactive proof systems, and, most recently, physical zero-knowledge proofs. We also discuss potential advancements of our method, for example "public virtual proofs" that function without exchanging witness objects between the verifier and the prover. Ulrich Rührmair, J. L. Martinez-Hurtado, Xiaolin Xu 0001, Christian Kraeh, Christian Hilgers, Dima Kononchuk, Jonathan J. Finley, Wayne P. Burleson |
IEEE Symposium on Security and Privacy | 3 |
| 2015 | Reliable Physical Unclonable Functions Using Data Retention Voltage of SRAM CellsabstractPhysical unclonable functions (PUFs) are circuits that produce outputs determined by random physical variations from fabrication. The PUF studied in this paper utilizes the variation sensitivity of static random access memory (SRAM) data retention voltage (DRV), the minimum voltage at which each cell can retain state. Prior work shows that DRV can uniquely identify circuit instances with 28% greater success than SRAM power-up states that are used in PUFs [1]. However, DRV is highly sensitive to temperature, and until now this makes it unreliable and unsuitable for use in a PUF. In this paper, we enable DRV PUFs by proposing a DRV-based hash function that is insensitive to temperature. The new hash function, denoted DRV-based hashing (DH), is reliable across temperatures because it utilizes the temperature-insensitive ordering of DRVs across cells, instead of using the DRVs in absolute terms. To evaluate the security and performance of the DRV PUF, we use DRV measurements from commercially available SRAM chips, and use data from a novel DRV prediction algorithm. The prediction algorithm uses machine learning for fast and accurate simulation-free estimation of any cell's DRV, and the prediction error in comparison to circuit simulation has a standard deviation of 0.35 mV. We demonstrate the DRV PUF using two applications-secret key generation and identification. In secret key generation, we introduce a new circuit-level reliability knob as an alternative to error correcting codes. In the identification application, our approach is compared to prior work and shown to result in a smaller false-positive identification rate for any desired true-positive identification rate. Xiaolin Xu 0001, Amir Rahmati, Daniel E. Holcomb, Kevin Fu, Wayne P. Burleson |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2014 | Efficient Power and Timing Side Channels for Physical Unclonable Functions
Ulrich Rührmair, Xiaolin Xu 0001, Jan Sölter, Ahmed Mahmoud, Mehrdad Majzoobi, Farinaz Koushanfar, Wayne P. Burleson |
CHES | 2 |
| 2014 | Hybrid side-channel/machine-learning attacks on PUFs: A new threat?abstractMachine Learning (ML) is a well-studied strategy in modeling Physical Unclonable Functions (PUFs) but reaches its limits while applied on instances of high complexity. To address this issue, side-channel attacks have recently been combined with modeling techniques to make attacks more efficient [25][26]. In this work, we present an overview and survey of these so-called “hybrid modeling and side-channel attacks” on PUFs, as well as of classical side channel techniques for PUFs. A taxonomy is proposed based on the characteristics of different side-channel attacks. The practical reach of some published side-channel attacks is discussed. Both challenges and opportunities for PUF attackers are introduced. Countermeasures against some certain side-channel attacks are also analyzed. To better understand the side-channel attacks on PUFs, three different methodologies of implementing side-channel attacks are compared. At the end of this paper, we bring forward some open problems for this research area. Xiaolin Xu 0001, Wayne P. Burleson |
DATE | 1 |
| 2013 | PUF Modeling Attacks on Simulated and Silicon DataabstractWe discuss numerical modeling attacks on several proposed strong physical unclonable functions (PUFs). Given a set of challenge-response pairs (CRPs) of a Strong PUF, the goal of our attacks is to construct a computer algorithm which behaves indistinguishably from the original PUF on almost all CRPs. If successful, this algorithm can subsequently impersonate the Strong PUF, and can be cloned and distributed arbitrarily. It breaks the security of any applications that rest on the Strong PUF's unpredictability and physical unclonability. Our method is less relevant for other PUF types such as Weak PUFs. The Strong PUFs that we could attack successfully include standard Arbiter PUFs of essentially arbitrary sizes, and XOR Arbiter PUFs, Lightweight Secure PUFs, and Feed-Forward Arbiter PUFs up to certain sizes and complexities. We also investigate the hardness of certain Ring Oscillator PUF architectures in typical Strong PUF applications. Our attacks are based upon various machine learning techniques, including a specially tailored variant of logistic regression and evolution strategies. Our results are mostly obtained on CRPs from numerical simulations that use established digital models of the respective PUFs. For a subset of the considered PUFs-namely standard Arbiter PUFs and XOR Arbiter PUFs-we also lead proofs of concept on silicon data from both FPGAs and ASICs. Over four million silicon CRPs are used in this process. The performance on silicon CRPs is very close to simulated CRPs, confirming a conjecture from earlier versions of this work. Our findings lead to new design requirements for secure electrical Strong PUFs, and will be useful to PUF designers and attackers alike. Ulrich Rührmair, Jan Sölter, Frank Sehnke, Xiaolin Xu 0001, Ahmed Mahmoud, Vera Stoyanova, Gideon Dror, Jürgen Schmidhuber, Wayne P. Burleson, Srini Devadas |
IEEE Trans. Inf. Forensics Secur. | 4 |