VLDB 2026 Research / reviewers in the wild / expert
Yunsi Fei
dblp:67/5667
· DBLP profile ↗
104ranked-venue papers
7as first author
29since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 70 · 5 first-author · 16 since 2021Security and privacy · 17 · 2 first-author · 9 since 2021Computer networks · 10 · 1 since 2021Software engineering, systems software and programming languages · 9 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 6 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Attack from Shadows: Unsupervised Side-channel Transfer Learning across Devices and ModalitiesabstractIn this work, we focus on unsupervised transfer learning attacks in side-channel analysis, where a passive adversary leverages only unlabeled traces from a target device to recover its secret key. This setting reflects a realistic assumption in practical scenarios where collecting labeled traces from the target device is infeasible, particularly for proprietary devices or cloud computing environments. Prior unsupervised deep learning side-channel analysis (DL-SCA) methods assume that a single domain-invariant model suffices across both source and target domains, but this assumption breaks down as device discrepancy increases, such as across different hardware platforms, side-channel modalities, or implementations. We address this gap with the Unsupervised Transfer Learning Attack (UTLA), which uses a separate encoder with a shared classifier to yield lower DL-SCA loss and improved key extraction fidelity. UTLA keeps the source classifier fixed while training a target-specific encoder, using Maximum Mean Discrepancy (MMD) regularization to align feature distributions. Unlike prior approaches that fail under significant domain mismatch, UTLA enables reliable key recovery across diverse scenarios: (a) different hardware platforms, from a Spartan-6 FPGA to an XMEGA MCU using only 61 traces; (b) different leakage modalities, from power measurements on an XMEGA MCU to cache-timing traces on an x86 processor using 435 traces; and (c) different implementations, from masked and shuffled AES on STM32 (ASCADv2) to masked AES on AVR (ASCADv1) using 1519 traces. Our results reveal a significant security implication: public side-channel datasets can serve as effective attack vectors against unseen devices implementing the same cryptographic algorithm. Saion Kumar Roy, A. Adam Ding, Yunsi Fei |
AsiaCCS | 4 |
| 2026 | Formal Methods-Assisted Chosen Ciphertext Attacks on PQC CRYSTALS-Kyber Using Electromagnetic EmanationsabstractNIST has released a set of post-quantum cryptography (PQC) standards that address the threat posed by the emergence of quantum computing. The standard includes a modular lattice-based key exchange mechanism (ML-KEM) based on the CRYSTALS-Kyber algorithm. Recent work has shown that Kyber is susceptible to electromagnetic (EM) and power side-channel attacks. A full understanding of the side-channel vulnerabilities in Kyber is of paramount importance for next-generation communication and computing infrastructures.In this study, we target a previously unexplored section of the Kyber algorithm and implement a chosen ciphertext side-channel attack. We focus our attack on the Barrett reduction operation in the decapsulation algorithm. Compared to previous attacks on Barrett reduction, which targeted variables after the Inverse-Number Theoretic Transform (INTT), we focus on Barrett reduction on NTT variables, allowing for more general chosen ciphertexts that can evade input sanity checking. We design a scheme that requires only a set of 12 ciphertexts and side-channel EM traces of the corresponding decapsulation processes, which can reveal distinct leakages under different key values. The secret key is retrieved by pattern matching of the EM leakages. We develop an algorithm that utilizes an SMT solver to automatically select a set of ciphertexts. We implement Kyber on an ARM Cortex M4-based microcontroller and launch this new EM side-channel attack. Our results show that the attack achieves a success rate of over 95% in recovering the secret key value. Yashaswini Makaram, Davis Ranney, A. Adam Ding, David Kaeli, Yunsi Fei |
DATE | 5 |
| 2026 | Exploring Side-Channel Protections in Hardware Implementations of PQC ML-KEM Verification
Davis Ranney, Yashaswini Makaram, A. Adam Ding, Yunsi Fei |
DSN | 4 |
| 2025 | EXAM: Exploiting Exclusive System-Level Cache in Apple M-Series SoCs for Enhanced Cache Occupancy AttacksabstractCache occupancy attacks exploit the shared nature of cache hierarchies to infer a victim's activities by monitoring overall cache usage, unlike access-driven cache attacks that focus on specific cache lines or sets.There exists some prior work that target the last-level cache (LLC) of Intel processors, which is inclusive of higher-level caches, and L2 caches of ARM systems.In this paper, we target the System-Level Cache (SLC) of Apple M-series SoCs, which is exclusive to higher-level CPU caches.We address the challenges of the exclusiveness and propose a suite of SLC-cache occupancy attacks, the first of its kind, where an adversary can monitor GPU and other CPU cluster activities from their own CPU cluster.We first discover the structure of SLC in Apple M1 SOC and various policies pertaining to access and sharing through reverse engineering.We propose two attacks against websites.One is a coarse-grained fingerprinting attack, recognizing which website is accessed based on their different GPU memory access patterns monitored through the SLC occupancy channel.The other attack is a fine-grained pixel stealing attack, which precisely monitors the GPU memory usage for rendering different pixels, through the SLC occupancy channel.Third, we introduce a novel screen capturing attack which works beyond webpages, with the monitoring granularity of 57 rows of pixels (there are 1600 rows for the screen).This significantly expands the attack surface, allowing the adversary to retrieve any screen display, posing a substantial new threat to system security.Our findings reveal critical vulnerabilities in Apple's M-series SoCs and emphasize the urgent need for effective countermeasures against cache occupancy attacks in heterogeneous computing environments. Tianhong Xu, A. Adam Ding, Yunsi Fei |
AsiaCCS | 3 |
| 2025 | MoEcho: Exploiting Side-Channel Attacks to Compromise User Privacy in Mixture-of-Experts LLMsabstractThe transformer architecture has become a cornerstone of modern AI, fueling remarkable progress across applications in natural language processing, computer vision, and multi-modal learning.As these models continue to scale explosively for performance, implementation efficiency remains a critical challenge.Mixtureof-Experts (MoE) architectures, selectively activating specialized subnetworks (experts), offer a unique balance between model accuracy and computational cost.However, the adaptive routing in MoE architectures-where input tokens are dynamically directed to specialized experts based on their semantic meaning-inadvertently opens up a new attack surface for privacy breaches.These inputdependent activation patterns leave distinctive temporal and spatial traces in hardware execution, which adversaries could exploit to deduce sensitive user data.In this work, we propose MoEcho (MoE-Echo), discovering a side-channel analysis-based attack surface that compromises user privacy on MoE-based systems.Specifically, in MoEcho, we introduce four novel architectural side-channels on different computing platforms, including Cache Occupancy Channels and Pageout+Reload on CPUs, and Performance Counter and TLB Evict+Reload on GPUs, respectively.Exploiting these vulnerabilities, we propose four attacks that effectively breach user privacy in large-language models (LLMs) and vision-language models (VLMs) based on MoE architectures: Prompt Inference Attack, Response Reconstruction Attack, Visual Inference Attack, and Visual Reconstruction Attack.We evaluate MoEcho on four open-source MoE-based models at different scales, with a specific focus on the DeepSeek architecture.Our end-to-end experiments on both CPUand GPU-deployed MoE models demonstrate a 99.8% success rate in inferring the patient's private inputs in healthcare records and 92.8% in reconstructing LLM responses.MoEcho is the first run-time * These authors contributed equally. Ruyi Ding, Tianhong Xu, A. Adam Ding, Yunsi Fei |
CCS | 5 |
| 2025 | Graph in the Vault: Protecting Edge GNN Inference with Trusted Execution EnvironmentabstractWide deployment of machine learning models on edge devices has rendered the model intellectual property (IP) and data privacy vulnerable. We propose GNNVault, the first secure Graph Neural Network (GNN) deployment strategy based on Trusted Execution Environment (TEE). GNNVault follows the design of “partition-before-training” and includes a private GNN rectifier to complement with a public backbone model. This way, both critical GNN model parameters and the private graph used during inference are protected within secure TEE compartments. Real-world implementations with Intel SGX demonstrate that GNNVault safeguards GNN inference against state-of-the-art link stealing attacks with a negligible accuracy degradation ($\lt 2 \%$). Ruyi Ding, Tianhong Xu, A. Adam Ding, Yunsi Fei |
DAC | 4 |
| 2025 | Probe-Me-Not: Protecting Pre-trained Encoders from Malicious Probing
Ruyi Ding, Tong Zhou 0002, Lili Su, A. Adam Ding, Xiaolin Xu 0001, Yunsi Fei |
NDSS | 6 |
| 2024 | Non-transferable Pruning
Ruyi Ding, Lili Su, A. Adam Ding, Yunsi Fei |
ECCV (86) | 4 |
| 2024 | GraphCroc: Cross-Correlation Autoencoder for Graph Structural ReconstructionabstractGraph-structured data is integral to many applications, prompting the development of various graph representation methods. Graph autoencoders (GAEs), in particular, reconstruct graph structures from node embeddings. Current GAE models primarily utilize self-correlation to represent graph structures and focus on node-level tasks, often overlooking multi-graph scenarios. Our theoretical analysis indicates that self-correlation generally falls short in accurately representing specific graph features such as islands, symmetrical structures, and directional edges, particularly in smaller or multiple graph contexts.To address these limitations, we introduce a cross-correlation mechanism that significantly enhances the GAE representational capabilities. Additionally, we propose the GraphCroc, a new GAE that supports flexible encoder architectures tailored for various downstream tasks and ensures robust structural reconstruction, through a mirrored encoding-decoding process. This model also tackles the challenge of representation bias during optimization by implementing a loss-balancing strategy. Both theoretical analysis and numerical evaluations demonstrate that our methodology significantly outperforms existing self-correlation-based GAEs in graph structure reconstruction. Shijin Duan, Ruyi Ding, A. Adam Ding, Yunsi Fei, Xiaolin Xu 0001 |
NeurIPS | 5 |
| 2024 | Side-Channel-Assisted Reverse-Engineering of Encrypted DNN Hardware Accelerator IP and Attack Surface ExplorationabstractDeep Neural Networks (DNNs) have revolutionized numerous application domains with their unparalleled performance. As the models become larger and more complex, hardware DNN accelerators are increasingly popular. Field-Programmable Gate Array (FPGA)-based DNN accelerators offer near-Application Specific Integrated Circuit (ASIC) efficiency and exceptional flexibility, establishing them as one of the primary hardware platforms for rapidly evolving deep learning implementations, particularly on edge devices. This prominence renders them lucrative targets for attackers. Existing attacks aimed at compromising the confidentiality of DNN models deployed on FPGA DNN accelerators often assume complete knowledge of the accelerators. However, this assumption does not hold for real-world, proprietary, high-performance FPGA DNN accelerators. In this study, we introduce a comprehensive and effective reverse-engineering methodology for demystifying FPGA DNN accelerator soft Intellectual Property (IP) cores. We demonstrate its application on the cutting-edge AMD-Xilinx Deep Learning Processing Unit (DPU). Our method relies on schematic analysis and, innovatively, electromagnetic (EM) side-channel analysis to reveal the data flow and scheduling of the DNN accelerators. To the best of our knowledge, this research is the first successful endeavor to reverse-engineer a commercial encrypted DNN accelerator IP. Moreover, we investigate attack surfaces exposed by the reverse-engineering findings, including the successful recovery of DNN model architectures and extraction of model parameters. These outcomes pose a significant threat to real-world commercial FPGA-DNN acceleration systems. We discuss potential countermeasures and offer recommendations for FPGA-based IP protection. Cheng Gongye, Yukui Luo, Xiaolin Xu 0001, Yunsi Fei |
SP | 4 |
| 2024 | Sub-6-GHz Energy-Detection-Based Fast On-Chip Analog Spectrum Sensing With Learning-Driven Signal ClassificationabstractCognitive communication utilizes transient openings in the spectrum to communicate opportunistically, which is a promising technique to enable more efficient spectrum usage in an increasingly congested spectrum environment. We aim to address two main challenges associated with cognitive communication: (i) spectrum sensing should be fast and energy efficient for processing a large bandwidth in a short time; (ii) the spectrum sensing approach should be able to simultaneously recognize multiple signals that are present. In this paper, we propose to address these challenges with a novel design framework that consists of a fast on-chip spectrum sensing in conjunction with a novel learning-based spectrum analysis model at the edge to enhance the optimizations for spectrum agility. We first utilize a model of a programmable analog-based high-quality factor (Q) on-chip spectrum sensor that is capable of scanning the sub-6 GHz band to detect the spectrum usage in less than 1μs. The proposed spectrum sensor also enhances the energy efficiency of the sensing. To complement the onchip spectrum sensor, a deep learning (DL) model is deployed for a fine-grained signal detection between channels in the 400 MHz to 6 GHz range, which is intended to be executed on edge devices. Simulation results show that the DL model can detect multiple different modulated signals with a mean Intersection-over-Union (IoU) of 86.8% in highly-variable bandwidth and center frequency scenarios. Finally, we present a system-level model of our framework to demonstrate the spectrum sensing and classification in the sub-6 GHz frequency band. Ankit Mittal, Milin Zhang 0002, Thomas Gourousis, Yunsi Fei, Marvin Onabajo, Francesco Restuccia 0001, Aatmesh Shrivastava |
IEEE Internet Things J. | 5 |
| 2024 | A High-Efficiency Power Obfuscation Switched-Capacitor DC-DC Converter ArchitectureabstractSide channel attacks (SCA) have been shown to be very effective in breaking cryptographic engines. In this paper, we present a new power obfuscation switched capacitor (POSC) DC-DC converter. To a first order approximation, it equalizes the charge such that the same amount of charge is drawn from the input power supply in each cycle. We evaluated the design by analyzing the power supply to an Advanced Encryption Standard (AES) unit powered by the proposed converter. CPA fails after evaluation with 10k traces. Two different topologies of the switched capacitor circuit are analyzed for their contribution to side channel power information leakage. The three phase POSC is designed with both switched capacitor converters (SCC1 and SCC2) and achieves efficiency of 77% and 70%. Nikita Mirchandani, Majid Sabbagh, Yunsi Fei, Aatmesh Shrivastava |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2023 | EMShepherd: Detecting Adversarial Samples via Side-channel LeakageabstractDeep Neural Networks (DNN) are vulnerable to adversarial perturbations — small changes crafted deliberately on the input to mislead the model for wrong predictions. Adversarial attacks have disastrous consequences for deep learning empowered critical applications. Existing defense and detection techniques both require extensive knowledge of the model, testing inputs and even execution details. They are not viable for general deep learning implementations where the model internal is unknown, a common ‘black-box’ scenario for model users. Inspired by the fact that electromagnetic (EM) emanations of a model inference are dependent on both operations and data and may contain footprints of different input classes, we propose a framework, EMShepherd, to capture EM traces of model execution, perform processing on traces and exploit them for adversarial detection. Only benign samples and their EM traces are used to train the adversarial detector: a set of EM classifiers and class-specific unsupervised anomaly detectors. When the victim model system is under attack by an adversarial example, the model execution will be different from executions for the known classes, and the EM trace will be different. We demonstrate that our air-gapped EMShepherd can effectively detect different adversarial attacks on a commonly used FPGA deep learning accelerator for both Fashion MNIST and CIFAR-10 datasets. It achieves a detection rate on most types of adversarial samples, which is comparable to the state-of-the-art ‘white-box’ software-based detectors. Ruyi Ding, Cheng Gongye, Siyue Wang, A. Adam Ding, Yunsi Fei |
AsiaCCS | 5 |
| 2023 | HammerDodger: A Lightweight Defense Framework against RowHammer Attack on DNNsabstractRowHammer attacks have become a serious security problem on deep neural networks (DNNs). Some carefully induced bit-flips degrade the prediction accuracy of DNN models to random guesses. This work proposes a lightweight defense framework that detects and mitigates adversarial bit-flip attacks. We employ a dynamic channel-shuffling obfuscation scheme to present moving targets to the attack, and develop a logits-based model integrity monitor with negligible performance loss. The parameters and architecture of DNN models remain unchanged, which ensures lightweight deployment and makes the framework compatible with commodity models. We demonstrate that our framework can protect various DNN models against RowHammer attacks. Cheng Gongye, Yukui Luo, Xiaolin Xu 0001, Yunsi Fei |
DAC | 4 |
| 2023 | Deep-Learning Model Extraction Through Software-Based Power Side-ChannelabstractDeep learning (DL) techniques have been increasingly applied across various applications, facing a growing number of security threats. One such threat is model extraction, an attack that steals the Intellectual Property of DL models, either by recovering the same functionality or retrieving high-fidelity models. Current model extraction methods can be categorized as learning-based or cryptanalytic, with the latter relying on model queries and computational methods to recover parameters. However, these are limited to shallow neural networks and are computationally prohibitive for deeper DL models. In this paper, we propose leveraging software-based power analysis, specifically the Intel Running Average Power Limit (RAPL) technique, for DL model extraction. RAPL allows us to measure power leakage of the most popular activation function, ReLU, through a software interface. Consequently, the ReLU branch direction can be leaked in the software power side-channel, a vulnerability common in many state-of-the-art DL frameworks. We introduce a novel methodology for model extraction Algorithm from input gradient assisted by side channel information. We implement our attack on the oneDNN framework, the most popular library on Intel processors. Compared to prior work, our model extraction, assisted by the software power side-channel, only requires 0.8% of the queries to retrieve as-layer MLP. We also successfully apply our method to a common Convolutional Neural Network (CNN) - Lenet-5. To the best of our knowledge, this is the first work that extracts CNN models with more than 5 layers based solely on queries and software. A. Adam Ding, Yunsi Fei |
ICCAD | 3 |
| 2023 | VertexSerum: Poisoning Graph Neural Networks for Link InferenceabstractGraph neural networks (GNNs) have brought superb performance to various applications utilizing graph structural data, such as social analysis and fraud detection. The graph links, e.g., social relationships and transaction history, are sensitive and valuable information, which raises privacy concerns when using GNNs. To exploit these vulnerabilities, we propose VertexSerum, a novel graph poisoning attack that increases the effectiveness of graph link stealing by amplifying the link connectivity leakage. To infer node adjacency more accurately, we propose an attention mechanism that can be embedded into the link detection network. Our experiments demonstrate that VertexSerum significantly outperforms the SOTA link inference attack, improving the AUC scores by an average of 9.8% across four real-world datasets and three different GNN structures. Furthermore, our experiments reveal the effectiveness of VertexSerum in both black-box and online learning settings, further validating its applicability in real-world scenarios. The source code is available at https://github.com/RollinDing/VertexSerum. Ruyi Ding, Shijin Duan, Xiaolin Xu 0001, Yunsi Fei |
ICCV | 4 |
| 2023 | A Guessing Entropy-Based Framework for Deep Learning-Assisted Side-Channel AnalysisabstractRecently deep-learning (DL) techniques have been widely adopted in side-channel power analysis. A DL-assisted SCA generally consists of two phases: a deep neural network (DNN) training phase and a follow-on attack phase using the trained DNN. However, currently the two phases are not well aligned, as there is no conclusion on what metric used in the training can result in the most effective attack in the second phase. When traditional loss functions such as negative log-likelihood (NLL) are used in training a DNN, the trained model does not yield optimal follow-on attack. Recently some information theoretical SCA leakage metrics are proposed, either as the validation metric to stop the DNN training with traditional loss functions, or as both the validation metric and the training loss function. None of those proposed metrics, however, directly measures the SCA effectiveness. We propose to conduct DNN training directly with a common SCA effectiveness metric, Guessing Entropy (GE). We overcome the prior practical difficulty of using GE in DNN training by utilizing the GEEA estimation algorithm introduced in CHES 2020. We show that using GEEA as either the validation metric or the loss function produces DNN models that lead to much more effective follow-on attacks. Our work consolidates the DL-assisted SCA framework with a consistent metric, which shows great potential to be adopted as the universal SCA-oriented DNN training framework. A. Adam Ding, Yunsi Fei |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2022 | A Cross-Platform Cache Timing Attack Framework via Deep LearningabstractWhile deep learning methods have been adopted in power side-channel analysis, they have not been applied to cache timing attacks due to the limited dimension of cache timing data. This paper proposes a persistent cache monitor based on cache line flushing instructions, which runs concurrently to a victim execution and captures detailed memory access patterns in high-dimensional timing traces. We discover a new cache timing side-channel across both inclusive and non-inclusive caches, different from the traditional “Flush+Flush” timing leakage. We then propose a non-profiling differential deep learning analysis strategy to exploit the cache timing traces for key recovery. We further propose a framework for cross-platform cache timing attack via deep learning. Knowledge learned from profiling a common reference device can be transferred to build models to attack many other victim devices, even in different processor families. We take the OpenSSL AES-128 encryption algorithm as an example victim and deploy an asynchronous cache attack. We target three different devices from Intel, AMD, and ARM processors. We examine various scenarios for assigning the teacher role to one device and the student role to other devices, and evaluate the cross-platform deep-learning attack framework. Experimental results show that this new attack is easily extendable to victim devices and is more effective than attacks without any prior knowledge. Ruyi Ding, Cheng Gongye, Yunsi Fei, A. Adam Ding |
DATE | 5 |
| 2022 | NNReArch: A Tensor Program Scheduling Framework Against Neural Network Architecture Reverse EngineeringabstractArchitecture reverse engineering has become an emerging attack against deep neural network (DNN) implementations. Several prior works have utilized side-channel leakage to recover the model architecture while the an DNN is executing on a hardware acceleration platform. In this work, we target an open-source deep-learning accelerator, Versatile Tensor Accelerator (VTA), and utilize electromagnetic (EM) side-channel leakage to comprehensively learn the association between DNN architecture configurations and EM emanations. We also consider the holistic system–including the low-level tensor program code of the VTA accelerator on a Xilinx FPGA, and explore the effect of such low-level configurations on the EM leakage. Our study demonstrates that both the optimization and configuration of tensor programs will affect the EM side-channel leakage.Gaining knowledge of the association between low-level tensor program and the EM emanations, we propose NNReArch, a lightweight tensor program scheduling framework against side-channel-based DNN model architecture reverse engineering. Specifically, NNReArch targets reshaping the EM traces of different DNN operators, through scheduling the tensor program execution of the DNN model so as to confuse the adversary. NNReArch is a comprehensive protection framework supporting two modes, a balanced mode that strikes a balance between the DNN model confidentiality and execution performance, and a secure mode where the most secure setting is chosen. We implement and evaluate the proposed framework on the open-source VTA with state-of-the-art DNN architectures. The experimental results demonstrate that NNReArch can efficiently enhance the model architecture security with a small performance overhead. In addition, the proposed obfuscation technique makes reverse engineering of the DNN architecture significantly harder. Yukui Luo, Shijin Duan, Cheng Gongye, Yunsi Fei, Xiaolin Xu 0001 |
FCCM | 4 |
| 2022 | Protected ECC Still Leaks: A Novel Differential-Bit Side-channel Power Attack on ECDH and CountermeasuresabstractOver the past decade, a few side-channel attacks (SCAs) and countermeasures against implementations of Elliptic-Curve Cryptography (ECC), commonly used in embedded systems and Internet-of- Things (IoT) devices, have been presented. This work discovers a new side-channel power leakage of an ECDH hardware implementation protected against existing attacks, where the power leakage is not directly related to the key bits, but related to the differential of two consecutive key bits. We propose an unsupervised differential-bit horizontal clustering attack and implement it against an ECDH FPGA implementation. We also comprehensively analyze the related operations and circuits, and identify the root cause of such leakage is due to the different arrival times of inputs to combinational circuits. Such leakage generally exists in ECC hardware implementations, including FPGA and ASIC. We further propose several effective countermeasures to address this new vulnerability and evaluate the implemetations. Tianhong Xu, Cheng Gongye, Yunsi Fei |
ACM Great Lakes Symposium on VLSI | 3 |
| 2022 | Ran$Net: An Anti-Ransomware Methodology based on Cache Monitoring and Deep LearningabstractRansomware has become a serious threat in the cyberspace. Existing software pattern-based malware detectors are specific for certain ransomware and may not capture new variants. Recognizing a common essential behavior of ransomware - employing local cryptographic software for malicious encryption and therefore leaving footprints on the victim machine's caches, this work proposes an anti-ransomware methodology, Ran$Net, based on hardware activities. It consists of a passive cache monitor to log suspicious cache activities, and a follow-on non-profiled deep learning analysis strategy to retrieve the secret cryptographic key from the timing traces generated by the monitor. We implement the first of its kind tool to combat an open-source ransomware and successfully recover the secret key. Ruyi Ding, Cheng Gongye, A. Adam Ding, Yunsi Fei |
ACM Great Lakes Symposium on VLSI | 6 |
| 2022 | High-Precision Nano-Amp Current Sensor and Obfuscation based Analog Trojan Detection CircuitabstractEmerging Analog Trojans such as A2, large-delay Trojans, and row-hammer have been shown to be more stealthier than previously known digital Trojans. They are smaller sized, do not rely on inputs for triggering, and the trigger for their payload can be made arbitrarily delayed, like a ticking time bomb. Furthermore, analog Trojans can easily evade detection due to their novel nature and incompatibility with the digital design and validation flow. In this paper, we propose a current signature-based detection scheme, which can effectively catch various analog Trojans at both run-time and production time validation. The paper includes techniques that advance Trojan detection method through incorporating detection of transient variation in the power supply current. Proposed current-sensor can sense currents down to 10s of nano-Amps improving over prior power sensing based techniques. Further, a configurable design of current sensor is developed to enable large range sensing capability. The design is also developed to be compatible with the digital design flow and can be logic obfuscated. This detection method can be used at run-time to potentially fence off activation of analog Trojans in the field through early warning signals. The commercial 65nm CMOS technology is utilized to verify the proposed idea. Mostafa Abedi, Tiancheng Yang, Yunsi Fei, Aatmesh Shrivastava |
ISCAS | 3 |
| 2022 | Masking Feedforward Neural Networks Against Power Analysis AttacksabstractAbstract Recent advances in machine learning have enabled Neural Network (NN) inference directly on constrained embedded devices. This local approach enhances the privacy of user data, as the inputs to the NN inference are not shared with third-party cloud providers over a communication network. At the same time, however, performing local NN inference on embedded devices opens up the possibility of Power Analysis attacks, which have recently been shown to be effective in recovering NN parameters, as well as their activations and structure. Knowledge of these NN characteristics constitutes a privacy threat, as it enables highly effective Membership Inference and Model Inversion attacks, which can recover information about the sensitive data that the NN model was trained on. In this paper we address the problem of securing sensitive NN inference parameters against Power Analysis attacks. Our approach employs masking, a countermeasure well-studied in the context of cryptographic algorithms. We design a set of gadgets, i.e., masked operations, tailored to NN inference. We prove our proposed gadgets secure against power attacks and show, both formally and experimentally, that they are composable, resulting in secure NN inference. We further propose optimizations that exploit intrinsic characteristics of NN inference to reduce the masking’s runtime and randomness requirements. We empirically evaluate the performance of our constructions, showing them to incur a slowdown by a factor of about 2–5. Konstantinos Athanasiou, Thomas Wahl, A. Adam Ding, Yunsi Fei |
Proc. Priv. Enhancing Technol. | 4 |
| 2021 | Intrinsic Examples: Robust Fingerprinting of Deep Neural Networks
Siyue Wang, Pu Zhao 0001, Xiao Wang 0028, Sang (Peter) Chin, Thomas Wahl, Yunsi Fei, Qi Alfred Chen, Xue Lin 0001 |
BMVC | 6 |
| 2021 | DeepStrike: Remotely-Guided Fault Injection Attacks on DNN Accelerator in Cloud-FPGAabstractAs Field-programmable gate arrays (FPGAs) are widely adopted in clouds to accelerate Deep Neural Networks (DNN), such virtualization environments have posed many new security issues. This work investigates the integrity of DNN FPGA accelerators in clouds. It proposes DeepStrike, a remotely-guided attack based on power glitching fault injections targeting DNN execution. We characterize the vulnerabilities of different DNN layers against fault injections on FPGAs and leverage time-to-digital converter (TDC) sensors to precisely control the timing of fault injections. Experimental results show that our proposed attack can successfully disrupt the FPGA DSP kernel and misclassify the target victim DNN application. Yukui Luo, Cheng Gongye, Yunsi Fei, Xiaolin Xu 0001 |
DAC | 3 |
| 2021 | Trident: A Hybrid Correlation-Collision GPU Cache Timing Attack for AES Key RecoveryabstractGiven the parallel processing capabilities of Graphics Processing Units (GPUs), many applications are exploiting GPUs and cryptographic systems have also begun to leverage GPUs to accelerate encryption/decryption. Recent work has identified how microarchitectural side-channel attacks can be carried out on AES (Advanced Encryption Standard) by exploiting the SIMT characteristics and memory coalescing of GPUs. In this work, we first show that previously proposed correlation-based side-channel attacks are not feasible on modern GPUs that support narrower data-cache accesses via a sectored-cache microarchitecture-resulting in memory accesses from different levels of the memory hierarchy. In comparison, we identify how negative timing correlation can occur in modern GPUs when data is fetched from different levels of the cache hierarchy. We then propose Trident - a hybrid cache-collision timing attack on GPUs that can fully recover all AES key bytes on modern GPUs. Cache collisions in GPUs present challenges due to the large number of threads and the number of samples required. To address these challenges, Trident consists of three different components - negative timing correlation, cache-collision attack, and chosen plaintext attack. We leverage the negative timing correlation to recover earlier key bytes of AES while exploiting cache-collision attacks for the latter AES key bytes. To enable GPU cache collision attacks, we exploit memory coalescing to control the number of memory accesses through chosen-plaintext attacks to significantly reduce the number of timing samples needed. Our proposed Trident attack results in over 10× reduction in the number of samples needed to recover the key bytes compared with prior work, while still being successful in full AES key recovery in modern GPUs. We also propose TridentShield - a latency-based countermeasure to the Trident attack that minimizes throughput degradation in GPUs. Jaeguk Ahn, Cheolgyu Jin, Minsoo Rhu, Yunsi Fei, David R. Kaeli, John Kim 0001 |
HPCA | 5 |
| 2021 | GPU Overdrive Fault Attacks on Neural NetworksabstractGraphics processing units (GPUs) are commonly used to accelerate training and inference of deep neural networks (DNNs). Modern cloud nodes are shared by multiple users to execute workloads concurrently. However, the reliability and security of sharing the heterogeneous CPU-GPU have not been carefully evaluated. In this paper, we thoroughly characterize fault injections and propagation in a victim convolutional neural network (CNN) on a GPU, and analyze the controllability of the attack. We successfully launch an end-to-end misclassification attack during CNN inferences with careful timing control. Majid Sabbagh, Yunsi Fei, David R. Kaeli |
ICCAD | 2 |
| 2021 | Introduction to the Special Issue on Emerging Challenges and Solutions in Hardware Securityabstractintroduction Introduction to the Special Issue on Emerging Challenges and Solutions in Hardware Security Share on Editors: Domenic Forte View Profile , Debdeep Mukhopadhyay View Profile , Ilia Polian View Profile , Yunsi Fei View Profile , Rosario Cammarota View Profile Authors Info & Claims ACM Journal on Emerging Technologies in Computing SystemsVolume 17Issue 3July 2021 Article No.: 29pp 1–4https://doi.org/10.1145/3464326Online:30 June 2021Publication History 0citation108DownloadsMetricsTotal Citations0Total Downloads108Last 12 Months108Last 6 weeks4 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Domenic Forte, Debdeep Mukhopadhyay, Ilia Polian, Yunsi Fei, Rosario Cammarota |
ACM J. Emerg. Technol. Comput. Syst. | 4 |
| 2021 | Large Delay Analog Trojans: A Silent Fabrication-Time Attack Exploiting Analog ModalitiesabstractThis article presents large delay-based analog Trojan circuits, a new class of analog Trojans that can be interfaced with digital and analog macros to launch fabrication-time hardware attacks. Two different circuit topologies of analog Trojan are presented, which can generate a delayed trigger output after two days and 60 ms, respectively, when implemented in 65-nm CMOS technology. The large delay is achieved using the transistor's gate-oxide leakage current or a diode's reverse saturation current in combination with the Miller capacitance-based circuits. The proposed analog Trojans can operate across multiple on-chip power domains and can be launched without any digital input signal, making their detection challenging. They show very limited variation in side-channel parameters, which makes them harder to detect through side-channel analysis. In addition, the proposed designs have a small area footprint of 55.5 μm2and 28 μm2, respectively, and can be easily concealed on-chip. We also demonstrate an attack launched using these Trojans to construct a “kill-switch” that disables the power management unit of an IC. Process and temperature variations were also investigated to assess their impact on the design. We implemented the thick-oxide gate leakage modeling to study the robustness of the proposed Trojan design. We also present the long-term potential threat of these Trojans where the output trigger signal is generated after an even larger delay. Tiancheng Yang, Ankit Mittal, Yunsi Fei, Aatmesh Shrivastava |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2020 | Reverse-Engineering Deep Neural Networks Using Floating-Point Timing Side-ChannelsabstractTrained Deep Neural Network (DNN) models have become valuable intellectual property. A new attack surface has emerged for DNNs: model reverse engineering. Several recent attempts have utilized various common side channels. However, recovering DNN parameters, weights and biases, remains a challenge. In this paper, we present a novel attack that utilizes a floating-point timing side channel to reverse-engineer parameters of multi-layer perceptron (MLP) models in software implementation, entirely and precisely. To the best of our knowledge, this is the first work that leverages a floating-point timing side-channel for effective DNN model recovery. Cheng Gongye, Yunsi Fei, Thomas Wahl |
DAC | 2 |
| 2020 | A Novel GPU Overdrive Fault AttackabstractGraphics processing units (GPUs) are widely used to accelerate applications including cryptographic operations. The reliability and security of GPUs have become a concern. Prior work reported power and timing side-channel attacks on GPUs. In this paper, we present the first-ever overdrive fault attack targeting modern GPUs. This attack exploits voltage-frequency scaling features present on most commercial GPUs to introduce random faults during kernel execution. We demonstrate an effective fault-based attack on an AMD GPU, recovering the AES keys in minutes. Such software-controlled fault injections also pose serious threats to data integrity and service availability in the cloud. Majid Sabbagh, Yunsi Fei, David R. Kaeli |
DAC | 2 |
| 2020 | Towards Secure Composition of Integrated Circuits and Electronic Systems: On the Role of EDAabstractModern electronic systems become evermore complex, yet remain modular, with integrated circuits (ICs) acting as versatile hardware components at their heart. Electronic design automation (EDA) for ICs has focused traditionally on power, performance, and area. However, given the rise of hardware-centric security threats, we believe that EDA must also adopt related notions like secure by design and secure composition of hardware. Despite various promising studies, we argue that some aspects still require more efforts, for example: effective means for compilation of assumptions and constraints for security schemes, all the way from the system level down to the "bare metal"; modeling, evaluation, and consideration of security-relevant metrics; or automated and holistic synthesis of various countermeasures, without inducing negative cross-effects.In this paper, we first introduce hardware security for the EDA community. Next we review prior (academic) art for EDA-driven security evaluation and implementation of countermeasures. We then discuss strategies and challenges for advancing research and development toward secure composition of circuits and systems. Johann Knechtel, Elif Bilge Kavun, Francesco Regazzoni 0001, Annelie Heuser, Anupam Chattopadhyay, Debdeep Mukhopadhyay, Soumyajit Dey, Yunsi Fei, Yaacov Belenky, Itamar Levi, Tim Güneysu, Patrick Schaumont, Ilia Polian |
DATE | 8 |
| 2020 | New Passive and Active Attacks on Deep Neural Networks in Medical ApplicationsabstractSecurity of deep neural network (DNN) inference engines, i.e., trained DNN models on various platforms, has become one of the biggest challenges in deploying artificial intelligence in domains where privacy, safety, and reliability are of paramount importance, such as in medical applications. In addition to classic software attacks such as model inversion and evasion attacks, recently a new attack surface---implementation attacks which include both passive side-channel attacks and active fault injection and adversarial attacks---is arising, targeting implementation peculiarities of DNN to breach their confidentiality and integrity. This paper presents several novel passive and active attacks on DNN we have developed and tested over medical datasets. Our new attacks reveal a largely under-explored attack surface of DNN inference engines. Insights gained during attack exploration will provide valuable guidance for effectively protecting DNN execution against reverse-engineering and integrity violations. Cheng Gongye, Hongjia Li 0003, Majid Sabbagh, Geng Yuan, Xue Lin 0001, Thomas Wahl, Yunsi Fei |
ICCAD | 8 |
| 2020 | Stealthy-Shutdown: Practical Remote Power Attacks in Multi - Tenant FPGAsabstractWith the deployment of artificial intelligent (AI) algorithms in a large variety of applications, there creates an increasing need for high-performance computing capabilities. As a result, different hardware platforms have been utilized for acceleration purposes. Among these hardware-based accelerators, the field-programmable gate arrays (FPGAs) have gained a lot of attention due to their re-programmable characteristics, which provide customized control logic and computing operators. For example, FPGAs have recently been adopted for on-demand cloud services by the leading cloud providers like Amazon and Microsoft, providing acceleration for various compute-intensive tasks. While the co-residency of multiple tenants on a cloud FPGA chip increases the efficiency of resource utilization, it also creates unique attack surfaces that are under-explored. In this paper, we exploit the vulnerability associated with the shared power distribution network on cloud FPGAs. We present a stealthy power attack that can be remotely launched by a malicious tenant, shutting down the entire chip and resulting in denial-of-service for other co-located benign tenants. Specifically, we propose stealthy-shutdown: a well-timed power attack that can be implemented in two steps: (1) an attacker monitors the realtime FPGA power-consumption detected by ring-oscillator-based voltage sensors, and (2) when capturing high power-consuming moments, i.e., the power consumption by other tenants is above a certain threshold, she/he injects a well-timed power load to shut down the FPGA system. Note that in the proposed attack strategy, the power load injected by the attacker only accounts for a small portion of the overall power consumption; therefore, such attack strategy remains stealthy to the cloud FPGA operator. We successfully implement and validate the proposed attack on three FPGA evaluation kits with running real-world applications. The proposed attack results in a stealthy-shutdown, demonstrating severe security concerns of co-tenancy on cloud FPGAs. We also offer two countermeasures that can mitigate such power attacks. Yukui Luo, Cheng Gongye, Shaolei Ren, Yunsi Fei, Xiaolin Xu 0001 |
ICCD | 4 |
| 2020 | Correlation Power Analysis and Higher-Order Masking Implementation of WAGE
Yunsi Fei, Guang Gong, Cheng Gongye, Kalikinkar Mandal, Raghvendra Rohit 0001, Tianhong Xu, Yunjie Yi, Nusa Zidaric |
SAC | 1 |
| 2020 | Exploiting Bank Conflict-based Side-channel Timing Leakage of GPUsabstractTo prevent information leakage during program execution, modern software cryptographic implementations target constant-time function, where the number of instructions executed remains the same when program inputs change. However, the underlying microarchitecture behaves differently when processing different data inputs, impacting the execution time of the same instructions. These differences in execution time can covertly leak confidential information through a timing channel. Given the recent reports of covert channels present on commercial microprocessors, a number of microarchitectural features on CPUs have been re-examined from a timing leakage perspective. Unfortunately, a similar microarchitectural evaluation of the potential attack surfaces on GPUs has not been adequately performed. Several prior work has considered a timing channel based on the behavior of a GPU’s coalescing unit. In this article, we identify a second finer-grained microarchitectural timing channel, related to the banking structure of the GPU’s Shared Memory. By considering the timing channel caused by Shared Memory bank conflicts, we have developed a differential timing attack that can compromise table-based cryptographic algorithms. We implement our timing attack on an Nvidia Kepler K40 GPU and successfully recover the complete 128-bit encryption key of an Advanced Encryption Standard (AES) GPU implementation using 900,000 timing samples. We also evaluate the scalability of our attack method by attacking an implementation of the AES encryption algorithm that fully occupies the compute resources of the GPU. We extend our timing analysis onto other Nvidia architectures: Maxwell, Pascal, Volta, and Turing GPUs. We also discuss countermeasures and experiment with a novel multi-key implementation, evaluating its resistance to our side-channel timing attack and its associated performance overhead. Zhen Hang Jiang, Yunsi Fei, David R. Kaeli |
ACM Trans. Archit. Code Optim. | 2 |
| 2019 | Fault Sneaking Attack: a Stealthy Framework for Misleading Deep Neural NetworksabstractDespite the great achievements of deep neural networks (DNNs), the vulnerability of state-of-the-art DNNs raises security concerns of DNNs in many application domains requiring high reliability. We propose the fault sneaking attack on DNNs, where the adversary aims to misclassify certain input images into any target labels by modifying the DNN parameters. We apply ADMM (alternating direction method of multipliers) for solving the optimization problem of the fault sneaking attack with two constraints: 1) the classification of the other images should be unchanged and 2) the parameter modifications should be minimized. Specifically, the first constraint requires us not only to inject designated faults (misclassifications), but also to hide the faults for stealthy or sneaking considerations by maintaining model accuracy. The second constraint requires us to minimize the parameter modifications (using ℓ0 norm to measure the number of modifications and ℓ2 norm to measure the magnitude of modifications). Comprehensive experimental evaluation demonstrates that the proposed framework can inject multiple sneaking faults without losing the overall test accuracy performance. Pu Zhao 0001, Siyue Wang, Cheng Gongye, Yanzhi Wang 0001, Yunsi Fei, Xue Lin 0001 |
DAC | 5 |
| 2019 | Side-channel Timing Attack of RSA on a GPUabstractTo increase computation throughput, general purpose Graphics Processing Units (GPUs) have been leveraged to accelerate computationally intensive workloads. GPUs have been used as cryptographic engines, improving encryption/decryption throughput and leveraging the GPU’s Single Instruction Multiple Thread (SIMT) model. RSA is a widely used public-key cipher and has been ported onto GPUs for signing and decrypting large files. Although performance has been significantly improved, the security of RSA on GPUs is vulnerable to side-channel timing attacks and is an exposure overlooked in previous studies. GPUs tend to be naturally resilient to side-channel attacks, given that they execute a large number of concurrent threads, performing many RSA operations on different data in parallel. Given the degree of parallel execution on a GPU, there will be a significant amount of noise introduced into the timing channel given the thousands of concurrent threads executing concurrently. In this work, we build a timing model to capture the parallel characteristics of an RSA public-key cipher implemented on a GPU. We consider optimizations that include using Montgomery multiplication and sliding-window exponentiation to implement cryptographic operations. Our timing model considers the challenges of parallel execution, complications that do not occur in single-threaded computing platforms. Based on our timing model, we launch successful timing attacks on RSA running on a GPU, extracting the private key of RSA. We also present an effective error detection and correction mechanism. Our results demonstrate that GPU acceleration of RSA is vulnerable to side-channel timing attacks. We propose several countermeasures to defend against this class of attacks. Yunsi Fei, David R. Kaeli |
ACM Trans. Archit. Code Optim. | 2 |
| 2019 | Comprehensive Side-Channel Power Analysis of XTS-AESabstractXTS-advanced encryption standard (AES) is an advanced mode of AES for data protection of sector-based devices. It features two secret keys instead of one, and an additional tweak for each data block. These characteristics make the mode not only resistant against cryptoanalysis attacks, but also more challenging for side-channel attack. In this paper, we comprehensively analyze the side-channel power leakage of various XTS-AES implementations and invent effective attacks. We first run a simple power analysis of a software implementation. For a hardware implementation on field-programmable gate array (FPGA), we analyze side-channel leakage of the particular modular multiplication in XTS-AES mode. In addition, we utilize the relationship between two consecutive block tweaks and propose a method to work around the masking of ciphertext by the tweak. These attacks are verified on an FPGA implementation of XTS-AES. The results show that XTS-AES is susceptible to side-channel power analysis attacks, and therefore dedicated protections are required for security of XTS-AES in storage devices. Yunsi Fei, A. Adam Ding, Pau Closas |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2019 | Efficient Nonprofiling 2nd-Order Power Analysis on Masked Devices Utilizing Multiple Leakage Pointsabstract2nd-order attacks utilize power values at two leakage points to break cryptographic systems protected by 1st-order random masking. Without profiling, the attacker do not know the exact location of the two leakage points. Standard 2nd-order attacks with an exhaustive search over two windows of size nw has computational complexity O(nw2) and does not scale well with the window size nw. We propose to apply a decision-combination attack, the majority vote (MV) attack, to combine 2nd order attacks at multiple candidate pairs of leakage points selected through two filters. The first filter pre-process the power traces with Fast Fourier Transformation (FFT) techniques and reduce the complexity to O(nwlog2(nw)). The second filter use an advanced statistical feature selection procedure, Higher Criticism (HC), to select leakage candidates that improve the effectiveness of decision-combination MV attack and other leakage-combination attacks. We derive theoretical success conditions of MV attacks as well as the typical maximum attack and a leakage-combination sum attack. The theoretical conditions are confirmed through performance comparisons of the attacks on synthetic data sets and on two real data sets, an FPGA implementation and a software implementation of masked AES. The proposed FF-HC-MV attack is data-adaptive, working well in all data sets. A. Adam Ding, Yunsi Fei, Pei Luo |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2018 | GPU acceleration of RSA is vulnerable to side-channel timing attacksabstractThe RSA algorithm [21] is a public-key cipher widely used in digital signatures and Internet protocols, including the Security Socket Layer (SSL) and Transport Layer Security (TLS). RSA entails excessive computational complexity compared with symmetric ciphers. For scenarios where an Internet domain is handling a large number of SSL connections and generating digital signatures for a large number of files, the amount of RSA computation becomes a major performance bottleneck. With the advent of general-purpose GPUs, the performance of RSA has been improved significantly by exploiting parallel computing on a GPU [9], [18], [23], [26], leveraging the Single Instruction Multiple Thread (SIMT) model. Yunsi Fei, David R. Kaeli |
ICCAD | 2 |
| 2018 | Effective simple-power analysis attacks of elliptic curve cryptography on embedded systemsabstractElliptic Curve Cryptography (ECC), initially proposed by Koblitz [17] and Miller [20], is a public-key cipher. Compared with other popular public-key ciphers (e.g., RSA), ECC features a shorter key length for the same level of security. For example, a 256-bit ECC cipher provides 128-bit security, equivalent to a 2048-bit RSA cipher [4]. Using smaller keys, ECC requires less memory for performing cryptographic operations. Embedded systems, especially given the proliferation of Internet-of-Things (IoT) devices and platforms, require efficient and low-power secure communications between edge devices and gateways/clouds. ECC has been widely adopted in IoT systems for authentication of communications, while RSA, which is much more costly to compute, remains the standard for desktops and servers. Yunsi Fei, David R. Kaeli |
ICCAD | 2 |
| 2018 | SCADET: a side-channel attack detection tool for tracking prime+probeabstractMicroarchitectural side-channel attacks have posed serious threats to many computing systems, ranging from embedded systems and mobile devices to desktop workstations and cloud servers. Such attacks exploit side-channel vulnerabilities stemming from fundamental microarchitectural performance features, including the most common caches, out-of-order execution (for the newly revealed Meltdown exploit), and speculative execution (for Spectre). Prior efforts have focused on identifying and assessing these security vulnerabilities, and designing and implementing countermeasures against them. However, the efforts aiming at detecting specific side-channel attacks tend to be narrowly focused, which can make them effective but also makes them obsolete very quickly. In this paper, we propose a new methodology for detecting microarchitectural side-channel attacks that has the potential for a wide scope of applicability, as we demonstrate using a case study involving the Prime+Probe attack family. Instead of looking at the side-effects of side-channel attacks on microarchitectural elements such as hardware performance counters, we target the high-level semantics and invariant patterns of these attacks. We have applied our method to different Prime+Probe attack variants on the instruction cache, data cache, and last-level cache, as well as several benign programs as benchmarks. The method can detect all of the Prime+Probe attack variants with a true positive rate of 100% and an average false positive rate of 7.4%. Majid Sabbagh, Yunsi Fei, Thomas Wahl, A. Adam Ding |
ICCAD | 2 |
| 2018 | A Timing Side-Channel Attack on a Mobile GPUabstractMobile devices are quickly becoming powerful computing platforms in many respects. Given the growing resource demands of applications, compute-heavy workloads on today's smartphone devices are offloaded to the on-board GPU for performance and power efficiency. Mobile devices carry a significant amount of sensitive and personal data, including credit/banking transactions, medical records and passwords. They are frequent targets for attackers, working to obtain an individual's personal information. Although there has been a significant amount of work focused on improving mobile device information security, there has been limited attention paid to the vulnerability of side-channel attacks on these devices, especially their on-board GPUs. In this paper, we present our work on timing side channel vulnerability, launched on a popular mobile device's GPU, exploiting its cache behavior. We target AES-128 encryption, and show that we can successfully recover the full encryption key when using known ciphertext by exploiting timing information. While we target a Qualcomm Snapdragon platform, our statistical analysis shows that our approach is a general method that can be applied to similar mobile platforms. Elmira Karimi, Zhen Hang Jiang, Yunsi Fei, David R. Kaeli |
ICCD | 3 |
| 2018 | Algebraic Fault Analysis of SHA-3 Under Relaxed Fault ModelsabstractAs the new hash standard, Keccak-based secure hash function (SHA-3) will be used in various cryptographic applications. Its security will be of paramount importance to the systems built on top of it. This paper proposes efficient algebraic fault analysis (AFA) methods, and for the first time, applies them to all four modes of SHA-3 under relaxed fault models. Our AFA utilizes the clear algebraic properties of Keccak operations and is very suitable for the fault analysis of SHA-3. Both our analysis and experimental results show that the proposed AFA method is more efficient than the traditional differential fault analysis (DFA) under the single-byte fault model, requiring much fewer faults to recover a whole internal state of the hashing computation. Meanwhile, as AFA is able to exploit all the information available, it can be applied to SHA-3 modes with shorter digests and under more relaxed fault models, where often times the DFA method fails. Our results show that AFA can successfully break all the four SHA-3 modes under a 16-bit fault model, and break SHA3-512 under an even more relaxed fault model, 32-bit fault, all within several minutes. The successful AFA on SHA-3 demonstrates the vulnerability of Keccak algorithms to fault analysis, calling for protections against fault injection and fault analysis. Pei Luo, Konstantinos Athanasiou, Yunsi Fei, Thomas Wahl |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2017 | Towards Sound and Optimal Leakage Detection Procedure
A. Adam Ding, François Durvaux, François-Xavier Standaert, Yunsi Fei |
CARDIS | 5 |
| 2017 | Algebraic fault analysis of SHA-3abstractThis paper presents an efficient algebraic fault analysis on all four modes of SHA-3 under relaxed fault models. This is the first work to apply algebraic techniques on fault analysis of SHA-3. Results show that algebraic fault analysis on SHA-3 is very efficient and effective due to the clear algebraic properties of Keccak operations. Comparing with previous work on differential fault analysis of SHA-3, algebraic fault analysis can identify the injected faults with much higher rates, and recover an entire internal state of the penultimate round with much fewer fault injections. Pei Luo, Konstantinos Athanasiou, Yunsi Fei, Thomas Wahl |
DATE | 3 |
| 2017 | Side-channel power analysis of XTS-AESabstractXTS-AES is an advanced mode of AES for data protection of sector-based devices. Compared to other AES modes, it features two secret keys instead of one, and an additional tweak for each data block. These characteristics make the mode not only resistant against cryptoanalysis attacks, but also more challenging for side-channel attack. In this paper, we propose two attack methods on XTS-AES overcoming these challenges. In the first attack, we analyze side-channel leakage of the particular modular multiplication in XTS-AES mode. In the second one, we utilize the relationship between two consecutive block tweaks and propose a method to work around the masking of ciphertext by the tweak. These attacks are verified on an FPGA implementation of XTS-AES. The results show that XTS-AES is susceptible to side-channel power analysis attacks, and therefore dedicated protections are required for security of XTS-AES in storage devices. Yunsi Fei, A. Adam Ding |
DATE | 2 |
| 2017 | A Novel Side-Channel Timing Attack on GPUsabstractTo avoid information leakage during program execution, modern software implementations of cryptographic algorithms target constant timing complexity, i.e., the number of instructions executed does not vary with different inputs. However, many times the underlying microarchitecture behaves differently when processing varying data inputs, which covertly leaks confidential information through the timing channel. In this paper, we exploit a novel fine-grained microarchitectural timing channel, stalls that occur due to bank conflicts in a GPU's shared memory. Using this attack surface, we develop a differential timing attack that can compromise table-based cryptographic algorithms. We implement our timing attack on an Nvidia Kepler K40 GPU, and successfully recover the complete 128-bit AES encryption key using 10 million samples. We also evaluate the scalability of our attack method by attacking a 8192-thread implementation of the AES encryption algorithm, recovering some key bytes using 1 million samples. Zhen Hang Jiang, Yunsi Fei, David R. Kaeli |
ACM Great Lakes Symposium on VLSI | 2 |
| 2017 | A novel cache bank timing attackabstractTo avoid information leakage through execution, modern software implementations of cryptographic algorithms target constant timing complexity, i.e., the number of instructions does not vary with different inputs. However, often times, the underlying microarchitecture behaves differently under different data inputs, which covertly leaks confidential information through the timing channel. Cache timing channel due to cache miss penalties has been explored in recent years to break system security. In this paper, we exploit a finer-grained L1 cache bank timing channel, the stalling delay due to cache bank conflicts, and develop a new timing attack against table lookup-based cryptographic algorithms. We implement the timing attack with three different methods on Sandy Bridge micro-architecture, and successfully recover the complete 128-bit AES encryption key. The most effective attack can achieve 50% success rate using 75,000 samples and 100% success rate using 200,000 samples. The whole attack process from collecting samples to recoverying all key bytes takes less than 3 minutes. We anticipate the new timing attack to be a threat to various platforms, including ARM-based smart phones and performance-critical accelerators like GPUs. Zhen Hang Jiang, Yunsi Fei |
ICCAD | 2 |
| 2017 | Compiler-Assisted Threshold Implementation against Power Analysis AttacksabstractSide-channel attack utilizes side-channel leakages to extract the secret in crypto systems. Various countermeasures for different algorithms and platforms have been proposed to protect crypto systems against such attacks. Manual countermeasure design requires deep understanding of the target algorithm and implementation, and oftentimes is platform-specific and error-prone. In this paper, we propose the construction of Threshold Implementation (TI), a provably secure countermeasure against power attacks, as an automated compiler pass in the open LLVM (Low Level Virtual Machine) framework. Attack results show that the automatically generated TI designs are secure against power attacks. As our proposed scheme implements the countermeasure at the intermediate representation (IR) level, our method can be applied to any cipher software in any programming language, and the generated implementations can be ported to different platforms and architectures. Pei Luo, Konstantinos Athanasiou, Zhen Hang Jiang, Yunsi Fei, A. Adam Ding, Thomas Wahl |
ICCD | 5 |
| 2017 | Embedded Device Forensics and SecurityabstractWhile the increasing digitalization of our society and amalgamation of embedded devices into the ever-increasing facets of our daily life (e.g., in smart and intelligent vehicles, smart cities and smart nations, and critical infrastructure sectors) have resulted in improved productivity and quality of life, the trend has also resulted in a trend of increasing frequency and sophistication of cyber exploitation and cyber threats. Hence, there is a need for coordinated efforts from the research community to address resulting concerns using both cryptographic and non-cryptographic solutions, such as those presented in this special section. Kim-Kwang Raymond Choo, Yunsi Fei, Yang Xiang 0001, Yu Yu 0001 |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2017 | Vehicle Speed Prediction by Two-Level Data Driven Models in Vehicular NetworksabstractVehicle speed prediction provides important information for many intelligent vehicular and transportation applications. Accurate on-road vehicle speed prediction is challenging, because an individual vehicle speed is affected by many factors, e.g., the traffic condition, vehicle type, and driver's behavior, in either deterministic or stochastic way. This paper proposes a novel data-driven vehicle speed prediction method in the context of vehicular networks, in which the real-time traffic information is accessible and utilized for vehicle speed prediction. It first predicts the average traffic speeds of road segments by using neural network models based on historical traffic data. Hidden Markov models (HMMs) are then utilized to present the statistical relationship between individual vehicle speeds and the traffic speed. Prediction for individual vehicle speeds is realized by applying the forward-backward algorithm on HMMs. To evaluate the prediction performance, simulations are set up in the SUMO microscopic traffic simulator with the application of a real Luxembourg motorway network and traffic count data. The vehicle speed prediction result shows that our proposed method outperforms other ones in terms of prediction accuracy. Bingnan Jiang, Yunsi Fei |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2017 | TARS: A Traffic-Adaptive Receiver-Synchronized MAC Protocol for Underwater Sensor NetworksabstractEfficient medium access control (MAC) is desirable for underwater sensor networks (UWSNs). However, designing an efficient underwater MAC protocol is challenging due to the long propagation delay of the underwater acoustic channel and spatial-temporal uncertainty. In this article, we propose a novel Traffic-Adaptive Receiver-Synchronized underwater MAC protocol, TARS, for throughput maximization. We divide time into equal-sized slots, each the size of one packet transmission time plus a guard time to cope with network dynamics. We adjust the packet transmission phase in a slot, determined by the sender-receiver distance, to align packet receptions for collision reduction. Both the sound propagation speed variation and the node mobility are considered in setting the transmission phase and slot size. We employ a queue-aware utility-optimization framework to determine the optimal transmission strategies dynamically, taking into account both the interference and data queue status. Extensive simulation results show that compared to the existing representative protocols, TARS achieves better performance with higher network throughput and lower packet delay (e.g., about 13%−146% higher in throughput and 13%−21% lower in delay than others in a mobile ad hoc network), as well as robustness under network mobility. Thus, TARS is highly suitable for mobile and traffic-varying UWSNs. Yunsi Fei |
ACM Trans. Sens. Networks | 2 |
| 2016 | Differential Fault Analysis of SHA3-224 and SHA3-256abstractThe security of SHA-3 against different kinds of attacks are of vital importance for crypto systems with SHA-3 as the security engine. In this paper, we look into the differential fault analysis of SHA-3, and this is the first work to conquer SHA3-224 and SHA3-256 using differential fault analysis. Comparing with one existing related work, we relax the fault models and make them realistic for different implementation architectures. We analyze fault propagation in SHA-3 under such single-byte fault models, and propose to use fault signatures at the observed output for analysis and secret retrieval. Results show that the proposed method can effectively identify the injected single-byte faults, and then recover the whole internal state of the input of last round χ operation (χi22) for both SHA3-224 and SHA3-256. Pei Luo, Yunsi Fei, A. Adam Ding |
FDTC | 2 |
| 2016 | Concurrent Error Detection for Reliable SHA-3 DesignabstractCryptographic systems are vulnerable to random errors and injected faults. Soft errors can inadvertently happen in critical cryptographic modules and attackers can inject faults into systems to retrieve the embedded secret. Different schemes have been developed to improve the security and reliability of cryptographic systems. As the new SHA-3 standard, Keccak algorithm will be widely used in various cryptographic applications, and its implementation should be protected against random errors and injected faults. In this paper, we devise different parity checking methods to protect the operations of Keccak. Results show that our schemes can be easily implemented and can effectively protect Keccak system against random errors and fault attacks. Pei Luo, Yunsi Fei |
ACM Great Lakes Symposium on VLSI | 3 |
| 2016 | A complete key recovery timing attack on a GPUabstractGraphics Processing Units (GPUs) have become mainstream parallel computing devices. They are deployed on diverse platforms, and an increasing number of applications have been moved to GPUs to exploit their massive parallel computational resources. GPUs are starting to be used for security services, where high-volume data is encrypted to ensure integrity and confidentiality. However, the security of GPUs has only begun to receive attention. Issues such as side-channel vulnerability have not been addressed. The goal of this paper is to evaluate the side-channel security of GPUs and demonstrate a complete AES (Advanced Encryption Standard) key recovery using known ciphertext through a timing channel. To the best of our knowledge, this is the first work that clearly demonstrates the vulnerability of a commercial GPU architecture to side-channel timing attacks. Specifically, for AES-128, we have been able to recover all key bytes utilizing a timing side channel in under 30 minutes. Zhen Hang Jiang, Yunsi Fei, David R. Kaeli |
HPCA | 2 |
| 2016 | SMARP: A Stochastic MAC Protocol with Randomized Power Control for Underwater Sensor NetworksabstractDesigning efficient medium access control (MAC) protocols for underwater sensor networks (UWSNs) is still a challenging issue, due to the long propagation delay and spatial-temporal uncertainty of underwater acoustic channel. In this paper, we make use of the capture effect in channel access, and propose a stochastic MAC protocol with randomized power control for UWSNs, called SMARP. Capture effect means when multiple packets arrive at the receiver simultaneously, it is possible that some packets can be decoded if its power strength is higher enough than other packets. We design a power control scheme by considering the non-negligible difference in acoustic propagation attenuation, and proactively create power captures at the receiver side to improve the network throughput. Fairness is maintained by randomly selecting the transmission power among a set of power levels. A utility-optimization framework is used to determine the optimal transmission strategy, which takes into account both the single-packet success probability and capture success probability. Extensive simulation results demonstrate that SMARP achieves higher network throughput and lower packet end-to-end delay than other representative underwater MAC protocols. Yunsi Fei, A. Adam Ding |
SECON | 2 |
| 2016 | DAP-MAC: A delay-aware probability-based MAC protocol for underwater acoustic sensor networks
Yunsi Fei |
Ad Hoc Networks | 2 |
| 2015 | Balance power leakage to fight against side-channel analysis at gate level in FPGAsabstractSide-channel attacks have been a serious threat to the security of embedded cryptographic systems, and various countermeasures have been devised to mitigate the leakages. Power balance technologies such as wave dynamic differential logic (WDDL) aim to balance the power by introducing differential logic. However, different routing length leads to different capacitance of wire, and this hampers the strength of the power balance countermeasure. In this paper, we further balance the power of differential signals by manipulating the lower level primitives and placement constraints on a Field Programmable Gate Array (FPGA). We choose Advanced Encryption Standard (AES) as the encryption algorithm and apply Hamming weight model to demonstrate the amount of leakage for different implementations. Results show that our method not only efficiently mitigates the side-channel leakage but also saves FPGA logic block resources and dynamic power consumption. Xin Fang 0001, Pei Luo, Yunsi Fei, Miriam Leeser |
ASAP | 3 |
| 2015 | Towards secure cryptographic software implementation against side-channel power analysis attacksabstractSide-channel attacks have been a real threat against many embedded cryptographic systems. A commonly used algorithmic countermeasure, random masking, incurs large execution delay and resource overhead. The other countermeasure, operation shuffling or permutation, can mitigate side-channel leakage effectively with minimal overhead. In this paper, we target automatically implementing operation shuffling in cryptographic algorithms to resist against side-channel power analysis attacks. We design a tool to detect independence among statements at the source code level and devise an algorithm for automatic operation shuffling. We test our algorithm on the new SHA3 standard, Keccak. Results show that the tool effectively implements operation-shuffling to reduce the side-channel leakage significantly, and therefore can guide automatic secure cryptographic software implementations against differential power analysis attacks. Pei Luo, Yunsi Fei, A. Adam Ding |
ASAP | 3 |
| 2015 | A Unified Metric for Quantifying Information Leakage of Cryptographic Devices Under Power Analysis Attacks
A. Adam Ding, Yunsi Fei, Pei Luo |
ASIACRYPT (2) | 3 |
| 2015 | Side-channel power analysis of a GPU AES implementationabstractGraphics Processing Units (GPUs) have been used to run a range of cryptographic algorithms. The main reason to choose a GPU is to accelerate the encryption/decryption speed. Since GPUs are mainly used for graphics rendering, and only recently have they become a fully-programmable parallel computing device, there has been little attention paid to their vulnerability to side-channel attacks. In this paper we present a study of side-channel vulnerability on a state-of-the-art graphics processor. To the best of our knowledge, this is the first work that attempts to extract the secret key of a block cipher implemented to run on a GPU. We present a side-channel power analysis methodology to extract all of the last round key bytes of a CUDA AES (Advanced Encryption Standard) implementation run on an NVIDIA TESLA GPU. We describe how we capture power traces and evaluate the power consumption of a GPU. We then construct an appropriate power model for the GPU. We propose effective methods to sample and process the GPU power traces so that we can recover the secret key of AES. Our results show that parallel computing hardware systems such as a GPU are highly vulnerable targets to power-based side-channel attacks, and need to be hardened against side-channel threats. Yunsi Fei, Pei Luo, Saoni Mukherjee, David R. Kaeli |
ICCD | 2 |
| 2015 | Traffic and vehicle speed prediction with neural network and Hidden Markov model in vehicular networksabstractAccurate on-road vehicle speed prediction is important for many intelligent vehicular and transportation applications. It is also challenging because the individual vehicle speed is affected by many factors, e.g., traffic speed, vehicle type, and driver's behavior, in either deterministic or stochastic ways. This paper proposes a novel vehicle speed prediction method in the context of vehicular networks, where the real-time traffic information is accessible. Traffic speeds of following road segments are first predicted by Neural Networks (NNs) based on historical traffic data. Hidden Markov models (HMMs) are trained by the Baum-Welch algorithm with historical traffic and vehicle data to present the statistical relationship between vehicle speed and traffic speed. The forward-backward algorithm is applied on HMMs to extract vehicle's speed on each road segment along the driving route. Simulation is set up on the SUMO microscopic traffic simulator with the application of a real Luxembourg highway network and traffic count data. The vehicle speed prediction result shows that our proposed method outperforms other ones in terms of prediction accuracy. Bingnan Jiang, Yunsi Fei |
Intelligent Vehicles Symposium | 2 |
| 2015 | TARS: A Traffic-Adaptive Receiver-Synchronized MAC Protocol for Underwater Sensor NetworksabstractEfficient medium access control (MAC) is desirable for underwater sensor networks (UWSNs). However, designing an efficient underwater MAC protocol is challenging due to the long propagation delay of the underwater acoustic channel and the spatial-temporal uncertainty. In this paper, we propose a novel Traffic-Adaptive Receiver-Synchronized underwater MAC protocol, TARS, a stochastic light-weight channel access scheme that addresses the spatial-temporal uncertainty for maximizing the network throughput. We adjust the packet transmission time (phase) in a slot, which is dependent on the sender-receiver distance, to align packet receptions for collision reduction. Both the sound propagation speed variation and the node mobility are considered in setting the optimal transmission phase and the slot size. We employ a queue-aware utility-optimization framework to determine the optimal traffic-adaptive transmission strategies dynamically, taking into account both the packet interference and the data queue status. Extensive simulation results show that compared to the existing representative underwater MAC protocols, TARS achieves better performance with higher network throughput and lower packet end-to-end delay. Yunsi Fei |
MASCOTS | 2 |
| 2014 | A Statistical Model for Higher Order DPA on Masked Devices
A. Adam Ding, Yunsi Fei, Pei Luo |
CHES | 3 |
| 2014 | On-road PHEV power management with hierarchical strategies in vehicular networksabstractIn plug-in hybrid electric vehicles (PHEVs), the power management system coordinates powertrain operations to achieve high energy efficiency. Conventional PHEV power management systems work in either an online or offline mode. Most online systems are based on some pre-set power balancing strategies without utilizing the driving cycle or route information. Offline management strategies solved from historical driving cycles are not optimal for real specific driving routes. With the rapid development of vehicular networks and proliferation of smartphones, real-time traffic information can be collected by smartphones from a vehicular network so as to facilitate online PHEV power management. This paper proposes an on-road PHEV power management cyber-physical system (CPS) with 2-level hierarchical optimizations to minimize the fuel consumption of a trip. The high-level online stochastic optimization generates a battery energy budget for each road at runtime according to the traffic prediction and trip information. The low-level powertrain policies are solved offline from historical driving cycles. During driving, the high-level battery energy budgets and low-level policies are combined to get the optimal power decisions according to current driving states. Simulation results show that the proposed method significantly outperforms other three methods in fuel savings. Bingnan Jiang, Yunsi Fei |
Intelligent Vehicles Symposium | 2 |
| 2014 | HiTS: A High Throughput Memory Scheduling Scheme to Mitigate Denial-of-Service Attacks in Multi-core SystemsabstractSharing DRAM memory by multiple cores in a computer system potentially exposes the running threads on cores to denial-of-service (DoS) attacks. This issue is usually addressed by memory scheduling schemes that rotate the memory service among threads according to a certain ranking mechanism. These ranking-based schemes, however, often incur many memory banks' row-buffer conflicts which reduce the throughput of DRAM and the entire system. This paper proposes a new ranking-based memory scheduling scheme, called HiTS, to mitigate DoS attacks in multicore systems with the lowest performance degradation. HiTS achieves these by ranking threads according to each thread's memory usage/requirement. HiTS then enforces the ranking in a way that minimum performance overhead would occur and fairness is also balanced. The effectiveness of HiTS is evaluated by simulations with 18 different workloads running on 8- and 16-core machines. The simulation results show up to 15.8% improvements in terms of unfairness reduction and 24.1% in system throughput compared with the best existing scheduling scheme. Mansour Shafaei, Yunsi Fei |
SBAC-PAD | 2 |
| 2013 | An adaptive routing protocol based on connectivity prediction for underwater disruption tolerant networksabstractUnderwater Sensor Networks (UWSNs) are a desirable networking technique to facilitate various aquatic applications. However, the adverse characteristics of underwater communications and high cost of underwater sensor nodes limit UWSNs to sparse deployment, resulting in intermittent connectivity and therefore calling for techniques for Delay/Disruption Tolerant Networks (DTNs). To cope with disruptions, extra efforts have to be made in the routing protocol to provide transparent and robust end-to-end connections to upper-layer applications. In this paper, we propose a novel adaptive and energy-efficient routing protocol for underwater DTNs. By exploiting underwater node mobility patterns with adaptive filters, sensor nodes are able to estimate future contact events with other nodes in addition to the average contact probabilities over a prediction window. The proposed protocol is based on a distributed machine learning technique, Q-learning, which aims to select the most promising forwarders so as to minimize the end-to-end delay. Extensive simulations of the proposed protocol are carried out, and the results have shown that our protocol yields significantly better network performances and energy efficiency compared to other existing DTN routing protocols. Tiansi Hu, Yunsi Fei |
GLOBECOM | 2 |
| 2013 | DSH-MAC: Medium Access Control based on Decoupled and Suppressed Handshaking for long-delay Underwater Acoustic Sensor NetworksabstractEfficient underwater networking is still a challenging issue due to its physical limitations, like long propagation delay. In this paper, we focus on medium access control (MAC) for underwater acoustic sensor networks (UW-ASNs). Considering that the handshaking process in traditional contention-based MACs is the main hurdle for improving the network channel utilization, we propose a novel MAC protocol with Decoupled and Suppressed Handshaking (DSH-MAC) in order to reduce the time overhead, and therefore achieve more efficient channel utilization. In DSH-MAC the conventional two-way handshaking is decoupled, and hence relevant nodes are able to perform other transmissions while control packets are propagating in water. DSH-MAC also suppresses unnecessary control packets with traffic prediction, further improving the channel utilization and throughput. Our proposed protocol has been proven to be channel-efficient with both theoretical analysis and intensive simulations. Tiansi Hu, Yunsi Fei |
LCN | 2 |
| 2013 | Leveraging speculative architectures for runtime program validation
Juan Carlos Martínez Santos, Yunsi Fei |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2012 | Static secure page allocation for light-weight dynamic information flow trackingabstractDynamic information flow tracking (DIFT) is an effective security countermeasure for both low-level memory corruptions and high-level semantic attacks. However, many software approaches suffer large performance degradation, and hardware approaches have high logic and storage overhead. We propose a flexible and light-weight hardware/software co-design approach to perform DIFT based on secure page allocation. Instead of associating every data with a taint tag, we aggregate data according to their taints, i.e., putting data with different attributes in separate memory pages. Our approach is a compiler-aided process with architecture support. The implementation and analysis show that the memory overhead is little, and our approach can protect critical information, including return address, indirect jump address, and system call IDs, from being overwritten by malicious users. Juan Carlos Martínez Santos, Yunsi Fei, Zhijie Jerry Shi |
CASES | 2 |
| 2012 | A Statistical Model for DPA with Novel Algorithmic Confusion Analysis
Yunsi Fei, Qiasi Luo, A. Adam Ding |
CHES | 1 |
| 2012 | MURAO: A multi-level routing protocol for acoustic-optical hybrid underwater wireless sensor networksabstractIn the past decade, underwater acoustic sensor networks (UW-ASNs) have been studied broadly in various aquatic applications, enabling humans to observe and explore the vast underwater domain. Although acoustic underwater communications are able to support long-range and low-bandwidth applications, the capabilities of UW-ASN are greatly limited by the long delay and low data rate of acoustic communications. Underwater free-space optical communication is a potential alternative solution. However, it has short transmission ranges and requires dense deployment. In this paper, we propose a novel acoustic-optical hybrid architecture for underwater wireless sensor networks, and a multi-level Q-learning based routing protocol, MURAO, for such networks. The network is physically partitioned into several groups and logically divided into two layers. The upper-layer group leaders supervise the routing in the lower layer, and the lower-layer group members carry out the actual data packet routing. Because upper-layer group leaders have a boarder view of the network and all the groups are able to carry out the learning process concurrently, the performance of routing is greatly improved compared to the flat Q-learning-based routing. The experiment results show that MURAO is more robust to changes of network topology, and achieves much higher delivery rates as well as shorter delays in a dynamic network than the flat Q-learning routing. Tiansi Hu, Yunsi Fei |
SECON | 2 |
| 2012 | Resource Sharing of Pipelined Custom Hardware Extension for Energy-Efficient Application-Specific Instruction Set Processor DesignabstractApplication-Specific Instruction set Processor (ASIP) has become an increasingly popular platform for embedded systems because of its high performance, flexibility, and short turn-around time. The hardware extension in ASIPs can speed-up program execution. However, it also incurs area overhead and extra static energy consumption. Traditional datapath merging techniques reduce the circuit overhead by reusing hardware modules for executing multiple operations. However, they introduce structural hazard for multiple custom instructions in sequence, and hence reduce the performance improvement. In this article, we introduce a pipelined configurable structure for the hardware extension in ASIPs, so that structural hazards can be remedied. With multiple subgraphs of operations selected, we design a novel operation-to-hardware mapping algorithm based on Integer Linear Programming (ILP) to automatically construct a resource-efficient pipelined configurable functional unit. Different resource sharing schemes would affect both the hardware overhead and the overall performance improvement. We analyze the design trade-offs between resource efficiency and performance improvement. At the end, we present our design space exploration results by setting the optimization objective to area, area and delay, and delay respectively. Hai Lin 0004, Yunsi Fei |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2012 | A Hardware/Software Cooperative Custom Register Binding Approach for Register Spill Elimination in Application-Specific Instruction Set ProcessorsabstractApplication-Specific Instruction set Processor (ASIP) has become an important design choice for embedded systems. It can achieve both high flexibility offered by the base processor core and high performance and energy efficiency offered by the dedicated hardware extensions. Although a lot of efforts have been devoted to computation acceleration, for example, automatic custom instruction identification and synthesis, limited on-chip data storage elements including the register file and data cache have become a potential performance bottleneck. For custom instructions that have more inputs and/or outputs than the generic register file I/O ports, custom registers are added in ASIPs to satisfy the need of additional inputs and outputs, and traditionally they are used only by custom instructions. In this article, we propose a hardware/software cooperative approach with a linear scan register allocation algorithm, which allows base instructions to utilize the existing custom registers in ASIPs for eliminating register spills of the program. The data traffic between the base processor and off-chip memory can be replaced with energy-efficient on-chip communications between the processor core and custom hardware extensions. Our experimental results demonstrate that a significant performance gain can be achieved, orthogonal to improvements by other techniques in ASIP design. Tiansi Hu, Yunsi Fei |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2012 | Hierarchical Design of an Application-Specific Instruction Set Processor for High-Throughput and Scalable FFT ProcessingabstractFast Fourier transformation (FFT), a kernel data processing task in communication systems, has been studied intensively for efficient software and hardware implementations. Nowadays, various orthogonal frequency division multiplexing (OFDM)-based wireless communication standards have raised more stringent requirements on both throughput and flexibility for FFT computation. Application-specific instruction set processor (ASIP) has emerged as a promising solution to meet these requirements. This paper presents a novel hierarchical design of an ASIP tailored for FFT. We reconstruct the FFT computation flow into a scalable array structure based on an 8-point butterfly unit (BU). The array structure can easily expand along both the horizontal and vertical dimensions for any-point FFT computation. We incorporate custom register files to reduce memory access and derive a regular data addressing rule accordingly. With the microarchitecture modifications, we extend the instruction set architecture (ISA) with new instructions to accelerate FFT operations. An FFT ASIP is implemented on Tensilica's reconfigurable processor platform. Our FFT ASIP achieves the data throughput of 405.7 Mb/s for 1 K-point FFT, which attains UWB-OFDM specifications. The area of our custom processor is 147 kilo gates and the total processor power consumption is 60.7 mW, which are acceptable compared to several other designs such as application specific integrated circuit, digital signal processing, field-programmable gate array, and other ASIP implementations. We also extend the implementation for up to 8 K-point FFTs, with degraded performance but still meeting the requirements of those communications standards that demand large-size FFT computations. Xuan Guan, Yunsi Fei, Hai Lin 0004 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2011 | Adaptive Extended Min-Sum Algorithm for Nonbinary LDPC DecodingabstractLow Density Parity Check (LDPC) codes over a Galois Field GF(q) can provide significantly better error- correcting quality than binary LDPC codes with moderate code length. However, the computational complexity and memory requirement of nonbinary codes are much higher, limiting the application of non- binary LDPC codes. In this paper, we propose an adaptive message truncation algorithm for non-binary LDPC decoding, guided by the estimated code error rates. Compared to the previous fixed message truncation method, it can cut the messages adaptively, and therefore provide better decoding quality and computation complexity reduction. To further reduce the computation, we propose another adaptive check node update algorithm, simplifying the decoding by reducing the number of check nodes updating. Our simulation results demonstrate that by combining these two algorithms together, the average message size can be reduced greatly and good decoding quality is achieved, at a little cost of iterations. Compared to the existing truncation algorithms, our approach can reduce the order of complexity to O(nalog2(na)) (where nais messages size), with less performance degradation. Xuan Guan, Yunsi Fei |
GLOBECOM | 2 |
| 2010 | A novel multi-objective instruction synthesis flow for application-specific instruction set processorsabstractApplication-Specific Instruction set Processor (ASIP) has become an increasingly popular platform for embedded systems. Traditional ASIP synthesis flows mainly target performance improvement, with other design metrics not being addressed appropriately. In this paper, we show that traditional custom instruction exploration algorithms and cost estimation methods for performance improvement only are not suitable for other design objectives, such as energy reduction and area minimization. We propose an ASIP design flow that can be adapted to different design objectives and achieve the balance between them. A novel design space exploration algorithm is developed to identify custom instructions for execution acceleration and energy reduction while reducing the hardware overhead. Hai Lin 0004, Yunsi Fei |
ACM Great Lakes Symposium on VLSI | 2 |
| 2010 | Exploring custom instruction synthesis for application-specific instruction set processors with multiple design objectivesabstractApplication-specific instruction set processor (ASIP) has become a promising platform for embedded system design in the past decade. Traditional custom instruction synthesis flows for ASIPs mainly target performance improvement. Other design metrics are not addressed appropriately. In this paper, we show that the existing custom instruction exploration algorithms and cost estimation methods for performance improvement only are not suitable for other important design objectives, such as increasing energy efficiency and reducing area overhead. We propose a holistic ASIP design flow that can be adapted to optimize performance, energy consumption, or area. We formulate the design space exploration problem into an operation scheduling process. Different algorithms are employed to find the corresponding best custom instruction set efficiently. Hai Lin 0004, Yunsi Fei |
ISLPED | 2 |
| 2010 | An Adaptive and Energy-efficient Routing Protocol Based on Machine Learning for Underwater Delay Tolerant NetworksabstractUnderwater Sensor Network (UWSN) is emerging as a promising networking technique for aquatic environment monitoring and exploration. However, because of the adverse characteristics of underwater communications, underwater sensor networks may get partitioned temporarily, and hence call for techniques for Delay/Disruption Tolerant Networks (DTNs). In this paper, we propose an adaptive and energy-efficient routing protocol based on a machine learning technique, Q-learning, for underwater DTNs. Extensive simulations of the proposed protocol are carried out, and the results have shown that our protocol can cope with dynamic disconnections and disruptions in underwater DTNs well and achieves a good trade-off between energy efficiency and end-to-end delay. Tiansi Hu, Yunsi Fei |
MASCOTS | 2 |
| 2010 | QELAR: A Machine-Learning-Based Adaptive Routing Protocol for Energy-Efficient and Lifetime-Extended Underwater Sensor NetworksabstractUnderwater sensor network (UWSN) has emerged in recent years as a promising networking technique for various aquatic applications. Due to specific characteristics of UWSNs, such as high latency, low bandwidth, and high energy consumption, it is challenging to build networking protocols for UWSNs. In this paper, we focus on addressing the routing issue in UWSNs. We propose an adaptive, energy-efficient, and lifetime-aware routing protocol based on reinforcement learning, QELAR. Our protocol assumes generic MAC protocols and aims at prolonging the lifetime of networks by making residual energy of sensor nodes more evenly distributed. The residual energy of each node as well as the energy distribution among a group of nodes is factored in throughout the routing process to calculate the reward function, which aids in selecting the adequate forwarders for packets. We have performed extensive simulations of the proposed protocol on the Aqua-sim platform and compared with one existing routing protocol (VBF) in terms of packet delivery rate, energy efficiency, latency, and lifetime. The results show that QELAR yields 20 percent longer lifetime on average than VBF. Tiansi Hu, Yunsi Fei |
IEEE Trans. Mob. Comput. | 2 |
| 2010 | Register file partitioning and recompilation for register file power reductionabstractRegister files in modern embedded processors contribute a substantial budget in the energy consumption due to their large switching capacitance and long working time. For some embedded processors, on average 25% of registers account for 83% of register file accessing time. This motivates us to partition the register file into hot and cold regions, with the most frequently used registers placed in the hot region, and the rarely accessed ones in the cold region. We employ the bit-line splitting and drowsy register cell techniques to reduce the overall register file accessing power. We propose a novel approach to partition the register in a way that can achieve the largest power saving. We formulate the register file partitioning process into a graph partitioning problem, and apply an effective algorithm to obtain the optimal result. We evaluate our algorithm for MiBench and SPEC2000 applications on the SimpleScalar PISA system, and an average saving of 58.3% and 54.4% over the nonpartitioned register file accessing power is achieved. The area overhead is negligible, and the execution time overhead is acceptable (5.5% for MiBench 2.4% for SPEC2000). Further evaluation for MiBench applications is performed on Alpha and X86 system. Xuan Guan, Yunsi Fei |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2010 | Register File Partitioning and Compiler Support for Reducing Embedded Processor Power ConsumptionabstractRegister file (RF) in modern embedded processors contributes a substantial budget in the energy consumption due to its large switching capacitance and long working time. For embedded processors, on average 25% of registers count for 83% of RF accessing time. This motivates us to partition the RF into hot and cold regions, with the most frequently used registers placed in the hot region, and the rarely accessed ones in the cold region. We employ the techniques of bit-line splitting and drowsy register cell to reduce the overall accessing power of RF. We propose a novel approach to partition the RF in a way that can achieve the largest power saving. We formulate the RF partitioning process into a graph partitioning problem, and apply an effective algorithm to obtain the optimal result. We evaluate our algorithm on MiBench and SPEC2000 applications, and an average saving of 58.3% and 54.4% over the non-partitioned RF accessing power is achieved for the SimpleScalar PISA system, respectively. The area overhead is negligible, and the execution time overhead is acceptable. Xuan Guan, Yunsi Fei |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2010 | Architectural Enhancement and System Software Support for Program Code Integrity Monitoring in Application-Specific Instruction-Set ProcessorsabstractProgram code in a computer system can be altered either by malicious security attacks or by various faults in microprocessors. At the instruction level, all code modifications are manifested as bit flips. In this paper, we present a generalized methodology for monitoring code integrity at run-time in application-specific instruction-set processors. We embed monitoring microoperations in machine instructions, so the processor is augmented with a hardware monitor automatically. The monitor observes the processor's execution trace at run-time, checks whether it aligns with the expected program behavior, and signals any mismatches. Since the monitor works at a level below the instructions, the monitoring mechanism cannot be bypassed by software or compromised by malicious users. We discuss the ability and limitation of such monitoring mechanism for detecting both soft errors and code injection attacks. We propose two different schemes for managing the monitor, the operating system (OS) managed and application controlled, and design the constituent components within the monitoring architecture. Experimental results show that with an effective hash function implementation, our microarchitectural support can detect program code integrity compromises at a high probability with small area overhead and little performance degradation. Hai Lin 0004, Yunsi Fei, Xuan Guan, Zhijie Jerry Shi |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2009 | Design of an application-specific instruction set processor for high-throughput and scalable FFTabstractVarious orthogonal frequency division multiplexing (OFDM)-based wireless communication standards have raised more stringent requirements on throughput and flexibility of fast Fourier transformation (FFT), a kernel data transformation task in communication systems. Application-specific instruction set processor (ASIP) has emerged as a promising solution to meet these requirements. In this paper, we propose a novel ASIP design tailored for FFT computation. We reconstruct the FFT computation flow into a scalable array structure based on an 8-point butterfly unit (BU). Any-point FFT computation can be carried out in the array structure which can easily expand along both the horizontal and vertical dimensions. We incorporate custom register files to reduce memory access. The data address for custom registers in each FFT stage is changed accordingly, and we derive a regular address changing (AC) rule. With the microarchitecture modifications, we extend the instruction set with three custom instructions correspondingly. Our FFT ASIP implementation achieves great performance improvement over the standard FFT software implementation, one TI DSP processor, and one commercial Xtensa ASIP, with the data throughput improvement as 866.5X, 5.9X, 2.3X, respectively. Meanwhile, the area and power consumption overhead of the custom hardware is negligible. Xuan Guan, Hai Lin 0004, Yunsi Fei |
DATE | 3 |
| 2009 | Resource sharing of pipelined custom hardware extension for energy-efficient application-specific instruction set processor designabstractApplication-Specific Instruction set Processor (ASIP) has become an increasingly popular platform for embedded systems because of its high performance and flexibility. Energy efficiency is critical for portable and embedded devices, and should be addressed separately from performance consideration. The hardware extension in ASIPs can speed-up program execution, but also incurs area overhead and static energy consumption of the processors. Traditional data path merging techniques reduce circuit overhead by reusing hardware resources for executing multiple custom instructions. However, they introduce structural hazard for custom instructions on extended processors, and hence reduce the performance improvement. In this paper, we introduce a pipelined configurable hardware structure for the hardware extension in ASIPs, so that structural hazards can be remedied. With multiple subgraphs of operations selected for custom hardware realization, we devise a novel operation-to-hardware mapping algorithm based on Integer Linear Programming (ILP) to automatically construct a resource-efficient pipelined configurable hardware extension. We demonstrate that different resource sharing schemes would affect both the hardware overhead and datapath delay of the custom instructions. We analyze the design tradeoffs between resource efficiency and performance improvement, and present the design space exploration results. Hai Lin 0004, Yunsi Fei |
ICCD | 2 |
| 2009 | A Hierarchical Design of an Application-specific Instruction Set Processor for High-throughput FFTabstractThis paper presents a novel hierarchical design of an application-specific instruction set processor (ASIP) tailored for fast Fourier transformation (FFT), a kernel data transformation task in digital communication systems, to meet the stringent requirements on throughput and flexibility. We reconstruct the FFT computation flow into a scalable array structure based on an 8-point butterfly unit (BU). The array can easily expand along both the horizontal and vertical dimensions for any-point FFT computation, and contains the same structure for each horizontal stage. We incorporate custom register files to reduce memory access, and derive a regular data addressing rule accordingly. With the microarchitecture modifications, we extend the instruction set with three custom instructions. Our FFT ASIP implementation achieves a data throughput improvement of 866.5times, 5.9times, 2.3times over the standard FFT software implementation, one TI DSP processor, and one commercial ASIP - Xtensa's implementation, respectively. Meanwhile, the area and power consumption overhead of the custom hardware is acceptable. Xuan Guan, Yunsi Fei, Hai Lin 0004 |
ISCAS | 2 |
| 2009 | Orchestrating Horizontal Parallelism and Vertical Instruction Packing of Programs to Improve System Overall EfficiencyabstractBoth performance and energy efficiency are critical concerns for embedded systems and portable devices. Multi-issue processors can exploit the instruction-level parallelism (ILP) of programs to improve the performance greatly, however, most of the time at a cost of energy and power consumption. How to reduce the energy consumption while maintaining the high performance of programs running on multi-issue processors remains a challenging problem. In this paper, we propose a novel approach to apply the instruction register file(IRF) technique from single-issue processor to VLIW architecture. Frequently executed instructions are selected to be placed in the on-chip IRF for fast and energy-efficient access in program execution. Violation of synchronization among VLIW instruction slots is avoided by introducing new instruction formats and microarchitectural support. The enhanced VLIW architecture is, thus, able to orchestrate the horizontal instruction parallelism and vertical instruction packing for programs to improve system overall efficiency. Our experimental results show that the proposed processor architecture achieves both the performance advantage provided by the VLIW architecture and high energy efficiency provided by the IRF-based instruction packing technique, e.g., the fetch energy consumption is reduced by 33.4 percent for a 4-way VLIW architecture with 16-entry IRFs for SPEC2000 testbenches. Hai Lin 0004, Yunsi Fei |
IEEE Trans. Computers | 2 |
| 2008 | Reducing power consumption of embedded processors through register file partitioning and compiler supportabstractAs embedded processors being widely used in specific application domains, such as communications, multimedia, and networking, the register file has contributed a substantial budget in embedded processor energy consumption due to its long working time for the data intensive computations and the large switching capacitance. It is found that 25% of registers can account for 83% of register file accessing time during many embedded application execution. This fact motivates us to reduce the register file power consumption by partitioning the registers to different regions according to their usage pattern. The most frequently used registers are put in the hot part, and the cold part of register file is rarely accessed. We employ the register file bitline splitting and the drowsy register cell techniques in our design to reduce the overall accessing power of the register file. We propose a novel approach to partition the register file in a way so that the largest power saving can be achieved. We formulate the register file partitioning process into a graph partitioning problem, and apply an effective algorithm to obtain the optimal result. We evaluate our algorithm on MiBench applications, and an average saving of 43.6% in the register file access power consumption over the original non-partitioned register file is achieved for the SimpleScalar PISA system. Xuan Guan, Yunsi Fei |
ASAP | 2 |
| 2008 | An efficient digital circuit for implementing Sequence Alignment algorithm in an extended processorabstractThe problem of Sequence Alignment (Edit Distance) between a pair of strings has been well studied in the field of computing algorithms. The classic dynamic programming-based algorithm, Needleman-Wunsch (O(n2)), has been widely used in practice, especially by biologists to find similarities between gene sequences. Any optimization in the implementation of this algorithm will have a significant practical impact on biological research. However, within the past several decades, not much has been done in improving the runtime of the algorithm in real implementations. Although algorithms based on systolic processor arrays and FPGAs were presented earlier to create custom hardware to aid in speed-up, their usage has been very limited due to their inherent synchronous design complexity and scalability issues. In view of this, we propose an efficient hardware implementation of the Sequence Alignment algorithm. We provide a simple and efficient asynchronous sequential design which can be readily implemented as an instruction in an extensible processor. Experimental results show that our circuit implementation can achieve a speed-up of 3.77X on average compared with the software counterpart, meanwhile reducing the area cost. Vamsi Kundeti, Yunsi Fei, Sanguthevar Rajasekaran |
ASAP | 2 |
| 2008 | Harnessing Horizontal Parallelism and Vertical Instruction Packing of Programs to Improve System Overall EfficiencyabstractMulti-issue processors can exploit the instruction level parallelism (ILP) of programs to improve the performance greatly. How to reduce the energy consumption while maintaining the high performance of programs running on multi- issue processors remains a challenging problem. In this paper, we propose a novel approach to apply the instruction register file (IRF) technique from single-issue processor to VLIW architecture. Frequently executed instructions are selected to be placed in the on-chip IRF for fast access in program execution. Violation of synchronization among VLIW instruction slots is avoided by introducing new instruction formats and microarchitectural support. The enhanced VLIW architecture is thus able to orchestrate the horizontal instruction parallelism and vertical instruction packing for programs to improve system overall efficiency. Our experimental results show that the proposed processor architecture achieves both the performance advantage provided by the VLIW architecture and high energy efficiency provided by the IRF-based instruction packing technique (e.g., 71.1% reduction in the fetch energy consumption for a 4-way VLIW architecture with 8-entry IRFs). Hai Lin 0004, Yunsi Fei |
DATE | 2 |
| 2008 | Leveraging speculative architectures for run-time program validationabstractProgram execution can be tampered by malicious attackers through exploiting software vulnerabilities. Changing the program behavior by compromising control data and decision data has become the most serious threat to computer systems security. Although several hardware approaches have been presented to validate program execution, they mostly suffer great hardware area or poor ambiguity handling. In this paper, we propose a new hardware-based approach by leveraging the existing speculative architectures for run-time program validation. The on-chip branch target buffer (BTB) is utilized as a cache of the legitimate control flow transfers stored in a secure memory region. In addition, the BTB is extended to store the correct program path information. At each indirect branch site, the BTB is used to validate the decision history of conditional branches before it, and more information about the future decision path is fetched to monitor the execution path at run-time. Implementation of this approach is transparent to the upper operating system and programs. Thus, it is applicable to legacy code. Due to good code locality of the executable programs and effectiveness of branch prediction, the frequency of run-time control flow validations against the secure off-chip memory is low. Our experimental results show a negligible performance penalty and small storage overhead with ambiguity reduced. Juan Carlos Martínez Santos, Yunsi Fei |
ICCD | 2 |
| 2008 | QELAR: A Q-learning-based Energy-Efficient and Lifetime-Aware Routing Protocol for Underwater Sensor NetworksabstractUnderwater sensor network (UWSN) has emerged as a promising network technique for various aquatic applications in recent years. Due to some constraints in UWSNs, such as high latency, low bandwidth and high energy consumption, it is challenging to build networking protocols for UWSNs. In this paper, we focus on addressing the routing issue in UWSNs. We propose an adaptive, energy-efficient, and lifetime-aware routing protocol based on reinforcement learning, QELAR. Our protocol assumes generic MAC protocols and aims at prolonging the lifetime of networks by making residual energy of sensor nodes more evenly distributed. The residual energy of each node as well as the energy distribution among a group is factored in throughout the routing process to calculate the reward function, which aids in selecting the adequate forwarders for packets. We have performed extensive simulations of the proposed protocol on the Aqua-sim platform, and compared with one existing routing protocol (VBF) in terms of packet delivery rate, energy efficiency, latency and lifetime. The results show that the QELAR protocol yields 20% longer lifetime on average than VBF. Tiansi Hu, Yunsi Fei |
IPCCC | 2 |
| 2008 | An energy-aware framework for dynamic software management in mobile computing systemsabstractEnergy efficiency has become a very important and challenging issue for resource-constrained mobile computers. In this article, we propose a novel dynamic software management (DSOM) framework to improve battery utilization. We have designed and implemented a DSOM module in user space, independent of the operating system (OS), which explores quality-of-service (QoS) adaptation to reduce system energy and employs a priority-based preemption policy for multiple applications to avoid competition for limited energy resources. Software energy macromodels for mobile applications are employed to predict energy demand at each QoS level, so that the DSOM module is able to select the best possible trade-off between energy conservation and application QoS; it also honors the priority desired by the user. Our experimental results for some mobile applications (video player, speech recognizer, voice-over-IP) show that this approach can meet user-specified task-oriented goals and significantly improve battery utilization. Yunsi Fei, Lin Zhong 0001, Niraj K. Jha |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2007 | Microarchitectural support for program code integrity monitoring in application-specific instruction set processorsabstractProgram code in a computer system can be altered either by malicious security attacks or by various faults in microprocessors. At the instruction level, all code modifications are manifested as bit flips. In this work, we present a generalized methodology for monitoring code integrity at run-time in application-specific instruction set processors (ASIPs), where both the instruction set architecture (ISA) and the underlying micro architecture can be customized for a particular application domain. We embed monitoring microoperations in machine instructions, thus the processor is augmented with a hardware monitor automatically. The monitor observes the processor's execution trace of basic blocks at run-time, checks whether the execution trace aligns with the expected program behavior, and signals any mismatches. Since microoperations are at a lower software architecture level than processor instructions, the microarchitectural support for program code integrity monitoring is transparent to upper software levels and no recompilation or modification is needed for the program. Experimental results show that our microarchitectural support can detect program code integrity compromises with small area overhead and little performance degradation Yunsi Fei, Zhijie Jerry Shi |
DATE | 1 |
| 2007 | Utilizing custom registers in application-specific instruction set processors for register spills eliminationabstractApplication-specific instruction set processor (ASIP) has become an important design choice for embedded systems. It can achieve both high flexibility offered by the base processor core and high performance and energy efficiency offered by the dedicated hardware extensions. Although a lot of efforts have been devoted to computation acceleration, e.g., automatic custom instruction identification and synthesis, the limited on-chip data storage elements, including the register file and data cache, have become a potential performance bottleneck. In this paper, we propose a hardware/software cooperative approach and a linear scan register allocation algorithm to utilize the existing custom registers in ASIPs for eliminating register spills. The data traffic between the processor and memory can be reduced through efficient on-chip communications between the base processor core and custom hardware extensions. Our experimental results demonstrate that a promising performance gain can be achieved, which is orthogonal to improvements by any other technique in ASIP design. Hai Lin 0004, Yunsi Fei |
ACM Great Lakes Symposium on VLSI | 2 |
| 2007 | Compiler-assisted architectural support for program code integrity monitoring in application-specific instruction set processorsabstract(ASIPs) are being increasingly used in mobile embedded systems, the ubiquitous networking connections have exposed these systems under various malicious security attacks, which may alter the program code running on the systems. In addition, soft errors in microprocessors can also change program code and result in system malfunction. At the instruction level, all code modifications are manifested as bit flips. In this work, we present a generalized methodology for monitoring code integrity at run-time in ASIPs, where both the instruction set architecture (ISA) and the underlying microarchitecture can be customized for a particular application domain. Based on the microoperation-based monitoring architecture that we have presented in previous work, we propose a compiler-assisted and application-controlled management approach for the monitoring architecture. Experimental results show that compared with the OS-managed scheme and other compiler-assisted schemes, our approach can detect program code integrity compromises with much less performance degradation. Hai Lin 0004, Xuan Guan, Yunsi Fei, Zhijie Jerry Shi |
ICCD | 3 |
| 2007 | Energy-optimizing source code transformations for operating system-driven embedded softwareabstractThis paper proposes four types of source code transformations for operating system (OS)-driven embedded software programs to reduce their energy consumption. Their key features include spanning of process boundaries and minimization of the energy consumed in the execution of OS services—opportunities which are beyond the reach of conventional compiler optimizations and source code transformations. We have applied the proposed transformations to several multiprocess benchmark programs in the context of an embedded Linux OS running on an Intel StrongARM processor. They achieve up to 37.9% (23.8%, on average) energy reduction compared to highly compiler-optimized implementations. Yunsi Fei, Srivaths Ravi 0001, Anand Raghunathan, Niraj K. Jha |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2004 | A hybrid energy-estimation technique for extensible processorsabstractIn this paper, we present an efficient and accurate methodology for estimating the energy consumption of application programs running on extensible processors. Extensible processors, which are getting increasingly popular in embedded system design, allow a designer to customize a base processor core through instruction set extensions. Existing processor energy macromodeling techniques are not applicable to extensible processors, since they assume that the instruction set architecture as well as the underlying structural description of the micro-architecture remain fixed. Our solution to the above problem is a hybrid energy macromodel suitably parameterized to estimate the energy consumption of an application running on the corresponding application-specific extended processor instance, which incorporates any custom instruction extension. Such a characterization is facilitated by careful selection of macromodel parameters/variables that can capture both the functional and structural aspects of the execution of a program on an extensible processor. Another feature of the proposed energy characterization flow is the use of regression analysis to build the macromodel. Regression analysis allows for in-situ characterization, thus allowing arbitrary test programs to be used during macromodel construction. We validated the proposed methodology by characterizing the energy consumption of a state-of-the-art extensible processor (Tensilica's Xtensa). We used the macromodel to analyze the energy consumption of several benchmark applications with custom instructions. The mean absolute error in the macromodel estimates is only 3.3%, when compared to the energy values obtained by a commercial tool operating on the synthesized register-transfer level (RTL) description of the custom processor. Our approach achieves an average speedup of three orders of magnitude over the commercial RTL energy estimator. Our experiments show that the proposed methodology also achieves good relative accuracy, which is essential in energy optimization studies. Hence, our technique is both efficient and accurate. Yunsi Fei, Srivaths Ravi 0001, Anand Raghunathan, Niraj K. Jha |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2004 | Register binding-based RTL power management for control-flow intensive designsabstractOne important way to reduce power consumption is to reduce the spurious switching activity in a circuit or circuit component, i.e., activity that is not required by its specified functionality. Given a scheduled behavior and functional unit binding, we show that spurious switching activity can be reduced through proper register binding using retentive multiplexers. Retentive multiplexers can preserve their previous select signal values in the control steps in which the select signals are don't cares. A functional unit, in which spurious switching activity is completely eliminated, is called perfectly power managed. We present a general sufficient condition for register binding to ensure a set of functional units to be perfectly power managed. This condition not only applies to data-flow intensive behaviors, but also to control-flow intensive behaviors. It leads to a straightforward power-managed (PM) register-binding algorithm, which uses this condition to preserve the previous values in the input registers of a functional unit during the states in which the unit is idle. The proposed algorithm is general and independent of the functional unit binding and scheduling algorithms. Hence, it can be easily incorporated into existing high-level synthesis systems. For the benchmarks we experimented with, an average 40.7% power reduction was achieved by our method at the cost of 6.9% average area overhead, compared to power-optimized register-transfer level circuits, which did not use PM register binding. Jiong Luo, Lin Zhong 0001, Yunsi Fei, Niraj K. Jha |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2003 | Energy Estimation for Extensible Processors
Yunsi Fei, Srivaths Ravi 0001, Anand Raghunathan, Niraj K. Jha |
DATE | 1 |
| 2003 | A comprehensive high-level synthesis system for control-flow intensive behaviorsabstractIn this paper, we describe a comprehensive high-level synthesis system for control-flow intensive as well as data-dominated behaviors. We propose a new control-data flow graph model to preserve the parallelism inherent in the application, as well as to facilitate high-level synthesis. Our algorithm, which is based on an iterative improvement strategy, performs clock selection, scheduling, module selection, resource allocation and assignment simultaneously to fully derive the benefits of design space exploration at the behavior level. The system can be used to optimize area, power or energy, by selecting the cost function accordingly. Experimental results show that for energy-optimized designs, energy is reduced by up to 79.4% (an average of 42.2%), with an average of 24.8% area overhead, compared to area-optimized designs. For power-optimized designs, power is reduced by up to 70.8% (an average of 56.7%), with an average of 25.2% area overhead, compared to area-optimized designs. No Vdd scaling is performed to obtain the above results. Tat Kee Tan, Jiong Luo, Yunsi Fei, Keith S. Vallerio, Lin Zhong 0001, Anand Raghunathan, Niraj K. Jha |
ACM Great Lakes Symposium on VLSI | 4 |
| 2002 | Register Binding Based Power Management for High-level Synthesis of Control-Flow Intensive BehaviorsabstractA circuit or circuit component that does not contain any spurious switching activity, i.e., activity that is not required by its specified functionality, is called perfectly power managed (PPM). We present a general sufficient condition for register binding to ensure that a given set of functional units is PPM. This condition not only applies to data-flow intensive (DFI) behaviors but also to control-flow intensive (CFI) behaviors. It leads to a straightforward power-managed (PM) register binding algorithm. The proposed algorithm is independent of the functional unit binding and scheduling algorithms. Hence, it can be easily incorporated into existing high-level synthesis systems. For the benchmarks we experimented with, an average 45.9% power reduction was achieved by our method at the cost of 7.7% average area overhead, compared to power-optimized register-transfer level (RTL) circuits which did not use PM register binding. Lin Zhong 0001, Jiong Luo, Yunsi Fei, Niraj K. Jha |
ICCD | 3 |