Patrick Schaumont

dblp:39/1269 · also Patrick R. Schaumont · DBLP profile ↗
← Back
118ranked-venue papers
23as first author
15since 2021 · last 2026
0000-0002-4586-5476ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 81 · 20 first-author · 10 since 2021Software engineering, systems software and programming languages · 24 · 5 first-authorSecurity and privacy · 22 · 1 first-author · 4 since 2021Computer networks · 7 · 1 since 2021Theory of computation · 3 · 2 first-authorArtificial intelligence and machine learning · 1Graphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Hierarchical EMFI Analysis on a RISC-V SoC
Dillibabu Shanmugam, Zhenyuan Liu 0005, Patrick Schaumont
ETS3
2026 Fault Analysis of Microscaling Formats on a RISC-V SoC
abstract
Microscaling (MX) formats share one exponent across an element block, creating a new fault surface: a single-bit flip in the shared exponent corrupts all block elements simultaneously. Five MX-compatible 8-bit encodings (MXINT8, MXFP8-E4M3, MXFP8-E5M2, LOG8-SUM, and LOG8-MAX) are compared through exhaustive single-bit weight fault injection across four workloads using gate-level simulation on the CAPRI1 RISC-V SoC, with the primary workload additionally validated on fabricated silicon. Format choice alone can substantially change vulnerability. Three MX-specific mechanisms explain this variation: block-shared exponent amplification, attention routing inversion from exponent-field faults, and log-domain fault filtering in the max-reduction variant. No single format dominates all axes: LOG8-MAX leads in speed and energy but is limited by accuracy on some attention workloads, MXINT8 provides the strongest fault resilience with full accuracy, and LOG8-SUM offers a balanced compromise across speed, accuracy, and resilience. MX element encoding should be treated not only as an accuracy-efficiency choice, but also as a security-relevant design decision.
Dillibabu Shanmugam, Patrick Schaumont
ACM Great Lakes Symposium on VLSI2
2025 Telescope: Top-Down Hierarchical Pre-silicon Side-channel Leakage Assessment in System-on-Chip Design
Zhenyuan Liu 0005, Andrew Malnicof, Arna Roy, Patrick Schaumont
AsiaCCS4
2025 SCAPEgoat: Side-channel Analysis Library
abstract
Side-channel analysis (SCA) is a growing field in hardware security where adversaries extract secret information from embedded devices by measuring physical observables like power consumption and electromagnetic emanation. SCA is a security assessment method used by governmental labs, standardization bodies, and researchers, where testing is not just limited to standardized cryptographic circuits, but it is expanded to AI accelerators, Post Quantum circuits, systems, etc. Despite its importance, SCA is performed on an ad hoc basis in the sense that its flow is not systematically optimized and unified among labs. As a result, the current solutions do not account for fair comparisons between analyses. Furthermore, neglecting the need for interoperability between datasets and SCA metric computation increases students’ barriers to entry. To address this, we introduce SCAPEgoat, a Python-based SCA library1with three key modules devoted to defining file format, capturing interfaces, and metric calculation. The custom file framework organizes side-channel traces using JSON for metadata, offering a hierarchical structure similar to HDF5 commonly applied in SCA, but more flexible and human-readable. The metadata can be queried with regular expressions, a feature unavailable in HDF5. Secondly, we incorporate memory-efficient SCA metric computations, which allow using our functions on resource-restricted machines. This is accomplished by partitioning datasets and leveraging statistics-based optimizations on the metrics. In doing so, SCAPEgoat makes the SCA more accessible to newcomers so that they can learn techniques and conduct experiments faster and with the possibility to expand on in the future.
Dev Mehta 0001, Trey Marcantino, Sam Karkache, Dillibabu Shanmugam, Patrick Schaumont, Fatemeh Ganji
VTS6
2024 Analysis of EM Fault Injection on Bit-sliced Number Theoretic Transform Software in Dilithium
abstract
Bitslicing is a software implementation technique that treats an N -bit processor datapath as N parallel single-bit datapaths. Bitslicing is particularly useful to implement data-parallel algorithms, algorithms that apply the same operation sequence to every element of a vector. Indeed, a bit-wise processor instruction applies the same logical operation to every single-bit slice. A second benefit of bitsliced execution is that the natural spatial redundancy of bitsliced software can support countermeasures against fault attacks. A k -redundant program on an N -bit processor then runs as N/k parallel redundant slices. In this contribution, we combine these two benefits of bitslicing to implement a fault countermeasure for the number-theoretic transform (NTT) . The NTT efficiently implements a polynomial multiplication. The internal symmetry of the NTT algorithm lends itself to a data-parallel implementation, and hence it is a good candidate for the redundantly bitsliced implementation. We implement a redundantly bitsliced NTT on an advanced 667MHz ARM Cortex-A9 processor, and study the fault coverage for the protected NTT under optimized electromagnetic fault injection (EMFI) . Our work brings two major contributions. First, we show for the first time how to develop a redundantly bitsliced version of the NTT. We integrate the protected NTT into a full Dilithium signature sequence. Second, we demonstrate an EMFI analysis on a prototype implementation of the Dilithium signature sequence on ARM Cortex-M9. We perform a detailed EM fault-injection parameter search to optimize the location, intensity and timing of injected EM pulses. We demonstrate that, under optimized fault injection parameters, about 10% of the injected faults become potentially exploitable. However, the redundantly bitsliced NTT design is able to catch the majority of these potentially exploitable faults, even when the remainder of the Dilithium algorithm as well as the control flow is left unprotected. To our knowledge, this is the first demonstration of a bitslice-redundant design of the NTT that offers distributed fault detection throughout the execution of the algorithm.
Richa Singh 0003, Saad Islam, Berk Sunar, Patrick Schaumont
ACM Trans. Embed. Comput. Syst.4
2023 Quantitative Fault Injection Analysis
Jakob Feldtkeller, Tim Güneysu, Patrick Schaumont
ASIACRYPT (4)3
2023 Lightning Talk: The Incredible Shrinking Black Box Model
abstract
A black box model is an assumption on the implementation of a cryptographic primitive to limit the capabilities of the attacker. Black boxes are a useful component in a proof of protocol correctness, but it is not obvious how to securely implement one in hardware. The current state of the art in tamper from open literature shows impressive efficiency and precision in prying open these black boxes. Secure hardware designers are compelled to shrink the black boxes with every new device generation, while making a careful assessment of the need to load, store and handle secrets in hardware.
Patrick Schaumont
DAC1
2022 SoK: Design Tools for Side-Channel-Aware Implementations
abstract
Side-channel attacks that leak sensitive information through a computing device's interaction with its physical environment have proven to be a severe threat to devices' security, particularly when adversaries have unfettered physical access to the device. Traditional approaches for leakage detection measure the physical properties of the device. Hence, they cannot be used during the design process and fail to provide root cause analysis. An alternative approach that is gaining traction is to automate leakage detection by modeling the device. The demand to understand the scope, benefits, and limitations of the proposed tools intensifies with the increase in the number of proposals. In this SoK, we classify approaches to automated leakage detection based on the model's source of truth. We classify the existing tools on two main parameters: whether the model includes measurements from a concrete device and the abstraction level of the device specification used for constructing the model. We survey the proposed tools to determine the current knowledge level across the domain and identify open problems. In particular, we highlight the absence of evaluation methodologies and metrics that would compare proposals' effectiveness from across the domain. We believe that our results help practitioners who want to use automated leakage detection and researchers interested in advancing the knowledge and improving automated leakage detection.
Ileana Buhan, Lejla Batina, Yuval Yarom, Patrick Schaumont
AsiaCCS4
2022 Signature Correction Attack on Dilithium Signature Scheme
abstract
Motivated by the rise of quantum computers, existing public-key cryptosystems are expected to be replaced by post-quantum schemes in the next decade in billions of devices. To facilitate the transition, NIST is running a standardization process which is currently in its final Round. Only three digital signature schemes are left in the competition, among which Dilithium and Falcon are the ones based on lattices. Besides security and performance, significant attention has been given to resistance against implementation attacks that target side-channel leakage or fault injection response. Classical fault attacks on signature schemes make use of pairs of faulty and correct signatures to recover the secret key which only works on deterministic schemes. To counter such attacks, Dilithium offers a randomized version which makes each signature unique, even when signing identical messages. In this work, we introduce a novel Signature Correction Attack which not only applies to the deterministic version but also to the randomized version of Dilithium and is effective even on constant-time implementations using AVX2 instructions. The Signature Correction Attack exploits the mathematical structure of Dilithium to recover the secret key bits by using faulty signatures and the public-key. It can work for any fault mechanism which can induce single bit-flips. For demonstration, we are using Rowhammer induced faults. Thus, our attack does not require any physical access or special privileges, and hence could be also implemented on shared cloud servers. Using Rowhammer attack, we inject bit flips into the secret key s1 of Dilithium, which results in incorrect signatures being generated by the signing algorithm. Since we can find the correct signature using our Signature Correction algorithm, we can use the difference between the correct and incorrect signatures to infer the location and value of the flipped bit without needing a correct and faulty pair. To quantify the reduction in the security level, we perform a thorough classical and quantum security analysis of Dilithium and successfully recover 1,851 bits out of 3,072 bits of secret key$s_{1}$for security level 2. Fully recovered bits are used to reduce the dimension of the lattice whereas partially recovered coefficients are used to to reduce the norm of the secret key coefficients. Further analysis for both primal and dual attacks shows that the lattice strength against quantum attackers is reduced from 2128to 281while the strength against classical attackers is reduced from 2141 to 289. Hence, the Signature Correction Attack may be employed to achieve a practical attack on Dilithium (security level 2) as proposed in Round 3 of the NIST post-quantum standardization process.
Saad Islam, Koksal Mus, Richa Singh 0003, Patrick Schaumont, Berk Sunar
EuroS&P4
2022 Leverage the Average: Averaged Sampling in Pre-Silicon Side-Channel Leakage Assessment
abstract
Pre-silicon side-channel leakage assessment is a useful tool to identify hardware vulnerabilities at design time, but it requires many high-resolution power traces and increases the power simulation cost of the design. By downsampling and averaging these high-resolution traces, we show that the power simulation cost can be considerably reduced without significant loss of side-channel leakage assessment quality. We introduce a theoretical basis for our claims. Our results demonstrate up to 6.5-fold power-simulation speed improvement on a gate-level side-channel leakage assessment of a RISC-V SoC. Furthermore, we clarify the conditions under which the averaged sampling technique can be successfully used.
Pantea Kiaei, Zhenyuan Liu 0005, Patrick Schaumont
ACM Great Lakes Symposium on VLSI3
2022 Threat Modeling and Risk Analysis for Miniaturized Wireless Biomedical Devices
abstract
The landscape of miniaturized wireless biomedical devices (MWBDs), including various injectables, ingestibles, implantables, and wearables, is rapidly expanding as proactive mobile healthcare proliferates. While the growth of MWBDs increases the flexibility of medical services, the adoption of these technologies poses privacy and security risks to their users. As a result, while being restricted in resources (size, power, processing, and storage), these devices require trust and must be at least minimally secure in the face of evolving threats. Making MWBDs secure begins with threat modeling. Therefore, this research reviews and summarizes the information on threat modeling applicable to MWBDs. Then, we propose a domain-specific qualitative-quantitative threat model that aims to help the designers and manufacturers of MWBDs to identify threats and embed security in their designs in the premarket phase of the lifecycle of an MWBD. This model is tailored to a wide range of MWBDs. Among the different stakeholders, this model focuses on the user. It also prioritizes noninvasive direct attacks against telemetry interfaces. To discuss the advantages and disadvantages of the proposed model, it is compared to some other threat models. To illustrate how the model can be adopted by a threat-modeling team, it is then applied to representative case studies from each category of MWBDs. The outcomes of the performed risk analysis reveal that the model is easy to apply and sufficient to disclose threats.
Vladimir Vakhter, Betul Soysal, Patrick Schaumont, Ulkuhan Guler 0001
IEEE Internet Things J.3
2022 ScatterVerif: Verification of Electronic Boards Using Reflection Response of Power Distribution Network
abstract
The globalization of electronic systems’ fabrication has made some of our most critical systems vulnerable to supply chain attacks. Implanting spy chips on the printed circuit boards (PCBs) or replacing genuine components with counterfeit/recycled ones are examples of such attacks. Unfortunately, conventional attack detection schemes for PCBs are ad hoc, costly, unscalable, and error prone. This work introduces a holistic physical verification framework for PCBs, called ScatterVerif , based on the characterization of the PCBs’ power distribution network. First, we demonstrate how scattering parameters, frequently used for impedance characterization of RF circuits, can characterize the entire PCB with a single measurement. Second, we present how a class of machine learning algorithms, namely the Gaussian mixture model, can be applied to the measurements to automatically classify/cluster the genuine and tampered/counterfeit PCBs. We show that these attacks affect the overall impedance of a PCB differently in various frequency ranges, hence the conventional impedance measurements using a constant-frequency electrical stimulus might leave the attack undetected. We conduct extensive experiments on counterfeit and tampered devices and demonstrate that these attacks can be detected with high confidence. Finally, we show that the acquired data from the power distribution network characterization can also be deployed for fingerprinting genuine PCBs.
Tahoura Mosavirik, Fatemeh Ganji, Patrick Schaumont, Shahin Tajik
ACM J. Emerg. Technol. Comput. Syst.3
2022 Benchmarking and Configuring Security Levels in Intermittent Computing
abstract
Intermittent computing derives its name from the intermittent character of the power source used to drive the computing, typically an energy harvester of ambient energy sources. Intermittent computing is characterized by frequent transitions between the powered and the non-powered state. To enable the processor to quickly recover from unexpected power loss, regular checkpoints store the run-time state of the program, including variables, control information, and machine state. In sensitive applications such as logged measurements, checkpoints must be secured against tamper and replay. We investigate the overhead of creating, securing, and restoring checkpoints with respect to the application. We propose a configurable checkpoint security setting that leverages application properties to reduce overhead of checkpoint security and implement the same using a secure checkpointing protocol. We discuss a prototype implementation for a FRAM-based micro-controller, and we characterize the cost of adding and configuring security to traditional checkpointing using a suite of embedded benchmark applications.
Archanaa S. Krishnan, Patrick Schaumont
ACM Trans. Embed. Comput. Syst.2
2021 Rewrite to Reinforce: Rewriting the Binary to Apply Countermeasures against Fault Injection
abstract
Fault injection attacks can cause errors in software for malicious purposes. Oftentimes, vulnerable points of a program are detected after its development. It is therefore critical for the user of the program to be able to apply last-minute security assurance to the executable file without having access to the source code. In this work, we explore two methodologies based on binary rewriting that aid in injecting countermeasures in the binary file. The first approach injects countermeasures by reassembling the disassembly whereas the second approach leverages a full translation to a high-level IR and lowering that back to the target architecture.
Pantea Kiaei, Cees-Bart Breunesse, Mohsen Ahmadi, Patrick Schaumont, Jasper Van Woudenberg
DAC4
2021 Socially-Distant Hands-On Labs for a Real-time Digital Signal Processing Course
abstract
Due to the pandemic, the vast majority of courses taught in the second half of 2020 proceeded online. We discuss the conversion of a Real-time Digital Signal Processing (DSP) course with traditional in-person hands-on labs to a hybrid (in-person and online) setting. This senior-level undergraduate course is built around the development and testing of DSP software. In the online version of the course, students use a small microcontroller kit and a USB oscilloscope, and they work in online teams. We describe the lab projects and we highlight the mechanisms that enable the students to collaborate on labs. We provide early feedback from students from the first offering of the modified course, and conclude with a list of planned improvements.
Patrick Schaumont
ACM Great Lakes Symposium on VLSI1
2020 ASHES 2020: 4th Workshop on Attacks and Solutions in Hardware Security
Chip-Hong Chang, Stefan Katzenbeisser 0001, Ulrich Rührmair, Patrick Schaumont
CCS4
2020 Using Universal Composition to Design and Analyze Secure Complex Hardware Systems
abstract
Modern hardware typically is characterized by a multitude of interacting physical components and software mechanisms. To address this complexity, security analysis should be modular: We would like to formulate and prove security properties of individual components, and then deduce the security of the overall design (encompassing hardware and software) from the security of the components. While this seems like an elusive goal, we argue that this is essentially the only feasible way to provide rigorous security analysis of modern hardware.This paper investigates the possibility of using the Universally Composable (UC) security framework towards this aim. The UC framework has been devised and successfully used in the theoretical cryptography community to study and formally prove security of arbitrarily interleaving cryptographic protocols. In particular, a sophisticated analytical toolbox has been developed using this framework. We provide an introduction to this frame-work, and investigate, via a number of examples, ways by which this framework can be used to facilitate a novel type of modular security analysis. This analysis applies to combined hardware and software systems, and investigates their security against attacks that combine both physical and digital steps.
Ran Canetti, Marten van Dijk, Hoda Maleki, Ulrich Rührmair, Patrick Schaumont
DATE5
2020 Towards Secure Composition of Integrated Circuits and Electronic Systems: On the Role of EDA
abstract
Modern electronic systems become evermore complex, yet remain modular, with integrated circuits (ICs) acting as versatile hardware components at their heart. Electronic design automation (EDA) for ICs has focused traditionally on power, performance, and area. However, given the rise of hardware-centric security threats, we believe that EDA must also adopt related notions like secure by design and secure composition of hardware. Despite various promising studies, we argue that some aspects still require more efforts, for example: effective means for compilation of assumptions and constraints for security schemes, all the way from the system level down to the "bare metal"; modeling, evaluation, and consideration of security-relevant metrics; or automated and holistic synthesis of various countermeasures, without inducing negative cross-effects.In this paper, we first introduce hardware security for the EDA community. Next we review prior (academic) art for EDA-driven security evaluation and implementation of countermeasures. We then discuss strategies and challenges for advancing research and development toward secure composition of circuits and systems.
Johann Knechtel, Elif Bilge Kavun, Francesco Regazzoni 0001, Annelie Heuser, Anupam Chattopadhyay, Debdeep Mukhopadhyay, Soumyajit Dey, Yunsi Fei, Yaacov Belenky, Itamar Levi, Tim Güneysu, Patrick Schaumont, Ilia Polian
DATE12
2020 TreeRNN: Topology-Preserving Deep Graph Embedding and Learning
abstract
General graphs are difficult for learning due to their irregular structures. Existing works employ message passing along graph edges to extract local patterns using customized graph kernels, but few of them are effective for the integration of such local patterns into global features. In contrast, in this paper we study the methods to transfer the graphs into trees so that explicit orders are learned to direct the feature integration from local to global. To this end, we apply the breadth first search (BFS) to construct trees from the graphs, which adds direction to the graph edges from the center node to the peripheral nodes. In addition, we proposed a novel projection scheme that transfer the trees to image representations, which is suitable for conventional convolution neural networks (CNNs) and recurrent neural networks (RNNs). To best learn the patterns from the graph-tree-images, we propose TreeRNN, a 2D RNN architecture that recurrently integrates the image pixels by rows and columns to help classify the graph categories. We evaluate the proposed method on several graph classification datasets, and manage to demonstrate comparable accuracy with the state-of-the-art on MUTAG, PTC-MR and NCI1 datasets.
Yecheng Lyu, Ming Li 0073, Xinming Huang 0001, Ulkuhan Guler 0001, Patrick Schaumont
ICPR5
2019 ASHES 2019: 3rd Workshop on Attacks and Solutions in Hardware Security
Chip-Hong Chang, Daniel E. Holcomb, Francesco Regazzoni 0001, Ulrich Rührmair, Patrick Schaumont
CCS5
2019 Secure Intermittent Computing Protocol: Protecting State Across Power Loss
abstract
Intermittent computing systems execute long-running tasks under a transient power supply such as an energy harvesting power source. During a power loss, they save intermediate program state as a checkpoint into write-efficient non-volatile memory. When the power is restored, the system state is reconstructed from the checkpoint, and the long-running computation continues. We analyze the security risks when power interruption is used as an attack vector, and we demonstrate the need to protect the integrity, authenticity, confidentiality, continuity, and freshness of checkpointed data. We propose a secure checkpointing technique called the Se-cure Intermittent Computing Protocol (SICP). The proposed protocol has the following properties. First, it associates every checkpoint with a unique power-on state to checkpoint replay. Second, every checkpoint is cryptographically chained to its predecessor, providing continuity, which enables the programmer to carry run-time security properties such as attested program images across power loss events. Third, SICP is atomic and resistant to power loss. We demonstrate a prototype implementation of SICP on an MSP430 microcontroller, and we investigate the overhead of SICP for several cryptographic kernels. To the best of our knowledge, this is the first work to provide a robust solution to secure intermittent computing.
Archanaa S. Krishnan, Charles Suslowicz, Daniel Dinu, Patrick Schaumont
DATE4
2019 A Secure Exception Mode for Fault-Attack-Resistant Processing
abstract
Fault attacks are a known threat to secure embedded implementations. We propose a generic technique to detect and react to fault attacks on embedded software. The countermeasure combines a micro-architecture extension in hardware with a secure trap in software. The combined extension leads to a secure exception mode to handle fault attacks. The microprocessor hardware uses a low-level hardware checkpointing mechanism to recover from fault injection. A high-level secure trap in software then enables an application-specific response. The trap is user-defined and can be co-developed with the application. The combination of hardware fault detection and recovery, with a high-level fault response policy in software leads to significantly lower overhead when compared to traditional redundancy-based techniques in hardware or software. We demonstrate a prototype implementation of the proposed secure exception mode. The prototype is based on a modified LEON3 processor and it is able to detect and respond to setup-time violation attacks. We have realized the design in a 180 nm standard cell ASIC with integrated memory. Using several driver application examples, we characterize the software and hardware overhead of the proposed solution, and we compare it to the conventional redundancy-based solutions. In our understanding this is the first proof-in-silicon processor to offer a comprehensive secure exception mode against fault-injection attacks.
Bilgiday Yuce, Chinmay Deshpande, Marjan Ghodrati, Abhishek Bendre, Leyla Nazhandali, Patrick Schaumont
IEEE Trans. Dependable Secur. Comput.6
2019 Editorial TVLSI Positioning - Continuing and Accelerating an Upward Trajectory
abstract
I. VLSI Systems: A Glance Into The Last Decades Since their inception in 1970s, VLSI systems have enabled several new technological capabilities and made them accessible to an unceasingly wider range of users, reaching a scale that has been exponentially increasing over the decades[1](seeFig. 1). Relentless integration of more complex systems has driven such remarkable evolution, as made possible by the inexorable miniaturization. As shown inFig. 1, more functionality has been crammed in a consistently smaller form factor, as exemplified by the physical volume shrinking of computers by 100 X/decade[2],[3]. At the same time, the energy per task has been decreasing at 10–100 X/decade, as shown inFig. 2, for several systems and system-on-chip subsystems[4]. This allowed packing more capabilities into the same power envelope, as generally observed in the electronic systems, even before the advent of the integrated circuit[5].
Massimo Alioto, Magdy S. Abadir, Tughrul Arslan, Chirn Chye Boon, Andreas Peter Burg, Chip-Hong Chang, Meng-Fan Chang, Yao-Wen Chang, Poki Chen, Pasquale Corsonello, Paolo Crovetti, Shiro Dosho, Rolf Drechsler, Ibrahim M. Elfadel, Ruonan Han 0001, Masanori Hashimoto, Chun-Huat Heng, Deuk Hyoun Heo, Tsung-Yi Ho, Houman Homayoun, Yuh-Shyan Hwang, Ajay Joshi, Rajiv V. Joshi, Tanay Karnik, Chulwoo Kim, Tony Tae-Hyoung Kim, Jaydeep P. Kulkarni, Volkan Kursun, Yoonmyung Lee, Hai Li 0001, Huawei Li 0001, Prabhat Mishra 0001, Baker Mohammad, Mehran Mozaffari Kermani, Makoto Nagata, Koji Nii, Partha Pratim Pande, Bipul Chandra Paul, Vasilis F. Pavlidis, José Pineda de Gyvez, Ioannis Savidis, Patrick Schaumont, Fabio Sebastiano, Anirban Sengupta 0003, Mingoo Seok, Mircea R. Stan, Mark Tehranipoor, Aida Todri, Marian Verhelst, Valerio Vignoli, Xiaoqing Wen, Jiang Xu 0001, Wei Zhang 0012, Zhengya Zhang, Jun Zhou 0017, Mark Zwolinski, Stacey Weber
IEEE Trans. Very Large Scale Integr. Syst.42
2018 Inducing local timing fault through EM injection
abstract
Electromagnetic fault injection (EMFI) is an efficient class of physical attacks that can compromise the immunity of secure cryptographic algorithms. Despite successful EMFI attacks, the effects of electromagnetic injection (EM) on a processor are not well understood. This paper presents a bottom-up analysis of EMFI effects on a RISC microprocessor. We study these effects at three levels: at the wire-level, at the chip-network level, and at the gate-level considering parameters such as EM-injection location and timing. We conclude that EMFI induces local timing errors implying current timing attack detection and prevention techniques can be adapted to overcome EMFI.
Marjan Ghodrati, Bilgiday Yuce, Surabhi Gujar, Chinmay Deshpande, Leyla Nazhandali, Patrick Schaumont
DAC6
2018 Eliminating timing side-channel leaks using program repair
abstract
We propose a method, based on program analysis and transformation, for eliminating timing side channels in software code that implements security-critical applications. Our method takes as input the original program together with a list of secret variables (e.g., cryptographic keys, security tokens, or passwords) and returns the transformed program as output. The transformed program is guaranteed to be functionally equivalent to the original program and free of both instruction- and cache-timing side channels. Specifically, we ensure that the number of CPU cycles taken to execute any path is independent of the secret data, and the cache behavior of memory accesses, in terms of hits and misses, is independent of the secret data. We have implemented our method in LLVM and validated its effectiveness on a large set of applications, which are cryptographic libraries with 19,708 lines of C/C++ code in total. Our experiments show the method is both scalable for real applications and effective in eliminating timing side channels.
Meng Wu 0001, Shengjian Guo, Patrick Schaumont, Chao Wang 0001
ISSTA3
2018 Special Section on Secure Computer Architectures
abstract
The papers in this special section focus on security in computer architectures. Computer architectures are profoundly affected by a new security landscape, caused by the dramatic evolution of information technology over the past decade. First, secure computer architectures have to support a wide range of security applications that extend well beyond the desktop environment, and that also include handheld, mobile, and embedded architectures, as well as high-end computing servers. Second, secure computer architectures have to support new applications of information security and privacy, as well as new information security standards. Third, secure computer architectures have to be protected and be tamperresistant at multiple abstraction levels, covering network, software, and hardware. This Special Section in Transactions on Computers aims to capture this evolving landscape of secure computing architectures, to build a vision of opportunities and unresolved challenges.
Patrick Schaumont, Ruby B. Lee, Ronald Perez, Guido Bertoni
IEEE Trans. Computers1
2017 Security in the Internet of Things: A challenge of scale
abstract
Technological scaling has offered a windfall of benefits to electronics design. Increased transistor density has offered an exponential increase in computing capabilities over time, but without a corresponding increase in system cost. Information security has its own success story with scaling. Cryptographic algorithms become exponentially harder to break through a mere linear increase in encryption complexity or in key-length. In the Internet of Things, scaling is as much a security liability as it is an advantage. These security liabilities are new, poorly understood and poorly regulated. Some examples include the following: privacy of IoT data in the cloud; the safety consequences of poor information security in cyber-physical systems; the liabilities of long-lifetime devices that use outdated or poorly tested information security; the performance-limited information security in devices that run on the outskirts of the IoT using nothing but harvested energy. In this contribution we consider the security landscape for IoT. We consider the technological consequences of securely extending the Internet into the physical world of things. We identify current limitations, ongoing research efforts, and open challenges for the design community.
Patrick Schaumont
DATE1
2017 Vector Instruction Set Extensions for Efficient Computation of Keccak
abstract
We investigate the design of a new instruction set for the KECCAK permutation, a cryptographic kernel for hashing, authenticated encryption, keystream generation and random-number generation. KECCAK is the basis of the SHA-3 standard and the newly proposed KEYAK and KETJE authenticated ciphers. We develop the instruction extensions for a 128-bit interface, commonly available in the vector-processing unit of many modern processors. We examine the trade-off between flexibility and efficiency, and we propose a set of six custom instructions to support a broad range of KECCAK-based cryptographic applications. We motivate our custom-instruction selections using a design space exploration that considers various methods of partitioning the state and the operations of the KECCAK permutation, and we demonstrate an efficient implementation of this permutation with the proposed instructions. To evaluate their performance, we integrate a simulation model of the proposed ARM NEON vector instructions into the GEM5 micro-architecture simulator. With this simulation model, we evaluate the performance improvement for several cryptographic operations that use the KECCAK permutation. Compared to a state-of-the-art NEON software implementation, we demonstrate a performance improvement of 2.2x for SHA-3. Compared to optimized 32-bit assembly programming, we demonstrate a performance improvement of 2.6x, 1.6x, and 1.4x for RIVER KEYAK, KETJESR and KETJEJR respectively. The proposed instructions require 4,658 gate-equivalent (GE) in 90 nm, which represents only a tiny fraction of the hardware cost of a modern processor.
Hemendra K. Rawat, Patrick Schaumont
IEEE Trans. Computers2
2017 Analyzing the Fault Injection Sensitivity of Secure Embedded Software
abstract
Fault attacks on cryptographic software use faulty ciphertext to reverse engineer the secret encryption key. Although modern fault analysis algorithms are quite efficient, their practical implementation is complicated because of the uncertainty that comes with the fault injection process. First, the intended fault effect may not match the actual fault obtained after fault injection. Second, the logic target of the fault attack, the cryptographic software, is above the abstraction level of physical faults. The resulting uncertainty with respect to the fault effects in the software may degrade the efficiency of the fault attack, resulting in many more trial fault injections than the amount predicted by the theoretical fault attack. In this contribution, we highlight the important role played by the processor microarchitecture in the development of a fault attack. We introduce the microprocessor fault sensitivity model to systematically capture the fault response of a microprocessor pipeline. We also propose Microarchitecture-Aware Fault Injection Attack (MAFIA). MAFIA uses the fault sensitivity model to guide the fault injection and to predict the fault response. We describe two applications for MAFIA. First, we demonstrate a biased fault attack on an unprotected Advanced Encryption Standard (AES) software program executing on a seven-stage pipelined Reduced Instruction Set Computer (RISC) processor. The use of the microprocessor fault sensitivity model to guide the attack leads to an order of magnitude fewer fault injections compared to a traditional, blind fault injection method. Second, MAFIA can be used to break known software countermeasures against fault injection. We demonstrate this by systematically breaking a collection of state-of-the-art software fault countermeasures. These two examples lead to the key conclusion of this work, namely that software fault attacks become much more harmful and effective when an appropriate microprocessor fault sensitivity model is used. This, in turn, highlights the need for better fault countermeasures for software.
Bilgiday Yuce, Nahid Farhady Ghalaty, Chinmay Deshpande, Harika Santapuri, Conor Patrick, Leyla Nazhandali, Patrick Schaumont
ACM Trans. Embed. Comput. Syst.7
2016 A design method for remote integrity checking of complex PCBs
Aydin Aysu, Shravya Gaddam, Harsha Mandadi, Carol Pinto, Luke Wegryn, Patrick Schaumont
DATE6
2016 Software Fault Resistance is Futile: Effective Single-Glitch Attacks
abstract
Fault attacks are a serious threat for the secure embedded software running on a wide spectrum of embedded devices. Fault attacks can be thwarted using countermeasures in software. Among them, instruction-level countermeasures provide a fine-grained protection by executing redundant copies of an assembly instruction, and verifying their results for fault detection. It is assumed that this fine-grained security can only be broken by injecting multiple faults with expensive tools. In this work, we break the security of state-of-the-art instruction-level countermeasures by injecting single clock glitches with a low-cost fault injection setup. We first analyze their vulnerabilities by considering micro-architectural aspects such as pipelining effects. Second, we experimentally demonstrate the feasibility of exploiting these vulnerabilities on a SAKURA-G board. Finally, as a case study, we apply a recent biased fault attack on a fault-resistant software implementation of LED block cipher, and retrieve its secret key.
Bilgiday Yuce, Nahid Farhady Ghalaty, Harika Santapuri, Chinmay Deshpande, Conor Patrick, Patrick Schaumont
FDTC6
2016 Lightweight Fault Attack Resistance in Software Using Intra-instruction Redundancy
Conor Patrick, Bilgiday Yuce, Nahid Farhady Ghalaty, Patrick Schaumont
SAC4
2016 Keymill: Side-Channel Resilient Key Generator, A New Concept for SCA-Security by Design - A New Concept for SCA-Security by Design
Mostafa M. I. Taha, Arash Reyhani-Masoleh, Patrick Schaumont
SAC3
2016 Compact and low-power ASIP design for lightweight PUF-based authentication protocols
abstract
There is a disconnection between the theory and the practice of lightweight physical unclonable function (PUF)‐based protocols. At a theoretical level, there exist several PUF‐based authentication protocols with unique features and novel efficiency claims, but most of these solutions lack real‐world implementations with simple performance figures. On the other hand, practical protocol implementations are ad‐hoc designs fixed to a specific functionality and with limited area optimisations. This work aims to bring these approaches on PUF protocols closer. The authors’ contribution is twofold. First, they provide a novel ASIP (application‐specific instruction set processor) that can efficiently execute PUF‐based authentication protocols. The key novelty of the proposed ASIP is optimisation for area without degrading the performance. Second, they demonstrate the capability of their ASIP by mapping three secure PUF‐based authentication protocols and benchmark their execution time, memory footprint, communication overhead, and power/energy consumption. Their results demonstrate the advantage of ASIP over dedicated architectures and also as opposed to general‐purpose programming on an MSP430. The results further demonstrate various efficiency metrics that can be used to compare PUF‐based protocol implementations.
Aydin Aysu, Ege Gulcan, Daisuke Moriyama, Patrick Schaumont
IET Inf. Secur.4
2016 Precomputation Methods for Hash-Based Signatures on Energy-Harvesting Platforms
abstract
Energy-harvesting techniques can be combined with wireless embedded sensors to obtain battery-free platforms with an extended lifetime. Although energy-harvesting offers a continuous supply of energy, the delivery rate is typically limited to a few Joules per day. This is a severe constraint to the achievable computing throughput on the embedded sensor node, and to the achievable latency obtained from applications running on those nodes. In this paper, we address these constraints with precomputation. The idea is to reduce the amount of computations required in response to application inputs, by partitioning the algorithm in an offline part, computed before the inputs are available, and an online part, computed in response to the actual input. We show that this technique works well on hash-based cryptographic signatures, which have a complex key generation for each new message that requires a signature. By precomputing the key-material, and by storing it as run-time coupons in non-volatile memory, there is a drastic reduction of the run-time energy needs for a signature, and a drastic reduction of the run-time latency to generate it. For a Winternitz hash-based scheme at 84-bit quantum security level on a MSP430 microcontroller, we measured a run-time energy reduction of 11.9$\times$and a run-time latency reduction of 23.5$\times$.
Aydin Aysu, Patrick Schaumont
IEEE Trans. Computers2
2015 End-To-End Design of a PUF-Based Privacy Preserving Authentication Protocol
Aydin Aysu, Ege Gulcan, Daisuke Moriyama, Patrick Schaumont, Moti Yung
CHES4
2015 Improving Fault Attacks on Embedded Software Using RISC Pipeline Characterization
abstract
A fault attack becomes more efficient when the fault behavior, the response of a device to a fault injection, is precisely understood. In this paper, we present a methodology for fault attacks and their analysis on pipelined RISC processors. For complex hardware structures such as microprocessor pipelines, modeling the fault behavior can become challenging. By analyzing the structure of the RISC pipeline, we obtain insight into the most likely faults, and we are able to pinpoint the most sensitive points during execution of a cryptographic software program. We use this result to apply a recent class of fault injection attacks, so-called biased fault injection attacks, to two different software implementations of AES. Our target microprocessor is a 7-stage pipeline LEON3, mapped into a Spartan6 FPGA. The paper explains the methodology, the fault injection setup, and the fault analysis on the embedded software design of AES. Our results are useful for embedded software designers who have a need to understand the fault attack sensitivity of their implementation, as well as for security engineers who are in charge of improving countermeasures, in hardware or in software, against fault attacks.
Bilgiday Yuce, Nahid Farhady Ghalaty, Patrick Schaumont
FDTC3
2015 Quantitative Masking Strength: Quantifying the Power Side-Channel Resistance of Software Code
abstract
Many commercial systems in the embedded space have shown weakness against power analysis-based side-channel attacks in recent years. Random masking is a commonly used technique for removing the statistical dependency between the sensitive data and the side-channel information. However, the process of designing masking countermeasures is both labor intensive and error prone. Furthermore, there is a lack of formal methods for quantifying the actual strength of a countermeasure implementation. Security design errors may therefore go undetected until the side-channel leakage is physically measured and evaluated. We show a better solution based on static analysis of C source code. We introduce the new notion of quantitative masking strength (QMS) to estimate the amount of information leakage from software through side channels. Once the user has identified the sensitive variables, the QMS can be automatically computed from the source code of a countermeasure implementation. Our experiments, based on measurement on real devices, show that the QMS accurately reflects the side-channel resistance of the software implementation.
Hassan Eldib, Chao Wang 0001, Mostafa M. I. Taha, Patrick Schaumont
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2015 The Future of Real-Time Security: Latency-Optimized Lattice-Based Digital Signatures
abstract
Advances in quantum computing have spurred a significant amount of research into public-key cryptographic algorithms that are resistant against postquantum cryptanalysis. Lattice-based cryptography is one of the important candidates because of its reasonable complexity combined with reasonable signature sizes. However, in a postquantum world, not only the cryptography will change but also the computing platforms. Large amounts of resource-constrained embedded systems will connect to a cloud of powerful server computers. We present an optimization technique for lattice-based signature generation on such embedded systems; our goal is to optimize latency rather than throughput. Indeed, on an embedded system, the latency of a single signature for user identification or message authentication is more important than the aggregate signature generation rate. We build a high-performance implementation using hardware/software codesign techniques. The key idea is to partition the signature generation scheme into offline and online phases. The signature scheme allows this separation because a large portion of the computation does not depend on the message to be signed and can be handled before the message is given. Then, we can map complex precomputation operations in software on a low-cost processor and utilize hardware resources to accelerate simpler online operations. To find the optimum hardware architecture for the target platform, we define and explore the design space and implement two design configurations. We realize our solutions on the Altera Cyclone-IV CGX150 FPGA. The implementation consists of a NIOS soft-core processor and a low-latency hash and polynomial multiplication engine. On average, the proposed low-latency architecture can generate a signature with a latency of 96 clock cycles at 40MHz, resulting in a response time of 2.4μs for a signing request. On equivalent platforms, this corresponds to a performance improvement of 33 and 105 times compared to previous hardware and software implementations, respectively.
Aydin Aysu, Bilgiday Yuce, Patrick Schaumont
ACM Trans. Embed. Comput. Syst.3
2015 Introduction for Embedded Platforms for Cryptography in the Coming Decade
abstract
editorial Free Access Share on Introduction for Embedded Platforms for Cryptography in the Coming Decade Editors: Patrick Schaumont Virginia Tech, USA Virginia Tech, USAView Profile , Maire O'Neill Queen's University Belfast, United Kingdom Queen's University Belfast, United KingdomView Profile , Tim Güneysu Ruhr University Bochum, Germany Ruhr University Bochum, GermanyView Profile Authors Info & Claims ACM Transactions on Embedded Computing SystemsVolume 14Issue 3May 2015 Article No.: 40pp 1–3https://doi.org/10.1145/2745710Published:21 April 2015Publication History 2citation284DownloadsMetricsTotal Citations2Total Downloads284Last 12 Months15Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my Alerts New Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
Patrick Schaumont, Máire O'Neill, Tim Güneysu
ACM Trans. Embed. Comput. Syst.1
2015 Key Updating for Leakage Resiliency With Application to AES Modes of Operation
abstract
Side-channel analysis (SCA) exploits the information leaked through unintentional outputs (e.g., power consumption) to reveal the secret key of cryptographic modules. The real threat of SCA lies in the ability to mount attacks over small parts of the key and to aggregate information over different encryptions. The threat of SCA can be thwarted by changing the secret key at every run. Indeed, many contributions in the domain of leakage resilient cryptography tried to achieve this goal. However, the proposed solutions were computationally intensive and were not designed to solve the problem of the current cryptographic schemes. In this paper, we propose a generic framework of lightweight key updating that can protect the current cryptographic standards and evaluate the minimum requirements for heuristic SCA-security. Then, we propose a complete solution to protect the implementation of any standard mode of Advanced Encryption Standard. Our solution maintains the same level of SCA-security (and sometimes better) as the state of the art, at a negligible area overhead while doubling the throughput of the best previous work.
Mostafa M. I. Taha, Patrick Schaumont
IEEE Trans. Inf. Forensics Secur.2
2014 QMS: Evaluating the Side-Channel Resistance of Masked Software from Source Code
abstract
Many commercial systems in the embedded space have shown weakness against power analysis based side-channel attacks in recent years. Designing countermeasures to defend against such attacks is both labor intensive and error prone. Furthermore, there is a lack of formal methods for quantifying the actual strength of a countermeasure implementation. Security design errors may therefore go undetected until the side-channel leakage is physically measured and evaluated. We show a better solution based on static analysis of C source code. We introduce the new notion of Quantitative Masking Strength (QMS) to estimate the amount of information leakage from software through side channels. The QMS can be automatically computed from the source code of a countermeasure implementation. Our experiments, based on side-channel measurement on real devices, show that the QMS accurately quantifies the side-channel resistance of the software implementation.
Hassan Eldib, Chao Wang 0001, Mostafa M. I. Taha, Patrick Schaumont
DAC4
2014 Analyzing and eliminating the causes of fault sensitivity analysis
abstract
Fault Sensitivity Analysis (FSA) is a new type of side-channel attack that exploits the relation between the sensitive data and the faulty behavior of a circuit, the so-called fault sensitivity. This paper analyzes the behavior of different implementations of AES S-box architectures against FSA, and proposes a systematic countermeasure against this attack. This paper has two contributions. First, we study the behavior and structure of several S-box implementations, to understand the causes behind the fault sensitivity. We identify two factors: the timing of fault sensitive paths, and the number of logic levels of fault sensitive gates within the netlist. Next, we propose a systematic countermeasure against FSA. The countermeasure masks the effect of these factors by intelligent insertion of delay elements. We evaluate our methodology by means of an FPGA prototype with built-in timing-measurement. We show that FSA can be thwarted at low hardware overhead. Compared to earlier work, our method operates at the logic-level, is systematic, and can be easily generalized to bigger circuits.
Nahid Farhady Ghalaty, Aydin Aysu, Patrick Schaumont
DATE3
2014 Differential Fault Intensity Analysis
abstract
Recent research has demonstrated that there is no sharp distinction between passive attacks based on side-channel leakage and active attacks based on fault injection. Fault behavior can be processed as side-channel information, offering all the benefits of Differential Power Analysis including noise averaging and hypothesis testing by correlation. This paper introduces Differential Fault Intensity Analysis, which combines the principles of Differential Power Analysis and fault injection. We observe that most faults are biased - such as single-bit, two-bit, or three-bit errors in a byte - and that this property can reveal the secret key through a hypothesis test. Unlike Differential Fault Analysis, we do not require precise analysis of the fault propagation. Unlike Fault Sensitivity Analysis, we do not require a fault sensitivity profile for the device under attack. We demonstrate our method on an FPGA implementation of AES with a fault injection model. We find that with an average of 7 fault injections, we can reconstruct a full 128-bit AES key.
Nahid Farhady Ghalaty, Bilgiday Yuce, Mostafa M. I. Taha, Patrick Schaumont
FDTC4
2014 Hardware-software co-design for heterogeneous multiprocessor sensor nodes
abstract
To meet the needs of innovative sensor network applications, sensor nodes have long evolved from underpowered single microcontroller designs into complex architectures that accommodate multiple processors and Field Programmable Gate Arrays (FPGAs). We address the problem of conceiving and implementing programs for such sensor node architectures. We rely on a universal, layered hardware/software interface that provides seamless interconnection between tasks running on a micro-controller and tasks running on a FPGA. Resource sharing is handled transparently through a shared-bus communication architecture. We demonstrate our methodology, through a heterogeneous sensor node simulator called SUNSHINE [1] for an application running the sensor nodes. We validate SUNSHINE by demonstrating the applications on a multiprocessor sensor node's testbed, which consists of a FPGA, a microcontroller and a radio front-end.
Srikrishna Iyer, Xiangwei Zheng 0002, Patrick Schaumont, Yaling Yang
GLOBECOM4
2014 SMT-Based Verification of Software Countermeasures against Side-Channel Attacks
Hassan Eldib, Chao Wang 0001, Patrick Schaumont
TACAS3
2014 Application design and performance evaluation for multiprocessor sensor nodes
abstract
Currently, the main components of most sensor nodes are a microcontroller and a radio. Their real-time and peak performance would be a bottleneck when executing computation-intensive tasks. Also, energy consumption may be high due to the long task-execution time. Several works [2]-[6] demonstrate that adding a coprocessor would be a solution to the problems. So far, no work have been done to analyze the performance of a multiprocessor sensor nodes with a FPGA coprocessor. In this paper, a design flow for heterogeneous multiprocessor sensor nodes is provided. Then, the performance comparison between multiprocessor and single processor sensor node's time and energy consumption is provided by executing applications on our in-house designed multiprocessor sensor node on a real testbed. The testbed results show that multiprocessor sensor node can either increase sensor node's execution speed or reduce the sensor node's energy consumption when the sensor node executes computation-intensive tasks.
Zhenhe Pan, Patrick Schaumont, Yaling Yang
WCNC3
2014 Formal Verification of Software Countermeasures against Side-Channel Attacks
abstract
A common strategy for designing countermeasures against power-analysis-based side-channel attacks is using random masking techniques to remove the statistical dependency between sensitive data and side-channel emissions. However, this process is both labor intensive and error prone and, currently, there is a lack of automated tools to formally assess how secure a countermeasure really is. We propose the first SMT-solver-based method for formally verifying the security of a masking countermeasure against such attacks. In addition to checking whether the sensitive data are masked by random variables, we also check whether they are perfectly masked , that is, whether the intermediate computation results in the implementation of a cryptographic algorithm are independent of the secret key. We encode this verification problem using a series of quantifier-free first-order logic formulas, whose satisfiability can be decided by an off-the-shelf SMT solver. We have implemented the proposed method in a software verification tool based on the LLVM compiler frontend and the Yices SMT solver. Our experiments on a set of recently proposed masking countermeasures for cryptographic algorithms such as AES and MAC-Keccak show the method is both effective in detecting power side-channel leaks and scalable for practical use.
Hassan Eldib, Chao Wang 0001, Patrick Schaumont
ACM Trans. Softw. Eng. Methodol.3
2014 The Impact of Aging on a Physical Unclonable Function
abstract
On-chip physical unclonable functions (PUFs) have shown promises to solve several security problems. A PUF's behavior needs to be robust against reversible as well as irreversible temporal variabilities in circuits so that noise in the PUF output is minimized. While the effect of the reversible temporal variabilities on PUFs is well studied, sufficient attention has not been given so far to analyze the effect of the irreversible temporal variabilities i.e., aging on PUFs. In this paper, we perform an accelerated aging test on a ring oscillator (RO) PUF and analyze how it affects the functionality of the PUF. With our experiment using a set of 90-nm field-programmable gate arrays, we observe that aging makes PUF responses unreliable. Additionally, simulations show that the randomness of PUF responses remains unaffected despite aging. We also show that a passive countermeasure technique using a configurable RO can mitigate aging effect on the PUF significantly.
Abhranil Maiti, Patrick Schaumont
IEEE Trans. Very Large Scale Integr. Syst.2
2013 Study of ASIC technology impact factors on performance evaluation of SHA-3 candidates
abstract
The main aim of NIST's SHA-3 competition is to choose the new secure hash standard which will become successor of SHA-2. Analysing the strength of candidates in terms of hardware performance is complicated due to presence of different platforms such as FPGA, ASIC. Within them, implementation in an ASIC platform can be challenging due to different choices available during implementation. In this study, we will present these choices and also present an approach to evaluate the impact of these choices on the performance characteristics of SHA-3 finalists. This article will help NIST as well as cryptographers in better understanding the ASIC benchmarking results presented by different research groups.
Meeta Srivastav, Yongbo Zuo, Xu Guo 0001, Leyla Nazhandali, Patrick Schaumont
ACM Great Lakes Symposium on VLSI5
2013 Using Virtual Secure Circuit to Protect Embedded Software from Side-Channel Attacks
abstract
Side-Channel Attacks (SCAs) can break a cryptographic implementation within a very short time, and therefore, has become a practical threat to embedded security. This work presents Virtual Secure Circuit (VSC) as a software countermeasure to SCA. VSC provides protection to software by emulating WDDL, an SCA-resistant hardware circuit style. VSC is algorithm independent. This enables designers to protect different cryptographic software with only one solution. This work proposes the concept of VSC together with two implementation schemes. One scheme is based on a custom-instruction single-core processor architecture and the other on a dual-core architecture. Correspondingly, we built two prototypes on FPGA systems. Experiments with real-world side-channel power and electromagnetic attacks demonstrate that, compared with the unprotected software, VSC on single-core processor provides 20 times security improvement. The experiments also show that, although VSC on dual-core processor does not thwart electromagnetic attacks, it offers more than 25 times security improvement against power attacks. We conclude that VSC is comparable in security improvement to WDDL, but is more flexible and has much lower hardware cost.
Zhimin Chen 0002, Ambuj Sinha, Patrick Schaumont
IEEE Trans. Computers3
2012 ASIC implementations of five SHA-3 finalists
abstract
Throughout the NIST SHA-3 competition, in relative order of importance, NIST considered the security, cost, and algorithm and implementation characteristics of a candidate [1]. Within the limited one-year security evaluation period for the five SHA-3 finalists, the cost and performance evaluation may put more weight in the selection of winner. This work contributes to the SHA-3 hardware evaluation by providing timely cost and performance results on the first SHA-3 ASIC in 0.13 μm IBM process using standard cell CMOS technology with measurements of all the five finalists using the latest Round 3 tweaks. This article describes the SHA-3 ASIC design from VLSI architecture implementation to the silicon realization.
Xu Guo 0001, Meeta Srivastav, Sinan Huang, Dinesh Ganta, Michael B. Henry, Leyla Nazhandali, Patrick Schaumont
DATE7
2012 A novel microprocessor-intrinsic Physical Unclonable Function
abstract
We present a novel Physical Unclonable Function (PUF) exploiting the variability existing in a microprocessor pipeline to uniquely identify the microprocessor chip. The PUF accepts a microprocessor instruction as a challenge and produces the delay in a data path or a control path in the microprocessor as the response. The delay value is captured by over-clocking the microprocessor. The entire mechanism can be controlled by the microprocessor itself. Moreover, this PUF requires no dedicated hardware resources. It is a microprocessor-intrinsic PUF solution. We demonstrate our proposed idea using the SPARC instruction set implemented in a 32-bit LEON3 processor in a Spartan 3E FPGA. Our implementation based on the characterization of a subset of five SPARC instructions can produce 37 secure response bits for authentication.
Abhranil Maiti, Patrick Schaumont
FPL2
2012 Efficient and side-channel-secure block cipher implementation with custom instructions on FPGA
abstract
The security threat of side-channel analysis (SCA) attacks has created a need for SCA countermeasures. While many countermeasures have been proposed, a key challenge remains to design a countermeasure that is effective, that is easy to integrate in existing cryptographic implementations, and that has low overhead in area and performance. We present our solution in the context of an embedded design flow for FPGA. We integrate an SCA-resistant custom instruction set on a soft-core CPU. The SCA resistance is based on dual-rail precharge logic. A balanced-interleaved data format, combined with a novel memory organization, ensures that we can support both logic operations as well as lookup tables. The resulting countermeasure applies to a broad class of block ciphers. We demonstrate our results on an Altera Cyclone-II FPGA with Nios-II/s processor for a 128-bit Advanced Encryption Standard (AES) T-box implementation. We show SCA improvement of more than 400× for a system-wide electro-magnetic attack that covers both the FPGA and offchip memory (SSRAM). This comes at an overhead of 2.7× in performance and 1.15× in area. Using comparisons with related work, we demonstrate that this represents an excellent trade-off between SCA resistance, (software and hardware) design complexity, performance, and circuit area cost.
Suvarna Mane, Mostafa M. I. Taha, Patrick Schaumont
FPL3
2012 A novel profiled side-channel attack in presence of high Algorithmic Noise
abstract
Understanding the nature of hardware designs is a vital element in a successful Side-Channel Analysis. The inherent parallelism of these designs adds excessive Algorithmic Noise in the power consumption trace, which makes it difficult to mount a successful power attack against it. In this paper, we address this high Algorithmic Noise with a novel profiled attack that is generic and independent of any specific cryptographic algorithm. We propose both a new profiling phase and two new insights in the attack phase. The proposed profiling technique takes the high design parallelism into consideration, which results in a more accurate power model. In the attack phase, we first define two new targeted regions in the power trace, then aggregate the attack results from each of them to get a more powerful attack phase. The proposed attack model has been tested on the 128bit AES of the widely known DPA Contest (V2) and achieved a stable 80% Global Success Rate (GSR) at 2755 traces.
Mostafa M. I. Taha, Patrick Schaumont
ICCD2
2012 Simulating power/energy consumption of sensor nodes with flexible hardware in wireless networks
abstract
Energy consumption and real-time performance are two important metrics for wireless sensor networks (WSNs). To estimate these metrics, a number of simulation environments have been developed. However, these environments were made specifically for sensor nodes with fixed architectures. The recent generation of sensor nodes often has flexible architectures through the use of programmable hardware components, i.e., Field-programmable gate arrays (FPGAs). So far, no simulators have been developed to evaluate the performance of such flexible nodes in wireless networks. In this paper, we present PowerSUNSHINE, a power- and energy-estimation tool that fills the void. PowerSUNSHINE is the first scalable power/energy estimation tool for WSNs that provides an accurate prediction for both fixed and flexible sensor nodes. In this paper, we first describe requirements and challenges of building PowerSUNSHINE. Then, we present power/energy models for both fixed and flexible sensor nodes. Two testbeds, a MicaZ platform and a flexible node consisting of a microcontroller, a radio and a FPGA based co-processor, are provided to demonstrate the simulation fidelity of PowerSUNSHINE. We also discuss several evaluation results based on simulation and testbeds to show that PowerSUNSHINE is a scalable simulation tool that provides accurate estimation of power/energy consumption for both fixed and flexible sensor nodes.
Srikrishna Iyer, Patrick Schaumont, Yaling Yang
SECON3
2012 A Robust Physical Unclonable Function With Enhanced Challenge-Response Set
abstract
A Physical Unclonable Function (PUF) is a promising solution to many security issues due its ability to generate a die unique identifier that can resist cloning attempts as well as physical tampering. However, the efficiency of a PUF depends on its implementation cost, its reliability, its resiliency to attacks, and the amount of entropy in it. PUF entropy is used to construct crypto graphic keys, chip identifiers, or challenge-response pairs (CRPs) in a chip authentication mechanism. The amount of entropy in a PUF is limited by the circuit resources available to build a PUF. As a result, generating longer keys or larger sets of CRPs may increase PUF circuit cost. We address this limitation in a PUF by proposing an identity-mapping function that expands the set of CRPs of a ring-oscillator PUF (RO-PUF) with low area cost. The CRPs generated through this function exhibit strong PUF qualities in terms of uniqueness and reliability. To introduce the identity-mapping function, we formulate a novel PUF system model that uncouples PUF measurement from PUF identifier formation. We show the enhanced CRP generation capability of the new function using a statistical hypothesis test. An implementation of our technique on a low-cost FPGA platform shows at least 2 times savings in area compared to the traditional RO-PUF. The proposed technique is validated using a population of 125 chips, and its reliability over varying environmental conditions is shown.
Abhranil Maiti, Inyoung Kim, Patrick Schaumont
IEEE Trans. Inf. Forensics Secur.3
2011 Data-oriented performance analysis of SHA-3 candidates on FPGA accelerated computers
abstract
The SHA-3 competition organized by NIST has triggered significant efforts in performance evaluation of cryptographic hardware and software. These benchmarks are used to compare the implementation efficiency of competing hash candidates. However, such benchmarks test the algorithm in an ideal setting, and they ignore the effects of system integration. In this contribution, we analyze the performance of hash candidates on a high-end computing platform consisting of a multi-core Xeon processor with an FPGA-based hardware accelerator. We implement two hash candidates, Keccak and SIMD, in various configurations of multi-core hardware and multi-core software. Next, we vary application parameters such as message length, message multiplicity, and message source. We show that, depending on the application parameter set, the overall system performance is limited by three possible performance bottlenecks, including limitations in computation speed, in communication band-width, and in buffer storage. Our key result is to demonstrate the dependency of these bottlenecks on the application parameters. We conclude that, to make sound system design decisions, selecting the right hash candidate is only half of the solution: one must also understand the nature of the data stream which is hashed.
Zhimin Chen 0002, Xu Guo 0001, Ambuj Sinha, Patrick Schaumont
DATE4
2011 Pre-silicon Characterization of NIST SHA-3 Final Round Candidates
abstract
The NIST SHA-3 competition aims to select a new secure hash standard. Hardware implementation quality is an important factor in evaluating the SHA-3 finalists. However, a comprehensive methodology to benchmark five final round SHA-3 candidates in ASIC is challenging. Many factors need to be considered, including application scenarios, target technologies and optimization goals. This work describes detailed steps in the silicon implementation of a SHA-3 ASIC. The plan of ASIC prototyping with all the SHA-3 finalists, as an integral part of our SHA-3 ASIC evaluation project, is motivated by our previously proposed methodology, which defines a consistent and systematic approach to move a SHA-3 hardware benchmark process from FPGA prototyping to ASIC implementation. We have designed the remaining five SHA-3 candidates in 0.13 μm IBM process using standard-cell CMOS technology. In this paper, we discuss our proposed methodology for SHA-3 ASIC evaluation and report the latest results based on post-layout simulation of the five SHA-3 finalists with Round 3 tweaks.
Xu Guo 0001, Meeta Srivastav, Sinan Huang, Dinesh Ganta, Michael B. Henry, Leyla Nazhandali, Patrick Schaumont
DSD7
2011 The Impact of Aging on an FPGA-Based Physical Unclonable Function
abstract
On-chip Physical Unclonable Functions (PUFs) are emerging as a powerful security primitive that can potentially solve several security problems. A PUF needs to be robust against reversible as well as irreversible temporal changes in circuits. While the effect of the reversible temporal changes on PUFs is well studied, it is equally important to analyze the effect of the irreversible temporal changes i.e. aging on PUFs. In this work, we perform an accelerated aging testing on an FPGA-based ring oscillator PUF (RO-PUF) and analyze how it affects the functionality of the PUF. Based on our experiment using a group of 90-nm Xilinx FPGAs, we observe that aging makes PUF responses unreliable. On the other hand, the randomness of PUF responses remains unaffected despite aging.
Abhranil Maiti, Logan McDougall, Patrick Schaumont
FPL3
2011 A Simulator for Flexible Sensor Nodes in Wireless Networks
abstract
Most current sensor nodes are composed of a microcontroller and a radio. Their real-time and peak performance would be a bottleneck when executing compute-intensive tasks. Several works demonstrate that adding a hardware co-processor could accelerate the execution speed of the sensor nodes. So far, no simulators can simulate these new sensor nodes in wireless networks. An extension to SUNSHINE [1], a hardware-software emulator is developed to fill the void.
Srikrishna Iyer, Patrick Schaumont, Yaling Yang
MSN3
2011 A software-hardware emulator for sensor networks
abstract
Simulators are important tools for analyzing and evaluating different design options for wireless sensor networks (sensornets) and hence, have been intensively studied in the past decades. However, existing simulators only support evaluations of protocols and software aspects of sensornet design. They cannot accurately capture the significant impacts of various hardware designs on sensornet performance. As a result, the performance/energy benefits of customized hardware designs are difficult to be evaluated in sensornet research. To fill in this technical void, in this paper, we describe the design and implementation of SUNSHINE (Sensor Unified aNalyzer for Software and Hardware in Networked Environments), a scalable hardware-software emulator for sensornet applications. SUNSHINE is the first sensornet simulator that effectively supports joint evaluation and design of sensor hardware and software performance in a networked context. SUNSHINE captures the performance of network protocols, software and hardware up to cycle-level accuracy through its seamless integration of three existing sensornet simulators: a network simulator TOSSIM, an instruction-set simulator SimulAVR and a hardware simulator GEZEL. SUNSHINE solves several sensornet simulation challenges, including data exchanges and time synchronizations across different simulation domains and simulation accuracy levels. SUNSHINE also provides hardware specification scheme for simulating flexible and customized hardware designs. Several experiments are given to illustrate SUNSHINE's simulation capability. Evaluation results are provided to demonstrate that SUNSHINE is an efficient tool for software-hardware co-design in sensornet research.
Sachin Hirve, Srikrishna Iyer, Patrick Schaumont, Yaling Yang
SECON5
2011 SUNSHINE extension: a hardware-software emulator for flexible sensor nodes in wireless networks
abstract
Most current sensor nodes are composed of a microcontroller and a radio. Their real-time and peak performance would be a bottleneck when executing compute-intensive tasks. Several works demonstrate that adding a hardware co-processor could accelerate the execution speed of the sensor nodes. So far, no simulators can simulate these new sensor nodes in wireless networks. An extension to SUNSHINE [1], a hardware-software emulator is developed to fill the void.
Srikrishna Iyer, Patrick Schaumont, Yaling Yang
SenSys3
2011 Improved Ring Oscillator PUF: An FPGA-friendly Secure Primitive
Abhranil Maiti, Patrick Schaumont
J. Cryptol.2
2011 A Parallel Implementation of Montgomery Multiplication on Multicore Systems: Algorithm, Analysis, and Prototype
abstract
The Montgomery Multiplication is one of the cornerstones of public-key cryptography, with important applications in the RSA algorithm, in Elliptic-Curve Cryptography, and in the Digital Signature Standard. The efficient implementation of this long-word-length modular multiplication is crucial for the performance of public-key cryptography. Along with the strong momentum of shifting from single-core to multicore systems, we present a parallel-software implementation of the Montgomery multiplication for multicore systems. Our comprehensive analysis shows that the proposed scheme, pSHS, partitions the task in a balanced way so that each core has the same amount of job to do. In addition, we also comprehensively analyze the impact of intercore communication overhead on the performance of pSHS. The analysis reveals that pSHS is high performance, scalable over different number of cores, and stable when the communication latency changes. The analysis also tells us how to set different parameters to achieve the optimal performance. We implemented pSHS on a prototype multicore architecture configured in a Field Programmable Gate Array (FPGA). Compared with the sequential implementation, pSHS accelerates 2,048-bit Montgomery multiplication by 1.97, 3.68, and 6.13 times on, respectively, two-core, four-core, and eight-core architectures with communication latency equal to 100 clock cycles.
Zhimin Chen 0002, Patrick Schaumont
IEEE Trans. Computers2
2010 Implementing virtual secure circuit using a custom-instruction approach
abstract
Although cryptographic algorithms are designed to resist at least thousands of years of cryptoanalysis, implementing them with either software or hardware usually leaks additional information which may enable the attackers to break the cryptographic systems within days. A Side Channel Attack (SCA) is such a kind of attack that breaks a security system at a low cost within a short time. SCA uses side-channel leakage, such as the cryptographic implementations' execution time, power dissipation and magnetic radiation. This paper presents a countermeasure to protect software-based cryptography from SCA by emulating the behavior of the secure hardware circuits. The emulation is done by introducing two simple complementary instructions to the processor and applying a secure programming style. We call the resulting secure software program a Virtual Secure Circuit (VSC). VSC inherits the idea of a secure logic circuit, a hardware SCA countermeasure. It not only maintains the secure circuits' generality without limitation to a specific algorithm, but also increases its flexibility. Experiments on a prototype implementation demonstrated that the new countermeasure considerably increases the difficulty of the attacks by 20 times, which is in the same order as the improvement achieved by the dedicated secure hardware circuits. Therefore, we conclude that VSC is an efficient way to protect cryptographic software.
Zhimin Chen 0002, Ambuj Sinha, Patrick Schaumont
CASES3
2010 pSHS: A scalable parallel software implementation of Montgomery multiplication for multicore systems
abstract
Parallel programming techniques have become one of the great challenges in the transition from single-core to multicore architectures. In this paper, we investigate the parallelization of the Montgomery multiplication, a very common and time-consuming primitive in public-key cryptography. A scalable parallel programming scheme, called pSHS, is presented to map the Montgomery multiplication to a general multicore architecture. The pSHS scheme offers a considerable speedup. Based on 2-, 4-, and 8-core systems, the speedup of a parallelized 2048-bit Montgomery multiplication is 1.98, 3.74, and 6.53, respectively. pSHS delivers stable performance, high portability, high throughput and low latency over different multicore systems. These make pSHS a good candidate for public-key software implementations, including RSA, DSA, and ECC, based on general multicore platforms. We present a detailed analysis of pSHS, and verify it on dual-core, quad-core and eight-core prototypes.
Zhimin Chen 0002, Patrick Schaumont
DATE2
2010 Guest Editorial
abstract
The four papers in this special section are extended versions of papers presented at the 2009 ACM-IEEE International Conference on Formal Methods and Models for Codesign (MEMOCODE).
Roderick Bloem, Patrick Schaumont
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2010 A Flexible Design Flow for Software IP Binding in FPGA
abstract
Software intellectual property (SWIP) is a critical component of increasingly complex field programmable gate arrays (FPGA)-based system-on-chip (SOC) designs. As a result, developers want to ensure that their Software Intellectual Property (SWIP) is protected from being exposed to or tampered with by unauthorized parties. By restricting the execution of SWIP to a single trusted FPGA platform, SWIP binding addresses developers' concerns about maintaining control of their intellectual property and the market position it affords. This work proposes a novel design flow for SWIP binding on a commodity FPGA platform lacking specialized hardcore security facilities. We accomplish this by leveraging the qualities of a Physical Unclonable Function (PUF) and a tight integration of hardware and software security features. A prototype implementation demonstrates our design flow's ability to successfully protect software by encryption using a 128 bit FPGA-unique key extracted from a PUF. Based on this proof of concept, a solution to perform secure remote software updates, a common challenge in embedded systems, is proposed to showcase the practicality and flexibility of the design flow.
Michael A. Gora, Abhranil Maiti, Patrick Schaumont
IEEE Trans. Ind. Informatics3
2010 Optimized System-on-Chip Integration of a Programmable ECC Coprocessor
abstract
Most hardware/software (HW/SW) codesigns of Elliptic Curve Cryptography have focused on the computational aspect of the ECC hardware, and not on the system integration into a System-on-Chip (SoC) architecture. We study the impact of the communication link between CPU and coprocessor hardware for a typical ECC design, and demonstrate that the SoC may become performance-limited due to coprocessor data- and instruction-transfers. A dual strategy is proposed to remove the bottleneck: introduction of control hierarchy as well as local storage. The performance of the ECC coprocessor can be almost independent of the selection of bus protocols. Besides performance, the proposed ECC coprocessor is also optimized for scalability. Using design space exploration of a large number of system configurations of different architectures, our proposed ECC coprocessor architecture enables trade-offs between area, speed, and security.
Xu Guo 0001, Patrick Schaumont
ACM Trans. Reconfigurable Technol. Syst.2
2009 Programmable and Parallel ECC Coprocessor Architecture: Tradeoffs between Area, Speed and Security
Xu Guo 0001, Junfeng Fan, Patrick Schaumont, Ingrid Verbauwhede
CHES3
2009 Optimizing the HW/SW boundary of an ECC SoC design using control hierarchy and distributed storage
abstract
Hardware/Software codesign of Elliptic Curve Cryptography has been extensively studied in recent years. However, most of these designs have focused on the computational aspect of the ECC hardware, and not on the system integration into a SoC architecture. We study the impact of the communication link between CPU and coprocessor hardware for a typical ECC design, and demonstrate that the SoC may become performance-limited due to coprocessor data- and instruction-transfers. A dual strategy is proposed to remove the bottleneck: introduction of local control as well as local storage in the coprocessor. We quantify the impact of this strategy on a prototype implementation for Field Programmable Gate Arrays (FPGA) and measured an average speed-up in the resulting design of 9.4 times over the baseline ECC system, while the resulting system area increases by a factor of 1.6. The optimal area-time product improvement of our ECC coprocessor is 4.3 times compared to that of the baseline ECC coprocessor. Using design space exploration of a large number of system configurations using the latest FPGA technology and tools, we show that the optimal choice of ECC coprocessor parameters is strongly dependent on the efficiency of system-level communication.
Xu Guo 0001, Patrick Schaumont
DATE2
2009 Impact and compensation of correlated process variation on ring oscillator based puf
abstract
A Physical Unclonable Function in silicon is a die-unique challenge-response function that exploits circuit variations. A PUF has the potential to become an important security solution due to its ability to generate volatile secret keys. Before a PUF can be integrated into a system, its critical quality factors, including uniqueness, reliability and resiliency to different types of attacks, must be ensured. The uniqueness of a PUF is determined by the random inter-die process variations. However, manufacturing process variations have another component, called correlated intra-die variations. This component becomes significant in deep-submicron device technology below 90nm. In this paper, we show that the quality of a ring oscillator (RO) based PUF is affected by correlated intra-die process variation. We present experiments on 90nm FPGA devices and analyze the experimental data to quantify the effect and also propose a method to improve the quality of an RO-PUF by minimizing the effect of correlated intra-die variations.
Abhranil Maiti, Patrick Schaumont
FPGA2
2009 Improving the quality of a Physical Unclonable Function using configurable Ring Oscillators
abstract
A silicon physical unclonable function (PUF), which is a die-unique challenge-response function, is an emerging hardware primitive for secure applications. It exploits manufacturing process variations in a die to generate unique signatures out of a chip. This enables chip authentication and cryptographic key generation. A ring oscillator (RO) based PUF is a promising solution for FPGA platforms. However, the quality factors of this PUF, which include uniqueness, reliability and attack resiliency, are negatively affected by environmental noise and systematic variations in the die. This paper proposes two methods to address these negative effects, and to achieve a higher reliability in an RO-based PUF. Both methods are empirically verified on a population of five FPGAs over varying environmental conditions, and demonstrate how practically useful RO-based PUF can be achieved.
Abhranil Maiti, Patrick Schaumont
FPL2
2009 Physical unclonable function and true random number generator: a compact and scalable implementation
abstract
Physical Unclonable Functions (PUF) and True Random Number Generators (TRNG) are two very useful components in secure system design. PUFs can be used to extract chip-unique signatures and volatile secret keys, whereas TRNGs are used for generating random padding bits, initialization vectors and nonces in cryptographic protocols.
Abhranil Maiti, Raghunandan Nagesh, Anand Reddy, Patrick Schaumont
ACM Great Lakes Symposium on VLSI4
2009 Guest Editors' Introduction to Security in Reconfigurable Systems Design
abstract
This special issue on Security in Reconfigurable Systems Design reports on recent research results in the design and implementation of trustworthy reconfigurable systems. Five articles cover topics including power-efficient implementation of public-key cryptography, side-channel analysis of electromagnetic radiation, side-channel resistant design, design of robust unclonable functions on an FPGA, and Trojan detection in an FPGA bitstream.
Patrick Schaumont, Alex K. Jones, Steven Trimberger
ACM Trans. Reconfigurable Technol. Syst.1
2008 Turning liabilities into assets: Exploiting deep submicron CMOS technology to design secure embedded circuits
abstract
This paper explores an unexpected link between system-level security considerations and deep-submicron CMOS circuits. Many deep-submicron effects including increased leakage power, process variability, noise-level, power-density and integration density are thought of to be liabilities for integrated design. However, we show how they may instead be an asset for certain types of secure circuits. These circuits are useful for secure embedded systems design, where stringent cost-, power- and implementation constraints, as well as the increased risk towards physical attacks, are among the design issues. We also conclude that not all deep sub-micron liabilities are secure-circuit assets, and point out some of the open challenges in secure circuit design.
Patrick Schaumont, David Hwang 0001
ISCAS1
2008 MEMOCODE 2008 Co-Design Contest
abstract
The second MEMOCODE hardware/software codesign contest invites participants to solve a practical hardware/software codesign problem within the time span of one month. The larger objective for this contest is to be a showcase of advances in co-design tools and methodologies, in combination with design ingenuity and creativity. In the second installment of the contest, we received 9 submissions. In this short writeup, we review this year's design problem, and we consider relevant contest statistics.
Patrick Schaumont, Krste Asanovic, James C. Hoe
MEMOCODE1
2007 Masking and Dual-Rail Logic Don't Add Up
Patrick Schaumont, Kris Tiri
CHES1
2007 Design methods for security and trust
abstract
The design of ubiquitous and embedded computers focuses on cost factors such as area, power-consumption, and performance. Security and trust properties, on the other hand, are often an afterthought. Yet the purpose of ubiquitous electronics is to act and negotiate on their owner s behalf, and this makes trust a first-order concern. We outline a methodology for the design of secure and trusted electronic embedded systems, which builds on identifying the secure-sensitive part of a system (the root-of-trust) and iteratively partitioning and protecting that root-of-trust over all levels of design abstraction. This includes protocols, software, hardware, and circuits. We review active research in the area of secure design methodologies
Ingrid Verbauwhede, Patrick Schaumont
DATE2
2007 VT Matrix Multiply Design for MEMOCODE '07
abstract
This design presents a system optimized for complex matrix multiplications on the XUP Virtex-II board. Utilizing the GEZEL HW/SW co-simulation environment, the resulting system achieves ~25x speedup over a standard software only implementation. Further system level optimization (with DMA) results in the same coprocessor being speedup by at least another order of magnitude.
Eric Simpson, Pengyuan Yu, Patrick Schaumont, Sumit Ahuja, Sandeep K. Shukla
MEMOCODE3
2006 Cross Layer Design to Multi-thread a Data-Pipelining Application on a Multi-processor on Chip
abstract
Data-Pipelining is a widely used model to represent streaming applications. Incremental decomposition and optimization of a data-pipelining application onto a multi-processor platform spans multiple design layers, including the application layer, the system software layer, the architecture layer and the micro-architecture layer. For best results, designers have to consider multiple design layers (vertical exploration) and multiple architecture options (horizontal exploration). By using a data-pipelining JPEG encoder as the application driver, this paper presents a comprehensive analysis of mapping a data-pipelined application through multiple design layers, to a shared-memory SMP (Symmetric Multi- Processing) system. It is shown that a single-layered optimization ends up with a 110% worse design if the system effects from other layers are not taken into account. Compared to the nominal case, with appropriate mapping of the application, we achieve 47.5% improvement for high performance design and 77.6% energy reduction for energy efficient design under constant performance.
Bo-Cheng Lai, Patrick Schaumont, Ingrid Verbauwhede
ASAP2
2006 Offline Hardware/Software Authentication for Reconfigurable Platforms
Eric Simpson, Patrick Schaumont
CHES2
2006 Design with race-free hardware semantics
abstract
Most hardware description languages do not enforce determinacy, meaning that they may yield races. Race conditions pose a problem for the implementation, verification, and validation of hardware. Enforcing determinacy at the modeling level provides a solution to this problem. In this paper, we consider a common model of computation for hardware modeling - a network of cycle-true finite-state-machines with datapaths (FSMDs) - and we identify the conditions under which such models are guaranteed to be race-free. We base our analysis on the Kahn principle and a formal framework to represent FSMD semantics. We present our conclusions as four simple and easy to enforce modeling rules. A hardware designer that applies those four modeling rules, will thus obtain race-free hardware
Patrick Schaumont, Sandeep K. Shukla, Ingrid Verbauwhede
DATE1
2006 Executing Hardware as Parallel Software for Picoblaze Networks
abstract
Multi-processor architectures have gained interest recently because of their ability to exploit programmable silicon parallelism at acceptable power-efficiency figures. Despite the potential benefit they offer over single-processor architectures, it is unresolved how one can write compact and efficient programs for multiple parallel cores. In this paper, we propose the use of a synchronous hardware description language to program a network of small PicoBlaze processors. The partitioning of a multiprocessor program over multiple cores is straightforward because the input specification is fully parallel. A systematic transformation process converts the parallel input specification into concurrent PicoBlaze programs. We demonstrate the mapping of a cryptographic design (AES) onto four PicoBlaze processors, showing almost linear speedup over an equivalent single-core design
Pengyuan Yu, Patrick Schaumont
FPL2
2006 Multilevel Design Validation in a Secure Embedded System
abstract
In this paper, we present the simulation-based validation approach that we used during the design of ThumbPod-2, a portable fingerprint authentication system. The particular nature of secure system design has considerable impact on the simulation requirements and design flow. We present two key contributions. We will first show that rigorous design of secure digital systems requires a multilevel validation approach, meaning validation at multiple steps in the design flow. Indeed, an attacker chooses the easiest entry point and does not stick with one abstraction level. Second, we show the use of a cosimulation and codesign environment called GEZEL that can support this type of multilevel validation. We will illustrate this multilevel design validation strategy with the verification of security of the ThumbPod-2 device.
Patrick Schaumont, David Hwang 0001, Shenglin Yang, Ingrid Verbauwhede
IEEE Trans. Computers1
2006 An interactive codesign environment for domain-specific coprocessors
abstract
Energy-efficient embedded systems rely on domain-specific coprocessors for dedicated tasks such as baseband processing, video coding, or encryption. We present a language and design environment called GEZEL that can be used for the design, verification and implementation of such coprocessor-based systems.The GEZEL environment creates a platform simulator by combining a hardware simulation kernel with one or more instruction-set simulators. The hardware part of the platform is programmed in GEZEL, a deterministic, cycle-true and implementation-oriented hardware description language. GEZEL designs are scripted, allowing the hardware configuration of the platform simulator to be changed quickly without going through lengthy recompiles. For this reason, we call the environment interactive. We present the execution ladder as an optimization framework to balance interactivity against simulation speed.We demonstrate our approach using several designs including an AES encryption coprocessor and a Viterbi decoding coprocessor. We discuss the advantages of our approach as opposed to more conventional approaches using SystemC and Verilog/VHDL.
Patrick Schaumont, Doris Ching, Ingrid Verbauwhede
ACM Trans. Design Autom. Electr. Syst.1
2005 Prototype IC with WDDL and Differential Routing - DPA Resistance Assessment
Kris Tiri, David Hwang 0001, Alireza Hodjat, Bo-Cheng Lai, Shenglin Yang, Patrick Schaumont, Ingrid Verbauwhede
CHES6
2005 Cooperative multithreading on 3mbedded multiprocessor architectures enables energy-scalable design
abstract
We propose an embedded multiprocessor architecture and its associated thread-based programming model. Using a cycle-true simulation model of this architecture, we are able to estimate energy savings for a threaded C program. The savings are obtained by voltage- and frequency-scaling of the individual processors. We port a fingerprint minutiae detection application onto this architecture, and show the resulting performance on single-, dual-, and quad-processor configurations. The energy-scaled quadprocessor version results in a 77% energy reduction over the single-processor non-scaled implementation, at only a 2.2% degradation in cycle count.
Patrick Schaumont, Bo-Cheng Lai, Ingrid Verbauwhede
DAC1
2005 A side-channel leakage free coprocessor IC in 0.18µm CMOS for embedded AES-based cryptographic and biometric processing
abstract
Security ICs are vulnerable to side-channel attacks (SCAs) that find the secret key by monitoring the power consumption and other information that is leaked by the switching behavior of digital CMOS gates. This paper describes a side-channel attack resistant coprocessor IC and its design techniques. The IC has been fabricated in 0.18µm CMOS. The coprocessor, which is used for embedded cryptographic and biometric processing, consists of four components: an Advanced Encryption Standard (AES) based cryptographic engine, a fingerprint-matching oracle, a template storage, and an interface unit. Two functionally identical coprocessors have been fabricated on the same die. The first, 'secure', coprocessor is implemented using a logic style called Wave Dynamic Digital Logic (WDDL) and a layout technique called differential routing. The second, 'insecure', coprocessor is implemented using regular standard cells and regular routing techniques. Measurement-based experimental results show that a differential power analysis (DPA) attack on the insecure coprocessor requires only 8,000 acquisitions to disclose the entire 128b secret key. The same attack on the secure coprocessor still does not disclose the entire secret key at 1,500,000 acquisitions. This improvement in DPA resistance of at least 2 orders of magnitude makes the attack de facto infeasible. The required number of measurements is larger than the lifetime of the secret key in most practical systems.
Kris Tiri, David Hwang 0001, Alireza Hodjat, Bo-Cheng Lai, Shenglin Yang, Patrick Schaumont, Ingrid Verbauwhede
DAC6
2005 Fast Dynamic Memory Integration in Co-Simulation Frameworks for Multiprocessor System on-Chip
abstract
The paper proposes a technique to integrate and simulate a dynamic memory in a multiprocessor framework based on C/C++/SystemC. Using the host machine's memory management capabilities, dynamic data processing is supported without compromising speed and accuracy of the simulation. A first prototype in a shared memory context is presented.
Oreste Villa, Patrick Schaumont, Ingrid Verbauwhede, Matteo Monchiero, Gianluca Palermo
DATE2
2005 Energy and Performance Analysis of Mapping Parallel Multithreaded Tasks for An On-Chip Multi-Processor System
abstract
Multiprocessor systems offer superior performance and potentially better energy-reduction than single-processor systems. It all depends, however, on how well the application can be mapped onto the architecture. Indeed, a careful tradeoff of energy and performance requires a thorough understanding of the energy consumption pattern of the application across the architecture. We develop a simulation platform, MultiPo-Sim, which returns the cycle-accurate performance and energy consumption of a multiprocessor system, for both hardware components and software primitives. On the hardware level, energy scaling techniques can be modeled and each processing core can operate at different energy modes. MultiPo-Sim achieves 331K cycles per second simulation speed for a four-processor system on a 3GHz, 512MByte Fedora-2 PC. On the software level, data parallelizing and task parallelizing are two common models of multi-thread programming. By using MultiPo-Sim, we show that they show different energy and performance characteristics when mapping onto a multi-processor system.
Bo-Cheng Lai, Patrick Schaumont, Ingrid Verbauwhede
ICCD2
2005 Extended abstract: a race-free hardware modeling language
abstract
We describe race-free properties of a hardware description language called GEZEL. The language describes networks of cycle-true finite-state-machines with datapaths (FSMDs). We derive a set of four rules under which a network of such FSMDs satisfies the Kahn principle. When applying those rules, GEZEL programs will be determinate and a designer will thus obtain race-free hardware. We define extended FSMD networks as FSMD networks for which some components are user-defined and not specified as FSMDs. An important result is that the determinate properties of the FSMD network are also valid for the extended FSMD network provided that the user-defined components are determinate. Most hardware description languages do not have this determinacy. Their simulation semantics are dependent on simulator implementation, and on a run-time race resolution mechanism. We therefore position GEZEL as a model of computation that RTL designers should have in mind while creating RTL models. In fact, we can generate SystemC and other HDL code from GEZEL models, thereby guaranteeing the determinacy in the generated HDL code.
Patrick Schaumont, Sandeep K. Shukla, Ingrid Verbauwhede
MEMOCODE1
2005 Platform-based design for an embedded-fingerprint-authentication device
abstract
Fingerprint authentication, in an embedded and portable context, requires complex signal, network, and security-protocol processing in a resource-constrained implementation. We present a platform-based design approach for this application, based on a hierarchy of virtual machines (VM). The fingerprint authentication is programmed in Java, C, and VHSIC hardware description language, and mapped onto a hierarchy of three machines, consisting of an embedded Java VM, an Sparc-V8 core, and an field programmable gate array. We show how our approach is able to cope with multiple concurrent design processes and multiple application domains, including biometrics signal processing, as well as security-protocol implementation. The platform-based design approach also deals with reuse requirements for embedded software and hardware. The formulation of a platform as a VM enables design exploration and incremental design validation throughout the design traject, and results in a specialized, but still programmable, platform. The Java bytecode of our fingerprint authentication takes less than 10 kB.
Patrick Schaumont, David Hwang 0001, Ingrid Verbauwhede
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2005 Skiing the embedded systems mountain
abstract
UCLA teaches students how to master the steep slopes of the embedded systems mountain. The EE201A graduate course connects high-level design specification to embedded implementation. There is a long-standing and wide culture gap between system designers that create those abstract specifications and the system architects that need to implement them. In industry, the culture gap has separated software from hardware teams, and platform creators from platform users. In an embedded context, where these are very tightly connected, this leads to large inefficiencies both in design time and design results. Our course takes students to both sides of the gap and lets them look at this problem from different perspectives. For a given application, it teaches how to select target architectures, tools, and design methods. The course covers a stepwise systematic design process. It includes specification, transformation, and refinement of an application. Specifications enable systematic and structured expression of an application. Transformations rework specifications into ones that are a better match for a given target architecture. Refinements lower the abstraction level toward the target architecture. The embedded systems mountain is traversed in two directions. A vertical refinement axis covers elements such as power-memory-reduction methods or fixed-point refinement. A horizontal exploration axis covers various architecture alternatives including application-specific integrated circuits (ASIC), domain-specific processors, digital signal processors (DSP), embedded cores, programmable processors, and system-on-chip (SOC). During the course, the students also go through an extensive design project to apply the methods learned in this course. A typical embedded application is used to drive the project. In this paper it is illustrated using an embedded version of an image encoder, more specifically a JPEG encoder. Several commercial tools, design environments, and platforms have been used as alternative implementation targets for this application.
Ingrid Verbauwhede, Patrick Schaumont
ACM Trans. Embed. Comput. Syst.2
2004 Java cryptography on KVM and its performance and security optimization using HW/SW co-design techniques
abstract
This paper describes a design approach to include and optimize Java based cryptographic applications into resource limited embedded devices. For easy prototyping and to be platform independent, the security applications are first developed in Java. Two Java cryptographic libraries, the Bouncy Castle API and the IAIK API are ported to a real embedded device for cost and performance evaluation. It requires 0.88Mbytes to 1.2Mbytes in the KVM footprint size and a few milliseconds to run secret key algorithms and message digests on a typical embedded device. In a second step, the performance critical components of the security applications are moved to hardware acceleration units. The GEZEL design environment is used for the hardware modeling and the co-simulation between software on KVM and the hardware co-processor. Moving the AES algorithm from the SH3-DSP microprocessor to a hardware co-processor shows a performance gain of 10.4x including the overhead in Java, C, and hardware interfaces. Then in a third step, the security critical components are realized by means of a special dynamic differential logic (DDL) style, which makes the secure modules resistant against side channel attacks. All key related actions and cryptographic algorithms are restricted to the secure co-processor. The overall performance gain is 25x compared to a pure Java implementation. Copyright 2004 ACM.
Yusuke Matsuoka, Patrick Schaumont, Kris Tiri, Ingrid Verbauwhede
CASES2
2004 Interactive Cosimulation with Partial Evaluation
abstract
We present a technique to improve the efficiency of hardware-software cosimulation, using design information known at simulator compile-time. The generic term for such optimization is partial evaluation. Our contribution is that we apply the optimization transparently to the user, and at multiple abstraction levels in the simulation. We use the technique to create an interactive codesign environment, and evaluate it on several designs including an AES encryption coprocessor and a Viterbi decoder, and for several instruction-set simulators. Compared to SystemC-based cosimulation, we achieve comparable cosimulation performance at only a fraction of the model-build time.
Patrick Schaumont, Ingrid Verbauwhede
DATE1
2004 Architectures and Design Techniques for Energy Efficient Embedded DSP and Multimedia Processing
abstract
Energy efficient embedded systems consist of a heterogeneous collection of very specific building blocks, connected together by a complex network of many dedicated busses and interconnect options. The trend to merge multiple functions into one device makes the design and integration of these "systems-on-chip" (SOC's) even more problematic. Yet, specifications and applications are never fixed and require the embedded units to be programmable. The topic of this paper is to give the designer architectures and design techniques to find the right balance between energy efficiency and flexibility. The key is to include programmability (or reconfiguration) at the right level of abstraction and tuned to the application domain. The challenge is to provide an exploration and programming environment for this heterogeneous architecture platform.
Ingrid Verbauwhede, Patrick Schaumont, Christian Piguet, Bart Kienhuis
DATE2
2004 Integrated Modeling and Generation of a Reconfigurable Network-on-Chip
abstract
Summary form only given. While a communication network is a critical component for an efficient system-on-chip multiprocessor, there are few approaches available to help with system-level architectural exploration of such a specialized interconnection network. We present an integrated modeling, simulation and implementation tool. A high level description of a network-on-chip can be simulated and converted into VHDL. The system simulation supports multiple instruction-set simulators, and obtains cycle-accurate performance metrics. This way, an optimal network configuration can be determined easily. We discuss our approach by designing a flexible network-on-chip and present implementation results after mapping into FPGA. The performance of our automatically generated network is comparable with a reference design directly developed in HDL.
Doris Ching, Patrick Schaumont, Ingrid Verbauwhede
IPDPS2
2004 Embedded Software Integration for Coarse-Grain Reconfigurable Systems
abstract
Summary form only given. Coarse-grain reconfigurable systems offer high performance and energy-efficiency, provided an efficient run-time reconfiguration mechanism is available. Using an embedded software vantage point, we define three levels of reconfigurability for such systems, each with a different degree of coupling between embedded software and reconfigurable hardware. We classify reconfigurable systems starting with tightly-coupled coprocessors and evolving to processor networks. This results in a gradual increase of energy-efficiency when compared to software-only systems, at the cost of increasing programming complexity. Using several sample applications including signal-, crypto-, and network-processing acceleration units, we demonstrate energy-efficiency improvements of 12 times over software for tightly-coupled systems up to 84 times for network-on-chip systems.
Patrick Schaumont, Kazuo Sakiyama, Alireza Hodjat, Ingrid Verbauwhede
IPDPS1
2003 Finding the best system design flow for a high-speed JPEG encoder
abstract
26 students at the University of California, Los Angeles (UCLA) studied system level design methodologies through the design of a high-speed JPEG encoder. The results produced by 5 different design flows onto various target platforms demonstrate the high impact of tools on design quality.
Kazuo Sakiyama, Patrick Schaumont, Ingrid Verbauwhede
ASP-DAC2
2003 Design flow for HW / SW acceleration transparency in the thumbpod secure embedded system
abstract
This paper describes a case study and design flow of a secure embedded system called ThumbPod, which uses cryptographic and biometric signal processing acceleration. It presents the concept of HW/SW acceleration transparency, a systematic method to accelerate Java functions in both software and hardware. An example of acceleration transparency for a Rijndael encryption function is presented. The embedded prototype hardware platform is also described. Acceleration transparency yields software and hardware performance gains of 333X.
David Hwang 0001, Bo-Cheng Lai, Patrick Schaumont, Kazuo Sakiyama, Shenglin Yang, Alireza Hodjat, Ingrid Verbauwhede
DAC3
2002 A Security Protocol for Biometric Smart Cards
David Hwang 0001, Bo-Cheng Lai, Patrick Schaumont, Ingrid Verbauwhede
CARDIS3
2002 Unlocking the design secrets of a 2.29 Gb/s Rijndael processor
abstract
This contribution describes the design and performance testing of an Advanced Encryption Standard (AES) compliant encryption chip that delivers 2.29 GB/s of encryption throughput at 56 mw of power consumption. We discuss how the high level reference specification in C is translated into a parallel architecture. Design decisions are motivated from a system level viewpoint. The prototyping setup is discussed.
Patrick Schaumont, Henry Kuo, Ingrid Verbauwhede
DAC1
2002 Techniques to Evolve a C++ Based System Design Language
abstract
Complex systems-on-chip present one of the most challenging design problems. To meet this challenge, new design languages capable of modelling such heterogeneous, dynamic systems are needed. For implementation of such a language, the use of an object oriented C++ class library has proven to be a promising approach, since new classes dealing with design- and platform-specific problems can be added in a conceptual and seamlessly reusable way. This paper shows the development of such an extension aimed to provide a platform-independent high-level structured storage object through hiding of the low-level implementation details. It results in a completely virtualised, user-extendible component, suitable for use in heterogeneous systems.
Robert Pasko, Serge Vernalde, Patrick Schaumont
DATE3
2002 Building a Virtual Framework for Networked Reconfigurable Hardware and Software Objects
Yajun Ha, Serge Vernalde, Patrick Schaumont, Marc Engels, Rudy Lauwereins, Hugo De Man
J. Supercomput.3
2001 Virtual Java/FPGA interface for networked reconfiguration
abstract
A virtual interface between Java and FPGA for networked reconfiguration is presented. Through the Java/FPGA interface, Java applications can exploit hardware accelerators with FPGAs for both functional flexibility and performance acceleration. At the same time, the interface is platform independent. It enables the networked application developers to design their applications with only one interface in mind when considering the interfacing issues. The virtual interface is part of our work to build a platform-independent deployment framework for the networked services. In the framework, both the software and hardware components of services can be platform independently described and deployed.
Yajun Ha, Geert Vanmeerbeeck, Patrick Schaumont, Serge Vernalde, Marc Engels, Rudy Lauwereins, Hugo De Man
ASP-DAC3
2001 Panel: The Next HDL: If C++ is the Answer, What was the Question?
abstract
The focus of this panel is on issues surrounding the use of C++ in modeling, integration of silicon IP and system-on-chip designs. In the last two years there have been several announcements promoting C++ based solutions and of multiple consortia (SystemC, Cynapps, Accellera, SpecC) that represent increasing commercial interest both from tool vendors as well as perhaps expression of genuine needs from the design houses. There are, however, serious questions about what value proposition does a C++ based design methodology bring to the IC or system designer? What has changed in the modeling technology (and/or available tools) that gives a new capability? Is synthesis the right target? or VAlidation? Tester modeling or testbench generation? This panel brings together advocates and opponents from the user community to highlight the achievements and the challenges that remain in use C++ for use in microelectronic circuits and systems.
Rajesh K. Gupta 0001, Shishpal Rawat, Ingrid Verbauwhede, Gérard Berry, Ramesh Chandra, Daniel Gajski, Kris Konigsfeld, Patrick Schaumont
DAC8
2001 A Quick Safari Through the Reconfiguration Jungle
abstract
Cost effective systems use specialization to optimize factors such as power consumption, processing throughput, flexibility or combinations thereof. Reconfigurable systems obtain this specialization at rim-time. System reconfiguration has a vertical, a horizontal and a time dimension. We organize this design space as the reconfiguration hierarchy, and discuss the design methods that deal with it. Finally, we survey existing commercial platforms that support reconfiguration and situate them in the reconfiguration jungle.
Patrick Schaumont, Ingrid Verbauwhede, Kurt Keutzer, Majid Sarrafzadeh
DAC1
2001 A SW/HW Interface API for Java/FPGA Co-Designed Applets
Yajun Ha, Patrick Schaumont, Serge Vernalde, Marc Engels, Rudy Lauwereins, Hugo De Man
FCCM2
2001 Development of a Design Framework for Platform-Independent Networked Reconfiguration of Software and Hardware
Yajun Ha, Bingfeng Mei, Patrick Schaumont, Serge Vernalde, Rudy Lauwereins, Hugo De Man
FPL3
2000 Standards for System-Level Design: Practical Reality or Solution in Search of a Question?
abstract
We address the issue of standards development for the system-level design space. System-level design IP re-use standards are key to the future of the VSIA. However, the concept of system-level standards has its share of sceptics: what role can standards play in this developing market segment? In response we present an overview of three standards in the system-level VC integration space, and describe two distinct industrial case studies to support their practicality.
Christopher K. Lennard, Patrick Schaumont, Gjalt G. de Jong, Anssi Haverinen, Pete Hardee
DATE2
1999 A 10 Mbit/s Upstream Cable Modem with Automatic equalization
abstract
A fully digital QAM16 burst receiver ASIC is presented.The B04 receiver demodulates at 10 Mbit/s and uses an advanced signal processing architecture that performs perburst automatic equalization.It is a critical building block in a broadband access system for HFC networks.The chip was designed using a C++ based flow and is implemented as a 80 Kgate 0.7~ CMOS standard cell design.
Patrick Schaumont, Radim Cmar, Serge Vernalde, Marc Engels
DAC1
1999 Hardware Reuse at the Behavioral Level
abstract
Standard interfaces for hardware reuse are currently defined at the structural level.In contrast to this, our contribution defines the reuse interface at the behavioral registertransfer (RT) level.This promotes direct reuse of functionality and avoids the integration problems of structural reuse.We present an object oriented reuse interface in C++ and show the use of it within two real-life designs.
Patrick Schaumont, Radim Cmar, Serge Vernalde, Marc Engels, Ivo Bolsens
DAC1
1999 A Methodology and Design Environment for DSP ASIC Fixed-Point Refinement
abstract
Complex signal processing algorithms are specified in floating point precision. When their hardware implementation requires fixed point precision, type refinement is needed. The paper presents a methodology and design environment for this quantization process. The method uses independent strategies for fixing MSB and LSB weights of fixed point signals. It enables short design cycles by combining the strengths of both analytical and simulation based methods.
Radim Cmar, Luc Rijnders, Patrick Schaumont, Serge Vernalde, Ivo Bolsens
DATE3
1999 A new algorithm for elimination of common subexpressions
abstract
The problem of an efficient hardware implementation of multiplications with one or more constants is encountered in many different digital signal-processing areas, such as image processing or digital filter optimization. In a more general form, this is a problem of common subexpression elimination, and as such it also occurs in compiler optimization and many high-level synthesis tasks. An efficient solution of this problem can yield significant improvements in important design parameters like implementation area or power consumption. In this paper, a new solution of the multiple constant multiplication problem based on the common subexpression elimination technique is presented. The performance of our method is demonstrated primarily on a finite-duration impulse response filter design. The idea is to implement a set of constant multiplications as a set of add-shift operations and to optimize these with respect to the common subexpressions afterwards. We show that the number of add/subtract operations can be reduced significantly this way. The applicability of the presented algorithm to the different high-level synthesis tasks is also indicated. Benchmarks demonstrating the algorithm's efficiency are included as well.
Robert Pasko, Patrick Schaumont, Veerle Derudder, Serge Vernalde, Daniela Duracková
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
1998 A Programming Environment for the Design of Complex High Speed ASICs
abstract
A C++ based programming environment for the design of complex high speed ASICs is presented. The design of a 75 Kgate DECT transceiv er is used as a driv er example. Compact descriptions, combined with efficient sim ulationand syn thesis strategies are essen tial for the design of such a complex system. It is sho wn how a C++ programming approach outperforms traditional HDL-based methods.
Patrick Schaumont, Serge Vernalde, Luc Rijnders, Marc Engels, Ivo Bolsens
DAC1
1997 Synthesis of pipelined DSP accelerators with dynamic scheduling
abstract
To construct complete systems on silicon, application specific DSP accelerators are needed to speed up the execution of high throughput DSP algorithms. In this paper, a methodology is presented to synthesize high throughput DSP functions into accelerator processors containing a datapath of highly pipelined, bit-parallel hardware units. Emphasis is put on the definition of a controller architecture that allows efficient run-time schedules of these DSP algorithms on such highly pipelined data paths. The methodology is illustrated by means of an image encoding filter bank.
Patrick Schaumont, Bart Vanthournout, Ivo Bolsens, Hugo De Man
IEEE Trans. Very Large Scale Integr. Syst.1