EDBT 2026 Demo / reviewers in the wild / expert
Rainer Buchty
dblp:51/6463
· DBLP profile ↗
22ranked-venue papers
1as first author
11since 2021 · last 2026
0009-0004-9413-2078ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 18 · 1 first-author · 10 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | IMS: Intelligent Hardware Monitoring System for Secure SoCsabstractIn the modern Systems-on-Chip (SoC), the Advanced eXtensible Interface (AXI) protocol exhibits security vulnerabilities, enabling partial or complete denial-of-service (DoS) through protocol-violation attacks. The recent counter-measures lack a dedicated real-time protocol semantic analysis and evade protocol compliance checks. This paper tackles this AXI vulnerability issue and presents an intelligent hardware monitoring system (IMS) for real-time detection of AXI protocol violations. IMS is a hardware module leveraging neural networks to achieve high detection accuracy. For model training, we perform DoS attacks through header-field manipulation and systematic malicious operations, while recording AXI transactions to build a training dataset. We then deploy a quantization-optimized neural network, achieving 98.7% detection accuracy with2.5 million inferences/s. We subsequently integrate this IMS into a RISC-V SoC as a memory-mapped IP core to monitor its AXI bus. For demonstration and initial assessment for later ASIC integration, we implemented this IMS on an AMD Zynq UltraScale+ MPSoC ZCU104 board, showing an overall small hardware footprint (9.04% look-up-tables (LUTs), 0.23% DSP slices, and 0.70% flip-flops) and negligible impact on the overall design’s achievable frequency. This demonstrates the feasibility of lightweight, security monitoring for resource-constrained edge environments. Wadid Foudhaili, Aykut Rencber, Anouar Nechi, Rainer Buchty, Mladen Berekovic, Saleh Mulhem |
DATE | 4 |
| 2026 | Fast Reinforcement Learning for Robust Beam Codebooks in Future Communication SystemsabstractMillimeter wave (mmWave) and terahertz (THz) MIMO systems typically rely on predefined beamforming codebooks for initial access and data transmission. However, these codebooks are often unoptimized for specific conditions, leading to large sizes and significant beam training overhead, thereby complicating support for highly mobile applications. This paper introduces a reinforcement learning framework that optimizes beam patterns using only receive power measurements, adapting to the environment, user distribution, and hardware constraints without prior channel knowledge. The framework explores three reinforcement learning algorithms: Deep Deterministic Policy Gradient (DDPG), Twin Delayed Deep Deterministic Policy Gradient (TD3), and Soft Actor-Critic (SAC). While reinforcement learning has shown promise for beamforming, a comprehensive comparative analysis of advanced RL algorithms under a combination of realistic challenges, such as Non-Line-of-Sight (NLoS) conditions and hardware impairments for adaptive beam codebook design in mmWave/THz systems has been largely unexplored. This paper presents the first such in-depth comparative study. Simulation results demonstrate the superiority of the SAC algorithm, achieving higher beamforming gain and faster convergence compared to DDPG and TD3 in various scenarios, including LoS and NLoS conditions, even with hardware impairments. Anouar Nechi, Zakaria Narjis, Rainer Buchty, Mladen Berekovic, Saleh Mulhem |
IEEE Trans. Commun. | 3 |
| 2025 | Nail: Not Another Fault-Injection Framework for Chisel-generated RTLabstractFault simulation and emulation are essential techniques for evaluating the dependability of integrated circuits, enabling early-stage vulnerability analysis and supporting the implementation of effective mitigation strategies. High-level hardware description languages such as Chisel facilitate the rapid development of complex fault scenarios with minimal modification to the design. However, existing Chisel-based fault injection (FI) frameworks are limited by their coarse-grained, instruction-level controllability, which restricts the precision of fault modeling. This work introduces Nail, a Chisel-based open-source FI framework that overcomes these limitations by introducing statebased faults. This approach allows fault scenarios based on specific system states instead of just instruction-level triggers, removing the need for precise timing of fault activation. For greater controllability, Nail allows users to arbitrarily modify internal trigger states via software at runtime. To support this, Nail automatically generates a software interface, offering straightforward access to the instrumented design. This enables fine-tuning of fault parameters during active fault-injection campaigns, a feature particularly beneficial for FPGA emulation, where synthesis is time-consuming. Utilizing these features, Nail narrows the gap between the high speed of emulation-based FI frameworks, the usability of software-based approaches, and the controllability achieved in simulation. We demonstrate Nail’s state-based fault injection and software framework by modeling a faulty general-purpose register in a RISC-V processor. Although this might appear straightforward, it requires statedependent fault injection and was previously impossible without fundamental changes to the design. The approach was validated in both simulation and FPGA emulation, where the addition of Nail introduced less than $1 \%$ resource overhead. Robin Sehm, Christian Ewert, Rainer Buchty, Mladen Berekovic, Saleh Mulhem |
DSD | 3 |
| 2025 | Cross Technology Prediction for SRAM Stability AnalysisabstractSRAM stability presents a significant challenge in technology scaling due to process variations. This paper introduces CrossNodeML, a machine learning-based approach that leverages device and bitcell simulations to predict SRAM behavior across new technology nodes using data from previous nodes. Our model is evaluated with three critical SRAM metrics: Access Disturb Margin (ADM), Write Margin (WRM), and Ireadmin. Comparisons with high-sigma verifier tool simulations demonstrate the accuracy of our predictions, significantly reducing the need for full statistical simulations in new nodes. This methodology revolutionizes design decisions and significantly enhances early stability assessments. Jihene Bouhlila, Rainer Buchty, Mladen Berekovic |
ISCAS | 2 |
| 2025 | A Lightweight Peripheral Design for RRAM-based LUTsabstractResistive random-access memory (RRAM) has garnered increasing interest due to its compact structure size, low energy requirements, and non-volatility. Recently, RRAM crossbars have shown potential for implementing lookup tables (LUTs) for reconfigurable logic. However, due to its analog properties, RRAM suffers from extensive peripherals, such as analog-to-digital converters (ADCs). In this paper, we tackle this problem by combining signal amplification, logic disjunction, and conversion to the digital domain in a single compact gate. As a result, the proposed RRAM-based LUT (R-LUT) design reduces the area per LUT by a factor of four and the peripheral overhead by more than six times. It is also almost three times smaller than an SRAM-based LUT implemented in the same technology. At the same time, it achieves a competitive access frequency of more than 3 GHz while consuming less than 240 fJ per inference, improving all major aspects of state-of-the-art R-LUT designs. Philipp Grothe, Christoph Hübner, Rainer Buchty, Mladen Berekovic, Saleh Mulhem |
ISCAS | 3 |
| 2025 | Lightweight Authenticated Integration and In-Field Secure Operation of System-in-PackageabstractSystem in Package (SiP) relies on integrating different chiplets potentially involving many third-party devices and chiplet foundries. This type of advanced packaging technology opens up numerous threat scenarios, especially: (a) the inauthentic and untraceable integration of chiplets into a SiP, (b) the insecure integration of malicious chiplets, which leads to a severe impact on the SiP security in the field. The current solutions require many hardware cryptographic primitives, making them costly and power-hungry. Therefore, a new lightweight solution is needed to ensure secure chiplet integration and secure SiP operation. In this article, we deal with these problems and introduce iTrustlet , as a combination of a physical unclonable function and an authenticated encryption scheme to ensure an authenticated and traceable chiplet integration. We propose a chiplet integration protocol based on iTrustlet and a classical root-of-trust (RoT) to ensure the integrated chiplets are unaltered and unreplaced. To guarantee SiP in-field security, iTrustlet with a hardware firewall (HWF) is proposed. Their interaction leads to two security features: (i) HWF provides a SiP protection mechanism, and (ii) iTrustlet secures the update of HWF rules. In particular, we provide a multilevel solution centralized around iTrustlet , focusing on lightweightness. The implementation results show that area and power overheads are 1.24% and 1.84% in the case of FPGA and 0.49% and 1.2% for ASIC implementation. Christian Ewert, Andrija Neskovic, Carsten Heinz, Felix Muuss, Alexander Treff, Marc Gourjon, Rainer Buchty, Thomas Eisenbarth 0001, Andreas Koch 0001, Mladen Berekovic, Saleh Mulhem |
ACM Trans. Design Autom. Electr. Syst. | 7 |
| 2025 | A Systematic Mapping Study on SystemC/TLM Modeling Capabilities in New Research DomainsabstractWith increasingly complex circuits and systems, the need for advanced design methodologies is growing. These methodologies shift the designers’ focus from technology-specific implementations to more abstract electronic system design (ESL). SystemC was developed to address this need. Being an open standard based on C++, SystemC facilitates hardware and software modeling across multiple levels of abstraction, with a particular emphasis on ESL. It is further enhanced by including the transaction-level modeling (TLM) layer, strengthening its capability to model communication between components, and even full-system simulators. Traditionally, SystemC/TLM has been deployed to provide hardware prototypes for software development early in the design process. However, surveys and literature reviews showing other capabilities of SystemC/TLM are scarce. Hence, it is essential to explore SystemC/TLM’s new capabilities in different domains such as in-circuit fault propagation, security assessment, and verification. In this article, we conduct a systematic mapping study (SMS) of SystemC/TLM modeling capabilities in certain research domains. We elaborate on the state-of-the-art ESL with an emphasis on SystemC/TLM-based system modeling. Subsequently, we present how such technologies can be applied to the new research domains within the field of circuit and system modeling, namely: (D.1) architecture exploration, (D.2) power estimation, (D.3) fault-injection analysis, (D.4) functional and security verification, and (D.5) side-channel analysis. This SMS highlights the advantages and disadvantages of the investigated SystemC/TLM capabilities and addresses the open challenges in these domains, concluding that SystemC/TLM offers significant potential in performance evaluation, verification, and security assessment of circuits and systems at ESL. Ahmed Mahmoudi, Andrija Neskovic, Celine Thermann, Robin Sehm, Christoph Hübner, Tavia Plattenteich, Rolf Meyer, Rainer Buchty, Mladen Berekovic, Saleh Mulhem |
ACM Trans. Design Autom. Electr. Syst. | 8 |
| 2024 | Machine Learning for SRAM Stability AnalysisabstractSRAM stability is a critical challenge in technology scaling due to process variations. In this paper, we introduce a cutting-edge approach leveraging machine learning based on device and bitcell simulation to predict SRAM behavior in high sigma local and global variations. Our focus includes both high-density (HDC) and Low Voltage Cell (LVC) analysis, revealing the Extreme Gradient Boosting Regressor (XGBR) as the top performer for both. This research demonstrates the superior accuracy of the XGBR regressor in predicting key SRAM metrics, such as Access Disturb Margin (ADM), Write Margin (WRM), and Ireadmin, offering a compelling alternative to traditional statistical simulations. The purpose of such prediction is to revolutionize the design process and speed up designers’ decisionmaking. Jihene Bouhlila, Felix Last, Rainer Buchty, Mladen Berekovic, Saleh Mulhem |
ISCAS | 3 |
| 2023 | SystemC Model of Power Side-Channel Attacks Against AI Accelerators: Superstition or not?abstractAs training artificial intelligence (AI) models is a lengthy and hence costly process, leakage of such a model's internal parameters is highly undesirable. In the case of AI accelerators, side-channel information leakage opens up the threat scenario of extracting the internal secrets of pre-trained models. Therefore, sufficiently elaborate methods for design verification as well as fault and security evaluation at the electronic system level are in demand. In this paper, we propose estimating information leakage from the early design steps of AI accelerators to aid in a more robust architectural design. We first introduce the threat scenario before diving into SystemC as a standard method for early design evaluation and how this can be applied to threat modeling. We present two successful side-channel attack methods executed via SystemC-based power modeling: correlation power analysis and template attack, both leading to total information leakage. The presented models are verified against an industry-standard netlist-level power estimation to prove general feasibility and determine accuracy. Consequently, we explore the impact of additive noise in our simulation to establish indicators for early threat evaluation. The presented approach is again validated via a model-vs-netlist comparison, showing high accuracy of the achieved results. This work hence is a solid step towards fast attack deployment and, subsequently, the design of attack-resilient AI accelerators. Andrija Neskovic, Saleh Mulhem, Alexander Treff, Rainer Buchty, Thomas Eisenbarth 0001, Mladen Berekovic |
ICCAD | 4 |
| 2023 | FPGA-based Deep Learning Inference Accelerators: Where Are We Standing?abstractRecently, artificial intelligence applications have become part of almost all emerging technologies around us. Neural networks, in particular, have shown significant advantages and have been widely adopted over other approaches in machine learning. In this context, high processing power is deemed a fundamental challenge and a persistent requirement. Recent solutions facing such a challenge deploy hardware platforms to provide high computing performance for neural networks and deep learning algorithms. This direction is also rapidly taking over the market. Here, FPGAs occupy the middle ground regarding flexibility, reconfigurability, and efficiency compared to general-purpose CPUs, GPUs, on one side, and manufactured ASICs on the other. FPGA-based accelerators exploit the features of FPGAs to increase the computing performance for specific algorithms and algorithm features. Filling a gap, we provide holistic benchmarking criteria and optimization techniques that work across several classes of deep learning implementations. This article summarizes the current state of deep learning hardware acceleration: More than 120 FPGA-based neural network accelerator designs are presented and evaluated based on a matrix of performance and acceleration criteria, and corresponding optimization techniques are presented and discussed. In addition, the evaluation criteria and optimization techniques are demonstrated by benchmarking ResNet-2 and LSTM-based accelerators. Anouar Nechi, Lukas Groth, Saleh Mulhem, Farhad Merchant, Rainer Buchty, Mladen Berekovic |
ACM Trans. Reconfigurable Technol. Syst. | 5 |
| 2022 | RemEduLa - Remote Education Laboratory for FPGA Design TechnologyabstractTeaching hardware design is both challenging for teachers and students as it typically requires direct access to the targeted hardware platform for final testing. In this paper, we introduce RemEduLa - Remote Educational Laboratory for FPGA design technology. The core idea is to provide students with a developing experience as close as possible to presence teaching as part of lab courses. Therefore, the physical FPGA board is connected to a hardware server enabling the virtual instrumentation via a web interface. An overlay design with virtual inputs and outputs serves as a gateway, offering the student full control over their FPGA development board. This includes peripherals such as buttons, LEDs, and external components (sensors, actuators) as well as real-time visual feedback via a video stream. This one-to-one mapping of real hardware and students allows for the reuse of exercises formerly conducted in the on-site lab time. Christopher Blochwitz, Philipp Grothe, Sven Dreier, Waiel Aljnabi, Rainer Buchty, Mladen Berekovic |
ISCAS | 5 |
| 2016 | A Scriptable Standard-Compliant Reporting and Logging Framework for SystemC
Rolf Meyer, Jan Wagner, Bastian Farkas, Sven Alexander Horsinka, Patrick Siegl, Rainer Buchty, Mladen Berekovic |
ACM Trans. Embed. Comput. Syst. | 6 |
| 2015 | Revealing Potential Performance Improvements by Utilizing Hybrid Work-Sharing for Resource-Intensive Seismic ApplicationsabstractHeterogeneous system architectures are becoming more and more of a commodity in the scientific community. While it remains challenging to fully exploit such architectures, the benefits in performance and hybrid speed-up, by using a host processor and accelerators in parallel in a non-monolithic matter, are significant. Hereby, the energy efficiency is becoming an increasingly critical challenge for future high-performance computing (HPC) systems, which do want to exceed the Exascale barrier with several competing architecture concepts ranging from high-performance CPUs, combined with GPUs acting as floating-point accelerators, to computationally weak CPUs, paired with dedicated and highly-perform ant FPGA-based accelerators. In this paper, we realize and evaluate a hybrid computing approach based on a two-dimensional seismic streaming algorithm with several heterogeneous system architectures, including conventional HPC approaches based on powerful CPUs and GPUs. Furthermore, we elaborate the effort on an embedded system platform claiming to be a "mini supercomputer" [1]. Several CPU and accelerator combinations are utilized in a manual work-sharing manner with the aim of achieving significant performance speed-ups and a detailed energy-efficiency study. Based on roofline models and experimental evaluations, the paper provides an insight into the fact that hybrid computing is mostly unconditionally beneficial for balanced systems regarding the performance as well as the energy efficiency, aiding the programmer in the decision whether or not costly, manually tuned, homogeneous implementations are worthwhile. Patrick Siegl, Rainer Buchty, Mladen Berekovic |
PDP | 2 |
| 2013 | Efficient Barrier Synchronization for OpenMP-Like Parallelism on the Intel SCCabstractThe continuous increase of the number of processing cores on die poses a new set of challenges to HPC applications programming including how to model, write, and verify software that has to use the full power of NoC-based manycore processors. Therefore, to simplify program development for the Single-chip Cloud Computer (SCC), it is desirable to have high-level, shared memory-based parallel programming abstractions (e.g., an OpenMP-like programming model). One of the key components of any similar programming model are barrier synchronization primitives, coordinating the work of parallel threads. To allow high-level barrier constructs to deliver good performance, we need an efficient implementation of the underlying synchronization algorithm. In this paper, we propose effective barrier synchronization implementations for shared-memory programming on non-cache-coherent cluster-on-chip represented by the Intel SCC. In particular, we present an extensive evaluation of the overhead associated with integrating barrier algorithms required for OpenMP runtime libraries on such a machine, validating several implementation variants that efficiently exploit the network topology and leveraging SCC-specific hardware. We provide a detailed evaluation of the performance achieved by different approaches by using micro-benchmarks. Hayder Al-Khalissi, Rainer Buchty, Mladen Berekovic |
ICPADS | 2 |
| 2013 | Safe Virtual Interrupts Leveraging Distributed Shared Resources and Core-to-Core Communication on Many-Core PlatformsabstractModern many-core platforms offer sufficient redundant resources for increasing availability and fault-tolerance of multiple applications, also of different criticality (mixed-criticality). A suitable platform must allow remapping applications and replacing peripherals dynamically. Mapping to distributed resources but also communication among resources ideally is transparent and flexible to allow changes at run time. Communication additionally has to be predictable, especially for safety-critical applications, and can be efficiently implemented by the use of interrupt requests. This paper presents a scalable interrupt translation mechanism supporting flexible and transparent communication among resources. Our contribution is of particular benefit for legacy applications but also eases development of new applications. A fast and predictable monitoring and control mechanism enforces specified behavior of applications and peripherals communicating with critical applications at run time. This significantly reduces integration effort for mixed-critical applications on a shared platform, and thus makes many-core platforms more attractive for embedded and safety-critical systems. Boris Motruk, Jonas Diemer, Philip Axer, Rainer Buchty, Mladen Berekovic |
PRDC | 4 |
| 2012 | A survey on hardware-aware and heterogeneous computing on multicore processors and acceleratorsabstractSUMMARY In the last few years, the landscape of parallel computing has been subject to profound and highly dynamic changes. The paradigm shift towards multicore and manycore technologies coupled with accelerators in a heterogeneous environment is offering a great potential of computing power for scientific and industrial applications. However, for one to take full advantage of these new technologies, holistic approaches coupling the expertise ranging from hardware architecture and software design to numerical algorithms are a pressing necessity. Parallel computing is no longer limited to supercomputers and is now much more diversified – with a multitude of technologies, architectures, and programming approaches leading to increased complexity for developers and engineers. In this work, we give – from the perspective of numerical simulation and applications – an overview of existing and emerging multicore and manycore technologies as well as accelerator concepts. We emphasize the challenges associated with high‐performance heterogeneous computing and discuss the interfaces needed to fill the gap between the hardware architecture and the implementation of efficient numerical algorithms. By means of this short survey – which stresses the necessity of hardware‐aware computing – we aim at giving assistance to users in scientific computing entering this fascinating field and help understanding associated issues and capabilities. Copyright © 2011 John Wiley & Sons, Ltd. Rainer Buchty, Vincent Heuveline, Wolfgang Karl, Jan-Philipp Weiss |
Concurr. Comput. Pract. Exp. | 1 |
| 2012 | Seamlessly portable applications: Managing the diversity of modern heterogeneous systemsabstractNowadays, many possible configurations of heterogeneous systems exist, posing several new challenges to application development: different types of processing units usually require individual programming models with dedicated runtime systems and accompanying libraries. If these are absent on an end-user system, e.g. because the respective hardware is not present, an application linked against these will break. This handicaps portability of applications being developed on one system and executed on other, differently configured heterogeneous systems. Moreover, the individual profit of different processing units is normally not known in advance. In this work, we propose a technique to effectively decouple applications from their accelerator-specific parts, respectively code. These parts are only linked on demand and thereby an application can be made portable across systems with different accelerators. As there are usually multiple hardware-specific implementations for a certain task, e.g., a CPU and a GPU version, a method is required to determine which are usable at all and which one is most suitable for execution on the current system. With our approach, application and hardware programmers can express the requirements and the abilities of the application and the hardware-specific implementations in a simplified manner. During runtime, the requirements and abilities are compared with regard to the present hardware in order to determine the usable implementations of a task. If multiple implementations are usable, an online-learning history-based selector is employed to determine the most efficient one. We show that our approach chooses the fastest usable implementation dynamically on several systems while introducing only a negligible overhead itself. Applied to an MPI application, our mechanism enables exploitation of local accelerators on different heterogeneous hosts without preliminary knowledge or modification of the application. Mario Kicherer, Fabian Nowak, Rainer Buchty, Wolfgang Karl |
ACM Trans. Archit. Code Optim. | 3 |
| 2011 | Cost-aware function migration in heterogeneous systemsabstractToday's approaches towards heterogeneous computing rely on either the programmer or dedicated programming models to efficiently integrate heterogeneous components. In this work, we propose an adaptive cost-aware function-migration mechanism built on top of a light-weight hardware abstraction layer. With this mechanism, the highly dynamic task of choosing the most beneficial processing unit will be hidden from the programmer while causing only minor variation in the work and program flow. The migration mechanism transparently adapts to the current workload and system environment without the necessity of JIT compilation or binary translation. Mario Kicherer, Rainer Buchty, Wolfgang Karl |
HiPEAC | 2 |
| 2007 | Optimizing Cache Performance of the Discrete Wavelet Transform Using a Visualization ToolabstractThe 2D DWT consists of two 1D DWT in both directions: horizontal filtering processes the rows followed by vertical filtering processes the columns. It is well known that a straightforward implementation of the vertical filtering shows quite different performance with various working set sizes. The only reasonable explanation for this has to be the access behavior of the cache memory. As known, vertical filtering has mapping conflicts in the cache with a working set size that is power of two. However, it is not clear how this conflict forms and whether cache problems exist with other data sizes. Such knowledge is the base for efficient code optimization. In order to acquire this knowledge and to achieve more accurate optimization potentials, we apply a cache visualization tool to examine the runtime cache activities of the vertical implementation. We find that besides mapping conflicts, vertical filtering also shows a large number of capacity misses. More specifically, the visualization tool allows us to detect the parameters related to the strategies. This guarantees the feasibility of the optimization. Our initial experimental results on several different architectures show an up to 215% gain in execution time compared to an already optimized baseline implementation. Jie Tao 0001, Asadollah Shahbahrami, Ben H. H. Juurlink, Rainer Buchty, Wolfgang Karl, Stamatis Vassiliadis |
ISM | 4 |
| 2006 | A network agent for diagnosis and analysis of real-time Ethernet networksabstractWithin the field of automation technology the use of Industrial Ethernet is rising. This in turn demands devices capable of precisely recording, analyzing, and manipulating communication data for diagnostic purposes. Existing solutions so far lack required flexibility or are unable to cope with sustained Gigabit-per-second data streams. This is especially true for general-purpose approaches employing ordinary network adapters and plain software-based analysis.In this paper we describe a flexible and lightweight network agent for real-time, high-performance networks. This agent is capable of handling sustained data rates up to 2x 1GBit/s while offering real-time event-triggers, 10ns-resolution timestamps, real-time filtering, and statistics functions. An auxiliary processing unit as well as a modular software environment allow customization for a variety of tasks. The agent is realized as a dual processor SoC design on a Xilinx Virtex-II Pro FPGA. Hans-Peter Löb, Rainer Buchty, Wolfgang Karl |
CASES | 2 |
| 2005 | CPU-independent Assembler in an FPGAabstractWe describe a system which enables FPGAs to generate machine code for various CPUs, similar to a conventional assembler. Such conversion from intermediate code to a CPU's native code can be used as the last step in just-in-time compilation for virtual machines like the Java Virtual Machine. The translation system itself and the FPGA logic are independent of the actual target CPU and can be used with both CISC and RISC CPUs. Due to an extended table lookup, the resulting code is very efficient and gains from pre-calculation of selected constants in the FPGA assembler. Georg Acher, Rainer Buchty, Carsten Trinitis |
FPL | 2 |
| 2003 | AES and the cryptonite crypto processorabstractCRYPTONITE is a programmable processor tailored to the needs of crypto algorithms. The design of CRYPTONITE was based on an in-depth application analysis in which standard crypto algorithms (AES, DES, MD5, SHA-1, etc) were distilled down to their core functionality. We describe this methodology and use AES as a central example. Starting with a functional description of AES, we give a high level account of how to implement AES efficiently in hardware, and present several novel optimizations (which are independent of CRYPTONITE).We then describe the CRYPTONITE architecture, highlighting how AES implementation issues influenced the design of the processor and its instruction set. CRYPTONITE is designed to run at high clock rates and be easy to implement in silicon while providing a significantly better performance/area/power tradeoff than general purpose processors. Dino Oliva, Rainer Buchty, Nevin Heintze |
CASES | 2 |