Tim Hotfilter

dblp:276/4253 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
6since 2021 · last 2023
0000-0001-9748-3149ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2023 ATLAS: An Approximate Time-Series LSTM Accelerator for Low-Power IoT Applications
abstract
Enabling the use of Deep Neural Networks (DNNs) for time-series-based applications on low-power devices such as wearables opens up a wide range of new features and services. However, inference requires an enormous amount of operations to be performed by the computing platform. In addition, Long Short-Term Memory (LSTM)-based networks require memory to store the internal cell state for future calculations. In this paper, we therefore propose a hardware/software co-design based low-power LSTM hardware accelerator architecture for Internet of Things (IoT) applications called ATLAS. The design is based on approximate computing techniques to reduce the power consumption and inference latency by achieving high accuracy. Exemplary, we investigate the impact of applying our proposed architecture to a DNN for handwriting recognition. Thereby, we can show that the accuracy decreases only slightly when the inference is executed on ATLAS. The low power consumption is achieved by a minimal design requiring 173 LUTs, 67 FFs, one DSP, and one BRAM on a Xilinx FPGA. As a result, ATLAS enables the efficient use of LSTM-based DNNs in IoT devices.
Fabian Kreß, Alexey Serdyuk, Micha Hiegle, Disnebio Waldmann, Tim Hotfilter, Julian Höfer, Tim Hamann, Jens Barth, Peter Kämpf, Tanja Harbaum, Jürgen Becker 0001
DSD5
2023 SiFI-AI: A Fast and Flexible RTL Fault Simulation Framework Tailored for AI Models and Accelerators
abstract
For AI-based systems in safety-critical domains, it is inevitable to understand the impact of random hardware faults affecting the target hardware accelerators. The high degree of data reuse makes Deep Neural Network (DNN) accelerators susceptible to significant fault propagation and hence hazardous predictions. Therefore, we present SiFI-AI, a simulation framework for fault injection in DNN accelerators. SiFI-AI proposes a hybrid simulation approach combining fast AI inference with cycle-accurate RTL simulation. Time-expensive RTL simulation is only used to accurately target registers in the hardware through condition-based fault injection. This enables to reveal vulnerable DNN layers and the related fault origin. In a resilience study with 1.5~M fault injection experiments, we analyze representative DNNs and a state-of-the-art DNN accelerator to identify vulnerable layers. The study only takes 1.15 days which is 7x faster than state-of-the-art. Our experiments show the high impact of control register faults and that narrow and deep layers are 10x more resilient compared to the wide and shallow layers of a DNN.
Julian Höfer, Fabian Kempf, Tim Hotfilter, Fabian Kreß, Tanja Harbaum, Jürgen Becker 0001
ACM Great Lakes Symposium on VLSI3
2023 A Hardware-Aware Sampling Parameter Search for Efficient Probabilistic Object Detection
Julian Höfer, Tim Hotfilter, Fabian Kreß, Tanja Harbaum, Jürgen Becker 0001
ICVS2
2023 CNNParted: An open source framework for efficient Convolutional Neural Network inference partitioning in embedded systems
Fabian Kreß, Vladimir Sidorenko, Patrick Schmidt 0003, Julian Höfer, Tim Hotfilter, Iris Fürst-Walter, Tanja Harbaum, Jürgen Becker 0001
Comput. Networks5
2022 Towards Reconfigurable Accelerators in HPC: Designing a Multipurpose eFPGA Tile for Heterogeneous SoCs
abstract
The goal of modern high performance computing platforms is to combine low power consumption and high throughput. Within the European Processor Initiative (EPI), such an SoC platform to meet the novel exascale requirements is built and investigated. As part of this project, we introduce an embedded Field Programmable Gate Array (eFPGA), adding flexibility to accelerate various workloads. In this article, we show our approach to design the eFPGA tile that supports the EPI SoC. While eFPGAs are inherently reconfigurable, their initial design has to be determined for tape-out. The design space of the eFPGA is explored and evaluated with different configurations of two HPC workloads, covering control and dataflow heavy applications. As a result, we present a well-balanced eFPGA design that can host several use cases and potential future ones by allocating 1% of the total EPI SoC area. Finally, our simulation results of the architectures on the eFPGA show great performance improvements over their software counterparts.
Tim Hotfilter, Fabian Kreß, Fabian Kempf, Jürgen Becker 0001, Juan Miguel De Haro Ruiz, Daniel Jiménez-González, Miquel Moretó, Carlos Álvarez 0001, Jesús Labarta, Imen Baili
DATE1
2022 Hardware-aware Partitioning of Convolutional Neural Network Inference for Embedded AI Applications
abstract
Embedded image processing applications like multicamera-based object detection or semantic segmentation are often based on Convolutional Neural Networks (CNNs) to provide precise and reliable results. The deployment of CNNs in embedded systems, however, imposes additional constraints such as latency restrictions and limited energy consumption in the sensor platform. These requirements have to be considered during hardware/software co-design of embedded Artifical Intelligence (AI) applications. In addition, the transmission of uncompressed image data from the sensor to a central edge node requires large bandwidth on the link, which must also be taken into account during the design phase.Therefore, we present a simulation toolchain for fast evaluation of hardware-aware CNN partitioning for embedded AI applications. This approach explores an efficient workload distribution between sensor nodes and a central edge node. Neither processing all layers close to the sensor nor transmitting all uncompressed raw data to the edge node is an optimal solution for each use case. Hence, our proposed simulation toolchain evaluates power and performance metrics for each reasonable partitioning point in a CNN. In contrast to the state of the art, our approach does not only consider the neural network architecture. In the evaluation, our simulation toolchain additionally takes into account hardware components such as special accelerators and memories that are implemented in the sensor node.Exemplary, we show the simulation results for three commonly used CNNs in embedded systems. Thereby, we identify advantageous partitioning points regarding inference latency and energy consumption. With the support of the toolchain, we are able to identify three beneficial partitioning points for FCN ResNet-50 and two for GoogLeNet as well as for SqueezeNet V1.1.
Fabian Kreß, Julian Höfer, Tim Hotfilter, Iris Fürst-Walter, Vladimir Sidorenko, Tanja Harbaum, Jürgen Becker 0001
DCOSS3