VLDB 2026 Research / reviewers in the wild / expert
Tobias Gemmeke
dblp:52/1021
· DBLP profile ↗
32ranked-venue papers
1as first author
25since 2021 · last 2026
0000-0003-1583-3411ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 25 · 1 first-author · 20 since 2021Artificial intelligence and machine learning · 5 · 4 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ConfASR: A Conformer Block Accelerator for Speech Recognition Optimized for Edge DevicesabstractAttention-based neural networks, like transformers, have significantly improved automatic speech recognition (ASR). Adding convolution operations to transformers results in conformers, which enable better learning of local dependencies and reduce word error rate. We introduce ConfASR, the first conformer block accelerator designed for efficient ASR inference on edge devices. Our system is optimized both in terms of algorithms and hardware to support all transformer operations and additional features required for conformers, including depthwise-separable convolution and learned positional encoding. We propose a hardware-friendly normalization, shared scaling factors for non-linear functions, and an efficient dataflow with a shared MAC array that keeps all activations on chip. Implemented in a 22 nm FDSOI technology, ConfASR operates at 250 MHz with a power consumption of 359 mW, and a die area of $1.19 \mathrm{~mm}^{2}$. It performs over 900 times faster than necessary for real-time streaming requirements. This makes the architecture suitable not only for ASR but also for other transformer-based applications. ConfASR reduces latency by over $4 \times$ and power consumption by $16 \times$ during real-time use compared to previous solutions, while supporting more functionality. Malte Wabnitz, Max Nilovic, Finn Scholz, Dominik Friedrich, Christian Lanius, Jie Lou, Tobias Gemmeke |
ASP-DAC | 7 |
| 2026 | GUPrecision: Group-Wise Uniform Precision Accelerator for Depthwise Separable Convolution using Hardware-Algorithm Co-DesignabstractQuantization and pruning are effective techniques for reducing neural network size and improving energy efficiency. Although fixed word length networks are well-suited for hardware acceleration, mixed-precision and pruned networks still suffer from efficient hardware support. Depthwise separable convolution (DSC) has become a key building block for resource-constrained devices; however, applying quantization and pruning to DSC models remains challenging. To address these challenges, we propose GUPrecision, a group-wise mixed-precision uniform quantization framework that inherently supports network pruning while enabling efficient hardware realization. GUPrecision achieves hardware compatibility by dividing channels into subgroups with a fixed total bit budget. Within each group, the mixed-precision multipliers in the PE array can dynamically adapt to varying word lengths using simple shifting and multiplexing operations. We evaluated GUPrecision on the MobileNetV1 model and implemented it using GlobalFoundries 22 nm FDSOI technology. The DSC accelerator operates at 1 GHz and 0.8 V after signoff, occupying an area of 0.71 mm2 . At 75% effective sparsity, it achieves a peak energy efficiency of 17.4 TOPS/W, with a corresponding throughput of 8136 GOPS and an area efficiency of 11459 GOPS/mm2. Jie Lou, Malte Wabnitz, Tobias Gemmeke |
ACM Great Lakes Symposium on VLSI | 4 |
| 2025 | A 1.27 fJ/B/transition Digital Compute-in-Memory Architecture for Non-Deterministic Finite Automata Evaluation
Christian Lanius, Florian Freye, Tobias Gemmeke |
ACM Great Lakes Symposium on VLSI | 3 |
| 2025 | A 22nm 96.83-TOPS/W Time-Domain Compute-in-Memory Engine Utilizing Mixed-Fidelity for Edge-AI Applications
Jie Lou, Florian Freye, Christian Lanius, Tobias Gemmeke |
ACM Great Lakes Symposium on VLSI | 4 |
| 2025 | Balanced and Efficient Spiking Neural NetworksabstractSpiking Neural Networks promise an efficient and robust alternative to traditional Artificial Neural Networks. When compared to the human brain, however, they are far from achieving their full potential, the reason unclear so far. Recent neuroscience research offers a clue to solve this puzzle - tight balance. The excitatory and inhibitory input currents flowing into individual neurons have been observed to be balanced very tightly on a millisecond scale. This substantially exceeds the traditional E/I balance that is established on population level. Accompanying theoretical works have shown that spiking networks operating in this tight regime can implement more efficient and robust codes in their spike trains. In this work, we consolidate these theoretical insights with modern neuromorphic computing approaches to apply them to complex classification problems. We show up to 99.9% reductions in firing rates in balanced networks compared to unbalanced ones on the neuroscience-inspired cue-based navigation task, and outperform prior work by achieving 97.60% accuracy. The resulting approach can easily be mapped to neuromorphic hardware as its core relies on a balanced initialization scheme only. Moreover, even complex variants only involve adaptions of neuron and synapse models. Ultimately, our work shows the importance of integrating novel insights from neuroscience research into modern spiking networks to increase their efficiency and robustness using neuromorphic methods. Tim Stadtmann, Janek Paeßens, Tobias Gemmeke |
IJCNN | 3 |
| 2025 | An All-Digital Time-Domain Compute-in-Memory Engine for Convolutional Neural Networks in 22nmabstractThis paper presents a standard cell (SC) based time-domain compute-in-memory (TDCIM) macro for convolutional neural networks (CNNs), supporting 4b×4b multiplication. A basic cell is proposed to enable bitwise multiplication with 1-bit weights and 2-bit activations, along with double-edge computing. An 8-stage successive approximation register time-to-digital converter (SAR-TDC) is employed to convert time-domain signals into the digital domain. We leverage the inherent features of the network to enhance throughput and have fabricated the TDCIM macro in a 22nm technology. We present the measured delay and variation of the basic cell, the integral nonlinearity (INL) of the TDC, and the total computation error. The proposed macro achieves an energy efficiency of 98.44 TOPS/W at 0.55V for 4b-input and 4b-weight MAC computations. Jie Lou, Florian Freye, Christian Lanius, Tobias Gemmeke |
ISCAS | 4 |
| 2025 | A Compact SHA256 Accelerator in 22nm for Energy Bounded Use-Cases with 8.2GHash/JabstractA 6.8•103μm2SHA256 hardware accelerator, achieving 8.2GHash/J and 3MHash/s throughput at 460mV, is fabricated in 22nm CMOS. Round-based dataflow with added pipelining and hybrid shift-FIFO structure provides 31% increase in clock frequency and 21% reduction in energy consumption. Replacing DFFs with pulsed D-latch can further increase clock frequency by 15% and decrease energy cost by 13%. Combination of high energy efficiency, high clock frequency and compact area makes the proposed SHA256 engines suitable for a wide range of applications spanning from IoT to bitcoin mining. Xin Fan 0002, Michael Gansen, Tobias Gemmeke |
ISCAS | 4 |
| 2024 | A DfT Strategy for Guaranteeing ReRAM's Quality after ManufacturingabstractAbstract Memristive devices have become promising candidates to complement the CMOS technology, due to their CMOS manufacturing process compatibility, zero standby power consumption, high scalability, as well as their capability to implement high-density memories and new computing paradigms. Despite these advantages, memristive devices are susceptible to manufacturing defects that may cause faulty behaviors not observed in CMOS technology, significantly increasing the challenge of testing these novel devices after manufacturing. This work proposes an optimized Design-for-Testability (DfT) strategy based on the introduction of a DfT circuitry that measures the current consumption of Resistive Random Access Memory (ReRAM) cells to detect not only traditional but also unique faults. The new DfT circuitry was validated using a case study composed of a 3x3 word-based ReRAM with peripheral circuitry implemented based on a 130 nm Predictive Technology Model (PTM) library. The obtained results demonstrate the fault detection capability of the proposed strategy with respect to traditional and unique faults. In addition, this paper evaluates the impact related to the DfT circuitry’s introduced overheads as well as the impact of process variation on the resolution of the proposed DfT circuitry. Thiago Copetti, Moritz Fieback, Tobias Gemmeke, Said Hamdioui, Letícia Maria Veiras Bolzani |
J. Electron. Test. | 3 |
| 2024 | An Energy Efficient All-Digital Time-Domain Compute-in-Memory Macro Optimized for Binary Neural NetworksabstractThe deployment of neural networks on edge devices has created a growing need for energy-efficient computing. In this paper, we propose an all-digital standard cell-based time-domain compute-in-memory (TDCIM) macro for binary neural networks (BNNs) that is compatible with commercial digital design flow. The TDCIM macro utilizes multiple computing chains that share one threshold chain, and supports double-edge operation, parallel computing and data reuse. Time-domain wave-pipelining technique is introduced to enhance throughput while preserving accuracy. Regular placement (RP) and custom routing (CR) are employed during place and route (P&R) to reduce systematic variations. We show computing delay, POOL computation accuracy, and network test accuracy at different voltages, indicating that the proposed TDCIM macro can maintain high accuracy under PVT variations. We implemented two versions of the TDCIM macro in 22nm FDSOI technology using foundry-provided delay cells DLY40 and DLY60, respectively. At a voltage of 0.5V, the TDCIM macro achieved an energy efficiency of 1.2 (1.05) POPS/W for DLY40 (DLY60), while maintaining a baseline accuracy of 98.9% on the MNIST dataset for both designs. Jie Lou, Florian Freye, Christian Lanius, Tobias Gemmeke |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2024 | Fully Digital, Standard-Cell-Based Multifunction Compute-in-Memory Arrays for Genome SequencingabstractThe rapid advancement in genome sequencing technology has led to a significant increase in the number of genomic reads in recent years. Due to the immense size of reference genomes, which can be up to 3 billion bases, finding optimal solutions for through approximate string matching proves to be computationally challenging. Current alignment algorithms address this by performing a preprocessing step to efficiently calculate likely matching regions and only aligning at the base level within these regions. This article demonstrates the acceleration of sorting and searching in memories, both crucial components of genome alignment algorithms. We designed a compute-in-memory (CIM) array using standard cells, which is capable of sorting datastreams blockwise, merging sorted blocks, as well as operating as a content addressable memory (CAM) while also being able to perform multiword logic operations. We address the problem of datasets not fitting into on-chip memory by reusing the CIM array for a merge sorting step, enabling arbitrarily sized sorting. Our 2.6-$\mu \text{m}^{2}$/bit design, fabricated using 22-nm fully depleted silicon-on-insulator (FDSOI) technology, yields a throughput of up to 4.28 GB/s at$f_{\text {max}}$and 4.97 nJ/sort at the minimum energy point (MEP) when executing sort operations. Christian Lanius, Tobias Gemmeke |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2024 | Stream Processing Architectures for Continuous ECG Monitoring Using Subsampling- Based ClassifiersabstractMonitoring of biomedical data, such as electrocardiogram (ECG) signals, requires accelerators, which can process data streams in a continuous manner. Especially, wearable monitoring systems require both ultralow power consumption and sufficiently complex deep neural network (DNN) classifiers to identify asymptomatic and critical health conditions, such as atrial fibrillation (AF). Such continuous data streams pose unique constraints on the processing pipeline for classification systems, which can be addressed in the design methodology of application-specific integrated circuits (ASICs). In this work, we identify specific constraints to define common operating conditions, which guide the design of ECG accelerators in an algorithm–hardware codesign methodology. In specific, we show that the input frame size and the number of classifications per time frame play a significant role for the computational complexity (CC) of the classifier, as well as the ECG accelerator executing the classifier in a continuous manner. As an example, the constraints are applied in a top-down algorithm–hardware codesign flow. Here, an ECG accelerator is designed starting from an AF classifier, while proposed constraints are considered in an early design stage to estimate costs for the hardware design. In the end, it is essential for future ECG accelerators to adhere to common constraints in the design process to handle increasingly complex DNN classifiers for continuous data streams with ultralow power targets. Johnson Loh, Tobias Gemmeke |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2023 | Lossless Sparse Temporal Coding for SNN-based Classification of Time-Continuous SignalsabstractUltra-low power classification systems using spiking neural networks (SNN) promise efficient processing for mobile devices. Temporal coding represents activations in an artificial neural network (ANN) as binary signaling events in time, thereby minimizing circuit activity. Discrepancies in numeric results are inherent to common conversion schemes, as the atomic computing unit, i.e. the neuron, performs algorithmically different operations and, thus, potentially degrading SNN's quality of service (QoS). In this work, a lossless conversion method is derived in a top-down design approach for continuous time signals using electrocardiogram (ECG) classification as an example. As a result, the converted SNN achieves identical results compared to its fixed-point ANN reference. The computations, implied by proposed method, result in a novel hybrid neuron model located in between the integrate-and-fire (IF) and conventional ANN neuron, which numerical result is equivalent to the latter. Additionally, a dedicated SNN accelerator is implemented in 22 nm FDSOI CMOS suitable for continuous real-time classification. The direct comparison with an equivalent ANN counterpart shows that power reductions of$2.32\times$and area reductions of$7.22\times$are achievable without loss in QoS. Johnson Loh, Tobias Gemmeke |
DATE | 2 |
| 2023 | Scalable Time-Domain Compute-in-Memory BNN Engine with 2.06 POPS/W Energy Efficiency for Edge-AI DevicesabstractTime-domain (TD) computing has attracted attention for its high computing efficiency and suitability for applications on energy-constrained edge devices. In this paper, we present a time-domain compute-in-memory (TDCIM) macro for binary neural networks (BNNs) realized by standard as well as custom delay cells. Multiply-and-accumulate (MAC) operations, batch normalization (BN) and binarization (Bin) are all processed in the time-domain, avoiding costly digital domain post-processing. In addition, it supports flexible mapping for different kernel sizes, achieving 100% utilization. Starting from a standard cell-based implementation, we propose two custom cells that provide interesting trade-offs between energy efficiency, area and accuracy. The two proposed custom designs can achieve 1.5 and 2.06 POPS/W energy efficiencies at 0.5V and 0.6V with less cell area while maintaining model test accuracy. Jie Lou, Florian Freye, Christian Lanius, Tobias Gemmeke |
ACM Great Lakes Symposium on VLSI | 4 |
| 2023 | NoisyDECOLLE: Robust Local Learning for SNNs on Neuromorphic HardwareabstractBased on their biological archetype, spiking neural networks (SNNs) promise substantial energy savings at massively increased performance. However, best-performing SNNs in supervised learning scenarios often fall short of their potential efficiency gains as they rely on backpropagation, which entails issues like weight update locking until forward and backward passes are finished, and critical resource & computation requirements. These challenges can be tackled by drawing inspiration from biological synapses which rely largely on more local information to adapt. Local learning algorithms adopt this idea by employing local classifiers in each layer of an SNN that use only spatially and temporally local information to update synaptic weights. However, mapping these algorithms to neuromorphic systems to unleash their potential can be impaired by various kinds of noise. In this work, we review prior art to derive realistic noise scenarios on neuromorphic systems. Based on these results, we introduce NoisyDEColle, a framework for applying various noise models on locally learned SNNs using the DECOLLE network architecture. We show that both noise-aware training and additional regularization techniques allow NoisyDECOLLE to reach competitive performance, even under challenging conditions as for example posed by resistive devices. Using quantization-aware training, NoisyDECOLLE reaches 98.6% (94.1%) accuracy on N-MNIST (DVSGesture) with 3b (8b) weights ($> 4-10\mathrm{x}$memory and energy savings compared to 32b). To analyze our results, we provide the first time-driven implementation of spike activation maps, aiding the explainability of neuromorphic computing. Tim Stadtmann, Benedikt Wahl, Tobias Gemmeke |
ICMLA | 3 |
| 2023 | Hardware Trojans in fdSOIabstractWith shortening turn-around times and increasing complexity for digital circuits, design reuse, third party IP and today even physical chiplets has increased. Malicious actors have more options to introduce hardware backdoors to packaged systems, which will leak data if triggered. In this work, we show two novel approaches to introduce such backdoors, that are possible due to the specifics of fully depleted silicon on insulator (fdSOI) technology. The first method relies on modifying the doping profile of an antenna cell to introduce a covert short between the back gate and logic signals. The second method constructs specific illegal states which are latched when the clock is running with the trigger frequency. Basic test structures have been designed such that they are DRC and STA clean. LVS does not reveal the hidden structure, while measurements in silicon confirm their operation. Christian Lanius, Florian Freye, Tobias Gemmeke |
ISLPED | 4 |
| 2023 | Automatic Generation of Structured Macros Using Standard Cells ‒ Application to CIMabstractRegularity can be exploited to efficiently describe, place and route logic blocks with a repetitive structure. We present a design flow to automatically generate regular, standard-cell based designs, which can be seamlessly integrated into a traditional digital flow in commercial EDA software. The generated arrays can be seamlessly integrated into a “sea of gates”. with no guard-rings or keep-out areas. The flow takes a description of a regular design as an input and generates netlist, placement, constraints, routing, initial parasitics estimates and timing information. We show that, in example designs, the run-time of EDA tooling is up to 2.5x faster, reduces the critical path by 47%, reduces the metal utilization by 45% and achieves a utilization of 93%. Christian Lanius, Jie Lou, Johnson Loh, Tobias Gemmeke |
ISLPED | 4 |
| 2023 | FPGA-based Acceleration of Lidar Point Cloud Processing and Detection on the EdgeabstractEdge nodes such as Intelligent Transportation System Stations are becoming increasingly relevant in the context of automated driving as they provide connected vehicles with additional information to support their automated driving functions. However, the power budget for these edge nodes is limited and data has to be processed in real-time to be of use to automated driving functions. In this work, we present a system for processing raw lidar data in real-time on an FPGA, resulting in a significant reduction in power consumption compared to conventional hardware. Our approach leads to a 42.4% reduction in power consumption while maintaining the quality of the results. Processing two 128-layer surround-view lidar point clouds takes 522 ms per frame and an average power consumption of 39.3 W for the CPU and 34.5W for the FPGA. Our optimizations surpass the state-of-the-art by up to 193 times. Cecilia Latotzke, Amarin Kloeker, Simon Schoening, Fabian Kemper 0002, Mazen Slimi, Lutz Eckstein, Tobias Gemmeke |
IV | 7 |
| 2023 | An Energy-Efficient and Area-Efficient Depthwise Separable Convolution Accelerator with Minimal On-Chip Memory AccessabstractDepthwise separable convolution (DSC) has emerged as a crucial building block for developing lightweight convolutional neural networks (CNNs). In this paper, we present a hardware accelerator for DSC that enables 100% utilization of the processing element (PE) array for depthwise convolution (DWC) and achieves up to 98% utilization for pointwise convolution (PWC), while also reducing latency. By partitioning the input feature map (ifmap) SRAM of the DWC into three banks, we minimize memory access and maximize data reuse. The input activations and weights only need to be loaded once from SRAM to PE for both DWC and PWC. Additionally, to support efficient operations across different layers, we present a layerwise matching method. The proposed DSC accelerator is implemented in 22nm FDSOI technology and validated using MobileNetV1 on the CIFAR10 dataset. The post-layout results demonstrate that the proposed accelerator can operate at 1GHz and achieve an energy efficiency of 5.07 (3.96) TOPS/W and an area efficiency of 519.2 (461.52) GOPS/mm2for DWC (PWC) at 0.8V. After scaling the supply voltage down to 0.5V, the energy efficiency for the proposed accelerator increases to 13.64 TOPS/W for DWC and 10.64 TOPS/W for PWC, respectively. Jie Lou, Christian Lanius, Florian Freye, Johnson Loh, Tobias Gemmeke |
VLSI-SoC | 6 |
| 2023 | A Digital Twin Network for Computational Neuroscience Simulators: Exploring Network Architectures for Acceleration of Biological Neural Network SimulationsabstractDespite recent advances in the domain of Computational Neuroscience (CN), the accelerated simulation of large-scale biological networks of natural density is daunting due to scalability issues and communication constraints of existing CN simulators. This challenge can be addressed by developing an ultra-low latency packet delivery service among the tightly coupled processing nodes of a CN simulator. Given that CN is a rapidly evolving field, constant updates are inevitable during the life of a CN simulator. Hence, a digital twin network for CN simulators is crucial as it can enable low-cost prototyping to keep their network up to date with ever-changing requirements caused by advances in CN. To this end, we have developed a framework to replicate the dynamic network behavior of potential CN simulators. The core of our framework is based on high-end FPGA boards with high data rate transceivers, enabling emulation of various network architectures with different topologies and executed protocols. The precise dynamic behavior of the FPGA-based digital twin provides accurate modeling and reliable assessments of real-time dynamics. To expedite the exploration and prototyping process, we use the FPGA cluster in conjunction with in-house software-based network simulators. Our initial evaluations based on the developed software network simulators resulted in an overestimation of network performance. However, after calibration, our results demonstrate the potential of our approach in addressing complex problems such as assessing large-scale networks. Vida Sobhani, Kevin Kauth, Tim Stadtmann, Tobias Gemmeke |
WoWMoM | 4 |
| 2022 | NEUROTEC I: Neuro-inspired Artificial Intelligence Technologies for the Electronics of the FutureabstractThe field of neuromorphic computing is approaching an era of rapid adoption driven by the urgent need of a substitute for the von Neumann computing architecture. NEUROTEC I: “Neuro-inspired Artificial Intelligence Technologies for the Elec-tronics of the Future” project is an initiative sponsored by the German Federal Ministry of Education and Research (BMBF for its initials in German), that aims to effectively advance the foundations for the utilization and exploitation of neuromorphic computing. NEUROTEC I stands at its successful “final stage” driven by the collaboration from more than 8 institutes from the Jiilich Research Center and the RWTH Aachen University, as well as collaboration from several high-tech industry partners. The NEUROTEC I project considers the field interplay among materials, circuits, design and simulation tools. This paper provides an overview of the project's overall structure and discusses the scientific achievements of its individual activities. Melvin Galicia, Stephan Menzel, Farhad Merchant, Maximilian Müller, Qing-Tai Zhao, Felix Cüppers, Abdur R. Jalil, Qi Shu, Peter Schüffelgen, Gregor Mussler, Carsten Funck, Christian Lanius, Stefan Wiefels, Moritz von Witzleben, Christopher Bengel, Nils Kopperberg, Tobias Ziegler 0005, R. Walied Ahmad, Alexander Krüger, Letícia Maria Veiras Bolzani, Regina Dittmann, Susanne Hoffmann-Eifert, Vikas Rana, Detlev Grützmacher, Matthias Wuttig, Dirk J. Wouters, Andrei Vescan, Tobias Gemmeke, Joachim Knoch, Max Christian Lemme, Rainer Leupers, Rainer Waser |
DATE | 29 |
| 2022 | Design of High-Throughput Mixed-Precision CNN Accelerators on FPGAabstractConvolutional Neural Networks (CNNs) reach high accuracies in various application domains, but require large amounts of computation and incur costly data movements. One method to decrease these costs while trading accuracy is weight and/or activation word-length reduction. Thereby, layer-wise mixed-precision quantization allows for more efficient results while inflating the design space. In this work, we present an in-depth quantitative methodology to efficiently explore the design space considering the limited hardware resources of a given FPGA. Our holistic exploration approach vertically traverses the various design entry levels from the architectural down to the logic level, and laterally covers optimization from processing elements to dataflow for an efficient mixed-precision CNN accelerator. Our resulting hardware accelerators implement truly mixed-precision operations that enable efficient execution of layer-wise and channel-wise quantized CNNs. Mapping feed-forward and identity-shortcut-connection mixed-precision CNNs result in competitive accuracy-throughout trade-offs: 245 frames/s with 87.48% Top-5 accuracy for ResNet-18 and 92.9% Top-5 accuracy with 1.13 TOps/s for ResNet-152, respectively. Thereby, the required memory footprint for parameters is reduced by 4.9 × and 9.4 × compared to the respective floating-point baseline. Cecilia Latotzke, Tim Ciesielski, Tobias Gemmeke |
FPL | 3 |
| 2022 | Post-Training Quantization for Energy Efficient Realization of Deep Neural NetworksabstractThe biggest challenge for the deployment of Deep Neural Networks (DNNs) close to the generated data on edge devices is their size, i.e., memory footprint and computational complexity. Both are significantly reduced with quantization. With the resulting lower word-length, the energy efficiency of DNNs increases proportionally. However, lower word-length typically causes accuracy degradation. To counteract this effect, the quantized DNN is retrained. Unfortunately, training costs up to 5000× more energy than the inference of the quantized DNN. To address this issue, we propose a post-training quantization flow without the need for retraining. For this, we investigated different quantization options. Furthermore, our analysis systematically assesses the impact of reduced word-lengths of weights and activations revealing a clear trend for the choice of word-length. Both aspects have not been systematically investigated so far. Our results are independent of the depth of the DNNs and apply to uniform quantization, allowing fast quantization of a given pre-trained DNN. We excel state-of-the-art for 6 bit by 2.2% Top-1 accuracy for ImageNet. Without retraining, our quantization to 8 bit surpasses floating-point accuracy. Cecilia Latotzke, Batuhan Balim, Tobias Gemmeke |
ICMLA | 3 |
| 2022 | Compiling All-Digital-Embedded Content Addressable Memories on Chip for Edge ApplicationabstractA spectrum of emerging applications, including edge artificial intelligence, advocates the precompute-and-search scheme with embedded small-size content addressable memory (CAM) for its hardware efficiency instead of repetitive arithmetic operations. However, the integration of the CAM macros that are conventionally implemented with full custom analog circuits renders design-space exploration and optimization to be difficult at system level. As an alternative, a complete design flow for compiling ternary CAM on chip using foundry-supplied digital standard cells is introduced in this article. Based on the novel CAM architecture and logic design, we leverage guided placement and routing with mainstream EDA tools for exploiting the inherent structure regularity of CAM-cell arrays in the layout. An analytical model is also presented, which allows us for a systematic investigation on the energy reduction by adapting our design to various presearch structures. Validated on a postlayout$32{\times }64$ternary CAM in 28 nm, our parallel-matching scheme performs at 2.6 GHz with 0.42 fJ/bit/search, and the (8-bit) presearch scheme achieves 0.19 fJ/bit/search at 1.1 GHz, both under the 0.9-V supply voltage. In addition to flexibility for tradeoffs between the search throughput and energy, our all-digital CAM design enables voltage scaling aggressively down to 0.45 V with a minimum energy consumption of 0.06 fJ/bit/search at 50 MHz. Xin Fan 0002, Niklas Meyer, Tobias Gemmeke |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2021 | Validating a DFT Strategy's Detection Capability regarding Emerging Faults in RRAMsabstractOver the last fifty years, Complementary Metal Oxide Semiconductor (CMOS) technology has been scaled down according to the predictions made by Gordon Moore in the 1960s, hence making the design of high-performance applications possible. However, there is a growing concern that device scaling will become infeasible below a certain feature size. In parallel, emerging applications present high demands regarding storage and computing capability, combined with challenging constraints in terms of size, power consumption, and response latency. Thus, memristive devices have become promising candidates to complement or replace the CMOS technology due to their CMOS manufacturing process compatibility, zero standby power consumption, high scalability and density, as well as their capability to implement high-density memories as well as new computing paradigms. Despite these advantages, memristive devices are also suscreptible to manufacturing defects that may cause faulty behaviors not observed in CMOS, significantly increasing the test complexity. This paper presents the validation of a Design-for-Testability (DFT) strategy for Resistive Random Access Memories (RRAMs). The proposed strategy, able to detect traditional and unique faults in RRAM cells, has been implemented using an X-Fab technology library and validated based on a simplified case study. The obtained results show that the idea of applying a predefined operating sequence in combination with electrical measurements can guarantee the detection of unique faults in RRAM cells. Thiago Copetti, Tobias Gemmeke, Letícia Maria Veiras Bolzani |
VLSI-SoC | 2 |
| 2021 | Review of Manufacturing Process Defects and Their Effects on Memristive DevicesabstractAbstract Complementary Metal Oxide Semiconductor (CMOS) technology has been scaled down over the last forty years making possible the design of high-performance applications, following the predictions made by Gordon Moore and Robert H. Dennard in the 1970s. However, there is a growing concern that device scaling, while maintaining cost-effective production, will become infeasible below a certain feature size. In parallel, emerging applications including Internet-of-Things (IoT) and big data applications present high demands in terms of storage and computing capability, combined with challenging constraints in terms of size, power consumption and response latency. In this scenario, memristive devices have become promising candidates to complement the CMOS technology due to their CMOS manufacturing process compatibility, great scalability and high density, zero standby power consumption and their capacity to implement high density memories as well as new computing paradigms. Despite these advantages, memristive devices are also susceptible to manufacturing defects that may cause unique faulty behaviors that are not seen in CMOS, increasing significantly the complexity of test procedures. This paper provides a review about the manufacturing process of memristives devices, focusing on Valence Change Mechanism (VCM)-based memristive devices, and a comparative analysis of the CMOS and memristive device manufacturing processes. Moreover, this paper identifies possible manufacturing failure mechanisms that may affect these novel devices, completing the list of the already known mechanisms, and provides a discussion about possible faulty behaviors. Note that the identification of these mechanisms provides insights regarding the possible memristive devices’ defective behaviors, enabling to derive more accurate fault models and consequently, more suitable test procedures. Letícia Maria Veiras Bolzani, Moritz Fieback, Susanne Hoffmann-Eifert, Thiago Copetti, E. Brum, Stephan Menzel, Said Hamdioui, Tobias Gemmeke |
J. Electron. Test. | 8 |
| 2020 | Low-Cost DNN Hardware Accelerator for Wearable, High-Quality Cardiac Arrythmia DetectionabstractThis work implements a digital signal processing (DSP) accelerator for ECG signal classification. Targeting the integration into wearable devices for 24/7 monitoring, low energy consumption per classification is a key requirement, while maintaining a high classification accuracy at the same time. Co-optimization on algorithm and hardware level led to an architecture consisting mostly of convolution operations in the processing pipeline. The realized discrete wavelet transform and convolutional neural network (CNN) is utilized for continuous time-sequence classification in a sliding-window approach moving away from sample/batch-based processing typical for CNNs. In contrast to previous hardware realizations in this domain, the proposed design was validated using the benchmark dataset from the demanding CinC challenge 2017. The architecture achieves a competitive 0.781 Fl-score with only 5597 trainable parameters reducing the computational complexity of state-of-the-art ECGDNN software solutions by three orders of magnitude. Synthesis in a 22-nm FDSOI-CMOS technology features 0.783 $\mu$J per solution meeting requirements for edge device operation at high-end classification performance. Johnson Loh, Jianan Wen, Tobias Gemmeke |
ASAP | 3 |
| 2020 | From Quantitative Analysis to Synthesis of Efficient Binary Neural NetworksabstractBinary Neural Networks (BNNs) offer an effective way to slash the cost of computation and memory accesses in inference. Recently, a plurality of ideas has been proposed, some of which are complementary while others are incompatible. This work presents a thorough review of state-of-the-art methods and an analysis of their computational cost based on the energy consumption of fixed-point, ternary and binary MAC vector operations. We derive an approach on how to systematically design a cost-efficient BNN. Our quantized LeNet and VGGNet architectures highlight the benefit of prudent capacity augmentation, with layer-wise ternarization providing best improvement of accuracy over μJ/dassificatíon in BNNs. Tim Stadtmann, Cecilia Latotzke, Tobias Gemmeke |
ICMLA | 3 |
| 2020 | Approximation of Transcendental Functions With Guaranteed Algorithmic QoS by Multilayer Pareto OptimizationabstractDesign space exploration of approximate computing is deemed tough. While restricting the optimization to reduce approximate errors at the individual-function level might simplify the problem to be tackled, the increased hardware cost should be well justified by the improvement of the quality of service (QoS) at the algorithm level. In light of the loose correlation between atomic errors and algorithmic QoS, it is imperative but computationally expensive to explore approximate design space incorporating a variety of alternatives for evaluation. Despite being addressed extensively in the literature, we consider the transcendental sigmoid (σ) and hyperbolic tangent (tanh) functions as typical examples manifesting the dilemma of approximate computing. In this work, we leverage Pareto-front optimization at three hierarchies, from the parameter layer up to the structure and algorithm layers, to effectively tailor the approximate design space of the σ and tanh functions for hardware-efficient as well as algorithm-feasible implementations. Our investigations are performed based on a comprehensive design library consisting of representative approximate schemes in radically different hardware structures featuring both linear and nonlinear approximations. As tested on MNIST with 99% accuracy and Wisconsin Breast Cancer data set with 96.6% accuracy, we identify a novel and compact shift-based approximation that directly applies to the two's complement numbers achieving a 1.5× area reduction and a 3.4× energy reduction compared with the prior art. Provided the flexibility in approximate functions is of concern, we also present a uniform yet concise structure for implementing the Chebyshev-polynomials-based approximation adaptive in silicon to arbitrary nonlinear functions and error constraints. Xin Fan 0002, Tobias Gemmeke |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2018 | Physical modeling of bitcell stability in subthreshold SRAMs for leakage-area optimization under PVT variationsabstractSubthreshold SRAM design is crucial for addressing the memory bottleneck in energy constrained applications. While statistical optimization can be applied based on Monte-Carlo (MC) simulation, exploration of bitcell design space is time consuming. This paper presents a framework for model-based design and optimization of subthreshold SRAM bitcells under random PVT variations. By incorporating key design and process features, a physical model of bitcell static noise margin (SNM) has been derived analytically. It captures intra-die SNM variations by the combination of a folded-normal distribution and a non-central chi-squared distribution. Validations with MC simulation show its accuracy of modeling SNM distributions down to 25mV beyond 6-sigma for typical bitcells in 28nm. Model-based tuning of subthreshold SRAM bitcells is investigated for design tradeoff between leakage, area and stability. When targeting a specific SNM constraint, we show that an optimal standby voltage exists which offers minimum bitcell leakage power – any deviation above or below increases the power consumption. When targeting a specific standby voltage, our design flow identifies bitcell instances of 12× less leakage power or 3× reductions in area as compared to the minimum-length design. Xin Fan 0002, Tobias Gemmeke |
ICCAD | 3 |
| 2015 | On the use of analytical techniques for reliability analysis in presence of hardware-induced errorsabstractReliability related to hardware-induced errors has become one of the important design issues in the design of modern digital systems, and is inherently in conflict with other goals, especially energy-efficiency. To derive reliability approaches that are cost-effective, accurate fault/error information should be provided for the system across the design layers. This information typically comes from system simulation, which is time-consuming due to the huge complexity of modern systems. This paper discusses the potential and the limitations of using analytical techniques as a more scalable way to study error propagation. Examples are given relating to the domain of communication systems. Georgia Psychou, Tobias Gemmeke, Tobias G. Noll |
INDIN | 2 |
| 2014 | Resolving the memory bottleneck for single supply near-threshold computingabstractThis paper focuses on a review of state-of-the-art memory designs and new design methods for near-threshold computing (NTC). In particular, it presents new ways to design reliable low-voltage NTC memories cost-effectively by reusing available cell libraries, or by adding a digital wrapper around existing commercially available memories. The approach is based on modeling at system level supported by silicon measurement on a test chip in a 40nm low-power processing technology. Advanced monitoring, control and run-time error mitigation schemes enable the operation of these memories at the same optimal near-Vtvoltage level as the digital logic. Reliability degradation is thus overcome and this opens the way to solve the memory bottleneck in NTC systems. Starting from the available 40 nm silicon measurements, the analysis is extended to future 14 and 10 nm technology nodes. Tobias Gemmeke, Mohamed M. Sabry, Jan Stuijt, Praveen Raghavan, Francky Catthoor, David Atienza 0001 |
DATE | 1 |
| 2005 | A parametrizable low-power high-throughput turbo-decoderabstractThe paper presents a high performance turbo decoder. Its major building blocks, the maximum-a-posteriori decoder and the interleaver, are optimized from architecture to layout level to achieve high-throughput at low-power. This includes a novel architecture for parallel interleaving, that sustains any interleaving scheme. Moreover, the key features of the major building blocks are analyzed and modeled for quick design space exploration e.g. achieving 760 Mb/s at 570 mW in a 0.13 /spl mu/m-CMOS-technology. Finally, the characterized implementations are benchmarked. Gordian Prescher, Tobias Gemmeke, Tobias G. Noll |
ICASSP (5) | 2 |