Lennart Bamberg

dblp:194/0911 · DBLP profile ↗
← Back
15ranked-venue papers
8as first author
5since 2021 · last 2023
0000-0003-4673-8310ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 15 · 8 first-author · 5 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author
YearPublicationVenuePosition
2023 Synapse Compression for Event-Based Convolutional-Neural-Network Accelerators
abstract
Manufacturing-viable neuromorphic chips require novel compute architectures to achieve the massively parallel and efficient information processing the brain supports so effortlessly. The most promising architectures for that are spiking/event-based, which enables massive parallelism at low complexity. However, the large memory requirements for synaptic connectivity are a showstopper for the execution of modern convolutional neural networks (CNNs) on massively parallel, event-based architectures. The present work overcomes this roadblock by contributing a lightweight hardware scheme to compress the synaptic memory requirements by several thousand times—enabling the execution of complex CNNs on a single chip of small form factor. A silicon implementation in a 12-nm technology shows that the technique achieves a total memory-footprint reduction of up to 374× compared to the best previously published technique at a negligible area overhead.
Lennart Bamberg, Arash Pourtaherian, Luc Waeijen, Anupam Chahar, Orlando Moreira
IEEE Trans. Parallel Distributed Syst.1
2022 ARTS: An adaptive regularization training schedule for activation sparsity exploration
abstract
Brain-inspired event-based processors have attracted considerable attention for edge deployment because of their ability to efficiently process Convolutional Neural Networks (CNNs) by exploiting sparsity. On such processors, one critical feature is that the speed and energy consumption of CNN inference are approximately proportional to the number of non-zero values in the activation maps. Thus, to achieve top performance, an efficient training algorithm is required to largely suppress the activations in CNNs. We propose a novel training method, called Adaptive-Regularization Training Schedule (ARTS), which dramatically decreases the non-zero activations in a model by adaptively altering the regularization coefficient through training. We evaluate our method across an extensive range of computer vision applications, including image classification, object recognition, depth estimation, and semantic segmentation. The results show that our technique can achieve 1.41 × to 6.00 × more activation suppression on top of ReLU activation across various networks and applications, and outperforms the state-of-the-art methods in terms of training time, activation suppression gains, and accuracy. A case study for a commercially-available event-based processor, Neuronflow, shows that the activation suppression achieved by ARTS effectively reduces CNN inference latency by up to 8.4 × and energy consumption by up to 14.1 ×.
Zeqi Zhu, Arash Pourtaherian, Luc Waeijen, Lennart Bamberg, Egor Bondarev, Orlando Moreira
DSD4
2021 Bridging the Frequency Gap in Heterogeneous 3D SoCs through Technology-Specific NoC Router Architectures
abstract
In heterogeneous 3D System-on-Chips (SoCs), NoCs with uniform properties suffer one major limitation; the clock frequency of routers varies due to different manufacturing technologies. For example, digital nodes allow for a higher clock frequency of routers than mixed-signal nodes. This large frequency gap is commonly tackled by complex and expensive pseudo-mesochronous or asynchronous router architectures. Here, a more efficient approach is chosen to bridge the frequency gap. We propose to use a heterogeneous network architecture. We show that reducing the number of VCs allows to bridge a frequency gap of up to 2x. We achieve a system-level latency improvement of up to 47% for uniform random traffic and up to 59% for PARSEC benchmarks, a maximum throughput increase of 50%, up to 68% reduced area and 38% reduced power in an exemplary setting combining 15-nm digital and 30-nm mixed-signal nodes and comparing against a homogeneous synchronous network architecture. Versus asynchronous and pseudo-mesochronous router architectures, the proposed optimization consistently performs better in area, in power and the average flit latency improvement can be larger than 51%.
Jan Moritz Joseph, Lennart Bamberg, Geonhwa Jeong, Ruei-Ting Chien, Rainer Leupers, Alberto García Ortiz, Tushar Krishna, Thilo Pionteck
ASP-DAC2
2021 NEWROMAP: mapping CNNs to NoC-interconnected self-contained data-flow accelerators for edge-AI
abstract
Conventional AI accelerators are limited by von-Neumann bottlenecks for edge workloads. Domain-specific accelerators (often neuromorphic) solve this by applying near/in-memory computing, NoC-interconnected massive-multicore setups, and data-flow computation. This requires an effective mapping of neural networks (i.e, an assignment of network layers to cores) to balance resources/memory, computation, and NoC traffic. Here, we introduce a mapping called Snake for the predominant convolutional neural networks (CNNs). It utilizes the feed-forward nature of CNNs by folding layers to spatially adjacent cores. We achieve a total NoC bandwidth improvement of up to 3.8X for MobileNet and ResNet vs. random mappings. Furthermore, NEWROMAP is proposed that continues to optimize Snake mapping through a meta-heuristic; it also simulates the NoC traffic and can work with TensorFlow models. The communication is further optimized with up to 22.52% latency improvement vs. pure snake mapping shown in simulations.
Jan Moritz Joseph, Murat Sezgin Baloglu, Rainer Leupers, Lennart Bamberg
NOCS5
2021 High-Performance Logic-on-Memory Monolithic 3-D IC Designs for Arm Cortex-A Processors
abstract
Monolithic 3-D IC (M3-D) is a promising solution to improve the performance and energy-efficiency of modern processors. But, designers are faced with challenges in design tools and methodologies, especially for power and thermal verifications. We developed a new physical design flow that optimally places and routes cache modules in one tier and logic gates in the other. Our tool also builds high-quality clock and power delivery networks targeting logic-on-memory M3-D designs. Finally, we developed a sign-off analysis tool flow to evaluate power, performance, area (PPA), thermal, and voltage-drop quality for given M3-D designs. Using our complete register transfer level (RTL)-to-Graphic Design System (GDS) tool flow, we designed commercial quality 2-D and M3-D implementation of Arm Cortex-A7 and Cortex-A53 processors in a commercial 28-nm technology. Experimental results show that our 3-D processors offer 20% (A7) and 21% (A53) performance gain, compared with their 2-D commercial counterparts. The voltage-drop degradation of our 3-D Cortex-A7 and Cortex-A53 processors is less than 3% of the supply voltage, while temperature increase is 10.71 °C and 13.04 °C, respectively.
Lingjun Zhu, Lennart Bamberg, Sai Pentapati, Kyungwook Chang, Francky Catthoor, Dragomir Milojevic, Manu Perumkunnil Komalan, Brian Cline, Saurabh Sinha 0001, Alberto García Ortiz, Sung Kyu Lim
IEEE Trans. Very Large Scale Integr. Syst.2
2020 Macro-3D: A Physical Design Methodology for Face-to-Face-Stacked Heterogeneous 3D ICs
abstract
Memory-on-logic and sensor-on-logic face-to-face stacking are emerging design approaches that promise a significant increase in the performance of modern systems-on-chip at reasonable costs. In this work, a netlist-to-layout design flow for such heterogeneous 3D systems is proposed. The proposed technique overcomes the severe limitations of existing 3D physical design methodologies. A RISC-V-based multi-core system, implemented in a commercial technology, is used as a case study to evaluate the proposed design flow. The case study is performed for modern/large and small cache sizes to show the superiority of the proposed methodology for a broad set of systems. While previous 3D design flows do not show to optimize performance against 2D baseline designs for processor systems with a significant memory area occupation, the proposed flow shows a performance and power improvement by 20.4-28.2% and 3.2-3.8%, respectively.
Lennart Bamberg, Alberto García Ortiz, Lingjun Zhu, Sai Pentapati, Da Eun Shim, Sung Kyu Lim
DATE1
2020 Misalignment-aware energy modeling of narrow buses for data encoding schemes
Amir Najafi 0001, Lennart Bamberg, Alberto García Ortiz
Integr.2
2019 System-Level Optimization of Network-on-Chips for Heterogeneous 3D System-on-Chips
abstract
For a system-level design of Networks-on-Chip for 3D heterogeneous System-on-Chip (SoC), the locations of components, routers and vertical links are determined from an application model and technology parameters. In conventional methods, the two inputs are accounted for separately; here, we define an integrated problem that considers both application model and technology parameters. We show that this problem does not allow for exact solution in reasonable time, as common for many design problems. Therefore, we contribute a heuristic by proposing design steps, which are based on separation of intralayer and interlayer communication. The advantage is that this new problem can be solved with well-known methods. We use 3D Vision SoC case studies to quantify the advantages and the practical usability of the proposed optimization approach. We achieve up to 18.8% reduced white space and up to 12.4% better network performance in comparison to conventional approaches.
Jan Moritz Joseph, Dominik Ermel, Lennart Bamberg, Alberto García Ortiz, Thilo Pionteck
ICCD3
2019 Crosstalk optimization for through-silicon vias by exploiting temporal signal misalignment
Lennart Bamberg, Jan Moritz Joseph, Thilo Pionteck, Alberto García Ortiz
Integr.1
2019 Edge effect aware low-power crosstalk avoidance technique for 3D integration
Lennart Bamberg, Amir Najafi 0001, Alberto García Ortiz
Integr.1
2019 Simulation environment for link energy estimation in networks-on-chip with virtual channels
Jan Moritz Joseph, Lennart Bamberg, Imad Hajjar, Robert Schmidt 0003, Thilo Pionteck, Alberto García Ortiz
Integr.2
2019 Coding-Based Low-Power Through-Silicon-Via Redundancy Schemes for Heterogeneous 3-D SoCs
abstract
Three-dimensional integration, employing through-silicon vias (TSVs), improves the system-on-chip (SoC) performance. However, redundancy schemes are required to cope with the relatively poor TSV manufacturing yield. Existing redundancy schemes do not exploit technological heterogeneity between the dies. Hardware costs can differ for the individual dies. This demands asymmetrical schemes with low complexity in costly mixed signal or RF dies. Furthermore, redundant TSVs are only used in the case of a defect. In the most probable case of correct manufacturing, they are unused. Another emerging technique using redundant lines is low-power coding (LPC). This paper presents a hybrid TSV redundancy technique based on coding, which can be used for LPC and for yield enhancement. Furthermore, the approach is strongly asymmetric. In case of a fault, a configuration is only required for the encoder or decoder located in the cheaper die, while in the costly die, a minimal set of XOR gates is sufficient. A case study for an existing heterogeneous SoC shows that the proposed technique decreases area overhead and power consumption compared to the best previous technique by over 69 % and 33%, respectively.
Lennart Bamberg, Alberto García Ortiz
IEEE Trans. Very Large Scale Integr. Syst.1
2018 Coding approach for low-power 3D interconnects
abstract
Through-silicon vias (TSVs) in 3D ICs show a significant power consumption, which can be reduced using coding techniques. This work presents an approach which reduces the TSV power consumption by a signal-aware bit assignment which includes inversions to exploit the MOS effect. The approach causes no overhead and results in a guaranteed reduction of the overall power consumption. An analysis of our technique shows a reduction in the TSV power consumption by up to 48 % for real correlated data streams (e.g. image sensor), and 11 % for low-power encoded random data streams.
Lennart Bamberg, Robert Schmidt 0003, Alberto García Ortiz
DAC1
2018 Edge effects on the TSV array capacitances and their performance influence
Lennart Bamberg, Amir Najafi 0001, Alberto García Ortiz
Integr.1
2017 High-Level Energy Estimation for Submicrometric TSV Arrays
abstract
The 3-D integration using through silicon vias (TSVs) is one of the most promising approaches to overcome the interconnect delay problem of current CMOS technologies. Nevertheless, the TSV energy consumption is not negligible due to the high capacitive coupling. This paper presents an abstract and yet accurate model to estimate the pattern-dependent energy consumption in arrays of TSVs; it is the first high-level model including the effects of the voltage-dependent metal-oxide-semiconductor (MOS) capacitances surrounding each TSV and a possible temporal misalignment between the input signals. We propose a regression method to estimate the dynamic size of the coupling capacitances as a function of the bit probabilities. Experimental results for real and synthetic data streams, a submicrometer 9-bit TSV array and a 65-nm technology show that the presented TSV energy model exhibits a maximum error of 5.53%, while the traditional high-level model shows errors of up to 79.77%. Furthermore, the new insights provided by our model reveal a possibility to easily boost the efficiency of existing low-power codes for TSV structures by over 10% without affecting the coding efficiency for the planar metal wires or the encoder complexity.
Lennart Bamberg, Alberto García Ortiz
IEEE Trans. Very Large Scale Integr. Syst.1