Thomas Boesch

dblp:179/1652 · DBLP profile ↗
← Back
7ranked-venue papers
0as first author
3since 2021 · last 2026
0009-0003-7927-8047ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CIM-FLEX: An Integer-Only Flexible Periphery for Distributed Compute-In-Memory Architectures
abstract
Distributed compute-in-memory (CIM) architectures are emerging as a path towards energy-efficient and highly parallel edge-class deep neural network (DNN) inference. However, the lack of a flexible, unified digital periphery limits the ability of distributed CIM systems to support integer-quantized DNN inference schemes, activation functions, and scalable tile-to-tile dataflow. We present CIM-FLEX, an integer-only periphery architecture that enables each CIM tile to autonomously perform partial-sum rescaling, asymmetric activation processing, and lightweight programmable activation functions. CIM-FLEX also supports mixed-precision integer multiply-and-accumulate (MAC) by decomposing higher-precision operations into uniform-precision partial MACs, without modifying existing CIM primitives. CIM-FLEX enables flexible tile-to-tile dataflow across distributed CIM tiles, exposing a tradeoff between parallel output generation and energy efficiency, where deeper accumulation chains reduce concurrent periphery utilization and improve overall efficiency. Numerical studies and TFLite-based LUT activation experiments validate a high signal-to-noise ratio (∼30 dB) for programmable activation functions, and a 14 nm CMOS implementation shows a 0.012 mm2 area for the CIM-FLEX.
Irem Sanli, Michele Rossi, Elena Ferro, Andreas Burg, Surinder Pal Singh, Thomas Boesch, Irem Boybat
ISLPED7
2025 Multi-Mode Borderguard Controllers for Efficient On-Chip Communication in Heterogeneous Digital/Analog Neural Processing Units
abstract
Driven by the growing demand for data-intensive parallel computation, particularly for Matrix-Vector Multiplications (MVMs), and the pursuit of high energy efficiency, Analog In-Memory Computing (AIMC) has garnered significant attention. AIMC addresses the data movement bottleneck by performing MVMs directly within memory, significantly reducing latency and enhancing energy efficiency. Integrating AIMC with digital units for non-MVM operations yields heterogeneous Neural Processing Units (NPUs) that can be combined in a tiled architecture to deliver promising solutions for end-to-end AI inference. Besides powerful heterogeneous NPUs, an efficient on-chip communication infrastructure is also pivotal for inter-node data transmission and efficient AI model execution. This paper introduces the Borderguard Controller (BG-CTRL), a multi-mode, path-through routing controller designed to support three distinct operating modes-time-scheduling, data-driven, and time-sliced data-driven (TSDD)-each offering varying levels of routing flexibility and energy efficiency depending on the data flow patterns and AI model complexity. To demonstrate the design, BG-CTRLs are integrated into a 9-node system of heterogeneous NPUs, arranged in a 3x3 grid and connected using a 2D mesh topology. The system is synthesized using STM 28nm FD-SOI technology. Experimental results show that the BG-CTRL cluster achieves an aggregate throughput of 983 Gb/s, with an energy efficiency of up to 0.41 pJ/B/hop at 0.64 GHz, and a minimal area overhead of 204 kGE.
Hong Pang, Carmine Cappetta, Riccardo Massa, Athanasios Vasilopoulos, Elena Ferro, Gamze Islamoglu, Angelo Garofalo, Francesco Conti 0001, Luca Benini, Irem Boybat, Thomas Boesch
DATE11
2025 NIMA: Near In-Memory High-Precision Accumulation Unit for Heterogeneous Analog/Digital Deep Learning Acceleration
abstract
Analog In-Memory Computing (AIMC) crossbars often face size limitations that hinder mapping entire neural network layers onto a single AIMC tile. To overcome this, tall layers are typically distributed across multiple tiles or within a single tile by multiplexing different sets of weights at various time intervals. However, this generates partial vector-matrix multiplication (VMM) results that need to be accumulated to produce the final output, underscoring the need for integrated accumulation capabilities in AIMC systems. In this work, we propose a Near In-Memory High-Precision Accumulation Unit (NIMA) with built-in internal and external tile accumulation functionality, positioned at the periphery of the AIMC crossbar. This unit leverages the affine correction capabilities of existing systems and ensures high precision during partial VMM accumulation. We have physically implemented NIMA in a 14nm CMOS technology, providing comprehensive performance and area evaluations and comparisons with prior art. We investigate and compare different layer mappings and accumulation schemes supported by the proposed unit, evaluating both latency and Mean Absolute Error (MAE). We further assess the precision of the proposed unit on large layers of ResNet9 and ResNet32 for image classification on CIFAR10/CIFAR100 datasets.
Irem Sanli, Elena Ferro, Athanasios Vasilopoulos, Thomas Boesch, Abu Sebastian, Irem Boybat
ISCAS4
2020 Runtime Design Space Exploration and Mapping of DCNNs for the Ultra-Low-Power Orlando SoC
abstract
Recent trends in deep convolutional neural networks (DCNNs) impose hardware accelerators as a viable solution for computer vision and speech recognition. The Orlando SoC architecture from STMicroelectronics targets exactly this class of problems by integrating hardware-accelerated convolutional blocks together with DSPs and on-chip memory resources to enable energy-efficient designs of DCNNs. The main advantage of the Orlando platform is to have runtime configurable convolutional accelerators that can adapt to different DCNN workloads. This opens new challenges for mapping the computation to the accelerators and for managing the on-chip resources efficiently. In this work, we propose a runtime design space exploration and mapping methodology for runtime resource management in terms of on-chip memory, convolutional accelerators, and external bandwidth. Experimental results are reported in terms of power/performance scalability, Pareto analysis, mapping adaptivity, and accelerator utilization for the Orlando architecture mapping the VGG-16, Tiny-Yolo(v2), and MobileNet topologies.
Ahmet Erdem, Cristina Silvano, Thomas Boesch, Andrea C. Ornstein, Surinder Pal Singh, Giuseppe Desoli
ACM Trans. Archit. Code Optim.3
2018 Design Space Exploration for Orlando Ultra Low-Power Convolutional Neural Network SoC
abstract
With the recent advances in machine learning, Deep Convolutional Neural Networks (DCNNs) represent state-of-the-art solutions especially in image and speech recognition and classification. The most important enabler factor of deep learning consists of the massive computing power offered by programmable GPUs for training DCNNs on large amounts of data. Even the complexity of DCNN deployment scenarios, where trained models are used for inference, have started to require powerful computing systems. Especially in the embedded systems domain, the computational requirements along with ultra low-power and memory constraints exacerbate the situation even further. The STM Orlando ultra low-power processor architecture with convolutional neural network acceleration targets exactly this class of problems. Orlando SoC integrates HW-accelerated blocks together with DSPs and on-chip memory resources to enable energy-efficient convolutions for future generations of DCNNs. Although Orlando platform provides flexibility with programmable DSPs, the large variety of DCNN applications make it challenging the design space exploration of the next generations of Orlando architecture. Many hardware design parameters affect the performance and energy-efficiency of the Orlando SoC. Given the huge size of the design space, design space exploration (DSE) and cost-performance tradeoff analysis is needed to select the best set of parameters for the target DCNN application. In this work, we present the exploration results and the tradeoff analysis carried out for the Orlando architecture on the vee - 16 case study.
Ahmet Erdem, Cristina Silvano, Thomas Boesch, Andrea C. Ornstein, Surinder Pal Singh, Giuseppe Desoli
ASAP3
2016 The Orlando Project: A 28 nm FD-SOI Low Memory Embedded Neural Network ASIC
Giuseppe Desoli, Valeria Tomaselli, Emanuele Plebani, Giulio Urlini, Danilo Pau, Viviana D'Alto, Tommaso Majo, Fabio De Ambroggi, Thomas Boesch, Surinder Pal Singh, Elio Guidetti, Nitin Chawla
ACIVS9
2016 Frame buffer-less stream processor for accurate real-time interest point detection
Gian Domenico Licciardo, Thomas Boesch, Danilo Pau, Luigi Di Benedetto
Integr.2