Floran de Putter

dblp:322/9484 · also Floran A. M. de Putter · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
7since 2021 · last 2025
0009-0008-9369-6532ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 2 first-author · 6 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2025 STEMS: Spatial-Temporal Mapping for Spiking Neural Networks
abstract
Spiking Neural Networks (SNNs) are event-driven bio-inspired neural networks. Recent research has trained SNN models with accuracy on par with Artificial Neural Networks (ANNs) on computer vision tasks. Due to their sparse, event-based computation, SNNs are particularly promising for energy-efficient processing, especially in event-based vision applications. However, neurons have internal states which evolve over time and keeping track of them can be costly. Hence, efficiently deploying them, especially on memory-constrained edge devices, requires careful mapping of their computation across both spatial and temporal dimensions.To address this issue, we introduce STEMS, Spatial-Temporal Mapping for SNNs. STEMS supports inter-layer mapping exploration, as well as loop tiling optimizations. By applying STEMS inter-layer exploration, we show up to 12× reduction in external memory traffic and up-to 5× reduction in energy consumption. Finally, we show that neuron states may not be needed in early SNN layers. By optimizing neuron states in one of our benchmarks, we reduced neuron states by 20x and improved energy performance by 1.4x saving without sacrificing accuracy.
Sherif Eissa, Sander Stuijk, Floran de Putter, Andrea Nardi-Dei, Federico Corradi, Henk Corporaal
IEEE Trans. Computers3
2023 PetaOps/W edge-AI $\mu$ Processors: Myth or reality?
abstract
With the rise of deep learning (DL), our world braces for artificial intelligence (AI) in every edge device, creating an urgent need for edge-AI SoCs. This SoC hardware needs to support high throughput, reliable and secure AI processing at ultra-low power (ULP), with a very short time to market. With its strong legacy in edge solutions and open processing platforms, the EU is well-positioned to become a leader in this SoC market. However, this requires AI edge processing to become at least 100 times more energy-efficient, while offering sufficient flexibility and scalability to deal with AI as a fast-moving target. Since the design space of these complex SoCs is huge, advanced tooling is needed to make their design tractable. The CONVOLVE project (currently in Inital stage) addresses these roadblocks. It takes a holistic approach with innovations at all levels of the design hierarchy. Starting with an overview of SOTA DL processing support and our project methodology, this paper presents 8 important design choices largely impacting the energy efficiency and flexibility of DL hardware. Finding good solutions is key to making smart-edge computing a reality.
Manil Dev Gomony, Floran de Putter, Anteneh Gebregiorgis, Gianna Paulin, Linyan Mei, Vikram Jain, Said Hamdioui, Victor Sanchez, Tobias Grosser, Marc Geilen, Marian Verhelst, Friedemann Zenke, Frank K. Gürkaynak, Barry de Bruin, Sander Stuijk, Simon Davidson, Sayandip De, Mounir Ghogho, Alexandra Jimborean, Sherif Eissa, Luca Benini, Dimitrios Soudris, Rajendra Bishnoi, Sam Ainsworth 0001, Federico Corradi, Ouassim Karrakchou, Tim Güneysu, Henk Corporaal
DATE2
2023 BOMP- NAS: Bayesian Optimization Mixed Precision NAS
abstract
Bayesian Optimization Mixed-Precision Neural Architecture Search (BOMP-NAS) is a method to quantizationaware neural architecture search that leverages both Bayesian optimization and mixed-precision quantization to efficiently search for compact, high performance deep neural networks. It is able to find neural networks that achieve state of the art accuracy with less search time. Compared to the closest related work, BOMP-NAS can find these neural networks in$\mathbf{6}\times$less search time.
David van Son, Floran de Putter, Sebastian Vogel, Henk Corporaal
DATE2
2023 QMTS: Fixed-point Quantization for Multiple-timescale Spiking Neural Networks
Sherif Eissa, Federico Corradi, Floran de Putter, Sander Stuijk, Henk Corporaal
ICANN (1)3
2023 BrainTTA: A 28.6 TOPS/W Compiler Programmable Transport-Triggered NN SoC
abstract
Accelerators designed for deep neural network (DNN) inference with extremely low operand widths, down to 1-bit, have become popular due to their ability to significantly reduce energy consumption during inference. This paper introduces a compiler-programmable flexible System-on-Chip (SoC) with mixed-precision support. This SoC is based on a Transport-Triggered Architecture (TTA) that facilitates efficient implementation of DNN workloads. By shifting the complexity of data movement from the hardware scheduler to the exposed-datapath compiler, DNN workloads can be implemented in an energy efficient yet flexible way. The architecture is fully supported by a compiler and can be programmed using C/C++/OpenCL. The SoC is implemented using 22nm FDX technology and achieves a peak energy efficiency of 28.6/14.9/2.47 TOPS/W for binary, ternary, and 8-bit precision, respectively, while delivering a throughput of 614/307/77 GOPS. Compared to state-of-the-art (SotA), this work achieves up to 3.3x better energy efficiency compared to other programmable solutions.
Maarten Molendijk, Floran de Putter, Manil Dev Gomony, Pekka Jääskeläinen, Henk Corporaal
ICCD2
2022 Quantization: how far should we go?
abstract
Machine learning, and specifically Deep Neural Networks (DNNs) impact all parts of daily life. Although DNNs can be large and compute intensive, requiring processing on big servers (like in the cloud), we see a move of DNNs into loT-edge based systems, adding intelligence to these systems. These systems are often energy constrained and too small for satisfying the huge DNN computation and memory demands. DNN model quantization may come to the rescue. Instead of using 32-bit floating point numbers, much smaller formats can be used, down to 1-bit binary numbers. Although this largely may solve the compute and memory problems, it comes with a huge price, model accuracy reduction. This problem spawned a lot of research into model repair methods, especially for binary neural networks. Heavy quantization triggers a lot of debate; we even see some movements of going back to higher precision using brainfloats. This paper therefore evaluates the trade-off between energy reduction through extreme quantization versus accuracy loss. This evaluation is based on ResNet-I8 with the ImageNet dataset, mapped to a fully programmable architecture with special support for 8-bit and 1-bit deep learning, the BrainTTA. We show that, after applying repair methods, the use of extremely quantized DNNs makes sense. They have superior energy efficiency compared to DNNs based on 8-bit precision of weights and data, while only having a slightly lower accuracy. There is still an accuracy gap, requiring further research, but results are promising. A side effect of the much lower energy requirements of BNNs is that external DRAM becomes more dominant. This certainly requires further attention.
Floran de Putter, Henk Corporaal
DSD1
2022 CELR: Cloud Enhanced Local Reconstruction from low-dose sparse Scanning Electron Microscopy images
abstract
Current Scanning Electron Microscopy (SEM) acquisition techniques are far too slow to capture large volumes in a feasible time. One solution is to use low-dose and sparse imaging. By computationally denoising and inpainting an image with acceptable quality can be approximated. This approach, however, requires significant compute resources. Therefore, this paper proposes CELR, a framework, that hides the computationally expensive workload of reconstructing low-dose sparse SEM images, such that (delayed) live reconstruction is possible. Live reconstruction is possible by using Convolutional Neural Networks (CNNs) that approximate a classical reconstruction algorithm like GOAL. The reconstruction by CNNs is done locally, while recurring training of CNNs is done in the cloud. Moreover, training labels are generated by GOAL in the cloud. Next to the framework, this paper evaluates and optimizes the CNN reconstruction throughput by employing Nvidia's TensorRT. This paper also touches upon open research questions about on-the-fly CNN training. The combination of CELR and TensorRT enables large volume acquisitions with a dwell-time of$\mathbf{1}\mu s$and 10% pixel coverage to be reconstructed on a single GPU.
Floran de Putter, Maurice Peemen, Pavel Potocek, Remco Schoenmakers, Henk Corporaal
DSD1