EDBT 2026 Demo / reviewers in the wild / expert
Nathan Laubeuf
dblp:255/5595
· DBLP profile ↗
7ranked-venue papers
2as first author
7since 2021 · last 2026
0000-0002-1592-755XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Evaluating Cross-Architecture Performance Modeling of Distributed ML Workloads Using StableHLOabstractPredicting the performance of large-scale distributed machine learning (ML) workloads across multiple accelerator architectures remains a central challenge in ML system design. Existing GPU and TPU focused simulators are typically architecture-specific, while distributed training simulators rely on workload-specific analytical models or costly post-execution traces, limiting portability and cross-platform comparison. This work evaluates whether MLIR’s StableHLO dialect can serve as a unified workload representation for cross-architecture and crossfidelity performance modeling of distributed ML workloads. The study establishes a StableHLO-based simulation methodology that maps a single workload representation onto multiple performance models, spanning analytical, profiling-based, and simulator-driven predictors. Using this methodology, workloads are evaluated across GPUs and TPUs without requiring access to scaled-out physical systems, enabling systematic comparison across modeling fidelities. An empirical evaluation covering distributed GEMM kernels, ResNet, and large language model training workloads demonstrates that StableHLO preserves relative performance trends across architectures and fidelities, while exposing accuracy trade-offs and simulator limitations. Across evaluated scenarios, prediction errors remain within practical bounds for early-stage design exploration, and the methodology reveals fidelity-dependent limitations in existing GPU simulators. These results indicate that StableHLO provides a viable foundation for unified, distributed ML performance modeling across accelerator architectures and simulators, supporting reusable evaluation workflows and crossvalidation throughout the ML system design process. Jonas Svedas, Nathan Laubeuf, Ryan Harvey, Changhai Man, Abubakr Nada, Tushar Krishna, James Myers, Debjyoti Bhattacharjee |
ISPASS | 2 |
| 2024 | Adaptive Block-Scaled GeMMs on Vector Processors for DNN Training at the EdgeabstractReduced precision datatypes have become essential to the efficient training and deployment of Deep Neural Networks (DNNs). A recent development in the field has been the emergence of block-scaled datatypes: tensor representation formats derived from floating-point, that share a common exponent across multiple elements. While these formats are being broadly adopted and optimised for by DNN-specific inference accelerators, the potential benefits for training workloads on general-purpose (GP) vector processors has yet to be thoroughly explored. This work proposes a benchmarked implementation of block-scaled general matrix multiplications (GeMM) for DNN training at the edge using commercially available vector instruction sets (ARM SVE). Using this implementation, we highlight an accuracy-speed trade-off involving the shape of shared exponent blocks - vectors or squares. We exploit this result to optimize the training of fully connected networks by dynamically adapting the shared exponent block shapes during training. This strategy yields on average around$1.95 \times$faster training with$2\times$lower memory footprint compared to standard IEEE 32-bit floating point (FP32), while achieving similar accuracy. Nitish Satya Murthy, Nathan Laubeuf, Debjyoti Bhattacharjee, Francky Catthoor, Marian Verhelst |
VLSI-SoC | 2 |
| 2022 | Learn to Learn on Chip: Hardware-aware Meta-learning for Quantized Few-shot Learning at the EdgeabstractRecent years have seen a growing trend of deploying deep neural network-based applications on edge devices. Many of these applications, such as biometric identification, activity tracking, user preference learning, etc., require fine-tuning of the trained networks for user personalization. One way to prepare these models to handle new, unseen tasks, is to pre-train them on a distribution of known tasks. This observation has led to increasing research into meta-learning based few-shot learning techniques. However, basic meta-learning approaches do not account for the limited memory and computational resources during on-chip training. We propose a modified meta-learning algorithm that enables quantized fine-tuning to optimally condition the models for on-chip few shot learning. The modification involves the inclusion of target hardware constraints upfront in the meta-learning process. Block floating point datatypes with low precision mantissa bits are utilized in the forward and backward passes, to allow hardware-friendly adaptation. Experiments show that our algorithm provides better initializations than conventional algorithms, more suitable for efficient quantized fine-tuning. This allows the few-shot learner to achieve better convergence, in terms of accuracy and speed. Extensive experiments are also performed to analyze the impact of initialization on quantized fine-tuning and further corroborate the benefits of our method. Nitish Satya Murthy, Peter Vrancx, Nathan Laubeuf, Peter Debacker, Francky Catthoor, Marian Verhelst |
SEC | 3 |
| 2022 | Analog Compute in Memory and Breaking Digital Number RepresentationsabstractAnalog in Memory Computing (AiMC) is a promising solution to achieve energy efficient deployment of Deep Neural Networks (DNNs). However, to this day, the accuracies of DNNs deployed using AIMC do not compare favorably with digital alternatives. In this work, we illustrate that these lacking performance can be partly attributed to mismatched arithmetic assumption. We show that correcting those assumptions results in more robust and accurate DNNs on analog platforms. By further leveraging insights from AiMC related methods, we question modern neural network quantization techniques, revealing, in the process, alternative multiplication-less modes of executions for DNNs. Nathan Laubeuf |
VLSI-SoC | 1 |
| 2022 | Dynamic Quantization Range Control for Analog-in-Memory Neural Networks AccelerationabstractAnalog in Memory Computing (AiMC) based neural network acceleration is a promising solution to increase the energy efficiency of deep neural networks deployment. However, the quantization requirements of these analog systems are not compatible with state-of-the-art neural network quantization techniques. Indeed, while the quantization of the weights and activations is considered by modern deep neural network quantization techniques, AiMC accelerators also impose the quantization of each Matrix Vector Multiplication (MVM) result. In most demonstrated AiMC implementations, the quantization range of MVM results is considered a fixed parameter of the accelerator. This work demonstrates that dynamic control over this quantization range is possible but also desirable for analog neural networks acceleration. An AiMC compatible quantization flow coupled with a hardware aware quantization range driving technique is introduced to fully exploit these dynamic ranges. Using CIFAR-10 and ImageNet as benchmarks, the proposed solution results in networks that are both more accurate and more robust to the inherent vulnerability of analog circuits than fixed quantization range based approaches. Nathan Laubeuf, Jonas Doevenspeck, Ioannis A. Papistas, Michele Caselli, Stefan Cosemans, Peter Vrancx, Debjyoti Bhattacharjee, Arindam Mallik, Peter Debacker, Diederik Verkest, Francky Catthoor, Rudy Lauwereins |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2021 | Noise tolerant ternary weight deep neural networks for analog in-memory inferenceabstractAnalog in memory computing (AiMC) is a promising hardware solution to efficiently perform inference with deep neural networks (DNNs). Similar to digital DNN accelerators, AiMC systems benefit from aggressively quantized DNNs. In addition, AiMC systems also suffer from noise on activations and weights. Training strategies to condition DNNs against weight noise can increase the efficiency of AiMC systems by enabling the use of more compact but more noisy weight memory devices. In this work, we utilize noise-aware training and introduce gradual noise training and network width scaling to increase the tolerance of DNNs against weight noise. Our results show that noise-aware training and gradual noise training drastically lowers the impact of weight noise without changing the network size. By utilizing network width scaling, the weight noise tolerance is increased even more with the penalty of more network parameters. Jonas Doevenspeck, Peter Vrancx, Nathan Laubeuf, Arindam Mallik, Peter Debacker, Diederik Verkest, Rudy Lauwereins, Wim Dehaene |
IJCNN | 3 |
| 2021 | Design-Technology Space Exploration for Energy Efficient AiMC-Based Inference AccelerationabstractExtremely energy-efficient convolutional neural network inference (CNN) is recently enabled by analog in-memory compute (AiMC). The integration of AiMC in a primarily digital inference system brings new challenges ranging from device specifications to defining novel system architectures. A novel framework to evaluate the impact of AiMC array at the system level is presented. The framework is used to model a SRAM-based 1024 × 512 prototype AiMC array, capable of energy efficiency of upto 675 TMACs/W. The proposed framework allows modelling different compute cells, array dimensions, operating voltage, and activation buffer energy models which can be used to determine the overall energy efficiency for various CNN workloads. Debjyoti Bhattacharjee, Nathan Laubeuf, Stefan Cosemans, Ioannis A. Papistas, Arindam Mallik, Peter Debacker, Myung Hee Na, Diederik Verkest |
ISCAS | 2 |