EDBT 2026 Demo / reviewers in the wild / expert
Pierpaolo Morì
dblp:327/1741
· DBLP profile ↗
9ranked-venue papers
4as first author
9since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 3 first-author · 7 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | LA-MTL: Latency-Aware Automated Multi-Task LearningabstractMulti-Task Learning (MTL) aims to unify a variety of tasks into a single network for improved training and inference efficiency. This is particularly attractive for real-time applications that require simultaneous execution of multiple workloads in resource-constrained embedded environments. However, most MTL approaches focus on enhancing parameters efficiency and overall tasks metrics, often lacking explicit inference latency awareness in the optimization loop. The design space exploration should not compromise on the parameters efficiency or task accuracy objectives in order to meet latency requirements. To address this, we propose LA-MTL, an automated layer-level MTL policy search that incorporates a novel analytical latency factor (ALF). By accounting for local and global latencies during the MTL policy search, we derive solutions that balance task metrics, parameters efficiency and latency constraints. LA-MTL search on ResNet34 yields solutions with up to 50% lower latency on the Jetson AGX Orin while maintaining competitive metrics in semantic segmentation and depth estimation tasks with a +/-2 p.p., on the CityScapes dataset. Additionally, we achieve a superior parameters efficiency, surpassing the state-of-theart MTL parameters reduction by over 20 p.p. Experiments on benchmark datasets (CityScapes, NYUv2) demonstrate the effectiveness of our approach across various backbones including ResNet34, MobileNetV2, and MobileOne in its expanded form. Code is available at https://github.com/shamvbs/LA-MTL.1 Shambhavi Balamuthu Sampath, Sami Sawani, Moritz Thoma, Lukas Frickenstein, Pierpaolo Morì, Nael Fasfous, Manoj Rohit Vemparala, Alexander Frickenstein, Ulf Schlichtmann, Claudio Passerone, Walter Stechele |
DAC | 5 |
| 2025 | SuperFast: Fast Supernet Training Using Initial KnowledgeabstractOnce-for-all based neural architecture search (NAS) proposes to train a supernet once and extract specialized subnets from it for efficient deployment. This decoupling between training and search enables easy multi-target deployment without retraining. Nevertheless, the initial training cost has remained extremely high, with SOTA approaches like ElasticViT and NASViT taking more than 72 and 83 GPU days respectively. While other approaches have tried to accelerate the training by warming up the largest model in the search space, we argue that this is suboptimal, and knowledge is easier scaled upward than downward. Hence, we propose SuperFast, a simple, plug and play workflow, that (I.) pretrains a subnet of the supernet search space, and (II.) distributes its knowledge within the supernet before the training. SuperFast offers a substantial acceleration in the supernet training, resulting in a significantly better accuracy vs. training-cost trade-off. Using SuperFast on both ElasticViT and NASViT supernets achieves the baseline’s accuracy $1.4 \times$ and $\mathbf{1. 8} \times$ faster on the ImageNet dataset. Moreover, for a given time budget, SuperFast improves accuracy vs. latency trade-offs for subnets, gaining 4.0 p.p. for the $20-50 \mathrm{~ms}$ range on Pixel 6. Code available in https://github.com/MoritzTho/SuperFast. Moritz Thoma, Emad Aghajanzadeh, Shambhavi Balamuthu Sampath, Pierpaolo Morì, Nael Fasfous, Alexander Frickenstein, Manoj Rohit Vemparala, Daniel Mueller-Gritschneder, Ulf Schlichtmann |
DAC | 4 |
| 2025 | HiFi-SAGE: High Fidelity GraphSAGE-Based Latency Estimators for DNN OptimizationabstractAs deep neural networks (DNNs) are increasingly deployed on resource-constrained edge devices, optimizing and compressing them for real-time performance becomes crucial. Traditional hardware-aware DNN search methods often rely on inaccurate proxy metrics, expensive latency lookup tables, or slow hardware-in-the-Iloop (HIL) evaluations. To address this, quasi-generalized latency estimators, typically meta-learning-based, were proposed to replace HIL evaluations and accelerate the search. These come with a one-time data collection and training cost and can adapt to new hardware with few measurements. However, they still have some drawbacks: (1) They increase complexity by trying to generalize across a range of diverse hardware types; (2) They depend on handcrafted hardware descriptors, which may fail to capture hardware characteristics; (3) They often perform poorly on new, unseen hardware that significantly differs from their initial training set. To overcome these challenges, this paper turns to the more straightforward platform-specific estimators that do not require hardware descriptors and can be easily trained on any hardware. We introduce HiFi-SAGE, a high fidelity GraphSAGE-based platform-specific latency estimator. When trained from scratch on only 100 latency measurements, our novel dual-head estimator design surpasses the state-of-the-art (SoTA) on the 10% error bound metric by up to 17.4 p.p. while achieving an impressive fidelity score of 99% on the diverse LatBench dataset. We demonstrate that applying HiFi-SAGE to a genetic algorithm-based DNN compression search, achieved a Pareto front comparable to real HIL feedback with a mean absolute percentage error (MAPE) of 2.54%, 2.48%, and 4.16%, for InceptionV3, DenseNet169, and ResNet50 respectively. Compared to existing platform-specific works, the lower number of latency measurements and higher fidelity scores positions HiFi-SAGE as an attractive alternative to replace expensive HIL setups. Code is available at: https://github.com/shamvbs/HiFi-SAGE * Shambhavi Balamuthu Sampath, Leon Hecht, Moritz Thoma, Lukas Frickenstein, Pierpaolo Morì, Nael Fasfous, Manoj Rohit Vemparala, Alexander Frickenstein, Walter Stechele, Daniel Mueller-Gritschneder, Claudio Passerone |
DATE | 5 |
| 2025 | HotShot: A Loss-Guided Data Augmentation and Curriculum Learning Technique for the Task of Semantic SegmentationabstractSemantic segmentation is an important computer vision task that requires costly pixel-level annotations to train deep neural networks (DNNs) for. Especially for applications like autonomous driving, precise pixel-level understanding of scenes is a decisive factor between success and failure of the application. It follows that every labeled sample of an existing dataset is highly valuable and should be optimally used during training to maximize its value. This is achieved using (1) augmentation of the same labeled sample to help the model learn it in different ways, and (2) curriculum learning to introduce training samples to the model in an strategic order to ease the learning process. In this work, we present HotShot, a loss-guided cropping technique that assesses the DNN's prediction capability during the training to derive probability scores of potential cropping regions. This effectively combines augmentation and curriculum learning in one technique, where a single sample is cropped (augmentation) in regions selected based on the DNN's loss throughout the training (curriculum learning). For UperNet using a ConvNeXt-tiny backbone and DeepLabV3+ architecture using a ResNet-50 backbone, applying HotShot provides a +0.41 p.p. and +0.43 p.p. mIoU improvement over randomly cropping regions on the CityScapes and BDD100K datasets respectively. More interestingly, the analysis shows HotShot primarily boosts the classes that are most challenging for the model. For example, the rider and motorcycle classes on the BDD100K dataset improve by 163% and 129% using DeepLabV3+ with a ResNet-50 backbone. HotShot achieves improved mIoU in almost all cases and normalizes imbalances in learning challenging classes in datasets. Lukas Frickenstein, Moritz Thoma, Pierpaolo Morì, Shambhavi Balamuthu Sampath, Nael Fasfous, Manoj Rohit Vemparala, Alexander Frickenstein, Christian Unger, Claudio Passerone, Walter Stechele |
IV | 3 |
| 2025 | ZuSE-KI-Mobil: AI Chip Design Platform for Automotive and Industrial Applications
Shaown Mojumder, Simon Friedrich, Emil Matús, Matthias Lüders, Martin Friedrich, Oliver Renke, Holger Blume, Markus Kock, Gregor Schewior, Darius Grantz, Jens Benndorf, Julian Höfer, Patrick Schmidt 0003, Jürgen Becker 0001, Nael Fasfous, Pierpaolo Morì, Hans-Jörg Vögel, Samira Ahmadifarsani, Leonidas Kontopoulos, Ulf Schlichtmann, Yun-Jin Li, Gerhard P. Fettweis |
IEEE Trans. Very Large Scale Integr. Syst. | 16 |
| 2024 | MATAR: Multi-Quantization-Aware Training for Accurate and Fast Hardware RetargetingabstractQuantization of deep neural networks (DNNs) reduces their memory footprint and simplifies their hardware arithmetic logic, enabling efficient inference on edge devices. Different hardware targets can support different forms of quantization, e.g. full 8-bit, or 8/4/2-bit mixed-precision combinations, or fully-flexible bit-serial solutions. This makes standard quantization-aware training (QAT) of a DNN for different targets challenging, as there needs to be careful consideration of the supported quantization-levels of each target at training time. In this paper, we propose a generalized QAT solution that results in a DNN which can be retargeted to different hardware, without any retraining or prior knowledge of the hardware's supported quantization policy. First, we present the novel training scheme which makes the model aware of multiple quantization strategies. Then we demonstrate the retargeting capabilities of the resulting DNN by using a genetic algorithm to search for layer-wise, mixed-precision solutions that maximize performance and/or accuracy on the hardware target, without the need of fine-tuning. By making the DNN agnostic of the final hardware target, our method allows DNNs to be distributed to many users on different hardware platforms, without the need for sharing the training loop or dataset of the DNN developers, nor detailing the hardware capabilities ahead of time by the end-users of the efficient quantized solution. Models trained with our approach can generalize on multiple quantization policies with minimal accuracy degradation compared to target-specific quantization counterparts. Pierpaolo Morì, Moritz Thoma, Lukas Frickenstein, Shambhavi Balamuthu Sampath, Nael Fasfous, Manoj Rohit Vemparala, Alexander Frickenstein, Walter Stechele, Daniel Mueller-Gritschneder, Claudio Passerone |
DATE | 1 |
| 2024 | Wino Vidi Vici: Conquering Numerical Instability of 8-bit Winograd Convolution for Accurate Inference Acceleration on EdgeabstractWinograd-based convolution can reduce the total number of operations needed for convolutional neural network (CNN) inference on edge devices. Most edge hardware accelerators use low-precision, 8-bit integer arithmetic units to improve energy efficiency and latency. This makes CNN quantization a critical step before deploying the model on such an edge device. To extract the benefits of fast Winograd-based convolution and efficient integer quantization, the two approaches must be combined. Research has shown that the transform required to execute convolutions in the Winograd domain results in numerical instability and severe accuracy degradation when combined with quantization, making the two techniques incompatible on edge hardware. This paper proposes a novel training scheme to achieve efficient Winograd-accelerated, quantized CNNs. 8-bit quantization is applied to all the intermediate results of the Winograd convolution without sacrificing task-related accuracy. This is achieved by introducing clipping factors in the intermediate quantization stages as well as using the complex numerical system to improve the transform. We achieve 2.8× and 2.1× reduction in MAC operations on ResNet-20-CIFAR-10 and ResNet-18-ImageNet, respectively, with no accuracy degradation. Pierpaolo Morì, Lukas Frickenstein, Shambhavi Balamuthu Sampath, Moritz Thoma, Nael Fasfous, Manoj Rohit Vemparala, Alexander Frickenstein, Christian Unger, Walter Stechele, Daniel Mueller-Gritschneder, Claudio Passerone |
WACV | 1 |
| 2023 | WinoTrain: Winograd-Aware Training for Accurate Full 8-bit Convolution AccelerationabstractEfficient inference is critical in realizing a low-power, real-time implementation of convolutional neural networks (CNNs) on compute and memory-constrained embedded platforms. Using quantization techniques and fast convolutional algorithms like Winograd, CNN inference can achieve benefits in latency and in energy consumption. Performing Winograd convolution involves (1) transforming the weights and activations to the Winograd domain, (2) performing element-wise multiplication on the transformed tensors, and (3) transforming the results back to the conventional spatial domain. Combining Winograd with quantization of all its steps results in severe accuracy degradation due to numerical instability. In this paper we propose a simple quantization-aware training technique, which quantizes all three steps of the Winograd convolution, while using a minimal number of scaling factors. Additionally, we propose an FPGA accelerator employing tiling and unrolling methods to highlight the performance benefits of using the full 8-bit quantized Winograd algorithm. We achieve 2× reduction in inference time compared to standard convolution on ResNet-18 for the ImageNet dataset, while improving the Top-1 accuracy by 55.7 p.p. compared to a standard post-training quantized Winograd variant of the network. Pierpaolo Morì, Shambhavi Balamuthu Sampath, Lukas Frickenstein, Manoj Rohit Vemparala, Nael Fasfous, Alexander Frickenstein, Walter Stechele, Claudio Passerone |
DAC | 1 |
| 2022 | Accelerating and pruning CNNs for semantic segmentation on FPGAabstractSemantic segmentation is one of the popular tasks in computer vision, providing pixel-wise annotations for scene understanding. However, segmentation-based convolutional neural networks require tremendous computational power. In this work, a fully-pipelined hardware accelerator with support for dilated convolution is introduced, which cuts down the redundant zero multiplications. Furthermore, we propose a genetic algorithm based automated channel pruning technique to jointly optimize computational complexity and model accuracy. Finally, hardware heuristics and an accurate model of the custom accelerator design enable a hardware-aware pruning framework. We achieve 2.44X lower latency with minimal degradation in semantic prediction quality (−1.98 pp lower mean intersection over union) compared to the baseline DeepLabV3+ model, evaluated on an Arria-10 FPGA. The binary files of the FPGA design, baseline and pruned models can be found in github.com/pierpaolomori/SemanticSegmentationFPGA Pierpaolo Morì, Manoj Rohit Vemparala, Nael Fasfous, Saptarshi Mitra, Sreetama Sarkar, Alexander Frickenstein, Lukas Frickenstein, Domenik Helms, Naveen Shankar Nagaraja, Walter Stechele, Claudio Passerone |
DAC | 1 |