EDBT 2026 Demo / reviewers in the wild / expert
Jordi Fornt
dblp:336/6150 · also Jordi Fornt Mas
· DBLP profile ↗
7ranked-venue papers
3as first author
7since 2021 · last 2025
0000-0002-2484-9688ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 3 first-author · 6 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | GAVINA: flexible aggressive undervolting for bit-serial mixed-precision DNN accelerationabstractVoltage overscaling, or undervolting, is an enticing approximate technique in the context of energy-efficient Deep Neural Network (DNN) acceleration, given the quadratic relationship between power and voltage. Nevertheless, its very high error rate has thwarted its general adoption. Moreover, recent undervolting accelerators rely on 8-bit arithmetic and cannot compete with state-of-the-art low-precision (<8b) architectures. To overcome these issues, we propose a new technique called Guarded Aggressive underVolting (GAV), which combines the ideas of undervolting and bit-serial computation to create a flexible approximation method based on aggressively lowering the supply voltage on a select number of least significant bit combinations. Based on this idea, we implement GAVINA (GAV mIxed-Precision Accelerator), a novel architecture that supports arbitrary mixed precision and flexible undervolting, with an energy efficiency of up to 89 TOP/sW in its most aggressive configuration. By developing an error model of GAVINA, we show that GAV can achieve an energy efficiency boost of 20% via undervolting, with negligible accuracy degradation on ResNet-18. Jordi Fornt, Pau Fontova, Adrian Gras, Omar Lahyani, Martí Caro, Jaume Abella 0001, Francesc Moll, Josep Altet |
ISLPED | 1 |
| 2025 | Mix-GEMM: Extending RISC-V CPUs for Energy-Efficient Mixed-Precision DNN Inference Using Binary SegmentationabstractEfficiently computing Deep Neural Networks (DNNs) has become a primary challenge in today's computers, especially on devices targeting mobile or edge applications. Recent progress on Post-Training Quantization (PTQ) and Quantization-Aware Training (QAT) has shown that the key to high energy efficiency lies in executing deep learning models with low- (8- to 5-bit) or ultra-low-precision (4- to 2-bit). Unfortunately, current Central Processing Unit (CPU) architectures and Instruction Set Architectures (ISAs) present severe limitations on the range of data sizes supported to compute DNN kernels. In this work, we presentMix-GEMM, a hardware-software co-designed architecture that enables RISC-V processors to efficiently compute arbitrary mixed-precision DNN kernels, supporting all data size combinations from 8- to 2-bit. By applyingbinary segmentation, our architecture can scale its throughput by decreasing the data size of the operands, resulting in a flexible approach capable of leveraging state-of-the-art QAT and PTQ to achieve high energy efficiency at a very low cost. Evaluating ourMix-GEMMarchitecture in a dual-issue in-order RISC-V processor shows that we are able to boost its performance and energy efficiency by up to$44\times$and$11\times$with respect to the baseline processor, with an area overhead of only 2%. This allows our extended processor to execute state-of-the-art DNNs with significantly higher performance and energy efficiency than the standard FP32 precision, while retaining almost the same model accuracy. Jordi Fornt, Enrico Reggiani, Pau Fontova, Narcís Rodas, Alessandro Pappalardo, Osman S. Unsal, Adrián Cristal, Josep Altet, Francesc Moll, Jaume Abella 0001 |
IEEE Trans. Computers | 1 |
| 2023 | Efficient Diverse Redundant DNNs for Autonomous DrivingabstractAutomotive applications with safety requirements must adhere to specific regulations such as ISO 26262, which imposes the use of diverse redundancy for the highest integrity levels (i.e., ASIL D). While this has been often achieved by means of Dual-Core LockStep (DCLS) for microcontrollers, it remains an open challenge how to realize diverse redundancy efficiently, i.e., without full duplication and preserving performance, for DNN-based safety-related tasks, such as object detection, needing accelerators for performance reasons.This paper proposes an architecture where the accelerator performing DNN inference is replicated, as in the case of DCLS for cores, but using a cheaper implementation for the replica. In particular, we build on the stochastic nature of DNN-based object detection to realize two redundant accelerators where the secondary accelerator uses smartly chosen lower precision arithmetic (e.g., dropping some bits of the original data) so that it provides diverse redundancy, it can keep the performance of the primary accelerator, does not require as much cost as full- precision replication, and can build on the very same data stream from memory used by the primary accelerator. With a simple heuristic, we show that such a diverse redundancy scheme is able to cope with faults restricting false positives and negatives to a few relatively small objects. Martí Caro, Jordi Fornt, Jaume Abella 0001 |
COMPSAC | 2 |
| 2023 | WFAsic: A High-Performance ASIC Accelerator for DNA Sequence Alignment on a RISC-V SoCabstractThe ever-increasing yields in genome sequence data production pose a computational challenge to current genome sequence analysis tools, jeopardizing the future of personalized medicine. Leveraging hardware accelerators (GPUs, FPGAs, and ASICs) to accelerate computationally-intensive algorithms like sequence alignment has become paramount. Recently, the wavefront alignment algorithm was introduced, significantly reducing the execution time to perform sequence alignment. This paper presents the first-ever ASIC accelerator of the WFA integrated into a RISC-V system-on-chip. Our designed chip greatly accelerates sequence alignment, delivering up to 1076 × better performance over the CPU implementation of the WFA running on the RISC-V core of the chip. Abbas Haghi, Lluc Alvarez, Jordi Fornt, Juan Miguel De Haro Ruiz, Roger Figueras, Max Doblas, Santiago Marco-Sola, Miquel Moretó |
ICPP | 3 |
| 2023 | An automotive case study on the limits of approximation for object detection
Martí Caro, Hamid Tabani, Jaume Abella 0001, Francesc Moll, Enric Morancho, Ramon Canal, Josep Altet, Antonio Calomarde, Francisco J. Cazorla, Antonio Rubio 0001, Pau Fontova, Jordi Fornt |
J. Syst. Archit. | 12 |
| 2023 | An Energy-Efficient GeMM-Based Convolution Accelerator With On-the-Fly im2colabstractSystolic array architectures have recently emerged as successful accelerators for deep convolutional neural network (CNN) inference. Such architectures can be used to efficiently execute general matrix–matrix multiplications (GeMMs), but computing convolutions with this primitive involves transforming the 3-D input tensor into an equivalent matrix, which can lead to an inflation of the input data, increasing the OFF-chip memory traffic which is critical for energy efficiency. In this work, we propose a GeMM-based systolic array accelerator that uses a novel data feeder architecture to perform ON-chip, on-the-fly convolution lowering (also known as im2col), supporting arbitrary tensor and kernel sizes as well as strided and dilated (or atrous) convolutions. By using our data feeder, we reduce memory transactions and required bandwidth on state-of-the-art CNNs by a factor of two, while only adding an area and power overhead of 4% and 7%, respectively. Application specific integrated circuit (ASIC) implementation of our accelerator in 22-nm technology fits in less than 1.1 mm 2 and reaches an energy efficiency of 1.10 TFLOP/sW with 16-bit floating-point arithmetic. Jordi Fornt, Pau Fontova, Martí Caro, Jaume Abella 0001, Francesc Moll, Josep Altet, Christoph Studer |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2021 | FlyDVS: An Event-Driven Wireless Ultra-Low Power Visual Sensor NodeabstractEvent-based cameras, also called dynamic vision sensors (DVS), inspired by the human vision system, are gaining popularity due to their potential energy-saving since they generate asynchronous events only from the pixels changes in the field of view. Unfortunately, in most current uses, data acquisition, processing, and streaming of data from event-based cameras are performed by power-hungry hardware, mainly high-power FPGAs. For this reason, the overall power consumption of an event-based system that includes digital capture and streaming of events, is in the order of hundreds of milliwatts or even watts, reducing significantly usability in real-life low-power applications such as wearable devices. This work presents FlyDVS, the first event-driven wireless ultra-low-power visual sensor node that includes a low-power Lattice FPGA and, a Bluetooth wireless system-on-chip, and hosts a commercial ultra-low-power DVS camera module. Experimental results show that the low-power FPGA can reach up to 874 efps (event-frames per second) with only 17.6mW of power, and the sensor node consumes an overall power of 35.5 mW (including wireless streaming) at 200 efps. We demonstrate FlyDVS in a real-life scenario, namely, to acquire event frames of a gesture recognition data set. Alfio Di Mauro, Moritz Scherer 0001, Jordi Fornt, Basile Bougenot, Michele Magno, Luca Benini |
DATE | 3 |