Martí Caro

dblp:320/0178 · DBLP profile ↗
← Back
7ranked-venue papers
5as first author
7since 2021 · last 2025
0000-0002-0856-1991ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2025 Leveraging Image-Based Transformations to Mitigate Adversarial Attacks in AI-Based Safety-Critical Systems
abstract
Dual (DMR) and Triple Modular Redundancy (TMR) are widely used techniques to provide fault detection and/or tolerance capabilities in safety-critical systems through – often diverse – redundancy. However, these systems remain vulnerable to adversarial attacks, which can mislead the AI models and lead to severe consequences. In this paper, we propose enhanced DMR and TMR implementations for image-based object detection leveraging image transformations during inference to mitigate the impact of adversarial attacks, hence addressing safety and security concerns simultaneously. Our approach achieves up to 12.9% and 12.2% higher accuracy in adversarial scenarios compared to state-of-the-art solutions in DMR and TMR configurations, respectively.
Martí Caro, Axel Brando, Jaume Abella 0001
IOLTS1
2025 GAVINA: flexible aggressive undervolting for bit-serial mixed-precision DNN acceleration
abstract
Voltage overscaling, or undervolting, is an enticing approximate technique in the context of energy-efficient Deep Neural Network (DNN) acceleration, given the quadratic relationship between power and voltage. Nevertheless, its very high error rate has thwarted its general adoption. Moreover, recent undervolting accelerators rely on 8-bit arithmetic and cannot compete with state-of-the-art low-precision (<8b) architectures. To overcome these issues, we propose a new technique called Guarded Aggressive underVolting (GAV), which combines the ideas of undervolting and bit-serial computation to create a flexible approximation method based on aggressively lowering the supply voltage on a select number of least significant bit combinations. Based on this idea, we implement GAVINA (GAV mIxed-Precision Accelerator), a novel architecture that supports arbitrary mixed precision and flexible undervolting, with an energy efficiency of up to 89 TOP/sW in its most aggressive configuration. By developing an error model of GAVINA, we show that GAV can achieve an energy efficiency boost of 20% via undervolting, with negligible accuracy degradation on ResNet-18.
Jordi Fornt, Pau Fontova, Adrian Gras, Omar Lahyani, Martí Caro, Jaume Abella 0001, Francesc Moll, Josep Altet
ISLPED5
2025 Semantic Diverse DMR and TMR for High-Integrity AI-Based Function Efficiency
abstract
Dual Modular Redundancy (DMR) and Triple Modular Redundancy (TMR), often with some form of diversity, are used in safety-critical systems to realize those functionalities at the highest integrity level providing fault detection and/or tolerance capabilities. Redundant executions are intended to provide bit-level identical results, and, upon any mismatch, an error is assumed and recovery actions taken as needed. In this article, we note that many emerging AI-based functionalities are intrinsically stochastic (e.g., camera-based object detection), and hence, their correctness must be judged semantically, with room for variations across correct outcomes (e.g., confidence must be above a given threshold, but how much it exceeds the threshold is irrelevant). Building on this observation, we propose strategies to create DMR and TMR implementations of AI-based functionalities that bring not only fault tolerance against random hardware faults but also against AI model inaccuracies. Those strategies, which can be realized with software-only means and ported to virtually any computing platform, build on input data modifications affecting the inference computations, but not the expected semantic output (e.g., introducing some controlled changes in the input data). Moreover, we provide our solution in the form of an open source tool for image and video processing aimed at facilitating the reproducibility of our evaluation results, and enabling others to use it and conduct further research on input transformations.
Martí Caro, Axel Brando, Jaume Abella 0001
ACM Trans. Cyber Phys. Syst.1
2023 Efficient Diverse Redundant DNNs for Autonomous Driving
abstract
Automotive applications with safety requirements must adhere to specific regulations such as ISO 26262, which imposes the use of diverse redundancy for the highest integrity levels (i.e., ASIL D). While this has been often achieved by means of Dual-Core LockStep (DCLS) for microcontrollers, it remains an open challenge how to realize diverse redundancy efficiently, i.e., without full duplication and preserving performance, for DNN-based safety-related tasks, such as object detection, needing accelerators for performance reasons.This paper proposes an architecture where the accelerator performing DNN inference is replicated, as in the case of DCLS for cores, but using a cheaper implementation for the replica. In particular, we build on the stochastic nature of DNN-based object detection to realize two redundant accelerators where the secondary accelerator uses smartly chosen lower precision arithmetic (e.g., dropping some bits of the original data) so that it provides diverse redundancy, it can keep the performance of the primary accelerator, does not require as much cost as full- precision replication, and can build on the very same data stream from memory used by the primary accelerator. With a simple heuristic, we show that such a diverse redundancy scheme is able to cope with faults restricting false positives and negatives to a few relatively small objects.
Martí Caro, Jordi Fornt, Jaume Abella 0001
COMPSAC1
2023 An automotive case study on the limits of approximation for object detection
Martí Caro, Hamid Tabani, Jaume Abella 0001, Francesc Moll, Enric Morancho, Ramon Canal, Josep Altet, Antonio Calomarde, Francisco J. Cazorla, Antonio Rubio 0001, Pau Fontova, Jordi Fornt
J. Syst. Archit.1
2023 An Energy-Efficient GeMM-Based Convolution Accelerator With On-the-Fly im2col
abstract
Systolic array architectures have recently emerged as successful accelerators for deep convolutional neural network (CNN) inference. Such architectures can be used to efficiently execute general matrix–matrix multiplications (GeMMs), but computing convolutions with this primitive involves transforming the 3-D input tensor into an equivalent matrix, which can lead to an inflation of the input data, increasing the OFF-chip memory traffic which is critical for energy efficiency. In this work, we propose a GeMM-based systolic array accelerator that uses a novel data feeder architecture to perform ON-chip, on-the-fly convolution lowering (also known as im2col), supporting arbitrary tensor and kernel sizes as well as strided and dilated (or atrous) convolutions. By using our data feeder, we reduce memory transactions and required bandwidth on state-of-the-art CNNs by a factor of two, while only adding an area and power overhead of 4% and 7%, respectively. Application specific integrated circuit (ASIC) implementation of our accelerator in 22-nm technology fits in less than 1.1 mm 2 and reaches an energy efficiency of 1.10 TFLOP/sW with 16-bit floating-point arithmetic.
Jordi Fornt, Pau Fontova, Martí Caro, Jaume Abella 0001, Francesc Moll, Josep Altet, Christoph Studer
IEEE Trans. Very Large Scale Integr. Syst.3
2022 At-scale evaluation of weight clustering to enable energy-efficient object detection
Martí Caro, Hamid Tabani, Jaume Abella 0001
J. Syst. Archit.1