Rafael Tornero

dblp:69/6781 · also Rafael Tornero Gavilá · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
4since 2021 · last 2024
0000-0001-5647-2232ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 3 first-author · 4 since 2021Software engineering, systems software and programming languages · 4 · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2024 Invited: Neuromorphic Vision Modalities in the NimbleAI 3D Chip
abstract
This paper provides an overview of the ongoing work to enable novel modalities of passive monocular neuromorphic vision in the NimbleAI sensing-processing architecture; namely, foveated and light-field event-driven vision with selective visual attention. The latter vision modality encodes 3D visual surroundings as sparse visual events in a 4D spatiotemporal domain, adding depth to current representation of visual information delivered by Dynamic Vision Sensors (DVS). The NimbleAI architecture implements hardware support for efficient execution of mainstream computer vision algorithms and AI models using these visual inputs. The architecture is designed to harness the latest advancements in 3D silicon integration, making it possible to squeeze sensing and spiking circuitry, memory, and processing engines into a miniature silicon volume.
Xabier Iturbe, Bernabé Linares-Barranco, Sio-Hoi Ieng, Arne Erdmann, Luca Peres, Oliver Rhodes, Rafael Tornero, Manolis Sifalakis, Marcel D. van de Burgwal, Amirreza Yousefzadeh, Maha Kooli, Riccardo Alidori, Pavel Zaykov
DAC7
2023 NimbleAI: Towards Neuromorphic Sensing-Processing 3D-integrated Chips
abstract
The NimbleAI Horizon Europe project leverages key principles of energy-efficient visual sensing and processing in biological eyes and brains, and harnesses the latest advances in$\mathbf{33D}$stacked silicon integration, to create an integral sensing-processing neuromorphic architecture that efficiently and accurately runs computer vision algorithms in area-constrained endpoint chips. The rationale behind the NimbleAI architecture is: sense data only with high information value and discard data as soon as they are found not to be useful for the application (in a given context). The NimbleAI sensing-processing architecture is to be specialized after-deployment by tunning system-level trade-offs for each particular computer vision algorithm and deployment environment. The objectives of NimbleAI are: (1)$\mathbf{100x}$performance per mW gains compared to state-of-the-practice solutions (i.e., CPU/GPUs processing frame-based video); (2)$\mathbf{50x}$processing latency reduction compared to CPU/GPUs; (3) energy consumption in the order of tens of mWs; and (4) silicon area of approx. 50 mm2.
Xabier Iturbe, Nassim Abderrahmane, Jaume Abella 0001, Sergi Alcaide, Eric Beyne, Henri-Pierre Charles, Christelle Charpin-Nicolle, Lars Chittka, Angélica Dávila, Arne Erdmann, Carles Estrada, Ander Fernández, Anna Fontanelli, José Flich, Gianluca Furano, Alejandro Hernán Gloriani, Erik Isusquiza, Radu Grosu, Carles Hernández 0001, Daniele Ielmini, Maha Kooli, Nicola Lepri, Bernabé Linares-Barranco, Jean-Loup Lachese, Eric Laurent, Menno Lindwer, Frank Linsenmaier, Mikel Luján, Karel Masarík, Nele Mentens, Orlando Moreira, Chinmay Nawghane, Luca Peres, Jean-Philippe Noël, Arash Pourtaherian, Christoph Posch, Peter Priller, Zdenek Prikryl, Felix Resch, Oliver Rhodes, Todor P. Stefanov, Moritz Storring, Michele Taliercio, Rafael Tornero, Marcel D. van de Burgwal, Geert Van der Plas, Elisa Vianello, Pavel Zaykov
DATE45
2023 Towards Efficient Neural Network Model Parallelism on Multi-FPGA Platforms
abstract
Nowadays, convolutional neural networks (CNN) are common in a wide range of applications. Their high accuracy and efficiency contrast with their computing requirements, leading to the search for efficient hardware platforms. FPGAs are suitable due to their flexibility, energy efficiency and low latency. However, the ever increasing complexity of CNNs demands higher capacity devices, forcing the need for multi-FPGA platforms. In this paper, we present a multi-FPGA platform with distributed shared memory support for the inference of CNNs. Our solution, in contrast with previous works, enables combining different model parallelism strategies applied to CNNs, thanks to the distributed shared memory support. For a four FPGA setting, the platform reduces the execution time of 2D convolutions by a factor of 3.95 when compared to single FPGA. The inference of standard CNN models is improved by factors ranging 3.63-3.87.
Rafael Tornero, José Flich
DATE2
2021 From a FPGA Prototyping Platform to a Computing Platform: The MANGO Experience
abstract
In this paper we describe the evolution of the FPGA-based prototype deployed in the MANGO project, from a hardware prototyping platform of HPC architectures to a computing platform targeting HPC and AI applications. Our main goal is to reinvest on the MANGO cluster by providing a duality in its use for both large-scale hardware prototyping and highperformance computation. From our experience we can reach several interesting conclusions about the complexities and hurdles that lay below FPGA technologies, and therefore, shedding some light onto the real complexities that difficult the adoption of FPGAs on either large-scale pure HPC systems or on hybrid systems (HPC + BigData/Ai).
José Flich, Rafael Tornero, Davide Russo, José Maria Martínez, Carles Hernández 0001
DATE2
2019 Challenges in Deeply Heterogeneous High Performance Systems
abstract
RECIPE (REliable power and time-ConstraInts-aware Predictive management of heterogeneous Exascale systems) is a recently started project funded within the H2020 FETHPC programme, which is expressly targeted at exploring new High-Performance Computing (HPC) technologies. RECIPE aims at introducing a hierarchical runtime resource management infrastructure to optimize energy efficiency and minimize the occurrence of thermal hotspots, while enforcing the time constraints imposed by the applications and ensuring reliability for both time-critical and throughput-oriented computation that run on deeply heterogeneous accelerator-based systems. This paper presents a detailed overview of RECIPE, identifying the fundamental challenges as well as the key innovations addressed by the project, which span run-time management, heterogeneous computing architectures, HPC memory/interconnection infrastructures, thermal modelling, reliability, programming models, and timing analysis. For each of these areas, the paper describes the relevant state of the art as well as the specific actions that the project will take to effectively address the identified technological challenges.
Giovanni Agosta, William Fornaciari, David Atienza 0001, Ramon Canal, Alessandro Cilardo, José Flich, Carles Hernández 0001, Michal Kulczewski, Giuseppe Massari, Rafael Tornero, Marina Zapater
DSD10
2017 MANGO: Exploring Manycore Architectures for Next-GeneratiOn HPC Systems
abstract
The Horizon 2020 MANGO project aims at exploring deeply heterogeneous accelerators for use in High-Performance Computing systems running multiple applications with different Quality of Service (QoS) levels. The main goal of the project is to exploit customization to adapt computing resources to reach the desired QoS. For this purpose, it explores different but interrelated mechanisms across the architecture and system software. In particular, in this paper we focus on the runtime resource management, the thermal management, and support provided for parallel programming, as well as introducing three applications on which the project foreground will be validated.
José Flich, Giovanni Agosta, Philipp Ampletzer, David Atienza 0001, Carlo Brandolese, Etienne Cappe, Alessandro Cilardo, Leon Dragic, Alexandre Dray, Alen Duspara, William Fornaciari, Gerald Guillaume, Ynse Hoornenborg, Arman Iranfar, Mario Kovac, Simone Libutti, Bruno Maitre, José Maria Martínez, Giuseppe Massari, Hrvoje Mlinaric, Ermis Papastefanakis, Tomás Picornell, Igor Piljic, Anna Pupykina, Federico Reghenzani, Isabelle Staub, Rafael Tornero, Marina Zapater, Davide Zoni
DSD27
2016 Enabling HPC for QoS-sensitive applications: The MANGO approach
José Flich, Giovanni Agosta, Philipp Ampletzer, David Atienza 0001, Carlo Brandolese, Alessandro Cilardo, William Fornaciari, Ynse Hoornenborg, Mario Kovac, Bruno Maitre, Giuseppe Massari, Hrvoje Mlinaric, Ermis Papastefanakis, Fabrice Roudet, Rafael Tornero, Davide Zoni
DATE15
2012 A multi-agent system for obtaining dynamic origin/destination matrices on intelligent road networks
abstract
Dynamic Origin/Destination matrices are one of the most important parameters for efficient and effective transportation system management. These matrices describe the vehicle flow between different points inside a region of interest for a given period of time. Usually, dynamic O/D matrices are estimated from link traffic counts, home interview and/or license plate surveys. Unfortunately, estimation methods take O/D flows as time invariant for a certain number of intervals of time, which cannot be suitable for some traffic applications. However, the advent of information and communication technologies (e.g., vehicle-to-infrastructure dedicated short range communications --V2I) to the transportation system domain has opened new data sources for computing O/D matrices. Taking the advantages of this technology, we propose in this paper a multi-agent system that computes the instantaneous and dynamic O/D matrix of any road network equipped with V2I technology for every time period and day in real-time. The implementation was done using JADE platform. The results show that the multi-agent system is able to obtain the instantaneous O/D matrix for any time period and day.
Rafael Tornero, Javier Martínez 0004, Joaquín Castelló
EATIS1
2012 An Optimal Control Approach to Power Management for Multi-Voltage and Frequency Islands Multiprocessor Platforms under Highly Variable Workloads
abstract
Reducing energy consumption in multi-processor systems-on-chip (MPSoCs) where communication happens via the network-on-chip (NoC) approach calls for multiple voltage/frequency island (VFI)-based designs. In turn, such multi-VFI architectures need efficient, robust, and accurate run-time control mechanisms that can exploit the workload characteristics in order to save power. Despite being tractable, the linear control models for power management cannot capture some important workload characteristics (e.g., fractality, non-stationarity) observed in heterogeneous NoCs, if ignored, such characteristics lead to inefficient communication and resources allocation, as well as high power dissipation in MPSoCs. To mitigate such limitations, we propose a new paradigm shift from power optimization based on linear models to control approaches based on fractal-state equations. As such, our approach is the first to propose a controller for fractal workloads with precise constraints on state and control variables and specific time bounds. Our results show that significant power savings (about 70%) can be achieved at run-time while running a variety of benchmark applications.
Paul Bogdan, Radu Marculescu, Rafael Tornero
NOCS4
2009 A multi-objective strategy for concurrent mapping and routing in networks on chip
abstract
The design flow of network-on-chip (NoCs) include several key issues. Among other parameters, the decision of where cores have to be topologically mapped and also the routing algorithm represent two highly correlated design problems that must be carefully solved for any given application in order to optimize several different performance metrics. The strong correlation between the different parameters often makes that the optimization of a given performance metric has a negative effect on a different performance metric. In this paper we propose a new strategy that simultaneously refines the mapping and the routing function to determine the Pareto optimal configurations which optimize average delay and routing robustness. The proposed strategy has been applied on both synthetic and real traffic scenarios. The obtained results show how the solutions found by the proposed approach outperforms those provided by other approaches proposed in literature, in terms of both performance and fault tolerance.
Rafael Tornero, Valentino Sterrantino, Maurizio Palesi, Juan M. Orduña
IPDPS1
2008 CART: Communication-Aware Routing Technique for Application-Specific NoCs
abstract
Networks on Chip (NoCs) have been shown as an efficient solution to the complex on-chip communication problems derived from the increasing number of processor cores. One of the key issues in the design of NoCs is the reduction of both area and power dissipation. As a result, two-dimensional meshes have become the preferred topology, since it offers low and constant link delay. Unfortunately, manufacturing defects or even real-time failures often make the resulting topology to become irregular, preventing the use of traditional routing algorithms. This scenario shows the need for topology-agnostic routing algorithms that provide a valid routing solution when applied over any topology. Moreover, in order to deal with run-time failures, the routing algorithm should be able to fit runtime constraints. This paper proposes a new communication-aware routing technique, referred to as CART, that optimizes the network performance for application-specific NoCs. CART combines a flexible, topology-agnostic routing algorithm with a communication-aware mapping technique that matches the traffic generated by the application with the available network bandwidth. Since the mapping technique can be pruned as needed in order to fit either quality function values or time constraints, CART can be adapted to fit with different computational costs. The evaluation results show that CART significatively improves network performance in terms of both latency and power consumption.
Rafael Tornero, Juan M. Orduña, Andres Mejia, José Flich, José Duato
DSD1
2008 A Communication-Aware Topological Mapping Technique for NoCs
Rafael Tornero, Juan M. Orduña, Maurizio Palesi, José Duato
Euro-Par1