Teresa Cervero

dblp:57/10313 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
5since 2021 · last 2026
0000-0001-7535-4821ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
YearPublicationVenuePosition
2026 Multi Partner Project: STRATUM, co-creation protocol and advanced smart GUI for a 3D neurosurgery supporting tool
abstract
STRATUM is a Horizon Europe multi-partner project developing a clinically validated, real-time 3D decision support tool for brain tumour surgery. The system integrates Hyperspectral Imaging (HSI), AI-based multimodal data fusion, and heterogeneous High-Performance Computing (HPC) architectures combining Graphics Processing Units (GPUs), Field-Programmable Gate Arrays (FPGAs), and Processing-In-Memory (PIM) technologies. A touchless augmented reality interface facilitates safe and intuitive intraoperative interaction. The distinguishing characteristic of STRATUM is its end-to-end co-designed approach, which integrates advanced computing, state-of-the-art imaging and clinical expertise into a unified Point-of-Care (PoC) platform. Utilising a structured co-creation methodology involving surgeons, engineers, and social scientists, the project ensures usability, safety and regulatory compliance from its early design stages to its clinical validation. The usability of STRATUM will be tested in three hospitals located in different European regions with diverse conditions and regulations. This will allow to collect advice and remarks from surgical staff in a continuous co-creation and co-tuning protocol. Beyond its clinical objectives, STRATUM contributes to the advancement of heterogeneous computing for real-time diagnostics, AI acceleration in critical medical environments and energy-efficient system integration. Furthermore, it delivers open datasets, validated AI pipelines, and performance benchmarks with a view to fostering future research and industrial innovation in digital surgery. The STRATUM project establishes a replicable model for intelligent, human-centred computing integrating microelectronics, AI and medicine.The paper presents an overview of the project in terms of aims, concepts and technologies and the description of the state of the work when approaching the end of the second of the five years planned. Specifically, the outcomes of the steps related to the collaboration with surgeons and medical staff (co-creation process) and the intelligent Graphical User Interface (GUI) development will be described. The latter allows for contactless interaction of the surgeon with several functions that have already been developed in the system.
Emanuele Torti, Himar Fabelo, Elisa Marenzi, Maria Luisa Alvarez-Male, Chrysanthi Bairaktari, Beatriz Noriega-Ortega, Raquel León, Santiago Marco, Asaf Badouh, Max Verbers, Javier Santana-Nunez, Yolanda Ramallo-Fariña, Christian Weis, Ana M. Wägner, Eduardo Juárez Martínez, Claudio Rial, Alfonso Lagares, Gustav Burström, Luis Jimenez-Roldan, Teresa Cervero, Miquel Moretó, Giovanni Danese, Svitlana Zinger, Francesca Manni, Miguel A. García-Bello, Lidia García, Jesús Morera, Juan F. Piñeiro, Bernardino Clavo, Francesco Leporati, Gustavo M. Callicó
DATE21
2024 3D Decision Support Tool for Brain Tumour Surgery: The STRATUM Project
abstract
Integrated digital diagnostics can support complex surgical procedures in many anatomical sites, brain tumour surgery being the most complex. STRATUM is a 5-year Horizon Europe funded project with the goal of developing an innovative 3D decision support tool for brain tumour surgeries, based on real-time multimodal data processing using artificial intelligence algorithms. The proposed tool is envisioned as an energy-efficient Point-of-Care computing system to be integrated within neurosurgical workflows to aid surgeons to make informed, efficient, and accurate decisions during surgical procedures. The expected long-term impact of STRATUM is to reduce the duration of surgical procedures, thus decreasing patients' risks, but also optimising the resources of European health care systems.
Himar Fabelo, Raquel León, Emanuele Torti, Santiago Marco, Max Verbers, Yann Falevoz, Yolanda Ramallo-Fariña, Christian Weis, Ana M. Wägner, Eduardo Juárez Martínez, Claudio Rial, Alfonso Lagares, Gustav Burström, Francesco Leporati, Elisa Marenzi, Teresa Cervero, Miquel Moretó, Giovanni Danese, Svitlana Zinger, Francesca Manni, Maria Luisa Alvarez-Male, Jesús Morera, Bernardino Clavo, Gustavo M. Callicó
DSD16
2024 Memory Sandbox: A Versatile Tool for Analyzing and Optimizing HBM Performance in FPGA
abstract
Main memory access has become an increasing performance bottleneck for traditional and High-Performance Computing (HPC) applications. High Bandwidth Memory (HBM) has emerged as an alternative to conventional DRAMs, offering higher bandwidth, lower power consumption, and greater integration capabilities to meet the escalating demands of contemporary applications. The transition to HBM of the most advanced Field Programmable Gate Arrays (FPGAs) marks a paradigm shift. However, users face substantial challenges due to the scarce technical documentation and tools supporting HBM on FPGAs. This paper introduces the Memory Sandbox, an open-source tool integrated into our FPGA-Shell1, designed to address the complexity of utilizing HBM in FPGAs. It allows developers and students to explore roofline performances while gaining insights to improve their designs. The Memory Sandbox enables users to configure various parameters such as the number of processing elements accessing memory, the access pattern (sequential, pseudo-random, or sparse-wise), and memory configurations, emulating multiple processor threads in diverse heterogeneous scenarios. It provides detailed analyses of memory access impacts in terms of latency and throughput for scenarios with accesses within and across HBM pseudo-channels; as well as HBM performance under concurrent access scenarios. Our results show that HBM achieves 99.99% of its nominal peak bandwidth with long sequential accesses but drops to 0.17% with random data access patterns. The tool also highlights the impact of multiple AXI ports targeting the same pseudo-channel, revealing that the aggregated throughput remains constant regardless of the pseudo-channel count. Furthermore, we validate the Memory Sandbox’s capabilities by effectively profiling complex access patterns like Sparse Matrix-Vector (SpMV), demonstrating its effectiveness in providing accurate performance insights.1https://github.com/MEEPproject/fpga_shell
Elias Perdomo, Xavier Martorell, Teresa Cervero, Behzad Salami 0001
SBAC-PAD3
2023 b8c: SpMV accelerator implementation leveraging high memory bandwidth
abstract
Sparse Matrix-Vector multiplication (SpMV), computing$y=A\times x$where$y, x$are dense vectors and$A$is a sparse matrix, is a key kernel in many HPC applications. Vitis Sparse Library's double precision SpMV (VSpMV) [1] is, to the best of our knowledge, the only performance-oriented, double-precision (64-bit) floating point implementation of SpMV on FPGAs equipped with High Bandwidth Memory (HBM).
José Oliver 0002, Carlos Álvarez 0001, Teresa Cervero, Xavier Martorell, John D. Davis, Eduard Ayguadé
FCCM3
2023 Accelerating SpMV on FPGAs Through Block-Row Compress: A Task-Based Approach
abstract
Sparse Matrix-Vector multiplication (SpMV), computing$y=\alpha\cdot A\times x+\beta\cdot y$where$y, x$are dense vectors,$\alpha, \beta$two scalar constants, and$A$is a sparse matrix, is a key kernel in many HPC applications. It exhibits a kind of memory access that is extremely hard to perform efficiently, due to its random access. In this paper, we present a new approach to accelerate SpMV on FPGAs. As FPGAs lack a default memory hierarchy, they can adapt to specific applications better. Also, an increasing number of FPGAs include High Bandwidth Memory (HBM), making the SpMV problem especially appealing to tackle on these kind of devices. We define a new sparse matrix encoding format (b8c) and its corresponding SpMV implementation using OmpSs@FPGA and HLS. This format allows us to leverage many of the FPGA strengths for intensive data processing, such as data streaming, customizable datapaths widths, parallel memory access for off-chip memory in the case of multiple memory channels (like in HBM), parallel memory access for on-chip memory and pipelining. We tested our proposal for both DDR and HBM memories to show the adaptability and scalability of our design. The presented b8c SpMV implementation is able to achieve higher performance than the state-of-the-art FPGA implementation of SpMV over all the matrices in the data set, achieving 3.52x performance on average with a minimum of 1.82x and a maximum of 6.28x even when running at 75% the frequency.
José Oliver 0002, Carlos Álvarez 0001, Teresa Cervero, Xavier Martorell, John D. Davis, Eduard Ayguadé
FPL3
2013 A Resource Manager for Dynamically Reconfigurable FPGA-Based Embedded Systems
abstract
FPGA-based embedded systems are gaining relevance for implementing a wide range of applications. Part of their success is due to their balanced compromise between performance and flexibility, but also because of their capability for exploiting the dynamic reconfiguration. However, the costly reconfiguration process and the lack of management support have prevented a broader use of the FPGAs. In order to contribute to solve these issues, in this paper we propose a software/hardware dynamic resource management system that combines scheduling and placement tasks, providing a complete management flow for supporting dynamically reconfigurable hardware designs. One of the advantages of the proposed model is the capability for running its scheduling and placement tasks in different nodes, as part of a distributed network. The results of our experiments demonstrate that our placement policy, specially designed for reconfigurable systems, achieves good results, in terms of reusability and performance, compared to other management approaches.
Teresa Cervero, Julio Dondo, Ana Gomez, Xerach Peña, Sebastián López, Fernando Rincón Calle, Roberto Sarmiento, Juan Carlos López 0001
DSD1
2011 Run-Time Scalable Architecture for Deblocking Filtering in H.264/AVC-SVC Video Codecs
abstract
Systems relying on fixed hardware components with a static level of parallelism can suffer from an under use of logical resources, since they have to be designed for the worst-case scenario. This problem is especially important in video applications due to the emergence of new flexible standards, like Scalable Video Coding (SVC), which offer several levels of scalability. In this paper, Dynamic and Partial Reconfiguration (DPR) of modern FPGAs is used to achieve run-time variable parallelism, by using scalable architectures where the size can be adapted at run-time. Based on this proposal, a scalable Deblocking Filter core (DF), compliant with the H.264/AVC and SVC standards has been designed. This scalable DF allows run-time addition or removal of computational units working in parallel. Scalability is offered together with a scalable parallelization strategy at the macro block (MB) level, such that when the size of the architecture changes, MB filtering order is modified accordingly.
Andrés Otero, Eduardo de la Torre, Teresa Riesgo, Teresa Cervero, Sebastián López, Gustavo M. Callicó, Roberto Sarmiento
FPL4
2011 A novel scalable Deblocking Filter architecture for H.264/AVC and SVC video codecs
abstract
A highly parallel and scalable Deblocking Filter (DF) hardware architecture for H.264/AVC and SVC video codecs is presented in this paper. The proposed architecture mainly consists on a coarse grain systolic array obtained by replicating a unique and homogeneous Functional Unit (FU), in which a whole Deblocking-Filter unit is implemented. The proposal is also based on a novel macroblock-level parallelization strategy of the filtering algorithm which improves the final performance by exploiting specific data dependences. This way communication overhead is reduced and a more intensive parallelism in comparison with the existing state-of-the-art solutions is obtained. Furthermore, the architecture is completely flexible, since the level of parallelism can be changed, according to the application requirements. The design has been implemented in a Virtex-5 FPGA, and it allows filtering 4CIF (704 × 576 pixels @30 fps) video sequences in real-time at frequencies lower than 10.16 Mhz.
Teresa Cervero, Andrés Otero, Sebastián López, Eduardo de la Torre, Gustavo M. Callicó, Roberto Sarmiento, Teresa Riesgo
ICME1