Nuno Roma

dblp:35/8006 · also Nuno Filipe Valentim Roma · DBLP profile ↗
← Back
68ranked-venue papers
6as first author
20since 2021 · last 2026
0000-0003-2491-4977ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 33 · 1 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 4 first-author · 1 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Artificial intelligence and machine learning · 2Computer networks · 1
YearPublicationVenuePosition
2026 S2VEC: Compiler-Driven Stream Specialization for Linearized Vectorization
abstract
The performance benefits of vectorization are often limited by data movement, which remains a dominant bottleneck in modern processors. To mitigate this bottleneck, data stream-based mechanisms have been recently proposed, allowing general-purpose architectures to specialize memory accesses, leading to a reduction of load-to-use latency and loop code optimizations. However, compiler support for stream-based execution remains limited. Accordingly, this paper presents S2VEC, an LLVM-based compilation toolchain that automatically extracts, represents, and optimizes memory access patterns as data streams to enable stream-specialized vector execution. At its core, the Stream-Dataflow IR models computation as a dataflow over parameterized data streams, bridging the gap between conventional compiler infrastructures and stream-oriented architectures. By transforming complex and indirect memory accesses into linearized stream representations, S2VEC simplifies vectorization and expands its applicability to complex memory access patterns. When targeting a RISC-V stream-vector ISA extension, S2VEC demonstrates improved vectorization coverage and produces code comparable to hand-optimized implementations. Conducted evaluations on a gem5-based in-order stream-vector model show an overall speedup of 5.9 × over an equivalently provisioned RVV-based in-order core and 1.5 × over a wide out-of-order processor, demonstrating the performance benefits of compiler-driven stream specialization.
Luís Crespo, Gabriel Falcão Paiva Fernandes, Pedro Tomás, Nuno Roma, Nuno Neves 0002
ICS5
2026 Reconfigurable FPU With Precision Auto-Tuning for Next-Generation Transprecision Computing
abstract
Recent advances in process technology have shifted the research focus from strict raw scaling to the conception of energy-efficient computing units, capable of adapting to the target application precision requirements. A key opportunity lies in floating-point arithmetic, where traditional fixed-precision formats (32/64-bit) often impose unnecessary resources and costs in both performance and power. To address this problem, this manuscript introduces an Automatic Precision Floating-point Unit (APFU) that extends prior reconfigurable designs by incorporating a hardware controller capable of autonomously tuning the operand precision and vectorization levels at runtime. The APFU supports IEEE-754 (double, single, half-precision), bfloat16, and DLFloat formats, and exploits unused datapath capacity to increase the throughput via vector operations. Unlike previous approaches, the proposed APFU introduces a runtime controller that autonomously determines the operand precision and vectorization levels based on operand characteristics and execution mode. The presented experimental evaluation considers an implementation using a 28-nm UMC technology, achieving up to 152 GOPS/W, and demonstrates a detailed analysis across operating frequencies, vectorization modes, and runtime precision adjustment.
Guilherme Dias, Luís Crespo, Timo Schlachter, Marc Andre Heller, Jens Krüger 0004, Pedro Tomás, Nuno Roma, Nuno Neves 0002
IEEE Trans. Circuits Syst. I Regul. Pap.7
2026 Real-Time ORB Accelerator for Embedded FPGA-Based SoCs With ROS Integration
abstract
As computer vision continues to expand across various application domains—including localization, mapping, object recognition, and 3-D reconstruction—feature extraction methods, such as oriented FAST and rotated BRIEF (ORB), have gained widespread adoption due to their rotation and scale invariance. However, existing efforts to accelerate these techniques through hardware implementations faced different challenges related to resource and power consumption demands, limiting their feasibility for low-power embedded devices. Accordingly, this article proposes a new scalable and efficient ORB accelerator, designed for low-power and resource-constrained environments. It introduces a novel and efficient architecture that exploits quantization of the feature orientation angle into discrete rotation sectors. A complete robot operating system (ROS) node based on the proposed ORB accelerator is also deployed, providing seamless integration with other computer vision-enabled systems. When compared to other state-of-the-art solutions, the proposed system, implemented on an embedded system-on-chip (SoC) with a low-cost field-programmable gate array (FPGA), offers energy efficiency improvements between$6.7\times $and$16.2\times $, while requiring fewer hardware resources.
Andre Costa, José Duarte Lopes, Pedro Tomás, Nuno Roma, Nuno Neves 0002
IEEE Trans. Very Large Scale Integr. Syst.4
2025 RVEBS: Event-Based Sampling on RISC-V
abstract
As RISC-V ISA continues to gain traction for both embedded and high-performance computing, the demand for advanced monitoring tools has become critical to fine-tuning the applications' performance. Current RISC-V hardware performance monitors already provide basic event counting but lack sophisticated features like event-based sampling, which are available in more established architectures such as x86 and ARM. This paper presents the first RISC-V Event-Based Sampling (RVEBS) system for comprehensive performance monitoring and application profiling. The proposed system builds upon existing RISC-V specifications, incorporating necessary modifications to enable the desired functionality. It also presents an OpenSBI extension to provide privileged software access to newly implemented control status registers that manage the sampling process. An implementation use case based on an OpenPiton processor featuring a CVA6 core on 28nm CMOS technology was presented. The results indicate that the proposed scheme is lightweight, highly accurate, and does not impact the processor's critical path while maintaining minimal impact on overall application performance.
Tiago Rocha, Nuno Neves 0002, Nuno Roma, Pedro Tomás, Leonel Sousa
DATE3
2025 URAM-Based Asynchronous FIFO Design for Improved Throughput and FPGA RAM Usage
abstract
First-In First-Out (FIFO) memories are essential for data transfers in Field Programmable Gate Array (FPGA) designs, providing buffering and communication facilities between different components and clock domains. Traditional implementations rely on Block-RAM (BRAM) cells, but transferring large data packets can lead to excessive BRAM usage, potentially limiting resources for other parts of the system. This paper proposes an alternative Ultra-RAM (URAM)-based asynchronous FIFO design, allowing to optimise the overall FPGA RAM usage for applications that require deep and wide FIFOs. Several implementation strategies are explored to ensure maximum data throughput and support different packet sizes. Experimental results obtained on an AMD Zynq UltraScale+ FPGA running at 500 MHz demonstrate the proposed FIFO's ability to successfully transfer complete packets (as large as 1024 64-bit words) between two distinct clock domains without introducing any stalls.
Martim Rosado, Pedro Tomás, Nuno Roma, André David
FPL3
2025 Real-Time ORB Accelerator with ROS Integration for Embedded FPGA SoCs
abstract
As computer vision continues to expand across various application domains - including localisation, mapping, object recognition, and 3D reconstruction - feature extraction methods such as Oriented FAST and Rotated BRIEF (ORB) gained widespread adoption due to their rotation and scale invariance. However, existing efforts to accelerate these techniques through hardware implementations faced challenges related to high resource and power consumption demands, limiting their feasibility for low-power embedded devices. Accordingly, this paper proposes a new scalable and efficient ORB accelerator, designed for low-power resource-constrained environments. It introduces a novel resource-efficient architecture that exploits quantisation of the feature orientation angle into discrete rotation sectors. A complete ROS node based on the proposed ORB accelerator is also deployed, providing seamless integration with other computer vision-enabled systems. When compared to other state-of-the-art solutions, the proposed system, implemented on an embedded System-on-Chip (SoC) with a low-cost FPGA, offers between 6.7x and 16.2x energy efficiency improvements, while requiring fewer hardware resources.
Andre Costa, Pedro Tomás, Nuno Roma, Nuno Neves 0002
ISCAS3
2025 Stream-Driven Acceleration for Embedded RISC-V SoCs
abstract
This paper proposes a stream-driven computational model that expands the recent stream vectorization paradigm into a full dataflow-driven computing model. It exploits spatial computation and time-multiplexing, while relying on streaming engines implementing the RISC-V UVE specification to manage data access patterns, thus streamlining memory operations and reducing latency. By abstracting the kernel loops into stream data-flow graphs and mapping them onto a processing element array, the conceived accelerator architecture exploits both spatial and temporal parallelism across a wide range of computational tasks. Experimental results, conducted on a synthesized 7nm implementation, demonstrate the proposed model’s potential to develop high-efficiency accelerators in data-intensive applications, offering performance gains of up to 6× compared with an ARM Cortex-A53 CPU with NEON and 15× compared with a scalar Rocket RISC-V CPU, along with 3.86× energy efficiency improvements.
João Maia, Ana Silveira, Gonçalo Midões, Nuno Neves 0002, Pedro Tomás, Nuno Roma
ISCAS6
2025 High-throughput packet aggregator for the back-end DAQ of CERN CMS HGCAL detector
abstract
The Phase 2 upgrade of CERN LHC accelerator requires the CMS detector to replace its endcap calorimeters with the new HGCAL. Each of the 96 back-end DAQ FPGAs will process LHC collision data from 108 input optical fibre pairs operating at 10 Gb/s and will have to route the aggregated data into 12 output optical fibre pairs operating at 24 Gb/s. This paper proposes an architecture of a multi-FIFO readout system, to be integrated into the highly space-constrained FPGA design of the back-end DAQ system of the HGCAL. This readout solution aggregates and conveys the processed data to the system’s outputs and minimises the detector dead time by load-balancing the data distribution between sources and destinations. Experimental results demonstrate that the proposed circuit is able to cope with the final detector’s bandwidth and configuration requirements, as well as several test systems needed for the assembly and commissioning of the detector.
Martim Rosado, Pedro Tomás, Nuno Roma, André David
ISCAS3
2025 MIDAS: A Mapping Infrastructure for Configurable, Data-Streaming Based Domain Specific Accelerators
abstract
As computational demands continue to grow in key application domains, Domain-Specific Accelerators (DSAs) have become a promising solution for delivering high performance with improved area and energy efficiency. However, traditional memory address generation in DSAs often consumes valuable resources that could otherwise be used to enhance performance. Data streaming mechanisms address this issue by eliminating the need for address generation nodes, but the exploitation of this paradigm remains underutilized in existing DSA design toolchains. On the other hand, despite fast kernel mapping and early feedback metrics being regarded as critical features in design space exploration for DSAs, most publicly available toolchains are either too slow, provide feedback only after a full compilation, or lack direct support for data streaming. This paper introduces a new Mapping Infrastructure for Data-Streaming-Based Accelerators (MIDAS) designed to rapidly map kernels onto configurable processing arrays equipped with an integrated streaming engine. When used as a co-design exploration tool, it provides early feedback on multiple metrics, facilitating architectural optimisation and pruning, and leading to a considerable improvement of hardware and energy efficiency. Moreover, the compute-only Data Flow Graphs (DFGs) are also more amenable to vectorisation, enabling further performance gains from data streaming. The obtained experimental results demonstrate that MIDAS achieves mapping speeds up to 73x faster than state-of-the-art tools like CGRA-ME’s CLUMAP, while maintaining competitive mapping quality. Under the data streaming paradigm, the implemented array achieves up to 8x performance improvement on a 4x4 PE array configuration compared to conventional architectures, with significantly reduced resource usage.
Martim Bento, Nuno Neves 0002, Pedro Tomás, Nuno Roma
SBAC-PAD4
2025 A Survey on Stream-Based Architectures: From Accelerators to CPUs
abstract
In the past few years, there has been a renewed effort to advance general-purpose architectures. In particular, to deliver performance and energy efficiency advantages, several techniques have been applied based on new forms of specialization while maintaining usability. As a result, data movement and communication have become the primary bottlenecks in computer systems. To overcome this, one of the most recent breakthroughs has been the introduction of data streaming mechanisms, just like those used in accelerators, into modern general-purpose processors (GPPs). This article comprehensively reviews stream-based architectures, tracing their development from accelerator solutions to their recent adoption in GPPs. This survey starts by introducing the fundamental principles of stream specialization, followed by a taxonomy for memory accesses, and formal mathematical models to represent them as data streams. Then, it categorizes different topologies of data stream specialization and examines them from a compiler’s perspective. Some of the most representative architectures proposed in the past few years, including instruction set architecture (ISA) and streaming engines, are described, followed by a comparative analysis that highlights their key features and presents quantitative evaluations. Then, we discuss some open challenges and suggest directions for future research in stream-based architectures.
Luís Crespo, Nuno Neves 0002, Pedro Tomás, Nuno Roma
Proc. IEEE4
2025 Improving Coding Efficiency of Massive Parallel Intra Prediction Using Alternative References
abstract
Exploring massive parallelism is a common strategy to mitigate the processing time of modern video encoding standards. Nonetheless, data dependencies challenge parallelism exploitation, especially during intra prediction, where the reconstructed adjacent blocks are used as references. Some works use the original frame samples as references to decouple adjacent blocks and allow parallelism. Still, the original samples are static and cannot model the nuances of different bitrates. In this context, this work seeks to improve the coding efficiency of parallel intra prediction implementations by using alternative reference samples based on low-pass filters that better represent the nuances of different bitrates for any partitioning structure. Variations in multiple aspects of the filters are considered, such as their dimension and also the precision and distribution of their coefficients. Experimental evaluations assessed the similarity of such alternative samples when compared to the regular ones, in addition to their impacts on coding efficiency and the processing overhead required to obtain such samples. The results from such experiments demonstrate that the alternative references improve coding efficiency when compared to the original samples, especially at lower bitrates. Furthermore, the additional filtering stage poses negligible timing overhead in most computing systems.
Iago Storch, Nuno Roma, Daniel Palomino 0001, Sergio Bampi
IEEE Trans. Circuits Syst. I Regul. Pap.2
2023 Supporting RISC-V Performance Counters Through Linux Performance Analysis Tools
abstract
Increased attention to RISC-V open Instruction Set Architecture (ISA), a base ISA with a variety of optional extensions, has fueled its move from embedded devices to the high-performance computing arena, with the proliferation of RISC-V-based accelerators. However, the absence of powerful performance monitoring tools often results in poorly optimized applications and, consequently, limited computing performance. While the RISC-V ISA already defines a hardware performance monitor (HPM), research and development on RISC-V-based devices have been more focused on architectures and compilers rather than tools to support monitoring performance. To overcome this limitation, a comprehensive set of extensions and modifications to the Performance analysis tools for Linux (perf/perf_events) are proposed in this paper, and a PAPI library interface is presented. These new extensions comprise not only the Linux kernel but also the OpenSBI interface, and aim to achieve full support for the RISC-V performance monitoring specification. The conducted testing and evaluation were carried out on a HiFive Unmatched board and on a CVA6 core, but the proposed extensions, and the corresponding implementation, are easily portable to other systems.
Joao Mario Domingos, Tiago Rocha, Nuno Neves 0002, Nuno Roma, Pedro Tomás, Leonel Sousa
ASAP4
2023 Trading Performance, Power, and Area on Low-Precision Posit MAC Units for CNN Training
abstract
The recently proposed Posit number system has been regarded as a particularly well-suited floating-point format to optimize the throughput and efficiency of low-precision computations in convolutional neural network (CNN) applications. In particular, the Posit format offers a balance between decimal accuracy and dynamic range, which results in a distribution of values that seems particularly interesting for deep learning applications. However, the adoption of the Posit still raises some concerns regarding hardware complexity, particularly when accounting for the overheads associated with the quire exact accumulator. Accordingly, this paper presents a holistic study on the model accuracy, performance, power, and area trade-offs when adopting low-precision Posit multiply-accumulate (MAC) units for the training of CNNs. In particular, 28nm ASIC implementations of a reference Posit MAC unit architecture demonstrate that the quire accounts for over 70% of the area and power utilization, and the obtained CNN training results showed that its use is only strictly required when considering mixed low-precision configurations. As a result, reducing the size of the quire results in an average reduction of area and power by 57% and 47%, without imposing visible training accuracy losses.
Luís Crespo, Pedro Tomás, Nuno Roma, Nuno Neves 0002
SBAC-PAD3
2022 Mode-Adaptive Subsampling of SAD/SSE Operations for Intra Prediction Cost Reduction
abstract
Modern video encoders, such as the recently proposed AV1 and VVC, offer significant encoding gains at the cost of a corresponding increase of the computational effort. This is the case of the adopted intra prediction techniques, comprehending an increased number of prediction modes and range. To mitigate this computational cost, the presented work proposes a new mode-adaptive algorithm that significantly reduces the number of SAD/SSE operations during intra prediction, by generating an optimized subsampling pattern adaptive to each prediction mode. The method can be applied to any video codec and, when applied to AV1, it led to an encoding time reduction and BD-BR impact of 15.36% and 0.6%, respectively, or 7.97% and –0.02%, depending on the selected subsampling parameters. When implemented in hardware, the proposed technique provides an effective reduction as high as 75% of both the area and power on the modified distortion calculation module.
Marcel Moscarelli Corrêa, Nuno Roma, Daniel Palomino 0001, Guilherme Corrêa 0001, Luciano Volcan Agostini
ISCAS2
2022 Early prototyping and testing of CERN LHC CMS high-granularity calorimeter slow-control system
abstract
The Compact Muon Solenoid (CMS) high-granularity calorimeter (HGCAL) upgrade for CERN's Large Hadron Collider (LHC) high-luminosity phase is a detector with more than 6 million channels that will provide precise sensing and measurement of position, timing, and energy of the particles produced in the collisions of the beams. The HGCAL electronics are a large and complex set of processing systems split into front-end and back-end. The front-end, located in the experimental cavern, consists of$\boldsymbol{\approx 150}$thousand radiation tolerant ASICs. The high-density FPGA-based back-end is housed away from the radiation area in a set of Advanced Telecommunications Computing Architecture (ATCA) boards and crates hosting$\boldsymbol{\approx 100}$FPGAs. Each ATCA back-end board will comprise one (or two) FPGAs, managing up to$\boldsymbol{\approx 120}$optical links, each providing a transmission rate of 10.24 Gb/s between the back-end and the front-end electronics. Each back-end FPGA is responsible for configuring and monitoring up to$\boldsymbol{\approx 3500}$front-end ASICs and will be controlled by software running on a back-end MPSoC that provides the entry point for the whole control procedure. This paper presents the design and implementation of the prototyping infrastructure deployed to test and validate the slow-control block of the HGCAL back-end electronics, together with the related interfaces with the controller MPSoC and the front-end transceiver ASICs. The required functionalities have been validated with a ZCU102 Xilinx Ultrascale+ development board, which emulated the back-end elements that are still under development and not yet available for this comprehensive test. This development board was connected to other custom ASIC development boards via optical links, emulating the front-end side of the system, also still under development. Besides providing reliable testing and validation of the operation of the whole infrastructure, the prototyping platform also allowed to attain the required software/hardware portability that ensures easy integration/replacement of all the (still) emulated components with their final implementations.
Martim Rosado, Stavros Mallios, Pedro Tomás, Nuno Roma, André David
RSP4
2022 gem5-ndp: Near-Data Processing Architecture Simulation From Low Level Caches to DRAM
abstract
Unlike standard accelerators, the performance of Near-Data Processing (NDP) devices highly depends on the operation of the surrounding system, namely, the Central Processing Unit (CPU) and the memory hierarchy. Therefore, to accurately evaluate the gain provided by such devices, the entire processing system must be considered. Recent proposals redesigned existing architectural simulators to estimate the performance of NDP devices. However, the conclusions that can be drawn from using these frameworks are limited, and they fail to provide full support to simulate these devices (e.g., most simulators do not allow simultaneous operation of the CPU and the NDP device). In this paper, a novel framework (called gem5-ndp) based on the gem5 architectural simulator is proposed, providing full support to the development, validation, and evaluation of novel NDP architectures. To illustrate the process of developing and integrating an NDP device with a processing system using the proposed framework, as well as to demonstrate its viability and benefits, two case studies are also proposed and thoroughly discussed. gem5-ndp significantly improves the performance evaluation confidence of NDP devices with results showing that classical approaches lead to a deviation of up to 54.9 % when compared with results obtained with gem5-ndp.
João Vieira, Nuno Roma, Gabriel Falcão Paiva Fernandes, Pedro Tomás
SBAC-PAD2
2022 Decoupling GPGPU voltage-frequency scaling for deep-learning applications
Francisco Mendes 0002, Pedro Tomás, Nuno Roma
J. Parallel Distributed Comput.3
2021 Positnn: Training Deep Neural Networks with Mixed Low-Precision Posit
abstract
Low-precision formats have proven to be an efficient way to reduce not only the memory footprint but also the hardware resources and power consumption of deep learning computations. Under this premise, the posit numerical format appears to be a highly viable substitute for the IEEE floating-point, but its application to neural networks training still requires further research. Some preliminary results have shown that 8-bit (and even smaller) posits may be used for inference and 16-bit for training, while maintaining the model accuracy. The presented research aims to evaluate the feasibility to train deep convolutional neural networks using posits. For such purpose, a software framework was developed to use simulated posits and quires in end-to-end training and inference. This implementation allows using any bit size, configuration, and even mixed precision, suitable for different precision requirements in various stages. The obtained results suggest that 8-bit posits can substitute 32-bit floats during training with no negative impact on the resulting loss and accuracy.
Gonçalo Raposo, Pedro Tomás, Nuno Roma
ICASSP3
2021 Unlimited Vector Extension with Data Streaming Support
abstract
Unlimited vector extension (UVE) is a novel instruction set architecture extension that takes streaming and SIMD processing together into the modern computing scenario. It aims to overcome the shortcomings of state-of-the-art scalable vector extensions by adding data streaming as a way to simultaneously reduce the overheads associated with loop control and memory access indexing, as well as with memory access latency. This is achieved through a new set of instructions that pre-configure the loop memory access patterns. These attain accurate and timely data prefetching on predictable access patterns, such as in multidimensional arrays or in indirect memory access patterns. Each of the configured data streams is associated to a general- purpose vector register, which is then used to interface with the streams. In particular, iterating over a given stream is simply achieved by reading/writing to the corresponding input/output stream, as the data is instantly consumed/produced. To evaluate the proposed UVE, a proof-of-concept gem5 implementation was integrated in an out-of-order processor model, based on the ARM Cortex-A76, thus taking into consideration the typical speculative and out-of-order execution paradigms found in high- performance computing processors. The evaluation was carried out with a set of representative kernels, by assessing the number of executed instructions, its impact on the memory bus and its overall performance. Compared to other state-of-the-art solutions, such as the upcoming ARM Scalable Vector Extension (SVE), the obtained results show that the proposed extension attains average performance speedups over 2.4 × for the same processor configuration, including vector length.
Joao Mario Domingos, Nuno Neves 0002, Nuno Roma, Pedro Tomás
ISCA3
2021 Compiler-Assisted Data Streaming for Regular Code Structures
abstract
The performance of modern processors is often limited by execution stalls resulting from long memory access latencies. Compile-time optimizations, deep cache hierarchies and prefetching mechanisms already provide significant performance gains, by performing memory accesses in parallel with computation. However, they are reaching a throughput improvement limit. Hence, new solutions that effectively exploit the memory access patterns to improve processing throughput are required. To achieve this objective, a new compiler-assisted data streaming method is proposed. It leverages static analysis and code transformations with an on-chip data streaming support as a viable alternative to prefetching mechanisms for regular code structures. Static analysis is used to identify and encode memory accesses with a dedicated representation. Then, a code transformation algorithm detaches data indexation and address calculation from computation, allowing for a significant code reduction. An on-chip data stream controller, attached to the L1 data cache, is used to autonomously generate memory accesses from the pattern representation and reorganize the data transfers in streams, with the aid of stream buffers. When compared with state-of-the-art prefetchers, the proposed solution provides up to 26 percent of code reduction, an IPC improvement of 2.4x, and an average performance improvement of 40 percent.
Nuno Neves 0002, Pedro Tomás, Nuno Roma
IEEE Trans. Computers3
2020 Reconfigurable Stream-based Tensor Unit with Variable-Precision Posit Arithmetic
abstract
The increased adoption of DNN applications drove the emergence of dedicated tensor computing units to accelerate multi-dimensional matrix multiplication operations. Although they deploy highly efficient computing architectures, they often lack support for more general-purpose application domains. Such a limitation occurs both due to their consolidated computation scheme (restricted to matrix multiplication) and due to their frequent adoption of low-precision/custom floating-point formats (unsuited for general application domains). In contrast, this paper proposes a new Reconfigurable Tensor Unit (RTU) which deploys an array of variable-precision Vector MultiplyAccumulate (VMA) units. Furthermore, each VMA unit leverages the new Posit floating-point format and supports the full range of standardized posit precisions in a single SIMD unit, with variable vector-element width. Moreover, the proposed RTU explores the Posit format features for fused operations, together with spatial and time-multiplexing reconfiguration mechanisms to fuse and combine multiple VMAs to map high-level and complex operations. The RTU is also supported by an automatic data streaming infrastructure and a pipelined data movement scheme, allowing it to accelerate the computation of most data-parallel patterns commonly present in vectorizable applications. The proposed RTU showed to outperform state-of-the-art tensor and SIMD units, present in off-the-shelf platforms, in turn resulting in significant energy-efficiency improvements.
Nuno Neves 0002, Pedro Tomás, Nuno Roma
ASAP3
2020 Processing Convolutional Neural Networks on Cache
abstract
With the advent of Big Data application domains, several Machine Learning (ML) signal-processing algorithms, such as Convolutional Neural Networks (CNNs), are required to process progressively larger datasets at a great cost in terms of both compute power and memory bandwidth. Although dedicated accelerators have been developed targeting this issue, they usually require moving massive amounts of data across the memory hierarchy to the processing cores and low-level knowledge of how data is stored in the memory devices to enable in-/near-memory processing solutions. In this paper, we propose and assess a novel mechanism that operates at cache level, leveraging both data-proximity and parallel processing capabilities, enabled by dedicated fully-digital vector Functional Units (FUs). We also demonstrate the integration of this mechanism in a conventional Central Processing Unit (CPU). The obtained results show that our engine provides performance improvements on CNNs ranging from 3.92× to 16.6×.
João Vieira, Nuno Roma, Gabriel Falcão Paiva Fernandes, Pedro Tomás
ICASSP2
2020 Exploiting Non-conventional DVFS on GPUs: Application to Deep Learning
abstract
The use of Graphics Processing Units (GPUs) to accelerate Deep Neural Networks (DNNs) training and inference is already widely adopted, allowing for a significant increase in the performance of these applications. However, this increase in performance comes at the cost of a consequent increase in energy consumption. While several solutions have been proposed to perform Voltage-Frequency (V-F) scaling on GPUs, these are still one-dimensional, by simply adjusting frequency while relying on default voltage settings. To overcome this, this paper introduces a methodology to fully characterize the impact of non-conventional Dynamic Voltage and Frequency Scaling (DVFS) in GPUs. The proposed approach was applied to an AMD Vega 10 Frontier Edition GPU. When applying this non-conventional DVFS scheme to DNNs, the obtained results show that it is possible to safely decrease the GPU voltage, allowing for a significant reduction of the energy consumption (up to 38%) and the Energy-Delay Product (EDP) (up to 41%) on the training of CNN models, with no degradation of the networks accuracy.
Francisco Mendes 0002, Pedro Tomás, Nuno Roma
SBAC-PAD3
2019 Flying tourist problem: Flight time and cost minimization in complex routes
Rafael Marques, Luís M. S. Russo, Nuno Roma
Expert Syst. Appl.3
2019 DVFS-aware application classification to improve GPGPUs energy efficiency
João Guerreiro 0004, Aleksandar Ilic, Nuno Roma, Pedro Tomás
Parallel Comput.3
2019 Modeling and Decoupling the GPU Power Consumption for Cross-Domain DVFS
abstract
Dynamic voltage and frequency scaling (DVFS) is a popular technique to improve the energy-efficiency of high-performance computing systems. It allows placing the devices into lower performance states when the computational demands are lower, opening the possibility for significant power/energy savings. This work presents a GPU power consumption model, used to predict the GPU power consumption of any application at different frequency levels. To obtain this model, an estimation algorithm is proposed, relying on careful benchmarking of the GPU architecture. The model can estimate the contribution of twelve different GPU components (FP32-ADD/MUL/FMA, FP64-ADD/MUL/FMA, INT, SF, CF units, shared memory, L2-cache, and DRAM) to the GPU power consumption. Different model use cases are evaluated (fixed-frequency, DVFS, and scaling-factors), which can obtain both the total or the per-component GPU power consumption. A technique to export models to a distinct GPU from the one it was estimated on is also proposed. These approaches were extensively validated on five different GPUs from the three most recent microarchitectures with a set of 42 standard benchmarks, achieving very accurate predictions. In particular, the scaling-factor power model achieves an average prediction error of 3.5 percent (Titan Xp), 4.6 percent (GTX Titan X), 3.1 percent (GTX 980) and 2.4 percent (Tesla K40c).
João Guerreiro 0004, Aleksandar Ilic, Nuno Roma, Pedro Tomás
IEEE Trans. Parallel Distributed Syst.3
2018 GPGPU Power Modeling for Multi-domain Voltage-Frequency Scaling
abstract
Dynamic Voltage and Frequency Scaling (DVFS) on Graphics Processing Units (GPUs) components is one of the most promising power management strategies, due to its potential for significant power and energy savings. However, there is still a lack of simple and reliable models for the estimation of the GPU power consumption under a set of different voltage and frequency levels. Accordingly, a novel GPU power estimation model with both core and memory frequency scaling is herein proposed. This model combines information from both the GPU architecture and the executing GPU application and also takes into account the non-linear changes in the GPU voltage when the core and memory frequencies are scaled. The model parameters are estimated using a collection of 83 microbenchmarks carefully crafted to stress the main GPU components. Based on the hardware performance events gathered during the execution of GPU applications on a single frequency configuration, the proposed model allows to predict the power consumption of the application over a wide range of frequency configurations, as well as to decompose the contribution of different parts of the GPU pipeline to the overall power consumption. Validated on 3 GPU devices from the most recent NVIDIA microarchitectures (Pascal, Maxwell and Kepler), by using a collection of 26 standard benchmarks, the proposed model is able to achieve accurate results (7%, 6% and 12% mean absolute error) for the target GPUs (Titan Xp, GTX Titan X and Tesla K40c).
João Guerreiro 0004, Aleksandar Ilic, Nuno Roma, Pedro Tomás
HPCA3
2018 Exploiting Compute Caches for Memory Bound Vector Operations
abstract
To reduce the average memory access time, most current processors make use of a multilevel cache subsystem. However, despite the proven benefits of such cache structures in the resulting throughput, conventional operations such as copy, simple maps and reductions still require moving large amounts of data to the processing cores. This imposes significant energy and performance overheads, with most of the execution time being spent moving data across the memory hierarchy. To mitigate this problem, a Cache Compute System (CCS) that targets memory-bound kernels such as map and reduce operations is proposed. The developed CCS takes advantage of long cache lines and data locality to avoid data transfers to the processor and exploits the intrinsic parallelism of vector compute units to accelerate a set of 48 operations commonly used in map and reduce patterns. The CCS was validated by integrating it with an MB-Lite soft-core in a Xilinx Virtex-7 VC709 Development Board. When compared to the MB-Lite core, the proposed CCS presents performance improvements in the execution of the commands ranging from 4x to 408x, and energy efficiency gains from 6x to 328x.
João Vieira, Nuno Roma, Pedro Tomás, Paolo Ienne, Gabriel Falcão Paiva Fernandes
SBAC-PAD2
2018 Highly parallel HEVC decoding for heterogeneous systems with CPU and GPU
abstract
The High Efficiency Video Coding HEVC standard provides a higher compression efficiency than other video coding standards but at the cost of an increased computational load, which makes hard to achieve real-time encoding/decoding for ultra high-resolution and high-quality video sequences. Graphics Processing Units GPU are known to provide massive processing capability for highly parallel and regular computing kernels, but not all HEVC decoding procedures are suited for GPU execution. Furthermore, if HEVC decoding is accelerated by GPUs, energy efficiency is another concern for heterogeneous CPU+GPU decoding. In this paper, a highly parallel HEVC decoder for heterogeneous CPU+GPU system is proposed. It exploits available parallelism in HEVC decoding on the CPU, GPU, and between the CPU and GPU devices simultaneously. On top of that, different workload balancing schemes can be selected according to the devoted CPU and GPU computing resources. Furthermore, an energy optimized solution is proposed by tuning GPU clock rates. Results show that the proposed decoder achieves better performance than the state-of-the-art CPU decoder, and the best performance among the workload balancing schemes depends on the available CPU and GPU computing resources. In particular, with an NVIDIA Titan X Maxwell GPU and an Intel Xeon E5-2699v3 CPU, the proposed decoder delivers 167 frames per second (fps) for Ultra HD 4K videos, when four CPU cores are used. Compared to the state-of-the-art CPU decoder using four CPU cores, the proposed decoder gains a speedup factor of 2.2×. When decoding performance is bounded by the CPU, a system wise energy reduction up to 36% is achieved by using fixed (and lower) GPU clocks, compared to the default dynamic clock settings on the GPU.
Biao Wang 0001, Diego F. de Souza, Mauricio Alvarez-Mesa, Chi Ching Chi, Ben H. H. Juurlink, Aleksandar Ilic, Nuno Roma, Leonel Sousa
Signal Process. Image Commun.7
2018 Stream data prefetcher for the GPU memory interface
Nuno Neves 0002, Pedro Tomás, Nuno Roma
J. Supercomput.3
2017 Energy-efficient motion estimation with approximate arithmetic
abstract
Energy efficiency has become a primary concern in the design of multimedia digital systems, particularly when targeting mobile devices. Approximate computing is a highly promising approach to address this challenge. This paper presents an architectural exploration in a variable block size motion estimation (VBSME) architecture using imprecise Lower-Part-OR Adders (LOA). These adders were applied to Sum of Absolute Differences units (SAD) in order to reduce the energy consumption while introducing a minimum impact on the coding efficiency. Three VBSME architectures with LOA operators were developed by considering different imprecision levels. The conducted evaluations, performed using the High-Efficiency Video Coding standard (HEVC) reference software, showed that this technique introduces a negligible impact on the coding efficiency (between 0.6% and 2.5% increase of the BD-Rate). Nevertheless, when the designed architectures were synthesized for a 45nm standard cells technology, significant power savings were observed (between 7% and 11.5%, depending on the used LOA version), demonstrating the viability and significant gains of the proposed approach.
Roger Endrigo Carvalho Porto, Luciano Volcan Agostini, Bruno Zatt, Marcelo Schiavon Porto, Nuno Roma, Leonel Sousa
MMSP5
2017 GHEVC: An Efficient HEVC Decoder for Graphics Processing Units
abstract
The high compression efficiency that is provided by the high efficiency video coding (HEVC) standard comes at the cost of a significant increase of the computational load at the decoder. Such an increased burden is a limiting factor to accomplish real-time decoding, specially for high definition video sequences (e.g., Ultra HD 4K). In this scenario, a highly parallel HEVC decoder for the state-of-the-art graphics processor units (GPUs) is presented, i.e., GHEVC. Contrasting to our previous contributions, the data-parallel GHEVC decoder integrates the whole decompression pipeline (except for the entropy decoding), both for intra- and interframes. Furthermore, its processing efficiency was highly optimized by keeping the decompressed frames in the GPU memory for subsequent inter frame prediction. The proposed GHEVC decoder is fully compliant with the HEVC standard, where explicit synchronization points ensure the correct HEVC module execution order. Moreover, the GPU-based HEVC decoder is experimentally evaluated for different GPU devices, an extensive range of recommended HEVC configurations and video sequences, where an average frame rate of 145, 318, and 605 frames per second for Ultra HD 4K, WQXGA, and Full HD, respectively, was obtained in the Random Access configuration with the NVIDIA GeForce GTX TITAN X GPU.
Diego F. de Souza, Aleksandar Ilic, Nuno Roma, Leonel Sousa
IEEE Trans. Multim.3
2017 Adaptive In-Cache Streaming for Efficient Data Management
abstract
The design of adaptive architectures is frequently focused on the sole adaptation of the processing blocks, often neglecting the power/performance impact of data transfers and data indexing in the memory subsystem. In particular, conventional address-based models, supported on cache structures to mitigate the memory wall problem, often struggle when dealing with memory-bound applications or arbitrarily complex data patterns that can be hardly captured by prefetching mechanisms. Stream-based techniques have proven to efficiently tackle such limitations, although not well-suited to handle all types of applications. To mitigate the limitations of both communication paradigms, an efficient unification is herein proposed, by means of a novel in-cache stream paradigm, capable of seamlessly adapting the communication between the address-based and stream-based models. The proposed morphable infrastructure relies on a new dynamic descriptor graph specification, capable of handling regular arbitrarily complex data patterns, which is able to improve the main memory bandwidth utilization through data reutilization and reorganization techniques. When compared with state-of-the-art solutions, the proposed structure offers higher address generation efficiency and achievable memory throughputs, and a significant reduction of the amount of data transfers and main memory accesses, resulting on average in 13 times system performance speedup and in 245 times energy-delay product improvement, when compared with the previous implementations.
Nuno Neves 0002, Pedro Tomás, Nuno Roma
IEEE Trans. Very Large Scale Integr. Syst.3
2016 Efficient HEVC decoder for heterogeneous CPU with GPU systems
abstract
The High Efficiency Video Coding (HEVC) standard provides higher compression efficiency than other video coding standards but at the cost of increased computational load, which makes it hard to achieve real-time encoding/decoding of high-resolution, high-quality video sequences. In this paper, we investigate how Graphics Processing Units (GPUs) can be employed to accelerate HEVC decoding. GPUs are known to provide massive processing capability for throughput computing kernels, but the HEVC entropy decoding kernel cannot be executed efficiently on GPUs. We therefore propose a complete HEVC decoding solution for heterogeneous CPU+GPU systems, in which the entropy decoder is executed on the CPU and the remaining kernels on the GPU. Furthermore, the decoder is pipelined such that the CPU and the GPU can decode different frames in parallel. The proposed CPU+GPU decoder achieves an average frame rate of 150 frames per second for Ultra HD 4K video sequences when four CPU cores are used with an NVIDIA GeForce Titan X GPU.
Biao Wang 0001, Mauricio Alvarez-Mesa, Chi Ching Chi, Ben H. H. Juurlink, Diego F. de Souza, Aleksandar Ilic, Nuno Roma, Leonel Sousa
MMSP7
2016 Multi-objective kernel mapping and scheduling for morphable many-core architectures
Nuno Neves 0002, Rui Ferreira Neves, Nuno Horta, Pedro Tomás, Nuno Roma
Expert Syst. Appl.5
2016 BowMapCL: Burrows-Wheeler Mapping on Multiple Heterogeneous Accelerators
abstract
The computational demand of exact-search procedures has pressed the exploitation of parallel processing accelerators to reduce the execution time of many applications. However, this often imposes strict restrictions in terms of the problem size and implementation efforts, mainly due to their possibly distinct architectures. To circumvent this limitation, a new exact-search alignment tool (BowMapCL) based on the Burrows-Wheeler Transform and FM-Index is presented. Contrasting to other alternatives, BowMapCL is based on a unified implementation using OpenCL, allowing the exploitation of multiple and possibly different devices (e.g., NVIDIA, AMD/ATI, and Intel GPUs/APUs). Furthermore, to efficiently exploit such heterogeneous architectures, BowMapCL incorporates several techniques to promote its performance and scalability, including multiple buffering, work-queue task-distribution, and dynamic load-balancing, together with index partitioning, bit-encoding, and sampling. When compared with state-of-the-art tools, the attained results showed that BowMapCL (using a single GPU) is 2 × to 7.5 × faster than mainstream multi-threaded CPU BWT-based aligners, like Bowtie, BWA, and SOAP2; and up to 4 × faster than the best performing state-of-the-art GPU implementations (namely, SOAP3 and HPG-BWT). When multiple and completely distinct devices are considered, BowMapCL efficiently scales the offered throughput, ensuring a convenient load-balance of the involved processing in the several distinct devices.
David Nogueira, Pedro Tomás, Nuno Roma
IEEE ACM Trans. Comput. Biol. Bioinform.3
2016 Adaptive Scheduling Framework for Real-Time Video Encoding on Heterogeneous Systems
abstract
To challenge real-time encoding of high-definition video sequences on heterogeneous desktop systems, a collaborative central processing units (CPU) + graphics processing unit (GPU) framework for interloop video encoding is proposed herein. The proposed framework considers the overall complexity of the collaborative interloop encoding as a unified optimization problem. Several functional blocks are integrated for simultaneous execution control, automatic data access management, performance characterization, and adaptive scheduling and load balancing. These blocks aim at fully exploiting the performance of heterogeneous devices, asymmetric bandwidth of communication links, and several levels of concurrency between computation and communication. To support a wide range of CPU and GPU architectures, a specific encoding library is developed with highly optimized algorithms for all interloop modules. The experimental results show that the proposed framework allows achieving a real-time encoding of full high-definition sequences in several CPU + GPU systems. It also delivers performance improvements of up to 61.2% over the state-of-the-art solution, while outperforming individual GPU and quad-core CPU executions by more than 2 and 5 times, respectively.
Aleksandar Ilic, Svetislav Momcilovic, Nuno Roma, Leonel Sousa
IEEE Trans. Circuits Syst. Video Technol.3
2015 Fast and Scalable Thread Migration for Multi-core Architectures
abstract
Heterogeneous computing is a promising approach to tackle the thermal, power and energy constraints posed by modern desktop and embedded computing systems. However, by also allowing the migration of application threads to the most appropriate cores, significant performance gains and energy efficiency levels can also be attained. Nevertheless, the considerably large overheads usually imposed by software-based thread migration procedures only allow exploiting migrations at a coarse-grained level, thus limiting the effectiveness of using such techniques. Accordingly, this paper proposes a fast and efficient hardware-based thread migration mechanism that can be easily plugged-in into any core architecture. To minimize the thread migration overhead and latency, the proposed approach considers both soft-and hard-migration procedures, and adopts a conventional "most recently used" prediction scheme to identify the cache blocks that should be migrated along with the thread context. Experimental results show that the proposed scheme is lightweight and requires limited hardware resources, while allowing to attain migration latencies below 100 clock cycles and to reduce post-migration overheads in up to 60%, making it particularly appropriate for exploiting short-lived application phases.
Miguel Rodrigues 0002, Nuno Roma, Pedro Tomás
EUC2
2015 Efficient data-stream management for shared-memory many-core systems
abstract
The design of most high-performance and heterogeneous processing platforms is usually solely focused on the computational part, while neglecting the power/performance impact of the data-management infrastructures. Moreover, such systems often struggle to achieve their potential performance when applications require fetching data with complex memory access patterns. To overcome these issues, a energy-efficient stream-based data-management infrastructure is herein proposed, relying on a novel tree-based descriptor specification. Such descriptors are decoded by a Descriptor Tree Controller (DTC) architecture, which allows simple and efficient management of arbitrarily complex memory access patterns. Moreover, a Stream Management Engine (SME) ensures energy-efficient data-reutilization through the application of automatic stream rerouting, splitting and merging techniques. The obtained results show that the proposed DTC architecture is capable of a highly efficient complex data-pattern generation, while significantly reducing the size occupied by the pattern description, when compared with state-of-the-art approaches. By also enabling the deployment of data-reuse techniques, a reduction of up to 85× in the number of accesses to the main shared memory is achieved, resulting in a decrease as high as 475× in the observed energy consumption.
Nuno Neves 0002, Pedro Tomás, Nuno Roma
FPL3
2015 Towards GPU HEVC intra decoding: Seizing fine-grain parallelism
abstract
To satisfy the growing demands on real-time video decoders for high frame resolutions, novel GPU parallel algorithms are proposed herein for fully compliant HEVC de-quantization, inverse transform and intra prediction. The proposed algorithms are designed to fully exploit and leverage the fine grain parallelism within these computationally demanding and highly data dependent modules. Moreover, the proposed approaches allow the efficient utilization of the GPU computational resources, while carefully managing the data accesses in the complex GPU memory hierarchy. The experimental results show that the real-time processing is achieved for all tested sequences and the most demanding QP, while delivering average fps of 118.6, 89.2 and 49.7 for Full HD, 2160p and Ultra HD 4K sequences, respectively.
Diego F. de Souza, Aleksandar Ilic, Nuno Roma, Leonel Sousa
ICME3
2015 High performance IP core for HEVC quantization
abstract
A new class of quantization architectures suitable for the realization of high performance and hardware efficient forward, inverse and unified quantizers for HEVC is presented. The proposed structures are based on a highly flexible and optimized integer datapath that can be configured to provide several pipelined and non-pipelined implementations, offering distinct trade-offs between performance and hardware cost, which makes them highly suitable for most video coding application domains. The experimental results obtained using a 90 nm CMOS process show that the proposed class of quantization architectures is able to process 4k UHDTV video sequences in real-time (3840 × 2160 @ 30fps), with a power consumption as low as 3.9 mW when the unified architecture is operated at 374 MHz.
Tiago Dias 0001, Nuno Roma, Leonel Sousa
ISCAS2
2015 Run-Time Machine Learning for HEVC/H.265 Fast Partitioning Decision
abstract
A novel fast Coding Tree Unit partitioning for HEVC/H.265 encoder is proposed in this paper. This method relies on run-time trained neural networks for fast Coding Units splitting decisions. Contrasting to state-of-the-art solutions, this method does not require any pre-training and provides a high adaptivity to the dynamic changes in video contents. By an efficient sampling strategy and a multi-thread implementation, the presented technique successfully mitigates the computational overhead inherent to the training process on both the overall processing performance and on the initial encoding delay. The experiments show that the proposed method successfully reduces the HEVC/H.265 encoding time for up to 65% with negligible rate-distortion penalties.
Svetislav Momcilovic, Nuno Roma, Leonel Sousa, Ivan Z. Milentijevic
ISM2
2015 Multi-kernel Auto-Tuning on GPUs: Performance and Energy-Aware Optimization
abstract
Prompted by their very high computational capabilities and memory bandwidth, Graphics Processing Units (GPUs) are already widely used to accelerate the execution of many scientific applications. However, programmers are still required to have a very detailed knowledge of the GPU internal architecture when tuning the kernels, in order to improve either performance or energy-efficiency. Moreover, different GPU devices have different characteristics, moving a kernel to a different GPU typically requires re-tuning the kernel execution, in order to efficiently exploit the underlying hardware. The procedure proposed in this work is based on real-time kernel profiling and GPU monitoring and it automatically tunes parameters from several concurrent kernels to maximize the performance or minimize the energy consumption. Experimental results on NVIDIA GPU devices with up to 4 concurrent kernels show that the proposed solution achieves near optimal configurations. Furthermore, significant energy savings can be achieved by using the proposed energy-efficiency auto-tuning procedure.
João Guerreiro 0004, Aleksandar Ilic, Nuno Roma, Pedro Tomás
PDP3
2015 Implementation and performance analysis of efficient index structures for DNA search algorithms in parallel platforms
abstract
Summary Because of the large datasets that are usually involved in deoxyribonucleic acid (DNA) sequence alignment, the use of optimal local alignment algorithms (e.g., Smith–Waterman) is often unfeasible in practical applications. As such, more efficient solutions that rely on indexed search procedures are often preferred to significantly reduce the time to obtain such alignments. Some data structures that are usually adopted to build such indexes are suffix trees, suffix arrays, and the hash tables of q‐mers. This paper presents a comparative analysis of highly optimized parallel implementations of index‐based search algorithms using these three distinct data structures, considering two different parallel platforms: a homogeneous multi‐core central processing unit (CPU) and a NVidia Fermi graphics processing unit (GPU). Contrasting to what happens with CPU implementations, the obtained experimental results reveal that GPU implementations clearly favor the suffix arrays, because of the achieved performance in terms of memory accesses. Furthermore, the results also reveal that both the suffix trees and suffix arrays outperform the hash tables of q‐mers when dealing with the largest datasets. When compared with a quad‐core CPU, the results demonstrate the possibility to achieve speedups as high as 65 with the GPU when considering a suffix‐array index, thus making it an adequate choice for high‐performance bioinfomatics applications.Copyright © 2012 John Wiley & Sons, Ltd.
Nuno Sebastião, Gustavo Encarnação, Nuno Roma
Concurr. Comput. Pract. Exp.3
2015 Multicore SIMD ASIP for Next-Generation Sequencing and Alignment Biochip Platforms
abstract
Targeting the development of new biochip platforms capable of autonomously sequencing and aligning biological sequences, a new multicore processing structure is proposed in this manuscript. This multicore structure makes use of a shared memory model and multiple instantiations of a novel application-specific instruction-set processor (ASIP) to simultaneously exploit both fine and coarse-grained parallelism and to achieve high performance levels at low-power consumption. The proposed ASIP is built by extending the instruction set architecture of a synthesizable processor, including both general and special-purpose single-instruction multiple-data instructions. This allows an efficient exploitation of fine-grained parallelism on the alignment of biological sequences, achieving over$30\times $speedup when compared with sequential algorithmic implementations. The complete system was prototyped on different field-programmable gate array platforms and synthesized with a 90-nm CMOS process technology. Experimental results demonstrate that the multicore structure scales almost linearly with the number of instantiated cores, achieving performances similar to a quad-core Intel Core i7 3820 processor, while using$25\times $less energy.
Nuno Neves 0002, Nuno Sebastião, David Martins de Matos, Pedro Tomás, Paulo F. Flores, Nuno Roma
IEEE Trans. Very Large Scale Integr. Syst.6
2014 Cooperative CPU+GPU deblocking filter parallelization for high performance HEVC video codecs
abstract
Heterogeneous platforms integrating several CPU cores and GPU accelerators have established in several application domains, from desktop, server and mobile. To take full advantage of such platforms, video encoders/decoders have to exploit a broader design space, by cooperatively executing in all the available CPU and GPU cores. To attain such objective, three novel contributions that aim the exploitation of the maximum parallelism level in an HEVC deblocking filter are presented: i) a highly optimized CPU parallel implementation, which outperforms the current state of the art; ii) the first known GPU implementation of the HEVC deblocking filter; and iii) an hybrid and load-balanced CPU+GPU implementation, where all the available resources cooperatively execute, in order to maximize the attained performance. The obtained experimental results demonstrated the ability to achieve processing times as low as 0.8 ms and 0.5 ms to filter 1080p I-type and B-type frames, respectively, corresponding to speedup factors as high as 17 and 9.
Diego F. de Souza, Nuno Roma, Leonel Sousa
ICASSP2
2014 Collaborative inter-prediction on CPU+GPU systems
abstract
In this paper we propose an efficient method for collaborative H.264/AVC inter-prediction in heterogeneous CPU+GPU systems. In order to minimize the overall encoding time, the proposed method provides stable and balanced load distribution of the most computationally demanding video encoding modules, by relying on accurate and dynamically built functional performance models. In an extensive RD analysis, an efficient temporary dependent prediction of the search area center is proposed, which allows dependency-aware workload partitioning and efficient GPU parallelization, while preserving high compression efficiency. The proposed method also introduces efficient communication-aware techniques, which maximize data reusing, and decrease the overhead of expensive data transfers in collaborative video encoding. The experimental results show that the proposed method is able of achieving real-time video encoding for very demanding video coding parameters, i.e. full HD video format, 64×64 pixels search area and the exhaustive motion estimation.
Svetislav Momcilovic, Aleksandar Ilic, Nuno Roma, Leonel Sousa
ICIP3
2014 FEVES: Framework for Efficient Parallel Video Encoding on Heterogeneous Systems
abstract
Lead by high performance computing potential of modern heterogeneous desktop systems and predominance of video content in general applications, we propose herein an autonomous unified video encoding framework for hybrid multi-core CPU and multi-GPU platforms. To fully exploit the capabilities of these platforms, the proposed framework integrates simultaneous execution control, automatic data access management, and adaptive scheduling and load balancing strategies to deal with the overall complexity of the video encoding procedure. These strategies consider the collaborative inter-loop encoding as a unified optimization problem to efficiently exploit several levels of concurrency between computation and communication. To support a wide range of CPU and GPU architectures, a specific encoding library is developed with highly optimized algorithms for all inter-loop modules. The obtained experimental results show that the proposed framework allows achieving a real-time encoding of full high-definition sequences in the state-of-the-art CPU+GPU systems, by outperforming individual GPU and quad-core CPU executions for more than 2 and 5 times, respectively.
Aleksandar Ilic, Svetislav Momcilovic, Nuno Roma, Leonel Sousa
ICPP3
2014 Cache-Oblivious parallel SIMD Viterbi decoding for sequence search in HMMER
abstract
BACKGROUND: HMMER is a commonly used bioinformatics tool based on Hidden Markov Models (HMMs) to analyze and process biological sequences. One of its main homology engines is based on the Viterbi decoding algorithm, which was already highly parallelized and optimized using Farrar's striped processing pattern with Intel SSE2 instruction set extension. RESULTS: A new SIMD vectorization of the Viterbi decoding algorithm is proposed, based on an SSE2 inter-task parallelization approach similar to the DNA alignment algorithm proposed by Rognes. Besides this alternative vectorization scheme, the proposed implementation also introduces a new partitioning of the Markov model that allows a significantly more efficient exploitation of the cache locality. Such optimization, together with an improved loading of the emission scores, allows the achievement of a constant processing throughput, regardless of the innermost-cache size and of the dimension of the considered model. CONCLUSIONS: The proposed optimized vectorization of the Viterbi decoding algorithm was extensively evaluated and compared with the HMMER3 decoder to process DNA and protein datasets, proving to be a rather competitive alternative implementation. Being always faster than the already highly optimized ViterbiFilter implementation of HMMER3, the proposed Cache-Oblivious Parallel SIMD Viterbi (COPS) implementation provides a constant throughput and offers a processing speedup as high as two times faster, depending on the model's size.
Miguel Ferreira, Nuno Roma, Luís M. S. Russo
BMC Bioinform.2
2014 Dynamic Load Balancing for Real-Time Video Encoding on Heterogeneous CPU+GPU Systems
abstract
The high computational demands and overall encoding complexity make the processing of high definition video sequences hard to be achieved in real-time. In this manuscript, we target an efficient parallelization and RD performance analysis of H.264/AVC inter-loop modules and their collaborative execution in hybrid multi-core CPU and multi-GPU systems. The proposed dynamic load balancing algorithm allows efficient and concurrent video encoding across several heterogeneous devices by relying on realistic run-time performance modeling and module-device execution affinities when distributing the computations. Due to an online adjustment of load balancing decisions, this approach is also self-adaptable to different execution scenarios. Experimental results show the proposed algorithm's ability to achieve real-time encoding for different resolutions of high-definition sequences in various heterogeneous platforms. Speed-up values of up to 2.6 were obtained when compared to the video inter-loop encoding on a single GPU device, and up to 8.5 when compared to a highly optimized multi-core CPU execution. Moreover, the proposed algorithm also provides an automatic tuning of the encoding parameters, in order to meet strict encoding constraints.
Svetislav Momcilovic, Aleksandar Ilic, Nuno Roma, Leonel Sousa
IEEE Trans. Multim.3
2013 BioBlaze: Multi-core SIMD ASIP for DNA sequence alignment
abstract
A new Application-Specific Instruction-set Processor (ASIP) architecture for biological sequences alignment is proposed in this manuscript. This architecture achieves high processing throughputs by exploiting both fine and coarse-grained parallelism. The former is achieved by extending the Instruction Set Architecture (ISA) of a synthesizable processor to include multiple specialized SIMD instructions that implement vector-vector and vector-scalar arithmetic, logic, load/store and control operations. Coarse-grained parallelism is achieved by using multiple cores to cooperatively align multiple sequences in a shared memory architecture, comprising proper hardware-specific synchronization mechanisms. To ease the programming, a compilation framework based on an adaptation of the GCC back-end was also implemented. The proposed system was prototyped and evaluated on a Xilinx Virtex-7 FPGA, achieving a 200MHz working frequency. A sequential and a state-of-theart SIMD implementations of the Smith-Waterman algorithm were programmed in both the proposed ASIP and an Intel Core i7 processor. When comparing the achieved speedups, it was observed that the proposed ISA achieves a 40x speedup, which contrasts with the 11x speedup provided by SSE2 in the Intel Core i7 processor. The scalability of the multi-core system was also evaluated and proved to scale almost linearly with the number of cores.
Nuno Neves 0002, Nuno Sebastião, Andre Patricio, David Martins de Matos, Pedro Tomás, Paulo F. Flores, Nuno Roma
ASAP7
2013 Scalable and high throughput biosensing platform
abstract
A novel multi-channel high performance embedded system capable of high throughput biological analysis is proposed in this paper. Despite other integrated lab-on-chip solutions based on magnetoresistive biochips have already been developed, they lack the scalability and computational resources to cope with new biochip designs featuring more than 1000 sensors. A new configurable acquisition and processing architecture is proposed, combining dedicated coprocessors to perform signal filtering and other computational demanding tasks, with a central processor controlling the whole system. The mapping of the architecture into a Zynq SoC demonstrated its ability to support 8 times more sensors, while ensuring a sampling frequency 1000+ times higher than the previous platforms. Furthermore, the Zynq reconfiguration abilities provide a mechanism to adapt the processing and maximize the biological sensitivity.
José M. Leitão, José A. Germano, Nuno Roma, Ricardo Chaves, Pedro Tomás
FPL3
2013 HotStream: Efficient Data Streaming of Complex Patterns to Multiple Accelerating Kernels
abstract
Designing accelerating kernels is a comprehensive task that requires efficient coupling of hardware and software. In particular, the structures responsible for handling data transfers in multi-core accelerator-based systems play a crucial role in the resulting performance. This paper proposes a data streaming accelerator framework that provides efficient data management facilities that are easily tailored for any application and data pattern. This is achieved through an innovative and fully programmable data management structure, implemented with two granularity levels. The obtained results show that the proposed framework is capable of efficient address generation and data fetch for complex streaming data patterns, while significantly reducing the size occupied by the pattern description. A large matrices multiplication case-study, based on a streaming architecture with four sub-block multiplication cores, demonstrates that, by enabling data re-use, the proposed framework increases the available bandwidth by 4.2x, resulting in a performance speedup of 2.1x. Furthermore, it reduces the Host memory requirements and its intervention by more than 40x.
Sergio Paiagua, Frederico Pratas, Pedro Tomás, Nuno Roma, Ricardo Chaves
SBAC-PAD4
2013 Configurable and scalable class of high performance hardware accelerators for simultaneous DNA sequence alignment
abstract
SUMMARY A new class of efficient and flexible hardware accelerators for DNA local sequence alignment based on the widely used Smith–Waterman algorithm is proposed in this paper. This new class of accelerating structures exploits an innovative technique that tracks the origin coordinates of the best alignment to allow a significant reduction of the size of the dynamic programming matrix that needs to be recomputed during the subsequent traceback phase, providing a considerable reduction of the resulting time and memory requirements. The significant performance of the enhanced class of accelerators is attained by also providing support for an additional level of parallelism: the capability to concurrently align several query sequences with one or more reference sequences, according to the specific application requisites. Moreover, the accelerator class also includes specially designed processing elements that improve the resource usage when implemented in a Field Programmable Gate Array (FPGA), and easily provide several different configurations in an Application Specific Integrated Circuit (ASIC) implementation. Obtained results demonstrated that speedups as high as 278 can be obtained in ASIC accelerating structures. A FPGA‐based prototyping platform, operating at a 40 times lower clock frequency and incorporating a complete alignment embedded system, still provides significant speedups as high as 27, compared with a pure software implementation.Copyright © 2012 John Wiley & Sons, Ltd.
Nuno Sebastião, Nuno Roma, Paulo F. Flores
Concurr. Comput. Pract. Exp.2
2012 High Performance Unified Architecture for Forward and Inverse Quantization in H.264/AVC
abstract
A new high-performance and reduced hardware architecture for the computation of the H.264/AVC forward and inverse quantization operations is presented in this paper. This architecture is based on a highly flexible processing structure that is suitable for very efficient implementations using both FPGA and ASIC technologies. Moreover, it offers several different configurations, in order to provide different trade-offs in terms of performance and hardware cost. Experimental results concerning implementations using a Xilinx Virtex-5 FPGA and a 90 nm CMOS process from UMC demonstrated that the proposed architecture can be used to compute, in real-time, the forward and inverse quantization operations for videos with resolutions up to the Digital Cinema format (4096x2048 @ 30fps).
Tiago Dias 0001, Luis Rosario, Nuno Roma, Leonel Sousa
DSD3
2012 System-level prototyping framework for heterogeneous multi-core architecture applied to biological sequence analysis
abstract
An event-driven prototyping and simulation framework to support the design and early development stages of an heterogeneous multi-core processing architecture is presented in this manuscript. The main focus of this parallel structure is to efficiently execute a set of widely used bio-informatics algorithms for DNA sequences alignment and processing. The conceived framework was entirely developed using the SystemC description language and allows a full parametrization of the prototyped multi-core architecture, such as the amount of computing nodes, the effective alignment throughput of each node, and the capacity and access time of the memory devices. The presented experimental results demonstrate that the conceived framework provides the system designer with a very useful preliminary evaluation of the prototyped architecture. In particular, the included evaluation demonstrates the relation between the number and performance of the computing nodes and the resulting alignment performance gain (speedup), as well as the inherent bus contention losses in the shared resources (bus and shared memory).
Nuno Roma, Pedro Magalhães
RSP1
2012 Integrated Hardware Architecture for Efficient Computation of the $n$-Best Bio-Sequence Local Alignments in Embedded Platforms
abstract
A flexible hardware architecture that implements a set of new and efficient techniques to significantly reduce the computational requirements of the commonly used Smith-Waterman sequence alignment algorithm is presented. Such innovative techniques use information gathered by the hardware accelerator during the computation of the alignment scores to constrain the size of the subsequence that has to be post-processed in the traceback phase using a general purpose processor (GPP). Moreover, the proposed structure is also capable of computing then-best local alignments according to the Waterman-Eggert algorithm, becoming the first hardware architecture that is able to simultaneously evaluate then-best alignments of a given sequence pair, by incorporating a set of ordering units that work in parallel with the systolic array. A complete alignment system was developed and implemented in a Virtex-4 FPGA, by integrating the proposed accelerator architecture with a Leon3 GPP. The obtained experimental results demonstrate that the proposed system is flexible and allows the alignment of large sequences in memory constrained systems. As an example, a speedup of 17 was obtained with the conceived system when compared with a regular implementation of the LALIGN35 program running on an Intel Core2 Duo processor running at a 40 × higher frequency.
Nuno Sebastião, Nuno Roma, Paulo F. Flores
IEEE Trans. Very Large Scale Integr. Syst.2
2011 A tutorial overview on the properties of the discrete cosine transform for encoded image and video processing
Nuno Roma, Leonel Sousa
Signal Process.1
2010 A Parallel Programming Framework for Multi-core DNA Sequence Alignment
abstract
A new parallel programming framework for DNA sequence alignment in homogeneous multi-core processor architectures is proposed. Contrasting with traditional coarse-grained parallel approaches, that divide the considered database in several smaller subsets of complete sequences to be aligned with the query sequence, the presented methodology is based on a slicing procedure of both the query and the database sequence under consideration in several tiles/chunks that are concurrently processed by the several cores available in the multi-core processor. The obtained experimental results have proven that significant accelerations of traditional biological sequence alignment algorithms can be obtained, reaching a speedup that is linear with the number of available processing cores and very close to the theoretical maximum.
Tiago Jose Barreiros Martins de Almeida, Nuno Roma
CISIS2
2010 p264: open platform for designing parallel H.264/AVC video encoders on multi-core systems
abstract
A highly modular and configurable platform for designing parallel H.264 video encoders on multi-core processors is presented. Departing from the H.264/AVC reference software, preliminary optimizations were conducted and new data structures were developed, in order to support the encoder's parallelization and to confer the developed platform with a flexible, user configurable and highly scalable characteristics in what concerns the number of available cores to be used in the target concretization. After a careful assessment using different instantiations of the platform, the experimental results have shown that significant and close to linear speedups in what concerns the achieved frame-rate can be obtained, by simultaneously exploiting the several different parallelization models that are made available by this platform.
António Rodrigues, Nuno Roma, Leonel Sousa
NOSSDAV2
2009 Distributed Software Platform for Automation and Control of General Anaesthesia
abstract
A parallel computer architecture and a distributed software platform for automation and control of general anesthesia is proposed in this paper. The system is a prototype research platform, intended to help on the development, simulation and test of new control algorithms for general anesthesia. It must be safe when used in real tests and flexible enough to allow the integration of new software modules. The system is composed by two computers, with the specific tasks of anesthesia control and process supervision. The platform makes use of TANGO, a specialized framework for distributed control systems, which provides software mechanisms useful to fulfill the project requirements. The architecture and the set of mechanisms proposed in this paper provide a high degree of flexibility to research on control algorithms, while ensuring the safeness of the whole procedure.
Gesner Passos, Nuno Roma, Bertinho Andrade da Costa, Leonel Sousa, João Miranda Lemos
ISPDC2
2008 Application Specific Programmable IP Core for Motion Estimation: Technology Comparison Targeting Efficient Embedded Co-Processing Units
abstract
The implementation of a recently proposed IP core of an efficient motion estimation co-processor is considered. Some significant functional improvements to the base architecture are proposed, as well as the presentation of a detailed description of the interfacing between the co-processor and the main processing unit of the video encoding system. Then, a performance analysis of two distinct implementations of this IP core is presented, considering two different target technologies: a high performance FPGA device, from the Xilinx Virtex-II Pro family, and an ASIC based implementation, using a 0.18um CMOS StdCell library. Experimental results have shown that the two alternative implementations have quite similar performance levels and allow the estimation of motion vectors in real-time.
Nuno Sebastião, Tiago Dias 0001, Nuno Roma, Paulo F. Flores, Leonel Sousa
DSD3
2006 Application Specific Instruction Set Processor for Adaptive Video Motion Estimation
abstract
Motion estimation is the most demanding operation of a video encoder, corresponding to at least 80% of the overall computational cost. With the proliferation of portable handheld devices that support digital video coding, data-adaptive motion estimation algorithms have been required to dynamically configure the search pattern not only to avoid unnecessary computations and memory accesses but also to save energy. This paper proposes an application specific instruction set processor (ASIP) to implement data-adaptive motion estimation algorithms, that is characterized by a specialized data-path and minimum and optimized instruction set. Due to its low-power nature, this architecture is specially adequate to develop motion estimators for portable, mobile and battery supplied devices. A cycle-based accurate simulator was also developed for the proposed ASIP and fast and data-adaptive search algorithms have been implemented, namely, the four-step search and the motion vector field adaptive search algorithms. Based on the proposed ASIP and the considered adaptive algorithms, several motion estimators were synthesized in 0.13mum CMOS technology. Experimental results show that very-low power adaptive motion estimators have been achieved to encode QCIF video sequences
Svetislav Momcilovic, Tiago Dias 0001, Nuno Roma, Leonel Sousa
DSD3
2005 Least squares motion estimation algorithm in the compressed DCT domain for H.26x/MPEG-x video sequences
abstract
A new compressed domain motion estimation algorithm that makes use of the DCT coefficients directly obtained from the H.26x or MPEG-x video stream is presented. The proposed algorithm is based on an iterative scheme that computes the new motion vectors by applying a least squares estimation technique. To reduce its computational effort, the algorithm may consider only an arbitrary subset of non-null DCT coefficients. The performance of the algorithm was assessed in a DCT domain H.263 video transcoder, where the obtained motion vectors provided the means to significantly enhance the quality of the temporal prediction scheme with a consequent reduction of the required bit-rate.
Nuno Roma, Leonel Sousa
AVSS1
2003 Customisable Core-Based Architectures for Real-Time Motion Estimation on FPGAs
Nuno Roma, Tiago Dias 0001, Leonel Sousa
FPL1
2003 Fast transcoding architectures for insertion of non-regular shaped objects in the compressed DCT-domain
Nuno Roma, Leonel Sousa
Signal Process. Image Commun.1
2002 Efficient and configurable full-search block-matching processors
abstract
Efficient VLSI architectures for motion estimation using the full-search block-matching algorithm are proposed in this paper. These structures are based on an improved and more efficient two-dimensional single-array architecture with minimum latency, maximum throughput, and full utilization of the hardware resources. This optimized architecture is extended to a class of fully parameterizable multiple array architectures that combine both pipelining and parallel processing techniques and provide the ability to configure the processors according to the setup parameters, the processing time and the circuit area specified limits. The development of a single-array processor in a single-chip based on a 0.25-/spl mu/m CMOS technology process proves the practical interest of the proposed architecture for implementing real-time motion estimators.
Nuno Roma, Leonel Sousa
IEEE Trans. Circuits Syst. Video Technol.1
1999 Low-power array architectures for motion estimation
abstract
This paper proposes new efficient low-power systolic architectures for full search-block matching (FS-BM) motion estimation. These architectures allow one to eliminate unnecessary computations, reducing the power consumption while preserving the optimal solution and the throughput. The new and traditional systolic architectures for motion estimation are compared with respect to required hardware and power consumption.
Leonel Sousa, Nuno Roma
MMSP2