EDBT 2026 Demo / reviewers in the wild / expert
Nuno Neves 0002
dblp:26/2819
· DBLP profile ↗
20ranked-venue papers
8as first author
13since 2021 · last 2026
0000-0003-0628-2259ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 16 · 7 first-author · 10 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | S2VEC: Compiler-Driven Stream Specialization for Linearized VectorizationabstractThe performance benefits of vectorization are often limited by data movement, which remains a dominant bottleneck in modern processors. To mitigate this bottleneck, data stream-based mechanisms have been recently proposed, allowing general-purpose architectures to specialize memory accesses, leading to a reduction of load-to-use latency and loop code optimizations. However, compiler support for stream-based execution remains limited. Accordingly, this paper presents S2VEC, an LLVM-based compilation toolchain that automatically extracts, represents, and optimizes memory access patterns as data streams to enable stream-specialized vector execution. At its core, the Stream-Dataflow IR models computation as a dataflow over parameterized data streams, bridging the gap between conventional compiler infrastructures and stream-oriented architectures. By transforming complex and indirect memory accesses into linearized stream representations, S2VEC simplifies vectorization and expands its applicability to complex memory access patterns. When targeting a RISC-V stream-vector ISA extension, S2VEC demonstrates improved vectorization coverage and produces code comparable to hand-optimized implementations. Conducted evaluations on a gem5-based in-order stream-vector model show an overall speedup of 5.9 × over an equivalently provisioned RVV-based in-order core and 1.5 × over a wide out-of-order processor, demonstrating the performance benefits of compiler-driven stream specialization. Luís Crespo, Gabriel Falcão Paiva Fernandes, Pedro Tomás, Nuno Roma, Nuno Neves 0002 |
ICS | 6 |
| 2026 | Reconfigurable FPU With Precision Auto-Tuning for Next-Generation Transprecision ComputingabstractRecent advances in process technology have shifted the research focus from strict raw scaling to the conception of energy-efficient computing units, capable of adapting to the target application precision requirements. A key opportunity lies in floating-point arithmetic, where traditional fixed-precision formats (32/64-bit) often impose unnecessary resources and costs in both performance and power. To address this problem, this manuscript introduces an Automatic Precision Floating-point Unit (APFU) that extends prior reconfigurable designs by incorporating a hardware controller capable of autonomously tuning the operand precision and vectorization levels at runtime. The APFU supports IEEE-754 (double, single, half-precision), bfloat16, and DLFloat formats, and exploits unused datapath capacity to increase the throughput via vector operations. Unlike previous approaches, the proposed APFU introduces a runtime controller that autonomously determines the operand precision and vectorization levels based on operand characteristics and execution mode. The presented experimental evaluation considers an implementation using a 28-nm UMC technology, achieving up to 152 GOPS/W, and demonstrates a detailed analysis across operating frequencies, vectorization modes, and runtime precision adjustment. Guilherme Dias, Luís Crespo, Timo Schlachter, Marc Andre Heller, Jens Krüger 0004, Pedro Tomás, Nuno Roma, Nuno Neves 0002 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 8 |
| 2026 | Real-Time ORB Accelerator for Embedded FPGA-Based SoCs With ROS IntegrationabstractAs computer vision continues to expand across various application domains—including localization, mapping, object recognition, and 3-D reconstruction—feature extraction methods, such as oriented FAST and rotated BRIEF (ORB), have gained widespread adoption due to their rotation and scale invariance. However, existing efforts to accelerate these techniques through hardware implementations faced different challenges related to resource and power consumption demands, limiting their feasibility for low-power embedded devices. Accordingly, this article proposes a new scalable and efficient ORB accelerator, designed for low-power and resource-constrained environments. It introduces a novel and efficient architecture that exploits quantization of the feature orientation angle into discrete rotation sectors. A complete robot operating system (ROS) node based on the proposed ORB accelerator is also deployed, providing seamless integration with other computer vision-enabled systems. When compared to other state-of-the-art solutions, the proposed system, implemented on an embedded system-on-chip (SoC) with a low-cost field-programmable gate array (FPGA), offers energy efficiency improvements between$6.7\times $and$16.2\times $, while requiring fewer hardware resources. Andre Costa, José Duarte Lopes, Pedro Tomás, Nuno Roma, Nuno Neves 0002 |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2025 | RVEBS: Event-Based Sampling on RISC-VabstractAs RISC-V ISA continues to gain traction for both embedded and high-performance computing, the demand for advanced monitoring tools has become critical to fine-tuning the applications' performance. Current RISC-V hardware performance monitors already provide basic event counting but lack sophisticated features like event-based sampling, which are available in more established architectures such as x86 and ARM. This paper presents the first RISC-V Event-Based Sampling (RVEBS) system for comprehensive performance monitoring and application profiling. The proposed system builds upon existing RISC-V specifications, incorporating necessary modifications to enable the desired functionality. It also presents an OpenSBI extension to provide privileged software access to newly implemented control status registers that manage the sampling process. An implementation use case based on an OpenPiton processor featuring a CVA6 core on 28nm CMOS technology was presented. The results indicate that the proposed scheme is lightweight, highly accurate, and does not impact the processor's critical path while maintaining minimal impact on overall application performance. Tiago Rocha, Nuno Neves 0002, Nuno Roma, Pedro Tomás, Leonel Sousa |
DATE | 2 |
| 2025 | Real-Time ORB Accelerator with ROS Integration for Embedded FPGA SoCsabstractAs computer vision continues to expand across various application domains - including localisation, mapping, object recognition, and 3D reconstruction - feature extraction methods such as Oriented FAST and Rotated BRIEF (ORB) gained widespread adoption due to their rotation and scale invariance. However, existing efforts to accelerate these techniques through hardware implementations faced challenges related to high resource and power consumption demands, limiting their feasibility for low-power embedded devices. Accordingly, this paper proposes a new scalable and efficient ORB accelerator, designed for low-power resource-constrained environments. It introduces a novel resource-efficient architecture that exploits quantisation of the feature orientation angle into discrete rotation sectors. A complete ROS node based on the proposed ORB accelerator is also deployed, providing seamless integration with other computer vision-enabled systems. When compared to other state-of-the-art solutions, the proposed system, implemented on an embedded System-on-Chip (SoC) with a low-cost FPGA, offers between 6.7x and 16.2x energy efficiency improvements, while requiring fewer hardware resources. Andre Costa, Pedro Tomás, Nuno Roma, Nuno Neves 0002 |
ISCAS | 4 |
| 2025 | Stream-Driven Acceleration for Embedded RISC-V SoCsabstractThis paper proposes a stream-driven computational model that expands the recent stream vectorization paradigm into a full dataflow-driven computing model. It exploits spatial computation and time-multiplexing, while relying on streaming engines implementing the RISC-V UVE specification to manage data access patterns, thus streamlining memory operations and reducing latency. By abstracting the kernel loops into stream data-flow graphs and mapping them onto a processing element array, the conceived accelerator architecture exploits both spatial and temporal parallelism across a wide range of computational tasks. Experimental results, conducted on a synthesized 7nm implementation, demonstrate the proposed model’s potential to develop high-efficiency accelerators in data-intensive applications, offering performance gains of up to 6× compared with an ARM Cortex-A53 CPU with NEON and 15× compared with a scalar Rocket RISC-V CPU, along with 3.86× energy efficiency improvements. João Maia, Ana Silveira, Gonçalo Midões, Nuno Neves 0002, Pedro Tomás, Nuno Roma |
ISCAS | 4 |
| 2025 | MIDAS: A Mapping Infrastructure for Configurable, Data-Streaming Based Domain Specific AcceleratorsabstractAs computational demands continue to grow in key application domains, Domain-Specific Accelerators (DSAs) have become a promising solution for delivering high performance with improved area and energy efficiency. However, traditional memory address generation in DSAs often consumes valuable resources that could otherwise be used to enhance performance. Data streaming mechanisms address this issue by eliminating the need for address generation nodes, but the exploitation of this paradigm remains underutilized in existing DSA design toolchains. On the other hand, despite fast kernel mapping and early feedback metrics being regarded as critical features in design space exploration for DSAs, most publicly available toolchains are either too slow, provide feedback only after a full compilation, or lack direct support for data streaming. This paper introduces a new Mapping Infrastructure for Data-Streaming-Based Accelerators (MIDAS) designed to rapidly map kernels onto configurable processing arrays equipped with an integrated streaming engine. When used as a co-design exploration tool, it provides early feedback on multiple metrics, facilitating architectural optimisation and pruning, and leading to a considerable improvement of hardware and energy efficiency. Moreover, the compute-only Data Flow Graphs (DFGs) are also more amenable to vectorisation, enabling further performance gains from data streaming. The obtained experimental results demonstrate that MIDAS achieves mapping speeds up to 73x faster than state-of-the-art tools like CGRA-ME’s CLUMAP, while maintaining competitive mapping quality. Under the data streaming paradigm, the implemented array achieves up to 8x performance improvement on a 4x4 PE array configuration compared to conventional architectures, with significantly reduced resource usage. Martim Bento, Nuno Neves 0002, Pedro Tomás, Nuno Roma |
SBAC-PAD | 2 |
| 2025 | A Survey on Stream-Based Architectures: From Accelerators to CPUsabstractIn the past few years, there has been a renewed effort to advance general-purpose architectures. In particular, to deliver performance and energy efficiency advantages, several techniques have been applied based on new forms of specialization while maintaining usability. As a result, data movement and communication have become the primary bottlenecks in computer systems. To overcome this, one of the most recent breakthroughs has been the introduction of data streaming mechanisms, just like those used in accelerators, into modern general-purpose processors (GPPs). This article comprehensively reviews stream-based architectures, tracing their development from accelerator solutions to their recent adoption in GPPs. This survey starts by introducing the fundamental principles of stream specialization, followed by a taxonomy for memory accesses, and formal mathematical models to represent them as data streams. Then, it categorizes different topologies of data stream specialization and examines them from a compiler’s perspective. Some of the most representative architectures proposed in the past few years, including instruction set architecture (ISA) and streaming engines, are described, followed by a comparative analysis that highlights their key features and presents quantitative evaluations. Then, we discuss some open challenges and suggest directions for future research in stream-based architectures. Luís Crespo, Nuno Neves 0002, Pedro Tomás, Nuno Roma |
Proc. IEEE | 2 |
| 2023 | Supporting RISC-V Performance Counters Through Linux Performance Analysis ToolsabstractIncreased attention to RISC-V open Instruction Set Architecture (ISA), a base ISA with a variety of optional extensions, has fueled its move from embedded devices to the high-performance computing arena, with the proliferation of RISC-V-based accelerators. However, the absence of powerful performance monitoring tools often results in poorly optimized applications and, consequently, limited computing performance. While the RISC-V ISA already defines a hardware performance monitor (HPM), research and development on RISC-V-based devices have been more focused on architectures and compilers rather than tools to support monitoring performance. To overcome this limitation, a comprehensive set of extensions and modifications to the Performance analysis tools for Linux (perf/perf_events) are proposed in this paper, and a PAPI library interface is presented. These new extensions comprise not only the Linux kernel but also the OpenSBI interface, and aim to achieve full support for the RISC-V performance monitoring specification. The conducted testing and evaluation were carried out on a HiFive Unmatched board and on a CVA6 core, but the proposed extensions, and the corresponding implementation, are easily portable to other systems. Joao Mario Domingos, Tiago Rocha, Nuno Neves 0002, Nuno Roma, Pedro Tomás, Leonel Sousa |
ASAP | 3 |
| 2023 | Trading Performance, Power, and Area on Low-Precision Posit MAC Units for CNN TrainingabstractThe recently proposed Posit number system has been regarded as a particularly well-suited floating-point format to optimize the throughput and efficiency of low-precision computations in convolutional neural network (CNN) applications. In particular, the Posit format offers a balance between decimal accuracy and dynamic range, which results in a distribution of values that seems particularly interesting for deep learning applications. However, the adoption of the Posit still raises some concerns regarding hardware complexity, particularly when accounting for the overheads associated with the quire exact accumulator. Accordingly, this paper presents a holistic study on the model accuracy, performance, power, and area trade-offs when adopting low-precision Posit multiply-accumulate (MAC) units for the training of CNNs. In particular, 28nm ASIC implementations of a reference Posit MAC unit architecture demonstrate that the quire accounts for over 70% of the area and power utilization, and the obtained CNN training results showed that its use is only strictly required when considering mixed low-precision configurations. As a result, reducing the size of the quire results in an average reduction of area and power by 57% and 47%, without imposing visible training accuracy losses. Luís Crespo, Pedro Tomás, Nuno Roma, Nuno Neves 0002 |
SBAC-PAD | 4 |
| 2021 | HEDAcc: FPGA-based Accelerator for High-order Epistasis DetectionabstractThe manifestation of important genetic diseases is often a consequence of the interactions between Single Nucleotide Polymorphisms (SNPs), also known as epistasis. Detecting epistasis for high-order interactions results in a huge computational complexity, as the number of SNP combinations to be evaluated exponentially grows with the interaction order. To address this challenge, state-of-the-art exhaustive search-based methods for epistasis detection rely on GPUs and FPGAs to provide high-performance solutions tailored for a specific order (second and rarely third-order interactions) and/or specific data-set sizes. In this paper, a novel parameterizable architecture is proposed that enables the deployment of FPGA-based accelerators targeting any order of interactions and any data-set size. By relying on a set of algorithmic and architecture optimizations, the proposed accelerator showed to outperform current FPGA accelerators for second and third-order interactions by as much as 4.6× and 9.5×, respectively. The proposed solution also showed comparable performance to current GPGPU third-order implementations, while consuming up to 8.2× less energy. Finally, the proposed architecture allowed for the implementation of a fourth-order epistasis detection accelerator in an FPGA platform. Gaspar Ribeiro, Nuno Neves 0002, Sergio Santander-Jiménez, Aleksandar Ilic |
FCCM | 2 |
| 2021 | Unlimited Vector Extension with Data Streaming SupportabstractUnlimited vector extension (UVE) is a novel instruction set architecture extension that takes streaming and SIMD processing together into the modern computing scenario. It aims to overcome the shortcomings of state-of-the-art scalable vector extensions by adding data streaming as a way to simultaneously reduce the overheads associated with loop control and memory access indexing, as well as with memory access latency. This is achieved through a new set of instructions that pre-configure the loop memory access patterns. These attain accurate and timely data prefetching on predictable access patterns, such as in multidimensional arrays or in indirect memory access patterns. Each of the configured data streams is associated to a general- purpose vector register, which is then used to interface with the streams. In particular, iterating over a given stream is simply achieved by reading/writing to the corresponding input/output stream, as the data is instantly consumed/produced. To evaluate the proposed UVE, a proof-of-concept gem5 implementation was integrated in an out-of-order processor model, based on the ARM Cortex-A76, thus taking into consideration the typical speculative and out-of-order execution paradigms found in high- performance computing processors. The evaluation was carried out with a set of representative kernels, by assessing the number of executed instructions, its impact on the memory bus and its overall performance. Compared to other state-of-the-art solutions, such as the upcoming ARM Scalable Vector Extension (SVE), the obtained results show that the proposed extension attains average performance speedups over 2.4 × for the same processor configuration, including vector length. Joao Mario Domingos, Nuno Neves 0002, Nuno Roma, Pedro Tomás |
ISCA | 2 |
| 2021 | Compiler-Assisted Data Streaming for Regular Code StructuresabstractThe performance of modern processors is often limited by execution stalls resulting from long memory access latencies. Compile-time optimizations, deep cache hierarchies and prefetching mechanisms already provide significant performance gains, by performing memory accesses in parallel with computation. However, they are reaching a throughput improvement limit. Hence, new solutions that effectively exploit the memory access patterns to improve processing throughput are required. To achieve this objective, a new compiler-assisted data streaming method is proposed. It leverages static analysis and code transformations with an on-chip data streaming support as a viable alternative to prefetching mechanisms for regular code structures. Static analysis is used to identify and encode memory accesses with a dedicated representation. Then, a code transformation algorithm detaches data indexation and address calculation from computation, allowing for a significant code reduction. An on-chip data stream controller, attached to the L1 data cache, is used to autonomously generate memory accesses from the pattern representation and reorganize the data transfers in streams, with the aid of stream buffers. When compared with state-of-the-art prefetchers, the proposed solution provides up to 26 percent of code reduction, an IPC improvement of 2.4x, and an average performance improvement of 40 percent. Nuno Neves 0002, Pedro Tomás, Nuno Roma |
IEEE Trans. Computers | 1 |
| 2020 | Reconfigurable Stream-based Tensor Unit with Variable-Precision Posit ArithmeticabstractThe increased adoption of DNN applications drove the emergence of dedicated tensor computing units to accelerate multi-dimensional matrix multiplication operations. Although they deploy highly efficient computing architectures, they often lack support for more general-purpose application domains. Such a limitation occurs both due to their consolidated computation scheme (restricted to matrix multiplication) and due to their frequent adoption of low-precision/custom floating-point formats (unsuited for general application domains). In contrast, this paper proposes a new Reconfigurable Tensor Unit (RTU) which deploys an array of variable-precision Vector MultiplyAccumulate (VMA) units. Furthermore, each VMA unit leverages the new Posit floating-point format and supports the full range of standardized posit precisions in a single SIMD unit, with variable vector-element width. Moreover, the proposed RTU explores the Posit format features for fused operations, together with spatial and time-multiplexing reconfiguration mechanisms to fuse and combine multiple VMAs to map high-level and complex operations. The RTU is also supported by an automatic data streaming infrastructure and a pipelined data movement scheme, allowing it to accelerate the computation of most data-parallel patterns commonly present in vectorizable applications. The proposed RTU showed to outperform state-of-the-art tensor and SIMD units, present in off-the-shelf platforms, in turn resulting in significant energy-efficiency improvements. Nuno Neves 0002, Pedro Tomás, Nuno Roma |
ASAP | 1 |
| 2018 | Stream data prefetcher for the GPU memory interface
Nuno Neves 0002, Pedro Tomás, Nuno Roma |
J. Supercomput. | 1 |
| 2017 | Adaptive In-Cache Streaming for Efficient Data ManagementabstractThe design of adaptive architectures is frequently focused on the sole adaptation of the processing blocks, often neglecting the power/performance impact of data transfers and data indexing in the memory subsystem. In particular, conventional address-based models, supported on cache structures to mitigate the memory wall problem, often struggle when dealing with memory-bound applications or arbitrarily complex data patterns that can be hardly captured by prefetching mechanisms. Stream-based techniques have proven to efficiently tackle such limitations, although not well-suited to handle all types of applications. To mitigate the limitations of both communication paradigms, an efficient unification is herein proposed, by means of a novel in-cache stream paradigm, capable of seamlessly adapting the communication between the address-based and stream-based models. The proposed morphable infrastructure relies on a new dynamic descriptor graph specification, capable of handling regular arbitrarily complex data patterns, which is able to improve the main memory bandwidth utilization through data reutilization and reorganization techniques. When compared with state-of-the-art solutions, the proposed structure offers higher address generation efficiency and achievable memory throughputs, and a significant reduction of the amount of data transfers and main memory accesses, resulting on average in 13 times system performance speedup and in 245 times energy-delay product improvement, when compared with the previous implementations. Nuno Neves 0002, Pedro Tomás, Nuno Roma |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2016 | Multi-objective kernel mapping and scheduling for morphable many-core architectures
Nuno Neves 0002, Rui Ferreira Neves, Nuno Horta, Pedro Tomás, Nuno Roma |
Expert Syst. Appl. | 1 |
| 2015 | Efficient data-stream management for shared-memory many-core systemsabstractThe design of most high-performance and heterogeneous processing platforms is usually solely focused on the computational part, while neglecting the power/performance impact of the data-management infrastructures. Moreover, such systems often struggle to achieve their potential performance when applications require fetching data with complex memory access patterns. To overcome these issues, a energy-efficient stream-based data-management infrastructure is herein proposed, relying on a novel tree-based descriptor specification. Such descriptors are decoded by a Descriptor Tree Controller (DTC) architecture, which allows simple and efficient management of arbitrarily complex memory access patterns. Moreover, a Stream Management Engine (SME) ensures energy-efficient data-reutilization through the application of automatic stream rerouting, splitting and merging techniques. The obtained results show that the proposed DTC architecture is capable of a highly efficient complex data-pattern generation, while significantly reducing the size occupied by the pattern description, when compared with state-of-the-art approaches. By also enabling the deployment of data-reuse techniques, a reduction of up to 85× in the number of accesses to the main shared memory is achieved, resulting in a decrease as high as 475× in the observed energy consumption. Nuno Neves 0002, Pedro Tomás, Nuno Roma |
FPL | 1 |
| 2015 | Multicore SIMD ASIP for Next-Generation Sequencing and Alignment Biochip PlatformsabstractTargeting the development of new biochip platforms capable of autonomously sequencing and aligning biological sequences, a new multicore processing structure is proposed in this manuscript. This multicore structure makes use of a shared memory model and multiple instantiations of a novel application-specific instruction-set processor (ASIP) to simultaneously exploit both fine and coarse-grained parallelism and to achieve high performance levels at low-power consumption. The proposed ASIP is built by extending the instruction set architecture of a synthesizable processor, including both general and special-purpose single-instruction multiple-data instructions. This allows an efficient exploitation of fine-grained parallelism on the alignment of biological sequences, achieving over$30\times $speedup when compared with sequential algorithmic implementations. The complete system was prototyped on different field-programmable gate array platforms and synthesized with a 90-nm CMOS process technology. Experimental results demonstrate that the multicore structure scales almost linearly with the number of instantiated cores, achieving performances similar to a quad-core Intel Core i7 3820 processor, while using$25\times $less energy. Nuno Neves 0002, Nuno Sebastião, David Martins de Matos, Pedro Tomás, Paulo F. Flores, Nuno Roma |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2013 | BioBlaze: Multi-core SIMD ASIP for DNA sequence alignmentabstractA new Application-Specific Instruction-set Processor (ASIP) architecture for biological sequences alignment is proposed in this manuscript. This architecture achieves high processing throughputs by exploiting both fine and coarse-grained parallelism. The former is achieved by extending the Instruction Set Architecture (ISA) of a synthesizable processor to include multiple specialized SIMD instructions that implement vector-vector and vector-scalar arithmetic, logic, load/store and control operations. Coarse-grained parallelism is achieved by using multiple cores to cooperatively align multiple sequences in a shared memory architecture, comprising proper hardware-specific synchronization mechanisms. To ease the programming, a compilation framework based on an adaptation of the GCC back-end was also implemented. The proposed system was prototyped and evaluated on a Xilinx Virtex-7 FPGA, achieving a 200MHz working frequency. A sequential and a state-of-theart SIMD implementations of the Smith-Waterman algorithm were programmed in both the proposed ASIP and an Intel Core i7 processor. When comparing the achieved speedups, it was observed that the proposed ISA achieves a 40x speedup, which contrasts with the 11x speedup provided by SSE2 in the Intel Core i7 processor. The scalability of the multi-core system was also evaluated and proved to scale almost linearly with the number of cores. Nuno Neves 0002, Nuno Sebastião, Andre Patricio, David Martins de Matos, Pedro Tomás, Paulo F. Flores, Nuno Roma |
ASAP | 1 |