João M. P. Cardoso

dblp:61/4222 · also João Manuel Paiva Cardoso · DBLP profile ↗
← Back
82ranked-venue papers
17as first author
11since 2021 · last 2026
0000-0002-7353-1799ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 56 · 15 first-author · 7 since 2021Software engineering, systems software and programming languages · 15 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-authorArtificial intelligence and machine learning · 3 · 1 since 2021Computer networks · 1Security and privacy · 1
YearPublicationVenuePosition
2026 Runtime-Adaptive Energy-Aware Selection of Configurations for Resource-Constrained Devices
Paulo J. S. Ferreira, João Mendes-Moreira 0001, João M. P. Cardoso
COMPSAC3
2025 Ph.D. Project: Holistic Partitioning and Optimization of CPU-FPGA Applications Through Source-to-Source Compilation
abstract
Critical performance regions of software applications are often accelerated by offloading them onto an FPGA. An efficient end result requires the judicious application of two processes: hardware/software (hw/sw) partitioning, which identifies the regions for offloading, and the optimization of those regions for efficient High-level Synthesis (HLS). Both processes are commonly applied separately, not relying on any potential interplay between them, and not revealing how the decisions made in one process could positively influence the other. This paper describes our primary efforts and contributions made so far, and our work-in-progress, in an approach that combines both hw/sw partitioning and optimization into a unified, holistic process, automated using source-to-source compilation. By using an Extended Task Graph (ETG) representation of a C/C++ application, and expanding the synthesizable code regions, our approach aims at creating clusters of tasks for offloading by a) maximizing the potential optimizations applied to the cluster, b) minimizing the global communication cost, and c) grouping tasks that share data in the same cluster.
João Bispo, João M. P. Cardoso
FCCM3
2025 On Improving the HLS Compatibility of Large C/C++ Code Regions
abstract
Heterogeneous CPU-FPGA C/C++ applications may rely on High-level Synthesis (HLS) tools to generate hardware for critical code regions. As typical HLS tools have several restrictions in terms of supported language features, to increase the size and variety of offloaded regions, we propose several code transformations to improve synthesizability. Such code transformations include: struct and array flattening; moving dynamic memory allocations out of a region; transforming dynamic memory allocations into static; and asynchronously executing host functions, e.g., printf(). We evaluate the impact of these transformations on code region size using three real-world applications whose critical regions are limited by non-synthesizable C/C++ language features.
João Bispo, João M. P. Cardoso, James C. Hoe
FCCM3
2024 A Flexible-Granularity Task Graph Representation and Its Generation from C Applications (WIP)
abstract
Modern hardware accelerators, such as FPGAs, allow offloading large regions of C/C++ code in order to improve the execution time and/or the energy consumption of software applications. An outstanding challenge with this approach, however, is solving the Hardware/Software (Hw/Sw) partitioning problem. Given the increasing complexity of both the accelerators and the potential code regions, one needs to adopt a holistic approach when selecting an offloading region by exploring the interplay between communication costs, data usage patterns, and target-specific optimizations. To this end, we propose representing a C application as an extended task graph (ETG) with flexible granularity, which can be manipulated through the merging and splitting of tasks. This approach involves generating a task graph overlay on the program's Abstract Syntax Tree (AST) that maps tasks to functions and the flexible granularity operations onto inlining/outlining operations. This maintains the integrity and readability of the original source code, which is paramount for targeting different accelerators and enabling code optimizations, while allowing the offloading of code regions of arbitrary complexity based on the data patterns of their tasks. To evaluate the ETG representation and its compiler, we use the latter to generate ETGs for the programs in Rosetta and MachSuite benchmark suites, and extract several metrics regarding data communication, task-level parallelism, and dataflow patterns between pairs of tasks. These metrics provide important information that can be used by Hw/Sw partitioning methods.
João Bispo, João M. P. Cardoso
LCTES3
2023 A CPU-FPGA Holistic Source-To-Source Compilation Approach for Partitioning and Optimizing C/C++ Applications
abstract
A common approach for improving performance uses FPGAs to accelerate critical code regions, which often involves two processes: hardware/software partitioning, which identifies regions to offload to the FPGA; and optimizing those regions (e.g., through HLS directives). As both processes are separate and usually applied in sequence, the interplay between them is unnatural, and it is unclear how the choices made in one step can benefit the choices made in the other step. This paper presents our work-in-progress for combining partitioning and optimization into a single holistic process. First, our source-to-source compiler builds a task-based representation from the input application. Then, a greedy algorithm builds clusters of tasks and assigns each cluster to either hardware (FPGA) or software (CPU). The algorithm iteratively refines the clusters and offloading decisions by: a) minimizing the communication costs between clusters by assigning tasks that work with shared data to the same cluster; b) reducing the global execution time by applying code optimizations to the tasks in each cluster. We show the impact of our holistic approach to a motivating edge detection example and compare the results when applying partitioning and code optimizations as independent steps. The results show that a holistic partitioning can lead to a speedup of up to$28.7\times$when compared to a simple offloading of the application to an FPGA.
João Bispo, João M. P. Cardoso
PACT3
2023 Preface ASAP 2023
abstract
keynote speeches and provide information regarding the ASAP'2023 organizing, steering and program committees, subreviewers, and sponsors.
João M. P. Cardoso, Alexandra Jimborean, Nele Mentens, José Gabriel F. Coutinho
ASAP1
2022 Pegasus: Performance Engineering for Software Applications Targeting HPC Systems
abstract
Developing and optimizing software applications for high performance and energy efficiency is a very challenging task, even when considering a single target machine. For instance, optimizing for multicore-based computing systems requires in-depth knowledge about programming languages, application programming interfaces (APIs), compilers, performance tuning tools, and computer architecture and organization. Many of the tasks of performance engineering methodologies require manual efforts and the use of different tools not always part of an integrated toolchain. This paper presents Pegasus, a performance engineering approach supported by a framework that consists of a source-to-source compiler, controlled and guided by strategies programmed in a Domain-Specific Language, and an autotuner. Pegasus is a holistic and versatile approach spanning various decision layers composing the software stack, and exploiting the system capabilities and workloads effectively through the use of runtime autotuning. The Pegasus approach helps developers by automating tasks regarding the efficient implementation of software applications in multicore computing systems. These tasks focus on application analysis, profiling, code transformations, and the integration of runtime autotuning. Pegasus allows developers to program their strategies or to automatically apply existing strategies to software applications in order to ensure the compliance of non-functional requirements, such as performance and energy efficiency. We show how to apply Pegasus and demonstrate its applicability and effectiveness in a complex case study, which includes tasks from a smart navigation system.
Pedro Pinto 0002, João Bispo, João M. P. Cardoso, Jorge G. Barbosa, Davide Gadioli, Gianluca Palermo, Jan Martinovic, Martin Golasowski, Katerina Slaninová, Radim Cmar, Cristina Silvano
IEEE Trans. Software Eng.3
2021 A methodology and framework for software memoization of functions
abstract
Enhancing performance is crucial when developing applications for high-performance and embedded computing. It requires sophisticated techniques and in-depth knowledge of the application domain and target architecture. Typically, developers prioritize the application's functional requirements over extra-functional requirements. Thus, a large part of the optimization effort is shifted to performance engineers, who rely on manual effort, alongside many analysis and optimization tools that need integration. This paper focuses on memoization, which caches results of pure computations and retrieves them if a function is called with repeating arguments. We propose a methodology for allowing developers and performance engineers to apply memoization straightforwardly by automating code analysis, code transformations, and memoization-specific profiling. It helps developers with no optimization expertise to quickly set up memoization and, simultaneously, it provides performance engineers with highly customizable analysis and memoization. We provide a concrete implementation supported by a DSL, a source-to-source compiler, and a memoization framework. We evaluate the methodology and framework with publicly available benchmarks. We show how one can analyze applications to select functions with performance improvement potential, which the experiments reveal might be challenging to find, and improve some applications with minimal effort.
Pedro Pinto 0002, João M. P. Cardoso
CF2
2021 On the Performance Effect of Loop Trace Window Size on Scheduling for Configurable Coarse Grain Loop Accelerators
abstract
By using Dynamic Binary Translation, instruction traces from pre-compiled applications can be offloaded, at runtime, to FPGA-based accelerators, such as Coarse-Grained Loop Accelerators, in a transparent way. However, scheduling onto coarse-grain accelerators is challenging, with two of current known issues being the density of computations that can be mapped, and the effects of memory accesses on performance. Using an in-house framework for analysis of instruction traces, we explore the effect of different window sizes when applying list scheduling, to map the window operations to a coarse-grain loop accelerator model that has been previously experimentally validated. For all window sizes, we vary the number of ALUs and memory ports available in the model, and comment how these parameters affect the resulting latency. For a set of benchmarks taken from the PolyBench suite, compiled for the 32-bit MicroBlaze softcore, we have achieved an average iteration speedup of 5.10x for a basic block repeated 5 times and scheduled with 8 ALUs and memory ports, and an average speedup of 5.46x when not considering resource constraints. We also identify which benchmarks contribute to the difference between these two speedups, and breakdown their limiting factors. Finally, we reflect on the impact memory dependencies have on scheduling.
Nuno Miguel Cardanha Paulino, João Bispo, João M. P. Cardoso, João Canas Ferreira
FPT4
2021 An ensemble of autonomous auto-encoders for human activity recognition
abstract
Human Activity Recognition is focused on the use of sensing technology to classify human activities and to infer human behavior. While traditional machine learning approaches use hand-crafted features to train their models, recent advancements in neural networks allow for automatic feature extraction. Auto-encoders are a type of neural network that can learn complex representations of the data and are commonly used for anomaly detection. In this work we propose a novel multi-class algorithm which consists of an ensemble of auto-encoders where each auto-encoder is associated with a unique class. We compared the proposed approach with other state-of-the-art approaches in the context of human activity recognition. Experimental results show that ensembles of auto-encoders can be efficient, robust and competitive. Moreover, this modular classifier structure allows for more flexible models. For example, the extension of the number of classes, by the inclusion of new auto-encoders, without the necessity to retrain the whole model.
Kemilly Dearo Garcia, Cláudio Rebelo de Sá, Mannes Poel, Tiago Carvalho 0001, João Mendes-Moreira 0001, João M. P. Cardoso, André C. P. L. F. de Carvalho, Joost N. Kok
Neurocomputing6
2021 Guest Editorial: IEEE TC Special Section on Compiler Optimizations for FPGA-Based Systems
abstract
The papers in this special section focus on compiler optimization for FPGA-based systems. Reconfigurable computing (RC) is growing in importance in many computing domains and systems, from embedded, mobile to cloud, and high-performance computing. We have witnessed important advancements regarding the programming of RC-based systems, but further improvements are needed, especially regarding efficient techniques for automatic mapping of computations described in high-level languages to the RC resources. The resources of high-end FPGAs allow these devices to implement complex Systemson- a-Chip (SoCs) and substantial computational components of software applications, e.g., when used as hardware accelerators and/or as more energy-efficient computing platforms. This, however, increases the continuous need for efficient compilers targeting FPGAs, and other RC platforms, from high-level programming languages.
João M. P. Cardoso, André DeHon, Laura Pozzi 0001
IEEE Trans. Computers1
2020 Executing ARMv8 Loop Traces on Reconfigurable Accelerator via Binary Translation Framework
abstract
Performance and power efficiency in edge and embedded systems can benefit from specialized hardware. To avoid the effort of manual hardware design, we explore the generation of accelerator circuits from binary instruction traces for several Instruction Set Architectures.
Nuno Miguel Cardanha Paulino, João Canas Ferreira, João Bispo, João M. P. Cardoso
FPL4
2020 Compilation of MATLAB computations to CPU/GPU via C/OpenCL generation
abstract
Summary In order to take advantage of the processing power of current computing platforms, programmers typically need to develop software versions for different target devices. This task is time‐consuming and requires significant programming and computer architecture expertise. A possible and more convenient alternative is to start with a single high‐level description of a program with minimum implementation details, and generate custom implementations according to the target platform. In this paper, we use MATLAB as a high‐level programming language and propose a compiler that targets CPU/GPU computing platforms by generating customized implementations in C and OpenCL. We propose a number of compiler techniques to automatically generate efficient C and OpenCL code from MATLAB programs. One of such compiler techniques relies on heuristics to decide when and how to use Shared Virtual Memory (SVM). The experimental results show that our approach is able to generate code that provides significant speedups (eg, geometric mean speedup of 11× for a set of simple benchmarks) using a discrete GPU over equivalent sequential C code executing on a CPU. With more complex benchmarks, for which only some code regions can be parallelized, and are thus offloaded, the generated code achieved speedups of up to 2.2×. We also show the impact of using SVM, specifically fine‐grained buffers, and the results show that the compiler is able to achieve significant speedups, both over the versions without SVM and with naïve aggressive SVM use, across three CPU/GPU platforms.
Luís Reis 0001, João Bispo, João M. P. Cardoso
Concurr. Comput. Pract. Exp.3
2020 Source-to-source compilation targeting OpenMP-based automatic parallelization of C applications
Hamid Arabnejad, João Bispo, João M. P. Cardoso, Jorge G. Barbosa
J. Supercomput.3
2019 An Efficient Scheme for Prototyping kNN in the Context of Real-Time Human Activity Recognition
Paulo J. S. Ferreira, Ricardo M. C. Magalhães, Kemilly Dearo Garcia, João M. P. Cardoso, João Mendes-Moreira 0001
IDEAL (1)4
2019 Supporting the Scale-Up of High Performance Application to Pre-Exascale Systems: The ANTAREX Approach
abstract
The ANTAREX project developed an approach to the performance tuning of High Performance applications based on an Aspect-oriented Domain Specific Language (DSL), with the goal to simplify the enforcement of extra-functional properties in large scale applications. The project aims at demonstrating its tools and techniques on two relevant use cases, one in the domain of computational drug discovery, the other in the domain of online vehicle navigation. In this paper, we present an overview of the project and of its main achievements, as well as of the large scale experiments that have been planned to validate the approach.
Cristina Silvano, Giovanni Agosta, Andrea Bartolini, Andrea Beccari, Luca Benini, Loïc Besnard, João Bispo, Radim Cmar, João M. P. Cardoso, Carlo Cavazzoni, Daniele Cesarini, Stefano Cherubin, Federico Ficarelli, Davide Gadioli, Martin Golasowski, Imane Lasri, Antonio Libri, Candida Manelfi, Jan Martinovic, Gianluca Palermo, Pedro Pinto 0002, Erven Rohou, Nico Sanna, Katerina Slaninová, Emanuele Vitali
PDP9
2019 Dynamic Partial Reconfiguration of Customized Single-Row Accelerators
abstract
The use of specialized accelerator circuits is a feasible solution to address performance and energy issues in embedded systems. This paper extends a previous field-programmable gate array-based approach that automatically generates pipelined customized loop accelerators (CLAs) from runtime instruction traces. Despite efficient acceleration, the approach suffered from high area and resource requirements when offloading a large number of kernels from the target application. This paper addresses this by enhancing the CLA with dynamic partial reconfiguration (DPR) support. Each kernel to accelerate is implemented as a variant of a reconfigurable area of the CLA which hosts all functional units and configuration memory. Evaluation of the proposed system is performed on a Virtex-7 device. We show, for a set of 21 kernels, that when comparing two CLAs capable of accelerating the same subset of kernels, the one which benefits from DPR can be up to$4.3\times $smaller. Resorting to DPR allows for the implementation of CLAs which support numerous kernels without a significant decrease in operating frequency and does not affect the initiation intervals at which kernels are scheduled. Finally, the area required by a CLA instance can be further reduced by increasing the IIs of the scheduled kernels.
Nuno Miguel Cardanha Paulino, João Canas Ferreira, João M. P. Cardoso
IEEE Trans. Very Large Scale Integr. Syst.3
2018 Autotuning and adaptivity in energy efficient HPC systems: the ANTAREX toolbox
abstract
Designing and optimizing applications for energy-efficient High Performance Computing systems up to the Exascale era is an extremely challenging problem. This paper presents the toolbox developed in the ANTAREX European project for autotuning and adaptivity in energy efficient HPC systems. In particular, the modules of the ANTAREX toolbox are described as well as some preliminary results of the application to two target use cases. 1
Cristina Silvano, Gianluca Palermo, Giovanni Agosta, Amir H. Ashouri, Davide Gadioli, Stefano Cherubin, Emanuele Vitali, Luca Benini, Andrea Bartolini, Daniele Cesarini, João M. P. Cardoso, João Bispo, Pedro Pinto 0002, Ricardo Nobre, Erven Rohou, Loïc Besnard, Imane Lasri, Nico Sanna, Carlo Cavazzoni, Radim Cmar, Jan Martinovic, Katerina Slaninová, Martin Golasowski, Andrea Beccari, Candida Manelfi
CF11
2018 SOCRATES - A seamless online compiler and system runtime autotuning framework for energy-aware applications
abstract
Configuring program parallelism and selecting optimal compiler options according to the underlying platform architecture is a difficult task. Tipically, this task is either assigned to the programmer or done by a standard one-fits-all policy generated by the compiler or runtime system. A runtime selection of the best configuration requires the insertion of a lot of glue code for profiling and runtime selection. This represents a programming wall for application developers. This paper presents a structured approach, called SOCRATES, based on an aspect-oriented language (LARA) and a runtime autotuner (mARGOt) to mitigate this problem. LARA has been used to hide the glue code insertion, thus separating the pure functional application description from extra-functional requirements. mARGOT has been used for the automatic selection of the best configuration according to the runtime evolution of the application.1
Davide Gadioli, Ricardo Nobre, Pedro Pinto 0002, Emanuele Vitali, Amir H. Ashouri, Gianluca Palermo, João M. P. Cardoso, Cristina Silvano
DATE7
2018 ANTAREX: A DSL-Based Approach to Adaptively Optimizing and Enforcing Extra-Functional Properties in High Performance Computing
abstract
The ANTAREX project relies on a Domain Specific Language (DSL) based on Aspect Oriented Programming (AOP) concepts to allow applications to enforce extra functional properties such as energy-efficiency and performance and to optimize Quality of Service (QoS) in an adaptive way. The DSL approach allows the definition of energy-efficiency, performance, and adaptivity strategies as well as their enforcement at runtime through application autotuning and resource and power management. In this paper, we present an overview of the ANTAREX DSL and some of its capabilities through a number of examples, including how the DSL is applied in the context of one of the project use cases.
Cristina Silvano, Giovanni Agosta, Andrea Bartolini, Andrea Beccari, Luca Benini, Loïc Besnard, João Bispo, Radim Cmar, João M. P. Cardoso, Carlo Cavazzoni, Stefano Cherubin, Davide Gadioli, Martin Golasowski, Imane Lasri, Jan Martinovic, Gianluca Palermo, Pedro Pinto 0002, Erven Rohou, Nico Sanna, Katerina Slaninová, Emanuele Vitali
DSD9
2018 Aspect composition for multiple target languages using LARA
Pedro Pinto 0002, Tiago Carvalho 0001, João Bispo, Miguel António Ramalho, João M. P. Cardoso
Comput. Lang. Syst. Struct.5
2017 Foreword to the special issue of the 18th IEEE international conference on computational science and engineering (CSE2015)
abstract
The Computational Science and Engineering (CSE) area has earned prominence through advances in electronic and integrated technologies. Advanced computing systems permeate our daily life and have an increasingly importance in many aspects and domains. CSE is shaping future research and development activities in academia and industry, ranging from engineering, science, finance, economics, healthcare, arts, and humanitarian fields. The IEEE International Conference on CSE has been providing a series of highly successful International Conferences on CSE. The 2015 edition of CSE, CSE2015 (http://www.fe.up.pt/cse2015), was held in Porto, Portugal on October 21–23, 2015. It brought together computer scientists, industrial engineers, and researchers to discuss and exchange experimental and theoretical results, work-in-progress, experiences, case studies, and trend-setting ideas, in the areas of advanced computing for solving problems in science and engineering applications. The six extended papers included have been selected from a preliminary set of 12 papers submitted to this special issue and are briefly described as follows. The article ‘Robust resource allocations through performance modeling with stochastic process algebra’ 1 presents a new resource allocation scheme that uses a stochastic process algebra for obtaining resource allocations. These allocations are robust with respect to unpredictable perturbations of the application or system characteristics during runtime. The key idea is to translate performance models into mathematical Markov chain descriptions that can be numerically evaluated without requiring the time consuming simulation process used in competing approaches. A comparison with previous studies shows that the proposed approach achieves competitive results. In addition, process algebras are easier to reproduce because they require neither the effort for learning a simulation framework nor setup or installation cost. Beyond that, the computational efficiency of the approach also allows for embedding the proposed process algebra model into the runtime systems of a model-based framework that re-evaluates the resource assignment whenever a system or application parameter changes at runtime. The article ‘Heterogeneous CPU + GPU Approaches for Mesh Refinement over Lattice-Boltzmann Simulations’ 2 investigates strategies for mapping Lattice-Boltzmann method (LBM) simulations to compute nodes with CPUs and GPUs. The particular challenge addressed is finding an efficient way so that adaptive mesh refinement strategies use both computing resources effectively. While parallelism is abundant in LBM simulations, the challenge is to structure the workload distribution and data access to perform well on both CPU and GPU, which inherently favor different granularities of parallelism. The authors propose two approaches, a multi-domain approach that uses a finer grid in domains where a higher resolution is required; and an irregular grid approach that uses a single Cartesian with non-uniform spacing. Both approaches (multi-domain and irregular grid) are implemented and evaluated for a system comprising a Xeon E5 CPU and an NVidia K20c GPU using either only the GPU or CPU + GPU for the LBM computation. The evaluation shows that the multi-domain approach allows for executing bigger simulations because it requires fewer lattice nodes. The irregular grid approach is easier to implement and delivers a higher throughput (million fluid lattice cell updates per second). For both methods, the CPU + GPU implementation outperforms the homogeneous GPU implementation by 10–30%. The article ‘Methods to Model and Simulate Super Carbon Nanotubes of Higher Order’ 3 presents a new, efficient approach based on graph algebra to simulate the mechanical behavior of super carbon nanotubes (SCNTs). Representing the SCNTs as directed graphs, the authors propose a new data structure that exploits the hierarchy of SCNTs for fast queries. In addition, they propose a novel, iterative solver using the conjugate gradient method. Exploiting the symmetry of SCNTs of order 0, the solver is able to drastically reduce the amount of required calculations and memory for small deformations. Further exploiting structural symmetry and adopting an improved proximity-aware Matrix–vector-Multiplication routine, the performance for SCNTs level 0 can be improved by an additional factor of 2. Up to 4.4 times speedup is achieved when running in parallel on a 16 core SMP system. The authors also explore optimizations for symmetry in SCNTs of order 1. Experimental results show that the new approach outperforms a compressed-row-storage-based reference solver, for SCNTs of order 0 and 1, regardless of deformation, and with much less memory consumption. Because in practice memory consumption is oftentimes the limiting factor for scaling, this approach can significantly expand the realm of feasible simulations for SCNTs. In the article, the readers can also find introductions to the basic mathematical formulations and algorithms for simulating SCNTs as well as a brief summary of previous results. The article ‘Combinatorial Optimization of DNA Sequence Analysis on Heterogeneous Systems’ 4 presents an experimental study of counting the occurrences of query patterns in DNA sequences in parallel. DNA sequence analysis has many important practical applications, and is both data and computation intensive. The authors parallelize the Aho Corasick pattern matching algorithm, and explore its execution on heterogeneous systems, such as Intel Xeon E5 with Xeon Phi as co-processor, for acceleration. To achieve maximal performance on heterogeneous systems, the authors employ simulated annealing to determine the number of threads, thread affinities, and data placement on the host and the accelerator as such configuration is critical to performance and system utilization. Using real-world DNA sequences, the authors evaluate the efficiency of their approach. They show that the average speedup achieved is 1.6 times compared against the host-only parallelization and 2 times against device-only parallelization. The article ‘Using Adaptive Runtime Filtering to Support an Event-based Performance Analysis’ 5 presents an approach to filter tracing data in the context of event-based monitoring for performance improvements. The approach is based on self-guided filters that automatically adapt to an application's runtime behavior and are able to reduce performance data to manageable sizes for large-scale parallel applications and long execution programs. The article presents four runtime filters, each one targeting a specific type of data redundancy. They evaluate their approach with five real-world applications from different scientific domains. Compared to the default settings, their filters achieve a data size reduction of two orders of magnitude while increasing execution time regarding the overhead of tracing by less than one percent on average. The article also examines the influence of filtering on the performance analysis and identifies its limitations and presents three schemes to help performance analysts to correctly interpret filtered traces or even reconstruct parts of a filtered trace. The article ‘Automatic source-to-source error compensation of floating-point programs: code synthesis to optimize accuracy and time’ 6 presents a source-to-source C compiler approach for automatically improving the numerical accuracy of floating-point programs without significantly increasing execution time. The approach is based on the automatic compensation of floating-point operations by applying error-free transformations, and on the synthesis of code for both accuracy and execution time criteria. The use of partial compensation is proposed in order to trade-off performance and accuracy. The article also presents a number of code transformations to increase accuracy and to tune the impact on execution time. In addition, the authors present a method to find the best transformation satisfying execution time or accuracy constraints. The approach presented is evaluated with a number of case studies using two target computing environments and is able to produce some compensated algorithms as accurate and efficient as the ones derived by hand. We would like to acknowledge the authors of the articles included in this special issue for the hard work on preparing high-quality papers, the work of the anonymous reviewers on providing very important insights and suggestions that undoubtedly helped authors to improve their papers, and the support of the CCPE editors, Geoffrey C. Fox and David W. Walker.
Christian Plessl, Guojing Cong, João M. P. Cardoso
Concurr. Comput. Pract. Exp.3
2017 Special issue on design of algorithms and architectures for signal and image processing
Marek Gorgon, João M. P. Cardoso, Diana Göhringer, Leandro Soares Indrusiak
J. Syst. Archit.2
2017 Introduction to the special issue on architecture of computing systems
Frank Hannig, João M. P. Cardoso, Dietmar Fey
J. Syst. Archit.2
2017 A MATLAB subset to C compiler targeting embedded systems
abstract
This paper describes MATISSE, a compiler able to translate a MATLAB subset to C targeting embedded systems. MATISSE uses LARA, an aspect-oriented programming language, to specify additional information and transformations to the input MATLAB code, for example, insertion of code for initialization of variables, and specification of types and shapes of variables. The compiler is being developed bearing in mind flexibility, multitarget and multitoolchain support, allowing for the generation of several implementations in C from the same reference code in MATLAB. In this paper, we also present a number of techniques being employed in MATLAB to C compilation, such as element-wise mapping operations, matrix views, weak types, and intrinsics. We validate these techniques using MATISSE and a set of representative benchmarks. More specifically, we evaluate the compiler with a set of 31 benchmarks using an embedded system board and a desktop computer. The results show speedups up to 1.8× by employing information provided by LARA aspects, when compared with C code generated without additional user information. When compared with the execution time of the original code running on MATLAB, the execution time of the generated C code achieved a geometric mean speedup of 13×. Copyright © 2016 John Wiley & Sons, Ltd.
João Bispo, João M. P. Cardoso
Softw. Pract. Exp.2
2017 Introduction to the Special Section on FPL 2015
abstract
No abstract available.
João M. P. Cardoso, Cristina Silvano
ACM Trans. Reconfigurable Technol. Syst.1
2017 The First 25 Years of the FPL Conference: Significant Papers
abstract
A summary of contributions made by significant papers from the first 25 years of the Field-Programmable Logic and Applications conference (FPL) is presented. The 27 papers chosen represent those which have most strongly influenced theory and practice in the field.
Philip H. W. Leong, Hideharu Amano, Jason Helge Anderson, Koen Bertels, João M. P. Cardoso, Oliver Diessel, Guy Gogniat, Mike Hutton, Wayne Luk, Patrick Lysaght, Marco Platzner, Viktor Prasanna 0001, Tero Rissa, Cristina Silvano, Hayden Kwok-Hay So, Yu Wang 0002
ACM Trans. Reconfigurable Technol. Syst.5
2017 Generation of Customized Accelerators for Loop Pipelining of Binary Instruction Traces
abstract
Many embedded applications process large amounts of data using regular computational kernels, amenable to acceleration by specialized hardware coprocessors. To reduce the significant design effort, the dedicated hardware may be automatically generated, usually starting from the application’s source or binary code. This paper presents a moduloscheduled loop accelerator capable of executing multiple loops and a supporting toolchain. A generation/scheduling procedure, which fully relies on MicroBlaze instruction traces, produces accelerator instances, customized in terms of functional units and interconnections. The accelerators support integer and single-precision floating-point arithmetic, and exploit instruction-level parallelism, loop pipelining, and memory access parallelism via two read/write ports. A complete implementation of the proposed architecture is evaluated in a Virtex-7 device. Augmenting a MicroBlaze processor with a tailored accelerator achieves a geometric mean speedup, over software-only execution, of 6.61$\times $for 13 floating-point kernels from the Livermore Loops set, and of 4.08$\times $for 11 integer kernels from Texas Instruments’ IMGLIB. The proposed customized accelerators are compared with ALU-based ones. The average specialized accelerator requires only 0.47$\times $the number of field-programmable gate array slices of an accelerator with four ALUs. A geometric mean speedup of 1.78$\times $over a four-issue very long instruction word (without floating-point support) was obtained for the integer kernels.
Nuno Miguel Cardanha Paulino, João Canas Ferreira, João M. P. Cardoso
IEEE Trans. Very Large Scale Integr. Syst.3
2016 Autotuning and adaptivity approach for energy efficient Exascale HPC systems: The ANTAREX approach
Cristina Silvano, Giovanni Agosta, Andrea Bartolini, Andrea Beccari, Luca Benini, João Bispo, Radim Cmar, João M. P. Cardoso, Carlo Cavazzoni, Jan Martinovic, Gianluca Palermo, Martin Palkovic, Pedro Pinto 0002, Erven Rohou, Nico Sanna, Katerina Slaninová
DATE8
2016 Towards a multi-softcore FPGA approach for the HOG algorithm
abstract
Object detection in images is a computing demanding task which usually needs to deal with the detection of different classes of objects, and thus requiring variations and adaptations easily provided by software solutions. Object detection algorithms are being part of real-time smarter embedded systems, such as automotive, medical, robotics and security systems. In most embedded systems, efficient implementations of object oriented algorithms need to provide high performance, low power consumption, and programmability to allow greater development flexibility. The Histogram of Oriented Gradients (HOG) is one of the most widely used algorithms for object detection in images. In this paper, we show our work towards mapping the HOG algorithm to an FPGA-based system consisting of multiple Nios II softcore processors and bearing in mind high-performance and programmability issues. We show how to reduce 19x the algorithms execution time by source to source transformations and specially avoiding redundant processing. Furthermore, we show how the use of pipelining processing using three Nios II processors allows a speedup of 49x compared to the embedded baseline application.
José A. M. de Holanda, João M. P. Cardoso, Eduardo Marques
INDIN2
2016 A graph-based iterative compiler pass selection and phase ordering approach
abstract
Nowadays compilers include tens or hundreds of optimization passes, which makes it difficult to find sequences of optimizations that achieve compiled code more optimized than the one obtained using typical compiler options such as -O2 and -O3. The problem involves both the selection of the compiler passes to use and their ordering in the compilation pipeline. The improvement achieved by the use of custom phase orders for each function can be significant, and thus important to satisfy strict requirements such as the ones present in high-performance embedded computing systems. In this paper we present a new and fast iterative approach to the phase selection and ordering challenges resulting in compiled code with higher performance than the one achieved with the standard optimization levels of the LLVM compiler. The obtained performance improvements are comparable with the ones achieved by other iterative approaches while requiring considerably less time and resources. Our approach is based on sampling over a graph representing transitions between compiler passes. We performed a number of experiments targeting the LEON3 microarchitecture using the Clang/LLVM 3.7 compiler, considering 140 LLVM passes and a set of 42 representative signal and image processing C functions. An exhaustive cross-validation shows our new exploration method is able to achieve a geometric mean performance speedup of 1.28x over the best individually selected -OX flag when considering 100,000 iterations; versus geometric mean speedups from 1.16x to 1.25x obtained with state-of-the-art iterative methods not using the graph. From the set of exploration methods tested, our new method is the only one consistently finding compiler sequences that result in performance improvements when considering 100 or less exploration iterations. Specifically, it achieved geometric mean speedups of 1.08x and 1.16x for 10 and 100 iterations, respectively.
Ricardo Nobre, Luiz G. A. Martins, João M. P. Cardoso
LCTES3
2016 Performance-driven instrumentation and mapping strategies using the LARA aspect-oriented programming approach
abstract
Summary The development of applications for high‐performance embedded systems is a long and error‐prone process because in addition to the required functionality, developers must consider various and often conflicting nonfunctional requirements such as performance and/or energy efficiency. The complexity of this process is further exacerbated by the multitude of target architectures and mapping tools. This article describes LARA, an aspect‐oriented programming language that allows programmers to convey domain‐specific knowledge and nonfunctional requirements to a toolchain composed of source‐to‐source transformers, compiler optimizers, and mapping/synthesis tools. LARA is sufficiently flexible to target different tools and host languages while also allowing the specification of compilation strategies to enable efficient generation of software code and hardware cores (using hardware description languages) for hybrid target architectures – a unique feature to the best of our knowledge not found in any other aspect‐oriented programming language. A key feature of LARA is its ability to deal with different models of join points, actions, and attributes. In this article, we describe the LARA approach and evaluate its impact on code instrumentation and analysis and on selecting critical code sections to be migrated to hardware accelerators for two embedded applications from industry. Copyright © 2014 John Wiley & Sons, Ltd.
João M. P. Cardoso, José Gabriel F. Coutinho, Tiago Carvalho 0001, Pedro C. Diniz, Zlatko Petrov, Wayne Luk, Fernando M. Gonçalves
Softw. Pract. Exp.1
2016 Clustering-Based Selection for the Exploration of Compiler Optimization Sequences
abstract
A large number of compiler optimizations are nowadays available to users. These optimizations interact with each other and with the input code in several and complex ways. The sequence of application of optimization passes can have a significant impact on the performance achieved. The effect of the optimizations is both platform and application dependent. The exhaustive exploration of all viable sequences of compiler optimizations for a given code fragment is not feasible. As this exploration is a complex and time-consuming task, several researchers have focused on Design Space Exploration (DSE) strategies both to select optimization sequences to improve the performance of each function of the application and to reduce the exploration time. In this article, we present a DSE scheme based on a clustering approach for grouping functions with similarities and exploration of a reduced search space resulting from the combination of optimizations previously suggested for the functions in each group. The identification of similarities between functions uses a data mining method that is applied to a symbolic code representation. The data mining process combines three algorithms to generate clusters: the Normalized Compression Distance, the Neighbor Joining, and a new ambiguity-based clustering algorithm. Our experiments for evaluating the effectiveness of the proposed approach address the exploration of optimization sequences in the context of the ReflectC compiler, considering 49 compilation passes while targeting a Xilinx MicroBlaze processor, and aiming at performance improvements for 51 functions and four applications. Experimental results reveal that the use of our clustering-based DSE approach achieves a significant reduction in the total exploration time of the search space (20× over a Genetic Algorithm approach) at the same time that considerable performance speedups (41% over the baseline) were obtained using the optimized codes. Additional experiments were performed considering the LLVM compiler, considering 124 compilation passes, and targeting a LEON3 processor. The results show that our approach achieved geometric mean speedups of 1.49 × , 1.32 × , and 1.24 × for the best 10, 20, and 30 functions, respectively, and a global improvement of 7% over the performance obtained when compiling with -O2.
Luiz G. A. Martins, Ricardo Nobre, João M. P. Cardoso, Alexandre C. B. Delbem, Eduardo Marques
ACM Trans. Archit. Code Optim.3
2015 Transparent acceleration of program execution using reconfigurable hardware
Nuno Miguel Cardanha Paulino, João Canas Ferreira, João Bispo, João M. P. Cardoso
DATE4
2015 A special-purpose language for implementing pipelined FPGA-based accelerators
abstract
A common use for Field-Programmable Gate Arrays (FPGAs) is the implementation of hardware accelerators. A way of doing so is to specify the internal logic of such accelerators by using Hardware Description Languages (HDLs). However, HDLs rely on the expertise of developers and their knowledge about hardware development with FPGAs. Regarding this, efforts have been focused on developing High-level Synthesis (HLS) tools in an attempt to increase the overall abstraction level required for using FPGAs. However, the solutions presented by such tools are commonly considered inefficient in comparison to the ones achieved by a specialized hardware designer. An alternative solution to program FPGAs is the use of Domain- Specific Languages (DSLs), as they can provide higher abstraction levels than HDLs still allowing the developers to deal with specific issues leading to more efficient designs and not always covered by HLS tools. In this paper we present our recent work on a DSL named LALP (Language for Aggressive Loop Pipelining), which has been developed focusing on the development of FPGAbased, aggressively pipelined, hardware accelerators. We present the recent LALP extensions and the challenges we are facing regarding to the compilation of LALP to FPGAs.
Cristiano Bacelar de Oliveira, Ricardo Menotti, João M. P. Cardoso, Eduardo Marques
FDL3
2015 Significant papers from the first 25 years of the FPL conference
abstract
The list of significant papers from the first 25 years of the Field-Programmable Logic and Applications conference (FPL) is presented in this paper. These 27 papers represent those which have most strongly influenced theory and practice in the field.
Philip H. W. Leong, Hideharu Amano, Jason Helge Anderson, Koen Bertels, João M. P. Cardoso, Oliver Diessel, Guy Gogniat, Mike Hutton, Wayne Luk, Patrick Lysaght, Marco Platzner, Viktor Prasanna 0001, Tero Rissa, Cristina Silvano, Hayden Kwok-Hay So, Yu Wang 0002
FPL5
2015 Reducing misses to external memory accesses in task-level pipelining
abstract
Recently, researchers have shown an increased interest in using task-level pipelining to accelerate the overall execution of applications mainly consisting of producer-consumer tasks. This paper proposes optimization techniques for enhancing our approach to pipeline the execution of producer-consumer tasks in FPGA-based multicore architectures with reductions in the number of accesses to external memory. Our approach is able to speedup the overall execution of successive, data-dependent tasks, by using multiple cores and specific customization features provided by FPGAs. We evaluate the impact in the performance of task-level pipelining when using different hash functions and optimization schemes in the inter stage buffer (ISB). The optimizations proposed in this paper were evaluated with FPGA implementations. The experimental results show the efficiency of a simple scheme to reduce external memory accesses and the suitability of the hash function being used. Furthermore, the results reveal noticeable performance improvements for the set of benchmarks being used.
Ali Azarian, João M. P. Cardoso
ISCAS2
2015 Programming Strategies for Contextual Runtime Specialization
abstract
Runtime adaptability is expected to adjust the application and the mapping of computations according to usage contexts, operating environments, resources availability, etc. However, extending applications with adaptive features can be a complex task, especially due to the current lack of programming models and compiler support. One of the run-time adaptability possibilities is the use of specialized code according to data workloads and environments. Traditional approaches use multiple code versions generated offline and, during runtime, a strategy is responsible to select a code version. Moving code generation to runtime can achieve important improvements but may impose unacceptable overhead. This paper presents an aspect-oriented programming approach for runtime adaptability. We focus on a separation of concerns (strategies vs. application) promoted by a domain-specific language for programming runtime strategies. Our strategies allow runtime specialization based on contextual information. We use a template-based runtime code generation approach to achieve program specialization. We demonstrate our approach with examples from image processing, which depict the benefits of runtime specialization and illustrate how several factors need to be considered to efficiently adapt the application.
Tiago Carvalho 0001, Pedro Pinto 0002, João M. P. Cardoso
SCOPES3
2015 Use of Previously Acquired Positioning of Optimizations for Phase Ordering Exploration
abstract
This paper presents a new approach to efficiently search for suitable compiler pass sequences, a challenge known as phase ordering. Our approach relies on information about the relative positions of compiler passes in compiler pass sequences previously generated for a set of functions when compiling for a specific processor. We enhanced two iterative compiler pass exploration schemes, one relying on simple sequential compiler pass insertion and other implementing an auto-tuned simulated annealing process, with a data structure that holds information about the relative positions of compiler sequences; in order to reduce the set of compiler passes considered for insertion in a given position of a given candidate compiler pass sequence to include only the passes that have a higher probability of performing well on that relative position in the compiler sequence, speeding up the exploration time as a result. We tested our approach with two different compilers and two different targets; the ReflectC and the LLVM compilers, targeting a MicroBlaze processor and a LEON3 processor, respectively. The experimental results show that we can considerably reduce the number of algorithm iterations by a factor of up to more than an order of magnitude when targeting the MicroBlaze or the LEON3, while finding compiler sequences that result in binaries that when executed on the target processor/simulator are able to outperform (i.e. use less CPU cycles) all the standard optimization levels (i.e., we compare against the most performing optimization level flag on each kernel, e.g. -O1, -O2 or -O3 in the case of LLVM) by a geometric mean performance improvement of 1.23x and 1.20x when targeting the MicroBlaze processor, and 1.94x and 2.65x when targetting the LEON3 processor; for each of the two exploration algorithms and two kernel sets considered.
Ricardo Nobre, Luiz G. A. Martins, João M. P. Cardoso
SCOPES3
2015 Guest Editorial FPL 2013
abstract
No abstract available.
João M. P. Cardoso, Pedro C. Diniz, Katherine Morrow
ACM Trans. Reconfigurable Technol. Syst.1
2015 Guest Editorial ARC 2014
abstract
No abstract available.
Diana Göhringer, Marco D. Santambrogio, João M. P. Cardoso, Koen Bertels
ACM Trans. Reconfigurable Technol. Syst.3
2015 A Reconfigurable Architecture for Binary Acceleration of Loops with Memory Accesses
abstract
This article presents a reconfigurable hardware/software architecture for binary acceleration of embedded applications. A Reconfigurable Processing Unit (RPU) is used as a coprocessor of the General Purpose Processor (GPP) to accelerate the execution of repetitive instruction sequences called Megablocks . A toolchain detects Megablocks from instruction traces and generates customized RPU implementations. The implementation of Megablocks with memory accesses uses a memory-sharing mechanism to support concurrent accesses to the entire address space of the GPP’s data memory. The scheduling of load/store operations and memory access handling have been optimized to minimize the latency introduced by memory accesses. The system is able to dynamically switch the execution between the GPP and the RPU when executing the original binaries of the input application. Our proof-of-concept prototype achieved geometric mean speedups of 1.60× and 1.18× for, respectively, a set of 37 benchmarks and a subset considering the 9 most complex benchmarks. With respect to a previous version of our approach, we achieved geometric mean speedup improvements from 1.22 to 1.53 for the 10 benchmarks previously used.
Nuno Miguel Cardanha Paulino, João Canas Ferreira, João M. P. Cardoso
ACM Trans. Reconfigurable Technol. Syst.3
2014 A clustering-based approach for exploring sequences of compiler optimizations
abstract
In this paper we present a clustering-based selection approach for reducing the number of compilation passes used in search space during the exploration of optimizations aiming at increasing the performance of a given function and/or code fragment. The basic idea is to identify similarities among functions and to use the passes previously explored each time a new function is being compiled. This subset of compiler optimizations is then used by a Design Space Exploration (DSE) process. The identification of similarities is obtained by a data mining method which is applied to a symbolic code representation that translates the main structures of the source code to a sequence of symbols based on transformation rules. Experiments were performed for evaluating the effectiveness of the proposed approach. The selection of compiler optimization sequences considering a set of 49 compilation passes and targeting a Xilinx MicroBlaze processor was performed aiming at latency improvements for 41 functions from Texas Instruments benchmarks. The results reveal that the passes selection based on our clustering method achieves a significant gain on execution time over the full search space still achieving important performance speedups.
Luiz G. A. Martins, Ricardo Nobre, Alexandre C. B. Delbem, Eduardo Marques, João M. P. Cardoso
IEEE Congress on Evolutionary Computation5
2014 Trace-Based Reconfigurable Acceleration with Data Cache and External Memory Support
abstract
This paper presents a binary acceleration approach based on extending a General Purpose Processor (GPP) with a Reconfigurable Processing Unit (RPU), both sharing an external data memory. In this approach repeating sequences of GPP instructions are migrated to the RPU. The RPU resources are selected and organized off-line using execution trace information. The RPU core is composed of Functional Units (FUs) that correspond to single CPU instructions. The FUs are arranged in stages of mutually independent operations. The RPU can enable several stages in tandem, depending on the data dependencies. External data memory accesses are handled by a configurable dual-port cache. A prototype implementation of the architecture on a Spartan-6 FPGA was validated with 12 benchmarks and achieved an overall geometric mean speedup of 1.91x.
Nuno Miguel Cardanha Paulino, João Canas Ferreira, João M. P. Cardoso
ISPA3
2014 Exploration of compiler optimization sequences using clustering-based selection
abstract
Due to the large number of optimizations provided in modern compilers and to compiler optimization specific opportunities, a Design Space Exploration (DSE) is necessary to search for the best sequence of compiler optimizations for a given code fragment (e.g., function). As this exploration is a complex and time consuming task, in this paper we present DSE strategies to select optimization sequences to both improve the performance of each function and reduce the exploration time. The DSE is based on a clustering approach which groups functions with similarities and then explore the reduced search space provided by the optimizations previously suggested for the functions in each group. The identification of similarities between functions uses a data mining method which is applied to a symbolic code representation of the source code. The DSE process uses the reduced set identified by clustering in two ways: as the design space or as the initial configuration. In both ways, the adoption of a pre-selection based on clustering allows the use of simple and fast DSE algorithms. Our experiments for evaluating the effectiveness of the proposed approach address the exploration of compiler optimization sequences considering 49 compilation passes and targeting a Xilinx MicroBlaze processor, and were performed aiming performance improvements for 41 functions. Experimental results reveal that the use of our new clustering-based DSE approach achieved a significant reduction on the total exploration time of the search space (18x over a Genetic Algorithm approach for DSE) at the same time that important performance speedups (43% over the baseline) were obtained by the optimized codes.
Luiz G. A. Martins, Ricardo Nobre, Alexandre C. B. Delbem, Eduardo Marques, João M. P. Cardoso
LCTES5
2014 A DSL for specifying run-time adaptations for embedded systems: an application to vehicle stereo navigation
André C. Santos, João M. P. Cardoso, Pedro C. Diniz, Diogo R. Ferreira, Zlatko Petrov
J. Supercomput.2
2013 An automatic tool flow for the combined implementation of multi-mode circuits
abstract
A multi-mode circuit implements the functionality of a limited number of circuits, called modes, of which at any given time only one needs to be realised. Using run-time reconfiguration of an FPGA, all the modes can be implemented on the same reconfigurable region, requiring only an area that can contain the biggest mode. Typically, conventional run-time reconfiguration techniques generate a configuration for every mode separately. To switch between modes the complete reconfigurable region is rewritten, which often leads to very long reconfiguration times. In this paper we present a novel, fully automated tool flow that exploits similarities between the modes and uses Dynamic Circuit Specialization to drastically reduce reconfiguration time. Experimental results show that the number of bits that is rewritten in the configuration memory reduces with a factor from 4.6× to 5.1× without significant performance penalties.
Brahim Al Farisi, Karel Bruneel, João M. P. Cardoso, Dirk Stroobandt
DATE3
2013 The MATISSE MATLAB compiler
abstract
This paper describes MATISSE, a MATLAB to C compiler targeting embedded systems that is based on Strategic and Aspect-Oriented Programming concepts. MATISSE takes as input: (1) MATLAB code and (2) LARA aspects related to types and shapes, code insertion/removal, and specialization based directives defining default variable values. In this paper we also illustrate the use of MATISSE in leveraging data types and shapes to generate customized C code suitable for high-level hardware synthesis tools. The preliminary experimental results presented here reveal the described approach to yield performance results for the resulting hardware and software references implementations that are comparable in terms of performance with hand-crafted solutions but derived automatically at a fraction of the cost.
João Bispo, Pedro Pinto 0002, Ricardo Nobre, Tiago Carvalho 0001, João M. P. Cardoso, Pedro C. Diniz
INDIN5
2013 Enriching MATLAB with aspect-oriented features for developing embedded systems
João M. P. Cardoso, João M. Fernandes 0001, Miguel P. Monteiro 0001, Tiago Carvalho 0001, Ricardo Nobre
J. Syst. Archit.1
2013 Transparent Trace-Based Binary Acceleration for Reconfigurable HW/SW Systems
abstract
This paper presents a novel approach to accelerate program execution by mapping repetitive traces of executed instructions, called Megablocks, to a runtime reconfigurable array of functional units. An offline tool suite extracts Megablocks from microprocessor instruction traces and generates a Reconfigurable Processing Unit (RPU) tailored for the execution of those Megablocks. The system is able to transparently movebcomputations from the microprocessor to the RPU at runtime. A prototype implementation of the system using a cacheless MicroBlaze microprocessor running code located in external memory reaches speedups from$2.2\times$to$18.2 \times$for a set of 14 benchmark kernels. For a system setup which maximizes microprocessor performance by having the application code located in internal block RAMs, speedups from$1.4 \times$to$2.8 \times$were estimated.
João Bispo, Nuno Miguel Cardanha Paulino, João M. P. Cardoso, João Canas Ferreira
IEEE Trans. Ind. Informatics3
2012 Controlling Hardware Synthesis with Aspects
abstract
The synthesis and mapping of applications to configurable embedded systems is a notoriously hard process. Tools have a wide range of parameters, which interact in very unpredictable ways, thus creating a large and complex design space. When exploring this space, designers must understand the interfaces to the various tools and apply, often manually, a sequence of tool-specific transformations making this an extremely cumbersome and error-prone process. This paper describes the use of aspect-oriented techniques for capturing synthesis strategies for tuning the performance of applications' kernels. We illustrate the use of this approach when designing application-specific architectures generated by a high-level synthesis tool. The results highlight the impact of the various strategies when targeting custom hardware and expose the difficulties in devising these strategies.
João M. P. Cardoso, Tiago Carvalho 0001, José Gabriel F. Coutinho, Pedro C. Diniz, Zlatko Petrov, Wayne Luk
DSD1
2012 Specifying Compiler Strategies for FPGA-based Systems
abstract
The development of applications for high-performance Field Programmable Gate Array (FPGA) based embedded systems is a long and error-prone process. Typically, developers need to be deeply involved in all the stages of the translation and optimization of an application described in a high-level programming language to a lower-level design description to ensure the solution meets the required functionality and performance. This paper describes the use of a novel aspect-oriented hardware/software design approach for FPGA-based embedded platforms. The design-flow uses LARA, a domain-specific aspect-oriented programming language designed to capture high-level specifications of compilation and mapping strategies, including sequences of data/computation transformations and optimizations. With LARA, developers are able to guide a design-flow to partition and map an application between hardware and software components. We illustrate the use of LARA on two complex real-life applications using high-level compilation and synthesis strategies for achieving complete hardware/software implementations with speedups of 2.5× and 6.8× over software-only implementations. By allowing developers to maintain a single application source code, this approach promotes developer productivity as well as code and performance portability.
João M. P. Cardoso, José C. Alves, Ricardo Nobre, Pedro C. Diniz, José Gabriel F. Coutinho, Wayne Luk
FCCM1
2012 Program and Aspect Metrics for MATLAB
Pedro Martins 0001, Paulo Lopes, João Paulo Fernandes, João Saraiva, João M. P. Cardoso
ICCSA (4)5
2011 A Domain-Specific Language for the Specification of Adaptable Context Inference
abstract
Context-aware mobile applications can benefit from context inference adaptation based on run-time operating conditions, such as battery life or sensor availability. Developing applications with such adaptable behavior, however, is notoriously cumbersome, as developers need to deal with low-level system interfacing and programming issues. In this paper we describe a domain-specific language (DSL) and a middleware infrastructure to support the specification, deployment and maintenance of run-time adaptable context inference processes. We illustrate the benefits of our approach via a case study, highlighting the new abstractions that facilitate the specification of adaptable behavior using different algorithms and the corresponding varying parameter settings, with a specific goal of minimizing the energy while maintanig acceptable end-application performance and accuracy.
André C. Santos, Pedro C. Diniz, João M. P. Cardoso, Diogo R. Ferreira
EUC3
2011 Fast placement and routing by extending coarse-grained reconfigurable arrays with Omega Networks
Ricardo S. Ferreira 0001, João M. P. Cardoso, Alex Damiany, Julio C. Goldner Vendramini, Tiago Teixeira
J. Syst. Archit.2
2010 On Identifying Segments of Traces for Dynamic Compilation
abstract
Typical computing systems based on general purpose processors (GPPs) are extended with coarse-grained reconfigurable arrays (CGRAs) to provide higher performance and/or energy savings. In order for applications to take advantage of these computing systems, efficient dynamic mapping techniques are required. Those dynamic mapping techniques will be responsible for automatically moving computations originally running in the GPP to the CGRA. The concept of dynamic compilation, widespread in the context of JIT compilation to GPPs, is receiving more attention by there configurable computing community. This paper presents our approach to dynamically map computations to CGRAs coupled to a GPP. Specifically, we present the identification of large sequences of instructions, MegaBlocks, being executed in a GPP. These MegaBlocks are then mapped to the target CGRA. We evaluate the potential of the MegaBlocks over Basic Blocks and Super Blocks to increase the IPC when targeting a CGRA and considering the execution of a number of representative benchmarks.
João Bispo, João M. P. Cardoso
FPL2
2010 On Identifying Patterns in Code Repositories to Assist the Generation of Hardware Templates
abstract
The identification of patterns on large repositories of code can be of paramount importance to guide the design of new hardware accelerators, to acquire the suitability of a certain hardware accelerator, and to generate application-specific architectures that maximize hardware reuse. This work intends to research and develop methods to both acquire the presence of a given pattern (map-suitability) and to identify common and highly similar patterns in code repositories (design-suggestions). The approach being proposed is based on a number of identification layers that refine the selections at each stage. We analyze two possible complementary options for a high-level layer. A first option is based on the representation of programs as a sequence of symbols and string matching and clustering algorithms are then used to expose similar patterns. A second option is based on tree matching techniques for identifying the presence of user's input patterns in the programs under inspection. We are evaluating our approach using the MiBench, MediaBench, UTDSP, and SNU code repositories. The results show the potential of our approach to identify approximate patterns that can be implemented by merging highly similar structures.
Adriano K. Sanches, João M. P. Cardoso
FPL2
2010 On identifying and optimizing instruction sequences for dynamic compilation
abstract
Typical computing systems based on general purpose processors (GPPs) can be extended with coarse-grained reconfigurable arrays (CGRAs) to provide higher performance and/or energy savings. In order for applications to take advantage of these computing systems, possibly including CGRAs varying in size, efficient dynamic compilation/mapping techniques are required. Dynamic mapping will be responsible for automatically moving computations originally running in the GPP to the CGRA. This paper presents our approach to dynamically map computations to CGRAs coupled to a GPP. Specifically, we evaluate the potential of the MegaBlock to accelerate the execution of a number of representative benchmarks when targeting an architecture based on a GPP and a CGRA. In addition, we show the impact on performance when using constant folding and propagation optimizations.
João Bispo, João M. P. Cardoso
FPT2
2010 A Query Processing Strategy for Conceptual Queries Based on Object-Role Modeling
abstract
There have been several authors asserting that conceptual query languages (CQLs) perform better for querying purposes than logical query languages such as SQL. This paper proposes a query mapping algorithm for the FConQuer system. FConQuer is a framework based on object-role modeling (ORM) schemas, which allow the end-user to formulate conceptual queries through the FConQuer language. Our mapping algorithm allows the FConQuer system to process conceptual queries based on ORM schemas. More precisely, our algorithm maps FConQuer queries to OQL.
António Rosado, João M. P. Cardoso
NSS2
2010 The Feasibility of Navigation Algorithms on Smartphones using J2ME
André C. Santos, Luís Tarrataca, João M. P. Cardoso
Mob. Networks Appl.3
2010 Providing user context for mobile and social networking applications
André C. Santos, João M. P. Cardoso, Diogo R. Ferreira, Pedro C. Diniz, Paulo Chainho
Pervasive Mob. Comput.2
2010 Preprocessing techniques for context recognition from accelerometer data
Davide Figo, Pedro C. Diniz, Diogo R. Ferreira, João M. P. Cardoso
Pers. Ubiquitous Comput.4
2009 Automatic generation of FPGA hardware accelerators using a domain specific language
abstract
This paper describes an alternative approach to direct mapping loops described in high-level languages onto FPGAs. Different from other approaches, this technique does not inherit from software pipelining techniques. The control is distributed over operations, thus a finite state machine is not necessary to control the order of operations, allowing efficient hardware implementations. The specification of a hardware block is done by means of LALP, a domain specific language specially designed to help the application of the techniques. While the language syntax resembles C, it contains certain constructs that allow programmer interventions to enforce or relax data dependences as needed, and so optimize the performance of the generated hardware blocks.
Ricardo Menotti, João M. P. Cardoso, Marcio Merino Fernandes, Eduardo Marques
FPL2
2009 LALP: A Novel Language to Program Custom FPGA-Based Architectures
abstract
Field-Programmable Gate Arrays (FPGAs) are becoming increasingly important in embedded and high-performance computing systems. They allow performance levels close to the ones obtained from Application-Specific Integrated Circuits (ASICs), while still keeping design and implementation flexibility. However, to efficiently program FPGAs, one needs the expertise of hardware developers and to master hardware description languages (HDLs) such as VHDL or Verilog. The attempts to furnish a high-level compilation flow (e.g., from C programs) still have open issues before efficient and consistent results can be obtained. Bearing in mind the FPGA resources, we have developed LALP, a novel language to program FPGAs. A compilation framework including mapping capabilities supports the language. The main ideas behind LALP is to provide a higher abstraction level than HDLs, to exploit the intrinsic parallelism of hardware resources, and to permit the programmer to control execution stages whenever the compiler techniques are unable to generate efficient implementations. In this paper we describe LALP, and show how it can be used to achieve high-performance computing solutions.
Ricardo Menotti, João M. P. Cardoso, Marcio Merino Fernandes, Eduardo Marques
SBAC-PAD2
2008 Combining Rewriting-Logic, Architecture Generation, and Simulation to Exploit Coarse-Grained Reconfigurable Architectures
abstract
In recent years, many coarse-grained reconfigurable architectures have been proposed as programmable accelerators for general purpose processors. The processing elements (PEs) of such architectures mainly differ on the computations they can directly support. Although different PEs and different interconnect resources among them are usually justified by the results presented, there have been few generic approaches able to exploit different PE computing structures while maintaining the same compilation flow. This paper shows our recent achievements concerning a design space exploration tool for an array of coarse-grained PEs. Our approach uses Rewriting Logic to map computations described by imperative software programming languages to the PEs of the target architecture, a VHDL generation step to prototype the architectures being exploited and a clock cycle-based simulator in order to achieve first assessments about the performance of the exploited architectures. Our approach can retarget different PE’s structures and complexities, and can be used to evaluate design solutions. In order to show the potential of our approach, we present results on exploiting a 1-D coarse-grained reconfigurable array as an accelerator and the effects of different PE’s structures and complexities.
Carlos Morra, João Bispo, João M. P. Cardoso, Jürgen Becker 0001
FCCM3
2007 A Data-Driven Approach for Pipelining Sequences of Data-Dependent Loops
abstract
Many video and image/signal processing applications can be structured as sequences of data-dependent tasks using a consumer/producer communication paradigm and are therefore amenable to pipelined execution. This paper presents an execution technique to speed-up the overall execution of successive, data-dependent tasks on a reconfigurable architecture. The technique pipelines sequences of data-dependent tasks by overlapping their execution subject to data-dependences. It decouples the concurrent data-path and control units and uses a custom, application data-driven, fine-grained synchronization and buffering scheme. In addition, the execution scheme allows for out-of- order, but data-dependent producer-consumer pairs not allowed by previous data-driven pipelining approaches. The approach has been exploited in the context of a high-level compiler targeting FPGAs. The preliminary experimental results reveal noticeable performance improvements and buffer size reductions for a number of benchmarks over traditional approaches.
Rui Rodrigues 0004, João M. P. Cardoso, Pedro C. Diniz
FCCM2
2007 Aggressive Loop Pipelining for Reconfigurable Architectures
abstract
In this work aims new techniques for mapping software loops to FPGAs. Extensive and aggressive use of pipelining techniques for achieving high performance solutions is the main goal. Those techniques are foreseen to effectively take advantage of the hardware synergies available in the current FPGA devices, especially the DSP blocks and the on-chip configurable memories.
Ricardo Menotti, Eduardo Marques, João M. P. Cardoso
FPL3
2007 Using Rewriting Logic to Match Patterns of Instructions from a Compiler Intermediate Form to Coarse-Grained Processing Elements
abstract
This paper presents a new and retargetable method to identify patterns of instructions with direct support in coarse-grained processing elements (PEs). The method uses a three-address code SSA (static single assignment) representation of the kernel being mapped and rewriting logic for template matching and algebraic optimizations. This approach is able to identify sets of SSA instructions that can be mapped to different PE complexities available in coarse-grained reconfigurable computing architectures. As a proof of concept, results of the approach with a number of benchmark kernels, as far as coverage of template instructions is concerned, are included.
Carlos Morra, João M. P. Cardoso, Jürgen Becker 0001
IPDPS2
2006 Regular expression matching for reconfigurable packet inspection
abstract
Recent intrusion detection systems (IDS) use regular expressions instead of static patterns as a more efficient way to represent hazardous packet payload contents. This paper focuses on regular expressions pattern matching engines implemented in reconfigurable hardware. A nondeterministic finite automata (NFA) based implementation was presented, which takes advantage of new basic building blocks to support more complex regular expressions than the previous approaches. The methodology is supported by a tool that automatically generates the circuitry for the given regular expressions, outputting VHDL representations ready for logic synthesis. Furthermore, techniques to reduce the area cost of our designs and maximize performance when targeting FPGAs were included. Experimental results show that our tool is able to generate a regular expression engine to match more than 500 IDS regular expressions (from the Snort ruleset) using only 25K logic cells and achieving 2 Gbps throughput on a Virtex2 and 2.9 on a Virtex4 device. Concerning the throughput per area required per matching non-meta character, our design is 3.4 and 10 times more efficient than previous ASIC and FPGA approaches, respectively
João Bispo, Ioannis Sourdis, João M. P. Cardoso, Stamatis Vassiliadis
FPT3
2006 A Methodology to Design FPGA-based PID Controllers
abstract
This paper presents a methodology to implement PID (proportional, integral, derivative) controllers in FPGAs (field-programmable gate arrays) using fixed-point numerical representation. The Matlab/Simulink environment is used for modeling, simulation and evaluation the performance provided by different fixed-point representations using a given control process. A static bit-width analyzer is used to give a specialized fixed-point representation for each operand/operator in the controller system. After bit-width analysis, a VHDL representation of the system is generated. Results show that the proposed methodology leads to shorten design cycles achieving important resource savings by employing specialized fixed-point representations.
João Miguel Gago Pontes de Brito Lima, Ricardo Menotti, João M. P. Cardoso, Eduardo Marques
SMC3
2005 On Estimations for Compiling Software to FPGA-based Systems
abstract
This paper presents recent advances in a compiler infrastructure to map algorithms described in a Java subset to FPGA-based platforms. We explain how delays and resources are estimated to guide the compiler through scheduling and temporal partitioning. The compiler supports complex analytical models to estimate resources and delays for each functional unit. The paper presents experimental results for a number of benchmarks. Those results also arise a question when performing temporal partitioning: shall we try to group as many computational structures in the same configuration or shall we have several configurations?.
João M. P. Cardoso
ASAP1
2005 An Infrastructure to Functionally Test Designs Generated by Compilers Targeting FPGAs
abstract
The paper presents an infrastructure to test the functionality of the specific architectures output by a highlevel compiler targeting dynamically reconfigurable hardware. It results in a suitable scheme to verify the architectures generated by the compiler, each time new optimization techniques are included or changes in the compiler are performed. We believe this kind of infrastructure is important to verify, by functional simulation, further research techniques, as far as compilation to field-programmable gate array (FPGA) platforms is concerned.
Rui Rodrigues 0004, João M. P. Cardoso
DATE2
2005 New challenges in computer science education
abstract
It is predicted that by the year 2010, 90% of the overall program code developed will be for embedded computing systems. This fact requires urgent changes in the organization of the current computer science curriculums, as advocated by a number of academics. The changes will help students deal with the idiosyncrasies of embedded systems, which requires knowledge about the computation engine, its energy consumption model, performance, interfaced artifacts, reconfigurable hardware programming, etc. This paper discusses some important issues to be included in modern computer science programs, in order to prepare students to be able to program future embedded computers. In particular, we present an approach we are attempting to implement at our institution. We also illustrate infrastructures that permit students to implement complex examples and gain deep knowledge about the topics being taught. Finally, with this paper we hope to foment a fruitful discussion on those issues.
João M. P. Cardoso
ITiCSE1
2004 An Environment for Exploring Data-Driven Architectures
Ricardo S. Ferreira 0001, João M. P. Cardoso, Horácio C. Neto
FPL2
2004 A Real Time Gesture Recognition System for Mobile Robots
Vanderlei Bonato, Adriano K. Sanches, Marcio Merino Fernandes, João M. P. Cardoso, Eduardo do Valle Simões, Eduardo Marques
ICINCO (2)4
2003 From C Programs to the Configure-Execute Model
João M. P. Cardoso, Markus Weinhardt
DATE1
2003 On Combining Temporal Partitioning and Sharing of Functional Units in Compilation for Reconfigurable Architectures
abstract
Resource virtualization on FPGA devices, achievable due to its dynamic reconfiguration capabilities, provides an attractive solution to save silicon area. Architectural synthesis for dynamically reconfigurable FPGA-based digital systems needs to consider the case of reducing the number of temporal partitions (reconfigurations) by enabling sharing of some functional units in the same temporal partition. This paper proposes a novel algorithm for automated datapath design from behavioral input descriptions (represented by an acyclic dataflow graph), which simultaneously performs temporal partitioning and sharing of functional units. The proposed algorithm attempts to minimize both the number of temporal partitions and the execution latency of the generated solution. Temporal partitioning, resource sharing, scheduling, and a simple form of allocation and binding are all integrated in a single task. The algorithm is based on heuristics and on a new concept of construction by gradually enlarging timing slots. Results show the efficiency and effectiveness of the algorithm when compared to existent approaches.
João M. P. Cardoso
IEEE Trans. Computers1
2002 Fast and Guaranteed C Compilation onto the PACT-XPP? Reconfigurable Computing Platform
abstract
We introduce the XPP-VC high-level compiler, which maps C programs to the coarse-grained XPP architecture. XPP-VC's main feature is the integration of pipeline vectorization and temporal partitioning techniques. The former provides high-throughput inner loop computations and the later allows the compilation of large programs or the use of fewer XPP processing elements. Although the preliminary results are very encouraging, improvements on the generation of configurations are still required. The evaluation we have conducted shows that only a few seconds are required to generate, from algorithms in C, the binaries to program the XPP. To our knowledge this compilation performance is unmatched by any other compiler targeting reconfigurable architectures. Moreover the compiler still achieves high-performance implementations. Since the XPP is delivered as an IP core or device to be coupled to a host processor, a future version of XPP-VC will consider co-compilation, i.e., compilation to hybrid microprocessor/XPP architectures.
João M. P. Cardoso, Markus Weinhardt
FCCM1
2002 XPP-VC: A C Compiler with Temporal Partitioning for the PACT-XPP Architecture
João M. P. Cardoso, Markus Weinhardt
FPL1
2001 Novel Algorithm Combining Temporal Partitioning and Sharing of Functional Units
João M. P. Cardoso
FCCM1
2001 Compilation Increasing the Scheduling Scope for Multi-memory-FPGA-Based Custom Computing Machines
João M. P. Cardoso, Horácio C. Neto
FPL1
1999 Macro-Based Hardware Compilation of Java(tm) Bytecodes into a Dynamic Reconfigurable Computing System
abstract
This paper presents a new approach to synthesize to reconfigurable hardware (HW) user-specified regions of a program, under the assumption of "virtual HW" support. The automation of this approach is supported by a compiler front-end and by an HW compiler under development. The front-end starts from the Java bytecodes and, therefore, supports any language that can be compiled to the JVM (Java Virtual Machine) model. It extracts from the bytecodes all the dependencies inside and between basic blocks. This information is stored in representation graphs more suitable to efficiently exploit the existent parallelism in the program than those typically used in high-level synthesis. From the intermediate representations the HW compiler exploits the temporal partitions at the behavior level, resolves memory access conflicts, and generates the VHDL descriptions at register-transfer level that will be mapped into the reconfigurable HW devices.
João M. P. Cardoso, Horácio C. Neto
FCCM1