VLDB 2026 Research / reviewers in the wild / expert
Fernando Gehm Moraes
dblp:m/FernandoGehmMoraes · also Fernando Moraes 0001
· DBLP profile ↗
76ranked-venue papers
4as first author
9since 2021 · last 2026
0000-0001-6126-6847ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 75 · 4 first-author · 9 since 2021Software engineering, systems software and programming languages · 8 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Design Space Exploration of RISC-V Vector Extension Targeting Embedded ProcessorsabstractThe RISC-V Vector Extension (RVV) offers flexibility and scalability for the growing demand for data-parallel workloads in embedded systems. However, RVV configurability introduces a large design space, the implications of which for resource-constrained processors remain insufficiently quantified. Understanding how parameters such as vector length (VLEN), maximum element size (ELEN), lane count, LMUL, and SEW affect system-level metrics enables efficient acceleration while preserving software portability and reducing verification effort. This work explores the RVV design space for embedded processors using a Zve32x subset and evaluates the impact of architectural parameters on performance, area, power, and energy efficiency. Annotated post-synthesis simulations across representative benchmarks enable quantitative trade-off analysis over multiple RVV configurations. Results show that larger VLENs and aggressive LMUL settings can yield substantial speedups but often increase area and power, whereas moderate configurations provide improved energy efficiency while meeting performance targets. The paper derives practical design guidelines for selecting RVV parameters in power- and areaconstrained embedded systems. The RISC-V core used in this work is publicly available athttps://github.com/gaph-pucrs/RS5. Willian Analdo Nunes, Antônio Vinicius Corrêa Dos Santos, Lucas Damo, Fernando Gehm Moraes |
IEEE Trans. Computers | 4 |
| 2025 | Accelerating Machine Learning with RISC-V Vector Extension and Auto-Vectorization TechniquesabstractConvolutional neural networks (CNNs) have played a significant role in the recent evolution of machine learning (ML) due to their feature extraction and pattern recognition capabilities. CNNs involve numerous Multiply and Accumulate (MAC) operations, which are computationally expensive and often require hardware acceleration to achieve acceptable performance. The RISC-V Vector (RVV) Extension is a candidate for accelerating vector processing operations, such as those found in CNNs. Several studies have explored the application of the RVV extension in various areas of ML, where it is often implemented as a coprocessor due to its complexity. This work presents an RVV implementation as a tightly coupled accelerator for a small RISC-V processor, RS5, which implements a subset of the RVV extension to maintain a non-prohibitive area overhead. The paper explores a case study of a 1-D CNN combined with the recently introduced auto-vectorization feature in GCC 14.1. Performance results show an average speedup of 1.88x compared to a scalar core, achieved without modifying the original code. Willian Analdo Nunes, Antônio Vinicius Corrêa Dos Santos, Fernando Gehm Moraes |
ISCAS | 3 |
| 2025 | Accelerating Machine Learning using RISC-V Vector Extension in a Manycore PlatformabstractThis work addresses the acceleration of convolutional neural network (CNN) inference in manycore architectures using coarse and fine-grain parallelism. Prior approaches focus on dedicated accelerators or modified NoCs, limiting flexibility. This work proposes integrating a RISC-V processor extended with the vector extensions (RVV) as general-purpose processing elements in a NoC-based manycore. The implementation applies depthwise convolution mapped across PEs and uses autovectorization provided by the compiler. Experiments on a $4 \times 4$ manycore running the first AlexNet layer achieved up to 5.70x speedup and reduced execution cycles by 82.45% compared to a scalar single-core baseline. Willian Analdo Nunes, Antônio Vinicius Corrêa Dos Santos, César A. M. Marcon, Fernando Gehm Moraes |
VLSI-SoC | 4 |
| 2025 | Conjunctive Merge Instruction to Accelerate Sparse Matrix - Dense Vector MultiplicationabstractSparse linear algebra is essential in many domains due to reduced computation and efficient memory usage. However, the irregularity of sparse data poses challenges for conventional software and hardware. While specialized accelerators offer performance gains, they lack general-purpose flexibility and rely on processor communication, creating bottlenecks. This work addresses these issues by proposing a tiling strategy to improve vector register usage and extending the RISC-V Vector (RVV) ISA with a custom merge instruction. Experiments using the gem5 simulator show that the tiled vector version achieved speedups of up to $1.30 \times(95 \%)$ sparsity) and $1.72 \times(65 \%)$. In contrast, the version with merge instructions reached up to $1.81 \times$ and $6.04 \times$, respectively, over a baseline implementation. Manuel Osterno, César A. M. Marcon, Jarbas Silveira, Fernando Gehm Moraes, Jardel Silveira |
VLSI-SoC | 4 |
| 2023 | Lightweight Authentication for Secure IO Communication in NoC-based Many-coresabstractNoC-based many-cores, with hundreds of IPs, are the current standard in the high-performance electronic industry. The attack surface on these systems increases at the same pace the complexity increases. Delegating security to software mechanisms does not guarantee system integrity, as it leaves the hardware exposed. Thus, adding hardware mechanisms in the many-core design is a requirement to execute applications safely. Proposals available in the literature include firewalls, spatial isolation, crypto cores, and PUFs, neglecting that applications communicate with peripherals, such as hardware accelerators and shared memories. This work presents a method to protect the communication between processing elements and peripherals by using a lightweight authentication process associated with hardware mechanisms for key generation and renewal. Our protocol protects the communication of applications with peripherals and detects attacks such as DoS, spoofing, and eavesdropping. Attack campaigns show the method's effectiveness in blocking such attacks without impairing the application's performance. Rafael Follmann Faccenda, Gustavo Comarú, Luciano L. Caimi, Fernando Gehm Moraes |
ISCAS | 4 |
| 2022 | Reliability Assessment of Many-Core Dynamic Thermal ManagementabstractPower density can limit the amount of energy a many-core can consume. A many-core running in its maximum performance can lead to safe temperature violation and reliability issues. The literature presents dynamic thermal management (DTM) techniques that guarantee system operation according to safe temperature restrictions. The state-of-art also targets dynamic reliability management (DRM), aiming for longer lifetime reliability, using the same actuation knobs used to control the temperature. This work assesses the reliability of DTM techniques applied in high and dynamic workloads, showing that DTMs keep the temperature within safe limits and increase system lifetime reliability. Alzemiro Henrique Lucas da Silva, Iaçanã I. Weber, Andre L. M. Martins, Fernando Gehm Moraes |
ISCAS | 4 |
| 2022 | A Fast, Accurate, and Comprehensive PPA Estimation of Convolutional Hardware AcceleratorsabstractConvolutional Neural Networks (CNN) are widely adopted for Machine Learning (ML) tasks, such as classification and computer vision. GPUs became the reference platforms for both training and inference phases of CNNs due to their tailored architecture to the CNN operators. However, GPUs are power-hungry architectures. A path to enable the deployment of CNNs in energy-constrained devices is adopting hardware accelerators for the inference phase. However, the literature presents gaps regarding analyses and comparisons of these accelerators to evaluate Power-Performance-Area (PPA) trade-offs. Typically, the literature estimates PPA from the number of executed operations during the inference phase, such as the number of MACs, which may not be a good proxy for PPA. Thus, it is necessary to deliver accurate hardware estimations, enabling design space exploration (DSE) to deploy CNNs according to the design constraints. This work proposes a fast and accurate DSE approach for CNNs using an analytical model fitted from the physical synthesis of hardware accelerators. The model is integrated with CNN frameworks, like TensorFlow, to generate accurate results. The analytic model estimates area, performance, power, energy, and memory accesses. The observed average error comparing the analytical model to the data obtained from the physical synthesis is smaller than 7%. Leonardo Juracy, Alexandre M. Amory, Fernando Gehm Moraes |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2021 | Dynamic Thermal Management in Many-Core Systems Leveraged by Abstract ModelingabstractFor years, transistor size reduction led to a linear decrease in power dissipation. This is the Dennard scaling law, which ended in 2000's technology nodes because the supply voltage no longer scales. The consequence is the increase of the power density in integrated circuits, leading to high temperatures, accelerating aging effects. Dynamic thermal management (DTM) is a technique adopted at runtime to act in the system using the components' current temperature to minimize hotspots and peak temperatures. This paper aims to present a DTM for NoC-based many-cores, using an abstract model, to estimate the temperature at the processing element level and task migration as the main actuation mechanism. Results show up to 12% of temperature reduction and almost 10oC of peak temperature reduction, highlighting the approach's effectiveness, compared to the pattering and spiral mappings. Alzemiro Henrique Lucas da Silva, Iaçanã I. Weber, Andre L. M. Martins, Fernando Gehm Moraes |
ISCAS | 4 |
| 2021 | A High-Level Modeling Framework for Estimating Hardware Metrics of CNN AcceleratorsabstractGPUs became the reference platform for both training and inference phases of Convolutional Neural Networks (CNN) due to their tailored architecture to the CNN operators. However, GPUs are power-hungry architectures. A path to enable the deployment of CNNs in energy-constrained devices is adopting hardware accelerators for the inference phase. The design space exploration of CNNs using standard approaches, such as RTL, is limited due to their complexity. Thus, designers need frameworks enabling design space exploration that delivers accurate hardware estimation metrics to deploy CNNs. This work proposes a framework to explore CNNs design space, providing power, performance, and area (PPA) estimations. The heart of the framework is a system simulator. The system simulator front-end is TensorFlow, and the back-end is performance estimations obtained from the physical synthesis of hardware accelerators, not only from components like multipliers and adders. The first set of results evaluate the CNN accuracy using integer quantization, the accelerators PPA after physical synthesis, and the benefits of using a system simulator. These results allow a rich design space exploration, enabling selecting the best set of CNN parameters to meet the design constraints. Leonardo Juracy, Matheus T. Moreira, Alexandre M. Amory, Alexandre F. Hampel, Fernando Gehm Moraes |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2020 | Lightweight Cryptographic Instruction Set Extension on Xtensa ProcessorabstractThe emerging popularity of the Internet of Everything makes the security an urgent issue, as well as the need for speed to cipher and decipher any information, which is essential for the embedded devices. Unlike many works in this field, where propositions considering application specific integrated circuits (ASICs), coprocessors, field-programmable gate arrays (FPGAs) or software were presented as alternatives to raise the efficiency of execution, we addressed the enhancement of the instruction set architecture (ISA) taking advantage of a hybrid design methodology to boost the performance as well as the design itself. We validate our ISA by measuring the area overhead, memory parameters and the speedup for different optimized implementations of AES, DES, 3DES and SHA using the Cadence LX7 Processor and Xtensa platform. The proposed architectures provided an excellent tradeoff with the area, memory and cycle count performance figures. Experimental results show that the proposed ISA can reduce cycle count between 1.76 and 10.99 with a cost of 6% in average of area overhead in a lightweight processor architecture. Gabriel H. Eisenkraemer, Fernando Gehm Moraes, Leonardo Londero de Oliveira, Everton Carara |
ISCAS | 2 |
| 2020 | Open-Source NoC-Based Many-Core for Evaluating Hardware Trojan Detection MethodsabstractIn many-cores based on Network-on-Chip (NoC), several applications execute simultaneously, sharing computation, communication and memory resources. This resource sharing leads to security and trust problems. Hardware Trojans (HTs) may steal sensitive information, degrade system performance, and in extreme cases, induce physical damages. Methods available in the literature to prevent attacks include firewalls, denial-of-service detection, dedicated routing algorithms, cryptography, task migration, and secure zones. The goal of this paper is to add an HT in an NoC, able to execute three types of attacks: packet duplication, block applications, and misrouting. The paper qualitatively evaluates the attacks' effect against methods available in the literature, and its effects showed in an NoC-based many-core. The resulting system is an open-source NoC-based many-core for researchers to evaluate new methods against HT attacks. Iaçanã I. Weber, Geaninne Marchezan, Luciano L. Caimi, César A. M. Marcon, Fernando Gehm Moraes |
ISCAS | 5 |
| 2020 | Modular and Distributed Management of Many-Core SoCsabstractMany-Core Systems-on-Chip increasingly require Dynamic Multi-objective Management (DMOM) of resources. DMOM uses different management components for objectives and resources to implement comprehensive and self-adaptive system resource management. DMOMs are challenging because they require a scalable and well-organized framework to make each component modular, allowing it to be instantiated or redesigned with a limited impact on other components. This work evaluates two state-of-the-art distributed management paradigms and, motivated by their drawbacks, proposes a new one called Management Application (MA) , along with a DMOM framework based on MA. MA is a distributed application, specific for management, where each task implements a management role. This paradigm favors scalability and modularity because the management design assumes different and parallel modules, decoupled from the OS. An experiment with a task mapping case study shows that MA reduces the overhead of management resources (-61.5%), latency (-66%), and communication volume (-96%) compared to state-of-the-art per-application management. Compared to cluster-based management (CBM) implemented directly as part of the OS, MA is similar in resources and communication volume, increasing only the mapping latency (+16%). Results targeting a complete DMOM control loop addressing up to three different objectives show the scalability regarding system size and adaptation frequency compared to CBM, presenting an overall management latency reduction of 17.2% and an overall monitoring messages’ latency reduction of 90.2%. Marcelo Ruaro, Anderson C. Sant'Ana, Axel Jantsch, Fernando Gehm Moraes |
ACM Trans. Comput. Syst. | 4 |
| 2019 | Distributed SDN architecture for NoC-based many-core SoCsabstractIn the Software-Defined Networking (SDN) paradigm, routers are generic and programmable forwarding units that transmit packets according to a given policy defined by a software controller. Recent research has shown the potential of such a communication concept for NoC management, resulting in hardware complexity reduction, management flexibility, real-time guarantees, and self-adaptation. However, a centralized SDN controller is a bottleneck for large-scale systems. Marcelo Ruaro, Nedison Velloso, Axel Jantsch, Fernando Gehm Moraes |
NOCS | 4 |
| 2019 | The power impact of hardware and software actuators on self-adaptable many-core systems
Andre L. M. Martins, Rafael Garibotti, Nikil Dutt, Fernando Gehm Moraes |
J. Syst. Archit. | 4 |
| 2019 | Hierarchical adaptive Multi-objective resource management for many-core systems
Andre L. M. Martins, Alzemiro Henrique Lucas da Silva, Amir-Mohammad Rahmani, Nikil Dutt, Fernando Gehm Moraes |
J. Syst. Archit. | 5 |
| 2019 | Self-Adaptive QoS Management of Computation and Communication Resources in Many-Core SoCsabstractProviding quality of service (QoS) for many-core systems with dynamic application admission is challenging due to the high amount of resources to manage and the unpredictability of computation and communication events. Related works propose a self-adaptive QoS mechanism concerned either in communication or computation resources, lacking, however, a comprehensive QoS management of both. Assuming a many-core system with QoS monitoring, runtime circuit-switching establishment, task migration, and a soft real-time task scheduler, this work fills this gap by proposing a novel self-adaptive QoS management. The contribution of this proposal comes with the following features in the QoS management: ( i ) comprehensiveness, by covering communication and computation resources; ( ii ) online, adopting the ODA (Observe, Decide, Act) runtime closed-loop adaptation; and ( iii ) reactive and proactive decisions, by using a dynamic application profile extraction technique, which enables the QoS management to be aware of the profile of running applications, allowing it to take proactive decisions based on a prediction analysis. The proposed QoS management adopts a decentralized organization by partitioning the system in clusters, each one managed by a dedicated processor, making the proposal scalable. Results show that the proactive feature accurately extracts the applications’ profile, and can prevent future QoS violations. The synergy of reactive and proactive decisions was able to sustain QoS, reducing the deadline miss rate by 99.5% with a severe disturbance in communication and computation levels, and avoiding deadline misses up to 70% of system utilization. Marcelo Ruaro, Axel Jantsch, Fernando Gehm Moraes |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2018 | Exploring the Impact of Soft Errors on NoC-based Multiprocessor SystemsabstractSoftware reliability is an essential design metric in emerging large-scale multiprocessor embedded systems. Designers should identify soft error susceptibility of multiple applications executing in parallel early in the design time to ensure reliable system operation. This work proposes a non-intrusive fault injection engine that enables to conduct bespoke soft error analysis, allowing to identify and understand the soft error propagation through the processing elements (PEs). The proposed fault injection campaign evaluates the impact of soft errors considering real benchmarks in an RTL model of a distributed-memory NoC-based multiprocessor. Experiments demonstrate that 19% of soft errors are propagated to other PEs, where 31.6% of them led to erroneous computation and 58.4% to a system crash. Thus, the fault analysis must consider not only its local effect on the processor and memory but also how the fault propagates to other system components. Felipe T. Bortolon, Geancarlo Abich, Sergio Bampi, Ricardo Augusto da Luz Reis, Fernando Gehm Moraes, Luciano Ost |
ISCAS | 5 |
| 2018 | An LSSD Compliant Scan Cell for Flip-FlopsabstractMost recent timing resilient templates are using asynchronous design techniques and integrating both flip-flops and latches in their design to enable more aggressive performance improvement and reduction in energy consumption. Despite these benefits, they impose challenges in terms of testability because both latches and flip-flops typically use different test protocols. This paper presents an optimized scan cell for flip-flops which is compatible with the protocol used by scannable latches. By using the proposed cell, it is possible to have latches and flip-flops in the same scan chain and the DfT flow fully automated by commercial EDA tools. Experimental results show that the proposed cell reduces silicon area, leakage, and dynamic power compared to the original cell. Leonardo Juracy, Matheus T. Moreira, Felipe A. Kuentzer, Fernando Gehm Moraes, Alexandre M. Amory |
ISCAS | 4 |
| 2018 | Software-Defined Networking Architecture for NoC-based Many-CoresabstractThe Software-Defined Networking (SDN) is a communication paradigm adopted in computer networking. The SDN assumes simple and programmable routers, removing the control logic from the routers' level, and assigning it to a high-level controller (software), which is responsible for defining the path of the communication flows at run-time. The controller can implement different communication rules to define the paths, as Quality-of-Service (QoS), fault-tolerance, and security. Many-cores may adopt the SDN paradigm due to its advantages: reduced hardware complexity, high reusability, and flexible management of communication policies. However, the challenge to apply the SDN may be the overhead for defining the paths in software against hardware-based approaches. The goal of this paper is to show that SDN can be a viable alternative for NoC management in many-core systems. This work proposes a generic SDN architecture for many-cores, detailing both the hardware and software designs. We compare the quality of the proposal with a state-of-the-art search path mechanism (hardware implemented), in a QoS case-study providing Circuit-Switching (CS) for applications. Results show that the SDN paradigm achieves similar performance than the hardware-based technique regarding path length. Hardware implemented mechanisms present a reduced latency to establish the paths. As the path establishment occurs once for each application flow, results show that the search path latency of SDN is not an actual drawback, as it could be expected. Marcelo Ruaro, Henrique Martins Medina, Alexandre M. Amory, Fernando Gehm Moraes |
ISCAS | 4 |
| 2017 | Activation of secure zones in many-core systems with dynamic reroutingabstractMany-core architectures provide massive parallelism and high performance to the users. They also introduce key challenges regarding security. The main threat rises from the resource sharing. Adoption of firewalls, encryption mechanisms and resource isolation (processor or memory) are common strategies to treat the security threats. The two first strategies present high hardware cost, while the resource isolation does not protect the communication infrastructure. This paper proposes to protect communication and computation simultaneously, creating continuous Secure Zones (SZ) at runtime. The SZ is an isolated area in the system, preventing traffic flows to cross the boundaries of the zone, reserving the Processing Elements (PEs) inside the SZ to execute a secure application (Appsec). The Appsec traffic is enclosed inside the SZ, and any other traffic crossing the SZ is rerouted to the outside of the SZ. Experiments evaluate the cost to close and open an SZ, the impact in the packet latency with and without the SZ in the presence of malicious flows, and the performance penalty in non-secure applications due to the rerouting process when an SZ is closed. Results show that: (1) the latency to start an Appsec is small, 1.2 μs for an SZ with 25 PEs; (2)the isolation process prevents Deny-of-Service (DoS), timing, and spoofing attacks and guarantees confidentiality and integrity;(3)the overhead in the non-secure applications' execution time is also small, being the worst-case 2.56% longer. Luciano L. Caimi, Vinicius Fochi, Eduardo Wächter, Daniel Munhoz, Fernando Gehm Moraes |
ISCAS | 5 |
| 2017 | Runtime energy management under real-time constraints in MPSoCsabstractThe workload of many-core systems includes real-time (RT) applications. Obtain energy savings while executing RT applications is a challenge due to the RT timing constraints. The techniques used to reduce the consumption, as dynamic voltage and frequency scaling (DVFS), usually delays the applications, leading to constraint violations. Most works in the literature ensure RT constraints based on design-time analysis of the expected workload by defining the voltage and frequency levels. Our main contribution is the proposal of a Runtime Energy Management for RT applications (RT-REM), using monitored data, and DVFS and task mapping as actuation policies. Supervising the application slack time enables to save energy of RT applications. Evaluations on many-core systems (up to 144 PEs), achieves 18% in energy savings, keeping constraint misses below 2.5%. Our proposal stands out from related works regarding scalability, realistic and accurate energy estimation, and absence of design-time analysis of the application set. Andre L. M. Martins, Marcelo Ruaro, Anderson C. Sant'Ana, Fernando Gehm Moraes |
ISCAS | 4 |
| 2017 | Demystifying the cost of task migration in distributed memory many-core systemsabstractTask migration plays a major role in the implementation of runtime adaptive techniques for many-core systems, as thermal and power management, load balancing, QoS, and fault tolerance. A fast task migration protocol contributes to implementing self-adaptive techniques with low overhead. State-of-the-art proposals still have limitations, with an important impact on the applications' execution time due to the latency to migrate tasks. Aware of the number of simultaneous task migrations required by self-adaptive techniques, this work proposes a low latency tasks migration protocol for many-core systems with distributed memory hierarchy. Our technique eliminates checkpoints, task code replication, enables simultaneous task migrations even in tasks of the same application and parallelizes the task migration with the application execution. These features induce a low latency in the task migration, demystifying the cost to adopt task migration in distributed memory systems. Results compare the proposed approach to the related works. Marcelo Ruaro, Fernando Gehm Moraes |
ISCAS | 2 |
| 2017 | Exploiting performance, dynamic power and energy scaling in full-system simulatorsabstractSummary Energy consumption constraints have become a critical issue in Multiprocessor Systems on Chip (MPSoC) designs. Whereas processor performance comes with a high power cost, there is an increasing interest in exploring the trade‐off between power and performance, taking into account the target application domain. Dynamic Voltage and Frequency Scaling (DVFS) techniques adaptively scale frequency or voltage level of CPU allowing it to reach just enough performance to process the system workload while meeting throughput constraints, and thereby, reducing the energy consumption. To explore this wide design space for energy efficiency and performance, hardware and software components, a system‐level simulation infrastructure must provide features to evaluate power savings mechanisms in early stages of the design. This paper presents an extension work of a framework for MPSoCs designs to support DVFS in MPSoCs simulators and evaluates three DVFS mechanisms. Our experiments show that applying DVFS in the system can save power and energy consumption, with negligible loss of performance. Copyright © 2016 John Wiley & Sons, Ltd. Liana Dessandre Duenha, Guilherme A. Madalozzo, Fernando Gehm Moraes, Rodolfo Azevedo |
Concurr. Comput. Pract. Exp. | 3 |
| 2016 | Dynamic Real-Time Scheduler for Large-Scale MPSoCsabstractLarge-scale MPSoCs requires a scalable and dynamic real-time (RT) task scheduler, able to handle non-deterministic computational behaviors. Current proposals for MPSoCs have limitations, as lack of scalability, complex static steps, validation with abstract models, or are not flexible to enable changes at runtime of the RT constraints. This work proposes a hierarchical task scheduler with monitoring features. The scheduler is dynamic, supporting changes in RT constraints at runtime. An API enables these features allowing to the application developer to reconfigure the tasks' period, deadline, and execution time by annotating the task code. At runtime, according to the task execution, the scheduler handles the API calls and adjust itself to ensure RT guarantees according to the new constraints. Scalability is ensured by dividing the scheduler into two hierarchical levels: LS (Local Schedulers), and CS (Cluster Schedulers). The LS runs at the processor level, using the LST (Least Slack-Time) algorithm. The CS runs at the cluster level, i.e., a group of processors controlled by a manager processor. The CS receives messages from the LSs, informing the processor slack-time, deadline violations, and RT changes. The CS implements an RT adaptation heuristic, triggering task migrations according to RT reconfiguration or deadline misses. Results show a negligible overhead in the applications' execution time and the fulfillment of the applications' RT constraints even with a high degree of resources sharing, in both processors and NoC. Marcelo Ruaro, Fernando Gehm Moraes |
ACM Great Lakes Symposium on VLSI | 2 |
| 2016 | Efficient traffic balancing for NoC routing latency minimizationabstractModern technologies of integrated circuits allow billions of transistors arranged into a single chip, enabling to implement complex systems, which need a scalable and parallel communication architecture. Network-on-Chip (NoC) is a natural candidate to fulfill such communication requirements, providing high performance when the communication demands are balanced. This work proposes a new static balancing method that uses the application's traffic pattern for NoC latency reduction. This method allows the generation of a deterministic routing algorithm with simplistic implementation and low latency. Experimental results compare four balancing methods, showing the improvement of the proposed static balancing concerning the average NoC latency. Joao Marcelo Ferreira, Jarbas Silveira, Jardel Silveira, Rodrigo Cataldo, Thais Webber, Fernando Gehm Moraes, César A. M. Marcon |
ISCAS | 6 |
| 2016 | DMNI: A specialized network interface for NoC-based MPSoCsabstractCurrent proposals of NoC-based MPSoC adopt an NI (Network Interface) interconnected to a DMA (Direct Memory Access) module to enable the communication between processors through the NoC. The adoption of both modules decouples computation from communication, and a standard interface at the NI provides an abstract way for designers to connect new cores. However, this architecture is inherited from bus-based architectures and can be optimized, by removing unnecessary interfaces, signals, and registers. This paper presents a specialized communication interface for NoC-based MPSoCs, called DMNI (Direct Memory Network Interface). The DMNI merges the functionalities of the DMA and the NI into a single component, directly connecting the NoC router with the processor memory. To avoid stalls in the communication, the design of the DMNI supports simultaneous packet reception and transmission. A simplified and generic programming interface exposes the DMNI services to the software layer. Results show a reduction in the silicon area and performance improvement in the packet transmission. Marcelo Ruaro, Felipe B. Lazzarotto, César A. M. Marcon, Fernando Gehm Moraes |
ISCAS | 4 |
| 2016 | MPSoCBench: A benchmark for high-level evaluation of multiprocessor system-on-chip tools and methodologies
Liana Dessandre Duenha, Guilherme A. Madalozzo, Thiago Santiago, Fernando Gehm Moraes, Rodolfo Azevedo |
J. Parallel Distributed Comput. | 4 |
| 2016 | Hierarchical energy monitoring for task mapping in many-core systems
Guilherme M. Castilhos, Marcelo Mandelli, Luciano Ost, Fernando Gehm Moraes |
J. Syst. Archit. | 4 |
| 2015 | Fault recovery protocol for distributed memory MPSoCsabstractFault handling mechanisms become more relevant as systems integrate more hardware logic. For instance, current multi-processor system-on-chips (MPSoCs) consists of hundreds of processors connected by an interconnection network. This type of system can only be cost effective if it can handle faults on its main components (i.e. processors and interconnect). Traditional fault recovery approaches for multi-processors were adapted from the domain of cluster of computers and might be more complex than required for common MPSoC applications domains. This paper presents a lightweight online fault recovery for embedded processors of MPSoCs based on distributed memory. This approach automatically restarts affected applications reallocating tasks to healthy processors. All steps are performed at the kernel level, without changing user application code. Results show very short recovery time, from 110 μs to 425 μs with a 100MHz clock, which are mostly dominated by the size of the reallocated tasks. Francisco F. S. Barreto, Alexandre M. Amory, Fernando Gehm Moraes |
ISCAS | 3 |
| 2015 | An integrated method for implementing online fault detection in NoC-based MPSoCsabstractThe continuing development of the silicon technology leads to systems with hundreds of processors interconnected by a network on chip (NoC-based MPSoCs). On one hand, the nanotechnology enables to develop such complex systems, but, on the other hand, the vulnerability to faults increases. The literature presents partial fault-tolerant approaches, targeting specific parts of the system, as high-level methods, router level, link level, and routing algorithms. There is an important gap in the literature, with an integrated method, from the fault detection at the router level up to the fault recovery and correct execution of applications in a real MPSoC. This is the goal of the present work, to present a method with fault-tolerant techniques from the physical to the transport layers. The MPSoC is modeled at the RTL level, using VHDL. A fault campaign injection (5 simultaneous injected faults) resulted in 2,000 simulated scenarios. Results demonstrated the effectiveness of the proposal, with most of the scenarios working correctly with routers operating in degraded mode, with an impact on the execution time below 1%. Vinicius Fochi, Eduardo Wächter, Augusto Erichsen, Alexandre M. Amory, Fernando Gehm Moraes |
ISCAS | 5 |
| 2015 | A context saving fault tolerant approach for a shared memory many-core architectureabstractMechanisms for runtime fault-tolerance in many-core architectures are mandatory to cope with transient and permanent faults. This issue is even more relevant with aggressive technology nodes due to process variability, aging effects, and susceptibility to upsets, among other factors. This work proposes to save periodically the context and to re-schedule tasks to the last reliable known state and avoid the faulty processor. This technique is implemented on an embedded multicore architecture named P2012. The proposed fault-tolerant approach induces a limited overhead of 9.37% in an industrial image processing application while guaranteeing a full-error recovery if any error is detected. Eduardo Wächter, Nicolas Ventroux, Fernando Gehm Moraes |
ISCAS | 3 |
| 2015 | Runtime Adaptive Circuit Switching and Flow Priority in NoC-Based MPSoCsabstractWith the significant increase in the number of processing elements in NoC-based MPSoCs, communication becomes, increasingly, a critical resource for performance gains and quality-of-service (QoS) guarantees. The main gap observed in the NoC-based MPSoCs literature is the runtime adaptive techniques to meet QoS. In the absence of such techniques, the system user must statically define, for example, the scheduling policy, communication priorities, and the communication switching mode of applications. The goal of this paper is to investigate the runtime adaptation of the NoC resources, according to the QoS requirements of each application running in the MPSoC. This paper adopts an NoC architecture with duplicated physical channels, adaptive routing, support to flow priorities and simultaneous packet and circuit switching. The monitoring and adaptation management is performed at the operating system level, ensuring QoS to the monitored applications. The QoS acts in the flow priority and the switching mode. Monitoring and QoS adaptation were implemented in software, resulting in flexibility to apply the techniques to other platforms or include other adaptive techniques, as task migration or DVFS. Applications with latency and throughput deadlines run concurrently with best-effort applications. Results with synthetic and real application reduced in average 60% the latency violations, ensuring smaller jitter and throughput. The execution time of applications is not penalized applying the proposed QoS adaptation methods. Marcelo Ruaro, Everton Carara, Fernando Gehm Moraes |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2014 | A monitored NoC with runtime path adaptationabstractNetworks-on-chip (NoCs) are already a common choice of communication infrastructure for complex systems-on-chip (SoCs) containing a large number of processing resources and with critical communication requirements. A NoC provides several advantages, such as higher scalability, efficient energy management, higher bandwidth and lower average latency, when compared to bus-based systems. Experiments with applications running on NoCs with more than 10% of bandwidth usage show that most of a typical message latency refers to buffered packets waiting to enter the NoC, while the latency portion that depends on packets traversing the NoC is often negligible. This work proposes a Monitored NoC called MoNoC, which is based on a monitoring mechanism and on the exchange of high-priority control packets. Practical experiments show that our fast adaptation method enables transmitting packets with smaller latencies, by using non-congested NoC areas, which reduces the most significant part of message latency. Edson I. Moreno, Thais Webber, César A. M. Marcon, Fernando Gehm Moraes, Ney Laert Vilar Calazans |
ISCAS | 4 |
| 2014 | Tool-set for NoC-based MPSoC debugging - A protocol view perspectiveabstractSoftware development becomes an important issue in today's MPSoC design. Due to the inherent non-deterministic behavior of MPSoCs, they are prone to concurrency bugs. Debugging tools for MPSoC may be grouped in the following classes: simulators, parallel software development environments, NoC debuggers. An important gap is observed concerning a complete NoC-based MPSoC: tools to inspect the traffic exchanged between processing elements in a higher abstraction level, and not simply as raw data. This is the goal of the paper: propose a new class of debugging tools, able to trace the messages exchanged between PEs, enabling debugging at the protocol level. Examples of protocols include communication between tasks, mapping heuristics, monitoring schemes for QoS, among others. The paper presents the proposed debug framework, as well as a task migration protocol as case study. Marcelo Ruaro, Everton Carara, Fernando Gehm Moraes |
ISCAS | 3 |
| 2014 | MoNoC: A monitored network on chip with path adaptation mechanism
Edson I. Moreno, Thais Webber, César A. M. Marcon, Fernando Gehm Moraes, Ney Laert Vilar Calazans |
J. Syst. Archit. | 4 |
| 2014 | Differentiated Communication Services for NoC-Based MPSoCsabstractThe adoption of Networks-on-Chip (NoCs) as the communication infrastructure for complex integrated systems is a fact, and has been promoted by the growing number of processing elements integrated in current MPSoCs. These are designed to execute several applications in parallel, with different communication requirements and distinct levels of required quality of service. To meet these restrictions, most designs customize the MPSoC at design time, using specific NoC communication services as adaptive routing algorithms, priorities, and connections. However, MPSoCs are increasingly used in embedded systems, where new applications may be added at runtime, characterizing dynamic workload scenarios. Such scenarios require adaptability at runtime, with applications having the possibility to select the most appropriate communication service according to their respective requirements. The goal of the present work is to link the hardware level of NoCs to the MPSoC application level, proposing the development of a communication API that exposes the communication services offered by the NoC to the application developer. Executing real and synthetic applications in two different MPSOCs, and using four different NoC communication services enabled to demonstrate the efficiency of the proposed approach to meet applications requirements. Everton Carara, Ney Laert Vilar Calazans, Fernando Gehm Moraes |
IEEE Trans. Computers | 3 |
| 2014 | Beware the Dynamic C-ElementabstractThe C-element is a well known component of asynchronous circuits. To overcome problems of current CMOS technologies, its use has even been extended to specific domains of the synchronous paradigm, such as clock generation, clock gating, and registers. An economical implementation of this component is the dynamic C-element. Its advantages over static implementations are reduced power, transition, and propagation delays as well as lower silicon area. Yet, research evaluating its electrical behavior, functionality, and robustness is scarce. This brief presents an in-depth analysis of the dynamic C-element electrical behavior. The analysis points to a constrained nature, which can lead to undefined output logic values, as well as excessive static power consumption. The brief also proposes a technique for robust design of such components that avoids such undefined values. Matheus T. Moreira, Fernando Gehm Moraes, Ney Laert Vilar Calazans |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2013 | Topology-agnostic fault-tolerant NoC routing methodabstractRouting algorithms for NoCs were extensively studied in the last 12 years, and proposals for algorithms targeting some cost function, as latency reduction or congestion avoidance, abound in the literature. Fault-tolerant routing algorithms were also proposed, being the table-based approach the most adopted method. Considering SoCs with hundred of cores in a near future, features as scalability, reachability, and fault assumptions should be considered in the fault-tolerant routing methods. However, the current proposals some have some limitations: (1) increasing cost related to the NoC size, compromising scalability; (2) some healthy routers may not be reached even if there is a source-target path; (3) some algorithms restricts the number of faults and their location to operate correctly. The present work presents a method, inspired in VLSI routing algorithms, to search the path between source-target pairs where the network topology is abstracted. Results present the routing path for different topologies (mesh, torus, Spidergon and Hierarchical-Spidergon) in the presence of faulty routers. The silicon area overhead and total execution time of the path computation is small, demonstrating that the proposed method may be adopted in NoC designs. Eduardo Wächter, Augusto Erichsen, Alexandre M. Amory, Fernando Gehm Moraes |
DATE | 4 |
| 2013 | Power-aware dynamic mapping heuristics for NoC-based MPSoCs using a unified model-based approachabstractThe mapping of tasks to processing elements of an MPSoC has critical impact on system performance and energy consumption. To cope with complex dynamic behavior of applications, it is common to perform task mapping during runtime so that the utilization of processors and interconnect can be taken into account when deciding the allocation of each task. This paper has two major contributions, one of them targeting the general problem of evaluating dynamic mapping heuristics in NoC-based MPSoCs, and another focusing on the specific problem of finding a task mapping that optimizes energy consumption in those architectures. Luciano Ost, Marcelo Mandelli, Gabriel Marchesan Almeida, Leandro Möller, Leandro Soares Indrusiak, Gilles Sassatelli, Pascal Benoit, Manfred Glesner, Michel Robert, Fernando Gehm Moraes |
ACM Trans. Embed. Comput. Syst. | 10 |
| 2012 | Proposal and evaluation of a task migration protocol for NoC-based MPSoCsabstractTask migration is a well-known strategy adopted in distributed systems for load balancing. but the adoption of such strategy in NoC-based MPSoC is scarce in the literature. This paper proposes a complete task migration protocol for NoC-based MPSoCs. The migration transfers the task code, data and context to another PE. The paper presents the communication strategy to ensure coherence in the messages delivery, the heuristic to compute the new task location, and the procedure to inform the new task position. Results evaluate the cost of the task migration using a real MPSoC (described in synthesizable VHDL), demonstrating that the cost to migrate a given task has a small impact in the system performance, enabling its use to improve the overall system performance. Fernando Gehm Moraes, Guilherme A. Madalozzo, Guilherme M. Castilhos, Everton Carara |
ISCAS | 1 |
| 2012 | A spectrum of MPSoC models for design space exploration and its useabstractImplementing on-chip multiprocessors is enabled by the use of deep submicron technologies and constitutes today a daunting task, due to the complexity of their design and verification. The development of such devices can be facilitated by the use of a carefully crafted set of models for each implementation step. This paper proposes the use of such a set of abstract models at several levels. These improve simulation speed and observability in one sense and level of detail and precision in the opposite sense. Initial development tasks such as software development and functionality specification refinement can evolve fast with very abstract models, while confidence in the final implementation can be achieved with lower level models. Our basic multi-processor system on a chip is configurable in several parameters, including number of processors, type of employed communication architecture and combination of abstraction levels used in the description of the distinct modules that compose the model. Carlos A. Petry, Eduardo Wächter, Guilherme M. Castilhos, Fernando Gehm Moraes, Ney Laert Vilar Calazans |
RSP | 4 |
| 2012 | Enabling Adaptive Techniques in Heterogeneous MPSoCs Based on VirtualizationabstractThis article explores the use of virtualization to enable mechanisms like task migration and dynamic mapping in heterogeneous MPSoCs, thereby targeting the design of systems capable of adapt their behavior to time-changing workloads. Because tasks may have to be mapped to target processors with different instruction set architectures, we propose the use of Low Level Virtual Machine (LLVM) to postcompile the tasks at runtime depending on their target processor. A novel dynamic mapping heuristic is also proposed, aiming to exploit the advantages of specialized processors while taking into account the overheads imposed by virtualization. Extensive experimental work at different levels of abstraction---FPGA prototype, RTL and system-level simulation---is presented to evaluate the proposed techniques. Luciano Ost, Sameer Varyani, Leandro Soares Indrusiak, Marcelo Mandelli, Gabriel Marchesan Almeida, Eduardo Wächter, Fernando Gehm Moraes, Gilles Sassatelli |
ACM Trans. Reconfigurable Technol. Syst. | 7 |
| 2011 | Evaluating energy consumption of homogeneous MPSoCs using spare tilesabstractThe yield of homogeneous network-on-chip based multi-processor chips can be improved with the addition of spare tiles. However, the impact of this reliability approach on the chip energy consumption is not documented. For instance, in a homogeneous MPSoC, application tasks can be placed onto any tile of a defect-free chip. On the other hand, a chip with defective tile needs a special task placement, where the faulty tile is avoided. This paper presents a task placement tool and the evaluation of energy consumption of homogeneous NoC-based MPSoCs with spare tiles. Results show NoC energy consumption overhead ranging from 1 to 10% when considering up to three faults randomly distributed over the tiles of a 3×4 mesh network. The results also indicate that faults on the central tiles typically have more impact on energy overhead. Alexandre M. Amory, Luciano Ost, César A. M. Marcon, Fernando Gehm Moraes, Marcelo Lubaszewski |
DATE | 4 |
| 2011 | Achieving composability in NoC-based MPSoCs through QoS management at software levelabstractMultiprocessors systems on chip (MPSoCs) have become the de-facto standard in embedded systems. The use of Networks-on-chip (NoCs) provides to these platforms scalability and support for parallel transactions. The computational power of these architectures enables the simultaneous execution of several applications, with different time constraints. However, as the number of applications executing simultaneously increases, the performance of such applications may be affected due to resources sharing. To ensure applications requirements are met, mechanisms are necessary for ensuring proper isolation. Such a feature is referred to as composability. As the NoC is the main shared component in NoC-based MPSoCs, quality-of-service (QoS) mechanisms are mandatory to meet application requirements in term of communication. In this work, we propose a hardware/software approach to achieve applications composability by means of QoS management mechanisms at the software level. The conducted experiments show the efficiency of the proposed method in terms of throughput, latency and jitter for a real time application sharing communication resources with best-effort applications. Everton Carara, Gabriel Marchesan Almeida, Gilles Sassatelli, Fernando Gehm Moraes |
DATE | 4 |
| 2011 | Dynamic Flow Reconfiguration Strategy to Avoid Communication Hot-SpotsabstractApplication-specific Network-on-Chip allows optimization for the interconnection to minimize its cost. When used with streaming applications, large flows of data can be predicted. However, these flows can be modified during the applications providing a dynamic flow graph. In that case, applying on off-line optimization leads to an over-sizing of the NoC. On the other hand, dynamic reconfiguration leads to unordered data deliveries with costly re-ordering units. In this paper, we propose a coarse grain dynamic reconfiguration which avoids data re-ordering requirement. We show that the proposed solution is efficient to deal with communication hot-spots, with a small area overhead, and can save up to 33% of latency. Romain Prolonge, Fabien Clermidy, Leonel Tedesco, Fernando Gehm Moraes |
DSD | 4 |
| 2011 | Predictive Dynamic Frequency Scaling for Multi-Processor Systems-on-ChipabstractThis paper proposes a novel strategy for optimizing resources in Multi-Processor Systems-on-Chip (MPSoC). The approach is based on using control-loop feedback mechanism to maximize the efficiency on exploiting available resources such as CPU time, operating frequency, etc. Each Processing Element (PE) in the architecture is equipped with a frequency scaling module responsible for tuning the frequency of processors at run-time according to the application requirements. Results show the system's capability of adapting to disturbing conditions. For validation purposes we have implemented a multi-threaded MJPEG decoder together with an ADPCM audio decoder and a FIR. Gabriel Marchesan Almeida, Rémi Busseuil, Everton Carara, Nicolas Hebert, Sameer Varyani, Gilles Sassatelli, Pascal Benoit, Lionel Torres, Fernando Gehm Moraes |
ISCAS | 9 |
| 2011 | Energy-aware dynamic task mapping for NoC-based MPSoCsabstractTo cope with the dynamic workload of actual NoC-based MPSoCs, dynamic mechanisms are required to guarantee the application requirements. Application mapping may drastically influence the system performance and the energy consumption, which can be crucial to the success (or failure) of a product, even more for battery-powered embedded systems. In this context, the current work presents an energy-aware dynamic task mapping heuristic, which was evaluated in a real NoC-based MPSoC platform. Results show that the proposed heuristic may reduces up to 22.8% of the communication energy consumption compared to other dynamic mapping heuristics. Marcelo Mandelli, Luciano Ost, Everton Carara, Guilherme Montez Guindani, Thiago Gouvea, Guilherme Medeiros, Fernando Gehm Moraes |
ISCAS | 7 |
| 2011 | A new test scheduling algorithm based on Networks-on-Chip as Test Access Mechanisms
Alexandre M. Amory, Cristiano Lazzari, Marcelo Lubaszewski, Fernando Gehm Moraes |
J. Parallel Distributed Comput. | 4 |
| 2011 | CAFES: A framework for intrachip application modeling and communication architecture design
César A. M. Marcon, Ney Laert Vilar Calazans, Edson I. Moreno, Fernando Gehm Moraes, Fabiano Hessel, Altamiro Amadeu Susin |
J. Parallel Distributed Comput. | 4 |
| 2010 | Improving QoS of Multi-layer Networks-on-Chip with Partial and Dynamic Reconfiguration of RoutersabstractNetworks-on-Chip (NoC) allow several data transfers to occur in parallel and are indeed the communication infra-structure of future hundred-cores Systems-on-Chip (SoCs). However, if specialized modules are sending data at full speed to the NoC, Quality of Service (QoS) can be no longer guaranteed. This work presents a multi-layer mesh NoC approach to improve the QoS of such communication hungry SoCs. While one mesh layer is fixed in the system for control purposes, other data layers can be configured at runtime to provide the desired data throughput required by the application. This is accomplished by partially and dynamically reconfiguring the data layer routers. Arbitration algorithms, routing algorithms and huge crossbars are removed from the data layer routers, because all data routers in the path a configured accordingly before its utilization. A SoC following this idea was prototyped in a Virtex-4 FPGA and the Early Access Partial Flow was used to partially and dynamically reconfigure the NoC. We show that 120 (5!) different configurations are needed for each reconfigurable router with 5 bidirectional ports. Each configuration requires 33KB of memory and occupies 32 CLBs of area. Leandro Möller, Peter Fischer 0003, Fernando Gehm Moraes, Leandro Soares Indrusiak, Manfred Glesner |
FPL | 3 |
| 2009 | HeMPS - a Framework for NoC-based MPSoC GenerationabstractMulti-processor systems-on-chip (MPSoCs) are increasingly popular in embedded systems. Due to their complexity and huge design space to explore for such systems, CAD tools and frameworks to customize MPSoCs are mandatory. Some academic and industrial frameworks are available to support bus-based MPSoCs, but few works target NoCs as underlying communication architecture. A framework targeting MPSoC customization must provide abstract models to enable fast design space exploration, flexible application mapping strategies, all coupled to features to evaluate the performance of running applications. This paper proposes a framework to customize NoC-based MPSoCs with support to static and dynamic task mapping and C/SystemC simulation models for processors and memories. A simple, specifically designed microkernel executes in each processor, enabling multitasking at the processor level. Graphical tools enable debug and system verification, individualizing data for each task. Practical results highlight the benefit of using dynamic mapping strategies (total execution time reduction) and abstract models (total simulation time reduction without losing accuracy). Everton Carara, Roberto P. de Oliveira, Ney Laert Vilar Calazans, Fernando Gehm Moraes |
ISCAS | 4 |
| 2009 | Increasing NoC power estimation accuracy through a rate-based modelabstractThis research work presents and compares two NoC power estimation models, one based on the volume of information transmitted in the network, and another based on the transmission rates of each router. Guilherme Montez Guindani, Cezar Reinbrecht, Thiago R. da Rosa, Fernando Gehm Moraes |
NOCS | 4 |
| 2009 | Crosstalk Fault Tolerant NoC: Design and Evaluation
Alzemiro Henrique Lucas da Silva, Alexandre M. Amory, Fernando Gehm Moraes |
VLSI-SoC | 3 |
| 2007 | Run-time mapping and communication strategies for Homogeneous NoC-Based MPSoCsabstractMultiprocessor systems-on-chip are becoming increasingly popular in embedded systems for the high degree of performance and flexibility they permit. While most MPSoCs are today highly heterogeneous for better fitting the target applications, homogeneous systems may become in a near future a viable alternative bringing other benefits such as run-time load balancing, high performance and low power consumption. The work presented in this paper relies on a homogeneous NoC-based MPSoC framework we developed which allows us to conduct cycle-accurate evaluations of 2 different techniques: proactive and reactive communications. Gilles Sassatelli, Nicolas Saint-Jean, Pascal Benoit, Lionel Torres, Michel Robert, Cristiane R. Woszezenki, Ismael Grehs, Fernando Gehm Moraes |
FCCM | 8 |
| 2007 | SCAFFI: An intrachip FPGA asynchronous interface based on hard macrosabstractBuilding fully synchronous VLSI circuits is becoming less viable as circuit geometries evolve. However, before the adoption of purely asynchronous strategies in VLSI design, globally asynchronous, locally synchronous (GALS) design approaches should take over. The design of circuits using complex field programmable components like state of the art FPGAs follows this same trend. In GALS design, a critical step is the definition of asynchronous interfaces between synchronous regions. This paper proposes SCAFFI, a new asynchronous interface to interconnect modules inside FPGAs. The interface is based on clock stretching techniques to avoid metastability. Differently from other interfaces, it can use both logic levels for stretching and do not require the use of arbiters. Also, compactness of the implementation is enhanced by the use of dedicated FPGA hard macros. A GALS version implementation of an RSA cryptography core demonstrates the use of SCAFFI. Julian J. H. Pontes, Rafael Soares, Ewerson Carvalho, Fernando Gehm Moraes, Ney Laert Vilar Calazans |
ICCD | 4 |
| 2007 | A Cryptographic Coarse Grain Reconfigurable Architecture Robust Against DPAabstractThis work addresses the problem of information leakage of cryptographic devices, by using the reconfiguration technique allied to an RNS based arithmetic. The information leaked by circuits, like power consumption, electromagnetic emissions and time to compute may be used to find cryptographic secrets. The results issue of prototyping shows that our coarse grained reconfigurable architecture is robust against power analysis attacks. Daniel Mesquita, Benoît Badrignans, Lionel Torres, Gilles Sassatelli, Michel Robert, Fernando Gehm Moraes |
IPDPS | 6 |
| 2007 | Evaluation of Algorithms for Low Energy Mapping onto NoCsabstractSystems on chip (SoCs) congregate multiple modules and advanced interconnection schemes, such as networks on chip (NoCs). One relevant problem in SoC design is module mapping onto a NoC targeting low energy. To date, few works are available on design and evaluation of mapping algorithms. The main goal of this work is to propose some algorithms and evaluate its results and performance with regard to low energy NoC mappings. These include exhaustive and stochastic search methods and heuristic approaches, and some combinations. The use of combined approaches compared to pure stochastic algorithms provides average reductions above 98% in execution time, while keeping energy saving within at most 5% of the best results. In addition, one heuristic provided average reductions in execution time above 90% when compared to pure stochastic algorithms, and obtained better energy saving than combined approaches. César A. M. Marcon, Edson I. Moreno, Ney Laert Vilar Calazans, Fernando Gehm Moraes |
ISCAS | 4 |
| 2007 | DfT for the Reuse of Networks-on-Chip as Test Access MechanismabstractThis paper presents new DfT modules required to use networks-on-chip as test access mechanism. The paper demonstrates that the proposed DfT modules can be also implemented on top of low cost networks-on-chip, i.e. networks without complex services. The DfT modules, which consist of test wrappers and test pin interfaces, are designed such that both the tester and CUTs transport test data unaware of the network. The DfT modules was analysed in terms of silicon area and test time, considering different network and test configurations. Alexandre M. Amory, Frederico Ferlini, Marcelo Lubaszewski, Fernando Gehm Moraes |
VTS | 4 |
| 2006 | Wrapper Design for the Reuse of Networks-on-Chip as Test Access MechanismabstractThis paper proposes a wrapper design for interconnects with guaranteed bandwidth and latency services and on-chip protocol. We demonstrate that these interconnects abstract the interconnect details and provide predictability in the data transfer, which are desirable not only for the functional domain but also for the test application. The proposed wrapper is implemented in VHDL and integrated to the Æthereal NoC. The results show the impact of of bandwidth in the core test time. The wrapper area and core test time are compared with a wrapper design for dedicated TAM. Alexandre M. Amory, Kees Goossens, Erik Jan Marinissen, Marcelo Lubaszewski, Fernando Gehm Moraes |
ETS | 5 |
| 2006 | A Leak Resistant Architecture Against Side Channel AttacksabstractHardware implementations of cryptographic algorithms may leak some information that can be used to recover cryptographic keys. This work combines reconfigurable techniques with the recently proposed leak resistant arithmetic (LRA) to thwart some side channel attacks (SCA). The introduced architecture outcomes the performance of classical implementation of modular multiplication, for key size exceeding 2048 bits, with a reasonable extra area overhead. Nevertheless, this is not a drawback, but a cost, since the main issue of the proposed architecture is the improved robustness in terms of security. Daniel Mesquita, Benoît Badrignans, Lionel Torres, Gilles Sassatelli, Michel Robert, Jean-Claude Bajard, Fernando Gehm Moraes |
FPL | 7 |
| 2006 | Reconfigurable Systems Enabled by a Network-on-ChipabstractA modern SoC design comprises dozens of dedicated IP cores for specialized tasks and processors for general-purpose tasks. Flexibility is the key feature of processors, since it is easy to modify their tasks behavior at runtime. However, most current SoCs have no capability to modify the hardware behavior or structure after system fabrication. On the other hand, to cope with current SoC internal communication complexity, suggestions to employ networks-on-chip (NoCs) are becoming widespread. This paper proposes to extend the inherent software flexibility to hard IP cores in SoCs using NoCs as the main internal communication resource. This is achieved by making IP cores reconfigurable. The paper advances two main contributions: first, a straightforward design flow for SoCs with reconfigurable IP cores; second, the proposition of a NoC, named Artemis, supporting IP core reconfiguration Leandro Möller, Ismael Grehs, Ney Laert Vilar Calazans, Fernando Gehm Moraes |
FPL | 4 |
| 2005 | MAIA: a framework for networks on chip generation and verificationabstractThe increasing complexity of SoCs makes networks on chip (NoC) a promising substitute for busses and dedicated wires interconnection schemes. However, new tools need to be developed to integrate NoC interconnection architectures and IP cores into SoCs. Such tools have to fulfill three main requirements: (i) automated NoC generation; (ii) automated production of NoC-IP core interfaces; (iii) seamless analysis of NoC traffic parameters. The objective of this paper is to present the MAIA framework, which includes functions to address all these requirements. NoCs generated by the MAIA framework have been used to successfully prototype SoCs in FPGAs. Luciano Ost, Aline Vieira de Mello, José Carlos S. Palma, Fernando Gehm Moraes, Ney Laert Vilar Calazans |
ASP-DAC | 4 |
| 2005 | Test Time Reduction Reusing Multiple Processors in a Network-on-Chip Based ArchitectureabstractThe increasing complexity and the short life cycles of embedded systems are pushing the current system-on-chip designs towards a rapid increase in the number of programmable processing units, while decreasing the gate count for custom logic. Considering this trend, this work proposes a test planning method capable of reusing available processors as test sources and sinks, and the on-chip network as the test access mechanism. Experimental results are based on ITC'02 benchmarks and on two open core processors, compliant with MIPS and SPARC instruction sets. The results show that the cooperative use of both the on-chip network and the embedded processors can increase the test parallelism and reduce the test time without additional cost in area and pins. Alexandre M. Amory, Marcelo Lubaszewski, Fernando Gehm Moraes, Edson I. Moreno |
DATE | 3 |
| 2005 | Exploring NoC Mapping Strategies: An Energy and Timing Aware TechniqueabstractComplex applications implemented as systems on chip (SoC) demand extensive use of system level modeling and validation. Their implementation gathers a large number of complex IP cores and advanced interconnection schemes, such as hierarchical bus architectures or networks on chip (NoC). Modeling applications involves capturing its computation and communication characteristics. Previously proposed communication weighted models (CWM) consider only the application communication aspects. This work proposes a communication dependence and computation model (CDCM) that can simultaneously consider both aspects of an application. It presents a solution to the problem of mapping applications on regular NoC while considering execution time and energy consumption. The use of CDCM is shown to provide estimated average reductions of 40% in execution time, and 20% in energy consumption, for current technologies. César A. M. Marcon, Ney Laert Vilar Calazans, Fernando Gehm Moraes, Altamiro Amadeu Susin, Igor M. Reis, Fabiano Hessel |
DATE | 3 |
| 2005 | A scalable test strategy for network-on-chip routersabstractNetwork-on-chip has recently emerged as alternative communication architecture for complex system chip and different aspects regarding NoC design have been studied in the literature. However, the test of the NoC itself for manufacturing faults has been marginally tackled. This paper proposes a scalable test strategy for the routers in a NoC, based on partial scan and on an IEEE 1500-compliant test wrapper. The proposed test strategy takes advantage of the regular design of the NoC to reduce both test area overhead and test time. Experimental results show that a good tradeoff of area overhead, fault coverage, test data volume, and test time is achieved by the proposed technique. Furthermore, the method can be applied for large NoC sizes and it does not depend on the network routing and control algorithms, which makes the method suitable to test a large class of network models Alexandre M. Amory, Eduardo Wenzel Brião, Érika F. Cota, Marcelo Lubaszewski, Fernando Gehm Moraes |
ITC | 5 |
| 2005 | Modeling the Traffic Effect for the Application Cores Mapping Problem onto NoCs
César A. M. Marcon, José Carlos S. Palma, Ney Laert Vilar Calazans, Fernando Gehm Moraes, Altamiro Amadeu Susin, Ricardo Augusto da Luz Reis |
VLSI-SoC | 4 |
| 2005 | Current Mask Generation: an Analog Circuit to Thwart DPA Attacks
Daniel Mesquita, Jean-Denis Techer, Lionel Torres, Michel Robert, Guy Cathébras, Gilles Sassatelli, Fernando Gehm Moraes |
VLSI-SoC | 7 |
| 2004 | MultiNoC: A Multiprocessing System Enabled by a Network on ChipabstractThe MultiNoC system implements a programmable onchip multiprocessing platform built on top of an efficient, low area overhead intra-chip interconnection scheme. The employed interconnection structure is a network on chip, or NoC. NoC are emerging as a viable alternative to increasing demands on interconnection architectures, due to the following characteristics: (i) energy efficiency and reliability; (ii) scalability of bandwidth, when compared to traditional bus architectures; (iii) reusability; (iv) distributed routing decisions. An external host computer feeds MultiNoC with application instructions and data. After this initialization procedure, MultiNoC executes some algorithm. After finishing execution of the algorithm, output data can be read back by the host. Sequential or parallel algorithms conveniently adapted to the MultiNoC structure can be executed. The main motivation to propose this design is to enable the investigation of current trends to increase the number of embedded processors in SoC, leading to the concept of "sea of processors" systems. Aline Vieira de Mello, Leandro Möller, Ney Laert Vilar Calazans, Fernando Gehm Moraes |
DATE | 4 |
| 2004 | FiPRe: An Implementation Model to Enable Self-Reconfigurable Applications
Leandro Möller, Ney Laert Vilar Calazans, Fernando Gehm Moraes, Eduardo Wenzel Brião, Ewerson Carvalho, Daniel Camozzato |
FPL | 3 |
| 2004 | HERMES: an infrastructure for low area overhead packet-switching networks on chip
Fernando Gehm Moraes, Ney Laert Vilar Calazans, Aline Vieira de Mello, Leandro Möller, Luciano Ost |
Integr. | 1 |
| 2003 | Development of a Tool-Set for Remote and Partial Reconfiguration of FPGAs
Fernando Gehm Moraes, Daniel Mesquita, José Carlos S. Palma, Leandro Möller, Ney Laert Vilar Calazans |
DATE | 1 |
| 2003 | Design of a fingerprint system using a hardware/software environmentabstractProcessing system of fingerprint are CPU time intensive, being normally implemented in software. This paper present a new algorithm for fingerprint features localization, that can be easily implemented in hardware (system-on-a-chip, FPGA). This algorithm is composed by 3 stages, first stage read a fingerprint image (255x255pixels, ash tones) and apply a Gaussian Filter, after this, apply a absolute difference mask (ADM) for detector the edges in the image filtered and the last stage look for fingerprint features into the image. The information showed by 3th stage are the coordinate X and Y for each feature detected, asked minutiae. For localization the minutiae, the system pursue the edge detected by ADM, this edge represent ridge edge, and analyzing the information from each pixel pursued is possible to locate the minutiae. The average time for localization all minutiae into the fingerprint image, implemented in hardware (FLEX10KE Family, Altera), was 306 milliseconds. Beyond hardware implementation be fast, is possible create embedded systems. Vanderlei Bonato, Rolf Fredi Molz, João Carlos Furtado, Marcos Flôres Ferrão, Fernando Gehm Moraes |
FPGA | 5 |
| 2003 | Propose of a Hardware Implementation for Fingerprint Systems
Vanderlei Bonato, Rolf Fredi Molz, João Carlos Furtado, Marcos Flôres Ferrão, Fernando Gehm Moraes |
FPL | 5 |
| 2003 | Software-Based Test for Non-Programmable Cores in Bus-Based System-on-Chip Architectures
Alexandre M. Amory, Leandro A. Oliveira, Fernando Gehm Moraes |
VLSI-SOC | 3 |
| 2003 | Are coarse grain reconfigurable architectures suitable for cryptography?
Daniel Mesquita, Lionel Torres, Fernando Gehm Moraes, Gilles Sassatelli, Michel Robert |
VLSI-SOC | 3 |
| 2003 | A Low Area Overhead Packet-switched Network on Chip: Architecture and Prototyping
Fernando Gehm Moraes, Aline Vieira de Mello, Leandro Möller, Luciano Ost, Ney Laert Vilar Calazans |
VLSI-SOC | 1 |