Rodrigo Cataldo

dblp:141/0659 · also Rodrigo Cadore Cataldo · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
3since 2021 · last 2023
0000-0003-4664-2909ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 2 first-author · 2 since 2021Computer networks · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Parallel and multicore computing · 38% Interconnection networks and networks-on-chip · 35% Processor architecture and microarchitecture · 27%

Topics — the 8 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Parallel and multicore computing
synchronization
0.822021
Subutai: Speeding Up Legacy Parallel Applications Through Data Synchronization · IEEE Trans. Parallel Distributed Syst. 2021
Subutai: distributed synchronization primitives in NoC interfaces for legacy parallel-applications · DAC 2018
Processor architecture and microarchitecture
multicore design
0.512021
Subutai: Speeding Up Legacy Parallel Applications Through Data Synchronization · IEEE Trans. Parallel Distributed Syst. 2021
Interconnection networks and networks-on-chip › routing algorithms
fault-tolerant routing
0.412020
Using Smart Routing for Secure and Dependable NoC-Based MPSoCs · IEEE/ACM Trans. Netw. 2020
Interconnection networks and networks-on-chip
routing algorithms
0.412020
Using Smart Routing for Secure and Dependable NoC-Based MPSoCs · IEEE/ACM Trans. Netw. 2020
Interconnection networks and networks-on-chip
network interface
0.312018
Subutai: distributed synchronization primitives in NoC interfaces for legacy parallel-applications · DAC 2018
Processor architecture and microarchitecture › multiprocessor architecture
synchronization hardware
0.312018
Subutai: distributed synchronization primitives in NoC interfaces for legacy parallel-applications · DAC 2018
Parallel and multicore computing › synchronization
synchronization mechanisms
0.312018
Subutai: distributed synchronization primitives in NoC interfaces for legacy parallel-applications · DAC 2018
Processor architecture and microarchitecture
chip multiprocessor
0.112018
Subutai: distributed synchronization primitives in NoC interfaces for legacy parallel-applications · DAC 2018

Methods — techniques the papers use, named apart from their topics

hardware-software co-design · 0.8POSIX threads · 0.5simulation · 0.4fault model · 0.4
YearPublicationVenuePosition
2023 The Last-Level-Cache Interference in Guest Performance: a Case-Study with Zephyr OS
abstract
Embedded systems are increasingly relying on virtual machines (VMs) to ensure portability and composability of services. To achieve high-performance and resources' isolation, the VM can be mapped to a dedicated core. In shared-memory multi/many cores architectures, the last level cache (LLC) is shared among cores. The state-of-the-art shows that interference on LLC can be a bottleneck for VM's performance. Such interference can depend on several factors, including the design of the applications, guest and host OSes, the hypervisor, and the architecture. Therefore, it is expected that studies analyze in-depth each of these factors. The goal of this work is focusing on the interference that cannot be mitigated by cache isolation techniques supported by the hypervisor, for instance, those caused by host OS applications. Such interference is unpredictable and can jeopardize the performance of the guest. The contribution is a new perspective about how cache interference affects the latency of the guest application and guest OS, crossing different performance metrics in comprehensive plots. We run our experiments on Arm Cortex-A53 processor, and, as guest OS, we employ the state-of-the-art Zephyr OS. Thus, a side contribution is to show Zephyr performance facing cache interference. Our results show that interference on LLC can affect in up to +8 × to slow down the guest application and up to +2.8 ×the subset of kernel functions which are involved for the application execution. Results also show that the guest is mostly affected when its data can fit on the LLC: after this point, the LLC saturation occurs and the host interference becomes insignificant from a guest perspective.
Marcelo Ruaro, Hadrien Barral, Matteo Bertolino, Rodrigo Cataldo, Roberto Medina 0001, Mohamed Karaoui, Etienne Borde
DSD4
2021 Using curved angular intra-frame prediction to improve video coding efficiency
Ramon Fernandes, Gustavo Sanchez, Rodrigo Cataldo, Luciano Volcan Agostini, César A. M. Marcon
J. Vis. Commun. Image Represent.3
2021 Subutai: Speeding Up Legacy Parallel Applications Through Data Synchronization
abstract
The decrease of the performance gain dictated by Moore's Law boosted the development of manycore architectures to replace single-core architectures. These new architectures must employ parallel applications and distribute its workload over a multitude of cores to reach the desired performance. Parallel applications are harder to develop than sequential ones since the developer must guarantee data integrity using synchronization primitives. While multiple novel solutions have been proposed to speed up parallel applications through handling one type of data synchronization primitive, exceptionally few works support multiple types of synchronization primitives and legacy code. This article proposes Subutai, a hardware/software co-design solution for accelerating multiple synchronization primitives without modifying the application source code. By providing a new user library, while retaining an existing synchronization API, legacy and novel applications can benefit from our solution. Our experimental evaluation, which provides a POSIX Threads implementation, demonstrates Subutai speeds up to 2.71× and 4.61× the execution of single- and multiple-application executions, respectively.
Rodrigo Cataldo, Ramon Fernandes, Kevin J. M. Martin, Jarbas Silveira, Gustavo Sanchez, Martha Johanna Sepúlveda, César A. M. Marcon, Jean-Philippe Diguet
IEEE Trans. Parallel Distributed Syst.1
2020 Broadcast Mechanism Based on Hybrid Wireless/Wired NoC for Efficient Barrier Synchronization in Parallel Computing
abstract
Parallel computing is essential to achieve the manycore architecture performance potential, since it utilizes the parallel nature provided by the hardware for its computing. These applications will inevitably have to synchronize its parallel execution: for instance, broadcast operations for barrier synchronization. Conventional network-on-chip architectures for broadcast operations limit the performance as the synchronization is affected significantly due to the critical path communications that increase the network latency and degrade the performance drastically. A Wireless network-on-chip offers a promising solution to reduce the critical path communication bottlenecks of such conventional architectures by providing hardware broadcast support. We propose efficient barrier synchronization support using hybrid wireless/wired NoC to reduce the cost of broadcast operations. The proposed architecture reduces the barrier synchronization cost up to 42.79% regarding network latency and saves up to 42.65% communication energy consumption for a subset of applications from the PARSEC benchmark.
Hemanta Kumar Mondal, Navonil Chatterjee, Rodrigo Cataldo, Jean-Philippe Diguet
ASP-DAC3
2020 Using Smart Routing for Secure and Dependable NoC-Based MPSoCs
abstract
The Internet-of-Things (IoT) boosted the building of computational systems that share computation, communication and storage resources for uncountable types of applications. MultiProcessor System-on-Chip (MPSoC) is a fundamental component of such systems offering large parallelism degree in an ocean of processors and memories connected through one or more Network-on-Chips (NoCs). Therefore, a massive quantity of sensitive information of several applications can share computation and communication resources of the MPSoCs demanding security mechanisms and policies. Besides, the advances of CMOS technologies increases the quantity of static and dynamic faults, requiring a dependable and resilient target architecture, which can be partially fulfilled by an effective and efficient NoC design. This work addresses fault tolerance and security at NoC level with SDR, a routing algorithm that includes the concept of security zones in the MPSoC while providing support for dependable routing avoiding faulty links. The proposed routing algorithm prioritizes communication paths deemed secure in 2D mesh NoCs with deadlock freedom. Experimental results employing realistic workload scenarios based on the NASA Numeric Aerodynamic Simulation (NAS) Parallel Benchmark (NPB) and a fault model for 65nm and 22nm CMOS fabrication technologies demonstrates the scalability, security, and dependability of SDR.
Ramon Fernandes, César A. M. Marcon, Rodrigo Cataldo, Martha Johanna Sepúlveda
IEEE/ACM Trans. Netw.3
2019 CDMA-based multiple multicast communications on WiNOC for efficient parallel computing
abstract
In this work, we introduce an hybrid WiNoC, which judicially uses the wired and wireless interconnects for broadcasting/multicasting of packets. A code division multiple access (CDMA) method is used to support multiple broadcast operations originating from multiple applications executed on the multiprocessor platform. The CDMA-based WiNoC is compared in terms of network latency and power consumption with wired-broadcast/multicast NoC.
Navonil Chatterjee, Hemanta Kumar Mondal, Rodrigo Cataldo, Jean-Philippe Diguet
NOCS3
2018 Subutai: distributed synchronization primitives in NoC interfaces for legacy parallel-applications
abstract
Parallel applications are essential for efficiently using the computational power of a Multiprocessor System-on-Chip (MPSoC). Unfortunately, these applications do not scale effortlessly with the number of cores because of synchronization operations that take away valuable computational time and restrict the parallelization gains. Moreover, synchronization is also a bottleneck due to sequential access to shared memory. We address this issue and introduce "Subutai", a hardware/software (HW/SW) architecture designed to distribute essential synchronization mechanisms over the Network-on-Chip (NoC). It includes Network Interfaces (NIs), drivers and a custom library of a NoC-based MPSoC architecture that speeds up the essential synchronization primitives of any legacy parallel application. Besides, we provide a fast simulation tool for parallel applications and a HW architecture of the NI. Experimental results with PARSEC benchmark show an average application speedup of 2.05 compared to the same architecture running legacy SW solutions for 36% overhead of HW architecture.
Rodrigo Cataldo, Ramon Fernandes, Kevin J. M. Martin, Martha Johanna Sepúlveda, Altamiro Amadeu Susin, César A. M. Marcon, Jean-Philippe Diguet
DAC1
2016 Efficient traffic balancing for NoC routing latency minimization
abstract
Modern technologies of integrated circuits allow billions of transistors arranged into a single chip, enabling to implement complex systems, which need a scalable and parallel communication architecture. Network-on-Chip (NoC) is a natural candidate to fulfill such communication requirements, providing high performance when the communication demands are balanced. This work proposes a new static balancing method that uses the application's traffic pattern for NoC latency reduction. This method allows the generation of a deterministic routing algorithm with simplistic implementation and low latency. Experimental results compare four balancing methods, showing the improvement of the proposed static balancing concerning the average NoC latency.
Joao Marcelo Ferreira, Jarbas Silveira, Jardel Silveira, Rodrigo Cataldo, Thais Webber, Fernando Gehm Moraes, César A. M. Marcon
ISCAS4