EDBT 2026 Demo / reviewers in the wild / expert
José Rodrigo Azambuja
dblp:21/1296 · also José Rodrigo Furlanetto Azambuja
· DBLP profile ↗
13ranked-venue papers
3as first author
9since 2021 · last 2025
0000-0002-2627-5075ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 3 first-author · 4 since 2021Software engineering, systems software and programming languages · 3 · 2 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Distributed Graph Neural Networks in Programmable Data PlanesabstractThe ability to redefine the data plane behavior with programmable network devices provides a plethora of novel possibilities for in-network computing. One of these possibilities is embedding Artificial Intelligence (AI) and Machine Learning (ML) techniques directly in the data plane. Motivations include reducing decision latency and closing the control loop-i.e., performing measurements, learning, decisions, and actions directly in the data plane. However, running entire AI/ML algorithms in a single device might be infeasible due to memory and computing constraints. This work addresses the research challenges of running a Graph Neural Network (GNN) in a set of devices of a programmable data plane. Our hypothesis is that by distributing the GNN processing across the devices, the GNN uses instantaneous snapshots of the global network state and can act more quickly. As a proof of concept, we trained and evaluated a distributed GNN to perform explicit congestion notifications based on Data Center Transmission Control Protocol (DCTCP). We verified the feasibility of GNN classification in the data plane through simulations and both software and hardware switch experiments with bmv2 and Intel Tofino. Ivan Peter Lamb, Pedro Arthur Pinheiro Rosa Duarte, Jonatas Adilson Marques, Marcelo Caggiani Luizelli, Luciano Paschoal Gaspary, Anderson Tavares, Ronaldo A. Ferreira, Ítalo S. Cunha, José Rodrigo Azambuja, Weverton Luis da Costa Cordeiro |
NOMS | 9 |
| 2024 | Multi-Tenant Programmable Switch Virtualization Leveraging Explicit Resource SharingabstractWith the migration of traditional computer networks to the Software-defined Networking paradigm, flexibility is a core feature that novel technologies must provide. In this context, virtualization is gaining traction in Programmable Data Planes (PDPs) as a means of achieving greater flexibility, with several solutions in the literature for instantiating virtual programmable switches on the same host device. Virtualization brings numerous advantages, enabling multi-tenancy in programmable data/research center networks and greater device resource utilization. Nevertheless, enabling a complete multitenant solution, in which the tenants have disjoint sets of virtual devices, requires management and security considerations not yet approached in previous investigations. Previous works focus mainly on the core underlying technology necessary to deploy multiple devices in the same physical host. This paper presents a PDP virtualization architecture based on program composition and access control for securely managing virtual switches from different tenants. Additionally, we define extensions to PDP programmability, allowing tenants to specify shared elements, such as tables, between their virtual devices. Our experiments highlight the ability to transparently manage multiple virtual switches hosted in the same physical device in networking scenarios with multiple tenants. Ivan Peter Lamb, Pedro Arthur Pinheiro Rosa Duarte, Marcelo Caggiani Luizelli, Luciano Paschoal Gaspary, José Rodrigo Azambuja, Weverton Luis sa Costa Cordeiro |
CNSM | 5 |
| 2024 | Spinner: Enabling In-network Flow Clustering Entirely in a Programmable Data PlaneabstractData plane programmability is redesigning the way we manage and operate forwarding devices. However, most of the algorithmic decisions performed by data planes are still deterministic and control-plane dependent. We argue that it is possible to break this dependency and make the data plane intelligent, so that it can learn the infrastructure state autonomously. Despite existing efforts to make data planes intelligent, little has been done to design unsupervised ML algorithms that fit the architectural constraints of programmable devices. Executing such approaches in the data plane has the potential to reduce the overall decision-making time, thus meeting packet processing deadlines (which are in the order of nanoseconds). In this paper, we propose Spinner, the first effort to deliver an unsupervised Machine Learning (ML) approach entirely in programmable devices. Spinner is a flow clustering algorithm designed to fit existing architectural constraints of SmartNICs, and that can reach line rate for most packet sizes with complexity O(k). To demonstrate the potential behind in-network clustering, we prototype and deploy Spinner in a programmable testbed and use it to enhance Explicit Congestion Notifications (ECN) at the server side. Spinner-enhanced TCP provides up to 2x higher throughput when comparing to de-facto TCP implementations. Luigi Cannarozzo, Thiago Bortoluzzi Morais, Paulo Silas Severo de Souza, Leonardo Gobatto, Ivan Peter Lamb, Pedro Arthur Pinheiro Rosa Duarte, José Rodrigo Azambuja, Arthur Francisco Lorenzon, Fábio D. Rossi, Weverton Luis da Costa Cordeiro, Marcelo Caggiani Luizelli |
NOMS | 7 |
| 2024 | Multi-Tenant Programmable Switch Virtualization ArchitectureabstractVirtualization is gaining traction in Programmable Data Planes (PDP), with several solutions in the literature for emulating or instantiating virtual programmable switches on the same host device. Virtualization has numerous advantages in this context, enabling multi-tenancy in programmable data center networks and greater device resource utilization. Nevertheless, enabling a complete multi-tenant solution requires management and security considerations not yet approached in previous investigations. This paper presents a PDP virtualization architecture based on program composition and access control for securely managing virtual switches from different tenants. Our experiments highlight the ability to transparently manage virtual switches hosted in the same physical device in networking scenarios with multiple tenants. Ivan Peter Lamb, Theo Facen, Pedro Arthur Pinheiro Rosa Duarte, José Rodrigo Azambuja, Weverton Luis da Costa Cordeiro |
NOMS | 4 |
| 2022 | G-GPU: A Fully-Automated Generator of GPU-like ASIC AcceleratorsabstractModern Systems on Chip (SoC), almost as a rule, require accelerators for achieving energy efficiency and high performance for specific tasks that are not necessarily well suited for execution in standard processing units. Considering the broad range of applications and necessity for specialization, the design of SoCs has thus become expressively more challenging. In this paper, we put forward the concept of G-GPU, a general-purpose GPU-like accelerator that is not application-specific but still gives benefits in energy efficiency and throughput. Furthermore, we have identified an existing gap for these accelerators in ASIC, for which no known automated generation platform/tool exists. Our solution, called GPUPlanner, is an open-source generator of accelerators, from RTL to GDSII, that addresses this gap. Our analysis results show that our automatically generated G-GPU designs are remarkably efficient when compared against the popular CPU architecture RISC- V, presenting speed-ups of up to 223 times in raw performance and up to 11 times when the metric is performance derated by area. These results are achieved by executing a design space exploration of the GPU-like accelerators, where the memory hierarchy is broken in a smart fashion and the logic is pipelined on demand. Finally, tapeout-ready layouts of the G-GPU in 65nm CMOS are presented. Tiago D. Perez, Marcio Gonçalves, Leonardo Gobatto, Marcelo Brandalero, José Rodrigo Azambuja, Samuel Nascimento Pagliarini |
DATE | 5 |
| 2022 | Improving Content-Aware Video Streaming in Congested Networks with In-Network ComputingabstractNetwork congestion and packet loss pose an ever-increasing challenge to video streaming. Despite the research efforts toward making video encoding schemes resilient to lossy network conditions, forwarding devices have not considered monitoring packet content to prioritize packets and minimize the impact of packet loss on video transmission. In this work, we advocate in favor of in-network computing employing a packet drop algorithm and an in-network hardware module to devise a solution for improving content-aware video streaming in congested network. Results show that our approach can reduce intra-predicted packet loss by over 80% at negligible resource usage and performance costs. Leonardo Gobatto, Mateus Saquetti, Cláudio Machado Diniz, Bruno Zatt, Weverton Luis da Costa Cordeiro, José Rodrigo Azambuja |
ISCAS | 6 |
| 2022 | Video Decoder Improvements with Near-Data Speculative Motion Compensation ProcessingabstractVideo decoder implementations are still evolving as they directly affect a large fraction of embedded systems nowadays. In this context, Versatile Video Coding (VVC) brings increased compression efficiency, which comes with extra over-head in terms of computational effort and energy consumption. At the same time, emerging Near-Data Processing (NDP) architectures promise drastic time and energy cuts for applications with data streaming behavior. In this paper, a speculative Motion Compensation (MC) is proposed to enable video decoders improvements through the exploitation of NDP. We adopted a large-vector SIMD-based NDP system (called VIMA) that provides high-performance operations over 2 K vectors. The proposed strategy leverages the correlation between the prediction modes and the motion data between spatially neighboring blocks within a frame to speculatively perform the MC for an entire region of 2Kx128 samples. MC interpolation kernels were implemented using VIMA and x86 AVX-256 SIMD libraries. Our NDP-based kernel implementation allows speedup of $1.9\times$ to $22\times$ compared to the x86 baseline solutions. Stepping forward, based on a coalescence estimation, our strategy can properly handle interpolation misses, achieving MC performance improvements from 7% to 64%. Garrenlus de Souza, José Rodrigo Azambuja, Bruno Zatt, Marco A. Z. Alves, Sergio Bampi, Felipe Sampaio |
ISCAS | 2 |
| 2022 | Protecting Virtual Programmable Switches from Cross-App Poisoning (CAP) AttacksabstractCross-App Poisoning (CAP) is an emerging class of network integrity attacks against Software Defined Networking (SDN). In a CAP attack, a malicious app poison shared data objects maintained by the controller, thus co-opting legitimate apps into carrying out bogus actions that the malicious app itself cannot perform due to insufficient privileges. Existing solutions, such as ProvSDN, demonstrated that Information Flow Control (IFC) can track and thus prevent such attacks. However, these solutions cannot prevent CAP attacks in networks where malicious apps can take advantage of programmable virtual switches to bypass IFC. In this paper, we propose Virtual Information Flow Control (vIFC), a solution for defending against CAP attacks that exploit virtual switches to obfuscate malicious information flow. vIFC has shown high effectivity while posing low performance overhead. We also propose a policy model that offers flexibility to the network manager to determine IFC between apps running on multiple controllers. Ivan Peter Lamb, Mateus Saquetti, Guilherme Bueno de Oliveira, José Rodrigo Azambuja, Weverton Luis da Costa Cordeiro |
NOMS | 4 |
| 2022 | Evaluating low-level software-based hardening techniques for configurable GPU architectures
Marcio Gonçalves, Josie E. Rodriguez Condia, Matteo Sonza Reorda, Luca Sterpone, José Rodrigo Azambuja |
J. Supercomput. | 5 |
| 2019 | A Knapsack Methodology for Hardware-based DMR Protection against Soft Errors in Superscalar Out-of-Order ProcessorsabstractHigh-performance superscalar processors have been adopted to satisfy the rising demand for processing applications of ever-growing complexity. This extra complexity, added to the increasing vulnerability of transistors due to technology scaling, poses a great challenge since these effects have also been proven to affect ground-level safety-critical applications. To increase microarchitectural resilience, designers may adopt Dual Modular Redundancy (DMR), which offers full fault detection. However, given that DMR incurs in high area and energy overheads, we propose a design-time methodology aiming to achieve the best tradeoff between resilience and area overhead, decreasing DMR costs and maintaining acceptable detection levels for such a complex design. This is done by adopting the Knapsack Problem (KSP) as a heuristic to identify the optimal micro-architectural structures that should be duplicated to achieve target resilience with the smallest possible area overhead. By injecting over 800k faults in 12 significant micro-architectural structures of different versions of the complex Berkeley Out-of-Order Machine (BOOM) superscalar processor modeled with RTL accuracy, we compare this optimal strategy against a greedy one, showing that 90% of vulnerability reduction may be achieved with 50.6% and 107.8% area overheads for the optimal and greedy strategies, respectively. Rafael Billig Tonetto, Douglas Maciel Cardoso, Marcelo Brandalero, Luciano Volcan Agostini, Gabriel L. Nazar, José Rodrigo Azambuja, Antonio Carlos Schneider Beck |
VLSI-SoC | 6 |
| 2013 | Algorithm transformation methods to reduce software-only fault tolerance techniques' overheadabstractThis paper introduces a framework that tackles the costs in area and energy consumed by methodologies like spatial or temporal redundancy with a different approach: given an algorithm, we find a transformation in which part of the computation involved is transformed into memory accesses. The precomputed data stored in memory can be protected then by applying traditional and well established ECC algorithms to provide fault tolerant hardware designs. At the same time, the transformation increases the performance of the system by reducing its execution time, which is then used by customized software-only fault tolerant techniques to protect the system without any degradation when compared to its original form. Application of this technique to key algorithms in a MP3 player, combined with a fault injection campaign, show that this approach increases fault tolerance up to 92%, without any performance degradation. José Rodrigo Azambuja, Gustavo Brown, Fernanda Lima Kastensmidt, Luigi Carro |
IOLTS | 1 |
| 2011 | Exploring the Limitations of Software-based Techniques in SEE Fault Coverage
José Rodrigo Azambuja, Samuel Nascimento Pagliarini, Lucas Rosa, Fernanda Lima Kastensmidt |
J. Electron. Test. | 1 |
| 2009 | Evaluating large grain TMR and selective partial reconfiguration for soft error mitigation in SRAM-based FPGAsabstractThis paper presents an innovative method that allows the use of dynamic partial reconfiguration combined with triple modular redundancy (TMR) in SRAM-based FPGAs fault-tolerant designs. The method combines large grain TMR with special voters capable of signalizing the faulty module and check point states that allow the sequential synchronization of the recovered module with the Xilinx TMR (XTMR) approach. As a result, only the faulty domain is reconfigured, minimizing time and energy spent in the process. In addition, the use of checkpoint states avoids system downtime, since the synchronization of the recovered module is performed while the others are kept running. Experimental results show that the method has a reduced fault recovery time compared to the standard TMR implementation, maintaining the compatible area overhead and performance. José Rodrigo Azambuja, Fernando Sousa, Lucas Rosa, Fernanda Lima Kastensmidt |
IOLTS | 1 |