Michael G. Jordan

dblp:206/8272 · also Michael Guilherme Jordan · DBLP profile ↗
← Back
14ranked-venue papers
4as first author
11since 2021 · last 2026
0000-0002-5776-2626ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 3 first-author · 9 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Energy-aware DVFS-driven workload provisioning in heterogeneous cloud FaaS architectures
abstract
Abstract Cloud Warehouses are evolving with diverse computational resources, including CPUs, GPUs, and accelerators, catering to a multitude of tenant applications. While this heterogeneity promises improved performance and energy efficiency, harnessing its full potential poses challenges due to dynamic workload characteristics and variable application demands. To address this, scheduling approaches combined with optimization techniques like Dynamic Voltage and Frequency Scaling (DVFS) are crucial. However, integrating these approaches effectively can be complex, potentially leading to conflicts and diminished benefits. This research proposes two frameworks, EAPECloud and EAPECloud-DVFS, designed for energy-aware collaborative provisioning in heterogeneous CPU-GPU cloud nodes. The first approach reduces energy consumption by selecting and maintaining a static combination of the best scheduler and V-F pair for most workloads. The second approach goes further by dynamically adjusting the V-F pair of each device using DVFS techniques while selecting the optimal scheduler. While the static approach delivers strong results in most cases, the dynamic strategy achieves even greater energy savings, albeit with an additional convergence time to determine the optimal V-F pair. Although each framework has distinct advantages and use cases, our findings demonstrate that both approaches effectively reduce energy consumption in heterogeneous environments, with EAPECloud-DVFS achieving up to a 126.33% performance improvement compared to the Linux CPU Governor, highlighting its efficiency and applicability in real-time systems.
Lucas Rister Machado, Gregory de Moraes Rossato, Antonio Carlos Schneider Beck, Michael G. Jordan, Mateus B. Rutzig
J. Supercomput.4
2024 Exploiting Virtual Layers and Pruning for FPGA-Based Adaptive Traffic Classification
abstract
Traffic classification is crucial to many network administration tasks, from resource management to QoS monitoring. While DNNs are the state-of-the-art method for classifying traffic, they are computationally intensive, posing challenges to their adoption within the networking infrastructure. One popular alternative is exploiting the FPGAs in the Smart Network Interface Cards (SmartNICs) to speed up the DNN inference processing. In this context, there are two main obstacles. First, modern networks experience high-volume and very volatile traffic flow, making the design of in-network accelerators difficult. Second, the use of FPGA-enabled SmartNICs involves reconfiguration when changing the classification task, which leads to significant time and energy overheads. In this work, we propose two different but complementary solutions to the aforementioned challenges: the use of pruning, which dynamically removes parts of the DNN to speed up its processing at the cost of controlled accuracy drops, therefore adapting the inference processing to the constantly changing traffic; and Hardware Virtual Layers (HWVL), which eliminate the need for FPGA reconfigurations for seamless and almost instantaneous task switching. Both approaches are combined in the Spyke Framework, improving throughput in up to 1.46× and reducing the energy per inference in up to 1.37× compared to a state-of-the-art FPGA accelerator.
Julio Costella Vicenzi, Guilherme Korol, Michael G. Jordan, Mateus B. Rutzig, Antonio Carlos Schneider Beck
DSD3
2023 Adaptive Inference on Reconfigurable SmartNICs for Traffic Classification
Julio Costella Vicenzi, Guilherme Korol, Michael G. Jordan, Mateus B. Rutzig, Antonio Carlos Schneider Beck
AINA (2)3
2023 Pruning and Early-Exit Co-Optimization for CNN Acceleration on FPGAs
abstract
The challenge of processing heavy-load ML tasks, particularly CNN-based ones at resource-constrained IoT devices, has encouraged the use of edge servers. The edge offers performance levels higher than the end devices and better latency and security levels than the Cloud. On top of that, the rising complexity of ML applications, the ever-increasing number of connected devices, and the current demands for energy efficiency require optimizing such CNN models. Pruning and early-exit are notable optimizations that have been successfully used to alleviate the computational cost of inference. However, these optimizations have not yet been exploited simultaneously: while pruning is usually applied at design time, which involves retraining the CNN before deployment, early-exit is inherently dynamic. In this work, we propose AdaPEx, a framework that exploits the intrinsic reconfigurable FPGA capabilities so both can be cooperatively employed. AdaPEx first explores the trade-off between pruning and early-exit at design-time, creating a design space never exploited in the state-of-the-art. Then, AdaPEx applies FPGA reconfiguration as a means to enable the combined use of pruning and early-exit dynamically. At runtime, this allows matching the inference processing to the current edge conditions and a user-configurable accuracy threshold. In a smart IoT application, AdaPEx processes up to 1.32× more inferences and improves EDP by up to 2.55× over the state-of-the-art FPGA-based FINN accelerator.
Guilherme Korol, Michael G. Jordan, Mateus B. Rutzig, Jerónimo Castrillón, Antonio Carlos Schneider Beck
DATE2
2023 MVSym: Efficient symbiotic exploitation of HLS-kernel multi-versioning for collaborative CPU-FPGA cloud systems
Michael G. Jordan, Bernardo Neuhaus Lignati, Guilherme Korol, Mateus B. Rutzig, Antonio Carlos Schneider Beck
Integr.1
2023 Energy-aware fully-adaptive resource provisioning in collaborative CPU-FPGA cloud environments
Michael G. Jordan, Guilherme Korol, Tiago Knorst, Mateus B. Rutzig, Antonio Carlos Schneider Beck
J. Parallel Distributed Comput.1
2022 AdaFlow: A Framework for Adaptive Dataflow CNN Acceleration on FPGAs
abstract
To meet latency and privacy requirements, resource-hungry deep learning applications have been migrating to the Edge, where IoT devices can offload the inference processing to local Edge servers. Since FPGAs have successfully accelerated an increasing number of deep learning applications (especially CNN-based ones), they emerge as an effective alternative for Edge platforms. However, Edge applications may present highly unpredictable workloads, requiring runtime adaptability in the inference processing. Although some works apply model switching on CPU and GPU platforms by exploiting different pruning rates at runtime, so the inference can adapt according to some quality-performance trade-off, FPGA-based accelerators refrain from this approach since they are synthesized to specific CNN models. In this context, this work enables model switching on FPGAs by adding to the well-known FINN accelerator an extra level of adaptability (i.e., flexibility) and support to the dynamic use of pruning via fast model switch on flexible accelerators, at the cost of some extra logic, or via FPGA reconfigurations of fixed accelerators. From that, we developed AdaFlow: a framework that automatically builds, at design time, a library from these new available versions (flexible and fixed, pruned or not) that will be used, at runtime, to dynamically select a given version according to a user-configurable accuracy threshold and current workload conditions. We have evaluated AdaFlow under a smart Edge surveillance application with two CNN models and two datasets, showing that AdaFlow processes, on average, 1.3× more inferences and increases, on average, 1.4× the power efficiency over state-of-the-art statically deployed dataflow accelerators.
Guilherme Korol, Michael G. Jordan, Mateus B. Rutzig, Antonio Carlos Schneider Beck
DATE2
2022 ConfAx: Exploiting Approximate Computing for Configurable FPGA CNN Acceleration at the Edge
abstract
The number of CNN-based applications executing at the Edge has been considerably increasing. Considering that CNNs are recognized error-resilient and the varied Edge conditions, we exploit hardware-level Approximate Computing to optimize FPGA-based CNN accelerators without any model retraining or other modifications. Given that, we propose ConfAx, a fully configurable multi-target Framework that navigates the accuracy-performance-resource trade-off to deploy different versions of CNN FPGA approximate accelerators. With an Edge case study (video surveillance), we show that ConfAx reduces power (up to $1.65\times)$ and energy (up to $ 1.44\times$) over a state-of-the-art accelerator at minor accuracy penalties (0.88% on average).
Guilherme Korol, Michael G. Jordan, Mateus B. Rutzig, Antonio Carlos Schneider Beck
ISCAS2
2021 Exploiting HLS-Generated Multi-Version Kernels to Improve CPU-FPGA Cloud Systems
abstract
Cloud Warehouses have been exploiting CPU-FPGA collaborative execution environments, where multiple clients share the same infrastructure to achieve to maximize resource utilization with the highest possible energy efficiency and scalability. However, the resource provisioning is challenging in these environments, since kernels may be dispatched to both CPU and FPGA concurrently in a highly variant scenario, in terms of available resources and workload characteristics. In this work, we propose MultiVers, a framework that leverages automatic HLS generation to enable further gains in such CPU-FPGA collaborative systems. MultiVers exploits the automatic generation from HLS to build libraries containing multiple versions of each incoming kernel request, greatly enlarging the available design space exploration passive of optimization by the allocation strategies in the cloud provider. Multivers makes both kernel multiversioning and allocation strategy to work symbiotically, allowing fine-tuning in terms of resource usage, performance, energy, or any combination of these parameters. We show the efficiency of MultiVers by using real-world cloud request scenarios with a diversity of benchmarks, achieving average improvements on makespan and energy of up to 4.62x and 19.04x, respectively, over traditional allocation strategies executing non-optimized kernels.
Bernardo Neuhaus Lignati, Michael G. Jordan, Guilherme Korol, Mateus B. Rutzig, Antonio Carlos Schneider Beck
ASP-DAC2
2021 FAIR: Fully-Adaptive Framework for Improving Resource Provisioning in Collaborative CPU-FPGA Cloud Environments
abstract
Cloud Warehouses have been exploiting CPU-FPGA collaborative environments to accelerate multi-tenant applications to achieve scalability and maximize resource utilization. However, resource provisioning is challenging in these environments since kernels may be dispatched to CPU and FPGA concurrently in a scenario with highly variant workloads and demands. The provisioning complexity is further aggravated due to diverse CPU and FPGA architectures being used at Cloud Warehouses (e.g., different FPGA/CPU devices between nodes). That means that the resource manager needs to consider the workload to be allocated and the characteristics of the Cloud infrastructure, which can be non-uniform. This paper shows that efficient resource provisioning in CPU-FPGA cloud environments requires different strategies depending on the demand, architecture, and workload. To provide the best use of resources in this complex environment, we propose FAIR, a Fully-Adaptive approach for Improving Resource provisioning in Collaborative CPU-FPGA Cloud. FAIR is end user-transparent and, in contrast to existing approaches, exploits the benefits of multiple provisioning strategies by dynamically selecting the most appropriate depending on the warehouse needs, workload properties, and target architecture. Over a varied set of scenarios, FAIR significantly improves the performance and energy efficiency of the environment compared to the use of fixed single strategies. On average, FAIR provides 32% performance improvements over the use of the best fixed single strategy. Compared to an Oracle that always selects the best energy strategies, FAIR achieves only 3% energy degradation.
Michael G. Jordan, Guilherme Korol, Mateus B. Rutzig, Antonio Carlos Schneider Beck
SBAC-PAD1
2021 Synergistically Exploiting CNN Pruning and HLS Versioning for Adaptive Inference on Multi-FPGAs at the Edge
abstract
FPGAs, because of their energy efficiency, reconfigurability, and easily tunable HLS designs, have been used to accelerate an increasing number of machine learning, especially CNN-based, applications. As a representative example, IoT Edge applications, which require low latency processing of resource-hungry CNNs, offload the inferences from resource-limited IoT end nodes to Edge servers featuring FPGAs. However, the ever-increasing number of end nodes pressures these FPGA-based servers with new performance and adaptability challenges. While some works have exploited CNN optimizations to alleviate inferences’ computation and memory burdens, others have exploited HLS to tune accelerators for statically defined optimization goals. However, these works have not tackled both CNN and HLS optimizations altogether; neither have they provided any adaptability at runtime, where the workload’s characteristics are unpredictable. In this context, we propose a hybrid two-step approach that, first, creates new optimization opportunities at design-time through the automatic training of CNN model variants (obtained via pruning) and the automatic generation of versions of convolutional accelerators (obtained during HLS synthesis); and, second, synergistically exploits these created CNN and HLS optimization opportunities to deliver a fully dynamic Multi-FPGA system that adapts its resources in a fully automatic or user-configurable manner. We implement this two-step approach as the AdaServ Framework and show, through a smart video surveillance Edge application as a case study, that it adapts to the always-changing Edge conditions: AdaServ processes at least 3.37× more inferences (using the automatic approach) and is at least 6.68× more energy-efficient (user-configurable approach) than original convolutional accelerators and CNN Models (VGG-16 and AlexNet). We also show that AdaServ achieves better results than solutions dynamically changing only the CNN model or HLS version, highlighting the importance of exploring both; and that it is always better than the best statically chosen CNN model and HLS version, showing the need for dynamic adaptability.
Guilherme Korol, Michael G. Jordan, Mateus B. Rutzig, Antonio Carlos Schneider Beck
ACM Trans. Embed. Comput. Syst.2
2020 MCEA: A Resource-Aware Multicore CGRA Architecture for the Edge
abstract
Modern IoT edge devices must address the unpredictability of applications with strict power and temperature constraints. In this scenario, heterogeneous multicore architectures have been driving many solutions due to their high energy efficiency and ability to exploit Task-Level Parallelism. However, while their performance is highly dependent on the quality of the scheduling, their adaptability and generality get restricted when they use fixed-size hardware accelerators. Considering that, this work proposes MCEA, a transparent and power-adaptive multicore reconfigurable architecture. MCEA dynamically adapts the hardware to the workload rather than migrating applications; and predicatively sizes its reconfigurable accelerators without prior knowledge of the applications' behaviors. For that, MCEA uses a synergistic and online profiling system with power gating, achieving performance levels near of homogeneous architectures with fixed and oversized reconfigurable fabric (within 99% on average) while presenting energy efficiency levels similar to heterogeneous architectures statically tuned to a specific workload (within 99% on average). Therefore, MCEA improves Energy-Delay Product in 1.55x and 1.21x when compared to their heterogeneous and homogeneous counterparts, and in 4.72x when compared to a multicore with OoO processors only. We also show that MCEA outperforms a state-of-the-art reconfigurable architecture for the edge under the same power envelope.
Guilherme Korol, Michael G. Jordan, Marcelo Brandalero, Michael Hübner 0001, Mateus B. Rutzig, Antonio Carlos Schneider Beck
FPL2
2019 Boosting SIMD Benefits through a Run-time and Energy Efficient DLP Detection
abstract
Data Level Parallelism has been improving performance-energy tradeoff of current processors by coupling SIMD engines, such as Intel AVX and ARM NEON. Special libraries and compilers are used to support DLP execution on such engines. However, timing overhead on hand coding is inevitable since most software developers are not skilled to extract DLP using unfamiliar libraries. In addition, DLP detection through compiler, besides breaking software compatibility, is limited to static code analysis, which compromises performance gains. In this work, we propose a runtime DLP detection named as Dynamic SIMD Assembler, which transparently identifies vectorizable code regions to execute in the ARM NEON engine. Due to its dynamic fashion, DSA keeps software compatibility and avoids timing overhead on software developing process. Results have shown that DSA outperforms ARM NEON auto-vectorization compiler by 32% since it covers wider vectorized regions, such as Dynamic Range, Sentinel and Conditional Loops. In addition, DSA outperforms hand-vectorized code using ARM library by 26% reducing 45% of energy consumption with no penalties over software development time.
Michael G. Jordan, Tiago Knorst, Julio Costella Vicenzi, Mateus B. Rutzig
DATE1
2017 A framework to automatically generate heterogeneous organization reconfigurable multiprocessing
abstract
Heterogeneous MPSoCs are vastly used in current embedded systems but they are highly dependent on special compilers. Dynamic reconfigurable systems are an alternative to overcome such drawback due to their adaptability. However, when such architectures are considered, one must concern about which hardware blocks should be heterogeneous and their degree of heterogeneity. In this work, we propose a framework that automatically generates heterogeneous reconfigurable multiprocessors that exploit the ideal ILP of parallel applications to improve performance/watt. Our generated system achieves, on average, 32% of performance improvements with 33% of energy savings over its manual generated counterpart, with equivalent chip area.
Josimar Sfreddo, Rafael Fao de Moura, Michael G. Jordan, Jeckson Dellagostin Souza, Antonio Carlos Schneider Beck, Mateus B. Rutzig
ISCAS3