EDBT 2026 Demo / reviewers in the wild / expert
Michael Frank 0008
dblp:42/4902-8
· DBLP profile ↗
8ranked-venue papers
0as first author
4since 2021 · last 2025
0009-0004-5963-2165ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 3 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Boosting Task Scheduling Data Locality with Low-latency, HW-accelerated Label PropagationabstractTask Scheduling is a popular technique for exploiting parallelism in modern computing systems.In particular, HW-accelerated Task Scheduling has been shown to be effective at improving the performance of fine-grained workloads by dynamically assigning tasks to cores based on their data dependencies with minimal overhead, allowing the handling of tasks with execution times in the order of thousands of cycles.However, the performance of applications assisted by accelerated Task Scheduling is limited by the fact that once a task has all its dependencies fulfilled, it is typically executed on the first available core, which might not be locality-optimal.We thus propose a novel approach to Task Scheduling that leverages HW-accelerated Label Propagation (LP), a graph clustering algorithm, to group tasks with intersecting data patterns such that they are executed on the same core.We show that our approach can significantly improve the performance of task-based applications, improving overall program execution times by up to 1.50× while simultaneously reducing average task sizes by up to 1.81×, augmenting both synthetic benchmarks and real-world applications running on a 24-core RISC-V processor mapped to the Alveo U55C FPGA.These gains rely heavily on the low-latency nature of our proposed label propagation accelerator, which will typically cluster dynamic task graphs in under 300 cycles, up to 581× faster than an equivalent software implementation.Furthermore, by ensuring that ideal placement predictions are used as a hint rather than a hard constraint, we allow the system to benefit from improved data locality for memory-intensive applications while also maintaining high core utilization in compute-bound scenarios.Our results hence demonstrate the potential of HW-accelerated label propagation to improve the performance of Task Scheduling systems with low-latency, dynamic data locality optimization. Lucas Morais, Juan Miguel De Haro Ruiz, Alfredo Goldman, Guido Araujo, Giacomo Pedretti, Jim Ignowski, Michael Frank 0008, Xavier Martorell, Daniel Jiménez-González, Carlos Álvarez 0001 |
MICRO | 7 |
| 2024 | Enabling HW-Based Task Scheduling in Large Multicore ArchitecturesabstractDynamic Task Scheduling is an enticing programming model aiming to ease the development of parallel programs with intrinsically irregular or data-dependent parallelism. The performance of such solutions relies on the ability of the Task Scheduling HW/SW stack to efficiently evaluate dependencies at runtime and schedule work to available cores. Traditional SW-only systems implicate scheduling overheads of around 30K processor cycles per task, which severely limit the (core count,task granularity) combinations that they might adequately handle. Previous work on HW-accelerated Task Scheduling has shown that such systems might support high performance scheduling on processors with up to eight cores, but questions remained regarding the viability of such solutions to support the greater number of cores now frequently found in high-end SMP systems.The present work presents an FPGA-proven, tightly-integrated, Linux-capable, 30-core RISC-V system with hardware accelerated Task Scheduling. We use this implementation to show that HW Task Scheduling can still offer competitive performance at such high core count, and describe how this organization includes hardware and software optimizations that make it even more scalable than previous solutions. Finally, we outline ways in which this architecture could be augmented to overcome inter-core communication bottlenecks, mitigating the cache-degradation effects usually involved in the parallelization of highly optimized serial code. Lucas Morais, Carlos Álvarez 0001, Daniel Jiménez-González, Juan Miguel De Haro Ruiz, Guido Araujo, Michael Frank 0008, Alfredo Goldman, Xavier Martorell |
IEEE Trans. Computers | 6 |
| 2023 | Tensor slicing and optimization for multicore NPUs
Rafael C. F. Sousa, Márcio Machado Pereira, Yongin Kwon, Namsoon Jung, Michael Frank 0008, Guido Araujo |
J. Parallel Distributed Comput. | 7 |
| 2021 | Efficient Tensor Slicing for Multicore NPUs using Memory Burst ModelingabstractAlthough code generation for Convolution Neural Network (CNN) models has been extensively studied, performing efficient data slicing and parallelization for highly-constrained Multicore Neural Processor Units (NPUs) is still a challenging problem. Given the size of convolutions' in-put/output tensors and the small footprint of NPU on-chip memories, minimizing memory transactions while maximizing parallelism and MAC utilization are central to any effective solution. This paper proposes a TensorFlow XLA/LLVM compiler optimization pass for Multicore NPUs, called Tensor Slicing Optimization (TSO), which: (a) maximizes convolution parallelism and memory usage across NPU cores; and (b) reduces data transfers between host and NPU on-chip memories by using DRAM memory burst time estimates to guide tensor slicing. To evaluate the proposed approach, a set of experiments was performed using the NeuroMorphic Processor (NMP), a multicore NPU containing 32 RISC-V cores extended with novel CNN instructions. Experimental results show that TSO is capable of identifying the best tensor slicing that minimizes execution time for a set of CNN models. Speed-ups of up to 21.7% result when comparing the TSO burst-based technique to a no-burst data slicing approach. Rafael C. F. Sousa, Byungmin Jung, Jaehwa Kwak, Michael Frank 0008, Guido Araujo |
SBAC-PAD | 4 |
| 2019 | Adding Tightly-Integrated Task Scheduling Acceleration to a RISC-V Multi-core ProcessorabstractTask Parallelism is a parallel programming model that provides code annotation constructs to outline tasks and describe how their pointer parameters are accessed so that they might be executed in parallel, and asynchronously, by a runtime capable of inferring and honoring their data dependence relationships. It is supported by several parallelization frameworks, as OpenMP and StarSs. Lucas Morais, Vitor Silva, Alfredo Goldman, Carlos Álvarez 0001, Jaume Bosch, Michael Frank 0008, Guido Araujo |
MICRO | 6 |
| 2019 | JetsonLEAP: A framework to measure power on a heterogeneous system-on-a-chip device
Tarsila Bessa, Christopher J. Gull, Pedro Quintão, Michael Frank 0008, José A. M. Nacif, Fernando Magno Quintão Pereira |
Sci. Comput. Program. | 4 |
| 2016 | ParallelME: A Parallel Mobile Engine to Explore Heterogeneity in Mobile Computing Architectures
Guilherme Andrade, Wilson de Carvalho, Renato Utsch, Pedro Caldeira, Alberto Albuquerque, Fabricio Ferracioli, Leonardo Rocha 0001, Michael Frank 0008, Dorgival O. Guedes, Renato Ferreira 0001 |
Euro-Par | 8 |
| 2016 | Task parallel programming model + hardware acceleration = performance advantageabstractPresents a collection of slides covering the following topics: task parallel programming model; hardware acceleration; SoC; multi-core heterogeneous computing architecture; software architecture; and task graph accelerator. Tamer Dallou, Divino Cesar S. Lucas, Guido Araujo, Lucas Morais, Eduardo Ferreira Barbosa, Michael Frank 0008, Richard Bagley, Raj Sayana |
Hot Chips Symposium | 6 |