EDBT 2026 Demo / reviewers in the wild / expert
Joshua Mack
dblp:175/4518
· DBLP profile ↗
12ranked-venue papers
5as first author
9since 2021 · last 2025
0000-0003-1066-5578ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 5 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Coarse-Grained Task Parallelization by Dynamic Profiling for Heterogeneous SoC-Based Embedded SystemabstractIn this study, we introduce a methodology for automatically transforming user applications written in C/C++ to a parallel representation consisting of coarse-grained tasks based on dynamic profiling. Such a parallel representation is suitable for mapping applications onto heterogeneous SoCs. We present our approach for instrumenting the user application binary during the compilation process with parallel primitives that enable the runtime system to schedule and execute independent computation-intensive coarse-grained tasks concurrently. We use the proposed compilation and code transformation methodology to retarget each application for execution on a heterogeneous SoC composed of processor cores and accelerators. We demonstrate the capabilities of our integrated compile time and runtime flow through task-level parallelization and functionally correct execution of real-world applications in the communication systems and radar processing domains. We demonstrate the functionality of our integrated system by executing six distinct applications with different degrees of parallelism on four different platforms: an eight-core general-purpose processor, a heterogeneous SoC simulator, and two heterogeneous SoCs utilizing the Xilinx Zynq UltraScale+ FPGA and the Nvidia Jetson AGX board. Our integrated approach offers a path forward for application developers to take full advantage of the target SoC without requiring users to become hardware or parallel programming experts. Liangliang Chang, Serhan Gener, Joshua Mack, Hasan Umut Suluhan, Ali Akoglu, Chaitali Chakrabarti |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2025 | RIMMS: Runtime Integrated Memory Management System for Heterogeneous ComputingabstractEfficient memory management in heterogeneous systems is increasingly challenging due to diverse compute architectures (e.g., CPU, GPU, and FPGA) and dynamic task mappings not known at compile time. Existing approaches often require programmers to manage data placement and transfers explicitly, or assume static mappings that limit portability and scalability. This article introduces RIMMS (Runtime Integrated Memory Management System), a lightweight, runtime-managed, hardware-agnostic memory abstraction layer that decouples application development from low-level memory operations. RIMMS transparently tracks data locations, manages consistency, and supports efficient memory allocation across heterogeneous compute elements without requiring platform-specific tuning or code modifications. We integrate RIMMS into a baseline runtime and evaluate with complete radar signal processing applications across CPU+GPU and CPU+FPGA platforms. RIMMS delivers up to 2.43× speedup on GPU-based and 1.82× on FPGA-based systems over the baseline. Compared to IRIS, a recent heterogeneous runtime system, RIMMS achieves up to 3.08X speedup and matches the performance of native CUDA implementations while significantly reducing programming complexity. Despite operating at a higher abstraction level, RIMMS incurs only 1–2 cycles of overhead per memory management call, making it a low-cost solution. These results demonstrate RIMMS’s ability to deliver high performance and enhanced programmer productivity in dynamic, real-world heterogeneous environments. Serhan Gener, Aditya Ukarande, Shilpa Mysore Srinivasa Murthy, Md Sahil Hassan, Joshua Mack, Chaitali Chakrabarti, Ümit Y. Ogras, Ali Akoglu |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2025 | Tutorial: A Novel Runtime Environment for Accelerator-Rich Heterogeneous ArchitecturesabstractAs the landscape of computing advances, system designers are increasingly exploring methodologies that leverage higher levels of heterogeneity to enhance performance within constrained size, weight, power, and cost parameters. CEDR (Compiler-integrated Extensible DSSoC Runtime) stands as an ecosystem facilitating productive and efficient application development and deployment across heterogeneous computing systems. It fosters the co-design of applications, scheduling heuristics, and accelerators within a unified framework. Our goal is to present CEDR as a promising environment for lifting the barriers to research on heterogeneous systems and addressing the broader challenges within domain-specific architectures. We introduce CEDR and discuss the evolutionary design decisions underlying its programming model. Subsequently, we explore its utility for a broad range of users through design sweeps on off-the-shelf heterogeneous platforms across scheduling heuristics, hardware compositions, and workload scenarios. Joshua Mack, Anish Krishnakumar, Ümit Y. Ogras, Ali Akoglu |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2024 | Cyclebite: Extracting Task Graphs From Unstructured Compute-ProgramsabstractExtracting portable performance in an application requires structuring that program into a data-flow graph of coarse-grained tasks (CGTs). Structuring applications that interconnect multiple external libraries and custom code (i.e., “Code From The Wild” (CFTW)) is challenging. When experts manually restructure a program, they trivialize the extraction of structure; however, this expertise is not broadly available. Automatic structuring approaches focus on the intersection of hot code and static loops, ignoring the data dependencies between tasks and significantly reducing the scope of analyzeable programs. This work addresses the problem of extracting the data-flow graph of CGTs from CFTW. To that end, we present Cyclebite. Our approach extracts CGTs from unstructured compute-programs by detecting CGT candidates in the simplified Markov Control Graph (MCG), and localizing CGTs in an epoch profile. Additionally, the epoch profile extracts the data dependence between CGTs required to build the data-flow graph of CGTs. Cyclebite demonstrates a robust selectivity for critical CGTs relative to the state-of-the-art (SoA), leading to a potential speedup of 12x on average and thread-scaling of 24x on average compared to modern compiler optimizers. We validate the results of Cyclebite and compare them to two SoA techniques using an input corpus of 25 open-source C/C++ libraries with 2,019 unique execution profiles. Benjamin R. Willis, Aviral Shrivastava, Joshua Mack, Shail Dave, Chaitali Chakrabarti, John S. Brunhaver |
IEEE Trans. Computers | 3 |
| 2023 | CEDR: A Compiler-integrated, Extensible DSSoC RuntimeabstractIn this work, we present a C ompiler-integrated, E xtensible D omain Specific System on Chip R untime (CEDR) ecosystem to facilitate research toward addressing the challenges of architecture, system software, and application development with distinct plug-and-play integration points in a unified compile time and runtime workflow. We demonstrate the utility of CEDR on the Xilinx Zynq MPSoC-ZCU102 for evaluating performance of pre-silicon hardware in the trade space of SoC configuration, scheduling policy and workload complexity based on dynamically arriving workload scenarios composed of real-life signal processing applications scaling to thousands of application instances with Fast Fourier Transform and matrix multiply accelerators. We provide insights into the tradeoffs present in this design space through a number of distinct case studies. CEDR is portable and has been deployed and validated on Odroid-XU3, X86, and Nvidia Jetson Xavier-based SoC platforms. Taken together, CEDR is a capable environment for enabling research in exploring the boundaries of productive application development, resource management heuristic development, and hardware configuration analysis for heterogeneous architectures. Joshua Mack, Md Sahil Hassan, Nirmal Kumbhare, Miguel Castro-Gonzalez, Ali Akoglu |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2022 | Enabling Software-Defined RF Convergence with a Novel Coarse-Scale Heterogeneous ProcessorabstractRF system development is traditionally constrained by a restrictive trade-off between power efficiency and programmatic flexibility. We outline a path towards achieving both, thereby enabling a range of new system concepts that better utilize limited resources. As an example, for many future applications, we consider RF convergence – reusing the same spectrum and waveforms to achieve multiple distributed system functions and goals, simultaneously. To enable this next step in processing, we develop a novel framework that includes both software and the system-on-chip (SoC) design. Daniel W. Bliss, Tutu Ajayi, Ali Akoglu, Ilkin Aliyev, Toygun Basaklar, Leul Belayneh, David T. Blaauw, John S. Brunhaver, Chaitali Chakrabarti, Liangliang Chang, Kuan-Yu Chen 0001, Ming-Hung Chen, Xing Chen 0004, Alex R. Chiriyath, Alhad Daftardar, Ronald G. Dreslinski, Arindam Dutta, Allen-Jasmin Farcas, Yukang Fu, A. Alper Goksoy, Xin He 0011, Md Sahil Hassan, Andrew Herschfelt, Jacob Holtom, Hun-Seok Kim, Anish Krishnakumar, Owen Ma, Joshua Mack, Saurav Mallik, Sumit K. Mandal, Radu Marculescu, Brittany M. McCall, Trevor N. Mudge, Ümit Y. Ogras, Vishrut Pandey, Saquib Ahmad Siddiqui, Yu-Hsiu Sun, Adarsh A. Venkataramani, Xiangdong Wei, Benjamin R. Willis, Hanguang Yu, Yufan Yue |
ISCAS | 29 |
| 2022 | A Hardware-based HEFT Scheduler Implementation for Dynamic Workloads on Heterogeneous SoCsabstractNon-uniform performance and power consumption across the processing elements (PEs) of heterogeneous SoCs increase the computation complexity of the task scheduling problem compared to homogeneous architectures. Latency of a software-based scheduler with the increased heterogeneity level in terms of number and types of PEs creates the necessity of deploying a scheduler as an overlay processor in hardware to be able to make scheduling decisions rapidly and enable deployment of real-life applications on heterogeneous SoCs. In this study we present the design trade-offs involved for implementing and deploying the runtime variant of the heterogeneous earliest finish time algorithm (HEFTRT) on the FPGA. We conduct performance evaluations on an SoC configuration emulated over the Xilinx Zynq ZCU102 platform. In a runtime environment we demonstrate hardware-based HEFTRT’s ability to make scheduling decisions with 9.144 ns latency on average, process 26.7% more tasks per second compared to its software counterpart, and reduce the scheduling latency by up to a factor of 183 based on workloads composed of a mixture of dynamically ×arriving real-life signal processing applications. Alexander Fusco, Md Sahil Hassan, Joshua Mack, Ali Akoglu |
VLSI-SoC | 3 |
| 2022 | Performant, Multi-Objective Scheduling of Highly Interleaved Task Graphs on Heterogeneous System on Chip DevicesabstractPerformance-, power-, and energy-aware scheduling techniques play an essential role in optimally utilizing processing elements (PEs) of heterogeneous systems. List schedulers, a class of low-complexity static schedulers, have commonly been used in static execution scenarios. However, list schedulers are not suitable for runtime decision making, particularly when multiple concurrent applications are interleaved dynamically. For such cases, the static task execution times and expectation of idle PEs assumed by list schedulers lead to inefficient system utilization and poor performance. To address this problem, we present techniques for optimizing execution of list scheduling algorithms in dynamic runtime scenarios via a family of algorithms inspired by the well-known heterogeneous earliest finish time (HEFT) list scheduler. Through dynamically arriving, realistic workload scenarios that are simulated in an open-source discrete event heterogeneous SoC simulator, we exhaustively evaluate each of the proposed algorithms across two SoCs modeled after the Xilinx Zynq Ultrascale+ ZCU102 and O-Droid XU3 development boards. Altogether, depending on the chosen variant in this family of algorithms, we are able to achieve an up to 39% execution time improvement, up to 7.24x algorithmic speedup, or up to 30% energy consumption improvement compared to the baseline HEFT implementation. Joshua Mack, Samet E. Arda, Ümit Y. Ogras, Ali Akoglu |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2021 | RANC: Reconfigurable Architecture for Neuromorphic ComputingabstractNeuromorphic architectures have been introduced as platforms for energy-efficient spiking neural network execution. The massive parallelism offered by these architectures has also triggered interest from nonmachine learning application domains. In order to lift the barriers to entry for hardware designers and application developers, we present RANC: a reconfigurable architecture for neuromorphic computing, an opensource highly flexible ecosystem that enables rapid experimentation with neuromorphic architectures in both software via C++ simulation and hardware via FPGA emulation. We present the utility of the RANC ecosystem by showing its ability to recreate behavior of IBM’s TrueNorth and validate with a direct comparison to IBM’s Compass simulation environment and published literature. RANC allows optimizing architectures based on application insights as well as prototyping future neuromorphic architectures that can support new classes of applications entirely. We demonstrate the highly parameterized and configurable nature of RANC by studying the impact of architectural changes on improving application mapping efficiency with quantitative analysis based on Alveo U250 FPGA. We present post routing resource usage and throughput analysis across implementations of synthetic aperture radar classification and vector matrix multiplication applications, and demonstrate a neuromorphic architecture that scales to emulating 259K distinct neurons and 73.3M distinct synapses. Joshua Mack, Ruben Purdy, Kris Rockowitz, Michael Inouye, Edward Richter, Spencer Valancius, Nirmal Kumbhare, Md Sahil Hassan, Kaitlin Lindsay Fair, John Mixter, Ali Akoglu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2020 | FPGA Based High-Throughput Real-Time Feature Extraction for Modulation ClassificationabstractThe spectral correlation density (SCD) function is a feature extraction method used in signal classification systems. Due to its computational complexity, SCD has not been a desirable method for systems under power and real-time constraints. In this study, we present results for a hardware implementation of key kernels of the SCD function on a Field Programmable Gate Array (FPGA). By analyzing profiling results for a state of the art GPU implementation, we developed a preliminary architecture that is able to accelerate the most computationally demanding aspects of the SCD algorithm. We find that this FPGA architecture is able to achieve a 2.03X speedup relative to state of the art GPU-based SCD implementations by coupling SCD's large-scale data-parallel nature with an architecture well suited for fine-grained control flow and data access patterns. Joshua Mack, Ali Akoglu |
FCCM | 1 |
| 2020 | DS3: A System-Level Domain-Specific System-on-Chip Simulation FrameworkabstractHeterogeneous systems-on-chip (SoCs) are highly favorable computing platforms due to their superior performance and energy efficiency potential compared to homogeneous architectures. They can be further tailored to a specific domain of applications by incorporating processing elements (PEs) that accelerate frequently used kernels in these applications. However, this potential is contingent upon optimizing the SoC for the target domain and utilizing its resources effectively at runtime. To this end, system-level design - including scheduling, power-thermal management algorithms and design space exploration studies - plays a crucial role. This article presents a system-level domain-specific SoC simulation (DS3) framework to address this need. DS3 enables both design space exploration and dynamic resource management for power-performance optimization of domain applications. We showcase DS3 using six real-world applications from wireless communications and radar processing domain. DS3, as well as the reference applications, is shared as open-source software to stimulate research in this area. Samet E. Arda, Anish Krishnakumar, A. Alper Goksoy, Nirmal Kumbhare, Joshua Mack, Anderson Luiz Sartor, Ali Akoglu, Radu Marculescu, Ümit Y. Ogras |
IEEE Trans. Computers | 5 |
| 2019 | Accelerated Shadow Detection and Removal MethodabstractShadows can have a negative effect on the ability of computer vision techniques for object detection, tracking, and recognition. Therefore, ability to remove shadows and byproducts of illumination is an important problem to enable effective object recognition actions. As applications move into levels of higher information extraction and higher required processing speeds, efficient and sophisticated shadow detection and removal becomes even more necessary. In this study we propose a shadow removal method, parallelize using a Tesla P100 GPU, and achieve a speedup of 21.67× on an 18 megapixel (MP) resolution image compared to the same method implemented in Matlab. Edward Richter, Ryan Raettig, Joshua Mack, Spencer Valancius, Burak Unal, Ali Akoglu |
AICCSA | 3 |