Md Sahil Hassan

dblp:259/0192 · DBLP profile ↗
← Back
9ranked-venue papers
1as first author
9since 2021 · last 2026
0000-0002-4574-9555ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 1 first-author · 8 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A portable framework with generalized runtime features for task graph execution and concurrent multi-application deployment on heterogeneous systems
Serhan Gener, Md Sahil Hassan, Hasan Umut Suluhan, Liangliang Chang, Chaitali Chakrabarti, Tsung-Wei Huang, Ümit Y. Ogras, Ali Akoglu
Future Gener. Comput. Syst.2
2025 Tutorial: CEDR: A Holistic Software and Hardware Design Environment for Hardware Agnostic Application Development and Deployment on FPGA-Integrated Heterogeneous Systems
abstract
This hands-on tutorial introduces CEDR, an open-source compilation and runtime framework designed to simplify application development and deployment on heterogeneous computing platforms composed of general purpose processors and hardware accelerators. CEDR unifies programming and execution on such platforms, enabling streamlined development workflows. Building on prior lectures and tutorials at ESWEEK 2023 and ISFPGA 2023/2025, this session aims to lower the barrier to research and prototyping in heterogeneous systems. Tailored for application developers and system architects, the tutorial in collaboration with the AMD University Program focuses on key capabilities of CEDR through hands-on programming exercises, including application development, functional verification, performance analysis, and design space exploration using emulation on FPGA-based SoCs.
Serhan Gener, Md Sahil Hassan, Ali Akoglu
CODES+ISSS2
2025 RIMMS: Runtime Integrated Memory Management System for Heterogeneous Computing
abstract
Efficient memory management in heterogeneous systems is increasingly challenging due to diverse compute architectures (e.g., CPU, GPU, and FPGA) and dynamic task mappings not known at compile time. Existing approaches often require programmers to manage data placement and transfers explicitly, or assume static mappings that limit portability and scalability. This article introduces RIMMS (Runtime Integrated Memory Management System), a lightweight, runtime-managed, hardware-agnostic memory abstraction layer that decouples application development from low-level memory operations. RIMMS transparently tracks data locations, manages consistency, and supports efficient memory allocation across heterogeneous compute elements without requiring platform-specific tuning or code modifications. We integrate RIMMS into a baseline runtime and evaluate with complete radar signal processing applications across CPU+GPU and CPU+FPGA platforms. RIMMS delivers up to 2.43× speedup on GPU-based and 1.82× on FPGA-based systems over the baseline. Compared to IRIS, a recent heterogeneous runtime system, RIMMS achieves up to 3.08X speedup and matches the performance of native CUDA implementations while significantly reducing programming complexity. Despite operating at a higher abstraction level, RIMMS incurs only 1–2 cycles of overhead per memory management call, making it a low-cost solution. These results demonstrate RIMMS’s ability to deliver high performance and enhanced programmer productivity in dynamic, real-world heterogeneous environments.
Serhan Gener, Aditya Ukarande, Shilpa Mysore Srinivasa Murthy, Md Sahil Hassan, Joshua Mack, Chaitali Chakrabarti, Ümit Y. Ogras, Ali Akoglu
ACM Trans. Embed. Comput. Syst.4
2023 A Novel Implementation Methodology for Error Correction Codes on a Neuromorphic Architecture
abstract
The Internet of Things infrastructure connects a massive number of edge devices with an increasing demand for intelligent sensing and inferencing capability. Such data-sensitive functions necessitate energy-efficient and programmable implementations of error correction codes (ECCs) and decoders. The algorithmic flow of ECCs with concurrent accumulation and comparison types of operations are innately exploitable by neuromorphic architectures for energy-efficient execution—an area that is relatively unexplored outside of machine learning applications. For the first time, we propose a methodology to map the hard-decision class of decoder algorithms on a neuromorphic architecture. We present the implementation of the Gallager B (GaB) decoding algorithm on a TrueNorth-inspired architecture that is emulated on the Xilinx Zynq ZCU102 MPSoC. Over this reference implementation, we propose architectural modifications at the neuron block level that result in a reduction of energy consumption by 31% with a negligible increase in resource usage while achieving the same error correction performance.
Md Sahil Hassan, Parker Dattilo, Ali Akoglu
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2023 CEDR: A Compiler-integrated, Extensible DSSoC Runtime
abstract
In this work, we present a C ompiler-integrated, E xtensible D omain Specific System on Chip R untime (CEDR) ecosystem to facilitate research toward addressing the challenges of architecture, system software, and application development with distinct plug-and-play integration points in a unified compile time and runtime workflow. We demonstrate the utility of CEDR on the Xilinx Zynq MPSoC-ZCU102 for evaluating performance of pre-silicon hardware in the trade space of SoC configuration, scheduling policy and workload complexity based on dynamically arriving workload scenarios composed of real-life signal processing applications scaling to thousands of application instances with Fast Fourier Transform and matrix multiply accelerators. We provide insights into the tradeoffs present in this design space through a number of distinct case studies. CEDR is portable and has been deployed and validated on Odroid-XU3, X86, and Nvidia Jetson Xavier-based SoC platforms. Taken together, CEDR is a capable environment for enabling research in exploring the boundaries of productive application development, resource management heuristic development, and hardware configuration analysis for heterogeneous architectures.
Joshua Mack, Md Sahil Hassan, Nirmal Kumbhare, Miguel Castro-Gonzalez, Ali Akoglu
ACM Trans. Embed. Comput. Syst.2
2022 Enabling Software-Defined RF Convergence with a Novel Coarse-Scale Heterogeneous Processor
abstract
RF system development is traditionally constrained by a restrictive trade-off between power efficiency and programmatic flexibility. We outline a path towards achieving both, thereby enabling a range of new system concepts that better utilize limited resources. As an example, for many future applications, we consider RF convergence – reusing the same spectrum and waveforms to achieve multiple distributed system functions and goals, simultaneously. To enable this next step in processing, we develop a novel framework that includes both software and the system-on-chip (SoC) design.
Daniel W. Bliss, Tutu Ajayi, Ali Akoglu, Ilkin Aliyev, Toygun Basaklar, Leul Belayneh, David T. Blaauw, John S. Brunhaver, Chaitali Chakrabarti, Liangliang Chang, Kuan-Yu Chen 0001, Ming-Hung Chen, Xing Chen 0004, Alex R. Chiriyath, Alhad Daftardar, Ronald G. Dreslinski, Arindam Dutta, Allen-Jasmin Farcas, Yukang Fu, A. Alper Goksoy, Xin He 0011, Md Sahil Hassan, Andrew Herschfelt, Jacob Holtom, Hun-Seok Kim, Anish Krishnakumar, Owen Ma, Joshua Mack, Saurav Mallik, Sumit K. Mandal, Radu Marculescu, Brittany M. McCall, Trevor N. Mudge, Ümit Y. Ogras, Vishrut Pandey, Saquib Ahmad Siddiqui, Yu-Hsiu Sun, Adarsh A. Venkataramani, Xiangdong Wei, Benjamin R. Willis, Hanguang Yu, Yufan Yue
ISCAS22
2022 A Hardware-based HEFT Scheduler Implementation for Dynamic Workloads on Heterogeneous SoCs
abstract
Non-uniform performance and power consumption across the processing elements (PEs) of heterogeneous SoCs increase the computation complexity of the task scheduling problem compared to homogeneous architectures. Latency of a software-based scheduler with the increased heterogeneity level in terms of number and types of PEs creates the necessity of deploying a scheduler as an overlay processor in hardware to be able to make scheduling decisions rapidly and enable deployment of real-life applications on heterogeneous SoCs. In this study we present the design trade-offs involved for implementing and deploying the runtime variant of the heterogeneous earliest finish time algorithm (HEFTRT) on the FPGA. We conduct performance evaluations on an SoC configuration emulated over the Xilinx Zynq ZCU102 platform. In a runtime environment we demonstrate hardware-based HEFTRT’s ability to make scheduling decisions with 9.144 ns latency on average, process 26.7% more tasks per second compared to its software counterpart, and reduce the scheduling latency by up to a factor of 183 based on workloads composed of a mixture of dynamically ×arriving real-life signal processing applications.
Alexander Fusco, Md Sahil Hassan, Joshua Mack, Ali Akoglu
VLSI-SoC2
2022 JITA4DS: Disaggregated Execution of Data Science Pipelines Between the Edge and the Data Centre
abstract
This paper targets the execution of data science (DS) pipelines supported by data processing, transmission and sharing across several resources executing greedy processes. Current data science pipelines environments provide various infrastructure services with computing resources such as general-purpose processors (GPP), Graphics Processing Units (GPUs), Field Programmable Gate Arrays (FPGAs) and Tensor Processing Unit (TPU) coupled with platform and software services to design, run and maintain DS pipelines. These one-fits-all solutions impose the complete externalization of data pipeline tasks. However, some tasks can be executed in the edge, and the backend can provide just in time resources to ensure ad-hoc and elastic execution environments.This paper introduces an innovative composable “Just in Time Architecture” for configuring DCs for Data Science Pipelines (JITA-4DS) and associated resource management techniques. JITA-4DS is a cross-layer management system that is aware of both the application characteristics and the underlying infrastructures to break the barriers between applications, middleware/operating system, and hardware layers. Vertical integration of these layers is needed for building a customizable Virtual Data Center (VDC) to meet the dynamically changing data science pipelines’ requirements such as performance, availability, and energy consumption. Accordingly, the paper shows an experimental simulation devoted to run data science workloads and determine the best strategies for scheduling the allocation of resources implemented by JITA-4DS.
Genoveva Vargas-Solar, Md Sahil Hassan, Ali Akoglu
J. Web Eng.2
2021 RANC: Reconfigurable Architecture for Neuromorphic Computing
abstract
Neuromorphic architectures have been introduced as platforms for energy-efficient spiking neural network execution. The massive parallelism offered by these architectures has also triggered interest from nonmachine learning application domains. In order to lift the barriers to entry for hardware designers and application developers, we present RANC: a reconfigurable architecture for neuromorphic computing, an opensource highly flexible ecosystem that enables rapid experimentation with neuromorphic architectures in both software via C++ simulation and hardware via FPGA emulation. We present the utility of the RANC ecosystem by showing its ability to recreate behavior of IBM’s TrueNorth and validate with a direct comparison to IBM’s Compass simulation environment and published literature. RANC allows optimizing architectures based on application insights as well as prototyping future neuromorphic architectures that can support new classes of applications entirely. We demonstrate the highly parameterized and configurable nature of RANC by studying the impact of architectural changes on improving application mapping efficiency with quantitative analysis based on Alveo U250 FPGA. We present post routing resource usage and throughput analysis across implementations of synthetic aperture radar classification and vector matrix multiplication applications, and demonstrate a neuromorphic architecture that scales to emulating 259K distinct neurons and 73.3M distinct synapses.
Joshua Mack, Ruben Purdy, Kris Rockowitz, Michael Inouye, Edward Richter, Spencer Valancius, Nirmal Kumbhare, Md Sahil Hassan, Kaitlin Lindsay Fair, John Mixter, Ali Akoglu
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.8