Serhan Gener

dblp:197/1239 · also Y. Serhan Gener · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
7since 2021 · last 2026
0000-0002-8163-1191ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 4 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 A portable framework with generalized runtime features for task graph execution and concurrent multi-application deployment on heterogeneous systems
Serhan Gener, Md Sahil Hassan, Hasan Umut Suluhan, Liangliang Chang, Chaitali Chakrabarti, Tsung-Wei Huang, Ümit Y. Ogras, Ali Akoglu
Future Gener. Comput. Syst.1
2025 Tutorial: CEDR: A Holistic Software and Hardware Design Environment for Hardware Agnostic Application Development and Deployment on FPGA-Integrated Heterogeneous Systems
abstract
This hands-on tutorial introduces CEDR, an open-source compilation and runtime framework designed to simplify application development and deployment on heterogeneous computing platforms composed of general purpose processors and hardware accelerators. CEDR unifies programming and execution on such platforms, enabling streamlined development workflows. Building on prior lectures and tutorials at ESWEEK 2023 and ISFPGA 2023/2025, this session aims to lower the barrier to research and prototyping in heterogeneous systems. Tailored for application developers and system architects, the tutorial in collaboration with the AMD University Program focuses on key capabilities of CEDR through hands-on programming exercises, including application development, functional verification, performance analysis, and design space exploration using emulation on FPGA-based SoCs.
Serhan Gener, Md Sahil Hassan, Ali Akoglu
CODES+ISSS1
2025 K-PACT: Kernel Planning for Adaptive Context Switching - A Framework for Clustering, Placement, and Prefetching in Spectrum Sensing
abstract
Efficient wideband spectrum sensing requires rapid evaluation and re-evaluation of signal presence and type across multiple subchannels. These tasks involve multiple hypothesis testing, where each hypothesis is implemented as decision tree workflow with compute-intensive kernels, including FFT, matrix operations, and signal-specific analyses. Given the dynamic nature of the spectrum environment, the ability to quickly switch between hypotheses is essential for maintaining low-latency, high-throughput operation. This work assumes a coarse-grained reconfigurable architecture consisting of an array of processing elements (PEs), each equipped with a local instruction memory (IMEM) capable of storing and executing kernels used in spectrum sensing applications. We propose a planner tool that efficiently maps hypothesis workflows onto this architecture to enable fast runtime context switching with minimal overhead. The planner performs two key tasks: clustering temporally non-overlapping kernels to share IMEM resources within a PE sub-array, and placing these clusters onto hardware to ensure efficient scheduling and data movement. By preloading kernels that are not simultaneously active into the same IMEM, our tool enables low-latency reconfiguration without runtime conflicts. It models the planning process as a multi-objective optimization, balancing trade-offs among context switch overhead, scheduling latency, and dataflow efficiency. We evaluate the proposed tool in simulated spectrum sensing scenario with 48 concurrent subchannels. Results show that our approach reduces off-chip binary fetches by 207.81×, lowers average switching time by 98.24×, and improves per-subband execution time by 132.92× over baseline without preloading. These improvements demonstrate that intelligent planning is critical for adapting to fast-changing spectrum environments in next-generation radio frequency systems.
Hasan Umut Suluhan, Serhan Gener, Chaitali Chakrabarti, Ümit Y. Ogras, Ali Akoglu
ICCAD3
2025 Coarse-Grained Task Parallelization by Dynamic Profiling for Heterogeneous SoC-Based Embedded System
abstract
In this study, we introduce a methodology for automatically transforming user applications written in C/C++ to a parallel representation consisting of coarse-grained tasks based on dynamic profiling. Such a parallel representation is suitable for mapping applications onto heterogeneous SoCs. We present our approach for instrumenting the user application binary during the compilation process with parallel primitives that enable the runtime system to schedule and execute independent computation-intensive coarse-grained tasks concurrently. We use the proposed compilation and code transformation methodology to retarget each application for execution on a heterogeneous SoC composed of processor cores and accelerators. We demonstrate the capabilities of our integrated compile time and runtime flow through task-level parallelization and functionally correct execution of real-world applications in the communication systems and radar processing domains. We demonstrate the functionality of our integrated system by executing six distinct applications with different degrees of parallelism on four different platforms: an eight-core general-purpose processor, a heterogeneous SoC simulator, and two heterogeneous SoCs utilizing the Xilinx Zynq UltraScale+ FPGA and the Nvidia Jetson AGX board. Our integrated approach offers a path forward for application developers to take full advantage of the target SoC without requiring users to become hardware or parallel programming experts.
Liangliang Chang, Serhan Gener, Joshua Mack, Hasan Umut Suluhan, Ali Akoglu, Chaitali Chakrabarti
ACM Trans. Embed. Comput. Syst.2
2025 RIMMS: Runtime Integrated Memory Management System for Heterogeneous Computing
abstract
Efficient memory management in heterogeneous systems is increasingly challenging due to diverse compute architectures (e.g., CPU, GPU, and FPGA) and dynamic task mappings not known at compile time. Existing approaches often require programmers to manage data placement and transfers explicitly, or assume static mappings that limit portability and scalability. This article introduces RIMMS (Runtime Integrated Memory Management System), a lightweight, runtime-managed, hardware-agnostic memory abstraction layer that decouples application development from low-level memory operations. RIMMS transparently tracks data locations, manages consistency, and supports efficient memory allocation across heterogeneous compute elements without requiring platform-specific tuning or code modifications. We integrate RIMMS into a baseline runtime and evaluate with complete radar signal processing applications across CPU+GPU and CPU+FPGA platforms. RIMMS delivers up to 2.43× speedup on GPU-based and 1.82× on FPGA-based systems over the baseline. Compared to IRIS, a recent heterogeneous runtime system, RIMMS achieves up to 3.08X speedup and matches the performance of native CUDA implementations while significantly reducing programming complexity. Despite operating at a higher abstraction level, RIMMS incurs only 1–2 cycles of overhead per memory management call, making it a low-cost solution. These results demonstrate RIMMS’s ability to deliver high performance and enhanced programmer productivity in dynamic, real-world heterogeneous environments.
Serhan Gener, Aditya Ukarande, Shilpa Mysore Srinivasa Murthy, Md Sahil Hassan, Joshua Mack, Chaitali Chakrabarti, Ümit Y. Ogras, Ali Akoglu
ACM Trans. Embed. Comput. Syst.1
2023 PyTorch and CEDR: Enabling Deployment of Machine Learning Models on Heterogeneous Computing Systems
abstract
The PyTorch programming interface enables efficient deployment of machine learning models, leveraging the parallelism offered by GPU architectures. In this study, we present the integration of the PyTorch framework with a compiler and runtime ecosystem . Our aim is to demonstrate the ability to deploy PyTorch-based models on FPGA-based SoC platforms, without requiring users to possess prior FPGA-based design experience. The proposed PyTorch model transformation approach expands the range of hardware architectures that PyTorch developers can target, enabling them to take advantage of the energy-efficient execution provided by heterogeneous computing systems. Our experiments involve compiling and executing real-life applications on heterogeneous SoC configurations emulated on the Xilinx Zynq Ultrascale+ ZCU102 system. We showcase our ability to deploy three distinct PyTorch applications, encompassing object detection, visual geometry group (VGG), and speech classification, using the integrated compiler and runtime system without loss of model accuracy. Furthermore, we extend our analysis by evaluating dynamically arriving workload scenarios, consisting of a mix of PyTorch models and non-PyTorch-based applications. Through these experiments, we vary the hardware composition and scheduling heuristics. Our findings indicate that when PyTorch-based applications coexist with unrelated applications, our integrated scheduler fairly dispatches tasks to the FPGA platform’s accelerator and CPU cores, without compromising the target throughput for each application.
Hasan Umut Suluhan, Serhan Gener, Alexander Fusco, H. Fatih Ugurdag, Ali Akoglu
AICCSA2
2022 GPGPU-based High Throughput Image Pre-processing Towards Large-Scale Optical Character Recognition
abstract
Studies have shown that pre-processing digital images through scaling, rotation and blurring type of operations allow optical character recognition (OCR) to focus on the key features in the image and result in improving recognition accuracy. We leverage the open-source Tesseract OCR and show that its accuracy can be improved through a pre-processing flow that includes thresholding, rotation, rescaling, erosion, dilation, and noise removal steps based on a dataset that is formed of 560 phone screen images. However, the serial CPU-based implementation of this flow introduces a latency of 48.32 ms per image on average. Even though time scale is low in the context of a single image, this latency poses as a barrier when processing millions of images with OCR. To address this, we parallelize the entire pre-processing flow on the Nvidia P100 GPU, implement a streaming based execution, and reduce the latency to 0.846 ms. This streaming-enabled implementation enables setting up a GPU based OCR engine to process large scale workloads.
Serhan Gener, Parker Dattilo, Dhruv Gajaria, Alexander Fusco, Ali Akoglu
AICCSA1
2019 Lossless Look-Up Table Compression for Hardware Implementation of Transcendental Functions
abstract
Look-Up Table (LUT) implementation of transcendental functions often offers lower latency compared to algebraic implementations at the expense of significant area penalty. MultiPartite table method (MP) can circumvent the area problem by breaking up the implementation into multiple smaller LUTs. However, even these smaller LUTs may be big in high accuracy MP designs. Lossless LUT compression can be applied to one or more of these LUTs to further improve area and even timing in some cases. The state-of-the-art 2T-TIV and 3T-TIV methods decompose the Table of Initial Values (TIV) of MP into a table of pivots and tables of differences from the pivots. Our technique, which we call Fully Random Access differential LUT (FR-dLUT), instead uses differences of consecutive elements and results in a smaller range of differences. We also propose a variant of FR-dLUT with variable length coding (Huffman) called FR-dLUT-VL, which introduces don't cares into the difference tables and lets logic synthesis optimize them out. We implemented Verilog generators of MP for sine and exponential, where TIV is a conventional LUT as well as 2T-TIV, 3T-TIV, FR-dLUT, and FR-dLUT-VL. We synthesized the generated designs on FPGA and found that our techniques produce around 10% improvement in area and timing beyond the state-of-the-art in large bit widths.
Serhan Gener, Sezer Gören 0001, H. Fatih Ugurdag
VLSI-SoC1
2017 Hardware Division by Small Integer Constants
abstract
This article studies the design of custom circuits for division by a small positive constant. Such circuits can be useful for specific FPGA and ASIC applications. The first problem studied is the Euclidean division of an unsigned integer by a constant, computing a quotient and remainder. Several new solutions are proposed and compared against the state-of-the-art. As the proposed solutions use small look-up tables, they match well with the hardware resources of an FPGA. The article then studies whether the division by the product of two constants is better implemented as two successive dividers or as one atomic divider. It also considers the case when only a quotient or only a remainder is needed. Finally, it addresses the correct rounding of the division of a floating-point number by a small integer constant. All these solutions, and the previous state-of-the-art, are compared in terms of timing, area, and area-timing product. In general, the relevance domains of the various techniques are different on FPGA and on ASIC.
H. Fatih Ugurdag, Florent de Dinechin, Serhan Gener, Sezer Gören 0001, Laurent-Stéphane Didier
IEEE Trans. Computers3