Gulsum Gudukbay Akbulut

dblp:263/7638 · also Gulsum Gudukbay · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
7since 2021 · last 2025
0000-0002-1767-2204ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 1 first-author · 5 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Multi-Dimensional ML-Pipeline Optimization in Cost-Effective Disaggregated Datacenter
Pingyi Huo, Anusha Devulapally, Hasan Al Maruf, Nandhini Chandramoorthy, Meena Arunachalam, Gulsum Gudukbay Akbulut, Mahmut T. Kandemir, Narayanan Vijaykrishnan
MICRO6
2025 FPGA-based accelerator for adaptive banded event alignment in nanopore sequencing data analysis
abstract
Adaptive Banded Event Alignment (ABEA) stands as a critical algorithmic component in sequence polishing and DNA methylation detection, employing dynamic programming to align raw Nanopore signal with reference reads. Motivated by the observation that, compared to CPUs and GPUs, cutting-edge FPGAs demonstrate—in certain cases—superior performance at a reduced cost and energy consumption, this paper presents an efficient FPGA-based accelerator for ABEA, leveraging the inherent high parallelism and sequential access pattern within ABEA. Our proposed FPGA-based ABEA accelerator significantly enhances ABEA performance compared to the original CPU-based implementation in Nanopolish as well as the state-of-art acceleration on GPU and FPGA platforms. Specifically, targeting Xilinx VU9P, our accelerator achieves an average throughput speedup of 10.05 $$\times$$ over the CPU-only implementation, an average 1.81 $$\times$$ speedup over the state-of-art GPU acceleration with only 7.2% of the energy, and a speedup of 10.11 $$\times$$ compared to an existing FPGA accelerator. Our work demonstrates that intensive genome analysis can benefit significantly from cutting-edge FPGAs, offering improvements in both performance and energy consumption.
Yilin Feng, Gulsum Gudukbay Akbulut, Narayanan Vijaykrishnan, Mahmut T. Kandemir, Chita R. Das
BMC Bioinform.3
2024 PIFS-Rec: Process-In-Fabric-Switch for Large-Scale Recommendation System Inferences
abstract
Deep Learning Recommendation Models (DLRMs) have become increasingly popular and prevalent in today's datacenters, consuming most of the AI inference cycles. The performance of DLRMs is heavily influenced by available band-width due to their large vector sizes in embedding tables and concurrent accesses. To achieve substantial improvements over existing solutions, novel approaches towards DLRM optimization are needed, especially, in the context of emerging interconnect technologies like CXL. This study delves into exploring CXL-enabled systems, implementing a process-in-fabric-switch (PIFS) solution to accelerate DLRMs while optimizing their memory and bandwidth scalability. We present an in-depth characterization of industry-scale DLRM workloads running on CXL-ready systems, identifying the predominant bottlenecks in existing CXL systems. We, therefore, propose PIFS-Rec, a PIFS-based scheme that implements near-data processing through downstream ports of the fabric switch. PIFS-Rec achieves a latency that is 3.89 x lower than Pond, an industry-standard CXL-based system, and also outperforms BEACON, a state-of-the-art scheme, by 2.03x.
Pingyi Huo, Anusha Devulapally, Hasan Al Maruf, Krishnakumar Nair, Meena Arunachalam, Gulsum Gudukbay Akbulut, Mahmut T. Kandemir, Narayanan Vijaykrishnan
MICRO7
2024 SpotVerse: Optimizing Bioinformatics Workflows with Multi-Region Spot Instances in Galaxy and Beyond
abstract
As demand for cloud computing in bioinformatics increases, various studies have explored options for running large-scale workloads with reduced costs, often leveraging spot instances in multi-region deployments. For example, spot instances offer lower prices but come with the risk of interruption, contrasting with regular (on-demand) instances. However, transitioning to regions with high interruption rates can undermine the benefits of spot instances, adversely affecting performance and cost efficiency. Additionally, regular instances sometimes outperform spot instances based on their specifications. Existing IaaS frameworks focus primarily on cost savings without adequately addressing performance stability in high-interruption regions. To address these challenges, we introduce SpotVerse, a framework designed to optimize cloud resource allocation for bioinformatics workloads, including those within Galaxy - an open-source, web-based platform widely used for managing bioinformatics workflows. SpotVerse efficiently manages long workloads at reduced costs while navigating the complexities of high-interruption regions and strategically selecting between on-demand and spot instances. Our experiments compare SpotVerse with traditional single-region deployments, on-demand instances, and other existing frameworks to evaluate its performance and cost efficiency. Through advanced algorithms for resilient workflows and heuristic resource management, SpotVerse minimizes disruption risks and showcases potential cost savings of up to 52% over traditional single-region deployments.
Myungjun Son, Gulsum Gudukbay Akbulut, Mahmut T. Kandemir
Middleware2
2023 Architecture-Aware Currying
abstract
In near-data computing (NDC), computation is brought into data, as opposed to bringing data to computation. While there is prior work focusing on different NDC opportunities, there is no study, to our knowledge, that investigates the importance of “neighborhood” in NDC. This paper explores the neighborhood concept in multithreaded programs that run on on-chip network-based manycore systems. We define the concept of “neighborhood”, in terms of on-chip network links, and use it to formulate the NDC problem. We propose a “generic” compiler algorithm, called “architecture-aware currying”, that uses the neighborhood concept to implement NDC. So, a core can perform some portions of computation with the nearby data and postpone the remainder of the computation until the remaining data become nearby. It can also perform computations - with nearby data - on behalf of other cores. Our experimental evaluation shows that the proposed compiler algorithm outperforms state-of-the-art data locality optimization strategies.
Mahmut T. Kandemir, Gulsum Gudukbay Akbulut, Wonil Choi, Mustafa Karaköy
PACT2
2023 Data Recomputation for Multithreaded Applications
abstract
Increasing dataset sizes put tremendous pressure on cache hierarchies of multicore and manycore systems, which requires going beyond the current hardware and compiler-based data locality optimization techniques. Data recomputation, which aims to eliminate costly data accesses by replacing each such access with multiple, less costly data accesses plus some computation, is one such technique. However, existing data recomputation techniques are single-thread centric and they do not take advantage of the recomputation opportunities that exist across threads. We propose a novel compiler-guided data recomputation approach that works across threads. Our fully-automated approach has two major components. The first component catches the data recomputation opportunities enabled by a multithreaded execution and takes advantage of them. The second component implements a novel compiler-guided cache replacement strategy that is “recomputation-aware”. The unique aspect of our strategy is that it makes its block/line replacement decisions in the cache based on not only recency information (as in the case of LRU) but also future data recomputation opportunities. Our proposed compiler algorithm improves application performance by an average of 13.25% over the conventional optimizations that do not use data recomputation and 7.68% over a single-thread centric data recomputation scheme. The corresponding improvements when also employing recomputation-conscious caching are 19.12% and 11.63%, respectively.
Gulsum Gudukbay Akbulut, Mahmut T. Kandemir, Mustafa Karaköy, Wonil Choi
ICCAD1
2023 License Forecasting and Scheduling for HPC
abstract
This work focuses on forecasting future license usage for high-performance computing environments and using such predictions to improve the effectiveness of job scheduling. Specifically, we propose a model that carries out both short-term and long-term license usage forecasting and a method of using forecasts to improve job scheduling. Our long-term forecasting model achieves a Mean Absolute Percentage Error (MAPE) as low as 0.26 for a 12-month forecast of daily peak license usage. Our job scheduling experimental results also indicate that wasted work from jobs with insufficient licenses can be reduced by up to 92% without increasing the average license-using job completion times, during periods of high license usage, with our proposed license-aware scheduler.
Ahmed Burak Gulhan, Gulsum Gudukbay Akbulut, Amit Amritkar, Jack Sampson, Vasant G. Honavar, Adam Focht, Chuck Pavloski, Mahmut T. Kandemir
MASCOTS2
2020 ResiRCA: A Resilient Energy Harvesting ReRAM Crossbar-Based Accelerator for Intelligent Embedded Processors
abstract
Many recent works have shown substantial efficiency boosts from performing inference tasks on Internet of Things (IoT) nodes rather than merely transmitting raw sensor data. However, such tasks, e.g., convolutional neural networks (CNNs), are very compute intensive. They are therefore challenging to complete at sensing-matched latencies in ultra-low-power and energy-harvesting IoT nodes. ReRAM crossbar-based accelerators (RCAs) are an ideal candidate to perform the dominant multiplication-and-accumulation (MAC) operations in CNNs efficiently, but conventional, performance-oriented RCAs, while energy-efficient, are power hungry and ill-optimized for the intermittent and unstable power supply of energy-harvesting IoT nodes. This paper presents the ResiRCA architecture that integrates a new, lightweight, and configurable RCA suitable for energy harvesting environments as an opportunistically executing augmentation to a baseline sense-and-transmit battery-powered IoT node. To maximize ResiRCA throughput under different power levels, we develop the ResiSchedule approach for dynamic RCA reconfiguration. The proposed approach uses loop tiling-based computation decomposition, model duplication within the RCA, and inter-layer pipelining to reduce RCA activation thresholds and more closely track execution costs with dynamic power income. Experimental results show that ResiRCA together with ResiSchedule achieve average speedups and energy efficiency improvements of 8× and 14× respectively compared to a baseline RCA with intermittency-unaware scheduling.
Keni Qiu, Nicholas Jao, Mengying Zhao, Cyan Subhra Mishra, Gulsum Gudukbay Akbulut, Sethu Jose, Jack Sampson, Mahmut T. Kandemir, Narayanan Vijaykrishnan
HPCA5