Ismail Akturk

dblp:76/7892 · DBLP profile ↗
← Back
15ranked-venue papers
6as first author
5since 2021 · last 2025
0000-0003-1970-2507ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 4 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Early Termination with Activation Sign Prediction for Energy-Efficient CNN Inference Using Sum-of-Power-of-Two Quantization
abstract
We propose two techniques to optimize CNN inference for both energy efficiency and accuracy. First, Sum-of-Power-of-Two (SPoT) quantization enhances logarithmic quantization by enabling shift-based multiplications with reduced accuracy loss. Second, activation sign prediction enables early termination by estimating pre-activation signs from the most significant power-of-two terms and skipping computations that ReLU would zero out. To validate these methods, we design a processing element (PE) with a shift-based MAC unit integrating SPoT and early termination.
Emir Eryilmaz, Selim Sandal, Ismail Akturk
ICCD3
2025 Design and Evaluation of an N-Trace Compliant Hardware Tracer for RISC-V Processors
abstract
Effective tracing at both hardware and software levels is critical for fault detection and debugging in processor design. This paper presents a hardware tracer architecture specifically designed for complex RISC-V processors, adhering to the recently standardized RISC-V N-Trace specification. As one of the earliest implementations targeting this specification, our architecture uniquely supports dynamic encoding, enabling the trace message size to adapt efficiently at runtime. The proposed design achieves compression footprint as low as 0.11 bits per instruction, averaging 2.06 bits per instruction across the evaluated benchmarks, demonstrating a very high compression rate with an area overhead of only 1.05%. Despite its compact footprint, the tracer maintains robust performance, consuming just 23 mW of power at 0.5 GHz using TSMC's 16nm FinFET technology. This combination of adaptability, efficiency, and low power dissipation underscores the potential of the proposed hardware tracer for high-performance fault detection and debugging in advanced RISC-V-based systems.
Omer Karslioglu, Ismail Akturk
ICCD2
2025 Optimizing Energy Efficiency in Heterogeneous Computing via Multi-objective Scheduling with Reinforcement Learning
Ezgi Nur Alisan, Ismail Akturk
JSSPP2
2025 Providing Edge to Cloud Continuum With Adaptive Model Selection and Operational Score
abstract
Advancements in Deep Neural Network (DNN) models and hardware accelerators have made edge intelligence a practical alternative to cloud-based intelligence. However, applicationspecific requirements, such as accuracy, latency, security, and privacy, as well as workload fluctuations, necessitate dynamic allocation of edge and cloud resources. To facilitate such dynamic allocation, we propose an adaptive model selection and switching framework that leverages operational performance scores. We evaluate the approach using various object classification models, demonstrating its ability to balance accuracy and inference time while ensuring scalability and efficient resource utilization.
Ismail Ari, Nurkan Fatih Altunel, Tugce Ozgirgin, Ezgi Nur Alisan, Emir Eryilmaz, Yarkin Dalgan, Ismail Akturk, Sandeep Pirbhulal, Habtamu Abie
VTC2025-Spring7
2025 The Case for Secure Miniservers Beyond the Edge
abstract
Beyond edge devicescan function off the power grid and without batteries, making them suitable for deployment in hard-to-reach environments. As the energy budget is extremely tight, energy-hungry long-distance communication required for offloading computation or reporting results to a server becomes a significant limitation. Based on the observation that the energy required for communication decreases with shorter distances, this paper makes a case for the deployment ofsecure beyond edge miniservers. These are strategically positioned, lightweight local servers designed to support beyond edge devices without compromising the privacy of sensitive information. We demonstrate that even for relatively small scale representative computations – which are more likely to fit into the tight power budget of a beyond edge device for local processing – deploying a beyond edge miniserver can lead to higher performance. To this end, we consider representative deployment scenarios of practical importance, including but not limited to agricultural systems or building structures, where beyond edge miniservers enable highly energy-efficient real-time data processing.
Salonik Resch, M. Hüsrev Cilasun, Zamshed I. Chowdhury, Masoud Zabihi, Yang Lv 0003, Jianping Wang 0006, Sachin S. Sapatnekar, Ismail Akturk, Ulya R. Karpuzcu
IEEE Trans. Computers8
2020 ACR: Amnesic Checkpointing and Recovery
abstract
Systematic checkpointing of the machine state makes restart of execution from a safe state possible upon detection of an error. The time and energy overhead of checkpointing, however, grows with the frequency of checkpointing. Considering the growth of expected error rates, amortizing this overhead becomes especially challenging, as checkpointing frequency tends to increase with increasing error rates. Based on the observation that due to imbalanced technology scaling, recomputing a data value can be more energy efficient than retrieving (i.e., loading) a stored copy, this paper explores how recomputation of data values (which otherwise would be read from a checkpoint from memory or secondary storage) can reduce the machine state to be checkpointed, and thereby, the checkpointing overhead. Even in a relatively small scale system, recomputation-based checkpointing can reduce the storage overhead by up to 23.91%; time overhead, by 11.92%; and energy overhead, by 12.53%, respectively.
Ismail Akturk, Ulya R. Karpuzcu
HPCA1
2017 AMNESIAC: Amnesic Automatic Computer
abstract
Due to imbalances in technology scaling, the energy consumption of data storage and communication by far exceeds the energy consumption of actual data production, i.e., computation. As a consequence, recomputing data can become more energy efficient than storing and retrieving precomputed data. At the same time, recomputation can relax the pressure on the memory hierarchy and the communication bandwidth. This study hence assesses the energy efficiency prospects of trading computation for communication. We introduce an illustrative proof-of-concept design, identify practical limitations, and provide design guidelines.
Ismail Akturk, Ulya R. Karpuzcu
ASPLOS1
2016 Accuracy Bugs: A New Class of Concurrency Bugs to Exploit Algorithmic Noise Tolerance
abstract
Parallel programming introduces notoriously difficult bugs, usually referred to as concurrency bugs. This article investigates the potential for deviating from the conventional wisdom of writing concurrency bug--free, parallel programs. It explores the benefit of accepting buggy but approximately correct parallel programs by leveraging the inherent tolerance of emerging parallel applications to inaccuracy in computations. Under algorithmic noise tolerance, a new class of concurrency bugs, accuracy bugs, degrade the accuracy of computation (often at acceptable levels) rather than causing catastrophic termination. This study demonstrates how embracing accuracy bugs affects the application output quality and performance and analyzes the impact on execution semantics.
Ismail Akturk, Riad Akram, Mohammad Majharul Islam, Abdullah Muzahid, Ulya R. Karpuzcu
ACM Trans. Archit. Code Optim.1
2014 Accordion: Toward soft Near-Threshold Voltage Computing
abstract
While more cores can find place in the unit chip area every technology generation, excessive growth in power density prevents simultaneous utilization of all. Due to the lower operating voltage, Near-Threshold Voltage Computing (NTC) promises to fit more cores in a given power envelope. Yet NTC prospects for energy efficiency disappear without mitigating (i) the performance degradation due to the lower operating frequency; (ii) the intensified vulnerability to parametric variation. To compensate for the first barrier, we need to raise the degree of parallelism - the number of cores engaged in computation. NTC-prompted power savings dominate the power cost of increasing the core count. Hence, limited parallelism in the application domain constitutes the critical barrier to engaging more cores in computation. To avoid the second barrier, the system should tolerate variation-induced errors. Unfortunately, engaging more cores in computation exacerbates vulnerability to variation further. To overcome NTC barriers, we introduce Accordion, a novel, light-weight framework, which exploits weak scaling along with inherent fault tolerance of emerging R(ecognition), M(ining), S(ynthesis) applications. The key observation is that the problem size not only dictates the number of cores engaged in computation, but also the application output quality. Consequently, Accordion designates the problem size as the main knob to trade off the degree of parallelism (i.e. the number of cores engaged in computation), with the degree of vulnerability to variation (i.e. the corruption in application output quality due to variation-induced errors). Parametric variation renders ample reliability differences between the cores. Since RMS applications can tolerate faults emanating from data-intensive program phases as opposed to control, variation-afflicted Accordion hardware executes fault-tolerant data-intensive phases on error-prone cores, and reserves reliable cores for control.
Ulya R. Karpuzcu, Ismail Akturk, Nam Sung Kim
HPCA2
2014 Application-Specific Heterogeneous Network-on-Chip Design
abstract
As a result of increasing communication demands, application-specific and scalable Network-on-Chips (NoCs) have emerged to connect processing cores and subsystems in Multiprocessor System-on-Chips. A challenge in application-specific NoC design is to find the right balance among different tradeoffs, such as communication latency, power consumption and chip area. We propose a novel approach that generates latency-aware heterogeneous NoC topology. Experimental results show that our approach improves the total communication latency up to 27% with modest power consumption.
Dilek Demirbas, Ismail Akturk, Ozcan Ozturk 0001, Ugur Güdükbay
Comput. J.2
2013 ILP-Based Communication Reduction for Heterogeneous 3D Network-on-Chips
abstract
Network-on-Chip (NoC) architectures and three-dimensional integrated circuits (3D ICs) have been introduced as attractive options for overcoming the barriers in interconnect scaling while increasing the number of cores. Combining these two approaches is expected to yield better performance and higher scalability. This paper explores the possibility of combining these two techniques in a heterogeneity aware fashion. We explore how heterogeneous processors can be mapped onto the given 3D chip area to minimize the data access costs. Our initial results indicate that the proposed approach generates promising results within tolerable solution times.
Ismail Akturk, Ozcan Ozturk 0001
PDP1
2013 Reliability-Aware Heterogeneous 3D Chip Multiprocessor Design
Ismail Akturk, Ozcan Ozturk 0001
J. Electron. Test.1
2013 Improving application behavior on heterogeneous manycore systems through kernel mapping
Omer Erdil Albayrak, Ismail Akturk, Ozcan Ozturk 0001
Parallel Comput.2
2010 Toward a Reliable Distributed Data Management System
abstract
Modern collaborative science has placed increasing burden on data management infrastructure to handle the increasingly large data archives generated. Beside functionality, reliability and availability are also key factors in delivering a data management system that can efficiently and effectively meet the challenges posed and compounded by the unbounded increase in the size of data archive generated by scientific applications. In this paper, we present our work on increasing and improving reliability and availability in the data management system we designed for the PetaShare project, we also discuss our work on benchmarking the performance and scalability of metadata management system in PetaShare project.
Ismail Akturk, Xinqi Wang, Tevfik Kosar
ISPDC1
2009 Cross-domain metadata management in data intensive distributed computing environment
abstract
As the the size of scientific datasets grows, it becomes imperative that cross-domain metadata management system needs to be developed to facilitate interdisciplinary scientific research. Three key issues need to be addressed: the development of a cross-domain metadata schema; the implementation of a metadata management system based on this schema; the integration of the metadata system into existing infrastructure with reasonable performance and scalability. In this paper, we give an overview of the research we have done as part of the PetaShare project to address the above mentioned problems.
Xinqi Wang, Ismail Akturk, Tevfik Kosar
CLUSTER2