EDBT 2026 Demo / reviewers in the wild / expert
Nihat Altiparmak
dblp:21/8166
· DBLP profile ↗
22ranked-venue papers
7as first author
5since 2021 · last 2025
0000-0002-7038-368XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 14 · 6 first-author · 3 since 2021Computer networks · 4 · 1 first-authorArtificial intelligence and machine learning · 1Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | File Systems' Impact on SSD Performance and Energy EfficiencyabstractThe rapid expansion of digital services, particularly generative AI and cloud computing, has significantly increased the energy demands of data centers. This growth poses serious threats to both the economy and the environment by straining power grids, driving up electricity costs, increasing greenhouse gas emissions, and exacerbating water scarcity. While energy-aware computing practices have emerged as a promising solution, the role of operating system components, especially file systems remains underexplored. In this paper, we investigate the performance and energy efficiency implications of three widely used Linux file systems: ext4,$xfs$, and$f2fs$. Using a high-performance NVMe SSD, a power meter, and a diverse set of synthetic and real-world 10 workloads, we conduct a systematic evaluation across varying request sizes and concurrency levels. Our results reveal that$f2fs$excels in low-intensity random write workloads due to its log-structured design, while$xfs$consistently delivers the best energy efficiency under high-concurrency scenarios. We highlight how file system design choices can significantly influence system-wide energy consumption and advocate for energy-aware system software design as a critical step toward sustainable cloud infrastructure. Sarah Massie, Stephanie Sithu, Jacob Higdon, Bryan Harris, Nihat Altiparmak |
CloudCom | 5 |
| 2023 | Energy Implications of IO Interface Design ChoicesabstractWith the availability of high performance storage technology, there is extra pressure on the efficiency of IO interfaces. In addition to the popular POSIX synchronous, POSIX asynchronous, and Linux asynchronous (libaio) IO interfaces, there are two recent interfaces, spdk and io_uring, that are increasingly attracting attention with their high performance asynchronous designs. While providing high performance IO is crucial, it is also essential to do so in an energy-aware manner. In this paper, we study the energy implications of IO interface design choices and how these choices impact a system's energy consumption. Our empirical evaluation using a power meter, an ultra-low latency storage device, and various workload behaviors including single and multiple thread scenarios allow us to lay out the most energy efficient design choices, with the goal of yielding energy-aware high-performance storage stack designs. Sidharth Sundar, William Simpson, Jacob Higdon, Caeden Whitaker, Bryan Harris, Nihat Altiparmak |
HotStorage | 6 |
| 2023 | Do we still need IO schedulers for low-latency disks?abstractThe performance of recent data storage devices has significantly improved over previous generations, with lower latency, greater throughput, and greater parallelism. Since we now have Ultra-Low Latency (ULL) data storage devices capable of providing data in less than 10 microseconds, in this paper we question the need for IO schedulers for better performance and energy efficiency. Specifically, we measure the latency costs of Linux IO scheduling algorithms and investigate their impact on overall performance and energy efficiency using a ULL storage device, a power meter, and various IO workloads. Our observations indicate that IO schedulers for ULL storage either do not help or significantly increase request latencies while also negatively impacting throughput and energy efficiency. Although we recognize the value of IO schedulers for slower devices or for other metrics such as fairness and QoS, we believe that IO schedulers have become unnecessary for ULL devices to improve performance or energy efficiency. Caeden Whitaker, Sidharth Sundar, Bryan Harris, Nihat Altiparmak |
HotStorage | 4 |
| 2022 | When poll is more energy efficient than interruptabstractPolling is commonly indicated to be a more suitable IO completion mechanism than interrupt for ultra-low latency storage devices. However, polling's impact on overall energy efficiency has not been thoroughly investigated. In this paper, contrary to common belief, we show that polling can also be more energy efficient than interrupt. To do so, we systematically investigate the energy efficiency of all available Linux IO completion mechanisms, including interrupt, classic polling, and hybrid polling using a real ultra-low latency storage device, a power meter, and various workload behaviors. Our experimental results indicate that although hybrid polling provides a good trade-off in CPU utilization, it is the least energy efficient, whereas classic polling is the most energy efficient for low latency IO requests. To the best of our knowledge, this is the first paper classifying polling as more energy efficient than interrupt for a real secondary storage device, and we hope that our observations will lead to more energy efficient IO completion mechanisms for new generation storage device characteristics. Bryan Harris, Nihat Altiparmak |
HotStorage | 2 |
| 2021 | Real-Time Characterization of Data Access CorrelationsabstractEfficient and accurate detection of data access correlations can provide various storage system optimizations. However, existing offline detection techniques are costly in terms of computation and memory, and they rely on previously recorded disk I/O traces, thus wasting additional storage space and causing further I/O. Moreover, due to their offline nature, they are inadequate for allowing automatic optimizations. In this paper, we propose a real-time data access characterization framework that eliminates the drawbacks of offline analysis with a slight compromise in accuracy. The proposed framework continuously monitors I/O requests and forms correlations by applying single-pass data analysis techniques and maintaining a synopsis data structure to efficiently characterize access behavior in various dimensions including spatial locality, temporal locality, and frequency. Our evaluation using synthetic and real world storage workloads indicates that the proposed framework can detect over 90% of data access correlations in real-time, using limited memory. The proposed framework is crucial for enabling future self-optimizing storage systems that can automatically react to I/O bottlenecks. Bryan Harris, Michael Marzullo, Nihat Altiparmak |
ISPASS | 3 |
| 2020 | Ultra-Low Latency SSDs' Impact on Overall Energy Efficiency
Bryan Harris, Nihat Altiparmak |
HotStorage | 2 |
| 2019 | Monte Carlo Based Server Consolidation for Energy Efficient Cloud Data CentersabstractThe growing energy consumption of data centers is a compelling global problem and effective server consolidation is at the heart of energy efficient cloud data centers. A variant of bin packing can be used to model the server consolidation problem, where the constraints are multidimensional and heterogeneous vectors rather than scalars and the goal is to satisfy the requested resource allocation using the minimum number physical servers. Since bin packing is NP-hard, we rely on heuristics for practical solutions. Variations of First Fit Decreasing (FFD) based heuristics have been shown to be effective both in theory and practice for the one dimensional homogeneous case. However, the multidimensional and heterogeneous aspects of the server consolidation problem make it more complicated, requiring additional research to adapt FFD to the server consolidation problem. In this paper, we present a new FFD-based server consolidation technique using a Monte Carlo method and Shannon entropy, which considers resource bottlenecks and dynamically adjusts to variance in the utilization of different resources. The proposed heuristic outperforms existing techniques in all scenarios, achieving within 2-5% of optimal on average for medium to high variance in resource utilization, and within 10% worse than optimal on average for all scenarios. Bryan Harris, Nihat Altiparmak |
CloudCom | 2 |
| 2018 | Towards Adaptive Parallel Storage SystemsabstractDisk I/O is a major bottleneck limiting the performance and scalability of data intensive applications. A common way to address disk I/O bottlenecks is using parallel storage systems and utilizing concurrent operation of independent storage components; however, achieving a consistently high parallel I/O performance is challenging due to static configurations. Modern parallel storage systems, especially in the cloud, enterprise data centers, and scientific clusters are commonly shared by various applications generating dynamic and coexisting data access patterns. Nonetheless, these systems generally utilize one-layout-fits-all data placement strategy frequently resulting in suboptimal I/O parallelism. Guided by association rule mining, graph coloring, bin packing, and network flow techniques, this paper proposes a general framework for adaptive parallel storage systems, with the goal of continuously providing a high-degree of I/O parallelism. Evaluation results indicate that the proposed framework is highly successful in adjusting to skewed parallel access patterns for both hard disk drive (HDD) based traditional storage arrays and solid-state drive (SSD) based all-flash arrays. In addition to the storage arrays, the proposed framework is sufficiently generic and can be tailored to various other parallel storage scenarios including but not limited to key-value stores, parallel/distributed file systems, and internal parallelism of SSDs. Erica Tomes, Everett Neil Rush, Nihat Altiparmak |
IEEE Trans. Computers | 3 |
| 2017 | Big Data Aware Virtual Machine Placement in Cloud Data CentersabstractWhile society continues to be transformed by insights from processing big data, the increasing rate at which this data is gathered is making processing in private clusters obsolete. A vast amount of big data already resides in the cloud, and cloud infrastructures provide a scalable platform for both the computational and I/O needs of big data processing applications. Virtualization is used as a base technology in the cloud; however, existing virtual machine placement techniques do not consider data replication and I/O bottlenecks of the infrastructure, yielding sub-optimal data retrieval times. This paper targets efficient big data processing in the cloud and proposes novel virtual machine placement techniques, which minimize data retrieval time by considering data replication, storage performance, and network bandwidth. We first present an integer-programming based optimal virtual machine placement algorithm and then propose two low cost data- and energy-aware virtual machine placement heuristics. Our proposed heuristics are compared with optimal and existing algorithms through extensive evaluation. Experimental results provide strong indications for the superiority of our proposed solutions in both performance and energy, and clearly outline the importance of big data aware virtual machine placement for efficient processing of large datasets in the cloud. Logan Hall, Bryan Harris, Erica Tomes, Nihat Altiparmak |
BDCAT | 4 |
| 2017 | A Comparative Study of HDD and SSD RAIDs' Impact on Server Energy ConsumptionabstractIn the US alone, data centers consumed around $20 billion (200 TWh) yearly electricity in 2016, and this amount doubles itself every five years. Data storage alone is estimated to be responsible for about 25% to 35% of data-center power consumption. Servers in data centers generally include multiple HDDs or SSDs, commonly arranged in a RAID level for better performance, reliability, and availability. In this study, we evaluate HDD and SSD based Linux (md) software RAIDs' impact on the energy consumption of popular servers. We used the Filebench workload generator to emulate three common server workloads: web, file, and mail, and measured the energy consumption of the system using the HOBO power meter. We observed some similarities and some differences in energy consumption characteristics of HDD and SSD RAIDs, and provided our insights for better energy-efficiency. We hope that our observations will shed light on new energy-efficient RAID designs tailored for HDD and SSD RAIDs' specific energy consumption characteristics. Erica Tomes, Nihat Altiparmak |
CLUSTER | 2 |
| 2016 | Dynamic Data Layout Optimization for High Performance Parallel I/OabstractStorage performance bottlenecks are one of the major threats limiting the scalability of I/O intensive applications. Parallel storage systems have the potential to alleviate I/O bottlenecks through concurrent operation of independent storage components if a parallelism-aware data layout can be continuously guaranteed. Existing systems use one-layout-fits-all data placement strategy that frequently results in sub-optimal I/O parallelism. Guided by association rule mining, graph coloring, bin packing, and network flow techniques, this paper proposes a general framework for self-optimizing parallel storage systems, with the goal of continuously providing a high-degree of I/O parallelism that is robust to changes in the parallel access patterns of applications and the coexistence of applications with different parallel access characteristics. Evaluation results indicate that the proposed framework is highly successful in adjusting to skewed parallel access patterns for both traditional hard disk drive (HDD) based storage arrays and solid-state drive (SSD) based all-flash arrays. In addition to the storage arrays, the proposed framework is sufficiently generic to be tailored to various other parallel storage scenarios including but not limited to key-value stores, parallel/distributed file systems, and internal parallelism of SSDs. Everett Neil Rush, Bryan Harris, Nihat Altiparmak, Ali Saman Tosun |
HiPC | 3 |
| 2016 | Exploiting Replication for Energy Efficiency of Heterogeneous Storage SystemsabstractAs a result of immense growth of digital data in the last decade, energy consumption has become an important issue in data storage systems. In the US alone, data centers were projected to consume $4 billion (40 TWh) yearly electricity in 2005. This cost had reached to $10 billion (100 TWh) in 2011, and expected to be around $20 billion (200 TWh) in 2016 by doubling itself every 5 years. In addition to the economic burden on companies and research institutions, these large scale data storage systems also have a negative impact on the environment. According to the EPA, generating 1 KWh of electricity in the US results in an average of 1.55 pounds of carbon dioxide emissions. Considering a projected 200 TWh energy requirement for 2016, energy-efficient data storage systems can have a huge economic and environmental impacts on society. This project exploits replication and heterogeneity existing in modern multi-disk storage systems and proposes an energy-efficient and performance-aware replica selection technique to reduce the energy consumption of data storage systems without negatively affecting their performance. Our proposed technique exploits the difference between active and idle energy consumption in heterogeneous disks holding the same replica and selects replicas by balancing energy and performance. Everett Neil Rush, Nihat Altiparmak |
MASCOTS | 2 |
| 2016 | Multithreaded Maximum Flow Based Optimal Replica Selection Algorithm for Heterogeneous Storage ArchitecturesabstractEfficient retrieval of replicated data from multiple disks is a challenging problem, especially for heterogeneous storage architectures. Recently, maximum flow based optimal replica selection algorithms were proposed guaranteeing the minimum retrieval time in heterogeneous environments. Although optimality of the retrieval schedule is an important property, execution time of the replica selection algorithm is also crucial since it might significantly affect the performance of the storage sub-system. Current replica selection mechanisms achieve the optimal response time retrieval schedule by performing multiple runs of a maximum flow algorithm in a black-box manner. Such black-box usage of a maximum flow algorithm results in unnecessary flow calculations since previously calculated flow values cannot be conserved. In addition, most new generation multi-disk storage architectures are powered with multi-core processors motivating the usage of multithreaded replica selection algorithms. In this paper, we propose multithreaded and integrated maximum flow based optimal replica selection algorithms handling heterogeneous storage architectures. Proposed and existing algorithms are evaluated using various homogeneous and heterogeneous multi-disk storage architectures. Experimental results show that proposed sequential integrated algorithm achieves 5X speed-up in homogeneous systems, and proposed multithreaded integrated algorithm achieves 21X speed-up using 16 threads in heterogeneous systems over the existing sequential black-box algorithm. Nihat Altiparmak, Ali Saman Tosun |
IEEE Trans. Computers | 1 |
| 2014 | Continuous Retrieval of Replicated Data from Heterogeneous Storage ArraysabstractReplicated declustering techniques reduce response times of disk requests by distributing data among multiple disks and retrieving in parallel. Efficient retrieval of replicated data from multiple disks is a challenging problem, especially for heterogeneous storage architectures receiving continuous disk requests. Existing techniques either do not consider the heterogeneity of the disks or handle the requests in a discrete manner assuming the storage system is idle. In this paper, we focus on continuous retrieval techniques in heterogeneous storage architectures to minimize the response time of disk requests considering waiting time and service time of the requests as well as the execution time of the retrieval algorithm. We investigate multiple trade-offs between these three factors affecting the response time of a disk request and propose a maximum flow based adaptive retrieval strategy. Performance of the proposed and existing continuous retrieval techniques are evaluated using simulations driven by real world traces and various homogeneous and heterogeneous multi-disk storage configurations. Nihat Altiparmak, Ali Saman Tosun |
MASCOTS | 1 |
| 2014 | Low-Cost Indoor Location Management for Robots Using IR Leds and an IR CameraabstractMany applications in wireless sensor networks can benefit from position information. However, existing accurate solutions for indoor environments are costly. Radio-Frequency (RF)-based approaches are not suitable for some indoor environments such as factory floors where heavy machinery can cause interference. We propose a low-cost and simple location management system using infrared (IR) leds and the Wii Remote Controller (WRC) which has an IR camera. The proposed solution is motivated by the need to find the location of a mobile robot used for data collection in a wireless sensor network. In the proposed schemes, the WRC is placed vertically on the mobile robot pointing upward and IR leds are placed irregularly on the ceiling. The mobile robot determines its position using the relative positions of the IR leds detected by the WRC. The WRC senses a few IR leds at a time, and they are differentiated using the irregularity among them. We analyze the problem theoretically and show that there exist limitations for covering large areas. We also discuss how to overcome these limitations. For small coverage areas, we provide optimal solutions using linear programming. The proposed scheme uses the resources efficiently and can cover a large area using a single WRC and multiple IR leds. We have simulation results including nonvertical placements of the WRC. The proposed scheme is easy to implement and requires minimal bandwidth for location management. Baris Tas, Nihat Altiparmak, Ali Saman Tosun |
ACM Trans. Sens. Networks | 2 |
| 2013 | Generalized Optimal Response Time Retrieval of Replicated Data from Storage ArraysabstractDeclustering techniques reduce query response times through parallel I/O by distributing data among parallel disks. Recently, replication-based approaches were proposed to further reduce the response time. Efficient retrieval of replicated data from multiple disks is a challenging problem. Existing retrieval techniques are designed for storage arrays with identical disks, having no initial load or network delay. In this article, we consider the generalized retrieval problem of replicated data where the disks in the system might be heterogeneous, the disks may have initial load, and the storage arrays might be located on different sites. We first formulate the generalized retrieval problem using a Linear Programming (LP) model and solve it with mixed integer programming techniques. Next, the generalized retrieval problem is formulated as a more efficient maximum flow problem. We prove that the retrieval schedule returned by the maximum flow technique yields the optimal response time and this result matches the LP solution. We also propose a low-complexity online algorithm for the generalized retrieval problem by not guaranteeing the optimality of the result. Performance of proposed and state of the art retrieval strategies are investigated using various replication schemes, query types, query loads, disk specifications, network delays, and initial loads. Nihat Altiparmak, Ali Saman Tosun |
ACM Trans. Storage | 1 |
| 2012 | Replication Based QoS Framework for Flash ArraysabstractThe increasing popularity of the storage cloud is leading organizations to move their applications and enterprise data into the cloud. It is desirable to move time-critical applications demanding high performance I/O operations. Flash based storage arrays have emerged to address the high performance I/O requirements, however, providing predictable Quality of Service (QoS) for applications with real time data requirements is a challenging open problem. This paper introduces a QoS framework for flash based storage arrays. Our framework provides deterministic and statistical response time guarantees through a combination of techniques including replication, data mining, and online retrieval. We evaluated the framework using synthetic and real-world traces. The QoS performance of the system is compared to the existing high-throughput RAID designs. Numerical results show that under the synthetic traces, QoS performance of the proposed system outperforms the existing high performance RAID designs. Real world traces indicate that the proposed QoS mechanism is tunable to support the guarantees required by various real world applications. Nihat Altiparmak, Ali Saman Tosun |
CLUSTER | 1 |
| 2012 | Integrated Maximum Flow Algorithm for Optimal Response Time Retrieval of Replicated DataabstractEfficient retrieval of replicated data from multiple disks is a challenging problem. Traditional retrieval techniques assume that replication is done at a single site using homogeneous disk arrays having no initial load or network delay. Recently, generalized retrieval algorithms are proposed to cover heterogeneous disk arrays, initial loads, and network delays. Generalized retrieval algorithms achieve the optimal response time retrieval schedule by performing multiple runs of a maximum flow algorithm. Since the maximum flow algorithm is used as a black box technique, flow values of the previous runs cannot be conserved to speed up the process. In this paper, we propose integrated maximum flow algorithms for the generalized optimal response time retrieval problem. Our first algorithm uses Ford-Fulkerson method and the second algorithm uses Push-relabel algorithm. Besides the sequential implementations, a multi-threaded version of the push-relabel algorithm is also implemented. Proposed algorithms are investigated using various replication schemes, query types, query loads, disk specifications, and system delays. Experimental results show that the sequential integrated push-relabel algorithm runs up to 2.5X faster than the black box version. Furthermore, parallel integrated push-relabel implementation achieves up to 1.7X speed up (~1.2X on average) over the sequential algorithm using two threads, which makes the integrated algorithm up to 4.25X (~3X on average) faster than its black box counterpart. Nihat Altiparmak, Ali Saman Tosun |
ICPP | 1 |
| 2012 | Equivalent Disk AllocationsabstractDeclustering techniques reduce query response times through parallel I/O by distributing data among multiple devices. Except for a few cases, it is not possible to find declustering schemes that are optimal for all spatial range queries. As a result of this, most of the research on declustering have focused on finding schemes with low worst case additive error. Number-theoretic declustering techniques provide low additive error and high threshold. In this paper, we investigate equivalent disk allocations and focus on number-theoretic declustering. Most of the number-theoretic disk allocations are equivalent and provide the same additive error and threshold. Investigation of equivalent allocations simplifies schemes to find allocations with desirable properties. By keeping one of the equivalent disk allocations, we can reduce the complexity of searching for good disk allocations under various criteria such as additive error and threshold. Using proposed scheme, we were able to collect the most extensive experimental results on additive error and threshold in 2, 3, and 4 dimensions. Nihat Altiparmak, Ali Saman Tosun |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2011 | Approximating the Number of Active Nodes Behind a NAT DeviceabstractNetwork Address Translation (NAT) is used for various reasons on the Internet and hides the IP address and number of nodes behind the NAT device. Although many applications benefit from the knowledge of number of active nodes behind a NAT device, existing schemes are limited. In this paper, we use TCP timestamp option to count the number of active nodes. Timestamp option includes current timestamp of the machine in the TCP packet. We propose an efficient scheme that counts the number of machines approximately using clustering of timestamps. We use least-squares line fit of timestamp values and convex hulls to efficiently maintain the crucial information about existing clusters. Proposed scheme is online and requires minimal resources. We have investigated various aspects of the scheme to improve its performance. Using a developed tool to send packets, we have observed that the proposed scheme approximates the number of machines that send more than threshold number of packets well. Real experiments validate the proposed scheme. Ali Tekeoglu, Nihat Altiparmak, Ali Saman Tosun |
ICCCN | 2 |
| 2011 | DoS resilience of real time streaming protocolabstractDenial of Service (DoS) attacks on a computer system or network cause loss of service to users typically by flooding a victim with many requests or by disrupting the connections between two machines. Although significant amount of work has been done on DoS attacks, DoS attacks on streaming video servers were not investigated in detail. In this paper, we investigate DoS resilience of Real Time Streaming Protocol (RTSP). We show that by using a simple command line tool that opens a large number of RTSP connections, we can launch DoS attacks on the server and the proxy. We discuss in detail how the CPU and the memory resources are affected by the attacks. We observe that, with the DoS attack we launch, clients can also keep the connections alive for a long period and maintain the resources allocated at the server. We propose a lightweight dynamic detection framework for the RTSP based DoS attacks. Nihat Altiparmak, Ali Tekeoglu, Ali Saman Tosun |
IPCCC | 1 |
| 2009 | Low cost indoor location management system using infrared leds and Wii Remote ControllerabstractMany applications in wireless sensor networks can benefit from position information. However, existing accurate solutions for indoor environments are costly. RF based approaches are not suitable for some indoor environments such as factory floors where heavy machinery can cause interference. In this paper, we propose a low cost and simple location management system using the Wii remote controller and infrared leds. Proposed solution is motivated by the need to find the location of a mobile robot used for data collection in a wireless sensor network. In proposed scheme, Wii remote controller is placed on the mobile robot pointing upward and several IR leds are placed on the ceiling. Proposed scheme uses the resources efficiently and can cover a large area using a single Wii remote controller and multiple IR leds. Proposed scheme is easy to implement and requires minimal bandwith for location management. Baris Tas, Nihat Altiparmak, Ali Saman Tosun |
IPCCC | 2 |