Samer Al-Kiswany

dblp:67/4560 · DBLP profile ↗
← Back
40ranked-venue papers
10as first author
13since 2021 · last 2026
0000-0002-6429-9983ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 23 · 9 first-author · 6 since 2021Software engineering, systems software and programming languages · 7 · 2 since 2021Computer networks · 4 · 2 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 DROPS: Managing Serverless Resource Pools in Microsoft Azure Functions
abstract
Azure Functions maintains pools of pre-warmed containers to avoid the high container-allocation latency. The size of a pool is important: a pool that is too small leads to high allocation latency, whereas a pool that is too large wastes resources and increases cost. Service providers typically oversize pools to meet service-level objectives (SLOs). Our findings indicate that the cost of maintaining pre-warmed container pools dominates the overall platform cost, motivating the need for effective pool management strategies.
Ahmed Alquraan, Abdelrahman Baba, Rafael Mendes da Silva, Sameh Elnikety, Paul Batum, Hamid Henry Safi, Seth Fine, Samer Al-Kiswany
EuroSys9
2026 Vectorized Sequence-Based Chunking for Data Deduplication
abstract
Data deduplication has gained wide acclaim as a mechanism to improve storage efficiency and conserve network bandwidth. Its most critical phase, data chunking, is responsible for the overall space savings achieved via the deduplication process. However, modern data chunking algorithms are slow and compute-intensive because they scan large amounts of data while simultaneously making data-driven boundary decisions. We present SeqCDC, a novel chunking algorithm that leverages lightweight boundary detection, content-defined skipping, and SSE/AVX acceleration to improve chunking throughput for large chunk sizes. Our evaluation shows that SeqCDC achieves 10× higher throughput than unaccelerated and 1.2×- 1.35× higher throughput than vector-accelerated data chunking algorithms while minimally affecting deduplication space savings.
Sreeharsha Udayashankar, Ali Assem Mahmoud, Samer Al-Kiswany
IEEE Trans. Parallel Distributed Syst.3
2025 VectorCDC: Accelerating Data Deduplication with Vector Instructions
Sreeharsha Udayashankar, Abdelrahman Baba, Samer Al-Kiswany
FAST3
2025 Measuring the Runtime Performance of C++ Code Written by Humans Using Github Copilot
abstract
GitHub Copilot is an artificially intelligent programming assistant used by many developers. While a few studies have evaluated the security risks of using Copilot, there has not been any study to show if it aids developers in producing code with better runtime performance. We evaluate the runtime performance of C++ code produced when developers use GitHub Copilot versus when they do not. To this end, we conducted a user study with 32 participants where each participant solved two C++ programming problems, one with Copilot and the other without it and measured the runtime performance of the participants' solutions on our test data. Our results suggest that using Copilot may produce$\mathbf{C + +}$code with (statistically significant) slower runtime performance.
Daniel Erhabor, Sreeharsha Udayashankar, Meiyappan Nagappan, Samer Al-Kiswany
ICSE4
2024 The Impact of Low-Entropy on Chunking Techniques for Data Deduplication
abstract
While numerous Content-Defined Chunking (CDC) algorithms exist for data deduplication, their relative performance has not been analyzed in the presence of low-entropy induced byte-shifting. This paper explores and evaluates hash-based and hashless CDC algorithms in the presence of low-entropy data regions, using synthetic datasets. Our evaluation shows that modern CDC algorithms are poor at handling low-entropy blocks when the block sizes are small and that their low-entropy detection ability depends upon the expected average chunk size. Contrary to previous studies focusing on conventional byte-shifting, hash-based algorithms achieve poor space savings compared to their hashless counterparts when low-entropy induced byte-shifting is involved. This can be explained by the greater variability in chunk sizes and the higher percentage of artificial boundaries they exhibit in the presence of these regions. All of these factors together highlight the need for specialized CDC algorithms to detect and eliminate low-entropy data blocks during the deduplication process.
Mu'men Al Jarah, Sreeharsha Udayashankar, Abdelrahman Baba, Samer Al-Kiswany
CLOUD4
2024 Draconis: Network-Accelerated Scheduling for Microsecond-Scale Workloads
abstract
We present Draconis, a novel scheduler for workloads in the range of tens to hundreds of microseconds. Draconis challenges the popular belief that programmable switches cannot house the complex data structures, such as queues, needed to support an in-network scheduler. Using programmable switches, Draconis achieves the low scheduling tail latency and high throughput needed to support these microsecond-scale workloads on large clusters. Furthermore, Draconis supports a wide range of complex scheduling policies, including locality-aware scheduling, priority-based scheduling, and resource-based scheduling.
Sreeharsha Udayashankar, Ashraf Abdel-Hadi, Ali José Mashtizadeh, Samer Al-Kiswany
EuroSys4
2024 Slicify: Fault Injection Testing for Network Partitions
abstract
Modern distributed systems are complex. They include hundreds of components that implement complex protocols such as scheduling, replication, and access control. These systems are expected to offer high availability and preserve their data even in the face of external environmental faults. Testing is the primary approach for improving system reliability. Testing against environmental faults such as hardware failures, memory corruption, and network problems is complicated since they can happen at any step in the protocol and affect any component.We present Slicify, a generic framework to test the network partition resilience of distributed systems. Slicify injects network partitions during unit tests to analyze system behavior in their presence. Slicify reduces the test space in an application-agnostic fashion with its novel connection tracking mechanism. We verify Slicify’s capabilities by reproducing previously documented failures in two production systems. In addition, we demonstrate its effectiveness by uncovering new failures in three popular distributed systems.
Seba Khaleel, Sreeharsha Udayashankar, Samer Al-Kiswany
MASCOTS3
2024 SeqCDC: Hashless Content-Defined Chunking for Data Deduplication
abstract
Data deduplication is critical to cloud storage providers and is widely employed to conserve server-side storage space. Data chunking is an important aspect of deduplication, being directly responsible for storage space savings and end-to-end system throughput. While deduplication systems deployed in production favor larger chunk sizes, existing data chunking algorithms are slow and offer minimal throughput increases with increasing chunk size.
Sreeharsha Udayashankar, Abdelrahman Baba, Samer Al-Kiswany
Middleware3
2024 LoLKV: The Logless, Linearizable, RDMA-based Key-Value Storage System
Ahmed Alquraan, Sreeharsha Udayashankar, Virendra J. Marathe, Bernard Wong 0001, Samer Al-Kiswany
NSDI5
2023 CASPR: Connectivity-Aware Scheduling for Partition Resilience
abstract
We present a comprehensive empirical study of the impact partial network partitions have on cluster managers in data analysis frameworks. Our study shows that modern scheduling approaches are vulnerable to partial network partitions. Partial partitions can lead to a complete cluster pause or a significant loss of performance. To overcome the shortcomings of the state-of-the-art sched-ulers, we design CASPR, a connectivity-aware scheduler. CASPR incorporates the current network connectivity information when making scheduling decisions to allocate fully connected nodes for a given application. CASPR effectively hides partial partitions from applications. Our evaluation of a CASPR prototype shows that it can tolerate partial network partitions, as well as eliminate application halting or significant loss of performance.
Sara Qunaibi, Sreeharsha Udayashankar, Samer Al-Kiswany
SRDS3
2023 Partial Network Partitioning
abstract
We present an extensive study focused on partial network partitioning. Partial network partitions disrupt the communication between some but not all nodes in a cluster. First, we conduct a comprehensive study of system failures caused by this fault in 13 popular systems. Our study reveals that the studied failures are catastrophic (e.g., lead to data loss), easily manifest, and are mainly due to design flaws. Our analysis identifies vulnerabilities in core systems mechanisms including scheduling, membership management, and ZooKeeper-based configuration management. Second, we dissect the design of nine popular systems and identify four principled approaches for tolerating partial partitions. Unfortunately, our analysis shows that implemented fault tolerance techniques are inadequate for modern systems; they either patch a particular mechanism or lead to a complete cluster shutdown, even when alternative network paths exist. Finally, our findings motivate us to build Nifty, a transparent communication layer that masks partial network partitions. Nifty builds an overlay between nodes to detour packets around partial partitions. Nifty provides an approach for applications to optimize their operation during a partial partition. We demonstrate the benefit of this approach through integrating Nifty with VoltDB, HDFS, and Kafka.
Basil Alkhatib, Sreeharsha Udayashankar, Sara Qunaibi, Ahmed Alquraan, Mohammed Alfatafta, Wael Al-Manasrah, Alex Depoutovitch, Samer Al-Kiswany
ACM Trans. Comput. Syst.8
2022 OrcBench: A Representative Serverless Benchmark
abstract
Serverless computing is rapidly growing area of research. No standardized benchmark currently exists for evaluating orchestration level decisions or executing large serverless workloads because of the limited data provided by cloud providers. Current benchmarks focus on other aspects, such as the cost of running general types of functions and their runtimes.We introduce OrcBench, the first orchestration benchmark based on the recently published Microsoft Azure serverless data set. OrcBench categorizes 8622 serverless functions into 17 distinct models, which represent 5.6 million invocations from the original trace.OrcBench also incorporates a time-series analysis to identify function chains within the dataset. OrcBench can use these to create workloads that mimic complete serverless applications, which includes simulating CPU and memory usage. The modeling allows these workloads to be scaled according to the target hardware configuration.
Ryan Hancock, Sreeharsha Udayashankar, Ali José Mashtizadeh, Samer Al-Kiswany
CLOUD4
2022 Accelerating Reads With In-Network Consistency-Aware Load Balancing
abstract
We present FLAIR, a novel approach for accelerating read operations in leader-based consensus protocols. FLAIR leverages the capabilities of the new generation of programmable switches to serve reads from follower replicas without compromising consistency. The core of the new approach is a packet-processing pipeline that can track client requests and system replies, identify consistent replicas, and at line speed, forward read requests to replicas that can serve the read without sacrificing linearizability. An additional benefit of FLAIR is that it facilitates devising novel consistency-aware load balancing techniques. Following the new approach, we designed FlairKV, a key-value store atop Raft. FlairKV implements the processing pipeline using the P4 programming language. We evaluate the benefits of the proposed approach and compare it to previous approaches using a cluster with a Barefoot Tofino switch. Our evaluation indicates that, compared to state-of-the-art alternatives, the proposed approach can bring significant performance gains: up to 42% higher throughput and 35–97% lower latency for most workloads. Furthermore, our evaluation shows that our novel load balancing techniques can cope with heterogeneous load and hardware to achieve higher performance, and that FLAIR can scale to support large data sets and clusters.
Ibrahim Kettaneh, Ahmed Alquraan, Hatem Takruri, Ali José Mashtizadeh, Samer Al-Kiswany
IEEE/ACM Trans. Netw.5
2020 FLAIR: Accelerating Reads with Consistency-Aware Network Routing
Hatem Takruri, Ibrahim Kettaneh, Ahmed Alquraan, Samer Al-Kiswany
NSDI4
2020 Toward a Generic Fault Tolerance Technique for Partial Network Partitioning
Mohammed Alfatafta, Basil Alkhatib, Ahmed Alquraan, Samer Al-Kiswany
OSDI4
2020 Scalable, NearZero Loss Disaster Recovery for Distributed Data Stores
abstract
This paper presents a new Disaster Recovery (DR) system, called Slogger, that differs from prior works in two principle ways: (i) Slogger enables DR for a linearizable distributed data store, and (ii) Slogger adopts the continuous backup approach that strives to maintain a tiny lag on the backup site relative to the primary site, thereby restricting the data loss window, due to disasters, to milliseconds. These goals pose a significant set of challenges related to consistency of the backup site's state, failures, and scalability. Slogger employs a combination of asynchronous log replication, intra-data center synchronized clocks, pipelining, batching, and a novel watermark service to address these challenges. Furthermore, Slogger is designed to be deployable as an "add-on" module in an existing distributed data store with few modifications to the original code base. Our evaluation, conducted on Slogger extensions to a 32-sharded version of LogCabin, an open source key-value store, shows that Slogger maintains a very small data loss window of 14.2 milliseconds which is near the optimal value in our evaluation setup. Moreover, Slogger reduces the length of the data loss window by 50% compared to incremental snapshotting technique without having any performance penalty on the primary data store. Furthermore, our experiments demonstrate that Slogger achieves our other goals of scalability, fault tolerance, and efficient failover to the backup data store when a disaster is declared at the primary data store.
Ahmed Alquraan, Alex Kogan, Virendra J. Marathe, Samer Al-Kiswany
Proc. VLDB Endow.4
2020 The Network-Integrated Storage System
abstract
We present NICE, a key-value storage system design that leverages new software-defined network capabilities to build cluster-based network-efficient storage system. NICE presents novel techniques to co-design network routing and multicast with storage replication, consistency, and load balancing to achieve higher efficiency, performance, and scalability. We implement the NICEKV prototype. NICEKV follows the NICE approach in designing four essential network-centric storage mechanisms: request routing, replication, consistency, and load balancing. Our evaluation shows that the proposed approach brings significant performance gains compared with the current systems design: up to 7× put/get performance improvement, up to 2× reduction in network load, 3× to 9× load reduction on the storage nodes, and the elimination of scalability bottlenecks present in current designs.
Ibrahim Kettaneh, Ahmed Alquraan, Hatem Takruri, Suli Yang, Andrea C. Arpaci-Dusseau, Remzi H. Arpaci-Dusseau, Samer Al-Kiswany
IEEE Trans. Parallel Distributed Syst.7
2018 COOL: A Cloud-Optimized Structure for MPI Collective Operations
abstract
We present COOL, a simple and generic structure for MPI collective operations. COOL enables highly efficient designs for all collective operations in the cloud. We then present a system design based on COOL that implements frequently used collective operations. Our design efficiently uses the intra-rack network while minimizing cross-rack communication, thus improving the application performance and scalability. We use recent software-defined networking capabilities to build optimal network paths for I/O intensive collective operations. Our analytical evaluation shows that our design imposes the least possible network overhead across racks. Furthermore, when compared with OpenMPI and MPICH, our design reduces the number of steps to only three, decreases the number of exchanged messages by a factor of N, the total number of processes, and reduces the network load by up to an order of magnitude. These significant improvements come at the cost of a modest increase in the computation load on a few processes.
Mohammed Alfatafta, Zuhair AlSader, Samer Al-Kiswany
IEEE CLOUD3
2018 An Analysis of Network-Partitioning Failures in Cloud Systems
Ahmed Alquraan, Hatem Takruri, Mohammed Alfatafta, Samer Al-Kiswany
OSDI4
2017 NICE: Network-Integrated Cluster-Efficient Storage
abstract
We present NICE, a key-value storage system design that leverages new software-defined network capabilities to build cluster-based network-efficient storage system. NICE presents novel techniques to co-design network routing and multicast with storage replication, consistency, and load balancing to achieve higher efficiency, performance, and scalability. We implement the NICEKV prototype. NICEKV follows the NICE approach in designing four essential network-centric storage mechanisms: request routing, replication, consistency, and load balancing. Our evaluation shows that the proposed approach brings significant performance gains compared to the current key-value systems design: up to 7× put/get performance improvement, up to 2× reduction in network load, 3× to 9× load reduction on the storage nodes, and the elimination of scalability bottlenecks present in current designs.
Samer Al-Kiswany, Suli Yang, Andrea C. Arpaci-Dusseau, Remzi H. Arpaci-Dusseau
HPDC1
2017 A cross-layer optimized storage system for workflow applications
Samer Al-Kiswany, Lauro Beltrão Costa, Hao Yang 0039, Emalayan Vairavanathan, Matei Ripeanu
Future Gener. Comput. Syst.1
2016 A Software-Defined Storage for Workflow Applications
abstract
We present a software-defined storage architecture with two key properties: firstly, it enables external control of storage operations, and, secondly, it allows extending the storage system with new, workload specific, optimizations, all without breaking the file system abstractions. We argue that this architecture is generic, we prototype FlexStore following this architecture, we instantiate it to support workflow applications, and we report on our preliminary experience.
Samer Al-Kiswany, Matei Ripeanu
CLUSTER1
2016 Support for Provisioning and Configuration Decisions for Data Intensive Workflows
abstract
System provisioning, resource allocation, and configuration decisions for I/O-intensive workflow applications are complex even for expert users. Users face choices at multiple levels: allocating resources to individual sub-systems (e.g., the application layer, the storage layer) as well as configuring each of these optimally (e.g., replication level, chunk size, caching policies in case of storage) all having a large impact on the overall application performance. This paper presents a solution to address the problem of supporting these provisioning, allocation and configuration decisions for workflow applications. To enable selecting a good choice in a reasonable time, we propose an approach that accelerates the exploration of the configuration space based on a low-cost performance predictor that estimates total execution time of a workflow application in a given setup. We evaluate the predictor in a number of different scenarios including the Montage application: a workflow composed of over 7,500 tasks structured in 10 different stages with varying characteristics. Our evaluation shows that: (i) the predictor is effective in identifying the desired system configuration, (ii) it can scale to model a complex workflow application run on a 100-node cluster, while (iii) using orders of magnitude less resources than running the actual application. Additionally, we extend the predictor to estimate the energy usage of the system, and we present our experience with incorporating it in the development process of a distributed storage system.
Lauro Beltrão Costa, Samer Al-Kiswany, Matei Ripeanu, Hao Yang 0039
IEEE Trans. Parallel Distributed Syst.2
2015 Split-level I/O scheduling
abstract
We introduce split-level I/O scheduling, a new framework that splits I/O scheduling logic across handlers at three layers of the storage stack: block, system call, and page cache. We demonstrate that traditional block-level I/O schedulers are unable to meet throughput, latency, and isolation goals. By utilizing the split-level framework, we build a variety of novel schedulers to readily achieve these goals: our Actually Fair Queuing scheduler reduces priority-misallocation by 28x; our Split-Deadline scheduler reduces tail latencies by 4x; our Split-Token scheduler reduces sensitivity to interference by 6x. We show that the framework is general and operates correctly with disparate file systems (ext4 and XFS). Finally, we demonstrate that split-level scheduling serves as a useful foundation for databases (SQLite and PostgreSQL), hypervisors (QEMU), and distributed file systems (HDFS), delivering improved isolation and performance in these important application scenarios.
Suli Yang, Tyler Caraza-Harter, Nishant Agrawal, Salini Selvaraj Kowsalya, Anand Krishnamurthy, Samer Al-Kiswany, Rini T. Kaushik, Andrea C. Arpaci-Dusseau, Remzi H. Arpaci-Dusseau
SOSP6
2015 The Case for Workflow-Aware Storage: An Opportunity Study
Lauro Beltrão Costa, Hao Yang 0039, Emalayan Vairavanathan, Abmar Barros, Ketan Maheshwari, Gilles Fedak, Daniel S. Katz, Michael Wilde, Matei Ripeanu, Samer Al-Kiswany
J. Grid Comput.10
2014 Supporting storage configuration for I/O intensive workflows
abstract
System provisioning, resource allocation, and system configuration decisions for I/O-intensive workflow applications are complex even for expert users. Users face choices at multiple levels: allocating resources to individual sub-systems (e.g., the application layer, the storage layer) and configuring each of these optimally (e.g., replication level, chunk size, caching policies in case of storage) all having a large impact on overall application performance. This paper presents our progress on addressing the problem of supporting these provisioning, allocation and configuration decisions for workflow applications. To enable selecting a good choice in a reasonable time, we propose an approach that accelerates the exploration of the configuration space based on a low-cost performance predictor that estimates total execution time of a workflow application in a given setup. Our evaluation shows that: (i) the predictor is effective in identifying the desired system configuration, (ii) it can scale to model a workflow application run on an entire cluster, while (iii) using over 2000x less resources (machines x time) than running the actual application.
Lauro Beltrão Costa, Samer Al-Kiswany, Hao Yang 0039, Matei Ripeanu
ICS2
2014 Physical Disentanglement in a Container-Based File System
Lanyue Lu, Thanh Do, Samer Al-Kiswany, Andrea C. Arpaci-Dusseau, Remzi H. Arpaci-Dusseau
OSDI4
2014 All File Systems Are Not Created Equal: On the Complexity of Crafting Crash-Consistent Applications
Thanumalayan Sankaranarayana Pillai, Vijay Chidambaram, Ramnatthan Alagappan, Samer Al-Kiswany, Andrea C. Arpaci-Dusseau, Remzi H. Arpaci-Dusseau
OSDI4
2013 Cost exploration of data sharings in the cloud
abstract
Enabling data sharing among mobile apps hosted in the same cloud infrastructure can provide a competitive advantage to the mobile apps by giving them access to rich information as well as increasing the revenue for the cloud provider. We introduce a costing tool that allows application owners (i.e., consumers) and the cloud service provider to assess the cost of a desired data sharing. The costing tool enables the consumers to effectively explore the cost space by choosing between alternative configurations of varying data qualities, specified by the staleness and the accuracy of the data sharing. In other words, staleness and accuracy requirements on the data sharing are used as levers for controlling costs. These capabilities are implemented in a What-if analysis tool, which has been integrated with a large data-sharing platform. We conducted extensive experiments on the integrated platform with a sharing ecosystem created around Twitter data and show the effectiveness of the results produced by the What-if tool.
Samer Al-Kiswany, Hakan Hacigümüs, Ziyang Liu 0001, Jagan Sankaranarayanan
EDBT1
2013 GPUs as Storage System Accelerators
abstract
Massively multicore processors, such as graphics processing units (GPUs), provide, at a comparable price, a one order of magnitude higher peak performance than traditional CPUs. This drop in the cost of computation, as any order-of-magnitude drop in the cost per unit of performance for a class of system components, triggers the opportunity to redesign systems and to explore new ways to engineer them to recalibrate the cost-to-performance relation. This project explores the feasibility of harnessing GPUs' computational power to improve the performance, reliability, or security of distributed storage systems. In this context, we present the design of a storage system prototype that uses GPU offloading to accelerate a number of computationally intensive primitives based on hashing, and introduce techniques to efficiently leverage the processing power of GPUs. We evaluate the performance of this prototype under two configurations: as a content addressable storage system that facilitates online similarity detection between successive versions of the same file and as a traditional system that uses hashing to preserve data integrity. Further, we evaluate the impact of offloading to the GPU on competing applications' performance. Our results show that this technique can bring tangible performance gains without negatively impacting the performance of concurrently running applications.
Samer Al-Kiswany, Abdullah Gharaibeh, Matei Ripeanu
IEEE Trans. Parallel Distributed Syst.1
2012 A Workflow-Aware Storage System: An Opportunity Study
abstract
This paper evaluates the potential gains a workflow-aware storage system can bring. Two observations make us believe such storage system is crucial to efficiently support workflow-based applications: First, workflows generate irregular and application-dependent data access patterns. These patterns render existing storage systems unable to harness all optimization opportunities as this often requires conflicting optimization options or even conflicting design decision at the level of the storage system. Second, when scheduling, workflow runtime engines make suboptimal decisions as they lack detailed data location information. This paper discusses the feasibility, and evaluates the potential performance benefits brought by, building a workflow-aware storage system that supports per-file access optimizations and exposes data location. To this end, this paper presents approaches to determine the application-specific data access patterns, and evaluates experimentally the performance gains of a workflow-aware storage approach. Our evaluation using synthetic benchmarks shows that a workflow-aware storage system can bring significant performance gains: up to 7× performance gain compared to the distributed storage system - MosaStore and up to 16× compared to a central, well provisioned, NFS server.
Emalayan Vairavanathan, Samer Al-Kiswany, Lauro Beltrão Costa, Zhao Zhang 0007, Daniel S. Katz, Michael Wilde, Matei Ripeanu
CCGRID2
2011 VMFlock: virtual machine co-migration for the cloud
abstract
This paper presents VMFlockMS, a migration service optimized for cross-datacenter transfer and instantiation of groups of virtual machine (VM) images that comprise an application-level solution (e.g., a three-tier web application). We dub these groups of related VM images VMFlocks. VMFlockMS employs two main techniques: first, data deduplication within the VMFlock to be migrated and between the VMFlock and the data already present at the destination datacenter, and, second, accelerated instantiation of the application at the target datacenter after transferring only a partial set of data blocks and prioritization of the remaining data based on previously observed access patterns originating from the running VMs. VMFlockMS is designed to be deployed as a set of virtual appliances which make efficient use of the available cloud resources to locally access and deduplicate the images and data in a distributed fashion with minimal requirements imposed on the cloud API to access the VM image repository. VMFlockMS provides an incrementally scalable and high-performance migration service. Our evaluation shows that VMFlockMS can reduce the data volumes to be transferred over the network to as low as 3% of the original VMFlock size, enables the complete transfer of the VM images belonging to a VMFlock over transcontinental link up to 3.5x faster than alternative approaches, and enables booting these VM images with as little as 5% of the compressed VMFlock data available at the destination.
Samer Al-Kiswany, Dinesh Subhraveti, Prasenjit Sarkar, Matei Ripeanu
HPDC1
2011 ThriftStore: Finessing Reliability Trade-Offs in Replicated Storage Systems
abstract
This paper explores the feasibility of a storage architecture that offers the reliability and access performance characteristics of a high-end system, yet is cost-efficient. We propose ThriftStore, a storage architecture that integrates two types of components: volatile, aggregated storage and dedicated, yet low-bandwidth durable storage. On the one hand, the durable storage forms a back end that enables the system to restore the data the volatile nodes may lose. On the other hand, the volatile nodes provide a high-throughput front-end. Although integrating these components has the potential to offer a unique combination of high throughput and durability at a low cost, a number of concerns need to be addressed to architect and correctly provision the system. To this end, we develop analytical and simulation-based tools to evaluate the impact of system characteristics (e.g., bandwidth limitations on the durable and the volatile nodes) and design choices (e.g., the replica placement scheme) on data availability and the associated system costs (e.g., maintenance traffic). Moreover, to demonstrate the high-throughput properties of the proposed architecture, we prototype a GridFTP server based on ThriftStore. Our evaluation demonstrates an impressive, up to 800 Mbps transfer throughput for the new GridFTP service.
Abdullah Gharaibeh, Samer Al-Kiswany, Matei Ripeanu
IEEE Trans. Parallel Distributed Syst.2
2010 A GPU accelerated storage system
abstract
Massively multicore processors, like, for example, Graphics Processing Units (GPUs), provide, at a comparable price, a one order of magnitude higher peak performance than traditional CPUs. This drop in the cost of computation, as any order-of-magnitude drop in the cost per unit of performance for a class of system components, triggers the opportunity to redesign systems and to explore new ways to engineer them to recalibrate the cost-to-performance relation.
Abdullah Gharaibeh, Samer Al-Kiswany, Sathish Gopalakrishnan, Matei Ripeanu
HPDC2
2009 GPU support for batch oriented workloads
abstract
This paper explores the ability to use graphics processing units (GPUs) as co-processors to harness the inherent parallelism of batch operations in systems that require high performance. To this end we have chosen bloom filters (space-efficient data structures that support the probabilistic representation of set membership) as the queries these data structures support are often performed in batches. Bloom filters exhibit low computational cost per amount of data, providing a baseline for more complex batch operations. We implemented BloomGPU a library that supports offloading bloom filter support to the GPU and evaluate this library under realistic usage scenarios. By completely offloading Bloom filter operations to the GPU, BloomGPU outperforms an optimized CPU implementation of the bloom filter as the workload becomes larger.
Lauro Beltrão Costa, Samer Al-Kiswany, Matei Ripeanu
IPCCC2
2009 Beyond Music Sharing: An Evaluation of Peer-to-Peer Data Dissemination Techniques in Large Scientific Collaborations
Samer Al-Kiswany, Matei Ripeanu, Adriana Iamnitchi, Sudharshan S. Vazhkudai
J. Grid Comput.1
2008 StoreGPU: exploiting graphics processing units to accelerate distributed storage systems
abstract
Today Graphics Processing Units (GPUs) are a largely underexploited resource on existing desktops and a possible cost-effective enhancement to high-performance systems. To date, most applications that exploit GPUs are specialized scientific applications. Little attention has been paid to harnessing these highly-parallel devices to support more generic functionality at the operating system or middleware level. This study starts from the hypothesis that generic middleware level techniques that improve distributed system reliability or performance (such as content addressing, erasure coding, or data similarity detection) can be significantly accelerated using GPU support.We take a first step towards validating this hypothesis, focusing on distributed storage systems. As a proof of concept, we design StoreGPU, a library that accelerates a number of hashing based primitives popular in distributed storage system implementations. Our evaluation shows that StoreGPU enables up to eight-fold performance gains on synthetic benchmarks as well as on a high-level application: the online similarity detection between large data files.
Samer Al-Kiswany, Abdullah Gharaibeh, Elizeu Santos-Neto, George Yuan, Matei Ripeanu
HPDC1
2008 enabling cross-layer optimizations in storage systems with custom metadata
abstract
Today, several data-storage systems allow applications to create and manage custom metadata to improve data search and navigability in large scale storage systems.
Elizeu Santos-Neto, Samer Al-Kiswany, Nazareno Andrade, Sathish Gopalakrishnan, Matei Ripeanu
HPDC2
2008 stdchk: A Checkpoint Storage System for Desktop Grid Computing
abstract
Checkpointing is an indispensable technique to provide fault tolerance for long-running high-throughput applications like those running on desktop grids. This article argues that a checkpoint storage system, optimized to operate in these environments, can offer multiple benefits: reduce the load on a traditional file system, offer high-performance through specialization, and, finally, optimize data management by taking into account checkpoint application semantics. Such a storage system can present a unifying abstraction to checkpoint operations, while hiding the fact that there are no dedicated resources to store the checkpoint data. We prototype stdchk, a checkpoint storage system that uses scavenged disk space from participating desktops to build a low-cost storage system, offering a traditional file system interface for easy integration with applications. This article presents the stdchk architecture, key performance optimizations, and its support for incremental checkpointing and increased data availability. Our evaluation confirms that the stdchk approach is viable in a desktop grid setting and offers a low cost storage system with desirable performance characteristics: high write throughput as well as reduced storage space and network effort to save checkpoint images.
Samer Al-Kiswany, Matei Ripeanu, Sudharshan S. Vazhkudai, Abdullah Gharaibeh
ICDCS1
2007 Are P2P Data-Dissemination Techniques Viable in Today's Data-Intensive Scientific Collaborations?
Samer Al-Kiswany, Matei Ripeanu, Adriana Iamnitchi, Sudharshan S. Vazhkudai
Euro-Par1