Abdullah Gharaibeh

dblp:64/6159 · DBLP profile ↗
← Back
11ranked-venue papers
8as first author
0since 2021 · last 2014
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 7 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
6 papers
Storage systems · 57% GPUs and heterogeneous computing · 37% Performance modeling and evaluation · 4%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%

Topics — the 13 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
GPUs and heterogeneous computing
GPU computing
0.322013
GPUs as Storage System Accelerators · IEEE Trans. Parallel Distributed Syst. 2013
A GPU accelerated storage system · HPDC 2010
Storage systems › data redundancy
replicated storage
0.222011
ThriftStore: Finessing Reliability Trade-Offs in Replicated Storage Systems · IEEE Trans. Parallel Distributed Syst. 2011
Exploring data reliability tradeoffs in replicated storage systems · HPDC 2009
Storage systems
storage reliability
0.222011
ThriftStore: Finessing Reliability Trade-Offs in Replicated Storage Systems · IEEE Trans. Parallel Distributed Syst. 2011
Exploring data reliability tradeoffs in replicated storage systems · HPDC 2009
GPUs and heterogeneous computing › CPU-GPU heterogeneous computing
GPU offloading
0.212013
GPUs as Storage System Accelerators · IEEE Trans. Parallel Distributed Syst. 2013
Storage systems
storage acceleration
0.212013
GPUs as Storage System Accelerators · IEEE Trans. Parallel Distributed Syst. 2013
Storage systems › storage hierarchy
hybrid storage
0.112011
ThriftStore: Finessing Reliability Trade-Offs in Replicated Storage Systems · IEEE Trans. Parallel Distributed Syst. 2011
GPUs and heterogeneous computing › GPU performance optimization
GPU application optimization
0.112010
Size Matters: Space/Time Tradeoffs to Improve GPGPU Applications Performance · SC 2010
Storage systems
distributed storage
0.112008
StoreGPU: exploiting graphics processing units to accelerate distributed storage systems · HPDC 2008
Storage systems
content-addressable storage
0.012013
GPUs as Storage System Accelerators · IEEE Trans. Parallel Distributed Syst. 2013
Storage systems
similarity detection
0.012013
GPUs as Storage System Accelerators · IEEE Trans. Parallel Distributed Syst. 2013
Bioinformatics and computational biology
sequence alignment
0.012010
Size Matters: Space/Time Tradeoffs to Improve GPGPU Applications Performance · SC 2010
Performance modeling and evaluation
workload characterization
0.012009
Exploring data reliability tradeoffs in replicated storage systems · HPDC 2009
Storage systems › storage reliability
erasure coding
0.012008
StoreGPU: exploiting graphics processing units to accelerate distributed storage systems · HPDC 2008

Methods — techniques the papers use, named apart from their topics

hashing · 0.2space/time tradeoff · 0.2memory representation optimization · 0.2simulation · 0.1analytical modeling · 0.1GPU acceleration · 0.1
YearPublicationVenuePosition
2014 DedupT: Deduplication for tape systems
abstract
Deduplication is a commonly-used technique on disk-based storage pools. However, deduplication has not been used for tape-based pools: tape characteristics, such as high mount and seek times combined with data fragmentation resulting from deduplication create a toxic combination that leads to unacceptably high retrieval times. This work proposes DedupT, a system that efficiently supports deduplication on tape pools. This paper (i) details the main challenges to enable efficient deduplication on tape libraries, (ii) presents a class of solutions based on graph-modeling of similarity between data items that enables efficient placement on tapes; and (iii) presents the design and evaluation of novel cross-tape and on-tape chunk placement algorithms that alleviate tape mount time overhead and reduce on-tape data fragmentation. Using 4.5 TB of real-world workloads, we show that DedupT retains at least 95% of the deduplication efficiency. We show that DedupT mitigates major retrieval time overheads, and, due to reading less data, is able to offer better restore performance compared to the case of restoring non-deduplicated data.
Abdullah Gharaibeh, Cornel Constantinescu, Maohua Lu, Ramani Routray, Prasenjit Sarkar, David Pease, Matei Ripeanu
MSST1
2013 On Graphs, GPUs, and Blind Dating: A Workload to Processor Matchmaking Quest
abstract
Graph processing has gained renewed attention. The increasing large scale and wealth of connected data, such as those accrued by social network applications, demand the design of new techniques and platforms to efficiently derive actionable information from large scale graphs. Hybrid systems that host processing units optimized for both fast sequential processing and bulk processing (e.g., GPUaccelerated systems) have the potential to cope with the heterogeneous structure of real graphs and enable high performance graph processing. Reaching this point, however, poses multiple challenges. The heterogeneity of the processing elements (e.g., GPUs implement a different parallel processing model than CPUs and have much less memory) and the inherent irregularity of graph workloads require careful graph partitioning and load assignment. In particular, the workload generated by a partitioning scheme should match the strength of the processing element the partition is allocated to. This work explores the feasibility and quantifies the performance gains of such low-cost partitioning schemes. We propose to partition the workload between the two types of processing elements based on vertex connectivity. We show that such partitioning schemes offer a simple, yet efficient way to boost the overall performance of the hybrid system. Our evaluation illustrates that processing a 4-billion edges graph on a system with one CPU socket and one GPU, while offloading as little as 25% of the edges to the GPU, achieves 2x performance improvement over state-of-the-art implementations running on a dual-socket symmetric system. Moreover, for the same graph, a hybrid system with dualsocket and dual-GPU is capable of 1.13 Billion breadth-first search traversed edge per second, a performance rate that is competitive with the latest entries in the Graph500 list, yet at a much lower price point.
Abdullah Gharaibeh, Lauro Beltrão Costa, Elizeu Santos-Neto, Matei Ripeanu
IPDPS1
2013 GPUs as Storage System Accelerators
abstract
Massively multicore processors, such as graphics processing units (GPUs), provide, at a comparable price, a one order of magnitude higher peak performance than traditional CPUs. This drop in the cost of computation, as any order-of-magnitude drop in the cost per unit of performance for a class of system components, triggers the opportunity to redesign systems and to explore new ways to engineer them to recalibrate the cost-to-performance relation. This project explores the feasibility of harnessing GPUs' computational power to improve the performance, reliability, or security of distributed storage systems. In this context, we present the design of a storage system prototype that uses GPU offloading to accelerate a number of computationally intensive primitives based on hashing, and introduce techniques to efficiently leverage the processing power of GPUs. We evaluate the performance of this prototype under two configurations: as a content addressable storage system that facilitates online similarity detection between successive versions of the same file and as a traditional system that uses hashing to preserve data integrity. Further, we evaluate the impact of offloading to the GPU on competing applications' performance. Our results show that this technique can bring tangible performance gains without negatively impacting the performance of concurrently running applications.
Samer Al-Kiswany, Abdullah Gharaibeh, Matei Ripeanu
IEEE Trans. Parallel Distributed Syst.2
2012 A yoke of oxen and a thousand chickens for heavy lifting graph processing
abstract
Large, real-world graphs are famously difficult to process efficiently. Not only they have a large memory footprint but most graph processing algorithms entail memory access patterns with poor locality, data-dependent parallelism, and a low compute-to- memory access ratio. Additionally, most real-world graphs have a low diameter and a highly heterogeneous node degree distribution. Partitioning these graphs and simultaneously achieve access locality and load-balancing is difficult if not impossible.
Abdullah Gharaibeh, Lauro Beltrão Costa, Elizeu Santos-Neto, Matei Ripeanu
PACT1
2012 CloudDT: Efficient tape resource management using deduplication in cloud backup and archival services
Abdullah Gharaibeh, Cornel Constantinescu, Maohua Lu, Ramani Routray, Prasenjit Sarkar, David Pease, Matei Ripeanu
CNSM1
2011 ThriftStore: Finessing Reliability Trade-Offs in Replicated Storage Systems
abstract
This paper explores the feasibility of a storage architecture that offers the reliability and access performance characteristics of a high-end system, yet is cost-efficient. We propose ThriftStore, a storage architecture that integrates two types of components: volatile, aggregated storage and dedicated, yet low-bandwidth durable storage. On the one hand, the durable storage forms a back end that enables the system to restore the data the volatile nodes may lose. On the other hand, the volatile nodes provide a high-throughput front-end. Although integrating these components has the potential to offer a unique combination of high throughput and durability at a low cost, a number of concerns need to be addressed to architect and correctly provision the system. To this end, we develop analytical and simulation-based tools to evaluate the impact of system characteristics (e.g., bandwidth limitations on the durable and the volatile nodes) and design choices (e.g., the replica placement scheme) on data availability and the associated system costs (e.g., maintenance traffic). Moreover, to demonstrate the high-throughput properties of the proposed architecture, we prototype a GridFTP server based on ThriftStore. Our evaluation demonstrates an impressive, up to 800 Mbps transfer throughput for the new GridFTP service.
Abdullah Gharaibeh, Samer Al-Kiswany, Matei Ripeanu
IEEE Trans. Parallel Distributed Syst.1
2010 A GPU accelerated storage system
abstract
Massively multicore processors, like, for example, Graphics Processing Units (GPUs), provide, at a comparable price, a one order of magnitude higher peak performance than traditional CPUs. This drop in the cost of computation, as any order-of-magnitude drop in the cost per unit of performance for a class of system components, triggers the opportunity to redesign systems and to explore new ways to engineer them to recalibrate the cost-to-performance relation.
Abdullah Gharaibeh, Samer Al-Kiswany, Sathish Gopalakrishnan, Matei Ripeanu
HPDC1
2010 Size Matters: Space/Time Tradeoffs to Improve GPGPU Applications Performance
abstract
GPUs offer drastically different performance characteristics compared to traditional multicore architectures. To explore the tradeoffs exposed by this difference, we refactor MUMmer, a widely-used, highly-engineered bioinformatics application which has both CPU- and GPU-based implementations. We synthesize our experience as three high-level guidelines to design efficient GPU-based applications. First, minimizing the communication overheads is as important as optimizing the computation. Second, trading-off higher computational complexity for a more compact in-memory representation is a valuable technique to increase overall performance (by enabling higher parallelism levels and reducing transfer overheads). Finally, ensuring that the chosen solution entails low pre- and post-processing overheads is essential to maximize the overall performance gains. Based on these insights, MUMmerGPU++, our GPU-based design of the MUMmer sequence alignment tool, achieves, on realistic workloads, up to 4× speedup compared to a previous, highly optimized GPU port.
Abdullah Gharaibeh, Matei Ripeanu
SC1
2009 Exploring data reliability tradeoffs in replicated storage systems
abstract
This paper explores the feasibility of a cost-efficient storage architecture that offers the reliability and access performance characteristics of a high-end system. This architecture exploits two opportunities: First, scavenging idle storage from LAN-connected desktops not only offers a low-cost storage space, but also high I/O throughput by aggregating the I/O channels of the participating nodes. Second, the two components of data reliability - durability and availability - can be decoupled to control overall system cost. To capitalize on these opportunities, we integrate two types of components: volatile, scavenged storage and dedicated, yet low-bandwidth durable storage. On the one hand, the durable storage forms a low-cost back-end that enables the system to restore the data the volatile nodes may lose. On the other hand, the volatile nodes provide a high-throughput front-end.
Abdullah Gharaibeh, Matei Ripeanu
HPDC1
2008 StoreGPU: exploiting graphics processing units to accelerate distributed storage systems
abstract
Today Graphics Processing Units (GPUs) are a largely underexploited resource on existing desktops and a possible cost-effective enhancement to high-performance systems. To date, most applications that exploit GPUs are specialized scientific applications. Little attention has been paid to harnessing these highly-parallel devices to support more generic functionality at the operating system or middleware level. This study starts from the hypothesis that generic middleware level techniques that improve distributed system reliability or performance (such as content addressing, erasure coding, or data similarity detection) can be significantly accelerated using GPU support.We take a first step towards validating this hypothesis, focusing on distributed storage systems. As a proof of concept, we design StoreGPU, a library that accelerates a number of hashing based primitives popular in distributed storage system implementations. Our evaluation shows that StoreGPU enables up to eight-fold performance gains on synthetic benchmarks as well as on a high-level application: the online similarity detection between large data files.
Samer Al-Kiswany, Abdullah Gharaibeh, Elizeu Santos-Neto, George Yuan, Matei Ripeanu
HPDC2
2008 stdchk: A Checkpoint Storage System for Desktop Grid Computing
abstract
Checkpointing is an indispensable technique to provide fault tolerance for long-running high-throughput applications like those running on desktop grids. This article argues that a checkpoint storage system, optimized to operate in these environments, can offer multiple benefits: reduce the load on a traditional file system, offer high-performance through specialization, and, finally, optimize data management by taking into account checkpoint application semantics. Such a storage system can present a unifying abstraction to checkpoint operations, while hiding the fact that there are no dedicated resources to store the checkpoint data. We prototype stdchk, a checkpoint storage system that uses scavenged disk space from participating desktops to build a low-cost storage system, offering a traditional file system interface for easy integration with applications. This article presents the stdchk architecture, key performance optimizations, and its support for incremental checkpointing and increased data availability. Our evaluation confirms that the stdchk approach is viable in a desktop grid setting and offers a low cost storage system with desirable performance characteristics: high write throughput as well as reduced storage space and network effort to save checkpoint images.
Samer Al-Kiswany, Matei Ripeanu, Sudharshan S. Vazhkudai, Abdullah Gharaibeh
ICDCS4