EDBT 2026 Demo / reviewers in the wild / expert
Abdullah Gharaibeh
dblp:64/6159
· DBLP profile ↗
11ranked-venue papers
8as first author
0since 2021 · last 2014
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 7 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
6 papers |
Storage systems · 57% GPUs and heterogeneous computing · 37% Performance modeling and evaluation · 4% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Bioinformatics and computational biology · 100% |
Topics — the 13 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
GPUs and heterogeneous computing
GPU computing |
0.3 | 2 | 2013 | GPUs as Storage System Accelerators · IEEE Trans. Parallel Distributed Syst. 2013 A GPU accelerated storage system · HPDC 2010 |
Storage systems › data redundancy
replicated storage |
0.2 | 2 | 2011 | ThriftStore: Finessing Reliability Trade-Offs in Replicated Storage Systems · IEEE Trans. Parallel Distributed Syst. 2011 Exploring data reliability tradeoffs in replicated storage systems · HPDC 2009 |
Storage systems
storage reliability |
0.2 | 2 | 2011 | ThriftStore: Finessing Reliability Trade-Offs in Replicated Storage Systems · IEEE Trans. Parallel Distributed Syst. 2011 Exploring data reliability tradeoffs in replicated storage systems · HPDC 2009 |
GPUs and heterogeneous computing › CPU-GPU heterogeneous computing
GPU offloading |
0.2 | 1 | 2013 | GPUs as Storage System Accelerators · IEEE Trans. Parallel Distributed Syst. 2013 |
Storage systems
storage acceleration |
0.2 | 1 | 2013 | GPUs as Storage System Accelerators · IEEE Trans. Parallel Distributed Syst. 2013 |
Storage systems › storage hierarchy
hybrid storage |
0.1 | 1 | 2011 | ThriftStore: Finessing Reliability Trade-Offs in Replicated Storage Systems · IEEE Trans. Parallel Distributed Syst. 2011 |
GPUs and heterogeneous computing › GPU performance optimization
GPU application optimization |
0.1 | 1 | 2010 | Size Matters: Space/Time Tradeoffs to Improve GPGPU Applications Performance · SC 2010 |
Storage systems
distributed storage |
0.1 | 1 | 2008 | StoreGPU: exploiting graphics processing units to accelerate distributed storage systems · HPDC 2008 |
Storage systems
content-addressable storage |
0.0 | 1 | 2013 | GPUs as Storage System Accelerators · IEEE Trans. Parallel Distributed Syst. 2013 |
Storage systems
similarity detection |
0.0 | 1 | 2013 | GPUs as Storage System Accelerators · IEEE Trans. Parallel Distributed Syst. 2013 |
Bioinformatics and computational biology
sequence alignment |
0.0 | 1 | 2010 | Size Matters: Space/Time Tradeoffs to Improve GPGPU Applications Performance · SC 2010 |
Performance modeling and evaluation
workload characterization |
0.0 | 1 | 2009 | Exploring data reliability tradeoffs in replicated storage systems · HPDC 2009 |
Storage systems › storage reliability
erasure coding |
0.0 | 1 | 2008 | StoreGPU: exploiting graphics processing units to accelerate distributed storage systems · HPDC 2008 |
Methods — techniques the papers use, named apart from their topics
hashing · 0.2space/time tradeoff · 0.2memory representation optimization · 0.2simulation · 0.1analytical modeling · 0.1GPU acceleration · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2014 | DedupT: Deduplication for tape systemsabstractDeduplication is a commonly-used technique on disk-based storage pools. However, deduplication has not been used for tape-based pools: tape characteristics, such as high mount and seek times combined with data fragmentation resulting from deduplication create a toxic combination that leads to unacceptably high retrieval times. This work proposes DedupT, a system that efficiently supports deduplication on tape pools. This paper (i) details the main challenges to enable efficient deduplication on tape libraries, (ii) presents a class of solutions based on graph-modeling of similarity between data items that enables efficient placement on tapes; and (iii) presents the design and evaluation of novel cross-tape and on-tape chunk placement algorithms that alleviate tape mount time overhead and reduce on-tape data fragmentation. Using 4.5 TB of real-world workloads, we show that DedupT retains at least 95% of the deduplication efficiency. We show that DedupT mitigates major retrieval time overheads, and, due to reading less data, is able to offer better restore performance compared to the case of restoring non-deduplicated data. Abdullah Gharaibeh, Cornel Constantinescu, Maohua Lu, Ramani Routray, Prasenjit Sarkar, David Pease, Matei Ripeanu |
MSST | 1 |
| 2013 | On Graphs, GPUs, and Blind Dating: A Workload to Processor Matchmaking QuestabstractGraph processing has gained renewed attention. The increasing large scale and wealth of connected data, such as those accrued by social network applications, demand the design of new techniques and platforms to efficiently derive actionable information from large scale graphs. Hybrid systems that host processing units optimized for both fast sequential processing and bulk processing (e.g., GPUaccelerated systems) have the potential to cope with the heterogeneous structure of real graphs and enable high performance graph processing. Reaching this point, however, poses multiple challenges. The heterogeneity of the processing elements (e.g., GPUs implement a different parallel processing model than CPUs and have much less memory) and the inherent irregularity of graph workloads require careful graph partitioning and load assignment. In particular, the workload generated by a partitioning scheme should match the strength of the processing element the partition is allocated to. This work explores the feasibility and quantifies the performance gains of such low-cost partitioning schemes. We propose to partition the workload between the two types of processing elements based on vertex connectivity. We show that such partitioning schemes offer a simple, yet efficient way to boost the overall performance of the hybrid system. Our evaluation illustrates that processing a 4-billion edges graph on a system with one CPU socket and one GPU, while offloading as little as 25% of the edges to the GPU, achieves 2x performance improvement over state-of-the-art implementations running on a dual-socket symmetric system. Moreover, for the same graph, a hybrid system with dualsocket and dual-GPU is capable of 1.13 Billion breadth-first search traversed edge per second, a performance rate that is competitive with the latest entries in the Graph500 list, yet at a much lower price point. Abdullah Gharaibeh, Lauro Beltrão Costa, Elizeu Santos-Neto, Matei Ripeanu |
IPDPS | 1 |
| 2013 | GPUs as Storage System AcceleratorsabstractMassively multicore processors, such as graphics processing units (GPUs), provide, at a comparable price, a one order of magnitude higher peak performance than traditional CPUs. This drop in the cost of computation, as any order-of-magnitude drop in the cost per unit of performance for a class of system components, triggers the opportunity to redesign systems and to explore new ways to engineer them to recalibrate the cost-to-performance relation. This project explores the feasibility of harnessing GPUs' computational power to improve the performance, reliability, or security of distributed storage systems. In this context, we present the design of a storage system prototype that uses GPU offloading to accelerate a number of computationally intensive primitives based on hashing, and introduce techniques to efficiently leverage the processing power of GPUs. We evaluate the performance of this prototype under two configurations: as a content addressable storage system that facilitates online similarity detection between successive versions of the same file and as a traditional system that uses hashing to preserve data integrity. Further, we evaluate the impact of offloading to the GPU on competing applications' performance. Our results show that this technique can bring tangible performance gains without negatively impacting the performance of concurrently running applications. Samer Al-Kiswany, Abdullah Gharaibeh, Matei Ripeanu |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2012 | A yoke of oxen and a thousand chickens for heavy lifting graph processingabstractLarge, real-world graphs are famously difficult to process efficiently. Not only they have a large memory footprint but most graph processing algorithms entail memory access patterns with poor locality, data-dependent parallelism, and a low compute-to- memory access ratio. Additionally, most real-world graphs have a low diameter and a highly heterogeneous node degree distribution. Partitioning these graphs and simultaneously achieve access locality and load-balancing is difficult if not impossible. Abdullah Gharaibeh, Lauro Beltrão Costa, Elizeu Santos-Neto, Matei Ripeanu |
PACT | 1 |
| 2012 | CloudDT: Efficient tape resource management using deduplication in cloud backup and archival services
Abdullah Gharaibeh, Cornel Constantinescu, Maohua Lu, Ramani Routray, Prasenjit Sarkar, David Pease, Matei Ripeanu |
CNSM | 1 |
| 2011 | ThriftStore: Finessing Reliability Trade-Offs in Replicated Storage SystemsabstractThis paper explores the feasibility of a storage architecture that offers the reliability and access performance characteristics of a high-end system, yet is cost-efficient. We propose ThriftStore, a storage architecture that integrates two types of components: volatile, aggregated storage and dedicated, yet low-bandwidth durable storage. On the one hand, the durable storage forms a back end that enables the system to restore the data the volatile nodes may lose. On the other hand, the volatile nodes provide a high-throughput front-end. Although integrating these components has the potential to offer a unique combination of high throughput and durability at a low cost, a number of concerns need to be addressed to architect and correctly provision the system. To this end, we develop analytical and simulation-based tools to evaluate the impact of system characteristics (e.g., bandwidth limitations on the durable and the volatile nodes) and design choices (e.g., the replica placement scheme) on data availability and the associated system costs (e.g., maintenance traffic). Moreover, to demonstrate the high-throughput properties of the proposed architecture, we prototype a GridFTP server based on ThriftStore. Our evaluation demonstrates an impressive, up to 800 Mbps transfer throughput for the new GridFTP service. Abdullah Gharaibeh, Samer Al-Kiswany, Matei Ripeanu |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2010 | A GPU accelerated storage systemabstractMassively multicore processors, like, for example, Graphics Processing Units (GPUs), provide, at a comparable price, a one order of magnitude higher peak performance than traditional CPUs. This drop in the cost of computation, as any order-of-magnitude drop in the cost per unit of performance for a class of system components, triggers the opportunity to redesign systems and to explore new ways to engineer them to recalibrate the cost-to-performance relation. Abdullah Gharaibeh, Samer Al-Kiswany, Sathish Gopalakrishnan, Matei Ripeanu |
HPDC | 1 |
| 2010 | Size Matters: Space/Time Tradeoffs to Improve GPGPU Applications PerformanceabstractGPUs offer drastically different performance characteristics compared to traditional multicore architectures. To explore the tradeoffs exposed by this difference, we refactor MUMmer, a widely-used, highly-engineered bioinformatics application which has both CPU- and GPU-based implementations. We synthesize our experience as three high-level guidelines to design efficient GPU-based applications. First, minimizing the communication overheads is as important as optimizing the computation. Second, trading-off higher computational complexity for a more compact in-memory representation is a valuable technique to increase overall performance (by enabling higher parallelism levels and reducing transfer overheads). Finally, ensuring that the chosen solution entails low pre- and post-processing overheads is essential to maximize the overall performance gains. Based on these insights, MUMmerGPU++, our GPU-based design of the MUMmer sequence alignment tool, achieves, on realistic workloads, up to 4× speedup compared to a previous, highly optimized GPU port. Abdullah Gharaibeh, Matei Ripeanu |
SC | 1 |
| 2009 | Exploring data reliability tradeoffs in replicated storage systemsabstractThis paper explores the feasibility of a cost-efficient storage architecture that offers the reliability and access performance characteristics of a high-end system. This architecture exploits two opportunities: First, scavenging idle storage from LAN-connected desktops not only offers a low-cost storage space, but also high I/O throughput by aggregating the I/O channels of the participating nodes. Second, the two components of data reliability - durability and availability - can be decoupled to control overall system cost. To capitalize on these opportunities, we integrate two types of components: volatile, scavenged storage and dedicated, yet low-bandwidth durable storage. On the one hand, the durable storage forms a low-cost back-end that enables the system to restore the data the volatile nodes may lose. On the other hand, the volatile nodes provide a high-throughput front-end. Abdullah Gharaibeh, Matei Ripeanu |
HPDC | 1 |
| 2008 | StoreGPU: exploiting graphics processing units to accelerate distributed storage systemsabstractToday Graphics Processing Units (GPUs) are a largely underexploited resource on existing desktops and a possible cost-effective enhancement to high-performance systems. To date, most applications that exploit GPUs are specialized scientific applications. Little attention has been paid to harnessing these highly-parallel devices to support more generic functionality at the operating system or middleware level. This study starts from the hypothesis that generic middleware level techniques that improve distributed system reliability or performance (such as content addressing, erasure coding, or data similarity detection) can be significantly accelerated using GPU support.We take a first step towards validating this hypothesis, focusing on distributed storage systems. As a proof of concept, we design StoreGPU, a library that accelerates a number of hashing based primitives popular in distributed storage system implementations. Our evaluation shows that StoreGPU enables up to eight-fold performance gains on synthetic benchmarks as well as on a high-level application: the online similarity detection between large data files. Samer Al-Kiswany, Abdullah Gharaibeh, Elizeu Santos-Neto, George Yuan, Matei Ripeanu |
HPDC | 2 |
| 2008 | stdchk: A Checkpoint Storage System for Desktop Grid ComputingabstractCheckpointing is an indispensable technique to provide fault tolerance for long-running high-throughput applications like those running on desktop grids. This article argues that a checkpoint storage system, optimized to operate in these environments, can offer multiple benefits: reduce the load on a traditional file system, offer high-performance through specialization, and, finally, optimize data management by taking into account checkpoint application semantics. Such a storage system can present a unifying abstraction to checkpoint operations, while hiding the fact that there are no dedicated resources to store the checkpoint data. We prototype stdchk, a checkpoint storage system that uses scavenged disk space from participating desktops to build a low-cost storage system, offering a traditional file system interface for easy integration with applications. This article presents the stdchk architecture, key performance optimizations, and its support for incremental checkpointing and increased data availability. Our evaluation confirms that the stdchk approach is viable in a desktop grid setting and offers a low cost storage system with desirable performance characteristics: high write throughput as well as reduced storage space and network effort to save checkpoint images. Samer Al-Kiswany, Matei Ripeanu, Sudharshan S. Vazhkudai, Abdullah Gharaibeh |
ICDCS | 4 |