EDBT 2026 Demo / reviewers in the wild / expert
Matei Ripeanu
dblp:93/24
· DBLP profile ↗
76ranked-venue papers
5as first author
6since 2021 · last 2024
0000-0001-9839-3866ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 54 · 4 first-author · 5 since 2021Security and privacy · 8Computer networks · 7 · 1 first-authorDatabases, data management, data science and information retrieval · 6 · 1 since 2021Artificial intelligence and machine learning · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Software engineering, systems software and programming languages · 2Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | (re)Assessing PiM Effectiveness for Sequence Alignment
Hamidreza Ramezani-Kebrya, Matei Ripeanu |
Euro-Par (2) | 2 |
| 2023 | Maximum Flow on Highly Dynamic GraphsabstractRecent advances in dynamic graph processing have enabled the analysis of highly dynamic graphs with change at rates as high as millions of edge changes per second. Solutions in this domain, however, have been demonstrated only for relatively simple algorithms like PageRank, breadth-first search, and connected components. Expanding beyond this, we explore the maximum flow problem, a fundamental, yet more complex problem, in graph analytics. We propose a novel, distributed algorithm for max-flow on dynamic graphs, and implement it on top of an asynchronous vertex-centric abstraction. We show that our algorithm can process both additions and deletions of vertices and edges efficiently at scale on fast-evolving graphs, and provide a comprehensive analysis by evaluating, in addition to throughput, two criteria that are important when applied to real-world problems: result latency and solution stability. Juntong Luo 0001, Scott Sallinen, Matei Ripeanu |
IEEE Big Data | 3 |
| 2023 | Real-Time PageRank on Dynamic GraphsabstractModern data generation has grown to enormous proportions, with events occurring at increasingly higher rates. Yet for graph analytics, this growth in scale and velocity has not been matched by improved algorithm or infrastructure techniques: most systems still focus on post-mortem or static analysis. This paper builds on an efficient graph processing abstraction that enables online analysis of dynamically evolving graphs at scale. Integral to this abstraction is that events tied to both graph topology changes as well as algorithmic maintenance occur and are processed asynchronously, concurrently, and autonomously (i.e., without shared state). Scott Sallinen, Juntong Luo 0001, Matei Ripeanu |
HPDC | 3 |
| 2023 | EdgeEngine: A Thermal-Aware Optimization Framework for Edge InferenceabstractHeterogeneous edge platforms enable the efficient execution of machine learning inference applications. These applications often have a critical constraint (such as meeting a deadline) and an optimization goal (such as minimizing energy consumption). To navigate this space, existing optimization frameworks adjust the platform's frequency configuration for the CPU, the GPU and/or the memory controller. However, existing optimization frameworks have two limitations. First, edge applications are frequently deployed in environments where they are exposed to ambient temperature variations. Second, a recent study has shown that temperature has a significant impact on edge platform characteristics. In this context, today's frequency optimization frameworks (which are thermal-oblivious) select frequency configurations that either violate the application's constraints, or are sub-optimal in terms of the optimization goal. Amirhossein Ahmadi, Hazem A. Abdelhafez, Karthik Pattabiraman, Matei Ripeanu |
SEC | 4 |
| 2022 | Characterizing Variability in Heterogeneous Edge Systems: A Methodology & Case StudyabstractThis study offers a methodology to characterize intra- and inter-node variability and applies it on two heterogeneous edge platforms (the NVIDIA Jetson AGX and Nano) for performance and power consumption. Firstly, we explore intra-node variability: investigate to what degree deployment decisions can limit it, highlight that it is unavoidable, and offer a scale so that one can compare to what other studies report. Secondly, we characterize inter-node variability by answering two questions: (i) Are the platforms we study statistically different in terms of the applications' power draw and runtime? and (ii) What is the magnitude of these differences? Finally, we attempt to answer the question of why is it paramount to characterize variability and take it into account? to achieve this, we discuss examples from the compiler and runtime optimization domains. Hazem A. Abdelhafez, Hassan Halawa, Amr Almoallim, Amirhossein Ahmadi, Karthik Pattabiraman, Matei Ripeanu |
SEC | 6 |
| 2021 | MIRAGE: Machine Learning-based Modeling of Identical Replicas of the Jetson AGX Embedded Platform
Hassan Halawa, Hazem A. Abdelhafez, Mohamed Osama Ahmed, Karthik Pattabiraman, Matei Ripeanu |
SEC | 5 |
| 2020 | HyGN: Hybrid Graph Engine for NUMAabstractModern shared-memory platforms embrace the Non-uniform Memory Access (NUMA) architecture - they have physically distributed, yet cache-coherent shared-memory. This paper explores the feasibility of a shared-memory graph processing engine for NUMA platforms inspired by designs that target zero-sharing platforms. This work exploits the characteristics of two processing modes, synchronous and asynchronous, in the context of the shared-memory NUMA platform. Depending on the algorithm, phase of execution, and graph topology, synchronous and asynchronous modes hold unique advantages over one another. We then explore a hybrid solution that combines synchronous and asynchronous processing within the same graph computation task and harness optimizations therein. An extensive evaluation using graphs with billions of edges and empirical comparisons with several state-of-the-art solutions demonstrate the performance advantages of our design. Tanuj Kr Aasawat, Tahsin Reza, Kazuki Yoshizoe, Matei Ripeanu |
IEEE BigData | 4 |
| 2020 | Approximate Pattern Matching in Massive Graphs with Precision and Recall GuaranteesabstractThere are multiple situations where supporting approximation in graph pattern matching tasks is highly desirable: (i) the data acquisition process can be noisy; (ii) a user may only have an imprecise idea of the search query; and (iii) approximation can be used for high volume vertex labeling when extracting machine learning features from graph data. We present a new algorithmic pipeline for approximate matching that combines edit-distance based matching with systematic graph pruning. We formalize the problem as identifying all exact matches for up to k edit-distance subgraphs of a user-supplied template. We design a solution which exploits unique optimization opportunities within the design space, not explored previously. Our solution is (i) highly scalable, (ii) supports arbitrary patterns and edit-distance, (iii) offers 100% precision and 100% recall guarantees, and (vi) supports a set of popular data analysis scenarios. We demonstrate its advantages through an implementation that offers good strong and weak scaling on massive real-world (257 billion edges) and synthetic (1.1 trillion edges) labeled graphs, respectively, and when operating on a massive cluster (256 nodes/9,216 cores), orders of magnitude larger than previously used for similar problems. Empirical comparison with the state-of-the-art highlights the advantages of our solution when handling massive graphs and complex patterns. Tahsin Reza, Matei Ripeanu, Geoffrey Sanders, Roger A. Pearce |
SIGMOD Conference | 2 |
| 2019 | BonVoision: leveraging spatial data smoothness for recovery from memory soft errorsabstractDetectable but Uncorrectable Errors (DUEs) in the memory subsystem are becoming increasingly frequent. Today, upon encountering a DUE, applications crash, and the recovery methods used incur significant performance, storage, and energy overheads. To mitigate the impact of these errors, we start from two high-level observations that apply to some classes of HPC applications (e.g., stencil computations on regular grids or irregular meshes): first, these applications, display a property we dub spatial data smoothness: i.e., data items that are nearby in the application's logical space are relatively similar. Second, since these data items are generally used together, programmers go to great lengths to place them in nearby memory locations to improve application's performance by improving access locality. Based on these observations we explore the feasibility of a roll-forward recovery scheme that leverages spatial data smoothness to repair the memory location corrupted by a DUE and continues the application execution. We present BonVoision, a run-time system that intercepts DUE events, analyzes the application binary at runtime to identify the data elements in the neighborhood of the memory location that generates a DUE, and uses them to fix the corrupted data. Our evaluation demonstrates that BonVoision is: (i) efficient - it incurs negligible overhead, (ii) effective - it is frequently successful in continuing the application with benign outcomes, and (iii) user friendly - as it does not require programmer input to expose the data layout or access to source code. We demonstrate that using BonVoision can lead to significant savings in the context of a checkpointing/restart schemes by enabling longer checkpoint intervals. Bo Fang 0002, Hassan Halawa, Karthik Pattabiraman, Matei Ripeanu, Sriram Krishnamoorthy |
ICS | 4 |
| 2019 | Incremental Graph Processing for On-line AnalyticsabstractModern data generation is enormous; we now capture events at increasingly fine granularity, and require processing at rates approaching real-time. For graph analytics, this explosion in data volumes and processing demands has not been matched by improved algorithmic or infrastructure techniques. Instead of exploring solutions to keep up with the velocity of the generated data, most of today's systems focus on analyzing individually built historic snapshots. Modern graph analytics pipelines must evolve to become viable at massive scale, and move away from static, post-processing scenarios to support on-line analysis. This paper presents our progress towards a system that analyzes dynamic incremental graphs, responsive at single-change granularity. We present an algorithmic structure using principles of recursive updates and monotonic convergence, and a set of incremental graph algorithms that can be implemented based on this structure. We also present the required middleware to support graph analytics at fine, event-level granularity. We envision that graph topology changes are processed asynchronously, concurrently, and independently (without shared state), converging an algorithm's state (e.g. single-source shortest path distances, connectivity analysis labeling) to its deterministic answer. The expected long-term impact of this work is to enable a transition away from offline graph analytics, allowing knowledge to be extracted from networked systems in real-time. Scott Sallinen, Roger A. Pearce, Matei Ripeanu |
IPDPS | 3 |
| 2018 | PruneJuice: pruning trillion-edge graphs to a precise pattern-matching solution
Tahsin Reza, Matei Ripeanu, Nicolas Tripoul, Geoffrey Sanders, Roger A. Pearce |
SC | 2 |
| 2018 | Accelerating Persistent Scatterer Pixel Selection for InSAR ProcessingabstractInterferometric Synthetic Aperture Radar (InSAR) is a remote sensing technology used for estimating the displacement of an object on the ground or the earth's surface itself. Persistent Scatterer-InSAR (PS-InSAR) is a category of time series algorithms enabling high resolution monitoring. PS-InSAR relies on successful selection of points that appear stable across a set of satellite images taken overtime. This paper presents PtSel, a new algorithm for selecting these points, a problem known as Persistent Scatterer Selection. The key advantage of PtSel over the key existing techniques is that it does not require model assumptions, yet preserves solution accuracy. Motivated by the abundance of parallelism the algorithm exposes, we have implemented it for GPUs. Our evaluation using real-world data shows that the GPU implementation not only offers superior performance but also scales linearly with GPU count and workload size. We compare the GPU implementation and a parallel CPU implementation: a consumer grade GPU offers 18x speedup over a 16-core Ivy Bridge Xeon System, while four GPUs offer 65x speedup. The GPU solution consumes 28x less energy than the CPU-only solution. Additionally, we present a comparison with the most widely used PS-interferometry software package StaMPS, in terms of point selection coverage and precision. Tahsin Reza, Aaron Zimmer, José Manuel Delgado Blasco, Parwant Ghuman, Tanuj Kr Aasawat, Matei Ripeanu |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2017 | Towards Practical and Robust Labeled Pattern Matching in Trillion-Edge GraphsabstractSubgraph pattern matching is fundamental to graph analytics and has wide applications. Unfortunately, high computational complexity limits the robustness guarantees of existing algorithms: they do not scale for modern large graph datasets and/or they have limitations in terms of accuracy or in terms of the intricacy of the patterns supported. We present algorithms, theory, and empirical evidence that iteratively eliminating vertices that do not meet local constraints dramatically reduces the search space for pattern matching in real-world graphs, and demonstrate a scalable implementation of our algorithms. We additionally identify the characteristics of patterns for which every non-eliminated vertex participates in a match. These techniques are an essential step to enable scalable, practical solutions for robust pattern matching in large-scale labeled graphs.We demonstrate the advantages of the proposed approach through strong and weak scaling experiments on massive-scale real-world (up to 257 billion edges) and synthetic (up to 2.2 trillion edges) graphs and at scales (256 compute nodes with 6,144 processors) orders of magnitude larger than those used in the past for similar problems. Tahsin Reza, Christine Klymko, Matei Ripeanu, Geoffrey Sanders, Roger A. Pearce |
CLUSTER | 3 |
| 2017 | NVIDIA Jetson Platform Characterization
Hassan Halawa, Hazem A. Abdelhafez, Andrew Boktor, Matei Ripeanu |
Euro-Par | 4 |
| 2017 | LetGo: A Lightweight Continuous Framework for HPC Applications Under FailuresabstractRequirements for reliability, low power consumption, and performance place complex and conflicting demands on the design of high-performance computing (HPC) systems. Fault-tolerance techniques such as checkpoint/restart (C/R) protect HPC applications against hardware faults. These techniques, however, have non negligible overheads particularly when the fault rate exposed by the hardware is high: it is estimated that in future HPC systems, up to 60% of the computational cycles/power will be used for fault tolerance. Bo Fang 0002, Qiang Guan, Nathan DeBardeleben, Karthik Pattabiraman, Matei Ripeanu |
HPDC | 5 |
| 2017 | A cross-layer optimized storage system for workflow applications
Samer Al-Kiswany, Lauro Beltrão Costa, Hao Yang 0039, Emalayan Vairavanathan, Matei Ripeanu |
Future Gener. Comput. Syst. | 5 |
| 2016 | A Software-Defined Storage for Workflow ApplicationsabstractWe present a software-defined storage architecture with two key properties: firstly, it enables external control of storage operations, and, secondly, it allows extending the storage system with new, workload specific, optimizations, all without breaking the file system abstractions. We argue that this architecture is generic, we prototype FlexStore following this architecture, we instantiate it to support workflow applications, and we report on our preliminary experience. Samer Al-Kiswany, Matei Ripeanu |
CLUSTER | 2 |
| 2016 | ePVF: An Enhanced Program Vulnerability Factor Methodology for Cross-Layer Resilience AnalysisabstractThe Program Vulnerability Factor (PVF) has been proposed as a metric to understand the impact of hardware faults on software. The PVF is calculated by identifying the program bits required for architecturally correct execution (ACE bits). PVF, however, is conservative as it assumes that all erroneous executions are a major concern, not just those that result in silent data corruptions, and it also does not account for errorsthat are detected at runtime, i.e., lead to program crashes. A more discriminating metric can inform the choice of the appropriate resilience techniques with acceptable performance and energy overheads. This paper proposes ePVF, an enhancement of the original PVF methodology, which filters out the crash-causing bits from the ACE bits identified by the traditional PVF analysis. The ePVF methodology consists of an error propagation model that reasons about error propagation in the program, and a crash model that encapsulates the platform-specific characteristics for handling hardware exceptions. ePVF reduces the vulnerable bits estimated by the original PVF analysis by between 45% and 67% depending on the benchmark, and has high accuracy (89% recall, 92% precision) in identifying the crash-causing bits. We demonstrate the utility of ePVF by using it to inform selectiveprotection of the most SDC-prone instructions in a program. Bo Fang 0002, Qining Lu, Karthik Pattabiraman, Matei Ripeanu, Sudhanva Gurumurthi |
DSN | 4 |
| 2016 | Harvesting the low-hanging fruits: defending against automated large-scale cyber-intrusions by focusing on the vulnerable populationabstractThe orthodox paradigm to defend against automated social-engineering attacks in large-scale socio-technical systems is reactive and victim-agnostic. Defenses generally focus on identifying the attacks/attackers (e.g., phishing emails, social-bot infiltrations, malware offered for download). To change the status quo, we propose to identify, even if imperfectly, the vulnerable user population, that is, the users that are likely to fall victim to such attacks. Once identified, information about the vulnerable population can be used in two ways. First, the vulnerable population can be influenced by the defender through several means including: education, specialized user experience, extra protection layers and watchdogs. In the same vein, information about the vulnerable population can ultimately be used to fine-tune and reprioritize defense mechanisms to offer differentiated protection, possibly at the cost of additional friction generated by the defense mechanism. Secondly, information about the user population can be used to identify an attack (or compromised users) based on differences between the general and the vulnerable population. This paper considers the implications of the proposed paradigm on existing defenses in three areas (phishing of user credentials, malware distribution and socialbot infiltration) and discusses how using knowledge of the vulnerable population can enable more robust defenses. Hassan Halawa, Konstantin Beznosov, Yazan Boshmaf, Baris Coskun, Matei Ripeanu, Elizeu Santos-Neto |
NSPW | 5 |
| 2016 | Graph colouring as a challenge problem for dynamic graph processing on distributed systemsabstractAn unprecedented growth in data generation is taking place. Data about larger dynamic systems is being accumulated, capturing finer granularity events, and thus processing requirements are increasingly approaching real-time. To keep up, data-analytics pipelines need to be viable at massive scale, and switch away from static, offline scenarios to support fully online analysis of dynamic systems. This paper uses a challenge problem, graph colouring, to explore massive-scale analytics for dynamic graph processing. We present an event-based infrastructure, and a novel, online, distributed graph colouring algorithm. Our implementation for colouring static graphs, used as a performance baseline, is up to an order of magnitude faster than previous results and handles massive graphs with over 257 billion edges. Our framework supports dynamic graph colouring with performance at large scale better than GraphLab's static analysis. Our experience indicates that online solutions are feasible, and can be more efficient than those based on snapshotting. Scott Sallinen, Keita Iwabuchi, Suraj Poudel, Maya B. Gokhale, Matei Ripeanu, Roger A. Pearce |
SC | 5 |
| 2016 | Íntegro: Leveraging victim prediction for robust fake account detection in large scale OSNs
Yazan Boshmaf, Dionysios Logothetis, Georgos Siganos, Jorge Lería, José Lorenzo, Matei Ripeanu, Konstantin Beznosov, Hassan Halawa |
Comput. Secur. | 6 |
| 2016 | Support for Provisioning and Configuration Decisions for Data Intensive WorkflowsabstractSystem provisioning, resource allocation, and configuration decisions for I/O-intensive workflow applications are complex even for expert users. Users face choices at multiple levels: allocating resources to individual sub-systems (e.g., the application layer, the storage layer) as well as configuring each of these optimally (e.g., replication level, chunk size, caching policies in case of storage) all having a large impact on the overall application performance. This paper presents a solution to address the problem of supporting these provisioning, allocation and configuration decisions for workflow applications. To enable selecting a good choice in a reasonable time, we propose an approach that accelerates the exploration of the configuration space based on a low-cost performance predictor that estimates total execution time of a workflow application in a given setup. We evaluate the predictor in a number of different scenarios including the Montage application: a workflow composed of over 7,500 tasks structured in 10 different stages with varying characteristics. Our evaluation shows that: (i) the predictor is effective in identifying the desired system configuration, (ii) it can scale to model a complex workflow application run on a 100-node cluster, while (iii) using orders of magnitude less resources than running the actual application. Additionally, we extend the predictor to estimate the energy usage of the system, and we present our experience with incorporating it in the development process of a distributed storage system. Lauro Beltrão Costa, Samer Al-Kiswany, Matei Ripeanu, Hao Yang 0039 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2016 | A Systematic Methodology for Evaluating the Error Resilience of GPGPU ApplicationsabstractThe wide adoption of graphics processing units (GPUs) as accelerators for general-purpose applications makes the end-to-end reliability implications of their use increasingly significant. Fault injection is a widely adopted method to evaluate the resilience of applications. However, building a fault injector for general-purpose GPU applications is challenging due to their massive parallelism, which makes it difficult to achieve representativeness while being time-efficient. This paper makes four key contributions. First, it presents a fault-injection methodology to evaluate the end-to-end reliability properties of application kernels running on GPUs. Second, it introduces GPU-Qin, a fault-injection tool that uses real GPU hardware and offers a tunable and efficient balance between the representativeness and the cost of a fault-injection campaign. Third, it characterizes the error resilience characteristics of seventeen application kernels. Finally, it provides preliminary insights on correlations between the algorithmic properties of applications and their error resilience. Bo Fang 0002, Karthik Pattabiraman, Matei Ripeanu, Sudhanva Gurumurthi |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2015 | Accelerating persistent scatterer pixel selection for InSAR processingabstractInterferometric Synthetic Aperture Radar (InSAR) is a remote sensing technology used for estimating displacement of the earth's surface. Phase unwrapping is the most important step in InSAR processing and relies on successful selection of points that appear stable across a set of satellite images taken over time. This paper presents a new algorithm for selecting these points, a problem known as persistent scatterer selection. The algorithm computes the temporal coherence on the wrapped phase derivative by subtracting phases of a pixel and one of its nearby neighbours. It does not require model assumptions, yet preserves accuracy. Motivated by the abundance of parallelism the algorithm exposes, we have implemented it for GPUs. Evaluation using real-world data shows that the GPU implementation not only offers widely superior performance but also scales linearly with GPU count and workload size. We compare the GPU implementation against a parallel CPU implementation: A consumer grade GPU offers an 18× speedup over a 16-core Ivy Bridge Xeon System, while four GPUs offer 65× speedup. Roofline analysis shows that, on a single GPU, our implementation achieves 83% of the peak FLOP-rate of the dual-CPU system. Additionally, the GPU-based solution consumes 29× less energy than the CPU-only solution. Tahsin Reza, Aaron Zimmer, Parwant Ghuman, Tanuj Kr Aasawat, Matei Ripeanu |
ASAP | 5 |
| 2015 | Integro: Leveraging Victim Prediction for Robust Fake Account Detection in OSNs
Yazan Boshmaf, Dionysios Logothetis, Georgos Siganos, Jorge Lería, José Lorenzo, Matei Ripeanu, Konstantin Beznosov |
NDSS | 6 |
| 2015 | Active Data: A programming model to manage data life cycle across heterogeneous systems and infrastructures
Anthony Simonet, Gilles Fedak, Matei Ripeanu |
Future Gener. Comput. Syst. | 3 |
| 2015 | The Case for Workflow-Aware Storage: An Opportunity Study
Lauro Beltrão Costa, Hao Yang 0039, Emalayan Vairavanathan, Abmar Barros, Ketan Maheshwari, Gilles Fedak, Daniel S. Katz, Michael Wilde, Matei Ripeanu, Samer Al-Kiswany |
J. Grid Comput. | 9 |
| 2014 | Evaluating the Error Resilience of Parallel ProgramsabstractAs a consequence of increasing hardware fault rates, HPC systems face significant challenges in terms of reliability. Evaluating the error resilience of HPC applications is an essential step for building efficient fault-tolerant mechanisms for these applications. In this paper, we propose a methodology to characterize the resilience of OpenMP programs using fault-injection experiments. We find that the error resilience of OpenMP applications depends on the program structure and thread model, hence, these need to be taken into account while characterizing error resilience. We also report preliminary results about the correlation between the application's error resilience and the algorithm(s) used in the application. Bo Fang 0002, Karthik Pattabiraman, Matei Ripeanu, Sudhanva Gurumurthi |
DSN | 3 |
| 2014 | Supporting storage configuration for I/O intensive workflowsabstractSystem provisioning, resource allocation, and system configuration decisions for I/O-intensive workflow applications are complex even for expert users. Users face choices at multiple levels: allocating resources to individual sub-systems (e.g., the application layer, the storage layer) and configuring each of these optimally (e.g., replication level, chunk size, caching policies in case of storage) all having a large impact on overall application performance. This paper presents our progress on addressing the problem of supporting these provisioning, allocation and configuration decisions for workflow applications. To enable selecting a good choice in a reasonable time, we propose an approach that accelerates the exploration of the configuration space based on a low-cost performance predictor that estimates total execution time of a workflow application in a given setup. Our evaluation shows that: (i) the predictor is effective in identifying the desired system configuration, (ii) it can scale to model a workflow application run on an entire cluster, while (iii) using over 2000x less resources (machines x time) than running the actual application. Lauro Beltrão Costa, Samer Al-Kiswany, Hao Yang 0039, Matei Ripeanu |
ICS | 4 |
| 2014 | GPU-Qin: A methodology for evaluating the error resilience of GPGPU applicationsabstractWhile graphics processing units (GPUs) have gained wide adoption as accelerators for general-purpose applications (GPGPU), the end-to-end reliability implications of their use have not been quantified. Fault injection is a widely used method for evaluating the reliability of applications. However, building a fault injector for GPGPU applications is challenging due to their massive parallelism, which makes it difficult to achieve representativeness while being time-efficient. This paper makes three key contributions. First, it presents the design of a fault-injection methodology to evaluate end-to-end reliability properties of application kernels running on GPUs. Second, it introduces a fault-injection tool that uses real GPU hardware and offers a good balance between the representativeness and the efficiency of the fault injection experiments. Third, this paper characterizes the error resilience characteristics of twelve GPGPU applications. Bo Fang 0002, Karthik Pattabiraman, Matei Ripeanu, Sudhanva Gurumurthi |
ISPASS | 3 |
| 2014 | DedupT: Deduplication for tape systemsabstractDeduplication is a commonly-used technique on disk-based storage pools. However, deduplication has not been used for tape-based pools: tape characteristics, such as high mount and seek times combined with data fragmentation resulting from deduplication create a toxic combination that leads to unacceptably high retrieval times. This work proposes DedupT, a system that efficiently supports deduplication on tape pools. This paper (i) details the main challenges to enable efficient deduplication on tape libraries, (ii) presents a class of solutions based on graph-modeling of similarity between data items that enables efficient placement on tapes; and (iii) presents the design and evaluation of novel cross-tape and on-tape chunk placement algorithms that alleviate tape mount time overhead and reduce on-tape data fragmentation. Using 4.5 TB of real-world workloads, we show that DedupT retains at least 95% of the deduplication efficiency. We show that DedupT mitigates major retrieval time overheads, and, due to reading less data, is able to offer better restore performance compared to the case of restoring non-deduplicated data. Abdullah Gharaibeh, Cornel Constantinescu, Maohua Lu, Ramani Routray, Prasenjit Sarkar, David Pease, Matei Ripeanu |
MSST | 8 |
| 2014 | Cheating in Online Games: A Social Network PerspectiveabstractOnline gaming is a multi-billion dollar industry that entertains a large, global population. One unfortunate phenomenon, however, poisons the competition and spoils the fun: cheating. The costs of cheating span from industry-supported expenditures to detect and limit it, to victims’ monetary losses due to cyber crime. This article studies cheaters in the Steam Community, an online social network built on top of the world’s dominant digital game delivery platform. We collected information about more than 12 million gamers connected in a global social network, of which more than 700 thousand have their profiles flagged as cheaters. We also observed timing information of the cheater flags, as well as the dynamics of the cheaters’ social neighborhoods. We discovered that cheaters are well embedded in the social and interaction networks: their network position is largely indistinguishable from that of fair players. Moreover, we noticed that the number of cheaters is not correlated with the geographical, real-world population density, or with the local popularity of the Steam Community. Also, we observed a social penalty involved with being labeled as a cheater: cheaters lose friends immediately after the cheating label is publicly applied. Most importantly, we observed that cheating behavior spreads through a social mechanism: the number of cheater friends of a fair player is correlated with the likelihood of her becoming a cheater in the future. This allows us to propose ideas for limiting cheating contagion. Jeremy Blackburn, Nicolas Kourtellis, John Skvoretz, Matei Ripeanu, Adriana Iamnitchi |
ACM Trans. Internet Techn. | 4 |
| 2013 | Graph-based Sybil detection in social and information systemsabstractSybil attacks in social and information systems have serious security implications. Out of many defence schemes, Graph-based Sybil Detection (GSD) had the greatest attention by both academia and industry. Even though many GSD algorithms exist, there is no analytical framework to reason about their design, especially as they make different assumptions about the used adversary and graph models. In this paper, we bridge this knowledge gap and present a unified framework for systematic evaluation of GSD algorithms. We used this framework to show that GSD algorithms should be designed to find local community structures around known non-Sybil identities, while incrementally tracking changes in the graph as it evolves over time. Yazan Boshmaf, Konstantin Beznosov, Matei Ripeanu |
ASONAM | 3 |
| 2013 | On Graphs, GPUs, and Blind Dating: A Workload to Processor Matchmaking QuestabstractGraph processing has gained renewed attention. The increasing large scale and wealth of connected data, such as those accrued by social network applications, demand the design of new techniques and platforms to efficiently derive actionable information from large scale graphs. Hybrid systems that host processing units optimized for both fast sequential processing and bulk processing (e.g., GPUaccelerated systems) have the potential to cope with the heterogeneous structure of real graphs and enable high performance graph processing. Reaching this point, however, poses multiple challenges. The heterogeneity of the processing elements (e.g., GPUs implement a different parallel processing model than CPUs and have much less memory) and the inherent irregularity of graph workloads require careful graph partitioning and load assignment. In particular, the workload generated by a partitioning scheme should match the strength of the processing element the partition is allocated to. This work explores the feasibility and quantifies the performance gains of such low-cost partitioning schemes. We propose to partition the workload between the two types of processing elements based on vertex connectivity. We show that such partitioning schemes offer a simple, yet efficient way to boost the overall performance of the hybrid system. Our evaluation illustrates that processing a 4-billion edges graph on a system with one CPU socket and one GPU, while offloading as little as 25% of the edges to the GPU, achieves 2x performance improvement over state-of-the-art implementations running on a dual-socket symmetric system. Moreover, for the same graph, a hybrid system with dualsocket and dual-GPU is capable of 1.13 Billion breadth-first search traversed edge per second, a performance rate that is competitive with the latest entries in the Graph500 list, yet at a much lower price point. Abdullah Gharaibeh, Lauro Beltrão Costa, Elizeu Santos-Neto, Matei Ripeanu |
IPDPS | 4 |
| 2013 | Design and analysis of a social botnet
Yazan Boshmaf, Ildar Muslukhov, Konstantin Beznosov, Matei Ripeanu |
Comput. Networks | 4 |
| 2013 | GPUs as Storage System AcceleratorsabstractMassively multicore processors, such as graphics processing units (GPUs), provide, at a comparable price, a one order of magnitude higher peak performance than traditional CPUs. This drop in the cost of computation, as any order-of-magnitude drop in the cost per unit of performance for a class of system components, triggers the opportunity to redesign systems and to explore new ways to engineer them to recalibrate the cost-to-performance relation. This project explores the feasibility of harnessing GPUs' computational power to improve the performance, reliability, or security of distributed storage systems. In this context, we present the design of a storage system prototype that uses GPU offloading to accelerate a number of computationally intensive primitives based on hashing, and introduce techniques to efficiently leverage the processing power of GPUs. We evaluate the performance of this prototype under two configurations: as a content addressable storage system that facilitates online similarity detection between successive versions of the same file and as a traditional system that uses hashing to preserve data integrity. Further, we evaluate the impact of offloading to the GPU on competing applications' performance. Our results show that this technique can bring tangible performance gains without negatively impacting the performance of concurrently running applications. Samer Al-Kiswany, Abdullah Gharaibeh, Matei Ripeanu |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2012 | A yoke of oxen and a thousand chickens for heavy lifting graph processingabstractLarge, real-world graphs are famously difficult to process efficiently. Not only they have a large memory footprint but most graph processing algorithms entail memory access patterns with poor locality, data-dependent parallelism, and a low compute-to- memory access ratio. Additionally, most real-world graphs have a low diameter and a highly heterogeneous node degree distribution. Partitioning these graphs and simultaneously achieve access locality and load-balancing is difficult if not impossible. Abdullah Gharaibeh, Lauro Beltrão Costa, Elizeu Santos-Neto, Matei Ripeanu |
PACT | 4 |
| 2012 | A Workflow-Aware Storage System: An Opportunity StudyabstractThis paper evaluates the potential gains a workflow-aware storage system can bring. Two observations make us believe such storage system is crucial to efficiently support workflow-based applications: First, workflows generate irregular and application-dependent data access patterns. These patterns render existing storage systems unable to harness all optimization opportunities as this often requires conflicting optimization options or even conflicting design decision at the level of the storage system. Second, when scheduling, workflow runtime engines make suboptimal decisions as they lack detailed data location information. This paper discusses the feasibility, and evaluates the potential performance benefits brought by, building a workflow-aware storage system that supports per-file access optimizations and exposes data location. To this end, this paper presents approaches to determine the application-specific data access patterns, and evaluates experimentally the performance gains of a workflow-aware storage approach. Our evaluation using synthetic benchmarks shows that a workflow-aware storage system can bring significant performance gains: up to 7× performance gain compared to the distributed storage system - MosaStore and up to 16× compared to a central, well provisioned, NFS server. Emalayan Vairavanathan, Samer Al-Kiswany, Lauro Beltrão Costa, Zhao Zhang 0007, Daniel S. Katz, Michael Wilde, Matei Ripeanu |
CCGRID | 7 |
| 2012 | CloudDT: Efficient tape resource management using deduplication in cloud backup and archival services
Abdullah Gharaibeh, Cornel Constantinescu, Maohua Lu, Ramani Routray, Prasenjit Sarkar, David Pease, Matei Ripeanu |
CNSM | 8 |
| 2012 | Branded with a scarlet "C": cheaters in a gaming social networkabstractOnline gaming is a multi-billion dollar industry that entertains a large, global population. One unfortunate phenomenon, however, poisons the competition and the fun: cheating. The costs of cheating span from industry-supported expenditures to detect and limit cheating, to victims' monetary losses due to cyber crime. This paper studies cheaters in the Steam Community, an online social network built on top of the world's dominant digital game delivery platform. We collected information about more than 12 million gamers connected in a global social network, of which more than 700 thousand have their profiles flagged as cheaters. We also collected in-game interaction data of over 10 thousand players from a popular multiplayer gaming server. We show that cheaters are well embedded in the social and interaction networks: their network position is largely indistinguishable from that of fair players. We observe that the cheating behavior appears to spread through a social mechanism: the presence and the number of cheater friends of a fair player is correlated with the likelihood of her becoming a cheater in the future. Also, we observe that there is a social penalty involved with being labeled as a cheater: cheaters are likely to switch to more restrictive privacy settings once they are tagged and they lose more friends than fair players. Finally, we observe that the number of cheaters is not correlated with the geographical, real-world population density, or with the local popularity of the Steam Community. Jeremy Blackburn, Ramanuja Simha, Nicolas Kourtellis, Xiang Zuo, Matei Ripeanu, John Skvoretz, Adriana Iamnitchi |
WWW | 5 |
| 2011 | The socialbot network: when bots socialize for fame and moneyabstractOnline Social Networks (OSNs) have become an integral part of today's Web. Politicians, celebrities, revolutionists, and others use OSNs as a podium to deliver their message to millions of active web users. Unfortunately, in the wrong hands, OSNs can be used to run astroturf campaigns to spread misinformation and propaganda. Such campaigns usually start off by infiltrating a targeted OSN on a large scale. In this paper, we evaluate how vulnerable OSNs are to a large-scale infiltration by socialbots: computer programs that control OSN accounts and mimic real users. We adopt a traditional web-based botnet design and built a Socialbot Network (SbN): a group of adaptive socialbots that are orchestrated in a command-and-control fashion. We operated such an SbN on Facebook---a 750 million user OSN---for about 8 weeks. We collected data related to users' behavior in response to a large-scale infiltration where socialbots were used to connect to a large number of Facebook users. Our results show that (1) OSNs, such as Facebook, can be infiltrated with a success rate of up to 80%, (2) depending on users' privacy settings, a successful infiltration can result in privacy breaches where even more users' data are exposed when compared to a purely public access, and (3) in practice, OSN security defenses, such as the Facebook Immune System, are not effective enough in detecting or stopping a large-scale infiltration as it occurs. Yazan Boshmaf, Ildar Muslukhov, Konstantin Beznosov, Matei Ripeanu |
ACSAC | 4 |
| 2011 | Failure Avoidance through Fault Prediction Based on Synthetic TransactionsabstractSystem logs are an important tool in studying the conditions (e.g., environment misconfigurations, resource status, erroneous user input) that cause failures. However, production system logs are complex, verbose, and lack structural stability over time. These traits make them hard to use, and make solutions that rely on them susceptible to high maintenance costs. Additionally, logs record failures after they occur: by the time logs are investigated, users have already experienced the failures' consequences. To detect the environment conditions that are correlated with failures without dealing with the complexities associated with processing production logs, and to prevent failure-causing conditions from occurring before the system goes live, this research suggests a three step methodology: (i) using synthetic transactions, i.e., simplified workloads, in pre-production environments that emulate user behavior, (ii) recording the result of executing these transactions in logs that are compact, simple to analyze, stable over time, and specifically tailored to the fault metrics of interest, and (iii) mining these specialized logs to understand the conditions that correlate to failures. This allows system administrators to configure the system to prevent these conditions from happening. We evaluate the effectiveness of this approach by replicating the behavior of a service used in production at Microsoft, and testing the ability to predict failures using a synthetic workload on a 650 million events production trace. The synthetic prediction system is able to predict 91% of real production failures using 50-fold fewer transactions and logs that are 10,000-fold more compact than their production counterparts. Mohammed Shatnawi, Matei Ripeanu |
CCGRID | 2 |
| 2011 | VMFlock: virtual machine co-migration for the cloudabstractThis paper presents VMFlockMS, a migration service optimized for cross-datacenter transfer and instantiation of groups of virtual machine (VM) images that comprise an application-level solution (e.g., a three-tier web application). We dub these groups of related VM images VMFlocks. VMFlockMS employs two main techniques: first, data deduplication within the VMFlock to be migrated and between the VMFlock and the data already present at the destination datacenter, and, second, accelerated instantiation of the application at the target datacenter after transferring only a partial set of data blocks and prioritization of the remaining data based on previously observed access patterns originating from the running VMs. VMFlockMS is designed to be deployed as a set of virtual appliances which make efficient use of the available cloud resources to locally access and deduplicate the images and data in a distributed fashion with minimal requirements imposed on the cloud API to access the VM image repository. VMFlockMS provides an incrementally scalable and high-performance migration service. Our evaluation shows that VMFlockMS can reduce the data volumes to be transferred over the network to as low as 3% of the original VMFlock size, enables the complete transfer of the VM images belonging to a VMFlock over transcontinental link up to 3.5x faster than alternative approaches, and enables booting these VM images with as little as 5% of the compressed VMFlock data available at the destination. Samer Al-Kiswany, Dinesh Subhraveti, Prasenjit Sarkar, Matei Ripeanu |
HPDC | 4 |
| 2011 | Authorization recycling in hierarchical RBAC systemsabstractAs distributed applications increase in size and complexity, traditional authorization architectures based on a dedicated authorization server become increasingly fragile because this decision point represents a single point of failure and a performance bottleneck. Authorization caching, which enables the reuse of previous authorization decisions, is one technique that has been used to address these challenges. This article introduces and evaluates the mechanisms for authorization “recycling” in RBAC enterprise systems. The algorithms that support these mechanisms allow making precise and approximate authorization decisions, thereby masking possible failures of the authorization server and reducing its load. We evaluate these algorithms analytically as well as using simulation and a prototype implementation. Our evaluation results demonstrate that authorization recycling can improve the performance of distributed-access control mechanisms. Jason Crampton, Konstantin Beznosov, Matei Ripeanu |
ACM Trans. Inf. Syst. Secur. | 4 |
| 2011 | ThriftStore: Finessing Reliability Trade-Offs in Replicated Storage SystemsabstractThis paper explores the feasibility of a storage architecture that offers the reliability and access performance characteristics of a high-end system, yet is cost-efficient. We propose ThriftStore, a storage architecture that integrates two types of components: volatile, aggregated storage and dedicated, yet low-bandwidth durable storage. On the one hand, the durable storage forms a back end that enables the system to restore the data the volatile nodes may lose. On the other hand, the volatile nodes provide a high-throughput front-end. Although integrating these components has the potential to offer a unique combination of high throughput and durability at a low cost, a number of concerns need to be addressed to architect and correctly provision the system. To this end, we develop analytical and simulation-based tools to evaluate the impact of system characteristics (e.g., bandwidth limitations on the durable and the volatile nodes) and design choices (e.g., the replica placement scheme) on data availability and the associated system costs (e.g., maintenance traffic). Moreover, to demonstrate the high-throughput properties of the proposed architecture, we prototype a GridFTP server based on ThriftStore. Our evaluation demonstrates an impressive, up to 800 Mbps transfer throughput for the new GridFTP service. Abdullah Gharaibeh, Samer Al-Kiswany, Matei Ripeanu |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2011 | The Small World of File SharingabstractWeb caches, content distribution networks, peer-to-peer file-sharing networks, distributed file systems, and data grids all have in common that they involve a community of users who use shared data. In each case, overall system performance can be improved significantly by first identifying and then exploiting the structure of community's data access patterns. We propose a novel perspective for analyzing data access workloads that considers the implicit relationships that form among users based on the data they access. We propose a new structure-the interest-sharing graph-that captures common user interests in data and justify its utility with studies on four data-sharing systems: a high-energy physics collaboration, the Web, the Kazaa peer-to-peer network, and a BitTorrent file-sharing community. We find small-world patterns in the interest-sharing graphs of all four communities. We investigate analytically and experimentally some of the potential causes that lead to this pattern and conclude that user preferences play a major role. The significance of small-world patterns is twofold: it provides a rigorous support to intuition and it suggests the potential to exploit these naturally emerging patterns. As a proof of concept, we design and evaluate an information dissemination system that exploits the small-world interest-sharing graphs by building an interest-aware network overlay. We show that this approach leads to improved information dissemination performance. Adriana Iamnitchi, Matei Ripeanu, Elizeu Santos-Neto, Ian T. Foster |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2010 | A GPU accelerated storage systemabstractMassively multicore processors, like, for example, Graphics Processing Units (GPUs), provide, at a comparable price, a one order of magnitude higher peak performance than traditional CPUs. This drop in the cost of computation, as any order-of-magnitude drop in the cost per unit of performance for a class of system components, triggers the opportunity to redesign systems and to explore new ways to engineer them to recalibrate the cost-to-performance relation. Abdullah Gharaibeh, Samer Al-Kiswany, Sathish Gopalakrishnan, Matei Ripeanu |
HPDC | 4 |
| 2010 | Efficient and Spontaneous Privacy-Preserving Protocol for Secure Vehicular CommunicationabstractThis paper introduces an efficient and spontaneous privacy-preserving protocol for vehicular ad-hoc networks based on revocable ring signature. The proposed protocol has three appealing characteristics: First, it offers conditional privacy-preservation: while a receiver can verify that a message issuer is an authorized participant in the system only a trusted authority can reveal the true identity of a message sender. Second, it is spontaneous: safety messages can be authenticated locally, without support from the roadside units or contacting other vehicles. Third, it is efficient: it offers fast message authentication and verification, cost-effective identity tracking in case of a dispute, and has low storage requirements. We use extensive analysis to demonstrate the merits of the proposed protocol and to compare it with previously proposed solutions. Hu Xiong, Konstantin Beznosov, Zhiguang Qin, Matei Ripeanu |
ICC | 4 |
| 2010 | Size Matters: Space/Time Tradeoffs to Improve GPGPU Applications PerformanceabstractGPUs offer drastically different performance characteristics compared to traditional multicore architectures. To explore the tradeoffs exposed by this difference, we refactor MUMmer, a widely-used, highly-engineered bioinformatics application which has both CPU- and GPU-based implementations. We synthesize our experience as three high-level guidelines to design efficient GPU-based applications. First, minimizing the communication overheads is as important as optimizing the computation. Second, trading-off higher computational complexity for a more compact in-memory representation is a valuable technique to increase overall performance (by enabling higher parallelism levels and reducing transfer overheads). Finally, ensuring that the chosen solution entails low pre- and post-processing overheads is essential to maximize the overall performance gains. Based on these insights, MUMmerGPU++, our GPU-based design of the MUMmer sequence alignment tool, achieves, on realistic workloads, up to 4× speedup compared to a previous, highly optimized GPU port. Abdullah Gharaibeh, Matei Ripeanu |
SC | 2 |
| 2010 | In search of simplicity: a self-organizing group communication overlayabstractAbstract Group communication primitives have broad utility as building blocks for distributed applications. The challenge is to create and maintain the distributed structures that support these primitives while accounting for volatile end‐nodes and variable network characteristics. Most solutions proposed to date rely on complex algorithms or on global information, thus limiting the scale of deployments and acceptance outside the academic realm. This article introduces a low‐complexity, self‐organizing solution for building and maintaining data dissemination trees, which we refer to as Unstructured Multi‐source Overlay (UMO). UMO uses traditional distributed systems techniques: layering, soft‐state, and passive data collection to adapt to the dynamics of the physical network and maintain data dissemination trees. The result is a simple, adaptive system with lower overheads than more complex alternatives. We implemented UMO and evaluated it on a 100‐node PlanetLab testbed and on up to 1024‐node emulated ModelNet networks. Extensive experimental evaluations demonstrate UMOs low overhead, efficient network usage compared with alternative solutions, and the ability to quickly adapt to network changes and to recover from failures. Copyright © 2009 John Wiley & Sons, Ltd. Matei Ripeanu, Adriana Iamnitchi, Ian T. Foster, Anne Rogers |
Concurr. Comput. Pract. Exp. | 1 |
| 2009 | Introduction
Thomas Fahringer, Alexandru Iosup, Marian Bubak, Matei Ripeanu, Xian-He Sun, Hong Linh Truong 0001 |
Euro-Par | 4 |
| 2009 | Exploring data reliability tradeoffs in replicated storage systemsabstractThis paper explores the feasibility of a cost-efficient storage architecture that offers the reliability and access performance characteristics of a high-end system. This architecture exploits two opportunities: First, scavenging idle storage from LAN-connected desktops not only offers a low-cost storage space, but also high I/O throughput by aggregating the I/O channels of the participating nodes. Second, the two components of data reliability - durability and availability - can be decoupled to control overall system cost. To capitalize on these opportunities, we integrate two types of components: volatile, scavenged storage and dedicated, yet low-bandwidth durable storage. On the one hand, the durable storage forms a low-cost back-end that enables the system to restore the data the volatile nodes may lose. On the other hand, the volatile nodes provide a high-throughput front-end. Abdullah Gharaibeh, Matei Ripeanu |
HPDC | 2 |
| 2009 | GPU support for batch oriented workloadsabstractThis paper explores the ability to use graphics processing units (GPUs) as co-processors to harness the inherent parallelism of batch operations in systems that require high performance. To this end we have chosen bloom filters (space-efficient data structures that support the probabilistic representation of set membership) as the queries these data structures support are often performed in batches. Bloom filters exhibit low computational cost per amount of data, providing a baseline for more complex batch operations. We implemented BloomGPU a library that supports offloading bloom filter support to the GPU and evaluate this library under realistic usage scenarios. By completely offloading Bloom filter operations to the GPU, BloomGPU outperforms an optimized CPU implementation of the bloom filter as the workload becomes larger. Lauro Beltrão Costa, Samer Al-Kiswany, Matei Ripeanu |
IPCCC | 3 |
| 2009 | Resource demand and supply in BitTorrent content-sharing communities
Nazareno Andrade, Elizeu Santos-Neto, Francisco Vilar Brasileiro, Matei Ripeanu |
Comput. Networks | 4 |
| 2009 | Beyond Music Sharing: An Evaluation of Peer-to-Peer Data Dissemination Techniques in Large Scientific Collaborations
Samer Al-Kiswany, Matei Ripeanu, Adriana Iamnitchi, Sudharshan S. Vazhkudai |
J. Grid Comput. | 2 |
| 2009 | The Globus Replica Location Service: Design and ExperienceabstractDistributed computing systems employ replication to improve overall system robustness, scalability, and performance. A replica location service (RLS) offers a mechanism to maintain and provide information about physical locations of replicas. This paper defines a design framework for RLSs that supports a variety of deployment options. We describe the RLS implementation that is distributed with the Globus toolkit and is in production use in several grid deployments. Features of our modular implementation include the use of soft-state protocols to populate a distributed index and Bloom filter compression to reduce overheads for distribution of index information. Our performance evaluation demonstrates that the RLS implementation scales well for individual servers with millions of entries and up to 100 clients. We describe the characteristics of existing RLS deployments and discuss how RLS has been integrated with higher-level data management services. Ann L. Chervenak, Robert Schuler, Matei Ripeanu, Muhammad Ali Amer, Shishir Bharathi, Ian T. Foster, Adriana Iamnitchi, Carl Kesselman |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2009 | Cooperative Secondary Authorization RecyclingabstractAs enterprise systems, Grids, and other distributed applications scale up and become increasingly complex, their authorization infrastructures--based predominantly on the request-response paradigm--are facing the challenges of fragility and poor scalability. We propose an approach where each application server recycles previously received authorizations and shares them with other application servers to mask authorization server failures and network delays. This paper presents the design of our cooperative secondary authorization recycling system and its evaluation using simulation and prototype implementation. The results demonstrate that our approach improves the availability and performance of authorization infrastructures. Specifically, by sharing authorizations, the cache hit rate--an indirect metric of availability--can reach 70 percent, even when only 10 percent of authorizations are cached. Depending on the deployment scenario, the average time for authorizing an application request can be reduced by up to a factor of two compared with systems that do not employ cooperation. Matei Ripeanu, Konstantin Beznosov |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2008 | StoreGPU: exploiting graphics processing units to accelerate distributed storage systemsabstractToday Graphics Processing Units (GPUs) are a largely underexploited resource on existing desktops and a possible cost-effective enhancement to high-performance systems. To date, most applications that exploit GPUs are specialized scientific applications. Little attention has been paid to harnessing these highly-parallel devices to support more generic functionality at the operating system or middleware level. This study starts from the hypothesis that generic middleware level techniques that improve distributed system reliability or performance (such as content addressing, erasure coding, or data similarity detection) can be significantly accelerated using GPU support.We take a first step towards validating this hypothesis, focusing on distributed storage systems. As a proof of concept, we design StoreGPU, a library that accelerates a number of hashing based primitives popular in distributed storage system implementations. Our evaluation shows that StoreGPU enables up to eight-fold performance gains on synthetic benchmarks as well as on a high-level application: the online similarity detection between large data files. Samer Al-Kiswany, Abdullah Gharaibeh, Elizeu Santos-Neto, George Yuan, Matei Ripeanu |
HPDC | 5 |
| 2008 | enabling cross-layer optimizations in storage systems with custom metadataabstractToday, several data-storage systems allow applications to create and manage custom metadata to improve data search and navigability in large scale storage systems. Elizeu Santos-Neto, Samer Al-Kiswany, Nazareno Andrade, Sathish Gopalakrishnan, Matei Ripeanu |
HPDC | 5 |
| 2008 | stdchk: A Checkpoint Storage System for Desktop Grid ComputingabstractCheckpointing is an indispensable technique to provide fault tolerance for long-running high-throughput applications like those running on desktop grids. This article argues that a checkpoint storage system, optimized to operate in these environments, can offer multiple benefits: reduce the load on a traditional file system, offer high-performance through specialization, and, finally, optimize data management by taking into account checkpoint application semantics. Such a storage system can present a unifying abstraction to checkpoint operations, while hiding the fact that there are no dedicated resources to store the checkpoint data. We prototype stdchk, a checkpoint storage system that uses scavenged disk space from participating desktops to build a low-cost storage system, offering a traditional file system interface for easy integration with applications. This article presents the stdchk architecture, key performance optimizations, and its support for incremental checkpointing and increased data availability. Our evaluation confirms that the stdchk approach is viable in a desktop grid setting and offers a low cost storage system with desirable performance characteristics: high write throughput as well as reduced storage space and network effort to save checkpoint images. Samer Al-Kiswany, Matei Ripeanu, Sudharshan S. Vazhkudai, Abdullah Gharaibeh |
ICDCS | 2 |
| 2008 | Authorization Using the Publish-Subscribe ModelabstractTraditional authorization mechanisms based on the request-response model are generally supported by point-to-point communication between applications and authorization servers. As distributed applications increase in size and complexity, an authorization architecture based on point-to-point communication becomes fragile and difficult to manage. This paper presents the use of the publish-subscribe (pub-sub) model for delivering authorization requests and responses between the applications and the authorization servers. Our analysis suggests that using the pub-sub architecture improves authorization system availability and reduces system administration overhead. We evaluate our design using a prototype implementation, which confirms the improvement in availability. Although the response time is also increased, this impact can be reduced by bypassing the pub-sub channel when returning authorizations or by caching coupled with local inference of authorization decisions based on previously cached authorizations. Matei Ripeanu, Konstantin Beznosov |
ISPA | 2 |
| 2008 | Authorization recycling in RBAC systemsabstractAs distributed applications increase in size and complexity, traditional authorization mechanisms based on a single policy decision point are increasingly fragile because this decision point represents a single point of failure and a performance bottleneck. Authorization recycling is one technique that has been used to address these challenges. This paper introduces and evaluates the mechanisms for authorization recycling in RBAC enterprise systems. The algorithms that support these mechanisms allow precise and approximate authorization decisions to be made, thereby masking possible failures of the policy decision point and reducing its load. We evaluate these algorithms analytically and using a prototype implementation. Our evaluation results demonstrate that authorization recycling can improve the performance of distributed access control mechanisms. Jason Crampton, Konstantin Beznosov, Matei Ripeanu |
SACMAT | 4 |
| 2007 | Are P2P Data-Dissemination Techniques Viable in Today's Data-Intensive Scientific Collaborations?
Samer Al-Kiswany, Matei Ripeanu, Adriana Iamnitchi, Sudharshan S. Vazhkudai |
Euro-Par | 2 |
| 2007 | Cooperative secondary authorization recyclingabstractAs distributed applications such as Grid and enterprise systems scale up and become increasingly complex, their authorization infrastructures-based predominantly on the request-response paradigm-are facing challenges in terms of fragility and poor scalability. We propose an approach where each application server caches previously received authorizations at its secondary decision point and shares them with other application servers to mask authorization server failures and network delays.This paper presents the design of our cooperative secondary authorization recycling system and its evaluation using simulation and prototype implementation. The results demonstrate that our approach improves the availability of authorization infrastructures while preserving their performance characteristics. Specifically, by sharing authorizations, the cache hit rate.an indirect metric of availability.can reach 70%, even when only 10% of authorizations are cached. Depending on the deployment scenario, the performance in terms of the average time for authorizing an application request can be reduced by up to 30%. Matei Ripeanu, Konstantin Beznosov |
HPDC | 2 |
| 2006 | The Design, Performance, and Use of DiPerF: An automated DIstributed PERformance evaluation Framework
Ioan Raicu, Catalin Dumitrescu, Matei Ripeanu, Ian T. Foster |
J. Grid Comput. | 3 |
| 2006 | A Layered Framework for Connecting Client Objectives and Resource CapabilitiesabstractIn large-scale, distributed systems such as Grids, an agreement between a client and a service provider specifies service level objectives both as expressions of client requirements and as provider assurances. From an application perspective, these objectives should be expressed in a high-level, service or application-specific manner rather than requiring clients to detail the necessary resources. Resource providers on the other hand, expect low-level, resource-specific performance criteria that are uniform across applications and can be easily interpreted and provisioned. This paper presents a framework for service management that addresses this gap between high-level specification of client performance objectives and existing resource management infrastructures. The paper identifies three levels of abstraction for resource requirements a service provider needs to manage, namely: detailed specification of raw resources, virtualization of heterogeneous resources as abstract resources, and performance objectives at an application level. The paper also identifies three key functions for managing service-level agreements, namely: translation of resource requirements across abstraction layers, arbitration in allocating resources to client requests, and aggregation and allocation of resources from multiple lower-level resource managers. One or more of these key functions may be present at each abstraction layer of a service-level manager. Thus, layering and the composition of these functions across abstraction layers enables modeling of a wide array of management scenarios. The framework we present uses service metadata and/or service performance models to map client requirements to resource capabilities, uses business value associated with objectives to arbitrate between competing requests, and allocates resources based on previously negotiated agreements. We instantiate this framework for three different scenarios and explain how the architectural principles we introduce are used in the real-word. Asit Dan, Kavitha Ranganathan, Catalin Dumitrescu, Matei Ripeanu |
Int. J. Cooperative Inf. Syst. | 4 |
| 2004 | Incentive mechanisms for large collaborative resource sharingabstractWe study the nature of sharing resources in distributed collaborations such as Grids and peer-to-peer systems. By applying the theoretical framework of the multi-person prisoner's dilemma to this resource sharing problem, we show that in the absence of incentive schemes, individual users are apt to hold back resources, leading to decreased system utility. Using both the theoretical framework as well as simulations, we compare and contrast three different incentive schemes aimed at encouraging users to contribute resources. Our results show that soft-incentive schemes are effective in incentivizing autonomous entities to collaborate, leading to increased gains for all participants in the system. Kavitha Ranganathan, Matei Ripeanu, A. Sarin, Ian T. Foster |
CCGRID | 2 |
| 2004 | Cache replacement policies revisited: the case of P2P trafficabstractPeer-to-peer (P2P) file-sharing applications generate a large part if not most of today's Internet traffic. The large volume of this traffic (thus the high potential benefits of caching) and the large cache sizes required (thus nontrivial costs associated with caching) only underline that efficient cache replacement policies are important in this case. P2P file-sharing traffic has several characteristics that distinguish it from well studied Web traffic and that require a focused study of efficient cache management policies. This paper uses trace driven simulations to compare traditional cache replacement policies with new policies that try to exploit characteristics of the P2P file-sharing traffic generated by applications using the FastTrack protocol. Adam Wierzbicki, Nathaniel Leibowitz, Matei Ripeanu, Rafal Wozniak |
CCGRID | 3 |
| 2004 | Globus and PlanetLab Resource Management Solutions Compared
Matei Ripeanu, Mic Bowman, Jeffrey S. Chase, Ian T. Foster, Milan Milenkovic |
HPDC | 1 |
| 2004 | Connecting client objectives with resource capabilities: an essential component for grid service managent infrastructuresabstractIn large-scale, distributed systs such as Grids, an agreent between a client and a service provider specifies service level objectives both as expressions of client requirents and as provider assurances. Ideally, these objectives are expressed in a high-level, service- or application-specific manner rather than requiring clients to detail the necessary resources. Resource providers on the other hand, expect low-level, resource specific performance criteria that are uniform across applications and can easily be interpreted and provisioned. Asit Dan, Catalin Dumitrescu, Matei Ripeanu |
ICSOC | 3 |
| 2004 | Small-World File-Sharing CommunitiesabstractWeb caches, content distribution networks, peer-to-peer file sharing networks, distributed file systems, and data grids all have in common that they involve a community of users who generate requests for shared data. In each case, overall system performance can be improved significantly if we can first identify and then exploit interesting structure within a community's access patterns. To this end, we propose a novel perspective on file sharing that considers the relationships that form among users based on the files in which they are interested. We propose a new structure that captures common user interests in data - the data-sharing graph - and justify its utility with studies on three data-distribution systems: a high-energy physics collaboration, the Web, and the Kazaa peer-to-peer network. We find small-world patterns in the data-sharing graphs of all three communities. We analyze these graphs and propose some probable causes for these emergent small-world patterns. The significance of small-world patterns is twofold: it provides a rigorous support to intuition and, perhaps most importantly, it suggests ways to design mechanisms that exploit these naturally emerging patterns. Adriana Iamnitchi, Matei Ripeanu, Ian T. Foster |
INFOCOM | 2 |
| 2002 | A Decentralized, Adaptive Replica Location MechanismabstractWe describe a decentralized, adaptive mechanism for replica location in wide-area distributed systems. Unlike traditional, hierarchical (e.g, DNS) and more recent (e.g., CAN, Chord, Gnutella) distributed search and indexing schemes, nodes in our location mechanism do not route queries, instead, they organize into an overlay network and distribute location information. We contend that this approach works well in environments where replica location queries are prevalent but the dynamic component of the system (e.g., node and network failures, replica add/delete operations) cannot be neglected. We argue that a replica location mechanism that combines probabilistic representations of replica location information with soft-state protocols and a flat overlay network of nodes brings important benefits: genuine decentralization, low query latency, and flexibility to introduce adaptive communication schedules. We support these claims in two ways. First, we provide a rough resource consumption evaluation: we show that, for environments similar to those encountered in large scientific data analysis projects, generated network traffic is limited and, more importantly, is comparable to the traffic generated by a request routing scheme. Second, we provide encouraging performance data from a prototype implementation. Matei Ripeanu, Ian T. Foster |
HPDC | 1 |
| 2002 | Giggle: a framework for constructing scalable replica location servicesabstractIn wide area computing systems, it is often desirable to create remote read-only copies (replicas) of files. Replication can be used to reduce access latency, improve data locality, and/or increase robustness, scalability and performance for distributed applications. We define a replica location service (RLS) as a system that maintains and provides access to information about the physical locations of copies. An RLS typically functions as one component of a data grid architecture. This paper makes the following contributions. First, we characterize RLS requirements. Next, we describe a parameterized architectural framework, which we name Giggle (for GIGa-scale Global Location Engine), within which a wide range of RLSs can be defined. We define several concrete instantiations of this framework with different performance characteristics. Finally, we present initial performance results for an RLS prototype, demonstrating that RLS systems can be constructed that meet performance goals. Ann L. Chervenak, Ewa Deelman, Ian T. Foster, Leanne Guy, Wolfgang Hoschek, Adriana Iamnitchi, Carl Kesselman, Peter Z. Kunszt, Matei Ripeanu, Robert Schwartzkopf, Heinz Stockinger, Kurt Stockinger, Brian Tierney |
SC | 9 |
| 2001 | Cactus Application: Performance Predictions in Grid Environments
Matei Ripeanu, Adriana Iamnitchi, Ian T. Foster |
Euro-Par | 1 |
| 2001 | Peer-to-Peer Architecture Case Study: Gnutella NetworkabstractDespite recent excitement generated by the P2P paradigm and despite surprisingly fast deployment of some P2P applications, there are few quantitative evaluations of P2P system behavior. Due to its open architecture and achieved scale, Gnutella is an interesting P2P architecture case study. Gnutella, like most other P2P applications, builds at the application level a virtual network with its own routing mechanisms. The topology of this overlay network and the routing mechanisms used have a significant influence on application properties such as performance, reliability, and scalability. We built a 'crawler' to extract the topology of Gnutella's application level network, we analyze the topology graph and evaluate generated network traffic. We find that although Gnutella is not a pure power-law network, its current configuration has the benefits and drawbacks of a power-law structure. These findings lead us to propose changes to the Gnutella protocol and implementations that bring significant performance and scalability improvements. Matei Ripeanu |
Peer-to-Peer Computing | 1 |
| 2001 | Supporting efficient execution in heterogeneous distributed computing environments with cactus and globusabstractImprovements in the performance of processors and networks make it both feasible and interesting to treat collections of workstations, servers, clusters, and supercomputers as integrated computational resources, or Grids. However, the highly heterogeneous and dynamic nature of such Grids can make application development difficult. Here we describe an architecture and prototype implementation for a Grid-enabled computational framework based on Cactus, the MPICH-G2 Grid-enabled message-passing library, and a variety of specialized features to support efficient execution in Grid environments. We have used this framework to perform record-setting computations in numerical relativity, running across four supercomputers and achieving scaling of 88% (1140 CPU's) and 63% (1500 CPUs). The problem size we were able to compute was about five times larger than any other previous run. Further, we introduce and demonstrate adaptive methods that automatically adjust computational parameters during run time, to increase dramatically the efficiency of a distributed Grid simulation, without modification of the application and without any knowledge of the underlying network connecting the distributed computers. Gabrielle Allen, Thomas Dramlitsch, Ian T. Foster, Nicholas T. Karonis, Matei Ripeanu, Edward Seidel, Brian R. Toonen |
SC | 5 |