Raja Appuswamy

dblp:57/1978 · DBLP profile ↗
← Back
38ranked-venue papers
13as first author
13since 2021 · last 2026
0000-0001-5887-4091ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 16 · 5 first-author · 4 since 2021Systems, architecture and hardware · 14 · 7 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 since 2021Security and privacy · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 X-BQSR: Holistic Acceleration of Base Quality Recalibration for Scalable Genomic Analysis
Eugenio Marinelli, Ivan Donchev Kabadzhov, Raja Appuswamy
Euro-Par (1)3
2025 COSE: Continuous Shapes Extraction from Dynamic Knowledge Graphs
Baya Dhouib, Raja Appuswamy
IEEE Big Data2
2025 BMPipe: Bubble-Memory Co-Optimization Strategy Planner for Very-Large DNN Training
abstract
Pipeline parallelism and activation recomputation are widely adopted optimization techniques, among others, to scale DNN training on large accelerator clusters. However, as DNNs grow in complexity and heterogeneity, it becomes increasingly difficult to determine the optimal combination of pipeline partitioning and recomputation strategies. Existing solutions either propose manual optimization approaches that do not scale or automated approaches that explore only a subset of optimization possibilities due to an explosion of search space. In this paper, we present BMPipe, a bubble-memory co-optimization planner that holistically optimizes computation imbalance, memory under utilization, redundant computation, and schedulinginduced preparation time. At its core, BMPipe uses symbolic representations that unify computation, memory, and bubbles into a single model that is solved by using an ILP-based planner. Using BMPipe, we perform a thorough experimental evaluation where we train several large, state-of-the-art DNN models on a 16K-NPU cluster. We show that BMPipe achieves up to$1.36 \times$speedup compared to the state-of-the-art solution Megatron. Against automatic planners PipeDream, Merak and AdaPipe,, it yields as$1.27 \times$speed-up. In addition, BMPipe boosts peak device-memory utilization by$\mathbf{1. 4 2} \times$compared with Megatron.
Ruiwen Wang, Chong Li 0003, Thibaut Tachon, Raja Appuswamy, Teng Su
CLUSTER4
2025 ManuMatic: Strategy Injection for Robust Automatic Hybrid Parallelism in Distributed DNN Training
Ruiwen Wang, Chong Li 0003, Raja Appuswamy, Yujie Yuan
NPC (2)4
2025 Adaptive Sampling For Storage Of Progressive Images On Dna
abstract
International audience
Xavier Pic, Nimesh Pinnamaneni, Raja Appuswamy
PCS3
2024 CMOSS: A Reliable, Motif-based Columnar Molecular Storage System
abstract
The surge in demand for cost-effective long-term archival media, coupled with density limitations of contemporary magnetic media, has resulted in synthetic DNA emerging as a promising new alternative. Despite its benefits, storing data in DNA poses several challenges as the technology used for reading/writing data on DNA are highly error prone. Thus, it is important to design pipelines that can efficiently use redundancy to mask errors without amplifying read/write cost. In this work, we present Columnar MOlecular Storage System (CMOSS), a novel, end-to-end DNA storage pipeline that can provide error-tolerant data storage at low read/write costs. CMOSS differs from state-of-the-art (SOTA) on three fronts (i) a motif-based, vertical layout in contrast to nucleotide-based horizontal layout, (ii) integrated consensus calling and decoding enabled by the vertical layout, and (iii) a flexible, block-based data organization for random access over DNA storage in contrast to object-based organization. Using an in-depth evaluation with several simulated and real wet lab experiments, we demonstrate the benefits of CMOSS design.
Eugenio Marinelli, Yiqing Yan, Lorenzo Tattini, Virginie Magnone, Pascal Barbry, Raja Appuswamy
SYSTOR6
2023 Towards Migration-Free Just-In-Case Data Archival for Future Cloud Data Lakes
abstract
Given the growing adoption of AI, cloud data lakes are facing the need to support cost-effective "just-in-case" data archival over long time periods to meet regulatory compliance requirements. Unfortunately, current media technologies suffer from fundamental issues that will soon, if not already, make cost-effective data archival infeasible. In this paper, we present a vision for redesigning the archival tier of cloud data lakes based on a novel, obsolescence-free storage medium-synthetic DNA. In doing so, we make two contributions: (i) we highlight the challenges in using DNA for data archival and list several open research problems, (ii) we outline OligoArchive-DSM (OA-DSM)-an end-to-end DNA storage pipeline that we are developing to demonstrate the feasibility of our vision.
Eugenio Marinelli, Yiqing Yan, Virginie Magnone, Marie-Charlotte Dumargne, Pascal Barbry, Thomas Heinis, Raja Appuswamy
Proc. VLDB Endow.7
2021 Universal Layout Emulation for Long-Term Database Archival
Raja Appuswamy, Vincent Joguin
CIDR1
2021 XJoin: Portable, parallel hash join across diverse XPU architectures with oneAPI
abstract
Modern server hardware is increasingly heterogeneous with a diverse mix of XPU architectures deployed across CPU, GPU, and FPGAs. However, till date, database developers have had to rely on either proprietary, architecture-specific solutions (like CUDA), or low-level, cross-architecture solutions that complicate development (like OpenCL). The lack of portable parallelism caused by the absence of a common high-level programming framework is one of the main reasons preventing a wider adoption of XPUs by database systems.
Eugenio Marinelli, Raja Appuswamy
DaMoN2
2021 Decoding Of Nanopore-Sequenced Synthetic DNA Storing Digital Images
abstract
Digital media explosion has led to an exponential increase of the amount of data generated worldwide and the need for new means of storage able to keep up with the current growth of digital information has become a critical challenge. During the last decade, DNA has been proven to be a potential candidate thanks to its biological properties allowing to store information at high density (215 petabytes in 1 gram) for centuries. In previous works we have presented an end-to-end storage workflow specifically designed for the efficient storage of images onto synthetic DNA and proven its feasibility in a wet-lab experiment in which sequencing was performed using the Illumina machine. In this work we are studying the sequencing using rather the MinION sequencer on the same data after being stored in a sealed capsule for two years. MinION is a very promising sequencer although introducing a much higher error rate in the process of reading. In this paper, we propose a solution to deal with the MinION sequencing noise allowing to recover the original stored data.
Eva Gil San Antonio, Melpomeni Dimopoulou, Marc Antonini, Pascal Barbry, Raja Appuswamy
ICIP5
2021 Generative DNA: Representation Learning for DNA-based Approximate Image Storage
abstract
Synthetic DNA has received much attention recently as a long-term archival medium alternative due to its high density and durability characteristics. However, most current work has primarily focused on using DNA as a precise storage medium. In this work, we take an alternate view of DNA. Using neural-network-based compression techniques, we transform images into a latent-space representation, which we then store on DNA. By doing so, we transform DNA into an approximate image storage medium, as images generated back from DNA are only approximate representations of the original images. Using several datasets, we investigate the storage benefits of approximation, and study the impact of DNA storage errors (substitutions, indels, bias) on the quality of approximation. In doing so, we demonstrate the feasibility and potential of viewing DNA as an approximate storage medium.
Giulio Franzese, Yiqing Yan, Giuseppe Serra 0003, Ivan D'Onofrio, Raja Appuswamy, Pietro Michiardi
VCIP5
2021 Accel-Align: a fast sequence mapper and aligner based on the seed-embed-extend method
abstract
BACKGROUND: Improvements in sequencing technology continue to drive sequencing cost towards $100 per genome. However, mapping sequenced data to a reference genome remains a computationally-intensive task due to the dependence on edit distance for dealing with INDELs and mismatches introduced by sequencing. All modern aligners use seed-filter-extend methodology and rely on filtration heuristics to reduce the overhead of edit distance computation. However, filtering has inherent performance-accuracy trade-offs that limits its effectiveness. RESULTS: Motivated by algorithmic advances in randomized low-distortion embedding, we introduce SEE, a new methodology for developing sequence mappers and aligners. While SFE focuses on eliminating sub-optimal candidates, SEE focuses instead on identifying optimal candidates. To do so, SEE transforms the read and reference strings from edit distance regime to the Hamming regime by embedding them using a randomized algorithm, and uses Hamming distance over the embedded set to identify optimal candidates. To show that SEE performs well in practice, we present Accel-Align an SEE-based short-read sequence mapper and aligner that is 3-12[Formula: see text] faster than state-of-the-art aligners on commodity CPUs, without any special-purpose hardware, while providing comparable accuracy. CONCLUSIONS: As sequencing technologies continue to increase read length while improving throughput and accuracy, we believe that randomized embeddings open up new avenues for optimization that cannot be achieved by using edit distance. Thus, the techniques presented in this paper have a much broader scope as they can be used for other applications like graph alignment, multiple sequence alignment, and sequence assembly.
Yiqing Yan, Nimisha Chaturvedi, Raja Appuswamy
BMC Bioinform.3
2021 Image storage onto synthetic DNA
Melpomeni Dimopoulou, Marc Antonini, Pascal Barbry, Raja Appuswamy
Signal Process. Image Commun.4
2020 Storing Digital Data Into DNA: A Comparative Study Of Quaternary Code Construction
abstract
The exponential increase of digital data that is being generated every year along with the capacity and durability limits of conventional storage devices are raising one of the greatest challenges for the field of data storage. The use of DNA for digital data archiving is a very promising alternative as the biological properties of the DNA molecule allow the storage of a huge amount of information into a very limited volume while also promising data longevity for centuries or even longer. In this paper we present a comparative study of our work with the state of the art solutions, and show that our solution is competitive.
Melpomeni Dimopoulou, Marc Antonini, Pascal Barbry, Raja Appuswamy
ICASSP4
2020 A system design for elastically scaling transaction processing engines in virtualized servers
abstract
Online Transaction Processing (OLTP) deployments are migrating from on-premise to cloud settings in order to exploit the elasticity of cloud infrastructure which allows them to adapt to workload variations. However, cloud adaptation comes at the cost of redesigning the engine, which has led to the introduction of several, new, cloud-based transaction processing systems mainly focusing on: (i) the transaction coordination protocol, (ii) the data partitioning strategy, and, (iii) the resource isolation across multiple tenants. As a result, standalone OLTP engines cannot be easily deployed with an elastic setting in the cloud and they need to migrate to another, specialized deployment. In this paper, we focus on workload variations that can be addressed by modern multi-socket, multi-core servers and we present a system design for providing fine-grained elasticity to multi-tenant, scale-up OLTP deployments. We introduce novel components to the virtualization software stack that enable on-demand addition and removal of computing and memory resources. We provide a bi-directional, low-overhead communication stack between the virtual machine and the hypervisor, which allows the former to adapt to variations coming both from the workload and the resource availability. We show that our system achieves NUMA-aware, millisecond-level, stateful and fine-grained elasticity, while it is not intrusive to the design of state-of-the-art, in-memory OLTP engines. We evaluate our system through novel use cases demonstrating that scale-up elasticity increases resource utilization, while allowing tenants to pay for actual use of resources and not just their reservation.
Angelos-Christos G. Anadiotis, Raja Appuswamy, Anastasia Ailamaki, Ilan Bronshtein, Hillel Avni, David Dominguez-Sal, Shay Goikhman, Eliezer Levy
Proc. VLDB Endow.2
2019 OligoArchive: Using DNA in the DBMS storage hierarchy
Raja Appuswamy, Kevin Le Brigand, Pascal Barbry, Marc Antonini, Olivier Madderson, Paul S. Freemont, James McDonald, Thomas Heinis
CIDR1
2019 Cold Storage Data Archives: More Than Just a Bunch of Tapes
abstract
The abundance of available sensor and derived data from large scientific experiments, such as earth observation programs, radio astronomy sky surveys, and high-energy physics already exceeds the storage hardware globally fabricated per year. To that end, cold storage data archives are the---often overlooked---spearheads of modern big data analytics in scientific, data-intensive application domains. While high-performance data analytics has received much attention from the research community, the growing number of problems in designing and deploying cold storage archives has only received very little attention.
Bunjamin Memishi, Raja Appuswamy, Marcus Paradies
DaMoN2
2019 Taster: Self-Tuning, Elastic and Online Approximate Query Processing
abstract
Current Approximate Query Processing (AQP) engines are far from silver-bullet solutions, as they adopt several static design decisions that target specific workloads and deployment scenarios. Offline AQP engines target deployments with large storage budget, and offer substantial performance improvement for predictable workloads, but fail when new query types appear, i.e., due to shifting user interests. To the other extreme, online AQP engines assume that query workloads are unpredictable, and therefore build all samples at query time, without reusing samples (or parts of them) across queries. Clearly, both extremes miss out on different opportunities for optimizing performance and cost. In this paper, we present Taster, a self-tuning, elastic, online AQP engine that synergistically combines the benefits of online and offline AQP. Taster performs online approximation by injecting synopses (samples and sketches) into the query plan, while at the same time it strategically materializes and reuses synopses across queries, and continuously adapts them to changes in the workload and to the available storage resources. Our experimental evaluation shows that Taster adapts to shifting workload and to varying storage budgets, and always matches or significantly outperforms the state-of-the-art performing AQP approaches (online or offline).
Matthaios Olma, Odysseas Papapetrou, Raja Appuswamy, Anastasia Ailamaki
ICDE3
2019 Hardware-Conscious Hash-Joins on GPUs
abstract
Traditionally, analytical database engines have used task parallelism provided by modern multisocket multicore CPUs for scaling query execution. Over the past few years, GPUs have started gaining traction as accelerators for processing analytical queries due to their massively data-parallel nature and high memory bandwidth. Recent work on designing join algorithms for CPUs has shown that carefully tuned join implementations that exploit underlying hardware can outperform naive, hardware-oblivious counterparts and provide excellent performance on modern multicore servers. However, there has been no such systematic analysis of hardware-conscious join algorithms for GPUs that systematically explores the dimensions of partitioning (partitioned versus non-partitioned joins), data location (data fitting and not fitting in GPU device memory), and access pattern (skewed versus uniform). In this paper, we present the design and implementation of a family of novel, partitioning-based GPU-join algorithms that are tuned to exploit various GPU hardware characteristics for working around the two main limitations of GPUs–limited memory capacity and slow PCIe interface. Using a thorough evaluation, we show that: i) hardware-consciousness plays a key role in GPU joins similar to CPU joins and our join algorithms can process 1 Billion tuples/second even if no data is GPU resident, ii) radix partitioning-based GPU joins that are tuned to exploit GPU hardware can substantially outperform non-partitioned hash joins, iii) hardware-conscious GPU joins can effectively overcome GPU limitations and match, or even outperform, state-of-the-art CPU joins.
Panagiotis Sioulas, Periklis Chrysogelos, Manos Karpathiotakis, Raja Appuswamy, Anastasia Ailamaki
ICDE4
2019 HetExchange: Encapsulating heterogeneous CPU-GPU parallelism in JIT compiled engines
abstract
Modern server hardware is increasingly heterogeneous as hardware accelerators, such as GPUs, are used together with multicore CPUs to meet the computational demands of modern data analytics work-loads. Unfortunately, query parallelization techniques used by analytical database engines are designed for homogeneous multicore servers, where query plans are parallelized across CPUs to process data stored in cache coherent shared memory. Thus, these techniques are unable to fully exploit available heterogeneous hardware, where one needs to exploit task-parallelism of CPUs and data-parallelism of GPUs for processing data stored in a deep, non-cache-coherent memory hierarchy with widely varying access latencies and bandwidth. In this paper, we introduce HetExchange-a parallel query execution framework that encapsulates the heterogeneous parallelism of modern multi-CPU-multi-GPU servers and enables the parallelization of (pre-)existing sequential relational operators. In contrast to the interpreted nature of traditional Exchange, HetExchange is designed to be used in conjunction with JIT compiled engines in order to allow a tight integration with the proposed operators and generation of efficient code for heterogeneous hardware. We validate the applicability and efficiency of our design by building a prototype that can operate over both CPUs and GPUs, and enables its operators to be parallelism- and data-location-agnostic. In doing so, we show that efficiently exploiting CPU-GPU parallelism can provide 2.8x and 6.4x improvement in performance compared to state-of-the-art CPU-based and GPU-based DBMS.
Periklis Chrysogelos, Manos Karpathiotakis, Raja Appuswamy, Anastasia Ailamaki
Proc. VLDB Endow.3
2017 The Case For Heterogeneous HTAP
Raja Appuswamy, Manos Karpathiotakis, Danica Porobic, Anastasia Ailamaki
CIDR1
2017 Analyzing the Impact of System Architecture on the Scalability of OLTP Engines for High-Contention Workloads
abstract
Main-memory OLTP engines are being increasingly deployed on multicore servers that provide abundant thread-level parallelism. However, recent research has shown that even the state-of-the-art OLTP engines are unable to exploit available parallelism for high contention workloads. While previous studies have shown the lack of scalability of all popular concurrency control protocols, they consider only one system architecture---a non-partitioned, shared everything one where transactions can be scheduled to run on any core and can access any data or metadata stored in shared memory. In this paper, we perform a thorough analysis of the impact of other architectural alternatives (Data-oriented transaction execution, Partitioned Serial Execution, and Delegation) on scalability under high contention scenarios. In doing so, we present Trireme, a main-memory OLTP engine testbed that implements four system architectures and several popular concurrency control protocols in a single code base. Using Trireme, we present an extensive experimental study to understand i) the impact of each system architecture on overall scalability, ii) the interaction between system architecture and concurrency control protocols, and iii) the pros and cons of new architectures that have been proposed recently to explicitly deal with high-contention workloads.
Raja Appuswamy, Angelos-Christos G. Anadiotis, Danica Porobic, Mustafa Iman, Anastasia Ailamaki
Proc. VLDB Endow.1
2016 More than a network: distributed OLTP on clusters of hardware islands
abstract
Multisocket multicores feature hardware islands - groups of cores that communicate fast among themselves and slower with other groups. With high speed networking becoming a commodity, clusters of hardware islands with fast networks are becoming a preferred platform for high end OLTP workloads. While behavior of OLTP on multisockets is well understood, multi-machine OLTP deployments have been studied only in the geo-distributed context where network is much slower. In this paper, we analyze the behavior of different OLTP designs when deployed on clusters of multisockets with fast networks.
Danica Porobic, Pinar Tözün, Raja Appuswamy, Anastasia Ailamaki
DaMoN3
2016 OLTP on a server-grade ARM: power, throughput and latency comparison
abstract
Although scaling out of low-power cores is an alternative to power-hungry Intel Xeon processors for reducing the power overheads, they have proven inadequate for complex, non-parallelizable workloads. On the other hand, by the introduction of the 64-bit ARMv8 architecture, traditionally low power ARM processors have become powerful enough to run computationally intensive server-class applications.
Utku Sirin, Raja Appuswamy, Anastasia Ailamaki
DaMoN2
2016 Cheap Data Analytics using Cold Storage Devices
abstract
Enterprise databases use storage tiering to lower capital and operational expenses. In such a setting, data waterfalls from an SSD-based high-performance tier when it is "hot" (frequently accessed) to a disk-based capacity tier and finally to a tape-based archival tier when "cold" (rarely accessed). To address the unprecedented growth in the amount of cold data, hardware vendors introduced new devices named Cold Storage Devices (CSD) explicitly targeted at cold data workloads. With access latencies in tens of seconds and cost/GB as low as $0.01/GB/month, CSD provide a middle ground between the low-latency (ms), high-cost, HDD-based capacity tier, and high-latency (min to h), low-cost, tape-based, archival tier. Driven by the price/performance aspect of CSD, this paper makes a case for using CSD as a replacement for both capacity and archival tiers of enterprise databases. Although CSD offer major cost savings, we show that current database systems can suffer from severe performance drop when CSD are used as a replacement for HDD due to the mismatch between design assumptions made by the query execution engine and actual storage characteristics of the CSD. We then build a CSD-driven query execution framework, called Skipper, that modifies both the database execution engine and CSD scheduling algorithms to be aware of each other. Using results from our implementation of the architecture based on PostgreSQL and OpenStack Swift, we show that Skipper is capable of completely masking the high latency overhead of CSD, thereby opening up CSD for wider adoption as a storage tier for cheap data analytics over cold data.
Renata Borovica, Raja Appuswamy, Anastasia Ailamaki
Proc. VLDB Endow.2
2015 Scaling the Memory Power Wall With DRAM-Aware Data Management
abstract
Improving the energy efficiency of database systems has emerged as an important topic of research over the past few years. While significant attention has been paid to optimizing the power consumption of tradition disk-based databases, little attention has been paid to the growing cost of DRAM power consumption in main-memory databases (MMDB).
Raja Appuswamy, Matthaios Olma, Anastasia Ailamaki
DaMoN1
2014 Towards Paravirtualized Network File Systems
Raja Appuswamy, Sergey Legtchenko, Antony I. T. Rowstron
HotStorage1
2014 Towards a Flexible, Lightweight Virtualization Alternative
abstract
In recent times, two virtualization approaches have become dominant: hardware-level and operating system-level virtualization. They differ by where they draw the virtualization boundary between the virtualizing and the virtualized part of the system, resulting in vastly different properties. We argue that these two approaches are extremes in a continuum, and that boundaries in between the extremes may combine several good properties of both. We propose abstractions to make up one such new virtualization boundary, which combines hardware-level flexibility with OS-level resource sharing. We implement and evaluate a first prototype.
David C. van Moolenbroek, Raja Appuswamy, Andrew S. Tanenbaum
SYSTOR2
2013 Scale-up vs scale-out for Hadoop: time to rethink?
abstract
In the last decade we have seen a huge deployment of cheap clusters to run data analytics workloads. The conventional wisdom in industry and academia is that scaling out using a cluster of commodity machines is better for these workloads than scaling up by adding more resources to a single server. Popular analytics infrastructures such as Hadoop are aimed at such a cluster scale-out environment.
Raja Appuswamy, Christos Gkantsidis, Dushyanth Narayanan, Orion Hodson, Antony I. T. Rowstron
SoCC1
2013 File-Level, Host-Side Flash Caching with Loris
abstract
As enterprises shift from using direct-attached storage to network-based storage for housing primary data, flash-based, host-side caching has gained momentum as the primary latency reduction technique. In this paper, we make the case for integration of flash caching algorithms at the file level, as opposed to the conventional block-level integration. In doing so, we will show how our extensions to Loris, a reliable, file-oriented storage stack, transform it into a framework for designing layout-independent, file-level caching systems. Using our Loris prototype, we demonstrate the effectiveness of Loris-based, file-level flash caching systems over their block-level counterparts, and investigate the effect of various write and allocation policies on the overall performance.
Raja Appuswamy, David C. van Moolenbroek, Sharan Santhanam, Andrew S. Tanenbaum
ICPADS1
2013 Cache, cache everywhere, flushing all hits down the sink: On exclusivity in multilevel, hybrid caches
abstract
Several multilevel storage systems have been designed over the past few years that utilize RAM and flash-based SSDs in concert to cache data resident in HDD-based primary storage. The low cost/GB and non-volatility of SSDs relative to RAM have encouraged storage system designers to adopt inclusivity (between RAM and SSD) in the caching hierarchy. However, in light of recent changes in hardware landscape, we believe that in the future, multilevel caches are invariably going to be hybrid caches where 1) all/most levels are physically collocated 2) the levels differ substantially only with respect to performance and not storage density, and 3) all levels are persistent. In this paper, we will investigate the design tradeoffs involved in building exclusive, persistent, direct-attached, multilevel storage caches. In doing so, we will first present a comparative evaluation of various techniques that have been proposed to achieve exclusivity in distributed storage caches in the context of a direct-attached, hybrid cache, and show the potential performance benefits of maintaining exclusivity. We will then investigate extensions to these demand-based, read-only data caching algorithms in order to address two issues specific to direct-attached hybrid caches, namely, handling writes and managing SSD lifetime.
Raja Appuswamy, David C. van Moolenbroek, Andrew S. Tanenbaum
MSST1
2013 Transaction-Based Process Crash Recovery of File System Namespace Modules
abstract
In this paper, we describe the emerging concept of namespace modules: operating system components that are responsible for constructing a hierarchical file system namespace based on one or more individual underlying file objects. We show that the likely presence of software bugs in such modules calls for the ability to recover from crashes, but that the current state of the art falls short of the desired behavior. We then introduce a crash recovery solution that is based on transactions, and detail the requirements for a system to implement this solution. We apply our solution to two different use cases: the primary namespace module for a storage stack, and an extension module that exposes the contents of scientific data files. Our evaluation shows that the transaction system has low overhead and significantly adds to the robustness of the namespace modules.
David C. van Moolenbroek, Raja Appuswamy, Andrew S. Tanenbaum
PRDC2
2012 Integrating flash-based SSDs into the storage stack
abstract
Over the past few years, hybrid storage architectures that use high-performance SSDs in concert with high-density HDDs have received significant interest from both industry and academia, due to their capability to improve performance while reducing capital and operating costs. These hybrid architectures differ in their approach to integrating SSDs into the traditional HDD-based storage stack. Of several such possible integrations, two have seen widespread adoption: Caching and Dynamic Storage Tiering. Although the effectiveness of these architectures under certain workloads is well understood, a systematic side-by-side analysis of these approaches remains difficult due to the range of design alternatives and configuration parameters involved. Such a study is required now more than ever to be able to design effective hybrid storage solutions for deployment in increasingly virtualized modern storage installations that blend several workloads into a single stream. In this paper, we first present our extensions to the Loris storage stack that transform it into a framework for designing hybrid storage systems. We then illustrate the flexibility of the framework by designing several Caching and DST-based hybrid systems. Following this, we present a systematic side-by-side analysis of these systems under a range of individual workload types and offer insights into the advantages and disadvantages of each architecture. Finally, we discuss the ramifications of our findings on the design of future hybrid storage systems in the light of recent changes in hardware landscape and application workloads.
Raja Appuswamy, David C. van Moolenbroek, Andrew S. Tanenbaum
MSST1
2012 Integrated System and Process Crash Recovery in the Loris Storage Stack
abstract
In this paper, we look at two important failure classes in the storage stack: system crashes, where the whole system shuts down unexpectedly, and process crashes, where a part of the storage stack software fails due to an implementation bug. We investigate these two problems in the context of the Loris storage stack. We show how restoring metadata consistency can provide a common first step for recovery from both types of crashes. In addition, we present fine-grained and corruption-resistant data resynchronization as the second step for system crash recovery, and an in-memory roll-forward log that can provide strong guarantees as the second step for process crash recovery in a microkernel setting. We implement our findings in our Loris prototype, and implement a new crash-resistant on-device layout as part of our proof of concept. The evaluation shows that our approach provides increased reliability at a reasonable performance cost.
David C. van Moolenbroek, Raja Appuswamy, Andrew S. Tanenbaum
NAS2
2011 Flexible, modular file volume virtualization in Loris
abstract
Traditional file systems made it possible for administrators to create file volumes, on a one-file-volume-per-disk basis. With the advent of RAID algorithms and their integration at the block level, this “one file volume per disk” bond forced administrators to create a single, shared file volume across all users to maximize storage efficiency, thereby complicating administration. To simplify administration, and to introduce new functionalities, file volume virtualization support was added at the block level. This new virtualization engine is commonly referred to as the volume manager, and the resulting arrangement, with volume managers operating below file systems, has been referred to as the traditional storage stack. In this paper, we present several problems associated with the compatibility-driven integration of file volume virtualization at the block level. In earlier work, we presented Loris, a reliable, modular storage stack, that solved several problems with the traditional storage stack by design. In this paper, we extend Loris to support file volume virtualization. In doing so, we first present “File pools”, our novel storage model to simplify storage administration, and support efficient file volume virtualization. Following this, we will describe how our single unified virtualization infrastructure, with a modular division of labor, is used to support several new functionalities like (1) instantaneous snapshoting of both files and file volumes, (2) efficient snapshot deletion through information sharing, and (3) open-close versioning of files. We then present “Version directories,” our unified interface for browsing file history information. Finally, we will evaluate the infrastructure, and provide an in-depth comparison of our approach with other competing approaches.
Raja Appuswamy, David C. van Moolenbroek, Andrew S. Tanenbaum
MSST1
2011 Efficient, Modular Metadata Management with Loris
abstract
With the amount of data increasing at an alarming rate, domain-specific user-level metadata management systems have emerged in several application areas to compensate for the shortcomings of file systems. Such systems provide domain-specific storage formats for performance-optimized metadata storage, search-based access interfaces for enabling declarative queries, and type-specific indexing structures for performing scalable search over metadata. In this paper, we highlight several issues that plague these user-level systems. We then show how integrating metadata management into the Loris stack solves all these problems by design. In doing so, we show how the Loris stack provides a modular framework for implementing domain-specific solutions by presenting the design of our own Loris-based metadata management system that provides 1) LSM-tree-based metadata storage, 2) an indexing infrastructure that uses LSM-trees for maintaining real-time attribute indices, and 3) scalable metadata querying using an attribute-based query language.
Richard van Heuven van Staereling, Raja Appuswamy, David C. van Moolenbroek, Andrew S. Tanenbaum
NAS2
2010 Block-level RAID Is Dead
Raja Appuswamy, David C. van Moolenbroek, Andrew S. Tanenbaum
HotStorage1
2010 Loris - A Dependable, Modular File-Based Storage Stack
abstract
The arrangement of file systems and volume management/RAID systems, together commonly referred to as the storage stack, has remained the same for several decades, despite significant changes in hardware, software and usage scenarios. In this paper, we evaluate the traditional storage stack along three dimensions: reliability, heterogeneity and flexibility. We highlight several major problems with the traditional stack. We then present Loris, our redesign of the storage stack, and we evaluate several aspects of Loris.
Raja Appuswamy, David C. van Moolenbroek, Andrew S. Tanenbaum
PRDC1