VLDB 2026 Research / reviewers in the wild / expert
Vasily Tarasov
dblp:73/5769
· DBLP profile ↗
37ranked-venue papers
6as first author
6since 2021 · last 2025
0000-0003-1424-9977ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 32 · 5 first-author · 5 since 2021Databases, data management, data science and information retrieval · 10 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Vela: A Virtualized LLM Training System with GPU Direct RoCEabstractVela is a cloud-native system designed for LLM training workloads built using off-the-shelf hardware, Linux KVM-based virtualization, and a virtualized RDMA over Converged Ethernet (RoCE) network. Vela virtual machines (VMs) support peer-to-peer DMA between the GPUs and SRIOV-based network interface. In this paper, we share Vela's key architectural aspects with details from an NVIDIA A100 GPU-based deployment in one of the IBM Cloud data centers. Throughout the paper, we share insights and experiences from designing, building, and operating the system over a ~2.5 year timeframe to highlight the capabilities of readily available software and hardware technologies and the improvement opportunities for future AI systems, thereby making AI infrastructure more accessible to a broader community. As we evaluated the system for performance at ~1500 GPU scale, we achieved ~80% of the ideal throughput while training a 50 billion parameter decoder model using model parallelism, and ~70% per GPU FLOPS compared to a single VM with the High-Performance Linpack benchmark. Apoorve Mohan, Robert Walkup, Bengi Karaçali, Ming-Hung Chen, Abdullah Kayi, Liran Schour, Shweta Salaria, Sophia Wen, I-Hsin Chung, Abdul Alim, Constantinos Evangelinos, Lixiang Luo, Marc Dombrowa, Laurent Schares, Ali Sydney, Pavlos Maniotis, Sandhya Koteshwara, Brent Tang, Joel Belog, Rei Odaira, Vasily Tarasov, Eran Gampel, Drew Thorstensen, Talia Gershon, Seetharami Seelam |
ASPLOS (2) | 21 |
| 2024 | An End-to-end High-performance Deduplication Scheme for Docker Registries and Docker Container Storage SystemsabstractThe wide adoption of Docker containers for supporting agile and elastic enterprise applications has led to a broad proliferation of container images. The associated storage performance and capacity requirements place a high pressure on the infrastructure of container registries that store and distribute images and container storage systems on the Docker client side that manage image layers and store ephemeral data generated at container runtime. The storage demand is worsened by the large amount of duplicate data in images. Moreover, container storage systems that use Copy-on-Write (CoW) file systems as storage drivers exacerbate the redundancy. Exploiting the high file redundancy in real-world images is a promising approach to drastically reduce the growing storage requirements of container registries and improve the space efficiency of container storage systems. However, existing deduplication techniques significantly degrade the performance of both registries and container storage systems because of data reconstruction overhead as well as the deduplication cost. We propose DupHunter, an end-to-end deduplication scheme that deduplicates layers for both Docker registries and container storage systems while maintaining a high image distribution speed and container I/O performance. DupHunter is divided into three tiers: registry tier, middle tier, and client tier. Specifically, we first build a high-performance deduplication engine at the registry tier that not only natively deduplicates layers for space savings but also reduces layer restore overhead. Then, we use deduplication offloading at the middle tier to eliminate the redundant files from the client tier and avoid bringing deduplication overhead to the clients. To further reduce the data duplicates caused by CoWs and improve the container I/O performance, we utilize a container-aware storage system at the client tier that reserves space for each container and arranges the placement of files and their modifications on the disk to preserve locality. Under real workloads, DupHunter reduces storage space by up to 6.9× and reduces the GET layer latency up to 2.8× compared to the state-of-the-art. Moreover, DupHunter can improve the container I/O performance by up to 93% for reads and 64% for writes. Muhui Lin, Hadeel Albahar, Arnab Kumar Paul, Zhijie Huan, Subil Abraham, Vasily Tarasov, Dimitrios Skourtis, Ali Anwar 0001, Ali Raza Butt |
ACM Trans. Storage | 8 |
| 2023 | F3: Serving Files Efficiently in Serverless ComputingabstractServerless platforms offer on-demand computation and represent a significant shift from previous platforms that typically required resources to be pre-allocated (e.g., virtual machines). As serverless platforms have evolved, they have become suitable for a much wider range of applications than their original use cases. However, storage access remains a pain point that holds serverless back from becoming a completely generic computation platform. Alex Merenstein, Vasily Tarasov, Ali Anwar 0001, Scott Guthridge, Erez Zadok |
SYSTOR | 2 |
| 2023 | InfiniStore: Elastic Serverless Cloud StorageabstractCloud object storage such as AWS S3 is cost-effective and highly elastic but relatively slow, while high-performance cloud storage such as AWS ElastiCache is expensive and provides limited elasticity. We present a new cloud storage service called ServerlessMemory, which stores data using the memory of serverless functions. ServerlessMemory employs a sliding-window-based memory management strategy inspired by the garbage collection mechanisms used in the programming language to effectively segregate hot/cold data and provides fine-grained elasticity, good performance, and a pay-per-access cost model with extremely low cost. We then design and implement InfiniStore, a persistent and elastic cloud storage system, which seamlessly couples the function-based ServerlessMemory layer with a persistent, inexpensive cloud object store layer. InfiniStore enables durability despite function failures using a fast parallel recovery scheme built on the auto-scaling functionality of a FaaS (Function-as-a-Service) platform. We evaluate InfiniStore extensively using both microbenchmarking and two real-world applications. Results show that InfiniStore has more performance benefits for objects larger than 10 MB compared to AWS ElastiCache and Anna, and InfiniStore achieves 26.25% and 97.24% tenant-side cost reduction compared to InfiniCache and ElastiCache, respectively. Benjamin Carver, Nicholas John Newman, Ali Anwar 0001, Lukas Rupprecht, Vasily Tarasov, Dimitrios Skourtis, Feng Yan 0001, Yue Cheng 0001 |
Proc. VLDB Endow. | 8 |
| 2021 | CNSBench: A Cloud Native Storage Benchmark
Alex Merenstein, Vasily Tarasov, Ali Anwar 0001, Deepavali Bhagwat, Julie Lee, Lukas Rupprecht, Dimitrios Skourtis, Erez Zadok |
FAST | 2 |
| 2021 | Large-Scale Analysis of Docker Images and Performance Implications for Container Storage SystemsabstractDocker containers have become a prominent solution for supporting modern enterprise applications due to the highly desirable features of isolation, low overhead, and efficient packaging of the application’s execution environment. Containers are created from images which are shared between users via a registry. The amount of data registries store is massive. For example, Docker Hub, a popular public registry, stores at least half a million public images. In this article, we analyze over 167 TB of uncompressed Docker Hub images, characterize them using multiple metrics and evaluate the potential of file-level deduplication. Our analysis helps to make conscious decisions when designing storage for containers in general and Docker registries in particular. For example, only 3 percent of the files in images are unique while others are redundant file copies, which means file-level deduplication has a great potential to save storage space. Furthermore, we carry out a comprehensive analysis of both small I/O request performance and copy-on-write performance for multiple popular container storage drivers. Our findings can motivate and help improve the design of data reduction and caching methods for images, pulling optimizations for registries, and storage drivers. Vasily Tarasov, Hadeel Albahar, Ali Anwar 0001, Lukas Rupprecht, Dimitrios Skourtis, Arnab Kumar Paul, Ali Raza Butt |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2020 | InfiniCache: Exploiting Ephemeral Serverless Functions to Build a Cost-Effective Memory Cache
Ali Anwar 0001, Lukas Rupprecht, Dimitrios Skourtis, Vasily Tarasov, Feng Yan 0001, Yue Cheng 0001 |
FAST | 7 |
| 2020 | Position: Can Microservices Drive a Renaissance in Workload-Aware Storage Management?
Pranav Bhandari, Avani Wildani, Dimitrios Skourtis, Vasily Tarasov, Deepavali Bhagwat, Lukas Rupprecht, Ali Anwar 0001 |
HotStorage | 4 |
| 2020 | The Case for Benchmarking Control Operations in Cloud Native Storage
Alex Merenstein, Vasily Tarasov, Ali Anwar 0001, Deepavali Bhagwat, Lukas Rupprecht, Dimitrios Skourtis, Erez Zadok |
HotStorage | 2 |
| 2020 | DupHunter: Flexible High-Performance Deduplication for Docker Registries
Hadeel Albahar, Subil Abraham, Vasily Tarasov, Dimitrios Skourtis, Lukas Rupprecht, Ali Anwar 0001, Ali Raza Butt |
USENIX ATC | 5 |
| 2019 | Bolt: Towards a Scalable Docker Registry via HyperconvergenceabstractDocker container images are typically stored in a centralized registry to allow easy sharing of images. However, with the growing popularity of containerized software, the number of images that a registry needs to store and the rate of requests it needs to serve are increasing rapidly. Current registry design requires hosting registry services across multiple loosely connected servers with different roles such as load balancers, proxies, registry servers, and object storage servers. Due to the various individual components, registries are hard to scale and benefits from optimizations such as caching are limited. In this paper we propose, implement, and evaluate BOLT-a new hyperconverged design for container registries. In BOLT, all registry servers are part of a tightly connected cluster and play the same consolidated role: each registry server caches images in its memory, stores images in its local storage, and provides computational resources to process client requests. The design employs a custom consistent hashing function to take advantage of the layered structure and addressing of images and to load balance requests across different servers. Our evaluation using real production workloads shows that BOLT outperforms the conventional registry design significantly and improves latency by an order of magnitude and throughput by up to 5x. Compared to state-of-the-art, BOLT can utilize cache space more efficiently and serve up to 35% more requests from its cache. Furthermore, BOLT scales linearly and recovers from failure recovery without significant performance degradation. Michael Littley, Ali Anwar 0001, Hannan Fayyaz, Zeshan Fayyaz, Vasily Tarasov, Lukas Rupprecht, Dimitrios Skourtis, Mohamed Mohamed 0001, Heiko Ludwig, Yue Cheng 0001, Ali Raza Butt |
CLOUD | 5 |
| 2019 | Slimmer: Weight Loss Secrets for Docker RegistriesabstractDue to their tight isolation, low overhead, and efficient packaging of the execution environment, Docker containers have become a prominent solution for deploying modern applications. Containers are created from images which are stored in a Docker registry. An image consists of a list of layers which can be shared among images. Docker registries store a large amount of images and with the increasing popularity of Docker, they continue to grow. For example, Docker Hub-a popular public registry-stores more than half a million public images. In this paper, we analyze over 167TB of uncompressed Docker images and evaluate the potential of file-level deduplication in the registry. Our analysis reveals that only 3% of the files in images are unique and Docker's existing layer sharing mechanism is not sufficient to eliminate this profound redundancy. We then present the design of Slimmer-a Docker registry with file deduplication support-and conduct a simulation-based analysis of its performance implications. Vasily Tarasov, Ali Anwar 0001, Lukas Rupprecht, Dimitrios Skourtis, Amit Warke, Mohamed Mohamed 0001, Ali Raza Butt |
CLOUD | 2 |
| 2019 | Agni: An Efficient Dual-access File System over Object StorageabstractObject storage is a low-cost, scalable component of cloud ecosystems. However, interface incompatibilities and performance limitations inhibit its adoption for emerging cloud-based workloads. Users are compelled to either run their applications over expensive block storage-based file systems or use inefficient file connectors over object stores. Dual access, the ability to read and write the same data through file systems interfaces and object storage APIs, has promise to improve performance and eliminate storage sprawl. Kunal Lillaney, Vasily Tarasov, David Pease, Randal C. Burns |
SoCC | 2 |
| 2019 | Large-Scale Analysis of the Docker Hub DatasetabstractDocker containers have become a prominent solution for supporting modern enterprise applications due to the highly desirable features of isolation, low overhead, and efficient packaging of the execution environment. Containers are created from images which are shared between users via a Docker registry. The amount of data Docker registries store is massive; for example, Docker Hub, a popular public registry, stores at least half a million public images. In this paper, we analyze over 167 TB of uncompressed Docker Hub images, characterize them using multiple metrics and evaluate the potential of file-level deduplication in Docker Hub. Our analysis helps to make conscious decisions when designing storage for containers in general and Docker registries in particular. For example, only 3% of the files in images are unique, which means file-level deduplication has a great potential to save storage space for the registry. Our findings can motivate and help improve the design of data reduction, caching, and pulling optimizations for registries. Vasily Tarasov, Hadeel Albahar, Ali Anwar 0001, Lukas Rupprecht, Dimitrios Skourtis, Amit Warke, Mohamed Mohamed 0001, Ali Raza Butt |
CLUSTER | 2 |
| 2019 | The Case for Dual-access File Systems over Object Storage
Kunal Lillaney, Vasily Tarasov, David Pease, Randal C. Burns |
HotStorage | 2 |
| 2019 | Performance and Resource Utilization of FUSE User-Space File SystemsabstractTraditionally, file systems were implemented as part of operating systems kernels, which provide a limited set of tools and facilities to a programmer. As the complexity of file systems grew, many new file systems began being developed in user space. Low performance is considered the main disadvantage of user-space file systems but the extent of this problem has never been explored systematically. As a result, the topic of user-space file systems remains rather controversial: while some consider user-space file systems a “toy” not to be used in production, others develop full-fledged production file systems in user space. In this article, we analyze the design and implementation of a well-known user-space file system framework, FUSE, for Linux. We characterize its performance and resource utilization for a wide range of workloads. We present FUSE performance and also resource utilization with various mount and configuration options, using 45 different workloads that were generated using Filebench on two different hardware configurations. We instrumented FUSE to extract useful statistics and traces, which helped us analyze its performance bottlenecks and present our analysis results. Our experiments indicate that depending on the workload and hardware used, performance degradation (throughput) caused by FUSE can be completely imperceptible or as high as −83%, even when optimized; and latencies of FUSE file system operations can be increased from none to 4× when compared to Ext4. On the resource utilization side, FUSE can increase relative CPU utilization by up to 31% and underutilize disk bandwidth by as much as −80% compared to Ext4, though for many data-intensive workloads the impact was statistically indistinguishable. Our conclusion is that user-space file systems can indeed be used in production (non-“toy”) settings, but their applicability depends on the expected workloads. Bharath Kumar Reddy Vangoor, Prafful Agarwal, Manu Mathew, Arun Ramachandran, Swaminathan Sivaraman, Vasily Tarasov, Erez Zadok |
ACM Trans. Storage | 6 |
| 2018 | Wharf: Sharing Docker Images in a Distributed File SystemabstractContainer management frameworks, such as Docker, package diverse applications and their complex dependencies in self-contained images, which facilitates application deployment, distribution, and sharing. Currently, Docker employs a shared-nothing storage architecture, i.e. every Docker-enabled host requires its own copy of an image on local storage to create and run containers. This greatly inflates storage utilization, network load, and job completion times in the cluster. In this paper, we investigate the option of storing container images in and serving them from a distributed file system. By sharing images in a distributed storage layer, storage utilization can be reduced and redundant image retrievals from a Docker registry become unnecessary. We introduce Wharf, a middleware to transparently add distributed storage support to Docker. Wharf partitions Docker's runtime state into local and global parts and efficiently synchronizes accesses to the global state. By exploiting the layered structure of Docker images, Wharf minimizes the synchronization overhead. Our experiments show that compared to Docker on local storage, Wharf can speed up image retrievals by up to 12x, has more stable performance, and introduces only a minor overhead when accessing data on distributed storage. Chao Zheng 0002, Lukas Rupprecht, Vasily Tarasov, Douglas Thain, Mohamed Mohamed 0001, Dimitrios Skourtis, Amit Warke, Dean Hildebrand |
SoCC | 3 |
| 2018 | Improving Docker Registry Design Based on Production Workload Analysis
Ali Anwar 0001, Mohamed Mohamed 0001, Vasily Tarasov, Michael Littley, Lukas Rupprecht, Yue Cheng 0001, Dimitrios Skourtis, Amit Warke, Heiko Ludwig, Dean Hildebrand, Ali Raza Butt |
FAST | 3 |
| 2018 | Towards Better Understanding of Black-box Auto-Tuning: A Comparative Analysis for Storage Systems
Vasily Tarasov, Sachin Tiwari, Erez Zadok |
USENIX ATC | 2 |
| 2018 | Cluster and Single-Node Analysis of Long-Term Deduplication PatternsabstractDeduplication has become essential in disk-based backup systems, but there have been few long-term studies of backup workloads. Most past studies either were of a small static snapshot or covered only a short period that was not representative of how a backup system evolves over time. For this article, we first collected 21 months of data from a shared user file system; 33 users and over 4,000 snapshots are covered. We then analyzed the dataset, examining a variety of essential characteristics across two dimensions: single-node deduplication and cluster deduplication. For single-node deduplication analysis, our primary focus was individual-user data. Despite apparently similar roles and behavior among all of our users, we found significant differences in their deduplication ratios. Moreover, the data that some users share with others had a much higher deduplication ratio than average. For cluster deduplication analysis, we implemented seven published data-routing algorithms and created a detailed comparison of their performance with respect to deduplication ratio, load distribution, and communication overhead. We found that per-file routing achieves a higher deduplication ratio than routing by super-chunk (multiple consecutive chunks), but it also leads to high data skew (imbalance of space usage across nodes). We also found that large chunking sizes are better for cluster deduplication, as they significantly reduce data-routing overhead, while their negative impact on deduplication ratios is small and acceptable. We draw interesting conclusions from both single-node and cluster deduplication analysis and make recommendations for future deduplication systems design. Zhen Jason Sun, Geoffrey H. Kuenning, Sonam Mandal, Philip Shilane, Vasily Tarasov, Nong Xiao 0001, Erez Zadok |
ACM Trans. Storage | 5 |
| 2018 | Challenges and Solutions for Tracing Storage Systems: A Case Study with Spectrum ScaleabstractIBM Spectrum Scale’s parallel file system General Parallel File System (GPFS) has a 20-year development history with over 100 contributing developers. Its ability to support strict POSIX semantics across more than 10K clients leads to a complex design with intricate interactions between the cluster nodes. Tracing has proven to be a vital tool to understand the behavior and the anomalies of such a complex software product. However, the necessary trace information is often buried in hundreds of gigabytes of by-product trace records. Further, the overhead of tracing can significantly impact running applications and file system performance, limiting the use of tracing in a production system. In this research article, we discuss the evolution of the mature and highly scalable GPFS tracing tool and present the exploratory study of GPFS’ new tracing interface, FlexTrace , which allows developers and users to accurately specify what to trace for the problem they are trying to solve. We evaluate our methodology and prototype, demonstrating that the proposed approach has negligible overhead, even under intensive I/O workloads and with low-latency storage devices. Marc-Andre Vef, Vasily Tarasov, Dean Hildebrand, André Brinkmann |
ACM Trans. Storage | 2 |
| 2017 | On the Performance Variation in Modern Storage Stacks
Vasily Tarasov, Hari Prasath Raman, Dean Hildebrand, Erez Zadok |
FAST | 2 |
| 2017 | To FUSE or Not to FUSE: Performance of User-Space File Systems
Bharath Kumar Reddy Vangoor, Vasily Tarasov, Erez Zadok |
FAST | 2 |
| 2017 | Introduction to the Special Issue on MSST 2016abstractNo abstract available. Carlos Maltzahn, Vasily Tarasov |
ACM Trans. Storage | 2 |
| 2016 | Using Hints to Improve Inline Block-layer Deduplication
Sonam Mandal, Geoffrey H. Kuenning, Dongju Ok, Varun Shastry, Philip Shilane, Sun Zhen, Vasily Tarasov, Erez Zadok |
FAST | 7 |
| 2016 | A long-term user-centric analysis of deduplication patternsabstractDeduplication has become essential in disk-based backup systems, but there have been few long-term studies of backup workloads. Most past studies either were of a small static snapshot or covered only a short period that was not representative of how a backup system evolves over time. For this paper, we collected 21 months of data from a shared user file system; 33 users and over 4,000 snapshots are covered. We analyzed the data set for a variety of essential characteristics. However, our primary focus was individual user data. Despite apparently similar roles and behavior in all of our users, we found significant differences in their deduplication ratios. Moreover, the data that some users share with others had a much higher deduplication ratio than average. We analyze this behavior and make recommendations for future deduplication systems design. Geoffrey H. Kuenning, Sonam Mandal, Philip Shilane, Vasily Tarasov, Nong Xiao 0001, Erez Zadok |
MSST | 5 |
| 2015 | Terra Incognita: On the Practicality of User-Space File Systems
Vasily Tarasov, Kumar Sourav, Sagar Trehan, Erez Zadok |
HotStorage | 1 |
| 2014 | Linux NFSv4.1 Performance Under a Microscope
Ming Chen 0013, Dean Hildebrand, Geoffrey H. Kuenning, Soujanya Shankaranarayana, Vasily Tarasov, Arun O. Vasudevan, Erez Zadok, Ksenia Zakirova |
LISA | 5 |
| 2013 | Virtual machine workloads: the case for new benchmarks for NAS
Vasily Tarasov, Dean Hildebrand, Geoffrey H. Kuenning, Erez Zadok |
FAST | 1 |
| 2013 | Improving I/O Performance Using Virtual Disk Introspection
Vasily Tarasov, Dean Hildebrand, Renu Tewari, Geoffrey H. Kuenning, Erez Zadok |
HotStorage | 1 |
| 2012 | Extracting flexible, replayable models from large block traces
Vasily Tarasov, Santhosh Kumar, Jack Ma, Dean Hildebrand, Anna Povzner, Geoffrey H. Kuenning, Erez Zadok |
FAST | 1 |
| 2012 | Generating Realistic Datasets for Deduplication Analysis
Vasily Tarasov, Amar Mudrankit, Will Buik, Philip Shilane, Geoffrey H. Kuenning, Erez Zadok |
USENIX ATC | 1 |
| 2011 | Benchmarking File System Benchmarking: It *IS* Rocket Science
Vasily Tarasov, Saumitra Bhanage, Erez Zadok, Margo I. Seltzer |
HotOS | 1 |
| 2011 | Static discovery and remediation of code-embedded resource dependenciesabstractMany enterprises perform data-center transformation, consolidation, and migration in order to improve the efficiency of their IT infrastructures. These transformation projects begin with the discovery of the existing infrastructure, in particular the dependencies between applications. These dependencies are needed in planning, in order to determine how components influence one another, and in relinking, so that the component names and addresses can be updated. Typically, dependency discovery is done by network monitoring and middleware configuration analysis. These existing approaches will often fail to detect dependencies expressed in the application code. In this paper, we present the first method and tool for automatically identifying code-embedded external dependencies in Java Enterprise Edition applications. In addition, our tool can automatically alter the application code to update the dependencies, or externalize them to configuration files. We analyzed over 1000 Java EE applications from three enterprise environments. The results demonstrate the prevalence of code-embedded dependencies that would otherwise have to be identified manually, often causing failures during user-acceptance testing. Nikolai Joukov, Vasily Tarasov, Joel Ossher, Birgit Pfitzmann, Sergej Chicherin, Marco Pistoia, Takaaki Tateishi |
Integrated Network Management | 2 |
| 2010 | Evaluating Performance and Energy in File System Server Workloads
Priya Sehgal, Vasily Tarasov, Erez Zadok |
FAST | 2 |
| 2010 | Optimizing energy and performance for server-class file system workloadsabstractRecently, power has emerged as a critical factor in designing components of storage systems, especially for power-hungry data centers. While there is some research into power-aware storage stack components, there are no systematic studies evaluating each component's impact separately. Various factors like workloads, hardware configurations, and software configurations impact the performance and energy efficiency of the system. This article evaluates the file system's impact on energy consumption and performance. We studied several popular Linux file systems, with various mount and format options, using the FileBench workload generator to emulate four server workloads: Web, database, mail, and fileserver, on two different hardware configurations. The file system design, implementation, and available features have a significant effect on CPU/disk utilization, and hence on performance and power. We discovered that default file system options are often suboptimal, and even poor. In this article we show that a careful matching of expected workloads and hardware configuration to a single software configuration—the file system—can improve power-performance efficiency by a factor ranging from 1.05 to 9.4 times. Priya Sehgal, Vasily Tarasov, Erez Zadok |
ACM Trans. Storage | 2 |
| 2009 | Energy and performance evaluation of lossless file data compression on server systemsabstractData compression has been claimed to be an attractive solution to save energy consumption in high-end servers and data centers. However, there has not been a study to explore this. In this paper, we present a comprehensive evaluation of energy consumption for various file compression techniques implemented in software. We apply various compression tools available on Linux to a variety of data files, and we try them on server class and workstation class systems. We compare their energy and performance results against raw reads and writes. Our results reveal that software based data compression cannot be considered as a universal solution to reduce energy consumption. Various factors like the type of the data file, the compression tool being used, the read-to-write ratio of the workload, and the hardware configuration of the system impact the efficacy of this technique. In some cases, however, we found compression to save substantial energy and improve performance. Rachita Kothiyal, Vasily Tarasov, Priya Sehgal, Erez Zadok |
SYSTOR | 2 |