EDBT 2026 Demo / reviewers in the wild / expert
Dimitrios Skourtis
dblp:115/5948 · also Dimitris Skourtis
· DBLP profile ↗
15ranked-venue papers
2as first author
4since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 4 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | An End-to-end High-performance Deduplication Scheme for Docker Registries and Docker Container Storage SystemsabstractThe wide adoption of Docker containers for supporting agile and elastic enterprise applications has led to a broad proliferation of container images. The associated storage performance and capacity requirements place a high pressure on the infrastructure of container registries that store and distribute images and container storage systems on the Docker client side that manage image layers and store ephemeral data generated at container runtime. The storage demand is worsened by the large amount of duplicate data in images. Moreover, container storage systems that use Copy-on-Write (CoW) file systems as storage drivers exacerbate the redundancy. Exploiting the high file redundancy in real-world images is a promising approach to drastically reduce the growing storage requirements of container registries and improve the space efficiency of container storage systems. However, existing deduplication techniques significantly degrade the performance of both registries and container storage systems because of data reconstruction overhead as well as the deduplication cost. We propose DupHunter, an end-to-end deduplication scheme that deduplicates layers for both Docker registries and container storage systems while maintaining a high image distribution speed and container I/O performance. DupHunter is divided into three tiers: registry tier, middle tier, and client tier. Specifically, we first build a high-performance deduplication engine at the registry tier that not only natively deduplicates layers for space savings but also reduces layer restore overhead. Then, we use deduplication offloading at the middle tier to eliminate the redundant files from the client tier and avoid bringing deduplication overhead to the clients. To further reduce the data duplicates caused by CoWs and improve the container I/O performance, we utilize a container-aware storage system at the client tier that reserves space for each container and arranges the placement of files and their modifications on the disk to preserve locality. Under real workloads, DupHunter reduces storage space by up to 6.9× and reduces the GET layer latency up to 2.8× compared to the state-of-the-art. Moreover, DupHunter can improve the container I/O performance by up to 93% for reads and 64% for writes. Muhui Lin, Hadeel Albahar, Arnab Kumar Paul, Zhijie Huan, Subil Abraham, Vasily Tarasov, Dimitrios Skourtis, Ali Anwar 0001, Ali Raza Butt |
ACM Trans. Storage | 9 |
| 2023 | InfiniStore: Elastic Serverless Cloud StorageabstractCloud object storage such as AWS S3 is cost-effective and highly elastic but relatively slow, while high-performance cloud storage such as AWS ElastiCache is expensive and provides limited elasticity. We present a new cloud storage service called ServerlessMemory, which stores data using the memory of serverless functions. ServerlessMemory employs a sliding-window-based memory management strategy inspired by the garbage collection mechanisms used in the programming language to effectively segregate hot/cold data and provides fine-grained elasticity, good performance, and a pay-per-access cost model with extremely low cost. We then design and implement InfiniStore, a persistent and elastic cloud storage system, which seamlessly couples the function-based ServerlessMemory layer with a persistent, inexpensive cloud object store layer. InfiniStore enables durability despite function failures using a fast parallel recovery scheme built on the auto-scaling functionality of a FaaS (Function-as-a-Service) platform. We evaluate InfiniStore extensively using both microbenchmarking and two real-world applications. Results show that InfiniStore has more performance benefits for objects larger than 10 MB compared to AWS ElastiCache and Anna, and InfiniStore achieves 26.25% and 97.24% tenant-side cost reduction compared to InfiniCache and ElastiCache, respectively. Benjamin Carver, Nicholas John Newman, Ali Anwar 0001, Lukas Rupprecht, Vasily Tarasov, Dimitrios Skourtis, Feng Yan 0001, Yue Cheng 0001 |
Proc. VLDB Endow. | 9 |
| 2021 | CNSBench: A Cloud Native Storage Benchmark
Alex Merenstein, Vasily Tarasov, Ali Anwar 0001, Deepavali Bhagwat, Julie Lee, Lukas Rupprecht, Dimitrios Skourtis, Erez Zadok |
FAST | 7 |
| 2021 | Large-Scale Analysis of Docker Images and Performance Implications for Container Storage SystemsabstractDocker containers have become a prominent solution for supporting modern enterprise applications due to the highly desirable features of isolation, low overhead, and efficient packaging of the application’s execution environment. Containers are created from images which are shared between users via a registry. The amount of data registries store is massive. For example, Docker Hub, a popular public registry, stores at least half a million public images. In this article, we analyze over 167 TB of uncompressed Docker Hub images, characterize them using multiple metrics and evaluate the potential of file-level deduplication. Our analysis helps to make conscious decisions when designing storage for containers in general and Docker registries in particular. For example, only 3 percent of the files in images are unique while others are redundant file copies, which means file-level deduplication has a great potential to save storage space. Furthermore, we carry out a comprehensive analysis of both small I/O request performance and copy-on-write performance for multiple popular container storage drivers. Our findings can motivate and help improve the design of data reduction and caching methods for images, pulling optimizations for registries, and storage drivers. Vasily Tarasov, Hadeel Albahar, Ali Anwar 0001, Lukas Rupprecht, Dimitrios Skourtis, Arnab Kumar Paul, Ali Raza Butt |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2020 | InfiniCache: Exploiting Ephemeral Serverless Functions to Build a Cost-Effective Memory Cache
Ali Anwar 0001, Lukas Rupprecht, Dimitrios Skourtis, Vasily Tarasov, Feng Yan 0001, Yue Cheng 0001 |
FAST | 6 |
| 2020 | Position: Can Microservices Drive a Renaissance in Workload-Aware Storage Management?
Pranav Bhandari, Avani Wildani, Dimitrios Skourtis, Vasily Tarasov, Deepavali Bhagwat, Lukas Rupprecht, Ali Anwar 0001 |
HotStorage | 3 |
| 2020 | The Case for Benchmarking Control Operations in Cloud Native Storage
Alex Merenstein, Vasily Tarasov, Ali Anwar 0001, Deepavali Bhagwat, Lukas Rupprecht, Dimitrios Skourtis, Erez Zadok |
HotStorage | 6 |
| 2020 | DupHunter: Flexible High-Performance Deduplication for Docker Registries
Hadeel Albahar, Subil Abraham, Vasily Tarasov, Dimitrios Skourtis, Lukas Rupprecht, Ali Anwar 0001, Ali Raza Butt |
USENIX ATC | 6 |
| 2019 | Bolt: Towards a Scalable Docker Registry via HyperconvergenceabstractDocker container images are typically stored in a centralized registry to allow easy sharing of images. However, with the growing popularity of containerized software, the number of images that a registry needs to store and the rate of requests it needs to serve are increasing rapidly. Current registry design requires hosting registry services across multiple loosely connected servers with different roles such as load balancers, proxies, registry servers, and object storage servers. Due to the various individual components, registries are hard to scale and benefits from optimizations such as caching are limited. In this paper we propose, implement, and evaluate BOLT-a new hyperconverged design for container registries. In BOLT, all registry servers are part of a tightly connected cluster and play the same consolidated role: each registry server caches images in its memory, stores images in its local storage, and provides computational resources to process client requests. The design employs a custom consistent hashing function to take advantage of the layered structure and addressing of images and to load balance requests across different servers. Our evaluation using real production workloads shows that BOLT outperforms the conventional registry design significantly and improves latency by an order of magnitude and throughput by up to 5x. Compared to state-of-the-art, BOLT can utilize cache space more efficiently and serve up to 35% more requests from its cache. Furthermore, BOLT scales linearly and recovers from failure recovery without significant performance degradation. Michael Littley, Ali Anwar 0001, Hannan Fayyaz, Zeshan Fayyaz, Vasily Tarasov, Lukas Rupprecht, Dimitrios Skourtis, Mohamed Mohamed 0001, Heiko Ludwig, Yue Cheng 0001, Ali Raza Butt |
CLOUD | 7 |
| 2019 | Slimmer: Weight Loss Secrets for Docker RegistriesabstractDue to their tight isolation, low overhead, and efficient packaging of the execution environment, Docker containers have become a prominent solution for deploying modern applications. Containers are created from images which are stored in a Docker registry. An image consists of a list of layers which can be shared among images. Docker registries store a large amount of images and with the increasing popularity of Docker, they continue to grow. For example, Docker Hub-a popular public registry-stores more than half a million public images. In this paper, we analyze over 167TB of uncompressed Docker images and evaluate the potential of file-level deduplication in the registry. Our analysis reveals that only 3% of the files in images are unique and Docker's existing layer sharing mechanism is not sufficient to eliminate this profound redundancy. We then present the design of Slimmer-a Docker registry with file deduplication support-and conduct a simulation-based analysis of its performance implications. Vasily Tarasov, Ali Anwar 0001, Lukas Rupprecht, Dimitrios Skourtis, Amit Warke, Mohamed Mohamed 0001, Ali Raza Butt |
CLOUD | 5 |
| 2019 | Large-Scale Analysis of the Docker Hub DatasetabstractDocker containers have become a prominent solution for supporting modern enterprise applications due to the highly desirable features of isolation, low overhead, and efficient packaging of the execution environment. Containers are created from images which are shared between users via a Docker registry. The amount of data Docker registries store is massive; for example, Docker Hub, a popular public registry, stores at least half a million public images. In this paper, we analyze over 167 TB of uncompressed Docker Hub images, characterize them using multiple metrics and evaluate the potential of file-level deduplication in Docker Hub. Our analysis helps to make conscious decisions when designing storage for containers in general and Docker registries in particular. For example, only 3% of the files in images are unique, which means file-level deduplication has a great potential to save storage space for the registry. Our findings can motivate and help improve the design of data reduction, caching, and pulling optimizations for registries. Vasily Tarasov, Hadeel Albahar, Ali Anwar 0001, Lukas Rupprecht, Dimitrios Skourtis, Amit Warke, Mohamed Mohamed 0001, Ali Raza Butt |
CLUSTER | 6 |
| 2018 | Wharf: Sharing Docker Images in a Distributed File SystemabstractContainer management frameworks, such as Docker, package diverse applications and their complex dependencies in self-contained images, which facilitates application deployment, distribution, and sharing. Currently, Docker employs a shared-nothing storage architecture, i.e. every Docker-enabled host requires its own copy of an image on local storage to create and run containers. This greatly inflates storage utilization, network load, and job completion times in the cluster. In this paper, we investigate the option of storing container images in and serving them from a distributed file system. By sharing images in a distributed storage layer, storage utilization can be reduced and redundant image retrievals from a Docker registry become unnecessary. We introduce Wharf, a middleware to transparently add distributed storage support to Docker. Wharf partitions Docker's runtime state into local and global parts and efficiently synchronizes accesses to the global state. By exploiting the layered structure of Docker images, Wharf minimizes the synchronization overhead. Our experiments show that compared to Docker on local storage, Wharf can speed up image retrievals by up to 12x, has more stable performance, and introduces only a minor overhead when accessing data on distributed storage. Chao Zheng 0002, Lukas Rupprecht, Vasily Tarasov, Douglas Thain, Mohamed Mohamed 0001, Dimitrios Skourtis, Amit Warke, Dean Hildebrand |
SoCC | 6 |
| 2018 | Improving Docker Registry Design Based on Production Workload Analysis
Ali Anwar 0001, Mohamed Mohamed 0001, Vasily Tarasov, Michael Littley, Lukas Rupprecht, Yue Cheng 0001, Dimitrios Skourtis, Amit Warke, Heiko Ludwig, Dean Hildebrand, Ali Raza Butt |
FAST | 8 |
| 2014 | Flash on Rails: Consistent Flash Performance through Redundancy
Dimitrios Skourtis, Dimitris Achlioptas, Noah Watkins, Carlos Maltzahn, Scott A. Brandt |
USENIX ATC | 1 |
| 2012 | QBox: guaranteeing I/O performance on black box storage systemsabstractMany storage systems are shared by multiple clients with different types of workloads and performance targets. To achieve performance targets without over-provisioning, a system must provide isolation between clients. Throughput-based reservations are challenging due to the mix of workloads and the stateful nature of disk drives, leading to low reservable throughput, while existing utilization-based solutions require specialized I/O scheduling for each device in the storage system. Dimitrios Skourtis, Shinpei Kato, Scott A. Brandt |
HPDC | 1 |