EDBT 2026 Demo / reviewers in the wild / expert
Manolis Marazakis
dblp:90/3704
· DBLP profile ↗
36ranked-venue papers
5as first author
9since 2021 · last 2025
0000-0002-4768-3289ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 26 · 4 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Artificial intelligence and machine learning · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | The ExaNeSt Prototype: Evaluation of Efficient HPC Communication Hardware in an ARM-based Multi-FPGA RackabstractWe present and evaluate the ExaNeSt prototype, which compactly packages 128 Xilinx ZU9EG MPSoCs, two TBytes of DRAM, and eight TBytes of SSD into a liquid-cooled rack, using a custom interconnection hardware based on 10 GB/s links. We developed this testbed in 2016–2019 in order to leverage the flexibility of FPGAs for experimenting with efficient hardware support for HPC communication among tens of thousands of processors and accelerators in the quest toward Exascale systems and beyond. In the years since then, we carefully studied this system, and we present our key design choices and insights resulting from our measurement and analysis. We developed this testbed, from architecture to the PCBs and the run-time software, within the ExaNeSt project. It is fully operational in configurations with up to 8 × 4 × 4 MPSoC nodes. It achieves high density through tight board design, while also leveraging state-of-the-art liquid cooling technology. In this article, we present a thorough architectural analysis, along with important aspects of our infrastructure development. Our custom interconnect includes a low-cost low-latency network interface, offering user-level, zero-copy RDMA, which we coupled with the ARMv8 processors in the MPSoCs. We further developed the corresponding runtimes that allow us to test real MPI applications on the large-scale testbed. We evaluated our platform through MPI microbenchmarks, mini application, and full MPI applications. Single-hop, one-way latency is 1.3 μs; approximately 0.47 μs out of these are attributed to network interface and the user-space library that exposes its functionality to the runtime. Latency over longer paths increases as expected, reaching 2.55 μs for a five-hop path. Bandwidth tests show that, for single-hop, link utilization reaches \(82\%\) of the theoretical capacity. Microbenchmarks based on MPI collectives reveal that broadcast latency scales as expected when the number of participating ranks increases. We also implemented a custom MPI_Allreduce accelerator in the network interface, which reduces the latency of such collectives by up to \(88\%\) . We assess performance scaling through weak and strong scaling tests for HPCG, LAMMPS, and the miniFE mini application; for all these tests, parallelization efficiency is at least \(69\%\) , or better. Manolis Ploumidis, Fabien Chaix, Nikolaos Chrysos, Marios Assiminakis, Nikolaos D. Kallimanis, Nikolaos Kossifidis, Michael Nikoloudakis, Nikolaos Dimou, Michalis Gianioudis, Giorgos Ieronymakis, Aggelos Ioannou, George Kalokerinos, Pantelis Xirouchakis, Astrinos Damianakis, Michael Ligerakis, Theocharis Vavouris, Manolis Katevenis, Vassilis Papaefstathiou, Manolis Marazakis, Iakovos Mavroidis |
ACM Trans. Reconfigurable Technol. Syst. | 19 |
| 2023 | eProcessor: European, Extendable, Energy-Efficient, Extreme-Scale, Extensible, Processor EcosystemabstractThe eProcessor project aims at creating a RISC-V full stack ecosystem. The eProcessor architecture combines a high-performance out-of-order core with energy-efficient accelerators for vector processing and artificial intelligence with reduced-precision functional units. The design of this architecture follows a hardware/software co-design approach with relevant application use cases from the high-performance computing, bioinformatics and artificial intelligence domains. Two eProcessor prototypes will be developed based on two fabricated eProcessor ASICs integrated into a computer-on-module. Lluc Alvarez, Abraham Ruiz, Arnau Bigas-Soldevilla, Pavel Kuroedov, Alberto González 0004, Hamsika Mahale, Noe Bustamante, Albert Aguilera, Francesco Minervini, Javier Salamero, Oscar Palomar, Vassilis Papaefstathiou, Antonis Psathakis, Nikolaos Dimou, Michalis Giaourtas, Iasonas Mastorakis, Giorgos Ieronymakis, Georgios-Michail Matzouranis, Vassilis Flouris, Nikolaos Kossifidis, Manolis Marazakis, Bhavishya Goel, Madhavan Manivannan, Ahsen Ejaz, Panagiotis Strikos, Mateo Vázquez, Ioannis Sourdis, Pedro Trancoso, Per Stenström, Jens Hagemeyer, Lennart Tigges, Nils Kucza, Jean-Marc Philippe, Ioannis Papaefstathiou |
CF | 21 |
| 2023 | Impact of Cache Coherence on the Performance of Shared-Memory based MPI Primitives: A Case Study for Broadcast on Intel Xeon Scalable ProcessorsabstractRecent processor advances have made feasible HPC nodes with high core counts, capable of hosting tens or even, hundreds of processes. Therefore, designing MPI collective operations at the intra-node level has received significant attention over the past years. Deriving efficient algorithms for modern HPC nodes, with complex internal topologies and memory hierarchies, is challenging. Moreover, the cache coherency protocol, and its impact on performance, further complicate algorithm design for MPI collectives. This latter concern is often only partially addressed. George Katevenis, Manolis Ploumidis, Manolis Marazakis |
ICPP | 3 |
| 2022 | LatEst: Vertical elasticity for millisecond serverless executionabstractCurrent state-of-the-art serverless frameworks can-not execute functions within a few milliseconds for bursty work-loads. The reason for that is that typically they rely on horizontal elasticity to cope with the varying demand for resources, which induces high overhead in the event of a cold -start. Recent literature has focused on minimizing the overhead of horizontal elasticity using mechanisms such as snapshots. However, the spawning of new function instances still requires several tens of milliseconds. This paper proposes vertical elasticity to scale resources of serverless functions to cope with bursting workloads. We design LatEst, a controller for serverless frameworks that adapts the allocated resources of active function instances. U sing vertical scaling, LatEst can adjust to bursts of function invocations within a few milliseconds. LatEst implements a feed-back control loop to: (1) predict the required resources during workload changes and (2) react rapidly and accurately to such changes. Moreover, LatEst spawns new instances for functions when the resources of the underlying server are reaching their limit. We evaluate LatEst as an extension of vHive [1] and find that LatEst can improve tail latency of serverless functions up to 25x compared to vHive. Yannis Sfakianakis, Manolis Marazakis, Christos Kozanitis, Angelos Bilas |
CCGRID | 2 |
| 2022 | A framework for hierarchical single-copy MPI collectives on multicore nodesabstractCollective operations are widely used by MPI applications to realize their communication patterns. Their efficiency is crucial for both performance and scalability of parallel applications. For deriving efficient MPI implementations, significant effort is put to keep pace with advances and capabilities of the underlying hardware and interconnect. Recent processor advances have led to nodes with higher core counts and complex internal structures and memory hierarchies. Such nodes are able to host tens to hundreds of processes and thus, performance of MPI collectives at the intra-node level becomes critical. In this work, we propose a framework for collective operations at the intra-node level, that aims to lower latency and increase bandwidth. Our approach utilizes knowledge of internal node structure to construct hierarchical algorithms, and XPMEM to achieve single-copy transfers. Pipelining is used to overlap communication at different levels of the hierarchy. We evaluate the proposed approach through several microbenchmarks and real-world MPI applications. For evaluation purposes, we compare the proposed approach with implementations of similar schemes from two recent studies. Our evaluation with microbenchmarks for Broadcast and Allreduce shows speedup up to 2.$5x$and$3x_{2}$respectively, over UCC and OpenMPI's default collectives implementation. Compared to recent research studies, we improve Broadcast by up to$5x_{2}$and Allreduce by up to$7x$. We reduce the time of three applications PiSvM, miniAMR and CNTK, by up to 12%, 52% and 12%, respectively, over the next best-performing alternative. George Katevenis, Manolis Ploumidis, Manolis Marazakis |
CLUSTER | 3 |
| 2021 | Skynet: Performance-driven Resource Management for Dynamic WorkloadsabstractA primary concern for cloud operators is to increase resource utilization while maintaining good performance for applications. This is particularly difficult to achieve for three reasons: users tend to overprovision applications, applications are diverse and dynamic, and their performance depends on multiple resources. In this paper, we present Skynet, an automated and adaptive cloud resource management approach that addresses all three concerns. Skynet uses performance level objectives (PLOs) to capture user intentions about required performance more accurately to remove the user from the resource allocation loop. Then, Skynet estimates the resources required to achieve the target PLO. For this purpose, we employ a Proportional Integral Derivative (PID) controller per application and adjust its parameters on the fly. Finally, to capture the dependence of applications on different or multiple resources, Skynet extends the traditional one-dimensional PID controller to estimate CPU, memory, I/O throughput, and network throughput. Essentially, Skynet builds a model on-the-fly to map target PLOs to resources for each application, taking into account multiple resources and changing input load. We implement Skynet as an end-to-end, custom scheduler in Kubernetes and evaluate it using real workloads on both a private cluster and AWS. Skynet decreases PLO violations by more than 7.4x and increases resource utilization by more than 2x, compared to Kubernetes. Essentially, Skynet builds a model on-the-fly to map target PLOs to resources for each application, taking into account multiple resources and changing input load. We implement Skynet as an end-to-end, custom scheduler in Kubernetes and evaluate it using real workloads on both a private cluster and AWS. Skynet decreases PLO violations by more than 7.4x and increases resource utilization by more than 2x, compared to Kubernetes. Yannis Sfakianakis, Manolis Marazakis, Angelos Bilas |
CLOUD | 2 |
| 2021 | IOTier: A Virtual Testbed to evaluate systems for IoT environmentsabstractInternet of Things (IoT) is an emerging field characterized by constrained resources, Internet-based communication, arbitrary topologies, geographical distance, and variable operational conditions. Additionally, IoT architectures typically exhibit at least three tiers: IoT devices, Edge gateways, Cloud servers. On top of challenging the design of networked systems, multiple tiers create a web of complexity that makes systems evaluation a challenging endeavor. This paper presents a framework for transforming a cluster of lab machines into a Virtual Testbed that provides views of how systems will perform in a tiered IoT environment. Experiments with constrained resources (CPU, memory, block device, network), multiple tiers, and programmables events are presented and discussed. Their effects are analyzed on the common path operation of micro-benchmarks and distributed key/value store. Fotios Nikolaidis, Manolis Marazakis, Angelos Bilas |
CCGRID | 2 |
| 2021 | Trace-Based Workload Generation and Execution
Yannis Sfakianakis, Eleni Kanellou, Manolis Marazakis, Angelos Bilas |
Euro-Par | 3 |
| 2021 | Memory-mapped I/O on steroidsabstractWith current technology trends for fast storage devices, the host-level I/O path is emerging as a main bottleneck for modern, data-intensive servers and applications. The need to improve I/O performance requires customizing various aspects of the I/O path, including the page cache and the method to access the storage devices. Anastasios Papagiannis, Manolis Marazakis, Angelos Bilas |
EuroSys | 2 |
| 2020 | DyRAC: Cost-aware Resource Assignment and Provider Selection for Dynamic Cloud WorkloadsabstractA primary concern for cloud users is how to minimize the total cost of ownership of cloud services. This is not trivial to achieve due to workload dynamics. Users need to select the number, size, type of VMs, and the provider to host their services based on available offerings. To avoid the complexity of re-configuring a cloud service, related work commonly approaches cost minimization as a packing problem that minimizes the resources allocated to services. However, this approach does not consider two problem dimensions that can further reduce cost: (1) provider selection and (2) VM sizing. In this paper, we explore a more direct approach to cost minimization by adjusting the type, number, size of VM instances, and the provider of a cloud service (i.e. a service deployment) at runtime. Our goal is to identify the limits in service cost reduction by online re-deployment of cloud services. For this purpose, we design DyRAC, an adaptive resource assignment mechanism for cloud services that, given the resource demands of a cloud service, estimates the most cost-efficient deployment. Our evaluation implements four different resource assignment policies to provide insight into how our approach works, using VM configurations of actual offerings from main providers (AWS, GCP, Azure). Our experiments show that DyRAC reduces cost by up to 33% compared to typical strategies. Yannis Sfakianakis, Manolis Marazakis, Angelos Bilas |
ICPADS | 2 |
| 2020 | Towards Communication Profile, Topology and Node Failure Aware Process PlacementabstractHPC systems need to keep growing in size to meet the ever-increasing demand for high levels of capability and capacity, often in tight time windows for urgent computation. However, increasing the size, complexity and heterogeneity of HPC systems also increases the risk and impact of system failures, that result in resource waste and aborted jobs. A major contributor to job completion time is the cost of interprocess communication. To address performance and energy efficiency, several prior studies have targeted improvements of communication locality. To meet this goal, they derive a mapping of MPI processes to system nodes in a way that reduces communication cost. However, such approaches disregard the effect of system failures. In this work, we propose a resource allocation approach for MPI jobs, considering both high performance and error resilience. Our approach, named Communication Profile, Topology and node Failure (CPTF), takes into account the application's communication profile, system topology and node failure probability for assigning job processes to nodes. We evaluate variants of CPTF through simulations of two MPI applications, one with a regular communication pattern (LAMMPS) and one with an irregular one (NPB-DT). In both cases, the variant of CPTF that strives to avoid failure-prone nodes and communication paths achieves lower time to complete job batches when compared to the default resource allocation policy of Slurm. It also exhibits the lowest ratio of aborted jobs. The average improvement in batch completion time is 67% for NPB-DT and 34% for LAMMPS. Ioannis Vardas, Manolis Ploumidis, Manolis Marazakis |
SBAC-PAD | 3 |
| 2020 | Optimizing Memory-mapped I/O for Fast Storage Devices
Anastasios Papagiannis, Giorgos Xanthakis, Giorgos Saloustros, Manolis Marazakis, Angelos Bilas |
USENIX ATC | 4 |
| 2019 | Towards Exascale: Measuring the Energy Footprint of Astrophysics HPC SimulationsabstractThe increasing amount of data produced in Astronomy by observational studies and the size of theoretical problems to be tackled in the next future pushes the need of HPC (High Performance Computing) resources towards the "Exascale". The HPC sector is undergoing a profound phase of transition, in which one of the toughest challenges to cope with is the energy efficiency that is one of the main blocking factors to the achievement of "Exascale". Since ideal peak-performance is unlikely to be achieved in realistic scenarios, the aim of this work is to give some insights about the energy consumption of contemporary architectures with real scientific applications in a HPC context. We use two state-of-the-art applications from the astrophysical domain, that we optimized in order to fully exploit the underlying hardware: a direct N-body code and a semi-analytical code for Cosmic Structure formation simulations. For these two applications, we quantitatively evaluate the impact of computation on the energy consumption when running on three different systems: one that represents the present of current HPC systems (an Intel-based cluster), one that (possibly) represents the future of HPC systems (a prototype of an Exascale supercomputer) and a micro-cluster based on Arm MPSoC. We provide a comparison of the time-to-solution, energy-to-solution and energy delay product (EDP) metrics, for different software configurations. ARM-based HPC systems have lower energy consumption albeit running ≈10 times slower. Giuliano Taffoni, Manolis Katevenis, Renato Panchieri, Gino Perna, Luca Tornatore, David Goz, Antonio Ragagnin, Sara Bertocco, Igor Coretti, Manolis Marazakis, Fabien Chaix, Manolis Ploumidis |
eScience | 10 |
| 2018 | Mainstream vs. Emerging HPC: Metrics, Trade-Offs and Lessons LearnedabstractVarious servers with different characteristics and architectures are hitting the market, and their evaluation and comparison in terms of HPC features is complex and multidimensional. In this paper, we share our experience of evaluating a diverse set of HPC systems, consisting of three mainstream and five emerging architectures. We evaluate the performance and power efficiency using prominent HPC benchmarks, High-Performance Linpack (HPL) and High Performance Conjugate Gradients (HPCG), and expand our analysis using publicly available specialized kernel benchmarks, targeting specific system components. In addition to a large body of quantitative results, we emphasize six usually overlooked aspects of the HPC platforms evaluation, and share our conclusions and lessons learned. Overall, we believe that this paper will improve the evaluation and comparison of HPC platforms, making a first step towards a more reliable and uniform methodology. Milan Radulovic, Kazi Asifuzzaman, Darko Zivanovic, Nikola Rajovic, Guillaume Colin de Verdière, Dirk Pleiter, Manolis Marazakis, Nikolaos D. Kallimanis, Paul M. Carpenter, Petar Radojkovic, Eduard Ayguadé |
SBAC-PAD | 7 |
| 2017 | Paving the Way Towards a Highly Energy-Efficient and Highly Integrated Compute Node for the Exascale Revolution: The ExaNoDe ApproachabstractPower consumption and high compute density are the key factors to be considered when building a compute node for the upcoming Exascale revolution. Current architectural design and manufacturing technologies are not able to provide the requested level of density and power efficiency to realise an operational Exascale machine. A disruptive change in the hardware design and integration process is needed in order to cope with the requirements of this forthcoming computing target. This paper presents the ExaNoDe H2020 research project aiming to design a highly energy efficient and highly integrated heterogeneous compute node targeting Exascale level computing, mixing low-power processors, heterogeneous co-processors and using advanced hardware integration technologies with the novel UNIMEM Global Address Space memory system. Alvise Rigo, Christian Pinto, Kevin Pouget, Daniel Raho, Denis Dutoit, Pierre-Yves Martinez, Chris Doran, Luca Benini, Iakovos Mavroidis, Manolis Marazakis, Valeria Bartsch, Guy Lonsdale, Antoniu Pop, John Goodacre, Annaik Colliot, Paul M. Carpenter, Petar Radojkovic, Dirk Pleiter, Dominique Drouin, Benoît Dupont de Dinechin |
DSD | 10 |
| 2016 | EUROSERVER: Share-anything scale-out micro-server design
Manolis Marazakis, John Goodacre, Didier Fuin, Paul M. Carpenter, John Thomson, Emil Matús, Antimo Bruno, Per Stenström, Jérôme Martin, Yves Durand, Isabelle Dor |
DATE | 1 |
| 2016 | The ExaNeSt Project: Interconnects, Storage, and Packaging for Exascale SystemsabstractExaNest is one of three European projects that support a ground-breaking computing architecture for exascale-class systems built upon power-efficient 64-bit ARM processors. This group of projects share an "everything-close" and "share-anything" paradigm, which trims down the power consumption -- by shortening the distance of signals for most data transfers -- as well as the cost and footprint area of the installation -- by reducing the number of devices needed to meet performance targets. In ExaNeSt, we will design and implement: (i) a physical rack prototype and its liquid-cooling subsystem providing ultra-dense compute packaging, (ii) a storage architecture with distributed (in-node) non-volatile memory (NVM) devices, (iii) a unified, low-latency interconnect, designed to efficiently uphold desired Quality-of-Service guarantees for a mix of storage with inter-processor flows, and (iv) efficient rack-level memory sharing, where each page is cacheable at only a single node. Our target is to test alternative storage and interconnect options on actual hardware, using real-world HPC applications. The ExaNeSt consortium brings together technology, skills, and knowledge across the entire value chain, from computing IP, packaging, and system deployment, all the way up to operating systems, storage, HPC, big data frameworks, and cutting-edge applications. Manolis Katevenis, Nikolaos Chrysos, Manolis Marazakis, Iakovos Mavroidis, Fabien Chaix, Nikolaos D. Kallimanis, Javier Navaridas, John Goodacre, Piero Vicini, Andrea Biagioni, Pier Stanislao Paolucci, Alessandro Lonardo, Elena Pastorelli, Francesca Lo Cicero, Roberto Ammendola, P. Hopton, P. Coates, Giuliano Taffoni, Stefano Cozzini, Martin L. Kersten, Julio Sahuquillo, Sergio Lechago, C. Pinto, Bernd Lietzow, D. Everett, Gino Perna |
DSD | 3 |
| 2014 | Vanguard: Increasing Server Efficiency via Workload Isolation in the Storage I/O PathabstractServer consolidation via virtualization is an essential technique for improving infrastructure cost in modern datacenters. From the viewpoint of datacenter operators, consolidation offers compelling advantages by reducing the number of physical servers, and reducing operational costs such as energy consumption. However, performance interference between co-located workloads can be crippling. Conservatively, and at significant cost, datacenter operators are forced to keep physical servers at low utilization levels (typically below 20%), to minimize adverse performance interactions. Yannis Sfakianakis, Stelios Mavridis, Anastasios Papagiannis, Spyridon Papageorgiou, Markos Fountoulakis, Manolis Marazakis, Angelos Bilas |
SoCC | 6 |
| 2014 | EUROSERVER: Energy Efficient Node for European Micro-ServersabstractEUROSERVER is a collaborative project that aims to dramatically improve data centre energy-efficiency, cost, and software efficiency. It is addressing these important challenges through the coordinated application of several key recent innovations: 64-bit ARM cores, 3D heterogeneous silicon-on-silicon integration, and fully-depleted silicon-on-insulator (FD SOI) process technology, together with new software techniques for efficient resource management, including resource sharing and workload isolation. We are pioneering a system architecture approach that allows specialized silicon devices to be built even for low-volume markets where NRE costs are currently prohibitive. The EUROSERVER device will embed multiple silicon "chiplets" on an active silicon interposer. Its system architecture is being driven by requirements from three use cases: data centres and cloud computing, telecom infrastructures, and high-end embedded systems. We will build two fully integrated full-system prototypes, based on a common micro-server board, and targeting embedded servers and enterprise servers. Yves Durand, Paul M. Carpenter, Stefano Adami, Angelos Bilas, Denis Dutoit, Alexis Farcy, Georgi Gaydadjiev, John Goodacre, Manolis Katevenis, Manolis Marazakis, Emil Matús, Iakovos Mavroidis, John Thomson |
DSD | 10 |
| 2014 | Jericho: Achieving scalability through optimal data placement on multicore systemsabstractAchieving high I/O throughput on modern servers presents significant challenges. With increasing core counts, server memory architectures become less uniform, both in terms of latency as well as bandwidth. In particular, the bandwidth of the interconnect among NUMA nodes is limited compared to local memory bandwidth. Moreover, interconnect congestion and contention introduce additional latency on remote accesses. These challenges severely limit the maximum achievable storage throughput and IOPS rate. Therefore, data and thread placement are critical for data-intensive applications running on NUMA architectures. In this paper we present Jericho, a new I/O stack for the Linux kernel that improves affinity between application threads, kernel threads, and buffers in the storage I/O path. Jericho consists of a NUMA-aware filesystem and a DRAM cache organized in slices mapped to NUMA nodes. The Jericho filesystem implements our task placement policy by dynamically migrating application threads that issue I/Os based on the location of the corresponding I/O buffers. The Jericho DRAM I/O cache, a replacement for the Linux page-cache, splits buffer memory in slices, and uses per-slice kernel I/O threads for I/O request processing. Our evaluation shows that running the FIO microbenchmark on a modern 64-core server with an unmodified Linux kernel results in only 5% of the memory accesses being served by local memory. With Jericho, more than 95% of accesses become local, with a corresponding 2x performance improvement. Stelios Mavridis, Yannis Sfakianakis, Anastasios Papagiannis, Manolis Marazakis, Angelos Bilas |
MSST | 4 |
| 2013 | FDIO: A Feedback Driven Controller for Minimizing Energy in I/O-Intensive Applications
Ioannis Manousakis, Manolis Marazakis, Angelos Bilas |
HotStorage | 2 |
| 2012 | Understanding Scalability and Performance Requirements of I/O-Intensive Applications on Future Multicore ServersabstractToday, there is increased interest in understanding the impact of data-centric applications on compute and storage infrastructures as datasets are projected to grow dramatically. In this paper, we examine the storage I/O behavior of twelve data-centric applications as the number of cores per server grows. We configure these applications with realistic datasets and examine configuration points where they perform significant amount of I/O. We propose using cycles per I/O (cpio) as a metric for abstracting many I/O subsystem configuration details. We analyze specific architectural issues pertaining to data-centric applications including the usefulness of hyperthreading, sensitivity to memory bandwidth, and the potential impact of disruptive storage technologies. Our results show that today's data-centric applications are not able to scale with the number of cores: moving from one to eight cores, results in 0% to 400% more cycles per I/O operation. These applications can achieve much of their performance with only 50% of the memory bandwidth available on modern processors. Hyper-threading is extremely effective for these applications and, on average, applications suffer only a 15% reduction in performance when hyper-threading is used instead of full cores. Further, DRAM-type persistent memory has the potential to solve scalability bottlenecks by reducing or eliminating idle and I/O completion periods and improving server utilization. We use a detailed methodology to project that in the year 2020, at 4096 processors, servers will require between 250-500 GB/s under optimistic scaling assumptions. We show that if the current trend in application scalability is not reversed, we will need about 2.5M servers that will consume 10 BKWh of energy to do a single pass over the projected 35 Zeta Bytes of data in 2020. Shoaib Akram 0001, Manolis Marazakis, Angelos Bilas |
MASCOTS | 2 |
| 2012 | Transparent Online Storage Compression at the Block-LevelabstractIn this work, we examine how transparent block-level compression in the I/O path can improve both the space efficiency and performance of online storage. We present ZBD , a block-layer driver that transparently compresses and decompresses data as they flow between the file-system and storage devices. Our system provides support for variable-size blocks, metadata caching, and persistence, as well as block allocation and cleanup. ZBD targets maintaining high performance, by mitigating compression and decompression overheads that can have a significant impact on performance by leveraging modern multicore CPUs through explicit work scheduling. We present two case-studies for compression. First, we examine how our approach can be used to increase the capacity of SSD-based caches, thus increasing their cost-effectiveness. Then, we examine how ZBD can improve the efficiency of online disk-based storage systems. We evaluate our approach in the Linux kernel on a commodity server with multicore CPUs, using PostMark, SPECsfs2008, TPC-C, and TPC-H. Preliminary results show that transparent online block-level compression is a viable option for improving effective storage capacity, it can improve I/O performance up to 80% by reducing I/O traffic and seek distance, and has a negative impact on performance, up to 34%, only when single-thread I/O latency is critical. In particular, for SSD-based caching, our results indicate that, in line with current technology trends, compressed caching trades off CPU utilization for performance and enhances SSD efficiency as a storage cache up to 99%. Yannis Klonatos, Thanos Makatos, Manolis Marazakis, Michail Flouris, Angelos Bilas |
ACM Trans. Storage | 3 |
| 2011 | Azor: Using Two-Level Block Selection to Improve SSD-Based I/O CachesabstractFlash-based solid state drives (SSDs) exhibit potential for solving I/O bottlenecks by offering superior performance over hard disks for several workloads. In this work we design Azor, an SSD-based I/O cache that operates at the block-level and is transparent to existing applications, such as databases. Our design provides various choices for associativity, write policies and cache line size, while maintaining a high degree of I/O concurrency. Our main contribution is that we explore differentiation of HDD blocks according to their expected importance on system performance. We design and analyze a two-level block selection scheme that dynamically differentiates HDD blocks, and selectively places them in the limited space of the SSD cache. We implement Azor in the Linux kernel and evaluate its effectiveness experimentally using a server-type platform and large problem sizes with three I/O intensive workloads: TPC-H, SPECsfs 2008, and Hammerora. Our results show that as the cache size increases, Azor enhances I/O performance by up to 14.02×, 1.63×, and 1.55× for each workload respectively. Additionally, our two-level block selection scheme further enhances I/O performance compared to a typical SSD cache by up to 95%, 16%, and 34% for each workload, respectively. Yannis Klonatos, Thanos Makatos, Manolis Marazakis, Michail Flouris, Angelos Bilas |
NAS | 3 |
| 2010 | Using transparent compression to improve SSD-based I/O cachesabstractFlash-based solid state drives (SSDs) offer superior performance over hard disks for many workloads. A prominent use of SSDs in modern storage systems is to use these devices as a cache in the I/O path. In this work, we examine how transparent, online I/O compression can be used to increase the capacity of SSD-based caches, thus increasing the costeffectiveness of the system. We present FlaZ, an I/O system that operates at the block-level and is transparent to existing file-systems. To achieve transparent, online compression in the I/O path and maintain high performance, FlaZ, provides support for variable-size blocks, mapping of logical to physical blocks, block allocation, and cleanup. FlaZ, mitigates compression and decompression overheads that can have a significant impact on performance by leveraging modern multicore CPUs. We implement FlaZ, in the Linux kernel and evaluate it on a commodity server with multicore CPUs, using TPC-H, PostMark, and SPECsfs. Our results show that compressed caching trades off CPU cycles for I/O performance and enhances SSD efficiency as a cache by up to 99%, 25%, and 11% for each workload, respectively. Thanos Makatos, Yannis Klonatos, Manolis Marazakis, Michail Flouris, Angelos Bilas |
EuroSys | 3 |
| 2010 | DARC: design and evaluation of an I/O controller for data protectionabstractLately, with increasing disk capacities, there is increased concern about protection from data errors, beyond masking of device failures. In this paper, we present a prototype I/O stack for storage controllers that encompasses two data protection features: (a) persistent checksums to protect data at-rest from silent errors and (b) block-level versioning to allow protection from user errors. Although these techniques have been previously used either at the device level (checksums) or at the host (versioning), in this work we implement these features in the storage controller, which allows us to use any type of storage devices as well as any type of host I/O stack. The main challenge in our approach is to deal with persistent metadata in the controller I/O path. Our main contribution is to show the implications of introducing metadata at this level and to deal with the performance issues that arise. Overall, we find that data protection features can be incorporated in the I/O path with a performance penalty in the range of 12% to 25%, offering much stronger data protection guarantees than today's commodity storage servers. Markos Fountoulakis, Manolis Marazakis, Michail Flouris, Angelos Bilas |
SYSTOR | 2 |
| 2010 | Data-Centric Privacy Protocol for Intensive Care GridsabstractModern e-Health systems require advanced computing and storage capabilities, leading to the adoption of technologies like the grid and giving birth to novel health grid systems. In particular, intensive care medicine uses this paradigm when facing a high flow of data coming from intensive care unit's (ICU) inpatients just like demonstrated by the ICGrid system prototyped by the University of Cyprus. Unfortunately, moving an ICU patient's data from the traditionally isolated hospital's computing facilities to data grids via public networks (i.e., the Internet) makes it imperative to establish an integral and standardized security solution to avoid common attacks on the data and metadata being managed. Particular emphasis must be put on the patient's personal data, the protection of which is required by legislations in many countries of the European Union and the world in general. In this paper, we extend our previous research with the following contributions: 1) a mandatory access control model to protect patient's metadata; 2) a major security revision to our previously proposed privacy protocol by contributing with a "quality of security" quantitative metric to improve fragmented data's assurance; and finally, 3) a set of early results to demonstrate that our protocol not only improves a patient personal data's security and privacy but also achieves a performance comparable with existing approaches. Jesus Luna, Marios D. Dikaiakos, Manolis Marazakis, Theodoros C. Kyprianou |
IEEE Trans. Inf. Technol. Biomed. | 3 |
| 2008 | Providing security to the Desktop Data GridabstractVolunteer computing is becoming a new paradigm not only for the computational grid, but also for institutions using production-level data grids because of the enormous storage potential that may be achieved at a low cost by using commodity hardware within their own computing premises. However, this novel "Desktop Data Grid" depends on a set of widely distributed and untrusted storage nodes, therefore offering no guarantees about neither availability nor protection to the stored data. These security challenges must be carefully managed before fully deploying desktop data grids in sensitive environments (such as eHealth) to cope with a broad range of storage needs, including backup and caching. In this paper we propose a cryptographic protocol able to fulfil the storage security requirements related with a generic desktop data grid scenario, which were identified after applying an analysis framework extended from our previous research on the data grid's storage services. The proposed protocol uses three basic mechanisms to accomplish its goal: (a) symmetric cryptography and hashing, (b) an information dispersal algorithm and the novel (c) "quality of security" (QoSec) quantitative metric. Although the focus of this work is the associated protocol, we also present an early evaluation using an analytical model. Our results show a strong relationship between the assurance of the data at rest, the QoSec of the volunteer storage client and the number of fragments required to rebuild the original file. Jesus Luna, Michail Flouris, Manolis Marazakis, Angelos Bilas |
IPDPS | 3 |
| 2007 | Optimization and bottleneck analysis of network block I/O in commodity storage systemsabstractBuilding commodity networked storage systems is an important architectural trend; Commodity servers hosting a moderate number of consumer-grade disks and interconnected with a high-performance network are an attractive option for improving storage system scalability and cost-efficiency. However, such systems incur significant overheads and are not able to deliver to applications the available throughput. We examine in detail the sources of overheads in such systems, using a working prototype to quantify the overheads associated with various parts of the I/O protocol. We optimize our base protocol to deal with small requests by batching them at the network level and without any I/O-specific knowledge. We also redesign our protocol stack to allow for asynchronous event processing, in-line, during send-path request processing. These techniques improve performance for a 8-disk SATA RAID0 array from 200 to 290 MBytes/s (45 % improvement). Using a ramdisk, peak performance improves from 320 to 474 MBytes/s (48 % improvement), which is 72 % of the maximum possible throughput in our experimental setup. We also analyze the remaining system bottlenecks, and find that although commodity storage systems have potential for building high-performance I/O subsystems, traditional network and I/O protocols are not fully capable of delivering this potential. Manolis Marazakis, Vassilis Papaefstathiou, Angelos Bilas |
ICS | 1 |
| 2006 | Experiences from Debugging a PCIX-based RDMA-capable NICabstractImplementing and debugging high-performance network subsystems is a challenging task. In this paper, we present our experiences from developing and debugging a network interface card (NIC). Our NIC targets networked storage subsystems (Marazakis et al., 2006). For this purpose it mainly provides support for remote direct-memory-access (RDMA) write, sender-side notification of RDMA write completion, and receiver-side interrupt generation. In our work we examine issues that arise during system implementation and debugging, both in terms of correctness as well as performance. We present an analysis of the individual problems we encounter and we discuss how we address each case. For most problems we encounter, it is not possible to rely on existing debugging tools. However, we find that most of the techniques we use in this process, rely on collecting some form of event records from software or hardware components. We believe that such capabilities can be provided for independent hardware or software components in isolation, a fairly straight-forward task, thus, significantly simplifying the debugging process in complex systems of this nature Manolis Marazakis, Vassilis Papaefstathiou, Giorgos Kalokairinos, Angelos Bilas |
CLUSTER | 1 |
| 2006 | Efficient remote block-level I/O over an RDMA-capable NICabstractModern storage systems are required to scale to large storage capacities and I/O throughput in a cost effective manner. For this reason, they are increasingly being built out of commodity components, mainly PCs equipped with large numbers of disks and interconnected of high-performance system area networks. A main issue in these efforts is to achieve high I/O throughput over commodity, low-cost system area networks and commodity operating systems. In this work, we examine in detail the performance of remote block-level storage I/O over commodity, RDMA-capable network interfaces and networks. We examine the support that is required from the network interface for achieving high throughput. We also examine in detail the overheads associated in kernel-level protocols for networked storage access. We find that base system performance is limited by (a) interrupt cost, (b) request size, and (c) protocol message size. We examine the impact of techniques to alleviate these factors and find that our techniques combined can improve throughput by up to 100 % over a simpler unoptimized configuration. Our current prototype is able to achieve a throughput of about 200 MBytes/s over a network that is capable of delivering about 500 MBytes/s. We identify major limiting factors, mostly at the I/O target-side. Manolis Marazakis, Konstantinos Xinidis, Vassilis Papaefstathiou, Angelos Bilas |
ICS | 1 |
| 2000 | Decentralized resource acquisition from autonomous markets in a QoS-capable environment
Spyros Lalis, Dimitris Papadakis, Manolis Marazakis |
Decis. Support Syst. | 3 |
| 1999 | Effects of an Asynchronous Resource Allocation Protocol on End-to-End Service ProvisionabstractIn this paper we present a protocol for acquiring resource bundles from autonomously operating markets and discuss results of experiments that were performed to quantify the protocol performance at high loads. A major finding from from our experiments is that an application class may be penalized by experiencing delays in accessing a shared resource as a result of overload on resources used by other application classes, which are unknown to this class. In turn, these delays may cause under-utilization of other resources used by this application class. Moreover, we illustrate how such effects can emerge by varying the relative occurrence frequencies application classes, but without changing aggregate load of the system. Spyros Lalis, Manolis Marazakis, Dimitris Papadakis |
ISADS | 2 |
| 1998 | Aurora: An Architecture for Dynamic and Adaptive Work Sessions in Open Environments
Manolis Marazakis, Dimitris Papadakis, Christos Nikolaou |
DEXA | 1 |
| 1998 | System Infrastructure for Digital Libraries: A Survey and Outlook
Christos Nikolaou, Manolis Marazakis |
SOFSEM | 2 |
| 1997 | Transaction Routing for Distributed OLTP Systems: Survey and Recent Results
Christos Nikolaou, Manolis Marazakis, G. Georgiannakis |
Inf. Sci. | 2 |