VLDB 2026 Research / reviewers in the wild / expert
Kostas Magoutis
dblp:99/3621
· DBLP profile ↗
35ranked-venue papers
4as first author
6since 2021 · last 2026
0000-0002-7288-5923ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 3 first-author · 1 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 3 since 2021Computer networks · 3 · 1 first-authorSecurity and privacy · 3Software engineering, systems software and programming languages · 3 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DEDALUS: A Quantum-Enhanced End-to-End Framework for Cost-Aware Join Order Optimization with Search Space Pruning
Emmanouil Limnaios, Markos Stergiopoulos, George T. Stamatiou, Efthymios Papageorgiou, Vasilis Efthymiou, Dimitrios Loupas, Dimitrios Tsourounis, Kostas Blekos, Aggelos Tsikas, Dimitris Plexousakis, Kostas Magoutis, Yannis Tzitzikas, Haridimos Kondylakis |
EDBT | 11 |
| 2026 | Architectural Foundations for Collaborative Machine Learning in Federated Data Spaces
Kostas Magoutis, Georgios Bouloukakis |
ICSA | 2 |
| 2025 | Proportional Fairness and Isolation for Serverless Applications over FaaS PlatformsabstractEffectively supporting multi-tenant application deployments in the emerging Function-as-a-Service (FaaS) (or serverless computing) model requires extending it with fairness and isolation mechanisms. Quality-of-service (QoS) concepts developed over time in the networking, storage, and virtualized infrastructure domains are currently being investigated in the space of serverless platforms. In this paper, we propose a two-level serverless QoS architecture that combines state-of-the-art scheduling algorithms and mechanisms with the unique characteristics of distributed serverless platforms, resulting into a system that provides proportional fairness for serverless applications with shared access to distributed and load-balanced FaaS platforms. The primary advantage of our approach is the use of higher-level scheduling mechanisms only, avoiding the need to manage low-level resources within underlying FaaS platforms (thus not requiring changes to them) for achieving fairness. We demonstrate the concrete benefits of our architecture using state-of-the-art benchmarks in experiments over AWS EC2. George Kelantonakis, Fallia Kourou, Kostas Magoutis |
ICPE | 3 |
| 2024 | Similarity Search based on Geo-footprints
Achilleas Michalopoulos, Konstantinos Lampropoulos 0002, George Kelantonakis, Chrysostomos Zeginis, Kostas Magoutis, Nikos Mamoulis |
EDBT | 5 |
| 2023 | SmartCityBus - A Platform for Smart Transportation SystemsabstractWith the growth of the Internet of Things (IoT), Smart(er) Cities have been a research goal of researchers, businesses and local authorities willing to adopt IoT technologies to improve their services. Among them, Smart Transportation [7,8], the integrated application of modern technologies and management strategies in transportation systems, refers to the adoption of new IoT solutions to improve urban mobility. These technologies aim to provide innovative solutions related to different modes of transport and traffic management and enable users to be better informed and make safer and 'smarter' use of transport networks. This talk presents SmartCityBus, a data-driven intelligent transportation system (ITS) whose main objective is to use online and offline data in order to provide accurate statistics and predictions and improve public transportation services in the short and medium/long term. Georgios Bouloukakis, Chrysostomos Zeginis, Kostas Magoutis, George Christodoulou 0005, Chrysanthi Kosyfaki, Konstantinos Lampropoulos 0002, Nikos Mamoulis |
WSDM | 4 |
| 2022 | Addressing the Read-Performance Impact of Reconfigurations in Replicated Key-Value StoresabstractRaw data are often orders of magnitude larger than main memory for many applications. As the performance of storage devices is still significantly slower than main memory, systems still rely on memory caching to improve performance. Data replication schemes are prevalent in data stores for high availability and reliability. In such schemes, while data updates are propagated to all replicas (either synchronously or in the background), reads are usually served by only a subset of replica group members (e.g., as in primary-backup and quorum systems). As a result, non-serving replicas cannot keep their memory cache state updated; thus, during a reconfiguration or a fail-over action, the system suffers from a high read-performance impact for a significant amount of time due to cold-cache misses. In our study we observed up to 70% hit after a reconfiguration due to cold cache misses, taking almost 18 minutes in some cases to fully restore to the pre-reconfiguration level of performance. In this article we propose a mechanism to maintain up-to-date read caches across replicas by sending read hints to the non-serving replicas to keep their caches warm. Thus the system is able to seamlessly achieve the same performance level even in the face of a replica group reorganization. This is especially important under the read-intensive workloads that are common today. Our evaluation shows that our mechanism has significant benefits during reconfigurations, with low performance impact under periods of resource strain. Given its advisory nature, the maintenance of read hints can be reduced or held off if needed during such periods. Antonis Papaioannou, Kostas Magoutis |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2020 | The Case for Better Integrating Scalable Data Stores and Stream-Processing SystemsabstractScalable stream processing systems require external storage systems for long-term storage of non-emphemeral state. Recent research have pointed to scalable in-memory key-value stores, such as Redis, as an efficient solution to external management of state [1], [2]. While such data stores have been interconnected with scalable streaming systems, they are currently managed independently, missing opportunities for optimizations, such as exploiting locality between stream partitions and table shards, as well as coordinating elasticity actions. Antonis Papaioannou, Chrysostomos Zeginis, Kostas Magoutis |
CLUSTER | 3 |
| 2019 | Towards Configurable Cloud Application SecurityabstractSecurity solutions for cloud applications usually exploit security tools as is by utilising their default configuration. On one hand, this can lead to a waste of resources. On the other hand, it can also lead to not properly protecting the different application components based on their diverse security requirements. To this end, this paper proposes a security solution for cross-cloud applications which is configurable according to the flexible configuration specification given by the devops. Such a specification conforms to a certain UML-based meta-model and is independent of the underlying security tools exploited. In this way, devops can enable to produce a varied security level per each application component that better suits its security requirements. We demonstrate the suitability of our solution through an evaluation showcasing that it can lead to reduced resource consumption without compromising the security of the components that it protects. Kyriakos Kritikos, Manos Papoutsakis, Sotiris Ioannidis, Kostas Magoutis |
CCGRID | 4 |
| 2019 | Decision-Making Approaches for Performance QoS in Distributed Storage Systems: A SurveyabstractDistributed storage systems designed to offer explicit performance quality-of-service (QoS) guarantees must regulate the allocation and use of resources to achieve a user-specified level of service. QoS-driven systems employ decision-making techniques to decide on appropriate actions to take during initial deployment or under variations in workload and/or system configuration. In this survey we cover both traditional approaches to decision-making for explicit performance QoS (control theory, multi-dimensional constrained optimization, policy-based techniques) as well as more recent approaches based on machine-learning, offering a broad perspective to the state-of-the-art in the field. As performance prediction is a central concept in decision-making, we also summarize research on performance prediction techniques used in this context. Flora Karniavoura, Kostas Magoutis |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2018 | Replica-Group Leadership Change as a Performance Enhancing Mechanism in NoSQL Data StoresabstractIn this paper we investigate replica-group reconfiguration as a way to mask performance bottlenecks on the primary node of a primary-backup replication group in a NoSQL data store. We investigate the benefit of changing replica-group leadership prior to resource-intensive background tasks such as LSM-tree compactions or data backups on the primary node, a method that can improve throughput by up to 23% during LSM-tree compactions and by 35% during backup tasks. Our implementation is based on MongoRocks (MongoDB 3.7 and RocksDB 5.7) using leveled compaction. We experimentally demonstrate the performance impact of compactions and data backups when they occur at replica-group primaries, and the benefits of targeted leadership-change actions. We evaluate our system using the Yahoo Cloud Serving Benchmark (YCSB) and compare to unmodified MongoRocks on dedicated infrastructure. Antonis Papaioannou, Kostas Magoutis |
ICDCS | 2 |
| 2017 | Incremental Elasticity for NoSQL Data StoresabstractElasticity actions in NoSQL data stores move large amounts of data over the network to take advantage of new resources. Here we propose incremental elasticity, a new mechanism for scheduling data transfers to a joining server, leading to smoother elasticity actions with a reduced performance impact. Antonis Papaioannou, Kostas Magoutis |
ICDCS | 2 |
| 2017 | Cross-layer management of a containerized NoSQL data storeabstractInternet-scale services increasingly rely on NoSQL data store technologies for scalable, highly available data persistence. To increase resource efficiency and deployment speed, such services are adopting new deployment models based on container technologies. Emerging container management systems (CMS) offer new models of interaction with containerized applications, exposing goal-oriented management capabilities. In this paper, we study the interaction of a containerized NoSQL data store, RethinkDB, with a popular open-source CMS, Kubernetes, demonstrating that events exposed by the CMS can be leveraged to drive adaptation actions on the NoSQL data store, improving its availability and quality of service. Typical adaptation actions in data stores often involve movement of data (e.g., migrating data replicas to a new node) and thus have high overhead. In this paper we demonstrate that lower-cost adaptation actions based on targeted reconfigurations of replica groups are possible and often offer rapid response to performance degradation (such as when a node experiences a temporary resource shortage). Such reconfigurations can be exercised by a cross-layer management system that bridges between CMS and NoSQL data store. We evaluate a prototype implementation of this management system in the context of Kubernetes and RethinkDB using the Yahoo Cloud Serving Benchmark on Google Container Engine. Our results demonstrate that infrastructure-level events provided by the CMS can drive proactive, low-cost data-store adaptation actions that improve overall system manageability, availability, and performance. Evdoxos Bekas, Kostas Magoutis |
IM | 2 |
| 2017 | A Measurement-Based Approach to Performance Prediction in NoSQL SystemsabstractIn this paper we evaluate a measurement-based approach to performance prediction of data-intensive applications over NoSQL systems. While the use of systematic measurements for building performance prediction models is a well studied topic, little attention has been paid so far on the application space of data-intensive systems using NoSQL databases. Measurement-based performance prediction approaches are often limited by a relatively narrow range of hardware characteristics available within each organization's private infrastructure. An opportunity to change this fact is the emergence of federated, publicly-accessible, large-scale research infrastructures, such as Fed4FIRE and GENI, featuring a variety of heterogeneous hardware. This paper demonstrates accurate measurement-based performance prediction modeling for NoSQL systems over such public infrastructures. We consider three prominent regression techniques: Multivariate adaptive regression splines (MARS), support vector regression (SVR), and artificial neural network (ANN) regression, applied to the YCSB data-intensive benchmark over the MongoDB NoSQL data store. Our measurements are drawn from 1-, 3-, and 5-node clusters of four node types. Performance prediction using MARS yields the best results with an average accuracy of 97.85% vs. 94.16% and 90.39% with SVR and ANN respectively. The approach can seamlessly extend to a wider range of hardware specifications available in federated research infrastructures. Flora Karniavoura, Kostas Magoutis |
MASCOTS | 2 |
| 2017 | Incremental Elasticity for NoSQL Data StoresabstractService elasticity, the ability to rapidly expand or shrink service processing capacity on demand, has become a first-class property in the domain of infrastructure services. Scalable NoSQL data stores are the de-facto choice of applications aiming for scalable, highly available data persistence. The elasticity of such data stores is still challenging, due to the complexity and performance impact of moving large amounts of data over the network to take advantage of new resources (servers). In this paper we propose incremental elasticity, a new mechanism that progressively increases processing capacity in a fine-grain manner during an elasticity action by making sub-sections of the transferred data available for access on the new server, prior to completing the full transfer. In addition, by scheduling data transfers during an elasticity action in sequence (rather than as simultaneous transfers) between each pre-existing server involved and the new server, incremental elasticity leads to smoother elasticity actions, reducing their overall impact on performance. Antonis Papaioannou, Kostas Magoutis |
SRDS | 2 |
| 2017 | A General-Purpose Architecture for Replicated Metadata Services in Distributed File SystemsabstractA large class of modern distributed file systems treat metadata services as an independent system component, separately from data servers. The availability of the metadata service is key to the availability of the overall system. Given the high rates of failures observed in large-scale data centers, distributed file systems usually incorporate high-availability (HA) features. A typical approach in the development of distributed file systems is to design and develop metadata services from the ground up, at significant cost in terms of complexity and time, often leading to functional shortcomings. Our motivation in this paper was to improve on this state of things by defining a general-purpose architecture for HA metadata services (which we call RMS) that can be easily incorporated and reused in new or existing file systems, reducing development time. Taking two prominent distributed file systems as case studies, PVFS and HDFS, we developed RMS variants that improve on functional shortcomings of the original HA solutions, while being easy to build and test. Our extensive evaluation of the RMS variant of HDFS shows that it does not incur an overall performance or availability penalty compared to the original implementation. Dimokritos Stamatakis, Nikos Tsikoudis, Eirini C. Micheli, Kostas Magoutis |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2016 | Towards Knowledge-Based Assisted IaaS SelectionabstractCurrent PaaS platforms enable single or hybrid cloud deployments. However, such deployment types cannot best cover the user application requirements as they do not consider the great variety of services offered by different cloud providers and the effects of vendor lock-in. On the other hand, multi-cloud deployment enables selecting the best possible service among equivalent ones providing the best trade-off between performance and cost. In addition, it avoids cases of service level deterioration due to service under-performance as main effects of vendor lock-in. While many multi-cloud application deployment research prototypes have been proposed, such prototypes do not examine the effect that deployment decisions have on application performance. As such, they blindly attempt to satisfy low-level hardware requirements by neglecting the impact of allocation decisions on higher-level requirements at the component or application level. To this end, this paper proposes a new IaaS selection algorithm which, apart from being able to satisfy both low and high level requirements of different types, it also exploits deployment knowledge offered via reasoning over previous application execution histories to take the best possible allocation decisions. The experimental evaluation clearly shows that by considering this extra knowledge, more optimal deployment solutions are derived, able to maintain the service levels requested by users, in less solving time. Kyriakos Kritikos, Kostas Magoutis, Dimitris Plexousakis |
CloudCom | 2 |
| 2015 | Cross-layer management of distributed applications on multi-cloudsabstractExisting cloud provisioning and deployment frameworks do not yet provide sufficient support for multi-cloud setups. In this paper we improve on the state of the art by adapting the SmartFrog framework to handle multi-cloud setups during the lifecycle (provisioning, deployment, change management, termination) of distributed applications. An administrator of a private cloud additionally needs to be aware of cross-layer dependencies between application components and the physical infrastructure to efficiently carry out a range of administration tasks. In this paper we propose an information repository and system to discover and store such cross-layer dependencies and use it to answer questions such as ”which physical machines are hosting particular application components (through hypervisors)” and vice versa. We demonstrate the use of our system in two cases of service-instance (SI) migration and replication on federated multi-cloud setups using the SPEC jEnterprise2010 application benchmark: a) migration of a SI from a private to a public cloud to support the need for maintaining service availability during downtime of physical hardware at the private cloud; b) replication of a SI from a large cloud provider to a regional (or ”near”) cloud provider to improve response time of clients that are geographically closer to the latter. Antonis Papaioannou, Damianos Metallidis, Kostas Magoutis |
IM | 3 |
| 2015 | Rethinking HBase: Design and Implementation of an Elastic Key-Value Store over Log-Structured Local VolumesabstractHBase is a prominent NoSQL system used widely in the domain of big data storage and analysis. It is structured as two layers: a lower-level distributed file system (HDFS)supporting the higher-level layer responsible for data distribution, indexing, and elasticity. Layered systems have in many occasions proven to suffer from overheads due to the isolation between layers, HBase is increasingly seen as an instance of this. To overcome this problem we designed, implemented, and evaluated HBase-BDB, an alternative to HBase that replaces the HDFS store with a thinner layer of a log-structured B+ tree key value store (Berkeley DB) operating over local volumes. We show that HBase-BDB overcomes HBase's performance bottlenecks (while retaining compatibility with HBase applications) without losing on elasticity features. We evaluate the performance of HBase and HBase-BDB using the Yahoo! Cloud Serving Benchmark (YCSB) and online transaction processing(OLTP) workloads on a commercial public Cloud provider. We find that HBase-BDB outperforms a tuned HBase configuration by up to 85% under a write-intensive workload due to HBase-BDB's reduced background-write activity. HBase-BDB's novel elasticity mechanisms operating over local volumes are shown to be as perform ant as HBase's equivalent features when stress-tested under TPC-C workloads. Giorgos Saloustros, Kostas Magoutis |
ISPDC | 2 |
| 2014 | ACaZoo: A Distributed Key-Value Store Based on Replicated LSM-TreesabstractIn this paper we describe the design and implementation of ACaZoo, a key-value store that combines strong consistency with high performance and high availability. ACaZoo supports the popular column-oriented data model of Apache Cassandra and HBase. It implements strongly-consistent data replication using primary-backup atomic broadcast of a write-ahead log, which records data mutations to a Log-structured Merge Tree (LSM-Tree). ACaZoo scales by horizontally partitioning the key space via consistent primary-key hashing on available replica groups (RGs). LSM-Tree compactions can hamper performance, especially when they take place at RG primaries. ACaZoo addresses this problem by changing RG leadership prior to heavy compactions, a method that can improve throughput by up to 40% in write-intensive workloads. We evaluate ACaZoo using the Yahoo Cloud Serving Benchmark (YCSB) and compare it to Oracle's NoSQL Database and to Cassandra providing serial consistency via an extension of the Paxos algorithm. Panagiotis Garefalakis, Panagiotis Papadopoulos, Kostas Magoutis |
SRDS | 3 |
| 2014 | Scalable entity-based summarization of web search results using MapReduce
Ioannis Kitsos, Kostas Magoutis, Yannis Tzitzikas |
Distributed Parallel Databases | 2 |
| 2013 | Managing Service Performance in the Cassandra Distributed Storage SystemabstractIn this paper we describe the architecture of a quality-of-service (QoS) infrastructure for achieving controlled application performance over the Cassandra distributed storage system. We present an implementation of our architecture and provide results from an evaluation using the Yahoo Cloud Serving Benchmark (YCSB) on the Amazon EC2 Cloud. A key focus of this paper is on a QoS-aware measurement-driven provisioning methodology. Our evaluation provides evidence that the methodology is effective in estimating application resource requirements and thus in achieving the type of controlled performance required by data intensive performance-critical applications. While our architecture is implemented and evaluated in the context of the Cassandra distributed storage system, its principles are general and can be applied to a variety of NoSQL systems. Maria Chalkiadaki, Kostas Magoutis |
CloudCom (1) | 2 |
| 2013 | An Architecture for Evaluating Distributed Application Deployments in Multi-cloudsabstractIn this paper we present an architecture for the modeling, collection, and evaluation of long-term histories of deployments of distributed multi-tier applications on federations of Clouds (Multi-Clouds). Our goal is to capture several aspects of application development and deployment lifecycle, including the evolving application structure, requirements, goals, and service level objectives, application deployment descriptions, runtime monitoring, and quality control, Cloud provider characteristics, and to provide a Cloud-independent resource classification scheme that is a key to reasoning about Multi-Cloud deployments of complex large-scale applications. Since our target is capturing the continuous evolution of applications and their deployments over time, we ensure that our metadata model is designed to optimize space usage. Additionally, we demonstrate that using the model and data collections over varying deployments of an application (using the SPEC jEnterprise2010 distributed benchmark as a case study) one can answer important questions about which deployment options work best in terms of performance, reliability, cost, and combinations thereof. Antonis Papaioannou, Kostas Magoutis |
CloudCom (1) | 2 |
| 2013 | Strengthening Consistency in the Cassandra Distributed Key-Value Store
Panagiotis Garefalakis, Panagiotis Papadopoulos, Ioannis Manousakis, Kostas Magoutis |
DAIS | 4 |
| 2013 | Topic 5: Parallel and Distributed Data Management - (Introduction)
María S. Pérez 0001, André Brinkmann, Stergios V. Anastasiadis, Sandro Fiore, Adrien Lèbre, Kostas Magoutis |
Euro-Par | 6 |
| 2012 | Scalability of Replicated Metadata Services in Distributed File Systems
Dimokritos Stamatakis, Nikos Tsikoudis, Ourania Smyrnaki, Kostas Magoutis |
DAIS | 4 |
| 2012 | Adapting data-intensive workloads to generic allocation policies in cloud infrastructuresabstractResource allocation policies in public Clouds are today largely agnostic to requirements that distributed applications have from their underlying infrastructure. As a result, assumptions about data-center topology that are built-into distributed data-intensive applications are often violated, impacting performance and availability goals. In this paper we describe a management system that discovers a limited amount of information about Cloud allocation decisions - in particular VMs of the same user that are collocated on a physical machine - so that data-intensive applications can adapt to those decisions and achieve their goals. Our distributed discovery process is based on either application-level techniques (measurements) or a novel lightweight and privacy-preserving Cloud management API proposed in this paper. Using the distributed Hadoop file system as a case study we show that VM collocation in a Cloud setup occurs in commercial platforms and that our methodologies can handle its impact in an effective, practical, and scalable manner. Ioannis Kitsos, Antonis Papaioannou, Nikos Tsikoudis, Kostas Magoutis |
NOMS | 4 |
| 2011 | CassMail: A Scalable, Highly-Available, and Rapidly-Prototyped E-Mail Service
Lazaros Koromilas, Kostas Magoutis |
DAIS | 2 |
| 2011 | CEC: Continuous eventual checkpointing for data stream processing operatorsabstractThe checkpoint roll-backward methodology is the underlying technology of several fault-tolerance solutions for continuous stream processing systems today, implemented either using the memories of replica nodes or a distributed file system. In this scheme the recovering node loads its most recent checkpoint and requests log replay to reach a consistent pre-failure state. Challenges with that technique include its complexity (typically implemented via copy-on-write), the associated overhead (exception handling under state updates), and limits to the frequency of checkpointing. The latter limit affects the amount of information that needs to be replayed leading to long recovery times. In this work we introduce continuous eventual checkpointing (CEC), a novel mechanism to provide fault-tolerance guarantees by taking continuous incremental state checkpoints with minimal pausing of operator processing. We achieve this by separating operator state into independent parts and producing frequent independent partial checkpoints of them. Our results show that our method can achieve low overhead fault-tolerance with adjustable checkpoint intensity, trading off recovery time with performance. Zoe Sebepou, Kostas Magoutis |
DSN | 2 |
| 2010 | Scalable storage support for data stream processingabstractContinuous data stream processing systems have offered limited support for data persistence in the past, for three main reasons: First, online, real-time queries examine current streaming data and (under the assumption of no server failures) do not require access to past data; second, stable storage devices are commonly thought to be constraining system throughput and response times when compared to main memory, and are thus kept off the common path; finally, the use of scalable storage solutions which would be required to sustain high data streaming rates have not been thoroughly investigated in the past. Our work advances the state of the art by providing data streaming systems with a scalable path to persistent storage. This path has low impact in the performance properties of a scalable streaming system and allows two fundamental enhancements to their capabilities: First, it allows stream persistence for reference/archival purposes (in other words, queries can now be applied on past data on-demand); second, fault tolerance is achievable by checkpointing and stream replay schemes that are not constrained by the size of main memory. Zoe Sebepou, Kostas Magoutis |
MSST | 2 |
| 2009 | Towards 100 gbit/s ethernet: multicore-based parallel communication protocol designabstractEthernet line rates are projected to reach 100 Gbits/s by as soon as 2010. While in principle suitable for high performance clustered and parallel applications, Ethernet requires matching improvements in the system software stack. In this paper we address several sources of CPU and memory system overhead in the I/O path at line rates reaching 80 Gbits/s (bi-directional), using multiple 10 Gbit/s links per system node. Key contributions of our work are the design of a parallel high-performance communication protocol that uses context-independent page-remapping to (a) reduce packet processing overheads; (b) reduce thread management and synchronization overheads; and (c) address affinity issues in NUMA multicore CPUs. Our design result in the full 40 Gbits/s of available one-way Ethernet bandwidth and in 57.6 Gbits/s (72%) of the 80 Gbits/s maximum bidirectional throughput (limited only by the memory system), while leaving ample CPU cycles for application processing. Stavros Passas, Kostas Magoutis, Angelos Bilas |
ICS | 2 |
| 2007 | Galapagos: Automatically Discovering Application-Data Relationships in Networked SystemsabstractIn large networked systems, relationships between applications and the data that they use through multiple tiers of middleware systems are often invisible. While the benefits of knowing such relationships are clear from a systems management perspective, discovery of such relationships is complicated by the widespread adoption of virtualization technologies and the tendency to view each middleware tier as an independent "domain" from a systems management perspective. In this paper we present a methodology and a system for automatic discovery of end-to-end application-data relationships. The key to the methodology is the modeling of data locations from which applications use data and of how middleware systems make data available to software layers above them. Kostas Magoutis, Murthy V. Devarakonda, Kiran Muniswamy-Reddy |
Integrated Network Management | 1 |
| 2005 | Memory Management Support for Multi-Programmed Remote Direct Memory Access (RDMA) SystemsabstractCurrent operating systems offer basic support for network interface controllers (NICs) supporting remote direct memory access (RDMA). Such support typically consists of a device driver responsible for configuring communication channels between the device and user-level processes but not involved in data transfer. Unlike standard NICs, RDMA-capable devices incorporate significant memory resources for address translation purposes. In a multi-programmed operating system (OS) environment, these memory resources must be efficiently shareable by multiple processes. For such sharing to occur in a fair manner, the OS and the device must cooperate to arbitrate access to NIC memory, similar to the way CPUs and OSes cooperate to arbitrate access to translation lookaside buffers (TLBs) or physical memory. A problem with this approach is that today's RDMA NICs are not integrated into the functions provided by OS memory management systems. As a result, RDMA NIC hardware resources are often monopolized by a single application. In this paper, I propose two practical mechanisms to address this problem: (a) Use of RDMA only in kernel-resident I/O subsystems, transparent to user-level software; (b) An extended registration API and a kernel upcall mechanism delivering NIC TLB entry replacement notifications to user-level libraries. Both options are designed to re-instate the multiprogramming principles that are violated in early commercial RDMA systems Kostas Magoutis |
CLUSTER | 1 |
| 2003 | Making the Most Out of Direct-Access Network Attached Storage
Kostas Magoutis, Salimah Addetia, Alexandra Fedorova, Margo I. Seltzer |
FAST | 1 |
| 2002 | Structure and Performance of the Direct Access File System
Kostas Magoutis, Salimah Addetia, Alexandra Fedorova, Margo I. Seltzer, Jeffrey S. Chase, Andrew J. Gallatin, Richard Kisley, Rajiv Wickremesinghe, Eran Gabber |
USENIX ATC, General Track | 1 |
| 2001 | Research Issues in No-Futz ComputingabstractAt the 1999 Workshop on Hot Topics in Operating Systems (HotOS VII), the attendees reached consensus that the most important issue facing the OS research community was "No-Futz" computing; eliminating the ongoing "futzing" that characterizes most systems today. To date, little research has been accomplished in this area. Our goal in this paper is to focus the research community on the challenges we face if we are to design systems that are truly futz-free, or even low-futz. David A. Holland, William K. Josephson, Kostas Magoutis, Margo I. Seltzer, Christopher A. Stein, Ada T. Lim |
HotOS | 3 |