EDBT 2026 Demo / reviewers in the wild / expert
Antonio Barbalace
dblp:132/0327
· DBLP profile ↗
34ranked-venue papers
6as first author
20since 2021 · last 2026
0000-0003-1641-0779ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 23 · 3 first-author · 12 since 2021Software engineering, systems software and programming languages · 11 · 3 first-author · 7 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | [Experiment, Analysis, and Benchmark] Systematic Evaluation of Plan-Based Adaptive Query ProcessingabstractUnreliable cardinality estimation remains a critical performance bottleneck in database management systems (DBMSs). Adaptive Query Processing (AQP) strategies address this limitation by providing a more robust query execution mechanism. Specifically, plan-based AQP achieves this by incrementally refining cardinality using feedback from the execution of sub-plans. However, the actual reason behind the improvements of plan-based AQP, especially across different storage architectures (on-disk vs. in-memory DBMSs), remains unexplored. This paper presents the first comprehensive analysis of state-of-the-art plan-based AQP. We implement and evaluate this strategy on both on-disk and in-memory DBMSs across two benchmarks. Our key findings reveal that while plan-based AQP provides overall speedups in both environments, the sources of improvement differ significantly. In the on-disk DBMS, PostgreSQL, performance gains primarily come from the query plan reorderings, but not the cardinality updating mechanism; in fact, updating cardinalities introduces measurable overhead. Conversely, in the in-memory DBMS, DuckDB, cardinality refinement drives significant performance improvements for most queries. We also observe significant performance benefits of the plan-based AQP compared to a state-of-the-art related-based AQP method. These observations provide crucial insights for researchers on when and why plan-based AQP is effective, and ultimately guide database system developers on the tradeoffs between the implementation effort and performance improvements. Pei Mu 0003, Anderson Chaves Carniel, Antonio Barbalace, Amir Shaikhha |
ICDE | 3 |
| 2026 | Photonic Quantum Computing on Spin Memory Architecture with Tree-Encoded Fusion
Yuexun Huang, Zhemin Zhang, Tsung-Yi Ho, Antonio Barbalace, Zhiding Liang |
ISCA | 6 |
| 2025 | Stramash: A Fused-Kernel Operating System For Cache-Coherent, Heterogeneous-ISA PlatformsabstractWe live in the world of heterogeneous computing. With specialised elements reaching all aspects of our computer systems and their prevalence only growing, we must act to rein in their inherent complexity. One area that has seen significantly less investment in terms of development is heterogeneous-ISA systems, specifically because of complexity. To date, heterogeneous-ISA processors have required significant software overheads, workarounds, and coordination layers, making the development of more advanced software hard, and motivating little further development of more advanced hardware. In this paper, we take a fused approach to heterogeneity, and introduce a new operating system (OS) design, the fused-kernel OS, which goes beyond the multiple-kernel OS design, exploiting cache-coherent shared memory among heterogeneous-ISA CPUs as a first principle -- introducing a set of new OS kernel mechanisms. We built a prototype fused-kernel OS, Stramash-Linux, to demonstrate the applicability of our design to monolithic OS kernels. We profile Stramash OS components on real hardware but tested them on an architectural simulator -- Stramash-QEMU, which we design and build. Our evaluation begins by validating the accuracy of our simulator, achieving an average of less than 4% errors. We then perform a direct comparison between our fused-kernel OS and state-of-the-art multiple-kernel OS designs. Results demonstrate speedups of up to 2.1× on NPB benchmarks. Further, we provide an in-depth analysis of the differences and trade-offs between fused-kernel and multiple-kernel OS designs. Tong Xing 0002, Cong Xiong, Tianrui Wei, April Sanchez, Binoy Ravindran, Jonathan Balkind, Antonio Barbalace |
ASPLOS (2) | 7 |
| 2025 | Rethinking Tiered Memory Management in Cloud Data CentersabstractCloud environments continue to experience substantial memory wastage due to inefficient resource sharing and workload variability. Emerging Compute Express Link (CXL) technology offers fabric-attached memory that can expand memory capacity despite added access latency, but existing VM allocation strategies in the Cloud have critical limitations. The static partitioning of local DRAM and CXL memory underutilize capacity, and cannot adapt to dynamic demands. Conversely, Host-managed tiering (software-managed placement) relies on page-table scans or sampling, which incur high CPU overhead. Tong Xing 0002, Jiaxun Yang, Javier Picorel, Antonio Barbalace |
SoCC | 4 |
| 2025 | A Scalable and Robust Compilation Framework for Emitter-Photonic Graph StateabstractQuantum graph states are critical resources for various quantum algorithms, and also determine essential interconnections in distributed quantum computing. There are two schemes for generating graph states - probabilistic scheme and deterministic scheme. While the all-photonic probabilistic scheme has garnered significant attention, the emitter-photonic deterministic scheme has been proved to be more scalable and feasible across several hardware platforms. This paper studies the GraphState-to-Circuit compilation problem in the context of the deterministic scheme. Previous research has primarily focused on optimizing individual circuit parameters, often neglecting the characteristics of quantum hardware, which results in impractical implementations. Additionally, existing algorithms lack scalability for larger graph sizes. To bridge these gaps, we propose a novel compilation framework that partitions the target graph state into subgraphs, compiles them individually, and subsequently combines and schedules the circuits to maximize emitter resource utilization. Furthermore, we incorporate local complementation to transform graph states and minimize entanglement overhead. Evaluation of our framework on various graph types demonstrates significant reductions in CNOT gates and circuit duration, up to 52% and 56%. Moreover, it enhances the suppression of photon loss, achieving improvements of up to $\times 1.9$. Yuexun Huang, Zhiding Liang, Antonio Barbalace |
DAC | 4 |
| 2025 | SmartNIC-Based Distributed Shared MemoryabstractAn emerging trend in the heterogeneous computing space is instruction-set-architecture (ISA) heterogeneity. Data center providers are increasingly incorporating ARM-based hardware into their high-end computing installations which are traditionally made up of x86-based servers. Another interesting trend is the emergence of “smart” I/O devices such as SmartNICs [2] and SmartSSDs which include full-featured SoCs. Hemanth Ramesh, Naarayanan Rao VSathish, Edson Horta, Antonio Barbalace, Binoy Ravindran |
FCCM | 4 |
| 2024 | CoSense: Compiler Optimizations using Sensor Technical SpecificationsabstractEmbedded systems are ubiquitous, but in order to maximize their lifetime on batteries there is a need for faster code execution – i.e., higher energy efficiency, and for reduced memory usage. The large number of sensors integrated into embedded systems gives us the opportunity to exploit sensors’ technical specifications, like a sensor’s value range, to guide compiler optimizations for faster code execution, small binaries, etc. We design and implement such an idea in COSENSE, a novel compiler (extension) based on the LLVM infrastructure, using an existing domain-specific language (DSL), NEWTON, to describe the bounds of and relations between physical quantities measured by sensors. COSENSE utilizes previously unexploited physical information correlated to program variables to drive code optimizations. COSENSE computes value ranges of variables and proceeds to overload functions, compress variable types, substitute code with constants and simplify the condition statements. We evaluated COSENSE using several microbenchmarks and two real-world applications on various platforms and CPUs. For microbenchmarks, COSENSE achieves 1.18× geomean speedup in execution time and 12.35% reduction on average in binary code size with 4.66% compilation time overhead on x86, and 1.23× geomean speedup in execution time and 10.95% reduction on average in binary code size with 5.67% compilation time overhead on ARM. For real-world applications, COSENSE achieves 1.70× and 1.50× speedup in execution time, 12.96% and 0.60% binary code reduction, 9.69% and 30.43% lower energy consumption, with a 26.58% and 24.01% compilation time overhead, respectively. Pei Mu 0003, Nikolaos Mavrogeorgis, Christos Vasiladiotis, Vasileios Tsoutsouras, Orestis Kaparounakis, Phillip Stanley-Marbell, Antonio Barbalace |
CC | 7 |
| 2024 | UNIFICO: Thread Migration in Heterogeneous-ISA CPUs without State TransformationabstractHeterogeneous-ISA processor designs have attracted considerable research interest. However, unlike their homogeneous-ISA counterparts, explicit software support for bridging ISA heterogeneity is required. The lack of a compilation toolchain ready to support heterogeneous-ISA targets has been a major factor hindering research in this exciting emerging area. For any such compiler “getting right” the mechanics involved in state transformation upon migration and doing this efficiently is of critical importance. In particular, any runtime conversion of the current program stack from one architecture to another would be prohibitively expensive. In this paper, we design and develop Unifico, a new multi-ISA compiler that generates binaries that maintain the same stack layout during their execution on either architecture. Unifico avoids the need for runtime stack transformation, thus eliminating overheads associated with ISA migration. Additional responsibilities of the Unifico compiler backend include maintenance of a uniform ABI and virtual address space across ISAs. Unifico is implemented using the LLVM compiler infrastructure, and we are currently targeting the x86-64 and ARMv8 ISAs. We have evaluated Unifico across a range of compute-intensive NAS benchmarks and show its minimal impact on overall execution time, where less than 6% overhead is introduced on average. When compared against the state-of-the-art Popcorn compiler, Unifico reduces binary size overhead from ∼200% to ∼10%, whilst eliminating the stack transformation overhead during ISA migration. Nikolaos Mavrogeorgis, Christos Vasiladiotis, Pei Mu 0003, Amir Khordadi, Björn Franke, Antonio Barbalace |
CC | 6 |
| 2024 | A Hardware-Aware Gate Cutting Framework for Practical Quantum Circuit KnittingabstractCircuit knitting emerges as a promising technique to overcome the limitation of the few physical qubits in near-term quantum hardware by cutting large quantum circuits into smaller subcircuits. Recent research in this area has been primarily oriented towards reducing subcircuit sampling overhead. Unfortunately, these works neglect hardware information during circuit cutting, thus posing significant challenges to the follow on stages. In fact, direct compilation and execution of these partitioned subcircuits yields low-fidelity results, highlighting the need for a more holistic optimization strategy. Antonio Barbalace |
ICCAD | 3 |
| 2024 | UTwinVM: Reliable hints on the effects of hypervisor updates on VMs in the CloudabstractWe investigate the problem of getting hints on the effects of virtualization system (aka hypervisor) updates impact on virtual machines (VMs). System administrators can be reluctant to apply updates due to vague hints regarding the updates' impact on running applications. The problem is challenging since VMs are black boxes by design, reducing the scope of the data that can be retrieved and analyzed. Additionally, cloning VMs is only sometimes possible for obvious legal and privacy concerns. Djob Mvondo, Tong Xing 0002, Antonio Barbalace |
Middleware | 3 |
| 2024 | Near-Storage Processing in FaaS Environments with FuncletsabstractServerless computing has disrupted how computation is performed in the Cloud. The ability to write Functions, and not care about infrastructure brings many benefits, including significantly lower deployment costs, improved developer workflow, scalability, resilience, and resource utilization. However, being stateless and not tied to specific machines, Functions need to access Cloud storage services to access data, which may require crossing the entire data center network incurring high overheads. Existing solutions either provide database APIs running on the storage servers to perform the data-intensive operations locally, or deploy entire Functions on storage servers. The former approach can not perform arbitrary computations locally on storage servers. The latter violates the principle of compute-storage disaggregation, resulting in poor scaling. We observe that allowing on-the-fly migration of I/O-intensive parts of Functions to storage nodes achieves both objectives. We propose a FaaS runtime that runs on both the compute servers and the servers running the storage services which introduces an efficient migration mechanism for Functions across machines to move I/O-intensive parts of Functions to the relevant storage node, with minimal code changes. This allows Functions to perform arbitrary computations on storage nodes, benefiting from the locality of data, without sacrificing the scalability offered by compute-storage disaggregation. We implement our approach in a state-of-the-art FaaS runtime and show that it improves latency and throughput in bandwidth-constrained FaaS workloads while making better utilization of idle CPU cycles on storage servers. Alan Nair, Raven Szewczyk, Donald Jennings, Antonio Barbalace |
Middleware | 4 |
| 2024 | Offloading Datacenter Jobs to RISC-V Hardware for Improved Performance and Power EfficiencyabstractThe end of Moore's Law has brought significant changes in the architecture of servers used in data centers, increasingly incorporating new ISAs beyond x86-64 as well as diverse accelerators. Further, single-board computers have become increasingly efficient and can run certain Linux applications at significantly lower equipment and energy costs compared to traditional servers. Past research has demonstrated that offloading applications at runtime from x86-based servers to ARM-based single-board computers can result in increases in throughput and energy efficiency. The RISC-V architecture has recently gained significant commercial interest, and OS-capable single-board computers with RISC-V cores are increasingly available at the commodity scale. Balvansh Heerekar, Cesar Philippidis, Ho-Ren Chuang, Pierre Olivier, Antonio Barbalace, Binoy Ravindran |
SYSTOR | 5 |
| 2024 | HEXO: Offloading Long-Running Compute- and Memory-Intensive Workloads on Low-Cost, Low-Power Embedded SystemsabstractOS-capable embedded systems exhibiting a very low power consumption are available at an extremely low price point. It makes them highly compelling in a datacenter context. We show that sharing long-running, compute-intensive datacenter workloads between a server machine and one or a few connected embedded boards of negligible cost and power consumption can yield significant performance and energy benefits. Our approach, named Heterogeneous EXecution Offloading (HEXO), selectively offloads Virtual Machines (VMs) from server-class machines to embedded boards. Our design tackles several challenges. We address the Instruction Set Architecture (ISA) difference between typical servers (x86) and embedded systems (ARM) through hypervisor and guest OS-level support for heterogeneous-ISA runtime VM migration. We cope with the low amount of resources in embedded systems by using lightweight VMs – unikernels – and by using the server's free RAM as remote memory for embedded boards through a transparent lightweight memory disaggregation mechanism for heterogeneous server-embedded clusters, called Netswap. VMs are offloaded based on an estimation of the slowdown expected from running on a given board. We build a prototype of HEXO and demonstrate significant increases in throughput (up to 67%) and energy efficiency (up to 56%) using benchmarks representative of compute-intensive long-running workloads. Pierre Olivier, A. K. M. Fazla Mehrab, Sandeep Errabelly, Stefan Lankes, Mohamed Lamine Karaoui, Robert Lyerly, Sang-Hoon Kim, Antonio Barbalace, Binoy Ravindran |
IEEE Trans. Cloud Comput. | 8 |
| 2023 | Maximizing VMs' IO Performance on Overcommitted CPUs with FairnessabstractTo improve resource utilization and reduce costs many Cloud providers adopt virtual machines (VMs) overcommitment. While effective, this strategy may lead to adverse outcomes, significantly affecting a VM IO performance when one virtual CPU (vCPU) is preempted by another vCPU within the same runqueue of the VM scheduler -- i.e., same physical CPU (pCPU). Additionally, the responsiveness of a VM is reduced during the inactive time of the vCPU, and it necessitates an extra schedule timeslice to react to any IO event. While such problems have been studied in academia and industry, no previous solution has been deployed in production. This is because for example certain solutions require modifications of the guest VM, which is in contrast with industry requirements. Tong Xing 0002, Cong Xiong, Chuan Ye, Javier Picorel, Antonio Barbalace |
SoCC | 6 |
| 2023 | Aggregate VM: Why Reduce or Evict VM's Resources When You Can Borrow Them From Other Nodes?abstractHardware resource fragmentation is a common issue in data centers. Traditional solutions based on migration or overcommitment are unacceptably slow, and modern commercial or research solutions like Spot VM may reduce or evict VM's resources anytime. We propose an alternative solution that does not suffer from these drawbacks, the Aggregate VM. We introduce a new distributed hypervisor design, the resource-borrowing hypervisor, which creates Aggregate VMs: distributed VMs that temporarily aggregate fragmented resources belonging to different host machines, which require mobility of virtual CPUs, memory and IO devices. We implement a prototype, FragVisor, which runs guest software transparently. We also propose minimal modifications to the guest OS that can enable significant performance gains. We evaluate FragVisor over a set of microbenchmarks and IaaS-style real applications. Although Aggregate VMs are not a perfect fit for every type of applications, some workloads enjoy significant speedups compared to overcommitted scenarios (up to 3.9x with 4 distributed vCPUs). We further demonstrate that FragVisor is faster than a state-of-the-art competitor, GiantVM (up to 2.5x). Ho-Ren Chuang, Karim Manaouil, Tong Xing 0002, Antonio Barbalace, Pierre Olivier, Balvansh Heerekar, Binoy Ravindran |
EuroSys | 4 |
| 2022 | Secure and Policy-Compliant Query Processing on Heterogeneous Computational Storage ArchitecturesabstractComputation Storage Architectures (CSA) are increasingly adopted in the cloud for near data processing, where the underlying storage devices/servers are now equipped with heterogeneous cores which enable computation offloading near to the data. While CSA is a promising high-performance architecture for the cloud, in general data analytics also presents significant data security and policy compliance (e.g., GDPR) challenges in untrusted cloud environments. In this paper, we present IronSafe, a secure and policy-compliant query processing system for heterogeneous computational storage architectures, while preserving the performance advantages of CSA in untrusted cloud environments. To achieve these design properties in a computing environment with heterogeneous host (x86) and storage system (ARM), we design and implement the entire hardware and software system stack from the ground-up leveraging hardware-assisted Trusted Execution Environments (TEEs): namely, Intel SGX and ARM TrustZone. More specifically, IronSafe builds on three core contributions: (1) a heterogeneous confidential computing framework for shielded execution with x86 and ARM TEEs and associated secure storage system for the untrusted storage medium; (2) a policy compliance monitor to provide a unified service for attestation and policy compliance; and (3) a declarative policy language and associated interpreter for concisely specifying and efficiently evaluating a rich set of polices. Our evaluation using the TPC-H SQL benchmark queries and GDPR anti-pattern use-cases shows that IronSafe is faster, on average by 2.3x than a host-only secure system, while providing strong security and policy-compliance properties. Harshavardhan Unnibhavi, David Cerdeira, Antonio Barbalace, Nuno Santos 0001, Pramod Bhatotia |
SIGMOD Conference | 3 |
| 2021 | Computational Storage: Where Are We Today?
Antonio Barbalace, Jaeyoung Do |
CIDR | 1 |
| 2021 | Tell me when you are sleepy and what may wake you up!abstractNowadays, there is a shift in the deployment model of Cloud and Edge applications. Applications are now deployed as a set of several small units communicating with each other - the microservice model. Moreover, each unit - a microservice, may be implemented as a virtual machine, container, function, etc., spanning the different Cloud and Edge service models including IaaS, PaaS, FaaS. A microservice is instantiated upon the reception of a request (e.g., an http packet or a trigger), and a rack-level or data-center-level scheduler decides the placement for such unit of execution considering for example data locality and load balancing. With such a configuration, it is common to encounter scenarios where different units, as well as multiple instances of the same unit, may be running on a single server at the same time. Djob Mvondo, Antonio Barbalace, Alain Tchana, Gilles Muller |
SoCC | 2 |
| 2021 | Xar-trek: run-time execution migration among FPGAs and heterogeneous-ISA CPUsabstractDatacenter servers are increasingly heterogeneous: from x86 host CPUs, to ARM or RISC-V CPUs in NICs/SSDs, to FPGAs. Previous works have demonstrated that migrating application execution at run-time across heterogeneous-ISA CPUs can yield significant performance and energy gains, with relatively little programmer effort. However, FPGAs have often been overlooked in that context: hardware acceleration using FPGAs involves statically implementing select application functions, which prohibits dynamic and transparent migration. We present Xar-Trek, a new compiler and run-time software framework that overcomes this limitation. Xar-Trek compiles an application for several CPU ISAs and select application functions for acceleration on an FPGA, allowing execution migration between heterogeneous-ISA CPUs and FPGAs at run-time. Xar-Trek's run-time monitors server workloads and migrates application functions to an FPGA or to heterogeneous-ISA CPUs based on a scheduling policy. We develop a heuristic policy that uses application workload profiles to make scheduling decisions. Our evaluations conducted on a system with x86-64 server CPUs, ARM64 server CPUs, and an Alveo accelerator card reveal 88%-l% performance gains over no-migration baselines. Edson Horta, Ho-Ren Chuang, Naarayanan Rao VSathish, Cesar Philippidis, Antonio Barbalace, Pierre Olivier, Binoy Ravindran |
Middleware | 5 |
| 2021 | H-Container: Enabling Heterogeneous-ISA Container Migration in Edge ComputingabstractEdge computing is a recent computing paradigm that brings cloud services closer to the client. Among other features, edge computing offers extremely low client/server latencies. To consistently provide such low latencies, services should run on edge nodes that are physically as close as possible to their clients. Thus, when the physical location of a client changes, a service should migrate between edge nodes to maintain proximity. Differently from cloud nodes, edge nodes integrate CPUs of different Instruction Set Architectures (ISAs), hence a program natively compiled for a given ISA cannot migrate to a server equipped with a CPU of a different ISA. This hinders migration to the closest node. We introduce H-Container, a system that migrates natively compiled containerized applications across compute nodes featuring CPUs of different ISAs. H-Container advances over existing heterogeneous-ISA migration systems by being (a) highly compatible – no user’s source-code nor compiler toolchain modifications are needed; (b) easily deployable – fully implemented in user space, thus without any OS or hypervisor dependency, and (c) largely Linux-compliant – it can migrate most Linux software, including server applications and dynamically linked binaries. H-Container targets Linux and its already-compiled executables, adopts LLVM, extends CRIU, and integrates with Docker. Experiments demonstrate that H-Container adds no overheads during program execution, while 10–100 ms are added during migration. Furthermore, we show the benefits of H-Container in real-world scenarios, demonstrating, for example, up to 94% increase in Redis throughput when client/server proximity is maintained through heterogeneous container migration. Tong Xing 0002, Antonio Barbalace, Pierre Olivier, Mohamed Lamine Karaoui, Wei Wang 0512, Binoy Ravindran |
ACM Trans. Comput. Syst. | 2 |
| 2020 | Edge computing: the case for heterogeneous-ISA container migrationabstractEdge computing is a recent computing paradigm that brings cloud services closer to the client. Among other features, edge computing offers extremely low client/server latencies. To consistently provide such low latencies, services need to run on edge nodes that are physically as close as possible to their clients. Thus, when a client changes its physical location, a service should migrate between edge nodes to maintain proximity. Differently from cloud nodes, edge nodes are built with CPUs of different Instruction Set Architectures (ISAs), hence a server program natively compiled for one ISA cannot migrate to another. This hinders migration to the closest node. Antonio Barbalace, Mohamed Lamine Karaoui, Wei Wang 0512, Tong Xing 0002, Pierre Olivier, Binoy Ravindran |
VEE | 1 |
| 2019 | kMVX: Detecting Kernel Information Leaks with Multi-variant ExecutionabstractKernel information leak vulnerabilities are a major security threat to production systems. Attackers can exploit them to leak confidential information such as cryptographic keys or kernel pointers. Despite efforts by kernel developers and researchers, existing defenses for kernels such as Linux are limited in scope or incur a prohibitive performance overhead. In this paper, we present kMVX, a comprehensive defense against information leak vulnerabilities in the kernel by running multiple diversified kernel variants simultaneously on the same machine. By constructing these variants in a careful manner, we can ensure they only show divergences when an attacker tries to exploit bugs present in the kernel. By detecting these divergences we can prevent kernel information leaks. Our kMVX design is inspired by multi-variant execution (MVX). Traditional MVX designs cannot be applied to kernels because of their assumptions on the run-time environment. kMVX, on the other hand, can be applied even to commodity kernels. We show our Linux-based prototype provides powerful protection against information leaks at acceptable performance overhead (20--50% in the worst case for popular server applications). Sebastian Österlund, Koen Koning, Pierre Olivier, Antonio Barbalace, Herbert Bos, Cristiano Giuffrida |
ASPLOS | 4 |
| 2019 | Rethinking Communication in Multiple-kernel OSes for New Shared Memory InterconnectsabstractFuture computer platforms will likely be built with a multitude of on-chip and off-chip processing units being potentially of different ISAs, OS-capable, and sharing memory with a form of consistency. Multiple-kernel OSes, from multikernels to single-system image OSes, have been demonstrated to mange such platforms efficiently, but they assume no shared memory between kernels as a founding principle. This position paper proposes a new multiple-kernel OS design, which leverages consistent shared memory across homogeneous and heterogeneous processing units in a machine. Among other benefits, this design enables porting commodity SMP OSes to such future platforms, capitalizing on their shared memory programming model, and extend them to multiple-kernel OSes. Herein we present such design, based on two new software primitives tackling the problem of sharing and data format differences between eventually heterogeneous computing units: typed shared memory and type-morphable executable code. We also describe an initial implementation built around Popcorn Linux for x86 and ARM. Antonio Barbalace, Pierre Olivier, Binoy Ravindran |
PLOS@SOSP | 1 |
| 2018 | AIRA: A Framework for Flexible Compute Kernel Execution in Heterogeneous PlatformsabstractHeterogeneous-ISA computing platforms have become ubiquitous, and will be used for diverse workloads which render static mappings of computation to processors inadequate. Dynamic mappings which adjust an application's usage in consideration of platform workload can reduce application latency and increase throughput for heterogeneous platforms. We introduce AIRA, a compiler and runtime for flexible execution of applications in CPU-GPU platforms. Using AIRA, we demonstrate up to a 3.78× speedup in benchmarks from Rodinia and Parboil, run with various workloads on a server-class platform. Additionally, AIRA is able to extract up to an 87 percent increase in platform throughput over a static mapping. Robert Lyerly, Alastair Murray, Antonio Barbalace, Binoy Ravindran |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2017 | Breaking the Boundaries in Heterogeneous-ISA DatacentersabstractEnergy efficiency is one of the most important design considerations in running modern datacenters. Datacenter operating systems rely on software techniques such as execution migration to achieve energy efficiency across pools of machines. Execution migration is possible in datacenters today because they consist mainly of homogeneous-ISA machines. However, recent market trends indicate that alternate ISAs such as ARM and PowerPC are pushing into the datacenter, meaning current execution migration techniques are no longer applicable. How can execution migration be applied in future heterogeneous-ISA datacenters? Antonio Barbalace, Robert Lyerly, Christopher Jelesnianski, Anthony Carno, Ho-Ren Chuang, Vincent Legout, Binoy Ravindran |
ASPLOS | 1 |
| 2017 | It's Time to Think About an Operating System for Near Data Processing Architecturesabstractresearch-article Share on It's Time to Think About an Operating System for Near Data Processing Architectures Authors: Antonio Barbalace Huawei German Research Center Huawei German Research CenterView Profile , Anthony Iliopoulos Huawei German Research Center Huawei German Research CenterView Profile , Holm Rauchfuss Huawei German Research Center Huawei German Research CenterView Profile , Goetz Brasche Huawei German Research Center Huawei German Research CenterView Profile Authors Info & Claims HotOS '17: Proceedings of the 16th Workshop on Hot Topics in Operating SystemsMay 2017 Pages 56–61https://doi.org/10.1145/3102980.3102990Published:07 May 2017Publication History 17citation1,298DownloadsMetricsTotal Citations17Total Downloads1,298Last 12 Months108Last 6 weeks13 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Antonio Barbalace, Anthony Iliopoulos, Holm Rauchfuss, Goetz Brasche |
HotOS | 1 |
| 2017 | A Distributed Operating System Network Stack and Device Driver for MulticoresabstractWith the advances in network speeds a single processor cannot cope anymore with the growing number of data streams from a single network card. Multicore processors come at a rescue but traditional SMP OSes, which integrate the software network stack, scale only to a certain extent,limiting an application's ability to serve more connections while increasing the number of cores. On the other hand, kernel bypass solutions seem to scale better, but limit resource flexibility and control. We propose attacking these problems with a distributed OS design, using multiple network stacks (one per kernel) and relying on multi-queue hardware and hardware flow steering. This creates a single-socket abstraction among kernels while minimizing inter-core communication. We introduce our design, consisting of a distributed network stack, a distributed device driver, and a load-balancing algorithm. We compare our prototype, NetPopcorn, with Linux, Affinity Accept, FastSocket. NetPopcorn accepts between 5 to 8 times more connections and reduces the tail latency compared to these competitors. We also compare NetPopcorn with mTCP and observe that for high core counts, mTCP accepts only 18% more connections yet with higher tail latency than NetPopcorn. Saif Ansary, Antonio Barbalace, Ho-Ren Chuang, Thomas Lazor, Binoy Ravindran |
ICDCS | 2 |
| 2017 | Transparent Fault-Tolerance Using Intra-Machine Full-Software-Stack Replication on Commodity Multicore HardwareabstractAs the number of processors and the size of the memory of computing systems keep increasing, the likelihood of CPU core failures, memory errors, and bus failures increases and can threaten system availability. Software components can be hardened against such failures by running several replicas of a component on hardware replicas that fail independently and that are coordinated by a State-Machine Replication protocol. One common solution is to replicate the physical machine to provide redundancy, and to rewrite the software to address coordination. However, a CPU core failure, a memory error, or a bus error is unlikely to always crash an entire machine. Thus, full machine replication may sometimes be an overkill, increasing resource costs. In this paper, we introduce full software stack replication within a single commodity machine. Our approach runs replicas on fault-independent hardware partitions (e.g., NUMA nodes), wherein each partition is software-isolated from the others and has its own CPU cores, memory, and full software stack. A hardware failure in one partition can be recovered by another partition taking over its functionality. We have realized this vision by implementing FT-Linux, a Linux-based operating system that transparently replicates race-free, multithreaded POSIX applications on different hardware partitions of a single machine. Our evaluations of FT-Linux on several popular Linux applications show a worst case slowdown (due to replication) by ≈20%. Giuliano Losa, Antonio Barbalace, Yuzhong Wen, Ho-Ren Chuang, Binoy Ravindran |
ICDCS | 2 |
| 2017 | Swift Birth and Quick Death: Enabling Fast Parallel Guest Boot and Destruction in the Xen HypervisorabstractThe ability to quickly set up and tear down a virtual machine is critical for today's cloud elasticity, as well as in numerous other scenarios: guest migration/consolidation, event-driven invocation of micro-services, dynamically adaptive unikernel-based applications, micro-reboots for security or stability, etc. Vlad Nitu, Pierre Olivier, Alain Tchana, Daniel Chiba, Antonio Barbalace, Daniel Hagimont, Binoy Ravindran |
VEE | 5 |
| 2016 | A flattened hierarchical scheduler for real-time virtualizationabstractMigrating legacy real-time software stacks to newer hardware platforms can be achieved with virtualization which allows several software stacks to run on a single machine. Existing solutions guarantee that deadlines of virtualized real-time systems are met but can only accommodate a reduced number of systems. Therefore, this paper introduces ExVM, a new scheduling framework to maximize the number of legacy uniprocessor real-time systems able to run on a single machine. Contrary to most existing solutions, ExVM uses a flattening approach where the host schedules the virtual machine which contains the task with the earliest deadline. The real-time characteristics of tasks are obtained through introspection during the execution. We implemented this framework using Linux's SCHED_DEADLINE real-time scheduling policy in the host. Simulations using an exact schedulability test show that ExVM is able to schedule 96% of randomly generated tasksets with a utilization of at least 0.8, while state-of-the-art solutions are only able to schedule 40% of the same tasksets. Experimental evaluations performed using synthetic benchmarks and production real-time applications show that ExVM always outperforms the existing solutions, always meeting more than 80% of deadlines while these solutions fall below 50% when the utilization increases. Michael Drescher, Vincent Legout, Antonio Barbalace, Binoy Ravindran |
EMSOFT | 3 |
| 2015 | Popcorn: bridging the programmability gap in heterogeneous-ISA platformsabstractThe recent possibility of integrating multiple-OS-capable, high-core-count, heterogeneous-ISA processors in the same platform poses a question: given the tight integration between system components, can a shared memory programming model be adopted, enhancing programmability? If this can be done, an enormous amount of existing code written for shared memory architectures would not have to be rewritten to use a new programming paradigm (e.g., code offloading) that is often very expensive and error prone. We propose a new software architecture that is composed of an operating system and a compiler framework to run ordinary shared memory applications, written for homogeneous machines, on OS-capable heterogeneous-ISA machines. Applications run transparently amongst different ISA processors while exploiting the most optimized instruction set for each code block. We have implemented and tested our system, called Popcorn, on a multi-core Intel Xeon machine with a PCIe Intel Xeon Phi to demonstrate the viability of our approach. Application execution on Popcorn demonstrates to be up to 52% faster than the most performant native execution on Linux, on either Xeon or Xeon Phi, while removing the burden of the programmer having to adopt a different programming model than shared memory on a heterogeneous system. When compared to an offloading programming model, Popcorn is shown to be up to 6.2 times faster. Antonio Barbalace, Marina Sadini, Saif Ansary, Christopher Jelesnianski, Akshay Ravichandran, Cagil Kendir, Alastair Murray, Binoy Ravindran |
EuroSys | 1 |
| 2015 | Thread Migration in a Replicated-Kernel OSabstractChip manufacturers continue to increase the number of cores per chip while balancing requirements for low power consumption. This drives a need for simpler cores and hardware caches. Because of these trends, the scalability of existing shared memory system software is in question. Traditional operating systems (OS) for multiprocessors are based on shared memory communication between cores and are symmetric (SMP). Contention in SMP OSes over shared data structures is increasingly significant in newer generations of many-core processors. We propose the use of the replicated-kernel OS design to improve scalability over the traditional SMP OS. Our replicated-kernel design is an extension of the concept of the multikernel. While a multikernel appears to application software as a distributed network of cooperating micro kernels, we provide the appearance of a monolithic, single-system image, task-based OS in which application software is unaware of the distributed nature of the underlying OS. In this paper we tackle the problem of thread migration between kernels in a replicated-kernel OS. We focus on distributed thread group creation, context migration, and address space consistency for threads that execute on different kernels, but belong to the same distributed thread group. This concept is embodied in our prototype OS, called Popcorn Linux, which runs on multicore x86 machines and presents a Linux-like interface to application software that is indistinguishable from the SMP Linux interface. By doing this, we are able to leverage the wealth of existing Linux software for use on our platform while demonstrating the characteristics of the underlying replicated-kernel OS. We show that a replicated-kernel OS scales as well as a multikernel OS by removing the contention on shared data structures. Popcorn, Barr elfish, and SMP Linux are compared on selected benchmarks. Popcorn is shown to be competitive to SMP Linux, and up to 40% faster. David Katz, Antonio Barbalace, Saif Ansary, Akshay Ravichandran, Binoy Ravindran |
ICDCS | 2 |
| 2014 | KairosVM: Deterministic introspection for real-time virtual machine hierarchical schedulingabstractConsolidation and isolation are key technologies that promoted the undisputed popularity of virtualization in most of the computer industry. This popularity has recently led to a growing interest in real-time virtualization, making this technology enter the real-time system market. However, it has several issues due to the strict timing guarantees contracted. Moreover supporting legacy software stacks adds another level of complexity when the software is a black box. We present KairosVM, a latency-bounded, real-time extension to Linux's KVM module. It aims to bridge the lack of communication of the real-time requirements between the guest scheduler and the host scheduler, exploiting virtual machine introspection. The hypervisor captures the real-time requirements of the guest by catching previously added undefined instructions, without the need to do any modification to the guests. Our evaluations show that KairosVM's overhead is negligible when compared to existing introspection solutions thus can be used in real-time. Kevin Burns, Antonio Barbalace, Vincent Legout, Binoy Ravindran |
ETFA | 2 |
| 2013 | HSG-LM: hybrid-copy speculative guest OS live migration without hypervisorabstractCurrent Virtual Machine (VM) live migration mechanisms only focus on providing a high availability service by offering minimal downtime to users. In this paper, we present a novel live migration technique called HSG-LM, which also aims to provide short waiting time to whoever is responsible for triggering the VM migration (e.g., the data center administrator). HSG-LM is implemented in the guest OS kernel in order to not rely on the hypervisor throughout the entire migration process. HSG-LM exploits a hybrid strategy that reaps the benefits of both pre-copy and post-copy mechanisms. Furthermore, HSG-LM integrates a speculation mechanism that improves the efficiency of handling post-copy page faults. From our evaluation on different real-world workloads (Sysbench, Apache, etc.), the results show that HSG-LM incurs minimal downtime as well as short total migration time. Moreover, compared with competitors, HSG-LM reduces the downtime by up to 55%, and reduces the total migration time by up to 27%. Antonio Barbalace, Binoy Ravindran |
SYSTOR | 2 |