VLDB 2026 Research / reviewers in the wild / expert
Alain Tchana
dblp:77/2224
· DBLP profile ↗
57ranked-venue papers
10as first author
23since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 34 · 4 first-author · 14 since 2021Software engineering, systems software and programming languages · 10 · 3 first-author · 4 since 2021Security and privacy · 6 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-authorComputer networks · 4 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Everything You Need to Know About Virtual Machine Live Migration Between Heterogeneous ProcessorsabstractThis paper focuses on the live migration of virtual machines (VMs) across semi-heterogeneous processors (SH), which share the same ISA but expose different features. Processor incompatibilities in such settings can block or break migrations, limiting resource utilization and operational flexibility. We provide the first detailed study of live migration feasibility across SH processors, yielding several findings valuable to cloud tenants, operators and hypervisor vendors. Building on these insights, we present MigCheck, a simulation tool that predicts the feasibility and outcome of VM migrations across SH processors without requiring costly real migrations. We demonstrate its accuracy and effectiveness through multiple use cases, including machine addition/removal, hypervisor updates, and VM platform adjustments. Kenta Ishiguro, Caleb Fonyuy-Asheri, Elouan Barraud, Renaud Lachaize, Yérom-David Bromberg, Alain Tchana |
EuroSys | 6 |
| 2026 | ZeroSwap: A Practical Solution to the Double Swapping Problem in Virtualized Environments
Luc Mahop, Kilian Kemgne, Baptiste Lepers, Fabienne Boyer, Alain Tchana |
ICDCS | 5 |
| 2026 | Contextual Cloud Steganography(CCS): Breaking the Capacity-Security Trade-OffabstractThe exponential growth of cloud storage adoption has introduced new attack surfaces and privacy concerns, necessitating advanced methods for secure and covert communication. Traditional steganographic techniques, particularly in distributed settings, often struggle with balancing undetectability, embedding capacity, and security. Recent indirect methods that avoid file modification have improved undetectability but remain constrained by low capacity and reliance on cryptographic secrecy alone. This paper introduces Contextual Cloud Steganography (CCS), a novel framework that replaces traditional base-based encoding with contextual protocols based on deterministic file ordering within cloud directories. CCS leverages pre-established sorting criteria (e.g., cryptographic hashes, creation dates) to embed secret data by selecting files from sorted lists, achieving a multiplicative increase in capacity of 3.2× to 4.3× over state-of-the-art base-B methods. The framework simultaneously enhances security through protocol-dependent protection, making message extraction computationally infeasible without the precise contextual protocol, even if the stego-folder is discovered. We provide formal security proofs under established steganographic models, detailed algorithms with cryptographic enhancements, and extensive experimental evaluation across major cloud platforms (Google Drive, Dropbox, OneDrive) over a 30-day period. Results demonstrate CCS's practical efficiency and robustness in dynamic environments, establishing it as a significant advancement in cloud-based covert communication that breaks the traditional capacity-security trade-off. Stéphane Willy Mossebo Tcheunteu, Leonel Moyou Metcheka, Hervé Talé Kalachi, Stephane Gael Raymond Ekodeck, René Ndoundam, Alain Tchana |
IEEE Trans. Cloud Comput. | 6 |
| 2025 | SigN: SIMBox Activity Detection Through Latency Anomalies at the Cellular EdgeabstractDespite their widespread adoption, cellular networks face growing vulnerabilities due to their inherent complexity and the integration of advanced technologies.One of the major threats in this landscape is Voice over IP (VoIP) to GSM gateways, known as SIMBox devices.These devices use multiple SIM cards to route VoIP traffic through cellular networks, enabling international bypass fraud with losses of up to $3.11 billion annually.Beyond financial impact, SIMBox activity degrades network performance, threatens national security, and facilitates eavesdropping on communications.Existing detection methods for SIMBox activity are hindered by evolving fraud techniques and implementation complexities, limiting their practical adoption in operator networks.This paper addresses the limitations of current detection methods by introducing SigN , a novel approach to identifying SIMBox activity at the cellular edge.The proposed method focuses on detecting remote SIM card association, a technique used by SIMBox appliances to mimic human mobility patterns.The method detects latency anomalies between SIMBox and standard devices by analyzing cellular signaling during network attachment.Extensive indoor and outdoor experiments demonstrate that SIMBox devices generate significantly higher attachment latencies, particularly during the authentication phase, where latency is up to 23 times greater than that of standard devices.We attribute part of this overhead to immutable factors such as LTE authentication standards and Internet-based communication protocols.Therefore, our approach offers a robust, scalable, and practical solution to mitigate SIMBox activity risks at the network edge. Josiane Kouam, Aline Carneiro Viana, Philippe Martins, Cédric Adjih, Alain Tchana |
AsiaCCS | 5 |
| 2025 | DISC: Backpressure Mitigation In Multi-tier Applications With Distributed Shared Connection
Brice Ekane, Djob Mvondo, Renaud Lachaize, Yérom-David Bromberg, Alain Tchana, Daniel Hagimont |
NSDI | 5 |
| 2024 | Battle of Wits: To What Extent Can Fraudsters Disguise Their Tracks in International bypass Fraud?abstractInternational bypass fraud, also known as SIMBox fraud, involves diverting international cellular voice traffic from regulated routes and rerouting it as local calls in the destination country. It has significantly affected cellular networks worldwide, generating $3.11 Billion of losses annually and threats to national security. Yet, SIMBox fraud remains an ongoing challenge, eluding operators detection due to the continual refinement of fraudulent behavior that is often overlooked in the design and validation of detection methods. Josiane Kouam, Aline Carneiro Viana, Alain Tchana |
AsiaCCS | 3 |
| 2024 | vPIM: Processing-in-Memory VirtualizationabstractData movement is the leading cause of performance degradation and energy consumption in modern data centers. Processing inmemory (PIM) is an architecture that addresses data movement by bringing computation inside the memory chips. This paper is the first to study the virtualization of PIM devices by designing and implementing vPIM, an open-source UPMEM-based virtualization system for the cloud. Our vPIM design considers four requirements: Compatibility such that no hardware and no hypervisor changes are needed; Multiplexing and isolation for a higher utilization ratio; Utilizability and transparency such that applications written for PIM can be efficiently run out-of-the-box, leading to rapid adoption; Minimalization of virtualization performance overhead. Dufy Teguia, Stella Bitchebe, Oana Balmau, Alain Tchana |
Middleware | 5 |
| 2024 | B-Side: Binary-Level Static System Call IdentificationabstractSystem call filtering is widely used to secure programs in multi-tenant environments, and to sandbox applications in modern desktop software deployment and package management systems. Filtering rules are hard to write and maintain manually, hence generating them automatically is essential. To that aim, analysis tools able to identify every system call that can legitimately be invoked by a program are needed. Existing static analysis works lack precision because of a high number of false positives, and/or assume the availability of program/libraries source code - something unrealistic in many scenarios such as cloud production environments. Gaspard Thévenon, Kevin Nguetchouang, Kahina Lazri, Alain Tchana, Pierre Olivier |
Middleware | 4 |
| 2024 | SVD: A Scalable Virtual Machine Disk FormatabstractContrary to CPU, memory, and network, disk virtualization is peculiar, for which virtualization through direct access is impossible. We study virtual disk utilization in a large-scale public cloud and observe the presence of long snapshot chains, sometimes composed of up to 1,000 files. We then demonstrate, through experimental measurements, that such long chains lead to virtualized storage performance and memory footprint scalability issues. To address these problems, we presentSVD, a new virtual disk format. We implementedSVDby extending Qcow2, a popular format, and its Qemu driver. We evaluated our prototype, demonstrating that it brings significant performance enhancements and memory footprint reduction. For example,SVDimproves the throughput of RocksDB by about 48% on a snapshot chain of length 500.SVDalso reduces the memory footprint by 15×. Kevin Nguetchouang, Stella Bitchebe, Théophile Dubuc, Mar Callau-Zori, Christophe Hubert, Pierre Olivier, Alain Tchana |
IEEE Trans. Cloud Comput. | 7 |
| 2023 | SGX Switchless Calls Made ConfiglessabstractIntel's software guard extensions (SGX) provide hardware enclaves to guarantee confidentiality and integrity for sensitive code and data. However, systems leveraging such security mechanisms must often pay high performance overheads. A major source of this overhead is SGX enclave transitions which induce expensive cross-enclave context switches. The Intel SGX SDK mitigates this with a switchless call mechanism for transitionless cross-enclave calls using worker threads. Intel's SGX switchless call implementation improves performance but provides limited flexibility: developers need to statically fix the system configuration at build time, which is error-prone and misconfigurations lead to performance degradations and waste of CPU resources. ZC-Switchless is a configless and efficient technique to drive the execution of SGX switchless calls. Its dynamic approach optimises the total switchless worker threads at runtime to minimise CPU waste. The experimental evaluation shows that ZC-Switchless obviates the performance penalty of misconfigured switchless systems while minimising CPU waste. Peterson Yuhala, Michael Paper, Timothée Zerbib, Pascal Felber, Valerio Schiavoni, Alain Tchana |
DSN | 6 |
| 2023 | SecV: Secure Code Partitioning via Multi-Language Secure ValuesabstractTrusted execution environments like Intel SGX provide enclaves, which offer strong security guarantees for applications. Running entire applications inside enclaves is possible, but this approach leads to a large trusted computing base (TCB). As such, various tools have been developed to partition programs written in languages such as C or Java into trusted and untrusted parts, which are run in and out of enclaves respectively. However, those tools depend on language-specific taint-analysis and partitioning techniques. They cannot be reused for other languages and there is thus a need for tools that transcend this language barrier. Peterson Yuhala, Pascal Felber, Hugo Guiroux, Jean-Pierre Lozi, Alain Tchana, Valerio Schiavoni, Gaël Thomas 0001 |
Middleware | 5 |
| 2023 | LSTM-based generation of cellular network trafficabstractDomain-wide recognized by their high value in human activity and network monitoring studies, cellular network traffic (i.e., Charging Data Records, named CDRs), however, present accessibility and usability issues, restricting their exploitation and research reproducibility. This paper tackles such challenges by modeling CDRs that fulfill real-world data attributes. Our designed framework, named Zen leverages LSTM to realistically model network users’ traffic behavior through a 4-stage generative pipeline. Results show that Zen’s models accurately capture individual and global distributions of a fully anonymized real-world traffic CDRs dataset. Finally, we validate Zen CDRs ability of reproducing daily cellular behaviors of the urban population and its usefulness in practical networking applications such as Radio Access Network’s power savings, and anomaly detection as compared to real-world CDRs. Josiane Kouam, Aline Carneiro Viana, Alain Tchana |
WCNC | 3 |
| 2023 | Networking in next generation disaggregated datacentersabstractSummary Nowadays, datacenters lean on a computer‐centric approach based on monolithic servers which include all necessary hardware resources (mainly CPU, RAM, network, and disks) to run applications. Such an architecture comes with two main limitations: (1) difficulty to achieve full resource utilization and (2) coarse granularity for hardware maintenance. Recently, many works investigated a resource‐centric approach called disaggregated architecture where the datacenter is composed of self‐content resource boards interconnected using fast interconnection technologies, each resource board including instances of one resource type. The resource‐centric architecture allows each resource to be managed (maintenance, allocation) independently. LegoOS is the first work which studied the implications of disaggregation on the operating system, proposing to disaggregate the operating system itself. They demonstrated the suitability of this approach, considering mainly CPU and RAM resources. However, they did not study the implication of disaggregation on network resources. We reproduced a LegoOS infrastructure and extended it to support disaggregated networking. We show that networking can be disaggregated following the same principles, and that classical networking optimizations such as DMA, DDIO, or loopback can be reproduced in such an environment. Our evaluations show the viability of the approach and the potential of future disaggregated infrastructures. Brice Ekane, Alain Tchana, Daniel Hagimont, Boris Teabe, Noel De Palma |
Concurr. Comput. Pract. Exp. | 2 |
| 2023 | HyperTP: A unified approach for live hypervisor replacement in datacenters
Tu Dinh Ngoc, Boris Teabe, Alain Tchana, Gilles Muller, Daniel Hagimont |
J. Parallel Distributed Comput. | 3 |
| 2022 | Odile: A scalable tracing tool for non-rooted and on-device Android phonesabstractAs Android’s popularity continues to grow among consumers and device manufacturers, it is also becoming a prime target for malware authors. Although static app analysis is quite simple to use and scale very well, it is inefficient when the app is obfuscated or the malicious code is dynamically downloaded at runtime. Runtime analysis of app behavior is thus becoming paramount for reverse engineers and app market maintainers (e.g., Google Play) to ensure that running apps do not include some malicious payload. Alain Tchana, Lavoisier Lavoisier Wapet, Yérom-David Bromberg |
RAID | 1 |
| 2022 | Out of Hypervisor (OoH): Efficient Dirty Page Tracking in Userspace Using Hardware Virtualization FeaturesabstractThis paper introduces Out of Hypervisor (OoH), a new virtualization research axis. Instead of emulating full virtual hardware inside a VM to support a hypervisor, the OoH principle is to individually expose current hypervisor-oriented hardware virtualization features to the guest OS. This way, guest's processes could also take benefit from those features. We illustrate OoH with Intel PML (Page Modification Logging), a feature that allows efficient dirty page tracking to improve VM live migration. Because dirty page tracking is at the heart of many essential tasks including process checkpointing (e.g., CRIU) and concurrent garbage collection (e.g, Boehm GC), OoH exposes PML to accelerate these tasks in the guest. We present two OoH solutions namely Shadow PML (SPML) and Extended PML (EPML) that we integrated into CRIU and Boehm GC. Evaluation results showed that EPML speeds up CRIU checkpointing by about 13 x and Boehm garbage collection by up to 6x compared to SPML, /proc, and userfaultfd while reducing their overhead on monitored applications by about 16 x. Stella Bitchebe, Alain Tchana |
SC | 2 |
| 2021 | Tell me when you are sleepy and what may wake you up!abstractNowadays, there is a shift in the deployment model of Cloud and Edge applications. Applications are now deployed as a set of several small units communicating with each other - the microservice model. Moreover, each unit - a microservice, may be implemented as a virtual machine, container, function, etc., spanning the different Cloud and Edge service models including IaaS, PaaS, FaaS. A microservice is instantiated upon the reception of a request (e.g., an http packet or a trigger), and a rack-level or data-center-level scheduler decides the placement for such unit of execution considering for example data locality and load balancing. With such a configuration, it is common to encounter scenarios where different units, as well as multiple instances of the same unit, may be running on a single server at the same time. Djob Mvondo, Antonio Barbalace, Alain Tchana, Gilles Muller |
SoCC | 3 |
| 2021 | Plinius: Secure and Persistent Machine Learning Model TrainingabstractWith the increasing popularity of cloud based machine learning (ML) techniques there comes a need for privacy and integrity guarantees for ML data. In addition, the significant scalability challenges faced by DRAM coupled with the high access-times of secondary storage represent a huge performance bottleneck for ML systems. While solutions exist to tackle the security aspect, performance remains an issue. Persistent memory (PM) is resilient to power loss (unlike DRAM), provides fast and fine-granular access to memory (unlike disk storage) and has latency and bandwidth close to DRAM (in the order of ns and GB/s, respectively). We present PLINIUS, a ML framework using Intel SGX enclaves for secure training of ML models and PM for fault tolerance guarantees. PLINIUS uses a novel mirroring mechanism to create and maintain (i) encrypted mirror copies of ML models on PM, and (ii) encrypted training data in byte-addressable PM, for near-instantaneous data recovery after a system failure. Compared to disk-based checkpointing systems, PLINIUS is 3.2× and 3.7× faster respectively for saving and restoring models on real PM hardware, achieving robust and secure ML model training in SGX enclaves. Peterson Yuhala, Pascal Felber, Valerio Schiavoni, Alain Tchana |
DSN | 4 |
| 2021 | OFC: an opportunistic caching system for FaaS platformsabstractCloud applications based on the "Functions as a Service" (FaaS) paradigm have become very popular. Yet, due to their stateless nature, they must frequently interact with an external data store, which limits their performance. To mitigate this issue, we introduce OFC, a transparent, vertically and horizontally elastic in-memory caching system for FaaS platforms, distributed over the worker nodes. OFC provides these benefits cost-effectively by exploiting two common sources of resource waste: (i) most cloud tenants overprovision the memory resources reserved for their functions because their footprint is non-trivially input-dependent and (ii) FaaS providers keep function sandboxes alive for several minutes to avoid cold starts. Using machine learning models adjusted for typical function input data categories (e.g., multimedia formats), OFC estimates the actual memory resources required by each function invocation and hoards the remaining capacity to feed the cache. We build our OFC prototype based on enhancements to the OpenWhisk FaaS platform, the Swift persistent object store, and the RAM-Cloud in-memory store. Using a diverse set of workloads, we show that OFC improves by up to 82 % and 60 % respectively the execution time of single-stage and pipelined functions. Djob Mvondo, Mathieu Bacou, Kevin Nguetchouang, Lucien Ngale, Stéphane Pouget, Josiane Kouam, Renaud Lachaize, Jinho Hwang, Timothy Wood 0001, Daniel Hagimont, Noel De Palma, Bernabe Batchakui, Alain Tchana |
EuroSys | 13 |
| 2021 | Mitigating vulnerability windows with hypervisor transplantabstractThe vulnerability window of a hypervisor regarding a given security flaw is the time between the identification of the flaw and the integration of a correction/patch in the running hypervisor. Most vulnerability windows, regardless of severity, are long enough (several days) that attackers have time to perform exploits. Nevertheless, the number of critical vulnerabilities per year is low enough to allow an exceptional solution. This paper introduces hypervisor transplant, a solution for addressing vulnerability window of critical flaws. It involves temporarily replacing the current datacenter hypervisor (e.g., Xen) which is subject to a critical security flaw, by a different hypervisor (e.g., KVM) which is not subject to the same vulnerability. Tu Dinh Ngoc, Boris Teabe, Alain Tchana, Gilles Muller, Daniel Hagimont |
EuroSys | 3 |
| 2021 | Montsalvat: Intel SGX shielding for GraalVM native imagesabstractThe popularity of the Java programming language has led to its wide adoption in cloud computing infrastructures. However, Java applications running in untrusted clouds are vulnerable to various forms of privileged attacks. The emergence of trusted execution environments (TEEs) such as Intel SGX mitigates this problem. TEEs protect code and data in secure enclaves inaccessible to untrusted software, including the kernel and hypervisors. To efficiently use TEEs, developers must manually partition their applications into trusted and untrusted parts, in order to reduce the size of the trusted computing base (TCB) and minimise the risks of security vulnerabilities. However, partitioning applications poses two important challenges: (i) ensuring efficient object communication between the partitioned components, and (ii) ensuring the consistency of garbage collection between the parts, especially with memory-managed languages such as Java. We present Montsalvat, a tool which provides a practical and intuitive annotation-based partitioning approach for Java applications destined for secure enclaves. Montsalvat provides an RMI-like mechanism to ensure inter-object communication, as well as consistent garbage collection across the partitioned components. We implement Montsalvat with GraalVM native-image, a tool for compiling Java applications ahead-of-time into standalone native executables that do not require a JVM at runtime. Our extensive evaluation with micro- and macro-benchmarks shows our partitioning approach to boost performance in real-world applications up to 6.6x (PalDB) and 2.2x (GraphChi) as compared to solutions that naively include the entire applications in the enclave. Peterson Yuhala, Jämes Ménétrey, Pascal Felber, Valerio Schiavoni, Alain Tchana, Gaël Thomas 0001, Hugo Guiroux, Jean-Pierre Lozi |
Middleware | 5 |
| 2021 | Extending Intel PML for hardware-assisted working set size estimation of VMsabstractIntel page modification logging (PML) is a hardware feature introduced in 2015 for tracking modified memory pages of virtual machines (VMs). Although initially designed to improve VMs checkpointing and live migration, we present in this paper how we can take advantage of this virtualization technology to efficiently estimate the working set size (WSS) of a VM. To this end, we first conduct a study of PML with the Xen hypervisor to investigate its performance impact on VMs and the accuracy of a WSS estimation system that relies on the current version of PML. Our three main findings are as follows. (1) PML reduces by up to 10.18% the time of both VM live migration and checkpointing. (2) PML slightly reduces the negative impact of live migration on application performance by up to 0.95%. (3) A WSS estimation system based on the current version of PML provides inaccurate results. Moreover, our experiments show that write-intensive applications are negatively impacted, with up to 34.9% of performance degradation, when using PML to estimate the WSS of a VM that runs these applications. Based on the aforementioned findings, we introduce page reference logging (PRL), an extended version of PML that allows both read and write memory accesses to be tracked without impacting user VMs, thus more suitable for WSS estimation. We propose a WSS estimation system that leverages PRL and show how it can be used in a data center exploiting memory overcommitment. We implement PRL and the underlying WSS estimation system under Gem5, a popular open-source computer architecture simulator. Evaluation results validate the accuracy of the WSS estimation system and show that PRL does not incur more performance degradation on user’s VMs. Stella Bitchebe, Djob Mvondo, Laurent Réveillère, Noel De Palma, Alain Tchana |
VEE | 5 |
| 2021 | (No)Compromis: paging virtualization is not a fatalityabstractNested/Extended Page Table (EPT) is the current hardware solution for virtualizing memory in virtualized systems. It induces a significant performance overhead due to the 2D page walk it requires, thus 24 memory accesses on a TLB miss (instead of 4 memory accesses in a native system). This 2D page walk constraint comes from the utilization of paging for managing virtual machine (VM) memory. This paper shows that paging is not necessary in the hypervisor. Our solution Compromis, a novel Memory Management Unit, uses direct segments for VM memory management combined with paging for VM's processes. This is the first time that a direct segment based solution is shown to be applicable to the entire VM memory while keeping applications unchanged. Relying on the 310 studied datacenter traces, the paper shows that it is possible to provision up to 99.99% of the VMs using a single memory segment. The paper presents a systematic methodology for implementing Compromis in the hardware, the hypervisor and the datacenter scheduler. Evaluation results show that Compromis outperforms the two popular memory virtualization solutions: shadow paging and EPT by up to 30% and 370% respectively. Boris Teabe, Peterson Yuhala, Alain Tchana, Fabien Hermenier, Daniel Hagimont, Gilles Muller |
VEE | 3 |
| 2020 | Fine-Grained Fault Tolerance for Resilient pVM-Based Virtual Machine MonitorsabstractVirtual machine monitors (VMMs) play a crucial role in the software stack of cloud computing platforms: their design and implementation have a major impact on performance, security and fault tolerance. In this paper, we focus on the latter aspect (fault tolerance), which has received less attention, although it is now a significant concern. Our work aims at improving the resilience of the "pVM-based" VMMs, a popular design pattern for virtualization platforms. In such a design, the VMM is split into two main components: a bare-metal hypervisor and a privileged guest virtual machine (pVM). We highlight that the pVM is the least robust component and that the existing fault-tolerance approaches provide limited resilience guarantees or prohibitive overheads. We present three design principles (disaggregation, specialization, and pro-activity), as well as optimized implementation techniques for building a resilient pVM without sacrificing end-user application performance. We validate our contribution on the mainstream Xen platform. Djob Mvondo, Alain Tchana, Renaud Lachaize, Daniel Hagimont, Noel De Palma |
DSN | 2 |
| 2020 | When FTM Discovered MUSIC: Accurate WiFi-based Ranging in the Presence of MultipathabstractThe recent standardization by IEEE of Fine Timing Measurement (FTM), a time-of-flight based approach for ranging has the potential to be a turning point in bridging the gap between the rich literature on indoor localization and the so-far tepid market adoption. However, experiments with the first WiFi cards supporting FTM show that while it offers meter-level ranging in clear line-of-sight settings (LOS), its accuracy can collapse in non-line-of-sight (NLOS) scenarios. We present FUSIC, the first approach that extends FTM's LOS accuracy to NLOS settings, without requiring any changes to the standard. To accomplish this, FUSIC leverages the results from FTM and MUSIC - both erroneous in NLOS - into solving the double challenge of 1) detecting when FTM returns an inaccurate value and 2) correcting the errors as necessary. Experiments in 4 different physical locations reveal that a) FUSIC extends FTM's LOS ranging accuracy to NLOS settings - hence, achieving its stated goal; b) it significantly improves FTM's capability to offer room-level indoor positioning. Kevin Jiokeng, Gentian Jakllari, Alain Tchana, André-Luc Beylot |
INFOCOM | 3 |
| 2020 | Fine-Grained Isolation for Scalable, Dynamic, Multi-tenant Edge Clouds
Yuxin Ren 0001, Guyue Liu, Vlad Nitu, Wenyuan Shao, Riley Kennedy, Gabriel Parmer, Timothy Wood 0001, Alain Tchana |
USENIX ATC | 8 |
| 2020 | Cacol: A zero overhead and non-intrusive double caching mitigation system
Grégoire Todeschi, Boris Teabe, Alain Tchana, Daniel Hagimont |
Future Gener. Comput. Syst. | 3 |
| 2019 | When eXtended Para - Virtualization (XPV) Meets NUMAabstractThis paper addresses the problem of efficiently virtualizing NUMA architectures. The major challenge comes from the fact that the hypervisor regularly reconfigures the placement of a virtual machine (VM) over the NUMA topology. However, neither guest operating systems (OSes) nor system runtime libraries (e.g., Hotspot) are designed to consider NUMA topology changes at runtime, leading end user applications to unpredictable performance. This paper presents eXtended Para-Virtualization (XPV), a new principle to efficiently virtualize a NUMA architecture. XPV consists in revisiting the interface between the hypervisor and the guest OS, and between the guest OS and system runtime libraries (SRL) so that they can dynamically take into account NUMA topology changes. The paper presents a methodology for systematically adapting legacy hypervisors, OSes, and SRLs. We have applied our approach with less than 2k line of codes in two legacy hypervisors (Xen and KVM), two legacy guest OSes (Linux and FreeBSD), and three legacy SRLs (Hotspot, TCMalloc, and jemalloc). The evaluation results showed that XPV outperforms all existing solutions by up to 304%. Vo Quoc Bao Bui, Djob Mvondo, Boris Teabe, Kevin Jiokeng, Patrick Lavoisier Wapet, Alain Tchana, Gaël Thomas 0001, Daniel Hagimont, Gilles Muller, Noel De Palma |
EuroSys | 6 |
| 2019 | Nested Virtualization Without the NestabstractWith the increasing popularity of containers, managing them on top of virtual machines becomes a common practice, called nested virtualization. This paper presents BrFusion and Hostlo, two solutions that address each of two networking issues of nested virtualization: network virtualization duplication and virtual machine-bounded pod deployments. The first issue lengthens network packet paths while the second issue leads to resource fragmentation. For instance, in respect with the first issue, we measured a throughput degradation of about 68% and a latency increase of about 31% in comparison with a single networking layer. We prototype BrFusion and Hostlo in Linux KVM/QEMU, Docker and Kubernetes systems. The evaluation results show that BrFusion leads to the same performance as a single-layer virtualization deployment. Concerning Hostlo, the results show that more than 11% of cloud clients see their cloud utilization cost reduced by down to 40%. Mathieu Bacou, Grégoire Todeschi, Alain Tchana, Daniel Hagimont |
ICPP | 3 |
| 2019 | Memory flipping: a threat to NUMA virtual machines in the CloudabstractvNUMA is the most recent technology used by hypervisors to deal with Non Uniform Memory Access (NUMA) machines, which currently composed most datacenters. vNUMA consists in presenting to the virtual machine (VM) the initial mapping (at boot time) of its virtual resources to physical resources. By this way, all NUMA optimizations implemented by almost all VM’s OS (e.g. Linux) can become effective. However, in order to be effective itself, vNUMA imposes that the initial resource mapping of the VM should remain unchanged during the VM lifetime. Current hypervisors enforce this requirement by avoiding virtual resource migration (between different NUMA nodes, in the same machine), VM migration (between different machines), and memory ballooning.However, we found that memory flipping the most efficient network virtualization approach violates the above requirement. In other words, a VM which performs network operations leads the hypervisor implicitly performs memory page migrations. In this paper, we show that violating this requirement can degrade performance by up to 18%. We present two solutions which mitigate the issue. We prototype these solutions in Xen hypervisor, a popular open source hypervisor, which is widely used by Amazon Web Services. The evaluation results, performed with well known benchmarks, show that our two solutions are able to almost cancel the issue, while keeping memory flipping effective. Djob Mvondo, Boris Teabe, Alain Tchana, Daniel Hagimont, Noel De Palma |
INFOCOM | 3 |
| 2019 | Drowsy-DC: Data Center Power Management SystemabstractIn a modern data center (DC), a large majority of costs arise from energy consumption. The most popular technique used to mitigate this issue is virtualization and more precisely virtual machine (VM) consolidation. Although consolidation may increase server usage by about 5-10%, it is difficult to actually witness server loads greater than 50%. By analyzing the traces from our cloud provider partner, confirmed by previous research work, we have identified that some VMs have sporadic moments of data computation followed by large periods of idleness. These VMs often hinder the consolidation system which cannot further increase the energy efficiency of the DC. In this paper we propose a novel DC power management system called Drowsy-DC, which is able to identify the aforementioned VMs which have matching patterns of idleness. These VMs can thus be colocated on the same server so that their idle periods are exploited to put the server to a low power mode (suspend to RAM) until some data computation is required. While introducing a negligible overhead, our system is able to significantly improve any VM consolidation system; evaluations showed improvements up to 81% and more when compared to OpenStack Neat. Mathieu Bacou, Grégoire Todeschi, Alain Tchana, Daniel Hagimont, Baptiste Lepers, Willy Zwaenepoel |
IPDPS | 3 |
| 2019 | Closer: A New Design Principle for the Privileged Virtual Machine OSabstractIn most of today's virtualized systems (e.g., Xen), the hypervisor relies on a privileged virtual machine (pVM). The pVM accomplishes work both for the hypervisor (e.g., VM life cycle management) and for client VMs (I/O management). Usually, the pVM is based on a standard OS (Linux). This is source of performance unpredictability, low performance, resource waste, and vulnerabilities. This paper presents Closer, a principle for designing a suitable OS for the pVM. Closer consists in respectively scheduling and allocating pVM's tasks and memory as close to the involved client VM as possible. By revisiting Linux and Xen hypervisor, we present a functioning implementation of Closer. The evaluation results of our implementation show that Closer outperforms standard implementations. Djob Mvondo, Boris Teabe, Alain Tchana, Daniel Hagimont, Noel De Palma |
MASCOTS | 3 |
| 2019 | Preventing the propagation of a new kind of illegitimate apps
Patrick Lavoisier Wapet, Alain Tchana, Giang Son Tran, Daniel Hagimont |
Future Gener. Comput. Syst. | 2 |
| 2018 | Welcome to zombieland: practical and energy-efficient memory disaggregation in a datacenterabstractIn this paper, we propose an effortless way for disaggregating the CPU-memory couple, two of the most important resources in cloud computing. Instead of redesigning each resource board, the disaggregation is done at the power supply domain level. In other words, CPU and memory still share the same board, but their power supply domains are separated. Besides this disaggregation, we make the two following contributions: (1) the prototyping of a new ACPI sleep state (called zombie and noted Sz) which allows to suspend a server (thus save energy) while making its memory remotely accessible; and (2) the prototyping of a rack-level system software which allows the transparent utilization of the entire rack resources (avoiding resource waste). We experimentally evaluate the effectiveness of our solution and show that it can improve the energy efficiency of state-of-the-art consolidation techniques by up to 86%, with minimal additional complexity. Vlad Nitu, Boris Teabe, Alain Tchana, Canturk Isci, Daniel Hagimont |
EuroSys | 3 |
| 2017 | Dealing with Performance Unpredictability in an Asymmetric Multicore Processor Cloud
Boris Teabe, Patrick Lavoisier Wapet, Alain Tchana, Daniel Hagimont |
Euro-Par | 3 |
| 2017 | The lock holder and the lock waiter pre-emption problems: nip them in the bud using informed spinlocks (I-Spinlock)abstractIn native Linux systems, spinlock's implementation relies on the assumption that both the lock holder thread and lock waiter threads cannot be preempted. However, in a virtualized environment, these threads are scheduled on top of virtual CPUs (vCPU) that can be preempted by the hypervisor at any time, thus forcing lock waiter threads on other vCPUs to busy wait and to waste CPU cycles. This leads to the well-known Lock Holder Preemption (LHP) and Lock Waiter Preemption (LWP) issues. Boris Teabe, Vlad Nitu, Alain Tchana, Daniel Hagimont |
EuroSys | 3 |
| 2017 | Swift Birth and Quick Death: Enabling Fast Parallel Guest Boot and Destruction in the Xen HypervisorabstractThe ability to quickly set up and tear down a virtual machine is critical for today's cloud elasticity, as well as in numerous other scenarios: guest migration/consolidation, event-driven invocation of micro-services, dynamically adaptive unikernel-based applications, micro-reboots for security or stability, etc. Vlad Nitu, Pierre Olivier, Alain Tchana, Daniel Chiba, Antonio Barbalace, Daniel Hagimont, Binoy Ravindran |
VEE | 3 |
| 2017 | StopGap: elastic VMs to enhance server consolidationabstractSummary Virtualized cloud infrastructures (also known as IaaS platforms) generally rely on a server consolidation system to pack virtual machines (VMs) on as few servers as possible. However, an important limitation of consolidation is not addressed by such systems. Because the managed VMs may be of various sizes (small, medium, large, etc.), VM packing may be obstructed when VMs do not fit available spaces. This phenomenon leaves servers with a set of unused resources (‘holes’). It is similar to memory fragmentation, a well‐known problem in operating system domain. In this paper, we propose a solution which consists in resizing VMs so that they can fit with holes. This operation leads to the management of what we call elastic VMs and requires cooperation between the application level and the IaaS level, because it impacts management at both levels. To this end, we propose a new resource negotiation and allocation model in the IaaS, calledHRNM. We demonstrate HRNM's applicability through the implementation of a prototype compatible with two main IaaS managers (OpenStack and OpenNebula). By performing thorough experiments with SPECvirt_sc2010 (a reference benchmark for server consolidation), we show that the impact of HRNM on customer's application is negligible. Finally, using Google data center traces, we show an improvement of about 62.5% for the traditional consolidation engines. Copyright © 2017 John Wiley & Sons, Ltd. Vlad Nitu, Boris Teabe, Leon Fopa, Alain Tchana, Daniel Hagimont |
Softw. Pract. Exp. | 4 |
| 2016 | Billing system CPU time on individual VMabstractIn virtualized cloud hosting centers, a virtual machine (VM) is generally allocated a fixed computing capacity. The virtualization system schedules the VMs and guarantees that each VM capacity is provided and respected. However, a significant amount of CPU time is consumed by the underlying virtualization system, which generally includes device drivers (mainly network and disk drivers). In today's virtualization systems, this CPU time consumed is difficult to monitor and it is not charged to VMs. Such a situation can have important consequences for both clients and provider: performance isolation and predictability for the former and resource management (and especially consolidation) for the latter. In this paper, we propose a virtualization system mechanism which allows estimating the CPU time used by the virtualization system on behalf of VMs. Subsequently, this CPU time is charged to VMs, thus removing the two previous side effects. This mechanism has been implemented in Xen. Its benefits have been evaluated using reference benchmarks. Boris Teabe, Alain Tchana, Daniel Hagimont |
CCGrid | 2 |
| 2016 | Application-specific quantum for multi-core platform schedulerabstractScheduling has a significant influence on application performance. Deciding on a quantum length can be very tricky, especially when concurrent applications have various characteristics. This is actually the case in virtualized cloud computing environments where virtual machines from different users are colocated on the same physical machine. We claim that in a multi-core virtualized platform, different quantum lengths should be associated with different application types. We apply this principle in a new scheduler called AQL_Sched. We identified 5 main application types and experimentally found the best quantum length for each of them. Dynamically, AQL_Sched associates an application type with each virtual CPU (vCPU) and schedules vCPUs according to their type on physical CPU (pCPU) pools with the best quantum length. Therefore, each vCPU is scheduled on a pCPU with the best quantum length. We implemented a prototype of AQL_Sched in Xen and we evaluated it with various reference benchmarks (SPECweb2009, SPECmail2009, SPEC CPU2006, and PARSEC). The evaluation results show that AQL_Sched outperforms Xen's credit scheduler. For instance, up to 20%, 10% and 15% of performance improvements have been obtained with SPECweb2009, SPEC CPU2006 and PARSEC, respectively. Boris Teabe, Alain Tchana, Daniel Hagimont |
EuroSys | 2 |
| 2016 | Mitigating performance unpredictability in the IaaS using the Kyoto principle
Alain Tchana, Vo Quoc Bao Bui, Boris Teabe, Vlad Nitu, Daniel Hagimont |
Middleware | 1 |
| 2016 | Software consolidation as an efficient energy and cost saving solution
Alain Tchana, Noel De Palma, Ibrahim Safieddine, Daniel Hagimont |
Future Gener. Comput. Syst. | 1 |
| 2015 | Roboconf: A Hybrid Cloud Orchestrator to Deploy Complex ApplicationsabstractThis paper presents Roboconf, an open-source distributed application orchestration framework for multi-cloud platforms, designed to solve challenges of current Autonomic Computing Systems in the era of Cloud computing. It provides a Domain Specific Language (DSL) which allows to describe applications and their execution environments (cloud platforms) in a hierarchical way in order to provide a fine-grained management. Roboconf implements an asynchronous and parallel deployment protocol which accelerates and makes resilient the deployment process. Intensive experiments with different type of applications over different cloud models (e.g. Private, hybrid, and multi-cloud) validate the genericity of Roboconf. These experiments also demonstrate its efficiency comparing to existing frameworks such as Right Scale, Scalr, and Cloudify. Linh Manh Pham, Alain Tchana, Didier Donsez, Noel De Palma, Vincent Zurczak, Pierre-Yves Gibello |
CLOUD | 2 |
| 2015 | VMcSim: A Detailed Manycore Simulator for Virtualized SystemsabstractDesigning a cloud infrastructure for HPC applications requires to correlate design choices for virtualization solutions, resources management strategies, and their implementation on specific hardware platforms. Such investigations can hardly be conducted without accurate simulation tools. This paper introduces VMcSim, the first micro architecture and virtualized many core simulation framework. VMcSim models a x86-based asymmetric many core micro architecture, including a virtualization layer which allows to model a hyper visor, a set of virtual machines with their virtual resources (CPU and memory), their operating system, and their applications. By simulating all the involved hardware and software layers, VMcSim outperforms other simulators. Simulation accuracy is evaluated through several benchmark (SPLASH-2). The simulator is demonstrated with the evaluation of virtual resources placement policies in multicore systems. Alain Tchana, Brice Ekane, Boris Teabe, Daniel Hagimont |
CLOUD | 1 |
| 2015 | Cooperative Resource Management in a IaaSabstractVirtualized IaaS generally rely on a server consolidation system to pack virtual machines (VMs) on as few servers as possible, for energy saving. However, two situations are not taken into account, and could enhance consolidation. First, since the managed VMs can be of various sizes (small, medium, large, etc.), VMs packing can be obstructed when sizes don't fit available spaces on servers. Therefore, we would need to "split" such VMs. Second, two VMs which host replicas of the same application server (for scalability) could be "fusion Ned" when they are located on the same physical server, in order to reduce virtualization overhead and VMs memory footprint. Split and fusion operations lead to the management of elastic VMs and requires cooperation between the application level and the provider level, as they impact management at both levels. In this paper, we propose a IaaS resource management system which implements elastic VMs based on split/fusion operations and cooperative management. We show its benefit with a set of experiments. Giang Son Tran, Alain Tchana, Daniel Hagimont, Noel De Palma |
AINA | 2 |
| 2015 | Software Consolidation as an Efficient Energy and Cost Saving Solution for a SaaS/PaaS Cloud Model
Alain Tchana, Noel De Palma, Ibrahim Safieddine, Daniel Hagimont, Bruno Diot, Nicolas Vuillerme |
Euro-Par | 1 |
| 2015 | Enforcing CPU allocation in a heterogeneous IaaS
Boris Teabe, Alain Tchana, Daniel Hagimont |
Future Gener. Comput. Syst. | 2 |
| 2015 | A self-scalable load injection serviceabstractLoad testing of applications is an important and costly activity for software provider companies. Classical solutions are very difficult to set up statically, and their cost is prohibitive in terms of both human and hardware resources. Virtualized cloud computing platforms provide new opportunities for stressing an application's scalability, by providing a large range of flexible and less expensive (pay-per-use model) computation units. On the basis of these advantages, load testing solutions could be provided on demand in the cloud. This paper describes a Benchmark-as-a-Service solution that automatically scales the load injection platform and facilitates its setup according to load profiles. Our approach is based on: (i) virtualization of the benchmarking platform to create self-scaling injectors; (ii) online calibration to characterize the injector's capacity and impact on the benched application; and (iii) a provisioning solution to appropriately scale the load injection platform ahead of time. We also report experiments on a benchmark illustrating the benefits of this system in terms of cost and resource reductions. Copyright © 2013 John Wiley & Sons, Ltd. Alain Tchana, Noel De Palma, Bruno Dillenseger, Xavier Etchevers |
Softw. Pract. Exp. | 1 |
| 2014 | Elastic Message QueuesabstractToday's systems are often distributed, and connecting their different components can be challenging. Message-Oriented-Middleware (MOM) is a popular tool to insure simple and reliable communication. With the ever growing loads of today's applications, MOMs needs to be scalable. But as the load changes, static scalability often underuses the resources it requires. This paper presents an elastic message queuing system leveraging cloud's on-demand resource provisioning, which allows the use of just enough resources to handle the current load. We will detail when and how provisioning decisions are made, and show the result of our system's evaluation on Amazon EC2 public cloud. This work is based on Joram, an open-source JMS compliant MOM and is now part of its distribution on OW2 consortium's website. Ahmed El-Rheddane, Noel De Palma, Alain Tchana, Daniel Hagimont |
IEEE CLOUD | 3 |
| 2014 | Coordinating self-sizing and self-repair managers for multi-tier systems
Soguy Mak Karé Gueye, Noel De Palma, Éric Rutten, Alain Tchana, Nicolas Berthier |
Future Gener. Comput. Syst. | 4 |
| 2014 | A Self-Scalable and Auto-Regulated Request Injection Benchmarking Tool for Automatic Saturation DetectionabstractSoftware applications providers have always been required to perform load testing prior to launching new applications. This crucial test phase is expensive in human and hardware terms, and the solutions generally used would benefit from further development. In particular, designing an appropriate load profile to stress an application is difficult and must be done carefully to avoid skewed testing. In addition, static testing platforms are exceedingly complex to set up. New opportunities to ease load testing solutions are becoming available thanks to cloud computing. This paper describes a Benchmark-as-a-Service platform based on: (i) intelligent generation of traffic to the benched application without inducing thrashing (avoiding predefined load profiles), (ii) a virtualized and self-scalable load injection system. The platform developed was experimented using two use cases based on the reference JEE benchmark RUBiS. This involved detecting bottleneck tiers, and tuning servers to improve performance. This platform was found to reduce the cost of testing by 50 percent compared to more commonly used solutions. Alain Tchana, Bruno Dillenseger, Noel De Palma, Xavier Etchevers, Jean-Marc Vincent, Nabila Salmi, Ahmed Harbaoui |
IEEE Trans. Cloud Comput. | 1 |
| 2013 | Dynamic Scalability of a Consolidation ServiceabstractIn the coming years, cloud environments will increasingly face energy saving issues. While consolidating the virtual machines running in a cloud is a well-accepted solution to reduce the energy consumption, ensuring the scalability of the consolidation service remains a challenging issue. In this paper, we propose an elastic consolidation service that scales according to the dynamic needs of the cloud environment. Our proposition is based on (i) virtualizing the consolidation manager, (ii) partitioning the consolidation work and (iii) regulating the consolidation scalability through an autonomic control loop. Our proposition has been tested and validated through several experiments. Ahmed El-Rheddane, Noel De Palma, Fabienne Boyer, Frédéric Dumont, Jean-Marc Menaud, Alain Tchana |
IEEE CLOUD | 6 |
| 2013 | A Scalable Benchmark as a Service Platform
Alain Tchana, Noel De Palma, Ahmed El-Rheddane, Bruno Dillenseger, Xavier Etchevers, Ibrahim Safieddine |
DAIS | 1 |
| 2013 | DVFS Aware CPU Credit Enforcement in a Virtualized System
Daniel Hagimont, Christine Mayap Kamga, Laurent Broto, Alain Tchana, Noel De Palma |
Middleware | 4 |
| 2013 | Self-scalable Benchmarking as a Service with Automatic Saturation Detection
Alain Tchana, Bruno Dillenseger, Noel De Palma, Xavier Etchevers, Jean-Marc Vincent, Nabila Salmi, Ahmed Harbaoui |
Middleware | 1 |
| 2013 | Two levels autonomic resource management in virtualized IaaS
Alain Tchana, Giang Son Tran, Laurent Broto, Noel De Palma, Daniel Hagimont |
Future Gener. Comput. Syst. | 1 |
| 2008 | Metamodeling Autonomic System Management Policies - Ongoing WorksabstractAutonomic computing is recognized as one of the most promising solution to address the increasingly complex task of distributed environments' administration. In this context, many projects relied on software components and architectures to organize such an autonomic management software. However, we observed that the interfaces of a component model are too low-level, difficult to use and still error prone. Therefore, we introduced higher-level languages for the modeling of deployment and management policies. These domain specific languages enhance simplicity and consistency of the policies. Our current work is to formally describe the metamodels and the semantics associated with these languages. Benoît Combemale, Laurent Broto, Alain Tchana, Daniel Hagimont |
COMPSAC | 3 |