EDBT 2026 Demo / reviewers in the wild / expert
Camille Coti
dblp:78/4708
· DBLP profile ↗
28ranked-venue papers
11as first author
12since 2021 · last 2026
0000-0002-1224-7786ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 17 · 7 first-author · 9 since 2021Software engineering, systems software and programming languages · 6 · 3 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-authorArtificial intelligence and machine learning · 2 · 1 first-authorTheory of computation · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Larger Cloud Servers, Fewer Hosts? on the Evolution of VM Sizes in IaaS Platforms
Pierre Jacquet, Camille Coti, Marcos Dias de Assunção |
CCGrid | 2 |
| 2026 | Untangling GPU Power Consumption: Job-Level Inference in Cloud Shared SettingsabstractAs the demand for AI-driven workloads increases, the energy consumption of Graphics Processing Units (GPUs) devices has come under intense scrutiny, particularly in hyperscale data centers where large numbers of accelerators are centralized and leased to diverse clients. Pierre Jacquet, Maxime Agusti, Eddy Caron, Camille Coti, Marcos Dias de Assunção, Laurent Lefèvre, Anne-Cécile Orgerie |
EuroSys | 4 |
| 2026 | Cinergy: Deterministic Power Monitoring for Carbon Accounting in the CloudabstractInternational audience Pierre Jacquet, Camille Coti, Marcos Dias de Assunção, Romain Rouvoy |
IEEE Trans. Cloud Comput. | 2 |
| 2025 | CINERGY: Reasoning Over the Worst Case Power Consumption of Cloud Virtual MachinesabstractEnergy consumption has become a critical concern in Information and Communication Technologies (ICT), pressing for more accurate measurements. While the power consumption of physical servers can be physically monitored, organizations are increasingly adopting virtual environments, such as cloud computing, rendering physical measurements impractical in operational contexts. The state-of-the-art approaches to estimating this ”virtual” consumption mostly consist of assigning server power consumption shares among hosted processes, guided by various system metrics. Unfortunately, such a bottom-up approach is highly sensitive in a multi-tenant environment, thus failing to report stable measurements to stakeholders. For example, the same activity performed by one Virtual Machine (VM) may lead to different power consumption traces, depending on the activity of the co-hosted VMs. As cloud customers have only control over their provisioned virtual resources, we propose a new method to model the power consumption of their virtual appliances, enabling contextagnostic tracking of their environmental impact. This framework, called CINERGY, is designed to be more predictable than the state-of-the-art power models, while still exposing the gains from consolidation. We evaluate its accuracy against ground-truth measurements, often lacking in the literature. We show that CINERGY is deterministic and accurate, with an average error of 6.6%. Pierre Jacquet, Camille Coti, Marcos Dias de Assunção, Romain Rouvoy |
CCGrid | 2 |
| 2025 | RVS-CUDA: A Real-Time Asynchronous Pipeline for Immersive View Synthesis
Enzo Di Maria, Hossein Pejman, Carlos Vázquez 0001, Stéephane Coulombe, Camille Coti |
PCS | 5 |
| 2024 | MQTT2EdgePeer: a Robust and Scalable Brokerless Peer-to-Peer Edge Middleware for Topic-Based Publish/SubscribeabstractThe topic-based publish-subscribe paradigm plays an important role among many applications, as it enables a seamless interconnection of heterogeneous client applications and devices, through easy-to-use high-level abstractions. While publish-subscribe systems are usually provided in a centralized (i.e., broker-based) model, a peer-to-peer (P2P) topology can bring several benefits, such as increased resiliency and scalability, better load distribution, and a better handling of hot topics that comprise a large amount of publishers and subscribers. The presence of hot topics can result in high outgoing bandwidth usage, which is an important consideration in edge-based deployments.This paper presents MQTT2EdgePeer, a robust and scalable P2P brokerless topic-based publish-subscribe middleware that provides compatibility with MQTT-based applications. MQTT2EdgePeer, which is built on top of a structured edge P2P overlay, provides two message delivery approaches that enable a trade-off between efficient bandwidth distribution among the nodes, and latency minimization. Our findings reveal that MQTT2EdgePeer provides improved scalability, better load distribution, and reduced latencies, in comparison with a state-of-the-art Distributed Single Root per Topic (DSRT) approach. Saeed Rahmani, Amir Ali Pour, Camille Coti, Julien Gascon-Samson |
CCGrid | 3 |
| 2022 | MARTINI: The Little Match and Replace Tool for Automatic Application Rewriting with Code Examples
Alister Johnson, Camille Coti, Allen D. Malony, Johannes Doerfert |
Euro-Par | 2 |
| 2022 | A Formal Model for Fault Tolerant Parallel Matrix FactorizationabstractAs exascale platforms are in sight, high-performance computing needs to take failures into account and provide fault-tolerant applications and environments. Checkpoint-restart approaches do not require modifying the application, but are expensive at large scale. Application-based fault tolerance is more specific to the application and is expected to achieve better performance. In this paper, we address fault-tolerant matrix factorization with algorithms that present good performance, both during failure-free executions and when failures happen. A challenge when designing fault-tolerant algorithms is to make sure they are resilient to any failure scenario. Therefore, we design a model for these algorithms and prove they can tolerate failures at any moment, as long as enough processes are still alive. Camille Coti, Laure Petrucci, Daniel Alberto Torres González |
ICECCS | 1 |
| 2021 | Fault-Tolerant LU Factorization Is Low Cost
Camille Coti, Laure Petrucci, Daniel Alberto Torres González |
Euro-Par | 1 |
| 2021 | DiPOSH: A portable OpenSHMEM implementation for short API-to-network pathabstractSummary In this article, we introduce DiPOSH, a multi‐network, distributed implementation of the OpenSHMEM standard. The core idea behind DiPOSH is to have an API‐to‐network software stack as slim as possible, in order to minimize the software overhead. Following the heritage of its non‐distributed parent POSH, DiPOSH's communication engine is organized around the processes' shared heaps, and remote communications are moving data from and to these shared heaps directly. This article presents its architecture and several communication drivers, including one that takes advantage of a helper process, called the Hub, for inter‐process communications. This architecture allows use to explore different options for implementing the communication drivers, from using high‐level, portable, optimized libraries to low‐level, close to the hardware communication routines. We present the perspectives opened by this additional component in terms of communication scheduling between and on the nodes. DiPOSH is available at https://github.com/coti/DiPOSH . Camille Coti, Allen D. Malony |
Concurr. Comput. Pract. Exp. | 1 |
| 2021 | A task-based approach to parallel parametric linear programming solving, and application to polyhedral computationsabstractSummary Parametric linear programming is a central operation for polyhedral computations, as well as in certain control applications. Here, we propose a task‐based scheme for parallelizing it, with quasi‐linear speedup over large problems. This type of parallel applications is challenging, because several tasks might be computing the same region. In this article, we are presenting the algorithm itself with a parallel redundancy elimination algorithm, and conducting a thorough performance analysis. Camille Coti, David Monniaux, Hang Yu 0005 |
Concurr. Comput. Pract. Exp. | 1 |
| 2021 | Quasi-optimal partial order reduction
Camille Coti, Laure Petrucci, César Rodríguez, Marcelo Sousa |
Formal Methods Syst. Des. | 1 |
| 2020 | On-the-fly Optimization of Parallel Computation of Symbolic Symplectic InvariantsabstractGroup invariants are used in high energy physics to define quantum field theory interactions. In this paper, we present the parallel algebraic computation of special invariants called symplectic and focus on one particular invariant that finds recent interest in physics. Our results will export to other invariants. The cost of performing basic computations on the multivariate polynomials evolves during the computation, as the polynomials get larger and/or have increasing numbers of terms. Interestingly, in some cases, they stay small. Traditionally, high-performance software is optimized by running experiments with sample data sets in order to profile and optimize expected behavior of workloads in practice. Since the (communication and computation) costs depend on the changing behavior of the symplectic invariant calculations, the standard optimization approach is insufficient. Thus, it is necessary to implement online performance tuning methods that can track the algorithm's progress and state, evaluate performance data in situ, and control the parallel resources during execution. Joseph Ben Geloun, Camille Coti, Allen D. Malony |
ISPDC | 2 |
| 2018 | Quasi-Optimal Partial Order ReductionabstractA dynamic partial order reduction (DPOR) algorithm is optimal when it always explores at most one representative per Mazurkiewicz trace. Existing literature suggests that the reduction obtained by the non-optimal, state-of-the-art Source-DPOR (SDPOR) algorithm is comparable to optimal DPOR. We show the first program with $$\mathop {\mathcal {O}} (n)$$ Mazurkiewicz traces where SDPOR explores $$\mathop {\mathcal {O}} (2^n)$$ redundant schedules (as this paper was under review, we were made aware of the recent publication of another paper [3] which contains an independently-discovered example program with the same characteristics). We furthermore identify the cause of this blow-up as an NP-hard problem. Our main contribution is a new approach, called Quasi-Optimal POR, that can arbitrarily approximate an optimal exploration using a provided constant k. We present an implementation of our method in a new tool called Dpu using specialised data structures. Experiments with Dpu, including Debian packages, show that optimality is achieved with low values of k, outperforming state-of-the-art tools. Huyen T. T. Nguyen, César Rodríguez, Marcelo Sousa, Camille Coti, Laure Petrucci |
CAV (2) | 4 |
| 2018 | One-Sided Communications for More Efficient Parallel State Space Exploration over RDMA Clusters
Camille Coti, Sami Evangelista, Laure Petrucci |
Euro-Par | 1 |
| 2018 | State Compression Based on One-Sided Communications for Distributed Model CheckingabstractWe propose a distributed implementation of the collapse compression technique used by explicit state model checkers to reduce memory usage. This adapatation makes use of lock-free distributed hash tables based on one-sided communication primitives provided by libraries such as OpenSHMEM. We implemented this technique in the distributed version of the model checker Helena. We report on experiments performed on the Grid'5000 cluster with an implementation over OpenMPI. These reveal that, for some models, this distributed implementation can altogether preserve the memory reduction provided by collapse compression and reduce execution times by allowing the exchanges of compressed states between processes. Camille Coti, Sami Evangelista, Laure Petrucci |
ICECCS | 1 |
| 2017 | Solving 0-1 Quadratic Problems with Two-Level Parallelization of the BiqCrunch SolverabstractIn this paper we present MLTBiqCrunch, a hierarchically parallelized version of the open-source solver BiqCrunch [1].More precisely, this version has two levels of parallelization: a coarse grain, assigning a thread to a node evaluation and a fine grain, parallelizing a node evaluation when some threads are not busy.We present experiments on some classical binary quadratic optimization problems with comparison of their scalability and raw performance.In particular, we obtain a superlinear speedup for some of the most difficult instances. Camille Coti, Étienne Leclercq, Frédéric Roupin, Franck Butelle |
FedCSIS | 1 |
| 2015 | Enhanced Distributed Behavioral Cartography of Parametric Timed Automata
Étienne André 0001, Camille Coti, Hoang Gia Nguyen |
ICFEM | 2 |
| 2012 | Fault Tolerance Logical Network Properties of Irregular Graphs
Christophe Cérin, Camille Coti, Michel Koskas |
ICA3PP (1) | 2 |
| 2011 | QCG-OMPI: MPI applications on grids
Emmanuel Agullo, Camille Coti, Thomas Hérault, Julien Langou, Sylvain Peyronnet, Ala Rezmerita, Franck Cappello, Jack J. Dongarra |
Future Gener. Comput. Syst. | 2 |
| 2010 | QR factorization of tall and skinny matrices in a grid computing environmentabstractPrevious studies have reported that common dense linear algebra operations do not achieve speed up by using multiple geographical sites of a computational grid. Because such operations are the building blocks of most scientific applications, conventional supercomputers are still strongly predominant in high-performance computing and the use of grids for speeding up large-scale scientific problems is limited to applications exhibiting parallelism at a higher level. We have identified two performance bottlenecks in the distributed memory algorithms implemented in ScaLAPACK, a state-of-the-art dense linear algebra library. First, because ScaLA-PACK assumes a homogeneous communication network, the implementations of ScaLAPACK algorithms lack locality in their communication pattern. Second, the number of messages sent in the ScaLAPACK algorithms is significantly greater than other algorithms that trade flops for communication. In this paper, we present a new approach for computing a QR factorization – one of the main dense linear algebra kernels – of tall and skinny matrices in a grid computing environment that overcomes these two bottlenecks. Our contribution is to articulate a recently proposed algorithm (Communication Avoiding QR) with a topology-aware middleware (QCG-OMPI) in order to confine intensive communications (ScaLAPACK calls) within the different geographical sites. An experimental study conducted on the Grid'5000 platform shows that the resulting performance increases linearly with the number of geographical sites on large-scale problems (and is in particular consistently higher than ScaLAPACK's). Emmanuel Agullo, Camille Coti, Jack J. Dongarra, Thomas Hérault, Julien Langou |
IPDPS | 2 |
| 2010 | PAR: a PARallel and distributed job crusherabstractUNLABELLED: Bioinformaticians are tackling increasingly computation-intensive tasks. In the meantime, workstations are shifting towards multi-core architectures and even massively multi-core may be the norm soon. Bag-of-Tasks (BoT) applications are commonly encountered in bioinformatics. They consist of a large number of independent computation-intensive tasks. This note introduces PAR, a scalable, dynamic, parallel and distributed execution engine for Bag-of-Tasks. PAR is aimed at multi-core architectures and small clusters. Accelerations obtained thanks to PAR on two different applications are shown. AVAILABILITY: PAR is released under the GNU General Public License version three and can be freely downloaded (http://download.savannah.gnu.org/releases/par/par.tgz). Francois Berenger, Camille Coti, Kam Y. J. Zhang |
Bioinform. | 2 |
| 2009 | Running Parallel Applications with Topology-Aware Grid MiddlewareabstractThe concept of topology-aware grid applications is derived from parallelized computational models of complex systems that are executed on heterogeneous resources, either because they require specialized hardware for certain calculations, or because their parallelization is flexible enough to exploit such resources. Here we describe two such applications, a multi-body simulation of stellar evolution, and an evolutionary algorithm that is used for reverse-engineering gene regulatory networks. We then describe the topology-aware middleware we have developed to facilitate the "modeling-implementing-executing" cycle of complex systems applications. The developed middleware allows topology-aware simulations to run on geographically distributed clusters with or without firewalls between them. Additionally, we describe advanced coallocation and scheduling techniques that take into account the applications topologies. Results are given based on running the topology-aware applications on the Grid'5000 infrastructure. Pavel Bar, Camille Coti, Derek Groen, Thomas Hérault, Valentin Kravtsov, Assaf Schuster, Martin T. Swain |
eScience | 2 |
| 2009 | MPI Applications on Grids: A Topology Aware Approach
Camille Coti, Thomas Hérault, Franck Cappello |
Euro-Par | 1 |
| 2009 | Kernels and learning curves for Gaussian process regression on random graphsabstractWe investigate how well Gaussian process regression can learn functions defined on graphs, using large regular random graphs as a paradigmatic example. Random-walk based kernels are shown to have some surprising properties: within the standard approximation of a locally tree-like graph structure, the kernel does not become constant, i.e.neighbouring function values do not become fully correlated, when the lengthscale $\sigma$ of the kernel is made large. Instead the kernel attains a non-trivial limiting form, which we calculate. The fully correlated limit is reached only once loops become relevant, and we estimate where the crossover to this regime occurs. Our main subject are learning curves of Bayes error versus training set size. We show that these are qualitatively well predicted by a simple approximation using only the spectrum of a large tree as input, and generically scale with $n/V$, the number of training examples per vertex. We also explore how this behaviour changes once kernel lengthscales are large enough for loops to become important. Peter Sollich, Matthew Urry, Camille Coti |
NIPS | 3 |
| 2008 | Grid Services for MPIabstractInstitutional grids consist of the aggregation of clusters belonging to different administrative domains to build a single parallel machine. To run an MPI application over an institutional grid, one has to address many challenges. One of the first problems to solve is the connectivity of the different nodes not belonging to the same administrative domain. Techniques based on communication relays, dynamic port opening, among others, have been proposed. In this work, we propose a set of Grid or Web Services to abstract this connectivity service, and we evaluate the performances of this new level of communication for establishing the connectivity of an MPI application over an experimental grid. Camille Coti, Thomas Hérault, Sylvain Peyronnet, Ala Rezmerita, Franck Cappello |
CCGRID | 1 |
| 2008 | Blocking vs. non-blocking coordinated checkpointing for large-scale fault tolerant MPI Protocols
Darius Buntinas, Camille Coti, Thomas Hérault, Pierre Lemarinier, Laurence Pilard, Ala Rezmerita, Eric Rodriguez, Franck Cappello |
Future Gener. Comput. Syst. | 2 |
| 2006 | MPI tools and performance studies - Blocking vs. non-blocking coordinated checkpointing for large-scale fault tolerant MPIabstractA long-term trend in high-performance computing is the increasing number of nodes in parallel computing platforms, which entails a higher failure probability. Fault tolerant programming environments should be used to guarantee the safe execution of critical applications. Research in fault tolerant MPI has led to the development of several fault tolerant MPI environments. Different approaches are being proposed using a variety of fault tolerant message passing protocols based on coordinated checkpointing or message logging. The most popular approach is with coordinated checkpointing. In the literature, two different concepts of coordinated checkpointing have been proposed: blocking and nonblocking. However they have never been compared quantitatively and their respective scalability remains unknown. The contribution of this paper is to provide the first comparison between these two approaches and a study of their scalability. We have implemented the two approaches within the MPICH environments and evaluate their performance using the NAS parallel benchmarks. Camille Coti, Thomas Hérault, Pierre Lemarinier, Laurence Pilard, Ala Rezmerita, Eric Rodriguez, Franck Cappello |
SC | 1 |