EDBT 2026 Demo / reviewers in the wild / expert
Stéphane Vialle
dblp:39/4045
· DBLP profile ↗
20ranked-venue papers
2as first author
5since 2021 · last 2025
0000-0001-6336-2269ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 4 since 2021Software engineering, systems software and programming languages · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Artificial intelligence and machine learning · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Data clustering on hybrid classical-quantum NISQ architecture with generative-based variational and parallel algorithms
Julien Rauch, Damien Rontani, Stéphane Vialle |
J. Syst. Archit. | 3 |
| 2024 | Pipeline for Semantic Segmentation of Large Railway Point Clouds
Hugo Gabrielidis, Filippo Gatti, Stéphane Vialle |
IDEAL (1) | 3 |
| 2022 | Parallel and accurate k-means algorithm on CPU-GPU architectures for spectral clusteringabstractSummary k ‐Means is a standard algorithm for clustering data. It constitutes generally the final step in a more complex chain of high‐quality spectral clustering. However, this chain suffers from lack of scalability when addressing large datasets. This can be overcome by applying also the k ‐means algorithm as a preprocessing task to reduce the input data instances. We propose parallel optimization techniques for the k ‐means algorithm on CPU and GPU. Particularly we use a two‐step summation method with package processing to handle the effect of rounding errors that may occur during the phase of updating cluster centroids. Our experiments on synthetic and real‐world datasets containing millions of instances exhibit a speedup up to 7 for the k ‐means iteration time on GPU versus 20/40 CPU threads using AVX units, and achieve double‐precision accuracy with single‐precision computations. Guanlin He, Stéphane Vialle, Marc Baboulin |
Concurr. Comput. Pract. Exp. | 2 |
| 2021 | Scalable Algorithms Using Sparse Storage for Parallel Spectral Clustering on GPU
Guanlin He, Stéphane Vialle, Nicolas Sylvestre, Marc Baboulin |
NPC | 2 |
| 2021 | Synthesis and feedback on the distribution and parallelization of FMI-CS-based co-simulations with the DACCOSIM platformabstractA co-simulation applied to Smart Grids consists of grouping in the same setting models of physical components (among other electrical ones) and models of control units (including communication devices). Combining these models needs to use a generic and robust co-simulation environment instead of developing a specific one. In this context, we developed the DACCOSIM 2017 co-simulation platform based on FMI-CS (Functional Mock-up Interface for CoSimulation) standard to simulate the physical components of a Smart Grid. These components represent the most CPU-consuming part of the co-simulation. However, the tasks of FMI-CS-based applications (FMUs) are exposed as heterogeneous gray boxes with no information concerning their computation and communication volumes. Moreover, all these FMUs frequently communicate with each other by sending a lot of small messages. Consequently, the deployment of an FMI-CS based co-simulation on a distributed architecture is a complex task carried out by DACCOSIM 2017. This paper introduces the development of DACCOSIM-2017, and its experiment on distributed architectures. Cherifa Dad, Jean-Philippe Tavella, Stéphane Vialle |
Parallel Comput. | 3 |
| 2016 | Toward an accurate and fast hybrid multi-simulation with the FMI-CS standardabstractMulti-simulation in the context of future smart electrical grids consists in associating components modeling different physical domains, but also their local or global control. Our DACCOSIM multi-simulation environment is based on the version 2.0 of the FMI-CS (Functional Mock-up Interface for Co-Simulation) standard maintained by the Modelica Association. It has been specifically designed to run large-scale and complex systems on a single PC or a cluster of multicore nodes. But it is quite challenging to accurately simulate FMUs-composed systems involving predictable and unpredictable events while preserving the system overall performance. This paper presents some additions to the FMI-CS standard aiming to improve the accuracy and the performance of distributed multi-simulations involving a mix of both time steps and various kinds of events. The proposed FMI-CS primitives are explained, as well as the Master Algorithm strategies to exploit them efficiently. Jean-Philippe Tavella, Mathieu Caujolle, Stéphane Vialle, Cherifa Dad, Charles Tan, Gilles Plessis, Mathieu Schumann, Arnaud Cuccuru, Sebastien Revol |
ETFA | 3 |
| 2016 | Scaling of Distributed Multi-simulations on Multi-core ClustersabstractDACCOSIM is a multi-simulation environment for continuous time systems, relying on FMI standard, making easy the design of a multi-simulation graph, and specially developed for multi-core PC clusters, in order to achieve speedup and size up. However, the distribution of the simulation graph remains complex and is still the responsibility of the simulation developer. This paper introduces DACCOSIM parallel and distributed architecture, and our strategies to achieve efficient multi-simulation graph distribution on multi-core clusters. Some performance experiments on two clusters, running up to 81 simulation components (FMU) and using up to 16 multi-core computing nodes, are shown. Performances measured on our faster cluster exhibit a good scalability, but some limitations of current DACCOSIM implementation are discussed. Cherifa Dad, Stéphane Vialle, Mathieu Caujolle, Jean-Philippe Tavella, Michel Ianotto |
WETICE | 2 |
| 2014 | Pricing derivatives on graphics processing units using Monte Carlo simulationabstractSUMMARY This paper is about using the existing Monte Carlo approach for pricing European and American contracts on a state‐of‐the‐art graphics processing unit (GPU) architecture. First, we adapt on a cluster of GPUs two different suitable paradigms of parallelizing random number generators, which were developed for CPU clusters. Because in financial applications, we request results within seconds of simulation, the sufficiently large computations should be implemented on a cluster of machines. Thus, we make the European contract comparison between CPUs and GPUs using from one up to 16 nodes of a CPU/GPU cluster. We show that using GPUs for European contracts reduces the execution time by ∼ 40 and diminishes the energy consumed by ∼ 50 during the simulation. In the second set of experiments, we investigate the benefits of using GPUs’ parallelization for pricing American options that require solving an optimal stopping problem and which we implement using the Longstaff and Schwartz regression method. The speedup result obtained for American options varies between two and 10 according to the number of generated paths, the dimensions, and the time discretization. Copyright © 2012 John Wiley & Sons, Ltd. Lokman A. Abbas-Turki, Stéphane Vialle, Bernard Lapeyre, Patrick P. Mercier |
Concurr. Comput. Pract. Exp. | 2 |
| 2012 | FT-GReLoSSS: A Skeletal-Based Approach towards Application Parallelization and Low-Overhead Fault ToleranceabstractFT-GReLoSSS (FTG) is a C++/MPI framework to ease the development of fault-tolerant parallel applications belonging to a SPMD family termed GReLoSSS. The originality of FTG is to rely on the MoLOToF programming model principles to facilitate the addition of an efficient checkpoint-based fault tolerance at the application level. Main features of MoLOToF encompass a structured application development based on fault-tolerant "skeletons" and lay emphasis on collaborations. The latter exist between the programmer, the framework and the underlying runtime middleware/environment. Together with the structured approach they contribute into achieving reduced checkpoint sizes, as well as reduced checkpoint and recovery overhead at runtime. This paper introduces the main principles of MoLOToF and the design of the FTG framework. To properly assess the framework's ease of use for a programmer as well as fault tolerance efficiency, a series of benchmarks were conducted up to 128 nodes on a multicore PC cluster. These benchmarks involved an existing parallel financial application for gas storage valuation, originally developed in collaboration with EDF company, and a rewritten version which made use of the FTG framework and its features. Experiments results display low-overhead compared to existing system-level counterparts. Constantinos Makassikis, Stéphane Vialle, Xavier Warin |
PDP | 2 |
| 2011 | A Javaspace-Based Framework for Efficient Fault-Tolerant Master-Worker Distributed ApplicationsabstractWe propose a framework built around a Java Space to ease the development of bag-of-tasks applications. The framework may optionally and automatically tolerate transient crash failures occurring on any of the distributed elements. It relies on check pointing and underlying middleware mechanisms to do so. To further improve check pointing efficiency, both in size and frequency, the programmer can introduce intermediate user-defined checkpoint data and code within the task processing program. The framework used without fault tolerance accelerates application development, does not introduce runtime overhead and yields to expected speedup. When enabling fault tolerance, our framework allows, despite failures, correct completion of applications with limited runtime and data storage overheads. Experiments run with up to 128 workers study the impact of some user-related and implementation-related on overall performance, and reveal good performances for classical Java Space-based master-worker application profiles. Virginie Galtier, Constantinos Makassikis, Stéphane Vialle |
PDP | 3 |
| 2010 | A Skeletal-Based Approach for the Development of Fault-Tolerant SPMD ApplicationsabstractDistributing applications over PC clusters to speed-up or size-up the execution is now commonplace. Yet efficiently tolerating faults of these systems is a major issue. To ease the addition of checkpoint-based fault tolerance at the application level, we introduce a {\em Model for Low-Overhead Tolerance of Faults}(MoLOToF) which is based on structuring applications using {\em fault-tolerant skeletons}. MoLOToF also encourages collaborations with the programmer and the execution environment. The skeletons are adapted to specific parallelization paradigms and yield what can be called {\em fault-tolerant algorithmic skeletons}. The application of MoLOToF to the SPMD parallelization paradigm results in our proposed FT-SPMD framework. Experiments show that the complexity for developing an application is small and the use of the framework has a small impact on performance. Comparisons with existing system-level checkpoint solutions, namely LAM/MPI and DMTCP, point out that FT-SPMD has a lower runtime overhead while being more robust when a higher level of fault tolerance is required. Constantinos Makassikis, Virginie Galtier, Stéphane Vialle |
PDCAT | 3 |
| 2009 | Implementation of the AdaBoost Algorithm for Large Scale Distributed Environments: Comparing JavaSpace and MPJabstractThis paper presents the parallelization of a machine learning method, called the AdaBoost algorithm. The parallel algorithm follows a dynamically load-balanced master-worker strategy, which is parameterized by the granularity of the tasks distributed to workers. We first show the benefits of this version with heterogeneous processors. Then, we study the application in a real, geographically distributed environment, hence adding network latencies to the execution. Performances of the application using more than a hundred processes are analyzed in both JavaSpace and P2P-MPI. We therefore present an head-to-head comparison of two parallel programming models. We study for each case the granularities yielding the best performance. We show that current network technologies enable to obtain interesting speedups in many situations for such an application, even when using a virtual shared memory paradigm in a large-scale distributed environment. Virginie Galtier, Stéphane Genaud, Stéphane Vialle |
ICPADS | 3 |
| 2009 | High dimensional pricing of exotic European contracts on a GPU Cluster, and comparison to a CPU clusterabstractThe aim of this paper is the efficient use of CPU and GPU clusters for a general path-dependent exotic European pricing, and their comparison in terms of speed and energy consumption. To reach our goal, we propose a parallel random number generator which is well suited to the parallelization paradigm, then, we implement a multidimensional Asian contract as a benchmark using g++/OpenMP/OpenMPI on CPUs and CUDA-nvcc/OpenMPI on GPUs. Finally, we give the detailed results of the two architectures for different size problems using 1-16 GPUs and 1-256 dual-core CPUs. Lokman A. Abbas-Turki, Stéphane Vialle, Bernard Lapeyre, Patrick P. Mercier |
IPDPS | 2 |
| 2009 | Large scale experiment and optimization of a distributed stochastic control algorithm. Application to energy management problemsabstractAsset management for the electricity industry leads to very large stochastic optimization problem. We explain in this article how to efficiently distribute the Bellman algorithm used, re-distributing data and computations at each time step, and we examine the parallelization of a simulation algorithm usually used after this optimization part. We focus on distributed architectures with shared memory multi-core nodes, and we design a multiparadigm parallel algorithm, implemented with both MPI and multithreading mechanisms. Then we lay emphasis on the serial optimizations carried out to achieve high performances both on a dual-core PC cluster and a Blue Gene/P IBM supercomputer with quad-core nodes. Finally, we introduce experimental results achieved on two large testbeds, running a 7-stocks and 10-state-variables benchmark, and we show the impact of multithreading and serial optimizations on our distributed application. Pascal Vezolle, Stéphane Vialle, Xavier Warin |
IPDPS | 2 |
| 2008 | Large scale distribution of stochastic control algorithms for gas storage valuationabstractThis paper introduces the distribution of a stochastic control algorithm which is applied to gas storage valuation, and presents its experimental performances on two PC clusters and an IBM Blue Gene/L supercomputer. This research is part of a French national project which gathers people from the academic world (computer scientists, mathematicians, ...) as well as people from the industry of energy and finance in order to provide concrete answers on the use of computational clusters, grids and supercomputers applied to problems of financial mathematics. The designed distribution allows to run gas storage valuation models which require considerable amounts of computational power and memory space while achieving both speedup and size-up: it has been successfully implemented and experimented on PC clusters (up to 144 processors) and on a Blue Gene supercomputer (up to 1024 processors). Finally, our distributed algorithm allows to use more computing resources in order to maintain constant the execution time while increasing the calculation accuracy. Constantinos Makassikis, Stéphane Vialle, Xavier Warin |
IPDPS | 2 |
| 2006 | A Fault Tolerant and Multi-Paradigm Grid Architecture for Time Constrained Problems. Application to Option Pricing in FinanceabstractThis paper introduces a Grid software architecture offering fault tolerance, dynamic and aggressive load balancing and two complementary parallel programming paradigms. Experiments with financial applications on a real multi-site Grid assess this solution. This architecture has been designed to run industrial and financial applications, that are frequently time constrained and CPU consuming, feature both tightly and loosely coupled parallelism requiring generic programming paradigm, and adopt client-server business architecture. Sebastien Bezzine, Virginie Galtier, Stéphane Vialle, Françoise Baude, Mireille Bossy, Viet Dung Doan, Ludovic Henrio |
e-Science | 3 |
| 2002 | A Design of Multi-Startegy Parallelization for an Entire Application of Document Categorization on Low-Cost Multiprocessor PCs
Stéphane Vialle, Guillaume Schaeffer, Michel Ianotto |
OPODIS | 1 |
| 1999 | A Library to Implement Neural Networks on MIMD Machines
Yann Boniface, Frédéric Alexandre, Stéphane Vialle |
Euro-Par | 3 |
| 1999 | A bridge between two paradigms for parallelism: neural networks and general purpose MIMD computersabstractHardware developments have led to the use of shared memory as an efficient parallel programming method. The main goals of the work reported here are to speed up executions and to decrease development time of parallel neural network implementations. To allow for such implementations, a library has been defined, as a bridge between neural networks and general purpose MIMD computer parallelisms. Yann Boniface, Frédéric Alexandre, Stéphane Vialle |
IJCNN | 3 |
| 1998 | Design and Implementation of a Parallel Cellular Language for MIMD Architectures
Stéphane Vialle, Yannick Lallement, Thierry Cornu |
Comput. Lang. | 1 |