EDBT 2026 Demo / reviewers in the wild / expert
Bruno Raffin
dblp:74/2662
· DBLP profile ↗
35ranked-venue papers
4as first author
5since 2021 · last 2024
0000-0002-7980-4946ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 19 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-authorHuman-computer interaction and ubiquitous computing · 6 · 2 first-authorArtificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Extreme-scale workflows: A perspective from the JLESC international community
Orcun Yildiz, Amal Gueroudji, Julien Bigot, Bruno Raffin, Rosa M. Badia, Tom Peterka |
Future Gener. Comput. Syst. | 4 |
| 2023 | Training Deep Surrogate Models with Large Scale Online LearningabstractThe spatiotemporal resolution of Partial Differential Equations (PDEs) plays important roles in the mathematical description of the world’s physical phenomena. In general, scientists and engineers solve PDEs numerically by the use of computationally demanding solvers. Recently, deep learning algorithms have emerged as a viable alternative for obtaining fast solutions for PDEs. Models are usually trained on synthetic data generated by solvers, stored on disk and read back for training. This paper advocates that relying on a traditional static dataset to train these models does not allow the full benefit of the solver to be used as a data generator. It proposes an open source online training framework for deep surrogate models. The framework implements several levels of parallelism focused on simultaneously generating numerical simulations and training deep neural networks. This approach suppresses the I/O and storage bottleneck associated with disk-loaded datasets, and opens the way to training on significantly larger datasets. Experiments compare the offline and online training of four surrogate models, including state-of-the-art architectures. Results indicate that exposing deep surrogate models to more dataset diversity, up to hundreds of GB, can increase model generalization capabilities. Fully connected neural networks, Fourier Neural Operator (FNO), and Message Passing PDE Solver prediction accuracy is improved by 68%, 16% and 7%, respectively. Lucas Thibaut Meyer, Marc Schouler, Robert Alexander Caulk, Alejandro Ribés, Bruno Raffin |
ICML | 5 |
| 2023 | High Throughput Training of Deep Surrogates from Large Ensemble RunsabstractRecent years have seen a surge in deep learning approaches to accelerate numerical solvers, which provide faithful but computationally intensive simulations of the physical world. These deep surrogates are generally trained in a supervised manner from limited amounts of data slowly generated by the same solver they intend to accelerate. We propose an open-source framework that enables the online training of these models from a large ensemble run of simulations. It leverages multiple levels of parallelism to generate rich datasets. The framework avoids I/O bottlenecks and storage issues by directly streaming the generated data. A training reservoir mitigates the inherent bias of streaming while maximizing GPU throughput. Experiment on training a fully connected network as a surrogate for the heat equation shows the proposed approach enables training on 8TB of data in 2 hours with an accuracy improved by 47% and a batch throughput multiplied by 13 compared to a traditional offline procedure. Lucas Thibaut Meyer, Marc Schouler, Robert Alexander Caulk, Alejandro Ribés, Bruno Raffin |
SC | 5 |
| 2021 | DEISA: Dask-Enabled In Situ AnalyticsabstractA widening performance gap is separating CPU performance and IO bandwidth on large scale systems. In some fields such as weather forecast and nuclear fusion, numerical models generate such amounts of data that classical post hoc processing is not feasible anymore due to the limits in both storage capacity and IO performance. In situ approaches are attractive to bypass disk accesses in these cases and fully leverage the HPC platform. They are however often complex to set up and can require to re-develop parallel versions of the analysis from scratch. In this paper we propose a hybrid model that is well suited for in situ workflows that combine regular simulations and irregular analytics. Our model couples the bulk synchronous parallel paradigm for simulation with a distributed task-based one for analysis. This reduces complexity and leverages the best of each of these two powerful paradigms. We validate the model with a prototype, called DEISA, that supports coupling MPI parallel codes with analyses written using Dask. This implementation requires minimal modifications of both the simulation and analysis codes compared to their post hoc counterpart. It give access to an already existing rich ecosystem to be used in situ such as the parallel versions of Numpy, Pandas and scikit-learn. Experiments in configurations up to 1024 cores show that DEISA can improve the simulation wallclock time (excluding analysis) by a factor up to 3 and the total experiment (including analysis) hour. core cost by a factor of up to 5 compared to parallel post hoc with plain Dask while requiring the modification of only two lines of python code, three of YAML, and none at all in a C simulation code already instrumented with PDI Data Interface. Amal Gueroudji, Julien Bigot, Bruno Raffin |
HiPC | 3 |
| 2021 | Preface - Special issue Advances on High Performance Computing for Artificial Intelligence
Marcos Dias de Assunção, Eduardo Rocha Rodrigues, Bruno Raffin |
J. Parallel Distributed Comput. | 3 |
| 2018 | Packed-Memory Quadtree: A cache-oblivious data structure for visual exploration of streaming spatiotemporal big data
Julio Toss, Cícero A. L. Pahins, Bruno Raffin, João Luiz Dihl Comba |
Comput. Graph. | 3 |
| 2017 | Automatic Data Filtering for In Situ WorkflowsabstractIn situ workflows contain tasks that exchange messages composed of several data fields. However, a consumer task may not necessarily need all the data fields from its producer. For example, a molecular dynamics simulation can produce atom positions, velocities, and forces; but some analyses require only atom positions. The user should decide whether to specialize the output of a producer task for a particular consumer and get better performance or to send more data than required by the consumer. The first option limits task portability, while the second wastes resources. In this paper, we introduce contracts for in situ tasks. A contract specifies for a producer each data field available for output and for a consumer the data fields needed as input. Comparing a producer and consumer contract allows automatic selection of the data fields a producer has to send for that consumer. We integrated our contracts mechanism within Decaf, a middleware for building and executing in situ workflows. Contracts enable to automatically extract at the producer the data the consumer needs. We evaluate the cost and performance of message extraction at runtime with both synthetic examples and a real scientific workflow coupling a molecular dynamics simulation with three different data analytics codes. Our contract-based automatic data extraction removes the need to specialize producers while entailing small overheads. Clément Mommessin, Matthieu Dreher, Bruno Raffin, Tom Peterka |
CLUSTER | 3 |
| 2017 | Melissa: large scale in transit sensitivity analysis avoiding intermediate filesabstractGlobal sensitivity analysis is an important step for analyzing and validating numerical simulations. One classical approach consists in computing statistics on the outputs from well-chosen multiple simulation runs. Simulation results are stored to disk and statistics are computed postmortem. Even if supercomputers enable to run large studies, scientists are constrained to run low resolution simulations with a limited number of probes to keep the amount of intermediate storage manageable. In this paper we propose a file avoiding, adaptive, fault tolerant and elastic framework that enables high resolution global sensitivity analysis at large scale. Our approach combines iterative statistics and in transit processing to compute Sobol' indices without any intermediate storage. Statistics are updated on-the-fly as soon as the in transit parallel server receives results from one of the running simulations. For one experiment, we computed the Sobol' indices on 10M hexahedra and 100 timesteps, running 8000 parallel simulations executed in 1h27 on up to 28672 cores, avoiding 48TB of file storage. Théophile Terraz, Alejandro Ribés, Yvan Fournier, Bertrand Iooss, Bruno Raffin |
SC | 5 |
| 2015 | Design and analysis of scheduling strategies for multi-CPU and multi-GPU architectures
João V. F. Lima, Vincent Danjean, Bruno Raffin, Nicolas Maillard |
Parallel Comput. | 4 |
| 2014 | A Flexible Framework for Asynchronous in Situ and in Transit Analytics for Scientific SimulationsabstractHigh performance computing systems are today composed of tens of thousands of processors and deep memory hierarchies. The next generation of machines will further increase the unbalance between I/O capabilities and processing power. To reduce the pressure on I/Os, the in situ analytics paradigm proposes to process the data as closely as possible to where and when the data are produced. Processing can be embedded in the simulation code, executed asynchronously on helper cores on the same nodes, or performed in transit on staging nodes dedicated to analytics. Today, software environnements as well as usage scenarios still need to be investigated before in situ analytics become a standard practice. In this paper we introduce a framework for designing, deploying and executing in situ scenarios. Based on a component model, the scientist designs analytics workflows by first developing processing components that are next assembled in a dataflow graph through a Python script. At runtime the graph is instantiated according to the execution context, the framework taking care of deploying the application on the target architecture and coordinating the analytics workflows with the simulation execution. Component coordination, zero-copy intra-node communications or inter-nodes data transfers rely on per-node distributed daemons. We evaluate various scenarios performing in situ and in transit analytics on large molecular dynamics systems simulated with Gromacs using up to 2048 cores. We show in particular that analytics processing can be performed on the fraction of resources the simulation does not use well, resulting in a limited impact on the simulation performance (less than 9%). Our more advanced scenario combines in situ and in transit processing to compute a molecular surface based on the Quick surf algorithm. Matthieu Dreher, Bruno Raffin |
CCGRID | 2 |
| 2013 | X-kaapi: A Multi Paradigm Runtime for Multicore ArchitecturesabstractThe paper presents X-KAAPI, a compact runtime for multicore architectures that brings multi parallel paradigms (parallel independent loops, fork-join tasks and dataflow tasks) in a unified framework without performance penalty. Comparisons on independent loops with OpenMP and on dense linear algebra with QUARK/PLASMA confirm our design decisions. Applied to EUROPLEXUS, an industrial simulation code for fast transient dynamics, we show that X-KAAPI achieves high speedups on multicore architectures by efficiently parallelizing both independent loops and dataflow tasks. Fabien Le Mentec, Vincent Faucher, Bruno Raffin |
ICPP | 4 |
| 2013 | XKaapi: A Runtime System for Data-Flow Task Programming on Heterogeneous ArchitecturesabstractMost recent HPC platforms have heterogeneous nodes composed of multi-core CPUs and accelerators, like GPUs. Programming such nodes is typically based on a combination of OpenMP and CUDA/OpenCL codes; scheduling relies on a static partitioning and cost model. We present the XKaapi runtime system for data-flow task programming on multi-CPU and multi-GPU architectures, which supports a data-flow task model and a locality-aware work stealing scheduler. XKaapi enables task multi-implementation on CPU or GPU and multi-level parallelism with different grain sizes. We show performance results on two dense linear algebra kernels, matrix product (GEMM) and Cholesky factorization (POTRF), to evaluate XKaapi on a heterogeneous architecture composed of two hexa-core CPUs and eight NVIDIA Fermi GPUs. Our conclusion is two-fold. First, fine grained parallelism and online scheduling achieve performance results as good as static strategies, and in most cases outperform them. This is due to an improved work stealing strategy that includes locality information; a very light implementation of the tasks in XKaapi; and an optimized search for ready tasks. Next, the multi-level parallelism on multiple CPUs and GPUs enabled by XKaapi led to a highly efficient Cholesky factorization. Using eight NVIDIA Fermi GPUs and four CPUs, we measure up to 2.43 TFlop/s on double precision matrix product and 1.79 TFlop/s on Cholesky factorization; and respectively 5.09 TFlop/s and 3.92 TFlop/s in single precision. João V. F. Lima, Nicolas Maillard, Bruno Raffin |
IPDPS | 4 |
| 2013 | Preliminary Experiments with XKaapi on Intel Xeon Phi CoprocessorabstractThis paper presents preliminary performance comparisons of parallel applications developed natively for the Intel Xeon Phi accelerator using three different parallel programming environments and their associated runtime systems. We compare Intel OpenMP, Intel CilkPlus and XKaapi together on the same benchmark suite and we provide comparisons between an Intel Xeon Phi coprocessor and a Sandy Bridge Xeon-based machine. Our benchmark suite is composed of three computing kernels: a Fibonacci computation that allows to study the overhead and the scalability of the runtime system, a NQueens application generating irregular and dynamic tasks and a Cholesky factorization algorithm. We also compare the Cholesky factorization with the parallel algorithm provided by the Intel MKL library for Intel Xeon Phi. Performance evaluation shows our XKaapi data-flow parallel programming environment exposes the lowest overhead of all and is highly competitive with native OpenMP and CilkPlus environments on Xeon Phi. Moreover, the efficient handling of data-flow dependencies between tasks makes our XKaapi environment exhibit more parallelism for some applications such as the Cholesky factorization. In that case, we observe substantial gains with up to 180 hardware threads over the state of the art MKL, with a 47% performance increase for 60 hardware threads. João V. F. Lima, François Broquedis, Bruno Raffin |
SBAC-PAD | 4 |
| 2012 | A hierarchical component model for large parallel interactive applications
Jean-Denis Lesage, Bruno Raffin |
J. Supercomput. | 2 |
| 2010 | Multi-GPU and Multi-CPU Parallelization for Interactive Physics Simulations
Everton Hermann, Bruno Raffin, François Faure, Jérémie Allard |
Euro-Par (2) | 2 |
| 2010 | A 3d data intensive tele-immersive gridabstractNetworked virtual environments like Second Life enable distant people to meet for leisure as well as work. But users are represented through avatars controlled by keyboards and mouses, leading to a low sense of presence especially regarding body language. Multi-camera real-time 3D modeling offers a way to ensure a significantly higher sense of presence. But producing quality geometries, well textured, and to enable distant user tele-presence in non trivial virtual environments is still a challenge today. Benjamin Petit, Thomas Dupeux, Benoît Bossavit, Joeffrey Legaux, Bruno Raffin, Emmanuel Melin, Jean-Sébastien Franco, Ingo Assenmacher, Edmond Boyer |
ACM Multimedia | 5 |
| 2010 | Binary Mesh Partitioning for Cache-Efficient VisualizationabstractOne important bottleneck when visualizing large data sets is the data transfer between processor and memory. Cache-aware (CA) and cache-oblivious (CO) algorithms take into consideration the memory hierarchy to design cache efficient algorithms. CO approaches have the advantage to adapt to unknown and varying memory hierarchies. Recent CA and CO algorithms developed for 3D mesh layouts significantly improve performance of previous approaches, but they lack of theoretical performance guarantees. We present in this paper a {\schmi O}(N\log N) algorithm to compute a CO layout for unstructured but well shaped meshes. We prove that a coherent traversal of a N-size mesh in dimension d induces less than N/B+{\schmi O}(N/M;{1/d}) cache-misses where B and M are the block size and the cache size, respectively. Experiments show that our layout computation is faster and significantly less memory consuming than the best known CO algorithm. Performance is comparable to this algorithm for classical visualization algorithm access patterns, or better when the BSP tree produced while computing the layout is used as an acceleration data structure adjusted to the layout. We also show that cache oblivious approaches lead to significant performance increases on recent GPU architectures. Marc Tchiboukdjian, Vincent Danjean, Bruno Raffin |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2008 | Grimage: 3D modeling for remote collaboration and telepresenceabstractReal-time multi-camera 3D modeling provides full-body geometric and photometric data on the objects present in the acquisition space. It can be used as an input device for rendering textured 3D models, and for computing interactions with virtual objects through a physical simulation engine. In this paper we present a work in progress to build a collaborative environment where two distant users, each one 3D modeled in real-time, interact in a shared virtual world. Benjamin Petit, Jean-Denis Lesage, Jean-Sébastien Franco, Edmond Boyer, Bruno Raffin |
VRST | 5 |
| 2007 | Work Stealing for Time-constrained Octree Exploration: Application to Real-time 3D ModelingabstractThis paper introduces a dynamic work balancing algorithm, based on work stealing, for time-constrained parallel octree carving. The performance of the algorithm is proved and confirmed by experimental results where the algorithm is applied to a real-time 3D modeling from multiple video streams. Compared to classical work stealing, the proposed algorithm enforces a relaxed width first octree carving that enables to stop computations at anytime while ensuring a balanced carving. Luciano P. Soares, Clément Ménier, Bruno Raffin, Jean-Louis Roch |
EGPGV | 3 |
| 2007 | A Hierarchical Programming Model for Large Parallel Interactive Applications
Jean-Denis Lesage, Bruno Raffin |
NPC | 2 |
| 2007 | Parallel Adaptive Octree Carving for Real-time 3D ModelingabstractInternational audience Luciano P. Soares, Clément Ménier, Bruno Raffin, Jean-Louis Roch |
VR | 3 |
| 2007 | Parallel graphics and visualization
Luís Paulo Santos, Bruno Raffin, Alan Heirich |
Parallel Comput. | 2 |
| 2006 | The GrImage Platform: A Mixed Reality Environment for InteractionsabstractIn this paper, we present a scalable architecture to compute, visualize and interact with 3D dynamic models of real scenes. This architecture is designed for mixed reality applications requiring such dynamic models, tele-immersion for instance. Our system consists in 3 main parts: the acquisition, based on standard firewire cameras; the computation, based on a distribution scheme over a cluster of PC and using a recent shape-from-silhouette algorithm which leads to optimally precise 3D models; the visualization, which is achieved on a multiple display wall. The proposed distribution scheme ensures scalability of the system and hereby allows control over the number of cameras used for acquisition, the frame-rate, or the number of projectors used for high resolution visualization. To our knowledge this is the first completely scalable vision architecture for real time 3D modeling, from acquisition to visualization through computation. Experimental results show that this framework is very promising for real time 3D interactions. Jérémie Allard, Jean-Sébastien Franco, Clément Ménier, Edmond Boyer, Bruno Raffin |
ICVS | 5 |
| 2006 | Distributed Physical Based Simulations for Large VR ApplicationsabstractWe present a novel software framework for developing highly animated virtual reality applications. Using a modular application design, our goal is to alleviate software engineering issues while yielding efficient execution on parallel machines. We target worlds involving numerous animated objects managed by physical based simulations. Mixing rigid objects, fluids, mass-spring or other deformable objects leads to complex interactions between them. Today no unified simulation algorithm with a reasonable complexity is available to manage all these types of objects. We propose a framework for coupling and distributing existing algorithms. We reuse and extend the data-flow model where an application is built from modules exchanging data through connections. The model relies on two main classes of modules, animators and interactors. Animators are responsible for updating objects’ states from forces applied to them. These forces are computed in parallel by interactors using the objects’ states they receive from animators. The network interconnecting modules can be progressively optimized. From a simple fully connected network enforcing a synchronous semantics, it can evolve towards an active network able to implement a bounding volume based dynamic routing or an asynchronous data re-sampling. As a result, we present an application managing interactions between rigid objects, mass-spring objects and a fluid. It is executed in real-time on a 54 processors cluster driving 5 cameras and 16 projectors for user interactions. Jérémie Allard, Bruno Raffin |
VR | 2 |
| 2006 | PC Clusters for Virtual RealityabstractIn the late 90’s the emergence of high performance 3D commodity graphics cards opened the way to use PC clusters for high performance Virtual Reality (VR) applications. Today PC clusters are broadly used to drive multi projector immersive environments. In this paper, we survey the different approaches that have been developed to use PC clusters for VR applications. We review the most common software tools that enable to take advantage of the power of clusters. We also discuss some new trends. Bruno Raffin, Luciano P. Soares, Tao Ni 0002, Robert Ball, Greg S. Schmidt, Mark A. Livingston, Oliver G. Staadt, Richard May 0001 |
VR | 1 |
| 2005 | A Shader-Based Parallel Rendering FrameworkabstractExisting parallel or remote rendering solutions rely on communicating pixels, OpenGL commands, scene-graph changes or application-specific data. We propose an intermediate solution based on a set of independent graphics primitives that use hardware shaders to specify their visual appearance. Compared to an OpenGL based approach, it reduces the complexity of the model by eliminating most fixed function parameters while giving access to the latest functionalities of graphics cards. It also suppresses the OpenGL state machine that creates data dependencies making primitive re-scheduling difficult. Using a retained-mode communication protocol transmitting changes between each frame, combined with the possibility to use shaders to implement interactive data processing operations instead of sending final colors and geometry, we are able to optimize the network load. High level information such as bounding volumes is used to setup advanced schemes where primitives are issued in parallel, routed according to their visibility, merged and re-ordered when received for rendering. Different optimization algorithms can be efficiently implemented, saving network bandwidth or reducing texture switches for instance. We present performance results based on two VTK applications, a parallel iso-surface extraction and a parallel volume renderer. We compare our approach with Chromium. Results show that our approach leads to significantly better performance and scalability, while offering easy access to hardware accelerated rendering algorithms. Jérémie Allard, Bruno Raffin |
IEEE Visualization | 2 |
| 2005 | Parallel graphics and visualization
Bruno Raffin, Han-Wei Shen, Dirk Bartz |
Parallel Comput. | 1 |
| 2004 | FlowVR: A Middleware for Large Scale Virtual Reality Applications
Jérémie Allard, Valérie Gouranton, Loïck Lecointre, Sébastien Limet, Bruno Raffin, Sophie Robert 0001 |
Euro-Par | 5 |
| 2003 | Coupling Parallel Simulation and Multi-display Visualization on a PC Cluster
Jérémie Allard, Bruno Raffin, Florence Zara |
Euro-Par | 2 |
| 2003 | Commodity Clusters for Virtual RealityabstractMultiprojection Immersive Environments are used in many applications ranging fromscience, engineering and art. Such VR-oriented systems have traditionally been powered byhigh-end graphics workstations or supercomputers, but, recently, clusters of commoditycomputers (such as PCs, Macs, and low cost workstations) have become a practical alternative.The advantages of a commodity cluster include low cost, flexibility, access to technology, andperformance scalability. The workshop will be divided in three parts. An introductory tutorialon VR clustering technologies will cover hardware and software issues and will emphasize freesoftware solutions. The second part will consist of 20-minute presentations on specificcommodity technologies as applied to virtual reality. An open panel discussion will close theworkshop. Bruno Raffin, Marcelo Knörich Zuffo, Hank Kerzmarski, Zhingeng Pan |
VR | 1 |
| 2002 | Net Juggler: Running VR Juggler with Multiple Displays on a Commodity Component ClusterabstractNet Juggler is an open source library that turns a commodity component cluster running the VR Juggler platform on each node into a single VR Juggler image cluster. Application parallelization is transparent to the user and leads to high performance executions even with limited bandwidth networks. Jérémie Allard, Valérie Gouranton, Loïck Lecointre, Emmanuel Melin, Bruno Raffin |
VR | 5 |
| 1999 | A Cost Model for Asynchronous and Structured Message Passing
Emmanuel Melin, Bruno Raffin, Xavier Rebeuf, Bernard Virot |
Euro-Par | 2 |
| 1998 | A Structured Synchronization and Communication Model Fitting Irregular Data Accesses
Emmanuel Melin, Bruno Raffin, Xavier Rebeuf, Bernard Virot |
J. Parallel Distributed Comput. | 2 |
| 1997 | SCL-chan: An Asynchronous Data-Parallel Language for Irregular AlgorithmsabstractParallelism suffers from a lack of programming languages both simple to handle and able to take advantage of the power of present parallel computers. If parallelism expression is too high level, compilers have to perform complex optimizations leading often to poor performances. One the other hand, too low level parallelism transfers difficulties toward the programmer. We propose a programming language that integrates both a synchronous data parallel programming model and an asynchronous execution model. The synchronous data parallel programming model allows safe program design. The asynchronous execution model yields an efficient execution on present MIMD architectures without any program transformation. Our language relies an logical instruction ordering exploited by specific send/receive communications. It allows one to express only the effective data dependences between processors. This ability is enforced by a possible send/receive unmatching, useful for irregular algorithms. A sparse vector computation exemplifies our language potentialities. Emmanuel Melin, Bruno Raffin, Xavier Rebeuf, Bernard Virot |
HIPS | 2 |
| 1995 | Learning and generalization with Minimerror, a temperature-dependent learning algorithmabstractWe study the numerical performances of Minimerror, a recently introduced learning algorithm for the perceptron that has analytically been shown to be optimal both on learning linearly and nonlinearly separable functions. We present its implementation on learning linearly separable boolean functions. Numerical results are in excellent agreement with the theoretical predictions. Bruno Raffin, Mirta B. Gordon |
Neural Comput. | 1 |