EDBT 2026 Demo / reviewers in the wild / expert
Philippe Swartvagher
dblp:272/7250
· DBLP profile ↗
7ranked-venue papers
2as first author
6since 2021 · last 2025
0000-0003-3786-7364ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 2 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | PALLAS: A Generic Trace Format for Large HPC Trace AnalysisabstractIdentifying performance bottlenecks in a parallel application is tedious, especially because it requires analyzing the behaviour of various software components, as bottlenecks may have several causes and symptoms. For example, a load imbalance may cause long MPI waiting times, or contention on disk may degrade the performance of I/O operations. Detecting a performance problem means investigating the execution of an application and applying several performance analysis techniques. To do so, one can use a tracing tool to collect information describing the behaviour of the application. At the end of the execution, a trace file in a specific format is available to the application user, which can be used to conduct a complete post-mortem investigation. Several challenges emerge from the generation and use of traces. Tracing applications may alter the performance of the application, and can create thousands of heavy trace files, especially at a large scale. Most importantly, the post-mortem analysis needs to load these thousands of trace files in memory, and process them. This quickly becomes impractical for large scale applications, as memory gets exhausted and the number of opened files exceeds the system capacity. In this paper, we propose PALLAS, a generic trace format tailored for conducting various post-mortem performance analysis of traces describing large executions of HPC applications. During the execution of the application, PALLAS collects events and detects their repetitions on-the-fly. When storing the trace to disk, PALLAS groups the data from similar events or groups of events together in order to later speed up trace reading. We demonstrate that the PALLAS online detection of the program structure does not significantly degrade the performance of the applications. Moreover, the PALLAS format allows faster trace analysis compared to other evaluated trace formats. Overall, the PALLAS trace format allows an interactive analysis of a trace that is required when a user investigates a performance problem. Catherine Guelque, Valentin Honoré, Philippe Swartvagher, Gaël Thomas 0001, François Trahay |
IPDPS | 3 |
| 2025 | Reproducibility Report for SC25 Paper Story of Two GPUs: Characterizing the Resilience of Hopper H100 and Ampere A100 GPUsabstractThis reproducibility report provides details about the artifact evaluation done with regards to the Artifact Description and Evaluation appendix of SC25 paper Story of Two GPUs: Characterizing the Resilience of Hopper H100 and Ampere A100 GPUs by Cui et al. The work was done as part of the Reproducibility Initiative of SC25. The author is a member of the SC25 Reproducibility Committee. Philippe Swartvagher |
SC | 1 |
| 2025 | Reproducibility Report for SC25 Paper HP-MDR: High-performance and Portable Data Refactoring and Progressive Retrieval with Advanced GPUsabstractThis reproducibility report provides details about the artifact evaluation done with regards to the Artifact Description and Evaluation appendix of SC25 paper HP-MDR: High-performance and Portable Data Refactoring and Progressive Retrieval with Advanced GPUs by Podhorszki et al. The work was done as part of the Reproducibility Initiative of SC25. The author is a member of the SC25 Reproducibility Committee. Philippe Swartvagher |
SC | 1 |
| 2024 | Improved Parallel Application Performance and Makespan by Colocation and Topology-aware Process MappingabstractIn modern, deeply hierarchical HPC systems shared resource congestion can hinder the efficient use of many cores by parallel applications and degrade performance. Such congestion is often caused when parallel processes within an application that execute similar operations share the same resources. Previous research suggests using fewer cores with better process-to-core mapping can improve applications’ performance but leaves many cores unused. To utilize these cores, we colocate additional applications and map them using a topology-aware process-to-core, application-agnostic mapping algorithm. We show that these mappings significantly impact memory bandwidth and communication latency. We evaluate our approach using eight parallel applications on an HPC system with 128-core nodes, demonstrating the performance effects of mappings combined with colocation. Our goal is to determine whether colocation with topology-aware mapping is a viable alternative to typical exclusive node allocation. Our results show makespan improvements of 2.4x over exclusive allocation in an HPC system, demonstrating the potential benefits of colocation with optimized mappings. Ioannis Vardas, Sascha Hunold, Philippe Swartvagher, Jesper Larsson Träff |
CCGrid | 3 |
| 2024 | Tracing task-based runtime systems: Feedbacks from the StarPU caseabstractSummary Given the complexity of current supercomputers and applications, being able to trace application executions to understand their behavior is not a luxury. As constraints, tracing systems have to be as little intrusive as possible in the application code and performances, and be precise enough in the collected data. In this article, we present how works the tracing system used by the task‐based runtime systemStarPU. We study the different sources of performance overhead coming from the tracing system and how to reduce these overheads. Then, we evaluate the accuracy of distributed traces with different clock synchronization techniques. Finally, we summarize our experiments and conclusions with the lessons we learned to efficiently trace applications, and the list of characteristics each tracing system should feature to be competitive. The reported experiments and implementation details comprise a feedback of integrating into a task‐based runtime system state‐of‐the‐art techniques to efficiently and precisely trace application executions. We highlight the points every application developer or end‐user should be aware of to seamlessly integrate a tracing system or just trace application executions. Alexandre Denis 0001, Emmanuel Jeannot, Philippe Swartvagher, Samuel Thibault |
Concurr. Comput. Pract. Exp. | 3 |
| 2021 | Interferences between Communications and Computations in Distributed HPC SystemsabstractInternational audience Alexandre Denis 0001, Emmanuel Jeannot, Philippe Swartvagher |
ICPP | 3 |
| 2020 | Using Dynamic Broadcasts to Improve Task-Based Runtime Performances
Alexandre Denis 0001, Emmanuel Jeannot, Philippe Swartvagher, Samuel Thibault |
Euro-Par | 3 |