VLDB 2026 Research / reviewers in the wild / expert
Arnaud Fiorini
dblp:330/2202
· DBLP profile ↗
3ranked-venue papers
1as first author
3since 2021 · last 2026
0000-0002-0295-636XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Transparent and Efficient Performance Analysis Approach to Enhance DPDK ObservabilityabstractIn recent years, the rapid growth of network traffic and the performance bottlenecks inherent in kernel networking stacks have driven the widespread adoption of userspace networking frameworks. While kernel-bypass solutions such as the Data Plane Development Kit (DPDK) effectively eliminate kernel overhead, they also limit observability for traditional monitoring tools, complicating fault diagnosis and performance tuning. This observability gap, coupled with the complexity of modern packet-processing software, makes diagnosing performance issues increasingly difficult. This paper presents a performance analysis framework tailored for DPDK-based applications. The framework leverages trace data collected through DPDK's native tracer to derive targeted performance metrics, which are visualized through interactive, domain-specific analyses in Trace Compass. By enabling fine-grained observability with minimal runtime overhead, the approach bridges the gap between low-level tracing and actionable performance insights. To ground our design in real-world needs, we surveyed 19 industry practitioners to validate our design choices and capture empirical evidence of the debugging challenges encountered when diagnosing DPDK-based applications. We further demonstrate how the proposed analyses can reveal and explain performance bottlenecks in a widely used software router. Adel Belkhiri, Arnaud Fiorini, Matthew Khouzam, Heng Li 0007 |
ICPE | 2 |
| 2025 | HybridRCA: Lightweight Critical-Path-Aware Hybrid Tracing for Root-Cause Analysis in Production Microservicesabstract[Context] Distributed cloud-native systems operated by our industrial partners, including Ericsson and Ciena, generate millions of trace spans daily. Capturing and analyzing this data at full granularity is infeasible due to excessive storage and computational overhead. [Objective] We aim to enable fast and accurate RCA with minimal trace volume and system overhead, quickly pinpointing the service causing a latency spike, making it practical for large-scale production environments. [Method] We present HybridRCA, a critical-path-aware RCA pipeline that (1) extracts the critical path of each request, (2) applies a PageRank-weighted spectrum analysis to identify suspicious spans, and (3) collects system metrics only for targeted spans. [Results] Across three microservice benchmarks (HotRod, TrainTicket, OnlineBoutique), HybridRCA improves recall by an average of$\text{0.45 \%}$over the best existing methods, while analyzing up to 22.6 % fewer spans and reducing kernel-level storage usage by over 99%. [Significance] HybridRCA addresses key observability challenges faced by our industry partners, enabling scalable, low-overhead RCA in real-world distributed systems. Maryam Ekhlasi, Arnaud Fiorini, Michel R. Dagenais, Naser Ezzati-Jivan, Maxime Lamothe |
ICSME | 2 |
| 2022 | Visualization of profiling and tracing in CPU-GPU programsabstractSummary As the complexity of the toolchain increases for heterogeneous CPU‐GPU systems, the needs for comprehensive tracing and debugging tools also grows. Heterogeneous platforms bring new possibilities but also new performance issues that are hard to detect. Some techniques that were used on CPU programs are now adapted to GPUs. However, there are some concepts specific to GPUs, like SIMD processing, and the effects of the close interactions between the CPUs and the GPUs, with shared virtual memory and user‐level queues. Multiple sources of data need to be extracted and correlated to obtain a more global view of the performance. In this article, we introduce a novel approach for measuring and visualizing performance defects inside CPU‐GPU programs by combining kernel events, compute kernel events, user API calls and memory transfers. We created two new views that combine this information, to help provide a global view. This framework uses the open source user queue system described in the HSA standard. It can easily be adapted to any user queue system for heterogeneous computing devices. We compare this framework with current existing tools and test it against the Rodinia benchmark. We look at how the execution behavior affects the tracing and profiling overhead and we use Trace Compass to visualize the resulting trace. Arnaud Fiorini, Michel R. Dagenais |
Concurr. Comput. Pract. Exp. | 1 |