VLDB 2026 Research / reviewers in the wild / expert
Fabien Gaud
dblp:17/8330
· DBLP profile ↗
5ranked-venue papers
2as first author
0since 2021 · last 2016
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 2 first-authorSoftware engineering, systems software and programming languages · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Software engineering, system software, and programming languages
3 papers |
Operating systems · 100% | |
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Cloud and datacenter computing · 38% Memory systems · 30% Performance modeling and evaluation · 19% |
Topics — the 9 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Operating systems › resource management
memory management |
0.4 | 2 | 2014 | Large Pages May Be Harmful on NUMA Systems · USENIX ATC 2014 Traffic management: a holistic approach to memory placement on NUMA systems · ASPLOS 2013 |
Operating systems › resource management › process management › CPU scheduling
thread scheduling |
0.2 | 1 | 2016 | The Linux scheduler: a decade of wasted cores · EuroSys 2016 |
Operating systems › resource management › memory management
huge pages |
0.2 | 1 | 2014 | Large Pages May Be Harmful on NUMA Systems · USENIX ATC 2014 |
Memory systems › non-uniform memory access
NUMA data placement |
0.2 | 1 | 2013 | Traffic management: a holistic approach to memory placement on NUMA systems · ASPLOS 2013 |
Cloud and datacenter computing
cluster resource management and scheduling |
0.1 | 1 | 2012 | A practical method for estimating performance degradation on multicore processors, and its application to HPC workloads · SC 2012 |
Cloud and datacenter computing › virtualization › virtual machine management
server consolidation |
0.1 | 1 | 2012 | A practical method for estimating performance degradation on multicore processors, and its application to HPC workloads · SC 2012 |
Operating systems
resource management |
0.1 | 1 | 2016 | The Linux scheduler: a decade of wasted cores · EuroSys 2016 |
Processor architecture and microarchitecture
multicore design |
0.1 | 1 | 2014 | Large Pages May Be Harmful on NUMA Systems · USENIX ATC 2014 |
Memory systems
non-uniform memory access |
0.1 | 1 | 2014 | Large Pages May Be Harmful on NUMA Systems · USENIX ATC 2014 |
Methods — techniques the papers use, named apart from their topics
trace-driven simulation · 0.3kernel implementation · 0.3scheduling visualization · 0.2online invariant checking · 0.2online estimation · 0.1machine learning · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2016 | The Linux scheduler: a decade of wasted coresabstractAs a central part of resource management, the OS thread scheduler must maintain the following, simple, invariant: make sure that ready threads are scheduled on available cores. As simple as it may seem, we found that this invariant is often broken in Linux. Cores may stay idle for seconds while ready threads are waiting in runqueues. In our experiments, these performance bugs caused many-fold performance degradation for synchronization-heavy scientific applications, 13% higher latency for kernel make, and a 14-23% decrease in TPC-H throughput for a widely used commercial database. The main contribution of this work is the discovery and analysis of these bugs and providing the fixes. Conventional testing techniques and debugging tools are ineffective at confirming or understanding this kind of bugs, because their symptoms are often evasive. To drive our investigation, we built new tools that check for violation of the invariant online and visualize scheduling activity. They are simple, easily portable across kernel versions, and run with a negligible overhead. We believe that making these tools part of the kernel developers' tool belt can help keep this type of bug at bay. Jean-Pierre Lozi, Baptiste Lepers, Justin R. Funston, Fabien Gaud, Vivien Quéma, Alexandra Fedorova |
EuroSys | 4 |
| 2014 | Large Pages May Be Harmful on NUMA Systems
Fabien Gaud, Baptiste Lepers, Jeremie Decouchant, Justin R. Funston, Alexandra Fedorova, Vivien Quéma |
USENIX ATC | 1 |
| 2013 | Traffic management: a holistic approach to memory placement on NUMA systemsabstractNUMA systems are characterized by Non-Uniform Memory Access times, where accessing data in a remote node takes longer than a local access. NUMA hardware has been built since the late 80's, and the operating systems designed for it were optimized for access locality. They co-located memory pages with the threads that accessed them, so as to avoid the cost of remote accesses. Contrary to older systems, modern NUMA hardware has much smaller remote wire delays, and so remote access costs per se are not the main concern for performance, as we discovered in this work. Instead, congestion on memory controllers and interconnects, caused by memory traffic from data-intensive applications, hurts performance a lot more. Because of that, memory placement algorithms must be redesigned to target traffic congestion. This requires an arsenal of techniques that go beyond optimizing locality. In this paper we describe Carrefour, an algorithm that addresses this goal. We implemented Carrefour in Linux and obtained performance improvements of up to 3.6 relative to the default kernel, as well as significant improvements compared to NUMA-aware patchsets available for Linux. Carrefour never hurts performance by more than 4% when memory placement cannot be improved. We present the design of Carrefour, the challenges of implementing it on modern hardware, and draw insights about hardware support that would help optimize system software on future NUMA systems. Mohammad Dashti 0002, Alexandra Fedorova, Justin R. Funston, Fabien Gaud, Renaud Lachaize, Baptiste Lepers, Vivien Quéma, Mark Roth |
ASPLOS | 4 |
| 2012 | A practical method for estimating performance degradation on multicore processors, and its application to HPC workloadsabstractWhen multiple threads or processes run on a multi-core CPU they compete for shared resources, such as caches and memory controllers, and can suffer performance degradation as high as 200%. We design and evaluate a new machine learning model that estimates this degradation online, on previously unseen workloads, and without perturbing the execution. Our motivation is to help data center and HPC cluster operators effectively use workload consolidation. Data center consolidation is about placing many applications on the same server to maximize hardware utilization. In HPC clusters, processes of the same distributed applications run on the same machine. Consolidation improves hardware utilization, but may sacrifice performance as processes compete for resources. Our model helps determine when consolidation is overly harmful to performance. Our work is the first to apply machine learning to this problem domain, and we report on our experience reaping the advantages of machine learning while navigating around its limitations. We demonstrate how the model can be used to improve performance fidelity and save energy for HPC workloads. Tyler Dwyer, Alexandra Fedorova, Sergey Blagodurov, Mark Roth, Fabien Gaud, Jian Pei 0001 |
SC | 5 |
| 2010 | Efficient Workstealing for Multicore Event-Driven SystemsabstractMany high-performance communicating systems are designed using the event-driven paradigm. As multicore platforms are now pervasive, it becomes crucial for such systems to take advantage of the available hardware parallelism. Event-coloring is a promising approach in this regard. First, it allows programmers to simply and progressively inject support for the safe, parallel execution of multiple event handlers through the use of annotations. Second, it relies on a workstealing algorithm to dynamically balance the execution of event handlers on the available cores. This paper studies the impact of the workstealing algorithm on the overall system performance. We first show that the only existing workstealing algorithm designed for event-coloring runtimes is not always efficient: for instance, it causes a 33% performance degradation on a Web server. We then introduce several enhancements to improve the workstealing behavior. An evaluation using both micro benchmarks and real applications, a Web server and the Secure File Server (SFS), shows that our system consistently outperforms a state-of-the-art runtime (Libasync-smp), with or without workstealing. In particular, our new workstealing improves performance by up to +25% compared to Libasync-smp without workstealing and by up to +73% compared to the Libasync-smp workstealing algorithm, in the Web server case. Fabien Gaud, Sylvain Geneves, Renaud Lachaize, Baptiste Lepers, Fabien Mottet, Gilles Muller, Vivien Quéma |
ICDCS | 1 |