Paul Delestrac

dblp:337/0913 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
5since 2021 · last 2026
0000-0002-7476-1422ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 3 first-author · 5 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 pHNSW: PCA-Based Filtering to Accelerate HNSW Approximate Nearest Neighbor Search
abstract
Hierarchical Navigable Small World (HNSW) has demonstrated impressive accuracy and low latency for high-dimensional nearest neighbor searches. However, its high computational demands and irregular, large-volume data access patterns present significant challenges to search efficiency. To address these challenges, we introduce pHNSW, an algorithm-hardware co-optimized solution that accelerates HNSW through Principal Component Analysis (PCA) filtering. On the algorithm side, we apply PCA filtering to reduce the dimensionality of the dataset, thereby lowering the volume of neighbor access and decreasing the computational load for distance calculations. On the hardware side, we design the pHNSW processor with custom instructions to optimize search throughput and energy efficiency. In the experiments, we synthesized the pHNSW processor RTL design with a 65nm technology node and evaluated it using DDR4 and HBM1.0 DRAM standards. The results show that pHNSW boosts Queries per Second (QPS) by $14.47 \times \sim 21.37 \times$ on a CPU and $5.37 \times \sim 8.46 \times$ on a GPU, while reducing energy consumption by up to $57.4 \%$ compared to standard HNSW implementation.
Guangyi Zeng, Paul Delestrac, Enyi Yao, Simei Yang
ASP-DAC3
2025 Modeling and Scheduling of Composable Instruction Set
abstract
State-of-the-art hardware accelerators with custom instruction set architectures (ISAs) are widely used in AI/ML applications. Despite the outstanding progress of modern accelerators, their performance potential is limited by their ISA. Recently, research around composable instruction sets (CIS) has emerged with the aim of improving the efficiency of modern accelerator ISAs. However, while CIS significantly outperforms existing ISAs with near-optimal PE utilization, no existing instruction scheduling approach supports the temporal composability requirements to correctly schedule CIS programs. The goal of this work is to propose a scalable scheduling algorithm that can meet CIS requirements. We propose a novel timing model structure built upon five fundamental concepts - operations, events, transformations, anchors, and constraints - that accurately capture CIS timing behavior and constraints. Using this model, we automatically formulate scheduling problems that can be solved using constraint programming (CP) solvers. In addition, we propose a methodology to synchronize scheduled instructions and generate functional CIS assembly code. Finally, we design experiments to study the scalability and quality of our scheduling approach. Across multiple dimensions of complexity, our approach scales linearly and produces near-optimal scheduling solutions.
Yu Yang 0020, Paul Delestrac, Ahmed Hemani
DSD2
2024 Analyzing GPU Energy Consumption in Data Movement and Storage
abstract
GPUs are the prevailing solution to execute high-performance tasks (e.g., machine learning training). As the peak performance of modern GPUs increases with each generation, so does their thermal design power (TDP). Hence, identifying energy bottlenecks in the GPU architecture is crucial to designing more efficient architectures in the future. However, due to the complex proprietary nature of modern GPU architectures, providing a detailed breakdown of the GPU energy consumption is not trivial. The goal of this work is to estimate a lower bound for the energy consumed by data movement and storage in modern GPU architectures, leveraging internal power sensors. We establish a basic energy model for modern GPUs, focused on data movement to/from the hardware-managed caches and software-managed memories. We propose a methodology to calibrate the energy model using microbenchmarks, performance counters, and the internal power sensor. We experimentally calibrate the model on an A100 NVIDIA GPU. Then, we challenge the consistency of the results by cross-validating with modified microbenchmarks with additional instructions. Finally, we use the calibrated energy model to evaluate breakdowns for workloads of increasing complexity (e.g., a ResNet-50 training iteration with different software optimizations). Our results show that data movement dominates the dynamic energy consumption of the GPU (up to 84%), with DRAM accesses being the main contributor.
Paul Delestrac, Jonathan Miquel, Debjyoti Bhattacharjee, Diksha Moolchandani, Francky Catthoor, Lionel Torres, David Novo
ASAP1
2024 Multi-Level Analysis of GPU Utilization in ML Training Workloads
abstract
Training time has become a critical bottleneck due to the recent proliferation of large-parameter ML models. GPUs continue to be the prevailing architecture for training ML models. However, the complex execution flow of ML frameworks makes it difficult to understand GPU computing resource utilization. Our main goal is to provide a better understanding of how efficiently ML training workloads use the computing resources of modern GPUs. To this end, we first describe an ideal reference execution of a GPU-accelerated ML training loop and identify relevant metrics that can be measured using existing profiling tools. Second, we produce a coherent integration of the traces obtained from each profiling tool. Third, we leverage the metrics within our integrated trace to analyze the impact of different software optimizations (e.g., mixed-precision, various ML frameworks, and execution modes) on the throughput and the associated utilization at multiple levels of hardware abstraction (i.e., whole GPU, SM subpartitions, issue slots, and tensor cores). In our results on two modern GPUs, we present seven takeaways and show that although close to 100% utilization is generally achieved at the GPU level, average utilization of the issue slots and tensor cores always remains below 50% and 5.2%, respectively.
Paul Delestrac, Debjyoti Bhattacharjee, Simei Yang, Diksha Moolchandani, Francky Catthoor, Lionel Torres, David Novo
DATE1
2022 Demystifying the TensorFlow Eager Execution of Deep Learning Inference on a CPU-GPU Tandem
abstract
Machine Learning (ML) frameworks are tools that facilitate the development and deployment of ML models. These tools are major catalysts of the recent explosion in ML models and hardware accelerators thanks to their high programming abstraction. However, such an abstraction also obfuscates the run-time execution of the model and complicates the understanding and identification of performance bottlenecks. In this paper, we demystify how a modern ML framework manages code execution from a high-level programming language. We focus our work on the TensorFlow eager execution, which remains obscure to many users despite being the simplest mode of execution in TensorFlow. We describe in detail the process followed by the runtime to run code on a CPU-GPU tandem. We propose new metrics to analyze the framework's runtime performance overhead. We use our metrics to conduct in-depth analysis of the inference process of two Convolutional Neural Networks (CNNs) (LeNet-5 and ResNet-50) and a transformer (BERT) for different batch sizes. Our results show that GPU kernels execution need to be long enough to exploit thread parallelism, and effectively hide the runtime overhead of the ML framework.
Paul Delestrac, Lionel Torres, David Novo
DSD1