VLDB 2026 Research / reviewers in the wild / expert
Chenle Yu
dblp:266/2453
· DBLP profile ↗
3ranked-venue papers
3as first author
2since 2021 · last 2024
0000-0002-1802-8680ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 2 · 2 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Parallel and multicore computing · 87% High-performance computing · 13% |
Topics — the 1 heaviest of 3, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Parallel and multicore computing › parallel programming models
task parallelism |
0.7 | 1 | 2023 | Taskgraph: A Low Contention OpenMP Tasking Framework · IEEE Trans. Parallel Distributed Syst. 2023 |
Methods — techniques the papers use, named apart from their topics
task dependency graph · 0.7record-and-replay execution · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Enhancing Heterogeneous Computing Through OpenMP and GPU GraphabstractModern computing platforms are increasingly heterogeneous, most of them include accelerators such as GPU. OpenMP as the de-facto standard to parallelize CPU applications, incorporates target construct allowing users to offload work onto such accelerators, to leverage the computing capability provided by these accelerators while maintaining a simple, directive based programming style. However, due to architectural differences between CPUs and GPUs, OpenMP is not always able to exploit the latest computing optimizations available on GPUs. Chenle Yu, Sara Royuela, Eduardo Quiñones |
ICPP | 1 |
| 2023 | Taskgraph: A Low Contention OpenMP Tasking FrameworkabstractOpenMP is the de-facto standard for shared memory systems in High-Performance Computing (HPC). It includes a tasking model that offers a high-level of abstraction to effectively exploit structured (loop-based) and highly dynamic unstructured (task-based) parallelism in an easy and flexible way. Unfortunately, the run-time overheads introduced to manage tasks are (very) high in most common OpenMP frameworks (e.g., GCC, LLVM), which defeats the potential benefits of the tasking model, and makes it suitable for coarse-grained tasks only. This paper presentstaskgraph, a framework that uses a task dependency graph (TDG) to represent a region of code implemented with OpenMP tasks in order to reduce the run-time overheads associated with the management of tasks, i.e., contention and parallel orchestration, including task creation and synchronization. The TDG avoids the overheads related to the resolution of task dependencies and greatly reduces those deriving from accesses to shared resources. Moreover, the taskgraph framework introduces in OpenMP therecord-and-replayexecution model that accelerates the taskgraph region from its second execution. Overall, the multiple optimizations presented in this paper allow exploiting fine-grained OpenMP tasks to cope with the trend in current applications pointing to leverage massive on-node parallelism, fine-grained and dynamic scheduling paradigms. The framework is implemented on LLVM 15.0. Results show that the taskgraph implementation outperforms the vanilla OpenMP system in terms of performance and scalability, for all structured and unstructured parallelism, and considering coarse and fine grained tasks. Furthermore, the proposed framework makes the tasking model a competitive alternative to the OpenMP thread model in most cases. Chenle Yu, Sara Royuela, Eduardo Quiñones |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2020 | OpenMP to CUDA graphs: a compiler-based transformation to enhance the programmability of NVIDIA devicesabstractHeterogeneous computing is increasingly being used in a diversity of computing systems, ranging from HPC to the real-time embedded domain, to cope with the performance requirements. Due to the variety of accelerators, e.g., FPGAs, GPUs, the use of high-level parallel programming models is desirable to exploit the performance capabilities of them, while maintaining an adequate productivity level. In that regard, OpenMP is a well-known high-level programming model that incorporates powerful task and accelerator models capable of efficiently exploiting structured and unstructured parallelism in heterogeneous computing. This paper presents a novel compiler transformation technique that automatically transforms OpenMP code into CUDA graphs, combining the benefits of programmability of a high-level programming model such as OpenMP, with the performance benefits of a low-level programming model such as CUDA. Evaluations have been performed on two NVIDIA GPUs from the HPC and embedded domains, i.e., the V100 and the Jetson AGX respectively. Chenle Yu, Sara Royuela, Eduardo Quiñones |
SCOPES | 1 |