VLDB 2026 Research / reviewers in the wild / expert
Nadharm Dhiantravan
dblp:360/3661
· DBLP profile ↗
2ranked-venue papers
0as first author
2since 2021 · last 2025
0009-0006-1869-9472ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Software engineering, system software, and programming languages
2 papers |
Runtime systems and virtual machines · 46% Compilers and program optimization · 40% Operating systems · 14% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Parallel and multicore computing · 100% |
Topics — the 6 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Runtime systems and virtual machines
binary translation |
0.9 | 1 | 2025 | Virtualization So Light, it Floats! Accelerating Floating Point Virtualization · HPDC 2025 |
Compilers and program optimization › parallelization
nested parallelism |
0.8 | 1 | 2024 | Compiling Loop-Based Nested Parallelism for Irregular Workloads · ASPLOS (2) 2024 |
Parallel and multicore computing › parallel programming models › task parallelism
fork-join parallelism |
0.2 | 1 | 2024 | Compiling Loop-Based Nested Parallelism for Irregular Workloads · ASPLOS (2) 2024 |
Parallel and multicore computing › task partitioning
granularity control |
0.2 | 1 | 2024 | Compiling Loop-Based Nested Parallelism for Irregular Workloads · ASPLOS (2) 2024 |
Parallel and multicore computing
parallel programming models |
0.2 | 1 | 2024 | Compiling Loop-Based Nested Parallelism for Irregular Workloads · ASPLOS (2) 2024 |
Parallel and multicore computing › parallel programming runtimes
runtime systems and scheduling |
0.2 | 1 | 2024 | Compiling Loop-Based Nested Parallelism for Irregular Workloads · ASPLOS (2) 2024 |
Methods — techniques the papers use, named apart from their topics
heartbeat scheduling · 1.5trap short-circuiting · 0.9kernel bypass · 0.9instruction sequence emulation · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Virtualization So Light, it Floats! Accelerating Floating Point VirtualizationabstractFloating point virtualization enables unmodified application binaries to utilize alternative arithmetic systems such as MPFR without code changes, but its performance overhead is a barrier to adoption. The existing trap-and-emulate model suffers from a significant virtualization bottleneck using general-purpose signal delivery mechanisms which take thousands of cycles. We introduce three techniques to reduce virtualization overhead. Trap short-circuiting bypasses general-purpose signal delivery for an 8x reduction in trap delegation overhead. Instruction sequence emulation amortizes trap costs by emulating multiple instructions per trap, achieving up to 32x reduction in trap frequency. Finally, kernel-bypass for correctness instrumentation eliminates traps and signals for correctness and reduces related overheads substantially. Our implementation within the FPVM system on x64/Linux demonstrates a 10x reduction in per-instruction overhead which, compared to the lower bound performance set by the alternative arithmetic system, drops virtualization overhead from up to 20x to 1.65x. This is for the alternative arithmetic system that is the worst case for virtualization overheads. More expensive systems, like MPFR, fare even better. Nick Wanninger, Nadharm Dhiantravan, Peter A. Dinda |
HPDC | 2 |
| 2024 | Compiling Loop-Based Nested Parallelism for Irregular WorkloadsabstractModern programming languages offer special syntax and semantics for logical fork-join parallelism in the form of parallel loops, allowing them to be nested, e.g., a parallel loop within another parallel loop. This expressiveness comes at a price, however: on modern multicore systems, realizing logical parallelism results in overheads due to the creation and management of parallel tasks, which can wipe out the benefits of parallelism. Today, we expect application programmers to cope with it by manually tuning and optimizing their code. Such tuning requires programmers to reason about architectural factors hidden behind layers of software abstractions, such as task scheduling and load balancing. Managing these factors is particularly challenging when workloads are irregular because their performance is input-sensitive. This paper presents HBC, the first compiler that translates C/C++ programs with high-level, fork-join constructs (e.g., OpenMP) to binaries capable of automatically controlling the cost of parallelism and dealing with irregular, input-sensitive workloads. The basis of our approach is Heartbeat Scheduling, a recent proposal for automatic granularity control, which is backed by formal guarantees on performance. HBC binaries outperform OpenMP binaries for workloads for which even entirely manual solutions struggle to find the right balance between parallelism and its costs. Yian Su, Mike Rainey, Nick Wanninger, Nadharm Dhiantravan, Jasper Liang, Umut A. Acar, Peter A. Dinda, Simone Campanoni |
ASPLOS (2) | 4 |