VLDB 2026 Research / reviewers in the wild / expert
Dani Voitsechov
dblp:148/9791
· DBLP profile ↗
4ranked-venue papers
4as first author
0since 2021 · last 2018
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 4 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
GPUs and heterogeneous computing · 37% Processor architecture and microarchitecture · 28% Parallel and multicore computing · 11% | |
| Software engineering, system software, and programming languages
1 paper |
Compilers and program optimization · 100% |
Topics — the 13 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
GPUs and heterogeneous computing
GPU architecture |
0.7 | 3 | 2018 | Software-Directed Techniques for Improved GPU Register File Utilization · ACM Trans. Archit. Code Optim. 2018 Control flow coalescing on a hybrid dataflow/von Neumann GPGPU · MICRO 2015 Single-graph multiple flows: Energy efficient design alternative for GPGPUs · ISCA 2014 |
Compilers and program optimization
register allocation |
0.3 | 1 | 2018 | Software-Directed Techniques for Improved GPU Register File Utilization · ACM Trans. Archit. Code Optim. 2018 |
Reconfigurable computing and FPGAs
coarse-grained reconfigurable architecture |
0.3 | 1 | 2018 | Inter-Thread Communication in Multithreaded, Reconfigurable Coarse-Grain Arrays · MICRO 2018 |
GPUs and heterogeneous computing
GPU programming |
0.3 | 1 | 2018 | Inter-Thread Communication in Multithreaded, Reconfigurable Coarse-Grain Arrays · MICRO 2018 |
Parallel and multicore computing › parallel computing › parallel communication
inter-thread communication |
0.3 | 1 | 2018 | Inter-Thread Communication in Multithreaded, Reconfigurable Coarse-Grain Arrays · MICRO 2018 |
Electronic design automation › high-level synthesis › resource binding
register allocation |
0.3 | 1 | 2018 | Software-Directed Techniques for Improved GPU Register File Utilization · ACM Trans. Archit. Code Optim. 2018 |
Processor architecture and microarchitecture
register file |
0.3 | 1 | 2018 | Software-Directed Techniques for Improved GPU Register File Utilization · ACM Trans. Archit. Code Optim. 2018 |
GPUs and heterogeneous computing
control flow divergence |
0.2 | 1 | 2015 | Control flow coalescing on a hybrid dataflow/von Neumann GPGPU · MICRO 2015 |
Processor architecture and microarchitecture
dataflow architecture |
0.2 | 1 | 2015 | Control flow coalescing on a hybrid dataflow/von Neumann GPGPU · MICRO 2015 |
Processor architecture and microarchitecture › dataflow architecture
hybrid dataflow/von neumann |
0.2 | 1 | 2015 | Control flow coalescing on a hybrid dataflow/von Neumann GPGPU · MICRO 2015 |
Processor architecture and microarchitecture › dataflow architecture › dataflow machine
dynamic dataflow |
0.2 | 1 | 2014 | Single-graph multiple flows: Energy efficient design alternative for GPGPUs · ISCA 2014 |
Energy-efficient computing
energy-efficient system design |
0.2 | 1 | 2014 | Single-graph multiple flows: Energy efficient design alternative for GPGPUs · ISCA 2014 |
Parallel and multicore computing
thread-level parallelism |
0.1 | 1 | 2014 | Single-graph multiple flows: Energy efficient design alternative for GPGPUs · ISCA 2014 |
Methods — techniques the papers use, named apart from their topics
simulation · 0.7compiler-driven optimization · 0.7dynamic scheduling · 0.2dataflow graph mapping · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2018 | Inter-Thread Communication in Multithreaded, Reconfigurable Coarse-Grain ArraysabstractTraditional von Neumann GPGPUs only allow threads to communicate through memory on a group-to-group basis. In this model, a group of producer threads writes intermediate values to memory, which are read by a group of consumer threads after a barrier synchronization. To alleviate the memory bandwidth imposed by this method of communication, GPGPUs provide a small scratchpad memory that prevents intermediate values from overloading DRAM bandwidth. In this paper we introduce direct inter-thread communications for massively multithreaded CGRAs, where intermediate values are communicated directly through the compute fabric on a point-to-point basis. This method avoids the need to write values to memory, eliminates the need for a dedicated scratchpad, and avoids workgroup global barriers. We introduce our proposed extensions to the programming model (CUDA) and execution model, as well as the hardware primitives that facilitate the communication. Our simulations of Rodinia benchmarks running on the new system show that direct inter-thread communication provides an average speedup of 2.8x (10.3x max) and reduces system power by an average of 5x (22x max), when compared to an equivalent Nvidia GPGPU. Dani Voitsechov, Oron Port, Yoav Etsion |
MICRO | 1 |
| 2018 | Software-Directed Techniques for Improved GPU Register File UtilizationabstractThroughput architectures such as GPUs require substantial hardware resources to hold the state of a massive number of simultaneously executing threads. While GPU register files are already enormous, reaching capacities of 256KB per streaming multiprocessor (SM), we find that nearly half of real-world applications we examined are register-bound and would benefit from a larger register file to enable more concurrent threads. This article seeks to increase the thread occupancy and improve performance of these register-bound applications by making more efficient use of the existing register file capacity. Our first technique eagerly deallocates register resources during execution. We show that releasing register resources based on value liveness as proposed in prior states of the art leads to unreliable performance and undue design complexity. To address these deficiencies, our article presents a novel compiler-driven approach that identifies and exploits last use of a register name (instead of the value contained within) to eagerly release register resources. Furthermore, while previous works have leveraged “scalar” and “narrow” operand properties of a program for various optimizations, their impact on thread occupancy has been relatively unexplored. Our article evaluates the effectiveness of these techniques in improving thread occupancy and demonstrates that while any one approach may fail to free very many registers, together they synergistically free enough registers to launch additional parallel work. An in-depth evaluation on a large suite of applications shows that just our early register technique outperforms previous work on dynamic register allocation, and together these approaches, on average, provide 12% performance speedup (23% higher thread occupancy) on register bound applications not already saturating other GPU resources. Dani Voitsechov, Arslan Zulfiqar, Mark Stephenson, Mark Gebhart, Stephen W. Keckler |
ACM Trans. Archit. Code Optim. | 1 |
| 2015 | Control flow coalescing on a hybrid dataflow/von Neumann GPGPUabstractWe propose the hybrid dataflow/von Neumann vector graph instruction word (VGIW) architecture. This data-parallel architecture concurrently executes each basic block's dataflow graph (graph instruction word) for a vector of threads, and schedules the different basic blocks based on von Neumann control flow semantics. The VGIW processor dynamically coalesces all threads that need to execute a specific basic block into a thread vector and, when the block is scheduled, executes the entire thread vector concurrently. The proposed control flow coalescing model enables the VGIW architecture to overcome the control flow divergence problem, which greatly impedes the performance and power efficiency of data-parallel architectures. Furthermore, using von Neumann control flow semantics enables the VGIW architecture to overcome the limitations of the recently proposed single-graph multiple-flows (SGMF) dataflow GPGPU, which is greatly constrained in the size of the kernels it can execute. Our evaluation shows that VGIW can achieve an average speedup of 3× (up to 11×) over an NVIDIA GPGPU, while providing an average 1.75× better energy efficiency (up to 7×). Dani Voitsechov, Yoav Etsion |
MICRO | 1 |
| 2014 | Single-graph multiple flows: Energy efficient design alternative for GPGPUsabstractWe present the single-graph multiple-flows (SGMF) architecture that combines coarse-grain reconfigurable computing with dynamic dataflow to deliver massive thread-level parallelism. The CUDA-compatible SGMF architecture is positioned as an energy efficient design alternative for GPGPUs. The architecture maps a compute kernel, represented as a dataflow graph, onto a coarse-grain reconfigurable fabric composed of a grid of interconnected functional units. Each unit dynamically schedules instances of the same static instruction originating from different CUDA threads. The dynamically scheduled functional units enable streaming the data of multiple threads (or graph flows, in SGMF parlance) through the grid. The combination of statically mapped instructions and direct communication between functional units obviate the need for afull instruction pipeline and a centralized register file, whose energy overheads burden GPGPUs. We show that the SGMF architecture delivers performance comparable to that of contemporary GPGPUs while consuming ~57% less energy on average. Dani Voitsechov, Yoav Etsion |
ISCA | 1 |