EDBT 2026 Demo / reviewers in the wild / expert
Yuri Dotsenko
dblp:89/3828
· DBLP profile ↗
7ranked-venue papers
2as first author
0since 2021 · last 2011
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 2 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
GPUs and heterogeneous computing · 31% High-performance computing · 31% Parallel and multicore computing · 26% |
Topics — the 9 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
High-performance computing › performance optimization
auto-tuning |
0.1 | 1 | 2011 | Auto-tuning of fast fourier transform on graphics processors · PPoPP 2011 |
GPUs and heterogeneous computing
GPU kernel optimization |
0.1 | 1 | 2011 | Auto-tuning of fast fourier transform on graphics processors · PPoPP 2011 |
Parallel and multicore computing › parallel computing
parallel programming languages |
0.1 | 2 | 2005 | An evaluation of global address space languages: co-array fortran and unified parallel C · PPoPP 2005 An emerging co-array fortran compiler · PPoPP 2003 |
High-performance computing
fast fourier transform |
0.1 | 1 | 2008 | High performance discrete Fourier transforms on graphics processors · SC 2008 |
GPUs and heterogeneous computing
GPU computing |
0.1 | 1 | 2008 | High performance discrete Fourier transforms on graphics processors · SC 2008 |
Memory systems › memory hierarchy
memory hierarchy optimization |
0.1 | 1 | 2008 | High performance discrete Fourier transforms on graphics processors · SC 2008 |
Parallel and multicore computing › parallel computing › parallel programming languages
unified parallel c |
0.1 | 1 | 2005 | An evaluation of global address space languages: co-array fortran and unified parallel C · PPoPP 2005 |
Parallel and multicore computing › parallel programming models › distributed memory programming models
global address space |
0.0 | 1 | 2003 | An emerging co-array fortran compiler · PPoPP 2003 |
Parallel and multicore computing
parallel programming models |
0.0 | 1 | 2003 | An emerging co-array fortran compiler · PPoPP 2003 |
Methods — techniques the papers use, named apart from their topics
search space pruning · 0.1profiling · 0.1auto-tuning · 0.1stockham formulation · 0.1modular arithmetic · 0.1mixed radix FFT · 0.1bluestein's algorithm · 0.1source-to-source translation · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2011 | Auto-tuning of fast fourier transform on graphics processorsabstractWe present an auto-tuning framework for FFTs on graphics processors (GPUs). Due to complex design of the memory and compute subsystems on GPUs, the performance of FFT kernels over the range of possible input parameters can vary widely. We generate several variants for each component of the FFT kernel that, for different cases, are likely to perform well. Our auto-tuner composes variants to generate kernels and selects the best ones. We present heuristics to prune the search space and profile only a small fraction of all possible kernels. We compose optimized kernels to improve the performance of larger FFT computations. We implement the system using the NVIDIA CUDA API and compare its performance to the state-of-the-art FFT libraries. On a range of NVIDIA GPUs and input sizes, our auto-tuned FFTs outperform the NVIDIA CUFFT 3.0 library by up to 38x and deliver up to 3x higher performance compared to a manually-tuned FFT. Yuri Dotsenko, Sara S. Baghsorkhi, Brandon Lloyd, Naga K. Govindaraju |
PPoPP | 1 |
| 2008 | Fast scan algorithms on graphics processorsabstractScan and segmented scan are important data-parallel primitives for a wide range of applications. We present fast, work-efficient algorithms for these primitives on graphics processing units (GPUs). We use novel data representations that map well to the GPU architecture. Our algorithms exploit shared memory to improve memory performance. We further improve the performance of our algorithms by eliminating shared-memory bank conflicts and reducing the overheads in prior shared-memory GPU algorithms. Furthermore, our algorithms are designed to work well on general data sets, including segmented arrays with arbitrary segment lengths. We also present optimizations to improve the performance of segmented scans based on the segment lengths. We implemented our algorithms on a PC with an NVIDIA GeForce 8800 GPU and compared our results with prior GPU-based algorithms. Our results indicate up to 10x higher performance over prior algorithms on input sequences with millions of elements. Yuri Dotsenko, Naga K. Govindaraju, Peter-Pike J. Sloan, Charles Boyd, John Manferdelli |
ICS | 1 |
| 2008 | High performance discrete Fourier transforms on graphics processorsabstractWe present novel algorithms for computing discrete Fourier transforms with high performance on GPUs. We present hierarchical, mixed radix FFT algorithms for both power-of-two and non-power-of-two sizes. Our hierarchical FFT algorithms efficiently exploit shared memory on GPUs using a Stockham formulation. We reduce the memory transpose overheads in hierarchical algorithms by combining the transposes into a block-based multi-FFT algorithm. For non-power-of-two sizes, we use a combination of mixed radix FFTs of small primes and Bluestein's algorithm. We use modular arithmetic in Bluestein's algorithm to improve the accuracy. We implemented our algorithms using the NVIDIA CUDA API and compared their performance with NVIDIA's CUFFT library and an optimized CPU-implementation (Intel's MKL) on a high-end quad-core CPU. On an NVIDIA GPU, we obtained performance of up to 300 GFlops, with typical performance improvements of 2–4× over CUFFT and 8–40× improvement over MKL for large sizes. Naga K. Govindaraju, Brandon Lloyd, Yuri Dotsenko, Burton Smith, John Manferdelli |
SC | 3 |
| 2007 | Scalability analysis of SPMD codes using expectationsabstractWe present a new technique for identifying scalability bottlenecks in executions of single-program, multiple-data (SPMD) parallel programs, quantifying their impact on performance, and associating this information with the program source code. Our performance analysis strategy involves three steps. First, we collect call path profiles for two or more executions on different numbers of processors. Second, we use our expectations about how the performance of executions should differ, e.g., linear speedup for strong scaling or constant execution time for weak scaling, to automatically compute the scalability of costs incurred at each point in a program's execution. Third, with the aid of an interactive browser, an application developer can explore a program's performance in a top-down fashion, see the contexts in which poor scaling behavior arises, and understand exactly how much each scalability bottleneck dilates execution time. Our analysis technique is independent of the parallel programming model. We describe our experiences applying our technique to analyze parallel programs written in Co-array Fortran and Unified Parallel C, as well as message-passing programs based on MPI. Cristian Coarfa, John M. Mellor-Crummey, Nathan Froyd, Yuri Dotsenko |
ICS | 4 |
| 2006 | Experiences with Sweep3D implementations in Co-array Fortran
Cristian Coarfa, Yuri Dotsenko, John M. Mellor-Crummey |
J. Supercomput. | 2 |
| 2005 | An evaluation of global address space languages: co-array fortran and unified parallel CabstractCo-array Fortran (CAF) and Unified Parallel C (UPC) are two emerging languages for single-program, multiple-data global address space programming. These languages boost programmer productivity by providing shared variables for inter-process communication instead of message passing. However, the performance of these emerging languages still has room for improvement. In this paper, we study the performance of variants of the NAS MG, CG, SP, and BT benchmarks on several modern architectures to identify challenges that must be met to deliver top performance. We compare CAF and UPC variants of these programs with the original Fortran+MPI code. Today, CAF and UPC programs deliver scalable performance on clusters only when written to use bulk communication. However, our experiments uncovered some significant performance bottlenecks of UPC codes on all platforms. We account for the root causes limiting UPC performance such as the synchronization model, the communication efficiency of strided data, and source-to-source translation issues. We show that they can be remedied with language extensions, new synchronization constructs, and, finally, adequate optimizations by the back-end C compilers. Cristian Coarfa, Yuri Dotsenko, John M. Mellor-Crummey, François Cantonnet, Tarek A. El-Ghazawi, Ashrujit Mohanti, Yiyi Yao, Daniel G. Chavarría-Miranda |
PPoPP | 2 |
| 2003 | An emerging co-array fortran compilerabstractNo abstract available. Cristian Coarfa, Yuri Dotsenko |
PPoPP | 2 |