Artem Chikin

dblp:234/2261 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
0since 2021 · last 2019
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
GPUs and heterogeneous computing · 87% Parallel and multicore computing · 13%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Compilers and program optimization
loop transformation
0.412019
Memory-access-aware Safety and Profitability Analysis for Transformation of Accelerator-bound OpenMP Loops · ACM Trans. Archit. Code Optim. 2019
Compilers and program optimization › memory optimization
memory access optimization
0.412019
Memory-access-aware Safety and Profitability Analysis for Transformation of Accelerator-bound OpenMP Loops · ACM Trans. Archit. Code Optim. 2019
GPUs and heterogeneous computing › CPU-GPU heterogeneous computing
GPU offloading
0.412019
Memory-access-aware Safety and Profitability Analysis for Transformation of Accelerator-bound OpenMP Loops · ACM Trans. Archit. Code Optim. 2019
GPUs and heterogeneous computing › GPU memory access
memory coalescing
0.412019
Memory-access-aware Safety and Profitability Analysis for Transformation of Accelerator-bound OpenMP Loops · ACM Trans. Archit. Code Optim. 2019
Parallel and multicore computing › parallel programming models › directive-based programming
OpenMP
0.112019
Memory-access-aware Safety and Profitability Analysis for Transformation of Accelerator-bound OpenMP Loops · ACM Trans. Archit. Code Optim. 2019

Methods — techniques the papers use, named apart from their topics

static analysis · 0.8iteration point difference analysis · 0.8
YearPublicationVenuePosition
2019 Memory-access-aware Safety and Profitability Analysis for Transformation of Accelerator-bound OpenMP Loops
abstract
Iteration Point Difference Analysis is a new static analysis framework that can be used to determine the memory coalescing characteristics of parallel loops that target GPU offloading and to ascertain safety and profitability of loop transformations with the goal of improving their memory access characteristics. This analysis can propagate definitions through control flow, works for non-affine expressions, and is capable of analyzing expressions that reference conditionally defined values. This analysis framework enables safe and profitable loop transformations. Experimental results demonstrate potential for dramatic performance improvements. GPU kernel execution time across the Polybench suite is improved by up to 25.5× on an Nvidia P100 with benchmark overall improvement of up to 3.2×. An opportunity detected in a SPEC ACCEL benchmark yields kernel speedup of 86.5× with a benchmark improvement of 3.3×. This work also demonstrates how architecture-aware compilers improve code portability and reduce programmer effort.
Artem Chikin, Taylor Lloyd, José Nelson Amaral, Ettore Tiotto
ACM Trans. Archit. Code Optim.1
2018 Automated GPU Grid Geometry Selection for OPENMP Kernels
abstract
Modern supercomputers are increasingly using GPUs to improve performance per watt. Generating GPU code for target regions in openMP 4.0, or later versions, requires the selection of grid geometry to execute the GPU kernel. Existing industrial-strength compilers use a simple heuristic with arbitrary numbers that are constant for all kernels. After characterizing the relationship between region features, grid geometry and performance, we built a machine-learning model that successfully predicts a suitable geometry for such kernels and results in a performance improvement with a geometric mean of 5% across the benchmarks studied. However, this prediction is impractical because the overhead of the predictor is too high. A careful study of the results of the predictor allowed for the development of a practical low-overhead heuristic that resulted in a performance improvement of up to 7 times with a geometric mean of 25.9%. This paper describes the methodology to build the machine-learning model, and the practical low-overhead heuristic that can be used in industry-strong compilers.
Taylor Lloyd, Artem Chikin, Sanket Kedia, Dhruv Jain, José Nelson Amaral
SBAC-PAD2