Albert Sidelnik

dblp:03/4794 · DBLP profile ↗
← Back
7ranked-venue papers
1as first author
1since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
GPUs and heterogeneous computing · 52% Parallel and multicore computing · 32% High-performance computing · 14%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%

Topics — the 15 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
GPUs and heterogeneous computing › GPU programming
dynamic parallelism
0.322016
Dynamic thread block launch: a lightweight execution mechanism to support irregular applications on GPUs · ISCA 2015
LaPerm: Locality Aware Scheduler for Dynamic Parallelism on GPUs · ISCA 2016
Compilers and program optimization
code generation
0.212016
Designing a Tunable Nested Data-Parallel Programming System · ACM Trans. Archit. Code Optim. 2016
High-performance computing › performance optimization
auto-tuning
0.212016
Designing a Tunable Nested Data-Parallel Programming System · ACM Trans. Archit. Code Optim. 2016
Parallel and multicore computing › parallel scheduling
locality-aware scheduling
0.212016
LaPerm: Locality Aware Scheduler for Dynamic Parallelism on GPUs · ISCA 2016
Parallel and multicore computing › data parallelism
nested data parallelism
0.212016
Designing a Tunable Nested Data-Parallel Programming System · ACM Trans. Archit. Code Optim. 2016
GPUs and heterogeneous computing › GPU scheduling
thread block scheduling
0.212016
LaPerm: Locality Aware Scheduler for Dynamic Parallelism on GPUs · ISCA 2016
Parallel and multicore computing › task scheduling
task graph scheduling
0.212024
CUDASTF: Bridging the Gap Between CUDA and Task Parallelism · SC 2024
GPUs and heterogeneous computing › GPU architecture
GPU execution model
0.212015
Dynamic thread block launch: a lightweight execution mechanism to support irregular applications on GPUs · ISCA 2015
Parallel and multicore computing
parallel programming models
0.212015
A collection-oriented programming model for performance portability · PPoPP 2015
High-performance computing › performance engineering
performance portability
0.212015
A collection-oriented programming model for performance portability · PPoPP 2015
Memory systems
data locality
0.112016
LaPerm: Locality Aware Scheduler for Dynamic Parallelism on GPUs · ISCA 2016
GPUs and heterogeneous computing › GPU programming
GPU code generation
0.112016
Designing a Tunable Nested Data-Parallel Programming System · ACM Trans. Archit. Code Optim. 2016
GPUs and heterogeneous computing › GPU memory
GPU memory hierarchy
0.112016
LaPerm: Locality Aware Scheduler for Dynamic Parallelism on GPUs · ISCA 2016
Parallel and multicore computing
graph processing
0.112015
Dynamic thread block launch: a lightweight execution mechanism to support irregular applications on GPUs · ISCA 2015
Parallel and multicore computing › parallel computing › parallel applications
irregular applications
0.112015
Dynamic thread block launch: a lightweight execution mechanism to support irregular applications on GPUs · ISCA 2015

Methods — techniques the papers use, named apart from their topics

auto-tuning · 0.7algorithmic skeletons · 0.5cycle-level simulation · 0.2nested kernel launching · 0.2nested data collections · 0.2code generation · 0.2
YearPublicationVenuePosition
2024 CUDASTF: Bridging the Gap Between CUDA and Task Parallelism
abstract
Organizing computation as asynchronous tasks with data-driven dependencies is a simple and efficient model for single- and multi-GPU programs. Sequential Task Flow (STF) is such a model that derives task graphs from data dependencies.We propose CUDASTF, a C++ library that implements STF over CUDA APIs, fostering easy creation of scalable and composable algorithms. Users may easily elect to use CUDA Graphs instead of streams, which improves performance of small kernels. Structured kernels are automatically spread over multiple devices and can exercise fine-grained affinity control. Implementationwise, CUDASTF makes a compelling argument for an event-based approach to asynchronous parallel libraries.We obtain up to a 1.8 x improvement over the cuSolverMg library on Cholesky decomposition. On a small weather simulation task we demonstrate near-optimal scalability of our multiGPU kernels; also, on a single GPU, CUDA Graphs improve performance by up to $30 \%$. Finally, we were able to author the first implementation of the CKKS fully homomorphic encryption scheme over multiple devices.
Cédric Augonnet, Andrei Alexandrescu, Albert Sidelnik, Michael Garland
SC3
2016 LaPerm: Locality Aware Scheduler for Dynamic Parallelism on GPUs
abstract
Recent developments in GPU execution models and architectures have introduced dynamic parallelism to facilitate the execution of irregular applications where control flow and memory behavior can be unstructured, time-varying, and hierarchical. The changes brought about by this extension to the traditional bulk synchronous parallel (BSP) model also creates new challenges in exploiting the current GPU memory hierarchy. One of the major challenges is that the reference locality that exists between the parent and child thread blocks (TBs) created during dynamic nested kernel and thread block launches cannot be fully leveraged using the current TB scheduling strategies. These strategies were designed for the current implementations of the BSP model but fall short when dynamic parallelism is introduced since they are oblivious to the hierarchical reference locality. We propose LaPerm, a new locality-aware TB scheduler that exploits such parent-child locality, both spatial and temporal. LaPerm adopts three different scheduling decisions to i) prioritize the execution of the child TBs, ii) bind them to the stream multiprocessors (SMXs) occupied by their parents TBs, and iii) maintain workload balance across compute units. Experiments with a set of irregular CUDA applications executed on a cycle-level simulator employing dynamic parallelism demonstrate that LaPerm is able to achieve an average of 27% performance improvement over the baseline round-robin TB scheduler commonly used in modern GPUs.
Jin Wang 0010, Norman Rubin, Albert Sidelnik, Sudhakar Yalamanchili
ISCA3
2016 Designing a Tunable Nested Data-Parallel Programming System
abstract
This article describes Surge, a nested data-parallel programming system designed to simplify the porting and tuning of parallel applications to multiple target architectures. Surge decouples high-level specification of computations, expressed using a C++ programming interface, from low-level implementation details using two first-class constructs: schedules and policies. Schedules describe the valid ways in which data-parallel operators may be implemented, while policies encapsulate a set of parameters that govern platform-specific code generation. These two mechanisms are used to implement a code generation system that analyzes computations and automatically generates a search space of valid platform-specific implementations. An input and architecture-adaptive autotuning system then explores this search space to find optimized implementations. We express in Surge five real-world benchmarks from domains such as machine learning and sparse linear algebra and from the high-level specifications, Surge automatically generates CPU and GPU implementations that perform on par with or better than manually optimized versions.
Saurav Muralidharan, Michael Garland, Albert Sidelnik, Mary W. Hall
ACM Trans. Archit. Code Optim.3
2015 Locality-Driven Dynamic GPU Cache Bypassing
abstract
This paper presents novel cache optimizations for massively parallel, throughput-oriented architectures like GPUs. L1 data caches (L1 D-caches) are critical resources for providing high-bandwidth and low-latency data accesses. However, the high number of simultaneous requests from single-instruction multiple-thread (SIMT) cores makes the limited capacity of L1 D-caches a performance and energy bottleneck, especially for memory-intensive applications. We observe that the memory access streams to L1 D-caches for many applications contain a significant amount of requests with low reuse, which greatly reduce the cache efficacy. Existing GPU cache management schemes are either based on conditional/reactive solutions or hit-rate based designs specifically developed for CPU last level caches, which can limit overall performance.
Chao Li 0004, Shuaiwen Song, Hongwen Dai, Albert Sidelnik, Siva Kumar Sastry Hari, Huiyang Zhou
ICS4
2015 Dynamic thread block launch: a lightweight execution mechanism to support irregular applications on GPUs
abstract
GPUs have been proven effective for structured applications that map well to the rigid 1D-3D grid of threads in modern bulk synchronous parallel (BSP) programming languages. However, less success has been encountered in mapping data intensive irregular applications such as graph analytics, relational databases, and machine learning. Recently introduced nested device-side kernel launching functionality in the GPU is a step in the right direction, but still falls short of being able to effectively harness the GPUs performance potential.
Jin Wang 0010, Norman Rubin, Albert Sidelnik, Sudhakar Yalamanchili
ISCA3
2015 A collection-oriented programming model for performance portability
abstract
This paper describes Surge, a collection-oriented programming model that enables programmers to compose parallel computations using nested high-level data collections and operators. Surge exposes a code generation interface, decoupled from the core computation, that enables programmers and autotuners to easily generate multiple implementations of the same computation on various parallel architectures such as multi-core CPUs and GPUs. By decoupling computations from architecture-specific implementation, programmers can target multiple architectures more easily, and generate a search space that facilitates optimization and customization for specific architectures. We express in Surge four real-world benchmarks from domains such as sparse linear-algebra and machine learning and from the same performance-portable specification, generate OpenMP and CUDA C++ implementations. Surge generates efficient, scalable code which achieves up to 1.32x speedup over handcrafted, well-optimized CUDA code.
Saurav Muralidharan, Michael Garland, Bryan Catanzaro, Albert Sidelnik, Mary W. Hall
PPoPP4
2012 Performance Portability with the Chapel Language
abstract
It has been widely shown that high-throughput computing architectures such as GPUs offer large performance gains compared with their traditional low-latency counterparts for many applications. The downside to these architectures is that the current programming models present numerous challenges to the programmer: lower-level languages, loss of portability across different architectures, explicit data movement, and challenges in performance optimization. This paper presents novel methods and compiler transformations that increase programmer productivity by enabling users of the language Chapel to provide a single code implementation that the compiler can then use to target not only conventional multiprocessors, but also high-throughput and hybrid machines. Rather than resorting to different parallel libraries or annotations for a given parallel platform, this work leverages a language that has been designed from first principles to address the challenge of programming for parallelism and locality. This also has the advantage of providing portability across different parallel architectures. Finally, this work presents experimental results from the Parboil benchmark suite which demonstrate that codes written in Chapel achieve performance comparable to the original versions implemented in CUDA on both GPUs and multicore platforms.
Albert Sidelnik, Saeed Maleki, Bradford L. Chamberlain, María Jesús Garzarán, David A. Padua
IPDPS1