Alex Skaletsky

dblp:69/4807 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
3since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 4 · 2 first-author · 3 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2026 GTPin: Enhancing Intel GPU Profiling With High-Level Binary Instrumentation
abstract
As GPUs are increasingly used for high-performance and data-intensive applications, optimizing workloads and memory behavior remains a significant challenge due to complex memory hierarchies, high parallelism, and limited runtime observability. Binary instrumentation has emerged as a key technique for analyzing and profiling GPU workloads under these constraints. This paper introduces new technology within GTPin - the binary instrumentation framework for Intel GPUs - adding support for high-level instrumentation alongside the previous low-level ISA insertion approach. The new mechanism enables developers to write analysis routines in a high-level programming language (OpenCL C) and dynamically invoke them at any instruction during kernel execution. This capability enables the implementation of sophisticated profiling tools for conducting advanced on-the-fly analyses and studies on code running on Intel GPUs, such as cache modeling, memory pattern inspection, race detection, and more, while maintaining reasonably low overhead on real hardware.In this work, we describe the design of the technology, present usage examples, provide a unique overhead breakdown analysis quantifying the performance impact of both high-level and lowlevel instrumentation, and compare it with other existing GPU instrumentation frameworks. This comparison shows that even with high-level instrumentation, on average GTPin has lower overhead. These capabilities bring GTPin to parity with leading GPU binary instrumentation frameworks, thus allowing the community to investigate, research, and analyze all three major GPU vendors in a uniform way.
Konstantin Levit-Gurevich, Alex Skaletsky, Ortal Geller, Moshe Goren, Rawan Ziadat, Rami Burstein, Tal Oved
ISPASS2
2022 Profiling Intel Graphics Architecture with Long Instruction Traces
abstract
In the process of developing software and hardware, profiling workloads is critical. Binary Instrumentation Technology plays a key role in this task for both x86 architecture and Intel Graphics Processing Units. The GTPin framework is the first tool that allows the profiling of graphics and compute kernels running on Intel GPUs. However, GTPin capabilities are less flexible than x86 profiling tools. In this paper, we introduce the concept of “gLIT” – Long Instruction Trace for Intel GPUs. Generated on real hardware, gLIT can be replayed on a simulator or an emulator running on the CPU device, and thus, can be easily profiled and analyzed “on the fly” with analysis tools of any complexity. Since the graphics devices are extremely parallel, the gLIT trace is, by definition, a multi-threaded trace, reflecting a kernel concurrently running hundreds of hardware threads. The ability to thoroughly profile and analyze workloads is critical for improving hardware and software readiness and creates new possibilities for academic research on Intel graphics devices.
Konstantin Levit-Gurevich, Alex Skaletsky, Michael Berezalsky, Yulia Kuznetcova, Hila Yakov
ISPASS2
2022 Flexible Binary Instrumentation Framework to Profile Code Running on Intel GPUs
abstract
Functional and performance profiling of workloads is critical in developing software and hardware. Binary Instrumentation Technology has played a key role in this task for many years in the world of x86 architecture. However, such capabilities have not been available until recently for graphics devices, especially in the Intel Graphics Processing Unit world. The GTPin framework is the only tool that supports profiling graphics and GP-GPU kernels running on extremely parallel Intel GPU devices. GTPin supports a wide range of capabilities for software and hardware developers. With GTPin, you can profile real-world graphics and compute applications at a level of performance close to real hardware. Such an ability is critical in accelerating hardware and software readiness.
Alex Skaletsky, Konstantin Levit-Gurevich, Michael Berezalsky, Yulia Kuznetcova, Hila Yakov
ISPASS1
2010 Dynamic program analysis of Microsoft Windows applications
abstract
Software instrumentation is a powerful and flexible technique for analyzing the dynamic behavior of programs. By inserting extra code in an application, it is possible to study the performance and correctness of programs and systems. Pin is a software system that performs run-time binary instrumentation of unmodified applications. Pin provides an API for writing custom instrumentation, enabling its use in a wide variety of performance analysis tasks such as workload characterization, program tracing, cache modeling, and simulation. Most of the prior work on instrumentation systems has focused on executing Unix applications, despite the ubiquity and importance of Windows applications. This paper identifies the Windows-specific obstacles for implementing a process-level instrumentation system, describes a comprehensive, robust solution, and discusses some of the alternatives. The challenges lie in managing the kernel/application transitions, injecting the runtime agent into the process, and isolating the instrumentation from the application. We examine Pin's overhead on typical Windows applications being instrumented with simple tools up to commercial program analysis products. The biggest factor affecting performance is the type of analysis performed by the tool. While the proprietary nature of Windows makes measurement and analysis difficult, Pin opens the door to understanding program behavior.
Alex Skaletsky, Tevi Devor, Nadav Chachmon, Robert S. Cohn, Kim M. Hazelwood, Vladimir Vladimirov, Moshe Bach
ISPASS1
2003 IA-32 Execution Layer: a two-phase dynamic translator designed to support IA-32 applications on Itanium-based systems
abstract
IA-32 execution layer (IA-32 EL) is a new technology that executes IA-32 applications on Intel Itanium processor family systems. Currently, support for IA-32 applications on Itanium-based platforms is achieved using hardware circuitry on the Itanium processors. This capability will be enhanced with IA-32 EL - software that will ship with Itanium-based operating systems and will convert IA-32 instructions into Itanium instructions via dynamic translation. In this paper, we describe aspects of the IA-32 execution layer technology, including the general two-phase translation architecture and the usage of a single translator for multiple operating systems. The paper provides details of some of the technical challenges such as precise exception, emulation of FP, MMX, and Intel streaming SIMD extension instructions, and misalignment handling. Finally, the paper presents some performance results.
Leonid Baraz, Tevi Devor, Orna Etzion, Shalom Goldenberg, Alex Skaletsky, Yigel Zemach
MICRO5