EDBT 2026 Demo / reviewers in the wild / expert
Farzad Khorasani
dblp:147/2060
· DBLP profile ↗
9ranked-venue papers
7as first author
0since 2021 · last 2019
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 7 first-authorSoftware engineering, systems software and programming languages · 2 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
5 papers |
GPUs and heterogeneous computing · 71% Processor architecture and microarchitecture · 11% Electronic design automation · 10% | |
| Software engineering, system software, and programming languages
1 paper |
Compilers and program optimization · 100% | |
| Theoretical computer science
1 paper |
Graph algorithms and graph theory · 100% |
Topics — the 15 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
GPUs and heterogeneous computing › GPU microarchitecture
GPU register file |
0.7 | 2 | 2019 | CORF: Coalescing Operand Register File for GPUs · ASPLOS 2019 RegMutex: Inter-Warp GPU Register Time-Sharing · ISCA 2018 |
Processor architecture and microarchitecture
register file |
0.4 | 1 | 2019 | CORF: Coalescing Operand Register File for GPUs · ASPLOS 2019 |
GPUs and heterogeneous computing
GPU kernel optimization |
0.3 | 1 | 2018 | In-Register Parameter Caching for Dynamic Neural Nets with Virtual Persistent Processor Specialization · MICRO 2018 |
GPUs and heterogeneous computing
GPU microarchitecture |
0.3 | 1 | 2018 | RegMutex: Inter-Warp GPU Register Time-Sharing · ISCA 2018 |
Electronic design automation › high-level synthesis › resource binding
register allocation |
0.3 | 1 | 2018 | RegMutex: Inter-Warp GPU Register Time-Sharing · ISCA 2018 |
GPUs and heterogeneous computing
control flow divergence |
0.2 | 1 | 2015 | Efficient warp execution in presence of divergence with collaborative context collection · MICRO 2015 |
GPUs and heterogeneous computing › GPU programming
GPU compiler optimization |
0.2 | 1 | 2015 | Efficient warp execution in presence of divergence with collaborative context collection · MICRO 2015 |
GPUs and heterogeneous computing
GPU graph processing |
0.2 | 1 | 2014 | CuSha: vertex-centric graph processing on GPUs · HPDC 2014 |
Parallel and multicore computing › graph processing
vertex-centric graph processing |
0.2 | 1 | 2014 | CuSha: vertex-centric graph processing on GPUs · HPDC 2014 |
Graph algorithms and graph theory
graph representation |
0.2 | 1 | 2014 | CuSha: vertex-centric graph processing on GPUs · HPDC 2014 |
Compilers and program optimization › register allocation
graph coloring register allocation |
0.1 | 1 | 2019 | CORF: Coalescing Operand Register File for GPUs · ASPLOS 2019 |
Compilers and program optimization
register allocation |
0.1 | 1 | 2019 | CORF: Coalescing Operand Register File for GPUs · ASPLOS 2019 |
Machine learning › Efficient and distributed learning
dynamic neural network |
0.1 | 1 | 2018 | In-Register Parameter Caching for Dynamic Neural Nets with Virtual Persistent Processor Specialization · MICRO 2018 |
Parallel and multicore computing › data parallelism
SIMD vectorization |
0.1 | 1 | 2015 | Efficient warp execution in presence of divergence with collaborative context collection · MICRO 2015 |
GPUs and heterogeneous computing › GPU performance analysis
GPU utilization |
0.1 | 1 | 2014 | CuSha: vertex-centric graph processing on GPUs · HPDC 2014 |
Methods — techniques the papers use, named apart from their topics
compiler-assisted register packing · 0.8bipartite edge frustration · 0.8virtual persistent processor specialization · 0.7CTA virtualization · 0.7concatenated windows · 0.4CUDA · 0.4register sharing · 0.3compiler-hardware co-design · 0.3collaborative context collection · 0.2code transformation · 0.2g-shards · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2019 | CORF: Coalescing Operand Register File for GPUsabstractThe Register File (RF) in GPUs is a critical structure that maintains the state for thousands of threads that support the GPU processing model. The RF organization substantially affects the overall performance and the energy efficiency of a GPU. For example, the frequent accesses to the RF consume a substantial amount of the dynamic energy, and port contention due to limited ports on operand collectors and register file banks affect performance as register operations are serialized. We present CORF, a compiler-assisted Coalescing Operand Register File which performs register coalescing by combining reads to multiple registers required by a single instruction, into a single physical read. To enable register coalescing, CORF utilizes register packing to co-locate narrow-width operands in the same physical register. CORF uses compiler hints to identify which register pairs are commonly accessed together. CORF saves dynamic energy by reducing the number of physical register file accesses, and improves performance by combining read operations, as well as by reducing pressure on the register file. To increase the coalescing opportunities, we re-architect the physical register file to allow coalescing reads across different physical registers that reside in mutually exclusive sub-banks; we call this design CORF++. The compiler analysis for register allocation for CORF++ becomes a form of graph coloring called the bipartite edge frustration problem. CORF++ reduces the dynamic energy of the RF by 17%, and improves IPC by 9%. Hodjat Asghari Esfeden, Farzad Khorasani, Hyeran Jeon, Daniel Wong 0001, Nael B. Abu-Ghazaleh |
ASPLOS | 2 |
| 2018 | RegMutex: Inter-Warp GPU Register Time-SharingabstractRegisters are the fastest and simultaneously the most expensive kind of memory available to GPU threads. Due to existence of a great number of concurrently executing threads, and the high cost of context switching mechanisms, contemporary GPUs are equipped with large register files. However, to avoid over-complicating the hardware, registers are statically assigned and exclusively dedicated to threads for the entire duration of the thread's lifetime. This decomposition takes into account the maximum number of live registers at any given point in the GPU binary although the points at which all the requested registers are used may constitute only a small fraction of the whole program. Therefore, a considerable portion of the register file remains under-utilized. In this paper, we propose a software-hardware co-mechanism named RegMutex (Register Mutual Exclusion) to share a subset of physical registers between warps during the GPU kernel execution. With RegMutex, the compiler divides the architected register set into a base register set and an extended register set. While physical registers corresponding to the base register set are statically and exclusively assigned to the warp, the hardware time-shares the remaining physical registers across warps to provision their extended register set. Therefore, the GPU programs can sustain approximately the same performance with the lower number of registers hence yielding higher performance per dollar. For programs that require a large number of registers for execution, RegMutex will enable a higher number of concurrent warps to be resident in the hardware via sharing their register allocations with each other, leading to a higher device occupancy. Since some aspects of register sharing orchestration are being offloaded to the compiler, RegMutex introduces lower hardware complexity compared to existing approaches. Our experiments show that RegMutex improves the register utilization and reduces the number of execution cycles by up to 23% for kernels demanding a high number of registers. Farzad Khorasani, Hodjat Asghari Esfeden, Amin Farmahini Farahani, Nuwan Jayasena, Vivek Sarkar |
ISCA | 1 |
| 2018 | In-Register Parameter Caching for Dynamic Neural Nets with Virtual Persistent Processor SpecializationabstractDynamic neural networks enable higher representation flexibility compared to networks with a fixed architecture and are extensively deployed in problems dealing with varying input-induced network structure, such as those in Natural Language Processing. One of the standard optimizations used in static net training is persistency of recurrent weights on the chip. In dynamic nets, possibly-inhomogeneous computation graph for every input prevents caching recurrent weights in GPU registers. Therefore, existing solutions suffer from excessive recurring off-chip memory loads as well as compounded kernel launch overheads leading to underutilization of GPU SMs. In this paper, we present a software system that enables persistency of weight matrices during the training of dynamic neural networks on the GPU. Before the training begins, our approach named Virtual Persistent Processor Specialization (VPPS) specializes a forward-backward propagation kernel that contains in-register caching and operation routines. VPPS virtualizes persistent kernel CTAs as CISC-like vector processors that can be guided to execute supplied instructions. VPPS greatly reduces the overall amount of off-chip loads by caching weight matrices on the chip, while simultaneously, provides maximum portability as it does not make any assumptions about the shape of the given computation graphs hence fulfilling dynamic net requirements. We implemented our solution on DyNet and abstracted away its design complexities by providing simple function calls to the user. Our experiments on a Volta micro-architecture shows that, unlike the most competitive solutions, VPPS shows excellent performance even in small batch sizes and delivers up to 6x speedup on training dynamic nets. Farzad Khorasani, Hodjat Asghari Esfeden, Nael B. Abu-Ghazaleh, Vivek Sarkar |
MICRO | 1 |
| 2016 | CuMAS: Data Transfer Aware Multi-Application Scheduling for Shared GPUsabstractRecent generations of GPUs and their corresponding APIs provide means for sharing compute resources among multiple applications with greater efficiency than ever. This advance has enabled the GPUs to act as shared computation resources in multi-user environments, like supercomputers and cloud computing. Recent research has focused on maximizing the utilization of GPU computing resources by simultaneously executing multiple GPU applications (i.e., concurrent kernels) via temporal or spatial partitioning. However, they have not considered maximizing the utilization of the PCI-e bus which is equally important as applications spend a considerable amount of time on data transfers. Mehmet Esat Belviranli, Farzad Khorasani, Laxmi N. Bhuyan, Rajiv Gupta 0001 |
ICS | 2 |
| 2016 | Eliminating Intra-Warp Load Imbalance in Irregular Nested Patterns via Collaborative Task EngagementabstractNested patterns are one of the most frequently occurring algorithmic themes in GPU applications where coarse-grained tasks are constituted from a number of fine-grained ones. However, efficient execution of irregular nested patterns, with coarse-grained tasks that substantially vary in size, has remained an open problem for the GPU's SIMT architecture. Existing methods rely on static task decomposition where one or a fixed number of threads inside the SIMD grouping (warp) carry out the fine-grained tasks. These approaches fail to provide portable performance across diversity of irregular inputs. Moreover, due to intra-warp load imbalance, they incur warp underutilization. In this paper, we introduce a novel software technique called Collaborative Task Engagement (CTE) that, unlike previous methods, achieves sustained high warp execution efficiencies across irregular inputs and provides portable performance. CTE assigns a group of coarse-grained tasks to the warp and allows threads inside the warp carry out the expanded list of fine-grained tasks collaboratively. In multiple rounds, all the warp threads perform mapping portion of fine-grained tasks and participate in a reduction phase with appropriate lanes to reduce calculated values. This scheme avoids over-subscription or under-subscription of threads while preserving the benefits of parallel reduction. We prepared a CUDA C++ device-side template library for developers to easily express nested patterns in GPU kernels using our technique. Our experiments show that CTE delivers up to 37% warp execution efficiency improvement and gives up to 1.51x speedup over sub-warp decomposition with the best sub-warp width. Farzad Khorasani, Bryan Rowe, Rajiv Gupta 0001, Laxmi N. Bhuyan |
IPDPS | 1 |
| 2015 | Stadium Hashing: Scalable and Flexible Hashing on GPUsabstractHashing is one of the most fundamental operations that provides a means for a program to obtain fast access to large amounts of data. Despite the emergence of GPUs as many-threaded general purpose processors, high performance parallel data hashing solutions for GPUs are yet to receive adequate attention. Existing hashing solutions for GPUs not only impose restrictions (e.g., inability to concurrently execute insertion and retrieval operations, limitation on the size of key-value data pairs) that limit their applicability, their performance does not scale to large hash tables that must be kept out-of-core in the host memory. In this paper we present Stadium Hashing (Stash) that is scalable to large hash tables and practical as it does not impose the aforementioned restrictions. To support large out-of-core hash tables, Stash uses a compact data structure named ticket-board that is separate from hash table buckets and is held inside GPU global memory. Ticket-board locally resolves significant portion of insertion and lookup operations and hence, by reducing accesses to the host memory, it accelerates the execution of these operations. Split design of the ticket-board also enables arbitrarily large keys and values. Unlike existing methods, Stash naturally supports concurrent insertions and retrievals due to its use of double hashing as the collision resolution strategy. Furthermore, we propose Stash with collaborative lanes (clStash) that enhances GPU's SIMD resource utilization for batched insertions during hash table creation. For concurrent insertion and retrieval streams, Stadium hashing can be up to 2 and 3 times faster than GPU Cuckoo hashing for in-core and out-of-core tables respectively. Farzad Khorasani, Mehmet Esat Belviranli, Rajiv Gupta 0001, Laxmi N. Bhuyan |
PACT | 1 |
| 2015 | Scalable SIMD-Efficient Graph Processing on GPUsabstractThe vast computing power of GPUs makes them an attractive platform for accelerating large scale data parallel computations such as popular graph processing applications. However, the inherent irregularity and large sizes of real-world power law graphs makes effective use of GPUs a major challenge. In this paper we develop techniques that greatly enhance the performance and scalability of vertex-centric graph processing on GPUs. First, we present Warp Segmentation, a novel method that greatly enhances GPU device utilization by dynamically assigning appropriate number of SIMD threads to process a vertex with irregular-sized neighbors while employing compact CSR representation to maximize the graph size that can be kept inside the GPU global memory. Prior works can either maximize graph sizes (VWC uses the CSR representation) or device utilization (e.g., CuSha uses the CW representation, however, CW is roughly 2.5x the size of CSR). Second, we further scale graph processing to make use of multiple GPUs while proposing Vertex Refinement to address the challenge of judiciously using the limited bandwidth available for transferring data between GPUs via the PCIe bus. Vertex refinement employs parallel binary prefix sum to dynamically collect only the updated boundary vertices inside GPUs' outbox buffers for dramatically reducing inter-GPU data transfer volume. Whereas existing multi-GPU techniques (Medusa, TOTEM) perform high degree of wasteful vertex transfers. On a single GPU, our framework delivers average speedups of 1.29x to 2.80x over VWC. When scaled to multiple GPUs, our framework achieves up to 2.71x performance improvement compared to inter-GPU vertex communication schemes used by other multi-GPU techniques (i.e., Medusa, TOTEM). Farzad Khorasani, Rajiv Gupta 0001, Laxmi N. Bhuyan |
PACT | 1 |
| 2015 | Efficient warp execution in presence of divergence with collaborative context collectionabstractGPU's SIMD architecture is a double-edged sword confronting parallel tasks with control flow divergence. On the one hand, it provides a high performance yet power-efficient platform to accelerate applications via massive parallelism; however, on the other hand, irregularities induce inefficiencies due to the warp's lockstep traversal of all diverging execution paths. In this work, we present a software (compiler) technique named Collaborative Context Collection (CCC) that increases the warp execution efficiency when faced with thread divergence incurred either by different intra-warp task assignment or by intra-warp load imbalance. CCC collects the relevant registers of divergent threads in a warp-specific stack allocated in the fast shared memory, and restores them only when the perfect utilization of warp lanes becomes feasible. We propose code transformations to enable applicability of CCC to variety of program segments with thread divergence. We also introduce optimizations to reduce the cost of CCC and to avoid device occupancy limitation or memory divergence. We have developed a framework that automates application of CCC to CUDA generated intermediate PTX code. We evaluated CCC on real-world applications and multiple scenarios using synthetic programs. CCC improves the warp execution efficiency of real-world benchmarks by up to 56% and achieves an average speedup of 1.69x (maximum 3.08x). Farzad Khorasani, Rajiv Gupta 0001, Laxmi N. Bhuyan |
MICRO | 1 |
| 2014 | CuSha: vertex-centric graph processing on GPUsabstractVertex-centric graph processing is employed by many popular algorithms (e.g., PageRank) due to its simplicity and efficient use of asynchronous parallelism. The high compute power provided by SIMT architecture presents an opportunity for accelerating these algorithms using GPUs. Prior works of graph processing on a GPU employ Compressed Sparse Row (CSR) form for its space-efficiency; however, CSR suffers from irregular memory accesses and GPU underutilization that limit its performance. In this paper, we present CuSha, a CUDA-based graph processing framework that overcomes the above obstacle via use of two novel graph representations: G-Shards and Concatenated Windows (CW). G-Shards uses a concept recently introduced for non-GPU systems that organizes a graph into autonomous sets of ordered edges called shards. CuSha's mapping of GPU hardware resources on to shards allows fully coalesced memory accesses. CW is a novel representation that enhances the use of shards to achieve higher GPU utilization for processing sparse graphs. Finally, CuSha fully utilizes the GPU power by processing multiple shards in parallel on GPU's streaming multiprocessors. For ease of programming, CuSha allows the user to define the vertex-centric computation and plug it into its framework for parallel processing of large graphs. Our experiments show that CuSha provides significant speedups over the state-of-the-art CSR-based virtual warp-centric method for processing graphs on GPUs. Farzad Khorasani, Keval Vora, Rajiv Gupta 0001, Laxmi N. Bhuyan |
HPDC | 1 |