Arslan Zulfiqar

dblp:138/4163 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
0since 2021 · last 2018
0009-0003-6240-5900ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
GPUs and heterogeneous computing · 34% Cloud and datacenter computing · 17% Memory systems · 14%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%
Artificial intelligence
1 paper
Efficient and distributed learning · 100%

Topics — the 14 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
GPUs and heterogeneous computing
GPU memory management
0.522016
vDNN: Virtualized deep neural networks for scalable, memory-efficient neural network design · MICRO 2016
Towards high performance paged memory for GPUs · HPCA 2016
Compilers and program optimization
register allocation
0.312018
Software-Directed Techniques for Improved GPU Register File Utilization · ACM Trans. Archit. Code Optim. 2018
GPUs and heterogeneous computing
GPU architecture
0.312018
Software-Directed Techniques for Improved GPU Register File Utilization · ACM Trans. Archit. Code Optim. 2018
Electronic design automation › high-level synthesis › resource binding
register allocation
0.312018
Software-Directed Techniques for Improved GPU Register File Utilization · ACM Trans. Archit. Code Optim. 2018
Processor architecture and microarchitecture
register file
0.312018
Software-Directed Techniques for Improved GPU Register File Utilization · ACM Trans. Archit. Code Optim. 2018
Machine learning › Efficient and distributed learning
memory-efficient training
0.212016
vDNN: Virtualized deep neural networks for scalable, memory-efficient neural network design · MICRO 2016
Cloud and datacenter computing › resource management › datacenter memory management
memory oversubscription
0.212016
Towards high performance paged memory for GPUs · HPCA 2016
Cloud and datacenter computing › virtualization
memory virtualization
0.212016
vDNN: Virtualized deep neural networks for scalable, memory-efficient neural network design · MICRO 2016
Memory systems › memory management › virtual memory
paged memory
0.212016
Towards high performance paged memory for GPUs · HPCA 2016
Interconnection networks and networks-on-chip
optical interconnection networks
0.212013
Wavelength stealing: an opportunistic approach to channel sharing in multi-chip photonic interconnects · MICRO 2013
Memory systems › memory management › virtual memory
address translation
0.112016
Towards high performance paged memory for GPUs · HPCA 2016
GPUs and heterogeneous computing
GPU training
0.112016
vDNN: Virtualized deep neural networks for scalable, memory-efficient neural network design · MICRO 2016
GPUs and heterogeneous computing › GPU communication
host-device data transfer
0.112016
Towards high performance paged memory for GPUs · HPCA 2016
Memory systems › memory management
virtual memory
0.112016
Towards high performance paged memory for GPUs · HPCA 2016

Methods — techniques the papers use, named apart from their topics

simulation · 0.9compiler-driven optimization · 0.7runtime memory manager · 0.5
YearPublicationVenuePosition
2018 Software-Directed Techniques for Improved GPU Register File Utilization
abstract
Throughput architectures such as GPUs require substantial hardware resources to hold the state of a massive number of simultaneously executing threads. While GPU register files are already enormous, reaching capacities of 256KB per streaming multiprocessor (SM), we find that nearly half of real-world applications we examined are register-bound and would benefit from a larger register file to enable more concurrent threads. This article seeks to increase the thread occupancy and improve performance of these register-bound applications by making more efficient use of the existing register file capacity. Our first technique eagerly deallocates register resources during execution. We show that releasing register resources based on value liveness as proposed in prior states of the art leads to unreliable performance and undue design complexity. To address these deficiencies, our article presents a novel compiler-driven approach that identifies and exploits last use of a register name (instead of the value contained within) to eagerly release register resources. Furthermore, while previous works have leveraged “scalar” and “narrow” operand properties of a program for various optimizations, their impact on thread occupancy has been relatively unexplored. Our article evaluates the effectiveness of these techniques in improving thread occupancy and demonstrates that while any one approach may fail to free very many registers, together they synergistically free enough registers to launch additional parallel work. An in-depth evaluation on a large suite of applications shows that just our early register technique outperforms previous work on dynamic register allocation, and together these approaches, on average, provide 12% performance speedup (23% higher thread occupancy) on register bound applications not already saturating other GPU resources.
Dani Voitsechov, Arslan Zulfiqar, Mark Stephenson, Mark Gebhart, Stephen W. Keckler
ACM Trans. Archit. Code Optim.2
2016 Towards high performance paged memory for GPUs
abstract
Despite industrial investment in both on-die GPUs and next generation interconnects, the highest performing parallel accelerators shipping today continue to be discrete GPUs. Connected via PCIe, these GPUs utilize their own privately managed physical memory that is optimized for high bandwidth. These separate memories force GPU programmers to manage the movement of data between the CPU and GPU, in addition to the on-chip GPU memory hierarchy. To simplify this process, GPU vendors are developing software runtimes that automatically page memory in and out of the GPU on-demand, reducing programmer effort and enabling computation across datasets that exceed the GPU memory capacity. Because this memory migration occurs over a high latency and low bandwidth link (compared to GPU memory), these software runtimes may result in significant performance penalties. In this work, we explore the features needed in GPU hardware and software to close the performance gap of GPU paged memory versus legacy programmer directed memory management. Without modifying the GPU execution pipeline, we show it is possible to largely hide the performance overheads of GPU paged memory, converting an average 2× slowdown into a 12% speedup when compared to programmer directed transfers. Additionally, we examine the performance impact that GPU memory oversubscription has on application run times, enabling application designers to make informed decisions on how to shard their datasets across hosts and GPU instances.
Tianhao Zheng, David W. Nellans, Arslan Zulfiqar, Mark Stephenson, Stephen W. Keckler
HPCA3
2016 vDNN: Virtualized deep neural networks for scalable, memory-efficient neural network design
abstract
The most widely used machine learning frameworks require users to carefully tune their memory usage so that the deep neural network (DNN) fits into the DRAM capacity of a GPU. This restriction hampers a researcher's flexibility to study different machine learning algorithms, forcing them to either use a less desirable network architecture or parallelize the processing across multiple GPUs. We propose a runtime memory manager that virtualizes the memory usage of DNNs such that both GPU and CPU memory can simultaneously be utilized for training larger DNNs. Our virtualized DNN (vDNN) reduces the average GPU memory usage of AlexNet by up to 89%, OverFeat by 91%, and GoogLeNet by 95%, a significant reduction in memory requirements of DNNs. Similar experiments on VGG-16, one of the deepest and memory hungry DNNs to date, demonstrate the memory-efficiency of our proposal. vDNN enables VGG-16 with batch size 256 (requiring 28 GB of memory) to be trained on a single NVIDIA Titan X GPU card containing 12 GB of memory, with 18% performance loss compared to a hypothetical, oracular GPU with enough memory to hold the entire DNN.
Minsoo Rhu, Natalia Gimelshein, Jason Clemons, Arslan Zulfiqar, Stephen W. Keckler
MICRO4
2013 Wavelength stealing: an opportunistic approach to channel sharing in multi-chip photonic interconnects
abstract
Silicon photonic technology offers seamless integration of multiple chips with high bandwidth density and lower energy-per-bit consumption compared to electrical interconnects. The topology of a photonic interconnect impacts both its performance and laser power requirements. The point-to-point (P2P) topology offers arbitration-free connectivity with low energy-per-bit consumption, but suffers from low node-to-node bandwidth. Topologies with channel-sharing improve inter-node bandwidth but incur higher laser power consumption in addition to the performance costs associated with arbitration and contention.
Arslan Zulfiqar, Pranay Koka, Herb Schwetman, Mikko H. Lipasti, Xuezhe Zheng, Ashok V. Krishnamoorthy
MICRO1