VLDB 2026 Research / reviewers in the wild / expert
Arslan Zulfiqar
dblp:138/4163
· DBLP profile ↗
4ranked-venue papers
1as first author
0since 2021 · last 2018
0009-0003-6240-5900ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
GPUs and heterogeneous computing · 34% Cloud and datacenter computing · 17% Memory systems · 14% | |
| Software engineering, system software, and programming languages
1 paper |
Compilers and program optimization · 100% | |
| Artificial intelligence
1 paper |
Efficient and distributed learning · 100% |
Topics — the 14 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
GPUs and heterogeneous computing
GPU memory management |
0.5 | 2 | 2016 | vDNN: Virtualized deep neural networks for scalable, memory-efficient neural network design · MICRO 2016 Towards high performance paged memory for GPUs · HPCA 2016 |
Compilers and program optimization
register allocation |
0.3 | 1 | 2018 | Software-Directed Techniques for Improved GPU Register File Utilization · ACM Trans. Archit. Code Optim. 2018 |
GPUs and heterogeneous computing
GPU architecture |
0.3 | 1 | 2018 | Software-Directed Techniques for Improved GPU Register File Utilization · ACM Trans. Archit. Code Optim. 2018 |
Electronic design automation › high-level synthesis › resource binding
register allocation |
0.3 | 1 | 2018 | Software-Directed Techniques for Improved GPU Register File Utilization · ACM Trans. Archit. Code Optim. 2018 |
Processor architecture and microarchitecture
register file |
0.3 | 1 | 2018 | Software-Directed Techniques for Improved GPU Register File Utilization · ACM Trans. Archit. Code Optim. 2018 |
Machine learning › Efficient and distributed learning
memory-efficient training |
0.2 | 1 | 2016 | vDNN: Virtualized deep neural networks for scalable, memory-efficient neural network design · MICRO 2016 |
Cloud and datacenter computing › resource management › datacenter memory management
memory oversubscription |
0.2 | 1 | 2016 | Towards high performance paged memory for GPUs · HPCA 2016 |
Cloud and datacenter computing › virtualization
memory virtualization |
0.2 | 1 | 2016 | vDNN: Virtualized deep neural networks for scalable, memory-efficient neural network design · MICRO 2016 |
Memory systems › memory management › virtual memory
paged memory |
0.2 | 1 | 2016 | Towards high performance paged memory for GPUs · HPCA 2016 |
Interconnection networks and networks-on-chip
optical interconnection networks |
0.2 | 1 | 2013 | Wavelength stealing: an opportunistic approach to channel sharing in multi-chip photonic interconnects · MICRO 2013 |
Memory systems › memory management › virtual memory
address translation |
0.1 | 1 | 2016 | Towards high performance paged memory for GPUs · HPCA 2016 |
GPUs and heterogeneous computing
GPU training |
0.1 | 1 | 2016 | vDNN: Virtualized deep neural networks for scalable, memory-efficient neural network design · MICRO 2016 |
GPUs and heterogeneous computing › GPU communication
host-device data transfer |
0.1 | 1 | 2016 | Towards high performance paged memory for GPUs · HPCA 2016 |
Memory systems › memory management
virtual memory |
0.1 | 1 | 2016 | Towards high performance paged memory for GPUs · HPCA 2016 |
Methods — techniques the papers use, named apart from their topics
simulation · 0.9compiler-driven optimization · 0.7runtime memory manager · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2018 | Software-Directed Techniques for Improved GPU Register File UtilizationabstractThroughput architectures such as GPUs require substantial hardware resources to hold the state of a massive number of simultaneously executing threads. While GPU register files are already enormous, reaching capacities of 256KB per streaming multiprocessor (SM), we find that nearly half of real-world applications we examined are register-bound and would benefit from a larger register file to enable more concurrent threads. This article seeks to increase the thread occupancy and improve performance of these register-bound applications by making more efficient use of the existing register file capacity. Our first technique eagerly deallocates register resources during execution. We show that releasing register resources based on value liveness as proposed in prior states of the art leads to unreliable performance and undue design complexity. To address these deficiencies, our article presents a novel compiler-driven approach that identifies and exploits last use of a register name (instead of the value contained within) to eagerly release register resources. Furthermore, while previous works have leveraged “scalar” and “narrow” operand properties of a program for various optimizations, their impact on thread occupancy has been relatively unexplored. Our article evaluates the effectiveness of these techniques in improving thread occupancy and demonstrates that while any one approach may fail to free very many registers, together they synergistically free enough registers to launch additional parallel work. An in-depth evaluation on a large suite of applications shows that just our early register technique outperforms previous work on dynamic register allocation, and together these approaches, on average, provide 12% performance speedup (23% higher thread occupancy) on register bound applications not already saturating other GPU resources. Dani Voitsechov, Arslan Zulfiqar, Mark Stephenson, Mark Gebhart, Stephen W. Keckler |
ACM Trans. Archit. Code Optim. | 2 |
| 2016 | Towards high performance paged memory for GPUsabstractDespite industrial investment in both on-die GPUs and next generation interconnects, the highest performing parallel accelerators shipping today continue to be discrete GPUs. Connected via PCIe, these GPUs utilize their own privately managed physical memory that is optimized for high bandwidth. These separate memories force GPU programmers to manage the movement of data between the CPU and GPU, in addition to the on-chip GPU memory hierarchy. To simplify this process, GPU vendors are developing software runtimes that automatically page memory in and out of the GPU on-demand, reducing programmer effort and enabling computation across datasets that exceed the GPU memory capacity. Because this memory migration occurs over a high latency and low bandwidth link (compared to GPU memory), these software runtimes may result in significant performance penalties. In this work, we explore the features needed in GPU hardware and software to close the performance gap of GPU paged memory versus legacy programmer directed memory management. Without modifying the GPU execution pipeline, we show it is possible to largely hide the performance overheads of GPU paged memory, converting an average 2× slowdown into a 12% speedup when compared to programmer directed transfers. Additionally, we examine the performance impact that GPU memory oversubscription has on application run times, enabling application designers to make informed decisions on how to shard their datasets across hosts and GPU instances. Tianhao Zheng, David W. Nellans, Arslan Zulfiqar, Mark Stephenson, Stephen W. Keckler |
HPCA | 3 |
| 2016 | vDNN: Virtualized deep neural networks for scalable, memory-efficient neural network designabstractThe most widely used machine learning frameworks require users to carefully tune their memory usage so that the deep neural network (DNN) fits into the DRAM capacity of a GPU. This restriction hampers a researcher's flexibility to study different machine learning algorithms, forcing them to either use a less desirable network architecture or parallelize the processing across multiple GPUs. We propose a runtime memory manager that virtualizes the memory usage of DNNs such that both GPU and CPU memory can simultaneously be utilized for training larger DNNs. Our virtualized DNN (vDNN) reduces the average GPU memory usage of AlexNet by up to 89%, OverFeat by 91%, and GoogLeNet by 95%, a significant reduction in memory requirements of DNNs. Similar experiments on VGG-16, one of the deepest and memory hungry DNNs to date, demonstrate the memory-efficiency of our proposal. vDNN enables VGG-16 with batch size 256 (requiring 28 GB of memory) to be trained on a single NVIDIA Titan X GPU card containing 12 GB of memory, with 18% performance loss compared to a hypothetical, oracular GPU with enough memory to hold the entire DNN. Minsoo Rhu, Natalia Gimelshein, Jason Clemons, Arslan Zulfiqar, Stephen W. Keckler |
MICRO | 4 |
| 2013 | Wavelength stealing: an opportunistic approach to channel sharing in multi-chip photonic interconnectsabstractSilicon photonic technology offers seamless integration of multiple chips with high bandwidth density and lower energy-per-bit consumption compared to electrical interconnects. The topology of a photonic interconnect impacts both its performance and laser power requirements. The point-to-point (P2P) topology offers arbitration-free connectivity with low energy-per-bit consumption, but suffers from low node-to-node bandwidth. Topologies with channel-sharing improve inter-node bandwidth but incur higher laser power consumption in addition to the performance costs associated with arbitration and contention. Arslan Zulfiqar, Pranay Koka, Herb Schwetman, Mikko H. Lipasti, Xuezhe Zheng, Ashok V. Krishnamoorthy |
MICRO | 1 |