EDBT 2026 Demo / reviewers in the wild / expert
Ari B. Hayes
dblp:146/3472
· DBLP profile ↗
10ranked-venue papers
4as first author
4since 2021 · last 2023
0009-0005-1965-325XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 3 first-author · 4 since 2021Software engineering, systems software and programming languages · 4 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Exploiting the Regular Structure of Modern Quantum Architectures for Compiling and Optimizing Programs with Permutable OperatorsabstractA critical feature in today's quantum circuit is that they have permutable two-qubit operators. The flexibility in ordering the permutable two-qubit gates leads to more compiler optimization opportunities. However, it also imposes significant challenges due to the additional degree of freedom. Our Contributions are two-fold. We first propose a general methodology that can find structured solutions for scalable quantum hardware. It breaks down the complex compilation problem into two sub-problems that can be solved at small scale. Second, we show how such a structured method can be adapted to practical cases that handle sparsity of the input problem graphs and the noise variability in real hardware. Our evaluation evaluates our method on IBM and Google architecture coupling graphs for up to 1,024 qubits and demonstrate better result in both depth and gate count - by up to 72% reduction in depth, and 66% reduction in gate count. Our real experiments on IBM Mumbai show that we can find better expected minimal energy than the state-of-the-art baseline. Yuwei Jin, Yan-Hao Chen, Ari B. Hayes, Chi Zhang 0041, Eddy Z. Zhang |
ASPLOS (4) | 4 |
| 2023 | A Pulse Generation Framework with Augmented Program-aware Basis Gates and Criticality AnalysisabstractNear-term intermediate-scale quantum (NISQ) devices are subject to considerable noise and short coherence time. Consequently, it is critical to minimize circuit execution latency and improve fidelity. Traditionally, each basis gate of a transpiled circuit is decoded into a fixed episode of the device control pulses. Recent studies investigate the merged pulse generation method for customized gates through quantum optimal control (QOC). In this work, we propose PAQOC, a novel QOC framework that can (i) exploit an augmented program-aware (APA) basis gate set for the tradeoff between compilation time and circuit performance, (ii) prune the search space based on a criticality-centric analytical model and experiment observations we learned from 150 benchmarks. Evaluations using seventeen applications show that PAQOC can achieve an average 54% reduction of the circuit latency, on average 43% reduction in compilation overhead, and a 1.27× improvement in fidelity. PAQOC is available on GitHub1. Yan-Hao Chen, Yuwei Jin, Ari B. Hayes, Ang Li 0006, Yunong Shi, Eddy Z. Zhang |
HPCA | 4 |
| 2021 | Time-optimal Qubit mappingabstractRapid progress in the physical implementation of quantum computers gave birth to multiple recent quantum machines implemented with superconducting technology. In these NISQ machines, each qubit is physically connected to a bounded number of neighbors. This limitation prevents most quantum programs from being directly executed on quantum devices. A compiler is required for converting a quantum program to a hardware-compliant circuit, in particular, making each two-qubit gate executable by mapping the two logical qubits to two physical qubits with a link between them. To solve this problem, existing studies focus on inserting SWAP gates to dynamically remap logical qubits to physical qubits. However, most of the schemes lack the consideration of time-optimality of generated quantum circuits, or are achieving time-optimality with certain constraints. In this work, we propose a theoretically time-optimal SWAP insertion scheme for the qubit mapping problem. Our model can also be extended to practical heuristic algorithms. We present exact analysis results by using our model for quantum programs with recurring execution patterns. We have for the first time discovered an optimal qubit mapping pattern for quantum fourier transformation (QFT) on 2D nearest neighbor architecture. We also present a scalable extension of our theoretical model that can be used to solve qubit mapping for large quantum circuits. Chi Zhang 0041, Ari B. Hayes, Longfei Qiu, Yuwei Jin, Yan-Hao Chen, Eddy Z. Zhang |
ASPLOS | 2 |
| 2021 | AutoBraid: A Framework for Enabling Efficient Surface Code Communication in Quantum ComputingabstractQuantum computers can solve problems that are intractable using the most powerful classical computer. However, qubits are fickle and error prone. It is necessary to actively correct errors in the execution of a quantum circuit. Quantum error correction (QEC) codes are developed to enable fault-tolerant quantum computing. With QEC, one logical circuit is converted into an encoded circuit. Yan-Hao Chen, Yuwei Jin, Chi Zhang 0041, Ari B. Hayes, Youtao Zhang, Eddy Z. Zhang |
MICRO | 5 |
| 2019 | Decoding CUDA BinaryabstractNVIDIA's software does not offer translation of assembly code to binary for their GPUs, since the specifications are closed-source. This work fills that gap. We develop a systematic method of decoding the Instruction Set Architectures (ISAs) of NVIDIA's GPUs, and generating assemblers for different generations of GPUs. Our framework enables cross-architecture binary analysis and transformation. Making the ISA accessible in this manner opens up a world of opportunities for developers and researchers, enabling numerous optimizations and explorations that are unachievable at the source-code level. Our infrastructure has already benefited and been adopted in important applications including performance tuning, binary instrumentation, resource allocation, and memory protection. Ari B. Hayes, Yan-Hao Chen, Eddy Z. Zhang |
CGO | 1 |
| 2018 | Locality-Aware Software Throttling for Sparse Matrix Operation on GPUs
Yan-Hao Chen, Ari B. Hayes, Chi Zhang 0041, Timothy Salmon, Eddy Z. Zhang |
USENIX ATC | 2 |
| 2017 | GPU Taint Tracking
Ari B. Hayes, Lingda Li, Mohammad Hedayati, Jia-Huan He, Eddy Z. Zhang |
USENIX ATC | 1 |
| 2016 | Tag-Split Cache for Efficient GPGPU Cache UtilizationabstractModern GPUs employ cache to improve memory system efficiency. However, large amount of cache space is underutilized due to irregular memory accesses and poor spatial locality which exhibited commonly in GPU applications. Our experiments show that using smaller cache lines could improve cache space utilization, but it also frequently suffers from significant performance loss by introducing large amount of extra cache requests. In this work, we propose a novel cache design named tag-split cache (TSC) that enables fine-grained cache storage to address the problem of cache space underutilization while keeping memory request number unchanged. TSC divides tag into two parts to reduce storage overhead, and it supports multiple cache line replacement in one cycle. TSC can also automatically adjust cache storage granularity to avoid performance loss for applications with good spatial locality. Our evaluation shows that TSC improves the baseline cache performance by 17.2% on average across a wide range of applications. It also out-performs other previous techniques significantly. Lingda Li, Ari B. Hayes, Shuaiwen Song, Eddy Z. Zhang |
ICS | 2 |
| 2016 | Orion: A Framework for GPU Occupancy Tuning
Ari B. Hayes, Lingda Li, Daniel G. Chavarría-Miranda, Shuaiwen Song, Eddy Z. Zhang |
Middleware | 1 |
| 2014 | Unified on-chip memory allocation for SIMT architectureabstractThe popularity of general purpose Graphic Processing Unit (GPU) is largely attributed to the tremendous concurrency enabled by its underlying architecture -- single instruction multiple thread (SIMT) architecture. It keeps the context of a significant number of threads in registers to enable fast ``context switches" when the processor is stalled due to execution dependence, memory requests and etc. The SIMT architecture has a large register file evenly partitioned among all concurrent threads. Per-thread register usage determines the number of concurrent threads, which strongly affects the whole program performance. Existing register allocation techniques, extensively studied in the past several decades, are oblivious to the register contention due to the concurrent execution of many threads. They are prone to making optimization decisions that benefit single thread but degrade the whole application performance. Ari B. Hayes, Eddy Z. Zhang |
ICS | 1 |