Samuel A. Stein

dblp:274/7675 · also Samuel Alexander Stein · DBLP profile ↗
← Back
13ranked-venue papers
7as first author
12since 2021 · last 2026
0000-0002-2655-8251ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 5 first-author · 7 since 2021Software engineering, systems software and programming languages · 5 · 3 first-author · 5 since 2021Theory of computation · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Kernpiler: Compiler Optimization for Quantum Hamiltonian Simulation with Partial Trotterization
abstract
Description This artifact contains the core implementation of Kernpiler, a compiler framework for optimizing quantum circuits, and supports full reproducibility of all experimental results presented in the associated paper. The artifact includes all code, data pipelines, and scripts required to regenerate Figures 5–11. System Used for Data Collection NVIDIA A100 GPU with 80GB memory AMD EPYC 9654P 96-core processor x86_64 Linux system Python 3.13.5 Experiments may be computationally intensive, but they are fully parallelizable across multiple devices. Installation Clone or download the repository and navigate to the project directory. Create a virtual environment: python3 -m venv validate source validate/bin/activate Install dependencies: python -m pip install -r requirements.txt Install torch-scatter: python -m pip install --no-cache-dir torch-scatter -f https://data.pyg.org/whl/torch-2.10.0+cu128.html Experiment Workflow All experiment scripts are located in: src/compiler/optimization_passes/experiments Each figure can be reproduced by running its corresponding data collection and graphing scripts: Figure 5exp_gatecount_datacollection.py→ graph using:graph_data_scripts/graph_absolute.py Figure 6exp_partition_scaling_datacollection_o1exp_partition_scaling_datacollection→ graph using:graph_data_scripts/graph_o1_o2_side_by_side.py Figure 7exp_runtime_per_pass.py→ graph using:graph_data_scripts/graph_runtime_per_pass.py Figure 8exp_partition_scaling_datacollectiono1_phoenixFT→ graph using:graph_data_scripts/graph_firstorder_scalingFT.py Figure 9exp_partitionalgvsrandom.py→ graph using:graph_data_scripts/graph_partition_vs_random.py Figure 10exp_scaling_data_rewriteradius.py→ graph using:graph_data_scripts/graph_scalingdata.py Figure 11exp_error_scaling_systemsize.py→ output generated directly (no additional graph script required) Execution Notes All experiments are independent Parallel execution is supported Runtime varies depending on system size and hardware
Ethan Decker, Lucas Goetz, Evan McKinney, Erik Gustafson, Junyu Zhou 0005, Alex K. Jones, Ang Li 0006, Alexander Schuckert, Samuel A. Stein, Eleanor Crane, Gushu Li
ISCA10
2026 Transpiler-Architecture Co-Design to Curb Clifford Costs in Fault-Tolerant Quantum Computing
abstract
Quantum Error Correction (QEC) codes form the foundation of Fault-Tolerant Quantum Computing (FTQC) and predominantly use the Clifford+T gate set. Recently, Clifford operations have become the key performance bottleneck in implementing QEC. While state-of-the-art approaches like Pauli-Based Compilation (PBC) reduce Clifford overhead by transforming Clifford gates into Pauli measurements, they do so at the cost of gate-level parallelism, inflating circuit depth and execution times. To overcome these limitations, we introduce TACO, a Transpiler-Architecture Co-design framework that tackles the Clifford bottleneck through circuit and architectural optimization. TACO uses FTQC insights to guide hardware-aware Clifford gate elimination and circuit restructuring, and leverages the resulting optimized circuits to refine architectural design. TACO applies FTQC-specific transformations to aggressively reduce Clifford overhead from rotation synthesis and Toffoli decompositions, while preserving gate-level parallelism. The resulting architecture is optimized for the locality and data-movement patterns of these circuits, enabling high-throughput, resource-efficient execution. Our evaluation across diverse benchmarks shows that TACO achieves up to 21.9x (mean 4.4x) reduction in execution time compared to the state-of-the-art baseline.
Meng Wang 0033, Samuel A. Stein, Yufei Ding 0001, Poulami Das 0005, Prashant J. Nair, Ang Li 0006
ISCA3
2026 STQS: A Unified System Architecture for Spatial Temporal Quantum Sensing
abstract
We present STQS, a unified system architecture for spatiotemporal quantum sensing that interlaces four key quantum components: sensing , memory , communication , and computation . By employing a comprehensive gate-based framework, we systemically explore the design space of quantum sensing schemes and probe the influence of noise at each state in a sensing workflow through simulation. We introduce a novel distance-based metric that compares reference states to sensing states and assigns a confidence level. We anticipate that the distance measure will serve as an intermediate step toward more advanced quantum signal processing techniques like quantum machine learning. To our knowledge, STQS is the first system-level framework to integrate quantum sensing within a coherent, unified architectural paradigm. STQS provides seamless avenues for unique state preparation, multi-user sensing requests, and addressing practical implementations. We demonstrate the versatility of STQS through evaluations of quantum radar and qubit-based dark matter detection. To highlight the near-term feasibility of our approach, we present results obtained from IBM’s Marrakesh and IonQ’s Forte devices, validating key STQS components on present day quantum hardware. We have made the simulation code and experimental data used in this work publicly available.
Anastashia Jebraeilli, Keyi Yin, Samuel A. Stein, Erik Lentz, Yufei Ding 0001, Ang Li 0006
ACM Trans. Quantum Comput.4
2025 HetEC: Architectures for Heterogeneous Quantum Error Correction Codes
abstract
Quantum Error Correction (QEC) is essential for future quantum computers due to its ability to exponentially suppress physical errors. The surface code is a leading error-correcting code candidate because of its local topological structure, experimentally achievable thresholds, and support for universal gate operations with magic states. However, its physical overhead scales quadratically with number of correctable errors. Conversely, quantum low-density parity-check (qLDPC) codes offer superior scaling but lack, on their own, a clear path to universal logical computation. Therefore, it is becoming increasingly evident that there are significant advantages to designing architectures using multiple codes. Heterogeneous architectures provide a clear path to universal logical computation as well as the ability to access different resource trade offs.
Samuel A. Stein, Shifan Xu, Andrew W. Cross, Theodore J. Yoder, Ali Javadi-Abhari, Zeyuan Zhou, Charlie Guinn, Yufei Ding 0001, Yongshan Ding 0001, Ang Li 0006
ASPLOS (2)1
2024 ARQUIN: Architectures for Multinode Superconducting Quantum Computers
abstract
Many proposals to scale quantum technology rely on modular or distributed designs wherein individual quantum processors, called nodes, are linked together to form one large multinode quantum computer (MNQC). One scalable method to construct an MNQC is using superconducting quantum systems with optical interconnects. However, internode gates in these systems may be two to three orders of magnitude noisier and slower than local operations. Surmounting the limitations of internode gates will require improvements in entanglement generation, use of entanglement distillation, and optimized software and compilers. Still, it remains unclear what performance is possible with current hardware and what performance algorithms require. In this article, we employ a systems analysis approach to quantify overall MNQC performance in terms of hardware models of internode links, entanglement distillation, and local architecture. We show how to navigate tradeoffs in entanglement generation and distillation in the context of algorithm performance, lay out how compilers and software should balance between local and internode gates, and discuss when noisy quantum internode links have an advantage over purely classical links. We find that a factor of 10–100× better link performance is required and introduce a research roadmap for the co-design of hardware and software towards the realization of early MNQCs. While we focus on superconducting devices with optical interconnects, our approach is general across MNQC implementations.
James Ang 0001, Gabriella Carini, Yanzhu Chen, Isaac L. Chuang, Michael DeMarco, Sophia E. Economou, Alec Eickbusch, Andrei Faraon, Kai-Mei Fu, Steven M. Girvin, Michael Hatridge, Andrew A. Houck, Paul Hilaire, Kevin Krsulich, Ang Li 0006, Yuan Liu 0023, Margaret Martonosi, David C. McKay, Jim Misewich, Mark B. Ritter, Robert J. Schoelkopf, Samuel A. Stein, Sara Sussman, Teague Tomesh, Norm M. Tubman, Nathan Wiebe, Yongxin Yao, Dillon Yost, Yiyu Zhou
ACM Trans. Quantum Comput.23
2023 Distributed Quantum Learning with co-Management in a Multi-tenant Quantum System
abstract
The rapid advancement of quantum computing has pushed classical designs into the quantum domain, breaking physical boundaries for computing-intensive and data-hungry applications Given its immense potential, quantum-based computing systems have attracted increasing attention with the hope that some systems may provide a quantum speedup. For example, variational quantum algorithms have been proposed for quantum neural networks to train deep learning models on qubits, achieving promising results. Existing quantum learning architectures and systems rely on single, monolithic quantum machines with abundant and stable resources, such as qubits. However, fabricating a large, monolithic quantum device is considerably more challenging than producing an array of smaller devices. In this paper, we investigate a distributed quantum system that combines multiple quantum machines into a unified system. We propose DQuLearn, which divides a quantum learning task into multiple subtasks. Each subtask can be executed distributively on individual quantum machines, with the results looping back to classical machines for subsequent training iterations. Additionally, our system supports multiple concurrent clients and dynamically manages their circuits according to the runtime status of quantum workers. Through extensive experiments, we demonstrate that DQuLearn achieves similar accuracies with significant runtime reduction, by up to 68.7% and an increase per-second circuit processing speed, by up to 3.99 times, in a 4-worker multi-tenant setting.
Anthony D'Onofrio Jr., Amir Hossain, Lesther Santana, Naseem Machlovi, Samuel A. Stein, Ang Li 0006, Ying Mao 0001
IEEE Big Data5
2023 Q-BEEP: Quantum Bayesian Error Mitigation Employing Poisson Modeling over the Hamming Spectrum
abstract
Quantum computing technology has grown rapidly in recent years, with new technologies being explored, error rates being reduced, and quantum processors' qubit capacity growing. However, near-term quantum algorithms are still unable to be induced without compounding consequential levels of noise, leading to non-trivial erroneous results. Quantum Error Correction (in-situ error mitigation) and Quantum Error Mitigation (post-induction error mitigation) are promising fields of research within the quantum algorithm scene, aiming to alleviate quantum errors. IBM recently published an article stating that Quantum Error Mitigation is the path to quantum computing usefulness. A recent work, namely HAMMER, demonstrated the existence of a latent structure regarding post-circuit induction errors when mapping to the Hamming spectrum. However, they assumed that errors occur solely in local clusters, whereas we observe that at higher average Hamming distances this structure falls away. In this work, we show that such a correlated structure is not only local but extends certain non-local clustering patterns which can be precisely described by a Poisson distribution model taking the input circuit, the device run time status (i.e., calibration statistics) and qubit topology into consideration. Using this quantum error characterizing model, we developed an iterative algorithm over the generated Bayesian network state-graph for post-induction error mitigation. Thanks to more precise modeling of the error distribution latent structure and the proposed iterative method, our Q-Beep approach provides state of the art performance and can boost circuit execution fidelity by up to 234.6% on Bernstein-Vazirani circuits and on average 71.0% on QAOA solution quality, using 16 practical IBMQ quantum processors. For other benchmarks such as those in QASMBench, a fidelity improvement of up to 17.8% is attained. Q-Beep is a light-weight post-processing technique that can be performed offline and remotely, making it a useful tool for quantum vendors to adopt and provide more reliable circuit induction results. Q-Beep is maintained at github.com/pnnl/qbeep
Samuel A. Stein, Nathan Wiebe, Yufei Ding 0001, James Ang 0001, Ang Li 0006
ISCA1
2023 HetArch: Heterogeneous Microarchitectures for Superconducting Quantum Systems
abstract
Noisy Intermediate-Scale Quantum Computing (NISQ) has dominated headlines in recent years, with the longer-term vision of Fault-Tolerant Quantum Computation (FTQC) offering significant potential albeit at currently intractable resource costs and quantum error correction (QEC) overheads. For problems of interest, FTQC will require millions of physical qubits with long coherence times, high-fidelity gates, and compact sizes to surpass classical systems. Just as heterogeneous specialization has offered scaling benefits in classical computing, it is likewise gaining interest in FTQC. However, systematic use of heterogeneity in either hardware or software elements of FTQC systems remains a serious challenge due to the vast design space and variable physical constraints.
Samuel A. Stein, Sara Sussman, Teague Tomesh, Charlie Guinn, Esin Tureci, Sophia Fuhui Lin, James Ang 0001, Srivatsan Chakram, Ang Li 0006, Margaret Martonosi, Fred Chong, Andrew A. Houck, Isaac L. Chuang, Michael DeMarco
MICRO1
2023 QASMBench: A Low-Level Quantum Benchmark Suite for NISQ Evaluation and Simulation
abstract
The rapid development of quantum computing (QC) in the NISQ era urgently demands a low-level benchmark suite and insightful evaluation metrics for characterizing the properties of prototype NISQ devices, the efficiency of QC programming compilers, schedulers and assemblers, and the capability of quantum system simulators in a classical computer. In this work, we fill this gap by proposing a low-level, easy-to-use benchmark suite called QASMBench based on the OpenQASM assembly representation. It consolidates commonly used quantum routines and kernels from a variety of domains including chemistry, simulation, linear algebra, searching, optimization, arithmetic, machine learning, fault tolerance, cryptography, and so on, trading-off between generality and usability. To analyze these kernels in terms of NISQ device execution, in addition to circuit width and depth, we propose four circuit metrics including gate density, retention lifespan, measurement density, and entanglement variance, to extract more insights about the execution efficiency, the susceptibility to NISQ error, and the potential gain from machine-specific optimizations. Applications in QASMBench can be launched and verified on several NISQ platforms, including IBM-Q, Rigetti, IonQ and Quantinuum. For evaluation, we measure the execution fidelity of a subset of QASMBench applications on 12 IBM-Q machines through density matrix state tomography, comprising 25K circuit evaluations. We also compare the fidelity of executions among the IBM-Q machines, the IonQ QPU and the Rigetti Aspen M-1 system. QASMBench is released at: http://github.com/pnnl/QASMBench .
Ang Li 0006, Samuel A. Stein, Sriram Krishnamoorthy, James Ang 0001
ACM Trans. Quantum Comput.2
2022 QuCNN: A Quantum Convolutional Neural Network with Entanglement Based Backpropagation
abstract
Quantum Machine Learning continues to be a highly active area of interest within Quantum Computing. Many of these approaches have adapted classical approaches to the quantum settings, such as QuantumFlow, etc. We push forward this trend, and demonstrate an adaption of the Classical Convolutional Neural Networks to quantum systems - namely QuCNN. QuCNN is a parameterised multi-quantum-state based neural network layer computing similarities between each quantum filter state and each quantum data state. With QuCNN, back propagation can be achieved through a single-ancilla qubit quantum routine. QuCNN is validated by applying a convolutional layer with a data state and a filter state over a small subset of MNIST images, comparing the backpropagated gradients, and training a filter state against an ideal target state.
Samuel A. Stein, Ying Mao 0001, James Ang 0001, Ang Li 0006
SEC1
2022 EQC: ensembled quantum computing for variational quantum algorithms
abstract
Variational quantum algorithm (VQA), which is comprised of a classical optimizer and a parameterized quantum circuit, emerges as one of the most promising approaches for harvesting the power of quantum computers in the noisy intermediate scale quantum (NISQ) era. However, the deployment of VQAs on contemporary NISQ devices often faces considerable system and time-dependant noise and prohibitively slow training speeds. On the other hand, the expensive supporting resources and infrastructure make quantum computers extremely keen on high utilization.
Samuel A. Stein, Nathan Wiebe, Yufei Ding 0001, Bo Peng 0024, Karol Kowalski, Nathan A. Baker, James Ang 0001, Ang Li 0006
ISCA1
2021 A Hybrid System for Learning Classical Data in Quantum States
abstract
Deep neural network powered artificial intelligence has rapidly changed our daily life with various applications. However, as one of the essential steps of deep neural networks, training a heavily-weighted network requires a tremendous amount of computing resources. Especially in the post Moore’s Law era, the limit of semiconductor fabrication technology has restricted the development of learning algorithms to cope with the increasing high intensity training data. Meanwhile, quantum computing has demonstrated its significant potential in terms of speeding up the traditionally compute-intensive workloads. For example, Google illustrated quantum supremacy by completing a sampling calculation task in 200 seconds, which is otherwise impracticable on the world’s largest supercomputers. To this end, quantum-based learning has become an area of interest, with the potential of a quantum speedup. In this paper, we propose GenQu, a hybrid and general-purpose quantum framework for learning classical data through quantum states. We evaluate GenQu with real datasets and conduct experiments on both simulations and real quantum computer IBM-Q. Our evaluation demonstrates that, compared with classical solutions, the proposed models running on GenQu framework achieve similar accuracy with a much smaller number of qubits, while significantly reducing the parameter size by up to 95.86% and converging speedup by 33.33% faster.
Samuel A. Stein, Ryan L'Abbate, Wenrui Mu, Betis Baheri, Ying Mao 0001, Qiang Guan, Ang Li 0006, Bo Fang 0002
IPCCC1
2020 A College Major Recommendation System
abstract
College students are required to select a major but are often provided with only a modest amount of support in making this important decision. A poor decision is detrimental to the student, since it may result in the student later switching to a different major with a delay in graduation—or even result in the student leaving the university. This also impacts the university since time to graduation and retention rate are used to evaluate the quality of a university. There is a general lack of research on recommender systems for college majors, with the most relevant systems focusing on course-level recommendations. This study describes and evaluates a recommender system for selecting an undergraduate major, utilizing nine years of historical student data from a large university. The system bases its recommendations on the courses that the student takes in the first few years of college, and how well they performed in these courses. The system is designed to recommend majors that the student is likely to be interested in and will perform well in. Recommendations are evaluated based on the likelihood that the student's actual major was in the top five recommended majors, and whether the student performed above average in that major. The recommendation system dramatically outperforms the baseline strategy of randomly selecting a major, and when the recommendation is followed the student is 12% more likely to perform above average in the major.
Samuel A. Stein, Gary Weiss 0001, Daniel D. Leeds
RecSys1