Tirthak Patel

dblp:208/1839 · DBLP profile ↗
← Back
59ranked-venue papers
20as first author
42since 2021 · last 2026
0000-0003-3127-5931ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 51 · 20 first-author · 34 since 2021Software engineering, systems software and programming languages · 9 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Security and privacy · 4 · 1 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 TuniQ: Autotuning Compilation Passes for Quantum Workloads at Scale for Effectiveness and Efficiency
abstract
Quantum processors are being integrated into HPC ecosystems as co-processors, where compilation of quantum circuits into hardware-executable form determines both output fidelity and runtime. Current compilers use a fixed pass sequence and ignore the fact that optimal pass selection varies with circuit, hardware, and noise conditions. We present TuniQ, a reinforcement learning-based system that selects compilation passes at each pipeline stage, adapting to circuit, backend, and current noise profile. TuniQ introduces several novel design components like a dual-encoder for stage-aware representation, shaped rewards for cross-stage credit assignment, and dynamic action masking for valid compilation. Evaluated across diverse quantum workloads on multiple IBM Quantum Cloud processors, TuniQ improves fidelity and reduces compilation time over the state-of-the-art IBM Qiskit transpiler, generalizes across backends without retraining, and scales strongly to utility-scale circuits with growing advantage.
Mohammad Abrarul Hasanat, Jason Ludmir, Tirthak Patel, Rohan Basu Roy
ICS3
2026 SpinTune: Improving the Reliability of Quantum Sensor Networks for Practical Quantum-Classical Utility
Jason Ludmir, Nicholas S. DiBrita, Jason Han, Tirthak Patel
ICS4
2026 QuantAid: A Quiz-Based Quantum Learning Platform for High-school and Undergraduate Students
abstract
Quantum computing (QC) is poised to transform multiple fields from information technology to medical science, but early exposure to its foundational concepts remains scarce, especially at the high school and undergraduate levels. Educational bottlenecks stem from limited course offerings, a global shortage of qualified instructors, and a steep conceptual and mathematical learning curve. Existing resources are often either too superficial or too advanced, exacerbating accessibility and impeding the development of a diverse, ''quantum-ready'' workforce. This paper presents a pilot study of QuantAid, an AI-enhanced, quiz-based learning platform designed to address these challenges by delivering personalized, scalable QC instruction grounded in constructivist pedagogy. Our platform integrates a light-weight, LLM-powered analogy engine that tailors quantum metaphors to students' prior knowledge and hobbies, scaffolds abstract quantum ideas to reduce cognitive load, and embeds interactive quizzes with immediate feedback to actively reinforce understanding. A conversational AI tutor, trained on vetted QC content, provides on-demand explanations that align with best practices in trust and transparency for educational AI. In a pilot with ?? = 21 students, users of the enhanced platform showed higher engagement and modest conceptual gains compared to peers using static materials. Usability (SUS) and user experience (UEQ) ratings were strong, with particular praise for the platform's analogical reasoning and interactivity. These findings suggest that a well-designed, AI-augmented constructivist tool can enable QC education at scale while preserving rigor, addressing the rising demand amid limited teaching resources.
Kevin Hernandez, Kaushani Patel, Tirthak Patel
SIGCSE (1)3
2026 Automated Tensor-Relational Decomposition for Large-Scale Sparse Tensor Computation
Zhiyuan Xin, Zhimin Ding, Daniel Bourgeois, Tirthak Patel, Chris Jermaine
Proc. VLDB Endow.6
2025 EnQode: Fast Amplitude Embedding for Quantum Machine Learning Using Classical Data
abstract
Amplitude embedding (AE) is essential in quantum machine learning (QML) for encoding classical data onto quantum circuits. However, conventional AE methods suffer from deep, variable-length circuits that introduce high output error due to extensive gate usage and variable error rates across samples, resulting in noise-driven inconsistencies that degrade model accuracy. We introduce EnQode, a fast AE technique based on symbolic representation that addresses these limitations by clustering dataset samples and solving for cluster mean states through a low-depth, machine-specific ansatz. Optimized to reduce physical gates and SWAP operations, EnQode ensures all samples face consistent, low noise levels by standardizing circuit depth and composition. With over 94% fidelity in data mapping, EnQode enables robust, high-performance QML on noisy intermediate-scale quantum (NISQ) devices. Our opensource solution provides a scalable and efficient alternative for integrating classical data with quantum models.
Jason Han, Nicholas S. DiBrita, Younghyun Cho, Hengrui Luo, Tirthak Patel
DAC5
2025 Quorum: Zero-Training Unsupervised Anomaly Detection using Quantum Autoencoders
abstract
Detecting mission-critical anomalous events and data is a crucial challenge across various industries, including finance, healthcare, and energy. Quantum computing has recently emerged as a powerful tool for tackling several machine learning tasks, but training quantum machine learning models remains challenging, particularly due to the difficulty of gradient calculation. The challenge is even greater for anomaly detection, where unsupervised learning methods are essential to ensure practical applicability. To address these issues, we propose Quorum, the first quantum anomaly detection framework designed for unsupervised learning that operates without requiring any training.
Jason Zev Ludmir, Sophia Rebello, Jacob Ruiz, Tirthak Patel
DAC4
2025 Quantum Neural Networks Need Checkpointing
abstract
Quantum Neural Networks (QNNs) harness quantum superposition and entanglement, offering promising advantages for machine learning tasks. However, noise in quantum computers frequently disrupts QNN training, wasting computational resources and extending queue times. This paper introduces the first QNN checkpointing framework to address this challenge. Through experiments on various quantum devices, we demonstrate that QNN behavior is fundamentally hardware-dependent, with the same model performing differently across platforms. This key finding shows that quantum checkpoints require additional metadata about hardware specifics and shot counts unique to quantum systems. Our framework requires minimal storage (only 186.6KB for a 100-qubit QNN) and negligible overhead, enabling frequent checkpointing to enhance training resilience and reproducibility in the NISQ era.
Christopher Kverne, Mayur Akewar, Yuqian Huo, Tirthak Patel, Janki Bhimani
HotStorage4
2025 Revisiting Noise-adaptive Transpilation in Quantum Computing: How Much Impact Does it Have?
abstract
Transpilation, particularly noise-aware optimization, is widely regarded as essential for maximizing the performance of quantum circuits on superconducting quantum computers. The common wisdom is that each circuit should be transpiled using up-to-date noise calibration data to optimize fidelity. In this work, we revisit the necessity of frequent noise-adaptive transpilation, conducting an in-depth empirical study across five IBM 127-qubit quantum computers and 16 diverse quantum algorithms. Our findings reveal novel and interesting insights: (1) noise-aware transpilation leads to a heavy concentration of workloads on a small subset of qubits, which increases output error variability; (2) using random mapping can mitigate this effect while maintaining comparable average fidelity; and (3) circuits compiled once with calibration data can be reliably reused across multiple calibration cycles and time periods without significant loss in fidelity. These results suggest that the classical overhead associated with daily, per-circuit noise-aware transpilation may not be justified. We propose lightweight alternatives that reduce this overhead without sacrificing fidelity – offering a path to more efficient and scalable quantum workflows.
Yuqian Huo, Jinbiao Wei, Christopher Kverne, Mayur Akewar, Janki Bhimani, Tirthak Patel
ICCAD6
2025 ResQ: A Novel Framework to Implement Residual Neural Networks on Analog Rydberg Atom Quantum Computers
abstract
Research in quantum machine learning has recently proliferated due to the potential of quantum computing to accelerate machine learning. An area of machine learning that has not yet been explored is neural ordinary differential equation (neural ODE) based residual neural networks (ResNets), which aim to improve the effectiveness of neural networks using the principles of ordinary differential equations. In this work, we present our insights about why analog Rydberg atom quantum computers are especially well-suited for ResNets. We also introduce ResQ, a novel framework to optimize the dynamics of Rydberg atom quantum computers to solve classification problems in machine learning using analog quantum neural ODEs.
Nicholas S. DiBrita, Jason Han, Tirthak Patel
ICCV3
2025 OpaQue: Program Output Obfuscation for Quantum Software Circuits in Quantum Clouds
abstract
Recent quantum software engineering efforts have made significant progress in testing and debugging quantum algorithms -however, providing confidentiality and privacy to quantum software in the cloud remains an unexplored critical area.OpaQue is the first solution to obfuscate quantum software and output to prevent the leaking of confidential information over the cloud.OpaQue implements a lightweight, scalable, and effective solution based on the unique principles of quantum computing to achieve this task.
Tirthak Patel, Aditya Ranjan, Daniel Silver, Harshitta Gandhi, William Cutler, Devesh Tiwari
ICS1
2025 GreenMix: Energy-Efficient Serverless Computing via Randomized Sketching on Asymmetric Multi-Cores
abstract
GreenMix is motivated by the renewed interest in asymmetric multi-core processors and the emergence of the serverless computing model. Asymmetric multi-cores offer better energy and performance trade-offs by placing different core types on the same die. However, existing serverless scheduling techniques do not leverage these benefits. GreenMix is the first serverless work to reduce energy and serverless keep-alive costs while meeting QoS targets by leveraging asymmetric multi-cores. GreenMix employs randomized sketching, tailored for serverless execution and keep-alive, to perform within 10% of the optimal solution in terms of energy efficiency and keep-alive cost reduction. GreenMix’s effectiveness is demonstrated through evaluations on clusters of ARM big.LITTLE and Intel Alder Lake asymmetric processors. It outperforms competing state-of-the-art schedulers, offering a novel approach for energy-efficient serverless computing.
Rohan Basu Roy, Tirthak Patel, Baolin Li 0001, Siddharth Samsi, Vijay Gadepally, Devesh Tiwari
SC2
2025 Enhancing Early Quantum Computing Education with QuantumAiEd: Bridging the Educational Gap
abstract
Quantum computing (QC) offers unprecedented computational power for solving complex problems in various domains. However, there is a significant gap in accessible QC educational resources for young learners, particularly at the high school and undergraduate levels. Early exposure to quantum concepts is crucial for cultivating future innovators. Moreover, QC education is often limited to institutions with higher resources. As QC is anticipated to become mainstream and disrupt many industries, the demand for a skilled, diverse workforce in QC will rise significantly. Current QC educational tools either only touch on surface-level concepts, relying on gamified concepts without depth, or are too advanced, requiring extensive mathematical and quantum physics backgrounds. To address this, QuantumAiEd bridges the educational gap by offering personalized, accessible QC education, ensuring inclusivity, and preparing a diverse future workforce. QuantumAiEd is an AI-based platform that bridges the educational gap by providing personalized, accessible QC education. AI chatbots have shown promise in personalized learning, but their application in QC education remains underexplored. QuantumAiEd offers interactive learning through AI-generated content, short quizzes, and immediate feedback, dynamically adapting to individual needs and making complex concepts more approachable. A preliminary study with 20 students showed that QuantumAiEd improves understanding and engagement, increasing performance after revisiting incorrectly answered concepts. The User Experience Questionnaire (UEQ-S) also indicated a positive user experience. This study demonstrates the potential of AI-powered platforms to democratize QC education and address existing educational gaps.
Kevin Hernandez, Tirthak Patel
SIGCSE (2)2
2024 ProxiML: Building Machine Learning Classifiers for Photonic Quantum Computing
abstract
Quantum machine learning has shown early promise and potential for productivity improvements for machine learning classification tasks, but has not been systematically explored on photonics quantum computing platforms. Therefore, this paper presents the design and implementation of ProxiML - a novel quantum machine learning classifier for photonic quantum computing devices with multiple noise-aware design elements for effective model training and inference. Our extensive evaluation on a photonic device (Xanadu's X8 machine) demonstrates the effectiveness of ProxiML machine learning classifier (over 90% accuracy on a real machine for challenging four-class classification tasks), and competitive classification accuracy compared to prior reported machine learning classifier accuracy on other quantum platforms - revealing the previously unexplored potential of Xanadu's X8 machine.
Aditya Ranjan, Tirthak Patel, Daniel Silver, Harshitta Gandhi, Devesh Tiwari
ASPLOS (3)2
2024 CodeCrunch: Improving Serverless Performance via Function Compression and Cost-Aware Warmup Location Optimization
abstract
Serverless computing has a critical problem of function cold starts. To minimize cold starts, state-of-the-art techniques predict function invocation times to warm them up. Warmed-up functions occupy space in memory and incur a keep-alive cost, which can become exceedingly prohibitive under bursty load. To address this issue, we design CodeCrunch, which introduces the concept of serverless function compression and exploits server heterogeneity to make serverless computing more efficient, especially under high memory pressure.
Rohan Basu Roy, Tirthak Patel, Rohan Garg 0001, Devesh Tiwari
ASPLOS (1)2
2024 ReCon: Reconfiguring Analog Rydberg Atom Quantum Computers for Quantum Generative Adversarial Networks
abstract
Quantum computing has shown theoretical promise of speedup in several machine learning tasks, including generative tasks using generative adversarial networks (GANs). While quantum computers have been implemented with different types of technologies, recently, analog Rydberg atom quantum computers have been demonstrated to have desirable properties such as reconfigurable qubit (quantum bit) positions and multi-qubit operations. To leverage the properties of this technology, we propose ReCon, the first work to implement quantum GANs on analog Rydberg atom quantum computers. Our evaluation using simulations and real-computer executions shows 33% better quality (measured using Frechet Inception Distance (FID)) in generated images than the state-of-the-art technique implemented on superconducting-qubit technology.
Nicholas S. DiBrita, Daniel Leeds, Yuqian Huo, Jason Ludmir, Tirthak Patel
ICCAD5
2024 Parallax: A Compiler for Neutral Atom Quantum Computers under Hardware Constraints
abstract
Among different quantum computing technologies, neutral atom quantum computers have several advantageous features, such as multi-qubit gates, application-specific topologies, movable qubits, homogenous qubits, and long-range interactions. However, existing compilation techniques for neutral atoms fall short of leveraging these advantages in a practical and scalable manner. This paper introduces PARALLAX, a zero-SWAP, scalable, and parallelizable compilation and atom movement scheduling method tailored for neutral atom systems, which reduces high-error operations by $25 \%$ and increases the success rate by $\mathbf{2 8 \%}$ on average compared to the state-of-the-art technique.
Jason Ludmir, Tirthak Patel
SC2
2024 LexiQL: Quantum Natural Language Processing on NISQ-era Machines
abstract
The rapid evolution of quantum hardware is propelling quantum computing to new frontiers. Nonetheless, the potential of natural language processing in the quantum paradigm (QNLP) is yet to be explored, including for Noisy Intermediate-Scale Quantum (NISQ) machines. To explore the QNLP frontier, we introduce LEXIQL, a novel noise-aware QNLP technique for text classification on NISQ quantum machines. LEXIQL employs an incremental data injection approach to process textual data in a quantum circuit. It also develops new and effective training methods, such as leveraging a diverse mix of expressible and shallow quantum circuits for the QNLP task of text classification. Our extensive evaluation using Yelp, IMDB, and Amazon datasets (along with synthetic QLNP datasets) demonstrates the effectiveness of LEXIQL’s noise-aware design in both ideal and noisy environments.
Daniel Silver, Aditya Ranjan, Rakesh Achutha, Tirthak Patel, Devesh Tiwari
SC4
2023 SLIQ: Quantum Image Similarity Networks on Noisy Quantum Computers
abstract
Exploration into quantum machine learning has grown tremendously in recent years due to the ability of quantum computers to speed up classical programs. However, these ef- forts have yet to solve unsupervised similarity detection tasks due to the challenge of porting them to run on quantum com- puters. To overcome this challenge, we propose SLIQ, the first open-sourced work for resource-efficient quantum sim- ilarity detection networks, built with practical and effective quantum learning and variance-reducing algorithms.
Daniel Silver, Tirthak Patel, Aditya Ranjan, Harshitta Gandhi, William Cutler, Devesh Tiwari
AAAI2
2023 Invited: Building Robust Quantum System Software for Technology-Specific Characteristics
abstract
This paper discusses the various technologies used for quantum computing and highlights the need for compiler and software stack solutions that are portable across different technologies (beyond superconducting qubit quantum computers) while providing a higher-level interface that allows scientists to run their programs in a technology-agnostic manner. To achieve this, quantum compilers and architecture designs must not be bound by classical-style standards and specifications. As a first step toward tackling this challenge, this paper then focuses on developing a compiler solution for neutral atom quantum computing technology, which has several potential benefits over superconducting qubit quantum computing technology.These benefits include a greater connectivity of qubits within the Rydberg interaction radius, which allows for fewer SWAP operations, and the ability to execute multi-qubit gates directly. However, neutral atom quantum computers have a different set of constraints and requirements, including interaction blockades, which can result in potential serialization of operations, reducing some of the gains due to the better connectivity of neutral atom quantum computers. The paper then concludes by stating that addressing these challenges requires further research and development in the field.
Tirthak Patel, Devesh Tiwari
DAC1
2023 ProPack: Executing Concurrent Serverless Functions Faster and Cheaper
abstract
The serverless computing model has been on the rise in recent years due to a lower barrier to entry and elastic scalability. However, our experimental evidence suggests that multiple serverless computing platforms suffer from serious performance inefficiencies when a high number of concurrent function instances are invoked, which is a desirable capability for parallel applications. To mitigate this challenge, this paper introduces ProPack, a novel solution that provides higher performance and yields cost savings for end users running applications with high concurrency. ProPack leverages insights obtained from experimental study to build a simple and effective analytical model that mitigates the scalability bottleneck. Our evaluation on multiple serverless platforms including AWS Lambda and Google confirms that ProPack can improve average performance by 85% and save cost by 66%. ProPack provides significant improvement (over 50%) over the state-of-the-art serverless workload manager such as Pywren, and is also, effective at mitigating the concurrency bottleneck for FuncX, a recent on-premise serverless execution platform for parallel applications.
Rohan Basu Roy, Tirthak Patel, Richmond Liew, Yadu N. Babuji, Ryan Chard, Devesh Tiwari
HPDC2
2023 MosaiQ: Quantum Generative Adversarial Networks for Image Generation on NISQ Computers
abstract
Quantum machine learning and vision have come to the fore recently, with hardware advances enabling rapid advancement in the capabilities of quantum machines. Recently, quantum image generation has been explored with many potential advantages over non-quantum techniques; however, previous techniques have suffered from poor quality and robustness. To address these problems, we introduce MosaiQ a high-quality quantum image generation GAN framework that can be executed on today’s Near-term Intermediate Scale Quantum (NISQ) computers.
Daniel Silver, Aditya Ranjan, Tirthak Patel, Harshitta Gandhi, William Cutler, Devesh Tiwari
ICCV3
2023 GRAPHINE: Enhanced Neutral Atom Quantum Computing using Application-Specific Rydberg Atom Arrangement
abstract
Multiple technologies for realizing quantum computing are currently under development. Neutral atom quantum computing is one such promising technology; it offers advantages such as the ability to perform long-distance interactions and gates consisting of more than two qubits. A particular advantage it provides is the flexibility to arrange the qubits in different topologies by customizing atom layouts. We design Graphine, which, to the best of our knowledge, is the first technique to leverage this flexibility to design application-specific topologies for different quantum algorithms based on the structural characteristics of the algorithm circuits. This enables Graphine to improve key performance metrics like the number of gates and pulses by up to 56% and the probability of error by up to 42% on average over widely-used topology designs.
Tirthak Patel, Daniel Silver, Devesh Tiwari
SC1
2023 Experimental Evaluation of Xanadu X8 Photonic Quantum Computer: Error Measurement, Characterization and Implications
abstract
Among the various types of quantum computers, photonic quantum computers have shown great potential due to their high degree of scalability. However, the development of photonic quantum computers is still in its infancy, and the characterization of their performance is of critical importance to guide further improvements. In this work, we present the first characterization and insights derived from Xanadu's X8 photonic quantum computer. Our work represents an important step toward the development of practical and scalable photonic quantum computers.
Aditya Ranjan, Tirthak Patel, Harshitta Gandhi, Daniel Silver, William Cutler, Devesh Tiwari
SC2
2022 QUILT: Effective Multi-Class Classification on Quantum Computers Using an Ensemble of Diverse Quantum Classifiers
abstract
Quantum computers can theoretically have significant acceleration over classical computers; but, the near-future era of quantum computing is limited due to small number of qubits that are also error prone. QUILT is a framework for performing multi-class classification task designed to work effectively on current error-prone quantum computers. QUILT is evaluated with real quantum machines as well as with projected noise levels as quantum machines become more noise free. QUILT demonstrates up to 85% multi-class classification accuracy with the MNIST dataset on a five-qubit system.
Daniel Silver, Tirthak Patel, Devesh Tiwari
AAAI2
2022 QUEST: systematically approximating Quantum circuits for higher output fidelity
abstract
We present QUEST, a procedure to systematically generate approximations for quantum circuits to reduce their CNOT gate count. Our approach employs circuit partitioning for scalability with procedures to 1) reduce circuit length using approximate synthesis, 2) improve fidelity by running circuits that represent key samples in the approximation space, and 3) reason about approximation upper bound. Our evaluation results indicate that our approach of "dissimilar" approximations provides close fidelity to the original circuit. Overall, the results indicate that QUEST can reduce CNOT gate count by 30-80% on ideal systems and decrease the impact of noise on existing and near-future quantum systems.
Tirthak Patel, Ed Younis, Costin Iancu, Wibe de Jong, Devesh Tiwari
ASPLOS1
2022 IceBreaker: warming serverless functions better with heterogeneity
abstract
Serverless computing, an emerging computing model, relies on "warming up" functions prior to its anticipated execution for faster and cost-effective service to users. Unfortunately, warming up functions can be inaccurate and incur prohibitively expensive cost during the warmup period (i.e., keep-alive cost). In this paper, we introduce IceBreaker, a novel technique that reduces the service time and the "keep-alive" cost by composing a system with heterogeneous nodes (costly and cheaper). IceBreaker does so by dynamically determining the cost-effective node type to warm up a function based on the function's time-varying probability of the next invocation. By employing heterogeneity, IceBreaker allows for more number of nodes under the same cost budget and hence, keeps more number of functions warm and reduces the wait time during high load. Our real-system evaluation confirms that IceBreaker reduces the overall keep-alive cost by 45% and execution time by 27% using representative serverless applications and industry-grade workload trace. IceBreaker is the first technique to employ and leverage the idea of mixing expensive and cheaper nodes to improve both service time and keep-alive cost for serverless functions -- opening up a new research avenue of serverless computing on heterogeneous servers for researchers and practitioners.
Rohan Basu Roy, Tirthak Patel, Devesh Tiwari
ASPLOS2
2022 MISO: exploiting multi-instance GPU capability on multi-tenant GPU clusters
abstract
GPU technology has been improving at an expedited pace in terms of size and performance, empowering HPC and AI/ML researchers to advance the scientific discovery process. However, this also leads to inefficient resource usage, as most GPU workloads, including complicated AI/ML models, are not able to utilize the GPU resources to their fullest extent - encouraging support for GPU multi-tenancy. We propose MISO, a technique to exploit the Multi-Instance GPU (MIG) capability on the latest NVIDIA datacenter GPUs (e.g., A100, H100) to dynamically partition GPU resources among co-located jobs. MISO's key insight is to use the lightweight, more flexible Multi-Process Service (MPS) capability to predict the best MIG partition allocation for different jobs, without incurring the overhead of implementing them during exploration. Due to its ability to utilize GPU resources more efficiently, MISO achieves 49% and 16% lower average job completion time than the unpartitioned and optimal static GPU partition schemes, respectively.
Baolin Li 0001, Tirthak Patel, Siddharth Samsi, Vijay Gadepally, Devesh Tiwari
SoCC2
2022 What does Inter-Cluster Job Submission and Execution Behavior Reveal to Us?
abstract
Modern High Performing Computing (HPC) facil-ities have multiple computing clusters that serve different pur-poses. These include large-scale computing clusters and smaller data visualization and analysis clusters, which are meant to shift the load of data analytics jobs from the large-scale systems. We perform the first in-depth characterization of cross-cluster behavior of users and jobs and provide an analysis of three inter-related systems at the Argonne Leadership Computing Facility (ALCF). Our analysis reveals interesting trends related to the resource utilization and predictability of user and job behavior across different clusters.
Tirthak Patel, Devesh Tiwari, Rajkumar Kettimuthu, William E. Allcock, Paul M. Rich, Zhengchun Liu
CLUSTER1
2022 OPTIC: A Practical Quantum Binary Classifier for Near-Term Quantum Computers
abstract
Quantum computers can theoretically speed up optimization workloads such as variational machine learning and classification workloads over classical computers. However, in practice, proposed variational algorithms have not been able to run on existing quantum computers for practical-scale problems owing to their error-prone hardware. We propose Optic, a framework to effectively execute quantum binary classification on real noisy intermediate-scale quantum (NISQ) computers.
Tirthak Patel, Daniel Silver, Devesh Tiwari
DATE1
2022 AI-Enabling Workloads on Large-Scale GPU-Accelerated System: Characterization, Opportunities, and Implications
abstract
Production high-performance computing (HPC) systems are adopting and integrating GPUs into their design to accommodate artificial intelligence (AI), machine learning, and data visualization workloads. To aid with the design and operations of new and existing GPU-based large-scale systems, we provide a detailed characterization of system operations, job characteristics, user behavior, and trends on a contemporary GPU-accelerated production HPC system. Our insights indicate that the pre-mature phases in modern AI workflow take up significant GPU hours while underutilizing GPUs, which opens up the opportunity for a multi-tier system. Finally, we provide various potential recommendations and areas for future investment for system architects, operators, and users.
Baolin Li 0001, Rohin Arora, Siddharth Samsi, Tirthak Patel, William Arcand, David Bestor, Chansup Byun, Rohan Basu Roy, Bill Bergeron, John T. Holodnak, Michael Houle 0001, Matthew Hubbell, Michael Jones 0001, Jeremy Kepner, Anna Klein, Peter Michaleas, Joseph McDonald, Lauren Milechin, Julie Mullen, Andrew Prout, Benjamin Price, Albert Reuther, Antonio Rosa, Matthew L. Weiss, Charles Yee, Daniel Edelman, Allan Vanterpool, Anson Cheng, Vijay Gadepally, Devesh Tiwari
HPCA4
2022 Geyser: a compilation framework for quantum computing with neutral atoms
abstract
Compared to widely-used superconducting qubits, neutral-atom quantum computing technology promises potentially better scalability and flexible arrangement of qubits to allow higher operation parallelism and more relaxed cooling requirements. The high performance computing (HPC) and architecture community is beginning to design new solutions to take advantage of neutral-atom quantum architectures and overcome its unique challenges.
Tirthak Patel, Daniel Silver, Devesh Tiwari
ISCA1
2022 Mashup: making serverless computing useful for HPC workflows via hybrid execution
abstract
This work introduces Mashup, a novel strategy to leverage serverless computing model for executing scientific workflows in a hybrid fashion by taking advantage of both the traditional VM-based cloud computing platform and the emerging serverless platform. Mashup outperforms the state-of-the-art workflow execution engines by an average of 34% and 43% in terms of execution time reduction and cost reduction, respectively, for widely-used HPC workflows on the Amazon Cloud platform (EC2 and Lambda).
Rohan Basu Roy, Tirthak Patel, Vijay Gadepally, Devesh Tiwari
PPoPP2
2022 Charter: Identifying the Most-Critical Gate Operations in Quantum Circuits via Amplified Gate Reversibility
abstract
When quantum programs are executed on noisy intermediate-scale quantum (NISQ) computers, they experience hardware noise; consequently, the program outputs are often erroneous. To mitigate the adverse effects of hardware noise, it is necessary to understand the effect of hardware noise on the program output and more fundamentally, understand the impact of hardware noise on specific regions within a quantum program. Identifying and optimizing regions that are more noise-sensitive is the key to expanding the capabilities of NISQ computers. Toward achieving that goal, we propose Charter, a novel technique to pinpoint specific gates and regions within a quantum program that are the most affected by the hardware noise and that have the highest impact on the program output. Using Charter's methodology, programmers can obtain a precise understanding of how different components of their code affect the output and optimize those components without the need for non-scalable quantum simulation on classical computers.
Tirthak Patel, Daniel Silver, Devesh Tiwari
SC1
2022 DayDream: Executing Dynamic Scientific Workflows on Serverless Platforms with Hot Starts
abstract
HPC applications are increasingly being designed as dynamic workflows for the ease of development and scaling. This work demonstrates how the serverless computing model can be leveraged for efficient execution of complex, real-world scientific workflows, although serverless computing was not originally designed for executing scientific workflows. This work characterizes, quantifies, and improves the execution of three real-world, complex, dynamic scientific workflows: ExaFEL (workflow for investigating the molecular structures via X-Ray diffraction), Cosmoscout-Vr(workflow for large scale virtual reality simulation), and Core Cosmology Library (a cosmology workflow for investigating dark matter). The proposed technique, DayDream, employs the hot start mechanism for warming up the components of the workflows by decoupling the runtime environment from the component function code to mitigate cold start overhead. DayDream optimizes the service time and service cost jointly to reduce the service time by 45% and service cost by 23% over the state-of-the-art HPC workload manager.
Rohan Basu Roy, Tirthak Patel, Devesh Tiwari
SC2
2021 Qraft: reverse your Quantum circuit and know the correct program output
abstract
Current Noisy Intermediate-Scale Quantum (NISQ) computers are useful in developing the quantum computing stack, test quantum algorithms, and establish the feasibility of quantum computing. However, different statistically significant errors permeate NISQ computers. To reduce the effect of these errors, recent research has focused on effective mapping of a quantum algorithm to a quantum computer in an error-and-constraints-aware manner. We propose the first work, QRAFT, to leverage the reversibility property of quantum algorithms to considerably reduce the error beyond the reduction achieved by effective circuit mapping.
Tirthak Patel, Devesh Tiwari
ASPLOS1
2021 Examining Failures and Repairs on Supercomputers with Multi-GPU Compute Nodes
abstract
Understanding the reliability characteristics of supercomputers has been a key focus of the HPC and dependability communities. However, there is no current study that analyzes both the failure and recovery characteristics over multiple generations of a GPU-based supercomputer with multiple GPUs on the same node. This paper bridges that gap and reveals surprising insights based on monitoring and analyzing the failures and repairs on the Tsubame-2 and Tsubame-3 supercomputers.
Amir Taherin, Tirthak Patel, Giorgis Georgakoudis, Ignacio Laguna, Devesh Tiwari
DSN2
2021 Operating Liquid-Cooled Large-Scale Systems: Long-Term Monitoring, Reliability Analysis, and Efficiency Measures
abstract
The past decade has seen a rise in the use of liquid cooling due to its energy efficiency. While many previous works have helped make progress toward improving data center cooling, a vast majority of them perform studies on a small system over a short span. The computer systems and HPC community lacks a long-term study highlighting the challenges and solutions in operating a liquid-cooled large-scale data center. We conduct the first detailed characterization of a petascale supercomputer, Mira, over a span of six years. The study is enabled by systematic monitoring of the environmental metrics, and discusses new research avenues, including coolant monitor failures.
Rohan Basu Roy, Tirthak Patel, Rajkumar Kettimuthu, William E. Allcock, Paul M. Rich, Adam Scovel, Devesh Tiwari
HPCA2
2021 SATORI: Efficient and Fair Resource Partitioning by Sacrificing Short-Term Benefits for Long-Term Gains*
abstract
Multi-core architectures have enabled data centers to increasingly co-locate multiple jobs to improve resource utilization and lower the operational cost. Unfortunately, naively co-locating multiple jobs may lead to only a modest increase in system throughput. Worse, some users may observe proportionally higher performance degradation compared to other users co-located on the same physical multi-core system. SATORI is a novel strategy to partition multi-core architectural resources to achieve two conflicting goals simultaneously: increasing system throughput and achieving fairness among the co-located jobs.
Rohan Basu Roy, Tirthak Patel, Devesh Tiwari
ISCA2
2021 Bliss: auto-tuning complex applications using a pool of diverse lightweight learning models
abstract
As parallel applications become more complex, auto-tuning becomes more desirable, challenging, and time-consuming. We propose, Bliss, a novel solution for auto-tuning parallel applications without requiring apriori information about applications, domain-specific knowledge, or instrumentation. Bliss demonstrates how to leverage a pool of Bayesian Optimization models to find the near-optimal parameter setting 1.64× faster than the state-of-the-art approaches.
Rohan Basu Roy, Tirthak Patel, Vijay Gadepally, Devesh Tiwari
PLDI2
2021 Systematically inferring I/O performance variability by examining repetitive job behavior
abstract
Monitoring and analyzing I/O behaviors is critical to the efficient utilization of parallel storage systems. Unfortunately, with increasing I/O requirements and resource contention, I/O performance variability is becoming a significant concern. This paper investigates I/O behavior and performance variability on a large-scale high-performance computing (HPC) system using a novel methodology that identifies similarity across jobs from the same application leveraging an I/O characterization tool and then, detects potential I/O performance variability across jobs of the same application. We demonstrate and discuss how our unique methodology can be used to perform temporal and feature analyses to detect interesting I/O performance variability patterns in production HPC systems, and their implications for operating/managing large-scale systems.
Emily Costa, Tirthak Patel, Benjamin Schwaller, Jim M. Brandt, Devesh Tiwari
SC2
2021 RIBBON: cost-effective and qos-aware deep learning model inference using a diverse pool of cloud computing instances
abstract
Deep learning model inference is a key service in many businesses and scientific discovery processes. This paper introduces Ribbon, a novel deep learning inference serving system that meets two competing objectives: quality-of-service (QoS) target and cost-effectiveness. The key idea behind Ribbon is to intelligently employ a diverse set of cloud computing instances (heterogeneous instances) to meet the QoS target and maximize cost savings. Ribbon devises a Bayesian Optimization-driven strategy that helps users build the optimal set of heterogeneous instances for their model inference service needs on cloud computing platforms - and, Ribbon demonstrates its superiority over existing approaches of inference serving systems using homogeneous instance pools. Ribbon saves up to 16% of the inference service cost for different learning models including emerging deep learning recommender system models and drug-discovery enabling models.
Baolin Li 0001, Rohan Basu Roy, Tirthak Patel, Vijay Gadepally, Karen Gettings, Devesh Tiwari
SC3
2021 Study of interconnect errors, network congestion, and applications characteristics for throttle prediction on a large scale HPC system
Saurabh Gupta 0002, Tirthak Patel, Michael Wilder, Weisong Shi, Song Fu, Christian Engelmann, Devesh Tiwari
J. Parallel Distributed Comput.3
2020 Making Disk Failure Predictions SMARTer!
Sidi Lu, Tirthak Patel, Yongtao Yao, Devesh Tiwari, Weisong Shi
FAST3
2020 GIFT: A Coupon Based Throttle-and-Reward Mechanism for Fair and Efficient I/O Bandwidth Management on Parallel Storage Systems
Tirthak Patel, Rohan Garg 0001, Devesh Tiwari
FAST1
2020 Uncovering Access, Reuse, and Sharing Characteristics of I/O-Intensive Files on Large-Scale Production HPC Systems
Tirthak Patel, Surendra Byna, Glenn K. Lockwood, Nicholas J. Wright, Philip H. Carns, Robert B. Ross, Devesh Tiwari
FAST1
2020 CLITE: Efficient and QoS-Aware Co-Location of Multiple Latency-Critical Jobs for Warehouse Scale Computers
abstract
Large-scale data centers run latency-critical jobs with quality-of-service (QoS) requirements, and throughput-oriented background jobs, which need to achieve high perfor-mance. Previous works have proposed methods which cannot co-locate multiple latency-critical jobs with multiple back-grounds jobs while: (1) meeting the QoS requirements of all latency-critical jobs, and (2) maximizing the performance of the background jobs. This paper proposes CLITE, a Bayesian Optimization-based, multi-resource partitioning technique which achieves these goals. CLITE is publicly available at https://github.com/GoodwillComputingLab/CLITE.
Tirthak Patel, Devesh Tiwari
HPCA1
2020 DisQ: A Novel Quantum Output State Classification Method on IBM Quantum Computers using OpenPulse
abstract
Superconducting quantum computing technology has ushered in a new era of computational possibilities. While a considerable research effort has been geared toward improving the quantum technology and building the software stack to efficiently execute quantum algorithms with reduced error rate, effort toward optimizing how quantum output states are defined and classified for the purpose of reducing the error rate is still limited. To this end, this paper proposes DisQ, a quantum output state classification approach which reduces error rates of quantum programs on NISQ devices.
Tirthak Patel, Devesh Tiwari
ICCAD1
2020 What does Power Consumption Behavior of HPC Jobs Reveal? : Demystifying, Quantifying, and Predicting Power Consumption Characteristics
abstract
As we approach exascale computing, large-scale HPC systems are becoming increasingly power-constrained, requiring them to run HPC workloads in an energy-efficient manner. The first step toward achieving this goal is to better understand, analyze, and quantify the power consumption characteristics of HPC jobs. However, there is a lack of understanding of the power consumption characteristics of HPC jobs which run on production HPC systems. Such characterization is required to guide the design of the next generation of power-aware resource management. To the best of our knowledge, we are the first study to open-source the data and analysis of power-consumption characteristics of HPC jobs and users from two medium-scale production HPC clusters.
Tirthak Patel, Adam Wagenhäuser, Christopher Eibel, Timo Hönig, Thomas Zeiser, Devesh Tiwari
IPDPS1
2020 Job characteristics on large-scale systems: long-term analysis, quantification, and implications
abstract
HPC workload analysis and resource consumption characteristics are the key to driving better operation practices, system procurement decisions, and designing effective resource management techniques. Unfortunately, the HPC community does not have easy accessibility to long-term introspective work-load analysis and characterization for production-scale HPC systems. This study bridges this gap by providing detailed long-term quantification, characterization, and analysis of job characteristics on two supercomputers: Intrepid and Mira. This study is one of the largest of its kind - covering trends and characteristics for over three billion compute hours, 750 thousand jobs, and spanning a decade. We confirm several long-held conventional wisdom, and identify many previously undiscovered trends and its implications. We also introduce a learning based technique to predict the resource requirement of future jobs with high accuracy, using features available prior to the job submission and without requiring any application-specific tracing or application-intrusive instrumentation.
Tirthak Patel, Zhengchun Liu, Rajkumar Kettimuthu, Paul M. Rich, William E. Allcock, Devesh Tiwari
SC1
2020 Experimental evaluation of NISQ quantum computers: error measurement, characterization, and implications
abstract
Noisy Intermediate-Scale Quantum (NISQ) computers are being increasingly used for executing early-stage quantum programs to establish the practical realizability of existing quantum algorithms. These quantum programs have uses cases in the realm of high-performance computing ranging from molecular chemistry and physics simulations to addressing NP-complete optimization problems. However, NISQ devices are prone to multiple types of errors, which affect the fidelity and reproducibility of the program execution. As the technology is still primitive, our understanding of these quantum machines and their error characteristics is limited. To bridge that understanding gap, this is the first work to provide a systematic and rich experimental evaluation of IBM Quantum Experience (QX) quantum computers of different scales and topologies. Our experimental evaluation uncovers multiple important and interesting aspects of benchmarking and evaluating quantum program on NISQ machines. We have open-sourced our experimental framework and dataset to help accelerate the evaluation of quantum computing systems.
Tirthak Patel, Abhay Potharaju, Baolin Li 0001, Rohan Basu Roy, Devesh Tiwari
SC1
2020 Veritas: accurately estimating the correct output on noisy intermediate-scale quantum computers
abstract
Noisy Intermediate-Scale Quantum (NISQ) machines are being increasingly used to develop quantum algorithms and establish use cases for quantum computing. However, these devices are highly error-prone and produce output, which can be far from the correct output of the quantum algorithm. In this paper, we propose VERITAS, an end-to-end approach toward designing quantum experiments, executing experiments, and correcting outputs produced by quantum circuits post their execution such that the correct output of the quantum algorithm can be accurately estimated.
Tirthak Patel, Devesh Tiwari
SC1
2020 UREQA: Leveraging Operation-Aware Error Rates for Effective Quantum Circuit Mapping on NISQ-Era Quantum Computers
Tirthak Patel, Baolin Li 0001, Rohan Basu Roy, Devesh Tiwari
USENIX ATC1
2019 What does Vibration do to Your SSD?
abstract
Vibration generated in modern computing environments such as autonomous vehicles, edge computing infrastructure, and data center systems is an increasing concern. In this paper, we systematically measure, quantify and characterize the impact of vibration on the performance of SSD devices. Our experiments and analysis uncover that exposure to both short-term and long-term vibration, even within the vendor-specified limits, can significantly affect SSD I/O performance and reliability.
Janki Bhimani, Tirthak Patel, Ningfang Mi, Devesh Tiwari
DAC2
2019 PERQ: Fair and Efficient Power Management of Power-Constrained Large-Scale Computing Systems
abstract
Large-scale computing systems are becoming increasingly more power-constrained, but these systems employ hardware over- provisioning to achieve higher system throughput because applications often do not consume the peak power capacity of nodes. Unfortunately, focusing on system throughput alone can lead to severe unfairness among multiple concurrently-running applications. This paper introduces PERQ, a new feedback-based principled approach to improve system throughput while achieving fairness among concurrent applications.
Tirthak Patel, Devesh Tiwari
HPDC1
2019 Revisiting I/O behavior in large-scale storage systems: the expected and the unexpected
abstract
Large-scale applications typically spend a large fraction of their execution time performing I/O to a parallel storage system. However, with rapid progress in compute and storage system stack of large-scale systems, it is critical to investigate and update our understanding of the I/O behavior of large-scale applications. Toward that end, in this work, we monitor, collect and analyze a year worth of storage system data from a large-scale production parallel storage system. We perform temporal, spatial and correlative analysis of the system and uncover surprising patterns which defy existing assumptions and have important implications for future systems.
Tirthak Patel, Surendra Byna, Glenn K. Lockwood, Devesh Tiwari
SC1
2018 Shiraz: Exploiting System Reliability and Application Resilience Characteristics to Improve Large Scale System Throughput
abstract
Large-scale applications rely on resilience mechanisms such as checkpoint-restart to make forward progress in the presence of failures. Unfortunately, this incurs huge I/O overhead and impedes productivity. To mitigate this challenge, this paper introduces a new technique, Shiraz, which demonstrates how to exploit differences in the checkpointing overhead among applications and knowledge of temporal characteristics of failures to improve both the overall system throughput and performance of individual applications.
Rohan Garg 0001, Tirthak Patel, Gene Cooperman, Devesh Tiwari
DSN2
2018 Understanding and Analyzing Interconnect Errors and Network Congestion on a Large Scale HPC System
abstract
Today's High Performance Computing (HPC) systems are capable of delivering performance in the order of petaflops due to the fast computing devices, network interconnect, and back-end storage systems. In particular, interconnect resilience and congestion resolution methods have a major impact on the overall interconnect and application performance. This is especially true for scientific applications running multiple processes on different compute nodes as they rely on fast network messages to communicate and synchronize frequently. Unfortunately, the HPC community lacks state-of-practice experience reports that detail how different interconnect errors and congestion events occur on large-scale HPC systems. Therefore, in this paper, we process and analyze interconnect data of the Titan supercomputer to develop a thorough understanding of interconnects faults, errors and congestion events. We also study the interaction between interconnect, errors, network congestion and application characteristics.
Saurabh Gupta 0002, Tirthak Patel, Michael Wilder, Weisong Shi, Song Fu, Christian Engelmann, Devesh Tiwari
DSN3
2018 Machine Learning Models for GPU Error Prediction in a Large Scale HPC System
abstract
GPUs are widely deployed on large-scale HPC systems to provide powerful computational capability for scientific applications from various domains. As those applications are normally long-running, investigating the characteristics of GPU errors becomes imperative for reliability. In this paper, we first study the system conditions that trigger GPU errors using six-month trace data collected from a large-scale, operational HPC system. Then, we use machine learning to predict the occurrence of GPU errors, by taking advantage of temporal and spatial dependencies of the trace data. The resulting machine learning prediction framework is robust and accurate under different workloads.
Bin Nie, Ji Xue, Saurabh Gupta 0002, Tirthak Patel, Christian Engelmann, Evgenia Smirni, Devesh Tiwari
DSN4
2017 Failures in large scale systems: long-term measurement, analysis, and implications
abstract
Resilience is one of the key challenges in maintaining high efficiency of future extreme scale supercomputers. Researchers and system practitioners rely on field-data studies to understand reliability characteristics and plan for future HPC systems. In this work, we compare and contrast the reliability characteristics of multiple large-scale HPC production systems. Our study covers more than one billion compute node hours across five different systems over a period of 8 years. We confirm previous findings which continue to be valid, discover new findings, and discuss their implications.
Saurabh Gupta 0002, Tirthak Patel, Christian Engelmann, Devesh Tiwari
SC2