VLDB 2026 Research / reviewers in the wild / expert
Ritvik Sharma
dblp:283/0097
· DBLP profile ↗
11ranked-venue papers
4as first author
10since 2021 · last 2026
0000-0002-5809-7031ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 1 first-author · 7 since 2021Software engineering, systems software and programming languages · 6 · 3 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FuseFlow: A Fusion-Centric Compilation Framework for Sparse Deep Learning on Streaming DataflowabstractAs deep learning models scale, sparse deep learning (DL) models that exploit sparsity in weights, activations, or inputs and specialized dataflow hardware have emerged as powerful solutions to address efficiency. We propose FuseFlow, a compiler that converts sparse machine learning models written in PyTorch to fused sparse dataflow graphs for reconfigurable dataflow architectures (RDAs). FuseFlow is the first compiler to support general cross-expression fusion of sparse operations. In addition to fusion across kernels (expressions), FuseFlow also supports optimizations like parallelization, dataflow ordering, and sparsity blocking. It targets a cycle-accurate dataflow simulator for microarchitectural analysis of fusion strategies. We use FuseFlow for design-space exploration across four real-world machine learning applications with sparsity, showing that full fusion (entire cross-expression fusion across all computation in an end-to-end model) is not always optimal for sparse models—fusion granularity depends on the model itself. FuseFlow also provides a heuristic to identify and prune suboptimal configurations. Using FuseFlow, we achieve performance improvements, including a ~2.7x speedup over an unfused baseline for GPT-3 with BigBird block-sparse attention. Rubens Lacouture, Nathan Zhang, Ritvik Sharma, Marco Siracusa, Fredrik Kjolstad, Kunle Olukotun, Olivia Hsu |
ASPLOS (2) | 3 |
| 2026 | CoTenN: Constrained Optimization with Tensor NetworksabstractSimulation of physics problems is one of the most important use cases of quantum computing. For this class of problems, the goal is typically to find the minimum energy state, or the ground state, of a physical system’s Hamiltonian. These problems frequently have constraints, such as symmetry conditions, which must also be satisfied. To solve such problems, researchers in computational physics use quantum-inspired algorithms that execute on classical computers. In particular, tensor network-based eigensolvers such as DMRG have become popular. However, to use these eigensolvers, the constrained optimization problem must first be encoded as a tensor network that implements a low-rank decomposition of the system’s Hamiltonian and state vector. These tensor network encodings are highly flexible, allowing for variables with ≥ 2 quantum states and supporting efficient constraint encodings that directly constrain the state vector. A critical challenge to developing tensor network-based encodings is that, currently, the encoding process is manual; significant effort is required to identify an efficient encoding for a new physics problem. In this work, we introduce a quantum constrained optimization problem (QCOP), a general abstraction for describing minimization problems over quantum variables that are subject to hard constraints. We present Masq, the first constraint programming language for QCOPs implementable with tensor networks, and CoTenN, a compiler that automatically maps QCOPs specified with Masq programs to tensor networks. To demonstrate the utility of Masq and CoTenN, we formulate two physics problems in Masq and then use CoTenN to find their ground states. We find the CoTenN-generated tensor networks generally outperform SOTA problem formulations, providing between 2.05×–53.32× total speedups across runs for QCOPs and yielding up to 2.49 · 10 7 × lower truncation errors for otherwise unconstrained problems. Ritvik Sharma, Siddharth Dangwal, Sara Achour |
Proc. ACM Program. Lang. | 1 |
| 2025 | A Probabilistic Perspective on Tiling Sparse Tensor AlgebraabstractSparse tensor algebra computations are often memory-bound due to irregular access patterns and low arithmetic intensity.We present D2T2 (Data-Driven Tensor Tiling), a framework that optimizes static coordinate-space tiling schemes to minimize memory traffic by identifying and leveraging relevant high-level statistics from input operands.For a given tensor algebra computation, D2T2 collects statistics from input tensors, builds a probability distribution-based model of the tensor computation, and uses it to predict traffic for various tiling configurations.It searches over tile shape and size configurations to minimize total traffic.We evaluate D2T2 against Tailors and DRT, two state of the art tiling schemes for sparse tensor algebra.We find that D2T2 achieves, on average, a 2.54× speedup over Tailors and a 1.13× lower memory bandwidth compared to DRT for sparse-sparse matrix multiplication (SpMSpM).We also achieve 1.22-48.94×lower bandwidth for SpMSpM and up to 34.31× lower bandwidth for tensor operations (TTM and MTTKRP) than conservative static tiling schemes.Unlike prior tiling techniques, D2T2 is deployable without specialized hardware support.On Opal, a 16nm sparse tensor algebra accelerator, D2T2 generated tiling configurations that achieve 1.23-3.34×speedups compared to their original hand-tuned configurations. Ritvik Sharma, Zi Yu Xue, Nathan Zhang, Rubens Lacouture, Fredrik Kjolstad, Sara Achour, Mark Horowitz |
MICRO | 1 |
| 2025 | Optimizing Ancilla-Based Quantum Circuits with SPAREabstractMany quantum algorithms instantiate and use ancillas, spare qubits that serve as temporary storage in a quantum circuit. In particular, many recently developed high-level and modular quantum programming languages (QPLs) use ancilla qubits to implement various programming constructs. These are lowered to circuits with nested/cascading compute-uncompute gate sequences that use ancilla qubits to track internal state. We present SPARE, a rewrite-based quantum circuit optimizer that restructures these compute-uncompute gate sequences, leveraging the ancilla qubit state information to optimize the circuit. In this work, we prove the correctness of SPARE’s rewrites and link SPARE’s gate-level transforms to language-level program rewrites, which may be performed on the input language. We evaluate SPARE on QPL-generated quantum circuits against Unqomp and Spire, two optimizing compilers for QPLs. SPARE achieves a reduction of up to 27.3 % in qubit count, 56.7 % in 2 -qubit gates, 68.2 % in 1-qubit gates and 73.9 % in depth against Unqomp, and up to 17.8 % in qubits, 67.3 % in 2-qubit gates, 61.4 % in 1-qubit gates and 59.9 % in depth against Spire. We also evaluate SPARE against the Quartz, Feynman, and PyZX circuit optimizers: SPARE achieves up to a 70.0 % reduction in two-qubit gates, up to a 53.6 % reduction in 1-qubit gates, and up to a 56.7 % reduction in depth compared to the best result from all the gate-level optimizers. Ritvik Sharma, Sara Achour |
Proc. ACM Program. Lang. | 1 |
| 2024 | Onyx: A Programmable Accelerator for Sparse Tensor Algebraabstract•Applications ranging from scientific computing to machine learning can have extremely sparse inputs Kalhan Koul, Maxwell Strange, Jackson Melchert, Alex Carsello, Yuchen Mei, Olivia Hsu, Taeyoung Kong, Huifeng Ke, Keyi Zhang, Qiaoyi Liu, Gedeon Nyengele, Akhilesh Balasingam, Jayashree Adivarahan, Ritvik Sharma, Zhouhua Xie, Christopher Torng, Joel S. Emer, Fredrik Kjolstad, Mark Horowitz, Priyanka Raina |
HCS | 15 |
| 2024 | Compilation of Qubit Circuits to Optimized Qutrit CircuitsabstractQuantum computers are a revolutionary class of computational platforms that are capable of solving computationally hard problems. However, today’s quantum hardware is subject to noise and decoherence issues that together limit the scale and complexity of the quantum circuits that can be implemented. Recently, practitioners have developed qutrit-based quantum hardware platforms that compute over ∣ 0 ⟩ , ∣ 1 ⟩ , and ∣ 2 ⟩ states, and have presented circuit depth reduction techniques using qutrits’ higher energy ∣ 2 ⟩ states to temporarily store information. However, thus far, such quantum circuits that use higher order states for temporary storage need to be manually crafted by hardware designers. We present D are , an optimizing compiler for qutrit circuits that implement qubit computations. D are deploys a qutrit circuit decomposition algorithm and a rewrite engine to construct and optimize qutrit circuits. We evaluate D are against hand-optimized qutrit circuits and qubit circuits, and find D are delivers up to 65 % depth improvement over manual qutrit implementations, and 43-75% depth improvement over qubit circuits. We also perform a fidelity analysis and find DARE-optimized qutrit circuits deliver up to 8.9 × higher fidelity circuits than their manually implemented counterparts. Ritvik Sharma, Sara Achour |
Proc. ACM Program. Lang. | 1 |
| 2023 | The Sparse Abstract MachineabstractWe propose the Sparse Abstract Machine (SAM), an abstract machine model for targeting sparse tensor algebra to reconfigurable and fixed-function spatial dataflow accelerators. SAM defines a streaming dataflow abstraction with sparse primitives that encompass a large space of scheduled tensor algebra expressions. SAM dataflow graphs naturally separate tensor formats from algorithms and are expressive enough to incorporate arbitrary iteration orderings and many hardware-specific optimizations. We also present Custard, a compiler from a high-level language to SAM that demonstrates SAM's usefulness as an intermediate representation. We automatically bind from SAM to a streaming dataflow simulator. We evaluate the generality and extensibility of SAM, explore the performance space of sparse tensor algebra optimizations using SAM, and show SAM's ability to represent dataflow hardware. Olivia Hsu, Maxwell Strange, Ritvik Sharma, Jaeyeon Won, Kunle Olukotun, Joel S. Emer, Mark Horowitz, Fredrik Kjolstad |
ASPLOS (3) | 3 |
| 2023 | APEX: A Framework for Automated Processing Element Design Space Exploration using Frequent Subgraph AnalysisabstractThe architecture of a coarse-grained reconfigurable array (CGRA) processing element (PE) has a significant effect on the performance and energy-efficiency of an application running on the CGRA. This paper presents APEX, an automated approach for generating specialized PE architectures for an application or an application domain. APEX first analyzes application domain benchmarks using frequent subgraph mining to extract commonly occurring computational subgraphs. APEX then generates specialized PEs by merging subgraphs using a datapath graph merging algorithm. The merged datapath graphs are translated into a PE specification from which we automatically generate the PE hardware description in Verilog along with a compiler that maps applications to the PE. The PE hardware and compiler are inserted into a flexible CGRA generation and compilation toolchain that allows for agile evaluation of CGRAs. We evaluate APEX for two domains, machine learning and image processing. For image processing applications, our automatically generated CGRAs with specialized PEs achieve from 5% to 30% less area and from 22% to 46% less energy compared to a general-purpose CGRA. For machine learning applications, our automatically generated CGRAs consume 16% to 59% less energy and 22% to 39% less area than a general-purpose CGRA. This work paves the way for creation of application domain-driven design-space exploration frameworks that automatically generate efficient programmable accelerators, with a much lower design effort for both hardware and compiler generation. Jackson Melchert, Kathleen Feng, Caleb Donovick, Ross Daly, Ritvik Sharma, Clark W. Barrett, Mark Horowitz, Pat Hanrahan, Priyanka Raina |
ASPLOS (3) | 5 |
| 2023 | CODEBench: A Neural Architecture and Hardware Accelerator Co-Design FrameworkabstractRecently, automated co-design of machine learning (ML) models and accelerator architectures has attracted significant attention from both the industry and academia. However, most co-design frameworks either explore a limited search space or employ suboptimal exploration techniques for simultaneous design decision investigations of the ML model and the accelerator. Furthermore, training the ML model and simulating the accelerator performance is computationally expensive. To address these limitations, this work proposes a novel neural architecture and hardware accelerator co-design framework, called CODEBench. It comprises two new benchmarking sub-frameworks, CNNBench and AccelBench, which explore expanded design spaces of convolutional neural networks (CNNs) and CNN accelerators. CNNBench leverages an advanced search technique, Bayesian Optimization using Second-order Gradients and Heteroscedastic Surrogate Model for Neural Architecture Search, to efficiently train a neural heteroscedastic surrogate model to converge to an optimal CNN architecture by employing second-order gradients. AccelBench performs cycle-accurate simulations for diverse accelerator architectures in a vast design space. With the proposed co-design method, called Bayesian Optimization using Second-order Gradients and Heteroscedastic Surrogate Model for Co-Design of CNNs and Accelerators, our best CNN–accelerator pair achieves 1.4% higher accuracy on the CIFAR-10 dataset compared to the state-of-the-art pair while enabling 59.1% lower latency and 60.8% lower energy consumption. On the ImageNet dataset, it achieves 3.7% higher Top1 accuracy at 43.8% lower latency and 11.2% lower energy consumption. CODEBench outperforms the state-of-the-art framework, i.e., Auto-NBA, by achieving 1.5% higher accuracy and 34.7× higher throughput while enabling 11.0× lower energy-delay product and 4.0× lower chip area on CIFAR-10. Shikhar Tuli, Chia-Hao Li, Ritvik Sharma, Niraj K. Jha |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2021 | A Crossbar Array of Analog-Digital-Hybrid Volatile Memory Synapse Cells for Energy-Efficient On-Chip LearningabstractConventional-silicon-transistor-based Volatile Memory (VM) synapse has been proposed as an alternative to Non Volatile Memory (NVM) synapse in crossbar-array-based neuromorphic/ in-memory-computing systems. Here, through SPICE simulations, we have designed an analog-digital-hybrid Volatile Memory Synapse Cell (VMSC) for such a crossbar array of VM synapses. In our VMSC, the transistor synapse stores nearly analog values of weight. But the other transistors, which carry out the weight update for the transistor synapse, are designed following the principle of static CMOS logic (digital), making our design energy-efficient. Through system-level study, we report classification accuracy, speed, and energy consumption for on-chip learning on the VMSC-based crossbar designed here, using popular machine learning data sets. We show that despite a low value of capacitance of our MOSFET synapses (low area- footprint hence), the weights are retained in them long enough for our VMSC-based crossbar to exhibit comparable accuracy as a NVM-synapse-based crossbar. Janak Sharda, Ritvik Sharma, Debanjan Bhowmik |
ISCAS | 2 |
| 2020 | A 4.2-pJ/Conv 10-b Asynchronous ADC with Hybrid Two-Tier Level-Crossing Event CodingabstractAn asynchronous continuous-time level-crossing analog-to-digital converter (LC-ADC) for high-throughput, high-resolution applications is presented. The proposed 10-bit ADC architecture comprises two stages of level-crossing ADCs, the first stage resolving for 5 MSBs and the second folded residue stage for 5 LSBs. Gray encoding of the output bits ensure single-bit transitions between adjacent digital outputs. Compared to uniform-sampling synchronous ADCs, LC-ADCs generate fewer samples for sparse signals, useful in many applications for biomedical signal acquisition, event-driven computer vision, etc. Unlike conventional LC-ADCs with a few comparators tuned for lower power consumption to acquire sparse signals, this two-tier LC-ADC is optimized for high-resolution tracking of continuous signals, like Electrocardiogram (ECG). Designed and fabricated in 0.18-μm CMOS technology, chip area of the proposed ADC is 1310 × 125 μm2. Operating at 1.8 V supply, the ADC consumes 160-426 μW for 1 Hz to 200 kHz input frequencies at full scale amplitude and achieves an energy efficiency figure-of-merit of 4.16-pJ/conv. Rajkumar Kubendran, Jongkil Park 0001, Ritvik Sharma, Chul Kim, Siddharth Joshi 0001, Gert Cauwenberghs, Sohmyung Ha |
ISCAS | 3 |