EDBT 2026 Demo / reviewers in the wild / expert
Catherine Graves
dblp:180/3847 · also Cat Graves, Catherine E. Graves
· DBLP profile ↗
7ranked-venue papers
1as first author
4since 2021 · last 2025
0000-0002-0907-583XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | RAGO: Systematic Performance Optimization for Retrieval-Augmented Generation ServingabstractRetrieval-augmented generation (RAG) is emerging as a popular approach for reliable LLM serving.However, efficient RAG serving remains an open challenge due to the rapid emergence of many RAG variants and the substantial differences in workload characteristics across them.This paper makes three fundamental contributions to advancing RAG serving.First, we introduce RAGSchema, a structured abstraction that captures the wide range of RAG algorithms, serving as a foundation for performance optimization.Second, we analyze several representative RAG workloads with distinct RAGSchema, revealing significant performance variability across these workloads.Third, to address this variability and meet diverse performance requirements, we propose RAGO (Retrieval-Augmented Generation Optimizer), a system optimization framework for efficient RAG serving.RAGO achieves up to a 2ˆincrease in QPS per chip and a 55% reduction in time-to-first-token latency compared to RAG systems built on LLM-system extensions. Wenqi Jiang 0001, Suvinay Subramanian, Catherine Graves, Gustavo Alonso, Amir Yazdanbakhsh, Vidushi Dadu |
ISCA | 3 |
| 2024 | CAMSHAP: Accelerating Machine Learning Model Explainability with Analog CAMabstractThe recent success of machine learning (ML) models has led to increasing demands for model explanations - why a result was given - along with model predictions. Tree-based ML models are considered more explainable than deep neural networks and higher performers in several domains. However, algorithms computing model explanations are irregular and scale poorly with model size. While many custom accelerators for training and inference have been proposed, little attention has been paid to accelerating model explanations. This lack of explanatory capability has limited the use of these models for real-time decision-making systems in critical fields such as healthcare, autonomous operation and cybersecurity. John Moon, Giacomo Pedretti, Pedro Bruel, Sergey Serebryakov, Omar Eldash, Luca Buonanno, Catherine Graves, Paolo Faraboschi, Jim Ignowski |
ICCAD | 7 |
| 2023 | ReRAM-based graph attention network with node-centric edge searching and hamming similarityabstractThe graph attention network (GAT) has demonstrated its advantages via local attention mechanism but suffered from low energy and latency efficiency when implemented on conventional von-Neumann hardware. This work proposes and experimentally demonstrates an algorithm-hardware co-designed GAT that runs efficiently and reliably in ReRAM-based hardware. The neighborhood information is retrieved from trained node embeddings stored on crossbars in a single time step, and attention is implemented by efficient hashing and hamming similarity for higher robustness. Our scaled simulation based on the experimentally-validated model shows only 0.9% accuracy loss with over 35,500x energy improvement on the Cora dataset compared with GPU, and 1.1% accuracy improvement with 2× energy improvement compared with state-of-the-art ReRAM-based GNN accelerator. Ruibin Mao, Xia Sheng, Catherine Graves, Can Li 0024 |
DAC | 3 |
| 2022 | A general tree-based machine learning accelerator with memristive analog CAMabstractDeep learning models have reached high accuracy in multiple classification tasks. However these models lack explainability, namely the capability of understanding why a certain class is chosen along with the class predicted. On the other hand, tree-based models are top performers in several applications, particularly when the training set is limited, while also being more explainable. However, tree-based models are difficult to accelerate with conventional digital hardware due to irregular memory access patterns. Here we show a tree-based ML accelerator based on a novel analog content addressable memory with memristor devices, capable of handling multiple types of bagging and boosting techniques common in tree-based algorithms. Our results show a large improvement of $\sim 60 \times $ lower latency and $160 \times $ reduced energy consumption compared to the state of the art, demonstrating the promise of our accelerator approach. Giacomo Pedretti, Sergey Serebryakov, John Paul Strachan, Catherine Graves |
ISCAS | 4 |
| 2019 | High performance, power efficient hardware accelerators: emerging devices, circuits and architecture co-designabstractGeneral-purpose digital systems have long benefited from favorable scaling, but performance improvements have slowed dramatically in the last decade. Computing is therefore returning to custom and specialized systems, frequently using heterogeneous accelerators. Particularly driven by the data-centric workloads of machine learning and deep learning, an intense development of conventional accelerators (GPUs, FPGAs, CMOS ASICs) but also unconventional accelerators using novel circuits and devices beyond CMOS is currently underway. In this talk, I will discuss some common characteristics of high-performance and power-efficient accelerators in this diverse space and the ecosystem development (such as new interconnects) needed for them to thrive. To illustrate accelerator characteristics and their potential, I will describe our group's efforts to co-design from algorithms and architectures down to novel devices for gains in speed and power. We have developed architectures leveraging the analog and non-volatile nature of memristors (tunable resistance switches) assembled in crossbar arrays to accelerate machine learning, image and signal processing. We have also developed new circuits and assembled architectures to accelerate Finite Automata, enabling rapid pattern matching used in applications from security to genomics. Significant improvements over CPUs, GPUs, and custom digital ASICs are forecasted in both such systems, highlighting the potential for unconventional accelerators in future high-performance computing systems. Catherine Graves |
CF | 1 |
| 2018 | Large Memristor Crossbars for Analog ComputingabstractMemristor with tunable non-volatile resistance offers in-memory computing capability that avoids the von-Neumann bottleneck. However, large-scale experimental demonstration to this end is yet to be implemented due to the immaturity of the device and integration technologies. Here in this paper we report our recent process in analog computing using analog-voltage-amplitude-vector input and analog-memristor-conductance matrix, with applications in signal and image processing. The vector matrix multiplication is processed in the memristor crossbars in one step, with 5-8 bit precision depending on the array size. The demonstration is made possible by high memristor yield (99.8%), stable multilevel memresistance states, linear current-voltage (IV) relation in the operation range, and low wire resistance between the cells. Can Li 0024, Yunning Li, Hao Jiang 0017, J. Joshua Yang, Qiangfei Xia, Miao Hu 0002, Eric Montgomery, Noraica Dávila, Catherine Graves, John Paul Strachan, R. Stanley Williams, Ning Ge 0001, Mark Barnell, Qing Wu 0002 |
ISCAS | 13 |
| 2016 | Dot-product engine for neuromorphic computing: programming 1T1M crossbar to accelerate matrix-vector multiplicationabstractVector-matrix multiplication dominates the computation time and energy for many workloads, particularly neural network algorithms and linear transforms (e.g, the Discrete Fourier Transform). Utilizing the natural current accumulation feature of memristor crossbar, we developed the Dot-Product Engine (DPE) as a high density, high power efficiency accelerator for approximate matrix-vector multiplication. We firstly invented a conversion algorithm to map arbitrary matrix values appropriately to memristor conductances in a realistic crossbar array, accounting for device physics and circuit issues to reduce computational errors. The accurate device resistance programming in large arrays is enabled by close-loop pulse tuning and access transistors. To validate our approach, we simulated and benchmarked one of the state-of-the-art neural networks for pattern recognition on the DPEs. The result shows no accuracy degradation compared to software approach (99 % pattern recognition accuracy for MNIST data set) with only 4 Bit DAC/ADC requirement, while the DPE can achieve a speed-efficiency product of 1,000× to 10,000× compared to a custom digital ASIC. Miao Hu 0002, John Paul Strachan, Emmanuelle M. Grafals, Noraica Dávila, Catherine Graves, Sity Lam, Ning Ge 0001, J. Joshua Yang, R. Stanley Williams |
DAC | 6 |