EDBT 2026 Demo / reviewers in the wild / expert
Emily Shriver
dblp:66/496
· DBLP profile ↗
14ranked-venue papers
0as first author
5since 2021 · last 2026
0009-0003-2135-6428ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 3 since 2021Software engineering, systems software and programming languages · 4 · 1 since 2021Computer networks · 1 · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Building A CSFQ-Inspired Transport for Switched CXL Memory Pooling
Zerui Guo, Emily Shriver, Ming Liu 0027 |
NSDI | 2 |
| 2026 | XPNet: Cross-FPGA Power Prediction From High-Level Language CodeabstractMachine learning (ML) has been successfully employed to estimate power consumption for FPGAs using features derived from the results of High Level Synthesis (HLS). However, such models trained on one FPGA cannot be directly applied to another FPGA, even within the same FPGA series. Training a model for a new FPGA is time-consuming due to the significant effort required for dataset preparation. Researchers have to invest significant effort (weeks) in constructing a sufficient dataset with power value annotations, to train an accurate model for a new target FPGA. Another challenge is that existing model construction methods depend on many features extracted late in the HLS process, which are tool-specific and cannot be transferred between tools from different vendors. To address these challenges, we propose a novel cross-FPGA power modeling methodology called XPNet. With only frontend features from HLS, XPNet combines Transfer-Learning with innovative data selection techniques that enable efficient fine-tuning for a new target FPGA. With XPNet, models trained on one FPGA can be quickly adapted to a new target FPGA and used to efficiently predict the power on this new FPGA with high accuracy. Experiments with Polybench, Machsuite and CHStone demonstrate an average error of only 8.40% (10.34% if cross-vendor) when less than 1% of designs are used for the fine-tuning to the new target FPGA. In comparison to best prior model (with full training and 5.74% error), XPNet yields 232x speed up in dataset preparation and training, and 5x speed up in inference on a new FPGA. Zhigang Wei, Allison Seigler, Sean Lowe, Emily Shriver, Aman Arora 0001, Lizy Kurian John |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2025 | ATAPP: Architecture and Technology Aware Power Predictor for Unseen FPGAS
Zhigang Wei, Aman Arora 0001, Emily Shriver, Lizy Kurian John |
FPL | 3 |
| 2024 | Cross-FPGA Power Estimation from High Level Synthesis via Transfer-LearningabstractMachine learning (ML) has been successfully employed to estimate power consumption for FPGAs using features derived from post High Level Synthesis (HLS). As a result, the power evaluation of the design bypasses time-consuming logic synthesis and implementation. However, such models have noticeable drawbacks. Firstly, the dataset preparation is time-consuming since researchers invest significant effort in constructing a sufficient dataset to train an accurate model for a target FPGA. Secondly, the model trained on one FPGA cannot be directly applied to another. Without prior knowledge about the architecture of the second FPGA, the model's power estimation on this new FPGA is of unknown confidence. To address these challenges, we propose a novel cross-FPGA power modeling methodology called XPNet that combines Transfer-Learning with an innovative data selection technique that enables efficient fine-tuning. We start by applying Transfer-Learning with our data selection methodology to adapt a GNN-based power model to a second FPGA using only 20 data samples, resulting in 6.53% error. We then explore if our approach works for lighter-weight ML-based models, such as multi-layer perception (MLP), and show less than a 1% degradation in accuracy. Additionally, we explore the impact of using Meta-Learning algorithm on our model and show that with only 40 data samples from the target FPGA, the model still manages an error of 6%. Zhigang Wei, Aman Arora 0001, Emily Shriver, Lizy Kurian John |
FPGA | 3 |
| 2023 | MQL: ML-Assisted Queuing Latency Analysis for Data Center NetworksabstractData center network (DCN) performance analysis is becoming increasingly critical due to the growing data center scale and proliferation of latency-critical applications. Packetlevel simulators, the de-facto performance evaluation tools, allow accurate modeling of the network and protocols, but they are extremely slow. Simulation of large-scale DCNs with thousands of nodes can take days, making meaningful design space exploration impractical. Analytical techniques, such as queuing theory, can mitigate the scalability problem and offer high accuracy when specific workload assumptions are satisfied. However, their accuracy may decline as these assumptions break, and execution times explode unless designed carefully. To address these challenges, we propose a novel and scalable performance analysis methodology that combines two powerful techniques. First, it uses queuing theory and the maximum entropy (ME) principle to approximate the waiting time in each queue in a DCN. It then finds the end-to-end latency of each flow using traffic input, routing algorithm, and network parameters. This ME-based queuing model can approximate the latency under generalized exponential input traffic and general service distributions. Since its accuracy can degrade as traffic diverges from input and service time assumptions, the second step of the proposed methodology learns and corrects the systematic errors using a regression tree. The resulting ML-assisted technique achieves less than 3% modeling error on average compared to ns-3 simulations. Moreover, the speedup over ns-3 ranges from 100× to 9000× on DCNs with 128 to 1024 nodes. Shruti Yadav Narayana, Jie Tong, Anish Krishnakumar, Nuriye Yildirim, Emily Shriver, Mahesh Ketkar, Ümit Y. Ogras |
ISPASS | 5 |
| 2020 | SimTrace: Capturing over Time Program Phase BehaviorabstractAs computers and the workloads they run have grown in size and complexity, it has become difficult to test the performance and power of future products under design. These products are often designed on simulators that are orders of magnitude slower than the final product. For this reason, industry and academia have developed methodologies to reduce run times. However, in order to study runtime adaptive techniques for performance and power/energy management, it is important to capture the over time phase behavior of workloads. One technique, SimPoint, has been demonstrated to capture average behavior accurately, but it is not known how well a sequence of SimPoints can capture over time program phase behavior. To explore this, we replay the sequence of SimPoints and evaluate the sequence's accuracy. Using SPEC CPU 2017 benchmarks as a case study, we discover good accuracy for the replayed sequence: with less than 5% performance error (Instructions Per Cycle) for four time-series metrics. Steven Flolid, Emily Shriver, Zachary Susskind, Benjamin Thorell, Lizy Kurian John |
ISPASS | 2 |
| 2019 | Application Performance Prediction and Optimization Under Cache Allocation TechnologyabstractMany applications running on high-performance computing systems share limited resources such as the last-level cache, often resulting in lower performance. Intel recently introduced a new control mechanism, called cache allocation technology (CAT), which controls the cache size used by each application. To intelligently utilize this technology for automated management, it is essential to accurately identify application performance behavior for different cache allocation scenarios. In this work, we show a novel approach which automatically builds a prediction model for application performance changes with CAT. We profile the workload characteristics based on Intel Top-down Microarchitecture Analysis Method (TMAM), and train the model using machine learning. The model predicts instructions per cycle (IPC) across available cache sizes allocated for the applications. We also design a dynamic cache management technique which utilizes the prediction model and intelligently partitions the cache resource to improve application throughput. We implemented and evaluated the proposed framework in Intel PMU profiling tool running on Xeon Platinum 8186 Skylake processor. In our evaluation, we show that the proposed model accurately predicts the IPC changes of applications with 4.7% error on average for different cache allocation scenarios. Our predictive online cache managements achieves improvements on application performance of up to 25% as compared to a prediction-agnostic policy. Yeseong Kim, Ankit More, Emily Shriver, Tajana Rosing |
DATE | 3 |
| 2019 | Hardware-Assisted Cross-Generation Prediction of GPUs Under DesignabstractThis paper introduces a predictive modeling framework for GPU performance. The key innovation underlying this approach is that performance statistics collected from representative workloads running on current generation GPUs can effectively predict the performance of next-generation GPUs. This is useful when simulators are available for the next-generation device, but simulation times are exorbitant, rendering early design space exploration of microarchitectural parameters and other features infeasible. When predicting performance across three Intel GPU generations (Haswell, Broadwell, Skylake), our models achieved impressively low out-of-sample-errors ranging from 7.45% to 8.91%, while running 29 481 to 44 214 times faster than cycle-accurate simulations. A detailed ranking of the most impactful features selected for these models provides an insight as to which microarchitectural subsystems have the greatest impact on performance from one generation to the next. Kenneth O'Neal, Philip Brisk, Emily Shriver, Michael Kishinevsky |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2017 | HALWPE: Hardware-Assisted Light Weight Performance Estimation for GPUsabstractThis paper presents a predictive modeling framework for GPU performance. The key innovation underlying this approach is that performance statistics collected from representative workloads running on current generation GPUs can effectively predict the performance of next-generation GPUs. This is useful when simulators are available for the next-generation device, but simulation times are exorbitant, rendering early design space exploration of microarchitectural parameters and other features infeasible. When predicting performance across three Intel GPU generations (Haswell, Broadwell, Skylake), our models achieved low out-of-sample-errors ranging from 7.45% to 8.91%, while running 30,000-45,000 times faster than cycle-accurate simulation. Kenneth O'Neal, Philip Brisk, Emily Shriver, Michael Kishinevsky |
DAC | 3 |
| 2017 | P4: Phase-based power/performance prediction of heterogeneous systems via neural networksabstractThe emergence of Internet of Things increases the complexity and the heterogeneity of computing platforms. Migrating workload between various platforms is one way to improve both energy efficiency and performance. Effective migration decisions require accurate estimates of its costs and benefits. To date, these estimates were done by either instrumenting the source code/binaries, thus causing high overhead, or by using power estimates from hardware performance counters, which work well for individual machines, but until now have not been accurate for predicting across different architectures. In this paper, we propose P4, a new Phase-based Power and Performance Prediction framework which identifies cross-platform application power and performance at runtime for heterogeneous computing systems. P4analyzes and detects machine-independent application phases by characterizing computing platforms offline with a set of benchmarks, and then builds neural network-based models to automatically identify and generalize the complex cross-platform relationships for each benchmark phase. It then leverages these models along with performance counter measurements collected at runtime to estimate performance and power consumption if it were running on a completely different computing platform, including a different CPU architecture, without ever having to run it on there. We evaluate the proposed framework on four commercial heterogeneous platforms, ranging from X86 servers to mobile ARM-based architecture, with 129 industry-standard benchmarks. Our experimental results show that P4can predict the power and performance changes with only 6.8% and 5.6% error, respectively, even for completely different architectures from the ones applications ran on. Yeseong Kim, Pietro Mercati, Ankit More, Emily Shriver, Tajana Rosing |
ICCAD | 4 |
| 2017 | GPU Performance Estimation using Software Rasterization and Machine LearningabstractThis paper introduces a predictive modeling framework to estimate the performance of GPUs during pre-silicon design. Early-stage performance prediction is useful when simulation times impede development by rendering driver performance validation, API conformance testing and design space explorations infeasible. Our approach builds a Random Forest regression model to analyze DirectX 3D workload behavior when executed by a software rasterizer, which we have extended with a workload characterizer to collect further performance information via program counters. In addition to regression models, this work produces detailed feature rankings which can provide valuable architectural insight, and accurate performance estimates for an Intel integrated Skylake generation GPU. Our models achieve reasonable out-of-sample-error rates of 14%, with an average simulation speedup of 327x. Kenneth O'Neal, Philip Brisk, Ahmed Abousamra, Zack Waters, Emily Shriver |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2012 | A hybrid and adaptive model for predicting register file and SRAM power using a reference designabstractThis paper presents a predictive SRAM power model that reduces the changes required to adapt existing models to handle new circuit topologies, process corners, and design space exploration. Analytical equations model the impact of varying common characteristics such as bit-width, entries, segmentation, gating, and sizing while topology specific characteristics are captured empirically from a reference design. On distinct topologies of multi-port read, single- and dual-ended writes, this approach demonstrates an error of 5% and 7% for leakage and dynamic power respectively. We show that for a specific topology, any reference configuration can be used for accurate prediction. Eric Donkoh, Alicia Lowery, Emily Shriver |
DAC | 3 |
| 2008 | A System Verilog Rewriting System for RTL Abstraction with Pentium Case StudyabstractThis paper presents a new tool for SystemVerilog RTL modifications with on-the-fly validation of local RTL changes. The tool, SV-rewrite, imports an initial version of SystemVerilog RTL and elaborates it into a hierarchical design description visualized as structural diagrams. From the design cockpit the user can select any set of visualized components, open a favorite text editor, modify then validate the new RTL description, and finally substitute this new rewritten RTL into the larger model to replace the originally selected components. This process of local validated rewrites can be repeated until the entire RTL is safely rewritten. We studied RTL abstraction using SV-rewrite to abstract the Pentium 80602 (P54CS) integer execution unit and register file. We have produced a significantly more readable RTL that is 2 to 3 times smaller than the original one. The abstracted RTL was validated by booting Linux on an FPGA-based emulation platform. Steve Haynal, Timothy Kam, Michael Kishinevsky, Emily Shriver |
MEMOCODE | 4 |
| 2002 | Power and CAD considerations for the 1.75mbyte, 1.2ghz L2 cache on the alpha 21364 CPUabstractA 1.75 MByte L2 cache has been designed and fabricated as part of the Alpha 21364 microprocessor[1] (Figure 1), in a .18m bulk CMOS process. The cache was designed to run at 1.2 GHz, and pass-1 samples confirm this. While Alpha CPUs are known primarily for high speed, the combination of package constraints and a tight schedule forced careful attention to the integrated whole of power expenditure and the interaction of CAD with design. The cache consumes only 7% of total die power. Joel Grodstein, Rachid Rayess, Tad Truex, Linda Shattuck, Sue Lowell, Dan Bailey, David Bertucci, Gabriel P. Bischoff, Daniel E. Dever, Michael K. Gowan, Roy Lane, Brian Lilly, Krishna Nagalla, Rahul Shah 0004, Emily Shriver, Shi-Huang Yin, Shannon V. Morton |
ACM Great Lakes Symposium on VLSI | 15 |