EDBT 2026 Demo / reviewers in the wild / expert
Stephen P. Crago
dblp:97/3819
· DBLP profile ↗
22ranked-venue papers
3as first author
2since 2021 · last 2023
0000-0002-5620-4227ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-authorSoftware engineering, systems software and programming languages · 2Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Energy-efficient computing · 26% Memory systems · 20% Processor architecture and microarchitecture · 20% | |
| Theoretical computer science
1 paper |
Mathematical optimization · 100% |
Topics — the 11 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Memory systems
processing-in-memory |
0.0 | 1 | 2003 | A Performance Analysis of PIM, Stream Processing, and Tiled Processing on Memory-Intensive Signal Processing Kernels · ISCA 2003 |
Memory systems › processing-in-memory
processor-in-memory architecture |
0.0 | 1 | 2003 | A Performance Analysis of PIM, Stream Processing, and Tiled Processing on Memory-Intensive Signal Processing Kernels · ISCA 2003 |
Distributed systems
stream processing |
0.0 | 1 | 2003 | A Performance Analysis of PIM, Stream Processing, and Tiled Processing on Memory-Intensive Signal Processing Kernels · ISCA 2003 |
Processor architecture and microarchitecture › data-parallel architecture
stream processor |
0.0 | 1 | 2003 | A Performance Analysis of PIM, Stream Processing, and Tiled Processing on Memory-Intensive Signal Processing Kernels · ISCA 2003 |
Parallel and multicore computing › parallel algorithms › parallel algorithm design
tiled algorithm |
0.0 | 1 | 2003 | A Performance Analysis of PIM, Stream Processing, and Tiled Processing on Memory-Intensive Signal Processing Kernels · ISCA 2003 |
Processor architecture and microarchitecture
tiled architecture |
0.0 | 1 | 2003 | A Performance Analysis of PIM, Stream Processing, and Tiled Processing on Memory-Intensive Signal Processing Kernels · ISCA 2003 |
Energy-efficient computing › voltage scaling
dynamic voltage scaling |
0.0 | 1 | 2002 | An optimal voltage synthesis technique for a power-efficient satellite application · DAC 2002 |
Energy-efficient computing
energy-aware scheduling |
0.0 | 1 | 2002 | A Fast Resource Synthesis Technique for Energy-Efficient Real-Time System · RTSS 2002 |
Energy-efficient computing
power management |
0.0 | 1 | 2002 | A Fast Resource Synthesis Technique for Energy-Efficient Real-Time System · RTSS 2002 |
Embedded and real-time systems
real-time scheduling |
0.0 | 1 | 2002 | A Fast Resource Synthesis Technique for Energy-Efficient Real-Time System · RTSS 2002 |
High-performance computing
scientific computing systems |
0.0 | 1 | 2003 | A Performance Analysis of PIM, Stream Processing, and Tiled Processing on Memory-Intensive Signal Processing Kernels · ISCA 2003 |
Methods — techniques the papers use, named apart from their topics
energy-aware resource allocation · 0.1EDF scheduling · 0.1simulation · 0.0performance measurement · 0.0optimization · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Quantpipe: Applying Adaptive Post-Training Quantization For Distributed Transformer Pipelines In Dynamic Edge EnvironmentsabstractPipeline parallelism has achieved great success in deploying large-scale transformer models in cloud environments, but has received less attention in edge environments. Unlike in cloud scenarios with high-speed and stable network inter-connects, dynamic bandwidth in edge systems can degrade distributed pipeline performance. We address this issue with QuantPipe, a communication-efficient distributed edge system that introduces post-training quantization (PTQ) to compress the communicated tensors. QuantPipe uses adaptive PTQ to change bitwidths in response to bandwidth dynamics, maintaining transformer pipeline performance while incurring limited inference accuracy loss. We further improve the accuracy with a directed-search analytical clipping for integer quantization method (DS-ACIQ), which bridges the gap between estimated and real data distributions. Experimental results show that QuantPipe adapts to dynamic bandwidth to maintain pipeline performance while achieving a practical model accuracy using a wide range of quantization bitwidths, e.g., improving accuracy under 2-bit quantization by 15.85% on ImageNet compared to naive quantization. Connor Imes, Souvik Kundu 0002, Peter A. Beerel, Stephen P. Crago, John Paul Walters |
ICASSP | 5 |
| 2022 | PipeEdge: Pipeline Parallelism for Large-Scale Model Inference on Heterogeneous Edge DevicesabstractDeep neural networks with large model sizes achieve state-of-the-art results for tasks in computer vision and natural language processing. However, such models are too compute- or memory-intensive for resource-constrained edge devices. Prior works on parallel and distributed execution primarily focus on training-rather than inference-using homogeneous accelerators in data centers. We propose PipeEdge, a distributed framework for edge systems that uses pipeline parallelism to both speed up inference and enable running larger, more accurate models that otherwise cannot fit on single edge devices. PipeEdge uses an optimal partition strategy that considers heterogeneity in compute, memory, and network bandwidth. Our empirical evaluation demonstrates that PipeEdge achieves 11.88× and 12.78× speedup using 16 edge devices for the ViT-Huge and BERT-Large models, respectively, with no accuracy loss. Similarly, PipeEdge improves throughput for ViT-Huge (which cannot fit in a single device) by 3.93× over a 4-device baseline using 16 edge devices. Finally, we show up to 4.16× throughput improvement over the state-of-the-art PipeDream when using a heterogeneous set of devices. Connor Imes, Xuanang Zhao, Souvik Kundu 0002, Peter A. Beerel, Stephen P. Crago, John Paul Walters |
DSD | 6 |
| 2017 | A Comparison of System Performance on a Private OpenStack Cloud and Amazon EC2abstractCloud computing is becoming increasingly pervasive and is being adopted even for high performance computing and mission critical applications. As cloud computing extends its usage, understanding of its performance becomes more important. In this paper, we present the system performance using Amazon EC2, representing a large public cloud platform, and OpenStack, representing the most popular open-source infrastructure-as-a-service cloud platform that can be used to implement both private and public clouds. Our study covers the performance of all significant cloud infrastructure resources - CPU, memory, storage, and network. We compare the performance of the public Amazon EC2 and a private OpenStack platform with similar configurations using several benchmarks. We compare and analyze diverse performance aspects based on different hypervisor, storage, and/or network configurations. For the OpenStack platform, we measure virtualization overhead for each resource. We also investigated the performance impact of more advanced features of each cloud, including faster storage options of AWS and the Infiniband network fabric with the Chameleon OpenStack platform. Mikyung Kang, Dong-In Kang 0001, John Paul Walters, Stephen P. Crago |
CLOUD | 4 |
| 2017 | Dynamically Improving Resiliency to Timing Errors for Stream Processing WorkloadsabstractLarge-scale data processing paradigms, such as stream processing, are widespread in academic and corporate workloads. These environments are commonly subject to real-time requirements, such as latency and throughput, and resiliency requirements to node or network failures. These requirements have generally been approached as separate problems. Intermittent timing delays due to factors such as garbage collection can further complicate the management of the stream processing workload. Insufficient resource allocations can also lead to poor performance. Currently, tuning these applications is done manually. We show that improper configuration can greatly affect performance. It is reported that even 100ms of increased latency in online sales platforms can potentially result in lower sales. In this paper we propose Dynamo, a framework and monitor that implements a methodology for addressing both the performance and timing error problems by increasing the resiliency of stream processing frameworks to timing delays. Dynamo autonomously adjusts the resource allocation by using a performance profile that is generated through application profiling. Dynamo partitions an application's allocated resources into active and passive partitions that are dynamically adjusted to match an application's multi-modal behavior. The distribution of resources determines the amount of computation that Dynamo can duplicate and process redundantly, thereby reducing the probability of timing errors that affect a tuple's total execution time. In our experiments, we observed improvements in the number of tuples with missed deadlines. Our results show that Dynamo is able to consistently improve the resiliency to timing errors over a number of differing occurrence rates. Furthermore, we show that the improvement in the number of missed deadlines increases with the amount of spare resources, with a 71.40% reduction in the best case. Geoffrey Phi C. Tran, John Paul Walters, Stephen P. Crago |
PDCAT | 3 |
| 2016 | Automated Demand-Based Vertical Elasticity for Heterogeneous Real-Time WorkloadsabstractCloud computing is revolutionizing the information technology field. However, clouds today have not yet addressed real-time applications. Current work has shown that real-time hypervisors are capable of allowing virtual machines to meet real-time requirements. Other work has looked at statically allocating resources to virtual guests, which allow those guests to meet deadlines. In this paper, we present DART-C (Demand-based Allocation for Real-Time Clouds), a dynamic real-time cloud framework that allows automated adaptation to these types of workloads. Many applications follow dynamic multi-modal execution patterns, with varying computational requirements over time. DART-C provides demand-based elasticity to support changing real-time requirements by enabling applications to report mode changes to a resource manager, which reallocates resources based on need. We also describe a prototype and demonstrate up to 62% in system resource utilization savings compared to a static allocation, when running a synthetic application set. Geoffrey Phi C. Tran, Yu-An Chen, Dong-In Kang 0001, John Paul Walters, Stephen P. Crago |
CLOUD | 5 |
| 2015 | Supporting High Performance Molecular Dynamics in Virtualized Clusters using IOMMU, SR-IOV, and GPUDirectabstractCloud Infrastructure-as-a-Service paradigms have recently shown their utility for a vast array of computational problems, ranging from advanced web service architectures to high throughput computing. However, many scientific computing applications have been slow to adapt to virtualized cloud frameworks. This is due to performance impacts of virtualization technologies, coupled with the lack of advanced hardware support necessary for running many high performance scientific applications at scale. Andrew J. Younge, John Paul Walters, Stephen P. Crago, Geoffrey C. Fox |
VEE | 3 |
| 2015 | Welcome to the IEEE Transactions on Big DataabstractPresents an editorial introducting the inaugural issue of this publication. Stephen P. Crago |
IEEE Trans. Big Data | 1 |
| 2014 | Bridging the Virtualization Performance Gap for HPC Using SR-IOV for InfiniBandabstractThis paper shows that using SRIOV for InfiniBand can enable virtualized HPC, but only if the NIC tunable parameters are set appropriately. In particular, contrary to common belief, our results show that the default policy of aggressive use of interrupt moderation can have a negative impact on the performance of InfiniBand platforms virtualized using SR-IOV. Careful tuning of interrupt moderation benefits both Native and VM platforms and helps to bridge the gap between native and virtualized performance. For some workloads, the performance gap is reduced by 15-30%. Malek Musleh, Vijay S. Pai, John Paul Walters, Andrew J. Younge, Stephen P. Crago |
IEEE CLOUD | 5 |
| 2014 | GPU Passthrough Performance: A Comparison of KVM, Xen, VMWare ESXi, and LXC for CUDA and OpenCL ApplicationsabstractAs more scientific workloads are moved into the cloud, the need for high performance accelerators increases. Accelerators such as GPUs offer improvements in both performance and power efficiency over traditional multi-core processors, however, their use in the cloud has been limited. Today, several common hypervisors support GPU passthrough, but their performance has not been systematically characterized. In this paper we show that low overhead GPU passthrough is achievable across 4 major hypervisors and two processor microarchitectures. We compare the performance of two generations of NVIDIA GPUs within the Xen, VMWare ESXi, and KVM hypervisors, and we also compare the performance to that of Linux Containers (LXC). We show that GPU passthrough to KVM achieves 98 -- 100\% of the base system's performance across two architectures, while Xen and VMWare achieve 96 -- 99\% of the base systems performance, respectively. In addition, we describe several valuable lessons learned through our analysis and share the advantages and disadvantages of each hypervisor/GPU passthrough solution. John Paul Walters, Andrew J. Younge, Dong-In Kang 0001, Ke Thia Yao, Mikyung Kang, Stephen P. Crago, Geoffrey C. Fox |
IEEE CLOUD | 6 |
| 2011 | Heterogeneous Cloud ComputingabstractCurrent cloud computing infrastructure typically assumes a homogeneous collection of commodity hardware, with details about hardware variation intentionally hidden from users. In this paper, we present our approach for extending the traditional notions of cloud computing to provide a cloud-based access model to clusters that contain a heterogeneous architectures and accelerators. We describe our ongoing work extending the Open Stack cloud computing stack to support heterogeneous architectures and accelerators, and our experiences running Open Stack on our local heterogeneous cluster testbed. Stephen P. Crago, Kyle Dunn, Patrick Eads, Lorin Hochstein, Dong-In Kang 0001, Mikyung Kang, Devendra Modium, Karandeep Singh, Jinwoo Suh, John Paul Walters |
CLUSTER | 1 |
| 2010 | Opportunities for concurrent dynamic analysis with explicit inter-core communicationabstractMulticore is now the dominant processor trend, and the number of cores is rapidly increasing. The paradigm shift to multicore forces the redesign of the software stack, which includes dynamic analysis. Dynamic analyses provide rich features to software in various areas, such as debugging, testing, optimization, and security. However, these techniques often suffer from excessive overhead, which make it less practical. Previously, this overhead has been overcome by improved processor performance as each generation gets faster, but the performance requirements of dynamic analyses in the multicore era cannot be fulfilled without redesigning for parallelism. Stephen P. Crago |
PASTE | 2 |
| 2007 | Evaluation of Stream Virtual Machine on Raw ProcessorabstractStream processing exploits the properties of stream applications such as parallelism and throughput-oriented nature of the applications. One of the most recent approaches is community-supported Morphware stable interface (MSI) used as a stable abstraction between high-level compilers (HLC) and low-level architecture-specific compilers (LLC). We focus on one part of the MSI, the stream virtual machine (SVM). We implemented a high-level compiler that produces SVM output renderings and SVM implementation. The SVM is implemented with the Raw compiler as the LLC and an accompanying library. We also implemented stream applications such as matrix multiplication, FIR bank, and ground moving target indicator (GMTI) using the implemented compilers. These applications are optimized and the results are analyzed. The results show that the SVM framework is generally suitable for streaming applications on Raw processor. Jinwoo Suh, Richard A. Lethin, Stephen P. Crago, Janice O. McMahon, Dong-In Kang 0001 |
IPDPS | 3 |
| 2007 | A Voltage and Resource Synthesis Technique for Energy-Aware Real-time SystemsabstractWe consider a resource synthesis technique for realtime systems where dynamic voltage scaling is supported, the energy budget is limited, and the performance of the system depends on how resources and energy are used. We propose a resource synthesis technique that derives both the supply voltages and the resource allocation of the tasks in the system to maximize system performance. The resulting system satisfies real-time schedulability and energy requirements. Dong-In Kang 0001, Stephen P. Crago, Jinwoo Suh, Janice O. McMahon |
RTCSA | 2 |
| 2006 | Design and evaluation of a hierarchical decoupled architecture
Won Woo Ro, Stephen P. Crago, Alvin M. Despain, Jean-Luc Gaudiot |
J. Supercomput. | 2 |
| 2003 | A Performance Analysis of PIM, Stream Processing, and Tiled Processing on Memory-Intensive Signal Processing KernelsabstractTrends in microprocessors of increasing die size and clock speed and decreasing feature sizes have fueled rapidly increasing performance. However, the limited improvements in DRAM latency and bandwidth and diminishing returns of increasing superscalar ILP and cache sizes have led to the proposal of new microprocessor architectures that implement processor-in- memory, stream processing, and tiled processing. Each architecture is typically evaluated separately and compared to a baseline architecture. In this paper, we evaluate the performance of processors that implement these architectures on a common set of signal processing kernels.The implementation results are compared with the measured performance of a conventional system based on the PowerPC with Altivec. The results show that these new processors show significant improvements over conventional systems and that each architecture has its own strengths and weaknesses. Jinwoo Suh, Eun-Gyu Kim, Stephen P. Crago, Lakshmi Srinivasan, Matthew C. French |
ISCA | 3 |
| 2002 | An optimal voltage synthesis technique for a power-efficient satellite applicationabstractThis paper presents an optimal voltage synthesis technique for a satellite application to maximize system performance subject to energy budget. A period of a satellite's orbit is partitioned into several independent regions with different characteristics such as type of computation, importance, performance requirements, and energy consumptions. Given a periodic energy recharge model, optimal voltages for the regions are synthesized such that the overall performance is maximized within the energy budget in the period. Dong-In Kang 0001, Jinwoo Suh, Stephen P. Crago |
DAC | 3 |
| 2002 | A Stream Processor Development PlatformabstractWe describe a hardware and software platform for developing streaming applications. Programmers write stream programs in high-level languages, and a set of software tools maps these programs to code that runs on a streaming hardware system. The hardware platform includes two Imagine stream processors, together providing 32 GFLOPS peak performance, and a high-speed onboard network to carry video and other data between peripherals and the Imagine processors. Ben Serebrin, John D. Owens, Chen H. Chen, Stephen P. Crago, Ujval J. Kapasi, Peter R. Mattson, Jinyung Namkoong, Scott Rixner, William J. Dally |
ICCD | 4 |
| 2002 | A Fast Resource Synthesis Technique for Energy-Efficient Real-Time SystemabstractWe consider a resource synthesis technique for real-time systems where the energy budget is limited and the performance of the system depends on how resources and energy are used. We consider two performance models for a task: (1) a task has variable execution time and performance of a task depends on the amount of execution time received, and (2) the execution time of a task is constant and the performance of a task depends on its frequency. We first propose an optimal resource synthesis technique which maximizes system performance without energy constraints. We prove its optimality with the earliest deadline first (EDF) scheduling policy when the performance function of a task is non-decreasing and concave. We propose an energy-aware resource allocation technique for systems with energy constraints using the same analytical framework. The energy-aware resource synthesis technique considers both resource usage and energy consumption to find a near optimal solution that maximizes system performance within the energy budget. Dong-In Kang 0001, Stephen P. Crago, Jinwoo Suh |
RTSS | 2 |
| 2001 | A PIM-based Multiprocessor SystemabstractThe growing gap in performance between processor and memory speeds has created a problem for data-intensive applications. A recent approach for solving this problem is to use processor-in-memory (PIM) technology. PIM technology integrates a processor on a DRAM memory chip, which increases bandwidth between the processor and memory. In this paper, we discuss a PIM-based multiprocessor system, the System Level Intelligent Intensive Computing (SLIIC) Quick look (QL) board. This system includes eight COTS PIM chips and two FPGA chips that implement a flexible interconnect network. The performance of the SLIIC QL board is measured and analyzed for the distributed corner-turn application. We show that the performance of the current SLIIC QL on the distributed corner turn application is better than a PowerPC-based multicomputer that consumes more power and occupies more area. This advantage, which can be achieved in a limited context, demonstrates that even limited COTS PIMs have some advantages for data-intensive computations. Jinwoo Suh, Stephen P. Crago, Changping Li, Robert Parker |
IPDPS | 2 |
| 2001 | Implementations of Real-time Data Intensive Applications on PIM-based Multiprocessor Systemsabstract∗ Effort sponsored by Defense Advanced Research Projects Agency (DARPA) through the Air Force Research Laboratory, USAF, under agreement number F30602-99-1-0521. The U.S. Government is authorized to reproduce and distribute reprints for governmental purposes notwithstanding any copyright annotation thereon. The views and conclusions contained herein are those of the authors and should not interpreted as necessarily representing the official policies or endorsement, either expressed or implied, of the Defense Advanced Research Projects Agency (DARPA), Air Force Research Laboratory, or the U.S. Government. † Richard Chau was working for Lockheed martin when this work was conducted. Abstract Jinwoo Suh, Changping Li, Stephen P. Crago, Stephen F. Shank, Richard H. Chau, Walter J. Mazur, Rick Pancoast |
IPDPS | 4 |
| 2000 | A Communication Scheduling Algorithm for Multi-FPGA SystemsabstractFor multiple FPGA systems, the limited number of I/O pins causes many problems. To solve these problems, efficient communication scheduling among FPGAs is crucial for obtaining high CLB utilization. We provide a heuristic for the NP-complete scheduling algorithm. Experimental results show that our algorithm generates excellent communication schedules: more than 90% of the randomly generated problem instances were scheduled with less than 20% overhead compared with an optimal algorithm. The execution time of the scheduling algorithm is two orders of magnitude less than the optimal scheduling algorithm. Jinwoo Suh, Dong-In Kang 0001, Stephen P. Crago |
FCCM | 3 |
| 1998 | SLAAC: A Distributed Architecture for Adaptive ComputingabstractSoftware tools, including debuggers and performance monitors, have been developed independently for specific adaptive systems. Consequently, a user has to learn a new set of tools when switching to a different architecture. Although at the level closest to the hardware the runtime system is necessarily different for specific systems, much of the runtime system software functionality is common across systems and could be standardized. In this paper, we propose the SLAAC (system level applications of adaptive computing) reference architecture to help address some of the issues in the adaptive computing community. Stephen P. Crago, Brian Schott, Robert Parker |
FCCM | 1 |