VLDB 2026 Research / reviewers in the wild / expert
Michael Laurenzano
dblp:29/5459 · also Michael A. Laurenzano
· DBLP profile ↗
33ranked-venue papers
7as first author
0since 2021 · last 2019
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 27 · 4 first-authorSoftware engineering, systems software and programming languages · 7 · 4 first-authorArtificial intelligence and machine learning · 1Human-computer interaction and ubiquitous computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
16 papers |
Cloud and datacenter computing · 32% Energy-efficient computing · 15% Performance modeling and evaluation · 14% | |
| Software engineering, system software, and programming languages
4 papers |
Compilers and program optimization · 70% Program analysis · 17% Runtime systems and virtual machines · 13% | |
| Artificial intelligence
4 papers |
Question answering and dialogue systems · 67% Speech recognition and synthesis · 12% Deep learning architectures and training · 8% |
Topics — the 30 heaviest of 50, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Cloud and datacenter computing › datacenter architecture
warehouse-scale computer |
1.2 | 5 | 2017 | Reining in Long Tails in Warehouse-Scale Computers with Quick Voltage Boosting Using Adrenaline · ACM Trans. Comput. Syst. 2017 Designing Future Warehouse-Scale Computers for Sirius, an End-to-End Voice and Vision Personal Assistant · ACM Trans. Comput. Syst. 2016 DjiNN and Tonic: DNN as a service and its implications for future warehouse scale computers · ISCA 2015 |
Energy-efficient computing
power management |
0.5 | 3 | 2016 | PowerChop: Identifying and Managing Non-critical Units in Hybrid Processor Architectures · ISCA 2016 Octopus-Man: QoS-driven task management for heterogeneous multicores in warehouse-scale computers · HPCA 2015 Making the Most of SMT in HPC: System- and Application-Level Perspectives · ACM Trans. Archit. Code Optim. 2014 |
Energy-efficient computing › power management
dynamic voltage and frequency scaling |
0.5 | 2 | 2017 | Reining in Long Tails in Warehouse-Scale Computers with Quick Voltage Boosting Using Adrenaline · ACM Trans. Comput. Syst. 2017 Adrenaline: Pinpointing and reining in tail queries with quick voltage boosting · HPCA 2015 |
Natural language and speech › Question answering and dialogue systems
intent detection |
0.4 | 1 | 2019 | An Evaluation Dataset for Intent Classification and Out-of-Scope Prediction · EMNLP/IJCNLP (1) 2019 |
Processor architecture and microarchitecture
instruction set architecture |
0.3 | 2 | 2016 | Concise loads and stores: The case for an asymmetric compute-memory architecture for approximation · MICRO 2016 Protean Code: Achieving Near-Free Online Code Transformations for Warehouse Scale Computers · MICRO 2014 |
GPUs and heterogeneous computing
deep learning on GPUs |
0.3 | 1 | 2017 | DeftNN: addressing bottlenecks for DNN execution on GPUs via synapse vector elimination and near-compute data fission · MICRO 2017 |
Hardware accelerators and domain-specific architectures
machine learning accelerator |
0.3 | 2 | 2015 | Sirius: An Open End-to-End Voice and Vision Personal Assistant and Its Implications for Future Warehouse Scale Computers · ASPLOS 2015 DjiNN and Tonic: DNN as a service and its implications for future warehouse scale computers · ISCA 2015 |
High-performance computing
performance optimization at scale |
0.3 | 2 | 2014 | Making the Most of SMT in HPC: System- and Application-Level Perspectives · ACM Trans. Archit. Code Optim. 2014 High-frequency simulations of global seismic wave propagation using SPECFEM3D_GLOBE on 62K processors · SC 2008 |
Performance modeling and evaluation
workload characterization |
0.3 | 3 | 2019 | Caliper: Interference Estimator for Multi-tenant Environments Sharing Architectural Resources · ACM Trans. Archit. Code Optim. 2019 Reining in Long Tails in Warehouse-Scale Computers with Quick Voltage Boosting Using Adrenaline · ACM Trans. Comput. Syst. 2017 Adrenaline: Pinpointing and reining in tail queries with quick voltage boosting · HPCA 2015 |
Compilers and program optimization
approximate computing |
0.2 | 1 | 2016 | Input responsiveness: using canary inputs to dynamically steer approximation · PLDI 2016 |
Compilers and program optimization
loop optimization |
0.2 | 1 | 2016 | Continuous shape shifting: Enabling loop co-optimization via near-free dynamic code rewriting · MICRO 2016 |
Compilers and program optimization › dynamic optimization
profile-guided optimization |
0.2 | 1 | 2016 | CrystalBall: Statically analyzing runtime behavior via deep sequence learning · MICRO 2016 |
Program analysis
static analysis |
0.2 | 1 | 2016 | CrystalBall: Statically analyzing runtime behavior via deep sequence learning · MICRO 2016 |
Emerging computing paradigms
approximate computing |
0.2 | 1 | 2016 | Concise loads and stores: The case for an asymmetric compute-memory architecture for approximation · MICRO 2016 |
Cloud and datacenter computing
datacenter architecture |
0.2 | 1 | 2016 | Designing Future Warehouse-Scale Computers for Sirius, an End-to-End Voice and Vision Personal Assistant · ACM Trans. Comput. Syst. 2016 |
Electronic design automation
hardware/software co-design |
0.2 | 1 | 2016 | PowerChop: Identifying and Managing Non-critical Units in Hybrid Processor Architectures · ISCA 2016 |
Processor architecture and microarchitecture › general-purpose processor architecture
hybrid processor |
0.2 | 1 | 2016 | PowerChop: Identifying and Managing Non-critical Units in Hybrid Processor Architectures · ISCA 2016 |
Energy-efficient computing
power gating |
0.2 | 1 | 2016 | PowerChop: Identifying and Managing Non-critical Units in Hybrid Processor Architectures · ISCA 2016 |
Performance modeling and evaluation
performance prediction |
0.2 | 2 | 2014 | Making the Most of SMT in HPC: System- and Application-Level Perspectives · ACM Trans. Archit. Code Optim. 2014 How Well Can Simple Metrics Represent the Performance of HPC Applications? · SC 2005 |
Cloud and datacenter computing
cluster resource management and scheduling |
0.2 | 1 | 2015 | Octopus-Man: QoS-driven task management for heterogeneous multicores in warehouse-scale computers · HPCA 2015 |
Embedded and real-time systems › real-time scheduling › multicore scheduling
heterogeneous multicore scheduling |
0.2 | 1 | 2015 | Octopus-Man: QoS-driven task management for heterogeneous multicores in warehouse-scale computers · HPCA 2015 |
Cloud and datacenter computing › datacenter architecture
server architecture |
0.2 | 1 | 2015 | Sirius: An Open End-to-End Voice and Vision Personal Assistant and Its Implications for Future Warehouse Scale Computers · ASPLOS 2015 |
Cloud and datacenter computing
warehouse-scale computing |
0.2 | 1 | 2015 | Adrenaline: Pinpointing and reining in tail queries with quick voltage boosting · HPCA 2015 |
Runtime systems and virtual machines
dynamic compilation |
0.2 | 1 | 2014 | Protean Code: Achieving Near-Free Online Code Transformations for Warehouse Scale Computers · MICRO 2014 |
Memory systems › cache management
cache contention mitigation |
0.2 | 1 | 2014 | Protean Code: Achieving Near-Free Online Code Transformations for Warehouse Scale Computers · MICRO 2014 |
Cloud and datacenter computing › datacenter operations
performance interference prediction |
0.2 | 1 | 2014 | SMiTe: Precise QoS Prediction on Real-System SMT Processors to Improve Utilization in Warehouse Scale Computers · MICRO 2014 |
Processor architecture and microarchitecture › multithreading
simultaneous multithreading |
0.2 | 1 | 2014 | Making the Most of SMT in HPC: System- and Application-Level Perspectives · ACM Trans. Archit. Code Optim. 2014 |
Performance modeling and evaluation › statistical analysis
statistical modeling |
0.2 | 1 | 2014 | Making the Most of SMT in HPC: System- and Application-Level Perspectives · ACM Trans. Archit. Code Optim. 2014 |
Natural language and speech › Speech recognition and synthesis
voice assistant |
0.1 | 2 | 2016 | Designing Future Warehouse-Scale Computers for Sirius, an End-to-End Voice and Vision Personal Assistant · ACM Trans. Comput. Syst. 2016 Sirius: An Open End-to-End Voice and Vision Personal Assistant and Its Implications for Future Warehouse Scale Computers · ASPLOS 2015 |
Machine learning › Deep learning architectures and training › neural network inference
DNN inference |
0.1 | 1 | 2017 | DeftNN: addressing bottlenecks for DNN execution on GPUs via synapse vector elimination and near-compute data fission · MICRO 2017 |
Methods — techniques the papers use, named apart from their topics
workload characterization · 0.7synapse vector elimination · 0.6data fission · 0.6intermediate representation analysis · 0.5deep sequence learning · 0.5AUROC · 0.5micro-experiment-based estimation · 0.4proactive boosting · 0.3per-query characteristics · 0.3simulation · 0.2runtime monitoring · 0.2phase identification · 0.2performance modeling · 0.2functional unit design · 0.2dynamic approximation steering · 0.2code generation · 0.2canary input · 0.2accelerator benchmarking · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2019 | An Evaluation Dataset for Intent Classification and Out-of-Scope PredictionabstractStefan Larson, Anish Mahendran, Joseph J. Peper, Christopher Clarke, Andrew Lee, Parker Hill, Jonathan K. Kummerfeld, Kevin Leach, Michael A. Laurenzano, Lingjia Tang, Jason Mars. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Stefan Larson, Anish Mahendran, Joseph Peper, Christopher Clarke, Andrew Lee 0001, Parker Hill, Jonathan K. Kummerfeld, Kevin Leach, Michael Laurenzano, Lingjia Tang, Jason Mars |
EMNLP/IJCNLP (1) | 9 |
| 2019 | Caliper: Interference Estimator for Multi-tenant Environments Sharing Architectural ResourcesabstractWe introduce Caliper , a technique for accurately estimating performance interference occurring in shared servers. Caliper overcomes the limitations of prior approaches by leveraging a micro-experiment-based technique. In contrast to state-of-the-art approaches that focus on periodically pausing co-running applications to estimate slowdown, Caliper utilizes a strategic phase-triggered technique to capture interference due to co-location. This enables Caliper to orchestrate an accurate and low-overhead interference estimation technique that can be readily deployed in existing production systems. We evaluate Caliper for a broad spectrum of workload scenarios, demonstrating its ability to seamlessly support up to 16 applications running simultaneously and outperform the state-of-the-art approaches. Ram Srivatsa Kannan, Michael Laurenzano, Jeongseob Ahn, Jason Mars, Lingjia Tang |
ACM Trans. Archit. Code Optim. | 2 |
| 2018 | Architectural support for convolutional neural networks on modern CPUsabstractA key focus of recent work in our community has been on devising increasingly sophisticated acceleration devices for deep neural network (DNN) computation, especially for networks driven by convolution layers. Yet, despite the promise of substantial improvements in performance and energy consumption offered by these approaches, general purpose computing is not going away because its traditional well-understood programming model and continued wide deployment. Therefore, the question arises as to what can be done, if anything, to evolve conventional CPUs to accommodate efficient deep neural network computation. Animesh Jain, Michael Laurenzano, Gilles Pokam, Jason Mars, Lingjia Tang |
PACT | 2 |
| 2018 | Proctor: Detecting and Investigating Interference in Shared DatacentersabstractCloud-scale datacenter management systems utilize virtualization to provide performance isolation while maximizing the utilization of the underlying hardware infrastructure. However, virtualization does not provide complete performance isolation as Virtual Machines (VMs) still compete for nonreservable shared resources (like caches, network, I/O bandwidth etc.) This becomes highly challenging to address in datacenter environments housing tens of thousands of VMs, causing degradation in application performance. Addressing this problem for production datacenters requires a non-intrusive scalable solution that 1) detects performance intrusion and 2) investigates both the intrusive VMs causing interference, as well as the resource(s) for which the VMs are competing for. To address this problem, this paper introduces Proctor, a real time, lightweight and scalable analytics fabric that detects performance intrusive VMs and identifies its root causes from among the arbitrary VMs running in shared datacenters across 4 key hardware resources - network, I/O, cache, and CPU. Proctor is based on a robust statistical approach that requires no special profiling phases, standing in stark contrast to a wide body of prior work that assumes pre-acquisition of application level information prior to its execution. By detecting performance degradation and identifying the root cause VMs and their metrics, Proctor can be utilized to dramatically improve the performance outcomes of applications executing in large-scale datacenters. From our experiments, we are able to show that when we deploy Proctor in a datacenter housing a mix of I/O, network, compute and cache-sensitive applications, it is able to effectively pinpoint performance intrusive VMs. Further, we observe that when Proctor is applied with migration, the application-level Quality-of-Service improves by an average of 2.2× as compared to systems which are unable to detect, identify and pinpoint performance intrusion and their root causes. Ram Srivatsa Kannan, Animesh Jain, Michael Laurenzano, Lingjia Tang, Jason Mars |
ISPASS | 3 |
| 2017 | DeftNN: addressing bottlenecks for DNN execution on GPUs via synapse vector elimination and near-compute data fissionabstractDeep neural networks (DNNs) are key computational building blocks for emerging classes of web services that interact in real time with users via voice, images and video inputs. Although GPUs have gained popularity as a key accelerator platform for deep learning workloads, the increasing demand for DNN computation leaves a significant gap between the compute capabilities of GPU-enabled datacenters and the compute needed to service demand. Parker Hill, Animesh Jain, Mason Hill, Babak Zamirai, Chang-Hong Hsu, Michael Laurenzano, Scott A. Mahlke, Lingjia Tang, Jason Mars |
MICRO | 6 |
| 2017 | Reining in Long Tails in Warehouse-Scale Computers with Quick Voltage Boosting Using AdrenalineabstractReducing the long tail of the query latency distribution in modern warehouse scale computers is critical for improving performance and quality of service (QoS) of workloads such as Web Search and Memcached. Traditional turbo boost increases a processor’s voltage and frequency during a coarse-grained sliding window, boosting all queries that are processed during that window. However, the inability of such a technique to pinpoint tail queries for boosting limits its tail reduction benefit. In this work, we propose Adrenaline , an approach to leverage finer-granularity (tens of nanoseconds) voltage boosting to effectively rein in the tail latency with query-level precision. Two key insights underlie this work. First, emerging finer granularity voltage/frequency boosting is an enabling mechanism for intelligent allocation of the power budget to precisely boost only the queries that contribute to the tail latency; second, per-query characteristics can be used to design indicators for proactively pinpointing these queries, triggering boosting accordingly. Based on these insights, Adrenaline effectively pinpoints and boosts queries that are likely to increase the tail distribution and can reap more benefit from the voltage/frequency boost. By evaluating under various workload configurations, we demonstrate the effectiveness of our methodology. We achieve up to a 2.50 × tail latency improvement for Memcached and up to a 3.03 × for Web Search over coarse-grained dynamic voltage and frequency scaling (DVFS) given a fixed boosting power budget. When optimizing for energy reduction, Adrenaline achieves up to a 1.81 × improvement for Memcached and up to a 1.99 × for Web Search over coarse-grained DVFS. By using the carefully chosen boost thresholds, Adrenaline further improves the tail latency reduction to 4.82 × over coarse-grained DVFS. Chang-Hong Hsu, Michael Laurenzano, David Meisner, Thomas F. Wenisch, Ronald G. Dreslinski, Jason Mars, Lingjia Tang |
ACM Trans. Comput. Syst. | 3 |
| 2016 | Lightweight, Early Identification of At-Risk CS1 StudentsabstractBeing able to identify low-performing students early in the term may help instructors intervene or differently allocate course resources. Prior work in CS1 has demonstrated that clicker correctness in Peer Instruction courses correlates with exam outcomes and, separately, that machine learning models can be built based on early-term programming assessments. This work aims to combine the best elements of each of these approaches. We offer a methodology for creating models, based on in-class clicker questions, to predict cross-term student performance. In as early as week 3 in a 12-week CS1 course, this model is capable of correctly predicting students as being in danger of failing, or not, for 70% of the students, with only 17% of students misclassified as not at-risk when at-risk. Additional measures to ensure more broad applicability of the methodology, along with possible limitations, are explored. Soohyun Nam Liao, Daniel Zingaro, Michael Laurenzano, William G. Griswold, Leo Porter 0001 |
ICER | 3 |
| 2016 | PowerChop: Identifying and Managing Non-critical Units in Hybrid Processor ArchitecturesabstractOn-core microarchitectural structures consume significant portions of a processor's power budget. However, depending on application characteristics, those structures do not always provide (much) performance benefit. While timeout-based power gating techniques have been leveraged for underutilized cores and inactive functional units, these techniques have not directly translated to high-activity units such as vector processing units, complex branch predictors, and caches. The performance benefit provided by these units does not necessarily correspond with unit activity, but instead is a function of application characteristics. This work introduces PowerChop, a novel technique that leverages the unique capabilities of HW/SW co-designed hybrid processors to enact unit-level power management at the application phase level. PowerChop adds two small additional hardware units to facilitate phase identification and triggering different power states, enabling the software layer to cheaply track, predict and take advantage of varying unit criticality across application phases by powering gating units that are not needed for performant execution. Through detailed experimentation, we find that PowerChop significantly decreases power consumption, reducing the leakage power of a hybrid server processor by 9% on average (up to 33%) and a hybrid mobile processor by 19% (up to 40%) while introducing just 2% slowdown. Michael Laurenzano, Lingjia Tang, Jason Mars |
ISCA | 1 |
| 2016 | Characterization and bottleneck analysis of a 64-bit ARMv8 platformabstractThis paper presents the first comprehensive study of the performance, power and energy consumption of the Applied-Micro X-Gene, the first commercially available 64-bit ARMv8 platform, for HPC workloads. Our study includes a detailed comparison of the X-Gene to three other architectural design points common in HPC systems. Across these platforms, we perform careful measurements across 400+ workloads, covering different application domains, parallelization models, floating-point precision models and memory intensities. We find that the X-Gene has an average of 1.2× better energy consumption than an Intel Sandy Bridge, a design commonly found in HPC installations, while the Sandy Bridge is an average of 2.3× faster than X-Gene. Precisely quantifying the causes of performance and energy differences between two platforms is an important but challenging problem that is often addressed via detailed simulation, an approach that has limited ability to scale up to full applications and broad workload mixes. Instead, this paper adopts a statistical framework called Partial Least Squares (PLS) Path Modeling to solve this problem. PLS Path Modeling allows us to capture complex cause-effect relationships and difficult-to-measure performance concepts relating to the effectiveness of architectural units and subsystems in improving application performance using readily available hardware counter measurements. We use PLS Path Modeling to quantify the causes of the performance differences between X-Gene and Sandy Bridge in the HPC domain, finding that the performance of the memory subsystem is the dominant cause of these differences. Michael Laurenzano, Ananta Tiwari, Allyson Cauble-Chantrenne, Adam Jundt, William A. Ward Jr., Roy L. Campbell, Laura Carrington |
ISPASS | 1 |
| 2016 | Concise loads and stores: The case for an asymmetric compute-memory architecture for approximationabstractCache capacity and memory bandwidth play critical roles in application performance, particularly for data-intensive applications from domains that include machine learning, numerical analysis, and data mining. Many of these applications are also tolerant to imprecise inputs and have loose constraints on the quality of output, making them ideal candidates for approximate computing. This paper introduces a novel approximate computing technique that decouples the format of data in the memory hierarchy from the format of data in the compute subsystem to significantly reduce the cost of storing and moving bits throughout the memory hierarchy and improve application performance. This asymmetric compute-memory extension to conventional architectures, ACME, adds two new instruction classes to the ISA - load-concise and store-concise - along with three small functional units to the micro-architecture to support these instructions. ACME does not affect exact execution of applications and comes into play only when concise memory operations are used. Through detailed experimentation we find that ACME is very effective at trading result accuracy for improved application performance. Our results show that ACME achieves a 1.3x speedup (up to 1.8x) while maintaining 99% accuracy, or a 1.1x speedup while maintaining 99.999% accuracy. Moreover, our approach incurs negligible area and power overheads, adding just 0.005% area and 0.1% power to a conventional modern architecture. Animesh Jain, Parker Hill, Shih-Chieh Lin, Muneeb Khan, Md. Enamul Haque, Michael Laurenzano, Scott A. Mahlke, Lingjia Tang, Jason Mars |
MICRO | 6 |
| 2016 | Continuous shape shifting: Enabling loop co-optimization via near-free dynamic code rewritingabstractThe class of optimizations characterized by manipulating a loop's interaction space for improved cache locality and reuse (i.e, cache tiling/blocking/strip mine and interchange) are static optimizations requiring a priori information about the microarchitectural and runtime environment of an application binary. However, particularly in datacenter environments, deployed applications face numerous dynamic environments over their lifetimes. As a result, this class of optimizations can result in sub-optimal performance due to the inability to flexibly adapt iteration spaces as cache conditions change at runtime. This paper introduces continuous shape shifiting, a compilation approach that removes the risks of cache tiling optimizations by dynamically rewriting (and reshaping) deployed, running application code. To realize continuous shape shifting, we present ShapeShifter, a framework for continuous monitoring of co-running applications and their runtime environments to reshape loop iteration spaces and pinpoint near-optimal loop tile configurations. Upon identifying a need for reshaping, a new tiling approach is quickly constructed for the application, new code is dynamically generated and is then seamlessly stitched into the running application with near-zero overhead. Our evaluation on a wide spectrum of runtime scenarios demonstrates that ShapeShifter achieves an average of 10-40% performance improvement (up to 2.4 χ) on real systems depending on the runtime environment compared to an oracle static loop tiling baseline. Animesh Jain, Michael Laurenzano, Lingjia Tang, Jason Mars |
MICRO | 2 |
| 2016 | CrystalBall: Statically analyzing runtime behavior via deep sequence learningabstractUnderstanding dynamic program behavior is critical in many stages of the software development lifecycle, for purposes as diverse as optimization, debugging, testing, and security. This paper focuses on the problem of predicting dynamic program behavior statically. We introduce a novel technique to statically identify hot paths that leverages emerging deep learning techniques to take advantage of their ability to learn subtle, complex relationships between sequences of inputs. This approach maps well to the problem of identifying the behavior of sequences of basic blocks in program execution. Our technique is also designed to operate on the compiler's intermediate representation (IR), as opposed to the approaches taken by prior techniques that have focused primarily on source code, giving our approach language-independence. We describe the pitfalls of conventional metrics used for hot path prediction such as accuracy, and motivate the use of Area Under the Receiver Operating Characteristic curve (AUROC). Through a thorough evaluation of our technique on complex applications that include the SPEC CPU2006 benchmarks, we show that our approach achieves an AUROC of 0.85. Stephen A. Zekany, Daniel Rings, Nathan Harada, Michael Laurenzano, Lingjia Tang, Jason Mars |
MICRO | 4 |
| 2016 | Input responsiveness: using canary inputs to dynamically steer approximationabstractThis paper introduces Input Responsive Approximation (IRA), an approach that uses a canary input — a small program input carefully constructed to capture the intrinsic properties of the original input — to automatically control how program approximation is applied on an input-by-input basis. Motivating this approach is the observation that many of the prior techniques focusing on choosing how to approximate arrive at conservative decisions by discounting substantial differences between inputs when applying approximation. The main challenges in overcoming this limitation lie in making the choice of how to approximate both effectively (e.g., the fastest approximation that meets a particular accuracy target) and rapidly for every input. With IRA, each time the approximate program is run, a canary input is constructed and used dynamically to quickly test a spectrum of approximation alternatives. Based on these runtime tests, the approximation that best fits the desired accuracy constraints is selected and applied to the full input to produce an approximate result. We use IRA to select and parameterize mixes of four approximation techniques from the literature for a range of 13 image processing, machine learning, and data mining applications. Our results demonstrate that IRA significantly outperforms prior approaches, delivering an average of 10.2× speedup over exact execution while minimizing accuracy losses in program outputs. Michael Laurenzano, Parker Hill, Mehrzad Samadi, Scott A. Mahlke, Jason Mars, Lingjia Tang |
PLDI | 1 |
| 2016 | The case for colocation of high performance computing workloadsabstractSummary The current state of practice in supercomputer resource allocation places jobs from different users on disjoint nodes both in terms of time and space. While this approach largely guarantees that jobs from different users do not degrade one another's performance, it does so at high cost to system throughput and energy efficiency. This focused study presents job striping, a technique that significantly increases performance over the current allocation mechanism by colocating pairs of jobs from different users on a shared set of nodes. To evaluate the potential of job striping in large‐scale environments, the experiments are run at the scale of 128 nodes on the state‐of‐the‐art Gordon supercomputer. Across all pairings of 1024 process network‐attached storage parallel benchmarks, job striping increases mean throughput by 26% and mean energy efficiency by 22%. On pairings of the real applications Gyrokinetic Toroidal Code (GTC), Large‐scale Atomic/Molecular Massively Parallel Simulator (LAMMPS), and MIMD Lattice Computation (MILC) at equal scale, job striping improves average throughput by 12% and mean energy efficiency by 11%. In addition, the study provides a simple set of heuristics for avoiding low performing application pairs. Copyright © 2013 John Wiley & Sons, Ltd. Alexander Dodd Breslow, Leo Porter 0001, Ananta Tiwari, Michael Laurenzano, Laura Carrington, Dean M. Tullsen, Allan Snavely |
Concurr. Comput. Pract. Exp. | 4 |
| 2016 | PMaC's green queue: a framework for selecting energy optimal DVFS configurations in large scale MPI applicationsabstractSummary This article presents Green Queue, a production quality tracing and analysis framework for implementing application aware dynamic voltage and frequency scaling (DVFS) for message passing interface applications in high performance computing. Green Queue makes use of both intertask and intratask DVFS techniques. The intertask technique targets applications where the workload is imbalanced by reducing CPU clock frequency and therefore power draw for ranks with lighter workloads. The intratask technique targets balanced workloads where all tasks are synchronously running the same code. The strategy identifies program phases and selects the energy‐optimal frequency for each by predicting power and measuring the performance responses of each phase to frequency changes. The success of these techniques is evaluated on 1024 cores on Gordon, a supercomputer at the San Diego Supercomputer Center built using Intel Xeon E5‐2670 (Sandybridge) processors. Green Queue achieves up to 21% and 32% energy savings for the intratask and intertask DVFS strategies, respectively. Copyright © 2013 John Wiley & Sons, Ltd. Joshua Peraza, Ananta Tiwari, Michael Laurenzano, Laura Carrington, Allan Snavely |
Concurr. Comput. Pract. Exp. | 3 |
| 2016 | Designing Future Warehouse-Scale Computers for Sirius, an End-to-End Voice and Vision Personal AssistantabstractAs user demand scales for intelligent personal assistants (IPAs) such as Apple’s Siri, Google’s Google Now, and Microsoft’s Cortana, we are approaching the computational limits of current datacenter (DC) architectures. It is an open question how future server architectures should evolve to enable this emerging class of applications, and the lack of an open-source IPA workload is an obstacle in addressing this question. In this article, we present the design of Sirius, an open end-to-end IPA Web-service application that accepts queries in the form of voice and images, and responds with natural language. We then use this workload to investigate the implications of four points in the design space of future accelerator-based server architectures spanning traditional CPUs, GPUs, manycore throughput co-processors, and FPGAs. To investigate future server designs for Sirius, we decompose Sirius into a suite of eight benchmarks (Sirius Suite) comprising the computationally intensive bottlenecks of Sirius. We port Sirius Suite to a spectrum of accelerator platforms and use the performance and power trade-offs across these platforms to perform a total cost of ownership (TCO) analysis of various server design points. In our study, we find that accelerators are critical for the future scalability of IPA services. Our results show that GPU- and FPGA-accelerated servers improve the query latency on average by 8.5× and 15×, respectively. For a given throughput, GPU- and FPGA-accelerated servers can reduce the TCO of DCs by 2.3× and 1.3×, respectively. Johann Hauswald, Michael Laurenzano, Hailong Yang 0002, Yiping Kang, Austin Rovinski, Arjun Khurana, Ronald G. Dreslinski, Trevor N. Mudge, Vinicius Petrucci, Lingjia Tang, Jason Mars |
ACM Trans. Comput. Syst. | 2 |
| 2015 | AREP: Adaptive Resource Efficient Prefetching for Maximizing Multicore PerformanceabstractModern processors widely use hardware prefetching to hide memory latency. While aggressive hardware prefetchers can improve performance significantly for some applications, they can limit the overall performance in highly-utilized multicore processors by saturating the offchip bandwidth and wasting last-level cache capacity. Co-executing applications can slowdown due to contention over these shared resources. This work introduces Adaptive Resource Efficient Prefetching (AREP) -- a runtime framework that dynamically combines software prefetching and hardware prefetching to maximize throughput in highly utilized multicore processors. AREP achieves better performance by prefetching data in a resource efficient way -- conserving offchip-bandwidth and last-level cache capacity with accurate prefetching and by applying cache-bypassing when possible. AREP dynamically explores a mix of hardware/software prefetching policies, then selects and applies the best performing policy. AREP is phase-aware and re-explores (at runtime) for the best prefetching policy at phase boundaries. A multitude of experiments with workload mixes and parallel applications on a modern high performance multicore show that AREP can increase throughput by up to 49% (8.1% on average). This is complemented by improved fairness, resulting in average quality of service above 94%. Muneeb Khan, Michael Laurenzano, Jason Mars, Erik Hagersten, David Black-Schaffer |
PACT | 2 |
| 2015 | Sirius: An Open End-to-End Voice and Vision Personal Assistant and Its Implications for Future Warehouse Scale ComputersabstractAs user demand scales for intelligent personal assistants (IPAs) such as Apple's Siri, Google's Google Now, and Microsoft's Cortana, we are approaching the computational limits of current datacenter architectures. It is an open question how future server architectures should evolve to enable this emerging class of applications, and the lack of an open-source IPA workload is an obstacle in addressing this question. In this paper, we present the design of Sirius, an open end-to-end IPA web-service application that accepts queries in the form of voice and images, and responds with natural language. We then use this workload to investigate the implications of four points in the design space of future accelerator-based server architectures spanning traditional CPUs, GPUs, manycore throughput co-processors, and FPGAs. Johann Hauswald, Michael Laurenzano, Austin Rovinski, Arjun Khurana, Ronald G. Dreslinski, Trevor N. Mudge, Vinicius Petrucci, Lingjia Tang, Jason Mars |
ASPLOS | 2 |
| 2015 | Adrenaline: Pinpointing and reining in tail queries with quick voltage boostingabstractReducing the long tail of the query latency distribution in modern warehouse scale computers is critical for improving performance and quality of service of workloads such as Web Search and Memcached. Traditional turbo boost increases a processor's voltage and frequency during a coarse-grain sliding window, boosting all queries that are processed during that window. However, the inability of such a technique to pinpoint tail queries for boosting limits its tail reduction benefit. In this work, we propose Adrenaline, an approach to leverage finer granularity, 10's of nanoseconds, voltage boosting to effectively rein in the tail latency with query-level precision. Two key insights underlie this work. First, emerging finer granularity voltage/frequency boosting is an enabling mechanism for intelligent allocation of the power budget to precisely boost only the queries that contribute to the tail latency; and second, per-query characteristics can be used to design indicators for proactively pinpointing these queries, triggering boosting accordingly. Based on these insights, Adrenaline effectively pinpoints and boosts queries that are likely to increase the tail distribution and can reap more benefit from the voltage/frequency boost. By evaluating under various workload configurations, we demonstrate the effectiveness of our methodology. We achieve up to a 2.50x tail latency improvement for Memcached and up to a 3.03x for Web Search over coarse-grained DVFS given a fixed boosting power budget. When optimizing for energy reduction, Adrenaline achieves up to a 1.81x improvement for Memcached and up to a 1.99x for Web Search over coarse-grained DVFS. Chang-Hong Hsu, Michael Laurenzano, David Meisner, Thomas F. Wenisch, Jason Mars, Lingjia Tang, Ronald G. Dreslinski |
HPCA | 3 |
| 2015 | Octopus-Man: QoS-driven task management for heterogeneous multicores in warehouse-scale computersabstractHeterogeneous multicore architectures have the potential to improve energy efficiency by integrating power-efficient wimpy cores with high-performing brawny cores. However, it is an open question as how to deliver energy reduction while ensuring the quality of service (QoS) of latency-sensitive web-services running on such heterogeneous multicores in warehouse-scale computers (WSCs). In this work, we first investigate the implications of heterogeneous multicores in WSCs and show that directly adopting heterogeneous multicores without re-designing the software stack to provide QoS management leads to significant QoS violations. We then present Octopus-Man, a novel QoS-aware task management solution that dynamically maps latency-sensitive tasks to the least power-hungry processing resources that are sufficient to meet the QoS requirements. Using carefully-designed feedback-control mechanisms, Octopus-Man addresses critical challenges that emerge due to uncertainties in workload fluctuations and adaptation dynamics in a real system. Our evaluation using web-search and memcached running on a real-system Intel heterogeneous prototype demonstrates that Octopus-Man improves energy efficiency by up to 41% (CPU power) and up to 15% (system power) over an all-brawny WSC design while adhering to specified QoS targets. Vinicius Petrucci, Michael Laurenzano, John Doherty, Daniel Mossé, Jason Mars, Lingjia Tang |
HPCA | 2 |
| 2015 | DjiNN and Tonic: DNN as a service and its implications for future warehouse scale computersabstractAs applications such as Apple Siri, Google Now, Microsoft Cortana, and Amazon Echo continue to gain traction, web-service companies are adopting large deep neural networks (DNN) for machine learning challenges such as image processing, speech recognition, natural language processing, among others. A number of open questions arise as to the design of a server platform specialized for DNN and how modern warehouse scale computers (WSCs) should be outfitted to provide DNN as a service for these applications. Johann Hauswald, Yiping Kang, Michael Laurenzano, Quan Chen 0002, Trevor N. Mudge, Ronald G. Dreslinski, Jason Mars, Lingjia Tang |
ISCA | 3 |
| 2014 | Characterizing the Performance-Energy Tradeoff of Small ARM Cores in HPC Computation
Michael Laurenzano, Ananta Tiwari, Adam Jundt, Joshua Peraza, William A. Ward Jr., Roy L. Campbell, Laura Carrington |
Euro-Par | 1 |
| 2014 | Modeling the Impact of Reduced Memory Bandwidth on HPC Applications
Ananta Tiwari, Anthony Collins Gamst, Michael Laurenzano, Martin Schulz 0001, Laura Carrington |
Euro-Par | 3 |
| 2014 | Protean Code: Achieving Near-Free Online Code Transformations for Warehouse Scale ComputersabstractRampant dynamism due to load fluctuations, co runner changes, and varying levels of interference poses a threat to application quality of service (QoS) and has limited our ability to allow co-locations in modern warehouse scale computers (WSCs). Instruction set features such as the non-temporal memory access hints found in modern ISAs (both ARM and x86) may be useful in mitigating these effects. However, despite the challenge of this dynamism and the availability of an instruction set mechanism that might help address the problem, a key capability missing in the system software stack in modern WSCs is the ability to dynamically transform (and re-transform) the executing application code to apply these instruction set features when necessary. In this work we introduce protean code, a novel approach for enacting arbitrary compiler transformations at runtime for native programs running on commodity hardware with negligible (<;1%) overhead. The fundamental insight behind the underlying mechanism of protean code is that, instead of maintaining full control throughout the program's execution as with traditional dynamic optimizers, protean code allows the original binary to execute continuously and diverts control flow only at a set of virtualized points, allowing rapid and seamless rerouting to the new code variants. In addition, the protean code compiler embeds IR with high-level semantic information into the program, empowering the dynamic compiler to perform rich analysis and transformations online with little overhead. Using a fully functional protean code compiler and runtime built on LLVM, we design PC3D, Protean Code for Cache Contention in Datacenters. PC3D dynamically employs non-temporal access hints to achieve utilization improvements of up to 2.8x (1.5x on average) higher than state-of-the-art contention mitigation runtime techniques at a QoS target of 98%. Michael Laurenzano, Lingjia Tang, Jason Mars |
MICRO | 1 |
| 2014 | SMiTe: Precise QoS Prediction on Real-System SMT Processors to Improve Utilization in Warehouse Scale ComputersabstractOne of the key challenges for improving efficiency in warehouse scale computers (WSCs) is to improve server utilization while guaranteeing the quality of service (QoS) of latency-sensitive applications. To this end, prior work has proposed techniques to precisely predict performance and QoS interference to identify 'safe' application co-locations. However, such techniques are only applicable to resources shared across cores. Achieving such precise interference prediction on real-system simultaneous multithreading (SMT) architectures has been a significantly challenging open problem due to the complexity introduced by sharing resources within a core. In this paper, we demonstrate through a real-system investigation that the fundamental difference between resource sharing behaviors on CMP and SMT architectures calls for a redesign of the way we model interference. For SMT servers, the interference on different shared resources, including private caches, memory ports, as well as integer and floating-point functional units, do not correlate with each other. This insight suggests the necessity of decoupling interference into multiple resource sharing dimensions. In this work, we propose SMiTe, a methodology that enables precise performance prediction for SMT co-location on real-system commodity processors. With a set of Rulers, which are carefully designed software stressors that apply pressure to a multidimensional space of shared resources, we quantify application sensitivity and contentiousness in a decoupled manner. We then establish a regression model to combine the sensitivity and contentiousness in different dimensions to predict performance interference. Using this methodology, we are able to precisely predict the performance interference in SMT co-location with an average error of 2.80% on SPEC CPU2006 and 1.79% on Cloud Suite. Our evaluation shows that SMiTe allows us to improve the utilization of WSCs by up to 42.57% while enforcing an application's QoS requirements. Michael Laurenzano, Jason Mars, Lingjia Tang |
MICRO | 2 |
| 2014 | Making the Most of SMT in HPC: System- and Application-Level PerspectivesabstractThis work presents an end-to-end methodology for quantifying the performance and power benefits of simultaneous multithreading (SMT) for HPC centers and applies this methodology to a production system and workload. Ultimately, SMT’s value system-wide depends on whether users effectively employ SMT at the application level. However, predicting SMT’s benefit for HPC applications is challenging; by doubling the number of threads, the application’s characteristics may change. This work proposes statistical modeling techniques to predict the speedup SMT confers to HPC applications. This approach, accurate to within 8%, uses only lightweight, transparent performance monitors collected during a single run of the application. Leo Porter 0001, Michael Laurenzano, Ananta Tiwari, Adam Jundt, William A. Ward Jr., Roy L. Campbell, Laura Carrington |
ACM Trans. Archit. Code Optim. | 2 |
| 2013 | Understanding the performance of stencil computations on Intel's Xeon PhiabstractAccelerators are becoming prevalent in high performance computing as a way of achieving increased computational capacity within a smaller power budget. Effectively utilizing the raw compute capacity made available by these systems, however, remains a challenge because it can require a substantial investment of programmer time to port and optimize code to effectively use novel accelerator hardware. In this paper we present a methodology for isolating and modeling the performance of common performance-critical patterns of code (so-called idioms) and other relevant behavioral characteristics from large scale HPC applications which are likely to perform favorably on Intel Xeon Phi. The benefits of the methodology are twofold: (1) it directs programmer efforts toward the regions of code most likely to benefit from porting to the Xeon Phi and (2) provides speedup estimates for porting those regions of code. We then apply the methodology to the stencil idiom, showing performance improvements of up to a factor of 4.7× on stencil-based benchmark codes. Joshua Peraza, Ananta Tiwari, Michael Laurenzano, Laura Carrington, William A. Ward Jr., Roy L. Campbell |
CLUSTER | 3 |
| 2011 | Reducing Energy Usage with Memory and Computation-Aware Dynamic Frequency Scaling
Michael Laurenzano, Mitesh R. Meswani, Laura Carrington, Allan Snavely, Mustafa M. Tikir, Stephen W. Poole |
Euro-Par (1) | 1 |
| 2011 | An idiom-finding tool for increasing productivity of acceleratorsabstractSuppose one is considering purchase of a computer equipped with accelerators. Or suppose one has access to such a computer and is considering porting code to take advantage of the accelerators. Is there a reason to suppose the purchase cost or programmer effort will be worth it? It would be nice to able to estimate the expected improvements in advance of paying money or time. We exhibit an analytical framework and tool-set for providing such estimates: the tools first look for user-defined idioms that are patterns of computation and data access identified in advance as possibly being able to benefit from accelerator hardware. A performance model is then applied to estimate how much faster these idioms would be if they were ported and run on the accelerators, and a recommendation is made as to whether or not each idiom is worth the porting effort to put them on the accelerator and an estimate is provided of what the overall application speedup would be if this were done. Laura Carrington, Mustafa M. Tikir, Catherine Mills Olschanowsky, Michael Laurenzano, Joshua Peraza, Allan Snavely, Stephen W. Poole |
ICS | 4 |
| 2010 | PEBIL: Efficient static binary instrumentation for LinuxabstractBinary instrumentation facilitates the insertion of additional code into an executable in order to observe or modify the executable's behavior. There are two main approaches to binary instrumentation: static and dynamic binary instrumentation. In this paper we present a static binary instrumentation toolkit for Linux on the x86/x86_64 platforms, PEBIL (PMaC's Efficient Binary Instrumentation Toolkit for Linux). PEBIL is similar to other toolkits in terms of how additional code is inserted into the executable. However, it is designed with the primary goal of producing efficient-running instrumented code. To this end, PEBIL uses function level code relocation in order to insert large but fast control structures. Furthermore, the PEBIL API provides tool developers with the means to insert lightweight hand-coded assembly rather than relying solely on the insertion of instrumentation functions. These features enable the implementation of efficient instrumentation tools with PEBIL. The overhead introduced for basic block counting by PEBIL is an average of 65% of the overhead of Dyninst, 41% of the overhead of Pin, 15% of the overhead of DynamoRIO, and 8% of the overhead of Valgrind. Michael Laurenzano, Mustafa M. Tikir, Laura Carrington, Allan Snavely |
ISPASS | 1 |
| 2009 | PSINS: An Open Source Event Tracer and Execution Simulator for MPI Applications
Mustafa M. Tikir, Michael Laurenzano, Laura Carrington, Allan Snavely |
Euro-Par | 2 |
| 2008 | High-frequency simulations of global seismic wave propagation using SPECFEM3D_GLOBE on 62K processorsabstractSPECFEM3D_GLOBE is a spectral-element application enabling the simulation of global seismic wave propagation in 3D anelastic, anisotropic, rotating and self-gravitating Earth models at unprecedented resolution. A fundamental challenge in global seismology is to model the propagation of waves with periods between 1 and 2 seconds, the highest frequency signals that can propagate clear across the Earth. These waves help reveal the 3D structure of the Earth's deep interior and can be compared to seismographic recordings. We broke the 2 second barrier using the 62K processor Ranger system at TACC. Indeed we broke the barrier using just half of Ranger, by reaching a period of 1.84 seconds with sustained 28.7 Tflops on 32K processors. We obtained similar results on the XT4 Franklin system at NERSC and the XT4 Kraken system at University of Tennessee Knoxville, while a similar run on the 28K processor Jaguar system at ORNL, which has better memory bandwidth per processor, sustained 35.7 Tflops (a higher flops rate) with a 1.94 shortest period. Laura Carrington, Dimitri Komatitsch, Michael Laurenzano, Mustafa M. Tikir, David Michéa, Nicolas Le Goff, Allan Snavely, Jeroen Tromp |
SC | 3 |
| 2005 | How Well Can Simple Metrics Represent the Performance of HPC Applications?abstractA systematic study of the effects of complexity of prediction methodology on its accuracy for a set of real applications on a variety of HPC systems is performed. Results indicate that the use of any single, simple synthetic metric to predict performance does an inadequate job, and the use of a linear combination of these simple metrics with optimized weights also performs poorly. Better, however, are methodologies that rely on the convolution of an application "transfer function" based on tracing information with system performance data measured by simple benchmarks. This latter methodology can predict performance with an average accuracy of 80%, based on the current work. Laura Carrington, Michael Laurenzano, Allan Snavely, Roy L. Campbell, Larry P. Davis |
SC | 2 |