VLDB 2026 Research / reviewers in the wild / expert
Allan Snavely
dblp:55/1134
· DBLP profile ↗
34ranked-venue papers
5as first author
0since 2021 · last 2016
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 31 · 5 first-authorSoftware engineering, systems software and programming languages · 3 · 2 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
17 papers |
Performance modeling and evaluation · 35% High-performance computing · 26% Memory systems · 14% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Environmental and earth informatics · 100% | |
| Software engineering, system software, and programming languages
4 papers |
Operating systems · 77% Compilers and program optimization · 23% |
Topics — the 30 heaviest of 46, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Performance modeling and evaluation
workload characterization |
0.3 | 5 | 2008 | Code coverage, performance approximation and automatic recognition of idioms in scientific applications · HPDC 2008 Precise and realistic utility functions for user-centric performance analysis of schedulers · HPDC 2007 Evaluating petascale - Evaluating petascale infrastructure systems: benchmarks, models, and applications · SC 2006 |
Performance modeling and evaluation
benchmarking |
0.2 | 4 | 2007 | A genetic algorithms approach to modeling the performance of memory-bound computations · SC 2007 Evaluating petascale - Evaluating petascale infrastructure systems: benchmarks, models, and applications · SC 2006 How Well Can Simple Metrics Represent the Performance of HPC Applications? · SC 2005 |
Performance modeling and evaluation
performance prediction |
0.2 | 3 | 2007 | A genetic algorithms approach to modeling the performance of memory-bound computations · SC 2007 How Well Can Simple Metrics Represent the Performance of HPC Applications? · SC 2005 A framework for performance modeling and prediction · SC 2002 |
Storage systems
flash and SSD |
0.1 | 2 | 2010 | DASH: a Recipe for a Flash-based Data Intensive Supercomputer · SC 2010 Understanding the Impact of Emerging Non-Volatile Memories on High-Performance, IO-Intensive Computing · SC 2010 |
Cloud and datacenter computing
job scheduling |
0.1 | 3 | 2006 | When Jobs Play Nice: The Case For Symbiotic Space-Sharing · HPDC 2006 Symbiotic jobscheduling with priorities for a simultaneous multithreading processor · SIGMETRICS 2002 Symbiotic Jobscheduling for a Simultaneous Multithreading Processor · ASPLOS 2000 |
High-performance computing
data-intensive computing |
0.1 | 1 | 2010 | DASH: a Recipe for a Flash-based Data Intensive Supercomputer · SC 2010 |
Memory systems
non-volatile memory |
0.1 | 1 | 2010 | Understanding the Impact of Emerging Non-Volatile Memories on High-Performance, IO-Intensive Computing · SC 2010 |
High-performance computing
performance optimization at scale |
0.1 | 2 | 2008 | High-frequency simulations of global seismic wave propagation using SPECFEM3D_GLOBE on 62K processors · SC 2008 A genetic algorithms approach to modeling the performance of memory-bound computations · SC 2007 |
High-performance computing › scientific computing systems
earthquake simulation |
0.1 | 1 | 2008 | High-frequency simulations of global seismic wave propagation using SPECFEM3D_GLOBE on 62K processors · SC 2008 |
High-performance computing
scientific computing systems |
0.1 | 1 | 2008 | High-frequency simulations of global seismic wave propagation using SPECFEM3D_GLOBE on 62K processors · SC 2008 |
High-performance computing › finite element method
spectral-element method |
0.1 | 1 | 2008 | High-frequency simulations of global seismic wave propagation using SPECFEM3D_GLOBE on 62K processors · SC 2008 |
Environmental and earth informatics
atmospheric modeling |
0.1 | 1 | 2007 | WRF nature run · SC 2007 |
Environmental and earth informatics › atmospheric modeling
numerical weather prediction |
0.1 | 1 | 2007 | WRF nature run · SC 2007 |
Cloud and datacenter computing › cluster resource management and scheduling
cluster scheduling |
0.1 | 1 | 2007 | Precise and realistic utility functions for user-centric performance analysis of schedulers · HPDC 2007 |
Memory systems
memory-bound computation |
0.1 | 1 | 2007 | A genetic algorithms approach to modeling the performance of memory-bound computations · SC 2007 |
High-performance computing › large-scale simulation
petascale simulation |
0.1 | 1 | 2007 | WRF nature run · SC 2007 |
Cloud and datacenter computing › job scheduling › economic scheduling
utility-based scheduling |
0.1 | 1 | 2007 | Precise and realistic utility functions for user-centric performance analysis of schedulers · HPDC 2007 |
Processor architecture and microarchitecture › multithreading
simultaneous multithreading |
0.1 | 2 | 2002 | Symbiotic jobscheduling with priorities for a simultaneous multithreading processor · SIGMETRICS 2002 Symbiotic Jobscheduling for a Simultaneous Multithreading Processor · ASPLOS 2000 |
High-performance computing › supercomputing
petascale computing |
0.1 | 1 | 2006 | Evaluating petascale - Evaluating petascale infrastructure systems: benchmarks, models, and applications · SC 2006 |
Parallel and multicore computing › parallel scheduling
space-sharing |
0.1 | 1 | 2006 | When Jobs Play Nice: The Case For Symbiotic Space-Sharing · HPDC 2006 |
Performance modeling and evaluation
trace compression |
0.1 | 1 | 2006 | Path Grammar Guided Trace Compression and Trace Approximation · HPDC 2006 |
Performance modeling and evaluation › simulation › discrete-event simulation
trace-driven simulation |
0.1 | 1 | 2006 | Path Grammar Guided Trace Compression and Trace Approximation · HPDC 2006 |
Performance modeling and evaluation
application performance modeling |
0.1 | 1 | 2005 | How Well Can Simple Metrics Represent the Performance of HPC Applications? · SC 2005 |
Memory systems › cache
cache performance |
0.1 | 1 | 2005 | Quantifying Locality In The Memory Access Patterns of HPC Applications · SC 2005 |
Performance modeling and evaluation › workload characterization
locality analysis |
0.1 | 1 | 2005 | Quantifying Locality In The Memory Access Patterns of HPC Applications · SC 2005 |
Memory systems
memory access patterns |
0.1 | 1 | 2005 | Quantifying Locality In The Memory Access Patterns of HPC Applications · SC 2005 |
Memory systems › data locality
spatial and temporal locality |
0.1 | 1 | 2005 | Quantifying Locality In The Memory Access Patterns of HPC Applications · SC 2005 |
Operating systems › resource management › process management
CPU scheduling |
0.0 | 2 | 2002 | Symbiotic Jobscheduling for a Simultaneous Multithreading Processor · ASPLOS 2000 Symbiotic jobscheduling with priorities for a simultaneous multithreading processor · SIGMETRICS 2002 |
Performance modeling and evaluation
analytical modeling |
0.0 | 1 | 2002 | A framework for performance modeling and prediction · SC 2002 |
Parallel and multicore computing › parallel scheduling
coscheduling |
0.0 | 1 | 2002 | Symbiotic jobscheduling with priorities for a simultaneous multithreading processor · SIGMETRICS 2002 |
Methods — techniques the papers use, named apart from their topics
performance measurement · 0.2idiom benchmarking · 0.2code coverage analysis · 0.2spectral methods · 0.1parallel i/o · 0.1genetic algorithm · 0.1performance evaluation · 0.1spectral-element method · 0.1MultiMAPS · 0.1Apex-MAPS · 0.1simulation · 0.0symbiosis prediction · 0.0sampling · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2016 | The case for colocation of high performance computing workloadsabstractSummary The current state of practice in supercomputer resource allocation places jobs from different users on disjoint nodes both in terms of time and space. While this approach largely guarantees that jobs from different users do not degrade one another's performance, it does so at high cost to system throughput and energy efficiency. This focused study presents job striping, a technique that significantly increases performance over the current allocation mechanism by colocating pairs of jobs from different users on a shared set of nodes. To evaluate the potential of job striping in large‐scale environments, the experiments are run at the scale of 128 nodes on the state‐of‐the‐art Gordon supercomputer. Across all pairings of 1024 process network‐attached storage parallel benchmarks, job striping increases mean throughput by 26% and mean energy efficiency by 22%. On pairings of the real applications Gyrokinetic Toroidal Code (GTC), Large‐scale Atomic/Molecular Massively Parallel Simulator (LAMMPS), and MIMD Lattice Computation (MILC) at equal scale, job striping improves average throughput by 12% and mean energy efficiency by 11%. In addition, the study provides a simple set of heuristics for avoiding low performing application pairs. Copyright © 2013 John Wiley & Sons, Ltd. Alexander Dodd Breslow, Leo Porter 0001, Ananta Tiwari, Michael Laurenzano, Laura Carrington, Dean M. Tullsen, Allan Snavely |
Concurr. Comput. Pract. Exp. | 7 |
| 2016 | PMaC's green queue: a framework for selecting energy optimal DVFS configurations in large scale MPI applicationsabstractSummary This article presents Green Queue, a production quality tracing and analysis framework for implementing application aware dynamic voltage and frequency scaling (DVFS) for message passing interface applications in high performance computing. Green Queue makes use of both intertask and intratask DVFS techniques. The intertask technique targets applications where the workload is imbalanced by reducing CPU clock frequency and therefore power draw for ranks with lighter workloads. The intratask technique targets balanced workloads where all tasks are synchronously running the same code. The strategy identifies program phases and selects the energy‐optimal frequency for each by predicting power and measuring the performance responses of each phase to frequency changes. The success of these techniques is evaluated on 1024 cores on Gordon, a supercomputer at the San Diego Supercomputer Center built using Intel Xeon E5‐2670 (Sandybridge) processors. Green Queue achieves up to 21% and 32% energy savings for the intratask and intertask DVFS strategies, respectively. Copyright © 2013 John Wiley & Sons, Ltd. Joshua Peraza, Ananta Tiwari, Michael Laurenzano, Laura Carrington, Allan Snavely |
Concurr. Comput. Pract. Exp. | 5 |
| 2011 | Reducing Energy Usage with Memory and Computation-Aware Dynamic Frequency Scaling
Michael Laurenzano, Mitesh R. Meswani, Laura Carrington, Allan Snavely, Mustafa M. Tikir, Stephen W. Poole |
Euro-Par (1) | 4 |
| 2011 | An idiom-finding tool for increasing productivity of acceleratorsabstractSuppose one is considering purchase of a computer equipped with accelerators. Or suppose one has access to such a computer and is considering porting code to take advantage of the accelerators. Is there a reason to suppose the purchase cost or programmer effort will be worth it? It would be nice to able to estimate the expected improvements in advance of paying money or time. We exhibit an analytical framework and tool-set for providing such estimates: the tools first look for user-defined idioms that are patterns of computation and data access identified in advance as possibly being able to benefit from accelerator hardware. A performance model is then applied to estimate how much faster these idioms would be if they were ported and run on the accelerators, and a recommendation is made as to whether or not each idiom is worth the porting effort to put them on the accelerator and an estimate is provided of what the overall application speedup would be if this were done. Laura Carrington, Mustafa M. Tikir, Catherine Mills Olschanowsky, Michael Laurenzano, Joshua Peraza, Allan Snavely, Stephen W. Poole |
ICS | 6 |
| 2011 | Automatic Recognition of Performance Idioms in Scientific ApplicationsabstractBasic data flow patterns that we call \textbf{performance idioms}, such as stream, transpose, reduction, random access and stencil, are common in scientific numerical applications. We hypothesize that a small number of idioms can cover most programming constructs that dominate the execution time of scientific codes and can be used to approximate the application performance. To check these hypotheses, we proposed an automatic idioms recognition method and implemented the method, based on the open source compiler Open64. With the NAS Parallel Benchmark (NPB) as a case study, the prototype system is about $90%$ accurate compared with idiom classification by a human expert. Our results showed that the above five idioms suffice to cover $100%$ of the six NPB codes (MG, CG, FT, BT, SP and LU). We also compared the performance of our idiom benchmarks with their corresponding instances in the NPB codes on two different platforms with different methods. The approximation accuracy is up to $96.6%$. The contribution is to show that a small set of idioms can cover more complex codes, that idioms can be recognized automatically, and that suitably defined idioms may approximate application performance. Jiahua He, Allan Snavely, Rob F. Van der Wijngaart, Michael A. Frumkin |
IPDPS | 2 |
| 2010 | PEBIL: Efficient static binary instrumentation for LinuxabstractBinary instrumentation facilitates the insertion of additional code into an executable in order to observe or modify the executable's behavior. There are two main approaches to binary instrumentation: static and dynamic binary instrumentation. In this paper we present a static binary instrumentation toolkit for Linux on the x86/x86_64 platforms, PEBIL (PMaC's Efficient Binary Instrumentation Toolkit for Linux). PEBIL is similar to other toolkits in terms of how additional code is inserted into the executable. However, it is designed with the primary goal of producing efficient-running instrumented code. To this end, PEBIL uses function level code relocation in order to insert large but fast control structures. Furthermore, the PEBIL API provides tool developers with the means to insert lightweight hand-coded assembly rather than relying solely on the insertion of instrumentation functions. These features enable the implementation of efficient instrumentation tools with PEBIL. The overhead introduced for basic block counting by PEBIL is an average of 65% of the overhead of Dyninst, 41% of the overhead of Pin, 15% of the overhead of DynamoRIO, and 8% of the overhead of Valgrind. Michael Laurenzano, Mustafa M. Tikir, Laura Carrington, Allan Snavely |
ISPASS | 4 |
| 2010 | Understanding the Impact of Emerging Non-Volatile Memories on High-Performance, IO-Intensive ComputingabstractEmerging storage technologies such as flash memories, phase-change memories, and spin-transfer torque memories are poised to close the enormous performance gap between disk-based storage and main memory. We evaluate several approaches to integrating these memories into computer systems by measuring their impact on IO-intensive, database, and memory-intensive applications. We explore several options for connecting solid-state storage to the host system and find that the memories deliver large gains in sequential and random access performance, but that different system organizations lead to different performance trade-offs. The memories provide substantial application-level gains as well, but overheads in the OS, file system, and application can limit performance. As a result, fully exploiting these memories' potential will require substantial changes to application and system software. Finally, paging to fast non-volatile memories is a viable option for some applications, providing an alternative to expensive, powerhungry DRAM for supporting scientific applications with large memory footprints. Adrian M. Caulfield, Joel Coburn, Todor I. Mollov, Arup De, Ameen Akel, Jiahua He, Arun Jagatheesan, Rajesh K. Gupta 0001, Allan Snavely, Steven Swanson |
SC | 9 |
| 2010 | DASH: a Recipe for a Flash-based Data Intensive SupercomputerabstractData intensive computing can be defined as computation involving large datasets and complicated I/O patterns. Data intensive computing is challenging because there is a five-orders-of-magnitude latency gap between main memory DRAM and spinning hard disks; the result is that an inordinate amount of time in data intensive computing is spent accessing data on disk. To address this problem we designed and built a prototype data intensive supercomputer named DASH that exploits flash-based Solid State Drive (SSD) technology and also virtually aggregated DRAM to fill the latency gap . DASH uses commodity parts including Intel® X25-E flash drives and distributed shared memory (DSM) software from ScaleMP®. The system is highly competitive with several commercial offerings by several metrics including achieved IOPS (input output operations per second), IOPS per dollar of system acquisition cost, IOPS per watt during operation, and IOPS per gigabyte (GB) of available storage. We present here an overview of the design of DASH, an analysis of its cost efficiency, then a detailed recipe for how we designed and tuned it for high data-performance, lastly show that running data-intensive scientific applications from graph theory, biology, and astronomy, we achieved as much as two orders-of- magnitude speedup compared to the same applications run on traditional architectures. Jiahua He, Arun Jagatheesan, Sandeep K. S. Gupta, Jeffrey Bennett, Allan Snavely |
SC | 5 |
| 2009 | PSINS: An Open Source Event Tracer and Execution Simulator for MPI Applications
Mustafa M. Tikir, Michael Laurenzano, Laura Carrington, Allan Snavely |
Euro-Par | 4 |
| 2008 | Code coverage, performance approximation and automatic recognition of idioms in scientific applicationsabstractBasic data flow patterns which we call idioms, such as stream, transpose, reduction, random access and stencil, are common in scientific numerical applications. We hypothesize that a small number of idioms can cover most programming constructs that dominate the execution time of scientific codes and can be used to approximate the application performance. In this paper, we start with a manual analysis of code coverage on the NAS Parallel Benchmark (NPB) and find that five idioms suffice to cover 100% of the NPB codes. We then compare the performance of our idiom benchmarks and their corresponding instances in different NPB codes on two different platforms and find that they differ by about 30%. To check the hypotheses with real applications further, we propose an automatic idioms recognition method, implement the method basing on the open source compiler Open64, and verify the prototype system with the previous manual analysis results. Jiahua He, Allan Snavely, Rob F. Van der Wijngaart, Michael A. Frumkin |
HPDC | 2 |
| 2008 | Accurate memory signatures and synthetic address traces for HPC applicationsabstractThough the performance of many scientific codes is dominated by memory behavior, our ability to describe, capture, compare, and recreate that behavior is quite limited. This inability underlies much of the complexity in the field of performance analysis: it is fundamentally difficult to relate benchmarks and applications or use realistic workloads to guide system design and procurement. An observable, reproducible, and machine-independent memory characterization is needed. Jonathan Weinberg, Allan Snavely |
ICS | 2 |
| 2008 | High-frequency simulations of global seismic wave propagation using SPECFEM3D_GLOBE on 62K processorsabstractSPECFEM3D_GLOBE is a spectral-element application enabling the simulation of global seismic wave propagation in 3D anelastic, anisotropic, rotating and self-gravitating Earth models at unprecedented resolution. A fundamental challenge in global seismology is to model the propagation of waves with periods between 1 and 2 seconds, the highest frequency signals that can propagate clear across the Earth. These waves help reveal the 3D structure of the Earth's deep interior and can be compared to seismographic recordings. We broke the 2 second barrier using the 62K processor Ranger system at TACC. Indeed we broke the barrier using just half of Ranger, by reaching a period of 1.84 seconds with sustained 28.7 Tflops on 32K processors. We obtained similar results on the XT4 Franklin system at NERSC and the XT4 Kraken system at University of Tennessee Knoxville, while a similar run on the 28K processor Jaguar system at ORNL, which has better memory bandwidth per processor, sustained 35.7 Tflops (a higher flops rate) with a 1.94 shortest period. Laura Carrington, Dimitri Komatitsch, Michael Laurenzano, Mustafa M. Tikir, David Michéa, Nicolas Le Goff, Allan Snavely, Jeroen Tromp |
SC | 7 |
| 2007 | Precise and realistic utility functions for user-centric performance analysis of schedulersabstractUtility functions can be used to represent the value users attach to job completion as a function of turnaround time. Most previous scheduling research used simple synthetic representations of utility, with the simplicity being due to the fact that real user preferences are difficult to obtain, and perhaps concern that arbitrarily complex utility functions could in turn make the scheduling problem intractable. In this work, we advocate a flexible representation of utility functions that can indeed be arbitrarily complex. We show that a genetic algorithm heuristic can improve global utility by analyzing these functions, and does so tractably. Since our previous work showed that users indeed have and can articulate complicated utility functions, the result here is relevant. We then provide a means to augment existing workload traces with realistic utility functions for the purpose of enabling realistic scheduling simulations. Cynthia Bailey, Allan Snavely |
HPDC | 2 |
| 2007 | WRF nature runabstractThe Weather Research and Forecast (WRF) model is a limited-area model of the atmosphere for mesoscale research and operational numerical weather prediction (NWP). A petascale problem is a WRF nature run that provides very high-resolution "truth" against which more coarse simulations or perturbation runs may be compared for purposes of studying predictability, stochastic parameterization, and fundamental dynamics. We carried out a nature run involving an idealized high resolution rotating fluid on the hemisphere to investigate scales that span the k-3 to k-5/3 kinetic energy spectral transition of the observed atmosphere using 65,536 processors of the BG/L machine at LLNL. We worked through issues of parallel I/O and scalability. The primary result is not just the scalability and high Tflops number, but an important step towards understanding weather predictability at high resolution. John Michalakes, Josh Hacker, Richard Loft, Michael O. McCracken, Allan Snavely, Nicholas J. Wright, Thomas E. Spelce, Brent C. Gorda, Robert Walkup |
SC | 5 |
| 2007 | A genetic algorithms approach to modeling the performance of memory-bound computationsabstractBenchmarks that measure memory bandwidth, such as STREAM, Apex-MAPS and MultiMAPS, are increasingly popular due to the "Von Neumann" bottleneck of modern processors which causes many calculations to be memory-bound. We present a scheme for predicting the performance of HPC applications based on the results of such benchmarks. A Genetic Algorithm approach is used to "learn" bandwidth as a function of cache hit rates per machine with MultiMAPS as the fitness test. The specific results are 56 individual performance predictions including 3 full-scale parallel applications run on 5 different modern HPC architectures, with various CPU counts and inputs, predicted within 10 % average difference with respect to independently verified runtimes. Mustafa M. Tikir, Laura Carrington, Erich Strohmaier, Allan Snavely |
SC | 4 |
| 2006 | Topic 2: Performance Prediction and Evaluation
Jesús Labarta, Bernd Mohr, Allan Snavely, Jeffrey S. Vetter |
Euro-Par | 3 |
| 2006 | Path Grammar Guided Trace Compression and Trace ApproximationabstractTrace-driven simulation is an important technique used in the evaluation of computer architecture innovations. However using it for studying parallel computers and applications is at best very challenging. Acquiring, representing and storing the traces are among the major issues. In this paper, we introduce path grammar guided trace compression (PGGTC) and effective address trace approximation (TA) to speedup compression and reduce trace sizes. PGGTC relies on static analysis to build rules and determine actions to guide online trace compression. Combined with gzip, PGGTC can compresses control flow traces over 330 times smaller than using gzip alone. Compared to the widely popular Sequitur algorithm alone, PGGTC with gzip is on average 40 times faster, while the traces are only 3 times bigger. PGGTC can be also used with Sequitur to double the compression ratios of Sequitur by itself and do it 14 times faster than Sequitur by itself. Address traces of parallel applications with significant randomness are often impossibly large even after being compressed with any lossless scheme including PGGTC. For effective address trace reduction, we introduce trace approximation (TA). Performance-wise similar effective addresses are generated based on very compact summaries of how the memory is accessed during each structure instance instead of compressing them. We demonstrate two approaches: selective dumping and memory signatures, to summarize the properties of effective address sequences. Both approaches are validated by feeding the generated approximate trace to cache simulators of 25 different configurations. The simulated results are very close to the simulation results based on full effective traces while the selective dumped address or memory signatures require several order of magnitude less disk space to store. In summary, we move trace-driven simulation into the realm of the feasible for larger parallel machines and applications Xiaofeng Gao 0003, Allan Snavely, Larry Carter |
HPDC | 2 |
| 2006 | When Jobs Play Nice: The Case For Symbiotic Space-SharingabstractUsing a large HPC platform, we investigate the effectiveness of "symbiotic space-sharing", a technique that improves system throughput by executing parallel applications in combinations and configurations that alleviate pressure on shared resources. We demonstrate that relevant benchmarks commonly suffer a 10-60% penalty in runtime efficiency due to memory resource bottlenecks and up to several orders of magnitude for I/O. We show that this penalty can be often mitigated, and sometimes virtually eliminated, by symbiotic space-sharing techniques and deploy a prototype scheduler that leverages these findings to improve system throughput by 20% Jonathan Weinberg, Allan Snavely |
HPDC | 2 |
| 2006 | User-guided symbiotic space-sharing of real workloadsabstractSymbiotic space-sharing is a technique that can improve system throughput by executing parallel applications in combinations and configurations that alleviate pressure on shared resources. We have shown prototype schedulers that leverage such techniques to improve throughput by 20% over conventional space-sharing schedulers when resource bottlenecks are known. Such evaluations have utilized benchmark workloads and proposed that schedulers be informed of resource bottlenecks by users at job submission time; in this work, we investigate the accuracy with which users can actually identify resource bottlenecks in real applications and the implications of these predictions for symbiotic space-sharing of production workloads. Using a large HPC platform, a representative application workload, and a sampling of expert users, we show that user inputs are of value and that for our chosen workload, user-guided symbiotic scheduling can improve throughput over conventional space-sharing by 15-22%. Jonathan Weinberg, Allan Snavely |
ICS | 2 |
| 2006 | Symbiotic Space-Sharing on SDSC's DataStar System
Jonathan Weinberg, Allan Snavely |
JSSPP | 2 |
| 2006 | Evaluating petascale - Evaluating petascale infrastructure systems: benchmarks, models, and applicationsabstractThis BOF is a venue for presentations and discussions of progress in the dual problems of evaluating the performance (including reliability) of petascale computing systems and of developing application codes that scale to run effectively on such systems. We therefore invite participation by representatives of large-system vendors, funding agencies, infrastructure operators, appplication-development groups, and others that have a stake in design, operation, and successful use of these systems.Technical topics to be addressed include: scaling properties of benchmark suites; scalable machine models; application modeling to predict scaling for future machines; the problem of preparing applications for petascale environments; and the role of all of these topics in petascale acquisitions.At this session we will start planning for a series of workshops on the subject. Robert J. Fowler, Allan Snavely, Daniel A. Reed |
SC | 2 |
| 2006 | 99% utilization - Is 99% utilization of a supercomputer a good thing?abstractThis BOF will continue debate revolving around productivity metrics for supercomputers. At several recent user forums, consensus emerged that it is not possible to develop petascale applications without interactive access to thousands of processors. But most large systems are managed via a batch scheduler with long (and unpredictable) queue wait times. Most batch scheduler policies assume high system utilization as "good". But high utilization dilates average queue wait time and increases wait-time unpredictability, both of which are "bad" for application developer's productivity. What are the options to address these conflicting implications for running a supercomputer at high system utilization? Is it possible to manage a supercomputer to meet the high-throughput demands of stable applications and the on-demand access requirements of large-scale code developers concurrently? Or do these two usage scenarios inherently conflict? Participants will explain and debate several creative solutions that could enable high throughput and high availability for program development. Allan Snavely, Jeremy Kepner |
SC | 1 |
| 2006 | A performance prediction framework for scientific applications
Laura Carrington, Allan Snavely, Nicole Wolter |
Future Gener. Comput. Syst. | 2 |
| 2006 | Special section: Large-scale system performance modeling and analysis
Adolfy Hoisie, Darren J. Kerbyson, Celso L. Mendes, Daniel A. Reed, Allan Snavely |
Future Gener. Comput. Syst. | 5 |
| 2005 | Performance Modeling: Understanding the Past and Predicting the Future
David H. Bailey, Allan Snavely |
Euro-Par | 2 |
| 2005 | Topic 2 - Performance Prediction and Evaluation
Allen D. Malony, Thomas Fahringer, Allan Snavely, Luís Silva |
Euro-Par | 3 |
| 2005 | How Well Can Simple Metrics Represent the Performance of HPC Applications?abstractA systematic study of the effects of complexity of prediction methodology on its accuracy for a set of real applications on a variety of HPC systems is performed. Results indicate that the use of any single, simple synthetic metric to predict performance does an inadequate job, and the use of a linear combination of these simple metrics with optimized weights also performs poorly. Better, however, are methodologies that rely on the convolution of an application "transfer function" based on tracing information with system performance data measured by simple benchmarks. This latter methodology can predict performance with an average accuracy of 80%, based on the current work. Laura Carrington, Michael Laurenzano, Allan Snavely, Roy L. Campbell, Larry P. Davis |
SC | 3 |
| 2005 | Quantifying Locality In The Memory Access Patterns of HPC ApplicationsabstractSeveral benchmarks for measuring the memory performance of HPC systems along dimensions of spatial and temporal memory locality have recently been proposed. However, little is understood about the relationships of these benchmarks to real applications and to each other. We propose a methodology for producing architecture-neutral characterizations of the spatial and temporal locality exhibited by the memory access patterns of applications. We demonstrate that the results track intuitive notions of locality on several synthetic and application benchmarks. We employ the methodology to analyze the memory performance components of the HPC Challenge Benchmarks, the Apex-MAP benchmark, and their relationships to each other and other benchmarks and applications. We show that this analysis can be used to both increase understanding of the benchmarks and enhance their usefulness by mapping them, along with applications, to a 2-D space along axes of spatial and temporal locality. Jonathan Weinberg, Michael O. McCracken, Erich Strohmaier, Allan Snavely |
SC | 4 |
| 2004 | Benchmark Probes for Grid AssessmentabstractSummary form only given. Like all computing platforms, grids are in need of a suite of benchmarks by which they can be evaluated, compared and characterized. As a first step towards this goal, we have developed a set of probes that exercise basic grid operations with the goal of measuring the performance and the performance variability of basic grid operations, as well as the failure rates of these operations. We present measurement data obtained by running our probes on a grid testbed that spans 5 clusters in 3 institutions. These measurements quantify compute times, network transfer times, and Globus middleware overhead. Our results help provide insight into the stability, robustness, and performance of our testbed, and lead us to make some recommendations for future grid development. Greg Chun, Holly Dail, Henri Casanova, Allan Snavely |
IPDPS | 4 |
| 2004 | Are User Runtime Estimates Inherently Inaccurate?
Cynthia Bailey, Yael Schwartzman, Jennifer Hardy, Allan Snavely |
JSSPP | 4 |
| 2002 | A framework for performance modeling and predictionabstractCycle-accurate simulation is far too slow for modeling the expected performance of full parallel applications on large HPC systems. And just running an application on a system and observing wallclock time tells you nothing about why the application performs as it does (and is anyway impossible on yet-to-be-built systems). Here we present a framework for performance modeling and prediction that is faster than cycle-accurate simulation, more informative than simple benchmarking, and is shown useful for performance investigations in several dimensions. Allan Snavely, Laura Carrington, Nicole Wolter, Jesús Labarta, Rosa M. Badia, Avi Purkayastha |
SC | 1 |
| 2002 | Symbiotic jobscheduling with priorities for a simultaneous multithreading processorabstractSimultaneous Multithreading machines benefit from jobscheduling software that monitors how well coscheduled jobs share CPU resources, and coschedules jobs that interact well to make more efficient use of those resources. As a result, informed coscheduling can yield significant performance gains over naive schedulers. However, prior work on coscheduling focused on equal-priority job mixes, which is an unrealistic assumption for modern operating systems.This paper demonstrates that a scheduler for an SMT machine can both satisfy process priorities and symbiotically schedule low and high priority threads to increase system throughput. Naive priority schedulers dedicate the machine to high priority jobs to meet priority goals, and as a result decrease opportunities for increased performance from multithreading and coscheduling. More informed schedulers, however, can dynamically monitor the progress and resource utilization of jobs on the machine, and dynamically adjust the degree of multithreading to improve performance while still meeting priority goals.Using detailed simulation of an SMT architecture, we introduce and evaluate a series of five software and hardware-assisted priority schedulers. Overall, our results indicate that coscheduling priority jobs can significantly increase system throughput by as much as 40%, and that (1) the benefit depends upon the relative priority of the coscheduled jobs, and (2) more sophisticated schedulers are more effective when the differences in priorities are greatest. We show that our priority schedulers can decrease average turnaround times for a random jobmix by as much as 33%. Allan Snavely, Dean M. Tullsen, Geoffrey M. Voelker |
SIGMETRICS | 1 |
| 2000 | Symbiotic Jobscheduling for a Simultaneous Multithreading ProcessorabstractSimultaneous Multithreading machines fetch and execute instructions from multiple instruction streams to increase system utilization and speedup the execution of jobs. When there are more jobs in the system than there is hardware to support simultaneous execution, the operating system scheduler must choose the set of jobs to coscheduleThis paper demonstrates that performance on a hardware multithreaded processor is sensitive to the set of jobs that are coscheduled by the operating system jobscheduler. Thus, the full benefits of SMT hardware can only be achieved if the scheduler is aware of thread interactions. Here, a mechanism is presented that allows the scheduler to significantly raise the performance of SMT architectures. This is done without any advance knowledge of a workload's characteristics, using sampling to identify jobs which run well together.We demonstrate an SMT jobscheduler called SOS. SOS combines an overhead-free sample phase which collects information about various possible schedules, and a symbiosis phase which uses that information to predict which schedule will provide the best performance. We show that a small sample of the possible schedules is sufficient to identify a good schedule quickly. On a system with random job arrivals and departures, response time is improved as much as 17% over a schedule which does not incorporate symbiosis. Allan Snavely, Dean M. Tullsen |
ASPLOS | 1 |
| 1998 | Multi-processor Performance on the Tera MTAabstractThe Tera MTA is a revolutionary commercial computer based on a multithreaded processor architecture. In contrast to many other parallel architectures, the Tera MTA can effectively use high amounts of parallelism on a single processor. By running multiple threads on a single processor, it can tolerate memory latency and to keep the processor saturated. If the computation is sufficiently large, it can benefit from running on multiple processors. A primary architectural goal of the MTA is that it provide scalable performance over multiple processors. This paper is a preliminary investigation of the first multi-processor Tera MTA. In a previous paper [1] we reported that on the kernel NAS 2 benchmarks [2], a single-processor MTA system running at the architected clock speed would be similar in performance to a single processor of the Cray T90. We found that the compilers of both machines were able to find the necessary threads or vector operations, after making standard changes to the random number generator. In this paper we update the single-processor results in two ways: we use only actual clock speeds, and we report improvements given by further tuning of the MTA codes. We then investigate the performance of the best single-processor codes when run on a two-processor MTA, making no further tuning effort. The parallel efficiency of the codes range from 77% to 99%. An analysis shows that the "serial bottlenecks" -- unparallelized code sections and the cost of allocating and freeing the parallel hardware resources -- account for less than a percent of the runtimes. Thus, Amdahl's Law needn't take effect on the NAS benchmarks until there are hundreds of processors running thousands of threads. Instead, the major source of inefficiency appears to be an imperfect network connecting the processors to the memory. Ideally, the network can support one memory reference per instruction. The current hardware has defects that reduce the throughput to about 85% of this rate. Except for the EP benchmark, the tuned codes issue memory references at nearly the peak rate of one per instruction. Consequently, the network can support the memory references issued by one, but not two, processors. As a result, the parallel efficiency of EP is near- perfect, but the others are reduced accordingly. Another reason for imperfect speedup pertains to the compiler. While the definition of a thread in a single processor or multi-processor mode is essentially the same, there is a different implementation and an associated overhead with running on multiple processors. We characterize the overhead of running "frays" (a collection of threads running on a single processor) and "crews" (a collection of frays, one per processor.) Allan Snavely, Larry Carter, Jay Boisseau, Amitava Majumdar 0001, Kang Su Gatlin, Nick Mitchell, John Feo, Brian D. Koblenz |
SC | 1 |