VLDB 2026 Research / reviewers in the wild / expert
Valerie Taylor 0001
dblp:50/1380 · also Valerie E. Taylor
· DBLP profile ↗
49ranked-venue papers
9as first author
9since 2021 · last 2026
0000-0001-6061-0191ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 38 · 7 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 5 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 3Artificial intelligence and machine learning · 2Graphics, computer vision, multimedia, augmented reality and games · 1Theory of computation · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SmartCap: Coordinated CPU-GPU Power Capping for Performance-Assurance Energy EfficiencyabstractPerformance prediction is essential for energy-efficient computing in heterogeneous computing systems that integrate CPUs and GPUs. However, traditional performance modeling methods often rely on exhaustive offline profiling, which becomes impractical due to the large setting space and the high cost of profiling large-scale applications. In this paper, we present OPEN, a framework consists of offline and online phases. The offline phase involves building a performance predictor and constructing an initial dense matrix. In the online phase, OPEN performs lightweight online profiling, and leverages the performance predictor with collaborative filtering to make performance prediction. We evaluate OPEN on multiple heterogeneous systems, including those equipped with A100 and A30 GPUs. Results show that OPEN achieves prediction accuracy up to 98.29\%. This demonstrates that OPEN effectively reduces profiling cost while maintaining high accuracy, making it practical for power-aware performance modeling in modern HPC environments. Overall, OPEN provides a lightweight solution for performance prediction under power constraints, enabling better runtime decisions in power-aware computing environments. Xingfu Wu, Valerie Taylor 0001, Michael E. Papka, Zhiling Lan |
ICS | 3 |
| 2026 | Beyond Throughput: Performance and Energy Insights of LLM Inference Across AI Accelerators
Giacomo Brunetta, Varuni Sastry 0001, Xingfu Wu, Valerie Taylor 0001, Michael E. Papka, Zhiling Lan |
IPDPS | 4 |
| 2025 | Evaluating Energy Efficiency of Ai Accelerators Using Two Mlperf BenchmarksabstractThe significantly increasing use of artificial intelligence (AI) has led to the availability of specialized AI accelerators, aiming to enhance the performance and energy efficiency of AI workloads. In this paper, we conduct an initial study to evaluate the energy requirements of four AI accelerators: Nvidia A100 GPUs, Intel Habana Gaudi Processing Units (HPUs), Graphcore Bow-Pod64 Intelligence Processing Units (IPUs), and GroqRack Language Processing Units (LPUs) using two popular MLPerf benchmarks: BERT-Large and ResNet50. We report the energy requirements for the two benchmarks to achieve a common MLPerfspecified target accuracy. The benchmarks and AI accelerators were chosen based on the following criteria: publicly available tools or libraries from the vendors to monitor power consumption, publicly available optimized models from the vendors, and access to the AI accelerators. Our experimental results indicate that for ResNet50, Intel Gaudi2 HPUs delivered the highest throughput, the lowest energy consumption for both training and inference, and the highest inference energy efficiency, while Graphcore demonstrated the highest training energy efficiency. For BERTLarge pre-training, Intel Gaudi2 outperformed both Nvidia A100 and Graphcore in terms of time, energy consumption, and training energy efficiency. However, for BERT-Large inference, Nvidia A100 achieved the shortest time and lowest energy consumption; Graphcore exhibited the highest throughput in both pre-training and inference, along with the highest inference energy efficiency. We discuss our observations and findings while exploring the associated tradeoffs. Farah Ferdaus, Xingfu Wu, Valerie Taylor 0001, Zhiling Lan, Sanjif Shanmugavelu, Venkatram Vishwanath, Michael E. Papka |
CCGrid | 3 |
| 2025 | ytopt: Autotuning Scientific Applications for Energy Efficiency at Large ScalesabstractABSTRACT As we enter the exascale computing era, efficiently utilizing power and optimizing the performance of scientific applications under power and energy constraints has become critical and challenging. We propose a low‐overhead autotuning framework to autotune performance and energy for various hybrid MPI/OpenMP scientific applications at large scales and to explore the tradeoffs between application runtime and power/energy for energy efficient application execution, then use this framework to autotune four ECP proxy applications—XSBench, AMG, SWFFT, and SW4lite. Our approach uses Bayesian optimization with a Random Forest surrogate model to effectively search parameter spaces with up to 6 million different configurations on two large‐scale HPC production systems, Theta at Argonne National Laboratory and Summit at Oak Ridge National Laboratory. The experimental results show that our autotuning framework at large scales has low overhead and achieves good scalability. Using the proposed autotuning framework to identify the best configurations, we achieve up to 91.59% performance improvement, up to 21.2% energy savings, and up to 37.84% EDP (energy delay product) improvement on up to 4096 nodes. Xingfu Wu, Prasanna Balaprakash, Michael Kruse, Jaehoon Koo, Brice Videau, Paul D. Hovland, Valerie Taylor 0001, Brad Geltz, Siddhartha Jana, Mary W. Hall |
Concurr. Comput. Pract. Exp. | 7 |
| 2023 | Utilizing ensemble learning for performance and power modeling and improvement of parallel cancer deep learning CANDLE benchmarksabstractAbstract Machine learning (ML) continues to grow in importance across nearly all domains in modeling to learn from data. Often a tradeoff exists between a model's ability to minimize bias and variance. In this article, we utilize ensemble learning to combine linear, nonlinear, and tree‐/rule‐based ML methods to cope with the bias‐variance tradeoff and result in more accurate models. We use the datasets collected for two parallel cancer deep learning CANDLE benchmarks, NT3 and P1B2, to build performance and power models based on hardware performance counters using single‐object and multiple‐objects ensemble learning to identify the most important counters for improvement on the Cray XC40 Theta at Argonne National Laboratory. Based on the insights from these models, we improve the performance and energy of P1B2 and NT3 by optimizing the deep learning environments TensorFlow, Keras, Horovod, and Python under the huge page size of 8 MB. Experimental results show that ensemble learning not only produces more accurate models but also provides more robust performance counter ranking. We achieve up to 61.15% performance improvement and up to 62.58% energy saving for P1B2 and up to 55.81% performance improvement and up to 52.60% energy saving for NT3 on up to 24,576 cores. Xingfu Wu, Valerie Taylor 0001 |
Concurr. Comput. Pract. Exp. | 2 |
| 2023 | Performance and power modeling and prediction using MuMMI and 10 machine learning methodsabstractSummary Energy‐efficient scientific applications require insight into how high performance computing system features impact the applications' power and performance. This insight can result from the development of performance and power models. In this article, we use the modeling and prediction tool MuMMI (Multiple Metrics Modeling Infrastructure) and 10 machine learning methods to model and predict performance and power consumption and compare their prediction error rates. We use an algorithm‐based fault‐tolerant linear algebra code and a multilevel checkpointing fault‐tolerant heat distribution code to conduct our modeling and prediction study on the Cray XC40 Theta and IBM BG/Q Mira at Argonne National Laboratory and the Intel Haswell cluster Shepard at Sandia National Laboratories. Our experimental results show that the prediction error rates in performance and power using MuMMI are less than 10% for most cases. By utilizing the models for runtime, node power, CPU power, and memory power, we identify the most significant performance counters for potential application optimizations, and we predict theoretical outcomes of the optimizations. Based on two collected datasets, we analyze and compare the prediction accuracy in performance and power consumption using MuMMI and 10 machine learning methods. Xingfu Wu, Valerie Taylor 0001, Zhiling Lan |
Concurr. Comput. Pract. Exp. | 2 |
| 2022 | Autotuning PolyBench benchmarks with LLVM Clang/Polly loop optimization pragmas using Bayesian optimizationabstractAbstract We develop a ytopt autotuning framework that leverages Bayesian optimization to explore the parameter space search and compare four different supervised learning methods within Bayesian optimization and evaluate their effectiveness. We select six of the most complex PolyBench benchmarks and apply the newly developed LLVM Clang/Polly loop optimization pragmas to the benchmarks to optimize them. We then use the autotuning framework to optimize the pragma parameters to improve their performance. The experimental results show that our autotuning approach outperforms the other compiling methods to provide the smallest execution time for the benchmarks syr2k, 3mm, heat‐3d, lu, and covariance with two large datasets in 200 code evaluations for effectively searching the parameter spaces with up to 170,368 different configurations. We find that the Floyd–Warshall benchmark did not benefit from autotuning. To cope with this issue, we provide some compiler option solutions to improve the performance. Then we present loop autotuning without a user's knowledge using a simple mctree autotuning framework to further improve the performance of the Floyd–Warshall benchmark. We also extend the ytopt autotuning framework to tune a deep learning application. Xingfu Wu, Michael Kruse, Prasanna Balaprakash, Hal Finkel, Paul D. Hovland, Valerie Taylor 0001, Mary W. Hall |
Concurr. Comput. Pract. Exp. | 6 |
| 2021 | A Dynamic Power Capping Library for HPC ApplicationsabstractAs HPC systems increase in scale and capability, the cost of supplying power to these systems grows significantly. This introduces the urgent need for energy efficient computing through the management of power consumption. The PowerStack initiative defines a holistic power management framework at three levels, i.e., at the cluster level, the job level, and the node level [1] –[3]. It is expected that a cluster will be given a system-wide power budget (aka an allocated power budget), which can be intelligently managed and distributed at the job and node levels. Our work provides a method for intelligently managing node-level power during execution. Zhiling Lan, Xingfu Wu, Valerie Taylor 0001 |
CLUSTER | 4 |
| 2021 | Increasing Diversity in Computing Education: Lesson LearnedabstractIt is recognized that diversity of perspectives results in solutions that serve a broad base. This is critical to the area of computing, which has applications to many areas including science, humanities, finance, as well as foundational work to advance the field of computing. Cultivating an environment that values and promotes diversity, however, is not an easy task. This is especially the case in higher education, with a focus on undergraduate and graduate computing education. In this talk, I will discuss lessons learned from the following CMD-IT programs: University Award, FLIP Alliance, and Academic Careers Workshops. The focus will be on effective strategies and challenges that require community engagement. Valerie Taylor 0001 |
SIGCSE | 1 |
| 2020 | Toward an End-to-End Auto-tuning Framework in HPC PowerStackabstractEfficiently utilizing procured power and optimizing performance of scientific applications under power and energy constraints are challenging. The HPC PowerStack defines a software stack to manage power and energy of high-performance computing systems and standardizes the interfaces between different components of the stack. This survey paper presents the findings of a working group focused on the end-to-end tuning of the PowerStack. First, we provide a background on the PowerStack layer-specific tuning efforts in terms of their high-level objectives, the constraints and optimization goals, layer-specific telemetry, and control parameters, and we list the existing software solutions that address those challenges. Second, we propose the PowerStack end-to-end auto-tuning framework, identify the opportunities in co-tuning different layers in the PowerStack, and present specific use cases and solutions. Third, we discuss the research opportunities and challenges for collective auto-tuning of two or more management layers (or domains) in the PowerStack. This paper takes the first steps in identifying and aggregating the important R&D challenges in streamlining the optimization efforts across the layers of the PowerStack. Xingfu Wu, Aniruddha Marathe, Siddhartha Jana, Ondrej Vysocky, Jophin John, Andrea Bartolini, Lubomir Riha, Michael Gerndt, Valerie Taylor 0001, Sridutt Bhalachandra |
CLUSTER | 9 |
| 2020 | Increasing Diversity in Computing: A Focus on RetentionabstractIt is recognized that diversity of perspectives results in solutions that serve a broad base. This is critical to the area of computing, which has applications to science, humanities, as well as foundational work to advance the field of computing. Cultivating an environment that values and retains diversity, however, is not an easy task. This is especially the case in the "classroom", whereby the students are in the space for only a few hours per week. Hence, in terms of retention, it is necessary to complement "classroom" activities to foster and retain diversity, with extra-curricular activities. The importance of retention of diverse students, which results in an increase in the graduation of diverse students, resulted in the Center for Minorities and People with Disabilities in IT (CMD-IT) establishing the CMD-IT University Award for the Retention of Minorities and People with Disabilities in Computing. In this talk, I will discuss lessons learned from the award as well as other organizations and alliances focused on retention of diverse students and present some challenges to the CS Education community. Valerie Taylor 0001 |
SIGCSE | 1 |
| 2019 | Performance, Energy, and Scalability Analysis and Improvement of Parallel Cancer Deep Learning CANDLE BenchmarksabstractTraining scientific deep learning models requires the significant compute power of high-performance computing systems. In this paper, we analyze the performance characteristics of the benchmarks from the exploratory research project CANDLE (Cancer Distributed Learning Environment) with a focus on the hyperparameters epochs, batch sizes, and learning rates. We present the parallel methodology that uses the distributed deep learning framework Horovod to parallelize the CANDLE benchmarks. We then use scaling strategies for both epochs and batch size with linear learning rate scaling to investigate how they impact the execution time and accuracy as well as the power, energy, and scalability of the parallel CANDLE benchmarks under conditions of strong scaling and weak scaling on the IBM Power9 heterogeneous system Summit at Oak Ridge National Laboratory and the Cray XC40 Theta at Argonne National Laboratory. This study provides insights into how to set the proper numbers of epochs, batch sizes, and compute resources for these benchmarks to preserve the high accuracy and to reduce the execution time of the benchmarks. We identify the data-loading performance bottleneck and then improve the performance and energy for better scalability. Results with the modified benchmarks on Summit indicate up to 78.25% in performance improvement and up to 78% in energy saving under strong scaling on up to 384 GPUs, and up to 79.5% in performance improvement and up to 77.11% in energy saving under weak scaling on up to 3,072 GPUs. On Theta, we achieve up to 45.22% performance improvement and up to 41.78% in energy saving under strong scaling on up to 384 nodes. Moreover, the modification dramatically reduces the broadcast overhead. Xingfu Wu, Valerie Taylor 0001, Justin M. Wozniak, Rick L. Stevens, Thomas S. Brettin, Fangfang Xia |
ICPP | 2 |
| 2019 | The New NSF Requirement for Broadening Participation in Computing (BPC) Plans: Community Advice and ResourcesabstractThe CISE directorate of the NSF is rolling out a requirement that all NSF grants include a Broadening Participation in Computing (BPC) plan (www.nsf.gov/cise/bpc/). This has the potential to drive important institutional change across CS departments in the U.S. This panel of BPC experts will offer their perspectives on meaningful BPC activities, talk about existing BPC programs, and share BPC-related resources that can help PIs and departments craft high-quality BPC plans. The panelists will offer contrasting perspectives on topics such as K-12 outreach, the allocation of department funds for BPC, faculty engagement, and first steps departments should take. Ultimately, NSF review panels made up of CISE community members will evaluate BPC plans, but we hope to spark productive conversations in the interest of fostering institutional change to achieve the social imperative of BPC. Tracy Camp, Wendy M. DuBow, Diane Levitt, Linda J. Sax, Valerie Taylor 0001, Colleen M. Lewis |
SIGCSE | 5 |
| 2015 | Transfer student pathways to engineering degrees: A multi-institutional study based in TexasabstractThis work in progress paper describes a mixed methods study designed to investigate transfer student pathways as a means to increase engineering degree production and broaden participation in engineering careers. The study explores the framework of transfer student capital and its relevance for engineering transfer students. In this paper, we provide an overview of a survey instrument developed to collect data from more than 6,000 students who successfully transferred to one of four 4-year institutions in Texas as new engineering students between 2007 and 2014. Our study expands the small body of literature on engineering transfer students and sheds light on specific policies and practices that impact transfer. Andrea M. Ogilvie, David B. Knight, Arturo A. Fuentes, Maura Borrego, Patricia A. Nava, Valerie Taylor 0001 |
FIE | 6 |
| 2013 | MuMMI: Multiple Metrics Modeling InfrastructureabstractThe MuMMI (Multiple Metrics Modeling Infrastructure) project is an infrastructure that facilitates systematic measurement, modeling, and prediction of performance, power consumption and performance-power tradeoffs for parallel systems. In this paper, we present the MuMMI framework, which consists of an Instrument or, Databases and Analyzer. The MuMMI instrument or provides for automatic performance and power data collection and storage with low overhead. The MuMMI Databases store performance, power and energy consumption and hardware performance counters' data. The MuMMI Analyzer entails performance and power modeling and performance-power tradeoff and optimizations. As part of the MuMMI project, we mainly focus on discussing the design and development of a MuMMI Instrument or to provide automatic performance and power data collection and storage with low overhead on multicore systems in detail, then utilize the MuMMI Instrument or to collect performance and power data for a hybrid MPI/OpenMP earthquake application to discuss application performance-power trade-off and optimizations. Our experimental results show that we reduce up to 8.5% the application execution time and lower up to 18.35% the energy consumption by applying Dynamic Voltage and Frequency Scaling (DVFS), Dynamic Concurrency Throttling (DCT) and loop optimizations. Xingfu Wu, Charles W. Lively, Valerie Taylor 0001, Hung-Ching Chang, Chun-Yi Su, Kirk W. Cameron, Shirley Moore, Daniel Terpstra, Vincent M. Weaver |
SNPD | 3 |
| 2013 | Performance Characteristics of Hybrid MPI/OpenMP Scientific Applications on a Large-Scale Multithreaded BlueGene/Q SupercomputerabstractIn this paper, we investigate the performance characteristics of five hybrid MPI/OpenMP scientific applications (two NAS Parallel benchmarks Multi-Zone SP-MZ and BT-MZ, an earthquake simulation PEQdyna, an aerospace application PMLB and a 3D particle-in-cell application GTC) on a large-scale multithreaded Blue Gene/Q supercomputer at Argonne National laboratory, and quantify the performance gap resulting from using different number of threads per node. We use performance tools and MPI profile and trace libraries available on the supercomputer to analyze and compare the performance of these hybrid scientific applications with increasing the number OpenMP threads per node, and find that increasing the number of threads to some extent saturates or worsens performance of these hybrid applications. For the strong-scaling hybrid scientific applications such as SP-MZ, BT-MZ, PEQdyna and PLMB, using 32 threads per node results in much better application efficiency than using 64 threads per node, and as increasing the number of threads per node, the FPU (Floating Point Unit) percentage decreases, and the MPI percentage (except PMLB) and IPC (Instructions per cycle) per core (except BT-MZ) increase. For the weak-scaling hybrid scientific application such as GTC, the performance trend (relative speedup) is very similar with increasing number of threads per node no matter how many nodes (32, 128, 512) are used. Xingfu Wu, Valerie Taylor 0001 |
SNPD | 2 |
| 2013 | Performance modeling of hybrid MPI/OpenMP scientific applications on large-scale multicore supercomputers
Xingfu Wu, Valerie Taylor 0001 |
J. Comput. Syst. Sci. | 2 |
| 2012 | Performance Characteristics of Hybrid MPI/OpenMP Implementations of NAS Parallel Benchmarks SP and BT on Large-Scale Multicore ClustersabstractThe NAS Parallel Benchmarks (NPB) are well-known applications with fixed algorithms for evaluating parallel systems and tools. Multicore clusters provide a natural programming paradigm for hybrid programs, whereby OpenMP can be used with the data sharing with the multicores that comprise a node, and MPI can be used with the communication between nodes. In this paper, we use Scalar Pentadiagonal (SP) and Block Tridiagonal (BT) benchmarks of MPI NPB 3.3 as a basis for a comparative approach to implement hybrid MPI/OpenMP versions of SP and BT. In particular, we can compare the performance of the hybrid SP and BT with the MPI counterparts on large-scale multicore clusters, Intrepid (BlueGene/P) at Argonne National Laboratory and Jaguar (Cray XT4/5) at Oak Ridge National Laboratory. Our performance results indicate that the hybrid SP outperforms the MPI SP by up to 20.76%, and the hybrid BT outperforms the MPI BT by up to 8.58% on up to 10 000 cores on Intrepid and Jaguar. We also use performance tools and MPI trace libraries available on these clusters to further investigate the performance characteristics of the hybrid SP and BT. Xingfu Wu, Valerie Taylor 0001 |
Comput. J. | 2 |
| 2009 | Performance projection of HPC applications using SPEC CFP2006 benchmarksabstractPerformance projections of high performance computing (HPC) applications onto various hardware platforms are important for hardware vendors and HPC users. The projections aid hardware vendors in the design of future systems, enable them to compare the application performance across different existing and future systems, and help HPC users with system procurement and application refinements. In this paper, we present a method for projecting the node level performance of HPC applications using published data of industry standard benchmarks, the SPEC CFP2006, and hardware performance counter data from one base machine. In particular, we project performance of eight HPC applications onto four systems, utilizing processors from different vendors, using data from one base machine, the IBM p575. The projected performance of the eight applications was within 7.2% average difference with respect to measured runtimes for IBM POWER6 systems and standard deviation of 5.3%. For two Intel based systems with different micro-architecture and instruction set architecture (ISA) than the base machine, the average projection difference to measured runtimes was 10.5% with standard deviation of 8.2%. Sameh Sharkawi, Don DeSota, Raj Panda, Rajeev Indukuru, Stephen Stevens, Valerie Taylor 0001, Xingfu Wu |
IPDPS | 6 |
| 2008 | Performance Analysis of Parallel Visualization Applications and Scientific Applications on an Optical GridabstractOne major challenge for grid environments is how to efficiently utilize geographically distributed resources given large communication latency introduced by wide area networks interconnecting different grid sites. In this paper, we use optical networks to connect four clusters from three different sites: Texas A&M University, University of Illinois at Chicago, and University of California at San Diego to form an optical grid as an OptIPuter testbed, and execute parallel visualization applications and scientific applications to analyze the performance of these applications on the optical grid. Our experiments indicate that on-demand scheduling for resource allocation on different grid sites ensures to run MPI programs on the grid successfully. We explore the dedicated optical grid for different purpose of usage with on-demand scheduling to address how the optical grid impacts the performance of parallel programs. To avoid the large across-site communication latency, the ideal applications for the optical grid are embarrassingly parallel. Xingfu Wu, Valerie Taylor 0001 |
CW | 2 |
| 2008 | A Methodology for Developing High Fidelity Communication Models for Large-Scale Applications Targeted on Multicore SystemsabstractResource sharing and implementation of software stack for emerging multicore processors introduce performance and scaling challenges for large-scale scientific applications, particularly on systems with thousands of processing elements. Traditional performance optimization, tuning and modeling techniques that rely on uniform representation of computation and communication requirements are only partially useful due to the complexity of applications and underlying systems and software architecture. In this paper, we propose a workload modeling methodology that allows application developers to capture and represent hierarchical decomposition and distribution of their applications thereby allowing them to explore and identify optimal mapping of a workload on a target system. We demonstrate the proposed methodology on a Teraflopsscale fusion application that is developed using message-passing (MPI) programming paradigm. Using our analysis and projection results, we obtain insight into the performance characteristics of the application on a quad-core system and also identify optimal mapping on a Teraflops-scale platform. 1. Charles W. Lively, Valerie Taylor 0001, Sadaf R. Alam, Jeffrey S. Vetter |
SBAC-PAD | 2 |
| 2006 | Performance Analysis, Modeling and Prediction of a Parallel Multiblock Lattice Boltzmann Application Using Prophesy SystemabstractRecently, the Lattice Boltzmann method is widely used in simulating fluid flows. In this paper, we present the performance analysis, modeling and prediction of a parallel multiblock Lattice Boltzmann application on up to 512 processors on three SMP clusters: two IBM SP systems at San Diego Supercomputing Center (DataStar-p655 and p690) and one IBM SP system at the DOE National Energy Research Scientific Computing Center (Seaborg) using the Prophesy system. By characterizing the performance of the Lattice Boltzmann application as the problem size and the number of processors increase, we can identify and eliminate performance bottlenecks, and predict the application performance. The experimental results indicate that the application with large problem sizes scales well across these three clusters, and performance models using the coupling method are accurate with less than 4.8% average relative prediction error. Xingfu Wu, Valerie Taylor 0001, Shane Garrick, Dazhi Yu, Jacques Richard |
CLUSTER | 2 |
| 2006 | A hybrid framework for design and analysis of fault-tolerant architecturesabstractIt is anticipated that self assembled ultra-dense nanomemories will be more susceptible to manufacturing defects and transient faults than conventional CMOS-based memories, thus the need exists for fault-tolerant memory architectures. The development of such architectures will require intense analysis in terms of achievable performance measures- power dissipation, area, delay and reliability. In this paper, we propose and develop a hybrid automation framework, called HMAN, that aids the design and analysis of fault-tolerant architectures for nanomemories. Our framework can analyze memory architectures at two different levels of the design abstraction, namely the system and circuit levels. To the best of our knowledge, this is the first such attempt at analyzing memory systems at different levels of abstraction and then correlating the different performance measures. We also illustrate the application of our framework to self-assembled crossbar architectures by analyzing a hierarchical fault-tolerant crossbar-based memory architecture that we have developed. Debayan Bhaduri, Sandeep K. Shukla, Deji Coker, Valerie Taylor 0001, Paul S. Graham, Maya B. Gokhale |
DATE | 4 |
| 2006 | DistDLB: Improving cosmology SAMR simulations on distributed computing systems through hierarchical load balancing
Zhiling Lan, Valerie Taylor 0001 |
J. Parallel Distributed Comput. | 2 |
| 2004 | Predicting application run times with historical information
Warren Smith, Ian T. Foster, Valerie Taylor 0001 |
J. Parallel Distributed Comput. | 3 |
| 2003 | Exploring cosmology applications on distributed environments
Zhiling Lan, Valerie Taylor 0001, Greg Bryan |
Future Gener. Comput. Syst. | 2 |
| 2002 | I/O Analysis and Optimization for an AMR Cosmology ApplicationabstractIn this paper we investigate the data access patterns and file I/O behaviors of a production cosmology application that uses the adaptive mesh refinement (AMR) technique for its domain decomposition. This application was originally developed using Hierarchical Data Format (HDF version 4) I/O library and since HDF4 does not provide parallel I/O facilities, the global file I/O operations were carried out by one of the allocated processors. When the number of processors becomes large, the I/O performance of this design degrades significantly due to the high communication cost and sequential file access. In this work, we present two additional I/O implementations, using MPI-IO and parallel HDF version 5, and analyze their impacts to the I/O performance for this typical AMR application. Based on the I/O patterns discovered in this application, we also discuss the interaction between user level parallel I/O operations and different parallel file systems and point out the advantages and disadvantages. The performance results presented in this work are obtained from an SGI Origin2000 using XFS, an IBM SP using GPFS, and a Linux cluster using PVFS. Wei-keng Liao, Alok N. Choudhary, Valerie Taylor 0001 |
CLUSTER | 4 |
| 2002 | Using Kernel Couplings to Predict Parallel Application PerformanceabstractPerformance models provide significant insight into the performance relationships between an application and the system used for execution. The major obstacle to developing performance models is the lack of knowledge about the performance relationships between the different functions that compose an application. This paper addresses the issue by using a coupling parameter, which quantifies the interaction between kernels, to develop performance predictions. The results, using three NAS parallel application benchmarks, indicate that the predictions using the coupling parameter were greatly improved over a traditional technique of summing the execution times of the individual kernels in an application. In one case the coupling predictor had less than 1% relative error in contrast the summation methodology that had over 20% relative error. Further, as the problem size and number of processors scale, the coupling values go through a finite number of major value changes that is dependent on the memory subsystem of the processor architecture. Valerie Taylor 0001, Xingfu Wu, Jonathan Geisler, Rick L. Stevens |
HPDC | 1 |
| 2002 | Performance Coupling: Case Studies for Improving the Performance of Scientific Applications
Jonathan Geisler, Valerie Taylor 0001 |
J. Parallel Distributed Comput. | 2 |
| 2002 | A novel dynamic load balancing scheme for parallel systems
Zhiling Lan, Valerie Taylor 0001, Greg Bryan |
J. Parallel Distributed Comput. | 2 |
| 2002 | Mesh Partitioning for Efficient Use of Distributed SystemsabstractMesh partitioning for homogeneous systems has been studied extensively; however, mesh partitioning for distributed systems is a relatively new area of research. To ensure efficient execution on a distributed system, the heterogeneities in the processor and network performance must be taken into consideration in the partitioning process; equal size subdomains and small cut set size, which results from conventional mesh partitioning, are no longer the primary goals. In this paper, we address various issues related to mesh partitioning for distributed systems. These issues include the metric used to compare different partitions, efficiency of the application executing on a distributed system, and the advantage of exploiting heterogeneity in network performance. We present a tool called PART, for automatic mesh partitioning for distributed systems. The novel feature of PART is that it considers heterogeneities in the application and the distributed system. Simulated annealing is used in PART to perform the backtracking search for desired partitions. While it is well-known that simulated annealing is computationally intensive, we describe the parallel version of simulated annealing that is used with PART. The results of the parallelization exhibit superlinear speedup in most cases and nearly perfect speedup for the remaining cases. Experimental results are also presented for partitioning regular and irregular finite element meshes for an explicit, nonlinear finite element application, called WHAMS2D, executing on a distributed system consisting of two IBM SPs with different processors. The results from the regular problems indicate a 33 to 46 percent increase in efficiency when processor performance is considered as compared to the conventional even partitioning. The results indicate a 5 to 15 percent increase in efficiency when network performance is considered as compared to considering only processor performance; this is significant given that the optimal improvement is 15 percent for this application. The results from the irregular problem indicate up to 36 percent increase in efficiency when processor and network performance are considered as compared to even partitioning. Jian Chen 0044, Valerie Taylor 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2001 | Dynamic Load Balancing for Structured Adaptive Mesh Refinement ApplicationsabstractAdaptive Mesh Refinement (AMR) is a type of multiscale algorithm that achieves high resolution in localized regions of dynamic, multidimensional numerical simulations. One of the key issues related to AMR is dynamic load balancing (DLB), which allows large-scale adaptive applications to run efficiently on parallel systems. In this paper we present an efficient DLB scheme for structured AMR (SAMR) applications. Our DLB scheme combines a grid-splitting technique with direct grid movements (e.g., direct movement from an overloaded processor to an underloaded proces sor), for which the objective is to efficiently redistribute workload among all the processors so as to reduce the parallel execution time. The potential benefits of our DLB scheme are examined by incorporating our techniques into a parallel, cosmological application that uses SAMR techniques. Experiments show that by using our scheme, the parallel execution time can be reduced by up to 47% and the quality of load-balancing can be improved by a factor of four. Zhiling Lan, Valerie Taylor 0001, Greg Bryan |
ICPP | 2 |
| 2001 | Dynamic load balancing of SAMR applications on distributed systems
Zhiling Lan, Valerie Taylor 0001, Greg Bryan |
SC | 2 |
| 2001 | Balancing Load versus Decreasing Communication: Parameterizing the Tradeoff
Valerie Taylor 0001, Eric J. Schwabe, Bruce K. Holmer, Michelle R. Hribar |
J. Parallel Distributed Comput. | 1 |
| 2001 | Implementing parallel shortest path for parallel transportation applications
Michelle R. Hribar, Valerie Taylor 0001, David E. Boyce |
Parallel Comput. | 2 |
| 2000 | Prophesy: An Infrastructure for Analyzing and Modeling the Performance of Parallel and Distributed ApplicationsabstractEfficient execution of applications requires insight into how the system features impact the performance of the application. For distributed systems, the task of gaining this insight is complicated by the complexity of the system features. This insight generally results from significant experimental analysis and possibly the development of performance models. This paper presents the Prophesy project, an infrastructure that aids in gaining this needed insight based upon experience. The core component of Prophesy is a relational database that allows for the recording of performance data, system features and application details. Xingfu Wu, Valerie Taylor 0001, Jonathan Geisler, Zhiling Lan, Rick L. Stevens, Mark Hereld, Ivan R. Judson |
HPDC | 2 |
| 2000 | Scheduling with Advanced ReservationsabstractSome computational grid applications have very large resource requirements and need simultaneous access to resources from more than one parallel computer. Current scheduling systems do not provide mechanisms to gain such simultaneous access without the help of human administrators of the computer systems. In this work, we propose and evaluate several algorithms for supporting advanced reservation of resources in supercomputing scheduling systems. These advanced reservations allow users to request resources from scheduling systems at specific times. We find that the wait times of applications submitted to the queue increases when reservations are supported and the increase depends on how reservations are supported. Further, we find that the best performance is achieved when we assume that applications can be terminated and restarted, backfilling is performed, and relatively accurate run-time predictions are used. Warren Smith, Ian T. Foster, Valerie Taylor 0001 |
IPDPS | 3 |
| 2000 | ParaPART: Parallel Mesh Partitioning Tool for Distributed SystemsabstractIn this paper, we present ParaPART, a parallel version of a mesh partitioning tool called PART for distributed systems. PART takes into consideration the heterogeneities in processor performance, network performance and application computational complexities to achieve a balanced estimate of execution time across the processors in the distributed system. Simulated annealing is used in PART to perform the backtracking search for desired partitions. ParaPART significantly improves the performance of PART by using the asynchronous multiple Markov chain approach of parallel simulated annealing. ParaPART is used to partition six irregular meshes into 8, 16 and 100 subdomains using up to 64 client processors on an IBM SP machine. The results show superlinear speedup in most cases and nearly perfect speedup for the remaining cases. Using the partitions from ParaPART, we ran an explicit, two-dimensional finite element code on two geographically distributed IBM SP machines. Results indicate that ParaPART produces results within 3% of PART. The execution time of the finite element code was reduced by 12% for eight processors as compared with partitions that consider only processor performance; this is significant given the theoretical upper bound of 15% reduction. Up to 35% reduction was achieved when up to 40 processors were used for the same problems. Copyright © 2000 John Wiley & Sons, Ltd. Jian Chen 0044, Valerie Taylor 0001 |
Concurr. Pract. Exp. | 2 |
| 1999 | Data Management for Large-Scale Scientific Computations in High Performance Distributed SystemsabstractWith the increasing number of scientific applications manipulating huge amounts of data, effective data management is an increasingly important problem. Unfortunately, so far the solutions to this data management problem either require deep understanding of specific storage architectures and file layouts (as in high-performance file systems) or produce unsatisfactory I/O performance in exchange for ease-of-use and portability (as in relational DBMSs). In this paper we present a new environment which is built around an active meta-data management system (MDMS). The key components of our three-tiered architecture are user application, the MDMS, and a hierarchical storage system (HSS). Our environment overcomes the performance problems of pure database-oriented solutions, while maintaining their advantages in terms of ease-of-use and portability. The high levels of performance are achieved by the MDMS, with the aid of user-specified directives. Our environment supports a simple, easy-to-use yet powerful user interface, leaving the task of choosing appropriate I/O techniques to the MDMS. We discuss the importance of an active MDMS and show how the three components, namely application, the MDMS, and the HSS, fit together. We also report performance numbers from our initial implementation and illustrate that significant improvements are made possible without undue programming effort. Alok N. Choudhary, Mahmut T. Kandemir, Harsha S. Nagesh, Jaechun No, Xiaohui Shen, Valerie Taylor 0001, Sachin More, Rajeev Thakur |
HPDC | 6 |
| 1998 | Mesh Partitioning for Distributed SystemsabstractDistributed systems, which consist of a collection of high performance systems interconnected via high performance networks (e.g. ATM), are becoming feasible platforms for execution of large-scale, complex problems. We address various issues related to mesh partitioning for distributed systems. These issues include the metric used to compare different partitions, efficiency of the application executing on a distributed system, the number of cut sets, and the advantage of exploiting heterogeneity in network performance. We present a tool called PART for automatic mesh partitioning for distributed systems. The novel feature of PART is that it considers heterogeneities in the application and the distributed system. The heterogeneities in the distributed system include processor and network performance; the heterogeneities in the application include computational complexities. Preliminary results are presented for partitioning regular and irregular finite element meshes for the WHAMS2D application executing on a distributed system consisting of two IBM SPs. The results from the regular problems indicate a 33-46% increase in efficiency when processor performance is considered as compared to the conventional even partitioning; the results also indicate an additional 5-16% increase in efficiency when network performance is considered. The result from the irregular problem indicate a 21% increase in efficiency when processor and network performance are considered as compared to even partitioning. Jian Chen 0044, Valerie Taylor 0001 |
HPDC | 2 |
| 1998 | Predicting Application Run Times Using Historical Information
Warren Smith, Ian T. Foster, Valerie Taylor 0001 |
JSSPP | 3 |
| 1998 | Termination Detection for Parallel Shortest Path Algorithms
Michelle R. Hribar, Valerie Taylor 0001, David E. Boyce |
J. Parallel Distributed Comput. | 2 |
| 1997 | PART: a partitioning tool for efficient use of distributed systemsabstractThe interconnection of geographically distributed supercomputers via high-speed networks allows users to access the needed compute power for large-scale, complex applications. For efficient use of such systems, the variance in processor performance and network (i.e., interconnection network versus wide area network) performance must be considered. In this paper, we present a decomposition tool, called PART, for distributed systems. PART takes into consideration the variance in performance of the networks and processors as well as the computational complexity of the application. This is achieved via the parameters used in the objective function of simulated annealing. The initial version of PART focuses on finite element based problems. The results of using PART demonstrate a 30% reduction in execution time as compared to using conventional schemes that partition the problem domain into equal-sized subdomains. Jian Chen 0044, Valerie Taylor 0001 |
ASAP | 2 |
| 1997 | Parallel Molecular Dynamics: Implications for Massively Parallel Machines
Valerie Taylor 0001, Rick L. Stevens, Kathryn E. Arnold |
J. Parallel Distributed Comput. | 1 |
| 1996 | A Decomposition Method For Efficient Use Of Distributed Supercomputers For Finite Element ApplicationsabstractThe interconnection of geographically distributed supercomputers via highspeed networks makes available the needed compute power for large-scale scientific applications, such as finite element applications. In this paper we propose a two-level data decomposition method for efficient execution of finite element applications on a network of supercomputers. Our method exploits the following features that may be different for each supercomputer in the system: processor speed, number of processors used from each supercomputer, local network performance, wide area network performance and wide area topology. Preliminary experiments involving a nonlinear, finite element application executed on a network of two supercomputers, one located at Argonne National Laboratory and the other one at the Cornell Theory Center, demonstrate a 20% reduction in execution time when the proposed decomposition is used as compared with naively applying conventional decompositions that are applicable to single supercomputers. Valerie Taylor 0001, Jian Chen 0044, Thomas Canfield, Rick L. Stevens |
ASAP | 1 |
| 1995 | SPAR: A New Architecture for Large Finite Element ComputationsabstractThe finite element method is a general and powerful technique for solving partial differential equations. The computationally intensive step of this technique is the solution of a linear system of equations. Very large and very sparse system matrices result from large finite-element applications. The sparsity must be exploited for efficient use of memory and computational components in executing the solution step. In this paper we propose a scheme, called SPAR, for efficiently storing and performing computations on sparse matrices. SPAR consists of an alternate method of representing sparse matrices and an architecture that efficiently executes computations on the proposed data structure. The SPAR architecture has not been built, but we have constructed a register-transfer level simulator and executed the sparse matrix computations used with some large finite element applications. The simulation results demonstrate a 95% utilization of the floating-point units for some 3D applications. SPAR achieves high utilization of memory, memory bandwidth, and floating-point units when executing sparse matrix computations.> Valerie Taylor 0001, Abhiram G. Ranade, David G. Messerschmitt |
IEEE Trans. Computers | 1 |
| 1994 | Practical Isuues of 2-D Parallel Finite Element AnalysisabstractThe use of parallel processors has made it possible to execute large scale applications such as finite element analysis. Generally, the speedup is limited by the interprocessor communication used by message-passing multicomputers. The major question addressed by users of message-passing machines is the identification, of the most efficient type of communication scheme for the. particular application. Practical considerations such as the following are frequently neglected: the range of message sizes for which a communication scheme is appropriate, the impact on the subdomain size, the. memory requirements versus the actual memory of the machine, and how the optimal communication method changes as a fixed size problem is scaled to a larger number of processors. We address these practical issues for 2-D finite, element problems executed on the Intel Delta machine. Michelle R. Hribar, Valerie Taylor 0001 |
ICPP (3) | 2 |
| 1992 | Sparse Matrix Computations: Implications for Cache DesignsabstractHigh-performance cache designs are studied for the class of sparse matrix computations, which are often excluded from the general programs used in previous cache studies. In particular, the data that should be stored in the cache are identified, and the cache organization is studied in terms of associativity, size, write operation, write policy, block size, and number of read and write ports. Simulation results demonstrate that a 1-kword or 8-kbyte (one word is equal to 64 b), direct-mapped cache produces good results with almost all of the misses occurring from first time accesses. This cache size can easily fit on a chip, with plenty of room to spare for other components.> Valerie Taylor 0001 |
SC | 1 |
| 1991 | Three-dimensional finite-element analyses: implications for computer architecturesabstractArticle Three-dimensional finite-element analyses: implications for computer architectures Share on Authors: Valerie E. Taylor Department of Electrical Engineering and Computer Sciences, University of California at Berkeley, Berkeley, CA Department of Electrical Engineering and Computer Sciences, University of California at Berkeley, Berkeley, CAView Profile , Abhiram Ranade Department of Electrical Engineering and Computer Sciences, University of California at Berkeley, Berkeley, CA Department of Electrical Engineering and Computer Sciences, University of California at Berkeley, Berkeley, CAView Profile , David G. Messerschmitt Department of Electrical Engineering and Computer Sciences, University of California at Berkeley, Berkeley, CA Department of Electrical Engineering and Computer Sciences, University of California at Berkeley, Berkeley, CAView Profile Authors Info & Claims Supercomputing '91: Proceedings of the 1991 ACM/IEEE conference on SupercomputingAugust 1991 Pages 786–795https://doi.org/10.1145/125826.126188Online:01 August 1991Publication History 0citation278DownloadsMetricsTotal Citations0Total Downloads278Last 12 Months0Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Valerie Taylor 0001, Abhiram G. Ranade, David G. Messerschmitt |
SC | 1 |